This is the full developer documentation for TypeGraph
# What is TypeGraph?
> A TypeScript-first embedded knowledge graph library
TypeGraph is a **TypeScript-first, embedded knowledge graph library** that brings property graph semantics and
ontological reasoning to applications using standard relational databases. Rather than introducing a separate graph
database, TypeGraph lives inside your application as a library, storing graph data in your existing SQLite or
PostgreSQL database.
## Architecture

## Core Capabilities
### 1. Type-Driven Schema Definition
Zod schemas are the single source of truth. From one schema definition, TypeGraph derives:
- Runtime validation rules
- TypeScript types (inferred, not duplicated)
- Database storage requirements
- Query builder type constraints
```typescript
const Person = defineNode("Person", {
schema: z.object({
fullName: z.string().min(1),
email: z.string().email().optional(),
dateOfBirth: z.date().optional(),
}),
});
```
### 2. Semantic Layer with Ontological Reasoning
Type-level relationships enable sophisticated inference:
| Relationship | Meaning | Use Case |
| -------------- | ----------------------------------------- | ---------------------- |
| `subClassOf` | Instance inheritance (Podcast IS-A Media) | Query expansion |
| `broader` | Hierarchical concept (ML broader than DL) | Topic navigation |
| `equivalentTo` | Same concept, different name | Cross-system mapping |
| `disjointWith` | Cannot be both (Person ≠ Organization) | Constraint validation |
| `implies` | Edge entailment (marriedTo implies knows) | Relationship inference |
| `inverseOf` | Edge pairs (manages/managedBy) | Bidirectional queries |
### 3. Self-Describing Schema (Homoiconic)
The schema and ontology are stored in the database as data, enabling:
- Runtime schema introspection
- Versioned schema history
- Self-describing exports and backups
- Migration tooling
### 4. Type-Safe Query Compilation
Queries compile to an AST before targeting SQL:
- Consistent semantics across SQLite and PostgreSQL
- Type-checked at compile time
- Query results have inferred types
### 5. Temporal and Bitemporal History
Every node and edge has a valid-time window (`validFrom` / `validTo`), so you
can ask what was true at a domain instant with `.temporal("asOf", T)` or
`store.asOf(T)`. Stores created with `{ history: true }` also capture recorded
time for TypeGraph-managed writes, the system-time axis that remembers when the
graph wrote each fact down. `store.asOfRecorded(T)` reconstructs what the graph
captured at a recorded instant, and
`store.asOf(validT).asOfRecorded(recordedT)` pins both axes independently.
Use it for audit trails, agent decision replay, effective-dated policies, and
breach forensics. See [Temporal queries](/queries/temporal) and the
[Bitemporal Time Travel](/examples/bitemporal-time-travel) example.
## Design Philosophy
### Embedded, Not External
TypeGraph is a library dependency, not a networked service. TypeGraph initializes with your application, uses your
database connection, and requires no separate deployment.
### Schema-First, Type-Driven
Define your schemas once with Zod, and TypeGraph handles validation, type inference, and storage.
No duplicate type definitions or manual synchronization.
### Explicit Over Implicit
TypeGraph favors explicit declarations:
- Relationships are declared, not inferred from foreign keys
- Semantic relationships are explicit in the ontology
- Cascade behavior is configured, not assumed
### Portable Abstractions
The query builder generates portable ASTs that can target different SQL dialects.
The same query code works with SQLite and PostgreSQL.
## What TypeGraph Is Not
TypeGraph deliberately excludes:
- **Broad graph analytics suites**: Focused PageRank, connectivity, and
deterministic label-propagation primitives are built in; modularity
optimization and most centrality measures are not
- **Distributed storage**: Single-database deployment only
These exclusions keep TypeGraph focused and maintainable.
Note: TypeGraph **does support** semantic search via native database vector
engines: pgvector for PostgreSQL, sqlite-vec for the local (better-sqlite3)
SQLite backend, and libSQL's built-in vectors for the libSQL / Turso backend.
See [Semantic Search](/semantic-search) for details.
Note: TypeGraph **does support** fulltext search — native BM25 on SQLite (FTS5)
and `tsvector` + GIN on PostgreSQL, with a query-builder
`n.$fulltext.matches()` predicate that composes with any other predicate.
Combine with semantic search for hybrid RAG retrieval. See
[Fulltext Search](/fulltext-search) for details.
Note: TypeGraph does support **variable-length paths** via `.recursive()` with
configurable depth limits, optional path/depth projection, and explicit cycle
policy. Cycle prevention is the default.
See [Recursive Traversals](/queries/recursive) for details.
Note: TypeGraph ships **Tier 1 graph algorithms** (shortest path, reachability,
neighborhoods, and degree) on `store.algorithms.*`. Traversal calls use a
set-based BFS frontier, while degree uses a single count query. See
[Graph Algorithms](/graph-algorithms) for details.
Note: TypeGraph supports **runtime schema induction** via graph
extensions. An LLM or ingestion agent can propose a typed schema as a
JSON-serializable document, an operator approves it, and `store.evolve()`
atomically commits a new schema version — no redeploy, full Zod
validation, restart parity. See [Graph Extensions](/graph-extensions)
for the agent-driven workflow.
Note: TypeGraph ships **graph merge** — fork a store into isolated working
copies, let many writers (parallel agents, importers, reviewers) edit
independently, then reconcile them into one canonical graph with deterministic
entity resolution (exact / blocking / fulltext / vector / hybrid), edge
repointing, conflict reporting, and provenance. `mergeIncremental()` folds new
sources into a *live* graph without creating duplicates — the primitive for
multi-agent knowledge-graph construction and continuous ingestion. See
[Graph Merge](/graph-merge) for the full guide.
Note: TypeGraph supports **bitemporal graph reads**. Valid time answers "when
was this fact true in the domain?" Recorded time answers "when did the graph
record it?" Together they reconstruct prior captured state after corrections, replay
agent decisions against the graph they actually saw, and traverse access graphs
at a breach instant for TypeGraph-managed writes. See
[Temporal queries](/queries/temporal) and the
[Agent Decision Replay](/examples/agent-decision-replay) example.
## Why TypeGraph?
### Compared to Graph Databases (Neo4j, Amazon Neptune)
Graph databases are powerful but come with operational overhead:
| Aspect | Graph Database | TypeGraph |
|--------|---------------|-----------|
| **Deployment** | Separate service to manage, scale, and monitor | Library in your app, uses existing database |
| **Network** | Additional latency for every query | In-process, no network hop |
| **Transactions** | Separate transaction scope from your SQL data | Same ACID transaction as your other data |
| **Learning curve** | New query language (Cypher, Gremlin) | TypeScript you already know |
| **Graph algorithms** | Broad suites (PageRank, shortest path, community detection) | Focused algorithms (shortest path, reachability, neighborhoods, degree, WCC, label propagation, PageRank/PPR) |
| **Scale** | Optimized for billions of nodes | Best for thousands to millions |
**Choose TypeGraph** when your graph is part of your application domain (knowledge bases, org
charts, content relationships) rather than a standalone analytical system.
### Compared to ORMs (Prisma, Drizzle, TypeORM)
ORMs model relations through foreign keys, which works well for simple associations but lacks graph semantics:
| Aspect | Traditional ORM | TypeGraph |
|--------|----------------|-----------|
| **Relationships** | Foreign keys, eager/lazy loading | First-class edges with properties |
| **Traversals** | Manual joins or N+1 queries | Fluent traversal API, compiled to efficient SQL |
| **Inheritance** | Table-per-class or single-table | Semantic `subClassOf` with query expansion |
| **Constraints** | Foreign key constraints | Disjointness, cardinality, implications |
| **Schema** | Migrations alter tables | Schema versioning, JSON properties |
**Choose TypeGraph** when you need to traverse relationships, model type hierarchies, or enforce
semantic constraints beyond what foreign keys provide.
### Compared to Triple Stores (RDF, SPARQL)
Triple stores and RDF provide rich ontological modeling but have practical challenges:
| Aspect | Triple Store | TypeGraph |
|--------|-------------|-----------|
| **Type safety** | Runtime validation, stringly-typed | Full TypeScript inference |
| **Query language** | SPARQL (powerful but verbose) | TypeScript fluent API |
| **Schema** | OWL/RDFS (complex specification) | Zod schemas (familiar, composable) |
| **Integration** | Separate system, data sync required | Embedded in your app |
| **Inference** | Full reasoning engines available | Precomputed closures, practical subset |
**Choose TypeGraph** when you want ontological concepts (subclass, disjoint, implies) without the
complexity of full semantic web stack.
### The TypeGraph Sweet Spot
TypeGraph is designed for applications where:
1. **The graph is your domain model** — not a separate analytical system
2. **You already use SQL** — and don't want another database to manage
3. **Type safety matters** — you want compile-time checking, not runtime surprises
4. **Semantic relationships help** — inheritance, implications, constraints add value
5. **Scale is moderate** — thousands to millions of nodes, not billions
## When to Use TypeGraph
TypeGraph is ideal for:
- **Knowledge bases** with typed entities and relationships
- **Organizational structures** with hierarchies and roles
- **Content graphs** with topics, articles, and references
- **Domain models** requiring semantic constraints
- **RAG applications** combining graph traversal with vector search
- **Multi-source ingestion & entity resolution** — reconcile parallel agent or
importer outputs into one canonical graph with [graph merge](/graph-merge)
- **Auditable AI systems and forensics** — reconstruct the graph an agent or
investigator saw at a recorded instant with [bitemporal reads](/queries/temporal#recorded-time-bitemporal)
TypeGraph is not ideal for:
- Large-scale graph analytics requiring distributed processing
- Social networks with billions of edges
- Real-time streaming graph data
- Applications requiring a broad graph-data-science suite such as community
detection or betweenness centrality (use Neo4j or a graph library; focused
algorithms—including PageRank and weighted shortest path—ship on
`store.algorithms.*`)
# Quick Start
> Set up TypeGraph and build your first knowledge graph
Get TypeGraph running in your project with this minimal example.
## 1. Install
```bash
npm install @nicia-ai/typegraph zod drizzle-orm better-sqlite3
npm install -D @types/better-sqlite3
```
> **Edge environments or libsql:** Skip `better-sqlite3` and use
> `@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql` with `@libsql/client`, or
> `@nicia-ai/typegraph/adapters/drizzle/sqlite` with your edge-compatible driver (D1, bun:sqlite).
> See [Backend Setup](/backend-setup#libsql--turso) and [Edge and Serverless](/integration#edge-and-serverless).
## 2. Create Your First Graph
```typescript
import { z } from "zod";
import { defineNode, defineEdge, defineGraph } from "@nicia-ai/typegraph";
import { createLocalSqliteStore } from "@nicia-ai/typegraph/sqlite/local";
// Define your schema
const Person = defineNode("Person", {
schema: z.object({ name: z.string(), role: z.string().optional() }),
});
const Project = defineNode("Project", {
schema: z.object({ name: z.string(), status: z.enum(["active", "done"]) }),
});
const worksOn = defineEdge("worksOn");
const graph = defineGraph({
id: "my_app",
nodes: { Person: { type: Person }, Project: { type: Project } },
edges: { worksOn: { type: worksOn, from: [Person], to: [Project] } },
});
// Provision an in-memory database and create the store
const store = await createLocalSqliteStore(graph);
// Use it!
const alice = await store.nodes.Person.create({ name: "Alice", role: "Engineer" });
const project = await store.nodes.Project.create({ name: "Website", status: "active" });
await store.edges.worksOn.create(alice, project, {});
// Query with full type safety
const results = await store
.query()
.from("Person", "p")
.traverse("worksOn", "e")
.to("Project", "proj")
.select((ctx) => ({ person: ctx.p.name, project: ctx.proj.name }))
.execute();
console.log(results); // [{ person: "Alice", project: "Website" }]
```
That's it! You have a working knowledge graph. Read on for the complete setup guide.
This managed entrypoint returns the complete typed `Store` while keeping its
public declaration surface independent of Drizzle. Use
`@nicia-ai/typegraph/postgres/pglite` for the same setup with in-process
PostgreSQL. If your application owns the database connection or needs direct
driver access, use the adapter entrypoints described below instead.
---
## Complete Setup Guide
This section covers production setup with SQLite and PostgreSQL in detail.
### Installation
```bash
npm install @nicia-ai/typegraph zod drizzle-orm better-sqlite3
npm install -D @types/better-sqlite3
```
> `better-sqlite3` is optional. For libsql/Turso, use `@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql`.
> For D1 or bun:sqlite, use `@nicia-ai/typegraph/adapters/drizzle/sqlite` with the matching Drizzle driver.
### SQLite Setup
TypeGraph provides two ways to set up SQLite:
#### Managed Store (Recommended)
Use the managed Store when TypeGraph should own the connection and provision
its schema:
```typescript
import { createLocalSqliteStore } from "@nicia-ai/typegraph/sqlite/local";
const store = await createLocalSqliteStore(graph, { path: "./my-app.db" });
// The Store owns the connection.
await store.close();
```
The return value is a `Store`: typed node and edge collections, queries,
algorithms, graph-owned transactions, schema evolution, and schema-derived
property types are all available. Adapter-native handles and caller-owned
transaction adoption are absent by design; opt into `AdapterStore` through a
Drizzle adapter entrypoint when application tables must share a transaction.
#### Quick Setup (Recommended for Development)
Use the backend wrapper when you also need the underlying Drizzle database or
want to choose how the Store is created.
> **Note:** `createLocalSqliteBackend` requires `better-sqlite3` and only works in Node.js.
> For edge environments, see [Manual Setup](#manual-setup-full-control) with
> `/adapters/drizzle/sqlite`.
```typescript
import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local";
// In-memory database (data lost on restart)
const { backend } = createLocalSqliteBackend();
// File-based database (persistent)
const { backend, db } = createLocalSqliteBackend({ path: "./my-app.db" });
```
The function returns both the `backend` (for use with `createStore`) and `db`
(the underlying Drizzle instance for direct SQL access if needed).
#### Manual Setup (Full Control)
For production deployments or when you need full control over the database configuration:
```typescript
import Database from "better-sqlite3";
import { drizzle } from "drizzle-orm/better-sqlite3";
import { createSqliteBackend, generateSqliteMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
// Create database connection
const sqlite = new Database("my-app.db");
// Run TypeGraph migrations (creates required tables)
sqlite.exec(generateSqliteMigrationSQL());
// Create Drizzle instance
const db = drizzle(sqlite);
// Create the backend
const backend = createSqliteBackend(db);
```
#### libsql / Turso Setup
For Turso, embedded replicas, or sharing a libsql connection with other libraries:
```typescript
import { createClient } from "@libsql/client";
import { createLibsqlBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql";
const client = createClient({ url: "file:app.db" });
const { backend } = await createLibsqlBackend(client);
```
`createLibsqlBackend` handles DDL automatically. The caller owns the client and
is responsible for closing it. See [Backend Setup](/backend-setup#libsql--turso) for
remote Turso URLs and caveats.
#### Edge-Compatible Setup (D1, bun:sqlite)
For Cloudflare Workers or Bun, use the driver-agnostic backend:
```typescript
import { drizzle } from "drizzle-orm/d1"; // or bun-sqlite
import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
// D1 example
const db = drizzle(env.DB);
const backend = createSqliteBackend(db);
```
Use [drizzle-kit managed migrations](/integration#drizzle-kit-managed-migrations-recommended)
to set up the schema.
#### Drizzle-Kit Managed Migrations
If you already use `drizzle-kit` for migrations, see [Drizzle-Kit Managed Migrations](/integration#drizzle-kit-managed-migrations-recommended)
for how to import TypeGraph's schema into your `schema.ts` file.
## Defining Your Schema
### Step 1: Define Node Types
Nodes represent entities in your graph. Each node type has a name and a Zod schema:
```typescript
import { z } from "zod";
import { defineNode } from "@nicia-ai/typegraph";
const Person = defineNode("Person", {
schema: z.object({
name: z.string().min(1),
email: z.string().email().optional(),
bio: z.string().optional(),
}),
});
const Project = defineNode("Project", {
schema: z.object({
name: z.string(),
description: z.string().optional(),
status: z.enum(["planning", "active", "completed"]),
}),
});
const Task = defineNode("Task", {
schema: z.object({
title: z.string(),
priority: z.enum(["low", "medium", "high"]),
completed: z.boolean().default(false),
}),
});
```
### Step 2: Define Edge Types
Edges represent relationships between nodes:
```typescript
import { defineEdge } from "@nicia-ai/typegraph";
const worksOn = defineEdge("worksOn", {
schema: z.object({
role: z.string().optional(),
since: z.string().optional(),
}),
});
const hasTask = defineEdge("hasTask", {
schema: z.object({}),
});
const assignedTo = defineEdge("assignedTo", {
schema: z.object({
assignedAt: z.string().optional(),
}),
});
// Unconstrained edge — connects any node to any node
const related = defineEdge("related");
```
### Step 3: Create the Graph Definition
Combine nodes, edges, and ontology into a graph:
```typescript
import { defineGraph, disjointWith } from "@nicia-ai/typegraph";
const graph = defineGraph({
id: "project_management",
nodes: {
Person: { type: Person },
Project: { type: Project },
Task: { type: Task },
},
edges: {
worksOn: { type: worksOn, from: [Person], to: [Project] },
hasTask: { type: hasTask, from: [Project], to: [Task] },
assignedTo: { type: assignedTo, from: [Task], to: [Person] },
related, // any→any
},
ontology: [
// A Person cannot be a Project or Task
disjointWith(Person, Project),
disjointWith(Person, Task),
disjointWith(Project, Task),
],
});
```
### Step 4: Create the Store
The store connects your graph definition to the database:
```typescript
import { createStore } from "@nicia-ai/typegraph";
const store = createStore(graph, backend);
```
#### Store Creation: Which Function to Use
| Function | Schema Handling | Use Case |
| -------------------------------- | ---------------------------------------------- | -------------------------------------------------- |
| `createLocalSqliteBackend` | Automatic | Quick start, development, tests (Node.js) |
| `createLibsqlBackend` | Automatic | libsql/Turso (Node.js, Workers, browser) |
| `createLocalPgliteBackend` | Automatic | In-process Postgres, embedded apps, pgvector tests |
| `createStore` + manual migration | None | When you manage migrations externally |
| `createStoreWithSchema` | Auto-creates tables, validates & auto-migrates | **Recommended for production** |
:::caution[Fulltext requires `createStoreWithSchema`]
If your graph has any `searchable()` fields, you must boot through
`createStoreWithSchema` once at startup. It durably materializes the
fulltext storage; bare `createStore()` is an attach-only path and throws
`StoreNotInitializedError` on the first fulltext operation. Graphs
without `searchable()` fields are unaffected.
:::
For production, use `createStoreWithSchema` to validate and auto-apply safe schema changes:
```typescript
import { createStoreWithSchema } from "@nicia-ai/typegraph";
const [store, result] = await createStoreWithSchema(graph, backend);
if (result.status === "initialized") {
console.log("Schema initialized at version", result.version);
} else if (result.status === "migrated") {
console.log(`Migrated from v${result.fromVersion} to v${result.toVersion}`);
}
// Other statuses: "unchanged", "pending", "breaking"
// See Schema Migrations for full details
```
#### Graph ID
Every graph has a unique `id` that scopes its data:
```typescript
const graph = defineGraph({
id: "my_app", // Scopes all nodes/edges to this graph
// ...
});
```
**Key behaviors:**
- All nodes and edges are stored with this `graph_id` in the database
- Multiple graphs can share the same database tables (isolated by `graph_id`)
- Changing the ID creates a new, empty graph (existing data is orphaned)
See [Multiple Graphs](/multiple-graphs) for multi-graph deployments.
## Working with Data
### Creating Nodes
```typescript
const alice = await store.nodes.Person.create({
name: "Alice Smith",
email: "alice@example.com",
});
const project = await store.nodes.Project.create({
name: "Website Redesign",
status: "active",
});
const task = await store.nodes.Task.create({
title: "Design mockups",
priority: "high",
});
```
### Creating Edges
Pass node objects directly to create edges:
```typescript
await store.edges.worksOn.create(alice, project, { role: "Lead Designer" });
await store.edges.hasTask.create(project, task, {});
await store.edges.assignedTo.create(task, alice, { assignedAt: new Date().toISOString() });
```
### Retrieving Nodes
```typescript
const person = await store.nodes.Person.getById(alice.id);
console.log(person?.name); // "Alice Smith"
```
### Updating Nodes
```typescript
const updated = await store.nodes.Task.update(task.id, { completed: true });
```
### Deleting Nodes
```typescript
await store.nodes.Task.delete(task.id);
```
## Querying Data
TypeGraph provides a fluent query builder:
```typescript
// Find all active projects
const activeProjects = await store
.query()
.from("Project", "p")
.whereNode("p", (p) => p.status.eq("active"))
.select((ctx) => ctx.p)
.execute();
// Find people working on a project
const teamMembers = await store
.query()
.from("Project", "p")
.traverse("worksOn", "e", { direction: "in" })
.to("Person", "person")
.select((ctx) => ({
project: ctx.p.name,
person: ctx.person.name,
}))
.execute();
// Multi-hop traversal: find tasks for a person
const myTasks = await store
.query()
.from("Person", "person")
.whereNode("person", (p) => p.name.eq("Alice Smith"))
.traverse("worksOn", "e1")
.to("Project", "project")
.traverse("hasTask", "e2")
.to("Task", "task")
.select((ctx) => ({
project: ctx.project.name,
task: ctx.task.title,
priority: ctx.task.priority,
}))
.execute();
```
`whereNode()` and `whereEdge()` constrain graph matches while traversal is
built. Use `.where((ctx) => ...)` when a condition should filter completed
rows, including optional or recursive results. Each successful traversal
combination is one row, so fanout can repeat a source entity; project an
identity and call relation `.distinct()` when the intended result is one row
per entity.
Traversal continues from the latest target by default. Reusable branching
fragments should state their source explicitly with `{ from: "alias" }` so
their behavior does not depend on which traversal preceded them. Direction,
ontology expansion, and temporal coordinates retain their ordinary query
defaults.
## Transactions
Group operations in transactions for atomicity:
```typescript
await store.transaction(async (tx) => {
const project = await tx.nodes.Project.create({
name: "New Feature",
status: "planning",
});
const task1 = await tx.nodes.Task.create({
title: "Research",
priority: "high",
});
const task2 = await tx.nodes.Task.create({
title: "Implementation",
priority: "medium",
});
await tx.edges.hasTask.create(project, task1, {});
await tx.edges.hasTask.create(project, task2, {});
});
```
## Error Handling
TypeGraph provides specific error types:
```typescript
import { ValidationError, NodeNotFoundError, DisjointError, RestrictedDeleteError } from "@nicia-ai/typegraph";
try {
await store.nodes.Person.create({ name: "" }); // Invalid: empty name
} catch (error) {
if (error instanceof ValidationError) {
console.log("Validation failed:", error.message);
}
}
try {
await store.nodes.Project.delete(project.id);
} catch (error) {
if (error instanceof RestrictedDeleteError) {
console.log("Cannot delete: edges exist");
}
}
```
## PostgreSQL Setup
TypeGraph also supports PostgreSQL for production deployments with better concurrency and JSON support.
For in-process Postgres during local development or tests, see
[PGlite in Backend Setup](/backend-setup#pglite-postgres-in-wasm).
### Installation
```bash
npm install @nicia-ai/typegraph zod drizzle-orm pg
npm install -D @types/pg
```
### Database Setup
```typescript
import { Pool } from "pg";
import { drizzle } from "drizzle-orm/node-postgres";
import { createPostgresBackend, generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
// Create connection pool
const pool = new Pool({
connectionString: process.env.DATABASE_URL,
max: 20, // Connection pool size
});
// Run TypeGraph migrations
await pool.query(generatePostgresMigrationSQL());
// Create Drizzle instance and backend
const db = drizzle(pool);
const backend = createPostgresBackend(db);
```
If you use `drizzle-kit` for migrations, see [Drizzle-Kit Managed Migrations](/integration#drizzle-kit-managed-migrations-recommended).
### PostgreSQL Advantages
- **JSONB**: Native JSON type with efficient indexing
- **Connection pooling**: Better concurrency handling
- **Partial indexes**: More efficient uniqueness constraints
- **Full transactions**: ACID guarantees across operations
### Using with Connection Pools
For production, always use connection pooling:
```typescript
import { Pool } from "pg";
const pool = new Pool({
connectionString: process.env.DATABASE_URL,
max: 20,
idleTimeoutMillis: 30000,
connectionTimeoutMillis: 2000,
});
// Graceful shutdown
process.on("SIGTERM", async () => {
await pool.end();
});
```
## Next Steps
- [Project Structure](/project-structure) - Organize your graph definitions as your project grows
- [Schemas & Types](/core-concepts) - Deep dive into nodes, edges, and schemas
- [Ontology](/ontology) - Learn about semantic relationships
- [Query Builder](/queries/overview) - Query patterns and traversals
- [Schemas & Stores](/schemas-stores) - Complete API documentation
# Schemas & Types
> Defining nodes, edges, and leveraging TypeScript inference
TypeGraph's power comes from its type system. Define your schema once with Zod, and get:
- **Runtime validation** on every create and update
- **TypeScript types** inferred automatically (no duplication)
- **Query builder constraints** that prevent invalid queries at compile time
## Contents
- [Nodes](#nodes) — Entities with properties and metadata
- [Defining Node Types](#defining-node-types)
- [Schema Features](#schema-features)
- [Node Operations](#node-operations)
- [Edges](#edges) — Relationships between nodes
- [Defining Edge Types](#defining-edge-types) (domain/range constraints)
- [Edge Constraints](#edge-constraints) (cardinality)
- [Edge Operations](#edge-operations)
- [Graph Definition](#graph-definition) — Combining nodes, edges, and ontology
- [Delete Behaviors](#delete-behaviors) — Restrict, cascade, disconnect
- [Uniqueness Constraints](#uniqueness-constraints) — Enforcing unique values
- [Type Inference](#type-inference) — Extracting TypeScript types from schemas
## Nodes
Nodes represent entities in your graph. Each node has:
- **Type**: The type of node (e.g., "Person", "Company")
- **ID**: A unique identifier within the graph
- **Props**: Properties defined by a Zod schema
- **Metadata**: Version, timestamps, and soft-delete state
### Defining Node Types
```typescript
import { z } from "zod";
import { defineNode } from "@nicia-ai/typegraph";
const Person = defineNode("Person", {
schema: z.object({
fullName: z.string().min(1),
email: z.string().email().optional(),
dateOfBirth: z.string().optional(),
tags: z.array(z.string()).default([]),
}),
description: "A person in the system", // Optional
});
```
### Schema Features
TypeGraph supports all Zod validation features:
```typescript
const Product = defineNode("Product", {
schema: z.object({
// Required string
name: z.string().min(1).max(200),
// Optional with default
status: z.enum(["draft", "active", "archived"]).default("draft"),
// Number with constraints
price: z.number().positive(),
// Array with items validation
categories: z.array(z.string()).min(1),
// Regex pattern
sku: z.string().regex(/^[A-Z]{2,4}-\d{4,8}$/),
// Nullable field
description: z.string().nullable(),
// Transform on validation
slug: z.string().transform((s) => s.toLowerCase().replace(/\s+/g, "-")),
}),
});
```
### Node Operations
```typescript
// Create with auto-generated ID
const node = await store.nodes.Person.create({ fullName: "Alice Smith" });
// Create with specific ID
const node = await store.nodes.Person.create({ fullName: "Alice Smith" }, { id: "person-alice" });
// Retrieve
const person = await store.nodes.Person.getById("person-alice");
// Update (partial)
const updated = await store.nodes.Person.update("person-alice", {
email: "alice@example.com",
});
// Delete (soft delete by default)
await store.nodes.Person.delete("person-alice");
// Hard delete (permanent removal) - use carefully!
await store.nodes.Person.hardDelete("person-alice");
```
### Node Object Shape
A node returned from the store has this structure:
```typescript
const alice = await store.nodes.Person.create({ name: "Alice", email: "a@example.com" });
// alice = {
// id: "01HX...", // Generated ULID (or your custom ID)
// kind: "Person", // The node type name
// name: "Alice", // Schema property (flattened to top level)
// email: "a@example.com", // Schema property
// meta: {
// version: 1,
// createdAt: "2024-01-15T10:30:00.000Z",
// updatedAt: "2024-01-15T10:30:00.000Z",
// deletedAt: undefined,
// validFrom: "2024-01-15T10:30:00.000Z", // defaults to createdAt when omitted,
// // unless a stated past validTo makes
// // the row "born already ended" (undefined)
// validTo: undefined,
// }
// }
```
Schema properties are flattened to the top level for ergonomic access (`alice.name` instead of
`alice.props.name`). System metadata lives under `meta`.
### Soft Delete vs Hard Delete
By default, `delete()` performs a **soft delete**—it sets the `deletedAt` timestamp but preserves the record:
```typescript
await store.nodes.Person.delete(alice.id); // Sets deletedAt, keeps the record
```
For permanent removal, use `hardDelete()`:
```typescript
await store.nodes.Person.hardDelete(alice.id); // Removes from database
```
**When to use each:**
| Method | Use Case |
|--------|----------|
| `delete()` | Standard deletions, audit trails, undo capability |
| `hardDelete()` | GDPR erasure, storage cleanup, removing test data |
**Warning:** `hardDelete()` is irreversible. It also removes associated uniqueness entries and
embeddings. Consider using soft delete for most use cases.
## Edges
Edges represent relationships between nodes. Each edge has:
- **Type**: The type of relationship (e.g., "worksAt", "knows")
- **ID**: A unique identifier
- **From**: Source node (type + ID)
- **To**: Target node (type + ID)
- **Props**: Properties defined by a Zod schema
### Defining Edge Types
```typescript
import { defineEdge } from "@nicia-ai/typegraph";
// Edge with properties
const worksAt = defineEdge("worksAt", {
schema: z.object({
role: z.string(),
startDate: z.string().optional(),
isPrimary: z.boolean().default(true),
}),
});
// Edge without properties
const knows = defineEdge("knows");
// Equivalent to: defineEdge("knows", { schema: z.object({}) })
```
#### Unconstrained Edges
Edges defined without `from` and `to` are **unconstrained** — they can connect any
node type to any node type. When used directly in `defineGraph`, they are automatically
allowed for all node types in the graph:
```typescript
const sameAs = defineEdge("sameAs");
const related = defineEdge("related", {
schema: z.object({ reason: z.string() }),
});
const graph = defineGraph({
id: "my_graph",
nodes: {
Person: { type: Person },
Company: { type: Company },
},
edges: {
sameAs, // any→any (Person↔Person, Person↔Company, Company↔Company)
related, // any→any, with properties
worksAt: { type: worksAt, from: [Person], to: [Company] }, // constrained
},
});
// All of these work:
await store.edges.sameAs.create(alice, bob, {}); // Person→Person
await store.edges.sameAs.create(alice, acme, {}); // Person→Company
await store.edges.sameAs.create(acme, alice, {}); // Company→Person
```
This is useful for semantic relationships like `sameAs`, `seeAlso`, `related`, or
`tagged` that apply broadly across node types.
#### Domain and Range Constraints
Edges can include built-in domain (source types) and range (target types) constraints
directly in their definition. This makes edge definitions self-contained and reusable:
```typescript
// Edge with built-in domain/range constraints
const worksAt = defineEdge("worksAt", {
schema: z.object({
role: z.string(),
startDate: z.string().optional(),
}),
from: [Person], // Domain: only Person can be the source
to: [Company], // Range: only Company can be the target
});
// Edge connecting multiple types
const mentions = defineEdge("mentions", {
from: [Article, Comment],
to: [Person, Company, Topic],
});
```
Any edge type can be used directly in `defineGraph` without an `EdgeRegistration`
wrapper. Constrained edges use their built-in `from`/`to`; unconstrained edges
allow all node types:
```typescript
const graph = defineGraph({
nodes: { Person: { type: Person }, Company: { type: Company } },
edges: {
worksAt, // Constrained - uses built-in from/to
sameAs, // Unconstrained - connects any node to any node
},
});
```
You can still use `EdgeRegistration` to narrow (but not widen) the constraints:
```typescript
const worksAt = defineEdge("worksAt", {
from: [Person],
to: [Company, Subsidiary], // Allows both Company and Subsidiary
});
const graph = defineGraph({
edges: {
// Narrow to only Subsidiary targets in this graph
worksAt: { type: worksAt, from: [Person], to: [Subsidiary] },
},
});
```
Attempting to widen beyond the edge's built-in constraints throws a `ConfigurationError`:
```typescript
const worksAt = defineEdge("worksAt", {
from: [Person],
to: [Company],
});
// This throws ConfigurationError - OtherEntity is not in the edge's range
defineGraph({
edges: {
worksAt: { type: worksAt, from: [Person], to: [OtherEntity] },
},
});
```
#### Source-Dependent Targets
An array-valued `to` allows every combination of the source and target types.
When the valid target depends on the source, use a map instead:
```typescript
const Employee = defineNode("Employee", { schema: z.object({ name: z.string() }) });
const Student = defineNode("Student", { schema: z.object({ name: z.string() }) });
const Department = defineNode("Department", { schema: z.object({ name: z.string() }) });
const Course = defineNode("Course", { schema: z.object({ name: z.string() }) });
const assignedTo = defineEdge("assignedTo", {
from: [Employee, Student],
to: {
Employee: [Department],
Student: [Course],
},
});
const graph = defineGraph({
id: "assignments",
nodes: {
Employee: { type: Employee },
Student: { type: Student },
Department: { type: Department },
Course: { type: Course },
},
edges: { assignedTo },
});
```
This permits `Employee → Department` and `Student → Course`. It rejects
`Employee → Course` and `Student → Department`. Using
`to: [Department, Course]` would permit all four combinations.
Map keys are the literal node kind names (`Employee.kind`), not aliases used to
register nodes in a graph. Every kind in `from` must have a map entry, no other
keys are allowed, and each target array must be nonempty. You can also use
computed keys such as `[Employee.kind]: [Department]`.
The map syntax works in an explicit graph registration too:
```typescript
edges: {
assignedTo: {
type: assignedTo,
from: [Employee],
to: { Employee: [Department] },
},
}
```
A registration may narrow the built-in allowed pairs, but it cannot introduce
new pairs. Replacing a map with arrays is valid only when every resulting
combination is already allowed by the edge definition.
At runtime, both endpoints must match the **same** declared pair, including
`subClassOf` assignability. A source matching several source entries can use the
targets allowed by any of those entries. An undeclared pair fails with
[`EndpointPairError`](/errors#endpointpairerror); an invalid source kind still
fails with `EndpointError`. Malformed declarations fail with `ConfigurationError`.
Typed collection writes preserve the source/target relationship; dynamic writes
and imports enforce it at runtime. Bulk writes reject invalid pairs atomically.
Import pair validation remains active even when reference validation is disabled;
imports retain their own documented error-handling and partial-success behavior.
See [collection types](/types#typededgecollectionr) for inference limits,
[graph extensions](/graph-extensions#edges) for runtime declarations, and
[schema management](/schema-management#endpoint-pair-changes) for schema changes.
### Edge Constraints
#### Cardinality
Control how many edges can exist:
```typescript
const graph = defineGraph({
edges: {
// Default: no limit
knows: { type: knows, from: [Person], to: [Person], cardinality: "many" },
// At most one edge of this type from any source node
currentEmployer: {
type: currentEmployer,
from: [Person],
to: [Company],
cardinality: "one",
},
// At most one edge between any (source, target) pair
rated: { type: rated, from: [Person], to: [Product], cardinality: "unique" },
// At most one active edge (valid_to IS NULL) from any source
currentRole: {
type: currentRole,
from: [Person],
to: [Company],
cardinality: "oneActive",
},
},
});
```
| Cardinality | Description |
|-------------|-------------|
| `"many"` | No limit (default) |
| `"one"` | At most one edge of this type from any source node |
| `"unique"` | At most one edge between any (source, target) pair |
| `"oneActive"` | At most one edge with `valid_to IS NULL` from any source |
#### Enforcement Timing
Cardinality constraints are checked at edge **creation time**, before the insert:
```typescript
// With cardinality: "one" on currentEmployer:
await store.edges.currentEmployer.create(alice, acme, {}); // OK
await store.edges.currentEmployer.create(alice, other, {}); // Throws CardinalityError
```
The check queries existing edges and throws `CardinalityError` if violated.
For `oneActive`, only edges with `validTo` unset count toward the limit.
### Edge Operations
```typescript
// Create edge - pass nodes directly
const edge = await store.edges.worksAt.create(alice, acme, { role: "Engineer" });
// Retrieve edge
const e = await store.edges.worksAt.getById(edge.id);
// Delete edge
await store.edges.worksAt.delete(edge.id);
```
## Graph Definition
The graph definition combines all components:
```typescript
import { defineGraph } from "@nicia-ai/typegraph";
const graph = defineGraph({
// Unique identifier for this graph
id: "my_application",
// Node registrations
nodes: {
Person: {
type: Person,
onDelete: "restrict", // Default behavior
},
Company: {
type: Company,
onDelete: "cascade",
},
Employment: {
type: Employment,
onDelete: "disconnect",
},
},
// Edge registrations
edges: {
worksAt: {
type: worksAt,
from: [Person],
to: [Company],
cardinality: "many",
},
employedAt: {
type: employedAt,
from: [Company],
to: [Employment],
cardinality: "many",
},
},
// Semantic relationships
ontology: [subClassOf(Company, Organization), disjointWith(Person, Company)],
});
```
## Delete Behaviors
Control what happens when nodes are deleted:
### Restrict (Default)
Blocks deletion if any edges are connected:
```typescript
nodes: {
Author: { type: Author }, // onDelete defaults to "restrict"
}
// This throws RestrictedDeleteError if Author has edges
await store.nodes.Author.delete(authorId);
```
### Cascade
Automatically deletes all connected edges:
```typescript
nodes: {
Book: { type: Book, onDelete: "cascade" },
}
// Deletes the book and all edges connected to it
await store.nodes.Book.delete(bookId);
```
### Disconnect
Soft-deletes edges (preserves history):
```typescript
nodes: {
Review: { type: Review, onDelete: "disconnect" },
}
// Marks connected edges as deleted (deleted_at is set)
await store.nodes.Review.delete(reviewId);
```
## Uniqueness Constraints
Ensure unique values within node types:
```typescript
const graph = defineGraph({
nodes: {
Person: {
type: Person,
unique: [
{
name: "person_email",
fields: ["email"],
where: (props) => props.email.isNotNull(),
scope: "kind",
collation: "caseInsensitive",
},
],
},
Company: {
type: Company,
unique: [
{
name: "company_ticker",
fields: ["ticker"],
scope: "kind",
collation: "binary",
},
],
},
},
});
```
### Scope Options
- `"kind"`: Unique within this exact type only
- `"kindWithSubClasses"`: Unique across this type and all subclasses
### Collation Options
- `"binary"`: Case-sensitive comparison
- `"caseInsensitive"`: Case-insensitive comparison
## Type Inference
TypeGraph infers TypeScript types from Zod schemas—you never duplicate type definitions.
### Extracting Types from Definitions
```typescript
import { z } from "zod";
import { defineNode, type Node, type NodeProps, type NodeId } from "@nicia-ai/typegraph";
const Person = defineNode("Person", {
schema: z.object({
name: z.string(),
email: z.string().email().optional(),
age: z.number().optional(),
}),
});
// For functions that work with full nodes (id, kind, metadata, props):
type PersonNode = Node;
// { id: NodeId; kind: "Person"; name: string; email?: string; version: number; createdAt: Date; ... }
// For functions that only need the property data:
type PersonProps = NodeProps;
// { name: string; email?: string; age?: number }
// For type-safe node IDs (prevents mixing IDs from different node types):
type PersonId = NodeId;
// string & { readonly [__nodeId]: typeof Person }
```
Use `Node` when your function needs the full node with metadata.
Use `NodeProps` when you only care about the schema properties (e.g., for form validation or API payloads).
### Typed Store Operations
```typescript
// Create returns a fully typed Node
const alice: Node = await store.nodes.Person.create({
name: "Alice",
email: "alice@example.com",
});
// TypeScript knows the structure
alice.id; // NodeId - branded string
alice.name; // string
alice.email; // string | undefined
alice.age; // number | undefined
alice.version; // number
alice.createdAt; // Date
// Type errors caught at compile time
await store.nodes.Person.create({
name: 123, // Error: Type 'number' is not assignable to type 'string'
invalid: "field", // Error: Object literal may only specify known properties
});
```
### Typed Query Results
```typescript
// Result type is inferred from your select projection
const results = await store
.query()
.from("Person", "p")
.select((ctx) => ({
name: ctx.p.name, // TypeScript knows: string
email: ctx.p.email, // TypeScript knows: string | undefined
id: ctx.p.id, // TypeScript knows: NodeId
}))
.execute();
// results: Array<{ name: string; email: string | undefined; id: NodeId }>
// Invalid property access is caught
.select((ctx) => ({
invalid: ctx.p.nonexistent, // TypeScript error!
}))
```
### Typed Edge Operations
Edge endpoints are constrained to valid node types:
```typescript
// Edge definition: worksAt goes from Person → Company
const graph = defineGraph({
// ...
edges: {
worksAt: { type: worksAt, from: [Person], to: [Company] },
},
});
// TypeScript enforces valid endpoints
await store.edges.worksAt.create(alice, acmeCorp, { role: "Engineer" }); // OK
await store.edges.worksAt.create(acmeCorp, alice, { role: "Engineer" });
// Error: Argument of type 'Node' is not assignable to parameter of type 'Node'
```
# Backend Setup
> Configure SQLite and PostgreSQL backends for TypeGraph
TypeGraph stores graph data in your existing relational database using Drizzle ORM adapters.
This guide covers setting up SQLite, PostgreSQL, and PGlite backends.
:::note[Custom indexes]
TypeGraph migrations create the core tables and built-in indexes. For application-specific indexes
on JSON properties (and Drizzle/drizzle-kit integration), see [Indexes](/performance/indexes).
:::
## SQLite
SQLite is ideal for development, testing, single-server deployments, and embedded applications.
### Quick Setup
For development and testing, use the convenience function that owns the
connection and provisions TypeGraph's base tables:
```typescript
import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local";
import { createStore } from "@nicia-ai/typegraph";
// In-memory database (resets on restart)
const { backend } = createLocalSqliteBackend();
const store = createStore(graph, backend);
// File-based database (persisted)
const { backend, db } = createLocalSqliteBackend({ path: "./app.db" });
const store = createStore(graph, backend);
```
The local backend owns its connection, so it applies performance pragmas at
open: `journal_mode=WAL`, `synchronous=NORMAL`, and a 5s `busy_timeout`. On
file databases this makes single-operation writes roughly 5× faster than the
driver defaults (rollback journal, `synchronous=FULL`). Override individual
values or opt out entirely:
```typescript
// Override one value, keep the other defaults
createLocalSqliteBackend({ path: "./app.db", pragmas: { busyTimeoutMs: 10_000 } });
// Keep better-sqlite3's driver defaults untouched
createLocalSqliteBackend({ path: "./app.db", pragmas: false });
```
:::caution[Fulltext and embeddings require `createStoreWithSchema`]
`createLocalSqliteBackend` creates the base tables but does not durably
materialize strategy-owned storage. If your graph has `searchable()` or
`embedding()` fields, boot with
`const [store] = await createStoreWithSchema(graph, backend);` instead of
bare `createStore()` — otherwise the first fulltext or embedding operation
throws `StoreNotInitializedError`.
:::
### Manual Setup
For full control over the database connection:
```typescript
import Database from "better-sqlite3";
import { drizzle } from "drizzle-orm/better-sqlite3";
import { createSqliteBackend, generateSqliteMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
import { createStoreWithSchema } from "@nicia-ai/typegraph";
// Create and configure the database
const sqlite = new Database("app.db");
sqlite.pragma("journal_mode = WAL"); // Recommended for performance
sqlite.pragma("foreign_keys = ON");
// Create Drizzle instance and backend
const db = drizzle(sqlite);
const backend = createSqliteBackend(db);
// createStoreWithSchema auto-creates tables on first run
const [store] = await createStoreWithSchema(graph, backend);
// Clean up when done
process.on("exit", () => sqlite.close());
```
For a fresh database whose DDL is managed externally, use
`generateSqliteMigrationSQL()` with `createStore()` instead:
```typescript
sqlite.exec(generateSqliteMigrationSQL());
const store = createStore(graph, backend);
```
The generated script is complete installation DDL and stamps the current
deployment-wide base-schema marker last; it is not an incremental upgrade
planner. Existing databases attached only through the zero-DDL runtime
factories must apply release-specific additive migrations through their
migration tool. See
[Upgrading deployment-wide base storage](#upgrading-deployment-wide-base-storage)
for the exact SQLite and PostgreSQL statements. A privileged
`createStoreWithSchema()` open adopts missing release storage once, then stamps
a deployment-wide base-schema marker. Warm opens read that marker and issue no
base-adoption DDL.
### SQLite with Vector Search
For semantic search, use the sqlite-vec extension. `createLocalSqliteBackend()` wires the
`sqliteVecStrategy` automatically when the extension loads. For a bring-your-own connection, load the
extension and pass the strategy explicitly:
```typescript
import Database from "better-sqlite3";
import { drizzle } from "drizzle-orm/better-sqlite3";
import { createSqliteBackend, generateSqliteMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
import { sqliteVecStrategy } from "@nicia-ai/typegraph";
const sqlite = new Database("app.db");
// Load sqlite-vec extension
sqlite.loadExtension("vec0");
// Run migrations (core tables)
sqlite.exec(generateSqliteMigrationSQL());
const db = drizzle(sqlite);
const backend = createSqliteBackend(db, { vector: sqliteVecStrategy });
```
sqlite-vec stores embeddings in `vec0` virtual tables and supports the `cosine` and `l2` metrics. Per-field
vector tables are provisioned by `createStoreWithSchema` at boot (not by the generated migration SQL), and the
runtime asserts a durable marker rather than issuing DDL on first write — see
[Database roles & least privilege](#database-roles--least-privilege).
See [Semantic Search](/semantic-search) for query examples.
### libsql / Turso
For edge deployments, shared-driver setups, or Turso cloud databases, use the first-class
libsql backend:
```bash
npm install @libsql/client
```
```typescript
import { createClient } from "@libsql/client";
import { createLibsqlBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql";
import { createStore } from "@nicia-ai/typegraph";
// Local file
const client = createClient({ url: "file:app.db" });
// Or remote Turso database
// const client = createClient({ url: "libsql://my-db.turso.io", authToken: "..." });
const { backend, db } = await createLibsqlBackend(client);
const store = createStore(graph, backend);
```
`createLibsqlBackend` handles DDL execution and configures the correct async
execution profile automatically. It returns both the `backend` and the underlying
Drizzle `db` instance for direct SQL access. The caller retains ownership of the
client and is responsible for closing it when done — this allows sharing a single
client across TypeGraph and other libraries. Its installation is complete: the
factory publishes the deployment-wide base-schema marker, and when it encounters
a pre-0.52 edge table it applies the focused match-identity storage adoption
before retrying the idempotent installation script. The local SQLite factory has
the same behavior.
The libsql backend has native vector and hybrid search, wired automatically via `libsqlVectorStrategy` — no
extension to load. It uses libSQL's built-in engine (`F32_BLOB(N)` storage, `vector_distance_cos` /
`vector_distance_l2`, and DiskANN approximate nearest neighbor via `libsql_vector_idx` + `vector_top_k`) and
supports the `cosine` and `l2` metrics. See [Semantic Search](/semantic-search) for query examples.
:::caution[In-memory databases and transactions]
libsql's `file::memory:` creates a separate database per connection. Since transactions
open a new connection, the original database is destroyed after a transaction completes
([tursodatabase/libsql-client-ts#229](https://github.com/tursodatabase/libsql-client-ts/issues/229)).
Use a file-based database (`file:path.db`) or remote URL when transactions are needed.
:::
### API Reference
#### `createLocalSqliteBackend(options?)`
Creates a SQLite backend with automatic database and schema setup.
```typescript
function createLocalSqliteBackend(options?: {
path?: string; // Database path, defaults to ":memory:"
tables?: SqliteTables;
/**
* Override the fulltext strategy. Defaults to `fts5Strategy` (SQLite's
* built-in FTS5 virtual table). Pass `false` to disable fulltext support
* entirely — the backend then advertises no `capabilities.fulltext` and
* omits the fulltext CRUD/search methods, and the managed installation
* never creates the fulltext table. Forwarded to both the installation
* DDL and `createSqliteBackend`.
*/
fulltext?: FulltextStrategy | false;
}): { backend: GraphBackend; db: BetterSQLite3Database };
```
#### `createSqliteBackend(db, options?)`
Creates a SQLite backend from an existing Drizzle database instance. Pass `vector` to enable vector search
(for example `sqliteVecStrategy` after loading the sqlite-vec extension).
```typescript
function createSqliteBackend(
db: BetterSQLite3Database,
options?: {
tables?: SqliteTables;
/**
* Override the fulltext strategy. Defaults to `fts5Strategy` (SQLite's
* built-in FTS5 virtual table). Pass `false` to disable fulltext
* support entirely — the backend then advertises no
* `capabilities.fulltext` and omits the fulltext CRUD/search methods,
* mirroring `vector` left unset. Required for a SQLite build without
* FTS5 compiled in.
*/
fulltext?: FulltextStrategy | false;
vector?: VectorStrategy;
capabilities?: BundledBackendCapabilityOverrides;
},
): GraphBackend;
```
Pass `{ fulltext: false }` on a SQLite build without FTS5 compiled in, or
whenever the graph has no `searchable()` fields and you would rather skip
the virtual table than carry it unused:
```typescript
const backend = createSqliteBackend(db, { fulltext: false });
```
#### `generateSqliteMigrationSQL()`
Returns complete fresh-installation SQL for creating TypeGraph tables and
stamping the current deployment-wide base-schema marker in SQLite.
```typescript
function generateSqliteMigrationSQL(
tables?: SqliteTables,
fulltextStrategy?: FulltextStrategy | false,
): string;
```
`generateSqliteDDL()` is the lower-level table/index statement array used by
backend bootstrap. It deliberately omits the deployment-wide marker row and is
therefore not a complete installation script. Use `generateSqliteMigrationSQL()`
when the resulting database will be opened through `createVerifiedStore()` or
the DML-only graph-template APIs.
#### `createLibsqlBackend(client, options?)`
Creates a SQLite backend from a `@libsql/client` instance. Runs DDL automatically.
The caller retains ownership of the client and is responsible for closing it.
```typescript
async function createLibsqlBackend(client: Client, options?: { tables?: SqliteTables }): Promise<{ backend: GraphBackend; db: LibSQLDatabase }>;
```
## PostgreSQL
PostgreSQL is recommended for production deployments with concurrent access, large datasets,
or when you need advanced features like pgvector.
`createPostgresBackend` is driver-agnostic. Pick the Drizzle adapter that matches your
runtime, and TypeGraph works the same way against each.
### Choosing a PostgreSQL driver
| Runtime | Recommended driver | Drizzle adapter |
| --------------------------------------------------------------------------------- | ---------------------------------------------------------------------- | -------------------------------------------------------- |
| Long-lived Node server (Fly, Render, Cloud Run, containers) | `pg` (node-postgres) or `postgres` (postgres-js) | `drizzle-orm/node-postgres` or `drizzle-orm/postgres-js` |
| Node serverless (Vercel Functions, AWS Lambda, Netlify Functions) | `postgres` (postgres-js) — faster cold start, lower per-query overhead | `drizzle-orm/postgres-js` |
| Bun server | `postgres` (postgres-js) or Bun's built-in SQL | `drizzle-orm/postgres-js` or `drizzle-orm/bun-sql` |
| Edge runtime (Cloudflare Workers, Vercel Edge, Netlify Edge) — needs transactions | `@neondatabase/serverless` Pool over WebSockets | `drizzle-orm/neon-serverless` |
| Edge runtime — single-statement reads/writes only | `@neondatabase/serverless` `neon(url)` over HTTP | `drizzle-orm/neon-http` |
| Cloudflare Hyperdrive | `pg` or `postgres` (through the Hyperdrive pooler) | `drizzle-orm/node-postgres` or `drizzle-orm/postgres-js` |
| Embedded apps, local development, Postgres dialect tests | `@electric-sql/pglite` | `drizzle-orm/pglite` |
:::note[Neon HTTP vs WebSocket]
Both Neon drivers work with TypeGraph. They have different tradeoffs:
- **`drizzle-orm/neon-http`** uses HTTP per statement. Lowest cold-start cost; survives Workers'
per-request isolation. **Cannot hold a session across statements**, so multi-statement transactions
are unavailable — TypeGraph auto-detects this driver and sets `capabilities.execution.interactiveTransactions = false`,
so `store.transaction(...)` refuses rather than pretending to provide rollback. Eligible
atomic-batch operations remain available when the transport is certified for them.
A schema-managed Store's write fuses its schema fence into the write's own statement when the
write fuses, and fails closed otherwise — see
[The guard every fused write shares](/limitations#the-guard-every-fused-write-shares) for which
writes fuse and the reasons a write that cannot refuses with.
- **`drizzle-orm/neon-serverless`** uses a WebSocket Pool. Holds a session, supports full transactional
semantics, but the WebSocket connection lifecycle needs care in serverless / per-request contexts
(you typically want a fresh Pool per request).
Pick HTTP for stateless reads and for the fused schema-managed writes. Pick WebSockets for schema
migrations, and for any write outside that fused envelope.
:::
### node-postgres (pg)
The default choice for long-lived Node servers. Widest ecosystem and most deployment
documentation.
```typescript
import { Pool } from "pg";
import { drizzle } from "drizzle-orm/node-postgres";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import { createStoreWithSchema } from "@nicia-ai/typegraph";
const pool = new Pool({
connectionString: process.env.DATABASE_URL,
max: 20,
});
const db = drizzle(pool);
const backend = createPostgresBackend(db);
const [store] = await createStoreWithSchema(graph, backend);
```
For a fresh database managed externally, use `generatePostgresMigrationSQL()` with `createStore()`:
```typescript
import { generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
await pool.query(generatePostgresMigrationSQL());
const store = createStore(graph, backend);
```
As with SQLite, this is complete installation DDL rather than an incremental
upgrade plan. Apply the
[base-schema upgrade](#upgrading-deployment-wide-base-storage)
to an existing database, or let a privileged `createStoreWithSchema()`
preparation adopt the storage before runtime workers use `createStore()`.
### postgres-js
A leaner Postgres client with lower per-query overhead and smaller bundle size. Good
default for Node serverless platforms and Bun. Fully tested against TypeGraph's adapter
and integration suites.
```bash
npm install postgres drizzle-orm
```
```typescript
import postgres from "postgres";
import { drizzle } from "drizzle-orm/postgres-js";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import { createStoreWithSchema } from "@nicia-ai/typegraph";
const sql = postgres(process.env.DATABASE_URL, {
max: 10,
idle_timeout: 30,
});
const db = drizzle(sql);
const backend = createPostgresBackend(db);
const [store] = await createStoreWithSchema(graph, backend);
```
Transactions go through `sql.begin(fn)`; TypeGraph handles this automatically via
Drizzle's `db.transaction()`. Isolation levels are honored the same way as with
node-postgres.
### Neon serverless (WebSockets)
For edge runtimes like Cloudflare Workers, Vercel Edge, and Netlify Edge — anywhere
native TCP sockets aren't available. Neon's `@neondatabase/serverless` driver speaks
the Postgres wire protocol over WebSockets and exposes a pg-Pool-compatible API.
```bash
npm install @neondatabase/serverless drizzle-orm
```
```typescript
import { Pool } from "@neondatabase/serverless";
import { drizzle } from "drizzle-orm/neon-serverless";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import { createStoreWithSchema } from "@nicia-ai/typegraph";
const pool = new Pool({ connectionString: env.NEON_DATABASE_URL });
const db = drizzle(pool);
const backend = createPostgresBackend(db);
const [store] = await createStoreWithSchema(graph, backend);
```
When running under Node.js (for local testing), install `ws` and configure it once
before connecting:
```typescript
import { neonConfig } from "@neondatabase/serverless";
import ws from "ws";
neonConfig.webSocketConstructor = ws;
```
Edge runtimes expose `WebSocket` globally and need no extra setup.
### Neon HTTP
For stateless edge workloads where you don't need transactional writes. The HTTP
driver issues one request per query — lowest cold-start cost, no session lifecycle
to manage. TypeGraph auto-detects this driver and sets
`capabilities.execution.interactiveTransactions` to `false` and
`capabilities.execution.unitOfWork` to `"batch"`. On a raw Store,
`store.transaction(...)` refuses rather than silently falling through to
sequential execution.
A schema-managed or verified Store's first write does not universally fail
closed here — it depends on whether the write fuses. A singleton node
create, update, `upsertById`, or delete fuses on a kind with no declared
unique constraint (a create takes a generated or a caller-supplied id) —
except a node delete, which fuses even when the kind DOES carry a declared
unique constraint, because the atomic delete program releases that claim in
the same statement. A singleton edge create fuses when the kind's
cardinality is `"many"`, and edge update and delete fuse the same way
(`EdgeCollection` has no `upsertById`). So do
`bulkInsert`/`bulkCreate`/`bulkDelete`/`bulkReplaceById`/`bulkUpsertById`,
and a constrained write inside an atomic program's claim envelope. Each of
these asserts the active schema version inside the statements neon-http
submits together, and `transaction(queries)` commits or rejects that
submission as a whole. A write that cannot fuse either fails closed with
`BATCH_WRITE_UNSUPPORTED` naming a proven reason (an interactive callback, a
probe-then-write constraint check, Operational Identity, history, or a
schema commit), or — for a write that simply doesn't fit the fused shape,
such as a singleton create, update, or `upsertById` on a uniquely-constrained
kind, or a supplied-id tombstone resurrection — fails closed with the plain
`SCHEMA_WRITE_FENCE_UNSUPPORTED` limitation and no named reason. See
[The guard every fused write shares](/limitations#the-guard-every-fused-write-shares)
for the shared guard and the full reason table.
Schema commits stay refused regardless: `commitSchemaVersion` and
`setActiveVersion` require holding one transaction across their
compare-and-swap read and activating write to eliminate the orphan-row crash
window they exist to fix, so they refuse with a typed `ConfigurationError` on
non-transactional backends. Run schema migrations from a process with a
transactional driver (`drizzle-orm/neon-serverless`, regular `pg`, etc.); the
edge worker can keep using neon-http for reads and for the fused writes
above. A raw `createStore()` remains available for writes outside that
envelope when the application explicitly accepts they are not fenced against
schema changes.
```bash
npm install @neondatabase/serverless drizzle-orm
```
```typescript
import { neon } from "@neondatabase/serverless";
import { drizzle } from "drizzle-orm/neon-http";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import { createStore } from "@nicia-ai/typegraph";
const sql = neon(env.NEON_DATABASE_URL);
const db = drizzle({ client: sql });
const backend = createPostgresBackend(db);
const store = createStore(graph, backend);
// backend.capabilities.execution.interactiveTransactions === false (auto-detected)
```
Use `neon-http` for reads and for the fused schema-managed writes listed
above. Run schema migrations, and any write outside that envelope, through
`neon-serverless`, regular `pg`, or another transactional driver.
### PGlite (Postgres-in-WASM)
[PGlite](https://pglite.dev/) is a full Postgres compiled to WebAssembly that runs
in-process — in Node, Bun, Deno, or the browser — with no server and no native
addon. It's ideal for local development, embedded apps, and running the real
Postgres dialect (including pgvector) in tests without Docker.
`@electric-sql/pglite` is an optional peer dependency. Vector support additionally
needs `@electric-sql/pglite-pgvector` (PGlite ≥ 0.5 ships pgvector as a separate
package):
```bash
npm install @electric-sql/pglite @electric-sql/pglite-pgvector
```
The batteries-included helper constructs the engine, loads pgvector, runs the
schema DDL, and returns a ready backend — the Postgres analog of
`createLocalSqliteBackend`:
```typescript
import { createLocalPgliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres/pglite";
import { createStore } from "@nicia-ai/typegraph";
// In-memory by default, with pgvector enabled.
const { backend, db, client } = await createLocalPgliteBackend();
const store = createStore(graph, backend);
// backend.close() disposes the PGlite engine.
```
```typescript
// Persistent on disk:
const { backend } = await createLocalPgliteBackend({ dataDir: "./pgdata" });
// No embeddings? Skip the extension (no pgvector dependency needed):
const { backend } = await createLocalPgliteBackend({ vector: false });
// Pass an explicit pgvector extension object:
import { vector } from "@electric-sql/pglite-pgvector";
const { backend } = await createLocalPgliteBackend({ vector });
```
If you construct PGlite yourself, pass its Drizzle database straight to
`createPostgresBackend` — the execution fast path detects PGlite and routes it
correctly:
```typescript
import { PGlite } from "@electric-sql/pglite";
import { vector } from "@electric-sql/pglite-pgvector";
import { drizzle } from "drizzle-orm/pglite";
import { createPostgresBackend, generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
const client = await PGlite.create({ extensions: { vector } });
await client.exec(generatePostgresMigrationSQL());
const backend = createPostgresBackend(drizzle(client));
```
PGlite is single-connection and serial: there is no pooling, so concurrent
`store.transaction()` calls queue rather than run in parallel. It complements,
rather than replaces, a Docker-based Postgres for CI — PGlite exercises the SQL
dialect and pgvector, but not driver-specific behavior (node-postgres statement
naming, postgres-js, pgbouncer, real concurrency).
### PostgreSQL with Vector Search
For semantic search, enable pgvector. `createPostgresBackend` defaults to `pgvectorStrategy`, so no extra
wiring is required:
```typescript
import { Pool } from "pg";
import { drizzle } from "drizzle-orm/node-postgres";
import { createPostgresBackend, generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
const pool = new Pool({ connectionString: process.env.DATABASE_URL });
// Migration SQL enables the pgvector extension
await pool.query(generatePostgresMigrationSQL());
// Runs: CREATE EXTENSION IF NOT EXISTS vector;
const db = drizzle(pool);
const backend = createPostgresBackend(db);
```
pgvector stores embeddings in per-field typed `vector(N)` tables (provisioned by `createStoreWithSchema` at boot
— the generated migration SQL creates no embedding table) with HNSW or IVFFlat indexes, and supports the
`cosine`, `l2`, and `inner_product` metrics.
See [Semantic Search](/semantic-search) for query examples.
### Refreshing planner statistics after bulk loads
`importGraph()` refreshes planner statistics automatically after an import
that created or updated rows, and `store.materializeIndexes()` does the
same on SQLite after creating indexes (pass `refreshStatistics: false` to
opt out). On PostgreSQL, `materializeIndexes()` builds with
`CREATE INDEX CONCURRENTLY` and skips the automatic refresh — call
`store.refreshStatistics()` after materializing.
`bulkCreate` and `bulkInsert` on nodes and edges also refresh
automatically when a single autocommit call writes 1,000 rows or more. Tune or disable this
with the `autoRefreshStatistics` store option:
```typescript
// Refresh after any autocommit bulkCreate of 5,000+ rows
const store = createStore(graph, backend, { autoRefreshStatistics: 5000 });
// Never refresh automatically after bulkCreate
const store = createStore(graph, backend, { autoRefreshStatistics: false });
```
Bulk writes inside a `store.transaction(...)` block never auto-refresh —
statistics collected mid-transaction cannot see the uncommitted rows —
so refresh manually after the transaction commits. The same applies to
loops of small `bulkCreate` batches that never individually reach the
threshold, and to backend-level batch inserts — the loop example below
covers that pattern.
PostgreSQL's query planner relies on table statistics to choose
between multi-column indexes on `typegraph_edges` (forward vs reverse vs
cardinality), and when those statistics are stale the planner can pick a
reverse-index scan with a filter — turning a 0.5ms forward traversal into a
5ms one. SQLite's planner is similarly sensitive: without `sqlite_stat1`
data, some FTS5 fulltext queries fall back to a plan that's roughly 30×
slower. Autovacuum / background statistics collection will catch up
eventually, but refreshing explicitly gives correct latencies immediately.
```typescript
for (const batch of batches) {
await store.nodes.Document.bulkCreate(batch);
}
await store.refreshStatistics();
```
The implementation runs `ANALYZE` against the TypeGraph-managed tables in
the configured backend — the call is safe regardless of custom table names
or fulltext / embedding configuration. Cloudflare D1 and Durable Object SQLite
reject the performance-only `PRAGMA analysis_limit` tuning statement through
their authorizer. TypeGraph recognizes only that `SQLITE_AUTH` failure and
continues with scoped `ANALYZE`; workerd permits `ANALYZE`, so planner statistics
are still refreshed but without bounded sampling. Unexpected PRAGMA or ANALYZE
failures stay visible through the existing caller warning or rejection. If you
need to bypass the API for an unusual deployment (for example issuing `ANALYZE`
over a separate admin connection), call `backend.execute()` with raw SQL as the
escape hatch.
### pgbouncer / transaction-pool mode
By default, the node-postgres / neon-serverless fast path issues server-side
prepared statements (`client.query({name, text, values})`) so PostgreSQL
caches the parsed plan per session. This is incompatible with pgbouncer in
transaction-pool mode: pgbouncer routes successive statements over different
backend connections, so a `name` registered on one connection isn't visible
on the next. Pass `prepareStatements: false` to fall back to unnamed
positional queries:
```typescript
const backend = createPostgresBackend(db, {
prepareStatements: false, // pgbouncer transaction-pool compatibility
});
```
The in-process cache that maps SQL text → statement name is LRU-bounded
(default 256 entries, override via `preparedStatementCacheMax`). Eviction
never recycles a name, because a live connection may still retain that name for
its original SQL. Therefore this setting does not bound server-side prepared
statement memory. For a high-cardinality stream of SQL text, use
`prepareStatements: false` instead.
### Adopted schema transactions
`store.withEvolvedTransaction(nativeTx, plan, callback, { waitBudgetMs })`
requires an initialized adapter Store and a live caller-owned transaction.
Plan the extension outside that transaction with `store.planEvolution()`.
For change plans, the default exclusive schema-fence wait budget is 5,000 ms; a `SchemaFenceTimeoutError`
requires rollback and retry of the complete application transaction. Omit
`waitBudgetMs` for no-op plans, which use ordinary adoption without the exclusive
fence and refuse that option.
Interactive PostgreSQL adapters validate the active session, retain the existing
schema advisory lock → schema row → recorded-write lock order, and use
transaction-scoped advisory locks. This lock lifetime is suitable for transaction
poolers such as Hyperdrive. The adapter restores temporary timeout settings before
the callback. Noninteractive HTTP drivers cannot adopt schema transactions.
SQLite schema adoption requires an active transaction on the backend's exact
native connection with an observable `inTransaction` state, as provided by
better-sqlite3. Drivers without that evidence refuse schema adoption; ordinary
transaction support alone does not imply support for this operation. A deferred
SQLite transaction acquires the writer slot before validating the schema plan.
Adapters default to a DML-only schema provisioning policy. Plans requiring new
vector slots or identity work refuse before taking a mutating fence, running
DDL, or changing schema rows. A privileged adapter configured with
`schemaProvisioning: "transactional"` can apply those plans: it revalidates
storage on the pinned caller session and provisions identity relations, vector
tables, and durable contribution markers inside that same transaction. The
caller must roll back the entire native transaction if any step fails.
```typescript
const backend = createPostgresBackend(db, {
schemaProvisioning: "transactional",
});
```
Use a connection with permission to run the required DDL for this adapter;
keep the default policy for a runtime role limited to DML.
Bootstrap base storage before this request path; missing bootstrap tables
refuse rather than being created lazily. Database permissions still determine
whether transactional DDL succeeds. Generic eager index materialization,
including concurrent PostgreSQL indexes, remains an explicit post-commit
operation on the refreshed Store.
Custom adapters must implement `adoptSchemaWriteTransaction` with the same
session-bound fencing, finite-wait, and CAS guarantees to support change plans.
See [Graph Extensions](/graph-extensions) for callback and receipt usage.
### Authoritative command sessions
Store create paths use the backend's `commands` port for writes whose
decision and mutation must share one command boundary. First-party paths pass
an explicit command context: a root port owns any internal transaction it
needs and cannot inherit caller coordination, while a transaction-scoped
backend uses the active caller or Store transaction. A
transaction command may additionally carry a coordination token only after it
has acquired the graph's advisory lock; the token is bound to that graph and
transaction session and cannot authorize work on another connection.
On PostgreSQL, the lock statement also observes the effective transaction
isolation and binds it to the same token. Match-key convergence therefore
accepts only read committed or serializable based on database state, not the
caller-requested option or the server's assumed default.
`GraphBackend.commands` is a required member as of the authoritative command
port release. Custom backends must expose `{ session, execute }` and implement
the `node.create`, `edge.create`, and `edge.converge-create` commands, or return
a typed `unsupported` result for dimensions they do not provide. The former
optional managed-create and specialized edge-insert hooks are no longer a
complete backend implementation; migrate those branches into the command
port before upgrading.
For a custom backend, the migration shape is:
```typescript
const commands: GraphCommandPort = {
session: "transaction", // use "root" for a single-statement backend
execute(command, context) {
// Apply every requested dimension, or explicitly refuse the command.
switch (command.kind) {
case "node.create": {
return { outcome: "unsupported", entity: "node", dimensions: ["claims"] };
}
case "edge.create": {
return {
outcome: "unsupported",
entity: "edge",
dimensions: ["endpointPredicate"],
};
}
case "edge.converge-create": {
return { outcome: "unsupported", entity: "edge", dimensions: ["convergence"] };
}
}
},
};
const backend: GraphBackend = { ...members, commands };
```
Every command port caller must provide the explicit context. TypeGraph-owned
write paths use the command helper, which verifies that any coordination token
belongs to the active graph and transaction session and carries a supported
effective isolation before executing convergence. The portable PostgreSQL
graph-lock path records that isolation automatically. A custom implementation
of `lockSchemaVersionAndGraphWrite` must return the normalized
`GraphCommandIsolation` observed by its combined lock statement. When
decorating a first-party backend with `deriveBackend`, a same-session
`commands` override retains the session identity. A wrapper that changes
session or forwards to a different connection is a new command boundary and
cannot reuse a token from the original port.
These are four different execution guarantees; do not use “atomic” as a
catch-all:
- **Interactive transaction** (`store.transaction(...)`) pins one session and
can make several Store operations commit or roll back together. The
`runOptionallyInTransaction` callback receives
`{ mode: "interactive-transaction" }` when this boundary was opened, or
`{ mode: "sequential" }` on a backend without transaction support.
- **Static internal adapter batch** is an adapter implementation detail (for
example, a D1 batch or a bind-budgeted multi-row insert). It may make one
precompiled set of statements atomic, but it is not a public Store
transaction and does not make an arbitrary sequence of Store calls atomic.
- **Certified atomic SQL program** is the backend-authoring transport seam for
a closed, ordered sequence of statements. A backend earns this capability by
passing the framework-agnostic conformance runner: result slots and bound
parameters must be preserved, a failure in a later statement must leave no
primary or sidecar writes, and an empty program must be a no-op. Certification
is separate from semantic mutation eligibility; a transport alone does not
authorize a mutation family. Bundled recognized PostgreSQL drivers provide
this boundary either through Neon HTTP's transaction batch or a pinned
interactive transaction; an unrecognized driver leaves it unavailable.
- **Authoritative one-statement command** is the `commands.execute` port. A
command returns a created/found/rejected/unsupported result after the
database statement itself owns the decision and mutation. It is the
transactionless path for eligible durable edge `matchIdentity` convergence;
it is not a promise that every command or side effect can be fused.
Operational Identity, single-edge claim/cardinality checks, and undeclared
dynamic `matchOn` convergence remain interactive-transaction contracts. A
custom or non-transactional backend must refuse those dimensions rather than
silently falling through to a sequence of independent statements. Eligible
direct edge batches on bundled roots are a narrower static-program contract:
the insert and cardinality sidecars execute in one native atomic exchange. A
declared durable edge `matchIdentity` is different for endpoint convergence:
its canonical key has a database arbiter, so the eligible root create/found
command can be authoritative in one statement.
Backend implementations may also expose the optional
`findEdgesByMatchIdentity` read capability for bounded merge planning. It must
match the complete `(graphId, kind, name, key)` tuple and return tombstoned
owners as well as active rows; omitting it keeps the portable full-clone path.
Custom Drizzle operation strategies can opt in by supplying the corresponding
owner-query builder. A strategy without that builder does not expose the
capability, so callers can detect and retain the portable path.
### Connection Pooling
For production, always use connection pooling:
```typescript
import { Pool } from "pg";
const pool = new Pool({
connectionString: process.env.DATABASE_URL,
max: 20, // Maximum pool size
idleTimeoutMillis: 30000, // Close idle connections after 30s
connectionTimeoutMillis: 2000, // Timeout for new connections
});
// Handle pool errors
pool.on("error", (err) => {
console.error("Unexpected pool error", err);
});
// Graceful shutdown
process.on("SIGTERM", async () => {
await pool.end();
process.exit(0);
});
```
### API Reference
#### `createPostgresBackend(db, options?)`
Creates a PostgreSQL backend adapter. Accepts any Drizzle PostgreSQL database
instance, regardless of the underlying driver. Tested with `drizzle-orm/node-postgres`,
`drizzle-orm/postgres-js`, `drizzle-orm/neon-serverless`,
`drizzle-orm/neon-http`, and `drizzle-orm/pglite`. The neon-http driver is auto-detected and
`capabilities.execution.interactiveTransactions` is set to `false` (HTTP can't hold a session); use
`drizzle-orm/neon-serverless` if you need transactional writes.
```typescript
function createPostgresBackend(
db: AnyPgDatabase,
options?: {
tables?: PostgresTables;
/**
* Override the fulltext strategy. Defaults to `tsvectorStrategy`.
* Pass a custom `FulltextStrategy` to swap the fulltext stack, or
* `false` to disable fulltext support entirely — the backend then
* advertises no `capabilities.fulltext` and omits the fulltext
* CRUD/search methods, mirroring `vector: false`.
*/
fulltext?: FulltextStrategy | false;
/**
* Override the vector search strategy. Defaults to
* `pgvectorStrategy`. Pass a custom `VectorStrategy` to change the
* storage / index engine, or `false` to disable vector support.
*/
vector?: VectorStrategy | false;
/**
* Override specific backend capabilities. Useful for HTTP-style
* drivers or test scenarios. neon-http already has
* `execution.interactiveTransactions: false` auto-applied — pass
* this to override that or to disable other capabilities for custom
* drivers.
*/
capabilities?: BundledBackendCapabilityOverrides;
/**
* Use server-side prepared statements on the node-postgres /
* neon-serverless fast path. Default `true`. Set to `false` when
* pooling through pgbouncer in transaction-pool mode (named
* statements are invisible across pooled connections).
*/
prepareStatements?: boolean;
/**
* LRU cap on the number of distinct SQL strings tracked for
* prepared-statement naming. Default 256. Worst-case server-side
* footprint is roughly `cap × pool size` prepared statements.
* Ignored when `prepareStatements` is `false`.
*/
preparedStatementCacheMax?: number;
},
): GraphBackend;
```
Pass `{ fulltext: false }` when the graph has no `searchable()` fields and
you would rather skip the fulltext table (`typegraph_node_fulltext`) and its
GIN index than carry them unused:
```typescript
const backend = createPostgresBackend(db, { fulltext: false });
```
#### `createPostgresTransactionBackend(tx, options?)`
Creates a full backend on a Drizzle PostgreSQL transaction opened by the
application. Use it when TypeGraph's tables share a transaction with other
application tables, especially when TypeGraph uses prefixed table names. Pass
the same `PostgresBackendOptions` as `createPostgresBackend`:
```typescript
import {
createPostgresTransactionBackend,
createPostgresTables,
} from "@nicia-ai/typegraph/adapters/drizzle/postgres";
const graphTables = createPostgresTables({ nodes: "app_graph_nodes" });
await db.transaction(async (tx) => {
const backend = createPostgresTransactionBackend(tx, {
tables: graphTables,
});
// Use the backend or a Store built from it within this callback.
});
```
The factory requires a transaction handle and serializes TypeGraph statements
on its single pinned connection, including concurrent reads started by the
same Store operation. Backends created for the same transaction handle share
one queue. The application owns commit and rollback and must await all work
using these backends before its transaction callback returns.
`createPostgresBackend(tx)` also routes a PostgreSQL transaction handle to the
transaction-scoped backend automatically. Use `createPostgresTransactionBackend`
when you want the transaction-scoped intent to be explicit; a regular database
handle passed to `createPostgresBackend(db)` still creates the pooled backend.
#### `createLocalPgliteBackend(options?)`
Creates an in-process PGlite backend with automatic engine construction,
schema DDL, and optional pgvector loading. The returned backend owns the PGlite
engine; call `backend.close()` when the process or test is done.
```typescript
async function createLocalPgliteBackend(options?: {
/**
* PGlite data directory. Omit for an in-memory database, pass a filesystem
* path for persistence, or use a runtime-specific scheme such as `idb://`.
*/
dataDir?: string;
tables?: PostgresTables;
/**
* Omit to load @electric-sql/pglite-pgvector, pass `false` to disable vector
* support, or pass a PGlite Extension object to control the extension import.
*/
vector?: false | Extension;
/**
* Override the fulltext strategy. Defaults to `tsvectorStrategy`. Pass
* `false` to disable fulltext support entirely — the backend then
* advertises no `capabilities.fulltext` and omits the fulltext CRUD/search
* methods, and the installation DDL never creates the fulltext table.
*/
fulltext?: FulltextStrategy | false;
}): Promise<{
backend: GraphBackend;
db: PgliteDatabase;
client: PGlite;
}>;
```
#### `generatePostgresMigrationSQL()`
Returns complete fresh-installation SQL for creating TypeGraph tables and
stamping the current deployment-wide base-schema marker in PostgreSQL. It
includes the pgvector extension. The vector-disabled local PGlite factory uses
the same installation builder internally while omitting only that extension.
```typescript
function generatePostgresMigrationSQL(
tables?: PostgresTables,
fulltextStrategy?: FulltextStrategy | false,
): string;
```
#### `generatePostgresDDL(tables?)`
Returns individual DDL statements (CREATE TABLE, CREATE INDEX) as an array. Useful when you
need per-statement control, for example to execute them in separate transactions or log them
individually. This low-level array deliberately omits the deployment-wide
marker row, so joining it does not produce a complete installation. Use
`generatePostgresMigrationSQL()` for a database that will be opened through
`createVerifiedStore()` or the DML-only graph-template APIs.
```typescript
function generatePostgresDDL(
tables?: PostgresTables,
fulltextStrategy?: FulltextStrategy | false,
): string[];
```
#### `generatePostgresDropSQL(tables?, fulltextStrategy?)`
Returns one `DROP TABLE IF EXISTS` statement for the base and fulltext tables
that `generatePostgresDDL()` would create. Use it to clean up an isolated,
prefixed PostgreSQL table set after closing every backend connected to it.
Pass the same tables and fulltext strategy used at installation. The statement
does not use `CASCADE`: PostgreSQL refuses the drop if an application-owned
object depends on one of these tables. It does not drop graph-scoped vector
tables materialized later at runtime, so a working copy using those tables
needs additional graph-scoped cleanup.
```typescript
function generatePostgresDropSQL(
tables?: PostgresTables,
fulltextStrategy?: FulltextStrategy | false,
): string;
```
### Upgrading deployment-wide base storage
Skip this section when `createStoreWithSchema()` or
`createAdapterStoreWithSchema()` owns schema preparation: the bundled SQLite
and PostgreSQL adapters adopt each numbered base-schema release automatically
on the first privileged open. No separate bootstrap command is needed. The
deployment invariant is ordering: that privileged open must finish before any
DML-only runtime worker starts. Base-schema version 1 includes the durable graph
template relation and edge match-identity storage. It is required even for
graphs without a `matchIdentity` declaration because every edge write names the
two nullable columns.
When database DDL is managed externally, apply the matching migration before a
runtime worker opens the new graph schema. Apply the marker write last: it is
the durable proof that every preceding step succeeded. The examples use the
default TypeGraph table names. Replace every occurrence consistently when the
adapter uses custom table names.
For SQLite, run this migration exactly once. SQLite has no portable `ADD COLUMN
IF NOT EXISTS`, so a migration tool must record whether it has already applied
the two `ALTER TABLE` statements. Fresh and published schemas include the
nullable-pair `CHECK` below. Privileged adoption accepts an externally managed
table that already has both columns without that defensive constraint: SQLite
does not expose structural CHECK metadata or support adding one without a full
table rebuild, while TypeGraph writes always bind both values or neither.
```sql
CREATE TABLE IF NOT EXISTS "typegraph_graph_templates" (
"template_id" TEXT PRIMARY KEY NOT NULL,
"schema_hash" TEXT NOT NULL,
"schema_doc" TEXT NOT NULL,
"created_at" TEXT NOT NULL
);
ALTER TABLE "typegraph_edges"
ADD COLUMN "match_identity_name" TEXT;
ALTER TABLE "typegraph_edges"
ADD COLUMN "match_identity_key" TEXT
CHECK (("match_identity_name" IS NULL) = ("match_identity_key" IS NULL));
CREATE UNIQUE INDEX IF NOT EXISTS "typegraph_edges_match_identity_uq"
ON "typegraph_edges" (
"graph_id", "kind", "match_identity_name", "match_identity_key"
);
CREATE TABLE IF NOT EXISTS "typegraph_base_schema_versions" (
"installation" INTEGER PRIMARY KEY NOT NULL,
"version" INTEGER NOT NULL,
"updated_at" TEXT NOT NULL,
CONSTRAINT "typegraph_base_schema_versions_singleton_check"
CHECK ("installation" = 1)
);
INSERT INTO "typegraph_base_schema_versions"
("installation", "version", "updated_at")
VALUES (1, 1, CURRENT_TIMESTAMP)
ON CONFLICT ("installation") DO UPDATE SET
"version" = excluded."version",
"updated_at" = excluded."updated_at"
WHERE "typegraph_base_schema_versions"."version" <= excluded."version";
```
For PostgreSQL, the adoption statements are idempotent:
```sql
CREATE TABLE IF NOT EXISTS "typegraph_graph_templates" (
"template_id" TEXT PRIMARY KEY NOT NULL,
"schema_hash" TEXT NOT NULL,
"schema_doc" JSONB NOT NULL,
"created_at" TIMESTAMPTZ NOT NULL
);
ALTER TABLE "typegraph_edges"
ADD COLUMN IF NOT EXISTS "match_identity_name" TEXT;
ALTER TABLE "typegraph_edges"
ADD COLUMN IF NOT EXISTS "match_identity_key" TEXT;
DO $$
BEGIN
IF NOT EXISTS (
SELECT 1
FROM pg_constraint
WHERE conrelid = to_regclass('"typegraph_edges"')
AND conname = 'typegraph_edges_match_identity_pair_check'
) THEN
ALTER TABLE "typegraph_edges"
ADD CONSTRAINT "typegraph_edges_match_identity_pair_check"
CHECK (
("match_identity_name" IS NULL) = ("match_identity_key" IS NULL)
);
END IF;
END $$;
CREATE UNIQUE INDEX IF NOT EXISTS "typegraph_edges_match_identity_uq"
ON "typegraph_edges" (
"graph_id", "kind", "match_identity_name", "match_identity_key"
);
CREATE TABLE IF NOT EXISTS "typegraph_base_schema_versions" (
"installation" INTEGER PRIMARY KEY NOT NULL,
"version" INTEGER NOT NULL,
"updated_at" TIMESTAMPTZ NOT NULL,
CONSTRAINT "typegraph_base_schema_versions_singleton_check"
CHECK ("installation" = 1)
);
INSERT INTO "typegraph_base_schema_versions"
("installation", "version", "updated_at")
VALUES (1, 1, NOW())
ON CONFLICT ("installation") DO UPDATE SET
"version" = excluded."version",
"updated_at" = excluded."updated_at"
WHERE "typegraph_base_schema_versions"."version" <= excluded."version";
```
The conditional update makes marker publication monotonic: replaying an older
migration can never claim that storage prepared by a newer TypeGraph release is
older. The fresh-installation generators use `DO NOTHING` instead because they
are not upgrade planners; an existing stale marker remains stale until the
numbered privileged adoption lifecycle runs.
`createVerifiedStore`, `assertSchemaCurrent`, and the DML-only graph-template
APIs read this marker and throw `BaseSchemaMigrationError` when it is missing,
stale, or newer than the running library. They never attempt repair. A plain
`createStore` remains a synchronous zero-I/O attach; if it reaches an edge
write on legacy storage, the write fails with `ConfigurationError` and
`details.code === "EDGE_MATCH_IDENTITY_STORAGE_UNAVAILABLE"` rather than a raw
missing-column error.
Provisioning the columns does not authorize re-keying existing data. Adding,
removing, renaming, or changing the fields of a declared `matchIdentity` remains
a breaking graph-schema change while that edge kind has any physical rows,
including tombstones. Export the affected edges, hard-delete them, publish the
new schema, and import them again so every row receives a key under the new
declaration.
### Base-schema version 5: byte-ordered `graph_id` indexes (PostgreSQL)
Version 5 adds one index to each relation `listGraphIds` seeks (`nodes`, `edges` and
`schema_versions`), ordering `graph_id` by bytes instead of by the database collation so that a page's
cursor, prefix and limit bound the walk. SQLite already keeps text indexes in byte order, so its step
only advances the marker. The index is a single-column `graph_id` index: PostgreSQL deduplicates the
repeated values, so it stays small (about 7 MB beside a 97 MB `nodes` heap of one million rows) and
adds 1 to 2% to writes on the relation it lands on (single creates and 1,000-row bulk writes alike).
The privileged open builds the three indexes with a plain `CREATE INDEX`, which blocks writes to the
table while it runs (about 0.1 second per million `nodes` rows on the measurement hardware). This
happens inline at boot even when `systemIndexes: "skip"` is set: that option only defers system index
materialization, not base-schema adoption. For a large deployment, build the indexes first with
`CONCURRENTLY`; the adoption step is `IF NOT EXISTS` and then finds them in place. Run each statement
outside a transaction, and never run the same concurrent build from two sessions at once. Use the
adapter's table names throughout:
```sql
CREATE INDEX CONCURRENTLY IF NOT EXISTS "typegraph_nodes_graph_id_bytes_idx"
ON "typegraph_nodes" ("graph_id" COLLATE "C");
CREATE INDEX CONCURRENTLY IF NOT EXISTS "typegraph_edges_graph_id_bytes_idx"
ON "typegraph_edges" ("graph_id" COLLATE "C");
CREATE INDEX CONCURRENTLY IF NOT EXISTS "typegraph_schema_versions_graph_id_bytes_idx"
ON "typegraph_schema_versions" ("graph_id" COLLATE "C");
```
`CREATE INDEX CONCURRENTLY` can leave an invalid index behind if it is interrupted; drop it and rerun.
`listGraphIds` checks only that each index exists and is valid, not its definition, so an index you
create by hand under one of these names with a different definition is trusted and makes the walk slow
rather than wrong. Create them exactly as shown.
Advancing the marker to 5 is a one-way step: a library release that predates version 5 refuses a
database stamped 5, so roll forward rather than back once any process has adopted it.
Externally managed DDL applies the same statements, then advances the marker to 5 with the
monotonic `INSERT ... ON CONFLICT` shown above. Until the indexes exist `listGraphIds` still returns
correct pages, by reading and de-duplicating the anchor relations instead of walking them.
## Drizzle-Free Entrypoints
TypeGraph keeps its public core and backend contracts independent of Drizzle:
- `@nicia-ai/typegraph/core` exports graph definition helpers and their
schema-derived types for packages that only define or share schemas.
- `@nicia-ai/typegraph/backend` exports the complete backend, dialect,
SQL-fragment, fulltext, and vector strategy contracts for adapter authors.
- `@nicia-ai/typegraph/sqlite/local` and
`@nicia-ai/typegraph/postgres/pglite` create managed Stores without exposing
adapter-native handles.
Application code can continue importing the complete portable Store API from
`@nicia-ai/typegraph`. Use the `/adapters/drizzle/...` entrypoints only when the
application deliberately owns a Drizzle connection or needs native transaction
interop.
Custom insert builders must apply the same born-ended validity rule as the
built-in adapters. Import its public owner instead of duplicating the bound
comparison:
```typescript
import { resolveStampedValidityLowerBound } from "@nicia-ai/typegraph/backend";
const validFrom = resolveStampedValidityLowerBound(
params.validFrom,
params.validTo,
writeInstant,
);
```
Use the same `writeInstant` for the decision and the row's creation/update
stamp. This keeps custom node and edge inserts, plus node resurrection paths
that reset the validity window, aligned with Store and interchange semantics at
the zero-width boundary. Edge resurrection retains its stored lower bound and
does not use this stamping helper.
## Managed Store Entrypoints
For local applications that do not need direct database access, TypeGraph can
own the connection, provision its schema, and return the complete typed Store:
- `@nicia-ai/typegraph/sqlite/local` — Node-only SQLite through the
native better-sqlite3 addon
- `@nicia-ai/typegraph/postgres/pglite` — in-process PostgreSQL through
PGlite's WebAssembly runtime
```typescript
import { createLocalSqliteStore } from "@nicia-ai/typegraph/sqlite/local";
import { createLocalPgliteStore } from "@nicia-ai/typegraph/postgres/pglite";
const sqliteStore = await createLocalSqliteStore(graph, { path: "./graph.db" });
const postgresStore = await createLocalPgliteStore(graph, { vector: false });
```
These entrypoints expose no adapter-native database handle. The returned
`Store` keeps the complete graph API, including graph-owned
`store.transaction(...)`, but intentionally omits `withTransaction` and
`withRecordedTransaction`, which require a caller-owned adapter handle. The
Store owns its connection, so call `store.close()` during shutdown. Its
declaration surface is safe for strict TypeScript consumers that do not install
unused database drivers.
PGlite vector support is enabled by default and loads the optional
`@electric-sql/pglite-pgvector` package. Install that package when using vector
fields, or pass `{ vector: false }` as above for a smaller non-vector setup.
Both managed entrypoints also accept `fulltext: false`, which skips the
fulltext table at bootstrap and returns a backend with no
`capabilities.fulltext`.
Both factories accept `store` and `schemaManagement` groups, so the managed
path supports the same hooks, history/revision tracking, custom SQL schema,
query defaults, and migration policy as `createStoreWithSchema`:
```typescript
import { createSqlSchema } from "@nicia-ai/typegraph";
const store = await createLocalSqliteStore(graph, {
path: "./graph.db",
pragmas: { busyTimeoutMs: 10_000 },
store: {
history: true,
schema: createSqlSchema({
nodes: "app_nodes",
edges: "app_edges",
fulltext: "app_fulltext",
uniques: "app_uniques",
}),
},
schemaManagement: { systemIndexes: "skip" },
});
```
When a custom SQL schema is supplied, the managed factory provisions those
same physical table names; no separate Drizzle table configuration is needed.
`drizzle-orm` is an optional peer dependency for these two managed
entrypoints: they load it only when their factory is called and, when it is
absent, reject with a typed `ConfigurationError` (`MISSING_PEER_DEPENDENCY`)
naming the package and the install command (`npm install drizzle-orm`) rather
than a bare module-resolution stack. The explicit `/adapters/drizzle/...`
entrypoints below expose Drizzle-native backends, connections, or schema
builders — or, for `/adapters/drizzle/engine`, the factory that assembles a
backend from a caller-supplied engine profile — and load `drizzle-orm` when
the module is evaluated. Importing one without the peer installed therefore
surfaces the raw module-resolution error, which names the same package.
## Drizzle Adapter Entrypoints
TypeGraph exposes Drizzle adapters through public entrypoints:
- `@nicia-ai/typegraph/adapters/drizzle/indexes` — Drizzle schema-builder helpers for TypeGraph index declarations
- `@nicia-ai/typegraph/adapters/drizzle/sqlite` — Generic SQLite adapter (any Drizzle SQLite driver)
- `@nicia-ai/typegraph/adapters/drizzle/sqlite/local` — Batteries-included better-sqlite3 wrapper (Node.js only)
- `@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql` — Batteries-included libsql wrapper (Node.js, Workers, browser)
- `@nicia-ai/typegraph/adapters/drizzle/postgres` — PostgreSQL adapter (any Drizzle Postgres driver)
- `@nicia-ai/typegraph/adapters/drizzle/postgres/working-copy` — PostgreSQL table-backed working-copy manager
- `@nicia-ai/typegraph/adapters/drizzle/postgres/pglite` — Batteries-included PGlite (Postgres-in-WASM) wrapper
- `@nicia-ai/typegraph/adapters/drizzle/engine` — `createSqlBackend`, `deriveEngineProfile`, the bundled builders, `SqlEngineProfile`
Import from the entrypoint matching your database:
```typescript
import { createSqliteBackend, tables } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local";
import { createLibsqlBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/libsql";
import { createPostgresBackend, tables } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import { createLocalPgliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres/pglite";
```
### Engine profiles
`createPostgresBackend` and `createSqliteBackend` are each `createSqlBackend`
applied to a profile built by `buildPostgresEngineProfile` /
`buildSqliteEngineProfile`, both exported alongside `createSqlBackend` and
`deriveEngineProfile` from the engine entrypoint:
```typescript
import {
buildPostgresEngineProfile,
createSqlBackend,
} from "@nicia-ai/typegraph/adapters/drizzle/engine";
const backend = createSqlBackend(buildPostgresEngineProfile(db, options));
```
Most callers adapting a bundled backend want `deriveEngineProfile`, which
builds a variant of a bundled profile — a different lock spelling, a looser
declared capability, a replaced resource-audit verdict — without hand-copying
every other field. A profile written from scratch is not constructible today:
the assembly constructor is unexported, and `createSqlBackend` refuses a
hand-built assembly. See [Authoring an engine profile](/backend-authoring) for
the derivable-field table, the refusals a custom profile can hit, a worked
example, and [what is not derivable yet](/backend-authoring#what-is-not-derivable-yet).
A profile owns everything that genuinely differs between engines: dialect
tokens, the execution adapter, transaction framing, its `fenceSql` lock
spelling (see [Write fence declaration](#write-fence-declaration-writefence)),
provisioning DDL, strategies, and limits. `createSqlBackend` owns everything that is the
same for every SQL engine: deriving the final capabilities, resolving the
write-fence decision once, assembling the mirrored member groups, auditing
the backend's resource shape, and applying the trust marks. A backend minted
this way earns the marks its own declarations back: the schema-fenced-insert
mark only when the resolved fence plan actually fences writers, the
root-autocommit mark only when the profile declares single-statement
durability, and the atomic-program registrations only when its capabilities
support root atomic batching — whether the profile is for a PostgreSQL-wire
engine with a different locking story or an embedded engine with a different
transaction model.
`createSqlBackend` refuses a profile whose resolved capabilities omit
`writeFence` — every mark and registration it applies assumes a
resolvable write-fence decision, and a profile that does not declare one
cannot back that decision soundly. `buildPostgresEngineProfile` and
`buildSqliteEngineProfile` are the reference profiles to read when modeling a
new one.
A profile's `provisioning.catalog` supplies the backend's optional `catalog`
member: physical-schema introspection — table and index existence, each
column's normalized type family and raw declared type (a `CatalogColumn` is
`{ name, kind, declaredType }`; `declaredType` is required, and every custom
`columnTypes` implementation must populate it), and this engine's
index-build facts — for the handful of store paths that need to read the
engine catalog directly instead of compiling a portable query: index
materialization (`store.materializeIndexes()` refuses only once its
empty-candidate short circuit and the status-table ensure step have already
run; `store.materializeSystemIndexes()`, which has no candidate short
circuit, refuses only once that same status-table ensure step has run), the
recorded-time schema check, and the recorded-time migration's column read. A
profile that leaves `catalog` unset builds a backend with no `catalog`
member at all; those paths then refuse with a `ConfigurationError` naming
`catalog` rather than guessing at engine-specific SQL.
A dialect also declares `subgraphMembershipStrategy`, naming a decision the
dialect adapter makes, not one a profile supplies directly — the dialect
adapters are a fixed record keyed by `SqlDialect`, and each adapter's
capabilities (`DialectCapabilities`) declares `subgraphMembershipStrategy`, so a
profile inherits whichever of the two dialects its own `dialect` field names. It
is the plan-shape choice behind `store.subgraph()`'s reachable-node filter:
`"materialized-ids"` fetches the traversal closure once and filters both the
node and edge queries against that fixed id list (the shape PostgreSQL uses,
trading one extra round trip for a parameter-driven plan), while `"inline-cte"`
embeds the recursive closure in each fetch instead (the shape SQLite uses, where
an in-process traversal is cheap and a growing parameter list would pressure the
bind budget). This is a control-flow and prepared-plan decision, not SQL text a
shared token could express identically on both shapes, so it lives on
`DialectCapabilities` rather than in the query compiler.
`instantiateStatement` — a member of the profile's `graphTemplateRuntime` bag, and so one of the
fields `deriveEngineProfile` can override — is a required builder cloning a durable schema template
into a fresh graph. Given the template and target graph's ids and schema hashes
(`InstantiateGraphTemplateSqlParams`: `templateId`, `templateSchemaHash`, `graphId`, `schemaHash`,
and the three physical table names it reads), it must return the statement that inserts the target
graph's `schema_versions` row from the template's stored document and copies the template's
contribution-marker rows into the target graph — taking the target graph's write lock, the same key
the schema-commit fence takes, co-atomically with the insert on an engine that fences with locks. An
engine whose dialect can compose a data-modifying CTE beside the schema INSERT (PostgreSQL) folds
the marker copy and the lock into that one statement; an engine that cannot (SQLite) instead
supplies the optional `copyContributionMarkers` dep, which runs the marker copy as a second
statement once the schema row is confirmed. The bundled
`postgresInstantiateGraphTemplateStatement` and `sqliteInstantiateGraphTemplateStatement` builders
(`graph-template-sql.ts`) are what `createPostgresBackend` and `createSqliteBackend` supply to
their own profiles; neither is exported, so a from-scratch profile reaches the same shape only by
copying a bundled profile and adapting its statement, while a derived profile can replace the whole
`graphTemplateRuntime` bag through `deriveEngineProfile`.
`FenceSql` (see [Write fence declaration](#write-fence-declaration-writefence))
declares `advisoryLockExpression` and `isolationFactExpression` as the two
composable, no-`SELECT` forms an `advisory`-mechanism backend author
supplies; TypeGraph derives the standalone-statement counterparts
(`acquireKeyed`, `acquireKeyedWithIsolation`, `isolationFact`) from them. A
statement that
must compose a lock or an isolation read INSIDE a larger query it builds
itself — a CTE, a data-modifying statement — embeds the bare expression
directly, rather than running the derived standalone form as its own
preceding statement. The schema write fence's fused schema + graph-write
statement (`postgres-schema-write-fence.ts`) is the one site that needs
this: it puts the expression in its own CTE's `SELECT ... AS "lock_token"`,
resolving the fence target's OWN `FenceSql` — the bundled `postgresFenceSql`
for a bundled backend, a derived profile's own override otherwise — so a
custom spelling backs this fused statement exactly as it backs every
ordinary lock site.
The graph-template instantiation statement's `locked AS (SELECT ...)` CTE is
a DIFFERENT case, not a `FenceSql` consumer at all: it composes the baked
single-argument `advisoryLockSingleExpression` directly.
The ONE lock form with no override point is `advisoryLockSingleExpression`,
the ONE-argument form on a bare key: PostgreSQL stores it in a lock space
distinct from every namespaced two-argument lock, and the schema-commit fence
and graph-template instantiation both take it on the same key so the two
mutually exclude. It is not a `FenceSql` member — both bundled builders bake it
in directly, and a custom profile has no way to replace it.
## Cloudflare D1
TypeGraph supports Cloudflare D1 for edge deployments, with some limitations.
Cloudflare D1 has no interactive transaction primitive, so it cannot commit
TypeGraph schema versions: `commitSchemaVersion` / `setActiveVersion` need to
hold one transaction across their compare-and-swap read and activating
write, and D1 has no session to hold it on. Apply the base DDL with Wrangler
/ drizzle-kit. `capabilities.execution.unitOfWork` reports `"batch"`. A
singleton node create, update, `upsertById`, or delete fuses on a kind with
no declared unique constraint (a create takes a generated or a
caller-supplied id) — except a node delete, which fuses even when the kind
DOES carry a declared unique constraint, because the atomic delete program
releases that claim in the same statement. A singleton edge create fuses
when the kind's cardinality is `"many"`, and edge update and delete fuse the
same way (`EdgeCollection` has no `upsertById`). So do
`bulkInsert`/`bulkCreate`/`bulkDelete`/`bulkReplaceById`/`bulkUpsertById`,
and a constrained write inside an atomic program's claim envelope. Each of
these asserts the active schema version inside the statements
`D1Database.batch()` runs together, so these succeed on a schema-managed
Store. A write that cannot fuse either fails closed with
`BATCH_WRITE_UNSUPPORTED` naming a proven reason (a probe-then-write
constraint check, an interactive callback, Operational Identity, history, or
a schema commit), or — for a write that simply doesn't fit the fused shape,
such as a singleton create, update, or `upsertById` on a uniquely-constrained
kind, or a supplied-id tombstone resurrection — fails closed with the plain
`SCHEMA_WRITE_FENCE_UNSUPPORTED` limitation and no named reason; see
[The guard every fused write shares](/limitations#the-guard-every-fused-write-shares)
for the full reason table. Use a raw `createStore()` only when the
application accepts unfenced writes for the remaining paths:
```typescript
import { drizzle } from "drizzle-orm/d1";
import { createStore } from "@nicia-ai/typegraph";
import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
export default {
async fetch(request: Request, env: Env) {
const db = drizzle(env.DB);
const backend = createSqliteBackend(db);
const store = createStore(graph, backend);
// Use store...
},
};
```
This raw Store does not validate or fence a committed TypeGraph schema version.
For schema commits and multi-statement schema-managed writes on Cloudflare, use
**Durable Objects** (below), whose SQLite storage exposes an interactive
transaction runner.
**Important:** D1 has no interactive transaction primitive
(`D1Database.batch(...)` is transactional, but batch-only — not an
interactive runner), so `store.transaction()` refuses on D1 before invoking
its callback. See
[Limitations](/limitations) for details. For a transactional Cloudflare
SQLite store, use **Durable Objects** (below) instead.
For the same reason, a write guarded by a **declared constraint** — edge
cardinality other than `many`, a `disjointWith` axiom, a shared-scope unique, or
dynamic `getOrCreateByEndpoints` convergence — is refused on D1 with
`CONSTRAINT_WRITE_FENCE_UNSUPPORTED` rather than committed unfenced. See
[Declared constraints require an interactive transaction](#declared-constraints-require-an-interactive-transaction).
## Cloudflare Durable Objects (SQLite)
A store backed by `drizzle(ctx.storage)` inside a Durable Object is
**auto-detected** as `transactionMode: "do-sqlite"` and reports
`capabilities.execution.interactiveTransactions: true` — no `executionProfile` hint needed.
Unlike D1, Durable Objects expose an interactive storage transaction runner,
so adapter stores can provide fully atomic `store.transaction()` and
`store.withTransaction()` operations.
The runtime authorizer forbids temporary tables, so the same profile reports
`capabilities.graphAnalytics.supported: false`. Traversal algorithms such as
`shortestPath`, `reachable`, and `weightedShortestPath` automatically use their
inline fallback; temporary-table-only analytics such as
`weaklyConnectedComponents` throw `UnsupportedBackendCapabilityError`.
The authorizer also rejects SQLite's `analysis_limit` tuning PRAGMA. Statistics
refresh catches that specific authorization error and still runs scoped
`ANALYZE`; this affects refresh cost only, not query results.
The same profile advertises Cloudflare's 100-bound-parameter query limit.
TypeGraph uses that hard ceiling for its managed write batches and
recorded-history flushes; capability overrides may lower it but cannot raise
it. Literal `.in()` and `.notIn()` query lists are packed into one JSON-bound
parameter, so the list itself does not exhaust the Durable Object budget.
```typescript
import { drizzle } from "drizzle-orm/durable-sqlite";
import { createAdapterStoreWithSchema } from "@nicia-ai/typegraph";
import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
export class MyObject {
constructor(private ctx: DurableObjectState) {}
async handle() {
const db = drizzle(this.ctx.storage);
const backend = createSqliteBackend(db);
// Boots schema/DDL outside any storage transaction (no DDL in the
// business transaction); the schema-version commit uses the
// do-sqlite runner.
const [store] = await createAdapterStoreWithSchema(graph, backend);
// Atomic across TypeGraph + the product's own relational tables:
await store.transaction(async (tx) => {
await tx.nodes.Document.update(documentId, props);
if (tx.sqlAvailability !== "available") {
throw new Error(`Native transaction unavailable: ${tx.sqlAvailability}`);
}
const sqlTx = tx.sql;
await sqlTx.insert(documentVersions).values(versionRow);
});
}
}
```
TypeGraph delegates to the async storage runner
`ctx.storage.transaction(async …)` (Drizzle's own `db.transaction()` on
Durable Objects is `ctx.storage.transactionSync` and cannot span an
`await`, so it is not used). See the
[Cross-Store Transactions recipe](/recipes#cross-store-transactions-drizzle--typegraph)
for the caller-owned (`withTransaction`) and graph-owned (`tx.sql`) shapes.
## Backend Capabilities
Check what features a backend supports:
```typescript
const backend = createSqliteBackend(db);
const store = createStore(graph, backend);
if (store.capabilities.execution.interactiveTransactions) {
await store.transaction(async (tx) => {
/* ... */
});
} else {
// Handle non-transactional execution
}
if (store.capabilities.vector?.supported) {
// Vector similarity queries available
}
```
`store.capabilities` is the portable runtime source of truth; adapter authors
can inspect the same object as `backend.capabilities`. The shape is:
| Field | Meaning |
| -------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- |
| `execution` | Execution boundaries: `interactiveTransactions`, exact-resource `atomicBatch` support, and derived `unitOfWork` |
| `windowFunctions` | SQL window functions such as `ROW_NUMBER()` are available |
| `orderedAggregates?` | Ordered scalar and record collection aggregation; absent means unsupported |
| `constraintClaims?` | The backend carries the claim relations that fence declared constraints without a lock (see below) |
| `durableEdgeMatchIdentity?` | Edge writes persist and atomically arbitrate a schema-declared endpoint/property identity |
| `graphAnalytics?.{supported,mathFunctions}` | Static support for whole-graph temporary-table iteration, plus availability of deferred transcendental-math algorithms |
| `vector?.metrics` / `vector?.indexTypes` / `vector?.maxDimensions` | Vector strategy capabilities (present once a vector strategy is configured) |
| `fulltext?.{supported,languages,phraseQueries,prefixQueries,highlighting}` | Fulltext strategy capabilities |
| `recursiveTraversal?.{supported,reason}` | Whether the engine can compute a bounded transitive closure of a relation in one round trip — a recursive CTE, or a graph-native expansion operator. **Absent means supported** |
| `writeFence?.{mechanism,drain}` | How this engine excludes concurrent writers, and how far a caller can drain a table lock — see [Write fence declaration](#write-fence-declaration-writefence) |
`expr.collect()` requires `orderedAggregates: true`. Bundled PostgreSQL supports it. Supported preparable synchronous
SQLite clients and the dedicated async `createLibsqlBackend()` factory are probed when the backend
is created; the query itself adds no discovery statement. Other unprobed SQLite
connections default to unsupported. If you have
verified that your engine supports aggregate-local ordering, declare
`capabilities: { orderedAggregates: true }` in the bundled backend options. Older or unsupported
engines must retain `false`; collection queries are refused before execution.
SQLite introduced `NULLS FIRST` / `NULLS LAST` ordering in version 3.30 and aggregate-local ordering
in [version 3.44](https://www.sqlite.org/releaselog/3_44_0.html).
Collection compilation avoids depending on JSON object subtype preservation during sorting.
Existing reads continue to work when ordered aggregates are unavailable.
Custom dialect adapters implement `orderedScalarJsonArray` with one required object argument:
`{ value, valueType, orderBy, filter }`. Migrate positional implementations by destructuring that
object, applying `filter` as an aggregate `FILTER (WHERE ...)` before wrapping the aggregate in the
empty-input `COALESCE`, and leaving it off when `filter` is `undefined`. The aggregate must preserve
included NULL operands and return `[]` for empty input. The `filter` key itself is required in the
adapter contract, even though its value may be `undefined`, which requires old positional
implementations to migrate explicitly.
The dialect adapter contract also requires `orderedRecordJsonArray` for
`expr.collect({ field: scalarExpression }, options)`. Custom adapters must add this method when
upgrading. It builds an ordered JSON array of flat records with the named projected scalar fields,
honors the required ordering tuple and optional aggregate-local filter, preserves admitted SQL NULL
fields within each record, and returns `[]` for empty input. The same `orderedAggregates` capability
governs scalar and record collections; a custom adapter must supply both SQL emitters before
declaring that capability.
The former top-level `capabilities.transactions` override is not interpreted
as an alias. Bundled factories refuse it with `LEGACY_CAPABILITY_OVERRIDE`,
including for JavaScript and already-compiled callers, because transaction
availability and atomic batching are now independent facts. Move the value to
`capabilities.execution.interactiveTransactions`; root `atomicBatch` support
is discovered from the bundled transport and cannot be claimed through factory
overrides. Bundled PostgreSQL transaction factories may expose
`atomicBatch: "session"` on the exact already-open transaction object. That
declaration is paired with fresh transport and semantic registrations and is
never inherited by an ordinary derived backend.
`execution.unitOfWork` is derived, never declared by a factory or override:
`"optimistic-retry"` when `interactiveTransactions` is `true` AND the resolved
write fence is `{ mechanism: "row", conflict: "commit-time" }` (see
[Write fence declaration](#write-fence-declaration-writefence) below);
`"interactive"` when `interactiveTransactions` is `true` otherwise; else
`"batch"` when `atomicBatch` is not `"none"` (an HTTP-only driver such as
`drizzle-orm/neon-http`, which cannot hold an open session but does support a
native atomic program); else `"none"`. Only the root capability derivation
ever resolves the write-fence plan needed for the `"optimistic-retry"` arm —
a derived or session-scoped capabilities object (a `store.transaction`
session, a projected backend) has no way to re-resolve that plan for
itself, but it carries the root's answer forward instead of losing it: it
reads whether its own source object was already `"optimistic-retry"` and
keeps the tier for as long as `interactiveTransactions` stays `true`,
falling back to `"interactive"` only for a capabilities object whose source
never carried the tier to begin with.
Two further internal readers key off the `"batch"` value: the batch-tier
write verdict (`resolveBatchWriteVerdict`) that produces
`BATCH_WRITE_UNSUPPORTED` refusals, and the autocommit single-statement
eligibility gate that decides whether a supplied-id singleton create can
fuse its schema fence into one statement. Even absent `"optimistic-retry"`,
`unitOfWork` exists so any caller can tell the execution shapes apart without
re-deriving the same distinction from `interactiveTransactions` and
`atomicBatch` separately.
Under `"optimistic-retry"`, every TypeGraph-owned transaction that acquires a
fence row replays a real commit-time conflict as a whole unit, up to
`OPTIMISTIC_RETRY_ATTEMPTS` (3) attempts, and only exhausting that budget (or
a non-retryable failure) surfaces `TransactionConflictError` to the caller —
see [Retrying on conflict](/schemas-stores#retrying-on-conflict). That covers
every store-owned write (collection create/update/delete, bulk paths,
`importGraph`, identity maintenance, contribution rebuild, index
materialization) as well as the two backend-owned transactions that acquire
the schema-commit fence row directly, outside the store's own write path:
graph-template instantiation and a schema commit (`commitSchemaVersion` and
its three siblings, via `runSchemaWriteTransaction`). A nested write running
inside an existing transaction (`store.transaction`, an adopted transaction)
never retries on its own: it cannot restart a transaction it does not own, so
its conflict propagates unchanged to the outermost store-owned write, or to
`store.transaction` itself. This tier therefore changes behavior only for a
transaction that opens its own top-level connection.
An `"optimistic-retry"` backend requires `node:async_hooks`' `AsyncLocalStorage`
to detect a retried unit nested inside another one; on a runtime where it is
unavailable, the first retried unit `runRetriedUnit` opens is refused with
`OPTIMISTIC_RETRY_REQUIRES_ASYNC_CONTEXT` rather than degrading to independent,
unsafe per-unit retries, while interactive backends are unaffected and keep
retrying (`store.transaction`'s own `retry` option) with no async-context
support at all.
`graphAnalytics.supported` describes the backend shape, not mutable PostgreSQL
session state. A hot standby or a role without the database `TEMP` privilege can
still reject the working-table transaction that the iterative graph algorithms
open: a standby refuses the read-write transaction itself, and a role without
`TEMP` refuses the `CREATE TEMP TABLE` inside it. Both refusals reach the caller
as `UnsupportedBackendCapabilityError`, with the PostgreSQL error retained as
its `cause`.
### Durable edge match identity capability
`capabilities.durableEdgeMatchIdentity: true` is a correctness promise. A
custom backend making it must provide all of these guarantees:
- Every edge write carrying `InsertEdgeParams.matchIdentity` stores both the
name and key with the row. They are either both absent or both present.
- A database constraint atomically owns uniqueness over `(graph_id, kind,
match_identity_name, match_identity_key)`. Soft deletion keeps the key;
physical hard deletion releases it.
- `commands.execute()` handles a durable `edge.converge-create` as one database
decision and returns the authoritative `created` or `found` row. Returning
`unsupported` fails closed with
`DURABLE_EDGE_MATCH_IDENTITY_COMMAND_UNSUPPORTED`; TypeGraph does not fall
back to a read-then-write race.
- Storage exists before runtime writes. Implement
`ensureEdgeMatchIdentityStorage` for privileged schema adoption, or provision
the columns, pair constraint, and unique arbiter independently before setting
the capability.
`insertEdgesDurableBatchReturning` is an optional throughput member. When
implemented, every input must carry a durable identity, conflicts are omitted
from the returned rows, and returned rows identify exactly which inputs were
created. Omitting it preserves correctness through per-row authoritative
commands, but loses the set-oriented bulk/import fast path.
`findEdgesByHeterogeneousEndpointSet` is likewise an optional set-read
optimization. An input carrying `opposite` requests an exact directed endpoint
pair, not every edge incident to the first endpoint. One call may contain only
incident inputs or only exact-pair inputs; mixing the two modes is refused. A
backend that omits the member retains the exact per-pair fallback.
### Validity-end clearing capability
Custom backends must advertise `capabilities.clearValidTo: true` only when both
`updateNode` and `updateEdge` apply `clearValidTo: true` by storing SQL `NULL` in
`valid_to`. The built-in SQLite and PostgreSQL adapters do. An explicit clear on
a backend without that promise is refused with `ConfigurationError` code
`CLEAR_VALID_TO_UNSUPPORTED` before coalescing or writes, so the result
does not depend on whether the target row is already open. Omission still means
preserve; custom backends that do not support clearing remain compatible with
all writes that omit the option.
### Recorded-table migration DDL (`recordedTableDdl`)
`GraphBackend.recordedTableDdl` is an optional, synchronous callback used only by the offline
timestamp-only recorded-time preview migration. The migration calls it twice, once with temporary
table names and once with the final names, and expects DDL for `recordedClock`, `recordedNodes`, and
`recordedEdges`. The backend owns this callback because table creation, indexes, and named
constraints are dialect-specific and must not pull Drizzle into portable entrypoints.
A custom backend can omit the callback unless it created data in the old preview schema. If
`migrateLegacyRecordedTime` discovers that schema and the callback is absent, it throws
`UnsupportedBackendCapabilityError` with `details.capability: "recordedTableDdl"`. When the engine
names primary-key constraints, the temporary and final callback results must either both name the
constraint or both omit it; a one-sided result throws `ConfigurationError` code
`RECORDED_DDL_CONSTRAINT_NAME_MISMATCH`.
The callback only describes DDL. It must not execute statements or inspect the catalog, because the
migration invokes it inside its transaction. See
[Migrating Preview Recorded Time](/schema-management#migrating-preview-recorded-time) for the
operator workflow.
### Recursive traversal capability
Both bundled backends declare `capabilities.recursiveTraversal: { supported: true }`. **Absent
means supported** — mirroring `returning`, not `constraintClaims`: every existing custom backend
already runs the six recursive-CTE emission sites unconditionally, so absence meaning unsupported
would refuse traversals that work today.
```typescript
const capabilities: Partial = {
recursiveTraversal: { supported: false, reason: "engine has no WITH RECURSIVE / equivalent" },
};
```
A backend that genuinely lacks the primitive declares `{ supported: false, reason }`. A factory
refuses a contradictory declaration — `supported: false` with no `reason`, or `supported: true`
with a dangling `reason` — with `ConfigurationError` details code `CAPABILITY_DECLARATION_CONTRADICTION`.
Five operations refuse when unsupported: variable-length (`traverse`) queries, `store.subgraph()`,
historical identity class reads, identity-expanded historical queries, and the identity
window-ledger read — each throwing `ConfigurationError` code `RECURSIVE_TRAVERSAL_UNSUPPORTED`
with `details.operation` naming the site and `details.reason` echoing the declaration.
`weightedShortestPath` is the one exception: on a backend with temporary statements but no
recursion, it **falls back** to a per-hop predecessor walk instead of refusing, issuing
`pathLength + 1` extraction statements for the path a recursive CTE would have returned in one
round trip. The unweighted `shortestPath` (along with `reachable`, `canReach`, and `neighbors`)
emits no recursive CTE at all — it routes through the iterative working-table or inline path
instead — so it neither refuses nor falls back regardless of this declaration.
### Write fence declaration (writeFence)
TypeGraph serializes a family of writes — Operational Identity's mutations, and the
TypeGraph-owned recorded-clock allocation behind `history` / `revisionTracking` — behind a
per-graph fence rather than trusting the engine's default isolation. `capabilities.writeFence`
declares what this backend can provide, as two independent facts, and `resolveWriteFencePlan` is
the one place that declaration turns into a plan every lock site consumes instead of re-deriving:
```typescript
const capabilities: Partial = {
writeFence: { mechanism: "advisory", drain: "table-lock" },
};
```
`mechanism` is how the backend excludes concurrent writers. `writeFence` is a discriminated union on
`mechanism`, and `drain` is a field of the `"advisory"` and `"row"` shapes only — `"engine-serialized"`
and `"caller-serialized"` declarations carry no `drain` key at all:
| `mechanism` | Meaning |
| --- | --- |
| `"advisory"` | A keyed `pg_advisory_xact_lock`-style lock a caller takes explicitly. Needs `fenceSql` (below) and a `drain`. |
| `"row"` | A keyed exclusion spelled by TypeGraph itself against the never-dropped fences relation, for an engine with no advisory-lock primitive. Needs a `drain` and a `conflict` (below); a `fenceSql.isolationFactExpression` is optional (absent means recorded capture and match-key convergence fail closed on an unknown isolation fact, exactly as they do for a target that supplies neither). |
| `"engine-serialized"` | The engine serializes writers by construction — SQLite's single writer slot. No lock statement, no `fenceSql`, no `drain`. |
| `"caller-serialized"` | A deployment-level promise, not an engine fact — see below. No lock statement, no `drain`; a `fenceSql` the backend still carries is used only for its isolation-fact read (recorded capture's isolation guard). |
`drain` (on `mechanism: "advisory"` or `"row"` only) is a separate fact: whether a caller that
already excluded other writers can additionally take a relation-wide lock on the table a drain site
protects:
| `drain` | Meaning |
| --- | --- |
| `"table-lock"` | Yes — a `LOCK TABLE`-style statement is available and the drain site takes it. |
| `"quiescent"` | The resource is already exclusive for some other reason (an advisory or row lock layered under a deployment's own `caller-serialized` promise, for instance), so the drain site takes NO statement — one it does not need rather than one it cannot spell. |
| `"none"` | Neither — a drain site refuses, naming the drain. |
`conflict` (on `mechanism: "row"` only) is the engine fact for what happens when two writers
acquire the SAME fence row:
| `conflict` | Meaning |
| --- | --- |
| `"wait"` | A lock-based engine — the second acquirer's statement blocks until the first commits, exactly like an advisory lock. |
| `"commit-time"` | An optimistic-concurrency engine — both acquirers proceed and the loser's COMMIT fails. Correctness comes from the retry owner replaying the whole unit, never from waiting, so `conflict: "commit-time"` derives the `"optimistic-retry"` execution tier (see [Backend Capabilities](#backend-capabilities) above) and requires it: declaring it on a non-interactive backend is refused the same way an out-of-place `drain` is. |
`"engine-serialized"` and `"caller-serialized"` satisfy every drain site unconditionally — a writer
slot and an in-process serialization promise are each already a stronger exclusion than any `drain`
value could add, so attaching one to either mechanism is refused (see **Runtime validation** below)
rather than silently ignored; attaching `conflict` to anything but `"row"` is refused the same way.
`resolveWriteFencePlan` resolves one of five plans:
- `{ kind: "lock", drain, sql }` — take the declared advisory lock (`sql`, the target's own
spelling), and, when `drain === "table-lock"`, the table lock a drain site needs.
- `{ kind: "row", drain, conflict, sql }` — take the SAME `sql.acquireKeyed` /
`sql.acquireKeyedWithIsolation` a `"lock"` plan's site calls, spelled instead against the fences
relation; `conflict` is the one fact a `"row"` site (and the execution tier) reads that a
`"lock"` site never needs.
- `{ kind: "engine-serialized" }` — no lock needed; the engine serializes writers by construction.
- `{ kind: "caller-serialized" }` — no lock needed; the deployment's own promise excludes concurrent
writers (see below).
- `{ kind: "unfenced" }` — no declaration at all. Every fence that guards a read-then-write across
statements refuses rather than running unfenced. Only a predicate carried *inside* the statement
it guards degrades, since one statement cannot race itself.
Resolution order: (1) the declared `writeFence` value, if present; (2) absent, AND the backend was
built by `createSqliteBackend` / `createPostgresBackend` — derived from `dialect`, which is exactly
what every lock site used to compute inline (this derivation never resolves `"row"`: it is the two
bundled dialects' own `"advisory"`/`"engine-serialized"` split); (3) absent on anything else —
`unfenced`, because an undeclared custom backend is by definition uncertified and inferring lock
support from `dialect` alone is the unsound inference this capability replaces.
The two bundled backends resolve exactly these declarations — copy the one matching your engine:
- PostgreSQL: `writeFence: { mechanism: "advisory", drain: "table-lock" }`
- SQLite: `writeFence: { mechanism: "engine-serialized" }` (no `drain`: the writer slot already
excludes every drain site's writer, so a drain site under it always takes no statement — the same
behavior `drain: "quiescent"` describes for `"advisory"`/`"row"`, without a `drain` field to spell it)
A backend that declares `mechanism: "advisory"` also supplies `fenceSql`: `lockTables` (only needed
when `drain: "table-lock"`) plus the two composable, no-`SELECT` forms `advisoryLockExpression` /
`isolationFactExpression` a statement embeds inside a larger query it builds itself (see the
schema-write-fence discussion above) — the complete `FenceSql` bag. `resolveWriteFencePlan`'s `lock`
arm derives the standalone-statement forms every ordinary lock site actually calls — `acquireKeyed`;
`acquireKeyedWithIsolation` (the lock plus the session's isolation-level fact, read in the same
statement it locks in); and `isolationFact` — from those two expressions, so a backend author never
spells both forms separately. `mechanism: "row"` needs no `advisoryLockExpression` at all: TypeGraph
spells its own `acquireKeyed` / `acquireKeyedWithIsolation` against the fences relation (see below),
and a `fenceSql.isolationFactExpression` — when supplied — rides the SAME acquisition statement, so a
`"row"` target's isolation fact is read on the exact connection that took the row. The bundled
PostgreSQL spelling is exported as `postgresFenceSql` from `@nicia-ai/typegraph/adapters/drizzle/postgres`
— pass it straight through as `fenceSql` when wrapping that backend (under either `"advisory"` or
`"row"`), or supply a custom `FenceSql` matching a different engine's lock syntax. A backend that
declares `mechanism: "advisory"` with a `fenceSql` missing a member the resolved `mechanism`/`drain`
combination needs is refused at construction with details code `WRITE_FENCE_SQL_UNAVAILABLE`, naming
the missing member; `"row"` is refused the same way only for `drain: "table-lock"` without
`lockTables` — its acquisition statement needs no author-supplied spelling at all, so a missing
`tableNames.fences` instead refuses the first time a keyed site actually acquires the row, not at
construction; `"engine-serialized"` and `"caller-serialized"` need no `fenceSql` to take a lock at
all.
#### The fences relation
A `"row"`-mechanism backend needs one relation, `typegraph_fences(key TEXT PRIMARY KEY, generation
BIGINT NOT NULL)` (`INTEGER NOT NULL` on SQLite) — part of TypeGraph's base schema on both bundled
dialects, so a fresh install already has it and `generateSqliteMigrationSQL` /
`generatePostgresMigrationSQL` add it to an existing database. It is **never dropped, never cleared
by `clear()`, and never row-deleted** — the same durability contract `schema_versions` and
`recorded_clock` carry. Every acquisition is one portable statement TypeGraph spells itself, never
the profile: `INSERT INTO typegraph_fences (key, generation) VALUES (key, 1) ON CONFLICT (key) DO
UPDATE SET generation = generation + 1 RETURNING generation`, keyed on `${namespace}:${key}` —
the SAME advisory namespaces and per-position keys an `"advisory"`-mechanism backend locks on,
verbatim, so the lock-order contract carries over unchanged to an engine using the fences relation
instead of `pg_advisory_xact_lock`. A custom backend supplies the relation's physical name through
`tableNames.fences` (defaulted to `typegraph_fences` by both bundled factories) exactly as it names
every other TypeGraph-owned table.
#### Runtime validation
TypeScript's discriminated union only holds a caller who goes through the type checker — a plain
JavaScript backend author, or a value round-tripped through JSON or a config file, can still supply
an unrecognized `mechanism` string, an unrecognized `drain` or `conflict` string, a `drain` attached
to `"engine-serialized"` / `"caller-serialized"`, or a `conflict` attached to anything but `"row"`.
`resolveWriteFencePlan` validates every declaration — whether it came from `capabilities.writeFence`
directly or from the first-party dialect fallback — before shaping a plan from it, and refuses with
`ConfigurationError` details code `WRITE_FENCE_DECLARATION_INVALID`, naming the invalid `field`
(`"mechanism"`, `"drain"`, or `"conflict"`) and, for an unrecognized value, the `accepted` list. An
unrecognized `drain` never falls through to behaving like `"quiescent"` — it is refused outright,
the same as an unrecognized `mechanism`. The same validator refuses `conflict: "commit-time"`
outright when the target's own `capabilities.execution.interactiveTransactions` is `false`: that
value is honored only by the `"optimistic-retry"` execution tier, which never derives without an
interactive transaction to replay inside, so accepting the declaration there would silently drop it
rather than apply it.
#### `caller-serialized`: the promise split into two halves
`caller-serialized` is for a deployment that knows its database has no other concurrent writer, but
whose engine is neither an advisory-lock engine nor a single-writer one — a PostgreSQL-wire engine
with no working `pg_advisory_xact_lock` / `LOCK TABLE`, for example. The promise has two halves,
and TypeGraph only enforces the first:
- **In process**, TypeGraph enforces it: every root member the backend classifies as mutating —
collection writes, `store.transaction` / `transactionWithNative`, schema commits, identity and
contribution maintenance, index materialization, table/DDL provisioning, `clearGraph`, import,
and the raw-SQL members (`execute`, `executeRaw`, `executeStatement`,
`executeTemporaryStatement`) that can carry an arbitrary write — runs through one per-backend
serialized queue, so two concurrent calls through one pool cannot race each other. A root write
awaited from inside a `store.transaction` callback is refused rather than left to deadlock behind
the transaction's own queue slot.
- **Outside the process**, the deployment enforces it: no other client writes to this database
while this backend is open. TypeGraph cannot see or verify that half; declaring
`caller-serialized` is asserting it.
Adopting an externally owned transaction (`store.withTransaction(externalTx)`, backed by
`adoptTransaction`) is refused outright on a `caller-serialized` backend, with `ConfigurationError`
details code `CALLER_SERIALIZED_REFUSES_ADOPTION`: an adopted transaction's lifetime belongs to the
caller, not to this backend's write-unit queue, so there is no honest way to hold a queue slot open
for it — queuing it would block every other queued write until the caller's own transaction ends,
and leaving it unqueued would let its writes interleave with the queue's own, silently breaking the
promise `caller-serialized` makes. Open the transaction through this backend's own `transaction()` /
`transactionWithNative()` instead, or do not declare `caller-serialized` on a backend that needs
cross-store adoption.
`createPostgresBackend` accepts `writeFence: { mechanism: "caller-serialized" }` — a claim about the
deployment — while still refusing `mechanism: "engine-serialized"` outright, because that value is
a claim about the *engine*, which this factory's own engine does not back.
Constructing Operational Identity, or `history: true` / `revisionTracking: true`, against an
`unfenced` backend is refused immediately at `createStore` — never mid-flush — with
`ConfigurationError` details code `IDENTITY_REQUIRES_WRITE_FENCE` (identity) or
`RECORDED_CLOCK_REQUIRES_WRITE_FENCE` (recorded-clock allocation), and the refusal message names
the exact declaration line to add.
A `lock` plan whose `drain` cannot back a site's `requires: "drain"` — `drain: "none"` — is
refused with details code `WRITE_FENCE_UNAVAILABLE`, naming `details.operation` and the drain that
could not be satisfied. `"engine-serialized"` and `"caller-serialized"` satisfy either `requires`
value (`"keyed"` or `"drain"`) without consulting `drain`.
The PostgreSQL schema fence refuses too, and it is worth knowing why it is not on the
degradable side. The per-graph advisory lock plus `SELECT ... FOR UPDATE` a schema commit takes,
and the `FOR SHARE` a managed write takes on that same row, each fence a read-then-write sequence
that spans **statements**: `commitSchemaVersion` reads the active version and then writes the
flip, and a managed write holds its `FOR SHARE` for the remainder of the transaction so the
version it asserted stays true through the writes that follow. Skipping those locks would not
give a slower-but-correct path; it would assert a version and then let the very change the
assertion was checking for land before the write. So a PostgreSQL-dialect backend that resolves
`unfenced` is refused at the schema commit with `WRITE_FENCE_UNAVAILABLE`, naming the operation.
The one part that *does* degrade is the fence folded into a managed insert's own statement. That
predicate is evaluated inside the INSERT that depends on it, and one statement cannot race
itself: with no locking clause the fence subquery still yields no row when the expected version
is no longer active, so the INSERT still writes nothing. This is how SQLite has always run the
path, on the strength of its writer slot.
This matters for a PostgreSQL-wire engine that implements neither `pg_advisory_xact_lock` nor the
`FOR UPDATE` / `FOR SHARE` clauses, and whose engine merges concurrent transactions rather than
serializing them.
`writeFence` has no arm for "no exclusion mechanism at all" — every `mechanism` value claims
something real. An engine with neither locks nor a writer slot nor a caller promise to make has
nothing honest to declare, and `unfenced` is how that shows up downstream — but neither bundled
factory will hand you that backend. `createPostgresBackend` and `createSqliteBackend` both build on
`createSqlBackend`, which refuses at construction, with `ConfigurationError` details code
`ENGINE_PROFILE_REQUIRES_WRITE_FENCE_DECLARATION`, when the resolved capabilities carry no
`writeFence` at all — including `capabilities: { writeFence: undefined }` passed to either factory,
which no longer builds a backend the way it once did:
```typescript
createPostgresBackend(db, {
capabilities: { writeFence: undefined },
});
// throws ConfigurationError: ENGINE_PROFILE_REQUIRES_WRITE_FENCE_DECLARATION
```
`unfenced` is reachable only outside that gate: a hand-assembled `GraphBackend` that never goes
through `createSqlBackend`, or a custom `SqlEngineProfile` whose `declaredCapabilities.writeFence`
some other override clears before it reaches a call site — never through either bundled factory.
Getting there is a way of admitting the engine truly has nothing to declare; TypeGraph then refuses
a schema-managed store built on it at `createStore`, naming the missing capability, rather than
running a fence the engine cannot enforce.
If the deployment instead knows it is the only writer of this database — a pool clamped to one
connection, or a single-writer topology otherwise enforced outside TypeGraph — declare
`writeFence: { mechanism: "caller-serialized" }` instead (see above): that is
the honest way to spell a deployment convention. Do not reach for `mechanism: "engine-serialized"`
for the same purpose — that declaration means the *engine* serializes writers by construction, and
a deployment convention is not a construction. `createPostgresBackend` refuses that particular
claim outright for this reason.
### Capability bundles
A **capability bundle** groups a set of `GraphBackend` members that one operation family needs
together, with one verdict resolver and one member accessor, so a caller never re-derives "does
this backend support X" from a scattered `undefined` check. Seven pilot bundles ship in this
release:
| Bundle | Kind | Disposition |
| ------------------------- | ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `claims` | gated | Bidirectional cross-check between the `constraintClaims` declaration and the core members; disagreement in either direction refuses with `CONSTRAINT_CLAIM_SURFACE_MISMATCH` |
| `statementExecution` | gated | Core `executeStatement` absent refuses with `IDENTITY_REQUIRES_STATEMENT_EXECUTION` |
| `recordedRevisionOrigins` | gated | Core `ensureRevisionOriginsTable` absent refuses with the operation's own typed error |
| `batchPointRead` | graduated | `getNodes` absent falls back to per-id `getNode`; `getEdges` absent falls back to per-id `getEdge` |
| `uniqueSidecarBatch` | graduated | `insertUniqueBatch` absent falls back to `issueClaimsIndividually`; `checkUniqueBatch` absent falls back to a per-key loop; `hardDeleteUniquesByNodeIds` absent refuses with the operation's own typed error |
| `contributionHealth` | graduated | `verifyContributions` / `repairContributions` / `rebuildContribution` absent each refuse with the operation's own typed error; `probeContributions` absent falls back to `{ entries: [] }` |
| `endpointSetRead` | graduated | `findEdgesByEndpointSet` absent refuses set-oriented `bulkFindFrom` / `bulkFindTo` with `ENDPOINT_SET_READ_UNSUPPORTED`; singleton reads remain available |
The port-mismatch rule that governs every bundle's member accessor is keyed to the disposition,
not blanket: a `refuse`-disposition row whose backend object cannot actually reach the member
throws that bundle's own `portSurfaceCode` (`CONSTRAINT_CLAIM_SURFACE_MISMATCH` for `claims`,
`BUNDLE_PORT_SURFACE_MISMATCH` for the other six); a `fallback`-disposition row whose port cannot
reach the member takes its declared fallback instead of throwing — the verdict said the member
was there, the object it binds against says otherwise, and a fallback row is defined to degrade
rather than assert.
This bundle model ships for **seven of the twenty-one** member-bearing operation families measured
in this workstream; the remaining fourteen are a named follow-up workstream, not a silent gap —
their members keep working exactly as before, unbundled, with an access-count ceiling that
prevents new scattered checks from accumulating ahead of that follow-up.
A backend author does not need to do anything for these seven bundles today: both bundled backends
already carry every core member each bundle's `dialects` scope requires. The atomic transport
conformance runner is the foundation for certifying a **third-party** backend: the author supplies
engine-specific statements, state observers, and exact-root provenance checks, while the runner
asserts the shared transport contract. Bundle verdicts remain a separate check against the declared
capabilities and the object the calls actually execute on. A backend must therefore make every
declared capability (`constraintClaims`, `contributions`, and execution support) truthful about
what the active backend object implements, not just which fields it sets.
Run the conformance fixture in the custom backend's own test suite, then pair
the earned declaration with the exact root transport in its factory:
```typescript
import {
decorateBackend,
registerAtomicMutationPrograms,
registerAtomicSqlProgram,
runAtomicMutationProgramConformance,
runAtomicTransportConformance,
} from "@nicia-ai/typegraph/backend";
const backend = createCustomBackend({
execution: {
interactiveTransactions: false,
atomicBatch: "root",
},
});
registerAtomicSqlProgram(backend, { executeAtomicBatch });
const authorCreatedWrapper = decorateBackend(backend, {});
await runAtomicTransportConformance({
...transportCases,
backend,
derivedBackends: [authorCreatedWrapper],
executeAtomicBatch,
});
registerAtomicMutationPrograms(backend, mutationPrograms);
const semanticCases = buildSemanticCases({ backend });
await runAtomicMutationProgramConformance({
backend,
derivedBackends: [authorCreatedWrapper],
equal: deepEqual,
cases: semanticCases,
});
```
Transport registration is exact-resource evidence only: a derived backend does
not inherit it, and a second registration on the same object is refused rather
than replacing the function production uses. A bundled PostgreSQL transaction
session earns a separate registration bound to its pinned client; it does not
inherit the root's registration.
Create wrappers with the exported `decorateBackend()` seam so the runner can
verify their lineage back to the registered root instead of accepting an
unrelated object as derivation evidence. The conformance fixture's mandatory
provenance checks prove registration, lineage, derived isolation, and—when
applicable—transaction isolation against the real objects supplied by the
backend author. A non-interactive root reports the
transaction-isolation check as skipped rather than claiming evidence it could
not obtain. Transport registration certifies mechanics, not graph semantics, and
therefore does not by itself opt a custom backend into any Store mutation
program.
The separate `registerAtomicMutationPrograms()` call is the semantic boundary:
each member declares one complete TypeGraph mutation family implemented by that
exact backend resource. Omitted families retain the portable path, and an empty profile or
a profile registered before its atomic transport is refused with
`ConfigurationError`.
The semantic executors must preserve the same schema fence, validation,
side-effect, refusal classification, rollback, postimage correlation, result
ordering, and bind-ceiling contracts as the bundled implementation. Registering
one family is not evidence for another. Derived and projected backends inherit
neither registration. An exact transaction session must be registered
independently before Store code can dispatch through it.
`runAtomicMutationProgramConformance()` is the executable semantic boundary.
For every reachable positive-limit variant in `mutationPrograms`, the fixture supplies
three real Store-level cases:
1. an ordered success whose return value and independently read committed state
both match their oracles;
2. a stale-schema-fence refusal that leaves the database unchanged; and
3. a family-specific typed refusal that either rolls back every sibling write
after native dispatch or explicitly refuses before dispatch without writing.
The runner resolves the profile from `backend`; it does not accept a detached
profile description, caller-supplied provenance verdict, or fixture-owned
dispatch counter. Before any fixture preparation can write, it validates the
complete case inventory and probes the author's actual derived backend objects.
It observes dispatch inside the exact registered executors and therefore refuses
a success that silently used the portable fallback,
a case bound to a different family claim, a missing or duplicate family case,
and a case that claims an unregistered family. A zero entry limit is an honest
opt-out and does not require an unreachable case. `mutateEdges` has separate
`resolvedSet` and `durableConvergence` variants because proving one does not
prove the other.
Every semantic case identifies the exact `backend` its callbacks use. The
runner checks that binding and the registered profile identity before any
preparation, again between preparation and execution, and after execution, so
a pre-dispatch refusal or a mid-run registry replacement cannot borrow another
root's certificate. The fixture callbacks should invoke public Store methods and inspect committed
rows through an independent database read. Supply at least one real wrapper or
derived backend created with `decorateBackend()`; the runner does not manufacture
a projection and mistake that tautology for author evidence. Do not instrument
or replace the registered executors—the runner owns dispatch evidence. Run
conformance with exclusive use of that exact root: unrelated same-variant writes
during the observation window cannot be distinguished from fixture traffic.
Mark each semantic refusal's `dispatch` as `"required"` or `"pre-dispatch"`
according to the Store contract, and do not use executor return rows as the
state oracle. Stale-fence cases always require native dispatch regardless of a
fixture value supplied by untyped JavaScript. Match
Store-level typed errors rather than raw driver sentinels. The runner is
framework-agnostic, so the same fixture runs in the custom backend's own test
suite. Pair it with the shared cross-backend Store integration suite; transport
conformance alone cannot prove graph semantics.
The profile is family-scoped:
| Member | Store operations authorized |
| ----------------------------- | ----------------------------------------------------------------------------------------- |
| `createNodes` / `createEdges` | Eligible direct `bulkInsert()` and `bulkCreate()` programs |
| `replaceNodes` | Eligible complete-document `nodes.bulkReplaceById()` programs |
| `deleteNodes` / `deleteEdges` | Eligible direct `bulkDelete()` programs |
| `updateNodes` / `updateEdges` | Eligible resolved update-only sets |
| `mutateNodes` / `mutateEdges` | Eligible mixed create/update sets; the edge family also owns durable endpoint convergence |
Executor limits such as `maxEntries`, `replaceNodes.maxEntries.plain`,
`replaceNodes.maxEntries.claimed`,
`createNodes.claimSupport.maxInputCostPerEntry`, and the two edge mutation
ceilings are part of the registration contract and must be nonnegative
integers; zero honestly declares that the backend's bind budget cannot admit
one member of that shape. TypeGraph validates those declarations before
publishing the exact-root profile. `claimSupport.families` explicitly
advertises `uniqueness` and/or `disjointness`; an empty list with a zero bound
honestly opts out of all claim work. The Store calls the exported
`atomicNodeClaimInputCost()` owner for each member and refuses the native path
when its complete normalized claim set exceeds the executor's declared bound.
Custom executors must use that same helper instead of reproducing its
dialect-reviewed bind formula. `deleteNodes.releasedClaimFamilies` similarly
declares which owner-side claim cleanup the delete program proves.
Bundled replacement executors also expose an `accepts(entries)` pre-dispatch
proof. It packs prepared members with the same bind-weighted planner used by
execution, so claimed batches are admitted by their actual work instead of an
unrelated fixed 32-entry ceiling; `false` is an explicit no-SQL fallback
verdict. Custom executors may provide the same exact admission seam when one
claimed-member ceiling would be needlessly pessimistic.
`replaceNodes.releasedClaimFamilies` declares which previous owner claims the
replacement releases before acquiring its complete postimage claims; the Store
does not infer that proof from `claimSupport`. Node
create/update/mutation executors advertise derived-storage support separately
through `projectionSupport.families`; omission or an empty list honestly opts
out, and the Store never infers projection safety from transport registration
alone. The supported families are `fulltext` and `embedding`.
On a transactionless root, dedicated
update-only and mixed mutation executors are independently reachable Store
families even when their entry ceilings are equal, so each requires its own
conformance evidence. On an interactive root, the collection-level
read/partition/write unit moves into a transaction and exact-root registration
does not follow; the root conformance inventory therefore excludes the mixed
variants while continuing to require direct create, delete, update, and durable
convergence evidence. Bundled PostgreSQL binds the same reviewed lowering to the
exact transaction session and exercises the mixed node and edge variants against
a real engine, including typed refusal rollback. The same session profile
registers `replaceNodes`, so a caller-owned PostgreSQL transaction keeps blind
replacement inside its savepoint-backed atomic program.
An exact `atomicBatch: "session"` conformance fixture includes those mixed
variants even though `interactiveTransactions` is true; nested-transaction
isolation is reported as inapplicable because the fixture resource is already
the open transaction. The transport runner accepts the same exact-session
resource and certifies its ordered slots, parameter preservation, rollback,
and empty-program behavior. Generic derived session objects still lose both
the declaration and the identity-bound registrations.
Backend authors implementing edge
refusal paths use the exported
`AtomicEdgeBatchEndpointRefusalError`,
`AtomicEdgeBatchCardinalityRefusalError`,
`AtomicEdgeConvergenceTombstoneRefusalError`, and
`AtomicEdgeDeleteIdentityRefusalError` signals; restricted node deletion uses
`AtomicNodeDeleteRestrictedRefusalError`. This preserves the Store's existing
typed diagnostic classification rather than exposing driver-specific sentinel
errors.
Execution support is intentionally not collapsed into one ordered “tier.” An
interactive transaction and an exact-root atomic batch are independent facts:
a backend may provide either, both, or neither. Each Store operation selects
the boundary its own semantics require instead of treating one mechanism as a
universal substitute for the other.
Endpoint-set reads have a small, independent conformance fixture for custom
backends. Import `runEndpointSetReadConformance` from the `backend` entrypoint
and provide the exact backend, one or more successful `FindEdgesByEndpointSetParams`
cases, and at least one refusal case. The runner checks that the backend exposes
`findEdgesByEndpointSet`, preserves the expected edge rows, and refuses invalid
requests using the adapter's typed error. It does not create a schema or assume
a driver, so the same fixture can run against any engine. A backend that omits
the member remains valid for singleton reads; Store bulk endpoint reads refuse
with `ENDPOINT_SET_READ_UNSUPPORTED`.
### Declared constraints require an interactive transaction
A **constrained write** — one whose correctness rests on a check-then-write that
no database key repeats at write time — runs its probe and its write under one
per-graph mutual exclusion. That fence is a transaction-scoped construct on both
dialects: SQLite's `BEGIN IMMEDIATE` writer slot, PostgreSQL's
`pg_advisory_xact_lock` (which outside a transaction is taken and dropped inside
its own implicit single-statement one, excluding nothing). A backend reporting
`capabilities.execution.interactiveTransactions: false` can supply neither, so such a write is
**refused** rather than run unfenced — a constraint enforced only when nothing
races is the defect the fence exists to close.
The refusal is a `ConfigurationError` with `details.code`
`CONSTRAINT_WRITE_FENCE_UNSUPPORTED`, and `details.constraint` naming which
class needed the fence, because the way forward differs per class:
| `details.constraint` | The write that needs the fence | Way forward without a transactional backend |
| --- | --- | --- |
| `edgeCardinality` | Creating or resurrecting an edge whose `cardinality` is `one`, `unique`, or `oneActive` | Declare the edge `cardinality: "many"` and enforce the limit in application code |
| `edgeMatchKeyConvergence` | `getOrCreateByEndpoints` using an undeclared dynamic `matchOn` key | Declare the edge registration's durable `matchIdentity`, or use `create` with a caller-chosen id |
| `nodeDisjointness` | Creating a node under a kind that participates in a `disjointWith` axiom | Drop the axiom and keep ids distinct across those kinds yourself |
| `nodeUniquenessClaim` | **Updating or resurrecting** a node whose kind declares any unique constraint, of any scope — a transition reserves the new key *before* the row write it gates, and only a transaction can undo the pair together | Drop the constraint, or run updates on a transactional backend. Plain **creates** under a `scope: "kind"` unique are unaffected: their claim follows the row |
| `nodeUniquenessScope` | Creating **or updating** a node under a `scope: "kindWithSubClasses"` unique that actually expands past the node's own kind | Scope the constraint to `"kind"`, which the uniques primary key enforces on its own |
`importGraph` / `importGraphStream` is refused on the same backends whenever any
node kind of the graph owes a claim ahead of its row — that is, declares **any**
unique constraint or has a disjoint partner — or any edge kind is non-`many`.
The import writes both creates and updates, so the widest of those placements is
what decides it.
This affects **Cloudflare D1**, **`drizzle-orm/neon-http`**, and any SQLite
backend built with `transactionMode: "none"`. Durable Objects are unaffected —
`do-sqlite` reports `capabilities.execution.interactiveTransactions: true` and fences normally.
Unconstrained writes on those backends are untouched and keep working exactly as
before: a `cardinality: "many"` edge created, updated and deleted; any node
delete, including one whose kind participates in a disjointness axiom (a delete
re-derives no cross-kind verdict); a node whose uniques are all `scope: "kind"`;
and an undeclared `getOrCreateByEndpoints` that *finds* an existing edge in the default
`ifExists: "return"` mode, or resurrects a `many` one — that resurrection is an
id-keyed `UPDATE` that re-derives nothing. With `coalesceUnchangedUpserts`
enabled, confirming that a single `ifExists: "update"` endpoint replay is
unchanged requires the endpoint match-key convergence fence and therefore
refuses on these backends. Outside the native durable-convergence envelope,
the bulk `getOrCreateByEndpoints` form returns an all-live default-`"return"`
batch from one set-oriented root read because that outcome writes nothing.
Inside the native envelope, the authoritative upsert runs first; it preserves
the logical `"found"` outcome in one exchange but may take incumbent-row locks
and produce write amplification. If any member may write, the whole batch
retains that refusal on transactionless roots unless it matches the narrow native
durable-convergence envelope: schema-declared
`matchIdentity`, `cardinality: "many"`, declared match fields, default
`ifExists: "return"`, and no temporal mutation. That eligible form is one
closed atomic exchange; dynamic match fields, update mode, constrained
cardinality, temporal options, and all transaction-scoped or derived roots
retain the refusal or fallback path required by their contracts.
An otherwise eligible tombstoned winner cannot use the native path: the native
attempt rolls back and transactionless convergence refuses with the typed
`CONSTRAINT_WRITE_FENCE_UNSUPPORTED` (`edgeMatchKeyConvergence`) error. Use a
transaction-capable backend when schema-aware resurrection is required.
### Claim relations, and what they do not promise
Underneath the lock, a declared constraint is also reserved in a **claim
relation** whose primary key admits one live claimant per axis: `uniques` (for
uniqueness scopes and `disjointWith` pairs) and `typegraph_edge_claims` (for
`cardinality: "one" | "unique" | "oneActive"`). Both bundled backends carry them
and report `capabilities.constraintClaims: true`. The claim is what makes those
constraints hold for TypeGraph writers that hold no per-graph lock at all —
`importGraph` is the one in the box. The protocol is application-maintained:
raw SQL that writes only `nodes` or `edges` bypasses the corresponding claim
write and can violate the declaration. An out-of-band writer is fenced only if
it participates in the same claim protocol in the same transaction.
Three properties of that mechanism are worth knowing before you rely on it:
- **A claim row's lock is held to the end of the transaction, including on
refusal.** A caller that catches a typed constraint error and keeps going —
import's per-row recovery, or your own `try`/`catch` inside
`store.transaction` — still holds the lock on the row it was refused at, and
any other writer of that axis waits until the transaction ends. This is
inherent to every row-lock fence, not specific to this one.
- **Above READ COMMITTED, PostgreSQL reports a serialization failure instead of
the typed error.** At `REPEATABLE READ` or `SERIALIZABLE`, `INSERT … ON
CONFLICT DO UPDATE` raises `40001` rather than resolving the conflict, so the
losing writer sees a serialization failure to retry rather than
`UniquenessError`. SQLite has no such mode. This is unchanged from earlier
versions, which already reserved single-kind uniqueness through the same
statement.
- **Pre-existing violations are neither repaired nor refused at boot.** A
database that already held two live claimants of one axis before the claim
relations existed keeps holding them; the next write that touches that axis is
refused with the ordinary typed error naming the incumbent.
`store.verifyConstraintFences()` is the read-only diagnostic that makes that
state legible ahead of time:
```typescript
for (const violation of await store.verifyConstraintFences()) {
// violation.target names the claim row two claimants contend for
console.warn(violation.family, violation.target.axis, violation.target.key);
}
```
It reports one entry per contended axis — `nodeUniqueness` and
`nodeDisjointness` carry the conflicting `owners` (each a `concrete_kind` /
`node_id` pair, because ids are unique only per kind), `edgeCardinality` carries
the conflicting `edgeIds`. It reads the nodes, edges and `uniques` relations, so
it finds violations that predate the claim tables; it writes nothing, and it
repairs nothing — choosing which claimant keeps the axis is a data-loss decision
that stays with you.
### SQLite ↔ PostgreSQL parity
The **query language is fully portable** between SQLite and PostgreSQL. Predicates (comparison, string/`ILIKE`,
null, `between`, array, object, JSON-path), fixed and variable-length (recursive) traversals, bounded
neighbor reads, per-edge-kind subgraph windows, one-statement query batches, aggregates
(`count`/`sum`/`avg`/`min`/`max` with `groupBy`/`having`), set operations (`UNION`/`UNION ALL`/`INTERSECT`/`EXCEPT`,
including traversal, subquery, `GROUP BY`/`HAVING`, and per-leaf `ORDER BY`/`LIMIT`/`OFFSET` leaves), ordering with
`NULLS FIRST`/`LAST`, cursor pagination, temporal queries, and the fulltext query modes (`websearch`, `phrase`,
`plain`, `raw`) all behave identically. A query you write against one backend compiles and runs the same way on the
other.
The remaining differences are **engine and runtime capability gaps** — they
stem from what each database or hosted authorizer implements, not from
TypeGraph choosing separate query semantics per backend:
| Capability | SQLite | PostgreSQL | Behavior on the unsupported side |
| ------------------------------------------------------ | ------------------------------------------------- | ------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------- |
| Whole-graph temporary-table analytics | ✓ standard connections / ✗ D1 and Durable Objects | ✓ connection-based drivers / ✗ `neon-http` | Throws `UnsupportedBackendCapabilityError`; traversal algorithms with an inline engine fall back automatically |
| Vector metric `inner_product` | ✗ | ✓ | Rejected at compile time on SQLite (`sqlite-vec`/`libsql-native` expose `cosine` + `l2`; `pgvector` adds `inner_product`) |
| Vector index type `ivfflat` | ✗ | ✓ | Index declaration is **skipped** on SQLite (`indexTypes`: `hnsw`/`none` vs `hnsw`/`ivfflat`/`none`) |
| Filtered approximate search **guarantees** a full page | ✓ `sqlite-vec` / ✗ `libsql-native` | ✗ (`pgvector` recovers, but is bounded) | Only `sqlite-vec` guarantees it; the others can return **fewer than `limit`** rows under heavy filtering — see below |
| Per-query fulltext `language` override | ✗ | ✓ | Throws on SQLite — FTS5's tokenizer is fixed at table-create time; `tsvector` accepts a regconfig per query |
| HNSW `efSearch` query tuning | ✗ | ✓ transactional HNSW drivers | Refused, never ignored: `UnsupportedBackendCapabilityError` with `details.capability` `vector.searchFrontierTuning` on **any** SQLite backend (vector and hybrid alike — neither `sqlite-vec`'s `vec0` KNN nor `libsql-native`'s DiskANN has a per-search frontier), and on transaction-less Postgres or a non-HNSW slot |
| Bounded planner-statistics sampling | ✓ standard connections / ✗ D1 and Durable Objects | Native `ANALYZE` sampling | Restricted SQLite skips `analysis_limit` but still attempts scoped `ANALYZE`. Performance only — same results |
| TypeGraph Identity Profile | ✓ transactional drivers | ✓ transactional drivers | Enabled graphs fail fast on non-atomic drivers; identity-disabled graphs retain their ordinary path |
| Constraint claim relations (`capabilities.constraintClaims`) | ✓ | ✓ | Identical relations and identical statements on both dialects. A third-party backend that omits them declares `constraintClaims` absent and keeps the per-graph lock as its only fence |
| Durable edge match identity (`capabilities.durableEdgeMatchIdentity`) | ✓ bundled adapters | ✓ bundled adapters | Both dialects persist the same canonical key and use a unique database arbiter. A custom backend must satisfy the full capability contract above or leave the capability absent |
| Managed node projection fusion | ✓ registered atomic bulk programs; singleton fallback | ✓ registered atomic bulk programs; singleton create fusion | Eligible node bulk creates and resolved updates group fulltext/vector transitions into the same atomic program as their row mutations on both dialects. PostgreSQL additionally fuses an eligible singleton generated-ID create into one SQL statement when every active strategy supplies an inserted-node builder |
| Managed node claim fusion (`capabilities.atomicNodeInsertClaims`) | ✗ portable transactional fallback | ✓ PostgreSQL/PGlite | SQLite keeps claim acquisition and insertion in the portable transaction. PostgreSQL transaction receivers fuse supported claim plans; a root non-transactional receiver is limited to exactly one generated-id, same-kind uniqueness claim with no other side effects |
| Managed edge cardinality fusion | ✗ portable transactional fallback | ✓ PostgreSQL/PGlite transaction receivers | SQLite keeps its guarded claim and edge insert in the portable transaction. PostgreSQL can combine endpoint liveness, one cardinality claim, and the insert in one statement after any required graph lock |
| Atomic SQL transport (`capabilities.execution.atomicBatch`) | ✓ on certified D1/libSQL roots; otherwise `none` | ✓ on bundled recognized PostgreSQL drivers, including neon-http | `root` means the exact backend owns the atomic boundary; `session` means the exact object is already bound to an open transaction and the outer transaction owns commit/rollback. Both require identity-keyed executor registration. Neon HTTP uses its native transaction batch; session-capable `pg`, postgres-js, neon-serverless, and PGlite drivers can execute programs on one pinned Drizzle transaction. Unrecognized drivers remain `none`. A custom backend must pass the framework-agnostic conformance runner before opting in; omitted support keeps the portable path |
| Eligible registered managed writes | ✓ bundled SQLite roots, including D1 and libSQL | ✓ bundled PostgreSQL roots, including neon-http | Eligible singleton generated-ID nodes and `cardinality: "many"` edges use one authoritative create statement. Eligible node updates may carry fulltext/vector replacements; unconstrained non-durable-identity edge updates, direct edge deletes, and plain restricted node deletes use one authoritative read/gate plus one registered atomic mutation. Generated-, caller-, or mixed-ID node `bulkInsert`/`bulkCreate` batches compose supported multi-claim/cross-scope claim sets with projections in one schema-fenced native program; direct edge programs also maintain durable match identity and cardinality claims. Direct edge `bulkDelete` and plain restricted node `bulkDelete` use the same mutation profile. Eligible mixed `bulkUpsertById` sets, including node projections, use the profile on serverless roots and on exact bundled PostgreSQL transaction sessions; a generic derived backend still loses the evidence. A custom backend may opt in per family only after registering its exact transport and semantic executor. Unregistered or otherwise ineligible families, projected/identity-enabled node deletes, over-budget claimed members, cascade/disconnect deletes, and other managed writes retain the existing path |
| Typed constraint error above READ COMMITTED | n/a (no such isolation mode) | ✗ at `REPEATABLE READ` / `SERIALIZABLE` | PostgreSQL raises `40001` from the claim's upsert instead of resolving the conflict, so the loser retries a serialization failure rather than reading `UniquenessError` |
| Claim row lock released before end of transaction | ✗ | ✗ | Held to commit/rollback on both dialects, refusal included — a caller that catches a constraint error blocks other writers of that axis for the rest of its transaction |
| Recursive traversal (`capabilities.recursiveTraversal`) | ✓ | ✓ | Identical on both bundled backends. A third-party backend declaring `{ supported: false, reason }` refuses the five recursion-dependent operations with `ConfigurationError` code `RECURSIVE_TRAVERSAL_UNSUPPORTED`; `weightedShortestPath` degrades to a predecessor walk instead — see above. Unweighted `shortestPath` is unaffected — it never emits a recursive CTE |
| Write fence (`capabilities.writeFence`) | ✓ `engine-serialized` (single writer slot) | ✓ `lock` (advisory + table locks) | Identical guarantee, different mechanism. A custom backend that declares no `writeFence` resolves `unfenced` and is refused at construction for Operational Identity or TypeGraph-owned recorded-clock allocation |
| Capability bundles (`CAPABILITY_BUNDLES`) | Identical | Identical | Both bundled backends implement every pilot bundle's core/extra members on both dialects it scopes to. A third-party backend with a port gap refuses (gated core, or a `refuse`-disposition extra) or degrades (a `fallback`-disposition extra) per that bundle's own registry row |
| Engine-native lineage (`backend.lineage`) | ✗ (recorded-relations lineage via `history`) | ✗ (recorded-relations lineage via `history`) | Neither bundled profile declares its own `lineage`. `resolveLineage` derives it from the store's recorded relations whenever `history: true` is on, identically on both dialects, so a graph-merge diff against such a store is pruned the same way regardless of backend. Without `history`, no lineage source resolves and the diff is full; the anchor is the durable revision anchor when `revisionTracking: true`, otherwise the compatibility content fingerprint |
| Engine-native recorded time (`backend.recordedTime`) | ✗ (TypeGraph-owned recorded relations via `history`) | ✗ (TypeGraph-owned recorded relations via `history`) | Neither bundled profile declares `recordedTime`, so `resolveRecordedTimeOwnership` derives `"typegraph-relations"` for both — `history: true` captures into TypeGraph's own recorded relations and clock, identically on both dialects, and every recorded-time integration suite and the parity snapshot run unchanged. A backend that supplies `recordedTime` (and the co-required `lineage`) reads and writes recorded time through its own engine instead: no TypeGraph capture, clock, or recorded relations, `RecordedInstant` anchors in the `e1:` form, and several TypeGraph-relation-specific surfaces refused — see [Engine-native recorded time](/queries/temporal#engine-native-recorded-time) and [Supplying `recordedTime`](/backend-authoring#supplying-recordedtime). No bundled backend implements this today; the capability is proven by a PostgreSQL-family simulation (shared by the always-running `tests/backends/postgres/pglite-engine-native-recorded-time.test.ts` and the `POSTGRES_URL`-gated `tests/backends/postgres/engine-native-recorded-time.test.ts`, both built on `engine-native-recorded-time-simulation.ts`) that dresses TypeGraph's own recorded relations as a temporal-table expression, labeled as a simulation rather than a real third engine |
Identity support also has a **driver** dimension inside each dialect:
| Driver | Atomic identity support | Behavior |
| ------------------------------------------------------------ | ----------------------- | ------------------------------------------------------------------------------------------------------------------- |
| Managed SQLite, libSQL, Durable Objects | ✓ | Full profile |
| PostgreSQL `node-postgres`, `postgres-js`, neon-serverless, PGlite | ✓ | Full profile; identity-affecting writes serialize per graph, limiting each graph to one identity writer at a time |
| Cloudflare D1 | ✗ | Enabled graphs fail at store construction with `ConfigurationError` details code `IDENTITY_REQUIRES_ATOMIC_BACKEND` |
| `drizzle-orm/neon-http` | ✗ | Same fail-fast error; identity-disabled graphs retain the ordinary single-statement path |
### Filtered approximate search
Every approximate (ANN) vector search carries at least one row filter: the liveness predicate that hides
soft-deleted and out-of-validity rows. A `.where(...)` predicate narrows it further. Engines differ in where they
apply that filter relative to the index traversal, which decides whether a page can come back short. Read it from
`backend.capabilities.vector.filteredApproximateSearch`:
```typescript
const filtered = backend.capabilities.vector?.filteredApproximateSearch;
if (filtered?.guaranteesFullPage !== true) {
// An approximate search here may return fewer than `limit` rows.
}
```
**Check `guaranteesFullPage`, not `mode`.** `mode` names the mechanism the strategy asks the engine for; only
`guaranteesFullPage` tells you whether a short page is possible.
| `mode` | Strategy | `guaranteesFullPage` | Meaning |
| ------------------- | --------------- | -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `"filter-pushdown"` | `sqlite-vec` | `true` | The filter constrains the `vec0` KNN candidate set itself. `limit` matching rows come back whenever `limit` exist. |
| `"iterative-scan"` | `pgvector` | `false` | The index is re-entered for more candidates (`hnsw.iterative_scan` / `ivfflat.iterative_scan`, applied automatically on pgvector ≥ 0.8). Much better recall than a post-filter, but **not a guarantee**: the scan stops at `hnsw.max_scan_tuples` / `ivfflat.max_probes`. And on **pgvector < 0.8** there is no iterative scan at all — the backend detects that, warns once, and the search stays `ef_search`-bounded. |
| `"post-filter"` | `libsql-native` | `false` | DiskANN's `vector_top_k` is a table function with no filter pushdown and no way to re-enter the index. TypeGraph over-fetches `4 × (limit + offset)` neighbors and filters afterwards, so once more than that headroom is filtered out **the search silently returns fewer than `limit` rows while more matches exist**. |
Heavy tombstone drift — routine in a temporal store — is what turns a bounded search from a theoretical caveat into
a short page. When a full page matters, use an exact search (`approximate: false`), which scans and so applies the
filter to every row; or declare the field's index as `"none"` so it is always brute-forced.
Vector and fulltext capabilities are populated from the configured strategy, so the matrix above reflects the
bundled strategies (`sqlite-vec`/`libsql-native`/`pgvector`, `fts5`/`tsvector`). A custom strategy advertising
different `metrics`/`indexTypes`/`filteredApproximateSearch`/`searchFrontierTuning` shifts these rows accordingly —
always check `backend.capabilities` at runtime rather than hard-coding the dialect.
`searchFrontierTuning` is **required** on a vector strategy's capabilities, so a strategy must state whether it has a
per-search ANN frontier knob rather than inheriting silence. It is a discriminated union: `{ tunable: true, parameter,
indexType, requiresTransactionScope }` names the engine parameter `efSearch` maps to (`pgvector`: `hnsw.ef_search`, on
an `hnsw` slot, needing a transaction to scope it), while `{ tunable: false, reason }` names why the engine has no such
knob and is what makes `efSearch` a typed refusal there. A hand-written strategy that omits the field no longer
compiles.
Both bundled backends advertise `windowFunctions: true`. Relation `topPerPartition()` refuses execution
with `UnsupportedBackendCapabilityError` when a custom backend sets `windowFunctions: false`.
Vector, fulltext, and hybrid relevance-ranking
queries use `ROW_NUMBER()` internally and throw `ConfigurationError` before SQL generation if a custom backend profile
sets `windowFunctions: false` — there the window output *is* the result (the relevance k-cutoff / rank ordinal), so
there is no correct fallback.
`bulkFindByIndex({ limitPerInput })` also uses `ROW_NUMBER()` when available, but it does **not** throw on a
windowless profile: the per-input cap is a transfer optimization with identical row semantics either way, so it
degrades gracefully — fetching all matching ids and capping per group in application code. The unbounded
`bulkFindByIndex` path needs no window and is always available.
:::note[JSON is native on both backends]
SQLite stores JSON as text and queries it with the built-in JSON functions (`json_extract`, `json_each`, …);
PostgreSQL uses native `JSONB`. The dialect layer hides this difference, so JSON-path predicates and **B-tree
expression indexes on scalar JSON properties** (`defineNodeIndex` / `defineEdgeIndex`) are at full parity. The one
JSON-related difference is performance, not capability: PostgreSQL can use a single GIN index to accelerate
array/object **containment** predicates (`contains()` / `containsAll()` / `hasKey()` / `pathEquals()`), whereas on
SQLite those run as `json_each()` scans — correct results, just not index-accelerated. See
[Indexes](/performance/indexes) for the full breakdown.
:::
:::note[Transactions are driver-dependent, not backend-dependent]
Both backends report `execution.interactiveTransactions: true` by default. The exception is symmetric and lives in
specific drivers:
Cloudflare D1 (SQLite) and `drizzle-orm/neon-http` (Postgres) are non-transactional, so they downgrade to
`execution.interactiveTransactions: false`. Operations that require atomicity (`commitSchemaVersion`,
`setActiveVersion`, Operational Identity) throw on those drivers regardless of backend. A
schema-managed Store's write that cannot fuse its schema fence into its own statement fails closed
the same way, because it has no other way to hold the transaction-scoped fence; `store.transaction()`
refuses on those roots regardless. See
[The guard every fused write shares](/limitations#the-guard-every-fused-write-shares) for which
writes fuse. Eligible operations
with a certified atomic SQL program remain available independently of this interactive transaction capability.
:::
:::note[Aggregate set operations are a builder limitation, not a parity gap]
`GROUP BY`/`HAVING` leaves are supported by the set-operation compiler on **both** backends, but the query builder
does not expose `.union()`/`.intersect()`/`.except()` on `.aggregate()` queries. That limit applies equally to SQLite
and PostgreSQL, so it is not a portability difference.
:::
## Connection Management
Connection ownership follows the entrypoint:
- **Managed Store factories** (`/sqlite/local` and `/postgres/pglite`) own the
connection and provisioned resources. `await store.close()` releases them.
- **Owned local backend factories** (`createLocalSqliteBackend` and
`createLocalPgliteBackend`) also own their resources. A Store delegates
`close()` to its backend, so `await store.close()` releases them.
- **Bring-your-own adapter factories** (`createSqliteBackend`,
`createPostgresBackend`, and `createLibsqlBackend`) leave connection ownership
with the caller. Their Store's `close()` does not close the supplied client or
pool.
When you bring your own connection, you are responsible for:
1. **Creating connections** with appropriate configuration
2. **Connection pooling** for production use
3. **Closing connections** on shutdown
```typescript
// You create the connection
const sqlite = new Database("app.db");
const db = drizzle(sqlite);
const backend = createSqliteBackend(db);
const store = createStore(graph, backend);
// You close the connection
process.on("exit", () => {
sqlite.close();
});
```
Here `store.close()` leaves `sqlite` open because the application supplied the
connection. Close the driver or pool through its own API.
### Serialized connections
Some drivers run every statement through **one** connection. Two long-lived
interchange streams cannot share such a connection — an export snapshot holds a
read transaction for the whole stream while an import writes one per chunk — so
TypeGraph refuses the second one with a typed error instead of letting it hang
(see
[Interchange serialized-connection guard codes](/errors#interchange-serialized-connection-guard-codes)).
Recognizing a serialized connection means recognizing the *driver*, from the
shape of the client object. That is deliberately conservative: a driver
TypeGraph cannot positively identify is left unmarked, because refusing a pooled
connection would refuse work that succeeds.
| Driver / configuration | Detected | Notes |
| --------------------------------------------------------------------------------------------- | ------------------ | ---------------------------------------------------------------------------------------------------------------- |
| better-sqlite3, bun:sqlite, sql.js, local libSQL (`file:` / `:memory:`), Durable Object storage | ✓ automatic | One handle, one connection |
| PGlite | ✓ automatic | One in-process WASM connection |
| Bare `pg` / neon-serverless `Client`, a checked-out `PoolClient` | ✓ automatic | One owned socket |
| `pg` `Pool` capped at one (`{ max: 1 }`, `{ max: "1" }`, `{ poolSize: "1" }`) | ✓ automatic | pg-pool does not coerce the cap, so the string forms are the same one-connection pool |
| postgres-js capped at one (`{ max: 1 }`, `?max=1`, `PGMAX=1`) | ✓ automatic | Same reasoning on the postgres-js side |
| Default-size pools, `neon-http`, D1, RDS Data API, remote libSQL (`http` / `ws`) | — deliberately not | Each statement gets an independent connection; refusing would refuse work that succeeds |
| `expo-sqlite`, `op-sqlite`, `sqlite-proxy`, `pg-proxy`, a bespoke adapter | ✗ **declare it** | Serialized in fact, but the client exposes no shape TypeGraph can attribute to a known driver |
| Bun `SQL` (Postgres) at `{ max: 1 }` | ✗ **declare it** | The cap is readable, but nothing identifies the driver, and a cap on an unknown client is not evidence |
| postgres-js with a non-numeric string cap other than one, e.g. `?max=5` | ✗ **declare it** | Opens exactly one connection today only because postgres-js does not coerce the value — marking it would encode an upstream bug that will one day be fixed |
For the rows marked **declare it**, tell TypeGraph what it cannot see. The
option is on `createSqliteBackend` and `createPostgresBackend` — the two
factories that resolve it. The batteries-included wrappers
(`createLibsqlBackend`, `createLocalSqliteBackend`, `createLocalPgliteBackend`)
do not take it, because each already detects its own connection.
`{ mode: "shared", resource: pool }` is incorrect for a `pg.Pool` that can open
multiple connections, even if several backends use that pool. Each transaction
checks out its own connection; marking the pool as one resource makes independent
snapshot exports and imports contend for a single lease and refuses concurrent
operations that the pool can run. Leave the declaration absent for such a pool.
```typescript
const sql = postgres(process.env.DATABASE_URL + "?max=5");
const backend = createPostgresBackend(drizzle(sql), {
// This client really does run every statement on one connection.
serializedResource: { mode: "shared", resource: sql },
});
```
Two backends that name the **same** object are one serialized resource, exactly
as two wrappers over a detected client are. Naming a *different* object than the
one TypeGraph detected is refused with a `ConfigurationError`
(`details.reason: "serialized-resource-conflict"`) rather than silently
preferred: two wrappers over one connection given two different sentinels would
stop being seen as a pair, which is the failure the guard exists to prevent.
The refusal names each side by constructor (`details.declaredKind` /
`details.detectedKind`) instead of carrying the two handles, because `details`
is what `toLogString()` serializes and a driver handle there would log whatever
that driver stores — a `pg.Pool` keeps its `connectionString`.
The reverse declaration escapes a detection that is wrong for your topology:
```typescript
const backend = createSqliteBackend(db, {
serializedResource: { mode: "independent" },
});
```
**Scope.** `{ mode: "independent" }` lifts the *shared-resource* refusal between
two distinct backend objects. It does not lift the object-identity refusal, under
which one SQLite backend exporting into **itself** is refused with
`INTERCHANGE_SAME_SQLITE_BACKEND_SNAPSHOT`. That one is a fact about a single
handle holding a single open snapshot transaction, not a claim about connection
topology, so no declaration can make it false — pass a second backend instead.
That surviving refusal is SQLite-only, so on PostgreSQL the declaration lifts
the refusal for one backend exporting into itself as well: a client that hands
out independent connections — which is exactly what the declaration claims —
runs the snapshot and the writes it contends with on different ones.
## Database roles & least privilege
`createStoreWithSchema()` and `createStore()` divide cleanly along DDL
privilege, so a production deployment can run its application under a
least-privilege, DML-only database role.
- **`createStoreWithSchema(graph, backend)` runs DDL.** It bootstraps the
base tables on a fresh database, applies safe auto-migrations, and
adopts release-added deployment-wide base storage on pre-provisioned
databases, even when the persisted graph schema is unchanged. The first
adoption creates or repairs the graph-template and edge match-identity
storage, then stamps a version marker; a warm base-schema check is one
`SELECT` with no base-adoption DDL. It also
durably materializes strategy-owned runtime storage — both fulltext and
each `embedding()` field's per-`(kind, field)` vector table, plus a
durable marker for each. It also brings TypeGraph's own base-relation
**system indexes** up to the running library version: bootstrap DDL only
runs on the very first boot, so an index shipped in a newer version
reaches an already-initialized database through this step (built with
`CREATE INDEX CONCURRENTLY` on PostgreSQL; a database whose indexes all
exist settles from the catalog with no DDL). Contribution and system-index
preparation have their own catalog checks and may still issue DDL, so the
role it runs under **must hold `CREATE` / DDL privileges**. Run it once at startup, outside request
handlers and transactions. (`store.evolve()` likewise provisions any
embedding field it introduces, so it too needs DDL privileges.)
Deployments that never run `createStoreWithSchema` (manual-DDL boot with
a plain `createStore` attach) adopt new system indexes by calling
`store.materializeSystemIndexes()` once under a DDL-capable role after
upgrading; deployments that must not run index builds inline at boot
(large tables behind a readiness probe) pass `systemIndexes: "skip"` to
`createStoreWithSchema` and materialize out-of-band the same way.
- **`createStore(graph, backend)` is a synchronous, zero-I/O attach.**
It does not create tables, repair DDL, or record that runtime storage
is materialized — it issues **no DDL ever**. Use it only to attach to a
database a prior `createStoreWithSchema` boot already initialized. A
fulltext operation or an **embedding write** against a database that was
never initialized — a `create({ embedding })` or embedding update/delete
— throws `StoreNotInitializedError` rather than silently emitting
`CREATE TABLE` on the hot path. (Vector *reads* are not marker-gated:
`store.search.vector`, `store.search.hybrid`, and a query-builder
`.similarTo()` predicate compile to SQL against the per-field table
directly, so on an un-provisioned database they surface the engine's own
missing-relation error instead — `no such table: tg_vec_…` on SQLite,
`relation … does not exist` on Postgres. Same cause, same fix; use
`createVerifiedStore` to catch it at attach rather than at first query.)
This is what lets a least-privilege role run vector ops: the table
already exists. Graphs with no `searchable()` or `embedding()` fields
are unaffected.
The Store is also raw and unversioned: its writes do not participate in
the schema-version fence. Direct backend writes have the same semantics.
Quiesce those writers yourself before changing schemas.
- **`createVerifiedStore(graph, backend)` is the same zero-DDL attach
with a verification gate.** It reads the active schema row, folds the
persisted graph extension, and refuses to construct the Store unless
the database is at the same schema version as the code graph. Throws
`BaseSchemaMigrationError` when deployment-wide base storage is missing,
stale, or newer than the library, `MigrationError` on graph-schema drift
(safe or breaking), `ConfigurationError` when no graph schema has been
initialized, and `StoreNotInitializedError`
when the schema is current but runtime-contribution markers are
missing. The runtime-side counterpart of `createStoreWithSchema` for
least-privilege deployments. If you only need the gate without
building a Store (e.g. a readiness probe), call `assertSchemaCurrent`.
Its managed writes require a transactional backend with the schema-write
fence; non-transactional and unsupported custom backends can attach for
reads but fail closed on the first write.
The adapter equivalents (`createAdapterStoreWithSchema` and
`createVerifiedAdapterStore`) carry the same managed metadata. So does
`createAdapterStore(..., { reconciled })` with a cached reconciliation snapshot,
and Stores returned by `evolve()` or rebound from an already-managed Store.
Check `store.introspect().schemaVersion !== undefined` at runtime. Calling
`store.clear()` deletes the schema rows and resets that Store to raw semantics;
reopen it through a managed factory before resuming version-fenced writes.
- **`store.verifyContributions()` diagnoses contribution storage;
`store.repairContributions()` repairs safe findings under a privileged
role.**
Every gate above trusts the marker row without probing the catalog, so
a database whose strategy-owned tables were dropped out of band opens
clean and fails at the first dependent read or write. This method compares each contribution
currently expected by the active graph and backend strategies with its
marker and the catalog. It does not audit retired marker rows, and a
never-attempted contribution with neither marker nor table is omitted, so
an empty result is not initialization proof. It is read-only (`SELECT`
only, no DDL) so the least-privilege role can run it, and it is deliberately
not part of any open path. For a readiness check, construct the Store with
`createVerifiedStore()` first and then run this diagnostic; otherwise use
it as an operator check. The repair method re-audits current declarations,
preserves data while repairing `missing-marker` and
`failed-materialization`, and reports `stale` or `orphaned-marker` as
`requires-rebuild`. Run repair through the DDL-capable migration role, not
the least-privilege runtime role. Follow the per-state table in
[The store opens clean but a fulltext or vector read fails](/troubleshooting#the-store-opens-clean-but-a-fulltext-or-vector-read-fails)
rather than applying one repair to every entry.
- **`store.probeContributions()` is the read-only readiness check;
`store.rebuildContribution()` is the destructive last resort.**
The two bracket `repairContributions()` into one escalation ladder:
probe (writes nothing) → repair (non-destructive) → rebuild
(destructive, but scoped to the calling graph). The probe reports one
`ready` / `degraded` entry per
search projection and is safe on a read path, on a replica, and under
the least-privilege role — it shares the detection logic of the other
two rather than reimplementing it, so it cannot disagree with the gate
the hot path actually consults. The rebuild is the only repair for a
`stale` contribution, whose table exists at a shape the current
`createDdl` no longer produces; it deletes and refills only the calling
graph's rows in the shared fulltext table, escalating to drop → recreate
when that table holds no other graph's rows (under a database-scoped DDL
advisory lock, since that DDL is database-global), and runs the whole
sequence inside one transaction under the schema-write fence. It refuses
with `ContributionRebuildUnsupportedError` for vector storage, whose
embeddings exist only in the table it would drop
(`reason: "vector-source-unavailable"`), and for a `stale` shape whose
storage still holds other graphs' rows
(`reason: "shared-storage-in-use"`, naming them in
`details.otherGraphIds`). Run rebuilds through
the DDL-capable migration role, in a maintenance window: the
transaction is held for the whole refill, and on PostgreSQL a drop's
`ACCESS EXCLUSIVE` lock blocks both searches and writes to any kind with
`searchable()` fields until it commits. Reach it from a `createStore()`
Store — the managed factory's boot step refuses to open while a
contribution is `stale`. See
[Contribution health: probe, repair, rebuild](/troubleshooting#contribution-health-probe-repair-rebuild).
Strategy contributions declare an ownership `scope`: `"graph"` (the
default for older custom strategies) provisions one physical contribution
per graph, while `"deployment"` provisions shared physical storage once
under TypeGraph's reserved deployment marker and records a separate
graph-local activation marker. The built-in full-text strategies use
deployment scope; vector slots remain graph-scoped. A subsequent graph open
reads both attestations and performs no DDL, so it can run under a
DML-only role without disabling full-text search. The reserved marker key is
exported as `DEPLOYMENT_CONTRIBUTION_GRAPH_ID`; graph definitions must not
use that id.
### Contribution capability parity
`backend.capabilities.contributions` declares how far up the ladder a
backend goes. Each rung is separate because a backend can genuinely stop
at any of them, and a rung a backend cannot serve refuses with a typed
error rather than returning something that looks like success.
| Backend | `supported` | `probe` | `rebuild` |
| --- | --- | --- | --- |
| SQLite (better-sqlite3, bun:sqlite, libSQL, Durable Objects) | ✅ | ✅ | ✅ |
| SQLite with `transactionMode: "none"` | ✅ | ✅ | ❌ no schema fence |
| PostgreSQL (`pg`, `postgres-js`, PGlite, `neon-serverless`) | ✅ | ✅ | ✅ |
| PostgreSQL over `neon-http` | ✅ | ✅ | ❌ no schema fence |
| Custom fulltext strategy without `dropDdl` | ✅ | ✅ | ❌ no teardown DDL |
| Fulltext disabled (`fulltext: false`) | ✅ | ✅ | ✅ with a schema fence |
`rebuild` requires two things at once: a fulltext strategy that declares
`dropDdl` on its contribution, and a transactional schema fence
(`schemaWriteTransaction`) to run the sequence under. The HTTP-only
PostgreSQL drivers cannot hold a session across statements, so they have
no fence — the same reason they already report
`capabilities.execution.interactiveTransactions === false`. A third-party strategy predating
`dropDdl` keeps working for every other operation and is reported as not
rebuildable rather than being dropped through a synthesized statement
TypeGraph guessed at. Vector contributions are never rebuildable on any
backend; that is a property of what TypeGraph stores, not of the engine. A
backend built with `fulltext: false` has no fulltext contribution at all, so
the first condition is vacuously satisfied and `rebuild` reduces to whether
the backend has the transactional schema fence — the same value it would
report if fulltext were still active on a driver with that fence.
**`fulltext: false` stops creating and maintaining the fulltext table; it
never drops one.** On a database that already carries fulltext rows,
disabling fulltext leaves them in place and unmaintained: a hard delete
performed while fulltext is off leaves an orphaned row behind in the
fulltext table, because `hardDeleteNode`'s cascade has no active strategy to
build a delete statement from. Re-enabling fulltext later therefore requires
the destructive contribution rebuild — `store.rebuildContribution("fulltext")`,
which drops and recreates the fulltext table — **not**
`store.search.rebuildFulltext()`: that method pages live nodes to recompute
their content, and a hard-deleted node has no row left in the node table for
it to page, so it never revisits, and therefore never clears, the orphan.
### Recommended deployment shape
Run schema/DDL changes as a **privileged, one-time migration step**, then
run the application under a **least-privilege runtime role** that holds
only `SELECT` / `INSERT` / `UPDATE` / `DELETE`:
```typescript
// 1. Migration step — privileged role with DDL/CREATE.
//
// createStoreWithSchema is mandatory here: it bootstraps tables,
// applies safe auto-migrations, commits the schema_versions row,
// and writes the durable contribution markers. The runtime gate
// checks all of those.
const [/* store */] = await createStoreWithSchema(graph, adminBackend);
// Optional prerequisite if you manage DDL externally with
// drizzle-kit. Generated SQL creates the tables but does NOT
// initialize the schema row or contribution markers — still run
// createStoreWithSchema afterwards (it skips bootstrap when tables
// already exist and commits the row + markers):
//
// import { generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
// await adminPool.query(generatePostgresMigrationSQL());
// await createStoreWithSchema(graph, adminBackend);
```
```typescript
// 2. Runtime — least-privilege, DML-only role. Zero DDL.
// createVerifiedStore fails fast if the privileged migrator is behind.
const runtimePool = new Pool({ connectionString: process.env.APP_DATABASE_URL });
const backend = createPostgresBackend(drizzle(runtimePool));
const [store] = await createVerifiedStore(graph, backend);
```
If the runtime role has no DDL privileges and you boot it with
`createStoreWithSchema()` anyway, the first cold boot fails with a
permission error on the bootstrap or contribution-marker DDL — see
[Troubleshooting](/troubleshooting).
## Environment-Specific Setup
### Development
```typescript
// In-memory for fast tests
const { backend } = createLocalSqliteBackend();
// Or file-based for persistence during development
const { backend } = createLocalSqliteBackend({ path: "./dev.db" });
```
### Testing
```typescript
// Fresh in-memory database per test
beforeEach(() => {
const { backend } = createLocalSqliteBackend();
store = createStore(graph, backend);
});
```
### Production
Single-role setup — `createStoreWithSchema` bootstraps and migrates on
boot, so the role needs DDL privileges. To run the application under a
least-privilege, DML-only role instead, split the migration step out as
described in [Database roles & least privilege](#database-roles--least-privilege).
```typescript
// PostgreSQL with pooling
const pool = new Pool({
connectionString: process.env.DATABASE_URL,
max: 20,
ssl: { rejectUnauthorized: false }, // For managed databases
});
const db = drizzle(pool);
const backend = createPostgresBackend(db);
const [store] = await createStoreWithSchema(graph, backend);
```
## Next Steps
- [Schemas & Types](/core-concepts) - Define your graph schema
- [Semantic Search](/semantic-search) - Vector embeddings and similarity search
- [Limitations](/limitations) - Backend-specific constraints
# Query Builder Overview
> A fluent, type-safe API for querying your graph
TypeGraph provides a fluent, type-safe query builder for traversing and filtering your graph. This
page introduces the query categories and how they compose together.
## Query Categories
Every query builder method falls into one of these categories:
| Category | Purpose | Key Methods |
|----------|---------|-------------|
| [Source](/queries/source) | Entry point - where to start | `from()` |
| [Filter](/queries/filter) | Reduce the result set | `whereNode()`, `whereEdge()` |
| [Traverse](/queries/traverse) | Navigate relationships | `traverse()`, `optionalTraverse()`, `to()` |
| [Recursive](/queries/recursive) | Variable-length paths | `recursive()` |
| [Shape](/queries/shape) | Transform output structure | `select()`, `project()`, `map()`, `aggregate()` |
| [Expressions](/queries/expressions) | Typed database calculations | `expr`, `project()`, expression callbacks |
| [Aggregate](/queries/aggregate) | Summarize data | `groupBy()`, `count()`, `sum()`, `avg()` |
| [Order](/queries/order) | Control result ordering/size | `orderBy()`, `limit()`, `offset()` |
| [Temporal](/queries/temporal) | Time-based queries | `temporal()` |
| [Compose](/queries/compose) | Reusable query parts | `pipe()`, `createFragment()` |
| [Combine](/queries/combine) | Set operations | `union()`, `intersect()`, `except()` |
| [Execute](/queries/execute) | Run and retrieve | `execute()`, `first()`, `count()`, `exists()`, `paginate()`, `stream()`, `batch()` |
## Query Flow
A typical query follows this flow:
```text
Source → Filter → Traverse → Filter → Shape → Order → Execute
↑__________________|
(repeat as needed)
```
Each step is optional except Source and Execute. You can filter, traverse, and filter again as many
times as needed before shaping and executing.
## Basic Example
```typescript
const results = await store
.query()
.from("Person", "p") // Source
.whereNode("p", (p) => p.status.eq("active")) // Filter
.traverse("worksAt", "e") // Traverse
.to("Company", "c") // Traverse (target)
.whereNode("c", (c) => c.industry.eq("Tech")) // Filter
.select((ctx) => ({ // Shape
person: ctx.p.name,
company: ctx.c.name,
role: ctx.e.role,
}))
.orderBy("p", "name", "asc") // Order
.limit(50) // Order
.execute(); // Execute
```
## Type Safety
The query builder is fully typed. TypeScript infers result types based on your schema and selection:
```typescript
// TypeScript infers: Array<{ name: string; email: string | undefined }>
const results = await store
.query()
.from("Person", "p")
.select((ctx) => ({
name: ctx.p.name, // string (required in schema)
email: ctx.p.email, // string | undefined (optional in schema)
}))
.execute();
// Invalid property access is caught at compile time:
.select((ctx) => ({
invalid: ctx.p.nonexistent, // TypeScript error!
}))
```
For new database-side projections and calculations, use typed
[database expressions](/queries/expressions). `project()` compiles its callback to SQL, while
`map()` transforms decoded rows in JavaScript. Existing `select()` callbacks retain their
compatibility behavior.
## When to Use Queries vs Store API
**Use the query builder** when you need:
- Filtering based on node properties
- Traversing relationships between nodes
- Aggregating data across multiple nodes
- Complex predicates with AND/OR logic
**Use the [Store API](/schemas-stores#store-api)** for simple operations:
- Get a node by ID
- Create a new node
- Update a node's properties
- Delete a node
## Predicates Reference
Predicates are the building blocks for filtering. Each data type has its own set of predicates:
| Type | Documentation |
|------|--------------|
| String | [String Predicates](/queries/predicates/#string) |
| Number | [Number Predicates](/queries/predicates/#number) |
| Date | [Date Predicates](/queries/predicates/#date) |
| Array | [Array Predicates](/queries/predicates/#array) |
| Object | [Object Predicates](/queries/predicates/#object) |
| Embedding | [Embedding Predicates](/queries/predicates/#embedding) |
## Performance Tips
### Filter Early
Apply predicates as early as possible to reduce the working set:
```typescript
// Good: Filter at source
store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.active.eq(true))
.traverse("worksAt", "e")
.to("Company", "c");
// Less efficient: Filter after traversal
store
.query()
.from("Person", "p")
.traverse("worksAt", "e")
.to("Company", "c")
.whereNode("p", (p) => p.active.eq(true));
```
### Be Specific with Kinds
Unless you need subclass expansion, use exact kinds:
```typescript
// More efficient: Exact kind
.from("Podcast", "p")
// Less efficient: Includes all subclasses
.from("Media", "m", { includeSubClasses: true })
```
### Always Paginate Large Results
```typescript
const page = await store
.query()
.from("Event", "e")
.orderBy("e", "date", "desc")
.limit(100)
.select((ctx) => ctx.e)
.execute();
```
## Next Steps
Start with the fundamentals:
1. [Source](/queries/source) - Starting queries with `from()`
2. [Filter](/queries/filter) - Reducing results with predicates
3. [Traverse](/queries/traverse) - Navigating relationships
4. [Shape](/queries/shape) - Transforming output with `select()`
# Temporal
> Time-based queries with temporal()
TypeGraph tracks temporal validity for all nodes and edges. Use temporal queries to view the graph
at a point in time, audit changes, or access historical data.
## Temporal Modes
The `temporal()` method controls which versions of data are returned:
| Mode | Description |
|------|-------------|
| `"current"` | Only currently valid data (default behavior) |
| `"asOf"` | Data as it existed at a specific timestamp |
| `"includeEnded"` | All versions, including historical |
| `"includeTombstones"` | All versions, including soft-deleted |
## Current State (Default)
By default, queries return only currently valid, non-deleted data:
```typescript
// Returns only current, non-deleted nodes
const currentPeople = await store
.query()
.from("Person", "p")
.select((ctx) => ctx.p)
.execute();
```
This is equivalent to:
```typescript
.temporal("current")
```
## Point-in-Time Queries (asOf)
Query the graph as it existed at a specific moment:
```typescript
const yesterday = new Date(Date.now() - 24 * 60 * 60 * 1000).toISOString();
const pastState = await store
.query()
.from("Article", "a")
.temporal("asOf", yesterday)
.whereNode("a", (a) => a.id.eq(articleId))
.select((ctx) => ctx.a)
.execute();
```
This returns nodes and edges that were valid at the specified timestamp, even if they've since been updated or deleted.
### Use Cases for asOf
- **Auditing**: See what data looked like at a specific time
- **Debugging**: Reproduce issues by querying historical state
- **Compliance**: Generate point-in-time reports
- **Recovery**: Find old values before an erroneous update
```typescript
// What did the user's profile look like last week?
const lastWeek = new Date(Date.now() - 7 * 24 * 60 * 60 * 1000).toISOString();
const historicalProfile = await store
.query()
.from("User", "u")
.temporal("asOf", lastWeek)
.whereNode("u", (u) => u.id.eq(userId))
.select((ctx) => ctx.u)
.first();
```
## Shared-Coordinate Views (store.asOf)
`.temporal("asOf", T)` pins a single query. When several reads should share one
temporal coordinate, pin it once with `store.asOf(T)` and reuse the returned
**read-only view** — TypeGraph's as-of database value, in the style of Datomic
`(d/as-of db t)` and SQL:2011 `FOR SYSTEM_TIME AS OF`.
```typescript
const past = store.asOf("2024-01-01T00:00:00.000Z");
// Every read on `past` observes the graph as it was valid at that instant.
const alice = await past.nodes.Person.getById(aliceId);
const jobs = await past.edges.worksAt.findFrom(alice);
const peers = await past.reachable(aliceId, { edges: ["knows"] });
const team = await past.subgraph(aliceId, { edges: ["reportsTo"] });
const names = await past
.query()
.from("Person", "p")
.whereNode("p", (p) => p.department.eq("Engineering"))
.select((ctx) => ctx.p.name)
.execute();
```
The view pins the `nodes` / `edges` collections (`getById`, `getByIds`, `find`,
`count`, `findFrom`, `findTo`), `query()`, `subgraph()`, and the graph
algorithms (`reachable`, `canReach`, `shortestPath`, `neighbors`, `degree`).
For the other modes, use `store.view({ mode, asOf })`:
```typescript
// A view over every version, including soft-deleted ones.
const audit = store.view({ mode: "includeTombstones" });
const everyVersion = await audit.nodes.Document.find();
```
A view is **read-only**: writes stay on the live `store`, and a view collection
rejects `create` / `update` / `delete` with a `ConfigurationError`. `search` is
refused on a non-`"current"` view (the fulltext / vector index reflects current
state only). `asOf` must be a canonical UTC ISO-8601 timestamp
(`YYYY-MM-DDTHH:mm:ss.sssZ`).
See the [`store.asOf` / `store.view`
reference](/schemas-stores#temporal-views-storeasof-and-storeview) for the full
surface.
## Recorded Time (Bitemporal)
The modes above query **valid time** — *when a fact was true in the world*
(`validFrom` / `validTo`). Recorded time (also called **system time**) is the
second axis — *when a fact was recorded by TypeGraph*. With the built-in
captured relation, TypeGraph can run **bitemporal graph reads** for
TypeGraph-managed writes: you can ask "what did TypeGraph reconstruct as true,
as of a captured commit instant?" — including seeing values that were later
corrected.
Recorded-time capture is **opt-in** per store, because it writes a history row
for every committed TypeGraph collection change:
```typescript
const store = createStore(graph, backend, { history: true });
```
With `history: true`, every committed TypeGraph node/edge write is captured into
recorded-time relations (`typegraph_recorded_nodes` /
`typegraph_recorded_edges`) stamped with a per-graph monotonic commit instant.
Enable it on a **fresh graph**: there is no backfill, so an entity that already
exists is first recorded the next time it is written through TypeGraph. Capture
requires a transactional backend with statement execution (the built-in SQLite /
PostgreSQL backends).
Advanced hosts can bind an already-populated recorded relation for reads without
using TypeGraph's writer wrapper:
```typescript
import { createSqlSchema, recordedRelation } from "@nicia-ai/typegraph";
const recordedRead = recordedRelation({
schema: createSqlSchema({
recordedNodes: "audit_nodes",
recordedEdges: "audit_edges",
}),
});
const store = createStore(graph, backend, { recordedRead });
```
That option only supplies the read source for `asOfRecorded(T)` reconstruction.
It does not capture writes, advance TypeGraph's recorded clock, or make
`store.recordedNow()` available. If TypeGraph should own capture, use
`history: true`. `recordedRead` must be created by `recordedRelation({ schema })`
with a `createSqlSchema(...)` schema; the store validates those factory
descriptors at runtime and rejects combining them with `history: true`.
### Reading at a recorded instant
`store.asOfRecorded(T)` reconstructs the graph as TypeGraph recorded it at
instant `T`. `T` is a `RecordedInstant`: a branded, versioned string containing
both a per-graph logical revision and a physical wall-time high-water mark. It
originates from `store.recordedNow()` (below), or from
`asRecordedInstant(...)` when an anchor previously returned by TypeGraph has
round-tripped through untyped storage:
```typescript
import {
asRecordedInstant,
recordedInstantWallTime,
} from "@nicia-ai/typegraph";
const recorded = store.asOfRecorded(
asRecordedInstant("r1:0000000000000042:2024-06-01T12:00:00.000Z"),
);
const doc = await recorded.nodes.Document.getById(docId);
const cited = await recorded.edges.cites.getByIds(citationIds);
const reachable = await recorded.reachable(docId, { edges: ["cites"] });
```
Recorded collections also expose `scan()` for complete snapshot reconstruction.
Each call returns at most 1,000 entities in canonical `id` order; use the opaque
`nextCursor` to continue without retaining a separate identity inventory:
```typescript
const first = await recorded.nodes.Document.scan({ limit: 500 });
const second =
first.nextCursor === undefined ?
undefined
: await recorded.nodes.Document.scan({
limit: 500,
after: first.nextCursor,
});
const citations = await recorded.edges.cites.scan({ limit: 500 });
```
Scan cursors are forward-only and bound to the graph, entity kind, and both
temporal coordinates. Passing a cursor to another collection or recorded-time
view throws a `ValidationError` instead of silently skipping data. Iterate each
declared node and edge kind to reconstruct a complete historical graph snapshot.
A raw wall-clock string — `store.asOfRecorded(new Date().toISOString())` — does
**not** type-check, by design. Wall time does not identify which commit to read
when several commits share a millisecond. The anchor's logical revision provides
that order; its ISO component records a non-decreasing physical wall-time
high-water mark. To pin "as things stand right now" deterministically, use
`store.recordedNow()` (the recorded high-water mark), then guard the `undefined`
case before passing it to `store.asOfRecorded()`.
```typescript
await store.nodes.Document.update(docId, { title: "Revised" });
const checkpoint = await store.recordedNow(); // a stable anchor for this state
if (checkpoint === undefined) throw new Error("expected a recorded checkpoint");
console.log(recordedInstantWallTime(checkpoint)); // canonical UTC wall time
// ...later, however much the graph has changed:
const asOfCheckpoint = store.asOfRecorded(checkpoint);
```
`recordedNow()` is **graph-global**, not scoped to any one caller or write. It is
the single high-water mark for the whole graph, advanced by *every* committed
capture from *any* writer. So a change in `recordedNow()` across two reads means
"something committed to this graph in between" — **not** "the write I just made
landed." Do not use a `recordedNow()` advance as a per-writer "did my write
succeed?" signal: under any concurrent writer to the same graph it both misses
dropped writes (another writer moved the clock) and misfires on no-op writes. To
confirm a specific write committed, observe the write itself (e.g. its return
value, or run it inside `store.transaction(...)` and act on success), not the
global clock.
#### Logical revision and physical time
The canonical encoding is
`r1:<16-digit revision>:`. Revisions are strict and
monotonic within one graph. The physical component is sampled from the
application clock and clamped to the previous anchor only when that clock moves
backward. It may repeat, but never decreases. TypeGraph does not add one
millisecond per commit, so throughput cannot push recorded wall time beyond the
greatest wall time the graph has actually observed. After a backward clock
correction, the component remains at its prior high-water mark until wall time
catches up.
This non-decreasing physical component preserves cumulative diagonal replay for
default validity timestamps: a later recorded anchor cannot pin valid time
before an earlier commit's default `valid_from`.
The fixed-width revision prefix makes anchors lexicographically sortable within
a graph and gives each captured transaction a distinct addressable state. Use
`compareRecordedInstants(a, b)` rather than manually comparing strings, and
only compare anchors from the same graph. Recorded relations store the revision
as an integer, so their open interval ceiling is independent of the `r1` API
encoding and PostgreSQL range scans do not depend on text collation. Recorded
clocks remain per graph, and TypeGraph does not provide one cross-graph recorded
anchor.
Batch related writes in `store.transaction(...)`: one transaction allocates one
recorded instant. For event logs, align transactions with durable replay or
checkpoint boundaries, and cap transaction size separately so an initial sync
does not hold a write lock or capture buffer without bound.
Direct `store.asOfRecorded(T)` is **diagonal** bitemporal sugar: it uses the
anchor's logical revision for the recorded-time axis and its physical wall-time
component for the valid-time axis. To pin the two axes independently — *what was
valid at one instant, as TypeGraph captured it at another* — chain from a
valid-time view:
```typescript
// The state valid on Jan 1, as TypeGraph recorded it on Jun 1
// (e.g. after a correction was entered later).
const corrected = store
.asOf("2024-01-01T00:00:00.000Z")
.asOfRecorded(
asRecordedInstant("r1:0000000000000042:2024-06-01T12:00:00.000Z"),
);
const asKnownThen = await corrected.nodes.Invoice.getById(invoiceId);
```
Use `recordedInstantRevision(T)` for diagnostics and
`recordedInstantWallTime(T)` for display or logging. Do not split the versioned
anchor string manually.
`store.view({ mode }).asOfRecorded(T)` composes recorded time with any
valid-time mode — e.g. `includeTombstones` to reconstruct soft-deleted rows at a
recorded instant.
### The recorded view surface
A `RecordedStoreView` is a **narrow, reconstructing** read lens. It exposes only
reads that can be faithfully rebuilt from the recorded relations:
- **Point reads** — `nodes..getById` / `getByIds`, and the edge equivalents
- **`query()`** — a sealed query builder over the recorded relations
- **`subgraph()`** and the graph algorithms — `reachable`, `canReach`,
`shortestPath`, `degree`
Broad collection reads (`find` / `count` / `findFrom` / …), `search`, and
fulltext / vector predicates are **refused** with a `ConfigurationError` /
`UnsupportedPredicateError`: the fulltext and vector indexes reflect *current*
state only, so they cannot answer a recorded-time question. `T` must use the
canonical versioned RecordedInstant encoding; a plain ISO timestamp is rejected.
:::caution[Preview-schema migration]
Timestamp-only anchors and recorded tables created by the initial preview need
an explicit offline migration. Run `migrateLegacyRecordedTime({ backend })`
before opening the upgraded store, then translate externally persisted
checkpoints with `migrateRecordedAnchor({ backend, graphId, anchor })`. See
[Migrating preview recorded time](/schema-management#migrating-preview-recorded-time).
:::
### Engine-native recorded time
Everything above describes **TypeGraph-owned** recorded time: `history: true`
captures into TypeGraph's own recorded relations and clock. A backend can
instead track recorded time itself — declare `EngineProvisioning.recordedTime`
on it — and `history: true` then reads and writes through the engine's own
temporal storage; TypeGraph's capture relations, clock, and write-fence-gated
clock allocation are never engaged.
Which ownership a store reads under is **derived**, never an option you set:
it is `"engine-native"` exactly when the backend declares `recordedTime`,
`"typegraph-relations"` otherwise. Neither bundled SQLite nor PostgreSQL
profile declares it, so every example on this page runs under
`"typegraph-relations"` as shown; see [Supplying
`recordedTime`](/backend-authoring#supplying-recordedtime) for what a
third-party engine implements to opt in.
Under engine-native ownership:
- `store.recordedNow()`, `store.revisionNow()`, and `TransactionReceipt.recorded`
all come from the engine's own revision instead of TypeGraph's clock — one
call per transaction, not per graph. `TransactionReceipt.recorded` is
stamped only when a graph node/edge/identity write inside the transaction
actually changed a row — a delete of a missing id, an
`insertNodeIfAbsent` that found the row, and a coalesced no-op upsert all
leave it `undefined`, matching a read-only transaction. A transaction whose
only effect is a raw `tx.sql` statement also leaves it `undefined` even
though the engine's revision advances underneath it; use a graph collection
write when you need `receipt.recorded` to reflect the change. To observe
those writes, a receipted engine-native transaction routes every write
through an observing wrapper, so `transactionWithReceipt` does not use
session-scoped atomic batching where a plain `transaction` on the same
store would.
- `RecordedInstant` anchors use the engine form
`e1::` rather than
`r1:<16-digit revision>:`. The revision is an opaque,
engine-assigned token, never parsed as a number, so ordering two `e1:`
anchors (`compareRecordedInstants`) falls back to the timestamp component
only — two engine revisions minted within the same millisecond compare
equal even though they are distinct commits, unlike a `r1:` anchor's strict
per-commit counter. `recordedInstantWallTime(instant)` works for either
form; `recordedInstantRevision(instant)` throws for an `e1:` anchor, since
there is no TypeGraph numeric revision to return.
- `store.asOfRecorded(instant)` requires an instant minted under the SAME
store's own ownership form. An engine-native store refuses an `r1:`
instant, and a TypeGraph-owned store refuses an `e1:` instant, both with a
`ConfigurationError` (`RECORDED_INSTANT_OWNERSHIP_MISMATCH`) — an anchor
from one ownership form is never valid against the other, even against a
different store over the same data.
- No recorded relation is read or written. (A profile built on the bundled
schema factories still creates the recorded tables as part of its base DDL
— they just stay empty.) `recordedRead: recordedRelation({ schema })`
(above) and `migrateLegacyRecordedTime` are both refused: neither has a
TypeGraph-owned recorded relation to bind or migrate.
- `revisionTracking: true` is refused whether or not `history: true` is also
requested — there is no TypeGraph clock for it to advance; the engine's
own revision is the only tracking engine-native has, and it is available
only under `history: true`.
- Reconstructing identity at a recorded coordinate — `store.identityAtCoordinate`
at a past instant, and any query that reaches the historical identity
traversal — is refused: identity history reads TypeGraph's own recorded
relations directly, which an engine-native backend does not populate. Read
identity at the current coordinate instead, or use a TypeGraph-owned store
for historical identity reconstruction.
Everything else on this page — `asOfRecorded`'s diagonal composition with
`asOf`, the recorded view surface's read shape, `includeTombstones`
composition — behaves the same under either ownership form; only the anchor
grammar, the write mechanics, and the refusals above differ. See [Lineage and
pruned diffs](/graph-merge#lineage-and-pruned-diffs) for how graph-merge
derives a change delta under engine-native ownership — from the engine's own
`lineage`, never from recorded relations, since none exist to derive one
from.
### Writing with history enabled
Capture flushes at transaction commit, so writes must go through the store's
typed collections — use `store.transaction(...)` as usual:
```typescript
await store.transaction(async (tx) => {
await tx.nodes.Document.create({ title: "Draft" });
});
```
#### Raw SQL under history capture
The portable `HistoryStore` exposes neither raw SQL nor caller-owned
transaction adoption. If the store was deliberately created through
`createAdapterStore(..., { history: true })`, raw `tx.sql` is still disabled
(it would bypass capture), and `store.withTransaction(externalTx)` is replaced
by the callback form
`store.withRecordedTransaction(externalTx, async (tx) => { ... })`, which gives
capture a flush point before your transaction commits. Out-of-band database
writes and row-returning raw SQL paths are not audited by the built-in capture
wrapper; use TypeGraph collection writes when the recorded relation is the
source of truth.
The adapter history store's `.backend` is a runtime and type-level
`HistoryStoreBackend` projection. Capture-wrapped graph reads and writes remain
available. `executeRaw`, `executeStatement`, `executeDdl`, `trustedImport`,
`clearGraph`, and nested `transaction` are absent because each can mutate live
rows without a corresponding capture flush. The full guarded backend remains
internal to TypeGraph's query and transaction implementation.
`store.withTransaction` on a history-enabled store is a **compile error** (the
`externalTx` argument is rejected with a message naming
`withRecordedTransaction`); the runtime guard still throws `ConfigurationError`
if suppressed. Inside an `AdapterHistoryStore.transaction(...)`, the typed
context omits `tx.sql`, and `tx.sqlAvailability` reports `"history"` (or
`"revisionTracking"`) so portable code can branch without touching the runtime
guard. Suppressed JavaScript or TypeScript access still throws — see the
`tx.sqlAvailability` guidance in
[Cross-Store Transactions](/recipes/).
Both guards carry a branchable `details.code`; see
[Recorded-capture guard codes](/errors/#recorded-capture-guard-codes).
To write your own relational tables atomically with graph writes on a history
store, pass your transaction handle to `withRecordedTransaction` and write your
tables through **that** handle (not `tx.sql`):
```typescript
await db.transaction(async (pgTx) => {
const { receipt } = await store.withRecordedTransaction(pgTx, async (tx) => {
await tx.nodes.Document.update(documentId, props); // graph write
});
await pgTx.insert(streamCursors).values(cursorRow); // your own table
}); // one COMMIT / ROLLBACK across both layers
```
`withRecordedTransaction` returns a
[`TransactionOutcome`](/schemas-stores/#transaction-receipts): destructure
`{ result, receipt }`. `receipt.writes` counts the graph writes (drop
detection) and `receipt.recorded` is this transaction's recorded commit instant
— the per-transaction replay anchor. When the callback runs user code that also
bookkeeps, scope a sub-receipt with `tx.measure((scoped) => ...)`: writes through
the `scoped` context are attributed to the sub-receipt, while the surrounding
bookkeeping written through `tx` is not.
This is separate from `recordedRead`: a store created with a `recordedRead`
binding can reconstruct from a relation populated by another system, but
TypeGraph is not responsible for making that relation complete or atomic with
live writes.
#### Write cost: batch under `history: true`
Each **un-batched** write under `history: true` becomes its own transaction —
it allocates a recorded commit instant under a per-graph clock lock and flushes
one history row at commit. So a tight loop of single `create`/`update`/`delete`
calls pays that fixed cost once per call. Wrapping the same writes in one
`store.transaction(...)` allocates **one** recorded instant for the whole batch
and amortizes the overhead to roughly nothing.
Measured per-op latency, identical workload with capture off vs on (history
off → on; N = 400; reproduce with
`pnpm --filter @nicia-ai/typegraph-benchmarks bench:recorded-write`):
| Workload | SQLite | PostgreSQL |
| ------------------------------ | -----: | ---------: |
| create — un-batched (per op) | ~2.5× | ~5.5× |
| create — **batched in one txn** | ~1.5× | ~1.0× |
| update — un-batched (per op) | ~2.8× | ~6× |
| soft delete — un-batched | ~1.7× | ~1.9× |
The takeaway: capture is opt-in and cheap when you batch. Under `history: true`,
prefer `store.transaction(...)` for bulk writes; a loop of individual
collection writes is the one pattern that pays the per-write multiple. (Stores
created without `history: true` are unaffected — graph writes never touch the
capture path.) Batching also reduces recorded-clock consumption: one captured
transaction advances the per-graph clock once, even when it contains many
writes. See [Logical revision and physical time](#logical-revision-and-physical-time)
for the anchor format.
> **Performance.** Recorded reads reconstruct from the history relations rather
> than the live tables, so they are slower than current-state reads — most
> noticeably for full-graph `subgraph` / algorithm reconstructions on
> PostgreSQL. Reach for `asOfRecorded` for audit and point-in-time
> reconstruction, not hot-path reads.
## Including Historical Data (includeEnded)
View all versions, including superseded records:
```typescript
const history = await store
.query()
.from("Article", "a")
.temporal("includeEnded")
.whereNode("a", (a) => a.id.eq(articleId))
.orderBy((ctx) => ctx.a.validFrom, "desc")
.select((ctx) => ({
title: ctx.a.title,
validFrom: ctx.a.validFrom,
validTo: ctx.a.validTo,
version: ctx.a.version,
}))
.execute();
// Result shows all versions:
// [
// { title: "Final Title", validFrom: "2024-03-01", validTo: undefined, version: 3 },
// { title: "Draft v2", validFrom: "2024-02-15", validTo: "2024-03-01", version: 2 },
// { title: "Initial Draft", validFrom: "2024-02-01", validTo: "2024-02-15", version: 1 },
// ]
```
### Audit Trail
Build a complete change history:
```typescript
async function getAuditTrail(nodeId: string) {
return store
.query()
.from("Document", "d")
.temporal("includeEnded")
.whereNode("d", (d) => d.id.eq(nodeId))
.select((ctx) => ({
version: ctx.d.version,
title: ctx.d.title,
status: ctx.d.status,
validFrom: ctx.d.validFrom,
validTo: ctx.d.validTo,
updatedAt: ctx.d.updatedAt,
}))
.orderBy("d", "version", "asc")
.execute();
}
```
## Including Soft-Deleted Data (includeTombstones)
Include records that have been soft-deleted:
```typescript
const allIncludingDeleted = await store
.query()
.from("User", "u")
.temporal("includeTombstones")
.select((ctx) => ({
id: ctx.u.id,
name: ctx.u.name,
deletedAt: ctx.u.deletedAt, // Will have a value for deleted records
}))
.execute();
```
### Filtering Deleted Records
```typescript
// Find only deleted records
const deletedUsers = await store
.query()
.from("User", "u")
.temporal("includeTombstones")
.whereNode("u", (u) => u.deletedAt.isNotNull())
.select((ctx) => ({
id: ctx.u.id,
name: ctx.u.name,
deletedAt: ctx.u.deletedAt,
}))
.execute();
```
## Temporal Metadata Fields
When querying with temporal context, these fields are available:
| Field | Type | Description |
|-------|------|-------------|
| `validFrom` | `string \| undefined` | When this version became valid (`undefined` on an **open-left** row — see below) |
| `validTo` | `string \| undefined` | When this version was superseded (undefined if current) |
| `createdAt` | `string` | When the node was first created |
| `updatedAt` | `string` | When this version was written |
| `deletedAt` | `string \| undefined` | Soft-delete timestamp (undefined if not deleted) |
| `version` | `number` | Optimistic concurrency version number |
### Open-left rows (`validFrom` is `undefined`)
A row may have **no lower bound at all**, which means "valid since forever, as
far as this store knows". `asOf` and `current` treat such a row as valid at
every instant strictly before its `validTo`, or every instant if it has no end.
These writes produce one:
- a Store create or resurrecting upsert stating `validFrom: null`;
- an interchange record stating `validFrom: null` — a source row confirmed to
have no lower bound, round-tripped rather than re-stamped;
- a **born-already-ended** write: one that CREATES a row, or RESETS its window,
while stating a `validTo` at or before its own instant and no `validFrom`. The
row's start is unknown rather than after its end, so no bound is stored and the
row reads back at every `asOf` before that end. A `validTo` in the *future* is
unaffected — it still stamps the write instant, so the row stays invisible at
instants before it existed. Every **node** path that resets the window
qualifies, and reaches the same stored shape: a create on a fresh id, a create
on a tombstoned one, and a resurrecting `upsertById` / `bulkUpsertById`. An
**edge** never does: an edge create cannot land on a tombstone (a taken id
raises `Edge already exists`), and the two paths that resurrect one —
`bulkUpsertById` and `getOrCreateByEndpoints` — RETAIN the bound the row
carries and judge the stated `validTo` against it.
#### Rows written by older versions
Before that rule existed, a born-already-ended write stored the write instant as
`valid_from`, leaving a window that runs backwards — a row readable at **no**
coordinate at all. Upgrading does not rewrite such rows; they keep their window
and stay invisible until an operator repairs them explicitly with
`repairInvertedValidityWindows`, which normalizes them to the open-left shape
above. Prefer `relations: "live-and-recorded"`: repairing only the live axis
leaves the recorded twin inverted, so `asOfRecorded` reads keep returning the
invisible shape. See
[Repairing inverted validity windows](/schema-management#repairing-inverted-validity-windows)
for the operator checklist — run it with writers stopped, and re-baseline merge
branches afterwards.
```typescript
.select((ctx) => ({
...ctx.a, // All node properties
validFrom: ctx.a.validFrom,
validTo: ctx.a.validTo,
createdAt: ctx.a.createdAt,
updatedAt: ctx.a.updatedAt,
deletedAt: ctx.a.deletedAt,
version: ctx.a.version,
}))
```
## Temporal Traversals
Temporal modes apply to traversals as well:
```typescript
// See who worked at a company last year
const lastYear = new Date("2023-01-01").toISOString();
const pastEmployees = await store
.query()
.from("Company", "c")
.temporal("asOf", lastYear)
.whereNode("c", (c) => c.name.eq("Acme Corp"))
.traverse("worksAt", "e", { direction: "in" })
.to("Person", "p")
.select((ctx) => ({
name: ctx.p.name,
role: ctx.e.role,
}))
.execute();
```
`store.subgraph()` and `store.algorithms.*` accept the same `temporalMode`
and `asOf` options, defaulting to `graph.defaults.temporalMode`. See
[Temporal Behavior](/graph-algorithms#temporal-behavior) for the algorithm
surface and [`store.subgraph()` options](/schemas-stores#storesubgraphrootid-options)
for subgraph.
## Real-World Examples
### Version Comparison
Compare two versions of a document:
```typescript
async function compareVersions(docId: string, v1: number, v2: number) {
const versions = await store
.query()
.from("Document", "d")
.temporal("includeEnded")
.whereNode("d", (d) => d.id.eq(docId))
.select((ctx) => ctx.d)
.execute();
const version1 = versions.find((v) => v.version === v1);
const version2 = versions.find((v) => v.version === v2);
return { version1, version2 };
}
```
### Compliance Reporting
Generate a report as of a specific date:
```typescript
async function generateQuarterlyReport(quarterEnd: string) {
const activeContracts = await store
.query()
.from("Contract", "c")
.temporal("asOf", quarterEnd)
.whereNode("c", (c) => c.status.eq("active"))
.traverse("belongsTo", "e")
.to("Customer", "cust")
.select((ctx) => ({
contractId: ctx.c.id,
value: ctx.c.value,
customer: ctx.cust.name,
}))
.execute();
return {
asOf: quarterEnd,
totalContracts: activeContracts.length,
totalValue: activeContracts.reduce((sum, c) => sum + c.value, 0),
contracts: activeContracts,
};
}
```
### Undo/Recovery
Find the previous value before an update:
```typescript
async function getPreviousVersion(nodeId: string) {
const versions = await store
.query()
.from("Document", "d")
.temporal("includeEnded")
.whereNode("d", (d) => d.id.eq(nodeId))
.select((ctx) => ctx.d)
.orderBy("d", "version", "desc")
.limit(2)
.execute();
return {
current: versions[0],
previous: versions[1],
};
}
```
## Next Steps
- [Filter](/queries/filter) - Filtering with predicates
- [Traverse](/queries/traverse) - Graph traversals
- [Execute](/queries/execute) - Running queries
- [Bitemporal Time Travel](/examples/bitemporal-time-travel) - Valid time plus
recorded time in one runnable example
- [Agent Decision Replay](/examples/agent-decision-replay) - Reconstruct the
exact graph an agent saw
- [Breach Forensics](/examples/breach-forensics) - Traverse a reconstructed
access graph at the breach instant
# Troubleshooting
> Solutions to common issues and frequently asked questions
This guide covers common issues and their solutions when working with TypeGraph.
## Installation Issues
### "Cannot find module '@nicia-ai/typegraph'"
**Cause:** Package not installed or using wrong package name.
**Solution:**
```bash
npm install @nicia-ai/typegraph zod drizzle-orm
```
### "better-sqlite3 compilation failed"
**Cause:** Native module compilation requires build tools.
**Solutions:**
**macOS:**
```bash
xcode-select --install
```
**Ubuntu/Debian:**
```bash
sudo apt-get install build-essential python3
```
**Windows:**
```bash
npm install --global windows-build-tools
```
**Alternative:** Use `sql.js` for pure JavaScript SQLite (no compilation needed).
### Missing optional `drizzle-orm` peer
**Cause:** The managed SQLite or PGlite Store entrypoint was called without the optional
`drizzle-orm` peer installed.
**Solution:** Install the peer in the application that uses the managed entrypoint:
```bash
npm install drizzle-orm
```
The root package and other portable entrypoints do not require Drizzle. Explicit
`@nicia-ai/typegraph/adapters/drizzle/...` entrypoints load Drizzle when the module is evaluated,
so a missing peer there appears as the runtime's raw module-resolution error instead of
`MISSING_PEER_DEPENDENCY`. See [Managed Store Entrypoints](/backend-setup#managed-store-entrypoints).
### "Module not found: drizzle-orm/better-sqlite3"
**Cause:** An explicit Drizzle adapter import is missing `drizzle-orm`, or the application imported
the wrong Drizzle subpath.
**Solution:** First install `drizzle-orm`, then ensure the import matches the driver:
```bash
npm install drizzle-orm
```
```typescript
// Correct
import { drizzle } from "drizzle-orm/better-sqlite3";
// Incorrect
import { drizzle } from "drizzle-orm";
```
## Schema Definition Errors
### "Node schema contains reserved property names"
**Cause:** Using reserved keys (`id`, `kind`, `meta`) in your Zod schema.
**Solution:** Rename your properties:
```typescript
// Bad - 'id' is reserved
const User = defineNode("User", {
schema: z.object({
id: z.string(), // Error!
name: z.string(),
}),
});
// Good - use a different name
const User = defineNode("User", {
schema: z.object({
externalId: z.string(),
name: z.string(),
}),
});
```
TypeGraph automatically provides `id`, `kind`, and `meta` on all nodes.
### "Edge type already has constraints defined"
**Cause:** Defining `from`/`to` constraints on both the edge type and graph registration.
**Solution:** Define constraints in one place only:
```typescript
// Option 1: On the edge type (reusable across graphs)
const worksAt = defineEdge("worksAt", {
from: [Person],
to: [Company],
});
const graph = defineGraph({
edges: {
worksAt: { type: worksAt }, // No from/to here
},
});
// Option 2: On the graph (flexible per-graph)
const worksAt = defineEdge("worksAt");
const graph = defineGraph({
edges: {
worksAt: { type: worksAt, from: [Person], to: [Company] },
},
});
```
## Runtime Errors
### ValidationError: "Invalid input"
**Cause:** Data doesn't match the Zod schema.
**Solution:** Check the error details for specific issues:
```typescript
try {
await store.nodes.Person.create({ name: "" });
} catch (error) {
if (error instanceof ValidationError) {
console.log(error.details.issues); // Zod issues array
}
}
```
### NodeNotFoundError
**Cause:** Attempting to read/update/delete a non-existent node.
**Solution:** Check if the node exists first or handle the error:
```typescript
const node = await store.nodes.Person.getById(someId);
if (!node) {
// Handle missing node
}
// Or use error handling
try {
await store.nodes.Person.update(someId, { name: "New" });
} catch (error) {
if (error instanceof NodeNotFoundError) {
console.log(`Node ${error.details.id} not found`);
}
}
```
### RestrictedDeleteError
**Cause:** Attempting to delete a node that has edges, with `onDelete: "restrict"` (the default).
**Solution:** Either delete the edges first or use a different delete behavior:
```typescript
// Option 1: Delete edges first. Include ended-but-not-deleted edges if you
// are cleaning up historical validity windows too.
const edges = await store.edges.worksAt.findFrom(person, {
temporalMode: "includeEnded",
});
for (const edge of edges) {
await store.edges.worksAt.delete(edge.id);
}
await store.nodes.Person.delete(person.id);
// Option 2: Use cascade delete in schema
const graph = defineGraph({
nodes: {
Person: { type: Person, onDelete: "cascade" },
},
});
```
### DisjointError
**Cause:** Creating a node with an ID that's already used by a disjoint type.
**Solution:** Ensure IDs are unique across disjoint types or don't use explicit IDs:
```typescript
// If Person and Organization are disjoint:
// Bad - same ID for different types
await store.nodes.Person.create({ name: "Alice" }, { id: "entity-1" });
await store.nodes.Organization.create({ name: "Acme" }, { id: "entity-1" }); // Error!
// Good - let TypeGraph generate unique IDs
await store.nodes.Person.create({ name: "Alice" });
await store.nodes.Organization.create({ name: "Acme" });
```
## Query Issues
### "Alias 'x' is already in use"
**Cause:** Using the same alias twice in a query.
**Solution:** Use unique aliases:
```typescript
// Bad
store.query().from("Person", "p").traverse("knows", "e").to("Person", "p"); // Error! 'p' already used
// Good
store.query().from("Person", "p1").traverse("knows", "e").to("Person", "p2");
```
### Empty results when expecting data
**Causes and solutions:**
1. **Type mismatch:** Ensure you're querying the correct node type
```typescript
// Check the node type name matches exactly
.from("Person", "p") // Must match defineNode("Person", ...)
```
2. **Missing includeSubClasses:** When querying a superclass
```typescript
.from("Content", "c", { includeSubClasses: true })
```
3. **Strict predicate:** Check your filters aren't too restrictive
```typescript
// Debug by removing filters temporarily
const all = await store
.query()
.from("Person", "p")
.select((c) => c.p)
.execute();
console.log(all.length); // How many total?
```
### Slow queries
**Solutions:**
1. **Use the query profiler:**
```typescript
import { QueryProfiler } from "@nicia-ai/typegraph/profiler";
const profiler = new QueryProfiler();
profiler.attachToStore(store);
// Run your queries...
const report = profiler.getReport();
console.log(report.recommendations);
```
2. **Add indexes** based on profiler recommendations:
```typescript
import { defineNodeIndex } from "@nicia-ai/typegraph/indexes";
const nameIndex = defineNodeIndex(Person, { fields: ["name"] });
```
3. **Limit results:**
```typescript
.limit(100)
// Or use pagination
.paginate({ first: 20 })
```
## Database Connection Issues
### "Database is locked" (SQLite)
**Cause:** Multiple processes accessing the same SQLite file without WAL mode.
**Solution:** Enable WAL mode:
```typescript
const sqlite = new Database("myapp.db");
sqlite.pragma("journal_mode = WAL");
```
### Connection pool exhausted (PostgreSQL)
**Cause:** Too many concurrent connections.
**Solution:** Configure pool limits:
```typescript
import { Pool } from "pg";
const pool = new Pool({
connectionString: process.env.DATABASE_URL,
max: 20, // Adjust based on your needs
idleTimeoutMillis: 30000,
});
```
### "relation 'typegraph_nodes' does not exist"
**Cause:** Migration not run.
**Solution:** Run the migration SQL:
```typescript
// PostgreSQL
import { generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
await pool.query(generatePostgresMigrationSQL());
// SQLite
import { generateSqliteMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
sqlite.exec(generateSqliteMigrationSQL());
```
### "permission denied" / cannot create relation on boot
**Cause:** `createStoreWithSchema()` is a privileged entry point. Its warm
base-schema check is one marker read with no base-adoption DDL, but bootstrap,
pending base adoption, graph migrations, contribution preparation, or system
index materialization can issue DDL. A DML-only role cannot safely own it.
**Solution:** Run schema/DDL changes as a privileged one-time migration
step, then attach at runtime with the zero-DDL
`createVerifiedStore()` (or `createStore()`) under the least-privilege
role. See
[Database roles & least privilege](/backend-setup#database-roles--least-privilege).
### `BaseSchemaMigrationError` from a zero-DDL runtime path
**Cause:** `createVerifiedStore`, `assertSchemaCurrent`, or graph-template
registration/instantiation found a missing, stale, or newer deployment-wide
base-schema marker. These paths deliberately do not repair physical storage.
**Solution:** For a missing or stale marker, run
`createStoreWithSchema(graph, adminBackend)` once under a DDL-capable role, or
apply the published external base-schema migration and stamp its marker last.
For a newer marker, deploy a TypeGraph release that supports that version.
The error details include `installedVersion`, `requiredVersion`, and `reason`.
### `MigrationError` from `createVerifiedStore` / `assertSchemaCurrent`
**Cause:** The runtime is using a code graph whose schema is ahead of
the database. The least-privilege runtime cannot migrate — by design,
it fails fast so requests don't run against a stale schema.
**Solution:** Run `createStoreWithSchema(graph, adminBackend)` under
the privileged role before promoting the new runtime build (apply any
generated migration SQL first if you manage DDL externally), then
restart the runtime. The thrown `MigrationError.message` includes the
diff summary and migration actions to apply.
### `ConfigurationError`: "no schema has been initialized"
**Cause:** A verifying attach (`createVerifiedStore` /
`assertSchemaCurrent`) ran before any privileged
`createStoreWithSchema()` boot — the database has no `schema_versions`
row (or no typegraph tables at all). The runtime deliberately refuses
to bootstrap under a least-privilege role. **Note:** running only the
generated migration SQL is not sufficient — it creates the tables but
does not write the schema row or contribution markers.
**Solution:** Run `createStoreWithSchema(graph, adminBackend)` once
under the privileged role. If you manage DDL externally with
drizzle-kit / `generatePostgresMigrationSQL()` /
`generateSqliteMigrationSQL()`, apply that first, then still run
`createStoreWithSchema()` to commit the schema row and contribution
markers. See
[Database roles & least privilege](/backend-setup#database-roles--least-privilege).
### `StoreNotInitializedError` on the first operation
**Cause:** The store was created with `createStore()` (a zero-I/O attach
that never materializes runtime storage) against a database that no
`createStoreWithSchema()` boot has initialized — commonly the runtime
started before the privileged migration step ran, or the wrong role/
database is configured. This covers fulltext operations and **embedding
writes**: a `store.nodes.*.create({ embedding })` (or embedding
update/delete) against an un-provisioned per-`(kind, field)` table throws
here rather than lazily issuing `CREATE TABLE` on the hot path. Vector
*reads* — `store.search.vector`, `store.search.hybrid`, and a
query-builder `.similarTo()` predicate — compile straight to SQL against
the per-field table, so they surface the engine's own missing-relation
error instead (`no such table: tg_vec_…` on SQLite, `relation … does not
exist` on Postgres) — same cause, same solution.
`createVerifiedStore()` catches every one of these cases at boot rather
than at the first hot-path operation.
A **`stale`** variant of this error on a vector field means something
different: the storage exists but was provisioned at a different shape —
typically the field's declared dimension changed after the table was
created. Boot deliberately leaves such a slot untouched (with a console
warning); run `store.reembedVectorField(kind, fieldPath)` to recreate the
storage at the new shape and re-embed.
**`ContributionUnavailableError`** with `state: "physical-storage-missing"`
means the physical fulltext table disappeared after initialization. Gated
fulltext operations preserve the driver error as `cause`, and transactional
backends roll back failed searchable writes. Query-builder fulltext predicates
compile directly to SQL and can still surface the engine's missing-relation
error. Run `store.rebuildContribution("fulltext")` to recreate the table and
repopulate it from the graph's nodes. A verified attach checks markers rather
than the physical catalog; use `probeContributions()` when startup must detect
out-of-band table loss.
**Solution:** Run `createStoreWithSchema(graph, adminBackend)` once
under the privileged role before the runtime attaches (it writes the
contribution markers that `createStore` / `createVerifiedStore` only
check), and prefer `createVerifiedStore()` over bare `createStore()` so
drift fails fast. See
[Database roles & least privilege](/backend-setup#database-roles--least-privilege).
A plain `createStore()` performs no reads and therefore cannot check the
base-schema marker at attach. If an edge write reaches legacy storage without
the match-identity columns, it throws `ConfigurationError` with
`details.code === "EDGE_MATCH_IDENTITY_STORAGE_UNAVAILABLE"` instead of leaking
the database driver's missing-column error. The remedy is the same privileged
base-schema adoption.
### `SCHEMA_WRITE_FENCE_UNSUPPORTED` on the first managed write
**Cause:** The Store carries committed schema metadata (for example, it came
from `createStoreWithSchema`, `createVerifiedStore`, an adapter equivalent, or a
cached `{ reconciled }` snapshot), but its backend cannot run transactions or
does not implement the schema-write fence. Common examples are Cloudflare D1,
`drizzle-orm/neon-http`, and incomplete custom backends. The attach can still
succeed for reads; writes fail closed rather than racing a schema change.
**Solution:** Use a transactional backend (`neon-serverless`, regular
PostgreSQL, SQLite with transactions, or Durable Objects SQLite). If raw,
unfenced writes are an explicit application decision, construct the Store with
`createStore()` / `createAdapterStore()` without `{ reconciled }` and quiesce
writers yourself around schema changes.
### A convergence or claim write refuses on a non-transactional backend
This is expected when the operation needs an interactive transaction. Static
adapter batches are not a public transaction, and a sequence of independent
requests cannot safely implement Operational Identity, claim/cardinality
checks, or undeclared dynamic `matchOn` convergence. Use a backend with
`capabilities.execution.interactiveTransactions === true` and call `store.transaction(...)` for
those operations.
A declared edge `matchIdentity` is the exception for the eligible root
`getOrCreateByEndpoints` path: its persisted canonical key is backed by a
database unique arbiter, so the authoritative one-statement command can
return `created` or `found` without an interactive transaction. This exception
does not cover claims, history/revision sidecars, or other writes in the same
application workflow.
### The store opens clean but a fulltext or vector read fails
**Cause:** The durable physical contribution marker still says `initialized`
while the table it names is gone — a partial restore, a
hand-run `DROP`, or a schema-scoped restore that missed the
strategy-owned tables. Nothing on the open path probes the catalog:
boot and the runtime asserts short-circuit on a per-instance signature
cache and then on durable marker rows alone, which keeps the hot path free
of catalog round trips. The cost is that this database opens
completely clean and fails at the first read or write that depends on the affected slot.
For deployment-scoped storage such as the shared fulltext table, readiness is
the conjunction of that physical marker and a graph-local activation marker.
Vector tables remain graph-scoped and use one graph-local physical marker.
**Diagnosis:** `store.verifyContributions()` reports detected drift or a
recorded failed attempt among contributions currently expected by the active
graph and backend strategies. Each entry carries `owner`, `logicalName`,
`physicalName`, a `state`, and — for vector slots — `kind` and
`fieldPath`. When the marker recorded an error against its last
attempt, `lastError` carries it: `state` tells you which repair to run,
`lastError` tells you why it broke, which is often a different
question. The call is read-only — one existence query per contribution
table, no DDL and no writes — so it is safe on a live store under a
least-privilege role. It is deliberately **not** a boot step; call it
from a health check or an operator script.
**Solution:** Run `repairContributions()` on a Store backed by the privileged
DDL-capable connection:
```typescript
const result = await adminStore.repairContributions();
for (const entry of result.results) {
if (entry.status === "failed") {
console.error(entry.diagnostic, entry.error);
}
if (entry.status === "requires-rebuild") {
console.warn("manual rebuild required", entry.diagnostic);
}
}
```
The method performs its own fresh audit and resolves contribution declarations
from the committed graph and the active backend strategies. It never accepts a
diagnostic, physical table name, or DDL from the caller. A Store opened before
another writer evolved the graph catches up before enumerating vector slots
instead of repairing from its stale in-memory graph snapshot.
| `state` | Result | Data behavior |
| --- | --- | --- |
| `missing-marker` | `repaired` or `failed` | Runs idempotent DDL and re-stamps the marker; existing rows are preserved |
| `failed-materialization` | `repaired` or `failed` | Retries the current idempotent contribution DDL |
| `orphaned-marker` | `requires-rebuild` | The table and its data are already gone |
| `stale` | `requires-rebuild` | The stored physical shape does not match the current declaration |
`remaining` is a fresh post-repair diagnostic pass. An empty `remaining` array
means no current declaration remains unhealthy after the pass. Once it is
empty, a second call is idempotent and returns no results.
For a vector `requires-rebuild` entry, use
`reembedVectorField(kind, fieldPath, { embed })`. It drops and recreates the
slot, so pass an `embed` callback or the field comes back with zero embeddings.
For a fulltext `requires-rebuild` entry, use
`rebuildContribution("fulltext")` — the third rung of the ladder, described
below. Do not hand-edit the marker or run backend-owned DDL directly.
`repairContributions()` intentionally does not use the public diagnostic as an
instruction list and does not force marker writes. A warm backend re-reads the
marker, and the normal signature guard still refuses to bless stale storage.
**An empty result does not mean everything was checked.** The diagnostic
enumerates only current declarations. It ignores retired marker rows and treats
an expected contribution with neither marker nor table as never attempted, so
`[]` is not proof of initialization. A backend that cannot probe its own catalog
throws `ConfigurationError` rather than reporting a clean bill of health, but
vector slots on a backend without vector support are skipped silently and
correctly — that backend never materialized them, so reporting them would be a
false positive on every store it opens. For a readiness check, first attach with
`createVerifiedStore()` to establish schema and marker initialization, then run
this diagnostic. Also assert `backend.capabilities.vector?.supported` when
embedding storage is required rather than treating an empty array as proof that
it is intact.
### Contribution health: probe, repair, rebuild
The three contribution maintenance operations form one escalation ladder.
Each rung does strictly more, and costs strictly more, than the one below
it. Start at the top of this table and stop as soon as the projection is
`ready`.
| Rung | Call | Writes | Use when |
| --- | --- | --- | --- |
| 1. Probe | `store.probeContributions()` | Nothing | You want to know whether search is coherent right now. Safe on a read path, on a replica, and under a least-privilege role |
| 2. Repair | `store.repairContributions()` | Marker rows and idempotent `CREATE ... IF NOT EXISTS` | The probe reports `degraded` and `verifyContributions()` says `missing-marker` or `failed-materialization` — storage is intact and only the bookkeeping is wrong |
| 3. Rebuild | `store.rebuildContribution("fulltext")` | **Deletes and refills this graph's rows**; drops and recreates the shared storage only when no other graph has rows in it | `verifyContributions()` says `stale` or `orphaned-marker`, which repair reports as `requires-rebuild` |
**Rung 1 — the read-only probe.** One entry per search projection the
graph declares, so a caller can decide whether to issue a query without
running a write operation first:
```typescript
const health = await store.probeContributions();
for (const entry of health.entries) {
if (entry.state !== "ready") {
console.warn(`${entry.contribution} search is ${entry.state}`, entry.detail);
}
}
```
`entries` is empty when there is nothing to assess — a graph with no
`searchable()` or `embedding()` fields, or a backend with no contribution
machinery. It is never empty because a check was skipped: a backend that
provisions contributions but cannot probe its catalog throws
`ConfigurationError`, and declares the gap as
`capabilities.contributions.probe === false`. Route on `state`; `detail`
is a human-readable summary and not a stable format, so call
`verifyContributions()` for the structured per-table findings behind it.
`graphRevision` stamps the durable revision the assessment was taken at,
placing the probe in the graph's committed history. It is graph-global
like the clock it reads: an advance between two probes means something
committed in between, not that a particular caller's write landed. It is
absent unless the Store is revision-tracked (`revisionTracking: true` or
`history: true`) and a tracked write has already anchored the clock. A
store with no revision clock has no revision to stamp, and substituting a
wall-clock timestamp or the schema version would be a weaker guarantee
wearing the name of a stronger one — the schema version in particular does
not advance on data writes, so it could not order anything.
`state: "building"` is reserved. No shipped path publishes it; the
destructive rebuild is atomic, so a concurrent probe observes the state
before or after it and never a partial one. Treat it as "not `ready`".
**Rung 3 — the destructive rebuild.** A `stale` fulltext contribution
means the table exists at the shape a *previous* `createDdl` produced. The
ordinary ensure path cannot fix it: its `CREATE ... IF NOT EXISTS` no-ops
against the existing table, and re-stamping the marker there would leave
it blessing storage whose shape is wrong — precisely what the drift guard
exists to prevent. Only a drop makes the recreate meaningful, so the drop
is its own named operation rather than a flag on the ensure path:
```typescript
const result = await adminStore.rebuildContribution("fulltext");
// { rebuilt: ["typegraph_node_fulltext"], processed, repopulated, skipped }
```
**Opening a Store to run it.** A `stale` contribution makes
`createStoreWithSchema()` refuse: its boot step materializes runtime
contributions, and the drift guard will not run the current DDL against a
table provisioned at another shape. That refusal is deliberate and
persistent — it repeats on every restart until the shape is fixed, and it
leaves the `stale` verdict intact rather than downgrading it to a state
whose repair would bless the wrong shape. Reach the rebuild from a Store
opened without that boot step, which `createStore()` and
`createVerifiedStore()` are (they run no DDL by contract):
```typescript
const adminStore = createStore(graph, backend);
await adminStore.probeContributions(); // degraded, detail names `stale`
await adminStore.rebuildContribution("fulltext");
// createStoreWithSchema() now opens normally again.
```
**The rebuild is scoped to the graph you call it on.** The fulltext
projection is one physical table holding every graph's rows keyed by
`graph_id`, so the default teardown is the same
`DELETE ... WHERE graph_id` that `clear()` issues — this graph's index
content and nothing else — followed by the current `createDdl`, a refill
from this graph's node rows, and the marker stamp, all in one transaction
under the same per-graph fence as a schema commit. That path takes **no
table lock at all**: the delete is transactional and touches only rows this
graph owns, so it never makes another graph's writers wait. An interrupted
rebuild rolls back to the state it started from rather than leaving storage
attested but empty.
It escalates to dropping and recreating that shared table only when the
table holds no other graph's rows — the case where the drop takes nothing
with it. That escalation is the one repair for storage provisioned at a
shape the current DDL no longer produces, and because the DDL it issues is
database-global it runs under a database-scoped advisory lock
(`typegraph:contribution-ddl`; a no-op on SQLite, whose fence already holds
the single writer slot) rather than only the per-graph fence.
**When the shared table is in use by another graph, a `stale` rebuild
refuses.** Only recreating the storage repairs a `stale` shape, so if that
storage still holds rows belonging to other graphs the call throws
`ContributionRebuildUnsupportedError` with
`reason: "shared-storage-in-use"` rather than destroying content it cannot
reconstruct — those rows are derived from other graphs' nodes through their
own schemas — or re-stamping this graph's marker over a physical shape
nothing verified. `details.otherGraphIds` names the graphs that are in the
way. The sanctioned repair is a maintenance window with every graph on that
database offline: drop the table out of band, then run
`store.rebuildContribution("fulltext")` once per graph, each run recreating
the table from the current DDL and refilling that graph's own rows.
What the *recreate* path costs, and why it is still the right trade: the
transaction is held for the whole refill, and on PostgreSQL the rebuild
takes `LOCK TABLE ... IN ACCESS EXCLUSIVE MODE` on the shared table and
keeps it until commit. It takes that lock **before** deciding to drop, not
merely as a side effect of the `DROP TABLE` — the verdict "no other graph
has rows here, so dropping this destroys nothing" is only as good as the
exclusion it was computed under. Ordinary fulltext writes take no advisory
lock, so the contribution DDL lock excludes other *rebuilds* and nothing
else: a neighbouring graph's `INSERT` could commit between an unlocked
probe and the drop, and be destroyed by a rebuild that had already decided
it was alone. The sequence is therefore probe → `ACCESS EXCLUSIVE` →
re-probe, and only the re-probe's verdict authorizes a drop. A verdict that
flips under the lock loses the drop and keeps the lock (PostgreSQL holds
locks until commit), which is the rare and safe direction to be wrong in.
The cheap unlocked probe ahead of it exists only to keep the graph-scoped
path off the relation lock, and can only err toward keeping the table.
The window blocks more than searches — every write to a kind with
`searchable()` fields maintains the same table, so those block too, for
every graph on the database. On SQLite the rebuild holds the write lock for
the same span, so concurrent writers wait out their busy timeout and then
fail. Run it in a maintenance window on a large graph.
When the storage *shape* is fine and only the content is stale — a field
gained `searchable()` after data was written, or a `language` changed —
`store.search.rebuildFulltext()` is the incremental, resumable pass that
transacts per page instead.
Nothing is permanently lost for the graph you rebuild: its searchable text
is derived from node properties TypeGraph already stores. Other graphs on
the same database are not in reach either — their rows are kept by the
graph-scoped delete, and the drop that would take them never runs (a
`stale` shape that could only be repaired by that drop refuses instead).
Nodes whose stored `props` cannot be read as an object are counted in
`skipped` and are absent from the rebuilt index;
`store.search.rebuildFulltext()` reports their ids individually.
**Vector contributions cannot be rebuilt, and the call refuses rather
than trying.** `rebuildContribution("vector")` always throws
`ContributionRebuildUnsupportedError` with
`reason: "vector-source-unavailable"`. TypeGraph stores the vectors
callers supply and never the inputs that produced them, so the embeddings
exist only in the storage a rebuild would drop — dropping anyway would
destroy them and hand back storage that looks healthy and returns
nothing. `reembedVectorField(kind, fieldPath, { embed })` is the
sanctioned destructive path for vector storage precisely because it takes
the callback that can regenerate what the drop discards.
The same typed error covers two wiring gaps, and both refuse before
anything is dropped: `reason: "no-drop-ddl"` when the active fulltext
strategy declares no `dropDdl` on its contribution, and
`reason: "no-schema-fence"` when the backend exposes no
`schemaWriteTransaction` to make the sequence atomic (the HTTP-only
PostgreSQL drivers, and SQLite with transactions disabled). Both are
declared ahead of time as `capabilities.contributions.rebuild === false`.
## Semantic Search Issues
### "Extension not found" / "vector type not available"
**Cause:** Vector extension not installed. Only applies to PostgreSQL
(pgvector) and SQLite (sqlite-vec). libSQL / Turso has a built-in
native vector engine — there is nothing to load and it is wired
automatically by `createLibsqlBackend`.
**PostgreSQL:**
```sql
CREATE EXTENSION IF NOT EXISTS vector;
```
**SQLite:**
```typescript
import * as sqliteVec from "sqlite-vec";
sqliteVec.load(sqlite); // Must be called before creating backend
```
### "Dimension mismatch"
**Cause:** Query embedding has different dimension than stored embeddings.
**Solution:** Use consistent embedding dimensions:
```typescript
// Schema defines 1536 dimensions
const Document = defineNode("Document", {
schema: z.object({
embedding: embedding(1536),
}),
});
// Query embedding must also be 1536
const queryEmbedding = await generateEmbedding(text);
console.log(queryEmbedding.length); // Should be 1536
```
### "Inner product not supported" (SQLite / libSQL)
**Cause:** `inner_product` is PostgreSQL-only. Neither sqlite-vec nor
libSQL support the inner product metric (cosine and l2 only). Check
`backend.capabilities.vector.metrics` for the active backend.
**Solution:** Use cosine or L2:
```typescript
// Instead of:
d.embedding.similarTo(query, 10, { metric: "inner_product" });
// Use:
d.embedding.similarTo(query, 10, { metric: "cosine" });
```
## TypeScript Issues
### "Property 'x' does not exist on type"
**Cause:** Accessing a property not defined in your schema.
**Solution:** Ensure the property is in your Zod schema:
```typescript
const Person = defineNode("Person", {
schema: z.object({
name: z.string(),
email: z.string().optional(),
}),
});
// Now both properties are available with correct types
const person = await store.nodes.Person.getById(id);
person?.name; // string
person?.email; // string | undefined
```
### Type inference not working in select
**Cause:** Complex generic inference limitations.
**Solution:** Use explicit typing or simplify:
```typescript
// If inference fails, be explicit
.select((ctx) => ({
name: ctx.p.name as string,
company: ctx.c.name as string,
}))
```
## Still Having Issues?
1. **Check the [Limitations](/limitations)** page for known constraints
2. **Review [Architecture](/architecture)** to understand how TypeGraph works
3. **Search [GitHub Issues](https://github.com/nicia-ai/typegraph/issues)** for similar problems
4. **Open a new issue** with a minimal reproduction case
# Errors
> Error types and handling in TypeGraph
TypeGraph uses typed errors to communicate specific failure conditions. All errors extend the base
`TypeGraphError` class and include categorization, contextual details, and actionable suggestions.
## Error Categories
Every error is categorized to help determine the appropriate response:
| Category | Description | Typical Response |
|----------|-------------|------------------|
| `user` | Invalid input or misuse of API | Fix the input and retry |
| `constraint` | Graph constraint violated | Handle as business logic violation |
| `system` | Internal or infrastructure error | Log, alert, potentially retry |
```typescript
import { isUserRecoverable, isConstraintError, isSystemError } from "@nicia-ai/typegraph";
try {
await store.nodes.Person.create(data);
} catch (error) {
if (isUserRecoverable(error)) {
// Show validation errors to user
return { error: error.toUserMessage() };
}
if (isConstraintError(error)) {
// Handle business rule violation
return { error: "This operation violates a constraint" };
}
if (isSystemError(error)) {
// Log and alert
console.error(error.toLogString());
throw error;
}
}
```
## Base Error
### `TypeGraphError`
Base error class for all TypeGraph errors.
```typescript
class TypeGraphError extends Error {
readonly code: string;
readonly category: ErrorCategory;
readonly details: Readonly>;
readonly suggestion?: string;
// Format error for end users (includes suggestion if available)
toUserMessage(): string;
// Format error for logging (includes code, category, and details)
toLogString(): string;
}
type ErrorCategory = "user" | "constraint" | "system";
```
**Properties:**
| Property | Type | Description |
|----------|------|-------------|
| `code` | `string` | Machine-readable error code |
| `category` | `ErrorCategory` | Error classification for handling |
| `details` | `Record` | Additional context about the error |
| `suggestion` | `string \| undefined` | Actionable guidance for resolution |
**Methods:**
| Method | Returns | Description |
|--------|---------|-------------|
| `toUserMessage()` | `string` | Human-readable message with suggestion |
| `toLogString()` | `string` | Detailed string for logging/debugging |
## Validation Errors
### `ValidationError`
Thrown when schema validation fails during node or edge creation/update. Includes structured issue
details with context about which entity failed.
```typescript
interface ValidationErrorDetails {
readonly issues: readonly ValidationIssue[];
readonly entityType?: "node" | "edge";
readonly kind?: string;
readonly operation?: "create" | "update";
readonly id?: string;
}
interface ValidationIssue {
readonly path: string;
readonly message: string;
readonly code?: string;
}
```
**Example:**
```typescript
try {
await store.nodes.Person.create({ name: "" }); // Empty name fails min(1)
} catch (error) {
if (error instanceof ValidationError) {
console.log(error.category); // "user"
console.log(error.details.kind); // "Person"
console.log(error.details.operation); // "create"
console.log(error.details.issues);
// [{ path: "name", message: "String must contain at least 1 character(s)" }]
console.log(error.toUserMessage());
// "Validation failed for Person create: name - String must contain at least 1 character(s)
//
// Suggestion: Check the data you're providing matches the schema..."
}
}
```
#### `INVERTED_VALIDITY_WINDOW`
A `ValidationError` whose issue carries the exported code
`INVERTED_VALIDITY_WINDOW` refused a valid-time window of negative width: the
write's `validTo` precedes the row's effective `validFrom`, so the row would have
stopped being true before it started and no `asOf` coordinate could observe it.
Branch on the code rather than on the message.
```typescript
import { INVERTED_VALIDITY_WINDOW_CODE, ValidationError } from "@nicia-ai/typegraph";
try {
// The stored validFrom is later than this end.
await store.edges.worksAt.update(edgeId, {}, { validTo: "2020-01-01T00:00:00.000Z" });
} catch (error) {
if (
error instanceof ValidationError &&
error.details.issues.some((issue) => issue.code === INVERTED_VALIDITY_WINDOW_CODE)
) {
// Supply an explicit validFrom for a historical window, or drop validTo.
}
}
```
Interchange import records the same refusal as a per-row error prefixed with the
code, so one bad row does not abort the import; trusted import refuses the whole
stream with `TrustedImportError` reason `invalid_stream`. A zero-width window
(`validTo === validFrom`) is legal and never raises this, and neither is a write
that STAMPS its own start while carrying only a historical `validTo`: any create,
and a node resurrection through `upsertById` / `bulkUpsertById`. Both store no
lower bound instead. An edge resurrection RETAINS the bound the row already
holds, so a `validTo` before that bound still raises this.
#### `IMMUTABLE_VALIDITY_LOWER_BOUND`
A `ValidationError` whose issue carries the exported code
`IMMUTABLE_VALIDITY_LOWER_BOUND` refused a `validFrom` the write could not apply.
A live row's lower bound is history: an in-place update never rewrites
`valid_from`, so a bound naming a different instant is refused rather than
accepted and silently dropped. The message names both instants — the one stated
and the one the row stores — so you can restate the stored bound without a
second read.
```typescript
import { IMMUTABLE_VALIDITY_LOWER_BOUND_CODE, ValidationError } from "@nicia-ai/typegraph";
try {
// The row is live and started at some other instant.
await store.nodes.Person.upsertById(id, props, { validFrom: "2020-01-01T00:00:00.000Z" });
} catch (error) {
if (
error instanceof ValidationError &&
error.details.issues.some(
(issue) => issue.code === IMMUTABLE_VALIDITY_LOWER_BOUND_CODE,
)
) {
// Omit validFrom, or restate the bound the row already holds.
}
}
```
What deliberately does not raise it:
- **Restating the stored bound.** Naming the instant the row already holds is
accepted; there is nothing to apply and nothing being ignored.
- **A create, or a resurrection.** Both write a fresh window, so a stated
`validFrom` is stored — that is the way to give a row a different lower bound.
- **`getOrCreateByEndpoints` returning an existing edge.** That branch performs
no write, so `validFrom` / `validTo` describe the row to create if none is
found. `clearValidTo` is refused on a live return-mode match because it names
a mutation; use `ifExists: "update"`.
- **A node upsert or endpoint-matched edge update with
`onImmutableLowerBound: "preserve"`.** This explicitly treats `validFrom` as
create/resurrection-only input. A live-row update keeps its stored lower
bound while still applying props and `validTo`; the default remains
`"refuse"` so an unqualified bound is never silently dropped. Edge updates
use the policy with `ifExists: "update"`; the bulk edge form sets it per item.
Under the default `"refuse"` policy, it reaches every path that accepts
`validFrom` against a live row: `upsertById`, `bulkUpsertById` (including a
repeated id in one batch, judged against the row the batch just queued),
`getOrCreateByEndpoints` / `bulkGetOrCreateByEndpoints` with
`ifExists: "update"`, and interchange import's `onConflict: "update"` legs —
where, as with the inverted-window refusal, it is recorded as a per-row error
prefixed with the code rather than aborting the import.
#### `ENTITY_ALREADY_EXISTS`
A `ValidationError` whose issue carries the exported code
`ENTITY_ALREADY_EXISTS` refused a create because the id is already taken.
`details.entityType` says whether a node or an edge was refused and
`details.kind` names its kind.
```typescript
import { ENTITY_ALREADY_EXISTS_CODE, ValidationError } from "@nicia-ai/typegraph";
try {
await store.nodes.Person.create({ name: "Alice" }, { id: takenId });
} catch (error) {
if (
error instanceof ValidationError &&
error.details.issues.some((issue) => issue.code === ENTITY_ALREADY_EXISTS_CODE)
) {
// Use a different id, or update the existing entity.
}
}
```
The code is the same whichever layer noticed, on either backend. A node create
finds out from its own existence probe — but the probe and the INSERT are two
statements, and PostgreSQL does not serialize two write transactions under its
default READ COMMITTED isolation, so a concurrent create of the same NEW id can
commit in between and the engine refuses the INSERT instead. (SQLite's
`BEGIN IMMEDIATE` gives the writer slot to one transaction at a time, so its probe
always sees the winner's row.) An edge create has no existence probe at all, so
the engine's refusal is always what reports a taken edge id. All of these raise the
same error, so a caller retrying a generated id needs one branch, not several.
`details.id` names the taken id, and is present for every single-entity create.
It is absent only when the refused statement inserted more than one row: the
engine reports that the statement collided without saying which row did, and its
transaction is already aborted, so there is nothing left to probe. No race is
needed to reach that — a bulk create of edges, whose ids you supplied and which
nothing probes, is refused this way on every backend. Treat `details.id` as
optional if you create in bulk.
This is about identity, not values. A conflict on a declared `unique` constraint
raises `UniquenessError` instead, and a violated `unique: true` index declaration
surfaces as the engine's own failure — neither is reshaped into this error.
### `DisjointError`
Thrown when attempting to create a node that violates a disjointness constraint.
```typescript
// If Person and Organization are disjoint:
await store.nodes.Person.create({ name: "Alice" }, { id: "entity-1" });
try {
// Same ID, different disjoint type
await store.nodes.Organization.create({ name: "Acme" }, { id: "entity-1" });
} catch (error) {
if (error instanceof DisjointError) {
console.log(error.category); // "constraint"
console.log(error.details);
// { nodeId: "entity-1", attemptedKind: "Organization", conflictingKind: "Person" }
console.log(error.suggestion);
// "Use a different ID for the new node, or delete the existing node first..."
}
}
```
### `IdentityContradictionError`
Thrown when an identity mutation would make the assertion ledger contradictory —
for example asserting two nodes are the same after they were asserted different,
folding a same-class pair the ontology forbids, or importing an archive whose
assertions conflict with the target graph. Only raised on identity-enabled
graphs.
```typescript
try {
await tx.identity.assertSame(alice, aliceCopy);
} catch (error) {
if (error instanceof IdentityContradictionError) {
console.log(error.code); // "IDENTITY_CONTRADICTION"
console.log(error.category); // "constraint"
console.log(error.details);
// {
// operation: "assertSame", // "assertSame" | "assertDifferent" | "fold" | "import"
// a: { kind: "Person", id: "..." },
// b: { kind: "Person", id: "..." },
// reason: "different-assertion", // "different-assertion" | "same-class" | "disjoint-kinds"
// conflictingAssertionId: "...", // present when an existing assertion conflicts
// conflictingKinds: ["Person", "Organization"], // present when reason is "disjoint-kinds"
// }
console.log(error.suggestion);
// "Retract the conflicting identity assertion or correct the graph ontology before retrying."
}
}
```
### Identity validity errors
`IdentityValidityWindowError` refuses a future start, future end, inverted
window, or a second non-identical open window for one current semantic pair.
Its code identifies the reason: `IDENTITY_VALIDITY_FUTURE_START`,
`IDENTITY_VALIDITY_FUTURE_END`, `IDENTITY_VALIDITY_INVERTED`, or
`IDENTITY_VALIDITY_OPEN_WINDOW_CONFLICT`.
`IdentityEndpointValidityError` (`IDENTITY_ENDPOINT_VALIDITY`) means an
explicit assertion window extends outside an endpoint node's own validity or
deletion bounds. Future or inverted identity windows are user-category input
errors. A second non-identical open window and an endpoint-window conflict are
constraint-category errors. Both classes are package-root exports.
### `IdentityMergeConflictError`
Detected at merge **plan time** when the branches being merged carry opposing
or otherwise contradictory identity truth: one branch asserts a pair `same`
while another asserts it `different` (directly, or transitively through a
chain of `same` assertions no single branch ever wrote), a branch retracts an
assertion that a different branch reasserts under a new id (a retract/reassert
race — a branch that reasserts a pair it *also* retracted itself is
convergent, not a conflict, and merges cleanly), or a branch asserts an
identity relation over a node another branch deleted. Extends `MergeError`, so
an `instanceof MergeError` catch covers it alongside the other merge failures.
`merge()` and `IdentityMergeConflictError` are both exported from
`@nicia-ai/typegraph/graph-merge`, not the package root. `merge()` takes an
array of branches and never throws a `MergeError` — it **returns** a
`Result`:
```typescript
import { merge, IdentityMergeConflictError, isErr } from "@nicia-ai/typegraph/graph-merge";
const result = await merge(store, [branch]);
if (isErr(result)) {
if (result.error instanceof IdentityMergeConflictError) {
console.log(result.error.code); // "GRAPH_MERGE_IDENTITY_CONFLICT"
console.log(result.error.details);
}
throw result.error;
}
```
### `MergeConstraintConflictError`
Returned when `merge()`, `mergeIncremental()`, or `applyMergePlan()` resolves a
plan whose final graph violates a deterministic store constraint. The store
remains the owner of constraint enforcement: the merge translates its typed
refusal only at the commit boundary, after the transaction has rolled back.
```typescript
import {
isErr,
merge,
MergeConstraintConflictError,
} from "@nicia-ai/typegraph/graph-merge";
const result = await merge(store, branches);
if (isErr(result) && result.error instanceof MergeConstraintConflictError) {
console.log(result.error.code); // "GRAPH_MERGE_CONSTRAINT_CONFLICT"
console.log(result.error.category); // "constraint"
console.log(result.error.details.constraintCode); // e.g. "CARDINALITY_ERROR"
console.log(result.error.details.edgeKind); // copied from the store error
console.log(result.error.cause); // the original CardinalityError, etc.
}
```
Cardinality, uniqueness, endpoint, disjointness, and restricted-delete
refusals share this surface when they arise from node or edge application.
The planner normally co-buckets nodes with the same declared unique key, but a
late store-owned uniqueness refusal uses the same completeness boundary rather
than falling back to a system error.
Identity truth conflicts retain `IdentityMergeConflictError`; backend,
environment, and stale-plan failures retain their existing system errors.
Constraint failure is atomic: neither graph writes nor merge provenance records
survive.
### Merge plan and evidence errors
The reviewable merge lifecycle also returns errors in its `Result` arm. It does
not throw them:
```typescript
import {
applyMergePlan,
isErr,
planMerge,
StaleMergePlanError,
} from "@nicia-ai/typegraph/graph-merge";
const planned = await planMerge(store, branches, options);
if (isErr(planned)) throw planned.error;
const applied = await applyMergePlan(store, planned.data);
if (isErr(applied)) {
if (applied.error instanceof StaleMergePlanError) {
// The reviewed artifact no longer describes the target. Plan and review again.
}
throw applied.error;
}
```
| Error | Code | Meaning |
| --- | --- | --- |
| `MergePlanCapabilityError` | `GRAPH_MERGE_PLAN_CAPABILITY` | The target cannot supply a durable revision fence for a cross-time plan. Enable `revisionTracking` or `history`; the contiguous `merge()` wrappers retain their documented compatibility behavior. |
| `MergePlanningStaleError` | `GRAPH_MERGE_PLANNING_STALE` | The target revision changed between the planner's opening and closing observations. This is an expected retry-and-replan outcome under concurrency: no artifact is returned, so recapture the target and create a new plan before retrying. |
| `StaleMergePlanError` | `GRAPH_MERGE_PLAN_STALE` | The target moved after planning, the plan already succeeded, or another concurrent application won. No plan writes committed. |
| `InvalidMergePlanError` | `GRAPH_MERGE_PLAN_INVALID` | The value failed the versioned plan schema or a semantic invariant. |
| `UnsupportedMergePlanVersionError` | `GRAPH_MERGE_PLAN_VERSION_UNSUPPORTED` | `formatVersion` is not supported by this TypeGraph version. |
| `MergePlanDigestMismatchError` | `GRAPH_MERGE_PLAN_DIGEST_MISMATCH` | Canonical plan content differs from the recorded digest. |
| `MergePlanTargetMismatchError` | `GRAPH_MERGE_PLAN_TARGET_MISMATCH` | The plan names a different graph id from the supplied target. |
| `MergePlanSchemaMismatchError` | `GRAPH_MERGE_PLAN_SCHEMA_MISMATCH` | The plan was resolved under a different active schema version or hash. |
| `MergePlanOriginMismatchError` | `GRAPH_MERGE_PLAN_ORIGIN_MISMATCH` | The target has an independently-created revision clock, even if its numeric revision happens to match. |
| `CandidateSourceError` | `GRAPH_MERGE_CANDIDATE_SOURCE` | A built-in candidate source failed. `details` identifies its source id, entity kind, and operation context. |
| `MatchEvidenceError` | `GRAPH_MERGE_EVIDENCE` | Candidate evidence is malformed or a score is non-finite. `NaN` and infinity are refused, never serialized or silently dropped. |
Plan validation and the target/schema/origin/revision fence run before canonical
writes. The revision check is inside the same transaction as apply, so two
concurrent attempts cannot both commit. A stale plan is not repaired or adapted:
create a new plan and obtain approval for its new `digest`.
Plans may contain the complete proposed application data. Their digest detects
content changes and gives approval systems a stable identity, but it is not a
signature and does not authenticate storage, authorize a caller, or prove who
created the artifact. Protect plan data and enforce those trust decisions in the
application before calling `applyMergePlan()`.
### `EndpointError`
Thrown when an edge is created with invalid endpoint types.
```typescript
// If worksAt only allows Person -> Company:
try {
await store.edges.worksAt.create(company, person, {}); // Wrong direction
} catch (error) {
if (error instanceof EndpointError) {
console.log(error.category); // "constraint"
console.log(error.suggestion);
// "Check the edge definition to see which node types are allowed..."
}
}
```
### `EndpointPairError`
Thrown when a [source-dependent edge](/core-concepts#source-dependent-targets)
receives a source/target combination that matches no declared pair. It extends
`TypeGraphError` directly, so catching `EndpointError` alone does not catch it.
An invalid source kind continues to produce `EndpointError`.
```typescript
import { EndpointPairError } from "@nicia-ai/typegraph";
try {
// Dynamic callers are checked at runtime, too.
// assignedTo allows Employee -> Department and Student -> Course.
await store.getEdgeCollection("assignedTo").create(employee, course, {});
} catch (error) {
if (error instanceof EndpointPairError) {
console.log(error.code); // "ENDPOINT_PAIR_ERROR"
console.log(error.category); // "constraint"
console.log(error.details);
// {
// edgeKind: "assignedTo", endpoint: "pair",
// fromKind: "Employee", toKind: "Course",
// allowedPairs: [
// { from: "Employee", to: "Department" },
// { from: "Student", to: "Course" },
// ],
// }
}
}
```
Malformed target maps and graph registrations that widen built-in constraints
fail at configuration time with `ConfigurationError`.
### `CardinalityError`
Thrown when a cardinality constraint is violated.
```typescript
// If worksAt has cardinality: "one" (person can only work at one company):
await store.edges.worksAt.create(alice, acme, { role: "Engineer" });
try {
await store.edges.worksAt.create(alice, otherCompany, { role: "Consultant" });
} catch (error) {
if (error instanceof CardinalityError) {
console.log(error.category); // "constraint"
console.log(error.details);
// { edgeKind: "worksAt", fromKind: "Person", fromId: "", cardinality: "one", existingCount: 1 }
console.log(error.suggestion);
// "Remove the existing edge before creating a new one, or update the existing edge..."
}
}
```
### `UniquenessError`
Thrown when a uniqueness constraint is violated.
```typescript
// If email has a unique constraint:
await store.nodes.Person.create({ name: "Alice", email: "alice@example.com" });
try {
await store.nodes.Person.create({ name: "Bob", email: "alice@example.com" });
} catch (error) {
if (error instanceof UniquenessError) {
console.log(error.category); // "constraint"
console.log(error.details);
// { constraintName: "unique_email", kind: "Person", existingId: "", newId: "", fields: ["email"] }
console.log(error.suggestion);
// "Use a different value for the unique field, or update the existing record..."
}
}
```
### `EdgeMatchIdentityConflictError`
Thrown when a direct edge create collides with the edge kind's declared
`matchIdentity`. Use `getOrCreateByEndpoints()` when the intended behavior is
to return the existing identity owner.
## Not Found Errors
### `NodeNotFoundError`
Thrown when a referenced node does not exist.
```typescript
try {
await store.nodes.Person.update("nonexistent-id", { name: "New Name" });
} catch (error) {
if (error instanceof NodeNotFoundError) {
console.log(error.category); // "user"
console.log(error.details); // { kind: "Person", id: "nonexistent-id" }
console.log(error.suggestion);
// "Verify the node ID is correct and the node hasn't been deleted..."
}
}
```
### `EdgeNotFoundError`
Thrown when a referenced edge does not exist.
```typescript
try {
await store.edges.worksAt.update("nonexistent-edge", { role: "Manager" });
} catch (error) {
if (error instanceof EdgeNotFoundError) {
console.log(error.category); // "user"
console.log(error.details); // { kind: "worksAt", id: "nonexistent-edge" }
console.log(error.suggestion);
// "Verify the edge ID is correct and the edge hasn't been deleted..."
}
}
```
### `KindNotFoundError`
Thrown when referencing a node or edge type that doesn't exist in the graph definition.
```typescript
try {
await store.query().from("NonExistentType", "n").execute();
} catch (error) {
if (error instanceof KindNotFoundError) {
console.log(error.category); // "user"
console.log(error.details); // { kindName: "NonExistentType", entity: "node" }
console.log(error.suggestion);
// "Check the graph definition to see which node and edge types are available..."
}
}
```
### `EndpointNotFoundError`
Thrown when an edge references a node that doesn't exist.
```typescript
try {
await store.edges.worksAt.create(
{ kind: "Person", id: "nonexistent" },
company,
{ role: "Engineer" }
);
} catch (error) {
if (error instanceof EndpointNotFoundError) {
console.log(error.category); // "user"
console.log(error.details);
// { edgeKind: "worksAt", endpoint: "from", nodeKind: "Person", nodeId: "nonexistent" }
console.log(error.suggestion);
// "Create the referenced node first, or verify the node ID is correct..."
}
}
```
## Delete Errors
### `RestrictedDeleteError`
Thrown when delete is blocked due to existing edges (when `onDelete: "restrict"`).
```typescript
// If Person has edges and onDelete is "restrict":
try {
await store.nodes.Person.delete(alice.id);
} catch (error) {
if (error instanceof RestrictedDeleteError) {
console.log(error.category); // "constraint"
console.log(error.details);
// { nodeKind: "Person", nodeId: "", edgeCount: 3, edgeKinds: ["worksAt", "authored"] }
console.log(error.suggestion);
// "Delete all edges connected to this node first, or change the delete behavior..."
}
}
```
## Configuration Errors
### `ConfigurationError`
Thrown when the store, backend, or schema definition is misconfigured.
```typescript
// Using transactions on D1 (which doesn't support them):
try {
await store.transaction(async (tx) => {
// ...
});
} catch (error) {
if (error instanceof ConfigurationError) {
console.log(error.category); // "system"
console.log(error.suggestion);
// "Check the backend documentation for supported features..."
}
}
```
#### Definition-time unique-constraint refusals
`defineGraph()` validates every node kind's `unique` constraints when the graph
is defined, rather than leaving a broken `where` clause to surface as odd
behavior on the first write. Three states are refused with `ConfigurationError`:
- A `where` callback that **does not return a predicate** — `details` carries
`kind` and `constraintName`.
- A predicate naming a **field the kind's schema does not declare** — `details`
adds `field` and `declaredFields`.
- A `where` clause on a kind whose **schema is not an object schema** (it
exposes no `.shape`, so there is no declared-field set to check the clause
against) — `details` carries `kind` and `constraintName`. Refused rather than
left unvalidated, because skipping the check silently would disable this guard
for exactly the untyped callers it exists for. A plain `unique: [{ fields }]`
on such a schema is *not* refused: it names props by key and evaluates fine
against a non-object schema.
All three carry only the class-level code `CONFIGURATION_ERROR`; match them by
class, not by a `details.code`. The equivalent invariant on the graph-extension
document path does have a stable code, `UNKNOWN_UNIQUE_WHERE_FIELD`.
A constraint built **outside** `defineGraph` never passed this gate, so the
non-predicate case is refused at evaluation too: `checkWherePredicate` throws the
same `ConfigurationError` (with `constraintName` and `fields`) on the write path
instead of treating a broken clause as one that applies to every row. All three
readers of a `where` clause — definition-time validation, per-write evaluation,
and persistence-time capture — now agree, because they read it through one
shared function.
Because the check evaluates the clause, a `where` callback now runs once at
definition time in addition to its per-write evaluations — keep it pure. The
check applies to node kinds whose schema exposes an object shape; edge `unique`
constraints are not validated here. Statically typed callers were already unable
to name an undeclared field, so this bites untyped or generated definitions.
#### Definition-time `__proto__` property refusal
`defineNode()` / `defineEdge()` refuse a schema that declares a property named
`__proto__` with a `ConfigurationError` carrying `details.conflicts` and a
`nodeType` / `edgeType` key. The name is **unstorable**, not merely reserved:
Zod accepts it in a shape but drops it from every parse result — reporting
success even when the field is required — so a value written to it is silently
lost.
It is only reachable through a computed key. `z.object({ __proto__: … })`
written literally sets the shape object's own prototype instead of creating an
entry, while `z.object({ ["__proto__"]: z.string() })` yields a shape whose
`Object.keys` really does contain it.
The graph-extension document path refuses the identical declaration with the
stable issue code `RESERVED_PROPERTY_NAME`, at any nesting depth — so a nested
object field named `__proto__` is refused on the same grounds as a top-level
one. Before this, the two authoring paths disagreed about the same field: a
typed refusal on the document path, silent data loss on the typed one.
#### `CONSTRAINT_WRITE_FENCE_UNSUPPORTED`
A write guarded by a declared constraint runs its probe and its write under one
per-graph mutual exclusion. That fence is transaction-scoped on both dialects
(SQLite's `BEGIN IMMEDIATE`, PostgreSQL's `pg_advisory_xact_lock`), so a backend
reporting `capabilities.execution.interactiveTransactions: false` — Cloudflare D1, `drizzle-orm/neon-http`,
any SQLite backend built with `transactionMode: "none"` — cannot hold it, and the
write is refused rather than run unfenced. Durable Objects are unaffected.
`details.constraint` names which class needed the fence, because "this backend
cannot fence constrained writes" is unusable advice while "your
`cardinality: 'one'` edge cannot be enforced here" is actionable. The
`suggestion` carries the per-class way forward.
| `details.constraint` | The write it describes |
| --- | --- |
| `edgeCardinality` | Creating or resurrecting an edge whose `cardinality` is `one`, `unique`, or `oneActive`. |
| `edgeMatchKeyConvergence` | Endpoint convergence that requires the portable transaction-scoped path: an undeclared dynamic `matchOn`, constrained cardinality, update or temporal options, derived/custom backends, or schema-aware resurrection of a tombstoned winner. A schema-declared durable `matchIdentity` removes this fence from eligible live single-item and bulk create/found paths. |
| `nodeDisjointness` | Creating a node under a kind that participates in a `disjointWith` axiom. Probed only where a node comes into existence, so deletes and in-place updates are not refused. |
| `nodeUniquenessScope` | Creating **or updating** a node under a `scope: "kindWithSubClasses"` unique that actually expands past the node's own kind. A `scope: "kind"` unique is backed by the uniques primary key and needs no fence. |
`details.graphId` names the graph. Unconstrained writes on the same backend are
untouched — see
[Declared constraints require an interactive transaction](/backend-setup#declared-constraints-require-an-interactive-transaction)
for what still works there.
`CONSTRAINT_TRANSACTION_NOT_WRITE_FENCED` is the corresponding refusal for a
caller-adopted SQLite transaction whose `DEFERRED` snapshot became stale before
the constrained write could take the writer slot. Roll back that transaction
and retry it with `BEGIN IMMEDIATE`; TypeGraph-owned transactions already use
that mode. The refusal happens before the constraint probe, so the write is
fenced or refused rather than allowed to rely on a stale decision.
#### `BATCH_WRITE_UNSUPPORTED`
A backend whose `capabilities.execution.unitOfWork` is `"batch"` (Cloudflare
D1's `batch()`, Neon HTTP's `transaction(queries)`) fixes every statement
before the first one runs and commits them together with no session in
between. Every fused write on such a backend — a static batch and a
certified atomic program alike — asserts the active schema version inside
the very statement that writes, so a stale version writes nothing and the
store reports `StaleVersionError`. See
[The guard every fused write shares](/limitations#the-guard-every-fused-write-shares).
A write that needs more than that one guarded statement refuses, but
`BATCH_WRITE_UNSUPPORTED` is not itself a top-level error code: the
enforcing gate keeps its own class and code
(`CONSTRAINT_WRITE_FENCE_UNSUPPORTED`, `UNSUPPORTED_BACKEND_CAPABILITY`,
`IDENTITY_REQUIRES_ATOMIC_BACKEND`, or a plain `ConfigurationError` for
`history` / `revisionTracking` / a schema commit) and nests
`{ code: "BATCH_WRITE_UNSUPPORTED", reason }` under `details.batchRefusal`,
naming what a closed batch cannot supply:
| `details.batchRefusal.reason` | What it needs | Raised by |
| --- | --- | --- |
| `interactive-callback` | Hold an interactive callback transaction open across several round trips. | `store.transaction(fn)` / `store.transactionWithReceipt(fn)` |
| `constraint-needs-probe` | Read a value it wrote earlier in the same write before deciding what to write next. | A declared constraint's probe-then-write (`CONSTRAINT_WRITE_FENCE_UNSUPPORTED`, above) |
| `identity` | Read and write Operational Identity's closure across several round trips inside one held transaction. | `Store` construction, or `requireAtomicIdentityBackend`, when `graph.identity` is declared |
| `history` | Hold the per-graph write lock and clock open across a whole write cascade. | `history: true` or `revisionTracking: true` |
| `schema-commit` | Hold one transaction across its compare-and-swap read and its activating write. | `commitSchemaVersion` / `setActiveVersion` |
`SCHEMA_WRITE_FENCE_UNSUPPORTED` — the portable schema-version fence an
ineligible write falls back to (see [Schema Migrations](/schema-management))
— does not carry `batchRefusal`. It is reached from many fuse failures that
are not specific to a batch-tier backend (an ineligible write kind, a
tombstone-resurrection write a supplied id falls through to, a derived
backend, a provenance mismatch), so it states its plain limitation without
guessing which of the reasons above, if any, applies.
#### Write-fence declaration codes
`capabilities.writeFence` resolves one of four write-fence plans a lock site
consumes — see
[Write fence declaration](/backend-setup#write-fence-declaration-writefence).
`ConfigurationError` codes name the ways a backend's fence declaration, or
its resolved plan, turns out not to cover what a write needs:
| `details.code` | Raised when |
| --- | --- |
| `WRITE_FENCE_DECLARATION_INVALID` | The declared `writeFence` fails runtime validation: an unrecognized `mechanism` string, an unrecognized `drain` string under `mechanism: "advisory"`, or a `drain` key present on `mechanism: "engine-serialized"` / `"caller-serialized"` (`drain` applies only to `"advisory"`). `details.field` names `"mechanism"` or `"drain"`; for an unrecognized value, `details.accepted` lists the allowed strings. Raised by `resolveWriteFencePlan` before any plan is shaped — an invalid `drain` never falls through to behaving like `"quiescent"`. |
| `WRITE_FENCE_SQL_UNAVAILABLE` | The resolved declaration's `mechanism` is `"advisory"` but the backend's `fenceSql` is missing the member that `mechanism`/`drain` combination needs to spell (`advisoryLockExpression`, `isolationFactExpression`, or, under `drain: "table-lock"`, `lockTables`) — or, independently of any lock plan, a session isolation-level read (recorded capture's isolation guard) finds no `fenceSql` at all. Raised at backend construction for the lock-plan case; at the point of the read for the session-fact case. |
| `RECORDED_CLOCK_REQUIRES_WRITE_FENCE` | The store is constructed with `history: true` or `revisionTracking: true` — TypeGraph-owned recorded-clock allocation — against a backend whose write-fence plan resolves `unfenced`. |
| `WRITE_FENCE_UNAVAILABLE` | A resolved plan cannot satisfy what a specific operation needs: either the plan is `unfenced` outright, or it is a `lock` plan whose `drain` is `"none"` meeting an operation whose `requires` is `"drain"`. `details.operation` names the operation and `details.requires` names which kind of exclusion (`"keyed"` or `"drain"`) it needed; a `drain: "none"` refusal also names the drain in the message. `"engine-serialized"` and `"caller-serialized"` satisfy either `requires` value without consulting `drain`. |
| `CALLER_SERIALIZED_REFUSES_ADOPTION` | `adoptTransaction` was called on a backend whose resolved write-fence plan is `caller-serialized`. An externally owned transaction's lifetime cannot be held by the backend's in-process write-unit queue, so `store.withTransaction(externalTx)` is refused rather than let its writes silently interleave with the queue's own. `details.member` names `"adoptTransaction"`. |
`RECORDED_CLOCK_REQUIRES_WRITE_FENCE` refuses at
`createStore`, never mid-flush, and the message names the exact declaration line to add.
`WRITE_FENCE_UNAVAILABLE` is not a `createStore`-time check: `requireWriteFence` is called from
every individual lock site (the identity graph lock, the identity-enablement drain, identity DDL,
trusted import, contribution DDL, recorded-clock allocation, schema-fence sites, graph-merge
provenance), so it fires wherever one of those runs — inside a live transaction, mid-operation,
not only at `createStore`. `WRITE_FENCE_SQL_UNAVAILABLE` and `WRITE_FENCE_DECLARATION_INVALID`
both refuse earlier, at backend construction for a `createSqlBackend`-built backend (or, for the
session-fact half of `WRITE_FENCE_SQL_UNAVAILABLE`, at the read that needed it), since they are
about the declaration itself rather than what a specific store option or operation requires of it.
`CALLER_SERIALIZED_REFUSES_ADOPTION` fires wherever `adoptTransaction` is actually called, which is
never at `createStore` time. `IDENTITY_REQUIRES_WRITE_FENCE` is another write-fence-related code —
see the Operational Identity guard codes table above — but is not in this table because it guards
identity construction, not recorded-clock allocation.
### Caller-serialized queue codes
The in-process queue a `writeFence: { mechanism: "caller-serialized" }` declaration builds
(`src/backend/serialized-execution-queue.ts`) raises two more `ConfigurationError` codes, both
naming `details.subject` — the SQLite dialect string for SQLite's own per-connection queue, or
`"caller-serialized"` for the write-unit queue a `caller-serialized` declaration builds:
| `details.code` | Raised when |
| --- | --- |
| `SERIALIZED_QUEUE_REENTRANT_SUBMISSION` | A queued operation was awaited from inside a transaction already running on the same queue — the transaction holds the queue's execution slot until it completes, so the nested operation could never run. Use the transaction-scoped context (`tx.nodes` / `tx.edges` / `tx.backend`) instead of the root store or backend inside a `store.transaction` callback, or move the operation outside the transaction. |
| `CALLER_SERIALIZED_REQUIRES_ASYNC_CONTEXT` | The queue's reentrancy detection depends on `node:async_hooks`' `AsyncLocalStorage`, which is unavailable on this runtime (or had not finished loading). A `caller-serialized` write-fence declaration's in-process promise depends on that detection actually working, so every submission is refused rather than run without it. SQLite's own per-connection queue never raises this code: it runs without detection instead of refusing when the context is unavailable. |
These codes are not part of `RECORDED_CAPTURE_GUARD_CODES` — that set is
closed to the three codes documented under
[Recorded-capture guard codes](#recorded-capture-guard-codes) below, and
`isRecordedCaptureGuardError` does not recognize any write-fence code.
### Optimistic-retry unit codes
The retry owner every `"optimistic-retry"`-tier unit of work runs through
(`src/backend/capabilities/retried-unit.ts`) raises one more
`ConfigurationError` code, naming `details.operation` — the same operation
name `TransactionConflictError` reports for the same unit:
| `details.code` | Raised when |
| --- | --- |
| `OPTIMISTIC_RETRY_REQUIRES_ASYNC_CONTEXT` | The unit's target is on the `"optimistic-retry"` execution tier (see [Backend Capabilities](/backend-setup#backend-capabilities)), and detecting a unit of work nested inside another one depends on `node:async_hooks`' `AsyncLocalStorage`, which is unavailable on this runtime. Running without that detection would let a nested unit's own independent retry commit against reads an outer attempt took before it ever conflicted, so the unit is refused, before its attempt ever runs, rather than run without it. A target on any other execution tier is unaffected: no nested owner exists there, so this code is never raised for it. |
#### Backend capability declaration codes
Custom backend declarations and capability bundles use stable `details.code` values when the
declared surface disagrees with what TypeGraph can safely execute:
| `details.code` | Raised when |
| --- | --- |
| `CAPABILITY_DECLARATION_CONTRADICTION` | `recursiveTraversal.supported` and its `reason` contradict each other: unsupported without a reason, or supported with a dangling reason. |
| `RECURSIVE_TRAVERSAL_UNSUPPORTED` | A backend declares recursive traversal unsupported and a recursive query, subgraph read, or historical identity operation needs it. `details.operation` names the refusing path and `details.reason` echoes the backend declaration. |
| `CONSTRAINT_CLAIM_SURFACE_MISMATCH` | The `constraintClaims` declaration and the claim members implemented by the backend disagree in either direction. |
| `BUNDLE_PORT_SURFACE_MISMATCH` | A non-claim capability bundle resolves a required member as present, but the backend port used by the operation cannot reach it. Fallback-disposition members degrade through their documented fallback instead of throwing this code. |
| `RECORDED_DDL_CONSTRAINT_NAME_MISMATCH` | `recordedTableDdl` names a primary-key constraint for only one of the temporary or final recorded-table name sets. |
The recorded-time preview migration also throws `UnsupportedBackendCapabilityError` with
`details.capability: "recordedTableDdl"` when a legacy schema needs rewriting and the custom
backend does not provide its DDL callback. See
[Migrating Preview Recorded Time](/schema-management#migrating-preview-recorded-time) and
[Capability bundles](/backend-setup#capability-bundles) for the corresponding migration and
backend-author guidance.
#### Approximate retrieval with a mismatched metric
`.similarTo(vector, k, { approximate: true, metric })` is refused with a
`ConfigurationError` when `metric` differs from the field's declared metric. An
ANN structure is built for one metric — `vec0` bakes `distance_metric` into the
virtual table, libSQL's DiskANN index is built with `metric=…`, pgvector's index
carries a per-metric operator class — so retrieving by the declared metric and
re-scoring under the override would return the declared metric's neighbors
wearing the override's scores. The two options state something that cannot both
hold, so the option is refused rather than downgraded to an exact scan behind the
caller's back.
`details` carries `nodeKind`, `fieldPath`, `requestedMetric`, `declaredMetric`,
and `indexType`; there is no stable `details.code`, so match by class and
`details`. A slot declared `indexType: "none"` is **not** refused — there is no
ANN structure to be bound to a metric, and the opt-in compiles to the exact scan,
a degradation stated on the `approximate` option itself. A mismatched metric with
no `approximate` is not refused on the query builder either; `store.search.vector`
and `store.search.hybrid` refuse every mismatched override on their own broader
rule. See
[Approximate retrieval](/semantic-search#approximate-retrieval-for-similarto-opt-in).
#### Durable edge match identity guard codes
Durable edge match identity uses stable `ConfigurationError` detail codes:
| `details.code` | Meaning |
| --- | --- |
| `EDGE_MATCH_IDENTITY_VALUE_NOT_SCALAR` | A declared identity field cannot be represented as a portable JSON scalar, or an untyped runtime value violated that declaration. |
| `EDGE_MATCH_IDENTITY_KEY_TOO_LARGE` | One complete durable identity tuple exceeds the portable 2,000-byte index budget. Normal import records this against the individual edge; trusted import is atomic and refuses the whole stream. |
| `EDGE_MATCH_IDENTITY_STORAGE_UNAVAILABLE` | The adapter declares durable identity support, but the database is missing its columns or unique arbiter. Initialize or migrate the schema before serving writes. |
| `EDGE_MATCH_IDENTITY_REQUIRES_ATOMIC_BACKEND` | Initial adoption needs an atomic empty-kind fence or materialization preflight that the custom backend does not implement. |
| `DURABLE_EDGE_MATCH_IDENTITY_COMMAND_UNSUPPORTED` | A custom backend declares durable identity support but refuses the authoritative convergence command. TypeGraph fails closed because the portable read-then-write fallback has no equivalent database arbiter. |
| `IMPORT_EDGE_BATCH_RETRY_REQUIRES_SAVEPOINT` | A durable import batch was refused without savepoint rollback protection, either because the backend is non-transactional or because its root/transaction statement-execution contract cannot serve savepoints. TypeGraph will not retry rows individually because that could double-attribute an already-written prefix. |
The last refusal deliberately differs from optional fused-command fallback:
the durable identity declaration delegates correctness to a database key, so a
backend that claims the feature but refuses its command cannot safely re-enter
the dynamic portable path.
#### Heterogeneous edge-read guard codes
`findEdgesByHeterogeneousEndpointSet` refuses mixed endpoint modes instead of
guessing how incident and exact-pair rows should be interpreted:
| `details.code` | Meaning |
| --- | --- |
| `EDGE_HETEROGENEOUS_READ_MIXED_ENDPOINT_MODES` | The request contains both incident-endpoint rows (without an opposite endpoint) and exact directed-pair rows. Supply an opposite endpoint for every row to request exact-pair matching. |
| `EDGE_HETEROGENEOUS_READ_BIND_BUDGET_EXCEEDED` | The endpoint set cannot fit within the backend's bind-parameter budget. Split the request into smaller calls. |
#### Operational Identity guard codes
Operational Identity lifecycle failures use stable `details.code` values on
`ConfigurationError`:
| `details.code` | Meaning |
| --- | --- |
| `IDENTITY_REQUIRES_ATOMIC_BACKEND` | The selected adapter cannot provide the interactive transaction required by identity writes. |
| `IDENTITY_REQUIRES_STATEMENT_EXECUTION` | The backend cannot execute the raw statements Operational Identity issues internally. |
| `IDENTITY_REQUIRES_WRITE_FENCE` | Operational Identity was constructed against a backend whose `capabilities.writeFence` resolves `unfenced` — declare the capability, matching the engine's real locking support. See [Write-fence declaration codes](#write-fence-declaration-codes). |
| `IDENTITY_NOT_ENABLED` | `store.identity`, `tx.identity`, `StoreView.identity`, or an identity-expanded query option was reached on a graph without `identity: { ... }` — normally caught at compile time; this is the runtime guard for a widened or `any`-typed handle. |
| `IDENTITY_STORAGE_MISSING` | An identity relation disappeared after enablement, or exists without this graph's fill. Restore ledgers, or recreate and rebuild the derived closure, before serving traffic. `details.reason: "unfilled"` marks the second case: the separation relation is present but holds no row for this graph while the ledger holds a live `different` assertion across two distinct identity classes — reopen the Store (the open runs the fill) or run `rebuildIdentityClosure(store)`. A Store handle opened while the relation did not exist keeps failing until it is reopened, which is deliberate: the alternative is a confident "not separated" the moment another graph's upgrade creates the shared relation. |
| `IDENTITY_UPGRADE_REQUIRES_ATOMIC_DDL` | The backend cannot publish the derived separation relation's upgrade — the `CREATE` and the fill — as one commit, on a graph that owes rows. `details.missingPorts` names what is absent: `schemaWriteTransaction` / `identityTableDdl` on the fenced path, or `executeSchemaDdl` on the schema-commit path. Refused rather than degraded, because a relation created empty and filled afterwards reads as "nothing is separated" in between. Both bundled Drizzle backends implement all three when transactions are enabled, so this is a custom-backend path. |
| `IDENTITY_ENABLEMENT_PENDING` | First enablement is pending because `autoMigrate` is disabled. |
| `IDENTITY_PROFILE_MIGRATION_PENDING` | A `sameIdAcrossKinds` change (a breaking `fold`↔`ignore` flip, or disabling identity) has not been applied — either it is breaking, or `autoMigrate` is disabled. |
| `IDENTITY_SCHEMA_MIGRATION_PENDING` | An identity-relevant ontology change is pending because `autoMigrate` is disabled. |
| `IDENTITY_SEPARATION_VIOLATION` | The derived separation relation refused a write that would place both endpoints of a current `different` assertion in one identity class. The database-level backstop beneath identity validation; reaching it means an earlier guard let a contradiction through. |
| `IDENTITY_TRANSACTION_NOT_WRITE_FENCED` | SQLite refused an identity write because the enclosing transaction was begun `DEFERRED` and another connection committed before it could take the writer slot. Only reachable through `store.withTransaction(externalTx)` / `store.withRecordedTransaction(externalTx)`, where the caller owns the `BEGIN` — TypeGraph's own transactions open `BEGIN IMMEDIATE` and hold the slot from the start. SQLite cannot upgrade a stale snapshot in place, so roll back and re-run the transaction, opening it with `BEGIN IMMEDIATE`. |
| `IDENTITY_SCHEMA_CONTRADICTION` | Existing nodes or assertions contradict the proposed identity profile or ontology, or the materialized closure disagrees with the assertions it was derived from. Run `rebuildIdentityClosure(store)` to recover from a closure mismatch. |
| `IDENTITY_IMPORT_REQUIRES_PROFILE` | An interchange document carries an `identity` section but the target graph does not have the profile enabled. |
| `IDENTITY_MERGE_REQUIRES_PROFILE` | A branch carries identity changes but the merge target graph does not have the profile enabled. |
| `IDENTITY_EXPORT_REQUIRES_TEMPORAL_FIELDS` | An identity-enabled export explicitly disabled temporal fields. Remove `includeTemporal` or set it to `true`; endpoint bounds are required to validate assertion windows on import. |
| `IDENTITY_IMPORT_ID_CONFLICT` | An imported assertion id already exists in the target ledger identifying different truth (relation, endpoints, or validity window). |
| `RECORDED_IDENTITY_SCHEMA_MISSING` | A `history: true` open of an identity-enabled graph could not find the recorded identity relation. Bundled backends provision it, so this is rare there and more likely on a custom backend. |
When an unapplied migration's **only** breaking change is the identity one, the
specific pending code above wins over the generic `MigrationError` (which is
attached as `cause`); a diff that also breaks nodes, edges, ontology, or
indexes raises the generic `MigrationError` enumerating all of them.
Identity import also raises `ValidationError` with one of these
`details.issues[].code` values when an interchange document's `identity`
section fails shape or integrity checks. Each issue carries the offending
assertion's id structurally in `details.issues[].assertionId`, and
`importGraph`/`importGraphStream` record these failures as
`entityType: "identity"` entries in `result.errors` (a self-assertion —
`IDENTITY_SELF_ASSERTION` — included) rather than throwing:
| Issue `code` | Meaning |
| --- | --- |
| `IDENTITY_IMPORT_UNKNOWN_KIND` | An assertion endpoint names a node kind not in the target graph's registry. |
| `IDENTITY_IMPORT_PAIR_NOT_NORMALIZED` | An assertion's `a`/`b` endpoints are not in code-point order. |
| `IDENTITY_STATE_IMPORT_ENDED_ASSERTION` | A `state`-mode import (the default) contains an already-ended assertion; use `identityMode: "archival"` on export to carry ended assertions. |
| `IDENTITY_IMPORT_FUTURE_VALID_FROM` | An open (current) assertion's `validFrom` is in the future, in either import mode. |
| `IDENTITY_IMPORT_FUTURE_VALID_TO` | An ended assertion's `validTo` is in the future. |
| `IDENTITY_IMPORT_INVALID_WINDOW` | An assertion's `validTo` precedes its `validFrom`. |
| `IDENTITY_IMPORT_ENDED_BY_WITHOUT_END` | An assertion names an `endedBy` cause but carries no `validTo`; only an ended assertion has a cause. |
| `IDENTITY_IMPORT_ENDED_BY_NOT_ENDPOINT` | An assertion's `endedBy` names a node that is not one of its own endpoints; a deletion cascade only ends assertions that touch the deleted node. |
| `IDENTITY_SELF_ASSERTION` | An assertion's `a` and `b` name the same node. |
#### Merge provenance sidecar codes
`persistProvenance: true` writes to a *sidecar* graph beside the merge target,
and `openProvenanceStore` refuses any sidecar graph id it cannot prove it owns.
Both refusals are `ConfigurationError`s with a stable `details.code`, and both
carry `details.graphId` (the sidecar id) and `details.targetGraphId`:
| `details.code` | Meaning |
| --- | --- |
| `GRAPH_MERGE_PROVENANCE_ID_COLLISION` | The sidecar graph id is occupied by something this library did not write. `details.reason` names which state was found, and the suggestion is specific to it. |
| `GRAPH_MERGE_PROVENANCE_CLAIM_UNFENCED` | The backend exposes no transactional schema fence (`schemaWriteTransaction`), so the id's emptiness check and its ownership-marker write cannot commit as one unit. Not a collision — the id may well be free. An already-owned sidecar still opens on such a backend, so read-only use of an existing sidecar stays available. |
The five `details.reason` values on `GRAPH_MERGE_PROVENANCE_ID_COLLISION`:
| `details.reason` | The state that was found |
| --- | --- |
| `application-graph` | The id holds rows (in any per-graph table) or a schema that is not the sidecar's, so it belongs to an application. When a pre-marker sidecar is classified, revision-change journal entries that record its own stored `Provenance` rows are not counted; every other journal entry is. Rename the colliding graph or point the merge elsewhere. |
| `empty-legacy-sidecar` | A pre-marker sidecar with no rows at all, which carries no evidence of authorship and is indistinguishable from an application graph of the same shape. |
| `unupgradeable-legacy-sidecar` | A pre-marker sidecar whose rows do not verify as provenance this library wrote for *this* target, so it cannot be upgraded to an owned sidecar. |
| `unowned-exact-schema-graph` | The current sidecar schema with no ownership marker. Because the marker is written *first*, this library cannot have produced this state; contents are not consulted, so an empty or provenance-shaped occupant is refused too. |
| `corrupt-ownership-marker` | A `ProvenanceOwner` row that is not a valid live claim for this target — soft-deleted, schema-invalid, naming a different target, or stored under a different row id. It is never overwritten or resurrected, because it may be an application's row. |
Under `persistProvenance: true` these arrive wrapped: the sidecar is opened and
claimed **before** the merge commits, and either code refuses the merge as an
`InvalidMergeOptionsError` (`details.option: "persistProvenance"`,
`details.provenanceErrorCode` echoing the code above, the `ConfigurationError`
as `cause`) with the target left unmodified. Only transient row-write failures
after the commit degrade to a `warnings` entry.
#### Interchange serialized-connection guard codes
Two long-lived interchange streams cannot share one serialized database
connection: an export snapshot holds a read transaction for the whole stream and
a streaming import writes a transaction per chunk on that same connection, so
the second one either nests a `BEGIN` or waits for a slot that never frees. The
lease is **exclusive** — one stream of any kind per connection — so all four
pairings refuse with a `ConfigurationError` rather than hanging:
| `details.code` | Raised when |
| --- | --- |
| `INTERCHANGE_SHARED_SERIALIZED_BACKEND_SNAPSHOT` | An export snapshot holds the connection, detected through the shared serialized resource the two backend wrappers were marked with. |
| `INTERCHANGE_SAME_SQLITE_BACKEND_SNAPSHOT` | The same condition, reported by the object-identity detector: one SQLite backend is exporting into itself. Worth telling apart because the fix differs — pass a second backend rather than await whatever else is running. |
| `INTERCHANGE_SERIALIZED_IMPORT_IN_PROGRESS` | A streaming import holds the connection, in either order of discovery. |
The code names *what holds the connection*; `details.requested` and
`details.heldBy` (each `"export-snapshot"` or `"import-stream"`) name which
pairing was actually refused, so a same-kind refusal is never reported as
something it is not. `details.graphId` names the graph the refused stream was
for.
`"import-stream"` is the kind of every long-lived import, not only
`importGraphStream`: `importGraph` holds the lease for the whole call, and
`trustedImportGraph` / `trustedImportGraphStream` hold it for the whole trusted
session — so those APIs throw this `ConfigurationError` as well as their own
`TrustedImportError`. Connections TypeGraph cannot observe are not refused: two
clients dialed at one server, or two SQLite handles on one file, are genuinely
independent. See
[Scaling branches and interchange](/graph-merge#scaling-branches-and-interchange)
for which drivers are recognized as serialized.
##### Declaring a connection the driver hides
Recognition is a duck-type over the client object, so a serialized driver
TypeGraph cannot identify (`expo-sqlite`, `op-sqlite`, `sqlite-proxy`,
`pg-proxy`, Bun `SQL`, a postgres-js client capped through a string it does not
coerce) is left unmarked and its stream pairs are not refused.
`createSqliteBackend` and `createPostgresBackend` accept a `serializedResource`
declaration for that gap — `{ mode: "shared", resource: client }` — and for the
reverse case, `{ mode: "independent" }`, when the detection is wrong for your
topology. See [Serialized connections](/backend-setup#serialized-connections).
The declaration is applied or refused, never quietly ignored:
| Declaration | Outcome |
| --- | --- |
| `{ mode: "shared", resource }` on a connection TypeGraph did not detect, or naming the client it did detect | The named object is the serialized resource; two backends naming the same object are one connection |
| `{ mode: "shared", resource }` naming a **different** object than the one detected | `ConfigurationError` (`code: "CONFIGURATION_ERROR"`) from the factory, with `details.reason: "serialized-resource-conflict"` and `details.declaredKind` / `details.detectedKind` naming what each side was |
| `{ mode: "independent" }` | Honored, whatever was detected — the documented escape hatch |
The conflict is refused rather than resolved because two wrappers over one
connection given two different sentinels would stop being seen as a pair, which
is precisely the refusal this guard exists to make.
The two `*Kind` details are constructor names (`"Database"`, `"BoundPool"`), not
the handles themselves: `details` is what `toLogString()` serializes, and a
driver handle there would print whatever that driver stores — a `pg.Pool` keeps
its `connectionString`, password included — into your logs.
`{ mode: "independent" }` lifts the shared-resource arm between two distinct
backend objects. It does **not** lift
`INTERCHANGE_SAME_SQLITE_BACKEND_SNAPSHOT`: one SQLite backend exporting into
itself holds the one snapshot transaction its own import writes through, which
is a fact about a single handle rather than a claim about connection topology.
Pass a second backend for that case. That surviving refusal is SQLite-only, so
on PostgreSQL a backend declared independent exporting into itself is not
refused either — a client that hands out independent connections is exactly
what the declaration claims.
#### `ExportStreamCancelledError`
An export stream whose `signal` fires settles with `ExportStreamCancelledError`
(`code: "INTERCHANGE_EXPORT_STREAM_ABORTED"`) rather than a silent end of stream,
so a consumer never mistakes a cancelled export for a complete one. It is thrown
only *after* the export has given back everything it took, so receiving it means
the connection is already free. What that was depends on the backend: a
transactional one rolls back the snapshot and releases the connection's stream
lease; one without transactions held neither and simply abandons its remaining
reads, its delivered chunks never having been a single snapshot. The message
says which.
`details.graphId` names the exported graph and `cause` carries the signal's own
`reason` when the caller supplied one. A signal that is already aborted refuses
the export before any transaction is opened. See
[Cancelling an export](/interchange#cancelling-an-export).
#### `ExportStreamIdleTimeoutError`
An `exportGraphStream` configured with `idleTimeoutMs` settles with
`ExportStreamIdleTimeoutError` (`code: "INTERCHANGE_EXPORT_STREAM_IDLE_TIMEOUT"`)
when its consumer does not request another chunk within that bound. The timeout
measures only the interval after a chunk is yielded; time spent waiting for the
backend to produce the next chunk does not count. `details.graphId` identifies
the graph and
`details.idleTimeoutMs` carries the configured bound. As with explicit
cancellation, a transactional export rolls its snapshot back and releases its
stream lease before the error is delivered; a non-transactional export held
neither and abandons its remaining reads. See
[Cancelling an export](/interchange#cancelling-an-export).
#### Recorded-capture guard codes
`ConfigurationError` is intentionally open-shaped, but the guards that fire on a
`history: true` / `revisionTracking: true` store carry a **stable, branchable
`details.code`** so a portable caller does not have to substring-match the
message. The three codes are exported as a set, `RECORDED_CAPTURE_GUARD_CODES`,
and reachable through the `isRecordedCaptureGuardError` type guard:
| `details.code` | Raised when |
|----------------|-------------|
| `RECORDED_CAPTURE_REQUIRES_CALLBACK_TRANSACTION` | `store.withTransaction(externalTx)` on a history-enabled store — it has no flush point before the caller commits. Use `store.withRecordedTransaction(externalTx, fn)`. (Also a compile error on an `AdapterHistoryStore`.) |
| `RECORDED_CAPTURE_RAW_SQL_DISABLED` | A raw SQL escape (`tx.sql`, `backend.executeStatement` / `executeDdl`) on a history-enabled store, where it would bypass recorded-time capture. |
| `REVISION_TRACKING_RAW_SQL_DISABLED` | The same raw SQL escape on a revision-tracked store, where it would bypass the revision anchor. |
Typed code cannot call `withTransaction` on an `AdapterHistoryStore`; use
`withRecordedTransaction` directly. The runtime code remains useful at
JavaScript and deliberately untyped boundaries. If one of those boundaries
throws, `isRecordedCaptureGuardError(error,
"RECORDED_CAPTURE_REQUIRES_CALLBACK_TRANSACTION")` narrows both the error and
its `details.code` without message matching.
Pass a specific code to narrow to one guard, or omit it to match any. The guard
narrows `error` to a `ConfigurationError` whose `details.code` is the passed
`RecordedCaptureGuardCode` (or the full union when no code is given), so no
untyped `details` spelunking is needed.
This composes with
[`tx.sqlAvailability`](/queries/temporal/#raw-sql-under-history-capture): the
discriminant tells a caller *why* `tx.sql` is unusable ahead of time
(`"history"` / `"revisionTracking"` vs. `"unavailable"` for a backend with no
transactions), while the guard code identifies a guard that has already thrown.
Between them, "history capture forbids raw SQL here" and "this backend has no
transactions" (which carries **no** guard code) are cleanly distinguishable
without catching-and-string-matching.
#### Engine-native recorded-time codes
A backend can track recorded (system) time itself by declaring
`GraphBackend.recordedTime` instead of using TypeGraph's own recorded
relations and clock — see [Engine-native recorded
time](/queries/temporal#engine-native-recorded-time) and [Supplying
`recordedTime`](/backend-authoring#supplying-recordedtime). Which ownership
form a store reads under is derived from that member's presence, never
declared separately, so there is no `recordedTimeOwnership` option to set.
Every refusal specific to that form carries a stable `details.code`:
| `details.code` | Raised when |
|----------------|-------------|
| `ENGINE_PROFILE_RECORDED_TIME_REQUIRES_LINEAGE` | An engine profile declares `recordedTime` without also declaring `lineage` — engine-native history keeps no recorded relations of its own for TypeGraph to derive a graph-merge change delta from. Raised at backend construction, naming both members. |
| `RECORDED_TIME_UNAVAILABLE` | A caller reached `requireRecordedTime` and found `recordedTime` absent on the backend it asked — store construction under `history: true` and the shared `recordedNow()`/`revisionNow()`/receipt-stamping read, both reached only once ownership has already resolved to `"engine-native"`. |
| `ENGINE_NATIVE_REVISION_TRACKING_UNSUPPORTED` | A store is constructed with `revisionTracking: true` against an engine-native backend, whether or not `history: true` is also requested — there is no TypeGraph clock for `revisionTracking` to advance; the engine's own revision is available only under `history: true`. |
| `ENGINE_NATIVE_RECORDED_READ_UNSUPPORTED` | A store is constructed with an external `recordedRead` binding against an engine-native backend — there is no TypeGraph recorded relation for one to populate. |
| `ENGINE_NATIVE_RECORDED_IDENTITY_UNSUPPORTED` | `store.identityAtCoordinate` at a past recorded instant, or the query compiler's historical identity traversal, is reached under engine-native recorded time — identity history reads TypeGraph's own recorded relations directly, which an engine-native backend does not populate. |
| `ENGINE_NATIVE_MIGRATE_RECORDED_TIME_UNSUPPORTED` | `migrateLegacyRecordedTime` is called against an engine-native backend — the migration rewrites TypeGraph's own recorded relations, which an engine-native backend does not have. |
| `RECORDED_INSTANT_OWNERSHIP_MISMATCH` | `store.asOfRecorded(instant)` receives an instant minted under the OTHER recorded-time ownership form — an `r1:` (TypeGraph-owned) instant against an engine-native store, or an `e1:` (engine-native) instant against a TypeGraph-owned store. |
The profile refusal fires at backend construction; the two store-option
refusals, and `RECORDED_TIME_UNAVAILABLE`'s construction arm, fire at
`createStore`; the remaining codes, and `RECORDED_TIME_UNAVAILABLE`'s read
arm, fire at the specific call that cannot be honored. None of these codes
are members of `RECORDED_CAPTURE_GUARD_CODES` above — that set stays closed
to the three TypeGraph-capture guards.
### `SchemaMismatchError`
Thrown when the database schema doesn't match the expected graph definition.
```typescript
try {
const [store] = await createStoreWithSchema(graph, backend);
} catch (error) {
if (error instanceof SchemaMismatchError) {
console.log(error.category); // "system"
console.log(error.details);
// { graphId: "my-graph", expectedHash: "", actualHash: "" }
console.log(error.suggestion);
// "Run migrations to update the database schema..."
}
}
```
### `MigrationError`
Thrown when schema migration fails due to breaking changes that require manual intervention.
The `details.reason` value `"edge-match-identity-rekey"` means a populated edge kind
changed or newly adopted its durable match identity. Existing rows cannot be assigned
new identity keys without choosing how conflicts converge. Export the affected edges,
hard-delete them, apply the schema migration, then reimport them so TypeGraph
materializes and arbitrates the new durable keys.
```typescript
try {
const [store] = await createStoreWithSchema(graph, backend);
} catch (error) {
if (error instanceof MigrationError) {
console.log(error.category); // "system"
console.log(error.details);
// { graphId: "my-graph", fromVersion: 3, toVersion: 4, reason: "Removed required field 'email' from Person" }
console.log(error.suggestion);
// "Review the breaking changes and perform manual migration if needed..."
}
}
```
### `BaseSchemaMigrationError`
Thrown by zero-DDL verified and graph-template entry points when the
deployment-wide physical TypeGraph schema has not been adopted to the version
required by the running library. This is separate from `MigrationError`, which
describes one graph's serialized schema evolution.
```typescript
try {
const [store] = await createVerifiedStore(graph, backend);
} catch (error) {
if (error instanceof BaseSchemaMigrationError) {
console.log(error.details);
// {
// installedVersion: undefined,
// requiredVersion: 1,
// reason: "missing"
// }
}
}
```
`reason` is `"missing"`, `"stale"`, or `"newer"`. For missing or stale
storage, run `createStoreWithSchema()` or `createAdapterStoreWithSchema()` once
under a DDL-capable role, or apply the published external base-schema migration.
A newer marker requires a TypeGraph release that supports that version.
## Query Errors
### `UnsupportedPredicateError`
Thrown when using a query predicate that isn't supported by the current backend.
```typescript
// Using vector similarity on a backend without vector support:
try {
await store
.query()
.from("Document", "d")
.whereNode("d", (d) => d.embedding.similarTo(queryVector, 10))
.select((ctx) => ctx.d)
.execute();
} catch (error) {
if (error instanceof UnsupportedPredicateError) {
console.log(error.category); // "system"
console.log(error.suggestion);
// "Use a backend that supports this predicate, or rewrite the query..."
}
}
```
## Transaction Errors
### `TransactionClosedError`
Thrown when a statement reaches a transaction-scoped backend after its
transaction boundary has already returned.
A transaction pins one database connection, which carries one statement at a
time. When `store.transaction(...)` resolves or rejects, the driver emits
`COMMIT` or `ROLLBACK` on that connection and hands it back to the pool. Any
statement still in flight then has nowhere safe to go — it would execute inside
somebody else's transaction — so TypeGraph refuses it.
The usual source is a callback that lets work escape it. `Promise.all` rejects
on its first rejection while its siblings keep running:
```typescript
await store.transaction(async (tx) => {
// If `a` fails, `b`'s remaining statements are orphaned.
await Promise.all([tx.nodes.Doc.create(a), tx.nodes.Doc.create(b)]);
});
```
You will normally never see this error: `Promise.all` has already rejected with
the original failure and discards the orphan's. It surfaces only if you await
the orphaned promise yourself. To avoid orphaning writes at all, use
`Promise.allSettled` and inspect the results, or await the writes in sequence.
`adoptTransaction()` never closes its queue — only the caller knows when their
transaction ends — so this error cannot arise there. It remains the caller's
job to await every graph write before committing.
### `TransactionConflictError`
Thrown when a transaction was aborted by a serialization failure or deadlock
on every attempt available to it. `details.operation` names the transaction
that failed, `details.attempts` the number tried, and `cause` is the last
attempt's driver error — PostgreSQL's own protocol for both conditions is to
re-run the whole transaction from the top, which is what this error reports
as exhausted.
`store.transaction()` and `store.transactionWithReceipt()` raise it with
`attempts: 1` for a conflict on their single attempt; passing
`retry: { attempts }` (see
[Retrying on conflict](/schemas-stores/#retrying-on-conflict)) raises it only
once every attempt has conflicted. Graph-merge's commit paths raise
`MergeError` on the same exhaustion, with a `TransactionConflictError` as its
`cause`.
```typescript
try {
await store.transaction(fn, { retry: { attempts: 3 } });
} catch (error) {
if (error instanceof TransactionConflictError) {
console.log(error.details.attempts); // 3
console.log(error.cause); // the last driver error
}
}
```
**The serialization covers TypeGraph's own statements, not `tx.sql`.** The raw
Drizzle handle you get for writing your own relational tables in the same
transaction shares the one pinned connection but bypasses the queue. Running a
raw statement concurrently with a graph write — or with another raw statement —
races two queries on that connection (the overlap `pg@9` removes), and the
boundary cannot drain a raw statement it never saw. Await each `tx.sql`
statement before the next write.
## Error Handling Patterns
### Using Error Utilities
TypeGraph provides utility functions for common error handling patterns:
```typescript
import {
isTypeGraphError,
isUserRecoverable,
isConstraintError,
isSystemError,
getErrorSuggestion,
} from "@nicia-ai/typegraph";
try {
await store.nodes.Person.create(data);
} catch (error) {
if (!isTypeGraphError(error)) {
// Not a TypeGraph error, handle differently
throw error;
}
// Get suggestion regardless of error type
const suggestion = getErrorSuggestion(error);
if (isUserRecoverable(error)) {
// User can fix this by providing different input
return {
error: error.toUserMessage(),
suggestion,
};
}
if (isConstraintError(error)) {
// Business rule violation
return {
error: "This operation violates a constraint",
details: error.details,
};
}
if (isSystemError(error)) {
// Infrastructure/configuration issue
console.error(error.toLogString());
throw error;
}
}
```
### Catch Specific Errors
```typescript
import {
ValidationError,
NodeNotFoundError,
DisjointError,
} from "@nicia-ai/typegraph";
try {
await store.nodes.Person.create(data);
} catch (error) {
if (error instanceof ValidationError) {
// Handle validation failure with contextual details
return {
error: "Invalid data",
issues: error.details.issues,
entity: error.details.kind,
};
}
if (error instanceof DisjointError) {
// Handle constraint violation
return { error: "ID already used by different type" };
}
throw error; // Re-throw unexpected errors
}
```
### Check Error Codes
```typescript
try {
await store.nodes.Person.update(id, data);
} catch (error) {
if (error instanceof TypeGraphError) {
switch (error.code) {
case "NODE_NOT_FOUND":
return { error: "Person not found" };
case "VALIDATION_ERROR":
return { error: "Invalid data", issues: error.details.issues };
default:
throw error;
}
}
throw error;
}
```
### Transaction Error Handling
```typescript
try {
await store.transaction(async (tx) => {
const person = await tx.nodes.Person.create({ name: "Alice" });
const company = await tx.nodes.Company.create({ name: "Acme" });
await tx.edges.worksAt.create(person, company, { role: "Engineer" });
});
} catch (error) {
// Transaction is automatically rolled back on any error
if (error instanceof ValidationError) {
console.log("Validation failed, transaction rolled back");
console.log("Failed on:", error.details.kind, error.details.operation);
}
throw error;
}
```
## Contextual Validation Utilities
For library authors or advanced use cases, validation utilities are available from the schema sub-export:
```typescript
import {
validateNodeProps,
validateEdgeProps,
wrapZodError,
createValidationError,
} from "@nicia-ai/typegraph/schema";
// Validate node properties with full context
const validated = validateNodeProps(PersonSchema, inputData, {
kind: "Person",
operation: "create",
});
// Wrap a Zod error with TypeGraph context
try {
schema.parse(data);
} catch (zodError) {
throw wrapZodError(zodError, {
entityType: "node",
kind: "Person",
operation: "update",
id: "person-123",
});
}
```
## Error Codes Reference
| Code | Error Class | Category | Description |
|------|-------------|----------|-------------|
| `VALIDATION_ERROR` | `ValidationError` | user | Schema validation failed |
| `DISJOINT_ERROR` | `DisjointError` | constraint | Disjointness constraint violated |
| `IDENTITY_CONTRADICTION` | `IdentityContradictionError` | constraint | Identity mutation would make the assertion ledger contradictory |
| `IDENTITY_VALIDITY_FUTURE_START` | `IdentityValidityWindowError` | user | Identity assertion starts after the operation clock |
| `IDENTITY_VALIDITY_FUTURE_END` | `IdentityValidityWindowError` | user | Identity assertion ends after the operation clock |
| `IDENTITY_VALIDITY_INVERTED` | `IdentityValidityWindowError` | user | Identity assertion ends before it starts |
| `IDENTITY_VALIDITY_OPEN_WINDOW_CONFLICT` | `IdentityValidityWindowError` | constraint | A different open window already represents the current semantic pair |
| `IDENTITY_ENDPOINT_VALIDITY` | `IdentityEndpointValidityError` | constraint | An endpoint does not cover the explicit assertion window |
| `GRAPH_MERGE_IDENTITY_CONFLICT` | `IdentityMergeConflictError` | system | Branches carry opposing identity truth |
| `GRAPH_MERGE_CONSTRAINT_CONFLICT` | `MergeConstraintConflictError` | constraint | The resolved merge would violate a store constraint |
| `ENDPOINT_ERROR` | `EndpointError` | constraint | Invalid edge endpoint types |
| `ENDPOINT_PAIR_ERROR` | `EndpointPairError` | constraint | Undeclared source/target combination |
| `CARDINALITY_ERROR` | `CardinalityError` | constraint | Cardinality constraint violated |
| `UNIQUENESS_VIOLATION` | `UniquenessError` | constraint | Uniqueness constraint violated |
| `EDGE_MATCH_IDENTITY_CONFLICT` | `EdgeMatchIdentityConflictError` | constraint | A direct edge write collided with its declared endpoint/property identity |
| `NODE_NOT_FOUND` | `NodeNotFoundError` | user | Referenced node doesn't exist |
| `EDGE_NOT_FOUND` | `EdgeNotFoundError` | user | Referenced edge doesn't exist |
| `KIND_NOT_FOUND` | `KindNotFoundError` | user | Unknown node/edge type |
| `ENDPOINT_NOT_FOUND` | `EndpointNotFoundError` | user | Edge endpoint node doesn't exist |
| `RESTRICTED_DELETE` | `RestrictedDeleteError` | constraint | Delete blocked by existing edges |
| `CONFIGURATION_ERROR` | `ConfigurationError` | system | Invalid configuration |
| `SCHEMA_MISMATCH` | `SchemaMismatchError` | system | Database schema mismatch |
| `MIGRATION_ERROR` | `MigrationError` | system | Migration failed |
| `BASE_SCHEMA_MIGRATION_REQUIRED` | `BaseSchemaMigrationError` | system | Deployment-wide base storage requires privileged adoption |
| `UNSUPPORTED_PREDICATE` | `UnsupportedPredicateError` | system | Predicate not supported |
| `UNSUPPORTED_BACKEND_CAPABILITY` | `UnsupportedBackendCapabilityError` | user | The backend does not advertise a capability the call needs. `details.capability` names it — for example `vector.searchFrontierTuning` for `efSearch` on any SQLite vector or hybrid search, where the engine has no per-search ANN frontier, with `details.reason` naming the limitation |
| `INTERCHANGE_EXPORT_STREAM_ABORTED` | `ExportStreamCancelledError` | user | An export stream's `signal` fired, after the export gave back everything it took. On a transactional backend that is the snapshot transaction and the connection's stream lease; on one without transactions the export held neither and simply abandoned its remaining reads. The message says which |
| `INTERCHANGE_EXPORT_STREAM_IDLE_TIMEOUT` | `ExportStreamIdleTimeoutError` | user | An export stream's consumer left a delivered chunk unacknowledged past its configured `idleTimeoutMs`; the export settled its snapshot and lease before reporting the timeout |
| `TRANSACTION_CONFLICT` | `TransactionConflictError` | system | A transaction was aborted by a serialization failure or deadlock on every attempt available to it. `details.attempts` is the number tried; `cause` is the last driver error |
# Architecture
> How TypeGraph works internally and the design decisions behind it
This page explains how TypeGraph works under the hood, the design decisions that shaped it, and why certain
tradeoffs were made.
## High-Level Architecture
```text
┌────────────────────────────────────────────────────────┐
│ Your Application │
│ │
│ ┌──────────────────────────────────────────────────┐ │
│ │ TypeGraph Library │ │
│ │ │ │
│ │ ┌────────────┐ ┌────────────┐ │ │
│ │ │ Schema │ │ Query │ │ │
│ │ │ DSL │ │ Builder │ │ │
│ │ └──────┬─────┘ └─────┬──────┘ │ │
│ │ │ │ │ │
│ │ └──────────────┴───────────────┘ │ │
│ │ │ │ │
│ │ ▼ │ │
│ │ ┌──────────────────┐ │ │
│ │ │ Ontology Layer │ │ │
│ │ └──────────────────┘ │ │
│ └─────────────────────────┬────────────────────────┘ │
│ │ │
│ ▼ │
│ ┌────────────────────────┐ │
│ │ TypeGraph Backend Port │ │
│ └────────────┬───────────┘ │
│ │ │
│ ┌──────▼───────┐ │
│ │ SQL Adapter │ │
│ │ (Drizzle) │ │
│ └──────┬───────┘ │
│ │ │
└───────────────────────────┼────────────────────────────┘
│
▼
┌─────────────────┐
│ Your Database │
└─────────────────┘
```
TypeGraph is an **embedded library**, not a database. It runs in your application
process and compiles queries to SQL. A managed local Store can own its SQLite or
PGlite connection; adapter integrations can instead use a connection your
application already owns.
### Capability boundaries
TypeGraph's public types match its supported runtime surfaces. The default
`Store` exposes graph operations; `AdapterStore` adds native transaction
interoperability; transaction contexts expose a read-only backend projection.
Every store exposes `store.capabilities`, the read-only runtime feature
descriptor used by its backend, so portable code can branch on atomicity,
vector, fulltext, or analytics support without reaching through the adapter
boundary.
Whenever a surface loses capabilities, TypeGraph constructs an explicit
allowlist projection. Proxy overlays are reserved for decorating a surface
without changing its capabilities.
The internal store and transaction ports are non-enumerable symbol properties
and are absent from public TypeScript contracts. They are not a JavaScript
security boundary: sufficiently reflective code can discover symbol properties,
as it can inspect internals in any in-process library. The guarantee applies to
all documented and supported access paths.
## Core Design Principles
### 1. Embedded, Not External
**Decision**: TypeGraph is a library dependency, not a separate service.
**Why**: Graph databases like Neo4j require managing another piece of infrastructure. For many use
cases—knowledge bases, organizational structures, content relationships—the graph is part of your application,
not a standalone system.
**Tradeoff**: You don't get Neo4j's broad graph-data-science suite (community
detection, betweenness centrality, and similar specialized analytics), but you
avoid:
- Additional deployment complexity
- Network latency between app and graph
- Separate scaling and monitoring
- Data synchronization challenges
### 2. Schema-First, Type-Driven
**Decision**: Zod schemas are the single source of truth. TypeScript types are inferred, not duplicated.
**Why**: In many graph systems, you define types in one place, validation in another, and database schemas in a
third. This leads to drift and bugs.
With TypeGraph:
```typescript
const Person = defineNode("Person", {
schema: z.object({
name: z.string().min(1),
email: z.string().email().optional(),
}),
});
// TypeScript type is inferred automatically
type PersonProps = z.infer;
// { name: string; email?: string }
```
The schema drives:
- Runtime validation on create/update
- TypeScript types for compile-time safety
- Database storage format
- Query builder type constraints
### 3. SQL as the Execution Engine
**Decision**: Compile graph queries to SQL, don't implement a custom query engine.
**Why**: SQLite and PostgreSQL are battle-tested, highly optimized query engines. Rather than building another one:
```typescript
// Your query
store.query()
.from("Person", "p")
.traverse("worksAt", "e")
.to("Company", "c")
.select((ctx) => ({ person: ctx.p.name, company: ctx.c.name }))
// Compiles to SQL with CTEs
WITH person_cte AS (
SELECT * FROM typegraph_nodes WHERE kind = 'Person' AND deleted_at IS NULL
),
edge_cte AS (
SELECT * FROM typegraph_edges WHERE kind = 'worksAt' AND deleted_at IS NULL
),
company_cte AS (
SELECT * FROM typegraph_nodes WHERE kind = 'Company' AND deleted_at IS NULL
)
SELECT
p.props->>'name' as person,
c.props->>'name' as company
FROM person_cte p
JOIN edge_cte e ON e.from_id = p.id
JOIN company_cte c ON c.id = e.to_id
```
This means:
- You get database-level query optimization
- Indexes work as expected
- Transactions are ACID
- You can analyze queries with EXPLAIN
### 4. Precomputed Ontology
**Decision**: Compute transitive closures at store initialization, not query time.
**Why**: Semantic relationships like `subClassOf` and `implies` form hierarchies. Computing "all subclasses of
Media" during every query would be expensive.
Instead, when you create a store:
```typescript
const store = createStore(graph, backend);
// ↑ Computes:
// - subClassOf closure: Media → [Media, Podcast, Article, Video]
// - implies closure: marriedTo → [marriedTo, partneredWith, knows]
// - disjoint sets: Person ⊥ Organization ⊥ Product
```
These closures are stored in the `TypeRegistry` and used during query compilation:
```typescript
.from("Media", "m", { includeSubClasses: true })
// At compile time, expands to: WHERE kind IN ('Media', 'Podcast', 'Article', 'Video')
```
**Tradeoff**: Changing the ontology requires recreating the store. But ontologies typically change rarely
compared to instance data.
### 5. Homoiconic Schema Storage
**Decision**: Store the graph schema and ontology as data in the database itself.
**Why**: Most ORMs and graph libraries define schemas only in application code. The database stores data but
has no record of what the data means. This creates problems:
- You can't understand the database without reading the application source
- Schema changes are invisible—no history, no diff, no audit trail
- Exports require the application to interpret the data
- Multiple applications can't share schema understanding
TypeGraph takes a different approach: the schema is data. When you initialize a store, the complete schema
(node types, edge types, property definitions, ontology relations, precomputed closures) is serialized to JSON
and stored in `typegraph_schema_versions`:
```sql
SELECT schema_doc FROM typegraph_schema_versions
WHERE graph_id = 'my_graph' AND is_active = TRUE;
```
From TypeScript,
[`getActiveSchema`](/schema-management#what-does-this-database-already-have)
runs this query and parses the result into a typed `SerializedSchema`.
The stored schema includes everything needed to understand the graph:
```typescript
{
graphId: "my_graph",
version: 3,
nodes: {
Person: { properties: { /* JSON Schema */ }, ... },
Company: { ... }
},
edges: {
worksAt: { fromKinds: ["Person"], toKinds: ["Company"], ... }
},
ontology: {
relations: [{ metaEdge: "subClassOf", from: "Engineer", to: "Person" }],
closures: {
subClassAncestors: { Engineer: ["Person"] },
// ... precomputed inference data
}
}
}
```
This enables:
| Capability | How It Works |
| ---------------------------- | ------------------------------------------------------------------------------------------------- |
| **Self-describing database** | Query the schema without application code—useful for debugging, admin tools, and data exploration |
| **Schema versioning** | Every schema change creates a new version; previous versions are preserved for auditing |
| **Change detection** | Compare stored schema to code schema to detect additions, removals, and breaking changes |
| **Portable exports** | The [interchange format](/interchange) is self-contained—importers know what the data means |
| **Runtime introspection** | Applications can query the schema at runtime for dynamic UI, validation, or documentation |
```typescript
import { getActiveSchema, getSchemaChanges } from "@nicia-ai/typegraph/schema";
// Query the active schema at runtime
const schema = await getActiveSchema(backend, "my_graph");
console.log("Node types:", Object.keys(schema.nodes));
console.log("Edge types:", Object.keys(schema.edges));
// Detect pending changes before deployment
const diff = await getSchemaChanges(backend, graph);
if (!diff.isBackwardsCompatible) {
console.error("Breaking changes require migration");
}
```
**Tradeoff**: Schema storage adds a small amount of database overhead (one JSON document per version). The
benefit is a database that explains itself.
## Data Model
### Storage Schema
TypeGraph uses two core tables:
```sql
-- Nodes table
CREATE TABLE typegraph_nodes (
graph_id TEXT NOT NULL,
kind TEXT NOT NULL,
id TEXT NOT NULL,
props JSON NOT NULL, -- Properties as JSON
version INTEGER NOT NULL, -- Optimistic concurrency
valid_from TEXT NOT NULL, -- Temporal validity
valid_to TEXT,
created_at TEXT NOT NULL,
updated_at TEXT NOT NULL,
deleted_at TEXT, -- Soft delete
PRIMARY KEY (graph_id, kind, id, valid_from)
);
-- Edges table
CREATE TABLE typegraph_edges (
graph_id TEXT NOT NULL,
kind TEXT NOT NULL,
id TEXT NOT NULL,
from_kind TEXT NOT NULL,
from_id TEXT NOT NULL,
to_kind TEXT NOT NULL,
to_id TEXT NOT NULL,
props JSON NOT NULL,
match_identity_name TEXT,
match_identity_key TEXT,
version INTEGER NOT NULL,
valid_from TEXT NOT NULL,
valid_to TEXT,
created_at TEXT NOT NULL,
updated_at TEXT NOT NULL,
deleted_at TEXT,
CHECK (
(match_identity_name IS NULL) = (match_identity_key IS NULL)
),
UNIQUE (graph_id, kind, match_identity_name, match_identity_key),
PRIMARY KEY (graph_id, kind, id, valid_from)
);
```
### Why JSON for Properties?
**Decision**: Store node/edge properties as JSON, not as columns.
**Why**:
1. **Schema flexibility**: Adding a property doesn't require ALTER TABLE
2. **Heterogeneous nodes**: Different node kinds have different schemas
3. **Query simplicity**: One table for all nodes, not one per kind
Both SQLite (JSON1 extension) and PostgreSQL (JSONB) have efficient JSON operators:
```sql
-- PostgreSQL
SELECT props->>'name' FROM typegraph_nodes WHERE props->>'status' = 'active';
-- SQLite
SELECT json_extract(props, '$.name') FROM typegraph_nodes WHERE json_extract(props, '$.status') = 'active';
```
**Tradeoff**: You can't create a B-tree index on a JSON property as easily as a column. For high-cardinality
filtering, consider:
- PostgreSQL: Expression indexes on JSONB paths
- SQLite: Expression indexes on `json_extract(...)` (or generated columns)
See [Indexes](/performance/indexes) for TypeGraph utilities to define and create these indexes.
### Durable edge match identity
An edge registration can promote one endpoint/property comparison from a
caller convention to a stored graph-schema contract:
```text
graph matchIdentity declaration
│
▼
canonical endpoint/property key
│
▼
authoritative edge command ──► database unique arbiter
│ │
└──────── created ◄───────┤
found ◄───────┘
```
The database unique constraint is the concurrency authority. A cache, an
outside read, or a process-local lock cannot replace it because serverless
requests may use different processes and connections. Bundled root backends can
therefore lower an eligible `getOrCreateByEndpoints()` call to one statement;
paths with cardinality claims, history, revision tracking, or caller-owned work
use the same arbiter inside their transaction.
Each decision has one owner:
| Decision | Owner |
| --- | --- |
| Which property fields form the identity | The graph registration's named `matchIdentity` |
| Whether a schema can activate or re-key it | The schema manager's all-physical-row emptiness fence |
| How endpoint/property values become a portable key | The shared canonical encoder |
| Which concurrent writer owns the identity | The database unique constraint |
| Whether a command created or found a row | The authoritative command result |
| Whether a failed import batch may retry rows | The semantic savepoint result |
The semantic savepoint is broader than raw SQL transaction control. Rolling it
back restores the database savepoint and TypeGraph's pending capture touches,
forced revisions, and graph-lock memo together. Consumers receive the resulting
decision instead of re-deriving it from an exception, row count, or follow-up
read. That is what keeps direct creates, convergence, bulk writes, import,
history, and operation hooks aligned.
### Temporal Model
Every node and edge tracks temporal validity:
```text
┌──────────────────────────────────────────────────────────────┐
│ Node: Article#123 │
├──────────────────────────────────────────────────────────────┤
│ Version 1: "Draft" │ valid_from: 2024-01-01 │
│ │ valid_to: 2024-01-15 │
├─────────────────────────┼────────────────────────────────────┤
│ Version 2: "Published" │ valid_from: 2024-01-15 │
│ │ valid_to: NULL (current) │
└─────────────────────────┴────────────────────────────────────┘
```
When you update a node:
1. The current row's `valid_to` is set to now
2. A new row is inserted with `valid_from = now`, `valid_to = NULL`
This enables:
- **Point-in-time queries**: "What did the graph look like on January 10th?"
- **Audit trails**: "What were all the versions of this article?"
- **Soft deletes**: `deleted_at` marks deletion without losing history
## Query Compilation
### The Query Pipeline
```text
Query Builder → Query AST → TypeGraph SQL Fragment → Adapter → Database
```
1. **Query Builder**: Fluent API that constructs a typed AST
2. **Query AST**: A data structure representing the query (nodes, edges, predicates, projections)
3. **SQL Generator**: Transforms the AST into TypeGraph's immutable,
database-independent SQL fragment representation
4. **Adapter**: Renders the fragment for SQLite or PostgreSQL and executes it
through the configured driver
### Common Table Expressions (CTEs)
TypeGraph compiles traversals to CTEs, which databases optimize well:
```typescript
store
.query()
.from("Person", "p")
.traverse("authored", "e")
.to("Document", "d")
.whereNode("d", (d) => d.status.eq("published"));
```
Becomes:
```sql
WITH
step_0 AS (
-- Start: all Person nodes
SELECT * FROM typegraph_nodes
WHERE graph_id = $1 AND kind = 'Person' AND deleted_at IS NULL
),
step_1 AS (
-- Traverse: follow 'authored' edges
SELECT e.*, s.id as _from_step
FROM typegraph_edges e
JOIN step_0 s ON e.from_id = s.id
WHERE e.kind = 'authored' AND e.deleted_at IS NULL
),
step_2 AS (
-- Arrive: at Document nodes
SELECT n.*, s.id as _edge_id
FROM typegraph_nodes n
JOIN step_1 s ON n.id = s.to_id
WHERE n.kind = 'Document' AND n.deleted_at IS NULL
)
SELECT
step_0.props->>'name' as person,
step_2.props->>'title' as document
FROM step_0
JOIN step_1 ON step_1._from_step = step_0.id
JOIN step_2 ON step_2._edge_id = step_1.id
WHERE step_2.props->>'status' = 'published';
```
### Recursive CTEs for Variable-Length Paths
For `recursive()` traversals with cycle prevention enabled (the default),
TypeGraph generates recursive CTEs like:
```sql
WITH RECURSIVE path AS (
-- Base case: starting nodes
SELECT id, 1 as depth, ARRAY[id] as path
FROM typegraph_nodes
WHERE kind = 'Person' AND id = $1
UNION ALL
-- Recursive case: follow edges
SELECT n.id, p.depth + 1, p.path || n.id
FROM path p
JOIN typegraph_edges e ON e.from_id = p.id
JOIN typegraph_nodes n ON n.id = e.to_id
WHERE e.kind = 'reportsTo'
AND p.depth < 10 -- Implicit cap for unbounded traversal
AND NOT n.id = ANY(p.path) -- Cycle detection
)
SELECT * FROM path;
```
When you opt into `cyclePolicy: "allow"` and do not project a path column,
TypeGraph can use a lighter recursive shape without path-array state and
cycle predicates.
## Vector Search Architecture
Semantic search with embeddings works across **all** backends — pgvector on PostgreSQL, sqlite-vec on
better-sqlite3, and libSQL/Turso's built-in vector engine. The behavior is selected by a pluggable
`VectorStrategy`, so adding a new backend is a single strategy object with no edits to the core.
### Pluggable Strategies
Each backend wires a strategy that knows how to store embeddings and compile similarity queries:
| Backend | Strategy | Storage / Index | Metrics |
| -------------- | --------------------------------------------------------- | ------------------------------------------------------------------- | ------------------------- |
| PostgreSQL | `pgvectorStrategy` (default) | typed `vector(N)` tables, HNSW / IVFFlat | cosine, l2, inner_product |
| better-sqlite3 | `sqliteVecStrategy` (when the sqlite-vec extension loads) | `vec0` virtual tables (KNN) | cosine, l2 |
| libSQL / Turso | `libsqlVectorStrategy` (wired automatically) | `F32_BLOB(N)`, DiskANN ANN via `libsql_vector_idx` + `vector_top_k` | cosine, l2 |
`createSqliteBackend` and `createPostgresBackend` accept a `vector?: VectorStrategy` option to override the
default. The strategies, `buildVectorCapabilities`, and the complete `VectorStrategy` / `VectorSlot` authoring
vocabulary are exported from `@nicia-ai/typegraph/backend`.
A backend advertises its vector support as data on `backend.capabilities.vector`:
```typescript
backend.capabilities.vector; // { supported, metrics, indexTypes, maxDimensions, ... }
```
### Storage
Embeddings are stored in **per-field typed tables**, one per `(graphId, nodeKind, fieldPath)`, each carrying
that field's fixed dimension. Tables are provisioned by the privileged migrator (`createStoreWithSchema`, and
`evolve()` for runtime-added fields), with a durable contribution marker; the runtime hot path asserts the
marker and never issues DDL. They are named `tg_vec___`. Graph-scoping the table name lets
multiple graphs in one database declare the same kind+field at different dimensions without collision.
```sql
-- PostgreSQL with pgvector: one table per (graphId, kind, field). The kind
-- and field are encoded in the table name, so rows only key by node.
CREATE TABLE tg_vec_my_graph_document_embedding (
graph_id TEXT NOT NULL,
node_id TEXT NOT NULL,
embedding vector(1536) NOT NULL, -- pgvector type, fixed dimension per field
created_at TIMESTAMPTZ NOT NULL,
updated_at TIMESTAMPTZ NOT NULL,
PRIMARY KEY (graph_id, node_id)
);
CREATE INDEX ON tg_vec_my_graph_document_embedding
USING hnsw (embedding vector_cosine_ops); -- HNSW index for fast similarity
```
`generatePostgresMigrationSQL()` runs `CREATE EXTENSION IF NOT EXISTS vector` but creates no embedding table —
the per-field tables are provisioned by `createStoreWithSchema` at boot (under the privileged role).
### Query Flow
The query API is storage-transparent and unchanged across backends:
```typescript
.whereNode("d", (d) => d.embedding.similarTo(queryVector, 10))
```
Compiles to a backend-specific nearest-neighbor query — for example, on PostgreSQL:
```sql
SELECT * FROM typegraph_nodes n
JOIN tg_vec_my_graph_document_embedding e
ON e.node_id = n.id AND e.graph_id = n.graph_id
ORDER BY e.embedding <=> $1 -- Cosine distance
LIMIT 10;
```
The backend's vector index (pgvector HNSW/IVFFlat, sqlite-vec `vec0`, or libSQL DiskANN) handles approximate
nearest neighbor search efficiently.
## Performance Characteristics
### What's Fast
- **Point lookups by ID**: O(1) with primary key index
- **Traversal frontiers**: Set-based SQL rounds with database-managed joins and de-duplication
- **Ontology expansion**: Precomputed at initialization, O(1) at query time
- **Semantic search**: ANN indexes (pgvector HNSW/IVFFlat, sqlite-vec `vec0`, libSQL DiskANN) provide sub-linear search
### What's Slower
- **Deep recursive traversals**: Recursive CTEs are more expensive than simple JOINs
- **Whole-graph algorithms**: WCC, label propagation, and PageRank iterate over every visible node
by default, or over an explicit `nodeKinds` induced subgraph, and their selected edges
- **Large property filtering without indexes**: JSON extraction is slower than column access
- **Cross-kind queries**: `includeSubClasses: true` increases the WHERE IN set
### Optimization Strategies
1. **Filter early**: Apply predicates as close to the source as possible
2. **Limit results**: Always paginate large result sets
3. **Use specific kinds**: Avoid `includeSubClasses` unless needed
4. **Index JSON paths**: For frequently-filtered properties, add expression indexes
5. **Batch writes**: Use transactions to reduce disk syncs and round-trips
### Mutation execution classes
TypeGraph classifies bulk writes by what must be known before SQL can be
submitted:
- A **closed mutation program** carries every input, fence, and refusal rule
needed for the database to decide the write. Eligible node/edge creates and
soft deletes use one exact-resource execution profile and can run as one native
atomic exchange on bundled serverless transports. The profile is attached to
the exact backend object. Derived backends do not inherit it accidentally;
an already-open PostgreSQL transaction earns a separate session-bound
registration.
- A **resolved mutation set** requires an authoritative database preimage and
application computation before its writes are known. `bulkUpsertById()` is
the canonical example: stored props are merged and Zod-validated, temporal
decisions are derived, and repeated IDs observe earlier batch items. An
eligible distinct-ID set can cross the exact registered boundary after resolution:
its guarded SQL carries the node versions or complete edge preimages that
justified the after-images. Update-only sets use one guarded set statement;
mixed create/update sets add a terminal database postimage assertion inside
the same native exchange, so an incomplete update aborts and rolls back its
creates before the transport commits. Complex sets resolve and execute inside
one interactive transaction. Neither shape is mislabeled as a read-free
program. On an interactive PostgreSQL root, the operation runs on the exact
collection-opened, caller-supplied, or adopted transaction and returns an
explicit `applied | unsupported` verdict. `unsupported` proves that no
program SQL ran before the complete portable path begins.
This boundary keeps transport optimization subordinate to Store semantics. A
new bulk optimization must either prove its mutation is closed or name the
authoritative resolution phase it preserves; it cannot read on one connection
and write on another, silently discard sidecars, or duplicate an eligibility
predicate beside the profile owner.
Backend authors can certify the transport boundary independently of mutation
eligibility with the framework-agnostic atomic transport conformance runner.
The runner supplies no dialect assumptions: the author provides statements,
state observers, and exact-root provenance checks, while the shared checks
verify ordered result slots, bound-parameter preservation, empty programs, and
all-or-nothing rollback across primary and sidecar writes.
## Why These Tradeoffs?
### Why Not a Native Graph Database?
Native graph databases (Neo4j, Amazon Neptune) excel at:
- Very deep traversals (10+ hops)
- Broad graph-data-science suites beyond the focused built-in algorithms
- Massive scale (billions of nodes)
TypeGraph is designed for:
- Knowledge bases with thousands to millions of nodes
- Shallow to medium traversals (1-5 hops typically)
- Applications that already use SQL databases
- Teams that want one database to manage
### Why a TypeGraph-Owned Backend Port?
The schema DSL, Store, query compiler, and SQL fragments belong to TypeGraph.
They do not import a database adapter's types. This keeps the public API stable
and prevents consumers from typechecking declarations for drivers and dialects
they never use.
Drizzle remains an implementation detail of the built-in SQLite and PostgreSQL
adapters:
1. **Driver integration**: Reuses mature SQLite and PostgreSQL connections
2. **Adapter-native access**: Bring-your-own-connection entrypoints retain
precise Drizzle database and transaction types
3. **Replaceable boundary**: The core depends on TypeGraph ports; adapters
translate fragments and operations at the edge
The package exports that boundary directly. Schema-only packages can import the
graph DSL and its schema-derived types from the Drizzle-free
`@nicia-ai/typegraph/core` entrypoint. Backend and search-strategy authors can
import the full Drizzle-free port vocabulary, including `GraphBackend`,
`AdapterBackend`, `DialectAdapter`, and `SqlFragment`, from
`@nicia-ai/typegraph/backend`.
### Why Zod for Schemas?
1. **Runtime validation**: Not just types, but actual validation
2. **Inference**: `z.infer` eliminates type duplication
3. **Composition**: Build complex schemas from simple ones
4. **Ecosystem**: Widely used, lots of integrations
## Next Steps
- [Performance](/performance/overview) - Benchmarks and optimization tips
- [Schemas & Stores](/schemas-stores) - Complete function signatures
- [Integration Patterns](/integration) - How to integrate with your stack
# Authoring an engine profile
> Derive a variant of a bundled SQL engine profile, and what building one from scratch still requires
[Backend Setup](/backend-setup) covers using the two bundled backends.
This page is for adapting one: changing a lock spelling, loosening a
declared capability, or swapping the resource-audit verdict without
hand-copying every other field a profile carries.
## What a profile is
A `SqlEngineProfile` is the data and dialect closures one SQL engine
contributes before any backend object exists: dialect tokens, the
execution adapter, transaction framing, DDL provisioning, strategies,
capability declarations, and an opaque `assembly` wrapping the
operation-backend builder. `createSqlBackend` is the one factory that
turns a profile into a `GraphBackend`, and it owns everything that is the
same for every engine:
- **Capability derivation** — running the shared capability tail
(atomic-batch detection, vector/fulltext capability shape,
contribution-rebuild support) over the profile's own
`declaredCapabilities`.
- **Fence resolution** — building the one write-fence target for the
whole backend and resolving its plan once, so every lock site and every
transaction-scoped handle agrees on the same decision.
- **Member assembly** — resolving the profile's `assembly` into its
operation-backend builder and late-member factory, then assembling the
mirrored member groups (contribution, identity, graph-template,
base-schema, index-materialization, kind-removal, schema-version).
- **Marks** — auditing the backend's resource shape and applying the
trust marks (root-autocommit eligibility, schema-fenced-insert
eligibility, first-party standing) that gate optimizations elsewhere.
`createPostgresBackend` and `createSqliteBackend` are each `createSqlBackend`
applied to a profile the bundled builders produce.
## The derivation path
```typescript
import {
buildPostgresEngineProfile,
createSqlBackend,
deriveEngineProfile,
} from "@nicia-ai/typegraph/adapters/drizzle/engine";
const baseProfile = buildPostgresEngineProfile(db, options);
const derivedProfile = deriveEngineProfile(baseProfile, {
// one or more of the derivable fields below
});
const backend = createSqlBackend(derivedProfile);
```
`buildPostgresEngineProfile` and `buildSqliteEngineProfile` are the
derivation base: they build a real profile against a real connection,
exactly the way `createPostgresBackend` / `createSqliteBackend` do
internally. `deriveEngineProfile(base, overrides)` returns a fresh
profile — `{...base, ...overrides}` — with `overrides` restricted to the
fields listed below. `createSqlBackend` then assembles a backend from
the result through the exact same path a bundled profile takes.
If your own module re-exports a derived profile as an inferred-typed
`const`, give it an explicit `SqlEngineProfile` annotation — the
opaque `assembly` field's internal brand is not itself exported, so
`tsc` cannot name it in a declaration file it has to infer.
## What you can override
Each field below is read directly off the profile object (or off the
`assembly`-derived context) by exactly one place in `createSqlBackend`,
with the one carve-out below — so overriding it changes the whole backend
consistently.
| Field | What overriding it changes |
| --- | --- |
| `declaredCapabilities` | The capabilities `finalizeEngineCapabilities` derives the rest of the backend's advertised capabilities from — for example, declaring `writeFence` differently changes which write-fence plan resolves. |
| `fenceSql` | The lock-statement spelling the resolved fence plan carries; pass `undefined` to remove it entirely (see [Removing `fenceSql`](#removing-fencesql) below). |
| `resourceAudit` | The serialized-resource verdict `createSqlBackend` records before the backend escapes. |
| `autocommit` | Whether a single statement outside an explicit transaction is durable — gates the root-autocommit mark. |
| `contributionRuntime` | Deps for the contribution-marker member group. |
| `identityRuntime` | Deps for the identity / recorded-relation member group. |
| `graphTemplateRuntime` | Deps for the graph-template member group. |
| `baseSchemaRuntime` | Deps for the base-schema lifecycle member group. |
| `indexMaterializationRuntime` | Deps for the index-materializations member group. |
| `kindRemovalRuntime` | Deps for the kind-removals member group. |
| `close` | The backend's `close` member. |
`DERIVABLE_ENGINE_PROFILE_KEYS` (exported alongside `DerivableEngineProfileKey`
and `DerivableEngineProfileOverrides`) is the exact set above, as an
`as const` array.
### The adapter-backed carve-out
`declaredCapabilities` and `resourceAudit` are otherwise freely derivable,
but `deriveEngineProfile` refuses an override that would change three of
their sub-fields — `declaredCapabilities.maxBindParameters`,
`declaredCapabilities.execution.interactiveTransactions`, and
`resourceAudit.kind` — away from the base profile's own value, naming the
sub-field (`ENGINE_PROFILE_OVERRIDE_UNSUPPORTED`). This check runs against
any base profile, PostgreSQL or SQLite, but it exists for the bundled
PostgreSQL builder: `buildPostgresEngineProfile` reads those exact three
sub-values to compute its execution adapter's own options before the
profile object exists, baking a copy of each into `profile.execution`,
which is not itself derivable. Deriving from a SQLite base refuses the same
override even though `buildSqliteEngineProfile`'s operation backend reads
`maxBindParameters` off the resolved capabilities directly and would honor
a changed value — the check does not distinguish the two dialects. Every
other sub-field on both objects — `writeFence`,
`windowFunctions`, `clearValidTo`, `returning`, `claims`, `graphAnalytics`,
`resourceAudit`'s `resource` / `identityLeaseResource`, and so on — stays
freely derivable.
## What you cannot override
Every other field is refused for one of these reasons: most are captured by
more than the profile's head alone, so overriding only the head would leave
`buildOperations`, `lateMembers`, or a member group they build reading the
value the base builder closed over; `dialect` and `assembly` are refused for
different reasons of their own (see the table).
| Field | Why it's refused |
| --- | --- |
| `dialect` | The operation backend literal hardcodes it. |
| `tableNames` | Captured by `buildOperations` and every transaction handle. |
| `execution` | Captured by `buildOperations` and every transaction handle. |
| `strategy` | Captured by `buildOperations` and every transaction handle. |
| `fulltext` | Captured by `buildOperations` and every transaction handle. |
| `vector` | Captured by `buildOperations` and every transaction handle. |
| `provisioning` | `ensureTable`, `catalog`, and `lineage` are all captured by migrations and transaction handles. |
| `assembly` | Opaque and bundled-only; a derived profile carries the base's `assembly` forward by reference, so it resolves to the identical `buildOperations` / `lateMembers` pair the base builder closed over. |
An override naming any of these throws `ConfigurationError` with code
`ENGINE_PROFILE_OVERRIDE_UNSUPPORTED`, naming the key, whether or not the
type would have allowed it — the check runs against the overrides
object's own keys at runtime, not only its declared type.
These refusals are `deriveEngineProfile`'s contract, not `createSqlBackend`'s.
A profile spread by hand (`{ ...base, execution: mine }`) carries the base's
`assembly` by reference, so `createSqlBackend` accepts it and applies the
override to some members while others keep the builder's value — exactly the
split the refusal exists to prevent. Derive through `deriveEngineProfile`.
## Worked example: a custom advisory-lock spelling
An engine that spells its advisory lock differently from the bundled
`pg_advisory_xact_lock(hashtext($namespace), hashtext($key))` form —
hashing one concatenated string instead of two separate arguments —
derives a `FenceSql` and passes it as an override:
```typescript
import {
buildPostgresEngineProfile,
createSqlBackend,
deriveEngineProfile,
} from "@nicia-ai/typegraph/adapters/drizzle/engine";
import { postgresFenceSql } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import type { FenceSql } from "@nicia-ai/typegraph/backend";
import { sql, type SqlFragment } from "@nicia-ai/typegraph";
function customAdvisoryLockExpression(
namespace: string,
key: string | number,
): SqlFragment {
const keyText = typeof key === "number" ? String(key) : key;
return sql`pg_advisory_xact_lock(hashtext(${namespace} || ':' || ${keyText}))`;
}
const customFenceSql: FenceSql = {
advisoryLockExpression: customAdvisoryLockExpression,
lockTables: postgresFenceSql.lockTables,
isolationFactExpression: postgresFenceSql.isolationFactExpression,
};
const baseProfile = buildPostgresEngineProfile(db, options);
const derivedProfile = deriveEngineProfile(baseProfile, {
fenceSql: customFenceSql,
});
const backend = createSqlBackend(derivedProfile);
```
`advisoryLockExpression` is the custom spelling here; `lockTables` and
`isolationFactExpression` are the bundled PostgreSQL builders, reused
because this example leaves them unchanged — a custom `FenceSql` need not
replace every member. TypeGraph derives the standalone-statement forms
every lock site actually calls (`acquireKeyed`, `acquireKeyedWithIsolation`,
`isolationFact`) from these two expressions, so `customFenceSql` never
spells a statement and its expression separately — the two cannot disagree
about what they lock or read. This is the same `customAdvisoryLockExpression`
pinned by `tests/engine-profile-derivation.test.ts` against a real
PostgreSQL connection, trimmed of the `customLockTables` /
`customIsolationFactExpression` coverage this example doesn't need.
Every write-fence lock site now spells its lock through `customFenceSql`
instead of the bundled one — including the recorded graph-write fence, which
fuses its lock into its own CTE (`buildLockSchemaVersionAndGraphWrite`) but
reads `advisoryLockExpression` / `isolationFactExpression` off the resolved
fence target rather than a hardcoded bundled spelling, so this derivation
reaches it too. The ONE exception, not reachable through `fenceSql`, is the
schema-commit fence (`acquireSchemaWriteFence` in `postgres.ts`): it emits a
standalone, single-argument `pg_advisory_xact_lock` call through
`advisoryLockSingleExpression`, baked directly into
`buildPostgresEngineProfile`'s closure. It deliberately occupies a different
lock space from every two-argument lock `fenceSql` spells, so it is not an
oversight `fenceSql` could close even if it were derivable — reaching it
needs a from-scratch profile (see
[What is not derivable yet](#what-is-not-derivable-yet)). The graph-template
instantiation statement is a different, already-reachable case: it is the
`instantiateStatement` member of `graphTemplateRuntime`, one of the fields
this same derivation can override (see the table above).
## Worked example: a portable `row`-mechanism fence
An engine with no advisory-lock primitive at all — a PostgreSQL-wire engine
with no working `pg_advisory_xact_lock` — declares `mechanism: "row"`
instead. TypeGraph spells the keyed acquisition itself against the
never-dropped fences relation, so this derivation needs no
`advisoryLockExpression` at all — only the declared mechanism and its two
facts, `drain` and `conflict`:
```typescript
const derivedProfile = deriveEngineProfile(baseProfile, {
declaredCapabilities: {
...baseProfile.declaredCapabilities,
writeFence: {
mechanism: "row",
drain: "quiescent",
conflict: "commit-time",
},
},
});
const backend = createSqlBackend(derivedProfile);
```
`conflict` states which of the two ways this engine resolves two writers of
one fence row: `"wait"` for a lock-based engine (the second acquirer's
statement blocks, exactly like `"advisory"`); `"commit-time"` for an
optimistic-concurrency engine, where both acquirers proceed and the loser's
COMMIT fails. Declaring `"commit-time"` on an interactive backend (as here)
derives `capabilities.execution.unitOfWork: "optimistic-retry"` — every
store-owned write this backend opens now replays a real commit-time conflict
as a whole unit, up to `OPTIMISTIC_RETRY_ATTEMPTS` (3) attempts, rather than
surfacing the raw driver error on the first one. `drain: "quiescent"` is the
simplest legal drain when nothing else needs a real table lock; pass
`fenceSql.lockTables` and declare `drain: "table-lock"` instead when this
engine has one. A `fenceSql.isolationFactExpression`, if this engine's wire
protocol supports reading it, rides the SAME acquisition statement — pass
`postgresFenceSql.isolationFactExpression` (or a custom one) as `fenceSql` to
keep recorded capture and match-key convergence trusting a real fact instead
of failing closed on an unknown one.
### Declaring a `serializationFailure` classifier
`isSerializationFailure` (the one predicate every retry owner consults)
recognizes PostgreSQL's own `40001` / `40P01` SQLSTATEs and their fixed
driver-message fallback. An engine whose commit-conflict shape is something
else entirely — a custom error class, a different code — declares
`execution.serializationFailure` so the SAME predicate recognizes it instead
of falling through to a raw, unretried failure:
```typescript
const profile = buildPostgresEngineProfile(db, options);
const backend = createSqlBackend({
...profile,
execution: {
...profile.execution,
serializationFailure: (error) =>
error instanceof Error && error.message.includes("CONFLICT_ON_COMMIT"),
},
});
```
This is deliberately NOT a `deriveEngineProfile` override: `execution` is
captured by `buildOperations` and every transaction handle (see
[What you cannot override](#what-you-cannot-override) below), so
`deriveEngineProfile` refuses it like every other field in that table.
Hand-spreading `execution` this way is safe for reading
`serializationFailure` itself, because `createSqlBackend` is the only reader
of `profile.execution.serializationFailure` — it registers the classifier
against the exact backend object it is about to return, once, at
construction — while every other `execution` member (`compile`, `execute`,
`runExclusive`, and so on) rides forward as the SAME function reference the
base builder closed over, spread unchanged. `createSqlBackend` consults the
registered classifier for `isSerializationFailure` calls that pass this
backend (or one of its transactions) as `target`; every store-owned unit
routed through `runRetriedUnit`, and `store.transaction`'s own retry, already
does.
That safety is narrow, and it does not extend to the profile object itself.
`{...profile, execution: {...}}` is a plain object literal — a different
object from the one `buildPostgresEngineProfile` returned — so
`isFirstPartyProfile` no longer recognizes it. `createSqlBackend` gates
every `markFirstPartyFactory` call on that check, so a hand-spread profile
loses standing to two optimizations, silently and with no functional
difference to catch in testing: the dialect-derivation write-fence fallback
(moot here, since the spread profile still carries `writeFence` declared)
and the lazy schema-fence lease
(`withTransactionSchemaFenceLease`, `src/store/operations/write-transaction.ts`),
which falls back to the conservative per-call fence instead. Accept that
trade for a one-off `serializationFailure` override; a backend meant to keep
first-party standing declares `serializationFailure` inside the builder
function that constructs `profile` in the first place, rather than spreading
the finished object afterward.
## Removing `fenceSql`
`fenceSql` is the one field a derived profile can clear: pass
`fenceSql: undefined` to drop the bundled spelling entirely. That alone
is not enough to reach a working profile — `createSqlBackend` still
resolves a write-fence plan eagerly, and a profile whose resolved
`writeFence.mechanism` is still `"advisory"` with no `fenceSql` refuses
with `WRITE_FENCE_SQL_UNAVAILABLE`. Pair it with a `declaredCapabilities`
override that stops claiming `"advisory"` (for example, declaring
`writeFence: { mechanism: "engine-serialized" }`
instead — no `drain` key: that field applies only to `mechanism: "advisory"`)
to actually resolve an `engine-serialized` plan that needs no
lock spelling at all.
## Supplying `lineage`
`EngineProvisioning.lineage` forwards onto the assembled backend's optional
`lineage` member unchanged, exactly like `provisioning.catalog` forwards onto
`catalog`. Neither bundled profile sets it: `buildPostgresEngineProfile` and
`buildSqliteEngineProfile` both leave it `undefined`, so a store built on a
bundled backend derives its `lineage` from its own recorded relations when
`history: true` is on, and has none otherwise (see
[Lineage and pruned diffs](/graph-merge#lineage-and-pruned-diffs)). An engine
whose storage layer already tracks a whole-database revision and can answer
"what changed in this graph since revision R" more cheaply than a full scan
supplies `lineage` directly.
Both `revision` and `changesSince` take a **session** as their first
argument. Run each read on that session (`session.execute` or
`session.executeRaw`); a connection captured by the strategy may see a
different snapshot. A caller planning outside a transaction passes the root
backend. A backend that also needs lineage inside transactions must thread
it through `EngineProvisioning.lineage` so transaction handles expose the
same capability.
`revision()` returns an opaque token comparable only with revisions from the
same lineage source. `changesSince()` must report every changed node and edge
key, including inserts, updates, deletes, and resurrection. Return
`{ kind: "unbounded" }` when the delta cannot be bounded. TypeGraph uses
lineage to prune branch diffs when it has a TypeGraph-owned revision anchor;
for stores without revision tracking, `base@V` uses a complete content
fingerprint regardless of engine lineage. That fingerprint covers current
identity assertions and is recomputed inside the target commit transaction.
Previously minted `engine:` base tokens are retired.
Test a new `lineage` against `tests/backends/integration/lineage-conformance.ts`'s
`registerLineageConformanceIntegrationTests` (registered per-dialect through
`createIntegrationTestSuite`, or called directly against your own backend,
via `{ getStore: () => ({ backend }) }`). It registers two describes: only
"lineage: recorded-relations conformance" is portable — it drives every case
through `resolveLineage`, the same path a real caller takes, and is the case
the bundled recorded-relations derivation passes: after N writes,
`changesSince(r0)` is exactly the touched keys, `changesSince(rN)` is empty, a
hard delete after a revision reports the deleted key once, and an unrecognized
revision is `unbounded`. "lineage: capture-completeness evidence" is
TypeGraph-specific — it exercises `recordedRelationsLineage` directly (the
per-revision evidence a bare engine revision has no equivalent gap for); an
engine profile's own suite should run against the conformance describe only
and skip the other.
## Supplying `recordedTime`
`EngineProvisioning.recordedTime` declares an engine that tracks recorded
(system) time itself, rather than through TypeGraph's own capture relations
and clock — a backend that declares it must also declare `lineage`
(engine-native history keeps no recorded relations for TypeGraph to derive a
change delta from; `createSqlBackend` refuses `recordedTime` without a
co-declared `lineage` with `ENGINE_PROFILE_RECORDED_TIME_REQUIRES_LINEAGE`).
`EngineRecordedTimeMembers` has two members, `source` and `revisionNow`, both
`this: void`.
`source(table, revision)` names the table expression `table` (`"nodes"` |
`"edges"` | `"identityAssertions"`) reads its recorded rows from, AS OF
`revision` — the engine's own temporal-table syntax, with the interval
already folded in (a system-time `AS OF` clause, or equivalent). It replaces
what a TypeGraph-relation-backed source spells as two members: the recorded
relation itself (`recordedNodesTable`/`recordedEdgesTable`) and a separate
`recorded_from <= r AND r < recorded_to` interval predicate. Because
`source`'s own expression already scopes every row to exactly one revision,
there is nothing left for a predicate to narrow — every recorded read this
member backs compiles with no interval clause at all. `revision` is an
opaque `{ revision, recordedAt }` pair minted by your own `revisionNow`
below; never parse `revision.revision` as a number; embed it in the AS OF
expression as an opaque token. `table` is never called with
`"identityAssertions"` today — a recorded identity read (`Store.
identityAtCoordinate` at a past instant, and the query compiler's historical
identity traversal) is refused outright under engine-native ownership before
any read compiles, so your implementation must still accept the shared
union without that arm ever running.
`revisionNow(session)` is called on two different kinds of session, and must
answer differently for each:
- **On a root backend** (`store.recordedNow()`, `store.revisionNow()`): the
engine's current COMMITTED revision.
- **On an open `transaction()` handle** (both places `TransactionReceipt.recorded`
is stamped, called before that transaction's own COMMIT): the revision at
which THIS transaction's writes will become visible once it commits — the
engine's pending/next revision for that session, not the last one committed
before it opened. TypeGraph stamps this still-uncommitted value straight
into the receipt it hands back to the caller once the transaction succeeds.
An engine that can only name its last-COMMITTED revision, never its own
pending one from inside an open transaction, cannot implement `recordedTime`:
stamping the last-committed value into a receipt would describe the state
*before* the write the receipt is reporting on, and there is no correct
point after COMMIT to read the right value from without reopening the race
`recordedTime` exists to close.
Declaring `recordedTime` changes what `history: true` means on your profile.
TypeGraph's own recorded relations, clock, and write-fence-gated clock
allocation are never engaged; `revisionTracking: true` is refused whether or
not `history: true` is also requested
(`ENGINE_NATIVE_REVISION_TRACKING_UNSUPPORTED` — there is no
TypeGraph clock for it to advance, and the engine's own revision is only
ever available under `history: true`); a `recordedRead` external binding is
refused (`ENGINE_NATIVE_RECORDED_READ_UNSUPPORTED` — there is no TypeGraph
recorded relation for one to populate); and `migrateLegacyRecordedTime`
refuses outright (`ENGINE_NATIVE_MIGRATE_RECORDED_TIME_UNSUPPORTED` — it
rewrites TypeGraph's own recorded relations, which your backend does not
have). `RecordedInstant` anchors from a `recordedTime`-declaring store use
the `e1::` form rather than TypeGraph's
`r1:<16-digit revision>:` form; `asOfRecorded` refuses an
anchor minted under the other ownership form with
`RECORDED_INSTANT_OWNERSHIP_MISMATCH`. See [Engine-native recorded
time](/queries/temporal#engine-native-recorded-time) for the full reader
contract.
## Refusals you may meet
| Code | When |
| --- | --- |
| `ENGINE_PROFILE_REQUIRES_WRITE_FENCE_DECLARATION` | The profile's resolved capabilities omit `writeFence` — `createSqlBackend` has no write-fence decision to resolve and refuses outright, naming the one capabilities line to add. |
| `WRITE_FENCE_SQL_UNAVAILABLE` | The resolved capabilities declare `mechanism: "advisory"` but the profile's `fenceSql` is missing the member that mechanism/drain combination needs; or `mechanism: "row"` with `drain: "table-lock"` but no `fenceSql.lockTables`. A `"row"` target missing `tableNames.fences` is NOT refused here — it refuses the first time a keyed site actually acquires the fence row. |
| `WRITE_FENCE_DECLARATION_INVALID` | The declared `writeFence` carries an unrecognized `mechanism`, `drain`, or `conflict` string; a `drain` key on a mechanism other than `"advisory"` / `"row"`; a `conflict` key on anything but `"row"`; or `conflict: "commit-time"` on a target whose own `capabilities.execution.interactiveTransactions` is `false` — that value is honored only by the `"optimistic-retry"` execution tier, which never derives without an interactive transaction to replay inside, so accepting it there would silently drop it rather than apply it. `resolveWriteFencePlan` validates the raw value (a plain-JavaScript author is not held to the discriminated-union type) before shaping a plan from it. |
| `CALLER_SERIALIZED_REFUSES_ADOPTION` | `adoptTransaction` was called on a backend whose resolved write-fence plan is `caller-serialized` — an externally owned transaction's lifetime cannot be held by the backend's in-process write-unit queue. |
| `CATALOG_UNAVAILABLE` | A store path that needs the backend's catalog probes (index materialization, the recorded-time schema check, the recorded-time migration's column read) finds `catalog` absent — a profile whose `provisioning.catalog` is unset builds a backend with no `catalog` member at all. |
| `LINEAGE_UNAVAILABLE` | A caller reached `requireLineage` and found `lineage` absent on the backend it asked. Every OUT-OF-TRANSACTION graph-merge caller consults `lineage` through `resolveLineage`, which already falls back to the recorded-relations lineage or to a full comparison rather than hitting this refusal. `assertTargetUnchanged`'s in-transaction re-validation reads the transaction handle's `lineage` ONLY — no fallback to the root — so this fires whenever a `lineage` that anchored the plan (found on the root at plan time) is not ALSO threaded onto the transaction handle that commits it; see "Supplying `lineage`" above for how to thread it correctly. |
| `ENGINE_PROFILE_RECORDED_TIME_REQUIRES_LINEAGE` | The profile declares `recordedTime` without also declaring `lineage` — engine-native history keeps no recorded relations of its own for TypeGraph to derive a graph-merge change delta from, so the engine's own `lineage` is the only source for one. Raised at `createSqlBackend` construction, naming both members. |
| `RECORDED_TIME_UNAVAILABLE` | A caller reached `requireRecordedTime` and found `recordedTime` absent on the backend it asked. Store construction under `history: true` and the shared `recordedNow()`/`revisionNow()`/receipt-stamping read are the only callers today, both reached only once `recordedTimeOwnership` has already resolved to `"engine-native"`, so this is defense-in-depth rather than a reachable misconfiguration on a bundled backend. |
| `ENGINE_NATIVE_REVISION_TRACKING_UNSUPPORTED` | A store was constructed with `revisionTracking: true` against a backend that declares `recordedTime`, whether or not `history: true` was also requested — engine-native has no TypeGraph clock for `revisionTracking` to advance on its own. |
| `ENGINE_NATIVE_RECORDED_READ_UNSUPPORTED` | A store was constructed with an external `recordedRead` binding against a backend that declares `recordedTime` — there is no TypeGraph recorded relation for one to populate; engine-native's own recorded reads are sourced from `recordedTime.source` instead. |
| `ENGINE_NATIVE_RECORDED_IDENTITY_UNSUPPORTED` | `Store.identityAtCoordinate` at a past recorded instant, or the query compiler's historical identity traversal, was reached under engine-native recorded time — identity history reads TypeGraph's own recorded relations directly, which an engine-native backend does not populate. |
| `ENGINE_NATIVE_MIGRATE_RECORDED_TIME_UNSUPPORTED` | `migrateLegacyRecordedTime` was called against a backend that declares `recordedTime` — the migration rewrites TypeGraph's own recorded relations, which an engine-native backend does not have. |
| `RECORDED_INSTANT_OWNERSHIP_MISMATCH` | `asOfRecorded(instant)` was called with an instant minted under the OTHER recorded-time ownership form — an `r1:` instant against an engine-native store, or an `e1:` instant against a TypeGraph-owned one. |
| `ENGINE_PROFILE_OVERRIDE_UNSUPPORTED` | `deriveEngineProfile`'s `overrides` names a key outside the derivable set, or one of the three adapter-backed sub-fields with a changed value (see [the carve-out](#the-adapter-backed-carve-out)). |
| `ENGINE_ASSEMBLY_UNRECOGNIZED` | The profile's `assembly` is not a value `assembleEngine` produced — a profile built by hand rather than obtained from a bundled builder (optionally adapted with `deriveEngineProfile`). |
## What non-first-party costs
First-party standing is bound to the exact profile object one of the two
bundled builders returned, not to a field — a derived profile is a new
object neither builder ever saw, so it never carries that standing
forward, even when every field is copied from a first-party profile
unchanged. That costs a derived profile's backend two things:
- **No dialect-derivation fallback.** `resolveWriteFencePlan`'s fallback
for a profile with no `writeFence` declared is sound only for the two
bundled dialects, so it never applies to a
derived profile regardless — irrelevant in practice as long as
`declaredCapabilities` is kept, since both bundled declarations already
carry a write-fence declaration explicitly.
- **No lazy per-transaction schema-fence lease.** The lease
`store/operations/write-transaction.ts` takes out under
`isFirstPartyFactory` is closed to a derived profile's backend; each
managed write takes its own fence instead.
Every gate `createSqlBackend` runs — the write-fence-declaration refusal, the
`mechanism: "advisory"` without `fenceSql` refusal, the schema-fenced-insert
and autocommit marks — still applies to a derived profile exactly as it does
to a bundled one.
A bundled profile object is frozen once its builder returns it: mutating a
field on that exact object throws, rather than silently drifting the
profile away from what first-party standing was granted to.
`deriveEngineProfile` is unaffected — it spreads `base`'s fields into a
new object literal, which does not freeze.
## What is not derivable yet
Building a profile from scratch — rather than deriving a variant of a
bundled one — needs an execution adapter, an operation strategy, and an
operation-backend assembly, none of which is exported today.
`SqlEngineProfile.assembly` is opaque, and its only constructor,
`assembleEngine`, is exported from no entrypoint: it is authoring a new
engine, not deriving a variant of an existing profile, and waits on a
future exported assembly constructor. Until then, derivation from a
bundled builder — changing a lock spelling, a capability declaration, a
resource-audit verdict, or a runtime dependency bag — is the supported
way to adapt a profile.
# TypeGraph vs. Neo4j, LadybugDB and pgGraph: Who Wins What
import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro";
My first pass at benchmarking TypeGraph against real graph databases ran the
seven LDBC "short read" queries (IS1–IS7) against Neo4j and LadybugDB.
TypeGraph on SQLite won every one of them, which should have made me
suspicious rather than happy: IS1–IS7 are point lookups and one-hop walks, and
an in-process engine doing direct index seeks can't lose that race to anything
that pays a network round trip.
So I rebuilt the harness around the queries a graph database is _supposed_ to
win (shortest paths, bounded neighborhood walks, complex multi-hop reads, and
whole-graph algorithms) and ran 17 queries against five engines at two
scales. The short version:
- **TypeGraph on SQLite still wins every point read, usually by 10–100x.**
- **The engines that keep an in-memory graph index win whole-graph algorithms
by three to four orders of magnitude**, and the gap gets wider as the data
grows.
I'm publishing both results rather than only the flattering one.
## The setup
All five engines run through one shared harness
([`packages/benchmarks/src/real/`](https://github.com/nicia-ai/typegraph/tree/bench/pggraph-comparison-v2/packages/benchmarks/src/real)):
| Engine | Version |
| ------------------------------ | ---------------------------------------------------------------------------- |
| SQLite | 3.53.2 (via `better-sqlite3` 12.11.1) |
| PostgreSQL (TypeGraph backend) | 18.1 (`pgvector/pgvector:pg18` image) |
| Neo4j | `neo4j:2026.05.0` server image, `neo4j-driver` 6.2.0, GDS plugin |
| LadybugDB | `@ladybugdb/core` 0.18.0 |
| pgGraph | Evokoa pgGraph 0.1.8 (`ghcr.io/evokoa/pggraph:0.1.8`, bundles PostgreSQL 17) |
pgGraph is the new entrant, and I think it's clever. Rather than being a
separate database, it's a Postgres extension that builds a derived CSR
(compressed sparse row) index over ordinary normalized tables and exposes
traversal and pathfinding as SQL functions. Its point reads are tuned
Postgres, and the CSR index only comes into play once a query traverses.
The 17 queries are the original IS1–IS7; IC13 (shortest path) and IC14
(weighted shortest path); BFS3, a bounded neighborhood walk; three complex
reads (IC2, IC8, IC9); and four graph-algorithm queries: GA_DEGREE, GA_WCC
(weakly connected components), and GA_BFS / GA_SSSP (whole-component
reachability and shortest-path depth from a seed).
I held the comparison to two rules. Every result is checked with a
value-level digest, per row, across every engine that runs the query, and
every run below passed, since a fast wrong answer shouldn't count. And nothing
is skipped silently: an engine without a comparable primitive for a query
reports a typed `gap`, so a `gap` cell means the engine can't run that query
in a comparable form, not that I didn't get around to it.
Scales are **SF1** (9,892 persons, 361K directed `knows` edges) and **SF10**
(65,645 persons, 3.88M `knows`, 21.9M comments). Each is one run on an EC2
host with shared vCPUs, so treat sub-millisecond cells and anything flagged
noisy as order-of-magnitude.
## SF1
p50 latency in milliseconds unless noted. Fastest engine per row in **bold**.
| Query | typegraph-sqlite | typegraph-postgres | neo4j | ladybugdb | pggraph |
| -------------------- | ---------------: | -----------------: | -------: | --------: | -------: |
| IS1 | **0.03** | 0.76 | 4.84 | 1.11 | 0.93 |
| IS2 | **1.85** | 21.5 | 39.9 | 76.0 | 19.9 |
| IS3 | **0.24** | 1.82 | 4.07 | 4.5 | 1.47 |
| IS4 | **0.02** | 0.98 | 3.01 | 0.39 | 0.91 |
| IS5 | **0.03** | 1.08 | 3.1 | 1.61 | 0.94 |
| IS6 | **0.07** | 1.97 | 3.23 | 3.93 | 1.89 |
| IS7 | **0.07** | 2.15 | 5.74 | 7.83 | 1.76 |
| IC13 (shortest path) | 3.88 | 22.3 | 3.24 | 9.91 | **2.18** |
| IC14 (weighted SP) | 5586 | **4860** | gap | gap | gap |
| BFS3 | **220** | 1139 | 468 | 1539 | 357 |
| IC2 | 51.8 | 523 | **36.2** | 89.1 | 347 |
| IC8 | **3.06** | 17.3 | 3.73 | 24.2 | 8 |
| IC9 | 2236 | 15069 | 3388 | **775** | 3498 |
| GA_DEGREE | **0.03** | 1.09 | 2.97 | 1.8 | 0.72 |
| GA_WCC | 7269 | 24237 | 18.7 | gap | **9.53** |
| GA_BFS | 221 | 1828 | **24.5** | gap | 288 |
| GA_SSSP | 220 | 1742 | **23.6** | gap | 284 |
### Point reads
TypeGraph on SQLite takes every IS row, often by one to two orders of
magnitude, because it's the only engine here that pays no network round trip
and no per-call query planning overhead. GA_DEGREE and
IC8 go the same way, because underneath the "algorithm" and "complex read"
labels they're point lookups too.
This is the case for an embedded graph. Most of what an application asks its
graph all day looks like IS1–IS7 (fetch this person, their recent posts,
who they know), and on IS1 Neo4j takes 4.84ms where SQLite takes 0.03ms.
### Graph algorithms
GA_WCC, GA_BFS, GA_SSSP, and IC13 are what a graph engine's specialized index
exists for, and here the CSR engines are in a different league:
- **pgGraph wins GA_WCC (9.53ms) and IC13 (2.18ms)**, running union-find and
shortest path directly over its CSR index.
- **Neo4j with the GDS plugin wins GA_BFS (24.5ms) and GA_SSSP (23.6ms).**
Without GDS, Neo4j answers these with Cypher path enumeration, which works
but is slow. With GDS it projects the graph into memory once and runs the
same kind of set-based traversal pgGraph does.
- **TypeGraph is three to four orders of magnitude slower on GA_WCC** (7,269ms
on SQLite, 24,237ms on Postgres, against pgGraph's 9.53ms).
That last gap isn't a bug I can fix. TypeGraph's algorithms run as rounds of
SQL, one window or aggregate query per round, however well indexed, while
pgGraph and GDS hold the graph in an in-memory structure built for this access
pattern, and query tuning won't make SQL iteration behave like a CSR
traversal.
## SF10
The SF10 run happened **before** I added the GDS plugin to the Neo4j setup, so
Neo4j's graph-algorithm rows show `gap` here even though Neo4j can run them.
For Neo4j on those queries, the SF1 table above is the current picture.
`s` = seconds; otherwise milliseconds.
| Query | typegraph-sqlite | typegraph-postgres | neo4j | ladybugdb | pggraph |
| -------------------- | ---------------: | -----------------: | ------: | --------: | -------: |
| IS1 | **0.03** | 1.06 | 5.22 | 1.12 | 0.93 |
| IS2 | **2.35** | 77.2 | 133 | 104 | 22.6 |
| IS3 | **0.42** | 6.23 | 53.8 | 15.7 | 1.87 |
| IS4 | **0.03** | 0.91 | 3.33 | 0.90 | 0.96 |
| IS5 | **0.04** | 8.67 | 5.52 | 2.68 | 0.96 |
| IS6 | **0.07** | 3.69 | 8.09 | 5.50 | 2.02 |
| IS7 | **0.07** | 7.29 | 8.23 | 8.38 | 1.80 |
| IC13 (shortest path) | 35.0 | 339 | 82.2 | 161 | **2.25** |
| IC14 (weighted SP) | 74.6s | **57.3s** | gap | gap | gap |
| BFS3 | **1.7s** | 7.2s | 3.9s | 9.3s | 2.2s |
| IC2 | 141 | 724 | **138** | 296 | 843 |
| IC8 | **5.24** | 220 | 92.1 | 47.6 | 15.3 |
| IC9 | 11.9s | 76.7s | 15.5s | **3.1s** | 25.4s |
| GA_DEGREE | **0.06** | 1.05 | 3.02 | 2.07 | 0.82 |
| GA_WCC | 119.9s | 506.0s | gap | gap | **69.0** |
| GA_BFS | 2.8s | 22.3s | gap | gap | **2.0s** |
| GA_SSSP | 2.8s | 22.3s | gap | gap | **1.9s** |
Point reads barely move at 10x the data, as you'd expect from an indexed
seek. The more interesting number is how GA_WCC grows:
| Engine | SF1 | SF10 | Growth |
| ------------------ | ------: | -----: | -----: |
| pggraph | 9.53ms | 69.0ms | ~7x |
| typegraph-sqlite | 7269ms | 119.9s | ~16x |
| typegraph-postgres | 24237ms | 506.0s | ~21x |
pgGraph grows a bit less than linearly with the data, while TypeGraph's
round-by-round algorithm grows faster than linearly because more edges also
means more rounds, so the gap gets bigger as the data grows. If you need
whole-graph analytics over millions of edges on a schedule, use a specialized
engine for that job.
### SQLite beats Postgres, even inside TypeGraph
Both TypeGraph backends run the same logical algorithms, and SQLite is 4–8x
faster on the heavy ones at SF10 (GA_WCC 119.9s vs 506.0s, GA_BFS 2.8s vs
22.3s, IC9 11.9s vs 76.7s). My first guess was network round trips, but that
doesn't hold up: GA_BFS already issues one `INSERT ... RETURNING` per round,
Postgres is on loopback, and the ratio stays about 8x at both scales, whereas
a fixed per-call cost would shrink as a share of a longer run. That points at
a per-row cost in how Postgres executes these queries, which I haven't tracked
down yet.
The one exception is IC14, the weighted shortest path, where Postgres wins.
Its set-based frontier expansion handles a large, unbounded Dijkstra better
than SQLite's row-at-a-time version.
## IC14: the one only TypeGraph ran
IC14 asks for the lowest-cost path between two people rather than the fewest
hops. In this lineup, nothing else could run it in a comparable form:
Neo4j's GDS has no stored `knows` weight to project, pgGraph's shortest path
counts hops only, and LadybugDB's weighted shortest path isn't wired into the
harness yet. TypeGraph answers it on both backends with
`store.algorithms.weightedShortestPath`, at real LDBC scale (5.6s / 4.9s at
SF1, 75s / 57s at SF10), with byte-identical results across the two.
Those times aren't fast, and the benchmark makes them look worse than real
use would, because it picks random pairs near the graph's diameter, which is
the worst case for single-source Dijkstra. Real weighted-path questions tend to be between
related, nearby entities, where the same algorithm stops early.
## Loading
Loading is where TypeGraph still trails. It improved a lot between my first
run and this one, because in the meantime 0.37 shipped a trusted initial
import, and the benchmark loader now uses it.
`importGraph` validates every row it writes: schema shape, edge endpoints,
cardinality, conflicts. That's the right default, since most imports come
from somewhere you don't fully trust. But when you're filling a brand-new
database from an export you produced and already validated, every one of
those checks is wasted work. `trustedImportGraph` and
`trustedImportGraphStream` skip them. They bypass the normal write pipeline,
drop secondary indexes, insert straight into the tables, then rebuild the
indexes and refresh statistics, all in one transaction.
Loading the same 200,000 nodes and 200,000 edges into a fresh SQLite
database both ways:
```text
run 1 — importGraph: 6783ms trustedImportGraphStream: 2455ms
run 2 — importGraph: 7253ms trustedImportGraphStream: 2206ms
run 3 — importGraph: 5048ms trustedImportGraphStream: 2116ms
```
Trusted import was 2.5–3x faster on every run. The streaming form takes a header, then node
chunks, then edge chunks, so the loader here reads the LDBC CSVs in two
bounded passes instead of holding a multi-million-row graph in memory.
Skipping validation needs a narrow contract. The node and edge tables must
be completely empty, and TypeGraph only checks stream order and kind names.
Property shapes, endpoints, and duplicate-free IDs are on you. Features whose
extra writes it would otherwise skip (history, uniqueness constraints,
`searchable()` and `embedding()` fields) are rejected outright. Point it at
a database with rows in it and it throws before touching anything:
```text
TrustedImportError: Trusted import requires globally empty TypeGraph node and edge tables.
details: { tables: ["typegraph_nodes", "typegraph_edges"], reason: "database_not_empty" }
```
For anything that isn't a one-time load of a fresh, dedicated database, use
`importGraph` (untrusted data, conflicts, history, search fields) or
collection `bulkInsert` (trusted data going into a database that isn't
empty).
Here's what it did to the benchmark:
| Engine | SF1 load | Earlier run | SF10 load |
| ------------------ | -------: | ----------: | --------: |
| ladybugdb | 46s | 41.6s | 373s |
| neo4j | 66s | 71.4s | 409s |
| pggraph | 120s | — | 1,119s |
| typegraph-sqlite | 164s | 587.1s | 2,199s |
| typegraph-postgres | 344s | 669.3s | 3,278s |
TypeGraph on SQLite went from 587s to 164s (~3.6x) and on Postgres from 669s
to 344s (~1.9x). SQLite is now within about 1.4x of pgGraph, while
Postgres is still about 3x slower than pgGraph at both scales, because pgGraph
loads through batched `INSERT`s tuned for its own schema and TypeGraph's
Postgres path still uses prepared statements instead of `COPY`. Switching to
`COPY` is next, and unlike most of what this benchmark turned up, it would
help every Postgres user.
## pgGraph as an accelerator
pgGraph is the most interesting result here, less for the rows it wins than
for how it works. It indexes tables that already look a lot like the ones
TypeGraph's Postgres backend writes (normalized nodes and edges in Postgres,
queryable over the same connection), which is why its point reads look like
tuned Postgres while its traversal numbers look like a graph engine's.
**None of this is built yet.** The pgGraph driver in this benchmark loads its own copy
of the data into its own schema. It doesn't sit on top of a TypeGraph-managed
database. But the numbers suggest the shape of a pairing: use TypeGraph for
the point reads and everyday writes it already wins, and when profiling finds
a real whole-graph workload (components, centrality, unweighted shortest path
at scale), build a pgGraph index over the same tables and send just that
query through it. I'm excited about that direction, because it keeps
everything in one database without giving up anything on the queries
applications run most.
## What the slow rows actually mean
It's tempting to read every slow cell as a TypeGraph bug. Mostly they aren't:
- **IC9 is a modeling artifact.** LDBC models a message's creator as an edge,
so ranking a feed means fetching a sort key for every candidate first: 1.2M
of them at SF1 to return the top 20. A real feed would put the owner on the
item and index `(owner, created_at)`, which `defineNodeIndex` already
supports, so the fix belongs in the schema rather than the engine.
- **The whole-graph algorithms are architectural.** Tuning has helped (a
delta-frontier rewrite of connected components nearly halved the Postgres
GA_WCC time during this work), but it won't close a 1,000x gap. Pairing with
something like pgGraph for those queries is the realistic answer.
- **IC14 is benchmark-amplified**, as above.
And the caveats: one run per scale on a shared-vCPU host, so several cells
are noisy and should be read as order-of-magnitude. What I'm confident in
regardless: every query passes value-level parity across every engine that
runs it, and the direction of every finding is far larger than run-to-run
noise.
## Try it
- [The harness](https://github.com/nicia-ai/typegraph/tree/bench/pggraph-comparison-v2/packages/benchmarks/src/real):
all five engine drivers and the EC2 runner
- Full results and investigation notes:
[`sf1-results.md`](https://github.com/nicia-ai/typegraph/blob/bench/pggraph-comparison-v2/packages/benchmarks/reports/sf1-results.md),
[`sf10-results.md`](https://github.com/nicia-ai/typegraph/blob/bench/pggraph-comparison-v2/packages/benchmarks/reports/sf10-results.md)
- [Trusted initial import](/interchange#trusted-initial-import): the full
contract
- [TypeGraph 0.35: Faster Almost Everywhere](/blog/typegraph-0-35-performance):
the fixes an earlier run of this benchmark turned up
# Bring Your Own Database
import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro";
The first versions of TypeGraph depended on Drizzle and assumed the database
underneath was either a `pg` pool or `better-sqlite3`. That held up until
people started running it on Cloudflare D1, inside Durable Objects, over
Neon's HTTP driver, in PGlite, and on engines that speak the Postgres wire
protocol but handle locking very differently from Postgres.
Each of those broke an assumption somewhere, usually as a SQL error from deep
inside a query the engine couldn't run, and over three releases (0.38, 0.51,
and 0.57) I've been removing those assumptions. This post covers all three:
how Drizzle became optional, how a backend can now declare what it can't do
before a query fails, and how an engine that locks differently can describe
that to TypeGraph.
## Drizzle is an adapter now
Before 0.38, every `Store` carried Drizzle's types whether your code
touched them or not. A strict TypeScript project that only ever called
`store.nodes.Person.create(...)` still had to resolve Drizzle's dialect
declarations to typecheck, which is a lot of ORM to pull into your type
checking just to create a node.
Now the portable `Store` has no Drizzle in it. The managed factories own
the connection and hand you a complete store:
```typescript
import { createLocalSqliteStore } from "@nicia-ai/typegraph/sqlite/local";
const store = await createLocalSqliteStore(graph, { path: "./graph.db" });
const alice = await store.nodes.Person.create({ name: "Alice" });
```
```text
created: LBPSjEoGPqI0P3C6kaJ5M Alice
capabilities.execution.interactiveTransactions: true
```
`createLocalPgliteStore` does the same for Postgres-in-WASM. The full graph
API is there, `store.transaction(...)` included.
When you do want to write your own tables on the same connection, you opt in
with `createAdapterStore`, and `tx.sql` hands you the native transaction:
```typescript
await store.transaction(async (tx) => {
await tx.nodes.Document.update(documentId, props);
if (tx.sqlAvailability !== "available") {
throw new Error(`Native transaction unavailable: ${tx.sqlAvailability}`);
}
await tx.sql.insert(documentVersions).values(versionRow);
});
```
`sqlAvailability` is a discriminant because raw SQL is sometimes off on
purpose: on a history-enabled store, writing around TypeGraph would skip
recorded-time capture, so `tx.sql` isn't available there.
## Drizzle is now an optional dependency
In 0.51 `drizzle-orm` became an optional peer dependency, and ten
entrypoints don't need it installed at all: the root, `backend`, `core`,
`schema`, `indexes`, `graph-extension`, `interchange`, `profiler`,
`graph-merge`, and `provenance`.
That list is enforced by tests: a fixture imports all ten with `drizzle-orm`
missing from `node_modules`, and source and build-output checks fail if an
import ever drags it back in. The last three routes to Drizzle
(recorded-time migration DDL, a claim comparison, and some removal-statement
builders) moved to portable code with golden tests pinning byte-identical SQL
on both dialects.
If you use a managed store or a `/adapters/drizzle/...` entrypoint you still
need Drizzle, and if your package manager skips optional peers you'll need to
install it yourself. The managed factories tell you so with a typed error
that includes the `npm install` command, instead of a module-resolution stack
trace.
## Backends say what they can't do
Removing the dependency was the easier part. The harder problem is that
TypeGraph emits SQL some engines can't run, and it used to find that out in
production, halfway through a query.
Recursive CTEs are the clearest case, since variable-length traversals,
subgraph extraction, and a few identity reads depend on them. An engine
without them can now say so:
```typescript
const capabilities: Partial = {
recursiveTraversal: {
supported: false,
reason: "engine has no WITH RECURSIVE / equivalent",
},
};
```
With that declared, the operations that need recursion throw a
`ConfigurationError` naming the operation, instead of sending the engine SQL
it can't parse. `weightedShortestPath` falls back to walking the path one hop
at a time, which returns the same answer in more round trips. Leaving the
capability out means it's supported, so every existing custom backend keeps
working as it did.
## How your engine keeps writers apart
Some writes (Operational Identity, and the recorded-time clock behind
`history` and `revisionTracking`) need exactly one writer per graph at a
time. TypeGraph used to pick the lock by checking which dialect it was
talking to, which went badly if your engine said "postgres" but had no
advisory locks.
Now the backend declares how it keeps writers apart, which for the two bundled
engines is one line each:
```typescript
// PostgreSQL
writeFence: { mechanism: "advisory", drain: "table-lock" }
// SQLite
writeFence: { mechanism: "engine-serialized" }
```
A custom backend that hosts identity or recorded history without declaring
one is rejected at construction, and the error message prints the line to add
for your dialect.
0.57 added two more mechanisms. `caller-serialized` is your promise that
nothing else writes to the database; TypeGraph enforces the in-process half
and trusts you with the rest. `row` is for engines with no advisory locks at
all: TypeGraph takes the lock by upserting a row in its own
`typegraph_fences` table. If your engine settles write conflicts at commit
time rather than blocking, declare `conflict: "commit-time"` and TypeGraph
retries its own transactions as a whole unit when they lose, up to three
attempts.
## Engine profiles
The piece I'm happiest about is that 0.57 made the two bundled backends
_data_. `createPostgresBackend` and `createSqliteBackend` are now the same
function, `createSqlBackend`, applied to an engine profile: the dialect, how
it executes, how it provisions, and what it declares. That means you can take
a bundled profile, change what's different about your engine, and get a real
backend:
```ts
import {
buildSqliteEngineProfile,
createSqlBackend,
deriveEngineProfile,
} from "@nicia-ai/typegraph/adapters/drizzle/engine";
import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local";
const { db } = createLocalSqliteBackend();
const base = buildSqliteEngineProfile(db);
const derived = deriveEngineProfile(base, {
declaredCapabilities: {
...base.declaredCapabilities,
writeFence: {
mechanism: "row",
drain: "quiescent",
conflict: "commit-time",
},
},
});
const backend = createSqlBackend(derived);
console.log(backend.capabilities.execution?.unitOfWork);
// "optimistic-retry"
```
(It's SQLite here only because that's what runs on a laptop.) You never set
`unitOfWork` yourself; TypeGraph works it out from what the engine declared
rather than guessing from the engine's name.
Derivation is narrow on purpose. You can override declared capabilities, the
lock SQL, and a handful of runtime hooks, but not `dialect` or `execution`,
because the bundled builders capture those in more than one place, and I'd
rather throw at construction than hand you a backend that's half one engine
and half another.
## One error for "you lost the race"
A Postgres serialization failure or deadlock used to surface as whatever the
driver felt like throwing. Now every conflict comes back as
`TransactionConflictError`, with the driver error as `cause`, and
`store.transaction()` can retry for you:
```ts
// SQLite never raises 40001; this stands in for what PostgreSQL would.
function serializationFailure(): Error {
return Object.assign(new Error("could not serialize access"), {
code: "40001",
});
}
let attempts = 0;
await store.transaction(
async (tx) => {
attempts += 1;
await tx.nodes.Account.create({ owner: "ada", balance: 100 });
if (attempts < 3) throw serializationFailure();
},
{ retry: { attempts: 3 } },
);
console.log(attempts, (await store.nodes.Account.find()).length);
// 3 1 — two rolled-back attempts left nothing behind
```
The callback reruns from the top, so it has to be safe to run more than once:
read inside it, don't reuse values from outside it, and don't cause side
effects that escape the transaction.
## Branches your host can copy
`branch()` normally copies a graph by streaming it into a fresh database.
0.57 also adds `forkedWorkingCopyStrategy`, which hands the copy to whatever
your host is good at, such as a file copy, `CREATE DATABASE ... TEMPLATE`, or
a provider's branch API. Because the fork is a physical copy of the database,
it keeps things a streamed copy can't, including recorded history, so a forked
branch can answer `asOfRecorded` queries from before the fork was taken.
## What isn't there yet
You can't build an engine from scratch yet. `deriveEngineProfile` adapts one
of the two bundled profiles, and building a profile from nothing needs pieces
that aren't exported. No third engine has been run through this in production
yet either: the commit-time retry path is tested by injecting conflicts into
real SQLite and PGlite transactions rather than against an engine that
produces them on its own. If you have one, I'd like to hear how it goes.
## Upgrading
From 0.37, the Drizzle-specific entrypoints moved under `/adapters/drizzle`
and the old paths are gone rather than aliased:
| 0.37 | 0.38 and later |
| ------------------ | ----------------------------------- |
| `/sqlite` | `/adapters/drizzle/sqlite` |
| `/sqlite/local` | `/adapters/drizzle/sqlite/local` |
| `/sqlite/libsql` | `/adapters/drizzle/sqlite/libsql` |
| `/postgres` | `/adapters/drizzle/postgres` |
| `/postgres/pglite` | `/adapters/drizzle/postgres/pglite` |
`/sqlite/local` and `/postgres/pglite` now mean the managed store factories.
If you read `tx.sql`, switch to `createAdapterStore` /
`createAdapterStoreWithSchema`; if you don't, nothing changes.
Custom backend authors: `capabilities.pessimisticLocks` is now
`capabilities.writeFence`, and conflicts should be matched as
`TransactionConflictError` rather than by SQLSTATE. The
[0.57.0 changelog](/changelog#0570) has the full list.
## Try it
- [Backend Setup](/backend-setup#drizzle-free-entrypoints): entrypoints,
capabilities, and the parity matrix
- [Write fence declaration](/backend-setup#write-fence-declaration-writefence):
every mechanism and what each error means
- [Authoring an engine profile](/backend-authoring): the derivable fields and
worked examples
- [GitHub](https://github.com/nicia-ai/typegraph)
# Merges That Wait
import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro";
[Graph Merge](/blog/graph-merge) plans before it applies: you fork a working
copy with `branch()`, stage changes on it, and `planMerge()` shows you the exact
write set before anything touches the target. That works well as long as
staging, planning and applying all happen in one run of one process.
The workflows that most want a review step don't look like that. Say a nightly
job proposes loyalty-point adjustments that a person approves the next
morning, a deploy lands in between, and the thing feeding the branch is a queue
that occasionally delivers the same message twice. An in-memory branch handle
and a plan that goes stale as soon as anything is written can't cope with any
of that.
It took three releases to fix, and I think the result is one of the more
unusual things TypeGraph can do. 0.56 made the review itself durable by storing
it as graph data, 0.67 made the working copy durable so a branch can be closed
in one process and reopened in another, and 0.68 made it safe to feed that
branch from an at-least-once queue.
## The review lives in the graph
A merge plan is tied to the target's revision, so if the target changes after
planning, applying the plan fails. That rule is what keeps a stale plan from
clobbering newer data, but it gets in the way if you want the review (the
proposal, the decision, who approved it) stored as graph data next to the
record it concerns, because _writing the review down_ is itself a write to the
target, and by the time anyone approves the plan it's stale.
0.56 resolves this by separating what was reviewed from the plan you
eventually apply. `planCandidateWriteSetReview()` captures the candidate
changes, the plan, the policy, and a baseline of the target as one immutable,
content-digested artifact:
```typescript
const review = unwrap(
await planCandidateWriteSetReview({
target: store,
makeBackend,
policy: {
id: "manual-acceptance-v1",
context: { requiredApprovals: 1, authorizedReviewers: ["reviewer:maya"] },
},
writeSet: {
formatVersion: 1,
sourceId: "catalog-review",
target: await captureCandidateWriteSetTarget(store),
nodes: [
{
kind: "Item",
id: proposal.id,
properties: { label: proposal.label, status: "accepted" },
},
],
edges: [],
},
}),
);
```
You store it as ordinary graph data, keyed by its own digest, along with the
reviewer's decision. `Artifact`, `Decision` and `evidence` here are ordinary
kinds the example defines itself, so this needs no special schema support:
```typescript
const artifact = await store.nodes.Artifact.create(
{ content: JSON.stringify(review) },
{ id: review.digest.value },
);
await store.edges.evidence.create(proposal, artifact, { note: "review" });
const decision = await store.nodes.Decision.create({
approved: true,
reviewDigest: review.digest.value,
reviewer: "reviewer:maya",
});
await store.edges.evidence.create(decision, artifact, { note: "approval" });
```
Applying the original plan at this point fails:
```typescript
const stale = await applyMergePlan(store, review.plan);
// stale.error is a StaleMergePlanError
```
Recording the review and the approval moved the target, so the plan is stale.
This is the part I like: durable review doesn't get an exemption from the
staleness check, because if recording an approval could quietly un-stale a
plan, the check would mean nothing.
Instead of reusing the plan, you call `revalidateCandidateWriteSetReview()`,
which reads the stored artifact, re-plans the retained candidate against the
current target, and tells you what it found:
| `checked.status` | What it means |
| ---------------- | ------------------------------------------------------------------------------------------------------------------ |
| `compatible` | Nothing that matters moved. You get a fresh `plan` and the original `reviewDigest`. |
| `changed` | Policy, options, baseline entities, or plan fields differ from what was reviewed. Get a new review and approval. |
| `incompatible` | Graph id, schema identity, or revision origin don't match. This approval can't be used against this target at all. |
In the example only the review and approval records were added, so the result
is `compatible`. I was careful to keep `compatible` meaning only that nothing
relevant moved; it says nothing about whether the caller is allowed to act, so
the example checks both before spending the plan:
```typescript
if (checked.status !== "compatible") {
throw new Error("A new review and approval are required");
}
if (checked.reviewDigest.value !== approval.reviewDigest) {
throw new Error("Approval does not identify the validated review");
}
const report = unwrap(await applyMergePlan(store, checked.plan));
```
Authenticating the stored decision and enforcing `authorizedReviewers` are
still your job. TypeGraph tells you whether the plan is safe to apply and
leaves the question of who may apply it to you.
## A working copy that outlives its process
The branch itself was still tied to one process, because a `branch()` result
holds its store and close handle in memory and disposing it deletes the fork.
0.67 adds `branchDurable()`, which forks a working copy that persists and hands
back a small JSON descriptor you can put on a queue. Here's the loyalty
ledger's nightly job:
```typescript
const created = unwrap(await branchDurable(base, host));
const staged = created.branch.store;
await staged.nodes.Account.update(ada.id, { points: 160 });
await staged.nodes.Account.create({ name: "Grace", points: 40 });
// Releases this process's connection and writer lease. The working copy stays.
await created.branch.close();
await queue.put(JSON.stringify(created.descriptor));
```
The descriptor is the only thing that leaves the process:
```json
{
"kind": "sqlite-file-host",
"version": 1,
"graphId": "loyalty",
"definitionHash": "ec9dd68d2fbf14e3",
"branchId": "gluQHM58QmB1LFjZIJEdA",
"base": "ec9dd68d2fbf14e3#s1\u0000revision:kHeqC6EE55ENVL3a_np2R:r1:0000000000000001:2026-09-20T19:16:21.492Z",
"store": { "id": "35da40a3-1820-4b9f-b9b4-2117e653ded8" },
"schemaAnchor": { "version": 1, "hash": "ec9dd68d2fbf14e3" }
}
```
TypeGraph doesn't ship a durable host. It defines the contract (a
`DurableWorkingCopyStrategy` with `create`, `seal`, `reopen`, `abort` and
`destroy`), and your host decides where a working copy lives, whether that's a
directory, a database, or a provider's branch API. The descriptor's `store`
field is the host's opaque locator, and everything else in it belongs to
TypeGraph. The examples here ran against a small file-backed SQLite host,
across separate `node` processes.
The next day a different process reads the descriptor and reopens the branch,
and what comes back is an ordinary `GraphBranch`, so everything after that is
the merge API you already know:
```typescript
const descriptor = JSON.parse(await queue.get());
const branch = unwrap(await reopenDurableBranch(graph, descriptor, host));
const plan = unwrap(await planMerge(base, [branch]));
const report = unwrap(
await applyDurableMergePlan({
target: base,
branch,
descriptor,
strategy: host,
plan,
}),
);
await branch.close();
unwrap(await destroyDurableBranch(descriptor, host));
```
```text
plan: 2 node upserts, 0 conflicts
merged: {"nodes":2,"edges":0,"identity":{"asserted":0,"retracted":0}}
base now: Grace=40, Ada=160
reopen after destroy refused: true
```
`applyDurableMergePlan()` applies the approved plan through the target Store
transaction. The former optional host-native merge hook was retired because a
database merge that commits internally can cross the transaction boundary
that checked the target revision. A host may still use native database
branches for its durable working copies.
A descriptor is a document your application stored and handed back later, so
TypeGraph treats it as untrusted input. At fork time the host seals the true
origin, and every reopen and destroy is checked against it, so relabeling the
branch id makes the reopen fail:
```text
Durable branch descriptor does not match the working copy the host attested for
its store locator: the descriptor's TypeGraph fences disagree with the origin
recorded at fork. This is a tampered, relabeled, or wrong-branch descriptor.
```
The same check stops you from destroying branch B with branch A's descriptor,
or reopening with a graph that reuses the id `"loyalty"` but defines
`Account` differently.
## The message that arrives twice
Suppose the branch is fed from a queue where "award Ada 25 points" can arrive
twice. The award has to apply exactly once, and whoever is downstream (a
notification, an audit log) has to hear about every applied award eventually,
even if the worker dies right after committing.
What that calls for is a transactional outbox scoped to the working copy, and
0.68 builds one into the durable-branch contract. You hand
`operateDurableBranch()` an idempotency key, a `mutation` describing the
change, and `metadata` to keep as evidence:
```typescript
const outcome = unwrap(
await operateDurableBranch(descriptor, host, {
idempotencyKey: "award-7731",
metadata: { source: "orders-queue", messageId: 7731 },
mutation: { op: "award", account: adaId, points: 25 },
}),
);
```
TypeGraph never interprets `mutation`; it digests it together with `metadata`,
hands the host the request, and validates what comes back. The host applies
the change and writes an evidence row in one database transaction, which in
this host is a single SQLite transaction on one connection. To check the
rollback, I made it throw after the graph write and before the evidence insert:
```text
DurableOperationError | GRAPH_MERGE_OPERATION | Durable operation failed: injected failure after the graph write
evidence: undefined
accounts: Grace=40, Ada=160
```
There's no evidence row, and Ada is still at the staged 160.
In the real run, the worker commits the award (Ada goes from 160 to 185) and
then dies before telling anyone. A fresh process picks up the descriptor, and
the queue redelivers the same message:
```text
recover pid 76873 | Ada on branch: 185
redelivered: replayed | Ada on branch: 185
```
`replayed` returns the evidence from the first run and applies nothing, so Ada
stays at 185 instead of 210. If you reuse the key with a different payload,
even just different metadata, the host rejects it without writing:
```text
changed payload: DurableOperationConflictError | GRAPH_MERGE_OPERATION_CONFLICT
Ada on branch: 185
```
The evidence rows act as the outbox. Each one starts with `delivered: false`,
and you can't destroy the branch while any are undelivered:
```text
destroy: DurableEvidenceUndeliveredError | GRAPH_MERGE_OPERATION_UNDELIVERED |
refusing to destroy "60367e1c-…": undelivered operation evidence remains
```
To deliver them, you scan for undelivered rows and mark each one after
publishing it:
```typescript
const page = unwrap(
await scanDurableOperations(descriptor, host, { limit: 100 }),
);
for (const evidence of page.operations) {
if (evidence.delivered) continue;
await publishDownstream(evidence); // your outbox consumer
unwrap(
await markDurableOperationDelivered(
descriptor,
host,
evidence.idempotencyKey,
),
);
}
unwrap(await destroyDurableBranch(descriptor, host));
```
A crash at any point in that sequence means, at worst, that a downstream
consumer hears about an award twice; the award itself is never applied twice
or silently dropped.
## Limits
- **There's no first-party host and no queue.** Whether your host's
transaction is atomic is up to your database. TypeGraph validates what the
host reports back but can't check your storage.
- **Exclusion is the host's job.** The example host's writer lease is a lock
file. When I left one behind, as a killed process would, reopening failed
until it was cleared. A real host wants a lease that expires.
- **Plan against a quiet branch.** Feeding operations into a branch while a
reviewer plans against it means planning against a moving target.
- **Delivery is at-least-once.** Make downstream writes idempotent on the key.
- **You probably don't need this for a one-shot job.** If you stage, plan and
apply in one run, `branch()` is unchanged and simpler.
## Try it
- [Durable candidate review](/graph-merge#durable-candidate-review-in-the-target-graph)
- [Durable host-native branches](/graph-merge#durable-host-native-branches)
and [atomic operations and evidence](/graph-merge#atomic-operations-and-immutable-evidence)
- [Example 27](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/27-durable-merge-review.ts):
the durable review, end to end
- [GitHub](https://github.com/nicia-ai/typegraph)
# The Cheapest Citation Lineage Isn't the Shortest One
import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro";
When I introduced TypeGraph I said there was no PageRank and no community
detection, and that if you needed those you wanted a real graph database.
That's still partly true (more on that at the end), but as of 0.38 the list
is a lot shorter.
Until now, `store.algorithms` answered questions about two nodes at a time,
like how to get from A to B or what's within three hops of A. Some questions
need the whole graph at once: whether a dataset is one connected body or
several islands, which nodes matter most structurally rather than by raw
count, or whether communities emerge from the topology on their own.
Answering those means running an algorithm round after round over the entire
graph, from one consistent snapshot, until it converges. That needs more
machinery than a single traversal, including a pinned transaction, a temporary
working table, and a clear rule for what happens when the rounds don't
settle. 0.37 added
`weaklyConnectedComponents` and `weightedShortestPath`. 0.38 added global and
personalized `pageRank` and deterministic `labelPropagation`, the same
algorithm the LDBC Graphalytics benchmark uses for community detection. All
five run as SQL against the store, so you don't export the graph to a separate
analytics engine and keep a second copy in sync.
They're also a lot of fun to play with, so the rest of this post runs them on
a small citation graph.
## The corpus
This is the same citation graph as the
[research-copilot example](/examples/research-copilot): 18 landmark ML
papers, 55 authors, 14 topics, and 37 real citation edges, from Rumelhart,
Hinton & Williams' 1986 backprop paper through LLaMA in 2023. That example
runs point queries; here I run the whole-graph algorithms over the same data.
```typescript
const backend = createExampleBackend();
const [store] = await createStoreWithSchema(graph, backend);
// Ingested 18 papers, 55 authors, 14 topics, 37 citation edges.
```
## One body of work, or islands?
Treat `cites` as undirected and partition by connectivity:
```typescript
const components = await store.algorithms.weaklyConnectedComponents({
edges: ["cites"],
nodeKinds: ["Paper"],
});
```
```text
18 papers partition into 1 component(s):
• component of 18 paper(s), rooted at "Adam: A Method for Stochastic Optimization"
```
Every paper reaches every other through some chain of citations. On a real,
messy dataset this is the sanity check to run first. `nodeKinds: ["Paper"]`
keeps authors and topics out of it, so more than one component would mean a
separate sub-literature rather than a lightly cited author hanging
off the edge.
## The cheapest lineage isn't the shortest one
This is my favorite result in the post, and getting to it took one false
start.
Citations always point from a newer paper to an older one, so the obvious edge
weight is `yearGap`, the number of years a citation reaches back. The trouble
is that every hop steps strictly backward in time, so the gaps along _any_
route from A to B telescope to exactly `A.year - B.year`. Every path ties, and
weighting by `yearGap` is just `shortestPath` with extra arithmetic.
Squaring the gap breaks the tie. `yearGapCost = yearGap²` is convex, so one
27-year leap costs `27² = 729` while the same span covered in several small
steps costs much less. The question becomes "what's the smoothest chain of
ideas between these two papers," and that's where `weightedShortestPath`
starts disagreeing with `shortestPath`:
```typescript
const hopPath = await store.algorithms.shortestPath(from.id, to.id, {
edges: ["cites"],
maxHops: 10,
});
const byConvexCost = await store.algorithms.weightedShortestPath(
from.id,
to.id,
{ edges: ["cites"], weightProperty: "yearGapCost" },
);
```
```text
transformer → backprop:
shortestPath (fewest hop): 2 hops transformer(2017) → adam(2014) → backprop(1986)
weighted by yearGapCost: 4 hops totalWeight=353 transformer(2017) → dropout(2014) → alexnet(2012) → lenet(1998) → backprop(1986) ◀── more hops, lower convex cost
clip → backprop:
shortestPath (fewest hop): 3 hops clip(2021) → simclr(2020) → dropout(2014) → backprop(1986)
weighted by yearGapCost: 7 hops totalWeight=359 clip(2021) → gpt2(2019) → bert(2018) → transformer(2017) → dropout(2014) → alexnet(2012) → lenet(1998) → backprop(1986) ◀── more hops, lower convex cost
```
The fewest-hops route from the Transformer paper to backprop makes a 28-year
jump through Adam. The convex-cost route takes four smaller steps (Dropout,
AlexNet, LeNet) and comes in at less than half the cost. From CLIP it walks
seven hops, almost straight down the history of deep learning. Both are
legitimate answers to "what's the best path" between the same two papers, and
I like that the convex-cost one reads like a syllabus.
Weights are checked before any traversal round runs, so a negative or
non-numeric `yearGapCost` anywhere in the selected edges throws
`InvalidEdgeWeightError` up front instead of producing a wrong answer.
## PageRank vs. counting citations
A citation count tells you how many papers cite this one. PageRank tells you
how much a paper matters given _who_ cites it: a citation from an important
paper is worth more, and that carries through the graph.
```typescript
const pageRankScores = await store.algorithms.pageRank({
edges: ["cites"],
nodeKinds: ["Paper"],
direction: "out", // random surfer follows citations forward, toward the classics
});
```
```text
PR-rank score cites raw-rank Δ title
────────────────────────────────────────────────────────────────
1 0.24209 6 1 · Learning representations by back-propagating errors
2 0.09060 3 7 +5 ImageNet Classification with Deep Convolutional N...
3 0.07559 3 6 +3 Efficient Estimation of Word Representations in V...
4 0.07424 4 2 -2 Attention Is All You Need
5 0.05868 4 3 -2 BERT: Pre-training of Deep Bidirectional Transfor...
6 0.05827 1 13 +7 Gradient-Based Learning Applied to Document Recog...
7 0.05818 3 5 -2 Dropout: A Simple Way to Prevent Neural Networks ...
8 0.05592 3 4 -4 Deep Residual Learning for Image Recognition
```
Backprop wins both rankings, which is no surprise since it's the root of the
whole corpus. The row I find interesting is 6th place. LeNet has exactly
**one** citation in this corpus, which puts it 13th by count, but PageRank
moves it up seven places because that one citation comes from AlexNet, which
is heavily cited itself. A plain count would treat it like any other
citation.
## What matters to CLIP, specifically
Personalized PageRank runs the same iteration, but instead of jumping to a
random node it keeps jumping back to a seed you choose. The question changes
from "important globally" to "important from where CLIP is standing":
```typescript
const personalized = await store.algorithms.personalizedPageRank({
edges: ["cites"],
nodeKinds: ["Paper"],
direction: "out",
seeds: [{ id: clip.id, kind: "Paper" }],
});
```
```text
PPR-rank score global-rank Δ title
──────────────────────────────────────────────────────────────────
1 0.25884 17 +16 Learning Transferable Visual Models From Natural ...
2 0.12804 1 -1 Learning representations by back-propagating errors
3 0.09396 5 +2 BERT: Pre-training of Deep Bidirectional Transfor...
4 0.07889 4 · Attention Is All You Need
5 0.06382 3 -2 Efficient Estimation of Word Representations in V...
6 0.05500 9 +3 Language Models are Unsupervised Multitask Learners
7 0.05500 15 +8 A Simple Framework for Contrastive Learning of Vi...
8 0.05500 16 +8 An Image is Worth 16x16 Words: Transformers for I...
```
CLIP itself jumps from 17th to 1st, since every jump lands back on it, and
SimCLR and ViT, both cited directly by CLIP and both well outside the global
top 10, climb eight places each. Changing only the seed gives you a ranking
for a different question, and it's the one I'd reach for in a "related work"
or recommendation feature.
## Do research communities fall out?
Label propagation finds communities by having every node adopt the most
common label among its neighbors, round after round, over the undirected
version of `cites`. Run it strictly first:
```typescript
const converged = await store.algorithms.labelPropagation({
edges: ["cites"],
nodeKinds: ["Paper"],
onMaxIterations: "throw", // default
});
```
```text
onMaxIterations: "throw" raised GraphAlgorithmConvergenceError —
the undirected citation graph oscillates (tree / even-cycle structure
that mirrors labels back and forth).
```
That error is expected. In synchronous label propagation a node doesn't vote
for itself, so a tree-shaped neighborhood (common once you flatten a citation
DAG into an undirected graph) can flip two labelings back and forth forever,
and more iterations won't fix it. The default throws rather than handing you
whatever labels the last round happened to land on. If you want that
fixed-round answer, which is what the LDBC Graphalytics benchmark specifies,
ask for it:
```typescript
const fixedRound = await store.algorithms.labelPropagation({
edges: ["cites"],
nodeKinds: ["Paper"],
onMaxIterations: "return",
});
```
```text
3 communities:
── community of 7 ──
Adam: A Method for Stochastic Optimization [Optimization]
Learning representations by back-propagating errors [Optimization, DeepLearning]
Dropout: A Simple Way to Prevent Neural Networks ... [DeepLearning, Optimization]
Gradient-Based Learning Applied to Document Recog... [CNN, ComputerVision]
Sequence to Sequence Learning with Neural Networks [RNN, NLP, DeepLearning]
Very Deep Convolutional Networks for Large-Scale ... [CNN, ComputerVision]
Efficient Estimation of Word Representations in V... [Embeddings, NLP]
── community of 7 ──
BERT: Pre-training of Deep Bidirectional Transfor... [Transformer, NLP, SelfSupervised]
Learning Transferable Visual Models From Natural ... [Contrastive, MultiModal, ComputerVision]
Chain-of-Thought Prompting Elicits Reasoning in L... [LanguageModel, Reasoning, NLP]
Language Models are Unsupervised Multitask Learners [Transformer, NLP, LanguageModel]
LLaMA: Open and Efficient Foundation Language Models [Transformer, LanguageModel, NLP]
Attention Is All You Need [Transformer, Attention, NLP]
An Image is Worth 16x16 Words: Transformers for I... [Transformer, ComputerVision, DeepLearning]
── community of 4 ──
ImageNet Classification with Deep Convolutional N... [CNN, ComputerVision, DeepLearning]
Momentum Contrast for Unsupervised Visual Represe... [Contrastive, SelfSupervised, ComputerVision]
Deep Residual Learning for Image Recognition [CNN, ComputerVision, DeepLearning]
A Simple Framework for Contrastive Learning of Vi... [Contrastive, SelfSupervised, ComputerVision]
```
The algorithm only saw undirected `cites` edges (the topic tags are printed
for you and were never fed to it), yet it separated the optimization and
classic-vision foundations, the transformer and language-model era, and the
contrastive self-supervised vision cluster purely from who cites whom.
## What's still missing
Shortest path (weighted and unweighted), reachability, neighborhoods,
degree, connected components, label propagation, and global and personalized
PageRank cover a lot of ground, but strongly connected
components, topological sort, betweenness/closeness/eigenvector centrality,
and Louvain/Leiden community detection aren't in `store.algorithms`. For
those, pull the edge list out with `.query().traverse()` or
`store.subgraph()` and hand it to an in-memory library like
[graphology](https://graphology.github.io/).
Scale matters too. These run as rounds of SQL, which keeps everything in one
database and is fine for graphs the size most applications have, but on
millions of edges the engines that hold the graph in a specialized in-memory
index are orders of magnitude faster at whole-graph work. The
[benchmark post](/blog/benchmarking-typegraph-neo4j-ladybugdb) has the
numbers, including the unflattering ones.
## Try it
- [Graph Algorithms](/graph-algorithms): every algorithm, shared options,
temporal behavior, and PageRank tolerance notes across backends
- [Research Copilot](/examples/research-copilot): the same corpus, run
through the point-query algorithms
- [Example 32](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/32-graph-analytics.ts):
the runnable source behind this post
- [GitHub](https://github.com/nicia-ai/typegraph)
# Merging Two Feeds That Disagree About the Same Patient
import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro";
Point two ingestion agents at overlapping data (an EHR export and a claims
feed, say) and tell them to "just write everything to the graph", and you end
up with two patient nodes for one person, each holding half the care history
and neither aware of the other. Most pipelines then add a nightly dedupe job
and hope nothing reads the graph in between.
I think append is the wrong default for graphs, which is why 0.31 ships
`@nicia-ai/typegraph/graph-merge`. It's the feature I've been most eager to get
into people's hands. You branch a store, let each writer work on its own copy,
and then fold the branches back in: entities are resolved, edges are repointed
onto the surviving nodes, disagreements are reported instead of silently
overwritten, and the merge records who contributed what.
## Branch, write, merge
`branch()` records the base store's current state and hands back an
isolated working copy. Writers use the ordinary store API against it.
`merge()` diffs every branch against the base and runs one pipeline to fold
them back in:
```text
stage (diff every branch)
→ generate candidates (exact unique · blocking key · similarity)
→ cluster (group nodes that are the same entity)
→ canonicalize (pick a survivor, union properties, resolve conflicts)
→ repoint + dedupe edges onto survivors
→ reconcile delete/modify and types
→ commit transactionally + build the report
```
The whole thing is deterministic: clusters resolve by stable keys and
conflicts are decided by an explicit `branchOrder` rather than by which branch
happened to arrive first, so merging the same branches in any order commits
the same graph.
## Two ways to be the same patient
[Example 18](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/18-fhir-graph-merge.ts)
runs this on a small FHIR-flavored care graph. An EHR branch and a claims
branch each record the same two patients, and each gets both identities
wrong in a different way:
- **Anna Rivera** (EHR) and **Ana Rivera** (claims) share `MRN-001`, and
because `mrn` is declared `unique`, they're matched regardless of how the
name is spelled, without any similarity threshold.
- **Mohammed Ali** (EHR, `MRN-204`) and **Mohamed Ali** (claims, `MRN-205`)
have _different_ MRNs, so nothing forces them together. They share a birth
date, which puts them in the same blocking bucket, and they collapse because
their fulltext name similarity clears the configured `0.78` threshold.
Landing in the same bucket only means two records get compared. The test
suite also covers the other side of the threshold with **Zoe Adams** and
**Quinn Webb**, who share a birth date and land in the same bucket but score
near zero on name similarity, so they stay two separate patients.
```typescript
const mergeOptions: MergeOptions = {
resolve: {
Patient: {
block: (node) => node.birthDate,
similarity: { kind: "fulltext", fields: ["name"] },
threshold: 0.78,
},
},
onPropertyConflict: "flag",
branchOrder: [EHR_BRANCH, CLAIMS_BRANCH],
provenance: true,
};
const result = await merge(base, [ehr, claims], mergeOptions);
```
Both pairs collapse to one canonical patient each:
```text
merged nodes: 9
merged edges: 10
entity resolutions: 2
```
That's six nodes staged on the EHR branch plus five on claims, minus the two
patients that folded into their counterparts, and all ten edges survive
without duplicates.
## Disagreements get flagged, not hidden
Merge doesn't quietly pick a value when branches disagree. The properties
that don't agree are flagged, and the flags also show which identity mechanism
was in play:
```text
conflicts:
- Patient.fhirId on patient-ana: claims-agent="Patient/claims-ana", ehr-agent="Patient/ehr-anna"
- Patient.name on patient-ana: claims-agent="Ana Rivera", ehr-agent="Anna Rivera"
- Patient.fhirId on patient-mohamed: claims-agent="Patient/claims-mohamed", ehr-agent="Patient/ehr-mohammed"
- Patient.mrn on patient-mohamed: claims-agent="MRN-205", ehr-agent="MRN-204"
- Patient.name on patient-mohamed: claims-agent="Mohamed Ali", ehr-agent="Mohammed Ali"
```
There's no `mrn` conflict for `patient-ana`, because Anna and Ana share
`MRN-001` exactly. The Mohammed/Mohamed pair does flag `mrn`, since their
merge was based on name similarity and never required the MRNs to agree.
## Every edge lands on the survivor
This is the part I like best: everything attached to either duplicate ends
up on the one canonical patient. Here it is read back through each survivor's
`forPatient` edges rather than a table scan:
```text
Ana Rivera (MRN-001, 1974-03-09)
- Encounter: Hypertension follow-up (2026-04-11T09:30:00-07:00)
- Encounter: Kidney function review (2026-04-14T10:00:00-07:00)
- MedicationRequest: Lisinopril 10 MG Oral Tablet - Take one tablet by mouth daily
- Observation: Blood pressure panel = 152/96 mmHg (high)
- Observation: Estimated glomerular filtration rate = 54 mL/min/1.73m2 (low)
Mohamed Ali (MRN-205, 1990-08-21)
- Encounter: Cardiology consult (2026-05-02T13:00:00-07:00)
- Observation: LDL cholesterol = 168 mg/dL (high)
```
Ana Rivera's five resources came from both branches (the EHR encounter and
medication, the claims encounter and lab result), and they now hang off a
single patient node that neither branch created on its own, so anyone
reading this record sees the full history. If an edge had been repointed
wrongly, it would be missing from the list.
## Who contributed what
With `provenance: true`, the merge reports which branch contributed each
committed node and edge:
```text
provenance:
- ehr-agent: 6 node(s), 6 edge(s)
- claims-agent: 5 node(s), 4 edge(s)
```
That report lives only as long as the call. Pass `persistProvenance: true`
and each contribution is also written as a durable
`{branch, sourceId} → canonical` row in a separate provenance graph on the
same backend, so "what did this provider ever contribute?" is a query you
can run next month, without adding anything to your domain schema.
## Merging into a graph that kept moving
`merge()` is a snapshot operation: every branch must fork from the target's
_current_ state, or it fails with `BaseVersionMismatchError` rather than risk
clobbering newer data. That works when you fork, write, and merge in one
round, but real ingestion keeps going, and
[Example 19](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/19-incremental-merge.ts)
covers folding new batches into a target that has already moved on.
In that example a company knowledge base already has `Acme Corp` (`acme.com`),
and a provider batch reports the same company under another spelling along
with one company that's new:
```text
Target before: [ 'Acme Corp (acme.com)' ]
Target after: [ 'Acme Corp (acme.com)', 'Globex (globex.io)' ]
No duplicate was created: the provider's "ACME Corporation" merged onto the
committed "Acme Corp" via the shared domain.
```
`mergeIncremental()` finds the already-committed row by its unique `domain`
and merges onto it instead of creating a duplicate. It's the same mechanism
as the shared-MRN case, matched against live data instead of another branch:
```typescript
const result = await mergeIncremental({
forkPoint, // the frozen ancestor the branch forked from
target, // the live committed graph, which may have advanced
branches: [provider],
options: {
resolve: {
Company: {
similarity: { kind: "fulltext", fields: ["name"] },
threshold: 0.9,
},
},
onPropertyConflict: "flag",
onBasePropertyConflict: "flag", // required: never let a stale branch value win
persistProvenance: true,
},
});
```
`onBasePropertyConflict: "flag"` is required so that a stale branch can't
overwrite something newer than the point it forked from. If the live target
changed the same row after the fork, the target's value wins and the
disagreement is reported.
## Try it
- [Graph Merge](/graph-merge): entity resolution, blocking, similarity
strategies, conflicts, scaling guards, and determinism
- [FHIR Graph Merge](/examples/fhir-graph-merge) and
[Incremental Graph Merge](/examples/incremental-merge): the docs
walkthroughs
- [Example 18](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/18-fhir-graph-merge.ts)
and
[Example 19](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/19-incremental-merge.ts):
the runnable source behind this post
- [GitHub](https://github.com/nicia-ai/typegraph)
# Embeddings Can't Find a SKU
import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro";
Type `PROD-1005-E` into a search box and you expect one result: the product
with that SKU. Vector search is bad at this in a way tuning won't fix: a rare
alphanumeric code barely moves an embedding, so cosine similarity ranks on
everything _else_ in the query, and the product whose exact code you typed
ends up under a pile of vaguely similar ones.
BM25 handles it easily, because it counts terms and a token that appears in
exactly one document wins however odd it looks. Rather than trying to make
embeddings better at exact matches, I wanted both signals fused in one place,
so 0.21 added native fulltext search and hybrid retrieval that combines it
with vector search, and 0.24 finished the job on SQLite.
## Fulltext is a field modifier
Mark a field `searchable()` and TypeGraph keeps a BM25 index in sync on every
write. On Postgres that's `tsvector` + GIN; on SQLite it's FTS5. You don't
need an Elasticsearch cluster for this, or a sync job to feed one.
```typescript
const Product = defineNode("Product", {
schema: z.object({
name: searchable({ language: "english" }),
description: searchable({ language: "english" }),
sku: searchable({ language: "english" }),
category: z.enum(["outerwear", "footwear", "accessories", "climbing"]),
embedding: embedding(16).optional(),
}),
});
```
Query it with `store.search.fulltext()`, or use `$fulltext.matches()` as a
predicate inside an ordinary query, where it combines with metadata filters
and traversals in the same SQL statement:
```typescript
const activeOuterwear = await store
.query()
.from("Product", "p")
.whereNode("p", (p) =>
p.$fulltext
.matches("lightweight", 10)
.and(p.status.eq("active"))
.and(p.category.eq("outerwear")),
)
.select((ctx) => ({ sku: ctx.p.sku, name: ctx.p.name }))
.execute();
```
## Hybrid, fused with RRF
`store.search.hybrid()` runs the vector search and the fulltext search and
merges them with Reciprocal Rank Fusion. RRF only looks at rank positions, so
it doesn't care that a cosine score and a BM25 score live on completely
different scales, which makes it the least fiddly fusion method I know of.
```typescript
const hybridHits = await store.search.hybrid("Product", {
limit: 5,
vector: { fieldPath: "embedding", queryEmbedding, metric: "cosine", k: 20 },
fulltext: { query: "waterproof shell", k: 20, includeSnippets: true },
fusion: { method: "rrf", k: 60, weights: { vector: 1, fulltext: 1.25 } },
});
```
Fulltext worked on both backends from 0.21, but hybrid needs the backend to
run a vector search, which SQLite couldn't do yet, so on SQLite the hybrid call
threw `ConfigurationError`. 0.24 gives SQLite a real vector search on top of
`sqlite-vec`, built to the same shape as the Postgres version, and the hybrid
call now takes the same options and returns the same results on both.
SQLite also turned out to be the faster of the two. On the project's search
benchmark (500 documents, 384 dimensions), SQLite hybrid runs in **0.8ms**
against Postgres's 2.5ms, partly because SQLite runs in-process and doesn't
pay for a round trip.
## Running it
[Example 15](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/15-fulltext-hybrid-search.ts)
seeds nine outdoor-gear products (parkas, shells, a climbing harness, ski
goggles), each with a searchable name, description, and SKU, plus a small
embedding. A BM25 query for `"waterproof jacket"` finds the Expedition Parka,
and the snippet shows why:
```text
Heavily insulated jacket for alpine expeditions and extreme
cold. Waterproof outer shell with down fill.
```
Then the SKU:
```text
Query: "PROD-1005-E" (looking up an exact SKU)
1. [PROD-1005-E] Climbing Harness Pro score=1.6328
```
It comes back as the only result, where a pure vector search on that string
ranks unrelated products higher because the SKU barely registers in the
embedding.
Hybrid is where it gets interesting. Here's `"waterproof shell"` with `k=20` on
each side, RRF `k=60`, and fulltext weighted at `1.25`:
```text
1. [PROD-1001-A] Expedition Parka score=0.0357 (v#3, f#3)
2. [PROD-1002-B] Arctic Shell score=0.0356 (v#6, f#1)
3. [PROD-9901-Z] Legacy Rain Shell score=0.0355 (v#5, f#2)
4. [PROD-1007-G] Compression Socks score=0.0164 (v#1, f—)
5. [PROD-1008-H] Hiking Daypack 25L score=0.0161 (v#2, f—)
```
The `(v#, f#)` tags are each hit's rank on each side. Arctic Shell was
fulltext's top pick but only sixth by vector, so pure vector search would have
buried it, and RRF lifts it to second because doing well on _either_ side
counts. Compression Socks went the other way: they were the vector side's top
hit, but with no fulltext match at all they end up at the bottom of the list.
## It composes with everything else
The same fusion is on the query builder as `.fuseWith()`, so a hybrid search
can carry ordinary predicates. Filtering to `status = "active"` drops the
discontinued Legacy Rain Shell in the same query, without a post-filter:
```typescript
const builderHybrid = await store
.query()
.from("Product", "p")
.whereNode("p", (p) =>
p.$fulltext
.matches("waterproof shell", 20)
.and(p.embedding.similarTo(queryEmbedding, 20))
.and(p.status.eq("active")),
)
.fuseWith({ k: 60, weights: { vector: 1, fulltext: 1.25 } })
.select((ctx) => ({ sku: ctx.p.sku, name: ctx.p.name }))
.limit(5)
.execute();
// → Arctic Shell, Expedition Parka, Hiking Daypack 25L, Ski Goggles UV400,
// Compression Socks
```
There are four query modes for what a search box actually receives:
`websearch` for Google-style syntax (`"ski goggles" -compression`), `phrase`
for exact adjacency, `plain` for all-terms-must-match, and `raw` when you want
the engine's native `tsquery` or FTS5 `MATCH` syntax.
And if the index falls behind (you added a `searchable()` field, or a bulk
write went around the store), `store.search.rebuildFulltext()` backfills it
page by page:
```text
Before rebuild: "waterproof" → 0 hits (fulltext rows cleared).
Rebuilt: kinds=Product processed=9 upserted=9 cleared=0 skipped=0
After rebuild: "waterproof" → 3 hits restored.
```
## Try it
- [Fulltext Search](/fulltext-search): RRF tuning, adding fulltext to
existing data, and alternate Postgres strategies (pg_trgm, ParadeDB,
pgroonga)
- [Semantic Search](/semantic-search): the vector half
- [Example 15](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/15-fulltext-hybrid-search.ts):
the runnable source behind this post
- [GitHub](https://github.com/nicia-ai/typegraph)
# An Infinite Supply of Graph Databases
import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro";
TypeGraph turns a SQLite (or PGlite, or Postgres) connection into a typed
property graph. If you point that connection at a Cloudflare Durable Object's
own storage, you **no longer have to provision the database at all**. Every
user, agent conversation, or workspace can have its own graph database without
anyone ever creating it, and I think that's a lot of fun.
I built a companion repo to show it off,
[cf-do-typegraph](https://github.com/nicia-ai/cf-do-typegraph). It wires
TypeGraph into a single Durable Object class and runs three real apps on
top, entirely on your laptop under `wrangler dev`. It's MIT-licensed.
## Naming is the provisioning step
Here's the whole multi-tenant boundary:
```ts
// the whole multi-tenant boundary, trimmed from src/worker.ts
const stub = env.GRAPH_DO.get(env.GRAPH_DO.idFromName(`${example}:${tenant}`));
return stub.fetch(request); // that tenant's graph, and nothing else, lives here
```
`idFromName` is effectively the provisioning API. The Durable Object it names,
and the private SQLite database inside it, comes into existence on the first
request and hibernates as soon as nothing is using it, so there's no
`CREATE DATABASE`, no row to add to a `tenants` table, and no migration to run
against a database that didn't exist five minutes ago. Giving every user,
agent, or document its own graph stops being a capacity-planning question,
because each one is just a name.
Minting a lot of them is just a loop:
```bash
curl -X POST "localhost:8787/api/spawn/notes?count=25&prefix=demo"
```
```json
{ "example": "notes", "seeded": ["demo-1", "demo-2", "...", "demo-25"] }
```
```bash
curl "localhost:8787/api/notes/demo-7/search?q=hibernation" # its own private graph
```
That endpoint forwards 25 `/seed` requests to 25 different names and lets
each Durable Object materialize when its request arrives. There's nothing
more to it than that.
## Isolation and shared transactions
Making each tenant a Durable Object has two consequences that matter more than
the provisioning trick.
**Isolation is physical.** Each tenant's SQLite file is its own `ctx.storage`
rather than a slice of a shared table. The usual SaaS approach is a
`tenant_id` column plus a promise that every query remembers its
`WHERE tenant_id = ?`, but here there's no shared table to put that column in,
so a bad join or a missing filter can only ever see one tenant's data.
**The graph and the app's data live in the same file and the same
transaction.** The Durable Object boots its store the same way on every
request:
```ts
// src/do/graph-do.ts
createSqliteBackend(drizzle(this.ctx.storage));
```
`ctx.storage` is the same SQLite the application would use for its regular
tables, and TypeGraph doesn't need a separate connection or service, so
writing an app row and the graph edge that describes it can be one
`store.transaction(...)`. There's no sync job between your data and your
graph, and nothing for them to drift apart on.
## Three apps, one Durable Object class
The repo runs the same `GraphDO` class three ways, distinguished only by
what you name the tenant. I picked each example because it needs a query
that's painful to hand-write in SQL.
### `authz`: a graph per workspace
Zanzibar-style permission checks are reachability questions, and
TypeGraph's ontology does the reasoning you'd otherwise hand-roll in a
`WITH RECURSIVE`:
```ts
// src/examples/authz/graph.ts
ontology: [
implies(owner, editor), // owner ⇒ editor ⇒ viewer
implies(editor, viewer),
inverseOf(memberOf, hasMember), // group membership, read backwards, zero stored reverse edges
],
```
A permission check is one `shortestPath` call over the edge kinds the
requested relation implies:
```bash
curl "localhost:8787/api/authz/acme/check?user=alice&relation=editor&doc=roadmap"
```
```json
{
"allowed": true,
"relation": "editor",
"subject": "user:alice",
"via": {
"grantSubject": "group:staff",
"grantRelation": "editor",
"grantTarget": "folder:engineering"
},
"pathNodeIds": [
"user:alice",
"group:eng",
"group:staff",
"folder:engineering",
"folder:specs",
"doc:roadmap"
],
"pathEdgeKinds": ["memberOf", "memberOf", "editor", "parentOf", "parentOf"]
}
```
Alice never got a direct grant on that doc, and the answer comes with the
path that proves she has access anyway: nested group membership (`alice` →
`eng` → `staff`), one group-level `editor` grant, and two levels of folder
inheritance, all inferred at query time from five stored edges without a
permissions table.
### `notes`: a graph per user
This one is a personal wiki with `[[wiki-links]]`, and the fun part is that
**FTS5 runs inside the Durable Object's own SQLite**, so there's no search
service to keep warm:
```bash
curl "localhost:8787/api/notes/me/search?q=reachability"
```
```json
{
"hits": [
{
"id": "note:reachability",
"title": "Reachability",
"score": "...",
"snippet": "Reachability asks whether you can get from one node to another..."
}
]
}
```
Backlinks don't need any stored rows. The graph only writes a forward `linksTo`
edge when a note contains `[[Some Target]]`, and one ontology declaration
lets a `linkedFrom` traversal read those same edges backwards:
```ts
// src/examples/notes/graph.ts
ontology: [inverseOf(linksTo, linkedFrom)],
```
There's no second edge to keep in sync, so backlinks can't drift from the
links that produced them.
### `agent-memory`: a graph per conversation
Each session is its own graph, seeded by replaying a hand-written
conversation one message at a time:
```text
"I'm Ada, and I'm PMing the Halo launch."
"Grace is my lead engineer — we've shipped three products together."
...
"New development: we're partnering with Northwind, and Grace is their main contact."
```
That eighth message is my favorite part of the repo. `Organization` isn't a
node kind in the compile-time schema, so it arrives at runtime, gets validated
the same way an LLM's proposed schema change would be, and is applied live:
```ts
// src/examples/agent-memory/replay.ts
const validation = validateGraphExtension(evolution.extension, {
strict: true,
});
current = await current.evolve(validation.data);
```
The graph grows a new node kind and a new `worksAt` edge kind in the middle
of a conversation, while the Durable Object is running. The change is
persisted, so it's still there the next time the object wakes from
hibernation. And because the store boots with `{ history: true }`, the same
graph can answer "what did the agent believe after message 4?", before
Northwind existed, by reading at a recorded point in time instead of the
live state.
## The same pattern without Cloudflare
The underlying idea, **a small, typed graph database per tenant, addressed by
name**, works well whenever tenants are numerous, mostly idle, and shouldn't
share a query surface. Durable Objects make it especially easy (SQLite at the
edge, isolated per object, and hibernating for free), but TypeGraph doesn't
know or care that one is underneath, and the same pattern works with:
- **A directory of SQLite files.** One file per tenant, opened through
`createLocalSqliteStore`. Provisioning is choosing a file path.
- **A pool of PGlite files.** Postgres-in-WASM via `createLocalPgliteStore`,
one file per tenant, no server process.
- **A database per Neon branch.** Real network Postgres, with the standard
`createPostgresBackend` pointed at a different connection string per
tenant, provisioned through Neon's API instead of `idFromName`.
I built the demo on Durable Objects because they need the least
infrastructure to try: there's no server to run, and you don't need an account
to develop against them.
## Tested against the real thing
Every example runs its tests against real Durable Objects via
`@cloudflare/vitest-pool-workers`, not mocks, including a cross-tenant
isolation test that seeds one tenant and checks that a same-shaped,
differently named tenant reads back empty. `pnpm dev` runs the whole thing
(landing page, live force-directed graph view, all three APIs) locally for
free. Workers AI, the one billable piece, is only used for optional
semantic search and entity extraction, and ships commented out.
## Try it
- [cf-do-typegraph](https://github.com/nicia-ai/cf-do-typegraph):
`pnpm install && pnpm dev`, then open `localhost:8787`
- [Bring Your Own Database](/blog/bring-your-own-database): the backend
abstraction that lets the same store run on a different SQLite underneath
- [Agent Memory That Knows Why It Believes Things](/blog/truth-maintenance-for-agent-memory):
the bitemporal history behind the `agent-memory` replay
- [An Agent That Grows Its Own Schema](/blog/runtime-schema-evolution): `store.evolve()`,
the mechanism behind the live `Organization` extension
- [GitHub](https://github.com/nicia-ai/typegraph)
# Introducing TypeGraph
import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro";
Every project I've worked on that needed real structure (a knowledge base, an
org chart, memory for an agent) ended up with the same stack: an ORM for the
actual data, a vector store added when someone wanted semantic search, and
eventually, once the relationships got interesting, a graph database off to
the side with a sync job keeping it roughly in line with the other two.
Each of those systems has its own idea of what a `Person` is and its own
consistency model, so the schema drifts between them. Meanwhile TypeScript
already has Zod, a schema language good enough to describe all of it, and none
of those systems treats it as the source of truth.
TypeGraph is the other option: keep the graph inside the application, as a
library, and store it in the database you already run.
## It runs in your process
TypeGraph writes through your existing SQLite or Postgres connection and
commits in the same transaction as the rest of your data. You don't deploy a
graph server or keep one running, and your app doesn't make a network hop to
reach its own graph.
## One schema
You describe your data once, in Zod:
```typescript
const Person = defineNode("Person", {
schema: z.object({ name: z.string(), role: z.string() }),
});
const worksOn = defineEdge("worksOn", {
schema: z.object({ since: z.string() }),
});
const graph = defineGraph({
id: "org",
nodes: { Person: { type: Person } },
edges: { worksOn: { type: worksOn, from: [Person], to: [Person] } },
});
```
From that one definition TypeGraph derives runtime validation, the TypeScript
types, the storage layout, and what the query builder will let you write. If
you've ever kept an ORM schema, a folder of hand-written interfaces, and a
Cypher cheat sheet in sync by hand, you know why I wanted this.
## Relationships that mean something
A foreign key tells you two rows are related, but it can't tell you that a
`Podcast` is a kind of `Media`, that `marriedTo` implies `knows`, or that a
`Person` and an `Organization` can never be the same thing. In most codebases
those rules live in a comment, or in the head of whoever wrote the migration.
In TypeGraph, edges are first-class and typed, and they carry their own
properties. The ontology layer (`subClassOf`, `implies`, `disjointWith`,
`equivalentTo`) does real work: a query for `Media` can include podcasts, a
`knows` traversal can pick up spouses, and a write that would make a person
and an organization share an identity fails. Traversals compile to SQL, so a
three-hop walk is one query rather than a loop of lookups.
## Vectors are just another field
Embeddings are a field type. When you declare one, TypeGraph stores and
indexes it on whichever backend you're running:
```typescript
const Document = defineNode("Document", {
schema: z.object({ title: z.string(), embedding: embedding(1536) }),
});
const similar = await store
.query()
.from("Document", "d")
.whereNode("d", (d) => d.embedding.similarTo(queryVector, 10))
.select((ctx) => ctx.d)
.execute();
```
`.similarTo()` is a predicate like any other, so it composes with filters and
traversals in the same query. You don't ask a vector database for a list of
IDs and then run a second query to find out what those IDs are.
## What it isn't
TypeGraph isn't trying to be Neo4j. There's no PageRank, no community
detection, no distributed storage, and at this stage no traversal algorithms
beyond the queries you write yourself. It's built for thousands to millions
of nodes living next to the rest of your data. If your graph _is_ the product
and it has billions of edges, you want a dedicated graph database.
## Where it stands
It's early, a handful of releases in. The DSL, the ontology layer, vector
search, and both backends are solid. What comes next depends on what people
build with it, so if you try it and hit a wall, open an issue.
## Try it
- [What is TypeGraph?](/overview) and the [Quick Start](/getting-started)
- [Ontology & Reasoning](/ontology): `subClassOf`, `implies`,
`disjointWith`, `equivalentTo`
- [Semantic Search](/semantic-search)
- [GitHub](https://github.com/nicia-ai/typegraph)
# Turning Two Agents' Event Streams Into One Canonical Graph
import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro";
Imagine a sales bot and a support bot that both talk to the same person
without knowing it. The sales bot knows her as "Jane Doe", VP Engineering, at
`jane@acme.com`, while the support bot has "J. Doe", VP Eng, at the same email,
along with two companies the sales bot has never heard of, one of which the
support bot retracts a day later. Both bots emit durable, at-least-once event
streams of what they've seen.
Event logs are great at delivery, ordering, and replay, but they don't resolve
entities, and you can't ask a Kafka topic what an agent believed last Tuesday.
That work happens in a layer between the log and whatever reads it, which is
what
[Materializing Event Logs](/materializing-event-logs) describes.
[`@nicia-ai/agent-stream-graph`](https://github.com/nicia-ai/agent-stream-graph)
is the reference implementation, built on 0.35's
`store.transactionWithReceipt()` and a couple of 0.36 additions. It pulls
together three things I'd built separately (bitemporal history, idempotent
writes, and graph merge), and seeing them click into one pipeline was a lot
of fun.
## Projecting events idempotently
Streams redeliver changes after a crash, a reconnect, or a replay, so the
second delivery of a change has to land on the same row as the first. In
practice that means `upsertById` for nodes and `getOrCreateByEndpoints` for
edges, and a bare `create` only when the source event carries its own unique
id that you pass through as the TypeGraph id.
```typescript
const project: Projector = async (belief, change) => {
switch (change.shape) {
case "person": {
if (change.operation === "delete") {
await belief.nodes.Person.delete(asNodeId(change.key));
return;
}
await belief.nodes.Person.upsertById(change.key, {
name: change.value.name,
email: change.value.email,
title: change.value.title,
});
return;
}
case "employment": {
await belief.edges.worksAt.getOrCreateByEndpoints(
{ kind: "Person", id: change.value.person },
{ kind: "Company", id: change.value.company },
{},
);
return;
}
}
};
```
## Resuming after a crash
`consume()` checkpoints the last processed offset and a recorded-time anchor
after every change, and uses `store.transactionWithReceipt()` to know whether
the projector actually wrote anything. The demo runs the sales bot's stream,
stops it after two of its four changes, then simulates the nastiest crash
window there is: a change gets projected, and the process dies before the
checkpoint lands.
```text
(b) Durable consumer — resume from checkpoint, replay safely after crash
consumer ran, then crashed after 2 messages
durable cursor: last offset = 002
sales-bot belief so far: 1 person, 1 company
crash window: projected 003, then died before checkpoint
uncheckpointed anchor existed: 2026-07-14T16:32:30.397Z
durable cursor is still: 002
belief already has worksAt edges: 1
restarted — replayed 003, then processed 004 (2 messages)
sales-bot belief now: 1 person, 1 company — Jane Doe (VP Eng & Product)
worksAt edges after replay: 1 (no duplicate edge)
re-run (at-least-once): 0 messages processed; belief unchanged: 1 person, 1 company
```
The cursor still says `002` even though `003` was already projected, which is
the crash window the demo sets up deliberately. On restart `003` replays,
`getOrCreateByEndpoints` finds the edge it already wrote, and processing
carries on. Running the whole stream a third time processes nothing because
the cursor is already at the end.
## Rebuilding from offset zero
A full rebuild is a different case: replay every event from the beginning
into a fresh belief graph, for recovery or a schema migration. Before 0.36,
that rewrote every row even when the replayed value matched what was already
there, so each rebuild cost a wasted write per event and grew recorded
history by the length of the log.
`createStore(graph, backend, { coalesceUnchangedUpserts: true })` makes a
value-identical redelivery a true no-op that writes nothing, adds no history
row, and doesn't advance the revision. A stream that actually changes a value
and later changes it back still writes both times, since those changes really
happened. 0.36
also adds [`tx.measure()`](/schemas-stores/#scoped-receipts-txmeasure),
which scopes a receipt to the writes made by one callback, so a
materializer's own cursor bookkeeping in the same transaction can't be
mistaken for projector output.
## What did each bot believe, and when?
Each bot's belief graph has history enabled, so `book.anchorFor(stream,
offset)` gives you a recorded-time coordinate for any processed offset. That
means you can read exactly what _this_ bot believed at _that_ point in its
own stream:
```text
(c) What did each agent believe, at which offset?
sales-bot's own belief graph:
@offset 001: people: Jane Doe (VP Engineering) | companies: —
@offset 004: people: Jane Doe (VP Eng & Product) | companies: Acme Corp (Series A) (title corrected)
support-bot's own belief graph (same person, different surface form):
@offset 004: people: J. Doe (VP Eng) | companies: Umbrella (unverified), Acme (Series B), Globex (Series C)
@offset 005: people: J. Doe (VP Eng) | companies: Acme (Series B), Globex (Series C) (Umbrella retracted)
→ Same email, but neither agent alone knows 'Jane Doe' and 'J. Doe'
are one person. The retracted company also remains visible in past belief.
```
The support bot's belief at offset 4 still shows "Umbrella (unverified)",
even though it was deleted at offset 5, because a read at a past offset
reconstructs the belief as it was then.
## One canonical graph
Each bot's belief exports through streaming interchange into a
[graph merge](/blog/graph-merge) branch, and `mergeIncremental()` folds it
into a canonical graph. It's the same entity resolution as in the graph merge
post, run once per stream as it arrives:
```typescript
await importGraphStream(
agentBranch.store,
exportGraphStream(belief, { includeTemporal: true }),
{ onConflict: "update" },
);
const result = await mergeIncremental({
forkPoint,
target: canonical,
branches: [agentBranch],
options: {
resolve: {
Person: {
similarity: { kind: "fulltext", fields: ["name"] },
threshold: 0.9,
},
Company: {
similarity: { kind: "fulltext", fields: ["name"] },
threshold: 0.9,
},
},
onPropertyConflict: "flag",
onBasePropertyConflict: "flag",
branchOrder: [branchId],
persistProvenance: true,
},
});
```
The sales bot merges first, as a clean append. The support bot merges
second, and that's where the same-person, different-spelling problem gets
resolved:
```text
[wave 1] merged sales-bot — conflicts: 0
[wave 2] merged support-bot — conflicts: 4
conflict: Company.name on c1: support-bot="Acme"; kept "Acme Corp"
conflict: Company.stage on c1: support-bot="Series B"; kept "Series A"
conflict: Person.name on p1: support-bot="J. Doe"; kept "Jane Doe"
conflict: Person.title on p1: support-bot="VP Eng"; kept "VP Eng & Product"
canonical now: 1 person, 2 companies — Jane Doe (VP Eng & Product)
provenance — sales-bot contributed to 3 canonical entities
provenance — support-bot contributed to 3 canonical entities
```
"Jane Doe" and "J. Doe" collapse into one canonical person, and every
disagreement is flagged instead of silently overwritten. Globex, which only
the support bot ever saw, joins as a new company. Umbrella, which the support
bot retracted before the merge, never shows up at all. And the canonical
graph has history too, so it time-travels across merge waves:
```text
canonical, time-travelled:
asOfRecorded(after wave 1): 1 person, 1 company
asOfRecorded(after wave 2): 1 person, 2 companies
```
So from two unreliable streams you end up with one canonical graph, and you
can still ask any of the three graphs what it believed at any point along the
way.
## A bigger example
The demo above is small on purpose: one person, two bots, five offsets. The
repo's main demo (`pnpm demo`) runs the same mechanics against a set of wiki
entries that mention overlapping concepts under different names, and
converges them into one canonical concept graph with source attribution for
every mention.
## Try it
- [Materializing Event Logs](/materializing-event-logs): idempotent
projectors, cursor bookkeeping, transaction receipts, and mapping event
time onto the two temporal axes
- [Graph Merge](/graph-merge): the entity resolution this builds on
- [`agent-stream-graph`](https://github.com/nicia-ai/agent-stream-graph):
`pnpm demo` and `pnpm demo:mechanics`
- [GitHub](https://github.com/nicia-ai/typegraph)
# Replacing the N+1 Loop With One Statement
import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro";
Most apps have a page like this one: a workspace lists its published
documents, and each row shows the title, how many versions the document has,
the note on the newest version, and the three most recent comments. Every
piece of that is a one-hop read from the document, so the first version anyone
writes is a loop:
```typescript
const documents = await store
.query()
.from("Document", "document")
.whereNode("document", (document) => document.status.eq("published"))
.orderBy("document", "title")
.select((ctx) => ctx.document)
.execute();
const rows = await Promise.all(
documents.map(async (document) => {
const versionEdges = await store.edges.hasVersion.findFrom(document);
const versions = await store.nodes.Version.getByIds(
versionEdges.map((edge) => edge.toId),
);
const commentEdges = await store.edges.hasComment.findFrom(document);
const comments = await store.nodes.Comment.getByIds(
commentEdges.map((edge) => edge.toId),
);
// Sort in JavaScript: newest version, version count, three latest comments.
return summarize(document, versions, comments);
}),
);
```
Counted at the driver, that loop issues 33 statements for eight documents,
one for the list and four per document. On in-memory SQLite you'd never
notice, but with the database across a network every one of those is a round
trip, the count grows with the length of the list, and you're loading every
comment to keep three and every version to keep one.
This is the N+1, the oldest performance bug in the book, and over releases
0.58 to 0.65 I went after it properly. The same page now costs **one SQL
statement**, and when TypeGraph can't do a batch in one statement it throws
before running anything instead of quietly issuing more.
(Every number here is a statement count, measured by wrapping
better-sqlite3's `Statement` and `exec` in the script that ran the code. None
of them are wall-clock times. What a round trip costs on your network is
yours to multiply in.)
## Step one: reads that say what they want
`store.neighbors()` joins each edge to the node on its far end in a single
statement, and applies ordering and the limit before hydrating anything, so
"the newest version" fetches one row. `store.countNeighbors()` is the matching
count and hydrates nothing at all.
```typescript
const [latest] = await store.neighbors(document, {
edges: ["hasVersion"],
orderBy: { by: "node", field: "sequence", direction: "desc" },
limit: 1,
});
const versions = await store.countNeighbors(document, {
edges: ["hasVersion"],
});
```
Two calls per document is 16 statements for the page, which is better than
33 but still grows with the length of the list.
## Step two: `batchOnce()`
`store.batchOnce()` is the part I'm happiest with. It takes any number of
independent reads and runs them as exactly one SQL statement: each read
becomes a CTE, and a single JSON envelope carries every result set back, in
input order. The callback can return a runtime-sized array, which means the
loop above translates almost literally:
```typescript
const results = await store.batchOnce((read) =>
documents.flatMap((document) => [
read.neighbors(document, {
edges: ["hasVersion"],
orderBy: { by: "node", field: "sequence", direction: "desc" },
limit: 1,
}),
read.countNeighbors(document, { edges: ["hasVersion"] }),
]),
);
// results[2 * i] is document i's newest version, results[2 * i + 1] its count.
```
That takes the page from sixteen statements to one. Inside a transaction,
`tx.batchOnce()` binds to the open connection, so the batch sees writes made
earlier in the same callback and is still one statement.
### It won't quietly fall back
A lot of "batch" APIs are really a `for` loop, and you only find out they
issued N queries when a latency graph shows it a month later. `batchOnce()`
has no fallback path, so if it can't do the job in one statement it throws
before executing anything:
```typescript
await store.batchOnce((read) =>
Array.from({ length: 501 }, () =>
read.countNeighbors(document, { edges: ["hasVersion"] }),
),
);
// ConfigurationError: store.batchOnce() accepts at most 500 reads in one statement.
```
Five hundred reads run as one statement, and 501 throws without issuing any.
It won't chunk to fit the backend's bind-parameter limit either, and reads that
can't be embedded (edge-collection `batchFind*` calls, for example) are
rejected instead of being run on the side. Since the name promises one
statement, I'd rather it fail loudly than break that promise.
The older `store.batch()`, by contrast, runs its queries one after another. It
always did, and the docs now say so plainly because people were putting it in
hot paths expecting a single round trip.
## Step three: shape it in SQL
A batch of `neighbors()` calls still has one read per parent, so eight
documents means sixteen reads inside the statement, and the 500-read ceiling
works out to 250 documents. When the question is "for every document in this
result", a relation answers it with a fixed number of reads however long the
list is.
0.61 added `project()`, derived relations, and aggregates; 0.63 added
`topPerPartition()`:
```typescript
const published = () =>
store
.query()
.from("Document", "document")
.whereNode("document", (document) => document.status.eq("published"));
const versions = published()
.traverse("hasVersion", "link")
.to("Version", "version")
.project((fields) => ({
documentId: fields.document.id,
versionId: fields.version.id,
sequence: fields.version.sequence,
note: fields.version.note,
}))
.asRelation();
const versionCounts = versions
.groupBy((columns) => [columns.documentId])
.aggregate((columns) => ({
documentId: columns.documentId,
versions: expr.count(columns.versionId),
}));
const latestVersions = versions.topPerPartition({
partitionBy: (columns) => [columns.documentId],
orderBy: (columns) => [
{ expression: columns.sequence, direction: "desc" },
{ expression: columns.versionId },
],
limit: 1,
});
```
`topPerPartition()` uses `ROW_NUMBER()`, so ties don't widen the limit. Both
orderings end in an id because the API can't know your ordering is unique,
and a repeatable winner needs a tiebreaker. The comments are the same shape
with a limit of three, plus an ordered collection that folds the winners into
one array per document:
```typescript
const recentComments = published()
.traverse("hasComment", "link")
.to("Comment", "comment")
.project((fields) => ({
documentId: fields.document.id,
commentId: fields.comment.id,
body: fields.comment.body,
postedAt: fields.comment.postedAt,
}))
.asRelation()
.topPerPartition({
partitionBy: (columns) => [columns.documentId],
orderBy: (columns) => [
{ expression: columns.postedAt, direction: "desc", nulls: "last" },
{ expression: columns.commentId },
],
limit: 3,
})
.groupBy((columns) => [columns.documentId])
.aggregate((columns) => ({
documentId: columns.documentId,
recent: expr.collect(columns.body, {
orderBy: [
{ expression: columns.postedAt, direction: "desc", nulls: "last" },
{ expression: columns.commentId },
],
}),
}));
```
Relations are batch members like any other read, so the whole page, titles
included, is one statement:
```typescript
const titles = published()
.orderBy("document", "title")
.select((ctx) => ({ id: ctx.document.id, title: ctx.document.title }));
const [documents, counts, latest, comments] = await store.batchOnce(
() => [titles, versionCounts, latestVersions, recentComments] as const,
);
```
Joined on `documentId`, a row comes out as:
```text
{ id: "1BMs3-6PEJYXtjB9BcrsU", versionCount: 3, latest: "v3 of doc 2",
recent: ["comment 5 on doc 2", "comment 4 on doc 2", "comment 3 on doc 2"] }
```
The script asserts that these eight rows are identical to what the loop
produced. Here's how the versions of the page compare:
| Page for eight documents | Statements |
| -------------------------------------------- | ---------: |
| Loop of `findFrom()` and `getByIds()` | 33 |
| Loop of `neighbors()` and `countNeighbors()` | 16 |
| `batchOnce()` of the `neighbors()` reads | 1 |
| `batchOnce()` of four relations | 1 |
`neighbors()` fits when you already hold a handful of sources, and relations
fit when you want "every parent in this result". `topPerPartition()` needs
window functions
and ordered `expr.collect()` needs ordered aggregates; a backend without them
throws a typed error before executing.
## The rest of the set-shaped toolkit
A few smaller pieces from the same stretch, all aimed at the same habit of
doing one thing per row:
- **`bulkFindFrom()` / `bulkFindTo()`** are `findFrom()` / `findTo()` for a
whole page of endpoints. Index `i` of the result holds the edges of input
`i`, an endpoint with no edges gets an empty array, and `limitPerInput`
caps each endpoint's fan-out. Fifty people's jobs cost one statement per
endpoint kind (split only when the bind-parameter budget requires it)
instead of fifty.
```typescript
const people = await store.nodes.Person.find({ limit: 50 });
const jobsPerPerson = await store.edges.worksAt.bulkFindFrom(people);
```
- **`store.bulkFindEdgesTo()`** does the inbound direction across several
node and edge kinds at once. Twelve documents and two edge kinds was one
statement; calling `findTo()` per kind per document was 24.
- **`updateWhere()`** is a set-based, transactional update that returns how
many rows changed. The selector is mandatory (`where`, `exists`, a
candidate query, or an explicit `all: true`), so you can't wipe out a whole
kind by forgetting a filter:
```typescript
const result = await store.nodes.Person.updateWhere({
patch: { active: false },
where: (person) => person.lastSeen.lt(cutoff),
exists: [
{
edgeKind: "worksAt",
direction: "out",
relatedKind: "Company",
whereRelated: (company) =>
company.field("status").string().eq("closed"),
},
],
});
// { affectedCount: number }
```
- **`in()` / `notIn()` take list parameters**, so `field.in(param("ids"))`
binds a runtime-sized list in a prepared query instead of forcing you to
rebuild the query per call.
- **Subgraphs got per-edge-kind windows** (keep only the newest N edges of a
noisy kind), and `subgraph()` works inside `batchOnce()`. Five roots
awaited in a loop cost 10 statements on SQLite; batched, one.
- **Mixed-kind cursor pages**: a query can start from several kinds
(`.from(["Person", "Team"], "entity")`), and a `.page()` can sit in a batch
next to unrelated reads, so a directory of people and teams can be read as
one ordered stream at one statement per page.
- **`executeChecked(version)`** folds the "has another isolate changed the
schema?" probe into the read itself, for serverless deployments that cache
the schema per isolate. Probing and then reading took 2 statements, and the
checked read takes 1.
A moved schema throws `SchemaChangedError`, even when the query would have
matched no rows.
## What one statement doesn't buy you
- **It saves round trips, but the database does the same work.** Every member
of a batch still runs its own plan. For one large closure on Postgres, the
direct `subgraph()` can beat the batched form, and `topPerPartition()`
bounds the rows returned, not necessarily the rows scanned.
- **Results are materialized.** `batchOnce()` returns JSON envelopes, doesn't
stream, and has no byte cap. Bound your reads with limits and projections.
- **The counts are SQLite driver counts.** They show the shape of the
improvement, and your actual latency depends on your network.
## Try it
- [`store.batchOnce()`](/schemas-stores#storebatchoncebuildreads-options):
the full contract, including what can and can't be embedded
- [Top-N per parent](/queries/relations#top-n-per-parent) and
[Ordered collections](/queries/relations#ordered-collections)
- [`updateWhere()`](/schemas-stores#updatewhereparams) and
[`bulkFindEdgesTo()`](/schemas-stores#storebulkfindedgestoparams-options)
- [GitHub](https://github.com/nicia-ai/typegraph)
# Tracking Which Records Are the Same Person
import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro";
Say a bank's KYC onboarding writes a `Person` node for each customer, keyed by
the bank's customer id, and its fraud pipeline writes a `CaseSubject` for
everyone named in an investigation. When the person under investigation is a
walk-in customer, the case system reuses the bank's id, so one human being
ends up as two nodes of different kinds, and nothing in the graph says they're
the same person.
Most graph libraries handle "these are the same thing" at the type level, and
TypeGraph used to as well. The `sameAs` and `differentFrom` ontology factories
related _kinds_ rather than rows, and `differentFrom` never checked anything
about actual data, which is of no use when a compliance team needs to say that
this particular customer is that particular case subject. Both factories are
deprecated now.
Their replacement, `store.identity`, is one of my favorite things in the
library. It's a ledger of claims about individual nodes, either that two of
them are the same entity or that they're provably not. You can query it at any
point in time and retract a claim without deleting it, the claims carry
through traversals and merges, and the database itself won't commit a set of
claims that contradicts itself.
## Folding on a shared id
Identity is opt-in per graph. Turning it on also decides what happens when
nodes of different kinds share an id:
```typescript
const graph = defineGraph({
id: "compliance",
nodes: {
Person: { type: Person },
CaseSubject: { type: CaseSubject },
Organization: { type: Organization },
},
edges: {},
ontology: [disjointWith(Person, Organization)],
identity: { sameIdAcrossKinds: "fold" },
});
```
With `"fold"`, live nodes of different kinds that share an id are the same
entity automatically. (`"ignore"` gives you the ledger without that implicit
join, for graphs where a shared id is a coincidence.) The two systems above
never have to coordinate:
```typescript
const person = await store.nodes.Person.create(
{ name: "Alex Rivera" },
{ id: "cust-8842" },
);
const subject = await store.nodes.CaseSubject.create(
{ note: "Named in fraud case FC-2201" },
{ id: "cust-8842" },
);
await store.identity.membersOf(person);
// [{ kind: "CaseSubject", id: "cust-8842" },
// { kind: "Person", id: "cust-8842" }]
```
A fold starts when the node actually exists, not at its `validFrom`. A node
created today with a backdated validity window shows up in historical reads,
but it doesn't retroactively fold anything in the past.
## Saying it explicitly
Most matches don't come with a shared id. Here an investigator links a
walk-in customer to an actor in an unrelated case, under two ids that have
nothing in common:
```typescript
const walkIn = await store.nodes.Person.create(
{ name: "J. Alvarez" },
{ id: "person-119" },
);
const caseActor = await store.nodes.CaseSubject.create(
{ note: "Signed the wire authorization in case FC-2214" },
{ id: "case-77" },
);
const linked = await store.identity.assertSame(walkIn, caseActor);
// { action: "created", assertion: { id, relation: "same", a, b, validFrom } }
await store.identity.assertSame(walkIn, caseActor);
// { action: "existing", assertion: } — idempotent
```
Asserting the same thing twice gives you the existing assertion back, so a
job that re-runs after a crash doesn't need to know whether it got that far
last time. There are bulk forms (`bulkAssertSame`, `bulkAssertDifferent`) that
return one result per input pair, in order.
Claims are also checked against the ontology before anything is stored.
Since `Person` is `disjointWith` `Organization`, folding the walk-in onto a
company fails:
```typescript
await store.identity.assertSame(walkIn, acmeFreight);
// throws IdentityContradictionError
// details: { operation: "assertSame", reason: "disjoint-kinds", a, b }
```
Claims can also be bounded in time. A merged account that was later split, or
a case subject that was only correct for the window an investigation covered,
gets a half-open validity window:
```typescript
await store.identity.assertSame(alice, legacyAlice, {
validFrom: "2020-01-01T00:00:00.000Z",
validTo: "2022-01-01T00:00:00.000Z",
});
```
Both endpoints have to exist for the whole window, and contradictions are
checked across every overlapping stretch of time, including through chains of
`same` claims.
## Changing your mind
Suppose forensic review later shows that `person-119` and `case-77` are two
different people who happened to share a wire-authorization signature. You
want to record that, but the ledger currently says they're the same entity,
and it won't hold both claims at once:
```typescript
await store.identity.assertDifferent(walkIn, caseActor);
// throws IdentityContradictionError
// details: { operation: "assertDifferent", reason: "same-class", a, b }
```
You retract the old claim first, explicitly:
```typescript
const ended = await store.identity.retractAssertion(linked.assertion.id);
// ended.validTo: "2026-08-21T18:03:11.442Z" — the fold ends, timestamped
await store.identity.assertDifferent(walkIn, caseActor);
// { action: "created", ... } — now recorded as provably different
await store.identity.areSame(walkIn, caseActor); // false
await store.identity.areDifferent(walkIn, caseActor); // true
```
Nothing was deleted: the retracted assertion and its `validTo` are still
readable, and a query at a time before the retraction still sees one entity,
so if an auditor asks why these two records were treated as one customer in
August, you can answer them.
Soft-deleting a node ends its current assertions too, and records the node as
the reason (`endedBy`), so a later reader can tell _why_ a claim ended.
## A backstop in the database
Everything so far is application code deciding whether a write is allowed,
and application code can be wrong. `assertSame` checks against state it just
read, so a bug somewhere else that writes identity rows directly could still
commit the contradiction the API rejects.
So identity also keeps a table of which identity classes are held apart by a
`different` claim, with a database `CHECK` constraint on it. When a
transaction merges two classes, it rewrites those rows in the same batch. If
two classes that are supposed to be separate get merged anyway, the rewrite
violates the constraint and the database aborts the transaction:
```typescript
// IdentitySeparationViolationError
// details: { graphId, enforcedBy: "database", classKey,
// assertionId, a, b }
```
The API catches contradictions first and gives you a useful error, and the
constraint is there for the case where the API has a bug.
## Traversals that understand it
Queries can expand through identity classes, per hop, and it's off by
default:
```typescript
const results = await store
.query()
.from("Person", "person")
.traverse("authored", "edge", { includeIdentityMembers: true })
.to("Document", "document")
.select((ctx) => ({ edge: ctx.edge, document: ctx.document }))
.execute();
```
That hop follows `authored` edges from the person _and_ from every other node
in their identity class, and returns the real edge and target rows with
duplicates removed.
At the current time, each hop looks classes up through an index on the
maintained closure, so the cost tracks the rows you start from and the size of
their classes, not how many classes the graph holds. Measured on SQLite, with
each `Person` folded to a `Company` and an `Alias` sharing its id, before and
after the closure became index-seekable:
| source rows | fan-out | matching edges | before | after |
| ----------- | ------- | -------------- | --------- | ----- |
| 250 | 1 | 250 | 67 ms | 6 ms |
| 1000 | 1 | 1000 | 1077 ms | 9 ms |
| 2000 | 1 | 2000 | 4616 ms | 19 ms |
| 1000 | 8 | 8000 | 8611 ms | 13 ms |
| 500 | 200 | 100,000 | 51,602 ms | 77 ms |
That closure only describes the present, which is worth planning around. A
hop at a _historical_ time has to rebuild classes from the assertion ledger
for the whole graph, once per statement, even if you only asked about one
node, so queries about the present are much cheaper than queries about the
past.
Identity claims also carry through interchange and graph merge. Exports carry
assertions (current ones by default, ended ones too if you ask for an archival
export), and graph merge treats identity as part of what it diffs. Two branches
that make opposing claims about the same pair, or that contradict each other
through a chain of `same` claims neither wrote alone, fail at plan time with
`IdentityMergeConflictError`, and the merge checks again inside its own
transaction before committing.
## Staging the duplicate you came to resolve
Ingestion is where identity gets messiest. Say a provider feed arrives with a
patient carrying MRN `MRN-4471` and the graph already has a `Patient` with
that MRN. That's expected, since deciding whether the two are the same person
is why the ingestion pass exists, but `Patient` declares `unique("mrn")`, so
the second row can't be written. If you stage it on an ordinary `branch()`,
staging fails at the duplicate before entity resolution has seen anything.
`ingestionBranch()` makes a working copy that defers node uniqueness and
nothing else, so schema validation, endpoint checks, disjointness and edge
cardinality all still apply while staging. The usual answer to this problem is
a "skip validation" flag on the import, which I didn't want, because it lets
every other kind of bad row in alongside the duplicates you meant to allow.
```typescript
const incoming = unwrap(
await ingestionBranch(base, makeBackend, { id: asBranchId("provider-a") }),
);
const imported = await importGraph(incoming, providerDocument, {
onConflict: "error",
onUnknownProperty: "error",
});
if (!imported.success) throw new Error("Provider import was rejected");
const alias = await incoming.nodes.Patient.getById(
asNodeId("incoming-patient"),
);
if (alias === undefined) throw new Error("Imported patient was not found");
// `canonicalPatient` was read from the base before forking.
await incoming.identity.assertSame(canonicalPatient, alias);
```
The duplicate MRN and the claim that explains it now sit on the same branch
and reach merge planning together. The handle only exposes the assertion
methods, without identity reads, retractions, transactions, or access to the
underlying store, so it can't be used to get around the deferred constraint.
At merge time the original uniqueness rule applies again. `applyMergePlan()`
checks uniqueness against the entire resolved write set inside the target
transaction, so a valid key handoff or swap goes through as one set (a
row-by-row check would see two owners in the middle and reject it). If
resolution still leaves two live owners of one MRN, the merge fails with
`MergeConstraintConflictError` and writes nothing, so the deferral only gives
you room to review the duplicate before merging.
## What it costs
A graph without identity enabled pays nothing: no identity SQL, locks, or
closure work. With it on:
- **One writer per graph for identity-affecting writes on Postgres.** They
serialize on a per-graph advisory lock. Other graphs and all reads are
unaffected.
- **First-time enablement is heavy.** It briefly takes a `SHARE` lock on the
shared nodes table, which blocks writes for every graph in that database,
and loads the whole graph to build the initial closure. Do it in a quiet
window.
- **Switching `fold` and `ignore` is a breaking schema change,** because it
changes every `areSame` answer on existing data. It needs an explicit
migration.
- **It needs real transactions.** The bundled SQLite and Postgres drivers
have them. Cloudflare D1 and `drizzle-orm/neon-http` don't, and reject an
identity-enabled graph at construction.
`store.identity` records and propagates the identity claims _you_ make, so
deciding that two records are the same person is still up to you or your
entity-resolution pass. What it adds is a durable record of those decisions,
including when each was made and when it stopped being true.
## Try it
- [Operational Identity](/identity): the full guide, including migrating off
`sameAs`/`differentFrom`
- [Constraint-aware ingestion branches](/graph-merge#constraint-aware-ingestion-branches)
- [Graph Merge](/graph-merge): branch, resolve, and merge identity back
- [GitHub](https://github.com/nicia-ai/typegraph)
# Five Bugs That Didn't Crash
import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro";
A bug that crashes is the easy kind, because it tells you where it is. The ones
I worry about in a database library are the ones that return a
reasonable-looking answer (zero rows, a successful write, a clean startup)
while the data is wrong, so you find out weeks later or not at all.
Over the last stretch of releases (0.44 through 0.50, plus one older fix) I've
fixed five of those in TypeGraph. None of them warrants a post of its own, but
together they show how I want the library to behave, and one of them you
could hit just by doing what TypeGraph's own error message told you to do.
## Rows that exist at no point in time
A valid-time window is half-open, so `asOf(t)` returns a row when
`valid_from <= t < valid_to`. If you backfilled a record you knew had already
ended, you'd pass a `validTo` in the past and no `validFrom`, and the write
stamped its own instant as `validFrom`, which put the start after the end. A
window that runs backwards contains no `t` at all, so the row was stored,
counted, and exported, but no temporal read could ever return it.
Since 0.48, a write that creates a row (or resets its window) with a `validTo`
at or before its own instant and no `validFrom` stores no lower bound instead,
so the row reads as "ended at T, start unknown", which is what you meant. A
future `validTo` behaves as before. Custom backends can call the same helper
the built-in ones use, `resolveStampedValidityLowerBound`, so the rule lives
in one place.
While I was in there, 0.47 added the thing people were faking with
delete-and-recreate: `clearValidTo: true` reopens a window you closed too
early, on the same row, with `oneActive` edges rechecked because reopening can
create a second active edge.
```typescript
await store.nodes.Employment.updateById(employmentId, {
patch: { department: "Research" },
clearValidTo: true,
});
```
### Upgrading doesn't repair old rows, on purpose
Older versions could store inverted windows, and upgrading leaves them alone.
I think that's the right call: an upgrade that made invisible rows start
appearing in historical queries would change your reports and replays without
telling you which rows had moved, which is its own quiet bug.
So the repair is an explicit operator action, with a dry run:
```typescript
import { repairInvertedValidityWindows } from "@nicia-ai/typegraph";
const report = await repairInvertedValidityWindows({
backend: anyBackend,
relations: "live-and-recorded",
mode: "report",
});
// report.counts.recordedNodes === undefined means NOT SCANNED, never "clean"
```
`mode: "apply"` rewrites what `report` counted. Stop writers first, pass the raw
backend on a history-enabled store, and repair `"live-and-recorded"` unless you
have a reason not to, because fixing only the live rows leaves the recorded
history carrying the same backwards window and `asOfRecorded` will keep serving
it. The [runbook](/schema-management#repairing-inverted-validity-windows) has
the details.
## Constraints that `importGraph` walked straight past
TypeGraph lets you declare hierarchy-wide uniqueness
(`scope: "kindWithSubClasses"`), `disjointWith(Person, Organization)`, and
edge cardinality (`one`, `unique`, `oneActive`). Plain per-kind uniqueness was
always backed by a real database key, but these three were enforced by a
writer that took the per-graph write lock, checked, and then wrote.
That works only as long as every writer takes the lock, and `importGraph`
didn't. It also skipped the disjointness and cardinality checks entirely, so
an import could commit a `Person` and an `Organization` with the same id, or
three edges on a `cardinality: "one"` relationship, and report success. Import
is also the path most likely to be carrying data you didn't write yourself.
As of 0.50, each of those constraints is also backed by a reservation row
whose primary key admits exactly one owner. Taking the constraint means
winning that insert, so a writer holding no lock still loses the race it
should lose. Import now enforces all three and reports rejected rows in its
`errors` like any other violation. Every store and import write also goes
through a single write pipeline now. Import had drifted because it had its own
hand-built write path, and nobody noticed which rules it was missing.
There are two caveats. This only covers writes that go through TypeGraph, so
raw SQL inserting into the node or edge tables skips the reservation. And
databases created before 0.50 need the new edge-claims table, which the normal
bootstrap or the generated migration SQL provides.
As with validity windows, the fix prevents new violations and leaves old ones
where they are. To find those:
```typescript
for (const violation of await store.verifyConstraintFences()) {
console.warn(violation.family, violation.target.axis, violation.target.key);
}
```
The audit reads the nodes and edges themselves rather than the new
reservation tables, because a database written before those tables existed
has no reservations in it, and an audit that only looked there would report
zero violations on exactly the data you're worried about. It writes and
repairs nothing, since picking which of two conflicting rows survives destroys
data either way, and that decision should be yours.
## A search index that said it was fine
Fulltext and vector search need their own tables. TypeGraph creates them on
first use and writes a marker row saying so, and from then on it trusted the
marker without checking that the tables were still there.
When one went missing out of band (a partial restore, a migration that
recreated a schema, or an edge runtime that lost a file), the store opened
cleanly and then failed on the first search with a raw driver error about a
missing relation. This mostly showed up on Cloudflare Durable Objects, where
[one SQLite database per tenant](/blog/infinite-graph-databases) means
thousands of small databases that nobody is watching individually.
That error is now a `ContributionUnavailableError` with
`state: "physical-storage-missing"` and rebuild guidance. The check only runs
on the error path, so healthy stores don't pay for it. There's also a
three-step ladder for when search is broken, from cheapest to most drastic:
| Step | Call | Writes |
| ------- | --------------------------------------- | ------------------------------------- |
| Probe | `store.probeContributions()` | Nothing. Safe on a replica |
| Repair | `store.repairContributions()` | Marker rows, `IF NOT EXISTS` |
| Rebuild | `store.rebuildContribution("fulltext")` | Deletes and refills this graph's rows |
Start at the top and stop when the probe says `ready`. Rebuild is the only
fix when the table exists in a shape the current code no longer produces, and
it needs a maintenance window. Vector storage can't be rebuilt, because the
embeddings exist only in the table a rebuild would drop, so TypeGraph won't
drop it. The
[troubleshooting guide](/troubleshooting#contribution-health-probe-repair-rebuild)
walks through each state.
## A migration that deleted your runtime kinds
This one stings. [Runtime schema evolution](/blog/runtime-schema-evolution)
lets an agent add node and edge kinds with `store.evolve()`, and those kinds
live in the stored schema rather than in your TypeScript graph definition.
`migrateSchema()` committed whatever graph you passed it, so if you migrated
with your compile-time graph, every kind added at runtime disappeared from the
schema while its rows stayed in the tables where nothing could reach them. The
usual way you'd end up calling `migrateSchema()` was the library's own error
message telling you to review the changes and run it, so following my advice
hid your data.
0.44 folds the stored runtime additions in before committing, the same way
store creation already did. It also throws if a migration would drop a kind
that still holds rows, unless you pass `{ discardDroppedKindRows: true }`.
Two races in the same area got fixed alongside it:
- A schema commit could land while another writer was mid-write against the
version being replaced. Managed writes now recheck their schema version
while holding a lock the commit also needs, so a stale write fails instead
of landing against a schema that no longer accepts it.
- Removing a kind and adding it back before its cleanup ran made the old
rows reappear next to the new ones, and the cleanup then skipped them
because the kind was live again. `evolve()` now won't re-add a kind while
its cleanup is pending, and cleanup rechecks the schema under the same lock.
## An `implies()` that folded nonsense rows
This is the older one, from 0.35. `implies(edgeA, edgeB)` means "a traversal
of `edgeB` can include `edgeA` edges", which only makes sense if `edgeA`'s
endpoints could stand in for `edgeB`'s. Nothing checked that, so this was
accepted:
```typescript
edges: {
authored: { type: authored, from: [Author], to: [Paper] },
covers: { type: covers, from: [Paper], to: [Topic] },
},
ontology: [implies(authored, covers)],
```
and a `covers` traversal with `expand: "implying"` would fold in rows that
start at an `Author`. Now it fails when the graph is built or loaded:
```text
ConfigurationError: implies("authored", "covers") is endpoint-incompatible:
from kind(s) [Author] declared on "authored" cannot be assigned to any of
"covers"'s from kind(s) [Paper].
```
The check also runs on schemas loaded from the database, so a bad relation
saved by an older version fails on first load after upgrading. The same
release made projected ids keep their `NodeId` brand through `.select()`
and gave every fixed-shape error class a typed `details`, so fewer mistakes
make it to runtime in the first place.
## What they have in common
All five fixes follow the rules I now hold the whole library to. If TypeGraph
can't do what you asked, it throws and tells you why rather than returning an
empty result that looks like a clean one, and it doesn't rewrite your data
behind your back, even to fix it. You get a report first and decide what to
do.
## Try it
- [Repairing inverted validity windows](/schema-management#repairing-inverted-validity-windows)
- [Claim relations, and what they do not promise](/backend-setup#claim-relations-and-what-they-do-not-promise)
- [Contribution health: probe, repair, rebuild](/troubleshooting#contribution-health-probe-repair-rebuild)
- [Schema Management](/schema-management): migrations and kind removal
# An Agent That Grows Its Own Schema
import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro";
If you follow a clinical trial through the literature, you see it registered
first, then published as a paper, and occasionally, years later, retracted. A
retraction notice is a kind of record that wasn't in anyone's schema when the
trial was registered, and usually wasn't there when the paper was indexed
either.
TypeGraph's schema normally lives in code: you declare `defineNode` and
`defineEdge` in TypeScript, you get full type inference, and adding a new kind
of thing means a code change and a deploy. I think that's the right default and
I'm keeping it, but it doesn't fit an ingestion pipeline pulling from an API you
don't control, a multi-tenant app where each tenant brings its own shape, or an
agent that discovers a new kind of record halfway through a corpus.
0.25 adds **graph extensions** for those cases. A schema change can be proposed
at runtime, validated as strictly as the compile-time path, and committed
atomically without a redeploy.
## Proposing a change
An extension is a plain JSON document (new node and edge kinds, their property
types, unique constraints, indexes) built with `defineGraphExtension` and
committed with `store.evolve()`:
```typescript
const proposal = defineGraphExtension({
nodes: {
Paper: {
properties: {
title: { type: "string", minLength: 1 },
doi: { type: "string", minLength: 1 },
year: { type: "number", int: true, min: 1900, max: 2100 },
},
unique: [{ name: "paper_doi_unique", fields: ["doi"] }],
},
},
});
const evolved = await store.evolve(proposal);
```
The type vocabulary is small on purpose: strings, numbers, booleans, enums,
and one level of array/object nesting. An LLM-written schema stays readable,
and it can't smuggle in a Zod refinement or a function. A malformed proposal
throws `GraphExtensionValidationError` with per-field issues before anything
touches the database. The commit is checked against the active schema version,
so two writers racing to extend the same graph get a `StaleVersionError` to
retry on, not a silent overwrite.
Reads get the same care. TypeScript can't see a kind that didn't exist at
compile time, so runtime kinds go through string-keyed versions of the query
builder, which check kind names against the live schema:
```typescript
const rows = await store
.query()
.fromDynamic("Paper", "p")
.traverseDynamic("authoredBy", "a")
.toDynamic("Author", "u")
.select((ctx) => ({ paper: ctx.p, author: ctx.u }))
.execute();
```
A typo in a kind name throws `KindNotFoundError`, where a schemaless store
would quietly return zero rows and leave you to find out later.
## Letting an agent drive
A toy example is fine for the API, but I wanted to know whether the loop holds
up against messy, real, multi-stage data. So I built
[`typegraph-clinical-demo`](https://github.com/pdlug/typegraph-clinical-demo):
the same machinery end to end, over real trial registrations from
ClinicalTrials.gov, their publications from PubMed, and retractions and
corrections from CrossRef. It's all public bibliographic metadata, with no
patient data.
An LLM agent watches the corpus arrive in three stages and proposes an
extension after each one. The proposals are **blind**: the agent sees sample
records and the names of kinds already in the graph, and I didn't give it any
hints about types, searchable fields, or which identifier should be unique.
Property types, optionality, searchability, and constraints are all inferred
from the samples.
- **Stage 1, registration.** The agent sees about 1,100 ClinicalTrials.gov
records and proposes `ClinicalTrial`.
- **Stage 2, publication.** PubMed records referencing those trials arrive, and
the agent proposes `Publication` with a `referencesTrial` edge back to stage 1.
- **Stage 3, retraction.** Retraction notices and corrections arrive. Nobody
designs for these up front, because most trials never get one. The agent
proposes `PublicationEvent` with a `correctsPublication` edge back to stage 2.
Each stage is a real `store.evolve()` against a real store, followed by a bulk
ingest under the new schema.
## When the proposal is valid but wrong
Validation catches malformed proposals, but it can't catch one that's
internally consistent and still wrong for the data, like a field the agent marked
required because the three samples it saw all happened to have it. So after
each accepted proposal the demo runs a **smoke test**: ingest the full sample
into a scratch store seeded with every earlier extension. Failures go back to
the agent in the same structured `{path, code, message}` form validation uses.
Stage 3 is where it fires, on a real run:
```text
STAGE 3: post-publication discourse
agent attempt=1
validator ACCEPTED +nodes=[PublicationEvent] +edges=[correctsPublication]
smoke test FAILED 2 issue(s) — proposal validates but doesn't fit the data:
[INGEST_INVALID_TYPE] artifactDoi: expected string, received undefined
[INGEST_INVALID_TYPE] date: expected string, received undefined
agent attempt=2
validator ACCEPTED +nodes=[PublicationEvent] +edges=[correctsPublication]
smoke test PASSED schema fits the sample
PublicationEvent nodes: 922/922 ingested (0 skipped)
```
The first proposal saw three samples, all `correction` events with
`artifactDoi` and `date` filled in, and made both required. The full-sample
ingest hit `comment` events without them and failed. On attempt 2 the agent
made both optional, the smoke test passed, and the graph committed. The whole
stage, repair included, takes about 16 seconds and one extra model call.
I love watching this loop work. Nobody explained the mistake to the model in
prose; it got the same machine-readable errors a developer would have seen and
fixed its own schema. Because the smoke test runs against a scratch store, a failed attempt
leaves no stray rows or half-claimed unique keys for the next attempt to trip
over. Only a proposal that survives gets committed to the real store.
## The payoff
After stage 3 the graph has three kinds and two edges that didn't exist when
the demo started, and one query walks all of them:
```typescript
const rows = await store
.query()
.fromDynamic("PublicationEvent", "event")
.traverseDynamic("correctsPublication", "correction")
.toDynamic("Publication", "pub")
.traverseDynamic("referencesTrial", "reference")
.toDynamic("ClinicalTrial", "trial")
.select((ctx) => ({
event: ctx.event,
publication: ctx.pub,
trial: ctx.trial,
}))
.execute();
```
On the real corpus it surfaces four documented retraction chains, including
SCIPIO (cardiac stem cells, Lancet 2011, retracted 2019) and a 2018 nilotinib
trial outcome, plus three more the corpus turned up on its own: the Anil Potti
genomic-predictor case and the Mehra et al. hydroxychloroquine retraction that
halted multiple registered trials in 2020 among them. Getting there took no
migrations, just three `store.evolve()` calls and a query written against kind
names that were plain strings until the agent defined them.
## You don't need a frontier model for this
The demo includes an eval harness (`pnpm eval`) that runs the same blind
three-stage pipeline across a lineup of models and scores whether each first
proposal survives validation and the smoke test. The cheapest model that
cleared all three stages on the first try was a 26B-parameter open-weight MoE
with 4B active parameters. It came in 4x cheaper than Gemini 3.1 Flash Lite
(which also went three for three) and 5x cheaper than GPT 5.4 nano (which
needed one repair).
The full three-stage run, repair included, takes under 30 seconds and costs
about **$0.003** in API calls. The loop needs a structured error channel and a
model that can read a JSON Schema error and try again, and when the schema
layer does its job, a small model handles that fine.
## Try it
- [Graph Extensions](/graph-extensions): the full reference
- [Agent-Driven Schema](/examples/agent-driven-schema): a minimal, in-repo
version of the same loop
- [`pdlug/typegraph-clinical-demo`](https://github.com/pdlug/typegraph-clinical-demo):
clone it, `pnpm demo`, and watch the schema grow
- [GitHub](https://github.com/nicia-ai/typegraph)
# Schema Changes That Roll Back With Everything Else
import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro";
[Runtime Schema Evolution](/blog/runtime-schema-evolution) ended with an agent
adding a kind for retraction notices to a live graph. This post uses a cut-down
version of that: a `Retraction` kind added to a graph of publications.
`store.evolve()` commits a change like that in a transaction of its own and
hands back a new Store. That's fine when the schema change is the whole job,
but not when your service keeps its own bookkeeping next to the graph, say a
`schema_audit` table recording which schema version each import ran under.
If you do it in the obvious order, `evolve()` commits first, then a second
transaction writes the retraction row and the audit row. When something
downstream throws in that second transaction, this is what's left:
```text
evolve(), then a failed import:
active schema version 2
schema_audit rows 0
Retraction node rows 0
```
The schema is at version 2 but nothing else moved, so there's a `Retraction`
kind that no ledger entry mentions and no row uses. Nothing is corrupt, but
your bookkeeping and the database now disagree about which schema exists,
because the schema change was the one step that couldn't join the transaction.
As of 0.62 it can join, and the rest of this post shows how.
## Plan outside, apply inside
The setup is a graph with one kind, the extension the agent proposed, and a
Drizzle table for the audit ledger in the same database.
```typescript
const Publication = defineNode("Publication", {
schema: z.object({ doi: z.string(), title: z.string() }),
});
const graph = defineGraph({
id: "trials",
nodes: { Publication: { type: Publication } },
edges: {},
});
const retractions = defineGraphExtension({
nodes: {
Retraction: {
properties: {
doi: { type: "string", minLength: 1 },
reason: { type: "string" },
},
},
},
});
const schemaAudit = sqliteTable("schema_audit", {
id: integer("id").primaryKey({ autoIncrement: true }),
schemaVersion: integer("schema_version").notNull(),
recordedAt: text("recorded_at"),
});
```
The change is split into planning and applying. `planEvolution()` runs before
any transaction opens, validates the extension against the active schema
without writing anything, and returns an immutable plan that names the schema
it starts from and the one it produces:
```typescript
const plan = await store.planEvolution(retractions);
```
```json
{
"status": "change",
"graphId": "trials",
"baseline": { "version": 1, "hash": "89654f9e5c1debfb" },
"result": { "version": 2, "hash": "aec6b8d50bf21bf8" },
"requirements": [
{ "kind": "new-kind", "entity": "node", "kindName": "Retraction" }
]
}
```
Doing the planning outside means the slow, fallible part (validation, diffing)
never holds a lock. For ordinary additions like a new kind or an optional
scalar field, applying the plan needs no entity scans and no DDL.
Applying takes _your_ transaction. better-sqlite3 is synchronous, so Drizzle's
`db.transaction()` can't take an async callback, and a small helper drives
`BEGIN`/`COMMIT`/`ROLLBACK` on the one connection. With node-postgres or libSQL
you'd pass the `nativeTx` from `db.transaction(async (nativeTx) => …)` instead;
[the cross-store transactions recipe](/recipes#cross-store-transactions-drizzle--typegraph)
covers both.
```typescript
async function inTransaction(
db: BetterSQLite3Database,
run: () => Promise,
): Promise {
db.run(sql`BEGIN`);
try {
const result = await run();
db.run(sql`COMMIT`);
return result;
} catch (error) {
db.run(sql`ROLLBACK`);
throw error;
}
}
```
```typescript
const outcome = await inTransaction(db, async () => {
const applied = await store.withEvolvedTransaction(db, plan, async (tx) => {
const notices = tx.getNodeCollection("Retraction");
if (notices === undefined) throw new Error("Retraction kind missing");
await notices.create({ doi: "10.1/a", reason: "fabricated" });
});
db.insert(schemaAudit)
.values({
schemaVersion: applied.receipt.schema.version,
recordedAt: String(applied.receipt.recorded),
})
.run();
return applied;
});
const current = await store.refreshSchema({
minVersion: outcome.receipt.schema.version,
});
```
Inside the callback, `tx` already sees the new schema: `Retraction` is
writable even though the committed schema doesn't have it yet. The receipt
carries the exact schema version and hash the transaction produced (and, on a
history-enabled store, the recorded-time anchor), so the audit row stores what
the import actually ran under instead of what the code assumed. Treat the
receipt as provisional until your outer `COMMIT` succeeds; after that,
`refreshSchema()` hands the new schema to the Store you keep around.
To check the rollback, I threw an error after the ledger insert and ran the
same plan twice, once failing and once clean:
```text
attempt 1 (fails after the callback):
receipt: schema v2, recorded r1:0000000000000002:2026-09-20T19:12:56.201Z
after rollback:
active schema version 1
schema_audit rows 0
Retraction node rows 0
attempt 2 (same plan):
receipt: schema v2, recorded r1:0000000000000002:2026-09-20T19:12:56.202Z
after commit:
active schema version 2
schema_audit rows 1
Retraction node rows 1
```
The failed attempt got as far as a receipt and a ledger row visible inside the
transaction, and none of it survived; when I checked the recorded-history table
directly it had no `Retraction` rows either. The retry reused the same plan, got
the same recorded revision, and committed, so the schema, the rows, the
history, and your own table commit or roll back together.
If another writer evolves the graph between planning and applying, the apply
throws `StaleVersionError` and the transaction is yours to roll back and
replan.
## No-ops, and schema-only checkpoints
A pipeline that runs the same wiring on every deploy will mostly produce no-op
plans. Plan an extension the Store already has and you get
`status: "noop"`, which doesn't need the exclusive schema lock, so an ordinary
`withRecordedTransaction()` is enough:
```typescript
const again = await current.planEvolution(retractions);
if (again.status === "noop") {
await inTransaction(db, () =>
current.withRecordedTransaction(db, async (tx) => {
const notices = tx.getNodeCollection("Retraction");
if (notices === undefined) throw new Error("Retraction kind missing");
await notices.create({ doi: "10.1/b", reason: "duplicate publication" });
}),
);
}
```
The opposite case is a schema change that is itself the event worth recording.
On a history-enabled store, `tx.requestRecordedRevision()` puts the change on
the recorded timeline even when no entities change: the receipt reports zero
writes, schema version 2, and a recorded anchor.
## Merges and reviewed writes that bring their own kinds
0.61 added `applyMergePlanInTransaction()` for applying an approved merge plan
next to your own SQL. But a merge plan is built against a specific schema, so a
merge that introduces a kind the target doesn't have yet couldn't share a
commit with the evolution that adds it. 0.62 closes that gap:
`branchForEvolution()` forks a branch that already has the planned kinds, and
`planMergeForEvolution()` plans against the resulting schema.
```typescript
const evolutionPlan = await target.planEvolution(retractions);
const futureBranch = unwrap(
await branchForEvolution(target, evolutionPlan, makeIsolatedBackend),
);
try {
await futureBranch.store
.getNodeCollectionOrThrow("Retraction")
.create({ doi: "10.1/a", reason: "fabricated" });
const mergePlan = unwrap(
await planMergeForEvolution(target, evolutionPlan, [futureBranch]),
);
await inTransaction(db, () =>
target.withEvolvedTransaction(db, evolutionPlan, (tx) =>
applyMergePlanInTransaction(target, tx, mergePlan),
),
);
} finally {
await futureBranch.close();
}
```
A plan built the ordinary way, against the old schema, is rejected inside the
evolved transaction with `MergePlanSchemaMismatchError` before anything is
written.
0.65 does the same for candidate write sets, the branch-free way to propose
changes as a JSON document a reviewer can read.
`planCandidateWriteSetForEvolution()` plans the agent's proposed `Retraction`
records against the schema the pending evolution will produce, and nothing is
written until the reviewed plan is applied inside `withEvolvedTransaction()`.
Because review takes time and the target can move in the meantime, there are
two different stale outcomes. If a write lands _while_ the plan is being built,
planning returns a `MergePlanningStaleError`, and you recapture the target and
replan. If the target changes after you already hold a finished plan, applying
throws `StaleMergePlanError` inside the outer transaction, and the schema
change rolls back with it. I injected a write after planning to check this, and
the active schema was still at version 1 afterwards.
## Limits
- **Plans stay in one process.** A plan is an in-memory token that can't be
serialized or reconstructed, so plan and apply in the same process.
- **The default adapter only does DML.** A plan that needs new storage (a
vector slot, identity work) is rejected before the callback runs, unless you
use a privileged adapter configured with
`schemaProvisioning: "transactional"`. Indexes
stay outside too: call `materializeIndexes()` after commit.
- **Only your database's writes are atomic.** The ledger row is atomic
with the schema change because they share a connection. A message you
publish to a queue from inside the callback is not.
- **Driver support varies.** SQLite adoption needs a native connection with an
observable transaction state (better-sqlite3 has one); HTTP-only drivers
can't adopt schema transactions at all.
Custom `StoreEvolution` implementations need `planEvolution()` and
`refreshSchema()`, and custom adapters must now declare `schemaProvisioning`
explicitly.
## Try it
- [Plan outside and apply inside a caller-owned transaction](/graph-extensions#plan-outside-and-apply-inside-a-caller-owned-transaction)
- [Merge after schema evolution in one caller transaction](/graph-merge#merge-after-schema-evolution-in-one-caller-transaction)
and [candidate write sets for a planned schema](/graph-merge#candidate-write-sets-for-a-planned-schema)
- [Adopted schema transactions](/backend-setup#adopted-schema-transactions):
per-driver requirements
- [GitHub](https://github.com/nicia-ai/typegraph)
# One Round Trip Per Write
import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro";
I turned on statement logging, turned off prepared statements, and created one
node. This is what went over the wire:
```
begin
SELECT ... schema-version fence probe
SELECT ... duplicate/endpoint check
INSERT ... (the write)
commit
```
That's five round trips to write one row. On a pooled connection sitting next
to the database you'd never notice, but a lot of people run TypeGraph from
edge functions against managed Postgres, where a round trip is more like 45ms
and each of those statements pays it in full, so the row takes roughly a
quarter of a second to exist.
When I captured a first-run provisioning flow for one graph, it issued 97
statements, and about 40 of them, spread across just seven writes, were
schema-version probes, duplicate checks, and `begin`/`commit` rather than
actual writes.
That was [issue #533](https://github.com/nicia-ai/typegraph/issues/533), and
closing it took two releases. The short version: an eligible write and every
check it depends on now go to the database together, as one request.
A note on the numbers: everything below is a count of requests crossing the
transport, which is what I measured. Latency figures are that count times an
assumed 45ms round trip, not wall-clock benchmarks.
## Checks and write, in one statement
The old path was a sequence of reads, each deciding whether the write could
proceed: whether the schema version still matched, whether the row was a
duplicate, whether the edge's endpoints existed, and whether the cardinality
constraint held. Each of those was another round trip, and the whole sequence
was wrapped in a transaction so nothing could change between the checks and
the insert.
In 0.52, for the common shape on Postgres (schema-managed, generated ID, no
history or identity tracking), all of those checks and the insert compile into
a single statement, and the database evaluates the conditions and performs
the write atomically.
That takes **five or six sequential requests down to one**, which at 45ms
removes roughly 180–225ms of waiting from every such write. On Neon's WebSocket
driver, a typical get-or-create that misses drops from five requests to one.
When a write needs something a single statement can't carry (history capture,
a caller-owned transaction, call-level `matchOn`), it takes the ordinary
transactional path, which is still there underneath for everything the fast
path doesn't cover.
## Batches on drivers without transactions
The bigger win is on Neon HTTP, Cloudflare D1, and libSQL, which don't give
you an interactive transaction at all, only one-shot atomic batches. Before
0.52, a bulk write on one of them ran statement by statement, so it was slow
and a partial failure left partial data behind.
Now TypeGraph compiles eligible bulk writes into a precompiled program that
those drivers execute as a single atomic batch. `nodes.bulkInsert()` becomes
one request, edge batches validate their endpoints inside the write instead of
reading candidate endpoints first, and batches with a cardinality constraint
(`one`, `unique`, `oneActive`) carry the constraint checks in the same batch.
Counting actual `execute()` / `batch()` calls on libSQL:
| Bulk edge write | Before | After |
| ----------------------------- | -----: | ----: |
| Unconstrained | 1 | 1 |
| With a durable match identity | 6 | 1 |
| With a cardinality constraint | 8 | 1 |
Since it's all one batch, a conflict anywhere rolls the whole thing back. I
wanted proof that the conflict check actually did something, so I wrote its
regression test by deleting the conflict clause, watching duplicates land,
and then putting the clause back and watching the same input get rejected.
## 0.53: the writes that have to look first
0.52 fused the writes that create things, which left the ones that need to
know what's already there: updates, upserts, and deletes that have to release
the uniqueness claims their row held. 0.53 moved those onto the same programs.
- **Single-row `update()` and `delete()`** are one read plus one guarded
write, rather than a transaction around both: 60–67% fewer requests.
- **`bulkUpsertById()`** is two requests: one batched read, one atomic write.
- **`bulkReplaceById()`** is new. Each item is a complete document, so there's
nothing to read first. Creates, replacements, resurrections of deleted
rows, uniqueness claims, and fulltext and vector index updates all go in one
request.
- **Large D1 upserts stay atomic.** D1 allows 100 bind parameters per
statement, so the program splits into many statements inside one atomic
batch. That raised the ceiling from 17 nodes and 6 edges to 512 nodes and
187 edges per call. Anything bigger falls back to the regular path instead
of building an unbounded request.
- **Postgres transaction sessions run the same programs** through a
savepoint, so a rejected program doesn't poison the transaction around it.
There's one trade-off to know about: the fused `update()` is optimistic. It
doesn't hold a lock between its read and its write, so if the row moved
underneath it, it re-reads and tries again, up to four times, and then throws
`DatabaseOperationError`. Under sustained contention on the same row, that can
fail a write that a transaction-capable backend used to serialize for you. I
think that's the right trade for most workloads, but not for something like a
hot counter.
## Letting the database decide: match identities
Fusing an edge get-or-create needed something the database could arbitrate on
its own, without a lock or an in-process cache: a real unique constraint. So
0.52 also added **durable edge match identities**. An edge can declare a named
set of fields that identify it:
```typescript
const worksAt = defineEdge("worksAt", {
schema: z.object({ role: z.string() }),
});
edges: {
worksAt: {
type: worksAt,
from: [Person],
to: [Company],
cardinality: "many",
matchIdentity: { name: "employment", fields: ["role"] },
},
},
```
TypeGraph stores that key on every edge row and backs it with a unique
index, on both SQLite and Postgres. Once that exists,
`getOrCreateByEndpoints` no longer has to read, decide, write, and hope
nothing raced, because it's a single conditional insert the database resolves
atomically, which is what makes the one-request path safe on a stateless edge
worker with nothing in memory to lean on.
Changing a match identity is a breaking schema change and is rejected while
the edge kind has rows. Call-level `matchOn` still works for matching you
don't want in the schema, through the regular transactional path.
## Upgrading
If you use a bundled backend (`createLocalSqliteBackend`,
`createPostgresBackend`, or one of the serverless factories), you don't need to
change any code.
There are two things to know. An existing or externally provisioned database has
to be opened once through `createStoreWithSchema()` (or get TypeGraph's
generated base-schema migration) before the zero-DDL verified-store paths will
run; until then they fail early with `BaseSchemaMigrationError`. And
`store.transaction()` now throws on a backend without interactive transactions
instead of quietly running your callback without one.
If you maintain a custom backend, `GraphBackend.commands` is now required and
a few capability flags changed shape. The
[authoritative command sessions](/backend-setup#authoritative-command-sessions)
section has the migration.
## Try it
- [Batch write patterns](/performance/overview#batch-write-patterns) and
[remote edge convergence](/performance/overview#remote-edge-convergence):
which writes are eligible, and the measured counts
- [Upgrading deployment-wide base storage](/backend-setup#upgrading-deployment-wide-base-storage)
- [Changelog](/changelog#0530) for 0.53.0, and [0.52.0](/changelog#0520)
- [GitHub](https://github.com/nicia-ai/typegraph)
# Letting an Edge's Targets Depend on Its Source
import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro";
This schema looks reasonable, but it allows more than you probably meant:
```typescript
const assignedTo = defineEdge("assignedTo", {
from: [Employee, Student],
to: [Department, Course],
});
```
You want employees assigned to departments and students assigned to courses,
but array-valued `from` and `to` declare the Cartesian product, so every
source may point at every target and both
`store.edges.assignedTo.create(student, department)` and
`create(employee, course)` succeed. Until 0.55 the only way to rule those two
combinations out was to split the edge into two kinds and query them
separately.
## Say which source gets which targets
`to` can now be a map from source kind to its own allowed targets:
```typescript
const assignedTo = defineEdge("assignedTo", {
from: [Employee, Student],
to: {
Employee: [Department],
Student: [Course],
},
});
const graph = defineGraph({
id: "assignments",
nodes: {
Employee: { type: Employee },
Student: { type: Student },
Department: { type: Department },
Course: { type: Course },
},
edges: { assignedTo },
});
```
`create(employee, department)` and `create(student, course)` still work, and
the other two combinations fail before anything is written:
```text
EndpointPairError: assignedTo: undeclared endpoint pair
{ edgeKind: "assignedTo", endpoint: "pair",
fromKind: "Employee", toKind: "Course",
allowedPairs: [
{ from: "Employee", to: "Department" },
{ from: "Student", to: "Course" },
] }
```
In typed code you won't get that far, because the collection's `create()`
signature narrows per source and handing it an employee and a course is a
compile error. The runtime check covers the paths a type checker can't see, such
as dynamic collections, bulk writes, and imports. An endpoint kind that isn't in
`from` at all still throws the existing `EndpointError`; `EndpointPairError` is
specifically for two kinds that are each valid but were never declared together.
The map itself is checked when you call `defineEdge()`: every kind in `from`
needs an entry, extra keys aren't allowed, and no entry may be empty. A
mistake throws `ConfigurationError` at that point rather than turning up later
as a confusing write failure.
## Narrowing, never widening
An edge's map is its outer bound, and a graph or a runtime
[graph extension](/blog/runtime-schema-evolution) can register a narrower
version:
```typescript
edges: {
assignedTo: {
type: assignedTo,
from: [Employee],
to: { Employee: [Department] }, // this graph has no students
},
}
```
A registration can't add a pair the edge never declared. That includes the
easy mistake of registering the old array form, `to: [Department, Course]`,
against an edge whose map only allows the correlated pairs, which would
quietly re-admit the cross-pairs the map is there to forbid, so it's rejected.
Subclasses are checked against the declared pairs. With `SubTask subClassOf
Task` and an edge that allows `Task: [Task]`, every mix of `Task` and `SubTask`
works, but a `SubTask` can't borrow a target that only some other source kind is
allowed.
## Bulk writes are all or nothing
```typescript
await store.edges.assignedTo.bulkCreate([
{ from: employeeA, to: departmentA }, // valid
{ from: employeeB, to: courseA }, // undeclared pair
]);
// throws EndpointPairError; store.edges.assignedTo.count() is still 0
```
A single undeclared pair rejects the whole batch, because committing the
valid half and quietly dropping the rest would leave you guessing which rows
made it.
## Pairs are stored in the schema
The pairs are stored in the serialized schema, and declaring them in a
different key order doesn't change the schema hash. Replacing an array `to`
with a map removes pairs even though it removes no kinds, so it counts as a
breaking change and follows the same migration rules as any other breaking
edge change. See
[endpoint pair changes](/schema-management#endpoint-pair-changes).
## Try it
- [Source-Dependent Targets](/core-concepts#source-dependent-targets): the
full reference
- [`EndpointPairError`](/errors#endpointpairerror)
- [Endpoint pair changes](/schema-management#endpoint-pair-changes)
- [GitHub](https://github.com/nicia-ai/typegraph)
# Look Before You Write
import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro";
Suppose you tighten the schema on an account directory so that `email` has to
be a real address, `plan` has to be one of three values, and `seats` has to be
positive. TypeGraph validates on write, so everything written from then on is
checked, but rows stored before the change are never looked at again, and you
have no idea how many of them fail the new rules. You'd like a background
agent (or yourself, in a REPL) to find and fix them while sales reps carry on
editing the same accounts.
To do that safely the agent needs a cheap overview of what's in the graph, a
way to find the rows that don't validate without reading everything back into
memory, and a write that won't clobber a change a person made after the agent
read the row.
0.54 added `store.describe()`, `store.validateStore()`, and `compareAndSet()`
for those three jobs, plus `planCandidateWriteSet()` for reviewing a whole
batch before it lands, and 0.64 lets the first two run inside a transaction.
It's the same loop [runtime schema evolution](/blog/runtime-schema-evolution)
uses when an agent proposes changes to the schema itself: the agent looks at
the current state, proposes a change, and the change only applies if nothing
has moved since it looked.
The graph here is 240 accounts and 90 people created through the store, 60
`memberOf` edges, and three legacy `Account` rows inserted through the backend
directly, the way a writer running the old rules would have stored them. The
outputs below come from running this code against it on SQLite.
```typescript
const Account = defineNode("Account", {
schema: z.object({
name: z.string().min(1),
email: z.email(),
plan: z.enum(["free", "team", "enterprise"]),
seats: z.number().int().positive(),
ownerId: z.string().optional(),
}),
});
```
## What's in the graph
```typescript
const { statistics } = await store.describe();
for (const kind of [...statistics.nodes, ...statistics.edges]) {
console.log(`${kind.entity} ${kind.kind}: ${kind.count}`);
for (const property of kind.properties) {
console.log(
` ${property.path} present=${property.presentCount} coverage=${property.coverage.toFixed(2)}`,
);
}
}
```
```text
node Account: 243
/email present=243 coverage=1.00
/name present=243 coverage=1.00
/ownerId present=60 coverage=0.25
/plan present=243 coverage=1.00
/seats present=243 coverage=1.00
node Person: 90
/name present=90 coverage=1.00
/title present=30 coverage=0.33
edge memberOf: 60
/role present=30 coverage=0.50
```
Every declared kind gets a row count, and every declared property gets a
presence count and a coverage figure. `/ownerId` at 0.25 says 60 of 243
accounts have an owner, which is enough for an agent to decide where to spend
its attention before it reads a single row.
The counting happens in the database, and on this graph `describe()` runs two
statements, one for all node kinds and one for all edge kinds. The result also
records the schema version and hash it was computed under, and TypeGraph
checks that they didn't change mid-way, so you never get statistics that
straddle a schema change.
## What no longer fits
```typescript
let cursor: string | undefined;
do {
const page = await store.validateStore({
entity: "node",
kind: "Account",
pageSize: 100,
...(cursor === undefined ? {} : { cursor }),
});
console.log(
`scanned ${page.scannedCount}, ${page.violations.length} violation(s)`,
);
for (const failure of page.violations) {
console.log(failure.id, failure.path, failure.reason);
}
cursor = page.nextCursor;
} while (cursor !== undefined);
```
```text
scanned 100, 0 violation(s)
scanned 100, 0 violation(s)
scanned 43, 3 violation(s)
acct_legacy_1 /email Invalid email address
acct_legacy_2 /plan Invalid option: expected one of "free"|"team"|"enterprise"
acct_legacy_2 /seats Too small: expected number to be >0
```
There are three violations across two records, because `acct_legacy_2` breaks
two rules. Each failure carries the record id, a JSON pointer to the property,
the Zod issue code, and the reason. Each page is one bounded statement, and the cursor is tied to the
schema it started under. If the schema changes mid-sweep, the next page throws
`StoreAnalysisCursorStaleError` instead of quietly checking the rest against
different rules.
The third legacy row, `acct_legacy_3`, isn't reported. It carries a
`salesforceId` the schema doesn't declare, and undeclared properties count as
healthy extra data, so you can sweep a graph whose shape is changing at runtime
without every extra field showing up as a defect.
## Fix a row without racing anyone
The obvious repair is to call `getById()`, decide what to change, and call
`update()`, but a sales rep's edit that lands between the read and the write
gets overwritten. `compareAndSet()` closes that gap by checking the expected
values and applying the update in one statement.
```typescript
const applied = await store.nodes.Account.compareAndSet(id, {
expected: { name: "Northwind", plan: "team", seats: 12 },
patch: { email: "ops@northwind.example" },
});
```
```text
applied=true email=ops@northwind.example version 1 -> 2
```
The interesting case is when someone else gets there first. Say the agent read
`acct_legacy_2` and planned a repair, but before it applied, a person fixed
the row by hand and changed the email along the way. The agent's guard names the email it read:
```typescript
const applied = await store.nodes.Account.compareAndSet(id, {
expected: { name: "Contoso", email: "it@contoso.example" },
patch: { plan: "team", seats: 5 },
});
```
```text
applied=false seats=20 version 2 -> 2 updatedAt changed=false
```
`false` means nothing was written, so the person's `seats=20` stands and the
version didn't move. It's an ordinary return value rather than an exception,
and the agent should re-read the row and plan again.
To require that a property is _missing_, use the exported
`compareAndSetAbsent` marker (`undefined` is too easy to lose while an object
is being built or serialized). Here are two agents trying to claim the same
unowned account:
```typescript
const claim = (ownerId: string) =>
store.nodes.Account.compareAndSet("acct_1", {
expected: { ownerId: compareAndSetAbsent },
patch: { ownerId },
});
console.log(await claim("rep-9"), await claim("rep-4"));
```
```text
true false
```
The first agent gets the account and the second finds out immediately,
without a lock or a retry loop.
Expected values are checked against the property's current schema, so you
can't guard on a value that's already invalid, which is why the repair above
guards on `name`, `plan` and `seats` rather than the bad email. The row also
has to be valid as a whole after the patch, so patching only `plan` on
`acct_legacy_2` fails on `seats` before anything is written.
## Review a batch before it lands
A guard works for one row, but a partner feed offering a batch of accounts
is more of a review problem, where you want to see what would change before
anything does. `planCandidateWriteSet()` takes a serializable batch of nodes
and edges, stages it on a throwaway copy, and hands back an ordinary merge plan
without ever writing to the target.
```typescript
const plan = unwrap(
await planCandidateWriteSet({
target: store,
makeBackend,
options, // two Accounts with the same email are the same account
writeSet: {
formatVersion: 1,
sourceId: "partner-feed",
target: await captureCandidateWriteSetTarget(store),
nodes: [partnerAccount7, partnerGlobex],
edges: [],
},
}),
);
```
```text
upserts: 2 accounts still 243
acct_7 name: keep "Account 7", feed says "Account 7 Holdings"
acct_7 plan: keep "team", feed says "enterprise"
acct_7 seats: keep 12, feed says 40
```
The plan has one new account and one match on `acct_7`, where the feed
disagrees on three properties. The existing values win by default and the
disagreements go on the review, so "enterprise, 40 seats" becomes a line item
for a person to look at instead of a silent overwrite.
If something unrelated writes to the graph before you apply, applying the
original plan fails:
```text
StaleMergePlanError: The target revision changed after this merge plan was created; the plan was not applied. accounts 243
```
Re-planning and applying takes the count to 244. The plan is tied to the
target's revision rather than only the rows it touches, so any write makes it
stale, which is cheap to recover from for a bounded batch. When a person's
approval has to sit between planning and applying, [durable review](/blog/durable-merge-branches)
handles it.
## One consistent snapshot
At the root store, each `describe()` statement and each `validateStore()` page
is its own read. TypeGraph catches schema changes between those reads but not
data changes, so when a sweep needs one consistent view, 0.64 puts both methods
on the transaction context:
```typescript
const summary = await store.transaction(
async (tx) => {
const { statistics } = await tx.describe();
// ...page through tx.validateStore() exactly as above
},
{ isolationLevel: "repeatable_read", accessMode: "read_only" },
);
```
```text
{ accounts: 244, scanned: 244, violations: [] }
```
The violations are gone because both repairs landed earlier in the run. Use
`repeatable_read` or `serializable` and consume every page inside the
callback.
## Limits
- **`describe()` covers directly addressable declared properties.** It
doesn't guess through `$ref`, unions, arrays, or conditionals.
`validateStore()` is the authority for those.
- **Both are current-state only.** There's no "what did the data look like
last month" analysis.
- **A guard can only protect what it can name.** You can't guard on an
invalid value, so if a person and an agent both repair the same bad field,
a guard on its neighbors won't tell them apart. The revision-fenced plan is
the alternative there.
- **Planning copies the whole target,** so its cost scales with the graph.
It's for bounded review workflows, not a hot path.
## Upgrading
Writing this post turned up two bugs in 0.68.0 and earlier.
`planCandidateWriteSet()` fails on a target that holds a row with undeclared
properties, like `acct_legacy_3` above
([#733](https://github.com/nicia-ai/typegraph/issues/733)), and
`compareAndSet()` throws a raw Zod error on a kind whose schema has an
object-level `.refine()`
([#734](https://github.com/nicia-ai/typegraph/issues/734)). Both are fixed in
0.68.1 ([#735](https://github.com/nicia-ai/typegraph/pull/735),
[#737](https://github.com/nicia-ai/typegraph/pull/737)), which the examples
here assume.
`BulkOperationHookContext["operation"]` now includes `"compareAndSet"`, so an
exhaustive `switch` over it needs the new case. If you run mixed versions
during a rollout, upgrade every process that writes schema versions to 0.54
before using graph-scoped annotations (also new in 0.54: metadata on the graph
itself, carried through extensions and returned by `describe()`). Older
writers drop fields they don't recognize.
## Try it
- [`describe()` and `validateStore()`](/graph-extensions#population-statistics-and-stored-data-validation)
- [`compareAndSet()`](/schemas-stores#compareandsetid-params) and
`compareAndSetAbsent`
- [`planCandidateWriteSet()`](/graph-merge#constraint-aware-ingestion-branches)
- [GitHub](https://github.com/nicia-ai/typegraph)
# Agent Memory That Knows Why It Believes Things
import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro";
Say a vuln-feed agent in your overnight security pipeline flags 60 services
as shipping a vulnerable library, blocks their deploys, opens fix PRs, and
pages the owning teams. In the morning someone notices that the feed had a bad
package-name mapping, which pinned a real CVE to the wrong package across a
whole class of entries. What should happen to those 60 flags?
Deleting all of them is wrong, because some were independently confirmed by a
SAST run or an SBOM check and those services really are vulnerable. Leaving
them is also wrong, because most were only ever the feed talking, and they're
blocking shipments and paging people for nothing. What you need is the blast
radius: which of the 60 flags depend only on the feed, and which have other
evidence behind them.
Most agent memory can't answer that, because "memory" today usually means
retrieval: embed everything, fetch the nearest neighbors, and let the model
sort it out. That works for recall, but a vector store only keeps what was
said. It has no record of _why_ something was concluded, or of what should
happen to a conclusion when one of its sources turns out to be wrong. You
could live with that when an agent was a single chat transcript, but not once
a fleet of agents writes into shared memory and acts on it.
Over the last few releases I built the layer that answers the question: a
second temporal axis (0.33), provenance-backed retraction on top of it
(0.34), and the hardening that makes both hold up under real load (0.39 and
0.40). I'm really happy with how this one came out.
## The blast radius, as a query
With the provenance layer, the support structure is right there in the
graph:
```plaintext
VulnFeed ──▶ Vulnerable(svc-14) ──▶ BlockDeploy(svc-14)
SASTRun ──▶ Vulnerable(svc-14) two sources, survives
VulnFeed ──▶ Vulnerable(svc-22) ──▶ BlockDeploy(svc-22)
feed only, dies
```
Retract the feed and `Vulnerable(svc-22)` loses its only source, so it goes
non-current. That takes away the premise of `BlockDeploy(svc-22)`, which goes
non-current too, and the deploy unblocks. `Vulnerable(svc-14)` survives
because the SAST run still backs it, so that block stands.
```typescript
const report = await provenance.retract({
kind: "VulnFeed",
id: badMappingSnapshotId,
});
// report.died:
// flags and deploy blocks that had only the feed behind them
//
// report.survivedVia:
// flags still confirmed by SAST or the SBOM check
```
Run that across all 60 services and `report` is your cleanup list: `died`
tells you which deploys to unblock and which PRs to close, and `survivedVia`
tells you which flags are still real. I wouldn't trust anyone to work that out
by hand at 7am.
The source doesn't have to be a feed; it can be another agent's run. If a
triage agent combined an overnight feed-ingest run, a SAST scan, and an SBOM
rebuild into block-and-page decisions, and the feed-ingest run turns out to
have trusted bad data, you retract that run, not the whole fleet's memory:
```typescript
const report = await provenance.retract({
kind: "AgentRun",
id: overnightFeedRunId,
});
```
Flags raised only by that run go non-current, while flags another run also
supports survive, so one agent's mistake gets cleaned up without disturbing
what the others established.
## How retraction works
The `@nicia-ai/typegraph/provenance` subpath maps your existing graph kinds
onto four roles (sources, justifications, facts, and the edges between them)
and gives you a `retract` that recomputes support instead of deleting
blindly:
```typescript
import { createRetractionCapability } from "@nicia-ai/typegraph/provenance";
const provenance = createRetractionCapability(store, {
source: { kinds: ["ScannerSource", "VendorSource"] },
justification: { kind: "Justification" },
fact: { kinds: ["Vulnerability", "DeployDecision"] },
premiseOf: { kind: "premiseOf" },
derives: { kind: "derives" },
});
```
A fact stays believed while at least one of its justifications has all its
premises still supported. Premises bottom out at sources. Retract a source
and every justification that leaned on it stops counting; a fact loses
currency only when it runs out of surviving justifications.
Two properties make this safe to use. Retraction is **scoped**, touching
only facts reachable from the sources that flipped, and it's **reversible**:
it changes whether a fact is believed without deleting anything, and the
fact's edges stay put, so `unRetract` is an exact inverse of `retract`.
None of this is new theory. The storage follows the JTMS shape from Doyle's
1979 _A Truth Maintenance System_: AND-justifications over premises, sources
as the base case, a fact believed only if some justification has all its premises supported. The
question `retract` answers (which facts keep support after a source drops
out) is the ATMS question from de Kleer's 1986 work, and all of it runs on ordinary
SQL. It covers the well-founded, monotonic part of classic truth maintenance
rather than all of it. The new part is where it lives: your agent keeps
writing ordinary graph data and gets retraction semantics without moving to a
dedicated reasoning engine.
## Replaying what the agent believed
Retraction is much more useful with history. You want "why did the agent
block svc-22 at 2am, and why doesn't it anymore?" to be a query rather than a
dig through logs, and recorded time is what makes that possible.
TypeGraph already had valid time: when a fact was true in the world, set
with `validFrom` / `validTo` and read with `store.asOf(T)`. 0.33 adds the
second axis, **recorded time**: when TypeGraph captured a fact, what the
system knew as of a commit. It's the SQL:2011 `FOR SYSTEM_TIME` axis, or
Datomic's system time. Turn it on per store:
```typescript
const store = createStore(graph, backend, { history: true });
```
Writes through that store are captured with a per-graph, monotonic commit
anchor. `store.asOfRecorded(T)` gives you a read-only view of the graph as
of that anchor, and `store.recordedNow()` hands you the current one. The two
axes compose:
```typescript
store.asOf(validTime).asOfRecorded(recordedTime);
```
Retraction is a normal write under `history: true`, so the whole before and
after replays:
```typescript
const before = await store.recordedNow();
await provenance.retract({
kind: "VulnFeed",
id: badMappingSnapshotId,
});
const after = await store.recordedNow();
await store.asOfRecorded(before).nodes.BlockDeploy.getById("svc-22"); // current
await store.asOfRecorded(after).nodes.BlockDeploy.getById("svc-22"); // non-current
```
One naming note, since it trips up people coming from other systems:
TypeGraph's `asOf` is valid time, the reverse of SQL:2011's
`FOR SYSTEM_TIME AS OF` and Datomic's `d/as-of`, where a bare "as of" means
system time. Valid-time reads are the common case here, so they got the
short name.
## Why plain SQL, and what it costs
There's no portable, engine-native system versioning across Postgres and
SQLite. Postgres needs an extension or an application-level pattern, and
SQLite has nothing. So TypeGraph stores history in its own tables and
reconstructs point-in-time views in the query compiler. One implementation
runs on both backends, so the same memory model works from a solo agent's
local SQLite file up to a fleet writing into shared Postgres.
That isn't free:
- **Only TypeGraph-managed writes are captured.** Raw `tx.sql` is disabled
on a history store, since it would bypass capture. This is an audit layer
for graph writes, not database-level CDC.
- **No backfill.** Enable history on a fresh graph. An entity that already
existed is first recorded the next time it's written.
- **Recorded reads are an audit and replay tool, not a hot path.** They
reconstruct from history, so they're slower than current-state reads, and
broad reads like `find`, `search`, and vector predicates aren't available
at a past anchor because those indexes only reflect the present.
- **Writes cost more.** Roughly 2.5–6× an uncaptured write when each write is
its own transaction, dropping to about 1–1.5× when writes are batched.
## Hardening it: 0.39 and 0.40
Once real consumers started reading this axis, two gaps showed up.
**Commits in the same millisecond.** 0.33 ordered commits by wall-clock
timestamp alone, which breaks exactly when writes get fast. Five
back-to-back writes:
```text
commit 0 anchor: r1:0000000000000001:2026-07-21T17:59:03.914Z
commit 1 anchor: r1:0000000000000002:2026-07-21T17:59:03.914Z
commit 2 anchor: r1:0000000000000003:2026-07-21T17:59:03.915Z
commit 3 anchor: r1:0000000000000004:2026-07-21T17:59:03.915Z
commit 4 anchor: r1:0000000000000005:2026-07-21T17:59:03.915Z
```
Commits 0 and 1 share `17:59:03.914Z`. With a timestamp-only anchor, "the
graph right after commit 0 but before commit 1" doesn't have an answer.
0.40 anchors are `r1::`: the revision is a strict
per-graph counter, so every commit gets its own position no matter how many
share a millisecond. The timestamp is still there for display, and it never
moves backwards, even if the system clock does. A raw ISO string no longer
type-checks as an anchor, on purpose; get anchors from `recordedNow()`.
Deployments that adopted the earlier format run `migrateLegacyRecordedTime()`
once before opening the upgraded store.
**Walking a whole snapshot.** 0.39 adds `scan()` to recorded collections:
bounded, deterministic pages (up to 1,000 per call) ordered by id, with a
cursor bound to the exact view it came from:
```typescript
const view = store.asOfRecorded(anchor);
const page1 = await view.nodes.Item.scan({ limit: 2 });
const page2 = await view.nodes.Item.scan({ limit: 2, after: page1.nextCursor });
```
Together those make "every row as of commit N, in order, one page at a time"
a supported operation for a replication tap, an audit export, or a search
index rebuild pinned to a point in history.
## Why agent memory needs this
A fact stored without provenance is organizational hearsay: the memory can't
tell you why it believes something or what should change when a source goes
bad. A single chat agent can get away with that because a lot of
inconsistency hides inside one transcript, but agents that share memory will
read stale feeds, trust bad mappings, and inherit each other's wrong
assumptions.
Memory that gates deploys, or maintains any long-lived model of a company,
needs to answer three questions: what do we believe, why do we believe it,
and what changes if this source turns out to be wrong. This layer is built to
answer those, and it runs on the SQL database you already have.
## Try it
- [Provenance and Retraction](/provenance) and the
[Provenance Retraction example](/examples/provenance-retraction/)
- [Bitemporal Time Travel example](/examples/bitemporal-time-travel/)
- [Logical revision and physical time](/queries/temporal#logical-revision-and-physical-time):
the anchor format and diagonal bitemporal reads
- [Migrating preview recorded time](/schema-management#migrating-preview-recorded-time)
- [GitHub](https://github.com/nicia-ai/typegraph)
If you've mapped JTMS/ATMS-style truth maintenance onto relational storage
elsewhere, or know of prior art that does, I'd like to hear about it. There's
plenty written on truth maintenance systems and not much on doing it
directly on ordinary SQL tables.
# TypeGraph 0.35: Faster Almost Everywhere
import NewsletterSignup from "../../../../components/landing/NewsletterSignup.astro";
A 2M-row bulk load into SQLite was still running after 4.5 hours with no sign
of finishing, and the cause turned out to be a statistics refresh. After a
large batch, `bulkCreate` and `bulkInsert` automatically run SQLite's `ANALYZE`
so the query planner has fresh numbers, but that `ANALYZE` was bare and
unscoped: it scanned every table in the database file (not just TypeGraph's),
with no limit, after every big batch. If you stream a load through repeated
`bulkInsert()` calls, each refresh costs more than the last and the total goes
quadratic.
It's now scoped to TypeGraph's own tables and bounded with `PRAGMA
analysis_limit`, the way the Postgres version already behaved. A 100k-row
reproduction of the same shape finishes in about 8 seconds, and the last
batch costs only about twice what the first one did.
0.35 has about 35 fixes like that one, and below are the ones that matter
most, with numbers.
## Bulk writes
These four changes stack on top of each other for anything that writes a lot
of rows:
- **SQLite pragmas at open.** `createLocalSqliteBackend` now turns on
`journal_mode=WAL`, `synchronous=NORMAL`, and a busy timeout by default.
better-sqlite3's own defaults pay a full fsync per write in
rollback-journal mode. Single-operation writes on file-backed databases are
roughly **5x faster**.
- **Real bind-parameter limits.** The backend used to assume SQLite's old
999-parameter limit on every driver. It now asks the driver (32,766 for
better-sqlite3, 100 for Cloudflare D1), so batches on better-sqlite3 use
**~33x fewer statements**: 111-row chunks became 3,640-row chunks.
- **`bulkCreate` / `bulkInsert` batched end to end.** Existence checks,
uniqueness checks, and fulltext/embedding writes each happen once per batch
instead of once per row: **~1,600 → ~4,100 rows/s (~2.6x)**.
- **`importGraph` batched the same way: ~26k → ~96k entities/s (~4x).** The
default `batchSize` also went from 100 to 1,000, and now actually applies.
A schema-parsing gap meant the old default was silently ignored. That change
alone took a 20k-node, 5k-edge Postgres import from 1,515ms to 781ms.
There's also a new `walAutocheckpointPages` option for tuning WAL checkpoints
during heavy loads. On its own it cut a 2M-row load's time by more than half
at the largest scale tested.
## The regression, and fixing it properly
This one's on me. 0.34 fixed a real correctness bug where "current" reads
compared against the _database's_ clock instead of the _application's_, so
when the two clocks drifted (app and database on separate hosts, which is
normal), a row you'd just created could be invisible to the very next read.
The fix was to bind the read instant fresh every time a query compiled, which
was correct but had two problems I didn't catch until this release.
**It was slow.** Every `execute()` recompiled the query from scratch,
including the `.prepare()`-once, `.execute()`-many pattern that exists
specifically to avoid that. A repeated point query cost about 47µs.
**It was also, briefly, worse than the bug it fixed.** The query builders
cached their compiled SQL text across calls, so a prepared query froze "now"
at the moment it first compiled. Every row created after that had a later
`valid_from` than the frozen instant, and stayed invisible to that query
forever. It reproduced in a single process on the very next insert after
preparing a query, whereas the clock-skew bug at least needed two hosts.
0.35 fixes both the same way: the SQL text is cached, and only the "now"
instant is re-bound as a parameter on each call. The repeated point query
went from **47µs to 2.4µs (~20x)**, and a row created after `.prepare()`
shows up on the next `.execute()`.
I'm not going to bury that in a changelog line. If you're on 0.34 and use
`.prepare()`, upgrade.
## Point operations and traversals
- **CRUD statements reuse the prepared-statement cache.** On synchronous
drivers, Drizzle's `db.all()` / `db.run()` re-prepared every statement. They
now go through the same cached path the query engine uses. Single creates:
**~18.3k → ~28.8k ops/s (~1.6x)**.
- **Cascade deletes remove edges in batches** instead of one statement per
edge. A 50-edge cascade on local Postgres: **24.4ms → 3.6ms**.
- **`degree()` can use the index again.** The direction filter compiled to a
shape neither edge index could seek. Now it matches the index prefix:
**0.30ms → 0.06ms** on Postgres 18, and on Postgres 17 and earlier it no
longer falls back to scanning the whole partition.
- **Subgraph extraction is ~4x faster on Postgres** (322ms → 82ms on a
depth-3 stress shape). The recursive traversal runs once instead of twice,
and the resulting ids go in as a single array parameter.
## Search
- **Hybrid search is one SQL statement** instead of two searches plus fusion
in JavaScript, with the candidate filter computed once and shared. Filtered
hybrid search at 5k documents: **26.5ms → 17.1ms**.
- **Postgres fulltext can use its GIN index.** The query referenced the
language from a per-row column, which kept the planner off the index. It's
now a constant: **12.9ms → 2.3ms** at 5,000 documents.
- **Exact `.similarTo()` on SQLite uses sqlite-vec's KNN.** It's brute force
in C, so results are identical, just faster: **489ms → 124ms** for top-10
over 50k 384-dimension embeddings.
- **Approximate vector search on Postgres uses the ANN index now.** A stray
`DISTINCT` kept the planner off the ordered index scan, and the inline path
wasn't applying the same pgvector tuning as the search API. Together:
**174ms → 2.1ms**, at 0.995 recall.
That last area also had a correctness bug. Exact `.similarTo()` was
**quietly approximate** whenever a matching ANN index existed, because
pgvector will happily answer an exact-looking `ORDER BY ... LIMIT k` from an
HNSW or IVFFlat index. Under a selective filter at 50k documents, measured
recall dropped as low as 0.000, which means the results were simply wrong.
The exact path now forces a true scan regardless of which indexes exist.
## Found by benchmarking against real graph databases
Two of these fixes came from running the LDBC Social Network Benchmark
against Neo4j and LadybugDB, not from profiling TypeGraph on its own. Edge
`bulkCreate` / `bulkInsert` had an N+1 endpoint check that made each batch
slower as the graph grew (~90ms → ~630ms per batch). And the default edge
traversal indexes were missing columns a join needed to stay index-only,
which didn't show up until the table outgrew the page cache and then hit a
multi-second latency cliff. The
[full benchmark writeup](/blog/benchmarking-typegraph-neo4j-ladybugdb)
covers where TypeGraph wins and where it doesn't.
**If you're upgrading an existing database**, note that the wider edge
indexes only appear on fresh databases. `CREATE INDEX IF NOT EXISTS` does
nothing when an index with that name exists, even with a different column
list, so an upgraded deployment keeps the narrow index until you rebuild it.
[Performance → Indexes](/performance/indexes) has the exact `DROP` /
`CREATE INDEX CONCURRENTLY` steps for both backends.
## Also in 0.35
- **`.aggregate({...}).orderBy(key, direction?)`**, so "top N groups by
count" no longer means fetching every group and sorting in JavaScript.
- **Breaking, ontology:** `implies(edgeA, edgeB)` now checks that the two
edges' endpoint kinds are compatible. This also runs when a persisted
schema loads, so check saved schemas before rolling out, not just source.
## Try it
- [Performance overview](/performance/overview) and
[Indexes](/performance/indexes)
- [Changelog](/changelog#0350): the full 0.35.0 entry
- [GitHub](https://github.com/nicia-ai/typegraph)
# Changelog
> Release notes for @nicia-ai/typegraph
Release notes for [`@nicia-ai/typegraph`](https://www.npmjs.com/package/@nicia-ai/typegraph). Generated from `packages/typegraph/CHANGELOG.md` on every build.
## 0.73.1
### Highlights
`updateWhere()` and `compareAndSet()` now preserve stored values for defaulted properties omitted from a patch, matching `update()`. Previously, changing one property could silently reset unrelated properties to their schema defaults on every matched row. Defaults are also no longer evaluated for omitted compare-and-set expectations.
### Upgrade notes
- After upgrading, remove workarounds that restate every defaulted property in `updateWhere()` or `compareAndSet()` patches. Omitted properties retain their stored values; to reset a property, supply the desired value explicitly.
- Check rows previously changed by `updateWhere()` or `compareAndSet()` on node kinds with defaulted properties, and restore any unintended resets from application history or backups. Upgrading prevents future resets but does not recover overwritten values.
### Patch Changes
- [#783](https://github.com/nicia-ai/typegraph/pull/783) [`d5560e5`](https://github.com/nicia-ai/typegraph/commit/d5560e5dd220e6682799e94e55c26927d7638310) Thanks [@pdlug](https://github.com/pdlug)! - Preserve omitted defaulted properties in `updateWhere()` and `compareAndSet()` patches. Validate only supplied patch and expected-state fields so defaults cannot silently overwrite stored values or run for omitted expectations.
## 0.73.0
### Highlights
The PostgreSQL working-copy manager now covers the whole branch lifecycle. `makeBackend` hands `branch`, `ingestionBranch`, candidate write-set planning, and evolution previews an empty, schema-mutable allocation recorded in the manager's ledger, so they no longer need a hand-rolled backend factory: closing the backend drops it, and `listUnsealedAllocations` and `abortAllocation` recover it after a crash. Durable working copies can now apply a host mutation and commit its operation evidence in one transaction (`operations: { graph, apply }`), with idempotent replay, commit-ordered scans, delivery marking, and a destroy fence that refuses while evidence is undelivered. Every allocation lives in one recorded schema, so a pooled connection whose `search_path` leads elsewhere can no longer scatter tables that cleanup never finds, and removal either drops everything the allocation owns or refuses with a `BranchError` naming where its relations went.
Operators can list the graphs in a database with `listGraphIds` and count one graph's rows per relation with `inspectGraphStorage`, on SQLite and PostgreSQL alike, without depending on TypeGraph's physical tables. `inspectGraphStorage` reports whether its counts came from one snapshot. The set of graph-scoped relations now has a single owner shared by `store.clear()`, namespace forks, working-copy cloning, and these reads, which fixed a few places that disagreed about it. On PostgreSQL, a new byte-ordered `graph_id` index keeps a page of `listGraphIds` at about 3 ms with 20,000 graphs in the database, and adds 1 to 2% to writes on the tables it covers.
Writes through an ephemeral PostgreSQL working copy of a history-enabled store work again; since 0.72.0 every create, update, or delete on such a copy failed with a `ConfigurationError`.
### Upgrade notes
- The base-schema marker advances from 4 to 5. Open each database once with `createStoreWithSchema` under a DDL-capable role, or apply the version-5 migration if you manage DDL externally, before starting workers that use `createVerifiedStore`, `assertSchemaCurrent`, or the DML-only graph-template APIs; until then they throw `BaseSchemaMigrationError`. Earlier releases refuse a database stamped 5, so do not roll back once any process has adopted it.
- On PostgreSQL, if `nodes` or `edges` see continuous writes, create the three `_graph_id_bytes_idx` indexes with `CREATE INDEX CONCURRENTLY IF NOT EXISTS` (as `backend-setup` shows) before upgrading. Otherwise the version-5 adoption builds them inline at boot, blocking writes to each table while it builds, even when `systemIndexes: "skip"` is set.
- Run the working-copy manager's `control` and every `connect` session as the same database role. A different role, including one that is merely a member of `control`'s role, is now refused with `WORKING_COPY_ROLE_MISMATCH` on every path. Drop by hand any tables that allocations created under a different role left behind.
- In `connect`, build the backend's tables with `createPostgresTables(names)` from the object `connect` receives, not a copy of it, use a driver that supports interactive transactions (not `drizzle-orm/neon-http`), and keep the allocation's schema on the connection's `search_path`.
- If you wrap the `control` or `connect` backend, forward the transaction `isolationLevel` option. Otherwise destroy, `abortAllocation`, closing ephemeral or `makeBackend` copies, and durable operations fail with `WORKING_COPY_ISOLATION_UNSUPPORTED` whenever the session's default isolation is not READ COMMITTED.
- Upgrade every process that shares a working-copy ledger before any of them creates or destroys a durable allocation. An earlier manager destroys allocations without checking their operation evidence and leaves the evidence table behind; the #777 entry below describes how to recover it.
### Minor Changes
- [#778](https://github.com/nicia-ai/typegraph/pull/778) [`de9054b`](https://github.com/nicia-ai/typegraph/commit/de9054bd86b13fc01356802203ed1717bc0369eb) Thanks [@pdlug](https://github.com/pdlug)! - Add `listGraphIds(backend, { prefix, after, limit })` and `inspectGraphStorage(store)` so operators can list the graphs in a database and count one graph's rows per relation, including per-field vector tables, without depending on TypeGraph's physical table layout. Both behave identically on SQLite and PostgreSQL: ids page in byte order regardless of database collation, only graphs a default `store.clear()` would empty are listed (a cleared graph stops being listed even though its contribution markers are kept), each page walks graph ids by index seek instead of scanning every row, bounded by the cursor, prefix and page size (about 3 ms a page on PostgreSQL at 20,000 graphs, against about 400 ms for the equivalent walk over the database collation), prefixes match as case-sensitive literal text, and relations that were never provisioned count as empty.
`inspectGraphStorage` reports whether its counts share one snapshot in a new `consistency` field, `"snapshot"` or `"per-statement"`. It requests a read-only `repeatable read` transaction, then reads the isolation level the counting session actually ran under inside its first count statement instead of trusting the request: SQLite transactions are always one snapshot, PostgreSQL reports `"snapshot"` only when the session was observed at `repeatable read` or `serializable`, and everything else reports `"per-statement"`: a backend without interactive transactions, a session observed at `read committed` (which is what a wrapper that drops the isolation option gets under a `read committed` default; under a `repeatable read` default the same wrapper still reports `"snapshot"`, because the level is observed rather than requested), or a backend that declares no session isolation read and so cannot be observed. With `"per-statement"`, a write between two counts can make the result describe a state that never existed. A graph with fewer than two count statements reports `"snapshot"`. The evidence proves the isolation of the session that ran the first count, so a backend wrapper that violates the transaction contract by handing the root pool through as its transaction backend can still run later counts on other sessions. It does not refuse on a weaker level.
The set of graph-scoped relations now has one owner. `store.clear()`, namespace forks, PostgreSQL working-copy clone policies, the provenance sidecar occupancy probe and the new reads all consume the same declaration, and a ratchet test fails when a bundled table gains a `graph_id` column without being classified. `SqlTableNames` now also carries the `indexMaterializations`, `contributionMaterializations`, `kindRemovals` and `reconciliationMarkers` names (optional, defaulting to the bundled names), and the SQLite backend reports them in `backend.tableNames` as PostgreSQL already did.
Inconsistencies the shared declaration exposed are fixed. `store.clear()` now removes the graph's revision-origin row whichever store clears it; before, a store that did not mint origin-namespaced tokens left a row that a revision-tracking store had minted, so the graph still read as occupied. `forkGraphNamespace` and `prepareNamespaceForkTarget` no longer fail with a raw missing-relation error on a PostgreSQL backend created with `fulltext: false`, and refuse a source and target whose fulltext storage differs with a `BranchError`. The provenance sidecar occupancy probe now counts revision-journal rows as occupancy, so a graph id whose only rows are journal entries is no longer treated as free.
PostgreSQL gains base-schema version 5: a byte-ordered (`COLLATE "C"`) single-column `graph_id` index on `nodes`, `edges` and `schema_versions`, named `_graph_id_bytes_idx`, which is what lets the listing bound its walk (SQLite already keeps text indexes in byte order, so its step only advances the marker). The privileged open adopts it on existing databases with a plain `CREATE INDEX IF NOT EXISTS` per relation, which blocks writes to that relation while it builds (about 0.1 second per million `nodes` rows). This happens inline at boot even when `systemIndexes: "skip"` is set, because that option does not defer base-schema adoption; for a large deployment, build the three indexes with `CREATE INDEX CONCURRENTLY IF NOT EXISTS` beforehand, as `backend-setup` describes. The index adds 1 to 2% to writes on the relations it lands on. The index names are reserved like the other system indexes, and a database whose indexes are absent still lists correct pages by de-duplicating the anchor relations instead of walking them.
## Breaking
- The base-schema marker advances from 4 to 5, so a database stamped 4 is stale until it is adopted: a store opened with `createStoreWithSchema` adopts it automatically under a DDL-capable role, while `createVerifiedStore`, `assertSchemaCurrent` and the DML-only graph-template APIs throw `BaseSchemaMigrationError` (`BASE_SCHEMA_MIGRATION_REQUIRED`) until it is. Open once with `createStoreWithSchema` before DML-only workers start, or apply the version-5 migration for externally managed DDL. Older library versions refuse the newer marker, so do not roll back to a release that predates version 5 once any process has adopted it.
- On PostgreSQL, build the three `graph_id_bytes_idx` indexes with `CREATE INDEX CONCURRENTLY IF NOT EXISTS` before upgrading a database whose `nodes` or `edges` tables see continuous writes, so the adoption step finds them in place and does not block writers.
- [#777](https://github.com/nicia-ai/typegraph/pull/777) [`38c3a21`](https://github.com/nicia-ai/typegraph/commit/38c3a2164dd9d8e27510e7ebb3b0baa2147deea2) Thanks [@pdlug](https://github.com/pdlug)! - Implement the durable operation capability on the bundled PostgreSQL working-copy manager. Passing `operations: { graph, apply }` to `createPostgresWorkingCopyManager` makes `durable.operations` apply the host's opaque mutation inside the transaction that commits its evidence, with idempotent replay, digest conflicts, commit-ordered scan cursors, delivery marking, and a destroy fence that refuses with `DurableEvidenceUndeliveredError` while undelivered evidence remains. Evidence coordinates carry `base` and, when the working copy tracks history or revisions, the engine `revision`.
Each durable allocation owns an evidence table, created in the allocation's schema and removed with the rest of the allocation (the destroy fence reads it only there, follows the manager's removal rule for allocations whose relations were moved or dropped, and runs only when the evidence relation still exists, so an allocation whose evidence relation is gone is removed like any other relation that exists nowhere), and the working-copy ledger gains an `operation_evidence` column. `operate`, `markDelivered`, destroy, and `abortAllocation` serialize per allocation on a transaction-scoped advisory lock keyed on the allocation id, in a namespace of its own, so the evidence sequence is commit order. Each requests READ COMMITTED and observes the isolation its session actually runs at; any other level is refused with a `ConfigurationError` (`details.code` `WORKING_COPY_ISOLATION_UNSUPPORTED`) before anything is read or written, so a `control` or `connect` wrapper must forward the transaction `isolationLevel` option. Every member checks `operations.graph` against the sealed allocation's graph id and definition hash before opening a connection. `apply` runs under the graph-wide write lock, which blocks tracked writes to the source graph and its sibling working copies for its duration.
Allocations created by earlier releases report `unsupported` (`evidenceStore`) for `operate` and no evidence for the read members, answered from the ledger row alone: one read-only ledger `SELECT` through `control`, with no DDL, no lock, and no `connect` call.
### Upgrade notes
- Upgrade every process that shares a working-copy ledger before any of them creates or destroys a durable allocation. Only managers on this version take the allocation lock and honor the evidence fence. A manager from an earlier release destroys an allocation without consulting its evidence, so undelivered evidence is lost with it, and it leaves the evidence table behind.
- If an earlier manager destroyed a durable allocation created by this version, a later allocation with the same id is refused because `op_evidence` exists without a ledger row; the refusal names the table. Read its undelivered rows (`WHERE NOT delivered`) and deliver them, then `DROP TABLE` the named relation and retry. TypeGraph does not drop it, because it may hold the only copy of undelivered evidence.
- A `control` or `connect` backend wrapper that drops the transaction `isolationLevel` option now fails destroy, `abortAllocation`, and durable operations with `WORKING_COPY_ISOLATION_UNSUPPORTED` when the session's default is not READ COMMITTED. Forward the option.
- The same check now runs wherever the manager drops an allocation, because dropping takes the allocation lock: closing an ephemeral working-copy store, closing a `makeBackend` backend, and cleanup after a failed allocation. The first two throw the refusal from `close()`. The cleanup swallows it so the allocation's original failure reaches the caller, and the allocation is left behind. List such orphans with `listUnsealedAllocations` and remove them with `abortAllocation` once `control` forwards the option.
- [#776](https://github.com/nicia-ai/typegraph/pull/776) [`e4a9419`](https://github.com/nicia-ai/typegraph/commit/e4a941998f9a311fda2189205eda92431b04a821) Thanks [@pdlug](https://github.com/pdlug)! - Add `makeBackend` to the PostgreSQL working-copy manager so `branch`, `ingestionBranch`, `planCandidateWriteSet`, `planCandidateWriteSetReview`, `branchForEvolution`, and `planCandidateWriteSetForEvolution` no longer need a hand-rolled PostgreSQL backend factory. Each call records a fresh, empty, schema-mutable allocation in the manager's ledger; closing the backend drops it, it is listed by `listUnsealedAllocations()` while live, and `abortAllocation()` removes it after a crash. Dropping any allocation now also removes the vector tables created under its reserved physical prefix after the ledger manifest was written, and declared graph indexes on a `makeBackend` allocation are scoped to it so they cannot collide with the source's or another allocation's index names. `connect` always receives the allocation vector strategy for `makeBackend` and must bind it or disable vector support. Backends derived from the returned one with `deriveBackend` inherit that index scoping.
Every allocation now lives in one explicit schema, the `control` session's current schema when it is allocated, recorded in a new `schema_name` ledger column (added to an existing ledger on first use; a row written by 0.72.0 carries no schema and is resolved through the removing session). Provisioning fixes its transaction to that schema, the connected backend runs the DDL it issues lazily with that schema leading its search path, the allocation's pgvector strategy creates and drops its tables and indexes schema-qualified, and clone inserts name their target relations through it, so a pooled `connect` connection whose `search_path` leads with another schema can no longer create relations that removal never finds. Removal searches the catalog across every schema for relations named with the allocation's reserved prefixes, drops those in the recorded schema schema-qualified, and deletes the ledger row in the same transaction, so a failed drop keeps the row and the allocation stays listed and recoverable. When such relations exist in another schema, removal refuses with a `BranchError` that keeps the row and carries `allocationId`, the recorded `schema`, the schemas found in `foundIn`, and a recovery `suggestion`: a renamed schema or moved tables read "not in its schema" (move them back or correct the row's `schema_name`); a stale copy left in another schema, such as a backup or restore schema, reads "also has relations" and blocks removal until that copy is dropped, and is only called a stale copy when every relation found elsewhere has a same-named relation in the recorded schema; a partial move (some relations moved, others stayed) reads "is split across schemas" and suggests dropping nothing, because the relations elsewhere may be the only copy; when they exist nowhere (the tables were dropped entirely) there is nothing to recover and the row is removed, so a crashed owner's allocation cannot stay listed forever. The same rule applies to a legacy row that records no schema. A backend built over a caller's own transaction never has that transaction's `search_path` rewritten: its lazy DDL (including extension DDL on the non-lock fence path), schema writes, and schema adoption run only when the session's current schema is the allocation's, and are refused with a `ConfigurationError` (`ALLOCATION_SCHEMA_SESSION_MISMATCH`) otherwise; extension installation under a lock fence makes no session check, running in its own transaction on a pooled backend and as a savepoint inside the caller's transaction on a backend built over one. The connected backend's catalog probes (`tablesExist`, `indexStates`, `columnTypes`) read the allocation's schema rather than the session's current schema, so a history-enabled allocation opens through a connection whose `search_path` leads with another schema. Trusted import drops and recreates the allocation's secondary indexes by the schema the catalog found them in, so a same-named index earlier on the connection's `search_path` is left alone.
**Breaking:** the PostgreSQL working-copy manager now supports one deployment shape, in which `control` and every `connect` session run as the same database role. The `ephemeral` and `durable` clone paths and durable reopen, which shipped in 0.72.0 without this check, now compare `current_user` on the two sessions and refuse a difference with a `ConfigurationError` (`details.code` `WORKING_COPY_ROLE_MISMATCH`). The comparison is by role name, so a `connect` role that is merely a member of `control`'s role, which worked before, is now refused too. `makeBackend` applies the same rule and refuses before it writes the ledger or any DDL. The clone paths and reopen run `connect` after the allocation tables exist, as before, so they refuse right after `connect` returns, before any clone or Store write: a refused clone removes the allocation it just created, and a refused reopen leaves the sealed allocation untouched.
**Breaking:** because the schema reaches a connected backend through the names `connect` receives, build the backend's tables with `createPostgresTables(names)` from that object, not a copy: a connection built over a copy is refused with a `BranchError` on every path (and a refused clone removes its allocation), and a `connect` driver that cannot hold an interactive transaction (`drizzle-orm/neon-http`) is refused with a `ConfigurationError` (`ALLOCATION_SCHEMA_REQUIRES_INTERACTIVE_TRANSACTIONS`). The connection's `search_path` must still include the allocation schema so it can resolve the allocation's tables.
**Migration:** run `control` and every `connect` session as the same database role, which needs `CREATE` on the schema. The reason is ownership: a connected Store creates tables and indexes that only their owner can drop, and `control` removes every allocation. If you connect as a different role today, allocations created that way can leave tables behind on close and `abortAllocation`; after switching roles, drop any such leftovers by hand.
### Patch Changes
- [#780](https://github.com/nicia-ai/typegraph/pull/780) [`e0be095`](https://github.com/nicia-ai/typegraph/commit/e0be095310f6c82f96bb36a17a907feaed07b7f8) Thanks [@pdlug](https://github.com/pdlug)! - Fix writes through an ephemeral PostgreSQL working copy of a history-enabled store. The copy's store was built over the allocation store's already capture-wrapped backend, so recorded capture wrapped twice and every create, update, or delete failed with a `ConfigurationError` from the raw-write guard. The ephemeral store is now built over the unwrapped owned backend, and its writes are captured in the copy's own recorded history.
## 0.72.0
### Highlights
PostgreSQL graphs using bundled table, tsvector, and pgvector storage can now use table-backed working copies. Each allocation owns its tables, indexes, and vector sidecars under a recovery ledger; graph-scoped cloning, durable reopen, and cleanup keep copies isolated from the source. Copies use a fixed schema, so migrate the source before allocating one when schema changes are needed.
On revision-tracked graphs, candidate merge planning stages the affected rows and their required identity, ontology, cardinality, and durable edge-identity dependencies instead of cloning the whole target. Candidate-scoped durable review is opt-in. Stores without a revision fence and custom backends missing the keyed reads needed for safe staging retain the complete-clone path.
Bulk node creation now reports a generated ID collision as `ValidationError` with `ENTITY_ALREADY_EXISTS_CODE` across the bundled SQLite and PostgreSQL drivers, including atomic batches and the portable fallback.
### Upgrade notes
- If a custom PostgreSQL table contribution participates in table-backed working copies, declare its `workingCopyClonePolicy`. Use `graphRows` or `graphDocument` only for graph-scoped content; use `freshSeed` or `rebuildAfterClone` for installation or physical status. An absent or unsupported policy now refuses allocation.
- If you handle generated ID collisions from `bulkCreate` or `bulkInsert`, branch on `ValidationError` with `ENTITY_ALREADY_EXISTS_CODE` rather than a driver error or `DatabaseOperationError`.
### Minor Changes
- [#759](https://github.com/nicia-ai/typegraph/pull/759) [`d3e74f8`](https://github.com/nicia-ai/typegraph/commit/d3e74f8aff5948038351539c3479ed3b3339989f) Thanks [@pdlug](https://github.com/pdlug)! - Plan candidate write sets on revision-tracked graphs, including Operational Identity and ontology graphs, from bounded candidate dependencies instead of cloning the complete target. Unsupported custom backend reads and candidate owners excluded from the clone projection retain full clone staging. Add opt-in candidate-scoped durable review evidence while retaining the existing whole-graph review default.
- [#757](https://github.com/nicia-ai/typegraph/pull/757) [`afe153b`](https://github.com/nicia-ai/typegraph/commit/afe153bd720c56ca17c67b527a4d981e8aefab66) Thanks [@pdlug](https://github.com/pdlug)! - Export `generatePostgresDropSQL()` for cleaning up isolated PostgreSQL table sets from the same contribution inventory used by installation DDL. Quote custom PostgreSQL table and index names consistently in generated DDL.
- [#762](https://github.com/nicia-ai/typegraph/pull/762) [`d793eff`](https://github.com/nicia-ai/typegraph/commit/d793efff201b78f2aff8fcf0b6be45b6ad9592a8) Thanks [@pdlug](https://github.com/pdlug)! - Support graph-declared PostgreSQL indexes in table-backed working copies with stable allocation-scoped physical names. B-tree, GIN, and trigram indexes retain their logical declaration names and schema hashes while copy allocation, durable reopen, retry, and cleanup use isolated physical indexes.
- [#761](https://github.com/nicia-ai/typegraph/pull/761) [`c6130a1`](https://github.com/nicia-ai/typegraph/commit/c6130a1dd05b3a0b19c13623b494069d1792527f) Thanks [@pdlug](https://github.com/pdlug)! - Add a PostgreSQL table-backed working-copy manager for graphs using bundled table and tsvector storage. It owns ephemeral and durable allocation, a persistent recovery ledger, origin-attested reopen and destroy, graph-scoped SQL cloning under source locks, and bounded inventory of unsealed allocations. Inventory rows may still be active, so callers confirm ownership before explicitly aborting one. Copies have a fixed schema and refuse evolution before mutation. Custom fulltext strategies remain available through host-level database forks.
- [#769](https://github.com/nicia-ai/typegraph/pull/769) [`ac8c781`](https://github.com/nicia-ai/typegraph/commit/ac8c781d80deb3ff51f061d89d15360d8c5753c7) Thanks [@pdlug](https://github.com/pdlug)! - Bound candidate merge planning and opt-in candidate-scoped review for `oneActive` graphs on bundled backends. An active-only source read excludes ended edge history while preserving the claim rule that an open edge counts even when its `validFrom` is in the future. Custom backends without the optional read continue to use complete-clone candidate planning.
- [#766](https://github.com/nicia-ai/typegraph/pull/766) [`c589e00`](https://github.com/nicia-ai/typegraph/commit/c589e00ad3b96960c89c58f8012decdedfdbb10a) Thanks [@pdlug](https://github.com/pdlug)! - Bound candidate merge planning on revision-tracked graphs with `one` or `unique` edge cardinality. The transient working copy now includes only cardinality peers for candidate sources or endpoint pairs, so staging preserves full-clone constraint decisions without reading unrelated edges. A `oneActive` graph uses complete-clone staging when its backend lacks the active-only keyed peer read.
- [#756](https://github.com/nicia-ai/typegraph/pull/756) [`ae813a5`](https://github.com/nicia-ai/typegraph/commit/ae813a55056c5eec6c72c7add53b59a1833e160a) Thanks [@pdlug](https://github.com/pdlug)! - Read graph rows across declared kinds in keyset pages for merge planning and review, avoiding an empty query for each unused kind. Reuse one row read when the target is both sides of a diff, and skip statistics refresh for disposable ingestion clones. Custom backends continue using the existing per-kind read path.
- [#768](https://github.com/nicia-ai/typegraph/pull/768) [`79bfa87`](https://github.com/nicia-ai/typegraph/commit/79bfa876765b84be6feb30e7f88237c7fbf3e0ca) Thanks [@pdlug](https://github.com/pdlug)! - Bound candidate merge planning on revision-tracked ontology graphs by reading live same-id peers across node kinds. This preserves full-clone disjointness and type-reconciliation decisions without scanning unrelated nodes, and extends opt-in candidate-scoped review evidence to ontology graphs.
- [#764](https://github.com/nicia-ai/typegraph/pull/764) [`a3c2d9c`](https://github.com/nicia-ai/typegraph/commit/a3c2d9c718dc6f7ccb4dfcaac7136a036aabe83a) Thanks [@pdlug](https://github.com/pdlug)! - PostgreSQL table-backed working copies now isolate pgvector sidecars under each allocation's ledger-reserved physical prefix, preserve their embeddings through clone and reopen, and remove their owned tables during destroy. Allocation claims and initial table/vector provisioning commit atomically.
- [#770](https://github.com/nicia-ai/typegraph/pull/770) [`7f69a44`](https://github.com/nicia-ai/typegraph/commit/7f69a4466ee36084f5866d0e3476794e687c1d4e) Thanks [@pdlug](https://github.com/pdlug)! - Add an optional backend read for exact durable edge match-identity owners, including tombstones. Candidate planning uses the bounded read to seed active owners into sparse working copies and falls back to full cloning for custom backends without the capability or when a durable owner is tombstoned.
- [#764](https://github.com/nicia-ai/typegraph/pull/764) [`a3c2d9c`](https://github.com/nicia-ai/typegraph/commit/a3c2d9c718dc6f7ccb4dfcaac7136a036aabe83a) Thanks [@pdlug](https://github.com/pdlug)! - Add `createPgvectorStrategy(namespace)` for allocation-scoped pgvector table and index names while preserving the default strategy's existing names.
### Patch Changes
- [#774](https://github.com/nicia-ai/typegraph/pull/774) [`14dacc4`](https://github.com/nicia-ai/typegraph/commit/14dacc4a1b057671fab06e56380448692a330c36) Thanks [@pdlug](https://github.com/pdlug)! - Preserve the typed duplicate-ID error when PostgreSQL rejects claim cleanup after a failed bulk node insert.
- [#760](https://github.com/nicia-ai/typegraph/pull/760) [`9198fa8`](https://github.com/nicia-ai/typegraph/commit/9198fa8b1c64e256d30665d913bc4ba82842f303) Thanks [@pdlug](https://github.com/pdlug)! - Compare cloned branch identity assertions with the base's current state so assertions ended before the fork do not appear as new branch retractions.
- [#775](https://github.com/nicia-ai/typegraph/pull/775) [`6820b04`](https://github.com/nicia-ai/typegraph/commit/6820b04f2f63b10abd0d17daf414369cf4085c47) Thanks [@pdlug](https://github.com/pdlug)! - Require explicit PostgreSQL working-copy clone policies for table contributions so physical status rows cannot be copied by column-shape inference.
- [#765](https://github.com/nicia-ai/typegraph/pull/765) [`d28c8cb`](https://github.com/nicia-ai/typegraph/commit/d28c8cba346b988c724e760a79f758d3c1b9a885) Thanks [@pdlug](https://github.com/pdlug)! - Refuse PostgreSQL working-copy allocation or durable reopen when the target connection changes the bundled fulltext strategy. This prevents copied fulltext projections from being exposed through a backend with different storage or disabled fulltext support.
- [#763](https://github.com/nicia-ai/typegraph/pull/763) [`857f578`](https://github.com/nicia-ai/typegraph/commit/857f578b402d35c9ac3f929a18f7d78405c838ba) Thanks [@pdlug](https://github.com/pdlug)! - Add opt-in candidate-scoped V2 merge review evidence for identity-enabled graphs. Revalidation expands the retained endpoint and assertion-ID scope again and detects connected identity changes while leaving V1's global baseline as the default.
## 0.71.1
### Patch Changes
- [#752](https://github.com/nicia-ai/typegraph/pull/752) [`2d40378`](https://github.com/nicia-ai/typegraph/commit/2d403780090d39e0aa8e5ef3e62cf229cfed3a14) Thanks [@pdlug](https://github.com/pdlug)! - Materialize class membership once while paging current identity classes. SQLite could otherwise repeat the full node scan inside the kind filter for every class, making `identity.classes()` much slower than the earlier in-memory read on graphs with many unrelated node kinds.
## 0.71.0
### Highlights
TypeGraph 0.71 makes identity groups easier to inspect. `identity.classes()` pages through visible classes, including singletons, at current or historical read coordinates. `identity.explainSame(a, b)` returns a shortest proof through persisted `same` assertions and eligible same-ID folds, so applications can show why two references belong together. Cursors are bound to the graph, coordinate, and kind scope, and remain compact as the graph schema grows.
Historical identity reads and traversals now agree on which kinds belong to the current graph. Explanations at a historical coordinate use only folds that existed then. Schema migration errors also identify changed validators with exact JSON Pointers and before-and-after patterns, making a blocked migration easier to diagnose.
### Upgrade notes
- Replace `IdentityReadSurface` with `IdentityReadFacade` and `IdentitySurface` with `IdentityFacade`. The surface aliases are no longer exported. If you implement `IdentityReadFacade` yourself, add `classes` and `explainSame`; these methods are also part of merge callback read contexts.
- Before removing a node kind that connects retained identities through `same` assertions, move the needed assertions to retained kinds if those identities should remain joined. Historical identity reads and identity-expanded traversals now exclude kinds absent from the current graph.
- To use current-coordinate `identity.classes()` with a custom backend, provide SQL window-function support and declare `capabilities.windowFunctions: true`. A backend profile that declares `false` raises `ConfigurationError` for this read.
### Minor Changes
- [#748](https://github.com/nicia-ai/typegraph/pull/748) [`d1c8322`](https://github.com/nicia-ai/typegraph/commit/d1c83228cb4a83c9a99eb6af2c0663dd7eaddc4f) Thanks [@pdlug](https://github.com/pdlug)! - Add `identity.classes({ kinds, cursor, limit })` to read visible identity classes in deterministic pages at the current or a historical coordinate. Pages include visible singleton classes and expose registered visible members of each matching class. Opaque cursors are bound to the graph, read coordinate, and kind filter.
- [#748](https://github.com/nicia-ai/typegraph/pull/748) [`d1c8322`](https://github.com/nicia-ai/typegraph/commit/d1c83228cb4a83c9a99eb6af2c0663dd7eaddc4f) Thanks [@pdlug](https://github.com/pdlug)! - Add `identity.explainSame(a, b)` to return a shortest path of persisted same assertions and implicit same-ID folds at the facade's read coordinate.
- [#750](https://github.com/nicia-ai/typegraph/pull/750) [`b305ad9`](https://github.com/nicia-ai/typegraph/commit/b305ad9e9b7f890a4d497135ad9fc38422ebf0e0) Thanks [@pdlug](https://github.com/pdlug)! - Make `IdentityReadFacade` and `IdentityFacade` the complete public identity surfaces, including `classes` and `explainSame`. Replace the exported `IdentityReadSurface` and `IdentitySurface` aliases with those facade types.
Historical identity reads and traversals now agree on registered kinds, and `explainSame` cites an implicit same-ID fold only when both nodes existed at the requested coordinate. Class cursors keep a fixed size as kind filters grow. Identity invariant errors include graph details and an appropriate current or historical recovery hint.
### Patch Changes
- [#748](https://github.com/nicia-ai/typegraph/pull/748) [`d1c8322`](https://github.com/nicia-ai/typegraph/commit/d1c83228cb4a83c9a99eb6af2c0663dd7eaddc4f) Thanks [@pdlug](https://github.com/pdlug)! - Report exact JSON Pointers and before-and-after pattern values in schema migration diagnostics so validator changes can be located and reviewed precisely.
## 0.70.0
### Highlights
TypeGraph 0.70 strengthens durable branch identity and recovery. Each allocation has its own ID, which is checked against the sealed host origin when a branch is reopened, destroyed, or merged. Hosts can persist the branch and allocation IDs before creation to reconcile an uncertain result. Durable merge plans now retain recorded fork points, and strategies can keep older locator formats readable while writing a new format.
PostgreSQL namespace forks now support the bundled pgvector storage. Embedding rows join the same verified snapshot as the rest of the graph, and owner-side preparation builds the target's vector tables and eligible indexes before the runtime copy. IVFFlat indexes are deferred until a post-copy `materializeIndexes()` call so they cluster the forked rows. Base-version tokens are now printable and can be stored directly in PostgreSQL text and JSON columns.
Incremental merges now preserve a node already committed by the target when another branch proposes the same entity. This prevents a second merge from trying to change the endpoints of existing committed edges. Local `@libsql/client` 0.18 clients also use transaction framing compatible with pooled connections.
### Upgrade notes
- Finish or remove durable branches created by an earlier release before upgrading, then create new descriptors and sealed host origins with allocation IDs. Earlier descriptors cannot be reopened, destroyed, merged, or used for evidence access in 0.70. Re-branch other work whose legacy `base@V` token is needed for merge planning; existing plans can still apply when their target fence has not moved.
- Update custom durable strategies: accept the allocation ID in `create()`, return `{ operations, cursor, hasMore }` from operation scans, and capture `forkRevision` inside `create()` only when it is atomic with allocation. Remove native `merge` implementations; durable plans now apply through the target Store transaction. Use `readableVersions` if a new strategy version must read older locator formats.
- Replace `installNamespaceForkLedger(target)` with owner-side `prepareNamespaceForkTarget(source, target)` before runtime namespace forks. Use matching vector storage on both backends for graphs with embedding fields, and call `materializeIndexes()` on the forked store after copying a graph with IVFFlat indexes.
- Re-branch or re-plan work that uses an old `engine:` anchor or untracked content token. New untracked tokens include the complete graph content and active schema version.
### Minor Changes
- [#745](https://github.com/nicia-ai/typegraph/pull/745) [`01b8149`](https://github.com/nicia-ai/typegraph/commit/01b814961c61b038fd71f6ad383bfdb0e350914b) Thanks [@pdlug](https://github.com/pdlug)! - Durable branch descriptors now carry a unique allocation ID, independent of the caller's branch ID. Reopen, destroy, and durable merge compare this ID with the host's sealed origin, so two copies using the same branch ID cannot be confused by a swapped locator. Callers may persist a stable branch ID and allocation ID before creation for host-side reconciliation after an uncertain result. A durable branch handle has the `DurableGraphBranch` type, which binds it to its allocation. `applyDurableMergePlan()` also carries the recorded fork point when comparing a branch with its descriptor, allowing plans for history-enabled durable branches to apply.
Strategies may declare `readableVersions` alongside the locator format `version` they write, allowing upgraded strategies to continue reading and managing older locator formats. The descriptor version is passed to every read, destroy, and evidence method. A strategy may supply a `forkRevision` captured atomically with allocation; when it cannot, TypeGraph uses a full diff to avoid missing writes between allocation and sealing. Durable operation scans now return `hasMore` and retain their cursor at the end of a page, so callers can resume after later commits; strategies must order evidence by a monotonic commit position.
Untracked stores now use the complete graph-content fingerprint and active schema version even when their backend offers lineage. The previous engine anchor could miss identity-only writes and did not read the planned graph state inside the commit transaction. The optional host-native `merge` strategy method is removed; `applyDurableMergePlan()` applies through the target Store transaction until a native merge contract can prove its target fence across the native operation's commit boundary.
### Upgrade notes
- Re-create durable branch descriptors and sealed host origins from earlier releases. They lack the required `allocationId` fence and are refused on reopen, destroy, merge, and evidence access. Keep the previous release available to finish or remove those branches before upgrading.
- Update `DurableOperationCapability.scan` implementations to return `{ operations, cursor, hasMore }`. The cursor must identify the last observed commit position even when `hasMore` is `false`; an empty page echoes `after`.
- Move any strategy revision capture into `create()` and return it as `forkRevision` only when it was captured atomically with the physical fork. Omit it when the host cannot prove that cut.
- Update `DurableWorkingCopyStrategy.create` to accept the allocation ID and refuse a duplicate until the host has explicitly reconciled it. Callers that need crash recovery should persist both IDs before calling `branchDurable` and pass them in options.
- Re-branch or re-plan work whose base version uses the retired `engine:` anchor or an older untracked content token. New untracked tokens fingerprint current identity assertions and include the active schema version.
- Remove `DurableWorkingCopyStrategy.merge` implementations and use `applyDurableMergePlan()`'s transactional apply. Database-native working-copy allocation remains supported.
- [#744](https://github.com/nicia-ai/typegraph/pull/744) [`3f7da34`](https://github.com/nicia-ai/typegraph/commit/3f7da34e23265941cffc52cb984aa60d657a2014) Thanks [@pdlug](https://github.com/pdlug)! - `forkGraphNamespace()` now forks graphs that use the bundled pgvector storage. Embedding rows are copied inside the same repeatable-read snapshot, included in the content digest that the copy, retries and `abort()` verify, and removed by `abort()`. A graph with embedding fields forks only between backends with the same vector storage, pgvector on both sides or `vector: false` on both; custom vector and fulltext strategies are still refused.
`prepareNamespaceForkTarget(source, target)` is the owner-side step. It installs the retry ledger, creates the graph's pgvector tables, and builds every index the source has materialized for the graph with the DDL the source used. It writes no graph rows and no materialization records, and the runtime fork still issues no DDL. IVFFlat indexes need the copied rows to cluster well, so preparation skips them, the fork does not copy their records, and `fork.store.materializeIndexes()` builds them after the copy.
`materializeIndexes()` now rebuilds an IVFFlat index that exists without a materialization record, for example one an aborted fork left behind, instead of keeping it with `IF NOT EXISTS`: it was clustered for other rows. A backend without `dropVectorIndex` keeps the previous behavior.
A materialized vector index no longer makes the fork refuse, and indexes whose build never completed on the source are neither built on nor required of the target.
### Upgrade notes
- Replace `installNamespaceForkLedger(target)` with `prepareNamespaceForkTarget(source, target)`, run with the schema owner role before the runtime fork. `installNamespaceForkLedger` is removed.
- Namespace-fork backends no longer need `vector: false`. For a graph with embedding fields, open source and target with the same vector storage: pgvector on both, or `vector: false` on both.
- After forking a graph that declares IVFFlat indexes, run `materializeIndexes()` on the forked store under the owner role to build them.
- [#743](https://github.com/nicia-ai/typegraph/pull/743) [`b69ec0b`](https://github.com/nicia-ai/typegraph/commit/b69ec0b9f4dc358b031e39eb8728b6ea11be871a) Thanks [@pdlug](https://github.com/pdlug)! - `base@V` tokens are now printable text. Their two components were joined by a NUL character, which PostgreSQL `text` and `jsonb` columns reject, so an application could not persist a durable branch descriptor, a recorded fork point, a merge plan, or durable operation evidence in PostgreSQL without re-encoding it. The separator is now `|`.
### Upgrade notes
- Merge or re-create branches and durable branches minted by an earlier release. Their `base@V` tokens are refused with `BaseVersionMismatchError` and `details.reason: "legacy-token-format"` when `planMerge()`, `merge()`, `planMergeIncremental()`, or `mergeIncremental()` validates the branch's base, and when an incremental plan starts from a persisted `RecordedForkPoint`. Reopening a durable branch still succeeds; planning a merge from it does not.
- Existing merge plans are unaffected. Applying a plan, including through `applyDurableMergePlan()`, validates the plan's target fence (graph id, schema, and revision anchor), not the format of the tokens recorded in its anchors. A plan whose target has not moved since planning still applies after the upgrade.
- Durable operation evidence stores the coordinates the host supplied and is not compared with newly minted tokens, so existing evidence remains readable.
- Code that stored tokens in a re-encoded form (base64, or JSON text in a `text` column) keeps working and may store them directly.
### Patch Changes
- [#740](https://github.com/nicia-ai/typegraph/pull/740) [`fe345b0`](https://github.com/nicia-ai/typegraph/commit/fe345b0e08213d414dc71321bc39bfe30c345e07) Thanks [@pdlug](https://github.com/pdlug)! - Update `nanoid` to 6.0, which requires Node.js 22 or later, matching the package's existing `engines` range. `ExportOptionsSchema.signal` keeps its declared `ZodCustom` type, so the published declarations stay valid across the whole `zod ^4.0.0` peer range.
- [#746](https://github.com/nicia-ai/typegraph/pull/746) [`6ad8c03`](https://github.com/nicia-ai/typegraph/commit/6ad8c030a5786ad09a73ddbc204bdc6c2d68f730) Thanks [@pdlug](https://github.com/pdlug)! - Incremental merges now keep a node the target committed after the fork point as the survivor when a branch proposes the same entity. Two branches forked from one point that both added an entity could previously fail on the second merge: when the second branch's node had the lexicographically smaller id, it won survivor selection, and the plan tried to repoint the committed edges of the first branch's node, which `applyMergePlan()` refused as an immutable-endpoint change. Merges that resolved through `blockIndex` or a unique constraint were not affected.
- [#741](https://github.com/nicia-ai/typegraph/pull/741) [`b8fd08e`](https://github.com/nicia-ai/typegraph/commit/b8fd08e2854d609325928038725c5502027b4b81) Thanks [@pdlug](https://github.com/pdlug)! - Support `@libsql/client` 0.18 local clients. From 0.18 a local client pools its connections and rolls back any transaction a single `execute()` leaves open, so the raw `BEGIN`/`COMMIT` framing used for local clients failed every transaction with "no transaction is active". `createLibsqlBackend()` now probes whether a local client's `execute()` calls share one session and frames transactions through `client.transaction()` when they do not; clients before 0.18 keep raw `BEGIN`/`COMMIT`.
## 0.69.0
### Highlights
TypeGraph 0.69 can copy one graph namespace into a separately allocated PostgreSQL database without discarding its recorded history. `forkGraphNamespace()` verifies a repeatable-read source snapshot against the target before commit and records a durable proof for exact retries. Durable branches can carry recorded fork points, allowing incremental merge planning to use changes since that point when lineage proves them complete.
Revision-tracked stores without recorded history can now use a revision-change journal for bounded changed-key lineage. The schema owner installs the journal, while runtime reads verify its readiness without running DDL. Disposable working copies can opt out with `revisionJournal: false`, including clones created by `branchForEvolution()`. PostgreSQL backends opened over transaction handles serialize statements on their pinned connection, and contribution-marker reads inside transactions use that same session.
Schema tooling can inspect a graph extension without opening a Store through `introspectGraphExtension()`. Linear traversal queries also carry their final hop directly into the projection.
### Upgrade notes
- Adopt base schema version 4 with the schema owner before deploying runtime roles. On existing PostgreSQL databases, run the generated migration or open once with privileged `createStoreWithSchema()` or `createAdapterStoreWithSchema()`. Older library versions refuse the newer base-schema marker.
- If a revision-tracked store without history needs journal-backed lineage, call `installRevisionChangesJournal()` once as the schema owner before runtime use. Without a ready journal, lineage raises `REVISION_JOURNAL_NOT_READY`; pass `revisionJournal: false` for a working copy that does not need it. Journal triggers capture every graph writing to their physical tables and rows have no automatic retention, so plan storage and retention before installing them on shared tables.
- Install the namespace fork ledger with `installNamespaceForkLedger()` on a private target before calling `forkGraphNamespace()`. Allocate an independent target database and size the worker for the largest copied relation; the source holds one repeatable-read snapshot for the full copy.
- `Store.clear()` now preserves graph-local contribution materialization markers by default. Pass `{ preserveContributionMaterializations: false }` when a full cutover purge must remove them.
- Custom engine profiles adopting base schema version 4 need revision-change table and index DDL. To enable journal-backed lineage, also provide trigger installation and a readiness probe.
### Minor Changes
- [#738](https://github.com/nicia-ai/typegraph/pull/738) [`8f26c78`](https://github.com/nicia-ai/typegraph/commit/8f26c78ec021668281e7f4dd44e743f32d779b73) Thanks [@pdlug](https://github.com/pdlug)! - Add history-preserving PostgreSQL graph namespace forks, store-free graph-extension introspection, recorded fork points for incremental merge, and bounded change enumeration for revision-tracked stores. Linear traversal queries now read their final hop directly. PostgreSQL transaction backends and bare client sessions serialize statements on their pinned connection; transaction marker checks read that same connection.
Install the revision-change journal with `installRevisionChangesJournal()` during privileged schema setup. Runtime lineage verifies that its table and triggers are ready without running DDL; short-lived clones, including `branchForEvolution()` working copies, may set `revisionJournal: false`. Install the namespace fork retry ledger with `installNamespaceForkLedger()` on the private target before runtime use. `Store.clear({ preserveContributionMaterializations: false })` also removes graph-local contribution markers for cutover purges.
### Upgrade notes
Adopt base schema version 4 with the schema owner before deploying runtime roles. Existing PostgreSQL installations need the generated migration or a privileged `createStoreWithSchema()` / `createAdapterStoreWithSchema()` open; the new revision-change table and index are part of that base schema. Install the optional revision-change function and triggers once with `installRevisionChangesJournal()` under the owner role. Runtime lineage only checks readiness and never runs DDL; a revision-tracked store without history throws `REVISION_JOURNAL_NOT_READY` when the journal is missing. Set `revisionJournal: false` for clones that do not need journal-backed lineage, including the fourth `branchForEvolution()` argument.
`Store.clear()` preserves graph-local contribution materialization markers by default; pass `{ preserveContributionMaterializations: false }` to remove them during a full cutover purge. Journal triggers attach to whole physical tables, so on shared tables they record writes for every graph using those tables, and journal rows have no automatic cleanup or retention policy. Avoid enabling the journal on shared tables unless that cross-graph capture and unbounded retention are acceptable.
`forkGraphNamespace()` holds one repeatable-read source transaction open for the entire copy, including row reads, target inserts, and digest checks. Long-running copies therefore retain the source snapshot until the copy finishes.
Custom engine profiles need revision-change table and index DDL for base-schema version 4 adoption, and trigger DDL plus a readiness probe to enable the change journal. Missing dependencies raise `ConfigurationError` when those operations are requested.
## 0.68.1
### Patch Changes
- [#735](https://github.com/nicia-ai/typegraph/pull/735) [`149db17`](https://github.com/nicia-ai/typegraph/commit/149db17e9042ee461908b9677f7c6c5c7f7bcbd6) Thanks [@pdlug](https://github.com/pdlug)! - `cloneWorkingCopyStrategy` now imports with `onUnknownProperty: "allow"`, so a working copy — and therefore `branch()`, `ingestionBranch()`, and `planCandidateWriteSet()` — can be seeded from live rows that carry undeclared properties `validateStore()` already reports as healthy. Incoming candidate write-set documents remain strict. A streamed interchange abort now names the failing entity and property in the thrown message instead of wrapping only a generic abort.
- [#737](https://github.com/nicia-ai/typegraph/pull/737) [`3428764`](https://github.com/nicia-ai/typegraph/commit/34287643aa708ed53caf56c8216c50042aef6441) Thanks [@pdlug](https://github.com/pdlug)! - `compareAndSet()` and `updateWhere()` no longer throw an untyped Zod error on node kinds whose schema has object-level refinements. Early field checks reconstruct a partial schema from `.shape` so refinements stay on the complete after-image, the same document `update()` already validates.
## 0.68.0
### Highlights
TypeGraph 0.68 adds atomic operations to durable graph-merge branches. `operateDurableBranch()` lets a durable host commit an opaque graph mutation and immutable evidence in one host transaction, so a process can recover and deliver committed work after a crash without inventing a second coordination protocol. Canonical request digests make retries exact: the same idempotency key replays its evidence, while a changed mutation or metadata payload conflicts without another write.
The new evidence lifecycle is inspectable and bounded. Applications can read or page committed evidence, mark delivery monotonically, and ask whether any evidence remains undelivered. TypeGraph validates every host-returned outcome and preserves a typed destruction fence until downstream delivery is complete; unsupported hosts execute no mutation, and existing durable strategies remain valid without the optional capability.
### Upgrade notes
- Existing `DurableWorkingCopyStrategy` implementations require no changes unless they opt into `operations`. To opt in, commit the host mutation and its immutable evidence in one transaction, attest the descriptor's sealed origin, enforce exact idempotency replay and digest conflicts, and serialize operations against destruction.
- Treat only `applied` and `replayed` outcomes from `operateDurableBranch()` as committed. An `unsupported` outcome guarantees that the host ran no mutation SQL; do not recreate the atomic guarantee with callbacks or a separate evidence write.
- Deliver committed evidence with `scanDurableOperations()` or `getDurableOperation()`, then call `markDurableOperationDelivered()` only after the downstream transaction commits. A first application must remain undelivered, and `destroyDurableBranch()` refuses while any evidence is undelivered.
### Minor Changes
- [#731](https://github.com/nicia-ai/typegraph/pull/731) [`f8e800b`](https://github.com/nicia-ai/typegraph/commit/f8e800bca25f7168747624d1ab4dea525ec736c0) Thanks [@pdlug](https://github.com/pdlug)! - Add atomic durable-branch operations. A `DurableWorkingCopyStrategy` may now expose an optional `operations` capability that commits an opaque host mutation and its immutable evidence in one host transaction, keyed by idempotency. New public orchestrators `operateDurableBranch()`, `getDurableOperation()`, `scanDurableOperations()`, `markDurableOperationDelivered()`, and `durableBranchHasUndeliveredEvidence()` wrap it, with typed `DurableOperationError` subclasses for conflicts, unsupported capabilities, malformed evidence, and the undelivered-evidence destroy fence.
## 0.67.1
### Patch Changes
- [#729](https://github.com/nicia-ai/typegraph/pull/729) [`0be1336`](https://github.com/nicia-ai/typegraph/commit/0be1336dbbfb0dedfbe5f473936dbabcd6e4be77) Thanks [@pdlug](https://github.com/pdlug)! - `store.clear()` now deletes the graph's durable contribution-materialization markers along with every other graph-scoped row. A cleared graph no longer leaves graph-local marker rows (full markers for graph-scoped contributions and activation markers for deployment-scoped ones) behind on SQLite or PostgreSQL; deployment-scoped physical markers are preserved because they attest shared storage the per-graph delete never touches, and the next privileged boot re-records the graph-local rows from them. On backends with interactive transactions, the delete runs in the same transaction as the rest of `clearGraph`; it also tolerates the marker table's absence on databases that never materialized a contribution.
## 0.67.0
### Highlights
TypeGraph 0.67 adds durable graph-merge branches for workflows that outlive one process. `branchDurable()` creates a persistent working copy and a JSON-safe descriptor that can cross a queue, deployment, or machine boundary; `reopenDurableBranch()` restores the ordinary `GraphBranch` used by merge planning, and `destroyDurableBranch()` explicitly removes or archives the host allocation. The existing reviewable plan/apply lifecycle remains the source of truth for accepted graph changes.
Durable strategies own host allocation and reconnection while TypeGraph validates the sealed graph, schema, branch, and base origin. Creation fences writes racing the allocation and accepts either an exact revision token or a complete semantic match when the persistent copy has its own revision namespace. Strategies can optionally attempt a proven-equivalent native merge; a mutation-free `unsupported` result returns to the complete portable apply path, while uncertain native failures never risk a second application.
### Upgrade notes
- Existing `branch()` and portable merge workflows require no changes. Use the durable APIs only when a working copy must survive closing its current backend or move between processes.
- Custom durable strategies must return a non-secret, JSON-safe locator and attest the complete sealed origin on reopen and destroy. Each opened working copy must provide either sound cross-client engine fencing or an allocation-wide exclusive writer lease; closing releases access but intentionally leaves the persistent allocation available until `destroyDurableBranch()` succeeds.
- Treat `DurableWorkingCopyStrategy.merge()` as an optional optimization. Return `unsupported` only when no merge SQL or host mutation ran. Return `applied` only after proving the complete host diff equals the approved TypeGraph write set and validating the target fence on the merged resource. Throw on failed or uncertain native outcomes; TypeGraph will not fall back after a possibly partial application. Apply callbacks and persisted provenance continue through the portable path.
### Minor Changes
- [#726](https://github.com/nicia-ai/typegraph/pull/726) [`595e6d9`](https://github.com/nicia-ai/typegraph/commit/595e6d9e5a68faf5ca118288bf921065b822ff61) Thanks [@pdlug](https://github.com/pdlug)! - Add durable graph-merge branches that can be serialized, reopened in a later process, and explicitly destroyed. Durable strategies attest the complete fork origin, prove the created working copy matches its stamped base, and declare either engine-level fencing or an allocation-wide exclusive writer lease. Approved plans can optionally use a strategy's proven-equivalent native database merge; unsupported native dimensions execute no host mutation and fall back to the complete portable plan application.
## 0.66.1
### Highlights
TypeGraph 0.66.1 fixes store-opening failures when upgrading older SQLite or PostgreSQL databases that never received the recorded-node and recorded-edge tables. Base-schema adoption now creates the missing tables and indexes while preserving existing graph data and custom table names, including when recorded history is not enabled.
### Upgrade notes
- If an earlier upgrade failed with a missing recorded-table error, deploy 0.66.1 and retry your normal privileged store-open or `backend.adoptBaseSchema()` path. Adoption resumes from the installed marker, including databases left at base-schema version 2.
- After upgrading and verifying successful adoption, remove any extra `bootstrapTables()` call added specifically to work around this failure. Keep bootstrap calls required by your normal provisioning workflow; runtime-only, least-privilege store opening still does not perform adoption.
- Roll affected deployments forward. Do not roll back to 0.56.0 after the base-schema marker has advanced beyond version 1: that release refuses the newer marker.
### Patch Changes
- [#724](https://github.com/nicia-ai/typegraph/pull/724) [`e3ebddd`](https://github.com/nicia-ai/typegraph/commit/e3ebddda8f37833255dcdd75cd38b5f24a02c2e5) Thanks [@pdlug](https://github.com/pdlug)! - Fix upgrades from legacy databases that lack recorded-node or recorded-edge tables. Version-3 base-schema adoption now creates these tables and their structural indexes before installing changed-since indexes on SQLite and PostgreSQL, preserving existing data and custom table names. Failed upgrades left at base-schema version 2 can retry through normal adoption without an explicit `bootstrapTables()` workaround.
## 0.66.0
### Highlights
TypeGraph 0.66 adds `tx.writeNodeUpsertBatch()` for recorded PostgreSQL transactions. Submit caller-assigned IDs spanning multiple plain node kinds and receive ordered postimages from one statement, while inserts, live updates, resurrections, history capture, and receipt counts remain atomic. The narrow envelope is designed for latency-sensitive heterogeneous node ingestion and preserves the existing per-entity pipeline for constrained writes.
Recorded transactions also lease their schema-fence evidence across managed writes and capture flushes. Reusing that evidence removes repeated fence probes from history-enabled adopted transaction loops while retaining conservative per-write fencing for non-history adopted transactions.
### Upgrade notes
- Use `tx.writeNodeUpsertBatch(entries)` only inside a recorded PostgreSQL transaction. The batch supports plain node kinds with caller-assigned IDs; it refuses operational-identity graphs, unique or disjointness claims, search or embedding projections, temporal options, unchanged-upsert coalescing, duplicate `(kind, id)` entries, stale schema fences, oversized batches, and unsupported backends. Split oversized inputs before retrying.
- Custom backends may implement the optional exact-session heterogeneous upsert capability to support this method. Backends that omit it retain the existing portable write paths, and `tx.writeNodeUpsertBatch()` refuses with a typed unsupported-capability error.
### Minor Changes
- [#722](https://github.com/nicia-ai/typegraph/pull/722) [`8ad6da1`](https://github.com/nicia-ai/typegraph/commit/8ad6da18fb352a7f593f0092f006e47b21ed7709) Thanks [@pdlug](https://github.com/pdlug)! - Add `tx.writeNodeUpsertBatch()` for one-statement, caller-ID upserts across plain node kinds inside a recorded PostgreSQL transaction.
## 0.65.0
### Highlights
TypeGraph 0.65 makes reviewed candidate writes evolution-aware. Plan candidate data against the schema produced by a pending evolution, then apply the schema and accepted writes together in one caller-owned transaction and recorded revision. If concurrent schema or data changes invalidate the planning snapshot, TypeGraph now reports an explicit retry-and-replan outcome.
Query composition gains cold cursor pages for `batchOnce()`, portable array-membership expressions, and native tuple comparisons for eligible keyset cursors. Cursor pages can execute independently or share one statement with companion reads while preserving the same results and cursor shape, and array membership can compare against candidate-row or correlated outer-row expressions.
Eligible existing rows in `bulkUpsertById()` now update as a version-gated batch while retaining history, uniqueness, full-text, and vector synchronization. Other cases continue through the portable row-wise path, while broader caching and set-oriented reads reduce repeated work across identity repair, query execution, candidate-scoped updates, and constrained-edge imports.
### Upgrade notes
- For candidate writes planned alongside an evolution, create the target with `captureCandidateWriteSetTargetForEvolution(target, evolutionPlan)` and plan it with `planCandidateWriteSetForEvolution()`. Apply the returned artifact inside the matching evolved transaction. Treat `MergePlanningStaleError` (`GRAPH_MERGE_PLANNING_STALE`) as a concurrency signal: discard the artifact, recapture the target, and replan.
- Custom dialect adapters that support `expr.arrayContains()` with expression operands should implement the optional `jsonArrayContainsExpression` hook with the documented JSON-array semantics. Adapters that omit it remain compatible, but compiling this expression refuses with a typed configuration error.
- Custom backends may implement the optional `updateResolvedNodesBatch` member to accelerate eligible `bulkUpsertById()` updates. Preserve the expected-version gate across the complete input and return no partial result when any row is ineligible; omitting the member retains the row-wise fallback.
### Minor Changes
- [#718](https://github.com/nicia-ai/typegraph/pull/718) [`8eb7ead`](https://github.com/nicia-ai/typegraph/commit/8eb7eada54f38503c129a11ae8d2a79c88ed9b31) Thanks [@pdlug](https://github.com/pdlug)! - Batch distinct existing-row updates in `bulkUpsertById()` while preserving version guards, recorded history, uniqueness claims, full-text indexes, and vector projections.
- [#717](https://github.com/nicia-ai/typegraph/pull/717) [`22af384`](https://github.com/nicia-ai/typegraph/commit/22af384cfb4e37f34c56abdcbeee69090836a860) Thanks [@pdlug](https://github.com/pdlug)! - Add cold cursor-page reads that execute independently or compose with other reads in one `batchOnce()` statement.
- [#714](https://github.com/nicia-ai/typegraph/pull/714) [`6036f98`](https://github.com/nicia-ai/typegraph/commit/6036f9841da1b3388a684596f15743a83ba721e2) Thanks [@pdlug](https://github.com/pdlug)! - Plan serializable candidate write sets against a pending schema evolution so the schema and accepted data can be applied in one adopted transaction and recorded revision. Document `MergePlanningStaleError` as a retry-and-replan concurrency outcome.
- [#715](https://github.com/nicia-ai/typegraph/pull/715) [`8fe206d`](https://github.com/nicia-ai/typegraph/commit/8fe206dfd4b7f3eb62bf8e4de4afeb28b093a1e8) Thanks [@pdlug](https://github.com/pdlug)! - Add expression-level array membership for candidate and correlated row values, and optimize safe keyset cursor comparisons with native row-value tuples.
### Patch Changes
- [#719](https://github.com/nicia-ai/typegraph/pull/719) [`a15e382`](https://github.com/nicia-ai/typegraph/commit/a15e382900bef8bd0d3ef84485327fa360a0be6b) Thanks [@pdlug](https://github.com/pdlug)! - Reduce repeated work in identity closure repair, projection and relation query execution, candidate-scoped updates, and constrained-edge imports. Reused queries now cache their compiled SQL templates while preserving fresh temporal bindings, and large identity or import batches avoid duplicate component expansion and per-key cardinality reads.
## 0.64.0
### Highlights
TypeGraph 0.64 lets set-based node updates take their candidates directly from the query DSL. Pass a same-Store or same-transaction query to `NodeCollection.updateWhere({ candidates, patch })` to reuse correlated cross-kind predicates, including relationships that are not stored as edges. TypeGraph projects the query back to root node identities, intersects it with any `where` or relationship selectors, and sends the result through the existing atomic update pipeline so validation, uniqueness, history, full-text, vectors, and revisions still succeed or roll back together.
Store analysis can now share the transaction snapshot that gives its results meaning. Transaction contexts expose `describe()` and `validateStore()`, allowing population statistics and every validation page to run on the pinned session. Use repeatable-read or serializable isolation and consume all pages inside the callback when concurrent writes must not change the dataset between statements.
Shared backend storage now has an explicit deployment lifecycle. Full-text tables are physically materialized once per deployment and activated independently for each graph, so later graphs can become ready without repeating privileged DDL; vector storage remains graph-scoped. Custom backends also gain the `endpointSetRead` capability bundle and `runEndpointSetReadConformance`, giving bulk endpoint reads one declared support verdict and a portable contract test.
### Upgrade notes
- Before serving full-text traffic through a DML-only role, run `createStoreWithSchema(graph, privilegedBackend)` after upgrading so TypeGraph can attest the deployment-scoped full-text table and activate each graph that uses it. `createStore()` remains a zero-I/O attach and does not repair missing markers. Custom table-contribution strategies may set `scope: "deployment"` only when one physical table is shared across graphs; omitting `scope` preserves the existing graph-scoped behavior.
- Custom backends that support `store.edges..bulkFindFrom()` or `bulkFindTo()` must expose `findEdgesByEndpointSet` on the executing backend object and should run `runEndpointSetReadConformance` in their adapter test suite. Backends that omit the member remain valid for singleton reads, while set-oriented endpoint reads refuse with `ENDPOINT_SET_READ_UNSUPPORTED`.
### Minor Changes
- [#712](https://github.com/nicia-ai/typegraph/pull/712) [`295f646`](https://github.com/nicia-ai/typegraph/commit/295f646c0add46fbd115654790c983ddd50972e7) Thanks [@pdlug](https://github.com/pdlug)! - Add deployment-scoped contribution ownership. Shared full-text storage is physically materialized once and separately activated per graph, allowing subsequent graph opens to run without DDL privileges while vector contributions remain graph-scoped.
- [#712](https://github.com/nicia-ai/typegraph/pull/712) [`295f646`](https://github.com/nicia-ai/typegraph/commit/295f646c0add46fbd115654790c983ddd50972e7) Thanks [@pdlug](https://github.com/pdlug)! - Add the `endpointSetRead` capability bundle and the framework-agnostic `runEndpointSetReadConformance` fixture for custom backends. Bulk endpoint reads now resolve one capability verdict and refuse with a typed error when set-oriented reads are unavailable.
- [#712](https://github.com/nicia-ai/typegraph/pull/712) [`295f646`](https://github.com/nicia-ai/typegraph/commit/295f646c0add46fbd115654790c983ddd50972e7) Thanks [@pdlug](https://github.com/pdlug)! - Allow `NodeCollection.updateWhere()` to take a same-graph, same-execution-target query as its candidate source. Candidate queries can use correlated cross-kind predicates without stored edges; TypeGraph forces their root-node identity projection and intersects it with existing `where` and relationship selectors before running the ordinary atomic set-update pipeline.
- [#712](https://github.com/nicia-ai/typegraph/pull/712) [`295f646`](https://github.com/nicia-ai/typegraph/commit/295f646c0add46fbd115654790c983ddd50972e7) Thanks [@pdlug](https://github.com/pdlug)! - Expose `describe()` and `validateStore()` on transaction contexts so population statistics and validation pages can run through the pinned transaction session. Callers can request repeatable-read or serializable isolation and consume all analysis work inside one callback when they need a stable data snapshot.
## 0.63.0
### Highlights
TypeGraph 0.63 can shape bounded child records for many parents in one query. `relation.topPerPartition({ partitionBy, orderBy, limit })` chooses up to N rows independently for each parent, and `expr.collect({ id, name }, { orderBy, filter })` assembles those rows into ordered, typed record arrays. The result composes with prepared queries and `batchOnce()`, avoiding a separate child query for each parent.
Record collections retain scalar codecs, including Date and Boolean fields, and decode admitted SQL NULL fields as `undefined`. Filter missing children inside `expr.collect()` to keep a parent with no matches and return `[]`. Ranking bounds the rows returned per partition; it can still scan and sort candidate rows, so provide a stable ordering key and measure the query on representative data.
### Upgrade notes
- Implement `orderedRecordJsonArray` in custom `DialectAdapter` implementations. It must preserve record field names and scalar values, apply collection-local ordering and filtering, and return `[]` for empty input.
- Update expression AST visitors that inspect `CollectExpressionNode.operand` to handle both scalar expressions and `CollectRecordOperand` (`kind: "record"`, with named `fields`).
- Advertise `capabilities.windowFunctions: true` on custom backends only when the active engine supports the window functions used by `topPerPartition()`; otherwise the new method refuses before SQL. For repeatable winners, include a stable final key in `orderBy`, and add relation ordering if the final result order matters.
### Minor Changes
- [#708](https://github.com/nicia-ai/typegraph/pull/708) [`9d4f328`](https://github.com/nicia-ai/typegraph/commit/9d4f3280e1e917ebdbffa8bcd9271c1b82cf34a1) Thanks [@pdlug](https://github.com/pdlug)! - Add ordered record collections with `expr.collect({ field: scalarExpression }, { orderBy, filter })`. Explicit flat record projections retain named scalar fields, decode Date and Boolean values, and preserve admitted SQL NULL fields as `undefined`. Record collections work through relation composition, prepared execution, and `batchOnce()`.
Custom `DialectAdapter` implementations must add the required `orderedRecordJsonArray` method when upgrading. This method emits the ordered, optionally filtered JSON record aggregate and returns `[]` for empty input.
Consumers inspecting expression ASTs must narrow `CollectExpressionNode.operand`: it can now be a scalar expression or a `CollectRecordOperand` with `kind: "record"` and named `fields`.
- [#709](https://github.com/nicia-ai/typegraph/pull/709) [`4a6568e`](https://github.com/nicia-ai/typegraph/commit/4a6568e6db6f69664b02b7d21a3731bbc5dd4a01) Thanks [@pdlug](https://github.com/pdlug)! - Add `relation.topPerPartition({ partitionBy, orderBy, limit })` to retrieve up to N rows per parent in one query. Explicit partition keys and ordering select winners independently for each parent, and the result can feed ordered record collections, prepared queries, and batches without losing scalar codecs.
Filters, distinctness, and ranges before the stage select its candidates; filters afterward remove winners without replacement. Include a stable final ordering key for repeatable winners and add relation ordering to control the final result order. Backends must advertise `windowFunctions: true`; unsupported profiles refuse execution before SQL.
## 0.62.0
### Highlights
TypeGraph 0.62 lets schema changes, graph writes, recorded history, and application SQL commit or roll back together in a caller-owned transaction. Prepare the change with `planEvolution()` before opening the transaction, then apply it through `withEvolvedTransaction()`. Planning stays outside the write fence, and ordinary additions of kinds or optional scalar fields avoid entity scans and provisioning DDL. No-op plans can use `withRecordedTransaction()` without acquiring the exclusive evolution fence.
Evolved callbacks operate against the resulting schema and return its exact version and hash in the transaction receipt. With TypeGraph-owned history, they also support schema-only recorded checkpoints through `requestRecordedRevision()`, so applications can persist an audit event even when no entities change. After the outer commit, `refreshSchema()` publishes the reconciled Store for subsequent work. Plans can move from a cached Store to a compatible `withBackend()` request Store within the same loaded TypeGraph module.
Schema evolution also composes with approved merges. `branchForEvolution()` and `planMergeForEvolution()` prepare work against the resulting schema, allowing a merge that introduces new kinds to join the same atomic commit. Privileged adapters can provision required identity storage and vector slots inside that transaction; the default DML-only policy refuses such work before mutation.
### Upgrade notes
- Bootstrap the normal TypeGraph storage before serving adopted evolution requests. Route plans requiring identity or vector provisioning to a privileged adapter configured with `schemaProvisioning: "transactional"`; bundled adapters default to `"dml-only"`. Run generic or concurrent index maintenance explicitly after commit with `materializeIndexes()` on the refreshed Store.
- Call `planEvolution()` outside the write transaction and keep its opaque token in memory. Do not serialize, clone, or reconstruct it; transfer it only between compatible Stores from the same loaded module. A `new-kind` requirement describes an addition, not queued removal work.
- Enter `withEvolvedTransaction()` before other TypeGraph callbacks on the same native transaction. Propagate callback failures so the caller rolls back, and finish all callback reads and writes before returning: escaped transaction contexts, deferred queries, and prepared batches refuse execution after callback completion.
- Handle `SchemaFenceTimeoutError` by rolling back and retrying the entire native transaction. Replan outside the transaction after a stale-baseline refusal. Change plans use a finite fence wait, defaulting to 5,000 ms; pass `waitBudgetMs` only for change plans, since no-op plans refuse an explicit budget.
- Treat `receipt.schema` and any recorded anchor as provisional until the outer commit succeeds. Then call `refreshSchema({ ref, minVersion: receipt.schema.version })` on the cached root Store. A matching cache skips SQL and does not probe for newer versions; omit `minVersion` when a fresh lookup is needed. Refresh performs no provisioning.
- Build resulting-schema merge plans with `planMergeForEvolution()` before opening the caller transaction, using `branchForEvolution()` when the branch needs new kinds. Apply the merge inside the evolved callback before other target graph writes; old-schema merge plans are refused.
- Add `planEvolution()` and `refreshSchema()` to custom `StoreEvolution` implementations. Declare `schemaProvisioning` explicitly on custom `AdapterBackend` and `SqlEngineProfile` implementations, choosing a policy that matches the connection's intended provisioning permissions.
- Offer `adoptSchemaWriteTransaction` only when a custom adapter can prove the active caller session and provide bounded schema fencing on that same transaction. Keep PostgreSQL advisory locks transaction-scoped. Native SQLite adoption requires observable transaction state on the actual connection; noninteractive adapters and SQLite drivers without that evidence cannot adopt schema changes.
### Minor Changes
- [#705](https://github.com/nicia-ai/typegraph/pull/705) [`acf5b47`](https://github.com/nicia-ai/typegraph/commit/acf5b47bc180520cf372dd1926285d00cbd55935) Thanks [@pdlug](https://github.com/pdlug)! - Plan schema evolution outside a write transaction with `store.planEvolution()`, then apply the version-bound plan alongside graph and application writes through `store.withEvolvedTransaction()`. Plans are opaque, nonserializable capability tokens that can move between compatible Stores from the same loaded module, including `withBackend()` request Stores. They expose `baseline` and `result` schema identities and a discriminated array of schema additions and apply-time requirements. A `new-kind` entry records a graph addition and does not imply a queued removal. No-op plans can use ordinary recorded transactions; metadata-only changes avoid entity scans and provisioning DDL. Schema fence waits are bounded and expose `SchemaFenceTimeoutError` for whole-transaction retry.
Evolved callbacks use the resulting schema, support recorded revision requests, and return exact schema version/hash metadata alongside the provisional recorded receipt. Escaped TypeGraph reads and writes refuse after callback completion. Publish root Store changes after outer commit through read-only `refreshSchema({ ref, minVersion })`; a matching cached version needs no reload.
Use `branchForEvolution()` to fork an isolated branch with the planned kind set and `planMergeForEvolution()` to prepare a merge for the resulting schema and apply it inside the evolved callback. Old-schema merge plans continue to refuse. Adapters default to a DML-only policy that refuses required identity or vector provisioning before mutation. Privileged adapters configured with `schemaProvisioning: "transactional"` provision identity storage, vector tables, and contribution markers on the caller's fenced transaction session, so outer rollback removes them with the schema and graph writes. Bootstrap storage is still required, and eager index maintenance runs explicitly after commit. SQLite schema adoption requires verifiable native transaction state, and noninteractive drivers remain unsupported.
Custom implementations of the `StoreEvolution` interface must add `planEvolution()` and `refreshSchema()`. Custom `SqlEngineProfile` and `AdapterBackend` implementations must declare their schema provisioning policy explicitly; the bundled adapters default to `"dml-only"`.
## 0.61.0
### Highlights
TypeGraph 0.61 expands the query DSL from graph matching into composable SQL result shaping. Schema-aware expressions power `project()`, completed-match filters, grouping, aggregates, and correlated subqueries; projected relations can be combined, filtered, deduplicated, prepared, and batched without returning intermediate rows to application code. `expr.collect(value, { orderBy, filter })` builds ordered scalar lists in SQL, including empty per-parent lists when optional matches are filtered inside the aggregate. `count()`, `exists()`, and selected-query `first()` make common terminal reads direct.
Fetching several subgraphs is now a straightforward latency optimization: `store.batchOnce(read => roots.map(root => read.subgraph(root.id, options)))` retrieves independent bounded subgraphs in one statement. Runtime-sized arrays, singleton batches, and empty batches are supported; empty batches submit no SQL. For overlapping roots with substantial shared payloads, opt-in `shareSubgraphs: true` can also share traversal and hydration work while preserving independent results. Benchmark that option against ordinary batching for your workload; disjoint or lightly overlapping roots may not benefit.
Queries can start from an explicit list of node kinds, such as `from(["Person", "Company"], "entity")`, and return one ordered stream with compatible shared fields and kind-discriminated results. Pagination now preserves nullable sort partitions and nodes whose IDs overlap across kinds. Native null ordering and directed node index keys let applications align indexes with their actual sort and identity columns. Recursive queries can compose multiple traversal stages, stop expansion explicitly, and return paths containing kind-qualified nodes and directed edge references.
Approved merge plans can join graph writes and application SQL in one caller-owned transaction through `applyMergePlanInTransaction()`. Workflows using TypeGraph-owned recorded history can also request a checkpoint with `requestRecordedRevision()` when no entity changes are needed; the completed capture receipt supplies the recorded anchor. These additions let applications commit their own receipts alongside TypeGraph work while retaining control of the outer commit and retry boundary.
### Upgrade notes
**Queries and pagination**
- Handle `undefined` for empty-input `sum`, `avg`, `min`, and `max` results. Equality and membership predicates require compatible operands, and `countDistinct` accepts scalar string, number, Boolean, or date values; project an explicit scalar key when replacing structured JSON or array distinct counts.
- Move cross-alias conditions out of staged `whereNode()` / `whereEdge()` predicates into completed-match `where()`, accounting for optional-row filtering. Use only compatible shared properties for polymorphic predicates, grouping, and ordering; query a specific kind when a field is not shared. Low-level composed `resultPredicate` ASTs must use database-expression predicates, optionally combined with AND/OR/NOT.
- Add an explicit query `limit()` when a ranked query must cap completed rows. Candidate `k` now bounds ranked candidates only; traversal fan-out can produce more than `k` result rows, including inside set operations.
- Restart saved multi-kind cursors that lack the new `kind` identity column, including cursors from subclass-expanded sources. For traversal fan-out, include traversed-row identities in the ordering when one source node produces multiple rows. Remove query-level `limit()` / `offset()` before cursor pagination and pass only one pagination direction; conflicting options are now refused.
- Keep one source per query, use unique aliases across nodes, edges, and recursive outputs, and pass non-negative safe integers for limits and offsets. Use integer subgraph depths from 0 through 1000 and supported traversal directions and cycle policies; invalid inputs are refused rather than ignored.
- Build batched reads and set-operation operands from the executing Store or transaction context. Split requests explicitly when a batch exceeds its planning or bind budget. Custom operands exposing only `toAst()` must also supply execution provenance; prefer library-created queries. Keep legacy `select()` callbacks pure because they may be probed; use `project()` for database expressions and `map()` for transformations of decoded projected rows.
**Transaction composition**
- Call `applyMergePlanInTransaction(target, tx, plan)` inside an active callback from the same Store, before other writes to the target graph. Propagate its thrown merge error so the caller rolls back, and retry the entire native transaction if needed; the helper opens no nested transaction and performs no local retry. Use a plan with `persistProvenance: false`: persisted merge provenance cannot join this atomic unit, while report-only provenance remains available.
- Request recorded checkpoints inside a writable callback with TypeGraph-owned history and obtain the anchor from its completed capture receipt; engine-native history and read-only transactions refuse explicit allocation. For atomic application receipt persistence, use `withRecordedTransaction()` and write the receipt through the same still-open native transaction before its outer commit. Do not infer a recorded revision number before capture flush.
**Custom backends and dialects**
- Implement the new `DialectAdapter` members `safeNumericConversion`, `unboundedLimit`, `textJsonArray`, `appendTextJsonArray`, and `orderedScalarJsonArray`. The last accepts one required `{ value, valueType, orderBy, filter }` argument: admit only SQL TRUE filter results, preserve included NULL elements, and return `[]` for empty input. Handle the dedicated `kind: "collect"` node in expression visitors; collection-only options are not ordinary aggregate-node fields.
- Enable `capabilities.orderedAggregates` only after verifying ordered and filtered aggregate support on the active engine. Bundled PostgreSQL declares support; supported SQLite factories probe it. An unprobed custom or remote SQLite connection must explicitly declare verified support before using `expr.collect()`. Existing non-collection reads do not require this capability.
- Supply engine serialization or session-bound READ COMMITTED isolation evidence through the write fence for composed merge application. Managed merge callbacks and adopted application now refuse missing or unsuitable evidence; a caller-serialization assertion alone is insufficient. Update custom backends and transaction test doubles that participate in these paths.
### Minor Changes
- [#702](https://github.com/nicia-ai/typegraph/pull/702) [`81fa27b`](https://github.com/nicia-ai/typegraph/commit/81fa27b26e182b4be276255452bc4b27f3f366b7) Thanks [@pdlug](https://github.com/pdlug)! - Add `applyMergePlanInTransaction()` so applications can apply an approved merge plan, record graph receipts, and write application SQL under one caller-owned transaction and recorded-time receipt.
Merge callbacks and adopted application now refuse custom backends without engine serialization or session-bound read-committed isolation evidence. Custom backends must expose that evidence through their write fence.
Fix constrained writes in adopted SQLite history transactions by acquiring the writer slot through the internal transaction-control path while retaining capture lifetime checks.
- [#701](https://github.com/nicia-ai/typegraph/pull/701) [`141deb4`](https://github.com/nicia-ai/typegraph/commit/141deb436af2818ca45288a647ede1fe7f61a6ff) Thanks [@pdlug](https://github.com/pdlug)! - Add directed node index keys that can interleave property and system columns, enabling B-tree indexes such as `(createdAt DESC, id ASC)` while keeping covering fields last. Export `NODE_SYSTEM_COLUMN_NAMES` as the readonly runtime companion to `NodeSystemColumnName` for config generation and validation.
- [#703](https://github.com/nicia-ai/typegraph/pull/703) [`0e41ee3`](https://github.com/nicia-ai/typegraph/commit/0e41ee3703df914154e04aef815c6c354643ab87) Thanks [@pdlug](https://github.com/pdlug)! - Query a nonempty explicit list of node kinds with `from(["Person", "Company"], "entity")`. Shared fields support the existing query composition APIs, and full-node results retain kind-discriminated properties.
Multi-kind cursor pagination and streaming now use both kind and ID to preserve rows when IDs overlap across kinds. Polymorphic sources refuse predicate, grouping, and ordering fields that are missing or incompatible across their kinds. Existing multi-kind cursors may need to be restarted because their identity columns now include kind.
- [#694](https://github.com/nicia-ai/typegraph/pull/694) [`3ba5ff4`](https://github.com/nicia-ai/typegraph/commit/3ba5ff4926b3ac224f95ad4dea9c6fe7dd0e845b) Thanks [@pdlug](https://github.com/pdlug)! - Add an optional database-expression `filter` to `expr.collect(value, { orderBy, filter })`. SQL TRUE includes an element, while false and SQL NULL exclude it. Aggregate-local filtering preserves parent groups from optional traversals, so missing children can produce `[]` without removing the parent row; included NULL operands still decode to `undefined` and keep the collection's inferred element type.
Keep scalar values and explicit nonempty ordering as the collection contract. `distinct` and aggregate-local `limit` remain unsupported, and collection expressions retain their dedicated `kind: "collect"` node.
Change custom dialect adapters to accept one required `{ value, valueType, orderBy, filter }` argument in `orderedScalarJsonArray`. Apply the optional filter inside the aggregate before empty-input coalescing, preserve included NULL operands, and return `[]` for empty input. The existing `orderedAggregates: true` capability remains the declaration for filtered collections.
- [#693](https://github.com/nicia-ai/typegraph/pull/693) [`6eb34ad`](https://github.com/nicia-ai/typegraph/commit/6eb34ad68c5b854a4f2af021fc52932392299f49) Thanks [@pdlug](https://github.com/pdlug)! - Add `expr.collect(value, { orderBy: [...] })` for ordered scalar collection aggregation. Project a relation, group by its parent columns, and collect string, number, Boolean, or date values into typed readonly arrays. Collection ordering is explicit and independent of result-row ordering; duplicates and nullable elements are preserved. Empty ungrouped collection aggregates return `[]`.
Export `CollectOptions` for reusable helpers and expose collection expressions as their own `kind: "collect"` node. Ordinary aggregate nodes do not carry collection-only options.
Collection results compose with preparation, projection, and one-statement batching. Structured equality restrictions continue to apply, and collections are materialized without implicit truncation. Object elements and aggregate-local limits are outside this scalar API.
Custom dialect adapters must implement `orderedScalarJsonArray()` with ordering and empty-input semantics. Collection reads require `capabilities.orderedAggregates: true`; bundled PostgreSQL declares support, and supported preparable synchronous SQLite clients and the async libSQL factory probe for support at construction. Other unprobed SQLite connections remain unsupported unless their capability is explicitly declared after verification. Existing reads are unaffected.
- [#691](https://github.com/nicia-ai/typegraph/pull/691) [`d80a10a`](https://github.com/nicia-ai/typegraph/commit/d80a10a21c10186e4286c0867890b28940304558) Thanks [@pdlug](https://github.com/pdlug)! - Add query `count()` and `exists()` terminals and selected-query `first()`. Scalar terminals count or test the current SQL relation, including grouping, limits, and offsets, without invoking result selectors. Chained `having()` conditions now accumulate with AND. Offset-only queries compile consistently on SQLite and PostgreSQL.
Allow `batchOnce()` to accept runtime-sized readonly arrays, singleton tuples, and empty arrays. Nonempty batches execute one statement or refuse before execution; empty batches execute no statement. Independent subgraphs retain their own roots, projections, traversal windows, and results. Batches now validate graph and execution-target provenance, window-function support, the request count, and the backend's declared bind budget.
Introduce schema-aware database expressions for SQL projection, predicates, ordering, grouping, and aggregates. `project()` builds SQL once, while `map()` transforms decoded rows; legacy `select()` keeps its compatibility behavior, including callback probing; keep selectors pure. Expressions include nested JSON paths, metadata, parameters, arithmetic, coalescing, conditions, and typed correlated `$exists()` / `$scalar()` subqueries with scope and temporal validation. Explicit projections support scalar terminals, preparation, and one-statement batching. Document `batchOnce(read => roots.map(root => read.subgraph(root.id, options)))` as the recommended pattern for reducing round trips across several independent, bounded subgraphs.
Add explicit SQL relation composition for projected and aggregated results. Combine visible columns with set operations, filter and aggregate derived results, deduplicate whole projections, order output columns, and execute prepared or batched relations through shared infrastructure. Typed preparation declarations preserve binding names and values across composition boundaries. Grouped relations apply input distinctness, ordering, limits, and offsets before grouping; repeated `groupBy()` calls accumulate. Identity-only `distinctNodes()` deduplicates node identities, and relation paging and streaming require a proven unique order.
Add scoped `where()` filters for completed graph matches, independently of optional-match and recursive hop constraints. Add `stopExpansion()` with an explicit stopping-node emission policy. Preserve these stages in prepared queries, batches, and logical plans, and document ranked candidates, fanout, and distinct-entity counting.
Ranked candidate `k` no longer implicitly caps completed rows after traversal fanout, including set-operation operands. Use an explicit query `limit()` to bound the final row count.
Add opt-in shared subgraph hydration with `batchOnce(build, { shareSubgraphs: true })`. Compatible reads share a multi-root traversal and hydrated entities while preserving per-request membership, projections, temporal coordinates, edge windows, and independent result objects. Default batching retains independent plans; benchmark overlapping, payload-heavy roots before enabling sharing. All current-time reads built inside a batch use one pinned instant.
Add qualified recursive paths with `path: { format: "qualified", alias: "route" }`. The output alternates kind-qualified node references and edge references with traversal direction. Existing `path: true` and string aliases still return node-ID arrays.
Compose multiple recursive traversal stages with separate depth, path, cycle, and stop state. Later stages expand upstream source identities and preserve prior row multiplicity; final filters and ranges apply after composition. Fixed-hop stages compose before and after recursion, retaining fixed-edge properties. The first recursive stage can be optional and preserves roots without eligible endpoints. Scalar recursive-edge projections remain unsupported. Ordered recursive reads now retain their sort columns when embedded in `batchOnce()`.
Add transaction-bound `query()`, `neighbors()`, `countNeighbors()`, `subgraph()`, and `batchOnce()` reads. Every read executes through the open transaction and observes earlier writes in the callback; `tx.subgraph()` and `tx.batchOnce()` each execute as exactly one statement.
### Upgrade notes
Equality and membership predicates now require compatible operands, and invalid dynamic literals are rejected before SQL execution. Aggregate results preserve scalar field types and represent empty-input SQL NULL as `undefined`; handle absent sum, average, minimum, and maximum results. `countDistinct` accepts only string, number, Boolean, and date operands; replace structured JSON or array distinct counts with an explicit portable scalar projection. Query sources cannot be replaced mid-chain, aliases must be unique across nodes, edges, and recursive outputs, and limits and offsets must be non-negative safe integers. Cursor pagination refuses query-level limits/offsets and conflicting direction options instead of ignoring them. Subgraph depths must be integers from 0 through 1000; unsupported traversal directions and cycle policies are refused. Build batch reads from the executing Store or transaction context, and split requests explicitly if a single statement exceeds its planning budget.
Custom objects exposing only `toAst()` are no longer accepted as legacy set-operation operands: operands must also supply execution provenance. Use queries created by the same Store or transaction so graph and execution-target compatibility can be verified.
Staged `whereNode()` and `whereEdge()` predicates now refuse cross-alias references instead of compiling incorrect comparisons; use completed-row `where()` for those conditions, accounting for its optional-row filtering behavior. Raw composed `resultPredicate` ASTs must use database-expression predicates, optionally combined with AND/OR/NOT.
- [#699](https://github.com/nicia-ai/typegraph/pull/699) [`f7f7376`](https://github.com/nicia-ai/typegraph/commit/f7f7376ae27a516d93816f70815b46d0d267c835) Thanks [@pdlug](https://github.com/pdlug)! - Add `requestRecordedRevision()` to history transaction contexts so applications can create a durable recorded-time checkpoint even when a transaction makes no entity changes. Repeated requests and entity changes in the same transaction allocate a single revision, exposed through the terminal receipt.
### Patch Changes
- [#700](https://github.com/nicia-ai/typegraph/pull/700) [`872ee09`](https://github.com/nicia-ai/typegraph/commit/872ee0997883eca0153c01d30c2eb9a1cb2430e2) Thanks [@pdlug](https://github.com/pdlug)! - Emit native null placement for field ordering so matching B-tree expression indexes can satisfy the primary sort without an added null-check key.
- [#698](https://github.com/nicia-ai/typegraph/pull/698) [`cefe0b1`](https://github.com/nicia-ai/typegraph/commit/cefe0b1b45b6250c9f42d6331456e78b721d162e) Thanks [@pdlug](https://github.com/pdlug)! - Fix cursor pagination and streaming across nullable sort values. Forward and backward pages now preserve rows on both sides of a NULL partition, including tied values and queries that omit the sort field from their selected result. Existing ordering defaults remain unchanged.
## 0.60.0
### Highlights
TypeGraph 0.60 brings the set-oriented read APIs introduced in 0.59 into transaction callbacks. `TransactionContext` now provides `query()`, `neighbors()`, `countNeighbors()`, `subgraph()`, and `batchOnce()`, all bound to the open transaction so read-modify-write workflows can observe earlier writes without leaving their atomic boundary.
The transaction forms preserve the physical guarantees that matter on a held connection: `tx.neighbors()` and `tx.countNeighbors()` each execute as one statement, while `tx.subgraph()` and `tx.batchOnce()` use exact-one-statement plans. Transaction-bound `batchOnce()` composes fluent reads from `tx.query()` with neighbor, count, and subgraph reads from its callback builder, returning independently typed results in tuple order.
The surface is consistent across managed, adapter, receipt-enabled, recorded-time, measurable, and adopted transaction contexts. `withCheckedReads()` remains a root-store API: transaction reads already execute against the transaction's bound session, and callers choose the required snapshot behavior through the transaction isolation options.
### Upgrade notes
- Update hand-built `TransactionContext`, adapter transaction-context, or measurable transaction-context mocks and wrappers with `query`, `neighbors`, `countNeighbors`, `subgraph`, and `batchOnce`. Contexts created by TypeGraph provide these methods automatically.
- Replace runtime feature checks and edge-read-plus-node-load fallbacks inside transaction callbacks with the typed transaction APIs. Type callback parameters as `TransactionContext` or the appropriate adapter/measurable variant rather than as `Store`; `withCheckedReads()` is intentionally unavailable on transaction contexts.
- On PostgreSQL, request `isolationLevel: "repeatable_read"` or `"serializable"` when several separate transaction reads must observe one stable database snapshot. `tx.subgraph()` and `tx.batchOnce()` remain one statement regardless of isolation level.
### Minor Changes
- [#689](https://github.com/nicia-ai/typegraph/pull/689) [`956fd56`](https://github.com/nicia-ai/typegraph/commit/956fd560d270dc58fab687f810b2c63abd42694a) Thanks [@pdlug](https://github.com/pdlug)! - Add transaction-bound `query()`, `neighbors()`, `countNeighbors()`, `subgraph()`, and `batchOnce()` reads. Every read executes through the open transaction and observes earlier writes in the callback; `tx.subgraph()` and `tx.batchOnce()` each execute as exactly one statement.
## 0.59.0
### Highlights
TypeGraph 0.59 adds an exact-one-statement read batch for latency-sensitive request assembly. `store.batchOnce((read) => [...])` embeds two or more independent fluent queries, set operations, and batch-scoped neighbor, neighbor-count, or subgraph reads into one SQL statement, then restores their independently typed results in tuple order. There is no sequential fallback: if a read shape cannot be embedded, TypeGraph refuses it before execution instead of weakening the statement-count contract.
Relationship reads no longer require applications to load every edge or hand-roll edge-plus-node joins. `store.neighbors()` returns each visible edge with its adjacent node in one statement, supports incoming, outgoing, and bidirectional reads, and can order by edge metadata or an adjacent-node property before applying a deterministic limit. `store.countNeighbors()` performs the matching aggregate without hydrating entities. `subgraph()` gains per-edge-kind windows with their own direction, ordering, and limit, so one traversal can follow different relationship kinds in different directions and retain only the top N edges for each oriented source.
Direct and composable graph reads now share one public model without parallel `*Query` APIs. Direct `store.neighbors()`, `store.countNeighbors()`, and `store.subgraph()` execute eagerly; the corresponding `read.*` forms inside `batchOnce()` defer the same logical read so it can be embedded. Direct subgraph extraction keeps its backend-tuned plan of two statements on SQLite and three on PostgreSQL, while the batch-scoped form uses one statement on both backends. Query hooks report every submitted statement for these paths.
`store.withCheckedReads(expectedSchemaVersion, fn)` extends schema-checked reads from one query to a fluent-query block. Every `.execute()` created through the scope checks the same expected active schema version, and a mismatch escapes through one callback boundary so an application can reload its schema and retry the whole read block.
### Upgrade notes
- Update hand-built `Store`, history-store, recorded-read-store, and adapter-store mocks or wrappers that expose the complete store surface with `batchOnce`, `neighbors`, `countNeighbors`, and `withCheckedReads`. Stores created by TypeGraph provide these methods automatically.
- Use `store.batchOnce()` only for two or more independent embeddable reads. Prepared queries, queued collection reads, pagination, streaming, and writes are intentionally excluded; keep using `store.batch()` for mixed queued reads and `store.transaction()` for atomic multi-operation work.
- When adopting `withCheckedReads`, catch `SchemaChangedError` outside the callback, reload the reconciled schema, and rebuild the whole block before retrying. The scope accepts ordinary fluent `.execute()` reads; aggregates, set operations, prepared queries, pagination, streaming, and `batchOnce()` refuse rather than run without the version check.
### Minor Changes
- [#687](https://github.com/nicia-ai/typegraph/pull/687) [`542e6a3`](https://github.com/nicia-ai/typegraph/commit/542e6a3de8e8aac001563fbd082bae4abdba1007) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.batchOnce()` for exact-one-statement independent reads, with a batch-scoped builder for composable neighbor, neighbor-count, and subgraph reads. Add one-statement `store.neighbors()` and `store.countNeighbors()` APIs with edge- or adjacent-node ordering, limits, and aggregates. Add per-edge-kind direction, ordering, and limits to `subgraph()` traversal and hydration. Direct `store.subgraph()` and batch-scoped `read.subgraph()` share result semantics while choosing backend-tuned and exact-one-statement physical plans, respectively. Add `store.withCheckedReads()` to bind an expected schema version once across a fluent-query read block.
## 0.58.0
### Highlights
TypeGraph 0.58 reduces database round trips in graph read paths. `store.bulkFindEdgesTo` and its pinned-view counterpart resolve inbound edges for a set of targets across node and edge kinds, mirroring `bulkFindEdgesFrom`. Callers can replace per-target lookups with a set-oriented read while retaining input order, repeated and empty target buckets, temporal visibility, and per-input limits. For a single edge kind, the existing `edges.Kind.bulkFindTo` remains available.
Whole-node and whole-edge selections now choose a full-row fetch before executing SQL, including nested and spread selections detected during planning. Previously, a fresh query instance could issue a projected query, discover that the selector needed the complete entity, and fetch again. These selections now avoid that extra statement without requiring applications to retain query instances between requests. Selectors whose field needs depend on row values keep the existing fallback.
`executeChecked(expectedSchemaVersion)` combines a relational read with an active schema-version check in one statement snapshot. It offers an explicit alternative to probing the committed version before fetching data: a mismatch raises `SchemaChangedError` before the selector runs, even when the query returns no rows. Applications can then reload the schema and rebuild the query before retrying. The check covers that statement; it does not pin later reads in the request or replace write fences.
### Upgrade notes
- Update hand-built `Store` mocks and wrappers exposing the full store surface with `bulkFindEdgesTo`; query wrappers exposing the full executable-query surface must also forward `executeChecked`. Library-created stores, pinned views, and queries provide the new methods automatically.
- When adopting `executeChecked`, catch `SchemaChangedError`, reload the reconciled schema, and rebuild the query before retrying. Start a new transaction if the old transaction holds a repeatable-read snapshot. An expected version of `undefined` means no active schema and is distinct from version zero.
- Use checked reads for relational queries with ordinary bound values. They fetch full rows and support traversals, ordering, offsets, and limits; recursive and relevance-ranked queries require a separate schema probe. Replace named `param()` references with bound values before building a checked query.
- Custom backends adopting checked reads must supply `tableNames.schemaVersions`, naming a relation with `graph_id`, `version`, and `is_active` columns whose active row agrees with `getActiveSchema`. A missing binding raises `ConfigurationError` before SQL execution. Bundled SQLite and PostgreSQL backends supply it automatically; this release requires no database migration.
### Minor Changes
- [#683](https://github.com/nicia-ai/typegraph/pull/683) [`ce35043`](https://github.com/nicia-ai/typegraph/commit/ce35043aa9356317229d85d0c4b9998284a36e1a) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.bulkFindEdgesTo` and its pinned-view counterpart for set-oriented inbound reads across edge kinds. Detect whole-node and whole-edge selections before issuing a projected query, avoiding a redundant fetch for fresh query instances. Add `executeChecked(expectedSchemaVersion)` for a relational read and committed-schema check in one statement, with `SchemaChangedError` on mismatch, including empty results.
## 0.57.1
### Patch Changes
- [#681](https://github.com/nicia-ai/typegraph/pull/681) [`d34dd51`](https://github.com/nicia-ai/typegraph/commit/d34dd515a75ab2c34451a77b113a70ca3939ad7c) Thanks [@pdlug](https://github.com/pdlug)! - Fix upgrades from older databases that lack recorded identity-assertion storage. Base-schema adoption now creates the missing table and its structural indexes before installing the version-3 changed-since indexes, preserving existing data and honoring custom table names.
## 0.57.0
### Highlights
TypeGraph 0.57 opens the backend boundary. Until now the library shipped two backends and hardcoded what it knew about them, so reaching a third engine meant editing the query compiler. This release replaces that with a declared contract. `createPostgresBackend` and `createSqliteBackend` are now the same `createSqlBackend` applied to a bundled `SqlEngineProfile`, and a new entrypoint, `@nicia-ai/typegraph/adapters/drizzle/engine`, exports both profile builders alongside `deriveEngineProfile` — which produces a variant of a bundled profile with a bounded set of fields replaced and a typed refusal for anything else. Decisions the code used to infer from a dialect comparison are now facts a backend states: `catalog` for physical-schema introspection, `fenceSql` for the lock a fence spells, `writeFence` for the exclusion primitive the engine actually provides.
Concurrency is the part of that contract with the longest reach. `capabilities.writeFence` is a discriminated union on `mechanism` — `advisory`, `engine-serialized`, `caller-serialized`, and the new `row`, a portable exclusion for an engine with no advisory-lock primitive, backed by a small `typegraph_fences` relation. An engine that resolves write conflicts at commit rather than by blocking declares `conflict: "commit-time"`, and TypeGraph then runs under an optimistic-retry tier: every store-owned transaction that takes a fence row replays as one whole unit on a real commit-time conflict, up to three attempts, invisible to hooks and to the caller. Conflicts everywhere now have one classifier and one typed error, `TransactionConflictError`, and `store.transaction()` accepts `retry: { attempts }` so an application can ask for the same replay on its own callback — under a documented replay contract, since such a callback runs more than once.
The third thread is for engines that already implement, in the database, what TypeGraph otherwise implements in software. A backend can declare `lineage` — an opaque whole-database revision, plus the rows of a graph that changed since it — and graph-merge prunes its diff to that set instead of scanning. A backend can declare `recordedTime` and answer temporal reads from its own system-versioned tables, in which case TypeGraph builds no capture relations and runs no clock at all; every recorded read in the library now resolves through a single `RecordedReadSource` seam, so TypeGraph's own capture, an externally bound relation, and an engine's native history are three interchangeable bindings rather than three spellings of the same interval predicate. And `branch()` gains a second bundled strategy: where `cloneWorkingCopyStrategy` streams a base through public interchange into a fresh backend, `forkedWorkingCopyStrategy` hands off to a host that can copy a database itself — a file copy, `CREATE DATABASE ... TEMPLATE`, a provider's branch API. Because a fork is the same physical database rather than a replay, it carries what interchange cannot: soft-delete tombstones, `created_at`/`updated_at`, the `version` column, and the base's recorded history, so a fork can answer `asOfRecorded` for instants from before it was taken.
Two fixes land regardless of which backend you run. `Store.clear()` now rotates the graph's durable revision-origin nonce in the same transaction as the clear. Previously a graph repopulated to look the same could mint a `base@V` token byte-identical to one from before the clear, and a branch forked against that older epoch would silently pass the merge precondition against entirely different content. And a PostgreSQL availability defect in constrained edge writes is repaired: the atomic edge-claim program built one predicate arm per proposed row, so a `bulkCreate` of a few thousand rows on a kind declaring a cardinality could run for minutes, grow past two gigabytes of server memory, and ignore cancellation. Those statements now drive from a single relation and are planned once.
Both bundled backends emit the same SQL, advertise the same capabilities, and behave exactly as they did in 0.56. Every new backend member above is optional, and neither bundled profile declares `recordedTime`, so recorded time on SQLite and PostgreSQL stays TypeGraph-owned; the bundled backends continue to derive their own `lineage` from their recorded relations.
### Upgrade notes
**Applications**
- `store.transaction()` and `store.transactionWithReceipt()` now throw `TransactionConflictError` (code `TRANSACTION_CONFLICT`) instead of the raw driver error on a serialization failure or deadlock. Replace matches on SQLSTATE, driver message, or a driver error class with `instanceof TransactionConflictError`, and read the original off its `cause`.
- `MergeError` raised once merge retries are exhausted now carries a `TransactionConflictError` as its `cause`, one link deeper than before. Code matching `mergeError.cause` against a driver error must match `mergeError.cause.cause`.
- Before passing `retry: { attempts }` to `store.transaction()`, check the callback against the replay contract: it must await all of its own work, read and write only values it creates fresh on each call, cause no effect outside its own transaction, and tolerate running more than once.
- A branch forked before `Store.clear()` now fails `merge()` with `BaseVersionMismatchError` — including when the graph was repopulated to look identical. Re-branch from the post-clear store rather than reusing a pre-clear branch.
**Graph-merge and working copies**
- `WorkingCopyStrategy.create` takes a second required parameter: `create(baseStore)` becomes `create(baseStore, base)`. `branch()` passes it automatically; a direct caller passes `await computeBaseVersion(baseStore)`.
- `GraphBranch` gains a required `close()`. A hand-built branch object — a structural mock or test fixture — must supply one.
- `Store` gains a required `workingCopyOptions` getter. A structural `Store` mock or wrapper not built through `createStore` / `createAdapterStore` / `createStoreWithSchema` must implement it.
**Recorded time**
- `RecordedInstant` admits a second anchor form, `e1:` (engine-native), beside `r1:`. `recordedInstantRevision()` now throws a `ValidationError` on an `e1:` anchor, and `compareRecordedInstants()` throws when handed two anchors of different ownership forms. Use `recordedInstantWallTime()` wherever an instant may be of either form. Two engine-native anchors minted in the same millisecond compare equal; a TypeGraph anchor's per-commit counter is strict.
- `ExternalRecordedReadSource` and `TypeGraphRecordedReadSource` renamed their string discriminant from `source` to `kind` (`"external"` / `"typegraph-capture"`), freeing `source` for the seam method. Update any pattern match on the old name. `RecordedReadBinding` now names the three-member binding union; `RecordedReadSource` names the shared seam the three implement.
**Custom backends and engine profiles**
- Replace `capabilities.recordedTimeOwnership` with `EngineProvisioning.recordedTime`: ownership is now derived from that member's presence rather than hand-declared. `ENGINE_NATIVE_RECORDED_TIME_NOT_IMPLEMENTED` is removed with no replacement. A profile declaring `recordedTime` must also declare `lineage`, or `createSqlBackend` refuses with `ENGINE_PROFILE_RECORDED_TIME_REQUIRES_LINEAGE`.
- Rename resolved-plan lock calls: `sql.advisoryLock(...)` becomes `sql.acquireKeyed(...)`, and `sql.advisoryLockWithIsolation(...)` becomes `sql.acquireKeyedWithIsolation(...)`. `sql.isolationFact(...)` is unchanged.
- `FenceSql`'s three members are now all optional, since a `row`-mechanism target supplies a different subset than an `advisory` one. Reach a lock through the resolved plan's accessors rather than the members directly, or narrow for `undefined` first.
- Add `fences` to any `ResolvedSqlTableNames` object literal built by hand. A caller that only overrides names through `createSqlSchema` or a bundled factory is unaffected.
- Exhaustive switches gain new cases: `WriteFencePlan["kind"]` gains `"row"`, and `capabilities.execution.unitOfWork` gains `"optimistic-retry"`.
- A profile that omits `provisioning.catalog` produces a backend with no `catalog`, and `store.materializeIndexes()`, `store.materializeSystemIndexes()`, the recorded-time schema check, and the recorded-time migration's column read each refuse with a `ConfigurationError` naming it. Supply `catalog`, or keep off those paths.
- Trusted import on a custom PostgreSQL backend now refuses before any statement runs when the resolved write fence is `unfenced` or carries `drain: "none"` — which now includes an advisory-only declaration that previously took the table lock anyway. Declare `{ mechanism: "advisory", drain: "table-lock" }` to restore the lock.
- `SqlEngineProfile` drops `firstParty` and replaces `buildOperations` / `lateMembers` with one opaque `assembly`. Build a profile through a bundled builder or `deriveEngineProfile`; a profile literal is no longer constructible, and first-party standing is bound to the object a bundled builder returned rather than to a field.
- Supply `SqlExecutionAdapter.serializationFailure` if the engine's commit-conflict shape is not PostgreSQL's `40001` / `40P01` SQLSTATE. A registered classifier adds to the standard rules rather than replacing them.
- An `"optimistic-retry"` backend requires `AsyncLocalStorage` to tell a nested write apart from an independent one. A runtime without `node:async_hooks` is refused with `OPTIMISTIC_RETRY_REQUIRES_ASYNC_CONTEXT` at the first retried unit; interactive backends are unaffected.
**Operators**
- The base schema gains a `typegraph_fences` relation on both dialects. Run the regenerated migration SQL (`generateSqliteMigrationSQL` / `generatePostgresMigrationSQL`) against an existing database before declaring `writeFence.mechanism: "row"` against it. The two bundled mechanisms, `advisory` and `engine-serialized`, need no migration and keep working unmigrated.
### Minor Changes
- [#626](https://github.com/nicia-ai/typegraph/pull/626) [`9b17a68`](https://github.com/nicia-ai/typegraph/commit/9b17a689e109c84500b6faf1e087db001c6b780f) Thanks [@pdlug](https://github.com/pdlug)! - `@nicia-ai/typegraph/adapters/drizzle/engine` now exports `buildPostgresEngineProfile` and `buildSqliteEngineProfile`, the bundled `SqlEngineProfile` builders, so a caller can derive a variant of one instead of only consuming a finished backend. It also exports `deriveEngineProfile` (with `DerivableEngineProfileOverrides`, `DerivableEngineProfileKey`, and `DERIVABLE_ENGINE_PROFILE_KEYS`), which builds a fresh profile from a bundled one with a bounded set of fields overridden — a lock spelling, a declared capability, a resource-audit verdict, or a runtime dependency bag — refusing any other field with a typed error. `SqlEngineProfile.firstParty` is removed; first-party standing is now bound to the exact profile object a bundled builder returned rather than to a field, so a copy or derived profile never carries it forward. `SqlEngineProfile.buildOperations` and `.lateMembers` are replaced by one opaque `assembly` field, constructible only by the two bundled builders. `BackendResourceAudit` is now public on the engine entrypoint. A derived profile's overridden `fenceSql` now also backs PostgreSQL's fused schema-version + recorded-graph-write statement, not only its standalone lock sites. `FenceSql` itself shrinks to three author-supplied members — `advisoryLockExpression`, `isolationFactExpression`, and `lockTables` — with the standalone-statement forms every ordinary lock site calls (`advisoryLock`, `advisoryLockWithIsolation`, `isolationFact`) now derived by TypeGraph from the two expressions, so a backend author never spells both forms separately. The bags a derived profile shares with its base by reference (`declaredCapabilities`, `resourceAudit`, `autocommit`, `tableNames`, `fenceSql`) are frozen so mutating one through the derived profile can no longer corrupt the base's own. This entrypoint is unreleased, so none of the above is a breaking change; the two bundled backends' emitted SQL, capabilities, marks, and behavior are unchanged.
See [Authoring an engine profile](https://typegraph.dev/backend-authoring) for the derivable-field table, the refusals a custom profile can hit, and a worked example.
- [#625](https://github.com/nicia-ai/typegraph/pull/625) [`e34d53c`](https://github.com/nicia-ai/typegraph/commit/e34d53cbddc8d5872a778b1e473f44fb6c75b019) Thanks [@pdlug](https://github.com/pdlug)! - `GraphBackend` and `TransactionBackend` gain an optional `catalog` member (`BackendCatalogProbes`):
`tableExists`, `tablesExist`, `indexStates`, `dropInvalidIndex`, `columnTypes`, and an
`indexBehavior` bag (`concurrentBuilds`, `hasInvalidIndexState`, `supportsGinFamily`). `columnTypes`
reports each column as a `CatalogColumn`, `{ name, kind, declaredType }`; `declaredType` is
required, and every custom `columnTypes` implementation must populate it alongside the normalized
`kind` a comparison classifies against. `dropInvalidIndex` is a root-backend operation on an engine
with an invalid-index state: a `transaction()`-scoped PostgreSQL catalog refuses it with
`CATALOG_DROP_INVALID_INDEX_REQUIRES_ROOT_BACKEND`, since PostgreSQL refuses `DROP INDEX
CONCURRENTLY` inside a transaction block, while SQLite has no invalid-index state and stays a no-op
in both scopes. `catalog` is the one physical-schema introspection surface a store path consults
directly instead of compiling a portable query, and four call sites across three modules require
it: `store.materializeIndexes()` refuses only once its empty-candidate short circuit and the
status-table ensure step have already run; `store.materializeSystemIndexes()`, which has no
candidate short circuit, refuses only once that same status-table ensure step has run; the
recorded-time schema check and the recorded-time migration's column read likewise need it.
`EngineProvisioning` gains a matching optional `catalog` field; a profile that builds one populates
the backend's member, and a profile that omits it produces a backend with no `catalog` — those four
call sites then refuse with a `ConfigurationError` naming `catalog` instead of reaching
engine-specific SQL with nothing to spell it. `createPostgresBackend` and `createSqliteBackend` both
supply `catalog`, each transaction reading its own session's uncommitted state rather than the root
connection's.
`DialectCapabilities` gains `subgraphMembershipStrategy` (`"materialized-ids" | "inline-cte"`),
naming the plan-shape decision `store.subgraph()`'s reachable-node filter already made per
dialect: fetch the traversal closure once and filter against a fixed id list, or embed the
recursive closure in each fetch. This capability replaces an inline dialect comparison in
`store/subgraph.ts`; emitted SQL, round-trip counts, and the resulting query's prepared-plan
shape are unchanged for both bundled backends.
The dialect-literal ESLint ban (previously scoped to the query compiler) now also covers
`src/backend` and `src/store`, behind a named, ratcheted exemption inventory
(`DIALECT_LITERAL_EXEMPTIONS` in `eslint.config.mjs`) asserted against the tree in both
directions by `tests/dialect-literal-inventory.test.ts`. Every remaining exemption is a decision
that is not query compilation (error classification, one-shot migrations, a driver-specific
resource audit, a SQLite-only transaction write-lock flag, or the write-fence planner's own
dialect-keyed lock semantics) and carries a reason and a site count. No bundled backend's emitted
SQL, capabilities, or behavior changes.
**Behavior change:** trusted import's PostgreSQL table lock now resolves the same write-fence plan
every other lock site does, instead of unconditionally taking `LOCK TABLE ... ACCESS EXCLUSIVE`.
Trusted import now refuses up front, before any statement runs, when a custom PostgreSQL backend's
`writeFence` declaration resolves `unfenced` (no declaration present) or resolves a `lock` plan
with `drain: "none"` — this now also catches an advisory-only declaration (`{ mechanism:
"advisory", drain: "none" }`), which previously took the table lock anyway. Every refusal names
the drain that could not be satisfied. `createPostgresBackend` itself rejects a `writeFence.
mechanism: "engine-serialized"` capability override at construction (`ConfigurationError`,
'PostgreSQL backend capability overrides cannot declare writeFence.mechanism: "engine-serialized"'),
so a declaration resolving `engine-serialized` is reachable only through a custom
`SqlEngineProfile` or a hand-built PostgreSQL-dialect backend for an engine that genuinely
serializes writers; for one, trusted import now takes no relation lock at all, where it previously
took `LOCK TABLE ... ACCESS EXCLUSIVE` — the declaration states the engine serializes writers, so
trusted import's own transaction is fence enough on its own. The `WRITE_FENCE_SQL_UNAVAILABLE`
code applies only to the narrower case of a `mechanism: "advisory"` declaration with no `fenceSql`
to spell the lock; every other refusal above is `WRITE_FENCE_UNAVAILABLE`. Declare `writeFence: {
mechanism: "advisory", drain: "table-lock" }` — the bundled `createPostgresBackend` default, which
also supplies `fenceSql` — to restore the lock.
**Author-facing:** `CommonOperationStrategy` no longer carries `dynamicEdgeConvergence`. The flag it
carried — whether a convergent edge create's non-durable match may inspect JSON match fields —
moved onto `OperationFusionHooks.dynamicEdgeConvergence`, which the bundled dialect factories pass
to `buildCommonOperationOptions`. Neither `OperationFusionHooks` nor `buildCommonOperationOptions`
is exported from any entrypoint. No action is required of a backend author: a
`CommonOperationStrategy` is not author-supplyable in this release. `strategy` is absent from
`DERIVABLE_ENGINE_PROFILE_KEYS`, so `deriveEngineProfile` refuses it, and `SqlEngineProfile.assembly`
— which replaced the `buildOperations`/`lateMembers` pair, see the derivable-profiles entry below —
is branded with a non-exported symbol, so a profile cannot be built from a literal either. The
bundled builders are the only source of a strategy.
`SqlEngineProfile.graphTemplateRuntime.instantiateStatement` is a required builder: given a
template and target graph's ids and schema hashes (`InstantiateGraphTemplateSqlParams` —
`templateId`, `templateSchemaHash`, `graphId`, `schemaHash`, and the three physical table names it
reads), it must return the statement that inserts the target graph's `schema_versions` row from the
template's stored document and copies the template's contribution-marker rows into the target
graph, taking the target graph's write lock — the same key the schema-commit fence takes —
co-atomically with the insert on an engine that fences with locks. An engine whose dialect can
compose a data-modifying CTE beside the schema INSERT (PostgreSQL) folds the marker copy and the
lock into that one statement; an engine that cannot (SQLite) instead supplies the optional
`copyContributionMarkers` dep, which runs the marker copy as a second statement once the schema row
is confirmed. The bundled `postgresInstantiateGraphTemplateStatement` and
`sqliteInstantiateGraphTemplateStatement` builders (`graph-template-sql.ts`) are what
`createPostgresBackend` and `createSqliteBackend` supply to their own profiles; neither is exported,
so a custom profile reaches the same shape only by copying a bundled profile and adapting its
statement, the same as every other engine-owned SQL a profile supplies.
The `adapters/drizzle/engine` authoring entrypoint that carries `SqlEngineProfile` and
`CommonOperationStrategy` ships for the first time in this release, so neither the removed
`dynamicEdgeConvergence` field nor the required `graphTemplateRuntime.instantiateStatement` builder
ever appeared in a published version; the notes above only affect authors building a custom profile
against `main`.
- [#656](https://github.com/nicia-ai/typegraph/pull/656) [`b32b7fc`](https://github.com/nicia-ai/typegraph/commit/b32b7fc51b626aadf8a76cb6d0b3774c9ba33b8a) Thanks [@pdlug](https://github.com/pdlug)! - `GraphBackend` gains an optional `recordedTime` member (`EngineRecordedTimeMembers`): an engine
that tracks recorded (system) time itself, rather than through TypeGraph's own capture relations
and clock. `source(table, revision)` names the table expression `"nodes"` / `"edges"` /
`"identityAssertions"` reads its recorded rows from AS OF an opaque `EngineRecordedRevision`
(`{ revision, recordedAt }`) — the engine's own temporal-table syntax, with the interval already
folded in — and `revisionNow(session)` reads `session`'s own recorded-time revision: the current
COMMITTED revision on a root backend, or the PENDING revision an open `transaction()` handle's
writes will land at once it commits (the position `TransactionReceipt.recorded` is stamped from).
`requireRecordedTime` is the typed refusal for a caller that needs it and finds it absent, in the
same style as `requireLineage`. `TransactionBackend`/`EngineProvisioning` gain the matching
optional member, threaded onto every `transaction()` handle both bundled dialects build, exactly
parallel to `lineage`. A profile that declares `recordedTime` must also declare `lineage`
(engine-native history keeps no recorded relations for TypeGraph to derive a graph-merge change
delta from); `createSqlBackend` refuses otherwise (`ENGINE_PROFILE_RECORDED_TIME_REQUIRES_LINEAGE`).
Neither bundled Drizzle profile declares `recordedTime`, so `resolveRecordedTimeOwnership` derives
`"typegraph-relations"` for both today, and every recorded-time integration suite and the parity
snapshot are unchanged — the engine-native path is proven by a PostgreSQL-family simulation
(`tests/backends/postgres/engine-native-recorded-time.test.ts`, `pglite-engine-native-recorded-time.test.ts`)
that dresses TypeGraph's own recorded relations as a temporal-table expression, labeled as a
simulation rather than a real third engine, since no bundled backend implements one.
Every recorded read — the query compiler's recorded arm, `recorded-read-service.ts`'s point reads
and scans, the historical identity readers — now goes through one `RecordedReadSource` seam
(`source(table, revision)` / `predicate(prefix, revision)` / `carriesInterval`) instead of each
spelling the recorded relation swap and the `recorded_from <= r AND r < recorded_to` interval
itself. TypeGraph's own capture binding and the external `recordedRelation({ schema })` binding
both implement it as the recorded relation plus the interval predicate (`carriesInterval: true`);
a new third binding kind, built only for a store whose backend declares `recordedTime`, implements
it as the engine's own `source` with `predicate` always `undefined` (`carriesInterval: false`) —
the engine's own expression already scopes every row to exactly one revision. Emitted SQL for both
bundled backends is unchanged: no query-compiler behavior differs for a `typegraph-relations` or
external-binding store, proven by the untouched parity snapshot and the full recorded-time
integration and property-law suites.
`RecordedInstant` widens to a two-form grammar: TypeGraph's own `r1:<16-digit revision>:`, and a new engine-native `e1::` minted internally
from a `revisionNow` result. `recordedInstantWallTime` works on either form; `recordedInstantRevision`
and `compareRecordedInstants` are narrower — see Breaking below. Store construction derives
`recordedTimeOwnership` once
(`resolveRecordedTimeOwnership(backend)`, `"engine-native"` exactly when `backend.recordedTime` is
declared) and branches only where engine-native genuinely differs from TypeGraph-owned capture:
`history: true` builds the engine-native read binding and leaves the backend unwrapped — no
capture relations, no clock, no write-fence-gated clock allocation; `revisionTracking: true` is
refused regardless of whether `history` is also requested
(`ENGINE_NATIVE_REVISION_TRACKING_UNSUPPORTED` — there is no TypeGraph clock for it to advance, and
the engine's own revision is available only under `history: true`); an external `recordedRead`
binding is refused (`ENGINE_NATIVE_RECORDED_READ_UNSUPPORTED`); `store.recordedNow()`,
`store.revisionNow()`, and both transaction-commit sites that stamp `TransactionReceipt.recorded`
now read the engine's revision through one owner, `#engineRecordedInstant(session)`, called once
per transaction on the actual committing handle — never once per graph, and never unless a graph
node/edge/identity write inside the transaction actually changed a row (a mutation witness watches
the write surface itself, not the collection-level write-intent counters `receipt.writes` is built
from, so a delete of a missing id, a found-not-created `insertNodeIfAbsent`, or a coalesced no-op
upsert all leave `recorded` undefined); and `store.asOfRecorded(instant)` refuses an instant minted
under the OTHER ownership form (`RECORDED_INSTANT_OWNERSHIP_MISMATCH`) before any read compiles.
`migrateLegacyRecordedTime` refuses under engine-native ownership
(`ENGINE_NATIVE_MIGRATE_RECORDED_TIME_UNSUPPORTED`): it rewrites TypeGraph's own recorded
relations, which an engine-native backend does not have. Reconstructing identity at a recorded
coordinate — `store.identityAtCoordinate` at a past instant, and the query compiler's historical
identity traversal — is refused under engine-native ownership
(`ENGINE_NATIVE_RECORDED_IDENTITY_UNSUPPORTED`): identity history reads TypeGraph's own recorded
relations directly, which an engine-native backend does not populate. `resolveLineage` under
engine-native ownership always answers with the backend's own `lineage` (the co-required member
above), never the recorded-relations one, since there are no recorded relations to derive it from.
Public exports beside `LineageMembers`: `EngineRecordedTimeMembers`, `EngineRecordedRevision`,
`RecordedTimeSession`, `RecordedTimeBackend`, `RecordedReadSource`, `RecordedSourceTable`. `Store`
gains a readonly `recordedTimeOwnership` property, the store-level reader of the derived ownership.
Documentation: [Engine-native recorded
time](/queries/temporal#engine-native-recorded-time) covers the reader-facing contract and the
`e1:`/`r1:` rule; [Supplying `recordedTime`](/backend-authoring#supplying-recordedtime) covers what
a profile implements; the [SQLite ↔ PostgreSQL parity
matrix](/backend-setup#sqlite--postgresql-parity) and [Engine-native recorded-time
codes](/errors#engine-native-recorded-time-codes) round it out.
## Breaking
- `capabilities.recordedTimeOwnership` is removed. It was hand-declared and could fall out of sync
with what a backend actually implemented; ownership is now derived from `backend.recordedTime`'s
presence. Declare `EngineProvisioning.recordedTime` instead — its presence alone makes
`resolveRecordedTimeOwnership(backend)` answer `"engine-native"`.
- `ENGINE_NATIVE_RECORDED_TIME_NOT_IMPLEMENTED` is removed. There is no replacement code: the
interim refusal it named no longer applies to any reachable construction path now that
engine-native construction is implemented.
- `RecordedReadBinding` widens from a two-member union
(`ExternalRecordedReadSource | TypeGraphRecordedReadSource`) to three members, adding
`EngineRecordedReadSource`. `RecordedReadSource` is repurposed and newly exported: it no longer
names the binding union (that role moved to `RecordedReadBinding`) and instead names the shared
seam shape (`source` / `predicate` / `carriesInterval`) all three binding kinds implement.
- `ExternalRecordedReadSource` (the type `recordedRelation({ schema })` returns) widens: it now
carries the `RecordedReadSource` seam's `source` / `predicate` / `carriesInterval` members
alongside its existing `schema` and brand, and its string discriminant is renamed from `source`
to `kind` (`"external"`) — the `source` name was freed for the seam method. The binding is
brand-gated and built only by `recordedRelation({ schema })`, so this affects only code that
pattern-matched the old `source` discriminant on a value it produced.
- `TypeGraphRecordedReadSource` (the type `history: true` binds internally) gets the same two
changes: it widens with the `RecordedReadSource` seam's members, and its string discriminant is
renamed from `source` to `kind` (`"typegraph-capture"`). The binding is brand-gated and built
only internally, so this affects only code that pattern-matched the old `source` discriminant.
- `RecordedInstant`'s grammar widens to admit the `e1:` form alongside `r1:`, and
`RecordedInstantParts` becomes a discriminated union (`kind: "typegraph" | "engine"`) instead of
a flat `{ revision: number; recordedAt: string }`. `recordedInstantRevision(instant)` now throws
a `ValidationError` for an `e1:` anchor — there is no TypeGraph numeric revision to return; use
`recordedInstantWallTime(instant)` for a value that works on both forms.
`compareRecordedInstants(a, b)` now throws when the two anchors were minted by different
ownership forms, and compares two engine-native (`e1:`) anchors by `recordedAt` only — document
the same-millisecond tie as a caveat in your own code if you compare engine-native anchors: two
distinct engine revisions minted within the same millisecond compare equal, unlike a
TypeGraph-owned anchor's strict per-commit counter.
- [#633](https://github.com/nicia-ai/typegraph/pull/633) [`f6d5387`](https://github.com/nicia-ai/typegraph/commit/f6d5387e5fecbea08db627c17aaeb3226e9e1db9) Thanks [@pdlug](https://github.com/pdlug)! - `@nicia-ai/typegraph/graph-merge` now exports `forkedWorkingCopyStrategy`, `ForkedWorkingCopyOptions`, and `ForkHandle` — a second bundled `WorkingCopyStrategy` for `branch()`, alongside the existing `cloneWorkingCopyStrategy`. Where the clone streams the base through public interchange into a fresh backend, `forkedWorkingCopyStrategy({ fork, connect })` targets a fork-capable host: `fork(baseStore)` calls the caller's own host-level fork API (a file copy, `CREATE DATABASE ... TEMPLATE`, a hosting provider's branch call) and returns a `TFork extends ForkHandle` (an optional `dispose`), and `connect(fork)` opens a `GraphBackend` on the result. The connected backend's `close` is composed with `dispose` so the working copy's single `close()` releases both the connection and the fork, and a `connect` failure disposes the fork before rethrowing.
`Store` gains a `workingCopyOptions` getter (returning the new `WorkingCopyOptions` type, also exported) — the one place a working-copy strategy reads a store's own hooks, upsert coalescing, SQL schema, auto-refresh-statistics threshold, query defaults, and externally-bound recorded-read relation, without re-deriving them from private state. A fork inherits the base's WHOLE such option set through it, plus `history`/`revisionTracking` matched to the base's own `historyEnabled`/`revisionTrackingEnabled`. This is safe because a fork is the SAME physical database as the base, so every one of those options names something the fork also carries. The clone strategy keeps its narrower, already-documented subset (`revisionTracking` only): its fresh backend is a distinct, empty database, so a schema naming the base's tables or an externally-bound recorded-read relation would misdirect it.
`WorkingCopyStrategy.create` gains a second parameter, `base: BaseVersion` — the token `branch()` already stamped off the base store, passed through so a strategy that needs to re-validate its working copy (the fork strategy) compares against the caller's own token instead of computing a second one. This is an additive parameter on a callback type callers implement; existing implementations that ignore the second argument are unaffected. Code that invokes a strategy's `create` directly (rather than going through `branch()`) must now pass the base token too, e.g. `strategy.create(baseStore, await computeBaseVersion(baseStore))`.
The strategy asserts `computeBaseVersion(forkStore) === base` right after attaching the store, and refuses with a typed `BranchError` — carrying `forkVersion`/`baseVersion` in `error.details`, closing the backend first — when they disagree; `branch()` returns that error as the `cause` of the `BranchError` it resolves with. This proves base-token equality (schema plus a revision anchor, or a live-content fingerprint) at the instant the fork was taken, not byte-for-byte physical identity — the untracked fingerprint deliberately omits tombstones, `created_at`/`updated_at`, the `version` column, and recorded history, and providing those unchanged is the fork mechanism's own contract, not something re-verified on every branch. That is still the right fence: the merge's lost-update guard reads `version` and the diff reads tombstones/timestamps straight off the fork, so a `fork` that is not a true physical copy breaks them regardless of what the content fingerprint agrees on. Unlike a clone, a fork is never rebuilt through `exportGraphStream`/`importGraphStream`, so it preserves soft-delete tombstones, `created_at`/`updated_at`, the `version` column, and — with `history: true` — the base's recorded relations, letting a fork answer `asOfRecorded` for instants before the fork was taken.
`create()` also refuses, before ever attaching a store, when `connect()`'s backend aliases the base's own backend: the same backend object, one derived from the other through backend derivation, or two wrappers sharing one underlying connection. Only the fork is disposed in that case — never the aliased backend, which the base still owns — and the refusal is a typed `BranchError` naming `connect()`. This cannot detect a fresh backend built over the base's own connection pool when that pool audits as independent (a default-size `pg.Pool`, for example); a pooled checkout genuinely is a different connection from the pool's own perspective.
See ["Forked working copies"](https://typegraph.dev/graph-merge#forked-working-copies) for a worked strategy and the suspend hazard on hosts that reclaim idle compute.
## Breaking
- `WorkingCopyStrategy.create` now takes a second, required parameter: `create(baseStore)` becomes `create(baseStore, base)`. `branch()` passes it automatically, so this only affects code that calls a strategy's `create` directly (rather than through `branch()`) — pass `await computeBaseVersion(baseStore)` for `base`.
- `GraphBranch` gains a required `close: () => Promise` member, releasing the branch's working-copy backend (composed with a forked working copy's host-level fork, when applicable). `branch()` and `ingestionBranch()` populate it; a hand-built `GraphBranch` object (a structural mock or test fixture) must now supply one too.
- `Store` gains a required `workingCopyOptions` getter (see above). A structural `Store` mock or wrapper — one that is not built through `createStore`/`createAdapterStore`/`createStoreWithSchema` — must now implement it too.
- [#617](https://github.com/nicia-ai/typegraph/pull/617) [`fa6b468`](https://github.com/nicia-ai/typegraph/commit/fa6b4683e4f39b140089fe369747d76c52620092) Thanks [@pdlug](https://github.com/pdlug)! - `createPostgresBackend` and `createSqliteBackend` accept `fulltext: false`,
mirroring the existing `vector: false` option. The backend then advertises
no `capabilities.fulltext` and omits the fulltext CRUD/search members
(`upsertFulltext`, `deleteFulltext`, `upsertFulltextBatch`,
`deleteFulltextBatch`, `fulltextSearch`) along with `hybridSearch` and
`fulltextStrategy` instead of stubbing them, and the generated DDL and
runtime contributions never create a fulltext table for that backend.
A fulltext predicate, a `searchable()` field, `store.search.fulltext`, and
hybrid search all refuse with `UnsupportedBackendCapabilityError` (reason
`fulltext_unsupported`) against a fulltext-off backend, instead of compiling
SQL against a table that does not exist. `hardDeleteNode`'s cascade skips
the fulltext delete for such a backend rather than issuing a statement
against a missing table; every other cascade step, and every configuration
that still has a fulltext strategy, is unchanged.
This refusal is not limited to the bundled backends: any `GraphBackend` —
including a third-party one — that omits the optional fulltext members now
refuses a write to a node kind with `searchable()` fields with the same
typed error, rather than silently skipping the fulltext index sync as it
did before.
The read path keys off `capabilities.fulltext` rather than the optional
members: a third-party `GraphBackend` that implements `fulltextSearch` and/or
sets `fulltextStrategy` but never declares `capabilities.fulltext` now has a
fulltext predicate and `store.search.fulltext`/hybrid refuse with
`UnsupportedBackendCapabilityError`, where before they compiled and ran. This
also replaces the `ConfigurationError` those two call sites previously threw
against a backend with no fulltext strategy at all — callers catching on the
old class or error code should switch to `UnsupportedBackendCapabilityError`.
`capabilities.contributions.rebuild` on a fulltext-off backend tracks only
the transactional-fence condition — with no fulltext contribution to fail
the check, that condition is vacuously satisfied — and a `rebuildContribution`
call naming the fulltext contribution on such a backend refuses with a typed
`ContributionRebuildUnsupportedError` instead of running DDL against nothing.
Soft-deleting or hard-deleting a `searchable()` node now succeeds on a
fulltext-off backend instead of refusing: a delete removes data rather than
accepting a write the backend cannot index, and this backend maintains no
fulltext sidecar to issue that removal against, so there is nothing to do.
Only create and update of a `searchable()` kind refuse.
`createLocalSqliteBackend`, `createLocalPgliteBackend`, `createLocalSqliteStore`,
and `createLocalPgliteStore` accept the same `fulltext?: FulltextStrategy | false`
option and forward it to both their installation DDL and the underlying
backend factory, so a batteries-included store can skip the fulltext table
the same way a hand-wired one can.
**Upgrade note:** `fulltext: false` stops creating and maintaining the
fulltext table; it never drops one. On a database that already has fulltext
rows, disabling fulltext leaves them in place and unmaintained — a hard
delete performed while fulltext is off leaves an orphaned row behind in the
fulltext table, because `hardDeleteNode`'s cascade has no active strategy to
build a delete statement from. Re-enabling fulltext later therefore requires
the destructive contribution rebuild, `store.rebuildContribution("fulltext")`
(which drops and recreates the fulltext table), not
`store.search.rebuildFulltext()`: that method pages live nodes to recompute
their content, and a hard-deleted node has no row left for it to page, so it
never revisits — and therefore never clears — the orphan.
- [#630](https://github.com/nicia-ai/typegraph/pull/630) [`7bb743a`](https://github.com/nicia-ai/typegraph/commit/7bb743a5a4ad6ea1b3baf2066b3298ab07189842) Thanks [@pdlug](https://github.com/pdlug)! - A backend whose `capabilities.execution.unitOfWork` is `"batch"` (Cloudflare D1's `batch()`, Neon
HTTP's `transaction(queries)`) fixes every statement before the first one runs and commits them
together with no session in between. `resolveBatchWriteVerdict` in
`src/backend/capabilities/batch-write-verdict.ts` is the one place that classifies a schema-managed
write's fitness for that tier: given a write's already-proven need (an interactive callback, a
probe-then-write constraint check, Operational Identity, history, or a schema commit), it either
defers (any other tier) or returns a refusal carrying a stable `BATCH_WRITE_UNSUPPORTED` code, the
reason, and a canonical explanation. Every enforcing gate that used to word its own batch-engine
limitation independently — the constrained-write fence, `store.transaction`, Operational Identity's
atomic-backend checks, recorded-time capture's transactionability guards, and each dialect's
schema-commit refusal — now asks this one verdict for its phrasing and nests
`{ code: "BATCH_WRITE_UNSUPPORTED", reason }` under `details.batchRefusal`, so every refusal on a
batch-tier backend names the same reason in the same words. The portable schema-version fence keeps
its plain, reasonless `SCHEMA_WRITE_FENCE_UNSUPPORTED` limitation for everything it reaches that
isn't one of those five proven needs — an ineligible write kind, a derived backend, a provenance
mismatch — rather than guessing which reason, if any, applies.
A singleton node `create` with a caller-supplied id now fuses its schema fence on a batch-tier
backend exactly as a generated id already did, provided the kind carries no declared unique
constraint: the id-generation gate that existed for an interactive root's autocommit durability no
longer excludes a batch program, which commits its one statement as a unit regardless of which id it
carries. `isAutocommitSingleStatementWrite` — the separate, stricter classifier for a bundled root's
transaction-free write — is deliberately not relaxed the same way: the fused supplied-id create
instead proves `insertNodeIfAbsentWithSchemaFence` through the ordinary hooked write plan, which
already selects the correct fenced statement per id. The tombstone-resurrection write a supplied id
can fall through to is fenced immediately before it runs, so it refuses on a batch-tier target rather
than writing the row unfenced.
`tests/batch-engine-harness.ts` adds a fake D1 client and a fake Neon HTTP client, each backed by a
real engine (better-sqlite3, PGlite) wrapped in a real transaction, so batch atomicity — a rollback
on a failing statement, a stale schema version writing nothing, every refusal reason reaching its
gate — is now proven against real engine behavior instead of a mocked response. Bundled interactive
behavior, emitted SQL, and the engine-profile-parity snapshot are unchanged.
- [#638](https://github.com/nicia-ai/typegraph/pull/638) [`d57098a`](https://github.com/nicia-ai/typegraph/commit/d57098a8b5fa8416842fb493de91dbd60d5ca9e6) Thanks [@pdlug](https://github.com/pdlug)! - `GraphBackend` gains an optional `lineage` member (`LineageMembers`): an opaque, whole-database
`revision(session)` an engine can report and compare, plus `changesSince(session, revision,
graphId)`, which names every node and edge of one graph that changed — inserted, updated,
deleted, or resurrected — since that revision, or admits `{ kind: "unbounded" }` when it cannot
bound the answer. It is a query surface over graph rows: a backend's own `lineage` writes nothing
at all, and the one carve-out on the bundled recorded-relations derivation below is a graph-identity
row, not a graph row. `requireLineage` is the typed refusal for a caller that needs it and finds it
absent, in the same style as
`requireCatalog`. `TransactionBackend` gains the same optional `lineage` member (through the new
`LineageBackend` member type, mirroring `CatalogBackend`), so a profile-supplied `lineage` is
visible on a `transaction()` handle exactly as `catalog` already was, not only on the root
backend. `EngineProvisioning` gains a matching optional `lineage` field, forwarded onto the
backend unchanged; neither bundled Drizzle profile supplies one, so a store's own
recorded-relations derivation backs the capability instead (below); the `lineage` member itself
emits no SQL. The recorded relations it derives from are already part of the schema regardless of
`history`, and a DDL-running boot (`createStoreWithSchema`, unless `systemIndexes: "skip"`) now
materializes two new system indexes on them, history on or off, plus a third structural index on
the recorded identity-assertions relation. A caller that opted out with `systemIndexes: "skip"`
gets the two system indexes on the next explicit `store.materializeSystemIndexes()` call instead
of at boot — see the parity-snapshot note below for exactly what moves.
`recordedRelationsLineage(store)` derives `lineage` from a store's own recorded relations for any
store constructed with `history: true`. `revision()` reports `:` — the graph's
durable, random revision-origin nonce (`typegraph_revision_origins`, minted on demand through
`Store.revisionOriginNow()`) joined to its recorded-time clock — never the bare clock value alone:
two independently created stores that happen to share a `graphId`, or the SAME store across a
`Store.clear()` boundary, mint numerically comparable clock values, and only the origin tells them
apart. `changesSince` refuses (`unbounded`) outright on an origin mismatch against the graph's LIVE
origin row, before comparing anything numeric. Otherwise it covers every write shape a recorded
relation can express — inserts, updates, soft deletes, hard deletes, and resurrections —
deduplicated, and proves completeness directly: every integer revision between the requested one
and the graph's current clock must carry direct evidence (a `recorded_from` or a non-sentinel
`recorded_to`) in one of the three recorded relations; anything short of that — one OTHER writer
sharing the graph's clock, a `revisionTracking`-only `Store` with no `history`, having advanced it
without capturing a row, at ANY point in the span, not only the most recent one — reports
`unbounded` rather than a delta missing that writer's rows. `resolveLineage(store)` is the one
place graph-merge (and any other caller) picks a `lineage` source: the backend's own when
declared, else this recorded-relations one when history is on, else `undefined`. A new system index,
`since_idx (graph_id, recorded_from)`, backs `changesSince`'s completeness scan on the two recorded
relations, and the recorded identity-assertions relation gains a matching `since_idx` of its own
(structural, created with the table, since it is not a `materializeIndexes`-managed system index) —
that same completeness scan folds the identity-assertions relation in alongside the two recorded
relations, so a graph whose earliest captured commit only asserted an identity is never mistaken
for a gap. A database already open when this ships adopts all three indexes on its NEXT open,
through the base-schema release-3 adoption step below (`"lineage-since-index"`) — the same lazy
backfill machinery a missing system index already goes through for any OTHER caller (the two
recorded-relation indexes; the identity-assertions index is adopted only through the base-schema
step, never through `store.materializeSystemIndexes()`), and immediately for the two
recorded-relation indexes when a caller opted out of boot-time materialization with
`systemIndexes: "skip"`, via its own explicit `store.materializeSystemIndexes()` call. The parity
snapshot moves by exactly these three index declarations, plus one extra version-marker
`INSERT`/`SELECT` round trip on each of four capture scenarios on both bundled backends (bootstrap
publishing the new base-schema release below) — no other statement, and no graph-data write SQL,
changes.
`GraphBackend` adopters that ship their own `EngineProvisioning` gain a required base-schema
release: `CURRENT_BASE_SCHEMA_VERSION` advances from 2 to 3, id `"lineage-since-index"`, adopting
the three `since_idx` indexes above through `CREATE INDEX IF NOT EXISTS` (idempotent, safe to run
concurrently, and a no-op on a fresh install whose generated DDL already carries them). The bump is
one-way — there is no downgrade path — and deployment-visible: a database already stamped 3 is
untouched, one stamped 2 is caught up in place on next open, and a store built against an
`EngineProvisioning` whose adoption-step registry stops at 2 fails to construct
(`CompilerInvariantError`, "adoption registry must end at the current version"). A zero-DDL
`createVerifiedStore` attach against a database still stamped 2 refuses with
`BaseSchemaMigrationError` until `adoptBaseSchema()` runs. A custom SQL engine profile must register
a version-3 adoption step (or accept the three indexes into its own fresh-install DDL and mark the
step `bootstrap: "covered-by-generated-ddl"`) before upgrading past this release.
`base@V`'s anchor gains a third form, `engine::`, chosen when a store has no
`revisionTracking`/`history` but its backend declares `lineage` directly (a capturing store's
recorded-relations lineage never reaches this form — capture also turns revision tracking on, so
the per-graph anchor wins first). `` is the SAME durable per-graph revision-origin nonce
the revision anchor carries (`typegraph_revision_origins`, ensured at mint time on the store's own
backend); the engine's own revision is whole-database, not per-graph, so pairing it with the
per-graph origin is what keeps two independent databases whose engines coincidentally report the
same bare revision string from minting indistinguishable anchors — without it, a branch forked
from one database could satisfy the base-version precondition of an unrelated database. The
precedence — revision anchor, then engine anchor, then the compatibility content fingerprint — is
documented once, in `base-version.ts`. Re-validating an engine anchor checks the live origin row
first (`revisionOriginMatch`, the same predicate the revision anchor's guard uses) and refuses with
`BaseVersionMismatchError` ("forked from a different store") on a mismatch before ever consulting
`changesSince`; once the origin matches, a raw revision mismatch is confirmed through `changesSince`
before refusing, since the engine's revision is whole-database and an unrelated graph's commit must
not fail this graph's merge — an empty delta is tolerated as unchanged, and a non-empty delta or
`unbounded` raises `BaseVersionMismatchError` with
`details: { expectedRevision, liveRevision, changedKeys? }`, where `changedKeys` (when present) is
capped to the first 20 node keys and first 20 edge keys plus each list's own total count, never the
raw unbounded delta. One known gap: `changesSince` names only node and edge keys, so a commit
touching only a graph's current identity assertions is invisible to an engine-anchored guard and
tolerated as unchanged — the content-fingerprint and revision-anchor forms do not share this gap.
`GraphBranch` gains an optional `forkRevision`, the fork's own `lineage.revision(session)`
captured by `branch()` right after the working copy is created, with the working copy's own root
backend as the session — an origin-bearing token for the recorded-relations source, so clearing
and repopulating the FORK itself to the same revision count `forkRevision` held is caught the same
way a cleared BASE store already is, rather than looking unchanged. `diffAgainstBase` takes an
optional `pruneTo` lineage delta: when
present, each node/edge kind is read by id set instead of a full keyset enumeration, restricted
to the union of what changed on the fork since `forkRevision` and on the base since its own
`base@V` anchor. A key absent from both deltas cannot have moved since the fork point, so pruning
cannot miss a change — it only narrows how much is read. Pruning applies only when both sides can
supply a bounded delta; a hand-built branch, a store with no `lineage`, an `unbounded` answer on
either side, or either side's `changesSince` REJECTING falls back to the full diff exactly as
before. Pruning is a pure optimization: it never changes what a merge decides, only how much of
the store it reads to decide it.
`LineageMembers`' `revision`/`changesSince` each take a **session** as their first argument — the
narrowest existing execution-target type a root backend and a `transaction()` handle both satisfy
(`LineageSession`, `Pick`). An implementation MUST
run its read on the session it is given, never on a connection it closed over instead:
`assertTargetUnchanged` (`graph-merge/merge.ts`) is the concrete caller this exists for — it
reads `lineage` off the pinned transaction handle and passes that SAME handle as the session, so
the read observes the transaction's own snapshot. The one documented exception is the bundled
recorded-relations derivation's origin resolution, which is a graph-identity row rather than a
transaction-scoped fact and is deliberately resolved off `session` entirely (see the `lineage`
capability's own doc). `requireLineage` now refuses with a
`ConfigurationError` (`LINEAGE_UNAVAILABLE`) when the transaction handle carries no `lineage` of
its own, with no fallback to the root backend's `lineage`; a `lineage` a custom backend wants
honored at commit time must be threaded through `EngineProvisioning.lineage` so it reaches every
`transaction()` handle, not attached only to the root object after construction. This is not
listed under Breaking below: `lineage` shipped on this same unreleased branch, so its signature
has never been part of a published release.
## Breaking
- `BaseSchemaRuntime` (and the `CreateBaseSchemaMembersDeps` it is derived from) gains a newly
required `sinceIndexDdl` field: `readonly [string, string, string]`, three `CREATE INDEX IF NOT
EXISTS` statements in `(recordedNodes, recordedEdges, recordedIdentityAssertions)` order, built
from a dialect's own physical table names via `sinceIndexAdoptionDdl` (`src/indexes/system.ts`).
A custom `SqlEngineProfile` that builds its own `baseSchemaRuntime` must supply this field.
- [#638](https://github.com/nicia-ai/typegraph/pull/638) [`d57098a`](https://github.com/nicia-ai/typegraph/commit/d57098a8b5fa8416842fb493de91dbd60d5ca9e6) Thanks [@pdlug](https://github.com/pdlug)! - `Store.clear()` now rotates the graph's durable revision-origin nonce
(`typegraph_revision_origins`) in the same transaction as the rest of the clear, for every store
able to mint either origin-namespaced `base@V` anchor form — a store with `revisionTracking` or
`history` enabled (the TypeGraph revision anchor), AND an engine-anchored store whose backend
declares `lineage` directly with tracking off (the engine anchor). Previously `clear()` reseeded
(or, under `history`, left unseeded) only the recorded clock and left an engine-anchored store's
origin untouched entirely, so a graph repopulated after `clear()` to look the same — the same
revision COUNT for a tracked store, or a coincidentally-matching engine revision for an
engine-anchored one — could mint a `base@V` token byte-identical to one minted before the clear,
and a branch forked before the clear would silently pass the merge precondition against a base
whose entire content had been replaced.
`computeBaseVersion` and `Store.revisionOriginNow()` also now read that origin row fresh on every
call instead of caching it per `Store` instance. Two live `Store` objects can legitimately observe
the same graph, and only one of them runs `clear()` at a time; the removed cache previously let the
OTHER instance keep minting anchors from its pre-clear origin until it happened to be recreated,
so every merge into it failed at commit for no reason visible to the caller.
This closes the BASE-side half of the epoch gap; the FORK side had an equivalent one of its own —
`recordedRelationsLineage`'s `revision()` used to report the bare recorded-clock value with no
origin, so `GraphBranch.forkRevision` carried nothing to catch a cleared-and-repopulated FORK
either. That half is closed the same way, by embedding the origin directly in the bundled
`EngineRevision` token every `revision()`/`changesSince()` call now compares — see the
`lineage`-capability changeset for the token format.
## Breaking
- A branch forked from a store BEFORE `Store.clear()` now correctly fails `merge()`'s `base@V`
precondition (`BaseVersionMismatchError`) once that store has been cleared, even when the
branch is later merged against a graph repopulated to look the same — for a revision-tracked
store, the same revision count; for an engine-anchored store, a coincidentally-matching engine
revision. This was always the documented intent — a cleared store is a new epoch a pre-clear
branch cannot merge into — and is now enforced for BOTH anchor forms. Re-branch from the
post-clear store instead of reusing one forked before the clear.
- [#628](https://github.com/nicia-ai/typegraph/pull/628) [`f74582c`](https://github.com/nicia-ai/typegraph/commit/f74582c99fb3587725255959665b27da7abc9a43) Thanks [@pdlug](https://github.com/pdlug)! - Transaction conflicts (PostgreSQL serialization failures and deadlocks) are now classified by one shared predicate everywhere the store recognizes them, and reported through a new typed error, `TransactionConflictError` (code `TRANSACTION_CONFLICT`, `details: { operation, attempts }`, `cause` the driver error), exported from the package root alongside its sibling `VersionConflictError`.
**Behavior change:** `store.transaction()` and `store.transactionWithReceipt()` now throw `TransactionConflictError` — not the raw driver error — when the backend reports a conflict, with `attempts: 1`. A caller that matched the previous driver-shaped error (by SQLSTATE, message, or `instanceof` on a driver error class) must instead match `TransactionConflictError` and read the same driver error off its `cause`.
`store.transaction()` and `store.transactionWithReceipt()` accept a new option, `retry: { attempts: number }`, to have TypeGraph itself re-run the whole callback on a conflict, up to `attempts` times total, with no delay before the second attempt and a short capped, jittered backoff after. A retried callback must satisfy a replay contract — await all of its own work, read and write only values it creates fresh on each call, perform no effect outside its own transaction, and tolerate being invoked more than once — documented on the option and on the transactions guide. `HookContext` gains an optional 1-based `attempt` field (absent means `1`, so a hook context built outside the store still typechecks) so `onOperationStart` / `onBulkOperationStart` / `onQueryStart` / `onError` can tell a replay from a new operation; a rolled-back attempt's completed operations report neither `onOperationEnd` nor `onError` of their own, and a retried `transactionWithReceipt()`'s receipt reflects only the committed attempt's writes.
Graph-merge's three commit paths (the public `apply`, its incremental variant, and internal plan commits) now go through the same retry owner as `store.transaction()`, with their existing budget of three attempts. `MergeError` on exhaustion now carries a `TransactionConflictError` as its `cause` (which itself carries the driver error), one link deeper than before — a caller matching `mergeError.cause` against the driver error directly must instead match `mergeError.cause.cause`.
`capabilities.execution` gains an optional derived field, `unitOfWork?: "interactive" | "batch" | "none"`, naming how a backend groups a multi-statement write: `"interactive"` when it can hold an open callback transaction, `"batch"` when it cannot but exposes a native atomic program (an HTTP-only driver such as `drizzle-orm/neon-http`), otherwise `"none"`. Both bundled backends derive and populate it, overwriting anything a profile declared; a custom `GraphBackend` may leave it absent. Nothing in the store consumes it yet.
- [#631](https://github.com/nicia-ai/typegraph/pull/631) [`aa0d599`](https://github.com/nicia-ai/typegraph/commit/aa0d5993a3f1feb4c20ae6a85a30bc13f4f71d40) Thanks [@pdlug](https://github.com/pdlug)! - `capabilities.writeFence` gains a third keyed mechanism, `"row"`: a portable exclusion for an engine with no advisory-lock primitive, backed by a new, never-dropped base-schema relation, `typegraph_fences(key TEXT PRIMARY KEY, generation BIGINT NOT NULL)` (`INTEGER NOT NULL` on SQLite). `{ mechanism: "row"; drain: "table-lock" | "quiescent" | "none"; conflict: "wait" | "commit-time" }` carries the same `drain` fact `"advisory"` already declares, plus `conflict` — the engine fact for two writers of one fence row: `"wait"` for a lock-based engine (the second acquirer's statement blocks, exactly like an advisory lock), `"commit-time"` for an optimistic-concurrency engine (both acquirers proceed and the loser's COMMIT fails). Every keyed lock site now shares one `case "lock": case "row":` body, spelling the acquisition through the resolved plan's `sql.acquireKeyed` / `sql.acquireKeyedWithIsolation` regardless of mechanism; the fences relation's key reuses each site's existing advisory namespace verbatim (`${namespace}:${key}`), so the lock-order contract carries over unchanged. The two bundled backends still resolve `"advisory"` / `"engine-serialized"` by default — emitted write SQL for both is unchanged, and only the bootstrap DDL gains the fences relation's `CREATE TABLE`. Existing databases add it by re-running the generated migration SQL (`generateSqliteMigrationSQL` / `generatePostgresMigrationSQL`), which now include it.
`capabilities.execution.unitOfWork` gains `"optimistic-retry"`, derived (never hand-set) when the backend is interactive and its resolved write-fence plan is `"row"` with `conflict: "commit-time"`. Under that tier, every TypeGraph-owned transaction that acquires a fence row replays a real commit-time conflict as one whole unit (open, prelude, reads, writes, commit) through the retry owner introduced for `store.transaction()`, up to `OPTIMISTIC_RETRY_ATTEMPTS` (3) attempts. That covers every store-owned write — collection create/update/delete, bulk paths, `importGraph`, identity maintenance, contribution rebuild, index materialization — as well as the two backend-owned transactions that acquire the schema-commit fence row directly: graph-template instantiation and a schema commit (`commitSchemaVersion` and its three siblings). A nested write running inside an existing transaction never retries on its own (it cannot restart a transaction it does not own), so its conflict propagates unchanged to the outermost store-owned write or to `store.transaction` itself. Under `"interactive"` this changes nothing: one attempt, as before. An `"optimistic-retry"` backend requires `node:async_hooks`' `AsyncLocalStorage` to tell a nested unit apart from an independent one; a runtime without it is refused with `OPTIMISTIC_RETRY_REQUIRES_ASYNC_CONTEXT` at the first retried unit, while interactive backends are unaffected.
`SqlExecutionAdapter` gains an optional `serializationFailure?: (error: unknown) => boolean` for an engine whose commit-conflict shape is not PostgreSQL's `40001` / `40P01` SQLSTATE (or its fixed message fallback). `isSerializationFailure(error, target?)` — the one predicate every retry owner consults — checks a classifier registered against `target` first, but the classifier only ever ADDS to the standard SQLSTATE/message rules: a registered classifier that recognizes `error` wins outright, while one that declines still falls through to those rules rather than having the final word, so there remains one predicate rather than a second inline check per engine.
`onOperationStart`'s `attempt` field counts caller-owned retries of `store.transaction` only. A store-owned unit's own internal replay under `"optimistic-retry"` (a create, an update, a bulk write, `importGraph`, and the rest) is invisible to that count: `onOperationStart` / `onOperationEnd` / `onError` each fire exactly once for the outer unit, for its one committed attempt, no matter how many attempts the retry owner spent internally to reach it.
## Breaking
- `FenceStatements`'s derived standalone-statement members are renamed to what a lock site actually asks for: `advisoryLock` → `acquireKeyed`, `advisoryLockWithIsolation` → `acquireKeyedWithIsolation`. `isolationFact` is unchanged. A caller consuming a resolved plan's `sql.advisoryLock(...)` / `sql.advisoryLockWithIsolation(...)` must call `sql.acquireKeyed(...)` / `sql.acquireKeyedWithIsolation(...)` instead.
- `WriteFencePlan` gains a `"row"` arm: `{ kind: "row"; drain; conflict; sql }`. An external exhaustive `switch` on `WriteFencePlan["kind"]` (or its `default` branch, if any) now sees this case too.
- The base schema gains the `typegraph_fences` relation on both dialects, and `ResolvedSqlTableNames` — the fully-resolved table-name set `createSqlSchema` returns and `GraphBackend.tableNames` exposes — gains its required `fences` member. Code that builds a complete `ResolvedSqlTableNames` object literal by hand, rather than through `createSqlSchema` or a bundled backend factory, must add it; the corresponding `SqlTableNames` input field stays optional and defaults, so a caller only overriding table names is unaffected. A database migrated before this release needs the regenerated migration SQL run against it before declaring `writeFence.mechanism: "row"` (the two bundled mechanisms, `"advisory"` and `"engine-serialized"`, do not need it and keep working unmigrated).
- `capabilities.execution.unitOfWork` gains `"optimistic-retry"` as a possible value — an external exhaustive `switch` on it now sees this case too.
- `SqlExecutionAdapter` gains an optional `serializationFailure` member — additive, but an object satisfying this interface structurally (rather than by declaring it) may need updating if it re-implements the full member list explicitly.
- `FenceSql`'s three members — `lockTables`, `advisoryLockExpression`, `isolationFactExpression` — are now all optional: which ones a `row`-mechanism target supplies differs from an `advisory`-mechanism one. A caller that reads one of these members directly, rather than through a resolved plan's `sql.acquireKeyed` / `sql.acquireKeyedWithIsolation` / `sql.isolationFact` / `sql.lockTables`, must narrow for `undefined` before calling it; the resolved-plan accessors already refuse with a named-member error when a mechanism does not supply one.
Bisect note: the intermediate commit adding the `row` write-fence mechanism is red on one construction-inventory ratchet, fixed by the commit that follows it in the same PR; bisecting between the two will find that known-red state.
- [#607](https://github.com/nicia-ai/typegraph/pull/607) [`e966b30`](https://github.com/nicia-ai/typegraph/commit/e966b3049755fb4945a7120319f5a9726abab1b0) Thanks [@pdlug](https://github.com/pdlug)! - Add a new entrypoint, `@nicia-ai/typegraph/adapters/drizzle/engine`, exporting
`createSqlBackend` and the `SqlEngineProfile` types. `createPostgresBackend`
and `createSqliteBackend` are now each `createSqlBackend` applied to a
profile built by `buildPostgresEngineProfile` / `buildSqliteEngineProfile`.
Emitted SQL, capabilities, marks, transaction framing, and error paths are
unchanged for every configuration the two factories accepted before.
Two construction-time narrowings apply to the bundled factories as well as
to third-party profiles, because both now run through `createSqlBackend`:
- A backend whose resolved capabilities carry no `writeFence`
declaration is refused with a `ConfigurationError`
(`ENGINE_PROFILE_REQUIRES_WRITE_FENCE_DECLARATION`) that prints the one
declaration line to add. Omitting `writeFence` from a `capabilities`
override is unaffected (the factory's own declaration applies); passing
`capabilities: { writeFence: undefined }` explicitly, which previously
built a backend that resolved every write fence through a dialect fallback,
now throws at construction.
- A backend whose declaration resolves `unfenced` no longer earns
the schema-fenced-insert eligibility mark, so a schema-managed first write
on it now refuses with `WRITE_FENCE_UNAVAILABLE` instead of fusing the
insert. Schema commits on such a backend already refused, so a working
configuration is unaffected.
- [#629](https://github.com/nicia-ai/typegraph/pull/629) [`02152da`](https://github.com/nicia-ai/typegraph/commit/02152da57cf26a00cf23c96e4e0a95e218fd1d06) Thanks [@pdlug](https://github.com/pdlug)! - `capabilities.writeFence` is the write-fence declaration, a discriminated union on `mechanism`:
`{ mechanism: "advisory"; drain: "table-lock" | "quiescent" | "none" }` |
`{ mechanism: "engine-serialized" }` | `{ mechanism: "caller-serialized" }`. `mechanism` is the
exclusion primitive a backend provides; `drain` — a field of the `"advisory"` shape only — is the
separate fact of whether a caller that already excluded other writers can additionally take a
relation-wide lock on a resource a few sites protect. Both bundled backends declare it directly
(`SQLITE_CAPABILITIES`: `{ mechanism: "engine-serialized" }`; `POSTGRES_CAPABILITIES`:
`{ mechanism: "advisory", drain: "table-lock" }`), so nothing built against them changes: same
emitted SQL, same resolved plan. `resolveWriteFencePlan` validates a declared `writeFence` at
runtime — an unrecognized `mechanism`, an unrecognized `drain`, or a `drain` attached to a
serialized mechanism — and refuses with `WRITE_FENCE_DECLARATION_INVALID` naming the field and
(where applicable) the accepted values, since a plain-JavaScript backend author is not held to the
discriminated-union type the way a TypeScript caller is.
A new arm, `{ kind: "caller-serialized" }`, joins `WriteFencePlan`'s union for a deployment-level
promise that no other client writes to the backend's database while it is open.
`createPostgresBackend` accepts `writeFence: { mechanism: "caller-serialized" }` — a claim about
the deployment, not the engine — while continuing to refuse `mechanism: "engine-serialized"`
outright, since that claims the engine itself serializes writers. The promise splits into two
halves. In process, TypeGraph enforces its own half: every root member the backend classifies in a
mutation-capable class — graph-entity and sidecar writes, backend-owned bulk import, derived-data
maintenance, schema commits, table/DDL provisioning, `clearGraph`, and the raw-SQL members that can
carry an arbitrary write (`execute`, `executeRaw`, `executeStatement`,
`executeTemporaryStatement`) — plus `transaction` and `transactionWithNative`, runs through one
per-backend serialized queue, so two concurrent calls through the same pool cannot race each other;
a root write awaited from inside a `store.transaction` callback is refused
(`SERIALIZED_QUEUE_REENTRANT_SUBMISSION`) rather than left to deadlock. Adopting an externally
owned transaction (`adoptTransaction`, backing `store.withTransaction(externalTx)`) is refused
outright (`CALLER_SERIALIZED_REFUSES_ADOPTION`): its lifetime belongs to the caller, not to this
backend's queue, so there is no honest way to hold a queue slot open for it. Outside the process,
the deployment still has to hold up its half (no other client writing to the same database) since
TypeGraph cannot observe that.
`requireWriteFence` takes `requires: "keyed" | "drain"`: `"keyed"` is satisfied by every
non-`unfenced` arm; `"drain"` refuses only when the resolved plan's `drain` is `"none"`, and
`"engine-serialized"` / `"caller-serialized"` satisfy it without consulting `drain` at all.
## Breaking
- `capabilities.pessimisticLocks` and its `PessimisticLockCapabilities` type are removed — declare
`capabilities.writeFence` instead.
- `requireWriteFence`'s `requires` parameter is renamed: `"advisory-lock"` becomes `"keyed"`,
`"table-lock"` becomes `"drain"`.
- `WriteFencePlan`'s `lock` arm drops `tableLocks` and `advisoryLocks` — read `drain` instead
(`"table-lock"` means what `tableLocks: true` used to).
- `WriteFencePlan`'s `unfenced` arm drops `reason` — declaring `writeFence` leaves no shape that
resolves `unfenced` for a reason other than an absent declaration, so there is nothing left to
distinguish.
- `WriteFencePlan` gained the `caller-serialized` arm as a permanent part of the union — an
external exhaustive switch on `WriteFencePlan["kind"]` must add a case for it (or its `default`
branch, if any, now sees it too).
- `WRITE_FENCE_DECLARATION_CONFLICT` is removed — `writeFence` is the only declaration, so no two
declarations can conflict.
- `WriteFenceDeclaration` is now a discriminated union on `mechanism`, not one flat shape: `drain`
is a field of `{ mechanism: "advisory" }` only. Declaring `drain` alongside
`mechanism: "engine-serialized"` or `mechanism: "caller-serialized"` — accepted (and ignored) by
earlier commits on this same feature branch — is now refused with
`WRITE_FENCE_DECLARATION_INVALID`. Read `declaration.drain` only after narrowing
`declaration.mechanism === "advisory"`.
- [#623](https://github.com/nicia-ai/typegraph/pull/623) [`04ce22c`](https://github.com/nicia-ai/typegraph/commit/04ce22cc87ffbf4f3cc44c47408d0ff18198dd31) Thanks [@pdlug](https://github.com/pdlug)! - `GraphBackend` gains an optional `fenceSql` member: the lock spelling a backend supplies
alongside `capabilities.writeFence`, as `FenceSql` — three builders,
`advisoryLockExpression`, `isolationFactExpression`, and `lockTables`. `resolveWriteFencePlan`'s
`lock` arm carries `sql: FenceStatements`: those three plus the standalone `advisoryLock`,
`advisoryLockWithIsolation`, and `isolationFact` statements, which `resolveFenceStatements`
derives from the two expressions so the portable lock sites and the fused recorded-write fence
always spell the same key. Every write-fence lock site consumes `fence.sql.(...)`
instead of hand-writing PostgreSQL lock syntax inline.
The bundled PostgreSQL spelling is exported as `postgresFenceSql` from
`@nicia-ai/typegraph/adapters/drizzle/postgres`. `createPostgresBackend` supplies it
automatically; `createSqliteBackend` supplies no `fenceSql` since its fence is
`engine-serialized` and takes no lock. A backend that declares
`capabilities.writeFence.mechanism: "advisory"` but supplies no `fenceSql` is now refused at
construction with a typed `ConfigurationError` (`WRITE_FENCE_SQL_UNAVAILABLE`) naming the member
to supply, rather than reaching a lock site with nothing to spell the statement.
For both bundled backends the locks taken, their order, and their modes are unchanged, and
the PostgreSQL statement text is equivalent: two advisory-lock sites now bind the lock
namespace as a parameter instead of an inline string literal (`hashtext` hashes the value
either way), and insignificant whitespace in three statements changed with the move.
Two behavior changes reach custom `dialect: "postgres"` backends. A backend declaring
`writeFence.mechanism: "advisory"` without `fenceSql` is refused at construction (above).
A backend declaring only `mechanism: "engine-serialized"` and no `fenceSql` is refused when a
history-capturing transaction reads its isolation level, which previously ran a hard-coded
`current_setting('transaction_isolation')` read; supply `fenceSql` (or `postgresFenceSql`) to
restore it.
### Patch Changes
- [#662](https://github.com/nicia-ai/typegraph/pull/662) [`5fc8e84`](https://github.com/nicia-ai/typegraph/commit/5fc8e84efc9faa45a10ba49375bd211b1b05d4a7) Thanks [@pdlug](https://github.com/pdlug)! - Fix a PostgreSQL availability defect in constrained edge writes: a `bulkCreate` or `bulkInsert` on an edge kind declaring `cardinality: "one"`, `"oneActive"`, or `"unique"` could run for minutes, grow past two gigabytes of server memory, and ignore cancellation. The three statements of the atomic edge-claim program were built with one predicate arm per proposed row — each arm carrying its own two `EXISTS` and one `NOT EXISTS` subquery — so a chunk the bind budget permits (thousands of rows) asked the executor to initialize and evaluate thousands of subplans, with per-arm cost that was not constant. A 200-arm statement took seconds to plan and execute against an empty match set; a 2000-arm statement did not finish.
Each statement now drives from a single `proposed` relation of the chunk's rows, so the engine plans it once and the per-row work is an index probe. The same shape lands on both bundled backends. The stale-claim release additionally becomes robust to the plan degradation reported under stale statistics after a bulk load: its driving relation is now the bounded proposed set joined to the claim relation on its primary key, rather than a claim scan whose correlated `NOT EXISTS` the planner could demote to a whole-graph nested-loop anti-join.
The claim predicates — what a competing live edge is, and whether a recorded claim holder still satisfies its axis — now have one owner each, rendered from an explicit value source so the single-row and batched statements cannot drift. The conditional takeover statement, which previously spelled the holder-liveness predicate a second time inline, resolves it through that owner.
**Author-facing:** `CommonOperationStrategy.buildDeleteStaleAtomicEdgeClaims`, `.buildAcquireAtomicEdgeClaims`, and `.buildAssertAtomicEdgeClaimsOwned` now return `readonly SQL[]` rather than `SQL`. A chunk renders one statement per distinct declared cardinality it contains — one statement in the ordinary single-kind case — because the endpoint terms an axis key covers and the liveness a holder must still satisfy are predicate shape, not values, and folding them into the relation as guarded terms would give back the index probes this change exists to gain. No action is required of a backend author: `CommonOperationStrategy` is visible on the engine entrypoint as part of `SqlEngineProfile`'s shape, but it is not author-supplyable — `strategy` is absent from `DERIVABLE_ENGINE_PROFILE_KEYS`, so `deriveEngineProfile` refuses it, and `SqlEngineProfile.assembly` is branded with a non-exported symbol, so a profile cannot be built from a literal either. The bundled builders are the only source of a strategy, and theirs changed with the code.
- [#627](https://github.com/nicia-ai/typegraph/pull/627) [`4cec2b0`](https://github.com/nicia-ai/typegraph/commit/4cec2b0096d5fad833d32ce8b5ed6fb3978d1dc2) Thanks [@pdlug](https://github.com/pdlug)! - The PostgreSQL backend's database-extension install now resolves the write-fence plan and spells its advisory lock through the backend's `fenceSql`, like every other lock site, instead of hardcoding `pg_advisory_xact_lock(hashtext(...), 0)` inline. A custom or derived PostgreSQL profile whose resolved plan is `engine-serialized` or `unfenced` installs extensions without taking that lock and relies solely on the duplicate-key retry, which was already the fence's correctness owner in that case. The bundled PostgreSQL backend's behavior and emitted SQL are unchanged.
## 0.56.0
### Highlights
TypeGraph 0.56 adds a durable review workflow for candidate write sets. `planCandidateWriteSetReview()` captures immutable, digest-checked evidence for the candidate, original plan, policy context, normalized options, and target baseline. After an application records approval, `revalidateCandidateWriteSetReview()` compares that evidence with current graph state and returns a structured compatible, changed, or incompatible result together with a fresh revision-fenced plan when execution remains safe.
Reviewed plan application can now share one protected transaction with application-owned checks and writes. Optional `beforeApply` and `afterApply` callbacks run after the target fence is validated, use transaction-bound typed collections, and roll back with the merge if any step fails. Transaction-conflict retries replay the complete operation, including both callbacks.
Generic code can dispatch edge operations by kind through `DynamicEdgeCollection` while preserving the selected edge's property and result types and validating endpoint pairs at runtime. Transactions gain `getEdgeCollection` and `getEdgeCollectionOrThrow`, valid-time views support pinned dynamic edge reads, and generic traversal factories now preserve declared target kinds across array targets, source-dependent maps, and unions.
### Upgrade notes
- Replace generic `tx.edges[kind]` access and uncorrelated edge collection casts with `tx.getEdgeCollectionOrThrow(kind)`. Concrete `.edges.` access retains compile-time endpoint-pair checking.
- Known-kind dynamic edge lookups now enforce their property schema at compile time. Parse unvalidated records before passing them, and update hand-authored transaction or view mocks with the new lookup methods.
- `beforeApply` and `afterApply` callbacks may run again after a transaction conflict, so callback behavior must be safe to retry. Callback failures roll back the combined transaction; optional provenance persistence remains a separate post-commit operation.
- Candidate review evidence is application-authenticated and V1 uses a conservative whole-target baseline. Applications must version opaque policy and callback dependencies and decide whether a compatible revalidation still satisfies their approval policy.
### Minor Changes
- [#615](https://github.com/nicia-ai/typegraph/pull/615) [`4574be7`](https://github.com/nicia-ai/typegraph/commit/4574be7f5b7fb91428639a9bf3056801f8f15bbf) Thanks [@pdlug](https://github.com/pdlug)! - Add durable candidate merge reviews with `planCandidateWriteSetReview()` and `revalidateCandidateWriteSetReview()`. Persist immutable review and approval evidence in the target graph, then compare the retained candidate against current state before applying a fresh revision-fenced plan. Structured compatibility results expose changes requiring review without weakening atomic apply-time concurrency or constraint checks.
- [#611](https://github.com/nicia-ai/typegraph/pull/611) [`cf126e3`](https://github.com/nicia-ai/typegraph/commit/cf126e3542f585eeb18d8bc784e769b224de4ed5) Thanks [@pdlug](https://github.com/pdlug)! - Support generic edge dispatch through `DynamicEdgeCollection` and graph-aware dynamic lookups. Edge property and result types are preserved while endpoint pairs are validated at runtime. Transactions now expose `getEdgeCollection` and `getEdgeCollectionOrThrow`, including scoped receipt accounting; valid-time views expose pinned dynamic edge reads.
Migration: replace generic `tx.edges[kind]` calls or uncorrelated collection casts with `tx.getEdgeCollectionOrThrow(kind)`. Concrete `.edges.` calls retain compile-time pair checking. Known-kind dynamic lookups now enforce their property schema at compile time; parse unvalidated records before passing them. Broad `EdgeRegistration` annotations continue to permit array or map targets; narrow the target shape when inspecting it. Hand-authored transaction and view mocks must provide the new lookup methods.
Fix outgoing traversal target inference in graph factories that accept generic node types. Preserve the declared target kinds for both array targets and source-dependent maps, including unions of edge kinds.
- [#614](https://github.com/nicia-ai/typegraph/pull/614) [`5232be5`](https://github.com/nicia-ai/typegraph/commit/5232be5dd1cc465dd7fc6569d391d9fd52e517eb) Thanks [@pdlug](https://github.com/pdlug)! - Compose reviewed merge-plan application with application-owned graph checks and writes using optional `beforeApply` and `afterApply` callbacks. Prechecks receive transaction-bound read-only collections after the target fence is validated; post-apply work uses typed graph operations before the same transaction commits. Failures roll back the combined operation, and transaction-conflict retries replay both callbacks.
## 0.55.0
### Highlights
TypeGraph 0.55 adds source-dependent edge targets to the existing `from`/`to` syntax. One edge kind can now permit Employee → Department and Student → Course without admitting the cross-pairs. TypeScript write inference preserves that relationship, runtime validation rejects invalid pairs with `EndpointPairError`, and restrictions survive schema serialization, imports, graph merges, and runtime graph extensions.
Store node and edge creation and upsert inputs now accept `validFrom: null` to explicitly request an open-left validity window. Snapshot and incremental graph merges preserve those windows through canonicalization, edge repointing, and serialized merge plans instead of narrowing them to the merge commit time. Bulk-upsert coalescing also distinguishes a confirmed absent lower bound from a creation timestamp that has not yet been assigned.
Startup identity repair now checks the schema version used to derive its registry before rebuilding closure. If a concurrent migration has committed newer semantics, the stale repair raises `StaleVersionError` and leaves the newer closure intact. This covers both existing-relation repair and creation plus population of missing derived relations.
### Upgrade notes
- Existing array-valued `to` declarations retain their Cartesian-product semantics. Removing allowed pairs from a source-dependent target map is classified as a breaking schema change.
- Omitted `validFrom` retains its existing defaults. Explicit `null` requests no lower bound, while returned metadata still uses `undefined` for an absent bound. Live-row upserts retain their immutable-bound rules; `validTo` continues to use its existing set/clear protocol.
- If startup identity repair raises `StaleVersionError` after a concurrent migration, reopen with the current graph definition before retrying.
### Minor Changes
- [#609](https://github.com/nicia-ai/typegraph/pull/609) [`39a65ef`](https://github.com/nicia-ai/typegraph/commit/39a65ef8eb544942f35f97fdbfbc2609d6284bba) Thanks [@pdlug](https://github.com/pdlug)! - Accept `validFrom: null` on Store node and edge creation and upsert inputs to explicitly request an open-left validity window. Omitted lower bounds keep their existing default behavior; live-row upserts still refuse changes to an immutable lower bound.
Preserve open-left staged node and edge windows through snapshot and incremental graph merges, edge repointing, and serialized merge plans instead of narrowing them to the merge commit time. Keep repeated bulk-upsert coalescing from confusing an unknown creation timestamp with a confirmed open-left bound.
- [#605](https://github.com/nicia-ai/typegraph/pull/605) [`f7dcd17`](https://github.com/nicia-ai/typegraph/commit/f7dcd1708be8dd3ef67ce2b32a1c2ee22536d2f6) Thanks [@pdlug](https://github.com/pdlug)! - Define source-dependent edge targets with a `to` map, such as `from: [Employee, Student]` with `to: { Employee: [Department], Student: [Course] }`. This allows one edge kind to connect specific source/target pairs without admitting every combination. Array-valued `to` declarations retain their existing Cartesian-product behavior.
Allowed pairs are preserved in typed writes, runtime validation, schema serialization, imports, graph merges, and runtime graph extensions. Invalid pairs produce `EndpointPairError`, removing allowed pairs is a breaking schema change, and ontology compatibility checks account for the pair relationship. See [source-dependent targets](https://typegraph.dev/core-concepts#source-dependent-targets) for examples and lifecycle rules.
### Patch Changes
- [#608](https://github.com/nicia-ai/typegraph/pull/608) [`5b2dcc0`](https://github.com/nicia-ai/typegraph/commit/5b2dcc07bd84d3fe97959a3c86cab4e29449c5f5) Thanks [@pdlug](https://github.com/pdlug)! - Prevent startup identity repair from overwriting a newer schema migration with closure data derived from an older schema. Repair now checks the observed schema version inside its write transaction and raises `StaleVersionError` if a concurrent migration advanced it, preserving the newer closure.
## 0.54.0
### Highlights
TypeGraph 0.54 makes runtime-evolved schemas first-class across metadata, typing, analysis, conflict planning, and guarded recovery. Graph-scoped annotations now flow through `defineGraph`, `defineGraphExtension`, `SerializedSchema`, and introspection. Schema-bound runtime-kind tokens preserve extension node and edge types through collections, bulk edge reads, and one-hop traversals without consumer-side widening adapters.
New current-state Store analysis APIs keep discovery work bounded as vocabularies and datasets grow. `store.describe()` computes per-kind population and declared-property coverage in SQL, splits node and edge work, and chunks wide schemas under a fixed result-column budget. `store.validateStore()` uses keyset-paginated scans to report declared-schema violations with record ids, JSON-pointer paths, and reasons while treating undeclared fields as healthy semi-structured data. Both APIs bracket their work with a stable schema coordinate; data is live, so concurrent writes can affect different statements or pages.
`planCandidateWriteSet()` brings the existing source-attributed merge planner to branch-free candidate writes without applying them. Node collections also gain scalar `compareAndSet()` and `compareAndSetAbsent()` guards that preserve ordinary validation, hooks, history, and version bookkeeping while enforcing the expected-current predicate atomically.
### Upgrade notes
- `BulkOperationHookContext["operation"]` now includes `"compareAndSet"`; exhaustive consumers must handle the new case.
- Store analysis is current-only and intentionally absent from `StoreView` and transaction callback facades. Calling `store.describe()` through an enclosing root Store inside a transaction callback does not enlist the analysis in that transaction.
- During a mixed-version rollout, upgrade every process that can write schema versions to 0.54 before enabling graph-scoped annotations. TypeGraph 0.54 preserves unknown top-level serialized-schema fields during reconciliation; older writers do not provide that guarantee.
### Minor Changes
- [#601](https://github.com/nicia-ai/typegraph/pull/601) [`f7d1ac4`](https://github.com/nicia-ai/typegraph/commit/f7d1ac44d105f6b1aedabb4c70917dbc57570c1a) Thanks [@pdlug](https://github.com/pdlug)! - Add first-class graph annotations, Store population statistics and validation,
branch-free candidate write-set conflict planning, schema-bound runtime-kind
tokens for typed collections and traversals, and guarded node compare-and-set.
This release also preserves unknown top-level serialized-schema fields across
reconciliation, reports no-op schema diffs without a phantom version delta,
ships the package changelog in the npm tarball, and clarifies base-schema
adoption and transaction-scoped mutation performance. The bulk-operation hook
context now also reports `"compareAndSet"`; exhaustive consumers of its
`operation` field must handle the new case.
## 0.53.0
### Highlights
TypeGraph 0.53 substantially reduces database exchanges for write-heavy applications running near Neon, Cloudflare D1, libSQL, or PostgreSQL. Compared with the former 5–6-exchange managed-write shape, eligible singleton updates and deletes now use one authoritative read or gate plus one atomic mutation submission, about 60–67% fewer exchanges. Eligible bulk creates, inserts, deletes, complete-document replacements, and durable edge convergence submit their complete mutation unit once; eligible upserts use one batched preimage read plus one atomic submission. These reductions apply to root writes outside `store.transaction()`; transaction-scoped writes remain on the interactive path and see no exchange-count change. These are transport-submission counts rather than wall-clock benchmarks. See the [performance guide](https://typegraph.dev/performance/overview) for the exact eligibility envelopes and fallback behavior.
The atomic programs now carry uniqueness and disjointness claims, full-text and vector projections, endpoint and schema fences, postimage assertions, and typed rollback diagnosis in the same submission. Large D1 upserts remain atomic across bind-sized statements, with supported ceilings raised to 512 nodes and 187 edges per call. PostgreSQL transaction sessions use the same programs through a savepoint, while unsupported shapes continue through the complete portable path.
This release also adds [`nodes..bulkReplaceById()`](https://typegraph.dev/schemas-stores#bulkreplacebyiditems) for read-free, complete-document replacement; makes graph-template registration and instantiation safe across serverless isolates; repairs externally provisioned pre-0.52 base storage through a numbered deployment-wide lifecycle; and publishes transport plus semantic conformance runners for custom atomic backends.
### Upgrade notes
- Existing or externally provisioned databases must be opened once with privileged `createStoreWithSchema()` / `createAdapterStoreWithSchema()`, or receive TypeGraph's generated base-schema migration, before any DML-only runtime opens. Applications that already use either privileged opening path adopt the base schema automatically during that open; no separate bootstrap command is needed. The deployment invariant is ordering: privileged adoption must finish before DML-only workers start. See [Upgrading deployment-wide base storage](https://typegraph.dev/backend-setup#upgrading-deployment-wide-base-storage).
- The removed top-level `capabilities.transactions` override is now refused instead of ignored. Use `capabilities.execution.interactiveTransactions`.
- `store.transaction()` and `store.transactionWithReceipt()` now refuse backends without interactive transaction support. Call ordinary Store write methods directly when the application intentionally owns partial-failure recovery.
- Eligible singleton updates now converge optimistically and retry a moved preimage up to four times. Under sustained same-row contention they can throw `DatabaseOperationError` where an interactive backend previously serialized the writers.
- A PostgreSQL-dialect backend that declares neither usable pessimistic locks nor serialized writers can no longer open a schema-managed store. TypeGraph now reports `WRITE_FENCE_UNAVAILABLE` rather than accepting a declaration it cannot honor. First-party PostgreSQL, Neon, PGlite, SQLite, D1, and libSQL configurations retain their existing fence posture. See [Write fence declaration](https://typegraph.dev/backend-setup#write-fence-declaration-pessimisticlocks).
### Minor Changes
- [#574](https://github.com/nicia-ai/typegraph/pull/574) [`e60f77c`](https://github.com/nicia-ai/typegraph/commit/e60f77cd2f2af0653193021a9dec7e6a6bf35aaf) Thanks [@pdlug](https://github.com/pdlug)! - Execute eligible durable `bulkGetOrCreateByEndpoints()` calls as one schema-fenced native atomic exchange on bundled Neon HTTP, Cloudflare D1, and libSQL roots. The eligible envelope is a schema-declared `matchIdentity` with `cardinality: "many"`, the declaration's match fields, default `ifExists: "return"`, and no temporal mutation. The program owns endpoint validation, identity arbitration, input-order restoration, and whole-call rollback. A tombstoned winner rolls the native attempt back and refuses transactionless convergence with the typed `CONSTRAINT_WRITE_FENCE_UNSUPPORTED` (`edgeMatchKeyConvergence`) error; use a transactional backend when schema-aware resurrection is required. Dynamic match fields, update mode, constrained cardinality, temporal options, caller-owned transactions, derived/custom backends, and history/revision stores retain their existing path. The existing read-only fast path remains available outside the native envelope: an all-live default-`"return"` batch completes from one set-oriented root read without opening a confirmation transaction. Inside the native envelope, the authoritative upsert program establishes the same logical `"found"` result in one exchange but can take incumbent-row locks and produce write amplification.
The libSQL transport inventory measures one `batch` submission and zero `execute` calls for a multi-item eligible call. This is a transport submission-count measurement, not a wall-clock RTT benchmark.
- [#588](https://github.com/nicia-ai/typegraph/pull/588) [`a8cd7b0`](https://github.com/nicia-ai/typegraph/commit/a8cd7b07c09e706542f0845d144f5a618b5f9488) Thanks [@pdlug](https://github.com/pdlug)! - Add `nodes..bulkReplaceById()` for complete-document replacement by ID. Eligible bundled SQLite, D1, libSQL, Neon HTTP, and PostgreSQL transaction-session roots execute missing-row creation, live-row replacement, tombstone resurrection, claims, and search projections as one read-free atomic submission; unsupported shapes retain the complete portable path.
- [#571](https://github.com/nicia-ai/typegraph/pull/571) [`b57fb90`](https://github.com/nicia-ai/typegraph/commit/b57fb90576349aa65586e8d439e251bc8661d636) Thanks [@pdlug](https://github.com/pdlug)! - Execute eligible mixed create/update `bulkUpsertById()` calls as one atomic mutation exchange after their batched authoritative read, preserving whole-set rollback on transactionless bundled roots.
- [#562](https://github.com/nicia-ai/typegraph/pull/562) [`57245c2`](https://github.com/nicia-ai/typegraph/commit/57245c22cb12da8e806b650c63c0caeedbc6928d) Thanks [@pdlug](https://github.com/pdlug)! - Execute eligible plain schema-managed `nodes.bulkInsert` and `nodes.bulkCreate` batches as one schema-fenced native atomic exchange on bundled Neon HTTP, Cloudflare D1, and libSQL roots. Claim-free nodes accept generated IDs, caller-supplied IDs, or mixed batches. A claimed node is eligible only for generated-ID batches with exactly one same-kind (`scope: "kind"`) uniqueness constraint, subject to the backend's claimed-member budget. `bulkCreate` restores rows in input order. Operational Identity, projections, history, revision, caller-supplied or mixed IDs in claimed batches, multiple or non-kind uniqueness constraints, disjointness, and other unsupported forms retain their existing transaction or fallback path.
The eligible program preserves live-duplicate rollback, tombstone resurrection validity semantics, bind-budget chunk rollback, and schema-fence errors. The libSQL transport inventory measures one `batch` submission and zero `execute` calls for both generated claim-free and generated single-claim node batches; this is a submission-count measurement, not a wall-clock RTT benchmark.
- [#582](https://github.com/nicia-ai/typegraph/pull/582) [`83b8bd0`](https://github.com/nicia-ai/typegraph/commit/83b8bd07fe0bece1a721c9c0418e6b39be1f9a59) Thanks [@pdlug](https://github.com/pdlug)! - Fold single disjointness claims and owner-side node claim cleanup into bundled atomic mutation programs. Eligible `nodes.bulkInsert()` and `nodes.bulkCreate()` calls for a `disjointWith` kind now retain one schema-fenced transport submission on Neon HTTP, Cloudflare D1, and libSQL, including legacy live-row detection and typed `DisjointError` rollback. Restricted node deletes release owned uniqueness and disjointness claims inside the same atomic program, and update-only upserts no longer fall back merely because their kind participates in disjointness.
Custom mutation executors advertise claim support explicitly through `claimSupport.families`, the per-member `claimSupport.maxInputCostPerEntry` bound, and `releasedClaimFamilies`; omitted claim families remain on the portable path and an empty family list with a zero bound is an honest opt-out.
- [#583](https://github.com/nicia-ai/typegraph/pull/583) [`0abfdc6`](https://github.com/nicia-ai/typegraph/commit/0abfdc6e5c05c5575fe5fe558ae9889450a0c29e) Thanks [@pdlug](https://github.com/pdlug)! - Fold eligible node fulltext and vector projection transitions into the same schema-fenced atomic programs as `bulkInsert()`, `bulkCreate()`, and resolved node mutations. Bundled Neon HTTP, Cloudflare D1, and libSQL roots keep projected creates and updates at one mutation exchange, while PostgreSQL session programs keep the complete row-and-sidecar unit on their pinned transaction. The atomic program proves the exact durable contribution markers in that same submission, including on a newly constructed backend. Custom atomic node executors can opt in per projection family through validated `projectionSupport` metadata; omitted support fails closed to the portable path.
- [#581](https://github.com/nicia-ai/typegraph/pull/581) [`2c8f97f`](https://github.com/nicia-ai/typegraph/commit/2c8f97ff2e108df33caa5985b23ca54d50b36007) Thanks [@pdlug](https://github.com/pdlug)! - Execute eligible mixed node and edge `bulkUpsertById()` mutation sets through atomic programs bound to the exact collection-opened, caller-supplied, or adopted PostgreSQL transaction. The operation now returns an explicit `applied | unsupported` verdict, where `unsupported` is emitted before program SQL and safely re-enters the complete portable path. Transaction-session programs use a savepoint so typed refusal diagnosis remains available without committing or poisoning the caller's surrounding transaction.
- [#578](https://github.com/nicia-ai/typegraph/pull/578) [`7c7e328`](https://github.com/nicia-ai/typegraph/commit/7c7e3281721443ee580ba113b3b2407f5a1999aa) Thanks [@pdlug](https://github.com/pdlug)! - Route eligible singleton node and edge updates and soft deletes through the same exact-root atomic mutation programs as resolved bulk writes. These operations preserve per-item hooks and fallback semantics while replacing explicit managed transactions with one authoritative read or gate plus one guarded atomic mutation exchange on bundled serverless roots.
Eligible singleton updates now converge optimistically rather than holding a
write transaction across their read and write. They retry a moved preimage up
to four times before throwing `DatabaseOperationError`; under sustained
same-row contention this can refuse an update that a transaction-capable
backend previously serialized. Transaction-scoped calls and ineligible shapes
retain the interactive transaction path.
- [#574](https://github.com/nicia-ai/typegraph/pull/574) [`e60f77c`](https://github.com/nicia-ai/typegraph/commit/e60f77cd2f2af0653193021a9dec7e6a6bf35aaf) Thanks [@pdlug](https://github.com/pdlug)! - Add independent execution capability declarations and a framework-agnostic atomic transport conformance runner for backend authors. It verifies ordered result slots, exact statement and parameter forwarding, empty-batch no-op behavior, later-statement rollback without primary or sidecar leakage, and caller-supplied exact-root provenance checks. Generic registration certifies transport mechanics but does not by itself authorize bundled node or edge mutation programs; semantic eligibility remains a separate fail-closed contract. Bundled factories now refuse the removed top-level `capabilities.transactions` override with migration guidance instead of silently retaining it as inert data; use `capabilities.execution.interactiveTransactions`.
- [#585](https://github.com/nicia-ai/typegraph/pull/585) [`2e878bc`](https://github.com/nicia-ai/typegraph/commit/2e878bcd740a6fc36856deb0745c24890293bd10) Thanks [@pdlug](https://github.com/pdlug)! - Batch committed-state diagnosis after a refused atomic node-claim program. Bundled backends now recover typed uniqueness and disjointness errors with set-oriented reads instead of sequential per-member probes, while custom scalar fallbacks use bounded windows, stop once the earliest refusal is known, and retain deterministic error selection.
- [#586](https://github.com/nicia-ai/typegraph/pull/586) [`51da9c0`](https://github.com/nicia-ai/typegraph/commit/51da9c05eeaddbdd45ee0283e48b81b8322f5929) Thanks [@pdlug](https://github.com/pdlug)! - Keep eligible large node and edge `bulkUpsertById()` sets on the atomic mutation-program path when they exceed one statement's bind budget. Resolved updates, mixed creates and updates, projection sidecars, per-chunk postimage assertions, and ordered result reads now execute as bind-sized statements inside one bounded atomic transport submission. On Cloudflare D1 this raises the previous 17-node and 6-edge batch-wide ceilings to 512 nodes and 187 edges without weakening rollback: a moved member in any chunk aborts every sibling chunk. Larger sets fail closed to the portable path rather than constructing an unbounded transport request.
- [#584](https://github.com/nicia-ai/typegraph/pull/584) [`69d5930`](https://github.com/nicia-ai/typegraph/commit/69d5930ef50a0db74c86db725e2c342b9bd478a3) Thanks [@pdlug](https://github.com/pdlug)! - Compose eligible node claim sets and fulltext/vector projections inside the same schema-fenced atomic `bulkInsert()` and `bulkCreate()` program. Multiple uniqueness constraints, hierarchy-wide uniqueness scopes, disjointness claims, caller/generated ID mixtures, and claim-plus-projection members now retain one Neon HTTP, Cloudflare D1, or libSQL transport submission and the same pinned PostgreSQL transaction program. Legacy claim-axis conflicts keep their typed errors and roll back every row and projection sidecar.
Claimed node programs now chunk row statements by actual per-member claim work inside one atomic submission instead of imposing a batch-wide claim ceiling. The custom-backend claim contract is correspondingly simplified from family-scoped batch ceilings to `claimSupport.families` plus `claimSupport.maxInputCostPerEntry`; backend authors use the exported `atomicNodeClaimInputCost()` owner rather than reproducing the compiled SQL cost model.
Claim refusal is enforced by a terminal database assertion so every row and projection chunk rolls back before failure-only committed-state reads recover the portable path's typed diagnostic.
- [#563](https://github.com/nicia-ai/typegraph/pull/563) [`589a38d`](https://github.com/nicia-ai/typegraph/commit/589a38d10d2c2d1a94002a5a12b4c19f5ce1fb32) Thanks [@pdlug](https://github.com/pdlug)! - Make graph-template registration DML-only after normal TypeGraph bootstrap, allow embedding-bearing templates on backends configured with `vector: false`, and clone durable runtime-contribution markers during instantiation so targets can be reopened through verified stores from a later serverless isolate. Vector-enabled backends continue to refuse schema-only templates that would require graph-scoped vector storage.
- [#577](https://github.com/nicia-ai/typegraph/pull/577) [`967c098`](https://github.com/nicia-ai/typegraph/commit/967c0985757bd774e420ba7b4f679744e8af43f8) Thanks [@pdlug](https://github.com/pdlug)! - Strengthen custom-backend atomic conformance so transport registration is immutable, runners bind the exact registered transport and semantic profile, observe dispatch inside property-preserving executor wrappers, verify author-created backend lineage, report inapplicable transaction checks as skipped, validate case bindings before fixture preparation, and distinguish native from pre-dispatch semantic refusals. Bundled libSQL now self-certifies every semantic mutation variant against complete committed database state: direct root families on the backend its public interactive factory returns, and resolved update or mixed families on a transactionless root where the Store can actually dispatch them.
- [#579](https://github.com/nicia-ai/typegraph/pull/579) [`8f890c7`](https://github.com/nicia-ai/typegraph/commit/8f890c7f3e55a595f5cba16c9c8bd27ca0b5fa25) Thanks [@pdlug](https://github.com/pdlug)! - Enable the registered atomic SQL and mutation-program profile on recognized interactive PostgreSQL drivers. TypeGraph now executes eligible bulk creates, bulk deletes, singleton updates/deletes, and durable edge convergence on one pinned transaction instead of falling back to the multi-step portable write plan. Neon HTTP keeps its existing native transaction-batch path, while unrecognized PostgreSQL drivers continue to fail closed to the portable implementation.
- [#565](https://github.com/nicia-ai/typegraph/pull/565) [`677ee91`](https://github.com/nicia-ai/typegraph/commit/677ee9161057b28e58ac23cc1c4515d53038e826) Thanks [@pdlug](https://github.com/pdlug)! - Existing databases must be opened once through privileged `createStoreWithSchema` / `createAdapterStoreWithSchema`, or receive the published base-schema migration, before zero-DDL verified and graph-template runtime paths are used. Those paths now fail early with `BaseSchemaMigrationError` until deployment-wide base storage is stamped at version 1.
Version deployment-wide base storage independently of per-graph schemas. A privileged open adopts the graph-template table and durable edge match-identity storage once, then stamps the marker; later warm opens perform only one marker read. This repairs externally provisioned 0.51 databases even when their graph schema is unchanged. Plain edge writes also classify legacy missing-column failures as `EDGE_MATCH_IDENTITY_STORAGE_UNAVAILABLE` instead of leaking driver errors.
Fresh SQLite and PostgreSQL installation SQL now publishes the current base-schema marker as its final statement, so zero-DDL verified stores and graph-template APIs can attach immediately after applying TypeGraph's generated migration. Concurrent adoption accepts a marker already advanced beyond the step it completed and never downgrades it.
- [#570](https://github.com/nicia-ai/typegraph/pull/570) [`5432598`](https://github.com/nicia-ai/typegraph/commit/5432598bc3ec851ac53c535f89d0db517336d09d) Thanks [@pdlug](https://github.com/pdlug)! - Reduce eligible update-only node and edge `bulkUpsertById()` calls on bundled serverless backends to one batched preimage read plus one guarded atomic set update. Consolidate fallback updates under one write plan and batch their authoritative reads when transaction semantics allow it.
- [#590](https://github.com/nicia-ai/typegraph/pull/590) [`c374a21`](https://github.com/nicia-ai/typegraph/commit/c374a210813e164ccd783add0c96e713fa37ba67) Thanks [@pdlug](https://github.com/pdlug)! - The PostgreSQL schema fence's three consumers now resolve a `WriteFencePlan` instead of emitting their locks unconditionally, bringing the last lock sites in `createPostgresBackend` into the model the rest of the codebase already reads: the per-graph schema-commit fence (`acquireSchemaWriteFence`: `pg_advisory_xact_lock` plus `SELECT ... FOR UPDATE` on the active schema row), the managed writer's `FOR SHARE` on that same row (`lockActiveSchemaVersion`), and the copy of that `FOR SHARE` the fused managed-insert programs carry inside their own statement (`schemaFenceInsertLockClause`). They are one `FOR UPDATE`/`FOR SHARE` contract, so they must never disagree about whether the engine honors row locks, and the inventory ratchet pins `resolveWriteFencePlan` at 14 call sites rather than 11.
`createPostgresBackend` already accepted and validated a `pessimisticLocks` override declaring no locks — it rewrites `serializedWriters` and throws only when that field claims a writer slot — and the eight sites consolidated previously already refuse under it. The schema fence emitted its locks regardless, so the backend accepted a stated capability and then contradicted it.
**The two cross-statement fences refuse an `unfenced` backend** rather than running without the lock, because both fence a read-then-write sequence that spans statements. `commitSchemaVersion` reads the row for the incoming version, reads the active version, and only then runs its `deactivateAll` / `activateVersion` pair; its own comment names this fence as what serializes that. A managed write HOLDS that `FOR SHARE` for the remainder of its transaction, which is what makes the version it just asserted binding through to the writes that follow — a concurrent schema commit's `FOR UPDATE` blocks on it. Skipping either lock does not leave a slower-but-correct path, it leaves a check-then-write window: the version is asserted, and then the flip the assertion was checking for lands before the write. The partial unique index on `(graph_id) WHERE is_active` still refuses a second active row, but it cannot order two commits that each read a version the other is about to replace. `engine-serialized` is exempt through `requireWriteFence` as everywhere else: the writer slot is the fence, which is why SQLite's `lockSchemaVersionForWrite` has always run as an ordinary read.
**The in-statement clause degrades**, and it is the only one that may, because its predicate is evaluated inside the INSERT that depends on it. One statement cannot race itself: with an empty clause the fence subquery still yields no row when the expected version is no longer active, so the INSERT still writes nothing. That is the posture SQLite has always run this path in. It is emitted as an empty clause rather than dropped, because dropping it (`undefined`) means "this backend has no schema-fenced insert program at all" and sends the Store down the unfused fallback path.
No first-party configuration changes behavior: `POSTGRES_CAPABILITIES` declares `{ advisoryLocks: true, tableLocks: true, serializedWriters: false }`, which resolves `lock`, and the emitted SQL is byte-for-byte what it was; SQLite already passed an empty clause. What changes is that a PostgreSQL-dialect backend declaring no usable write fence is now refused at the schema commit with `WRITE_FENCE_UNAVAILABLE`, naming the operation, instead of silently running a schema-managed store without the serialization that store's correctness assumes.
- [#575](https://github.com/nicia-ai/typegraph/pull/575) [`5e30ee4`](https://github.com/nicia-ai/typegraph/commit/5e30ee4ec05e40f5fd7419df55261d26fef29837) Thanks [@pdlug](https://github.com/pdlug)! - Open exact-root atomic Store mutation programs to custom backends through an explicit per-family semantic registration profile. Transport-only registration continues to enable no Store fast path; derived and transaction-scoped backends inherit neither proof, and malformed or out-of-order registrations fail with typed configuration errors.
- [#576](https://github.com/nicia-ai/typegraph/pull/576) [`5622ff4`](https://github.com/nicia-ai/typegraph/commit/5622ff48490208ae471fc8196b2ea6d7037d6f72) Thanks [@pdlug](https://github.com/pdlug)! - Add a framework-agnostic semantic conformance runner for custom atomic mutation programs. The runner requires ordered Store results, independently observed committed state, stale-fence no-write behavior, typed semantic-refusal rollback, exact executor dispatch, and exact-root provenance for every registered family variant.
- [#574](https://github.com/nicia-ai/typegraph/pull/574) [`e60f77c`](https://github.com/nicia-ai/typegraph/commit/e60f77cd2f2af0653193021a9dec7e6a6bf35aaf) Thanks [@pdlug](https://github.com/pdlug)! - Make `store.transaction()` and `store.transactionWithReceipt()` fail closed on backends without transaction support instead of invoking callbacks with non-atomic write semantics. Applications that intentionally relied on the old D1 or Neon HTTP fallthrough must call ordinary Store write methods directly and own partial-failure recovery explicitly.
- [#569](https://github.com/nicia-ai/typegraph/pull/569) [`ed66e47`](https://github.com/nicia-ai/typegraph/commit/ed66e4703e9a4f9cd6094e8887a0a2e0b8a8e693) Thanks [@pdlug](https://github.com/pdlug)! - Unify bundled root create and delete optimizations behind one exact-root mutation-program profile. Eligible node and edge `bulkDelete()` calls now run as one schema-fenced atomic exchange on bundled Neon HTTP, Cloudflare D1, and libSQL roots; the programs preserve edge collection identity, node restricted-delete semantics, stale-schema refusal, bind-budget chunking, and whole-call rollback, while node kinds that owe unique, disjointness, identity, projection, or capture sidecars retain the transactional path. Portable edge bulk deletion now replaces per-ID reads and writes with one batched authoritative read and set-based soft-delete chunks when the backend exposes the existing batch ports.
### Patch Changes
- [#580](https://github.com/nicia-ai/typegraph/pull/580) [`43f2e34`](https://github.com/nicia-ai/typegraph/commit/43f2e3427e5f3dc35c0407c563deb5727695f70b) Thanks [@pdlug](https://github.com/pdlug)! - Harden atomic edge-write refusal diagnosis. Bundled backends now diagnose missing endpoints with set-oriented reads instead of a sequential read per edge; custom backends without the batch-point-read capability use bounded-concurrency windows that still cover the complete input. If endpoint or cardinality state changes after the atomic rollback and no current violation can be found, TypeGraph preserves the driver cause in a typed `DatabaseOperationError` instead of leaking an unclassified transport error. Native singleton edge deletes now report the same authoritative `written` hook outcome as node deletes.
- [#587](https://github.com/nicia-ai/typegraph/pull/587) [`a3a4315`](https://github.com/nicia-ai/typegraph/commit/a3a4315f2fd87f39d3fc328085f99bf512123f5d) Thanks [@pdlug](https://github.com/pdlug)! - Eliminate the cold contribution-marker read before eligible atomic node projection writes. Bundled atomic programs now prove the exact fulltext/vector marker identities and strategy signatures with additional SQL statements inside the same database submission as the row and projection changes; missing, stale, failed, or unmaterialized evidence rolls the whole program back and is diagnosed through the existing typed contribution errors. This keeps projected writes at one mutation exchange even when a Cloudflare Worker constructs a fresh Neon HTTP, D1, or libSQL backend for each request, while retaining the server-side evidence check.
- [#568](https://github.com/nicia-ai/typegraph/pull/568) [`cda8bdd`](https://github.com/nicia-ai/typegraph/commit/cda8bdd72b0ea8568bda892f09a666e3b55a790f) Thanks [@pdlug](https://github.com/pdlug)! - Harden deployment-wide base-schema adoption under concurrent upgrades, make managed SQLite and libSQL installations repair pre-0.52 edge storage without bypassing numbered lifecycle steps, and add ratchets that bind ordered physical DDL and supported provisioning paths to executable match-identity storage rather than trusting the version marker alone.
- [#589](https://github.com/nicia-ai/typegraph/pull/589) [`09cc34f`](https://github.com/nicia-ai/typegraph/commit/09cc34f957c8114a45b3252d58fb493a61657fb5) Thanks [@pdlug](https://github.com/pdlug)! - Fence constrained writes inside caller-adopted SQLite transactions by taking the writer slot before decision-driving reads, acquire graph-merge locks in the canonical schema-first order, and compile temporal-system `orderBy` fields against their physical columns when the schema does not declare a same-named property. Declared properties retain precedence so filtering and ordering use the same field.
Bulk node and edge wrappers now have regression coverage that pins one durable revision advance per public bulk call rather than one advance per member.
## 0.52.0
### Highlights
TypeGraph 0.52 introduces the authoritative command and atomic SQL-program foundations later expanded in 0.53. Eligible schema-managed creates can fold their fences, constraint decisions, projections, and row write into one authoritative statement, while eligible bulk node and edge creates execute as one bounded atomic submission on bundled Neon HTTP, Cloudflare D1, and libSQL roots. Unsupported dimensions continue through the complete interactive-transaction path or a typed refusal; the optimization does not silently omit requested behavior.
Edges can now declare a durable graph-local `matchIdentity`. TypeGraph persists and arbitrates the canonical endpoint-and-property key, maintains it through ordinary and import writers, and uses it to converge eligible `getOrCreateByEndpoints()` calls without a read-then-insert race. This release also adds durable schema-only graph templates through `registerGraphTemplate()` and `instantiateGraphTemplate()`.
### Upgrade notes
- Custom `GraphBackend` implementations must provide the required `commands: { session, execute }` port and move managed node create, edge create, and edge convergence behavior out of the removed specialized hooks. A command that cannot honor a requested dimension must return its typed `unsupported` result before executing SQL.
- `OptionalTransactionExecution.atomic` is replaced by the discriminated `execution.mode: "interactive-transaction" | "sequential"`. Custom `lockSchemaVersionAndGraphWrite` implementations must now return the effective `GraphCommandIsolation` observed by the same pinned session that acquired the lock.
- Adding, removing, or changing an edge `matchIdentity` is a breaking schema change. The affected edge kind must be empty during activation; export and hard-delete its rows, migrate the schema, then import them so every row receives the durable key.
- Inside `store.transaction()`, issue work through the callback-scoped Store. Calls through an enclosing root Store do not join the callback's transaction and may observe or create a different execution boundary.
### Minor Changes
- [#558](https://github.com/nicia-ai/typegraph/pull/558) [`48532f8`](https://github.com/nicia-ai/typegraph/commit/48532f87de7ce5a769c5e090b5afddae325dbdc5) Thanks [@pdlug](https://github.com/pdlug)! - Eligible schema-managed generated-id node creates and `cardinality: "many"` edge creates on bundled root backends can now execute as one authoritative statement, including Neon HTTP and Cloudflare D1, where all required claims, projections, and side effects are either absent or fused into that statement. History/revision and Operational Identity work, plus other managed writes, continue to require the interactive transaction or an explicit typed refusal.
Clarify and harden execution boundaries for managed writes. The authoritative command helper now validates command/result correlation once, with typed node, edge, and convergence overloads; first-party Store consumers no longer duplicate that check, while recorded-capture retains its direct transaction-wrapper assertion. `OptionalTransactionExecution` is now a discriminated `{ mode: "interactive-transaction" | "sequential" }` value; migrate custom consumers from `execution.atomic` to `execution.mode`.
Document the distinction between interactive Store transactions, static internal adapter batches, and authoritative one-statement commands. Durable edge `matchIdentity` convergence may qualify for the one-statement root command because its canonical key has a database arbiter; claims/cardinality, undeclared dynamic `matchOn`, history/revision sidecars, and Operational Identity remain interactive-transaction contracts. The static native-batch adapter foundation remains internal; no new public Store batching API is implied.
- [#556](https://github.com/nicia-ai/typegraph/pull/556) [`d2c2557`](https://github.com/nicia-ai/typegraph/commit/d2c25576fd4fb32b4db5cfe2af33b92e309060ae) Thanks [@pdlug](https://github.com/pdlug)! - Replace the optional managed-create hook with a required semantic command port that carries its root or transaction session and, only after an advisory graph lock is acquired, a graph- and session-bound coordination token. On PostgreSQL that lock statement records the effective transaction isolation in the same token, so convergence never trusts a requested option or assumed server default and adds no isolation-probe round trip. Transparent `deriveBackend` command wrappers retain the underlying session identity; a wrapper for another connection cannot reuse its token. Under that lock, PostgreSQL endpoint get-or-create folds the match-key read, endpoint validation, and insert into one statement, returning either the created edge or the existing winner without another application-level read. Adopted transactions may return an existing match at any isolation, but the create leg refuses repeatable read.
Breaking change and migration: custom `GraphBackend` implementations must add a `commands` member with `{ session, execute(command, context) }`. Move managed-create and specialized edge-insert behavior into the `node.create`, `edge.create`, and `edge.converge-create` command cases, and return the typed `unsupported` result for dimensions the backend cannot apply. The optional combined-fence hook `lockSchemaVersionAndGraphWrite` now returns `Promise` instead of `Promise`; custom implementations must return the normalized effective isolation observed by the same pinned-session statement that acquires both locks. Built-in adapters already provide this port.
- [#554](https://github.com/nicia-ai/typegraph/pull/554) [`ddf851d`](https://github.com/nicia-ai/typegraph/commit/ddf851d99774b5d7be9a962a7994bf405188d7ee) Thanks [@pdlug](https://github.com/pdlug)! - Introduce authoritative node create commands and reduce first-party PostgreSQL write round trips by folding schema and graph fences, uniqueness and disjointness verdicts, endpoint and cardinality checks, and generated fulltext/vector projections into atomic statements. Managed Store transactions now lease one schema fence across their writes without sacrificing the fused first statement, and endpoint get-or-create decisions are confirmed from transaction-scoped evidence on caching transports.
The backend planning API now uses one required semantic command port for node and edge creates. Commands carry an explicit root or transaction session and, when an advisory graph lock was actually acquired, a graph- and port-bound coordination token; custom backends implement that same contract rather than silently falling back to a second decision path.
Breaking change and migration: add `commands: { session, execute }` to custom `GraphBackend` objects and route node/edge create plans through it. A backend that cannot honor a requested plan dimension must return its typed `unsupported` result; it must not silently ignore the dimension. See the authoritative command sessions section of the backend setup guide.
- [#554](https://github.com/nicia-ai/typegraph/pull/554) [`ddf851d`](https://github.com/nicia-ai/typegraph/commit/ddf851d99774b5d7be9a962a7994bf405188d7ee) Thanks [@pdlug](https://github.com/pdlug)! - Replace the three specialized edge-insert backend hooks with the shared semantic command port. Managed edge creates now compile endpoint validation, an optional schema fence, and an optional cardinality claim into one all-or-nothing `edge.create` command with an explicit result. Custom backends implement the required command contract and must apply or refuse every requested dimension.
Breaking change and migration: custom backends should remove their old specialized edge-insert hook wiring and implement `commands.execute` for `edge.create` (and `edge.converge-create` when convergence is supported). Return the typed `unsupported` dimension result when endpoint, schema-fence, cardinality, or convergence behavior is unavailable.
- [#561](https://github.com/nicia-ai/typegraph/pull/561) [`cdcb814`](https://github.com/nicia-ai/typegraph/commit/cdcb8144df618b4cd2294ef39969f637df831ad2) Thanks [@pdlug](https://github.com/pdlug)! - Execute eligible durable-match and cardinality-constrained `edges.bulkInsert` and `edges.bulkCreate` calls as one schema-fenced atomic exchange on bundled Neon HTTP, Cloudflare D1, and libSQL roots. The program maintains stale-claim takeover and legacy-incumbent detection, preserves typed endpoint, match-identity, and cardinality refusals, and rolls back every edge row when any constraint sidecar fails.
- [#557](https://github.com/nicia-ai/typegraph/pull/557) [`6799b65`](https://github.com/nicia-ai/typegraph/commit/6799b65cff2e667aa858b625b0e0236e1bcfac28) Thanks [@pdlug](https://github.com/pdlug)! - Add graph-local durable edge match identities. An edge registration can declare one named, canonical property-field set; TypeGraph persists its directed endpoint/property key on every edge row, maintains it across normal and trusted import writers, refuses ordinary updates to identity fields, retains it across soft deletion, and releases it on hard deletion. SQLite and PostgreSQL provision and idempotently upgrade the edge relation with a pair-null check and unique arbiter.
Schema-managed root `getOrCreateByEndpoints` calls using a declared identity now compile endpoint validation, the schema fence, conflict arbitration, and the created/found result into one database statement on bundled SQLite and PostgreSQL backends. Dynamic call-level `matchOn` remains available through the transaction-fenced compatibility path, and a supplied field list on a declared edge must exactly match the declaration. Bulk endpoint candidate reads use the set-oriented heterogeneous endpoint member instead of one read per endpoint pair on bundled backends.
Direct creates use the same durable arbiter at every cardinality. Built-in bulk creates preserve set-oriented insertion through a conflict-arbitrated batch command rather than falling back to one managed write per row.
Operation-end hooks now report `outcome: "written" | "unchanged" | "unknown"`. An authoritative get-or-create command that finds an incumbent completes as `"unchanged"` and does not fire `onError`; the same explicit outcome prevents revision/history churn for the no-write leg. Commands without an authoritative physical-write verdict report `"unknown"` instead of guessing from success.
Normal import uses the same set-oriented durable command for claimless slices and savepoint-protected batch recovery for exceptional conflicts, including on history-enabled stores. Non-transactional backends refuse ambiguous per-row retry after a failed batch rather than re-inserting a possibly committed prefix.
Adding, removing, or changing a match identity is a breaking schema change. The initial migration contract refuses activation while the affected edge kind holds rows; export and hard-delete those rows, migrate the schema, then import them so every row receives the new durable key.
- [#539](https://github.com/nicia-ai/typegraph/pull/539) [`2dcae2f`](https://github.com/nicia-ai/typegraph/commit/2dcae2f074c00a46f521f2c45f73e49f12bfee2d) Thanks [@pdlug](https://github.com/pdlug)! - Add durable schema-only graph templates with idempotent v1 instantiation.
- [#560](https://github.com/nicia-ai/typegraph/pull/560) [`da56b5f`](https://github.com/nicia-ai/typegraph/commit/da56b5f08d67fb5c867b895b8d6d988efdaecd96) Thanks [@pdlug](https://github.com/pdlug)! - Execute eligible unconstrained `edges.bulkInsert` and `edges.bulkCreate` calls as one schema-fenced native atomic exchange on bundled Neon HTTP, Cloudflare D1, and libSQL roots. The closed program validates live endpoints at the write boundary, preserves input-order results, and rolls back every bind-budget chunk when any statement fails. Cardinality claims, durable match identity, history, revision tracking, caller-owned transactions, derived backends, and unproven custom backends retain the existing transaction or fallback path.
Cast closed-program CTE values to their destination column types so PostgreSQL accepts JSON and temporal values in both node and edge native batches.
- [#559](https://github.com/nicia-ai/typegraph/pull/559) [`ff42eb4`](https://github.com/nicia-ai/typegraph/commit/ff42eb4c1eced3a66628ec6bd71ebfb273443394) Thanks [@pdlug](https://github.com/pdlug)! - Execute eligible schema-managed, generated-ID `nodes.bulkInsert` batches as one schema-fenced native atomic exchange on bundled Neon HTTP, Cloudflare D1, and libSQL roots. This first closed-program slice excludes claims, Operational Identity, projections, history, revision, and caller-supplied IDs; unsupported shapes retain their existing transaction or fallback path.
Skip the guaranteed-empty existence-priming read for generated-ID bulk node operations while preserving caller-ID existence and resurrection checks.
### Patch Changes
- [#537](https://github.com/nicia-ai/typegraph/pull/537) [`57ea6ac`](https://github.com/nicia-ai/typegraph/commit/57ea6ac74aef5a8ce22043f4cbf066d8fcac834f) Thanks [@pdlug](https://github.com/pdlug)! - Compute reciprocal-rank-fusion scores with floating-point division.
## 0.51.1
### Patch Changes
- [#526](https://github.com/nicia-ai/typegraph/pull/526) [`dbe8ee6`](https://github.com/nicia-ai/typegraph/commit/dbe8ee6a6d0f41ccb43204c0c3433e5dbe653e17) Thanks [@pdlug](https://github.com/pdlug)! - Read persisted unique constraints that omit `scope` or `collation` by applying the documented `"kind"` and `"binary"` defaults. Schema-management APIs can now inspect databases written with those omitted fields instead of reporting a malformed schema document.
## 0.51.0
### Highlights
TypeGraph 0.51 removes `drizzle-orm` from the dependency graph of its portable entrypoints. Applications using the root, backend, core, schema, indexes, graph-extension, interchange, profiler, graph-merge, or provenance entrypoints can now install and run TypeGraph without Drizzle; managed SQLite and PGlite Stores and explicit `/adapters/drizzle/...` entrypoints still use it.
Custom backend behavior is now resolved through explicit capability declarations and shared bundles instead of scattered optional-member checks. The first bundle set covers claims, statement execution, recorded revision origins, batch point reads, unique-sidecar batching, and contribution health. Recursive traversal and write-fence support are also explicit: unsupported engines can refuse recursive operations with a stable reason, while stateful features require a fence plan backed by real locks or engine-serialized writers.
### Upgrade notes
- Applications using managed SQLite or PGlite Store entrypoints, or any explicit Drizzle adapter entrypoint, must keep `drizzle-orm` installed. Managed Store factories report `MISSING_PEER_DEPENDENCY` with the installation command when it is absent.
- A custom backend hosting Operational Identity, `history: true`, or `revisionTracking: true` must declare truthful `capabilities.pessimisticLocks`. TypeGraph refuses an unfenced declaration rather than assuming safety from the dialect name.
- A custom backend without recursive SQL or an equivalent graph-native operation should declare `capabilities.recursiveTraversal: { supported: false, reason }`. Omission retains the pre-0.51 assumption that recursive traversal is supported.
- A custom backend that created the timestamp-only recorded-time preview schema must implement `recordedTableDdl(tableNames)` before running `migrateLegacyRecordedTime`; bundled SQLite and PostgreSQL backends already provide it.
### Minor Changes
- [#521](https://github.com/nicia-ai/typegraph/pull/521) [`da251ef`](https://github.com/nicia-ai/typegraph/commit/da251ef021c8b2b61717f0c8fa7c07535a28599a) Thanks [@pdlug](https://github.com/pdlug)! - Custom backend capabilities now resolve through six shared bundles instead of scattered `undefined` checks. This makes each operation family consistently choose one of three outcomes: use the declared member, take its documented fallback, or refuse with a typed error.
The pilot covers `claims`, `statementExecution`, `recordedRevisionOrigins`, `batchPointRead`, `uniqueSidecarBatch`, and `contributionHealth`. Batch point reads and supported sidecar operations degrade to their existing per-item implementations when the batch member is absent. Operations without a safe fallback keep their existing typed refusal. A backend whose capability declaration disagrees with the members reachable on its execution port now refuses with `CONSTRAINT_CLAIM_SURFACE_MISMATCH` for claims or `BUNDLE_PORT_SURFACE_MISMATCH` for the other bundles.
`CAPABILITY_BUNDLES`, the six named definitions, their verdict and binding types, and the `resolveBundle`, `bindCore`, `bindExtra`, and `bindExtraIfReachable` helpers are public for backend conformance tooling. Both bundled backends already implement every required member, so their behavior is unchanged.
This pilot covers six of twenty-one member-bearing operation families; the other fifteen continue to work through their existing paths. See [Capability bundles](https://typegraph.dev/backend-setup#capability-bundles) for the complete member and fallback matrix.
- [#520](https://github.com/nicia-ai/typegraph/pull/520) [`b6478f6`](https://github.com/nicia-ai/typegraph/commit/b6478f6e67d57767784c59691268049ce8585f75) Thanks [@pdlug](https://github.com/pdlug)! - `drizzle-orm` is now optional when using TypeGraph's portable entrypoints. Applications that import only the root, backend, core, schema, indexes, graph-extension, interchange, profiler, graph-merge, or provenance entrypoints no longer need Drizzle installed.
Applications using a managed SQLite or PGlite Store, or an explicit `/adapters/drizzle/...` entrypoint, must still install it with `npm install drizzle-orm` when their package manager does not install optional peers automatically. Managed Store factories report a typed `MISSING_PEER_DEPENDENCY` error with that command; explicit Drizzle adapters retain the runtime's raw module-resolution error. See [Managed Store Entrypoints](https://typegraph.dev/backend-setup#managed-store-entrypoints) and [installation troubleshooting](https://typegraph.dev/troubleshooting#missing-optional-drizzle-orm-peer).
- [#520](https://github.com/nicia-ai/typegraph/pull/520) [`b6478f6`](https://github.com/nicia-ai/typegraph/commit/b6478f6e67d57767784c59691268049ce8585f75) Thanks [@pdlug](https://github.com/pdlug)! - Portable entrypoints no longer reach Drizzle through recorded-time migration, claim comparison, or removal-statement builders. This completes the separation that lets consumers use the ten portable entrypoints without installing `drizzle-orm`; source and packaged-output checks now prevent those imports from returning.
Custom backends that need to migrate the timestamp-only recorded-time preview schema must implement the new optional `GraphBackend.recordedTableDdl(tableNames)` member. It returns backend-owned table and index DDL for the temporary and final recorded-relation names. A migration that reaches the legacy rewrite without this member throws `UnsupportedBackendCapabilityError` with `details.capability: "recordedTableDdl"` instead of importing Drizzle or crashing. See [Migrating Preview Recorded Time](https://typegraph.dev/schema-management#migrating-preview-recorded-time).
- [#521](https://github.com/nicia-ai/typegraph/pull/521) [`da251ef`](https://github.com/nicia-ai/typegraph/commit/da251ef021c8b2b61717f0c8fa7c07535a28599a) Thanks [@pdlug](https://github.com/pdlug)! - Custom backends can now declare whether they support recursive traversal with `capabilities.recursiveTraversal`. Absence means supported for backward compatibility; an engine without recursive SQL or a graph-native equivalent declares `{ supported: false, reason }`.
The decision is resolved through a branded `RecursiveTraversalVerdict` that only `resolveRecursiveTraversal` can construct, so a caller cannot forge one by writing `{ supported: true }` inline. Also exported: `assumeRecursiveTraversalSupported` (the one sanctioned way to obtain a verdict without a backend, used by the query compiler's no-backend entry point), `assertRecursiveTraversal`, `recursiveTraversalUnsupportedError`, and the `RecursiveTraversalCapability` type itself.
Variable-length queries, `store.subgraph()`, and the three recursion-dependent historical identity reads now refuse an unsupported declaration with `ConfigurationError` code `RECURSIVE_TRAVERSAL_UNSUPPORTED`; `details.operation` and `details.reason` identify the affected path and engine limitation. `weightedShortestPath` keeps working when temporary statements are available, reconstructing the same path through `pathLength + 1` predecessor reads instead of one recursive extraction statement.
`CompileQueryOptions` gains an optional `recursiveTraversal`, threaded by `propagateOptions` into every set-operation sub-compile so a `union()`/`intersect()`/`except()` operand carries the same verdict as its parent query.
Bundled factories refuse contradictory declarations — unsupported without a reason or supported with a dangling reason — using `CAPABILITY_DECLARATION_CONTRADICTION`. The bundled SQLite and PostgreSQL backends declare support, so their query behavior is unchanged.
Custom backend note: the factory-owned clone of `backend.capabilities` is now deep-frozen, so mutating it after construction throws. Objects supplied through factory options remain caller-owned and mutable. See [Recursive traversal capability](https://typegraph.dev/backend-setup#recursive-traversal-capability) and the [recursive query guide](https://typegraph.dev/queries/recursive#backend-support).
- [#521](https://github.com/nicia-ai/typegraph/pull/521) [`da251ef`](https://github.com/nicia-ai/typegraph/commit/da251ef021c8b2b61717f0c8fa7c07535a28599a) Thanks [@pdlug](https://github.com/pdlug)! - **Custom backend migration:** a backend not built by `createSqliteBackend` or `createPostgresBackend` must now declare `capabilities.pessimisticLocks` before hosting Operational Identity, `history: true`, or `revisionTracking: true`. Store construction refuses an undeclared backend immediately instead of risking an unfenced concurrent write. PostgreSQL backends normally declare `pessimisticLocks: { advisoryLocks: true, tableLocks: true, serializedWriters: false }`; SQLite backends normally declare `pessimisticLocks: { advisoryLocks: false, tableLocks: false, serializedWriters: true }`. Verify those values against the engine's actual guarantees rather than copying them for a different topology.
`BackendCapabilities` gains an optional `pessimisticLocks` field (`{ advisoryLocks, tableLocks, serializedWriters }`) declaring how an engine serializes concurrent writers, if at all. `resolveWriteFencePlan` is the one place that declaration turns into a `WriteFencePlan` (`lock` / `engine-serialized` / `unfenced`) every lock site now consumes instead of re-deriving from `dialect` inline, and `requireWriteFence` is the one place an operation's specific lock requirement (`"advisory-lock"` / `"table-lock"`) is checked against the resolved plan, refusing with `WRITE_FENCE_UNAVAILABLE` when it cannot be met. This consolidates eight call sites that used to spell the same dialect-keyed decision independently.
`BackendCapabilities` also gains an optional `recordedTimeOwnership` field (`"typegraph-relations"` | `"engine-native"`) naming who allocates recorded-time revisions. Absent means `"typegraph-relations"` — today's behavior for every existing backend. Declaring `"engine-native"` together with `history`/`revisionTracking` is refused at construction with `ENGINE_NATIVE_RECORDED_TIME_NOT_IMPLEMENTED` as an interim measure, independently of the write-fence plan, because the engine-native read/write path does not exist yet; a later release lifts this refusal with that path.
Both bundled backends already declare their write-fence support, so shipped configurations keep their existing behavior. See [Write fence declaration](https://typegraph.dev/backend-setup#write-fence-declaration-pessimisticlocks), [recorded-time ownership](https://typegraph.dev/backend-setup#recorded-time-ownership-recordedtimeownership), and the [stable error codes](https://typegraph.dev/errors#write-fence-declaration-codes).
### Patch Changes
- [#521](https://github.com/nicia-ai/typegraph/pull/521) [`da251ef`](https://github.com/nicia-ai/typegraph/commit/da251ef021c8b2b61717f0c8fa7c07535a28599a) Thanks [@pdlug](https://github.com/pdlug)! - **For contributors:** CI now compares the public `etc/*.api.md` snapshots with the last published tag through `test:api-surface`. The check fails when an external consumer would lose a member, see an optional member become required, or need to supply a newly required member through a contravariant API position. This adds no runtime or published API; it makes breaking surface changes visible before release. See the [release verification commands](https://github.com/nicia-ai/typegraph/blob/main/docs/RELEASE.md#pre-release-verification).
## 0.50.0
### Highlights
TypeGraph 0.50 completes database-arbitrated enforcement across the declared constraint families. Hierarchy-wide uniqueness, `disjointWith`, and edge `one`, `unique`, and `oneActive` cardinality now remain fenced under concurrent writers, including interchange import. Claim ownership is the concrete `(kind, id)` pair, so namesake nodes in different kinds cannot take over one another's reservations and lifecycle cleanup releases only the claims a node owns.
`store.verifyConstraintFences()` audits uniqueness, disjointness, and cardinality violations that predate these fences without modifying data. Internally, every Store and interchange write now runs through one typed write-plan/session pipeline; this architectural consolidation does not intentionally change public statement order, lock scope, or error behavior beyond the constraint corrections described above.
### Upgrade notes
- Databases created before 0.50 need the new `typegraph_edge_claims` relation before their first constrained edge write. Run the normal idempotent bootstrap path or apply the SQL emitted by `generatePostgresMigrationSQL()` / `generateSqliteMigrationSQL()`; a missing relation is reported as `EDGE_CLAIM_RELATION_MISSING`.
- The new fences prevent future conflicting writes but do not choose winners among violations already stored. Run `store.verifyConstraintFences()` after upgrading and resolve any reported owners or edge ids deliberately.
- Custom backends that implement constraint claims should add the `edgeClaims` table name and the edge-cardinality claim members, then declare matching `capabilities.constraintClaims`. Omission remains a supported opt-out; a declaration/member mismatch is refused.
- Custom `deleteUnique` implementations must scope release by `concreteKind` and `nodeId`, and apply the `nodeKind` predicate only when `params.nodeKind` is present. Keeping the old unconditional predicate can leak lifecycle claims permanently.
### Minor Changes
- [#490](https://github.com/nicia-ai/typegraph/pull/490) [`02c0370`](https://github.com/nicia-ai/typegraph/commit/02c03708dc9ebe81cb3f2aa4ec3b91d6352106a7) Thanks [@pdlug](https://github.com/pdlug)! - Issue a node's claims on the side of the row write their placement names, and refuse the writes a backend with no transactions cannot undo.
A uniqueness claim whose axis spans kinds beyond the writer's own is the **only** fence for that axis — the nodes primary key is `(graph_id, kind, id)`, so an `Employee`'s insert does not collide with a `Contractor`'s — and a fence issued after the write it fences is not a fence. Those claims now precede the row insert they gate, with the reservations given back if that insert does not land, so a refusal leaves zero net effect. On the import path this is what turns a violation into a refusal instead of a committed row: `importGraph` recovers per row and takes no per-graph lock, so a claim written after the row it was supposed to refuse let the row commit.
A claim whose axis **is** the writer's own kind keeps the position it has today, after the row: the uniques primary key at that axis is already the complete fence for it, moving it would buy nothing, and it would cost a refusal on backends with no transactions. Placement is decided once, per claim, from the one fact both readings turn on — does this claim's axis span kinds beyond the writer's own? — and carried as data through the entry, the claim seam, the refusal and the lock projection. Which writes take the per-graph advisory lock, and the reason each reports, are unchanged.
Two new refusals follow, both on `transactions: false` backends (Cloudflare D1, `drizzle-orm/neon-http`, `transactionMode: "none"` SQLite), and both `ConfigurationError` / `CONSTRAINT_WRITE_FENCE_UNSUPPORTED`:
- **`importGraph` / `importGraphStream` into a graph any of whose node kinds declares a unique constraint of any scope, or any of whose edge kinds is non-`many`.** Import writes claim rows like every other writer but is not covered by the write-transaction refusal, so it would have written reservations with nothing to roll them back. The refusal is computed up front, before the first chunk, so a streamed import cannot commit _k-1_ chunks and then fail. Disjointness owes no claim yet — that fence lands in a later batch — so a disjoint-only graph is not refused here.
- **A node UPDATE or RESURRECT whose kind declares only `scope: "kind"` unique constraints**, reason `nodeUniquenessClaim`. This closes an existing hole rather than paying for a new one: the transition seam already claims before its gated row write for every scope, so that path already wrote a reservation with nothing to undo it. The matching **create** is not refused — its claim stays after the row — which is the pair that makes the rule legible: same kind, same constraint, opposite verdicts, decided only by placement.
`ConstraintFenceReason` gains `nodeUniquenessClaim` for that refusal. It is never returned by the lock projection, so it cannot widen the set of writes that take the per-graph lock.
Claim statements within each placement group are issued in one canonical order — code-point on `(relation, graph, axis, constraint, key)` — and the pre-insert group is always issued first, so two writers touching the same claim rows for one row acquire them in the same order instead of deadlocking. For a node create this is observable as statement order: a kind owing only own-axis claims emits exactly what it emits today, a kind owing a cross-kind claim emits it ahead of the row insert, and a kind owing both emits two claim statements, one on each side.
- [#490](https://github.com/nicia-ai/typegraph/pull/490) [`02c0370`](https://github.com/nicia-ai/typegraph/commit/02c03708dc9ebe81cb3f2aa4ec3b91d6352106a7) Thanks [@pdlug](https://github.com/pdlug)! - Tolerate the concurrent `CREATE EXTENSION` race when materializing trigram indexes, and give every extension install one owner.
`method: "trigram"` needs `pg_trgm`, the extension is database-global, and the claim that serializes an index build is keyed per index — so two materializers building different trigram indexes both reach `CREATE EXTENSION IF NOT EXISTS pg_trgm`. That statement is not a concurrency primitive on PostgreSQL: its existence check cannot see another session's uncommitted `pg_extension` row, so the loser waited for the winner and was handed SQLSTATE 23505 instead of a notice, reporting `failed` for an extension the winner had already installed.
`GraphBackend` gains an optional `ensureExtension(name)` member — the single owner of "install a database-global extension idempotently" — which the bundled PostgreSQL backend implements with both fences: a transaction advisory lock keyed on the extension, so same-key installers never raise at all, and the concurrent-DDL retry its table and column creates already use, which clears the 23505 an installer that did NOT take that lock can still hand it (a peer on an older version, or a `capabilities.transactions: false` backend with no transaction to hang the lock on). The name is validated against the exported `DATABASE_EXTENSION_NAMES` allowlist rather than interpolated freely.
`GraphBackend.ensureTrigramExtension`, the `pg_trgm`-only member added in 0.47, is deprecated in favour of `ensureExtension` and now says exactly the same thing: the bundled PostgreSQL backend implements it by delegating, and index materialization consults it only after `ensureExtension`, so a backend written against 0.47 keeps its fence unchanged. A backend implementing neither keeps issuing the bare statement with materialization's own one-shot retry, so a third-party trigram index is still materialized. The advisory-lock key changed from `typegraph:pg-trgm-ddl` to `typegraph:extension-ddl:` when the fence generalized to any allowlisted extension — a 0.47 peer therefore takes a different key, which is exactly why the retry is retained on the locked path too.
Closes [#446](https://github.com/nicia-ai/typegraph/issues/446).
- [#490](https://github.com/nicia-ai/typegraph/pull/490) [`02c0370`](https://github.com/nicia-ai/typegraph/commit/02c03708dc9ebe81cb3f2aa4ec3b91d6352106a7) Thanks [@pdlug](https://github.com/pdlug)! - Fence `disjointWith` with a claim on the declared pair, and enforce it under `importGraph`.
Disjointness was probed and never fenced. The nodes primary key is `(graph_id, kind, id)`, so `Person "X"` and `Company "X"` are two different rows by construction — the exact collision the axiom forbids is the one the database cannot refuse — and the probe was only as good as the serialization around it. `importGraph` takes no per-graph lock and, until now, ran no disjointness probe at all: an import could commit both halves of a violating pair.
A create of a kind with a `disjointWith` partner now reserves one claim row per partner, in the same relation uniqueness claims use, at the pair's own axis with the node's id as the key. Both kinds of a pair fold to one axis through the registry's own canonical pair label, so their claims contend for one row and its primary key refuses the second writer. The claim precedes the row it gates and is given back if that row does not land, so a refusal leaves zero net effect. Because the two families arrive through one list of claim sites, every path that already maintained uniqueness reservations — create, batch create, delete, import — maintains disjointness reservations too. A **resurrect** — a soft-deleted node revived by `.create()` on its tombstoned id, `upsertById`, `upsertByIdFromRecord`, `bulkUpsertById`, or `getOrCreateByConstraint` — reserves the same claim, since reviving a tombstone re-introduces a live id under a kind exactly as a create does; the resurrect leg no longer has a window where it could revive a node under an id a disjoint partner already holds live.
`importGraph` gains the per-row disjointness probe both node paths were missing, and the per-row recovery it sits in is widened from `UniquenessError` to every declared-constraint refusal. Behavior deltas:
- **Import now enforces disjointness.** A payload containing a `Person` and a `Company` with the same id, in one batch or in sequence, refuses the second **row** and reports it in `errors` while the import continues — the accepted rows commit. Previously both committed silently. A _concurrent_ violation, taken by another writer between this row's probe and the batch's claim, still surfaces from the claim and aborts the import; that asymmetry is what import already does for uniqueness.
- **`importGraph` / `importGraphStream` is refused on a `transactions: false` backend when any node kind has a disjoint partner**, joining the unique-constraint and non-`many`-cardinality cases (`ConfigurationError` / `CONSTRAINT_WRITE_FENCE_UNSUPPORTED`). A disjoint create's claim precedes its row, and without a transaction a failure between the two would leave a reservation with no repair path.
- **A kind or unique constraint name containing `U+001E` is refused at `defineNode` / `defineGraph`.** That code point builds the axes that are not kinds, so a name carrying it could spell the reserved disjointness axis. New refusal on an input no real schema carries.
The refusal a caller sees is the family's own `DisjointError`, with the same payload whichever layer produced it: the probe reads the partner's node row, the claim reads its reservation's owner, and both name the holder's concrete kind.
- [#490](https://github.com/nicia-ai/typegraph/pull/490) [`02c0370`](https://github.com/nicia-ai/typegraph/commit/02c03708dc9ebe81cb3f2aa4ec3b91d6352106a7) Thanks [@pdlug](https://github.com/pdlug)! - Fence edge cardinality with a claim relation, and enforce it under `importGraph`.
Declared edge cardinality was probed and never fenced. `one` and `oneActive` are predicates over `(kind, from)` and `unique` is one over `(kind, from, to)`, while the edges relation's only uniqueness is its `(graph_id, id)` primary key — so two writers could both count zero sibling edges and both commit, and nothing in the schema re-decided at write time. `importGraph` made it worse by running no cardinality probe at all: a payload could commit any number of edges a `cardinality: "one"` declaration forbids.
A new relation, `typegraph_edge_claims`, keyed `(graph_id, axis, key)`, is the fence. Each constrained edge write reserves the axis its declaration spans (`:`) against the endpoint identity the declaration covers, in two statements: a decision-free create-or-lock that reports the committed holder, then — only when the holder is a different edge — a conditional takeover that succeeds exactly when that holder is no longer an edge the axis and key describe. Deciding inside one upsert would read the pre-lock snapshot of the edges relation under READ COMMITTED and accept both writers; the split is what makes the second one lose.
The claim needs no release path. A holder that is soft-deleted, hard-deleted, or (for `oneActive`) ended fails the takeover's liveness predicate and is replaced in place, so no delete, end, cascade or kind-removal path participates in the fence. The holder is identified by its kind and source endpoints as well as its id, because edge ids are caller-suppliable: a reused id would otherwise read as a live holder and block its axis forever. `EDGE_CARDINALITY_SPECS` is the one table both the TypeScript probe and the takeover's SQL read for which endpoints an axis covers, whether an edge born already ended claims at all, and what a holder must still be — so the probe and the fence cannot drift apart.
Behavior deltas:
- **Import now enforces edge cardinality** (`one` / `unique` / `oneActive`). Two edges from one source in one payload refuse the second **row** and report it in `errors` while the import continues; the accepted rows commit. Import's edge slice reuses the store's own in-batch cardinality accounting to make that per-row rather than a whole-slice abort. A _concurrent_ violation, taken by another writer between a row's probe and the slice's claim, still surfaces from the claim and aborts the import — the same asymmetry import already has for uniqueness.
- **`PostgresTableNames` / `SqliteTableNames` / `SqlTableNames` gain `edgeClaims`**, and `BackendCapabilities` gains an **optional** `constraintClaims`. Absent means `false`: a backend that predates the claim relations keeps every fence it has today and is never refused for the absence. Both bundled dialects declare `true` and implement every claim member. A backend whose declaration and surface disagree in either direction is refused with `ConfigurationError` / `CONSTRAINT_CLAIM_SURFACE_MISMATCH` rather than silently unfenced.
- **`GraphBackend` gains optional `claimEdgeCardinality`, `claimEdgeCardinalityBatch` and `purgeEdgeClaims`.** Additive; a custom backend that omits them declares `constraintClaims: false` and keeps working.
- **A database bootstrapped before this release needs the new table.** It is emitted by the existing idempotent boot path and by `generatePostgresMigrationSQL` / `generateSqliteMigrationSQL`. A store reaching a missing relation on its first constrained edge write is refused with a typed `ConfigurationError` (`EDGE_CLAIM_RELATION_MISSING`) naming the relation and the migration to run, instead of an opaque driver failure.
The refusal a caller sees is the family's own `CardinalityError`, built by the same functions the probe calls, so it is indistinguishable from the serial refusal it replaces.
- [#490](https://github.com/nicia-ai/typegraph/pull/490) [`02c0370`](https://github.com/nicia-ai/typegraph/commit/02c03708dc9ebe81cb3f2aa4ec3b91d6352106a7) Thanks [@pdlug](https://github.com/pdlug)! - Scope uniqueness-claim releases to the node that owns the claim.
A `typegraph_node_uniques` row records both the axis it fences on (`node_kind`) and the node that owns it (`concrete_kind`, `node_id`). Releases keyed on the axis alone could not tell those apart: a soft delete gave up whatever row sat at the node's own kind, and a kind removal deleted every claim whose _axis_ was the removed kind. Both readings are wrong the moment an axis and a concrete kind differ — they leak a claim that blocks its key forever, and they delete a surviving sibling's claim.
Release now has three explicitly different shapes, each with one owner: a **lifecycle** release gives up every claim the node holds for a constraint and key at whatever axis it sits on (soft delete, an update's key-change release, the resurrect diff); a **compensating** release undoes exactly the row a failed write claimed, at the axis it claimed on; and **kind reaping** removes every claim the removed kind's nodes own, through the new `buildHardDeleteUniquesByConcreteKind` builder that `materializeRemovals` and the new optional `hardDeleteUniquesByConcreteKind` backend member both compile.
`DeleteUniqueParams` gains the owner pair `concreteKind` / `nodeId`, and its `nodeKind` becomes optional — present selects the compensating shape, absent the lifecycle one. A third-party `GraphBackend` that implements `deleteUnique` must make BOTH changes: add `concrete_kind` and `node_id` to its predicate, and make the `node_kind` term _conditional_ on `params.nodeKind` being present. Doing neither does not leave the old behavior in place: the lifecycle release now passes no `nodeKind`, so a predicate that still spells `node_kind = :nodeKind` unconditionally compares against NULL, matches zero rows, and releases nothing — every soft delete and every key-change release leaks its claim, and the key stays blocked forever. TypeScript cannot catch it either, since an `undefined` bound into a SQL template is accepted silently.
- [#490](https://github.com/nicia-ai/typegraph/pull/490) [`02c0370`](https://github.com/nicia-ai/typegraph/commit/02c03708dc9ebe81cb3f2aa4ec3b91d6352106a7) Thanks [@pdlug](https://github.com/pdlug)! - Fence `scope: "kindWithSubClasses"` uniqueness on a shared claim axis, and decide claim ownership by `(concrete_kind, node_id)`.
A shared-scope unique constraint used to reserve its key under the writer's OWN kind, while the probe walked the whole hierarchy. Sibling kinds therefore reserved rows that could never collide — the `typegraph_node_uniques` primary key was structurally incapable of refusing the second writer, and only the per-graph lock stood between two concurrent creates and a duplicate ([#436](https://github.com/nicia-ai/typegraph/issues/436)). The claim is now written at the scope's **axis**: the code-point minimum of the connected `subClassOf` component, which every kind in that component computes identically, so two writers of two kinds contend for one row and the primary key is the fence. Under multiple inheritance or multiple roots this is stricter than before — the component is what the old "walk one root's descendants" reading was documented to mean — and the probe still visits every kind in scope, so rows written before the upgrade are still read and no data migration is required.
Ownership of a claim is the pair `(concrete_kind, node_id)`, not the id alone. Ids are unique only per kind, so `Employee "X"` and `Contractor "X"` are two different nodes; comparing ids let the second one match the "this row is already mine" arm of the upsert, rewrite the incumbent's `concrete_kind`, and read its own id back as proof it had won. Both upsert builders now compare the pair and return it, the accept/refuse test compares the pair, and the batch-validation cache remembers the pair — so a live claim held by a namesake under another kind is a refusal where it was previously a silent takeover. In one import batch, that refusal is now reported per row (with the earlier rows committed) instead of aborting the whole batch at the flush.
Two payload/behavior corrections come with it. `UniquenessError.kind` now names the **holder's own kind** rather than the `node_kind` the row was found under, which is the same value on a single-kind scope and the meaningful one on a shared scope. And the cross-kind lookups behind `findByConstraint` and `getOrCreateByConstraint` state their preference explicitly — axis first, then the remaining kinds in code-point order, live rows preferred over tombstoned ones — so a database carrying both a pre-upgrade and a post-upgrade row for one key resolves deterministically instead of by iteration order.
The per-graph write lock is unchanged: which writes take it, and the reason each one reports, are byte-identical, now derived from the same claim-site classification the claim itself is written from.
- [#490](https://github.com/nicia-ai/typegraph/pull/490) [`02c0370`](https://github.com/nicia-ai/typegraph/commit/02c03708dc9ebe81cb3f2aa4ec3b91d6352106a7) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.verifyConstraintFences()`, the read-only audit of constraint violations that predate the fence.
The claim relations refuse the second live claimant of an axis from the first write after upgrade onward, but they repair nothing that is already there. A database that carried two live siblings sharing a `scope: "kindWithSubClasses"` key, an id live under both kinds of a `disjointWith` pair, or two live `cardinality: "one"` edges from one source keeps carrying them: the next write that touches such an axis is refused with the ordinary typed error naming the incumbent, and until then nothing says so. This is the diagnostic that says so.
It reads the relation each constraint is **declared over**, never a claim relation's primary key. A claim key admits one row per axis by construction, and a database written before the claim tables existed holds no edge claims at all, so a claim scan would report zero violations on precisely the data the audit exists to find. Uniqueness is read from the live `uniques` rows and folded onto the axis each row's `node_kind` belongs to — which is how a pre-upgrade duplicate sitting at two different `node_kind`s is found at all — restricted to constraint names the graph declares, so disjointness claims (whose `node_kind` is a pair label, not a kind) are audited from the nodes relation instead. Contention is counted in distinct **owner pairs** (`concrete_kind`, `node_id`), not rows, so one node legitimately holding its key at a legacy axis and at the current one is not reported.
Each entry names the claim row two claimants contend for — built by the same functions the fence writes with — plus the conflicting `owners` (uniqueness, disjointness) or `edgeIds` (cardinality). It writes nothing and repairs nothing: choosing which claimant keeps an axis is a data-loss decision that belongs to the operator.
- **`GraphBackend` gains an optional `readConstraintFenceViolations`.** Additive; both bundled dialects implement it through one shared statement per family. `store.verifyConstraintFences()` refuses with `ConfigurationError` / `CONSTRAINT_FENCE_AUDIT_UNSUPPORTED` on a backend without it, rather than returning an empty report a caller would read as "clean".
- **`KindRegistry` gains `disjointKindPairs()`**, the declared pairs as kind pairs — the inverse of the internal pair label, so an enumerating caller never spells the label's form itself.
The parity matrix in `backend-setup.md` gains the three rows this mechanism owes a reader: the `constraintClaims` capability, PostgreSQL's `40001` in place of the typed error above READ COMMITTED, and the claim row's lock being held to end-of-transaction on both dialects — refusal included, so a caller that catches a constraint error and continues blocks other writers of that axis for the rest of its transaction.
### Patch Changes
- [#490](https://github.com/nicia-ai/typegraph/pull/490) [`02c0370`](https://github.com/nicia-ai/typegraph/commit/02c03708dc9ebe81cb3f2aa4ec3b91d6352106a7) Thanks [@pdlug](https://github.com/pdlug)! - Restore a merged node's `disjointWith` reservations after a resolved merge write set commits.
A resolved node write set (the graph-merge apply path, and the set update it shares its preflight with) validates the whole after-image, then clears the affected nodes' sidecar rows so its upserts can take the approved keys in any order, then rebuilds them once at the end — the rebuild is what keeps a coalesced, otherwise side-effect-free upsert from leaving its key unreserved.
The clear is keyed on the claim's OWNER, so it takes every reservation the affected nodes hold, and since 0.48 that includes their `disjointWith` claims as well as their uniqueness claims. The rebuild now goes through the same claim writer an ordinary create uses, which restores whatever the row's kind owes rather than the uniqueness slice alone; previously a merged node came out of every merge with its disjointness axis unreserved, leaving it unfenced against a disjoint namesake for the rest of the graph's life.
- [#491](https://github.com/nicia-ai/typegraph/pull/491) [`2009e6c`](https://github.com/nicia-ai/typegraph/commit/2009e6cb39bd5bc306a34acc61a81e2619e007bc) Thanks [@pdlug](https://github.com/pdlug)! - Close the write-pipeline seam: the row-work read projection is now the type
every write path actually uses. The preparation helpers, constraint and
uniqueness probes, batch validation caches and identity hooks the migrated
modules reach are re-typed off `GraphBackend | TransactionBackend` onto the
narrow handle a write frame hands out, so the counted `unfencedTarget` widening
falls from seventeen call sites to one. That one is structural rather than
migration debt — the bulk `getOrCreateByEndpoints` legs re-enter the executor
against their enclosing frame's target, and re-entry mints a session — and the
ratchet now records it as a reasoned floor with a second escape failing the
build.
Three seams are stated instead of implied along the way. `IdentityTarget` is an
explicit facet composition of what an identity statement needs (reads plus the
optional raw-statement port) rather than the whole backend union, with the
service context's `backend` named for what it is: the handle the service opens
its own write frames on. `ConstraintContext` and the uniqueness probe carry
read facets, which states in the type that no check in either module writes.
The executor's overlaid-session mint takes the READS to answer rather than a
backend to write through, so row work can no longer hand the session an
arbitrary backend. `src/store/operations/index.ts` publishes the seam
(`runWritePlan`, `WritePlan`, `WriteSession`) and still re-exports no step or
sidecar module, and the write-pipeline files come off knip's ignore list. No
public API, behavior, error type, statement or lock scope changes.
- [#491](https://github.com/nicia-ai/typegraph/pull/491) [`2009e6c`](https://github.com/nicia-ai/typegraph/commit/2009e6cb39bd5bc306a34acc61a81e2619e007bc) Thanks [@pdlug](https://github.com/pdlug)! - Route every edge write through the write pipeline: the raw `insertEdge` /
`updateEdge` / `deleteEdge` / `hardDeleteEdge` calls move into a new
`edge-write-pipeline.ts` step module and the insert dispatch, reached through
the session's seven edge methods under an edge write plan. All nine edge entry
points — including the bulk `getOrCreateByEndpoints` batch — now declare their
constraint probe as plan data instead of spelling it at the transaction call,
and an edge update states its asserted identity and validity bound as a fence
record whose keys are required, so a partially stated fence is a type error
rather than a silently unfenced write. The zero-row diagnosis
(`withUnmatchedEdgeUpdateRefusal`) moves to `edge-write-fences.ts` and stays
caller-applied, because the store's converge-or-refuse reading and interchange
import's report-and-continue reading are genuinely different recovery policies.
No public API, behavior, error type, statement, or lock scope changes.
- [#491](https://github.com/nicia-ai/typegraph/pull/491) [`2009e6c`](https://github.com/nicia-ai/typegraph/commit/2009e6cb39bd5bc306a34acc61a81e2619e007bc) Thanks [@pdlug](https://github.com/pdlug)! - Route interchange import through the write pipeline: its hand-built write
context and its own `runInWriteTransaction` call are gone, and all six write
legs — the batched and per-row node creates, the node update, the batched and
per-row edge creates, and the edge update — now run as session calls under one
write plan whose identity participation the executor acquires. Import's
hand-rolled `insertNodesBatch === undefined` / `insertEdgesBatch === undefined`
probes converge on the insert dispatch that already owns that decision, and its
edge update states the five immutable identity components and the window
guard's stored lower bound as a fence record with required keys instead of a
spread convention. The write-pipeline exemption list has no migration debt left:
every remaining entry is a step, sidecar or reasoned carve-out. No public API,
behavior, error type, statement or lock scope changes.
- [#491](https://github.com/nicia-ai/typegraph/pull/491) [`2009e6c`](https://github.com/nicia-ai/typegraph/commit/2009e6cb39bd5bc306a34acc61a81e2619e007bc) Thanks [@pdlug](https://github.com/pdlug)! - Route every node write except the set update through the write pipeline: the
eight managed entry points in `node-operations.ts` now compose a write plan and
run through the executor, and their row and sidecar writes are the session's
fused units rather than hand-paired calls. The identity-participation decision
moves from eight inline conditions to one declaration per plan, and the update
path's validity lower-bound fence becomes a required argument instead of a
spread convention. No public API, behavior, error type, or lock scope changes.
- [#491](https://github.com/nicia-ai/typegraph/pull/491) [`2009e6c`](https://github.com/nicia-ai/typegraph/commit/2009e6cb39bd5bc306a34acc61a81e2619e007bc) Thanks [@pdlug](https://github.com/pdlug)! - Add the internal write-pipeline seam: a total, disjoint classification of every
`GraphBackend` member, a typed write plan, per-kind write fences with total
applier maps, the fused write session, and the executor that is the single
sanctioned caller of the write transaction. An ESLint rule now bans direct
backend mutation calls outside the step and sidecar modules that own them, with
a declared exemption list a ratchet holds equal to the tree. No public API,
behavior, statement order, or lock scope changes.
- [#491](https://github.com/nicia-ai/typegraph/pull/491) [`2009e6c`](https://github.com/nicia-ai/typegraph/commit/2009e6cb39bd5bc306a34acc61a81e2619e007bc) Thanks [@pdlug](https://github.com/pdlug)! - Route the set-based node update through the write pipeline: `updateWhere`'s
transaction body is now `applyNodeSetUpdate`, a node write step, reached through
`session.reviseNodeSet` under a write plan. The uniqueness drop it performs
moves to the uniqueness sidecar module, and the fence the set UPDATE has no
field to carry is now refused by name instead of being absent from the call.
With this, no module outside the declared step and sidecar modules calls a
backend mutation member for a node write. No public API, behavior, error type,
or lock scope changes.
## 0.49.0
### Minor Changes
- [#488](https://github.com/nicia-ai/typegraph/pull/488) [`0db8f9c`](https://github.com/nicia-ai/typegraph/commit/0db8f9c68ae0a13c9363f393c889d56c48114f4c) Thanks [@pdlug](https://github.com/pdlug)! - Allow `importGraph` and `importGraphStream` to stage interchange data directly into an opaque ingestion branch while preserving its deferred-uniqueness boundary.
- [#486](https://github.com/nicia-ai/typegraph/pull/486) [`b82d436`](https://github.com/nicia-ai/typegraph/commit/b82d436a5c151ee6fa230591861785f82aaee302) Thanks [@pdlug](https://github.com/pdlug)! - Let constraint-aware ingestion branches stage Operational Identity same and different assertions through a conditional assertion-only facade, so duplicate unique aliases and their identity evidence can reach merge planning together.
### Patch Changes
- [#485](https://github.com/nicia-ai/typegraph/pull/485) [`184fc96`](https://github.com/nicia-ai/typegraph/commit/184fc965e20973b6196fb1349e91e4139c3e6b27) Thanks [@pdlug](https://github.com/pdlug)! - Stop publishing a private workspace ESLint config in package metadata, preventing lockfile-refreshing pnpm installs from resolving an unpublished package.
- [#489](https://github.com/nicia-ai/typegraph/pull/489) [`f4e31d9`](https://github.com/nicia-ai/typegraph/commit/f4e31d91d04faf93033b124bc351c39458ec16d3) Thanks [@pdlug](https://github.com/pdlug)! - Translate missing fulltext storage failures from Cloudflare Durable Objects SQLite into `ContributionUnavailableError` while preserving the underlying database error and transactional rollback.
## 0.48.0
### Highlights
TypeGraph 0.48 adds constraint-aware ingestion branches. Applications can stage overlapping records without prematurely enforcing node uniqueness, then resolve duplicates and validate the complete merge write set atomically against the target graph.
Writes with an implicit start and a historical end now preserve an unknown lower validity bound, so an already-ended record remains readable before its end. The new `repairInvertedValidityWindows()` utility reports or repairs older rows whose stored start is later than their end. PostgreSQL connection detection also recognizes additional single-connection configurations, preventing streaming interchange from waiting indefinitely on a connection it already holds.
### Upgrade notes
- Existing inverted validity windows are not repaired automatically. Run `repairInvertedValidityWindows()` in report mode first; apply repairs with writers stopped, preferably across both live and recorded relations, then re-baseline outstanding merge branches.
- Rows created with an implicit start and a historical `validTo` can now return `meta.validFrom: undefined`. Custom backends should use `resolveStampedValidityLowerBound` for insert and node-resurrection stamping.
- String-valued PostgreSQL single-connection settings now trigger the same interchange guards as numeric `max: 1`. Working-copy clones on those connections use a materialized in-memory export; use the `serializedResource` declaration when connection topology cannot be inferred correctly.
### Minor Changes
- [#476](https://github.com/nicia-ai/typegraph/pull/476) [`a58ba03`](https://github.com/nicia-ai/typegraph/commit/a58ba031b439f3befbcc073e016f896b178af00e) Thanks [@pdlug](https://github.com/pdlug)! - Store no validity lower bound for a write that would otherwise be born already
ended. A write that stamps a `valid_from` the caller did not state now stores
the write instant only when doing so leaves a window some coordinate can read:
if a stated `validTo` falls at or before that instant, the row is stored with no
lower bound at all — "ended at T, start unknown" — and reads back at every `asOf`
before its end instead of at none ([#407](https://github.com/nicia-ai/typegraph/issues/407)). The decision lives in the SQL builders,
so it holds for every `GraphBackend` caller, including interchange import and
trusted import, not only for the store paths.
Three consequences worth naming:
- `meta.validFrom` is `undefined` for such a row, where it used to be an instant
no query could match.
- A resurrecting `upsertById` / `bulkUpsertById` that names a lone historical
`validTo` on a tombstoned node **no longer refuses**: it reaches the same
stored shape a `create` on the same id reaches. One stated window, one outcome,
whichever entry point resets the window.
- A `validTo` in the future is unchanged — it still stamps the write instant, so
a scheduled-end row stays invisible before it existed.
Rows already stored with an inverted window are not rewritten; they stay
invisible at every coordinate until repaired.
Custom `GraphBackend` implementations should route node/edge insert stamping
and node resurrection stamping through `resolveStampedValidityLowerBound`, now
exported from `@nicia-ai/typegraph/backend`, so adapter-specific builders cannot
drift from the shared validity-window contract. Edge resurrection retains its
stored lower bound and therefore does not stamp one.
- [#477](https://github.com/nicia-ai/typegraph/pull/477) [`09318d6`](https://github.com/nicia-ai/typegraph/commit/09318d672e05f50bc3b1c61a4c9037b5f6f0c1be) Thanks [@pdlug](https://github.com/pdlug)! - Recognize every spelling of a one-connection Postgres cap that its driver
actually honors, so interchange refuses the pairs that would otherwise hang.
Serialized-connection detection previously required a numeric `max: 1`, which
missed three configurations that really do run every statement on one
connection:
1. `new Pool({ max: "1" })` and the legacy `new Pool({ poolSize: "1" })` — the
shape `max: process.env.PG_MAX` produces. pg-pool never coerces the value, so
the cap stays a string, and its own `_clients.length >= options.max` check
then coerces it: the pool really is capped at one.
2. `postgres(url + "?max=1")` — postgres-js resolves `max` from the URL and does
not coerce it either.
3. `PGMAX=1` with postgres-js — the same cap through the environment.
On those three backends, a streaming export/import pair now throws
`INTERCHANGE_SHARED_SERIALIZED_BACKEND_SNAPSHOT` where it previously hung, and
two concurrent streaming imports now throw
`INTERCHANGE_SERIALIZED_IMPORT_IN_PROGRESS` where they previously interleaved
and succeeded slowly — the lease is exclusive across all four pairings, so this
one conservative refusal is inherited verbatim from the existing numeric
`max: 1` behavior. `store.withWorkingCopy` and branch cloning also switch from a
streamed clone to a fully materialized in-memory export on those backends, a
memory-profile change on large graphs.
Deliberately still unmarked: a postgres-js client given a non-numeric string cap
other than one (`?max=5`), which opens exactly one connection today only because
postgres-js does not coerce it — marking that would be marking on an upstream
bug, and would refuse legitimate concurrent work the day the driver fixes it.
A `pg` pool given `max: "5"` genuinely opens five connections and is likewise
unmarked. Existing correctly-detected backends see no change: same marks, same
refusal codes, same messages.
Shipping in the same release, so nobody surprised by a new refusal is stuck:
`createSqliteBackend` and `createPostgresBackend` gain an optional
`serializedResource` declaration. `{ mode: "shared", resource: client }` marks a
connection TypeGraph cannot detect — the `?max=5` shape above, Bun `SQL`,
`expo-sqlite`, `op-sqlite`, `sqlite-proxy`, `pg-proxy` — and two backends naming
the same object are one serialized resource. `{ mode: "independent" }` escapes a
detection that is wrong for your topology. A `"shared"` declaration naming a
different object than the one detected is refused with a `ConfigurationError`
carrying `details.reason: "serialized-resource-conflict"` and a constructor-name
description of each side (`details.declaredKind` / `details.detectedKind`, never
the handles themselves — `details` is what `toLogString()` serializes, and a
driver handle there would log the credentials that driver stores) rather than
silently preferred. `"independent"` lifts the shared-resource refusal between
two distinct backends — one SQLite backend exporting into itself still reports
`INTERCHANGE_SAME_SQLITE_BACKEND_SNAPSHOT`, which is a fact about one handle
rather than a claim about connection topology. That surviving refusal is
SQLite-only, so on PostgreSQL a backend declared independent may export into
itself.
- [#479](https://github.com/nicia-ai/typegraph/pull/479) [`e285a71`](https://github.com/nicia-ai/typegraph/commit/e285a717a5d1df20868c3bc7a2c0f3a7914f5d40) Thanks [@pdlug](https://github.com/pdlug)! - Add constraint-aware ingestion branches that defer node uniqueness during staging and validate the resolved merge write set atomically.
- [#476](https://github.com/nicia-ai/typegraph/pull/476) [`a58ba03`](https://github.com/nicia-ai/typegraph/commit/a58ba031b439f3befbcc073e016f896b178af00e) Thanks [@pdlug](https://github.com/pdlug)! - Add `repairInvertedValidityWindows`, the explicit operator action that makes
rows an older version stored with a backwards window (`valid_from > valid_to`)
observable again. Such a row is readable at no coordinate at all; upgrading
deliberately rewrites nothing, so repairing is a decision an operator takes
rather than a side effect of a deploy.
`mode: "report"` counts and writes nothing — it reads through `execute`, a
required backend member, so detection works on every backend including a
history-capturing one and one with no statement-execution support.
`mode: "apply"` normalizes the rows it counted to no lower bound ("ended at T,
start unknown"), the shape today's write paths store, and is idempotent and
convergent. `relations` is required and takes `"live"` or `"live-and-recorded"`;
`"live-and-recorded"` is recommended, because repairing only the live axis
leaves the recorded twin inverted and re-materializes the invisible row at any
`asOfRecorded` coordinate.
The repair mints no revision, bumps no `version` and does not move `updated_at`:
it normalizes storage for rows that were never observable, so it is not a
logical write. Run it with writers stopped, and re-baseline outstanding merge
branches afterwards — `valid_from` is part of the `base@V` content fingerprint.
`apply` refuses rather than guessing on the states it cannot honor: a backend
without statement execution, a recorded-capture backend, and (on SQLite, where
bounds compare as text) a relation holding non-canonical bounds it cannot
classify.
## 0.47.0
### Highlights
TypeGraph 0.47 separates merge planning from execution with JSON-serializable `MergePlanArtifact` values. Applications can inspect the resolved write set and entity-resolution evidence, then apply the reviewed artifact against its target, schema, and revision fence without rerunning candidate generation or conflict callbacks. Existing one-call merge APIs retain their behavior.
Operational Identity assertions gain explicit half-open validity windows, temporal contradiction checks, endpoint coverage validation, and window-aware interchange and merge behavior. Node and edge update/upsert APIs gain `clearValidTo: true` to reopen ended records, while dynamic collection results can participate directly in identity operations. Streaming exports also gain an optional consumer-idle timeout that releases their snapshot transaction and connection lease.
### Upgrade notes
- Custom similarity scorers must return finite numbers; `NaN` and infinity now raise `MatchEvidenceError`. Candidate-source failures use `CandidateSourceError`, and deterministic merge constraint failures use `MergeConstraintConflictError` with the original store error as their cause.
- With `coalesceUnchangedUpserts` enabled, an unchanged endpoint edge get-or-create returns `"found"`, not `"updated"`. The coalescing path requires the endpoint convergence fence and refuses with `CONSTRAINT_WRITE_FENCE_UNSUPPORTED` on backends that cannot provide it.
- Compiled-query metadata timestamps now use canonical fixed-width UTC ISO 8601 across drivers. Consumers comparing serialized timestamps should expect the same rendering as collection reads.
- Reopening a validity window rechecks applicable constraints, including `oneActive` cardinality. Custom backends that cannot honor `clearValidTo` refuse it explicitly.
### Minor Changes
- [#475](https://github.com/nicia-ai/typegraph/pull/475) [`4abc7ba`](https://github.com/nicia-ai/typegraph/commit/4abc7ba34897f3cec1ed876321f0eaba0e2829c9) Thanks [@pdlug](https://github.com/pdlug)! - Add an opt-in `idleTimeoutMs` safety bound to `exportGraphStream`. The timeout
measures how long a delivered chunk remains unacknowledged, then rolls back the
snapshot transaction, releases the serialized stream lease, and reports the new
typed `ExportStreamIdleTimeoutError`. Time spent waiting for the backend does
not count as consumer idleness, and existing `AbortSignal` and cooperative
`break`/`return` cancellation behavior is unchanged.
- [#468](https://github.com/nicia-ai/typegraph/pull/468) [`c53c006`](https://github.com/nicia-ai/typegraph/commit/c53c0067f8a349e307f4f6d05e316172bcd726c2) Thanks [@pdlug](https://github.com/pdlug)! - Add public snapshot and incremental merge planning APIs that return stable,
JSON-serializable `MergePlanArtifact` values. Plans bind the reviewed write set
to the target graph, active schema, durable revision origin and revision, carry a
content digest, and can be applied exactly once with `applyMergePlan`. Applying a
plan validates the artifact and checks its fence atomically without re-running
candidate generation, scoring, embeddings, canonical selection, or conflict
callbacks. Existing `merge` and `mergeIncremental` entry points retain their
one-call compatibility behavior while sharing the same resolution, validation,
and mechanical write owners.
Explain entity resolution with deterministic decisive edges, complete built-in
candidate-source attribution, and scored strategy/score/threshold evidence while
keeping definitional matches distinct from similarity scores. Add opt-in,
deterministically bounded accepted/rejected candidate diagnostics. Default
evidence excludes raw compared values.
Custom similarity scorers that return `NaN` or infinity now fail with
`MatchEvidenceError`; non-finite values cannot be represented faithfully in a
serialized evidence artifact. Candidate-source failures now use the more
specific `CandidateSourceError`; legacy `details.source` remains available
alongside `details.sourceId` for base-source configuration failures.
- [#473](https://github.com/nicia-ai/typegraph/pull/473) [`44d1486`](https://github.com/nicia-ai/typegraph/commit/44d1486938661bbde6d171bd0c99a0b269b597c3) Thanks [@pdlug](https://github.com/pdlug)! - Allow nodes returned by runtime string-keyed collections to participate directly
in Operational Identity reads, assertions, bulk operations, and pair
retractions. Identity result types now honestly include runtime-evolved members
alongside compile-time graph references.
- [#475](https://github.com/nicia-ai/typegraph/pull/475) [`4abc7ba`](https://github.com/nicia-ai/typegraph/commit/4abc7ba34897f3cec1ed876321f0eaba0e2829c9) Thanks [@pdlug](https://github.com/pdlug)! - Add explicit half-open validity windows to scalar and bulk Operational Identity assertions, with bounded temporal contradiction checks, endpoint coverage validation, archival interchange support, node-window integrity guards, scalable branch-merge validation, and report-visible window reconciliation.
- [#469](https://github.com/nicia-ai/typegraph/pull/469) [`8e50bdb`](https://github.com/nicia-ai/typegraph/commit/8e50bdbda614942b2848c6355b9cfb11c6468d2f) Thanks [@pdlug](https://github.com/pdlug)! - Add `clearValidTo: true` across node and edge update/upsert APIs so applications can reopen an ended valid-time window without changing entity identity. Built-in SQLite and PostgreSQL backends apply the clear, unchanged replays coalesce, `oneActive` relationships are rechecked when reopening, unsupported custom backends refuse explicitly, and graph merge carries branch-authored reopenings while rejecting delete-and-resurrect window artifacts.
- [#475](https://github.com/nicia-ai/typegraph/pull/475) [`4abc7ba`](https://github.com/nicia-ai/typegraph/commit/4abc7ba34897f3cec1ed876321f0eaba0e2829c9) Thanks [@pdlug](https://github.com/pdlug)! - Return `MergeConstraintConflictError` when a resolved graph merge would violate a deterministic store constraint, preserving the typed store error as its cause and exposing actionable constraint details.
### Patch Changes
- [#470](https://github.com/nicia-ai/typegraph/pull/470) [`0b0022c`](https://github.com/nicia-ai/typegraph/commit/0b0022cf2b3c13483b468536404be4f35a8dfe39) Thanks [@pdlug](https://github.com/pdlug)! - Coalesce unchanged endpoint edge get-or-create updates when `coalesceUnchangedUpserts` is enabled. A coalesced replay now returns action `"found"`; `"updated"` means an update actually ran.
The coalescing check needs the endpoint match-key convergence fence. On a backend without top-level transactions, such as Cloudflare D1 or `neon-http`, an otherwise unchanged endpoint replay now refuses with `CONSTRAINT_WRITE_FENCE_UNSUPPORTED` instead of running unfenced. The bulk endpoint form and the create leg already required this fence.
This option does not coalesce node `getOrCreateByConstraint` updates; use `upsertById` for replay projectors that need unchanged node writes to avoid history churn.
- [#475](https://github.com/nicia-ai/typegraph/pull/475) [`4abc7ba`](https://github.com/nicia-ai/typegraph/commit/4abc7ba34897f3cec1ed876321f0eaba0e2829c9) Thanks [@pdlug](https://github.com/pdlug)! - Canonicalize node and edge metadata timestamps returned by compiled queries.
All supported database drivers now expose the same fixed-width UTC ISO 8601
rendering through compiled-query projections and store collection reads.
- [#475](https://github.com/nicia-ai/typegraph/pull/475) [`4abc7ba`](https://github.com/nicia-ai/typegraph/commit/4abc7ba34897f3cec1ed876321f0eaba0e2829c9) Thanks [@pdlug](https://github.com/pdlug)! - Serialize PostgreSQL `pg_trgm` extension installation across concurrent trigram index materializers.
- [#475](https://github.com/nicia-ai/typegraph/pull/475) [`4abc7ba`](https://github.com/nicia-ai/typegraph/commit/4abc7ba34897f3cec1ed876321f0eaba0e2829c9) Thanks [@pdlug](https://github.com/pdlug)! - Preserve PostgreSQL vector index build failures throughout serial-fallback preparation and durable `parallel_workers` cleanup, report the exact manual repair when cleanup fails, and reset built-in pgvector tables before every materialization attempt so recovery survives backend recreation without mutating custom strategy storage.
## 0.46.2
### Patch Changes
- [#459](https://github.com/nicia-ai/typegraph/pull/459) [`0e2afe2`](https://github.com/nicia-ai/typegraph/commit/0e2afe2d76381b5bc8309485d3a7e84c7bda092e) Thanks [@pdlug](https://github.com/pdlug)! - Extend `onImmutableLowerBound: "preserve"` to endpoint-matched edge writes.
`getOrCreateByEndpoints` accepts the policy in its options, and
`bulkGetOrCreateByEndpoints` accepts it per item alongside `validFrom` and
`validTo`. The policy applies a stated `validFrom` on create or resurrection,
while a live `ifExists: "update"` preserves the stored lower bound and still
applies properties and `validTo`. Strict refusal remains the default.
- [#462](https://github.com/nicia-ai/typegraph/pull/462) [`44f60cf`](https://github.com/nicia-ai/typegraph/commit/44f60cfbfaa6290323b02fb33fad2ca3541b1127) Thanks [@pdlug](https://github.com/pdlug)! - Surface lost fulltext contribution storage on gated operations as a typed `ContributionUnavailableError` with `state: "physical-storage-missing"` and rebuild guidance. Healthy operations retain the cached marker fast path; the error path translates only a missing-relation failure whose same driver error names the declared fulltext table.
## 0.46.1
### Patch Changes
- [#456](https://github.com/nicia-ai/typegraph/pull/456) [`a091902`](https://github.com/nicia-ai/typegraph/commit/a091902264fcbcd8336179c893a1e0a16eab528c) Thanks [@pdlug](https://github.com/pdlug)! - Add an explicit event-materializer policy for node upserts. Passing
`onImmutableLowerBound: "preserve"` applies `validFrom` when the upsert creates
or resurrects a row, but preserves a live row's stored lower bound while still
applying props and `validTo`. The strict `IMMUTABLE_VALIDITY_LOWER_BOUND`
refusal remains the default. The policy is available on `upsertById`,
`upsertByIdFromRecord`, and each `bulkUpsertById` item, including unchanged
coalescing replays.
Widen the optional `better-sqlite3` peer range through 13.x and exercise 13.0.3
in this repository. Correct the event-log projector guidance to update existing
endpoint-matched edges, document historical replay window requirements, and
clarify that `MergeReport.validityEnds` only reports inherited-row claims.
## 0.46.0
### Highlights
TypeGraph 0.46 introduces Operational Identity: an opt-in profile for asserting that graph nodes represent the same or different entities. Typed Store, transaction, and temporal-view APIs expose identity membership, representatives, assertions, retractions, and history. Applications can choose same-ID folding across kinds or assertion-only identity, and use identity-expanded traversal without physically merging the underlying nodes.
Identity truth travels through interchange and graph merge, with transactional contradiction checks and derived closure storage on SQLite and PostgreSQL. Merge handling also preserves inherited validity-window changes, retains parallel edges unless repointing causes a collision, and gives committed rows precedence when a collapse selects a survivor.
This release completes the contribution maintenance sequence with read-only `probeContributions()`, non-destructive `repairContributions()`, and explicit `rebuildContribution()`. Validity-window validation is shared across write paths, and streaming exports support cancellation while holding a consistent snapshot.
### Upgrade notes
- Audit graph definitions and persisted extension schemas before upgrading. Duplicate ontology relations, hierarchical self-loops, disjointness contradictions, incompatible inverse endpoints, multiple inverse partners, and unresolved extension edge names now fail validation on both construction and reload.
- Operational Identity requires a supported transactional backend. Use ordinary interchange import for identity-bearing data; trusted import refuses identity-enabled targets and identity-bearing streams. Exports now use format `2.0`, while imports still accept `1.0`.
- Custom backend table-name resolutions and `SqlSchema` subclasses must provide the identity assertion, recorded identity assertion, and closure relations. Restore missing assertion ledgers from backup; a missing derived closure can be recreated and rebuilt before traffic resumes.
- `StoreView` and `RecordedStoreView` retain construction and `instanceof` support but can no longer be subclassed. Exhaustive merge-report and import-error consumers must handle the new `"identity"` entity kind, and tooling that assumes `FORMAT_VERSION` is the literal `"1.0"` must be updated.
- Writes now refuse inverted validity windows and live-row updates that state a different immutable `validFrom`. Trusted imports require canonical UTC validity timestamps. Node creation or upsert on a soft-deleted same-kind ID now resurrects the row instead of exposing a storage constraint error.
- Changing `identity.sameIdAcrossKinds` requires an explicit schema migration, which rebuilds closure with the schema commit. The type-level `sameAs` and `differentFrom` factories are deprecated in favor of Operational Identity.
### Minor Changes
- [#427](https://github.com/nicia-ai/typegraph/pull/427) [`facef56`](https://github.com/nicia-ai/typegraph/commit/facef560d14c38607d6414818c65e43dc65a88d2) Thanks [@pdlug](https://github.com/pdlug)! - Add `signal` to `exportGraphStream`, and keep a streaming import's statistics refresh inside its connection lease.
An export stream holds one `repeatable_read` / `read_only` transaction for its whole life, and on a single-connection backend it holds that connection's exclusive interchange-stream lease with it. Every cooperative exit already settled both, because each runs the generator's `finally`: `break` or `throw` out of a `for await`, and an explicit `iterator.return()`. A consumer that pulls `next()` and then simply DROPS the iterator — the `Promise.race([iterator.next(), timeout])` pattern — has no cooperative exit, because async-generator `finally` blocks do not run on garbage collection. That leaked the snapshot transaction for the life of the process, and with it the lease, so every later export and every later import on that connection was refused for a stream nobody was reading.
`ExportOptionsSchema` now accepts `signal?: AbortSignal`, so both `exportGraph` and `exportGraphStream` take it, on both capability arms — a non-transactional export holds no snapshot and no lease, but it still owes its consumer an answer rather than a silent stall, and the same cancellation path gives it one. On a transactional backend, aborting rolls the snapshot transaction back and releases the lease whether or not anyone is waiting on `next()`; the pull that is in flight when the abort lands — and a pull from a consumer that walked away and came back — rejects with the new `ExportStreamCancelledError` (`code: "INTERCHANGE_EXPORT_STREAM_ABORTED"`), carrying the signal's own `reason` as `cause`. A signal that is already aborted refuses the export before any transaction is opened or any lease claimed. The listener is subscribed before anything is claimed or opened and `signal.aborted` is re-checked immediately after subscribing, so an abort at any instant is either seen by that re-check or delivered to the listener — including one raised synchronously by a driver inside `backend.transaction(...)`, which an `AbortSignal` never replays to a listener that arrives later. Everything else is unchanged: an export without a signal behaves exactly as before, and a cooperative exit still reports a clean end rather than a cancellation.
There is deliberately no garbage-collection fallback. A `FinalizationRegistry` on the iterable cannot work here — not merely unreliably, but never: the producer is interruptible only where it is parked waiting for the consumer, so any cleanup state able to settle an abandoned stream must reach the stream's internal channel, and a registry holds its held value strongly, so holding anything that reaches that channel keeps the abandoned stream permanently reachable and the entry can never fire. The signal is the mechanism, and it is a contract rather than a hint.
Separately, `importGraphStream` now holds its target connection's stream lease across the trailing planner-statistics refresh instead of releasing it when the chunk loop ends. That `ANALYZE` is a write like the chunks were, and running it outside the lease left it to be stranded by an export snapshot opening in that window — swallowed as a warning, because the refresh is best-effort. `importGraph` never had the hole (`withImportStreamLease` spans its whole call), so this also removes a divergence between the two import surfaces. The lease is still released on every exit, including the error paths.
Also fixes a silent cross-kind edge overwrite in import. Edge ids are unique per graph but the import's existence probe (`getEdge` / `getEdges`) is keyed on `(graph_id, id)` with no kind comparison, so a document edge of kind A whose id was already held by a kind-B row matched that row: `onConflict: "update"` wrote A's properties onto the kind-B row with nothing in `result.errors`, and `onConflict: "skip"` counted the document's edge as already present when no edge of its kind existed. Both are now reported as a per-row `ImportError` prefixed `INTERCHANGE_EDGE_KIND_CONFLICT`, naming the stored kind and the stated one, with the stored row left untouched — the check runs before the conflict strategy, so all three strategies answer alike. `backend.updateEdge` is additionally called with `kind`, which `UpdateEdgeParams` documents as MUST-apply, so the predicate lives in the UPDATE's own `WHERE` and the check cannot be raced by a concurrent hard-delete-and-recreate; a write that consequently matches no row is reported as the same per-row error rather than aborting the import. Nodes were never affected — their probe is kind-scoped.
The export's snapshot guarantee is now stated as the capability-scoped fact it is, in the API docs, the option docs, the error class, and the abort message: a backend reporting `capabilities.transactions` reads the whole export inside one repeatable-read transaction, while one without (SQLite `transactionMode: "none"`, session-less HTTP Postgres drivers) paginates statement by statement and can show a mid-stream write in later pages. `ExportStreamCancelledError`'s message now says which of the two it is describing, so a cancelled non-transactional export no longer claims to have rolled back a snapshot it never opened.
- [#417](https://github.com/nicia-ai/typegraph/pull/417) [`9d3014c`](https://github.com/nicia-ai/typegraph/commit/9d3014c05f2936fcdadb3fa50950445a0d8e2652) Thanks [@pdlug](https://github.com/pdlug)! - graph-merge: judge the edge fold's property union against base, and report a
target-precedence window discard
Two adjacent gaps in the edge repoint/window path. One is a bug fix, the other
adds an optional field to a report type, so this ships at the higher `minor`
bump and covers both.
The repoint fold's property union had no base to compare against, unlike the
node path's three-way merge, so a staged copy of an INHERITED edge contributed
its whole fork property bag as first-class `(branch, value)` claims — including
the values it never touched. Under any rank-based `onPropertyConflict` an
untouched base value could therefore outvote a value a branch actually authored,
decided by whichever branch label happened to ride on the untouched copy. The
window-only carrier made it observable: an inherited row whose only change is
its end-of-validity is staged solely to give that ending somewhere to ride, its
properties ARE the base's, and its branch is merely whichever sorted first in
staging. The union now filters every contributor to the properties it CHANGED
from its own base — a branch-created edge has no base, so everything it carries
stays a full claim — which means a carrier contributes no claim and raises no
conflict at any rank. Genuine disagreements are unaffected: two members that
changed one property differently still conflict, over their real values alone.
Filtering claims does not erase content: the folded row commits the same property
set as before, and a key only a non-survivor carries keeps the value held by the
member with the minimum edge ID — the row, never the branch label riding on it,
since for these keys no branch claimed anything and an arbitrary label deciding
the committed value is the very thing being fixed.
`MergeReport.validityEnds` now also reports the window claims that target
precedence discards. When the incremental target had already moved an inherited
row's end, the reconciler took the row out of the resolution and the branch
claims vanished from the report entirely — less visible than a claim that merely
lost the least-claim rule, which stays named in `claimedBy`. Such a row now gets
a resolution naming the target's own committed instant, its discarded claimants,
and the new optional `ValidityEndResolution.precedence` field set to the
exported `VALIDITY_END_TARGET_PRECEDENCE`. The field is absent on every entry
the merge itself decided, so existing consumers read what they always read; no
write is staged and no provenance credit is minted for such a row, and a row no
branch claimed still produces no entry at all.
- [#364](https://github.com/nicia-ai/typegraph/pull/364) [`fb29816`](https://github.com/nicia-ai/typegraph/commit/fb2981664c77d704b7f78933b8f887222c796091) Thanks [@pdlug](https://github.com/pdlug)! - Add optional `topK` to `pageRank()` and `personalizedPageRank()`, and optional
`minComponentSize` to `weaklyConnectedComponents()`. Both bound only result
extraction: the limit and the inclusive component-size filter are applied in
extraction SQL after the existing deterministic ordering, so bounded rows never
reach the driver. Default results and ordering are unchanged, and the graph
computation itself still runs over the whole visible induced subgraph.
- [#415](https://github.com/nicia-ai/typegraph/pull/415) [`b68e643`](https://github.com/nicia-ai/typegraph/commit/b68e6437a337cdcd2e3c166754deb87008e25152) Thanks [@pdlug](https://github.com/pdlug)! - Refuse non-canonical validity-window timestamps in trusted import.
`trustedImportGraph` / `trustedImportGraphStream` accept a pre-typed stream and
never re-parse it, so a `validFrom` / `validTo` that TypeScript types as `string`
but is not canonical fixed-width UTC ISO 8601 used to flow straight to SQL. Every
temporal filter compares those values AS TEXT against an `asOf` coordinate, so a
stored `"2021-01-01"`, `"...T00:00:00Z"`, `"...:00.1Z"` or `"...+01:00"` mis-sorts
and silently includes or excludes the wrong rows — and it mis-decided the
negative-width window check that the same path performs on the way in.
Both window fields of every streamed node and edge are now format-checked with
the same `isCanonicalIsoDate` decision the untrusted import schema and the store's
own writes make. A violation refuses the WHOLE stream with a `TrustedImportError`
carrying the existing reason `invalid_stream`, naming the offending field, row and
value; the session's transaction rolls back, so chunks already streamed are not
left behind. This is a behavior change: a stream that previously imported and
stored an unsortable timestamp now fails loudly. Convert such values with
`new Date(value).toISOString()`. The check is format-only — trusted import still
skips property, reference and conflict validation — and it leaves an absent field
and an explicitly `null` (confirmed open-left) `validFrom` untouched.
Also documents a pre-existing bulk-API limitation, with no behavior change:
`bulkUpsertById` groups every create ahead of every update, so one batch cannot
hand a constrained value from one row to another (releasing a `unique` value or a
`oneActive` edge slot and claiming it in the same batch throws `UniquenessError` /
`CardinalityError`, where the equivalent sequential upserts succeed). The
workaround is two batches, or sequential upserts.
- [#374](https://github.com/nicia-ai/typegraph/pull/374) [`fadf932`](https://github.com/nicia-ai/typegraph/commit/fadf93297df40bd619a1ca45b165edc04ef6ebfe) Thanks [@pdlug](https://github.com/pdlug)! - graph-merge: stage cascade retractions with their cause instead of inferring intent
A node soft-delete ends every open identity assertion touching the node, so a
branch that deletes a node stages retractions it never asked for. The merge
previously separated those cascade endings from a branch's own retraction with
a conservative branch-level heuristic, which deliberately over-dropped the
same-branch case: a branch that retracted an assertion and LATER deleted one of
its endpoints looked exactly like a pure cascade, so its retraction was dropped
whenever the deletion was overruled — silently keeping truth the branch had
explicitly ended.
The soft-delete cascade now ends assertions at the deleted node's own
`deleted_at`, which makes the cause derivable: the state-diff compares each
retracted assertion's end instant to the deletion instants of its endpoints and
stages the retraction as either a cascade naming the deleted node or the
branch's own act. The merge planner drops a retraction only when EVERY branch
staged it as the cascade of a deletion that delete/modify resolution then
overruled, so an explicit retraction survives even when it comes from the
deleting branch. Two cases stay conservative because nothing distinguishes them
at the stored resolution — a hard delete (which removes the assertion rows) and
a retraction issued in the same millisecond as the delete that followed it.
- [#387](https://github.com/nicia-ai/typegraph/pull/387) [`71361d7`](https://github.com/nicia-ai/typegraph/commit/71361d70b1fda83ad8539228c8e133abd0ce57f9) Thanks [@pdlug](https://github.com/pdlug)! - Complete the contribution health lifecycle with a read-only readiness probe and
an explicit destructive rebuild, so the three maintenance operations form one
escalation ladder: `probeContributions()` (writes nothing) →
`repairContributions()` (non-destructive, already shipped) →
`rebuildContribution()` (drops and recreates storage).
`store.probeContributions()` answers "is search coherent with the graph right
now" without mutating anything — safe on a read path, on a replica, and under a
least-privilege role. It returns one `ready` / `degraded` entry per search
projection plus the durable `graphRevision` the assessment was taken at on a
revision-tracked Store. It shares the detection logic of
`verifyContributions()` rather than reimplementing it, so a health check can
never disagree with the gate the hot path actually consults. A projection with
no declared contributions is omitted rather than reported `ready`, and a backend
that provisions contributions but cannot probe its catalog refuses instead of
answering — "assessed and healthy" and "never looked" never share a return
value.
`store.rebuildContribution("fulltext")` is the repair that was missing for a
`stale` contribution, whose table exists at a shape the current `createDdl` no
longer produces: the ensure path's `CREATE ... IF NOT EXISTS` no-ops against it,
so re-stamping the marker would leave it blessing storage of the wrong shape.
The rebuild drops the storage, recreates it, reconstructs the content from the
node rows, and stamps the marker inside one transaction under the schema-write
fence, so an interrupted rebuild rolls back rather than leaving storage attested
but empty. It is reachable only by name — never from `repairContributions()`,
which continues to report these findings as `requires-rebuild`.
Vector contributions are not rebuildable, and the call refuses with
`ContributionRebuildUnsupportedError` rather than dropping them: TypeGraph
stores the vectors callers supply and never the inputs that produced them, so
the embeddings exist only in the storage a rebuild would destroy.
`reembedVectorField(kind, fieldPath, { embed })` remains the sanctioned
destructive path, because it takes the callback that can regenerate them. The
same typed error covers a fulltext strategy that declares no `dropDdl` and a
backend with no transactional schema fence; all three refuse before anything is
dropped, and all three are declared ahead of time on the new
`backend.capabilities.contributions` capability.
Fixes the drift guard so the ladder can actually be climbed: when the guard
refused a shape change it recorded the failed attempt at the _new_ signature,
overwriting the only evidence of the shape the table really had. The verdict
then read as `missing-marker` rather than `stale`, so `repairContributions()`
reported it repaired — re-stamping the marker over the unchanged old-shape table
— and the next boot skipped the guard entirely. The guard now preserves the
recorded signature, so a `stale` contribution stays `stale` across restarts,
`repairContributions()` keeps reporting `requires-rebuild`, and the refusal
persists until `rebuildContribution("fulltext")` fixes the shape. Reach that
call from a `createStore()` / `createVerifiedStore()` Store: the managed
factory's boot step is what the guard refuses.
Also adds optional `dropDdl` to `TableContribution` — declared by both bundled
fulltext strategies — which is what opts a strategy into the rebuild.
- [#376](https://github.com/nicia-ai/typegraph/pull/376) [`8c3a8e6`](https://github.com/nicia-ai/typegraph/commit/8c3a8e6af5ec2813a26d0aa13bf58da2c50fbaa3) Thanks [@pdlug](https://github.com/pdlug)! - Add a database-level contradiction backstop for Operational Identity.
A `different` assertion and a `same` assertion that would place both of its
endpoints in one identity class are a contradiction, and until now only
application code stood between such a write and a committed graph: the
plan-time simulation and the identity applier's validation both decide by
reading state and comparing, so a bug in either commits the contradiction
silently.
Identity now also maintains a derived **separation relation** — one row per
pair of identity classes a current `different` assertion holds apart, keyed by
the two class keys under a `CHECK (class_key_low < class_key_high)`
constraint. Every transaction that fuses two classes relabels the affected
separation rows in the same statement batch, so fusing two separated classes
relabels both sides of their shared row to one key and the database aborts the
transaction. A write that reached the ledger through a path that skipped
identity validation can no longer commit a contradictory graph; it fails with
the new typed `IdentitySeparationViolationError`.
The relation is derived and requires no application changes: it is maintained
wherever the identity closure is (assert, retract, fold, delete, merge,
import, rebuild), `rebuildIdentityClosure(store)` recomputes it from the
assertion ledger, and store-open identity validation checks it against that
recomputation.
Upgrading an existing identity-enabled database needs no manual step. A store
opened with `createStoreWithSchema` / `createAdapterStoreWithSchema` creates
the new `typegraph_identity_separation` relation through the same idempotent
identity DDL path as the other identity relations and recomputes it from the
ledger once, before anything reads it. A missing assertion ledger or closure
relation is still refused as data loss.
Custom backend authors: the resolved table-name types (`ResolvedSqlTableNames`,
`SqliteTableNames`, `PostgresTableNames`) and the `ensureIdentityTables`
parameter each gained a required `identitySeparation` entry, so an
implementation that builds one of those objects needs the new name added. Code
that only reads `backend.tableNames`, or that passes a partial name override to
`createSqliteTables` / `createPostgresTables` / `createSqlSchema`, is
unaffected — an omitted name still resolves to the default
`typegraph_identity_separation`.
- [#397](https://github.com/nicia-ai/typegraph/pull/397) [`d2b935e`](https://github.com/nicia-ai/typegraph/commit/d2b935ed071fc2eb8e4310cfb1bdeecca64072b0) Thanks [@pdlug](https://github.com/pdlug)! - Graph Merge now keeps the **inherited** edge when a repoint-induced collapse folds a
committed row together with a branch-created one, instead of keeping the
lexicographically-minimal edge id.
Previously the survivor of such a collapse was whichever edge id sorted lowest. A
collapse rewrites the row it keeps and ends none of the rows folded into it, so when a
branch-created id sorted below the committed one, the merge wrote the branch's row as a
new edge and left the committed edge live beside it at its pre-merge properties — two
live rows for one folded relationship, the edit staged for the committed row never
written, and `merged.edges` counting one of them. Which of the two you got depended on
an id sort, so it was not behavior a caller could depend on.
The surviving edge id reported in `PropertyConflict.entityId`, window resolutions and
provenance is consequently the inherited row's id whenever the collapse involved one.
That is the id of the row that actually persists, and it no longer moves with
branch-created id lexicographics. Collapses among branch-created edges alone are
unchanged, as is the property/window reconciliation applied to the survivor.
- [#427](https://github.com/nicia-ai/typegraph/pull/427) [`facef56`](https://github.com/nicia-ai/typegraph/commit/facef560d14c38607d6414818c65e43dc65a88d2) Thanks [@pdlug](https://github.com/pdlug)! - `store.rebuildContribution("fulltext")` is now scoped to the graph it is called on. The fulltext projection is one physical table holding every graph's rows keyed by `graph_id`, while the rebuild runs under the per-graph schema fence — so the old unconditional `DROP TABLE` destroyed every other graph's search index on the same database, with no concurrency required: a neighbouring graph's `fulltext` marker survived the drop (markers are keyed by `graph_id`), so `probeContributions()` and `verifyContributions()` went on reporting it `ready` while every search it served returned nothing. The rebuild now removes only the calling graph's rows, through the same `DELETE ... WHERE graph_id` statement `clear()` uses — one exported builder both call — and escalates to dropping and recreating the shared table only when that table holds no other graph's rows. That drop remains the one repair for storage provisioned at a shape the current DDL no longer produces, and the lock scope now matches the decision it authorizes. Two locks, protecting different resources: a constant-keyed advisory lock (`typegraph:contribution-ddl`) serializes the contribution's DDL across graphs — it survives the drop and exists even when the table does not, which is what a relation lock cannot do — and, on the path that may drop, `LOCK TABLE ... IN ACCESS EXCLUSIVE MODE` excludes ordinary writers, which take no advisory lock at all and could otherwise commit a row between the probe and the `DROP TABLE` that the probe had already decided was safe. The verdict is re-established under that lock before any drop; the cheap unlocked probe ahead of it exists only to keep the graph-scoped path off the relation lock, and can only err toward keeping the table. Both are no-ops on SQLite, whose `BEGIN IMMEDIATE` fence already holds the database's single writer slot from probe through commit.
When the recorded shape is `stale` — the state only a recreate repairs — and the storage that would have to be recreated holds another graph's rows, the rebuild refuses with `ContributionRebuildUnsupportedError` and the new reason `shared-storage-in-use` instead of either destroying content it cannot reconstruct or re-stamping this graph's marker over a physical shape nothing verified. The refusal names the other graph ids and the sanctioned maintenance-window sequence. Vector contributions are unaffected: their storage is per-`(graph, kind, field)`, so no other graph's data is ever in reach.
- [#427](https://github.com/nicia-ai/typegraph/pull/427) [`facef56`](https://github.com/nicia-ai/typegraph/commit/facef560d14c38607d6414818c65e43dc65a88d2) Thanks [@pdlug](https://github.com/pdlug)! - Harden the Operational Identity release and adjacent write paths found during its adversarial review. Identity interchange now exports one repeatable-read snapshot, uses target-bound keyset pagination pinned to code-point order via the dialect adapter's `binaryText` seam so a `base@V` content token minted on 0.45 still matches its recomputation on PostgreSQL, cancels cleanly, and refuses streams that would deadlock a serialized connection — a PGlite connection, a bare `pg`/neon `Client` (including a checked-out `PoolClient`), a `Pool` explicitly configured with `max: 1`, a postgres-js client built with `{ max: 1 }`, a better-sqlite3 handle, a `bun:sqlite` database, a sql.js database, a local (`file:`/`:memory:`) libSQL client, or Cloudflare Durable Object storage, whose transaction frame is ambient on the storage object. The refusal is one EXCLUSIVE long-lived-stream lease per serialized resource, not a one-time observation and not a cross-kind-only exclusion: at most one interchange stream of any kind holds a given connection, so all four pairings are refused rather than only the two that mix kinds — an import behind an export snapshot (including through a user-wrapped stream that no longer identifies its source backend), an export snapshot behind a streaming import, and now export-behind-export and import-behind-import too, which previously reached the driver as a nested `BEGIN` after chunks had already committed. Whichever long-lived stream starts second gets a typed `ConfigurationError` instead of both hanging: `details.code` names the condition that holds the connection (`INTERCHANGE_SHARED_SERIALIZED_BACKEND_SNAPSHOT` behind an export snapshot, `INTERCHANGE_SAME_SQLITE_BACKEND_SNAPSHOT` when the object-identity detector is what answered, or the new `INTERCHANGE_SERIALIZED_IMPORT_IN_PROGRESS` behind another import), while the new `details.requested` and `details.heldBy` name which pairing was actually refused, so a same-kind refusal is never reported as something it is not. `"import-stream"` is the kind of EVERY long-lived import, not only the chunk-streaming one: `importGraph` takes the lease for the whole call and `trustedImportGraphStream` / `trustedImportGraph` hold it for the whole trusted session, so both APIs can now throw these serialized-connection `ConfigurationError` codes — a new error TYPE on a trusted-import surface that previously threw only `TrustedImportError`. Every exit releases the lease, including a mid-stream producer failure and a synchronous throw out of `backend.transaction(...)` (a closed handle, a refused pool checkout), which would otherwise strand the connection for the life of the process. Relatedly, SQLite's manually framed transactions no longer let a failing `ROLLBACK` mask the failure that caused it: SQLite auto-rolls-back on `SQLITE_FULL` / `SQLITE_IOERR` / `SQLITE_NOMEM`, so the unwinding `ROLLBACK` can itself fail with "cannot rollback - no transaction is active" — the caller now receives the ORIGINAL error and the rollback failure is warned instead of thrown. Graph merge uses injective composite keys, preflights provenance sidecar collisions, and refuses merge options it cannot honor instead of ignoring them.
Edge identity checks now include kind and endpoint kind on every create, delete, and get-or-create path, including tombstoned rows, and that check is carried by the write statement itself rather than re-derived beside it: `UpdateEdgeParams`, `DeleteEdgeParams`, and `HardDeleteEdgeParams` each gain an optional `kind`, and a backend that receives it MUST scope the statement to that kind. Both bundled Drizzle backends satisfy that contract through one shared predicate, but a hand-written `GraphBackend` has to honor it or it will silently widen a write it was told to narrow. Because a kind-scoped statement that also requires `deleted_at IS NULL` is its own recheck, the redundant in-transaction re-read that used to precede each edge delete and hard delete is gone — the statement either matches the edge the caller named or affects nothing. Node `bulkDelete` remains one atomic, hookless bulk operation exactly as in 0.45. Edge `bulkDelete` changes behavior in 0.46: 0.45.x looped single deletes and fired per-item `onOperationStart`/`onOperationEnd` hooks for each one, and 0.46 makes it one atomic single-transaction batch that emits NO hook events at all — neither per-item nor bulk, since `onBulkOperationStart`/`onBulkOperationEnd` fire only for node `updateWhere` and no bulk-hook coverage for deletes exists yet — so a consumer that relied on those per-item events for audit or metrics must either keep deleting individually (single `delete` still fires per-item hooks) or capture the deletions another way, such as from the ids it passes in and the rows it reads back; an id in the batch that belongs to another edge kind is refused with `ValidationError` carrying `EDGE_IDENTITY_MISMATCH_CODE`, rolling back every delete already applied earlier in the same batch.
Constrained writes no longer take their decision from a read the write cannot vouch for. Every write whose correctness rests on a check-then-write — edge cardinality `one`, `unique`, and `oneActive` (including the create and resurrect legs of `getOrCreateByEndpoints`, single and bulk), node-kind disjointness on create, and a `kindWithSubClasses` uniqueness constraint that actually expands to more than one kind (a scope covering a single kind probes exactly the row the uniques table's own primary key then reserves, so that key IS its fence) — now runs its probe and its write under the same per-graph mutual exclusion, whether or not the store enables `history` or `revisionTracking`. That exclusion previously arrived only as a SIDE EFFECT of recorded capture's advisory lock, so the DEFAULT PostgreSQL store — no history, no revision tracking — raced: two writers each probed a graph that satisfied the constraint and each committed, producing exactly the duplicate the constraint exists to prevent, with no error on either side. On PostgreSQL the fence is that same transaction-scoped advisory lock, now taken for the constraint's sake rather than the clock's; on SQLite it is the `BEGIN IMMEDIATE` writer slot the backend already holds. Writes with nothing to check — an unconstrained create, any delete, a cardinality-`many` edge — take no lock at all, so the cost tracks the constraints a graph actually declares rather than becoming a blanket serialization.
A backend running WITHOUT transactions has no fence to take, and a constrained write there is now REFUSED rather than run unfenced. Both halves of the fence are transaction-scoped constructs — SQLite's `BEGIN IMMEDIATE`, PostgreSQL's `pg_advisory_xact_lock`, which outside a transaction is acquired and dropped inside its own implicit single-statement one and excludes nothing — so "can this backend fence" and "does this write run inside a transaction" are the same question, and the refusal is keyed on exactly that. It is a `ConfigurationError` with `details.code` `CONSTRAINT_WRITE_FENCE_UNSUPPORTED` and `details.constraint` naming WHICH declared constraint needed the fence — `edgeCardinality`, `edgeMatchKeyConvergence`, `nodeDisjointness`, or `nodeUniquenessScope` — because "this backend cannot fence constrained writes" is unusable advice while "your `cardinality: 'one'` edge cannot be enforced here" is actionable; the `suggestion` carries the per-class way forward. The blast radius is Cloudflare D1 (auto-detected as `transactionMode: "none"`), `drizzle-orm/neon-http`, and any SQLite backend explicitly built with `transactionMode: "none"` — Durable Objects are NOT affected, since `do-sqlite` reports `capabilities.transactions: true` and fences normally. On those three, a declared-constraint write that previously raced silently now throws: a `cardinality` other than `many` cannot be created or resurrected, a `disjointWith` kind cannot be created, a `kindWithSubClasses` unique that actually expands past one kind refuses on create AND update, and `getOrCreateByEndpoints` can no longer take its CREATE leg (a call that FINDS an existing edge still returns it, and a `many` resurrection is an id-keyed UPDATE re-deriving no verdict, so both keep working; the BULK form fences its whole batch and therefore refuses whatever the outcome would have been). Everything unconstrained is untouched — a `many` edge created, updated and deleted, any node delete including one whose kind participates in a disjointness axiom, and a `scope: "kind"` unique whose uniques primary key IS its fence — so this is not a blanket loss of write access on those engines. Refused rather than degraded, per the accepted-or-refused rule: a constraint enforced only when nothing races is the exact defect the fence exists to close, and reporting it as enforced would make the invariant above false precisely where it matters.
A store created with `coalesceUnchangedUpserts` likewise stopped letting an optimization change an answer. A single-node `upsertById` decided to skip its write from an autocommit read, so a writer committing between that read and the skip left the caller told its props were stored while the store in fact held the other writer's — a DIFFERENT outcome from the same call with the flag off, where the update's own in-transaction re-read merges the caller's props over whatever it finds. The skip is now taken only on evidence re-read inside a transaction after that first observation, and a losing verdict falls through to the ordinary write path; only an upsert that is about to coalesce pays the second read, so a store without the flag, or one whose props differ, keeps the single read and single write it always had.
`getOrCreateByEndpoints` now converges rather than retrying once and hoping. Its single-shot retry became a bounded loop of three attempts whose ordinary case is cheaper than before — under the fence a losing writer learns of the winner from its own in-transaction lookup instead of from a `CardinalityError` — and whose exhaustion is a typed refusal rather than a livelock or a stray constraint error leaking out of a lost race: a competitor that repeatedly creates and removes the same match key ends in a `DatabaseOperationError` naming that pattern and telling the caller to serialize or retry. The bulk edge `getOrCreateByEndpoints` and bulk node `getOrCreate` paths, which had no retry at all and surfaced a concurrent winner as a raw constraint violation, gained one.
`efSearch` is now refused everywhere it cannot be applied, not only on PostgreSQL. The guarantee that vector search never silently drops an accepted option held only for the PostgreSQL path: every SQLite backend accepted `efSearch` and ignored it, on the vector path and the hybrid path alike (the hybrid path dropped it in a second place, while rebuilding the vector parameters), so a caller tuning recall got the default frontier and no indication. All of them now ask one owner, and an engine with no per-search ANN frontier refuses with `UnsupportedBackendCapabilityError` — `details.capability` `vector.searchFrontierTuning`, `details.reason` naming the limitation (`sqlite-vec`'s `vec0` KNN takes only `k`, the page size; libSQL's DiskANN `vector_top_k` fixes `search_l` at index-creation time). PostgreSQL is unchanged, including its existing refusals for a non-HNSW slot and a driver that cannot scope `SET LOCAL`, which now come from that same owner instead of from a second spelling of the same decision.
Separately, sql.js backends could not execute a compiled query at all. The compiled-execution adapter recognized any client exposing `prepare()` and then called `all()` on the resulting statement, but sql.js's `Statement` has no `all()` — it is a `bind`/`step`/`getAsObject`/`free` cursor — so the first prepared statement threw. Client detection is now shape-specific, sql.js is excluded from the compiled path, and it runs through Drizzle's own session, which drives that cursor correctly; `bun:sqlite`, whose statement DOES expose `all()`, keeps the compiled path.
The merge-provenance sidecar now claims its graph id MARKER-FIRST instead of inferring ownership from circumstantial evidence: the durable `ProvenanceOwner` marker is the sidecar's FIRST write of any kind, committed inside the schema fence (`schemaWriteTransaction` — the same per-graph fence every schema commit and schema-managed write already respects) and BEFORE the sidecar schema is registered, which is possible because the marker is a plain node row needing no per-graph DDL. A competing writer therefore either commits first and is seen, or waits until the claim has committed; a refused open leaves the occupant byte-identical, writing no schema row, no marker, and no provenance row. Because the marker precedes the schema, the resumable interrupted state is marker-WITHOUT-schema (or marker beside a pre-marker legacy schema), which resumes by registering or migrating the schema — while a graph carrying the exact current sidecar schema with NO marker is a state this module cannot produce and is refused UNCONDITIONALLY as `unowned-exact-schema-graph`, contents never consulted: empty and provenance-shaped occupants are refused exactly like any other, since contents an application could have written are not evidence of authorship. Freedom is judged by occupancy across EVERY per-graph row table the backend names through its `tableNames` port — nodes and edges, but equally recorded-time history, the revision clock and origins, identity assertions with their recorded ledger, closure and separation, fulltext, and unique keys — so a graph id holding only, say, identity or fulltext rows is occupied and refused, and an unregistered schema is never taken as evidence of a free namespace. Only the exact validated live marker counts as ownership: a tombstoned, malformed, wrong-target, or non-canonical `ProvenanceOwner` row refuses with the reason `corrupt-ownership-marker` and is never overwritten or resurrected. Refusals report one of five typed reasons under `GRAPH_MERGE_PROVENANCE_ID_COLLISION` — `application-graph`, `empty-legacy-sidecar`, `unupgradeable-legacy-sidecar`, `unowned-exact-schema-graph`, or `corrupt-ownership-marker` — each carrying remediation specific to the state actually found, and a backend that exposes no schema fence refuses an unclaimed sidecar with `GRAPH_MERGE_PROVENANCE_CLAIM_UNFENCED` rather than claiming without atomicity (an already-owned sidecar needs no claim and still opens there). One writer class takes neither the per-graph advisory lock nor the active schema row — a schema-LESS raw `createStore` writer, or a direct `backend.insertNode` / `insertEdge` — and at PostgreSQL's READ COMMITTED its insert could commit between the claim's fenced re-inspection and the claim's commit, leaving the marker on a graph id an application had just made its own. That window is closed rather than accepted: on PostgreSQL the claim issues `LOCK TABLE , IN SHARE ROW EXCLUSIVE MODE` inside the fence and before the re-inspection, draining in-flight row writers and holding new ones off until the marker commits. `SHARE ROW EXCLUSIVE` and not `SHARE` because the mode must be SELF-exclusive: two concurrent claims on different sidecar ids hold different advisory locks, so under `SHARE` both would acquire it and then both request `ROW EXCLUSIVE` for their own marker INSERT — a lock-upgrade deadlock PostgreSQL resolves by aborting one. The cost is real and bounded: while a claim runs, every node and edge write on the whole DATABASE waits, for the duration of a few probes and one INSERT with no caller code inside — and the lock is taken only when a sidecar is created, upgraded from the pre-marker schema, or resumed after a crash, never on the common path where an already-owned sidecar opens with no fence at all. SQLite takes no such lock; `BEGIN IMMEDIATE` already owns the single writer slot.
`persistProvenance: true` is now honored or refused, never dropped: the sidecar is opened and claimed PRE-COMMIT, so an occupied sidecar graph id or a backend that cannot fence the claim refuses the whole merge as `InvalidMergeOptionsError` (`details.option` `"persistProvenance"`, with the originating `ConfigurationError` as `cause` and its code echoed as `details.provenanceErrorCode`) and leaves the target unmodified — where 0.45 would have committed the merge and reported the same configuration verdict as a `warnings` entry. Those verdicts are as true before the merge as after it, so reporting them post-commit left the caller with a committed graph and a stated option TypeGraph had silently ignored. The post-commit best-effort warning path survives only for what is genuinely transient — a row write failing against a sidecar this library already owns. One visible consequence of claiming early: after a `persistProvenance` merge the sidecar (marker and schema) exists even if the merge itself later fails, holding an owned, empty sidecar and no target change. `mergeIncremental`'s refusal of a non-`"flag"` `onBasePropertyConflict` is now a typed `InvalidMergeOptionsError` (`MERGE_ERROR_CODES.invalidOptions`, category `user`) instead of a plain `MergeError`, so it is catchable the same way every other refused merge option is. `mergeIncremental` additionally refuses a fork point that moved under it. Its plan is a set of diffs against one `base@V`, and only the TARGET was ever allowed to advance while the merge ran — but nothing checked, so a write landing on the fork-point store mid-call left the commit applying diffs against an ancestor that no longer existed. The fork point is now frozen for the duration of the call: the version read before planning is carried into the commit and re-compared as the first act of the commit transaction, and a mismatch raises `BaseVersionMismatchError` (`GRAPH_MERGE_BASE_VERSION_MISMATCH`) naming the expected and live fork-point bases rather than committing. And `branch()` no longer leaks the working copy's backend when the post-transfer schema-anchor read fails — ownership of that engine transfers to `branch()` on the strategy's success path, and a `branch()` that reports failure as `err(...)` hands the caller no handle to close.
Operational Identity's derived separation relation is never published in a state that under-reports separations. The relation is created INSIDE the transaction that fills it — the fenced path issues its DDL under `schemaWriteTransaction`, the schema-commit path returns the DDL as data for the commit transaction to issue — so a commit refused by the `IDENTITY_PROFILE_MIGRATION_PENDING` gate, a stale CAS, or a contradiction now creates nothing at all, where previously it stranded a readable, empty relation that the next open skipped because "present" was what suppressed the rebuild. What that cannot undo, a per-graph predicate heals: the fill decision is "does THIS graph hold live `different` assertions and no separation rows", not "does the table exist", so a relation left empty by an older version or by another graph's provisioning is rebuilt at the next open of the graph that owns the assertions. The predicate is exact in both directions — it shares the fill's registry kind filter and additionally requires the assertion's two endpoints to resolve to DIFFERENT identity classes, since a contradicted ledger projects to a degenerate pair the relation's CHECK refuses, which is its own fault with its own error rather than an unfilled relation. Identity DDL is serialized database-wide by a constant-keyed advisory lock (`typegraph:identity-ddl`; a no-op on SQLite, whose writer slot already serializes the database), taken inside the per-graph schema fence and outside the per-graph identity locks, because the identity relations are shared by every graph while the fence is not. A backend that cannot publish that upgrade atomically — missing `schemaWriteTransaction` or `identityTableDdl` on the fenced path, or `executeSchemaDdl` on the commit path — is refused with the new `ConfigurationError` code `IDENTITY_UPGRADE_REQUIRES_ATOMIC_DDL` naming the missing ports, but only when a fill is actually owed; both bundled Drizzle backends implement all three when transactions are enabled. Finally, `isSeparated` no longer trusts an empty read: a graph with zero separation rows whose ledger holds a live, kind-filtered `different` assertion across two distinct classes raises `IDENTITY_STORAGE_MISSING` with the new `details.reason` `"unfilled"` rather than answering "not separated". That is the state a Store handle opened while the relation did not exist would otherwise slide into the moment another graph's upgrade created the shared relation mid-session; the remedy is to reopen the Store, which runs the fill, and the error says so. That proof is taken once per Store handle rather than once per read: it settles a property of the graph, not of the pair, and the assertion ledger has no index that answers it cheaply — proving it per read cost a 200-assertion same-only import 32% and eight concurrent `assertSame` 56% on SQLite, on the workload class whose separation relation is legitimately empty forever. A handle opened while the relation was absent still refuses, because a handle that cannot read the relation never records a proof; what a kept proof no longer re-detects is a relation truncated out of band midway through one handle's life, which `validateIdentity()` reports and the CHECK constraint still refuses at the next fusing write.
Three smaller guards join that one. The recorded revision clock can no longer move backward: its upsert advances the stored row only WHERE the stored revision is strictly less than the one being written, so a late allocation cannot rewind a clock other readers have already passed, and a caller that supplies an explicit stale `previousRevision` now gets a `ConfigurationError` stating that the write would have moved the graph's revision clock backward — carrying the graph id and both revisions — instead of quietly winning. Operational Identity's SQLite writes now name the one state they cannot recover from: an identity mutation inside a transaction the CALLER began and TypeGraph adopted (`store.withTransaction(externalTx)` / `store.withRecordedTransaction(externalTx)`) can find its read snapshot invalidated before it ever takes the writer slot if that transaction was opened `BEGIN DEFERRED` and another connection committed first, and SQLite cannot upgrade a stale snapshot in place. That surfaces as a `ConfigurationError` with `details.code` `IDENTITY_TRANSACTION_NOT_WRITE_FENCED` and `details.sqliteCode` `SQLITE_BUSY_SNAPSHOT`, telling the caller to roll back and reopen with `BEGIN IMMEDIATE`; TypeGraph's own transactions already open that way, so the state is unreachable without an adopted frame. And the three PostgreSQL error shapes a concurrent `CREATE ... IF NOT EXISTS` race can take — SQLSTATE `23505`, `42701`, and `XX000` carrying "tuple concurrently updated" — are now classified by one shared predicate rather than by each call site's own partial spelling of the set, so a race one site tolerated is no longer a hard failure at the next.
Edge `matchOn` composite-key construction and its per-field match comparison, embedding/fulltext field extraction, and the uniqueness path — both the unique-key computation and the `where` predicate's evaluation — now read a props bag by declared own key rather than plain property access, so a field named after an `Object.prototype` member (`toString`, `constructor`, `valueOf`) can no longer resolve to the inherited prototype member instead of the field's actual (absent) value: a unique constraint over an absent field named `toString` keyed on the inherited function, producing the empty key under `binary` collation and throwing `TypeError` under `caseInsensitive`, where it must key as absent like every other missing value. The remaining NUL-joined cache and bucket keys in edge and node operations, and legacy provenance record ids, are now built with the same injective tuple encoding already used elsewhere, closing the last collision-prone key constructions.
That own-key discipline now extends past reads of a props bag. A prototype-named field the schema DECLARES is projected and returned as its stored value instead of being short-circuited as prototype noise, so selecting a field called `toString` or `valueOf` answers with what was written rather than with the inherited function; the field tracker asks the schema introspector whether the name is declared before deciding, and an UNDECLARED prototype name still resolves exactly as it always did. Schema canonicalization builds its sorted form on a null-prototype bag, so a schema carrying a `__proto__` property is no longer canonicalized — and therefore hashed and diffed — identically to one where the property is absent; the output is byte-identical for every schema without such a key, so no existing schema's hash moves. Graph merge's property bags got the same treatment, which is what lets a fork-side DELETION of a `__proto__` property record as a deletion: the deletion marker is an assignment, and on an ordinary object literal `Object.prototype`'s setter swallowed it, silently reverting the delete to the base value. A graph-extension document that declares a PROPERTY named `__proto__` is refused outright with `RESERVED_PROPERTY_NAME`, because schema validation cannot carry it — at any depth, so a NESTED object field named `__proto__` is refused on the same grounds rather than only a top-level one. `defineNode` / `defineEdge` refuse the identical declaration at definition time (`ConfigurationError`, `details.conflicts`) — at ANY nesting depth, walking nested object schemas and every wrapper (optional, nullable, default, arrays, records, unions, lazy) structurally through Zod's public `def`, with a dotted path in the error — so the two authoring paths no longer disagree about the same unstorable field: it was a typed refusal on the document path and silent data loss on the typed one. It is reachable only through a computed key — `z.object({ __proto__: … })` written literally sets the shape object's prototype instead of creating an entry, while `z.object({ ["__proto__"]: z.string() })` yields a shape whose `Object.keys` really does contain it — and it is UNSTORABLE rather than merely reserved, because Zod drops the key from every parse result and reports success even when the field is required.
The same misreading has a WRITE side, and it is now closed as a class rather than case by case. `bag[key] = value` on a `{}` literal does not create an entry when `key` is `__proto__`: it invokes `Object.prototype`'s `__proto__` setter, which reparents the bag for an object value and does nothing at all for a primitive, so the value is dropped and every later own-key read agrees the writer never wrote it. Kind names (`isValidKindName` admits `__proto__` exactly as it admits `toString`), schema property names, JSON-Schema keywords, query aliases and `JSON.parse`d document keys are all data, and all of them admit it. `normalizeEdges` in `defineGraph` and every other data-keyed accumulator in the tree now build through one owner, `createDataKeyedBag`, so an EDGE kind named `__proto__` survives `defineGraph`, schema serialization, and a live store round trip instead of vanishing between the config and the registration. Because the class had already recurred twice from an incomplete enumeration, it is made self-enforcing: a ratchet test scans `src/**` for statement-position `{}` initializations and fails on any that is not allowlisted with a stated reason. Behavior note for callers: none of this is observable on returned values — every record TypeGraph hands back (serialized schema maps, aggregate rows, select contexts, migration counters, extension documents) accumulates on a null-prototype bag internally and is spread into an ordinary object at the public boundary, which preserves an own `__proto__` alias as data while restoring `Object.prototype`, so `row.toString()` and `record instanceof Object` behave exactly as before. Relatedly, the field tracker and the selective projection now track a DECLARED field named `toJSON` as the stored data it is, instead of exempting the name unconditionally; the exemption survives, unchanged, for kinds that do NOT declare it, where it exists only to keep an incidental `JSON.stringify` of the tracking context from being recorded as a field access.
Two remaining leaks of that internal null prototype are closed, and one of them was a behavior that depended on which query plan ran. A smart-selected alias object and its `meta` are guarded PROXIES, and a proxy's target is caller-observable — `instanceof`, `Object.getPrototypeOf` and every other internal method resolve against it, and no `get` trap can disguise it — so `ctx.p instanceof Object` answered `false` under a selective projection and `true` under the full mapper, for the same query. The boundary spread now happens inside the guard, so both mappers hand back objects rooted at `Object.prototype` while a projected field named `__proto__` survives as an own key; the tracking context handed to the `select` callback on the field-tracking pass got the same treatment, so the probe and the engine agree. `TransactionReceipt`'s `writes.nodes` and `writes.edges` are likewise ordinary objects now, matching `writes.identity`, which always was — a `__proto__` kind still reads back as an own key with its count. And the builder handed to an index `where:` callback is no longer built on a null prototype either.
`defineNode` / `defineEdge` now REFUSE a schema containing a `z.lazy()` whose getter cannot run yet, with a `ConfigurationError` naming the kind and the dotted path. This is reachable from one shape: a mutually recursive pair declared AROUND the definition call, so the second `z.object()` const is still in its temporal dead zone when the first one's getter fires. Previously that branch was skipped, and skipping it was a fail-open — a definition is validated exactly once, so a `__proto__` nested under the unreadable subtree was accepted at definition time and then silently dropped by every parse, which is the precise outcome the unstorable-name refusal exists to prevent. Recursion itself is not refused: declaring both consts before the definition — which the error message asks for — resolves every getter, and the walk then reports the real conflict at its full nested path. Note that a `z.lazy` property field is typed `unknown` by the query introspector regardless, so predicates over it degrade; recursive property schemas are not a supported shape, and this makes the one silently-wrong case loud.
An import UPDATE now asserts every component its verdict read, closing the temporal half of the class rounds 6 and 7 closed for edge kind and endpoints. `UpdateNodeParams` and `UpdateEdgeParams` each gain an optional `expectedValidFrom`, with the same MUST-apply contract as `UpdateEdgeParams.kind` — a backend that receives it has to put it in the statement's own `WHERE`, and the three states are distinct: omitted asserts nothing, `null` asserts `IS NULL` (an open-left window), a string asserts equality. Both bundled Drizzle backends satisfy it through one shared NULL-safe predicate builder; a hand-written `GraphBackend` must honor it or it will silently widen a write it was told to narrow. All four import update legs — node and edge, batched and per-row — now state the bound they validated the document's window against, so a concurrent hard-delete-and-recreate between the probe and the write matches no row instead of ignoring a `validFrom` the document stated or persisting a `validTo` below the new row's `validFrom`. They state it on exactly the terms the store paths do, because it is the same verdict object: a document naming neither `validFrom` nor `validTo` made the verdict read no bound, so its properties update is fenced on identity and liveness alone and a concurrent recreate that only moved the bound no longer refuses it. A node write that consequently matches nothing is reported per row with the new message prefix `INTERCHANGE_NODE_UPDATE_TARGET_CHANGED`; the edge equivalent keeps the published `INTERCHANGE_EDGE_KIND_CONFLICT` prefix, with the validity bound added to its message text. `ImportError` still carries no `code` field, so the prefix is the branchable token. Relatedly, a node update's row write and its uniqueness transition now commit or fail as ONE unit: the new keys are claimed before the row write (the claim upsert reports the key's final owner, so it IS the conflict gate, and a transaction holding the key cannot lose it to a peer), the row write follows, and the old keys are released only once it lands — with the claims compensated away if it does not. Import is the reason this has to hold on its own terms rather than on the transaction's: `onConflict: "update"` catches a per-row `UniquenessError`, records it, and commits everything else, so an ordering where the claim can fail AFTER the row changed reported `updated: 0` for a row whose props HAD changed, whose old reservation was released, and whose new reservation belonged to another node. The fulltext and embedding syncs still run after the row write, since a write that lands on nothing must not re-derive them.
The same fence now covers the STORE update paths, which read the probed row's `valid_from` for exactly the same verdict. `store.nodes.*.update` / `upsertById` and `store.edges.*.update` / `upsertById` carry `expectedValidFrom` into the statement's own `WHERE` — but only when the window verdict actually consulted the row's bound, which is when the caller stated a `validFrom` to compare against it or a lone `validTo` to invert against it. A plain `update({ props })` names no window, reads no bound, and is fenced by nothing extra; that conditionality is not an optimization but the same "only what it asserted" rule the edge identity components already followed, since predicating a write on a component the caller never claimed refuses writes that are legitimate. The decision has one owner, and it hands over the predicate rather than a flag: `assertWritableValidityWindow` now RETURNS the `expectedValidFrom` fence its verdict obliges the write to carry — empty when the verdict read no stored bound — so the answer comes from the branches the guard actually took and there is nothing left for a caller to re-derive. Interchange import had re-derived it, asserting the probed `valid_from` unconditionally and over-fencing exactly the props-only updates the store paths left alone; that second spelling is gone rather than corrected. When the assertion does catch a replaced row the update CONVERGES rather than failing: it re-reads, re-merges the caller's partial props over the current props, and re-judges the window against the current bound, so a stated window that no longer fits is refused with the same typed `ValidationError` it would have raised on the first attempt, and one that still fits is applied to the row that really exists. Convergence is bounded at one retry; a peer that keeps replacing the row ends in a `DatabaseOperationError` naming the contention instead of a livelock or a false "not found". Two adjacent defects in the same paragraph of code are fixed with it: `applyNodeResurrect` reserved its uniqueness keys before the gating `deleted_at IS NOT NULL` update and kept them when the gate refused, so a resurrection that lost its race left reservations behind for a revival it never performed (it now runs through the same claim/gate/release transition as `applyNodeUpdate`, which gives the reservations back when the gate matches nothing); and the bulk `getOrCreateByConstraint` decided whether to resurrect from the uniques row its batch probe captured, while the single-item path decided from the node row it was about to write — one decision with two owners, now read from the node row on both.
Relatedly, the builder a uniqueness `where` clause names fields on now answers for every declared field rather than only the fields the props bag happens to carry — which is what its type has always promised (`-?` makes every schema field required on the builder, precisely so a partial constraint can ask whether an OPTIONAL field is present). Naming an absent field previously hit the builder object's prototype instead: an everyday partial constraint over an absent optional field threw `TypeError: Expected a defined value` for every node written without it, and a field named `toString` found `Object.prototype.toString` and threw `isNull is not a function`. Such a field now evaluates as null, which is what a partial constraint means by absent — the same builder shape schema serialization has always captured a `where` clause with. `defineGraph` also refuses a `where` clause it can already see is broken, at definition time rather than on the first write it distorts: a callback that returns something other than a predicate, or a predicate naming a field the kind's schema does not declare, throws a `ConfigurationError` naming the kind, the constraint, and — for the undeclared field — the fields the kind actually declares. A statically typed caller could express neither mistake, so this bites generated or untyped definitions, where the old behavior was a constraint that quietly matched every row or partitioned on a field that was absent forever. A third state joins those two: a constraint carrying a `where` on a kind whose schema exposes no `.shape` — not an object schema, so there is no declared-field set to check the clause against — is REFUSED rather than left unvalidated, because skipping the check silently would disable the guard for exactly the untyped callers it was written for, who are also the only callers able to put a non-`ZodObject` there. Narrowly so: a plain `unique: [{ fields }]` on such a schema needs no shape to be meaningful and still works. And the same malformed clause is refused at EVALUATION too, not only at definition — `checkWherePredicate` throws the equivalent `ConfigurationError` for a callback that returns a non-predicate, so the third reader of a `where` clause now agrees with the other two (definition-time validation and persistence-time capture, all three reading the clause through one owner) instead of quietly treating a broken constraint as one that applies to every row. That matters for constraints built outside `defineGraph`, which never passed the definition-time gate. Because the check evaluates the clause, a `where` callback now runs one extra time when the graph is defined, so it must be pure — which it already had to be, since the uniqueness path evaluates it per write. This validates node kinds whose schema exposes an object shape; edge `unique` constraints are not covered.
This adds public API surface — `InvalidMergeOptionsError`, `ExportStreamCancelledError`, `EDGE_IDENTITY_MISMATCH_CODE`, `MERGE_ERROR_CODES.invalidOptions`, the `GraphBackend.identityTableDdl` port with its `IdentityTableNames` type, the `VectorSearchFrontierTuning` type, `ExportOptionsSchema`'s `signal?: AbortSignal` (accepted by both `exportGraph` and `exportGraphStream`), the optional `kind` on `UpdateEdgeParams` / `DeleteEdgeParams` / `HardDeleteEdgeParams` from the `@nicia-ai/typegraph/backend` entry point plus the four optional endpoint assertions `fromKind` / `fromId` / `toKind` / `toId` on `UpdateEdgeParams` alone (one assertion that moves together or not at all, under the same MUST-apply contract as `kind`), the `"shared-storage-in-use"` member of `ContributionRebuildRefusal`, `ValidationErrorDetails.operation` widened with `"delete"` and `"hardDelete"`, the `details.requested` / `details.heldBy` pairing on every serialized-connection interchange refusal, and the new error codes: `INTERCHANGE_EXPORT_STREAM_ABORTED` on `ExportStreamCancelledError`, the `vector.searchFrontierTuning` value of `UnsupportedBackendCapabilityError`'s `details.capability`, and the interchange, identity, and provenance `ConfigurationError` codes (`INTERCHANGE_SERIALIZED_IMPORT_IN_PROGRESS`, `INTERCHANGE_SHARED_SERIALIZED_BACKEND_SNAPSHOT`, `INTERCHANGE_SAME_SQLITE_BACKEND_SNAPSHOT`, `IDENTITY_UPGRADE_REQUIRES_ATOMIC_DDL`, `IDENTITY_TRANSACTION_NOT_WRITE_FENCED`, `IDENTITY_STORAGE_MISSING`'s new `details.reason` `"unfilled"`, `GRAPH_MERGE_PROVENANCE_ID_COLLISION` with its five refusal reasons, `GRAPH_MERGE_PROVENANCE_CLAIM_UNFENCED`, and `CONSTRAINT_WRITE_FENCE_UNSUPPORTED`, whose `details.constraint` carries one of `edgeCardinality` / `edgeMatchKeyConvergence` / `nodeDisjointness` / `nodeUniquenessScope` — the four members of the internal `ConstraintFenceReason` union, which is not itself exported; branch on the string values).
Three refusals join the surface without a stable `details.code`, so match them by class plus `details` rather than by code. `assertApproximateMetricSupported` throws a `ConfigurationError` when `similarTo(..., { approximate: true })` is combined with a `metric` override that differs from the slot's declared metric, carrying `details` `{ nodeKind, fieldPath, requestedMetric, declaredMetric, indexType }`. `defineNode` / `defineEdge` throw a `ConfigurationError` for a schema property named `__proto__`, carrying `details.conflicts` and a `nodeType` / `edgeType` key. And a per-row import failure whose message is prefixed `INTERCHANGE_EDGE_KIND_CONFLICT` appears in `result.errors` when an interchange edge's id belongs to another kind — `ImportError` has no `code` field, so the message prefix is the branchable token, following the existing `windowErrorOf` idiom used by the validity-window import errors.
Two of those are BREAKING for callers who reach past the bundled implementations. `VectorCapabilities.searchFrontierTuning` is REQUIRED, not optional: a hand-written vector strategy must now state whether its engine has a per-search ANN frontier knob — `{ tunable: true, parameter, indexType, requiresTransactionScope }` or `{ tunable: false, reason }` — rather than inheriting silence, which is the exact defect the field closes, and a strategy that omits it no longer compiles. And a hand-written `GraphBackend` must apply the new `kind` on the three edge params when it is present, and the four endpoint assertions on `UpdateEdgeParams` alongside it; a backend that accepts and ignores either turns a write the caller narrowed into an unscoped one, and nothing above it re-reads to catch that any more. Kind alone is not enough for the update: an edge's endpoints are immutable for a given row but its id is not, so a concurrent hard-delete-and-recreate under the SAME kind with DIFFERENT endpoints satisfies a kind-only predicate, and an upsert that resolved the id BY endpoints would write to an edge pointing somewhere it never looked. The endpoint fields are stated only by a write that actually checked them — a plain `update` on a kind-scoped collection resolved the edge by id and kind and states none of them, because predicating on endpoints it never checked would refuse legitimate writes. `tests/edge-write-self-verification.test.ts` asserts the contract against the bundled backends.
It also changes behavior callers can observe: edge `bulkDelete`'s hooks; `importGraph` / `trustedImportGraph` / `trustedImportGraphStream` newly throwing the serialized-connection `ConfigurationError` codes (a new error type on the trusted-import surface); `persistProvenance`'s new pre-commit refusal turning a merge that previously committed-and-warned into one that refuses without touching the target; `mergeIncremental` refusing a commit whose fork point moved mid-call, where it previously committed diffs against a vanished ancestor; a stale explicit `previousRevision` now refused instead of rewinding the recorded clock; a SQLite vector or hybrid search that supplies `efSearch` now throwing `UnsupportedBackendCapabilityError` where it previously searched at the default frontier and said nothing; `defineGraph` throwing on a unique `where` clause that names an undeclared field, returns a non-predicate, or sits on a kind whose schema is not an object schema, and evaluating every such callback once at definition time; a graph-extension property named `__proto__` refused with `RESERVED_PROPERTY_NAME` at any nesting depth, and the same name in a `defineNode` / `defineEdge` schema refused with a `ConfigurationError`; a constrained write on D1, `neon-http`, or a `transactionMode: "none"` SQLite backend now refused with `CONSTRAINT_WRITE_FENCE_UNSUPPORTED` where it previously committed unfenced; `similarTo` with `approximate: true` and a mismatched `metric` override now refused where it previously served the exact scan and dropped one of the two options silently (`store.search.vector` / `hybrid` already refused every mismatched override on their own broader rule, and the builder's EXACT path stays deliberately wider — only the silent half is closed); an interchange edge whose id belongs to a row with a different kind OR different endpoints now reported as a per-row `INTERCHANGE_EDGE_KIND_CONFLICT` error naming the mismatched components, under `onConflict: "update"` (which previously overwrote the other row's properties while silently keeping its endpoints) and `"skip"` (which previously counted it present and silently lost it) — and the import's `updateEdge` statements carry the full identity assertion in their own `WHERE`, closing the concurrent-recreate window for endpoints exactly as it was closed for kind; and `MergeIncrementalArgs.options` narrowed to `Omit` — a compile error for code that passed `options.target` to `mergeIncremental`, and a runtime `InvalidMergeOptionsError` for untyped callers that still do, where the named `target` argument previously won and `options.target` was silently ignored. So this release ships as a `minor`, not a `patch`.
- [#426](https://github.com/nicia-ai/typegraph/pull/426) [`eb4b9c1`](https://github.com/nicia-ai/typegraph/commit/eb4b9c1076eb67d54cf3a7692ebe8f9a3b203453) Thanks [@pdlug](https://github.com/pdlug)! - Refuse a `validFrom` that a live row's update cannot store, instead of accepting
it and writing without it.
An in-place update never rewrites `valid_from`; only a resurrection does. Stating
a bound that named a different instant used to block coalescing, so the upsert
wrote — bumping the version and capturing a history row — while the bound itself
was dropped at the SQL builder and the row's window never moved. It now raises a
`ValidationError` whose issue carries the new exported code
`IMMUTABLE_VALIDITY_LOWER_BOUND`, naming both the stated instant and the one the
row holds so the caller can restate it without a second read.
This reaches every path that accepts `validFrom` against a live row: `upsertById`
and `bulkUpsertById` (nodes and edges, including a repeated id in one batch, which
is judged against the row the batch just queued), `getOrCreateByEndpoints` and
`bulkGetOrCreateByEndpoints` with `ifExists: "update"` — which previously dropped
the option before it reached any guard — and interchange import's
`onConflict: "update"` legs, where it is recorded as a per-row error prefixed with
the code rather than aborting the import.
What stays legal: restating the bound a row already holds (nothing to apply, so
nothing is ignored); a create or a resurrection, both of which store a stated
bound and are the way to give a row a different one; zero-width windows; and
`getOrCreateByEndpoints` returning an existing edge, which performs no write at
all.
Previously-accepted writes now refuse, so this is a MINOR bump — the same
precedent as the window refusals in the two releases before it.
Note for temporal imports: replaying an `includeTemporal: true` export over rows
that were created separately now reports those rows instead of updating their
props under a lower bound it ignored. Omit `validFrom` from the update document,
export with `includeTemporal: false`, or import into a fresh graph.
- [#383](https://github.com/nicia-ai/typegraph/pull/383) [`bab0752`](https://github.com/nicia-ai/typegraph/commit/bab075213b480d61c91be06a2c833e510ab61a18) Thanks [@pdlug](https://github.com/pdlug)! - Merge an inherited row's end-of-validity instead of discarding it
`update(id, {}, { validTo })` on a branch is an ordinary write, but the merge
silently dropped it: modification detection compared properties only, so a
branch that ended an inherited node's or edge's validity merged as a no-op.
There was no workaround preserving row identity and history — deleting the row
was the only statement the merge honored, and it is a strictly stronger one.
An end-of-validity is now treated as a **sibling of deletion**:
- one branch ends a row → that end is committed, including a _later_ end that
extends the window;
- several branches end it differently → no conflict; the **earliest** end wins
(a fixed, commutative rule, so the merge stays order-independent);
- `mergeIncremental()`'s target already ended it → the target's end stands, the
same committed-target precedence identity survivors already get;
- one branch ends it and another deletes it → deleted, with **no**
`DeleteModifyConflict` — the stronger statement absorbs the weaker one;
- a branch re-states the end the target holds → nothing is staged at all: no
write, no version bump, no history row.
`MergeReport` gains `validityEnds`, listing every row whose end the merge
changed and the branches that claimed it — the arbitration is silent by design,
so this is how a caller sees it happened. Window deltas the commit cannot apply
to a live row (a fork `validFrom` divergence, or a `validTo` cleared back to
open — both reachable only by soft-delete + resurrect inside a fork) are now
reported in `dropped` with reason `"window-not-applicable"` instead of being
ignored.
**Behavior change.** Merges where a branch ended an inherited row now write that
end, so new version bumps, history rows, and recorded-time entries appear where
a no-op used to be. There is no opt-out flag: a permanent knob for "does the
merge lose data" is worse than this note. Nothing that previously succeeded now
fails.
Also hardens `coalesceUnchangedUpserts`: the requested and stored valid-time
bounds are compared as instants rather than as raw text, so the decision cannot
come to depend on a dialect's timestamp rendering. A bound that is not a
representable instant still counts as a change, so it reaches the write path
that rejects it rather than being coalesced away.
- [#426](https://github.com/nicia-ai/typegraph/pull/426) [`eb4b9c1`](https://github.com/nicia-ai/typegraph/commit/eb4b9c1076eb67d54cf3a7692ebe8f9a3b203453) Thanks [@pdlug](https://github.com/pdlug)! - Report a non-canonical validity bound whether or not `coalesceUnchangedUpserts`
is on.
A parseable-but-non-canonical bound equal to the stored instant —
`"2100-06-01T00:00:00Z"` against a stored `"2100-06-01T00:00:00.000Z"` — compared
as "unchanged" with coalescing on, so the write was skipped and the
`ValidationError` the same call raises with coalescing off was swallowed. An
unrelated performance flag decided whether malformed input was reported.
A non-canonical REQUESTED bound now counts as a window change, so it reaches the
write path and raises identically either way. Re-stating a window in canonical
form still coalesces, including against a driver that renders the stored value as
an equivalent zoned string: only the stored side needs canonicalizing, because
the requested side is held to canonical form by the write validation this no
longer hides.
- [#394](https://github.com/nicia-ai/typegraph/pull/394) [`6dd43c4`](https://github.com/nicia-ai/typegraph/commit/6dd43c40bbb988fc8d112f177f03caab86c18359) Thanks [@pdlug](https://github.com/pdlug)! - graph-merge: scope the edge repoint/dedupe fold to collisions repointing caused
A TypeGraph store is a multigraph: nothing enforces uniqueness on
`(from, kind, to)`, `create()` makes a parallel edge, and
`getOrCreateByEndpoints()` is the opt-in set-semantics accessor. The merge's edge
fold nevertheless grouped **every** staged edge by `(from, kind, to)` and collapsed
each group onto its lowest-sorting edge id, so a branch that created a parallel
edge lost one of the two rows — and _which_ one it lost depended on how the
branch-created id happened to sort against the existing one.
The fold is now restricted to what it was designed for. It groups staged edges by
the endpoint pair they named **before** repointing, and collapses one row per pair —
so a collision the canonicalization itself induced (`x → a` and `x → b` both becoming
`x → c*`) still folds to a single edge, keeping the existing min-id-survivor,
property-reconciliation, and end-of-validity behavior. Edges that already shared
their endpoints are no longer folded together: each distinct edge id commits as its
own parallel row, and a valid-time end lands on the row whose author claimed it
rather than migrating to an unrelated survivor.
A group that mixes the two folds only **across** the pairs. A repointed `x → b`
joining two parallel `x → a` rows merges into one of them, and the other row still
commits with the edit its author made — repointing said nothing about the rows that
were already there. Previously the whole group collapsed, which dropped that edit
silently: a folded-away row is never rewritten.
What makes two staged edges "the same row" is their **edge id**, not equal
properties. One inherited edge staged by several branches still folds into a single
write with its property disagreements reconciled; a branch-created edge is a new
parallel row even when its properties coincide with an existing one's.
Merges that previously collapsed parallel edges will now commit both, and the
spurious `PropertyConflict` those collapses reported between two rows that were
never the same row is gone.
- [#406](https://github.com/nicia-ai/typegraph/pull/406) [`248f56a`](https://github.com/nicia-ai/typegraph/commit/248f56ab2c8bc849fe9385157f02344c7ceb612d) Thanks [@pdlug](https://github.com/pdlug)! - Refuse valid-time windows of negative width. A write whose `validTo` precedes
the row's effective `validFrom` describes a row that stopped being true before
it started — observable at no `asOf` coordinate, and unrepairable by any later
write — and it used to be accepted silently on every path except a node update.
It now raises a `ValidationError` whose issue carries the new exported code
`INVERTED_VALIDITY_WINDOW`.
This is a behavior change: writes that previously succeeded now fail. Two shapes
refuse where they did not before.
- A stated `validFrom` / `validTo` PAIR must be ordered, on node and edge
`create`, `upsertById`, `bulkUpsertById`, `getOrCreateByEndpoints` and its bulk
form, and on an imported document. `getOrCreateByEndpoints` judges the pair
before its existence probe, so whether a call is valid no longer depends on
whether the edge happens to exist yet.
- An UPDATE's lone `validTo` must not precede the lower bound the row carries.
Nodes already enforced this; edges did not, which is how a graph merge could
hand a committed edge an end predating its start and still report success. This
covers a resurrecting write too: an edge RETAINS its `valid_from` across
resurrection, so reviving one into a window that closed before it began now
means restating the start — pass `validFrom` alongside `validTo`. Landing a
revived edge in the ENDED state is otherwise unchanged.
`getOrCreateByEndpoints` and its bulk form now honor `validFrom` on the
`"resurrected"` branch, where they previously accepted it and silently dropped
it — which is what left the refusal above with no way to satisfy it. As the
backend has always documented for a resurrecting write, naming `validFrom`
asserts the COMPLETE window, so an accompanying `validTo` is applied and an
omitted one REOPENS the revived row rather than leaving the tombstoned
incarnation's end in place. A `"found"` or `"updated"` live edge is unaffected:
its stored lower bound is history and still stays put.
Interchange import records the refusal as a per-row error prefixed with
`INVERTED_VALIDITY_WINDOW`, so one bad row does not abort the import; its
`onConflict: "update"` legs are held to the existing row's `valid_from` exactly
as a direct `update` is. Trusted import refuses the whole stream with reason
`invalid_stream`.
Two shapes stay legal, deliberately. A ZERO-width window
(`validTo === validFrom`) is what a same-instant retraction produces at
millisecond precision, so the store's own output still round-trips. An INSERT
carrying a lone historical `validTo` still means "born already ended": the write
instant stamped as `valid_from` is a storage convention rather than a caller
assertion, and such a row is read back through `includeEnded`.
- [#380](https://github.com/nicia-ai/typegraph/pull/380) [`2c5dd29`](https://github.com/nicia-ai/typegraph/commit/2c5dd29d9c23024119589d5360b265a0c3ab49da) Thanks [@pdlug](https://github.com/pdlug)! - Store the cause of an identity assertion's ending instead of deriving it.
The identity assertion relation and its recorded mirror gain nullable
`ended_by_kind` / `ended_by_id` columns. A node soft-delete cascade stamps the
deleted node's `(kind, id)` onto every assertion it ends, in the same statement
that closes the row; `NULL` means the row was retracted explicitly. Graph
merge's `RetractionCause` now reads that column instead of comparing an
assertion's `valid_to` against a deleted endpoint's `deleted_at`.
This removes the derivation's same-millisecond residue: a retraction issued in
the same millisecond as the delete that followed it is now classified as
`explicit` and survives a merge whose deletion is overruled, where the
timestamp comparison could only read the tie as a cascade and drop the
branch's intent. The hard-delete residue remains by design — a hard delete
removes the assertion rows outright, so no evidence survives to read.
Archival interchange carries the cause as an optional `endedBy` on each
assertion, so an export/import round-trip preserves why an assertion ended.
Import rejects an `endedBy` on an open assertion
(`IDENTITY_IMPORT_ENDED_BY_WITHOUT_END`) or one naming a node that is not an
endpoint of the assertion (`IDENTITY_IMPORT_ENDED_BY_NOT_ENDPOINT`); a CHECK
constraint on the relation backs both rules at the database.
Operational Identity has not shipped in a release, so the relation changes
shape with no migration path.
- [#268](https://github.com/nicia-ai/typegraph/pull/268) [`9721ba2`](https://github.com/nicia-ai/typegraph/commit/9721ba2854f6e5504018ee8c3b7a0eaaf87314bb) Thanks [@pdlug](https://github.com/pdlug)! - Add the opt-in TypeGraph Identity Profile with typed store, transaction, and
temporal-view APIs; configurable same-ID folding or assertion-only identity;
kind-branded and hydrated member reads; idempotent assertion receipts; ended
assertion retraction results; assertion history; interchange and graph
merge propagation; identity-expanded traversal; cross-backend closure storage;
and fail-fast capability errors for non-transactional D1 and neon-http drivers.
Harden ontology construction and reload validation: propagate disjointness
through interleaved subclass and equivalence closure, validate inverse endpoint
compatibility and partner uniqueness, reject unresolved extension edge names in
`inverseOf` and `implies` while retaining absolute external IRIs, recompute
serialized closures, and deprecate the type-level `sameAs` and `differentFrom`
factories in favor of Operational Identity.
**Behavior changes.** Ontology and registry validation is now stricter and runs
both at graph construction and when a persisted schema is loaded, so a few
patterns earlier versions silently accepted now throw a `ConfigurationError`:
duplicate ontology relations, hierarchical self-loops, disjointness
contradictions (a kind disjoint with itself, with a subclass ancestor, a common
subclass of two disjoint parents, or a kind declared both `equivalentTo` and
`disjointWith`), multiple distinct `inverseOf` partners for one edge, inverse
endpoint incompatibility, and unresolved extension edge names in `inverseOf`
or `implies`. To
recover, fix the graph definition; for a persisted extension document, correct
the stored document before upgrading (or rewrite it through the previous minor,
which still accepts it). Interchange documents remain readable across versions —
`1.0` documents are still accepted on import, and exports write `2.0`.
Trusted import rejects identity-enabled target stores (`identity_unsupported`)
and identity-bearing input (`invalid_stream`) rather than silently dropping
assertions or leaving the derived closure empty; use `importGraphStream` for an
export that carries identity truth.
Bundled SQLite and PostgreSQL backends provision the three identity relations,
including effective custom `SqlSchema` names, before first-enable preflight. An
already-enabled graph with missing identity storage instead fails with
`IDENTITY_STORAGE_MISSING`; restore missing ledgers from backup, or recreate a
missing derived closure and rebuild it before serving traffic.
`create()`/`upsertById()` of a soft-deleted same-`(kind, id)` row now resurrects that
row on every graph (properties replaced, validity window reset so `validFrom`
becomes the resurrection instant) rather than leaking a storage constraint
error. These are additive-strictness and semantics-pinning changes on top of
the new opt-in profile, hence the minor bump.
**Type-level breaking notes for backend and tooling authors.**
1. `ResolvedSqlTableNames` gained three required fields
(`identityAssertions`, `recordedIdentityAssertions`, `identityClosure`).
Out-of-tree `GraphBackend` implementations must supply them; the
`SqlTableNames` input type keeps these optional, so only the resolved
type is total.
2. `SqlSchema` (the abstract class) gained three abstract members
(`identityAssertionsTable`, `identityClosureTable`,
`recordedIdentityAssertionsTable`). External subclasses must add them;
the `createSqlSchema` factory path is unaffected.
3. `FORMAT_VERSION`'s literal type changed from `"1.0"` to `"2.0"`.
Comparisons like `FORMAT_VERSION === "1.0"` are now type errors; both
versions remain accepted on import.
**Behavioral note.** `revisionNow()` now returns
`Promise` (a branded string, assignable to
`string`; use `asRecordedInstant` to round-trip).
**Review-hardening pass (same release).**
- `store.identity` and the read-only view `identity` surfaces now use the same
conditional presence as `tx.identity`: the property does not exist on
identity-disabled graph types, so misuse is a compile error. The
`IdentityFacadeFor` / `IdentityReadFacadeFor` helper aliases and the
duplicate `IdentityNodeRef` type are gone (use `IdentityFacade`,
`IdentityReadFacade`, and `GraphNodeReference`); the loose input type
formerly named `GraphNodeRef` is now `IdentityNodeRefInput`.
- `StoreView` and `RecordedStoreView` are now type aliases over an
implementation class plus `ViewIdentityAccess`, exported alongside a
construction-compatible `const`. `new StoreView(...)` and
`instanceof StoreView` keep working; subclassing them does not.
- `MergeReport.merged` gained an `identity: { asserted, retracted }` section
(`MergedCounts`), and `DroppedItem` is now a discriminated union
(`kind: "node" | "edge" | "identity"`) so dropped identity assertions are
enumerable in the report.
- Identity merge conflicts — including transitive `same`/`different`
contradictions, retract/reassert races, and assertions over merge-deleted
nodes — are detected at plan time and surface as `IdentityMergeConflictError`
(`GRAPH_MERGE_IDENTITY_CONFLICT`) through `merge()`'s returned `Result`.
Convergent edits (the re-asserting branch itself also retracted the pair)
merge cleanly.
- `ImportError.entityType` widened to `"node" | "edge" | "identity"`; identity
import failures are recorded in `result.errors` instead of throwing.
Archival identity imports now bound validity windows (`validTo` must not be
in the future for ended rows, `validFrom` must not be for open rows) with
`IDENTITY_IMPORT_FUTURE_VALID_TO` / `IDENTITY_IMPORT_FUTURE_VALID_FROM`.
- Changing `identity.sameIdAcrossKinds` is now classified a breaking schema
change requiring explicit migration; explicit `migrateSchema()` rebuilds the
identity closure atomically with the schema commit, and an unapplied
identity-only breaking change surfaces `IDENTITY_PROFILE_MIGRATION_PENDING`
rather than a generic `MigrationError`.
**Performance.** Current-coordinate identity reads (`membersOf`, `areSame`,
`areDifferent`, `representativeOf`, `nodesOf`) were O(total graph size) on
SQLite — the class-members lookup defeated the closure's class index and the
planner scanned every live node per read. The rewritten statement is
O(class size): ~40x faster on a populated graph (0.013 ms vs 0.56 ms per
read at ~6,000 nodes), with a smaller improvement on PostgreSQL.
**Follow-up hardening (same release).** `ValidationIssue` gained an optional
`assertionId` field carrying the offending identity assertion structurally;
identity import failures (self-assertions included) attribute their
`result.errors` entries by that id, never by message parsing. The identity
enablement preflight is derived inside `initializeSchema()` itself, so every
public first-commit path — bare `ensureSchema`/`initializeSchema` included —
builds and validates the closure atomically with version 1. Identity reads on
`includeTombstones` views hydrate soft-deleted rows the coordinate makes
visible instead of silently dropping them.
**Import error attribution.** The import coordinator tags rethrown errors
with the id of the assertion it was applying, so `ImportResult.errors`
attribution for contradictions and missing endpoints identifies the failing
assertion rather than the first assertion sharing its endpoints.
**The identity preflight is not substitutable.** `initializeSchema()` and
`SchemaManagerOptions` no longer accept a schema-commit preflight callback —
a no-op callback could suppress the mandatory closure build at version 1.
Both (and `MigrateSchemaOptions`) instead accept the effective `SqlSchema`
(`schema`), and every identity-enabled schema commit derives the closure
preflight internally from it.
**One schema source in the batteries-included constructors.** The nested
`schemaManagement` option no longer accepts `schema` (typed out and stripped
at runtime): the effective `SqlSchema` has exactly one source, `store.schema`,
which also drives physical table provisioning — a second schema could name
tables that were never created. The manager brand-validates the `schema`
option with `requireSqlSchema()` before any DDL or version commit, so a
schema-shaped plain object is rejected (`INVALID_SQL_SCHEMA`) instead of
committing a closure into tables the Store never reads.
**Historical bridges must exist; plan-time simulation knows the profile.**
Archival identity imports now require every ended assertion's endpoints to
exist structurally (soft-deleted rows qualify; the store's own exports
already satisfy this), so a hand-built document can no longer conduct
historical identity through a node that never existed. The graph-merge
plan-time contradiction check now simulates the target's identity semantics
— implicit same-id folds under `sameIdAcrossKinds: "fold"` and ontology
`disjointWith` between class member kinds — so those contradictions surface
as `GRAPH_MERGE_IDENTITY_CONFLICT` at plan time instead of a generic commit
failure. Counterfeit schema objects are rejected before any identity DDL
runs, on fresh and already-enabled graphs alike.
**Assertion-free nodes join the plan-time simulation.** The merge planner's
contradiction check now seeds its universe with every post-merge canonical
node and the live target peers sharing their ids (one kind-free indexed
probe, only under `sameIdAcrossKinds: "fold"`), so a node no assertion
names — newly created, retyped, or an existing same-id peer — can no longer
fold into a disjoint-kind class undetected and fail at commit as a generic
merge error.
**Universe seeding, precisely.** The plan-time simulation seeds retyped
canonical nodes under the kind the commit writes (not their pre-retype
kind), the live same-id peer probe reads the merge TARGET when it differs
from the diff source (`mergeAgainstBase`, `mergeIncremental`), and the
incremental commit revalidates the probed peer set inside its transaction —
a same-id peer landing in the plan→commit window is refused as the same
typed replan error the other window guards raise.
**The window guard ranges over the committed plan.** The incremental
fold-peer revalidation compares only ids the final plan folds on —
commit-ready canonical nodes and remapped assertion endpoints — so a window
row at an id canonicalization dropped is tolerated as an ordinary target
advance instead of raising a spurious replan error.
**The window guard is class-transitive.** The incremental fold-peer guard
also snapshots each final seed's structural identity class at plan time and
revalidates the fingerprints inside the commit transaction — a window row
or assertion that joins a seed's class through another member (leaving the
seed's direct same-id peers untouched) is refused as the typed replan
error, and a rerun surfaces the contradiction as a plan-time
`GRAPH_MERGE_IDENTITY_CONFLICT`.
**A validated baseline, exactly.** The incremental identity guard now
re-probes and snapshots the final seeds' classes AFTER planning and re-runs
the identity simulation against that exact snapshot — its members join the
simulation universe unlinked, with connectivity rebuilt from the
deletion-filtered fresh ledger and fold unions — so drift landing between
planning and the snapshot fails as a typed plan-time conflict instead of
becoming the guard's baseline. Fingerprints are structurally encoded (injective for ids
containing any character) and carry a liveness bit, so a planned assertion
endpoint deleted in the commit window is refused as the typed replan error
rather than failing generically.
**Negative truth in the baseline.** The post-plan identity recheck consumes
the target's FRESH assertion ledger (not the pre-planning staging capture),
and the transaction guard carries a deterministic fingerprint of the
`different` assertions touching the guarded universe — a `different`
committed in either window is refused typed instead of surfacing as a
generic commit failure.
**The identity guard covers both profiles.** The incremental identity
baseline, class/liveness fingerprints, and negative-ledger guard run for
every identity-enabled merge — under `sameIdAcrossKinds: "ignore"` too,
where explicit assertions still change plan legality. Only the same-id
fold expansion stays profile-gated; the plan-time simulation additionally
models the profile-independent create-time constraint that one id cannot
be shared by ontology-disjoint kinds, and the direct-peer window check
refuses a disjoint same-id arrival under `"ignore"` while tolerating a
benign one.
**Replacement is legal.** Planned node deletions are excluded from both
sides of the incremental identity guard (peers, liveness, class members,
and the ledger slice), and `applyMergePlan` soft-deletes nodes BEFORE the
node writes — so a plan replacing a node with a disjoint same-id one (the
order the create-time constraint permits, and the order the same
operations run directly on a store) commits instead of being falsely
rejected or failing at apply.
**Deleting a bridge splits the class.** The incremental recheck derives
connectivity from the deletion-filtered fresh ledger and the checker's
fold unions — never by pre-linking the old closure's filtered member
lists — so a plan that deletes an identity bridge and asserts its former
ends `different` commits instead of being falsely rejected. Snapshot class
members still join the simulation universe (unlinked) so fold links at
unprobed ids keep participating.
**The transaction re-derives legality.** The incremental commit guard's
final step re-runs the full identity simulation on transaction reads —
fresh deletion-filtered ledger, snapshot members, fold unions — so drift
that leaves every fingerprint unchanged (a redundant `same(a, b)` that
becomes the surviving link once the plan removes the pair's bridge) is
refused as the typed replan error instead of failing generically at apply.
**One assertion id, one truth — validated where it can be typed.** The
planner refuses one id staged for two different complete truths and any
staged id already identifying different truth among the target's stored
rows (ended included, exactly the set the import coordinator compares);
the commit transaction revalidates every planned id against transaction
reads (both commit modes), so a window row reusing a planned id — even with
endpoints entirely outside the guarded universe — refuses as the typed
replan error instead of a generic id-conflict at apply.
**Retractions carry their complete truth.** A merge plan's identity
retractions are full expected rows, never bare ids: the planner validates
each one against the row its id identifies on the target and SKIPS —
reported as `identity:retraction-target-mismatch` in `dropped` — a
retraction whose id the target reuses for different truth, instead of
ending a row the branch never saw. The commit transaction revalidates the
surviving retractions (and every planned assertion id) by id in BOTH
commit modes; snapshot commits need this explicitly because the legacy
base@V token fingerprints only CURRENT assertions, so an ended window row
claiming a planned id would otherwise slip through to a generic apply
failure. The raw staged assertions are also checked one-id-one-truth
BEFORE the semantic survivor dedupe, closing the validity-only collision
(same id, same pair, different `validFrom`) that dedupe used to collapse
silently while the report listed the id as both applied and dropped.
**The applier is the completeness backstop, typed.** Any identity refusal
that still escapes the commit — an invariant the plan-time simulation
does not (yet) mirror — is translated into the typed
`IdentityMergeConflictError` with the applier's error as its cause,
instead of surfacing as the generic merge wrapper. Identity-typed
environment errors (missing profile, non-atomic backend) pass through
unchanged. A property-based law suite additionally quantifies the merge
contract over randomized identity histories on both backends: refusals
are always typed, a committed ledger is internally consistent, pre-merge
truth survives unless a branch retracted it or deleted an endpoint, and
the report never lists an id as both dropped-as-duplicate and newly
current.
**Truth replacement is visible to the diff.** The identity diff compares
ids present on both sides by COMPLETE truth, not presence: a branch that
hard-deletes an assertion's endpoint (physically removing the row),
recreates it, and imports the same id for different truth used to diff as
empty — the merge silently kept the base truth the branch had replaced.
The replacement now stages as a retraction plus a new assertion, and
because the applier never reuses an ended row's id, the merge refuses
typed instead of silently preserving either side.
**Identity semantics extracted; translation at the applier boundary.**
The plan-time identity derivation, contradiction simulation, and commit
guards now live in `graph-merge/merge-identity.ts` with a one-directional
dependency from the merge orchestrator (functions take a structural
`IdentityPlanSlice`, never the full plan type). The typed-conflict
translation wraps exactly the identity-apply call inside the commit, so
it also classifies refusals whose identity code lives in nested
validation issues (`details.issues[].code`) and — because only identity
rows are applied at that boundary — a missing-node error there can only
mean a vanished assertion endpoint, which now translates too instead of
surfacing as the generic wrapper. Exact-duplicate staging (two branches
importing one identical row) no longer reports the id as dropped while
applying it.
**Five laws, three lanes.** The property suite now also holds every
successful merge to BRANCH-EFFECT accounting — every truth a branch holds
is applied with equal complete truth, enumerated as dropped, retracted,
or invalidated by an endpoint deletion; silent loss is a law violation —
and runs the whole law set in three lanes: snapshot `merge()` under both
identity profiles (with hard-delete/recreate and same-id fold peers in
the operation alphabet) and `mergeIncremental()` against a target that
ADVANCED after the fork, where branch truth meets independently-moved
target truth. Truth-preservation and branch-effect exclusions are
truth-aware: a retraction excuses a row's death only when the retracted
COMPLETE truth matches, and a hard-delete/recreate excuses exactly the
rows it physically killed, not everything ever touching the node. A
dropped-as-duplicate id must never be current post-merge. The generator
skips only expected semantic refusals (contradiction, missing node); any
other error fails the run rather than silently emptying the histories.
Independent-target merge semantics are now documented in the identity
guide.
**The survivor pick respects committed truth.** The law suite caught its
first live defect within a day: a branch-minted assertion id could win
the semantic-pair dedupe against the target's own committed row — the
applier (idempotent per pair) then skipped the write, so the report
claimed an id as applied that never landed while listing the target's
committed row as dropped. Ids already committed on the target with the
exact staged truth now always win the survivor pick, pinned by a
deterministic incremental test alongside the law.
**The simulation uses the plan's REAL canonical map.** Both closure
re-runs (post-plan and in-transaction) previously reconstructed the
member→survivor map from the report-shaped resolutions, which drops pure
ontology-retype clusters and mis-keys mixed-kind members — degrading the
decisive in-transaction backstop into judging endpoints at pre-merge
identities (a false negative) and enabling an unresolvable replan loop (a
false refusal). The plan now carries the exact `canonicalOf` map the
commit repoints edges with, and the reconstruction is deleted. The
simulated base ledger is also deletion-filtered inside the checker
itself, so all three call sites share one post-deletion rule.
**An overruled deletion no longer ends identity truth.** A node
soft-delete cascades — it ends every open assertion touching the node —
so the deleting branch's diff stages those endings as retractions
indistinguishable from intent. When the delete/modify resolution keeps
the modification (the default `"flag"` and `"modifyWins"` policies), the
node survives, and the cascaded retraction is now dropped with it —
reported as `identity:deletion-overruled` — instead of ending the
resurrected node's assertions anyway.
**Identity-only merges advance the revision clock.** The interchange
import records capture touches through its own recorded binding, so a
merge whose only effect was creating assertions never marked the mutation
as written: the durable revision clock stayed unmoved and every base@V
token went stale, letting a later commit's target-unchanged guard pass
against a target that DID move. The apply now marks the write from the
import summary, with a regression test on a revision-tracking store.
**Guard structure hardened.** The by-id freshness check is invoked
directly by BOTH commit paths (never through the peer-probe guard's early
return), the environment-code passthrough covers the identity
environment/corruption codes that must never be translated into replan
advice, and one id staged as both a new assertion and a retraction — an
applier-refusing shape currently unreachable through any supported
staging path — refuses typed defensively at plan time.
**External-review hardening (cross-model pass).** An independent review
with a different model produced six verified fixes: (1) the
deletion-overruled retraction filter is provenance-aware — a retraction
is dropped only when EVERY contributing branch is explained by an
overruled endpoint deletion, so a branch that retracted independently
keeps its effect (the earlier filter silently suppressed it); (2)
committed-row precedence in the survivor dedupe is RE-DERIVED after
endpoint canonicalization, closing the collision the first fix missed
when reconciliation collapses a branch pair onto a committed target
pair; (3) merges refuse, typed, any branch whose store ran a schema
operation after forking (its committed schema hash no longer matches the
fork source's) — schema side effects can no longer be smuggled into a
data merge as bare identity changes; (4) a kind-dropping
`migrateSchema()` now cascades the assertion ledger exactly as
`Store.removeKinds()` does, instead of stranding current assertions on
unregistered kinds where a later "no-op" merge would end them; (5) a
staged survivor's valid-time window travels with the commit write, so a
branch-authored — possibly already ended — window survives resurrection
instead of being reset to merge time and silently joining a live fold
class; (6) `merged.identity` reports rows the applier actually created
and ended (idempotent skips excluded) instead of planned intents, and
the replan-vs-conflict error suggestions are path-specific. Temporal
windows on MODIFIED inherited nodes remain outside merge state — a
documented boundary.
**Second cross-model pass: the fixes' own compositions.** A follow-up
external review of the previous round's fixes produced seven more
verified corrections. The schema-drift guard now anchors on the branch's
AT-FORK `(version, hash)` row — a round-trip migration that restores the
document hash still advances the monotonic version and is refused, and
unmanaged fork sources are no longer falsely rejected; revision-anchored
`base@V` tokens bake in the active schema version, fencing the same
round-trip on the target side (the legacy content fingerprint already
covered it). Kind-dropping schema operations cascade the assertion
ledger even when identity is DISABLED at drop time (the ledger, not the
schema profile, is the signal), and first enablement purges assertions
naming unregistered kinds, so historical orphans cannot be adopted into
a fresh closure. Node resurrection carries `validFrom` through the
internal update path (a branch-authored ended window no longer inverts
into merge-time-start), merged edges carry their staged windows exactly
as nodes do, and when the live incremental target itself contributed the
surviving member, the TARGET's committed window wins over a branch
re-window. Canonicalization that would move a COMMITTED assertion's own
endpoints refuses with a specific typed conflict (committed rows cannot
be rewritten), and window-identical upserts coalesce again instead of
rewriting version and history state.
**Final pre-merge pass.** A last scoped external review of the previous
hardening commit returned three refinements, all applied: the
disabled-identity cascade's outside-transaction emptiness probe is
skipped when THIS commit is the one disabling identity (writers on the
still-enabled prior schema could otherwise slip an assertion in between
probe and lock — the locked cascade always runs for that shape); node
writes validate the EFFECTIVE validity lower bound, so a lone historical
`validTo` on a resurrecting upsert refuses typed instead of persisting a
born-inverted, permanently invisible window (edge resurrection keeps its
sanctioned resurrect-as-ended contract — edges retain their stored lower
bound, so the node-side corruption cannot arise there); and bulk edge
coalescing compares explicit windows against the stored window, so
no-op incremental merges stop rewriting byte-identical target edges.
The property law lanes carry explicit five-minute test budgets sized
for coverage-instrumented CI shards.
- [#375](https://github.com/nicia-ai/typegraph/pull/375) [`fc6075d`](https://github.com/nicia-ai/typegraph/commit/fc6075ddca02de4c0fa9d15167c124868ab947b5) Thanks [@pdlug](https://github.com/pdlug)! - **The merge commit proves its own identity result.** After a merge commit's
identity DML, and inside the same transaction, the applier now re-derives the
identity classes the merge TOUCHED and refuses a contradiction there — a class
whose member kinds the ontology declares disjoint, or a current `different`
assertion whose endpoints share a class. Both commit modes run it, seeded from
the planned assertion and retraction endpoints plus (under
`sameIdAcrossKinds: "fold"`) the node identities the commit writes, so the cost
is proportional to the affected classes rather than the graph.
This makes a committed identity ledger correct independently of the plan-time
simulation, which reasons about state read before any write. The simulation and
the commit-window fingerprints remain as the diagnosability layer: they refuse
early, before anything is written, naming exactly what drifted.
Because the scans resolve classes through the materialized closure — the same
authority every current identity read uses — a closure that lags its ledger can
hide a contradiction as easily as invent one. On any inconsistency the closure
is rebuilt from the base relations inside the commit transaction and the scans
re-run against it: a clean second pass means the closure was stale and is now
repaired atomically with the merge, while a repeated contradiction aborts the
whole merge. There is no partial commit either way.
`IdentityContradictionErrorDetails.operation` gained a `"merge"` member for
this refusal, which reaches callers as the existing
`IdentityMergeConflictError` (`GRAPH_MERGE_IDENTITY_CONFLICT`) with the
contradiction as its cause.
### Patch Changes
- [#418](https://github.com/nicia-ai/typegraph/pull/418) [`5795127`](https://github.com/nicia-ai/typegraph/commit/57951275ee6310bcbfebadb1e3bdc46204769052) Thanks [@pdlug](https://github.com/pdlug)! - store: coalesce a bulk upsert that re-states a row's own validity window
`bulkUpsertById` now decides whether a requested `validFrom` / `validTo` is a
change the same way `upsertById` does, through one shared comparison, so a batch
and the same items applied one at a time write the same rows.
Two defects met in that comparison. The node bulk path refused to coalesce
whenever an item named a bound AT ALL, so any caller that re-stated a row
together with the window it already holds — a merge commit, or any
read-modify-write loop that round-trips `meta.validFrom` — bumped the row's
version and wrote a history and revision entry for a row that did not change.
The edge bulk path did compare, but compared the bounds as DRIVER TEXT: a
Postgres driver that renders `timestamptz` as a zoned string rather than a
`Date` yields text that is equivalent to the caller's canonical ISO bound
without being identical to it, so the same batch could coalesce on one backend
and write on another. Both paths now compare INSTANTS, and an unrepresentable
bound still counts as a change so the write path raises the `ValidationError`
the caller is owed rather than coalescing it away.
The bulk paths also track the window each queued write leaves behind, so a
repeated id in one batch is compared against the batch's own pending state
rather than the once-prefetched row. Previously an edge item that re-stated the
window the row held BEFORE the batch was read as unchanged and skipped, dropping
a write the sequential path performs. A later copy that re-states the window a
queued write established coalesces; one that names a bound the backend was left
to stamp (an omitted `validFrom` on a create) writes, since that instant is not
knowable batch-locally.
- [#363](https://github.com/nicia-ai/typegraph/pull/363) [`cdc904b`](https://github.com/nicia-ai/typegraph/commit/cdc904b574d4aa5a4ef10f3378b1ca4079373209) Thanks [@pdlug](https://github.com/pdlug)! - Chunk iterative graph-algorithm node-kind initialization within each backend's bind-parameter budget.
- [#396](https://github.com/nicia-ai/typegraph/pull/396) [`994c7da`](https://github.com/nicia-ai/typegraph/commit/994c7da9aff06d995680381e3b43c4a05bece5a3) Thanks [@pdlug](https://github.com/pdlug)! - Reach the candidate edge of a current-coordinate identity-expanded traversal by
an equi-join instead of a correlated membership scan
An identity-expanded hop at the current coordinate read the materialized closure
from inside a correlated `EXISTS`, so nothing in the join condition linked the
frontier row to the edge row. Both engines were free to enumerate
_frontier rows × edges of the matching kind_ and probe the closure per pair, which
cost quadratically in graph size.
Each traversal step now widens its frontier onto the closure's class members with
an outer join, so the candidate edge is reached by the same ordinary indexed
equality a traversal without identity expansion uses. One compiler path serves
both coordinates and both emitters. On SQLite a hop over 100,000 matching edges
from a 500-row frontier drops from 51.6 s to 77 ms, and `EXPLAIN QUERY PLAN`
seeks `typegraph_edges_from_idx` where it used to scan every matching edge per
source row; PostgreSQL drops from 9.2 s to 61 ms.
A traversal at a **historical** coordinate reaches its candidate edge through the
same step, so it gains the same join order: on SQLite an `asOf` hop over 100,000
matching edges drops from 31.3 s to 241 ms. That coordinate's own remaining cost
is the ledger reconstruction, still tracked in typegraph#310.
Results are unchanged at every coordinate: physical edges stay deduplicated, and
member visibility, the `sameIdAcrossKinds` profile and the read instant are all
resolved exactly where they were. The class members a current-coordinate step
joins are reached by seeking the closure from the frontier row, so the cost of
the widening tracks the frontier and its classes rather than the identity
population — see the follow-up changeset, which replaced the graph-wide relation
this change first shipped with that seek.
- [#400](https://github.com/nicia-ai/typegraph/pull/400) [`05af68d`](https://github.com/nicia-ai/typegraph/commit/05af68d122371ddeab03269536f2c7484ccf1a74) Thanks [@pdlug](https://github.com/pdlug)! - graph-merge: record each provenance contribution once, so the persisted count is
the rows actually written
Several planning phases legitimately observe the same
`(role, canonical, branch, source)` contribution. An inherited edge is credited
once when its modification survives delete/modify and again when the repoint
fold reads it as a source, and a fold set's `mergedIds` carries one entry per
staged copy — so a row staged by several branches re-offered each of its
branches once per copy. The tuple is exactly the sidecar row's identity, so
those re-observations were never new information: they inflated
`provenancePersisted.count`, and because a single `bulkUpsertById` batch cannot
create the same id twice, the over-count was the milder half: with
`persistProvenance: true`, a merge in which a single branch modified one
inherited edge failed the whole best-effort persist, so `provenancePersisted`
came back absent, a `provenance persistence failed …` warning was reported, and
NO provenance rows were written at all.
Contributions are now collapsed at the single recording funnel, so the record
list, the in-memory `provenance.byBranch` index and the reported count all
speak about distinct contributions. `persistProvenanceRecords` additionally
collapses records that hash to one id before the batch, which makes its
documented "row count written" true for any caller's record list. Every
genuinely distinct contributing branch is still credited.
- [#384](https://github.com/nicia-ai/typegraph/pull/384) [`cc7af7b`](https://github.com/nicia-ai/typegraph/commit/cc7af7b1ef53458296f27b73484cd799667df13c) Thanks [@pdlug](https://github.com/pdlug)! - Report the real active schema version in `StaleVersionError.details.actual` when
a PostgreSQL schema-managed write loses to a concurrent schema commit.
The write fence takes a `FOR SHARE` lock on the active schema row. At `read
committed`, a locking read that blocks behind an in-flight schema commit
rechecks only the row versions its own statement snapshot saw — so once the
winner marked the old row inactive, the fence saw no active row at all and
reported `actual: 0`, misrepresenting the database as having no active schema.
The fence now settles an empty locked read with a non-locking read, which
observes the committed winner, and reports `0` only when a graph genuinely has
no active version. The write itself was always correctly rejected; only the
error metadata was wrong.
- [#427](https://github.com/nicia-ai/typegraph/pull/427) [`facef56`](https://github.com/nicia-ai/typegraph/commit/facef560d14c38607d6414818c65e43dc65a88d2) Thanks [@pdlug](https://github.com/pdlug)! - Bound current-coordinate identity expansion by the frontier instead of the
identity population
An identity-expanded traversal at the current coordinate built its class relation
by self-joining the whole identity closure into a materialized CTE, before any
frontier predicate applied. The relation's size is the sum of the squares of
every class in the graph, so a hop from a single start row paid for identity
classes it never touched: nine unrelated classes of 501 members materialize
2,259,009 seed/member pairs, and the hop measured 564 ms on SQLite and 568 ms on
PostgreSQL where the equivalent traversal without expansion costs microseconds.
Doubling an unrelated class quadrupled the cost.
Each step now seeks the closure from its own frontier rows — the frontier row's
class through the closure primary key, that class's members through the class
index, each member's node for its visibility — so the peer relation is never
built for classes the query does not touch. The same hop measures 0.5 ms on
SQLite and 2.4 ms on PostgreSQL, and PostgreSQL's `EXPLAIN (ANALYZE)` reports 18
rows visited against 4,522,557. A single-start-row hop over 50,000 folded triples
drops from 387 ms to 1 ms on SQLite. Wide-frontier hops are unchanged: 500 source
rows over 100,000 matching edges measures 325 ms against 331 ms, because that
shape was already paying for a population it used.
The **historical** coordinate keeps its hoisted, materialized relation. Its rows
come from a recursive fixed point over the assertion ledger that no frontier row
narrows, so evaluating it once per statement is still the win, and the two
coordinates are now deliberately different strategies behind one interface rather
than one relation with two sources. Both remain a single compilation path across
dialects.
Results are unchanged at both coordinates: physical edges stay deduplicated,
member visibility is still resolved against the read instant, and a frontier row
in no class still expands to itself.
- [#391](https://github.com/nicia-ai/typegraph/pull/391) [`8eafebd`](https://github.com/nicia-ai/typegraph/commit/8eafebdae8d28fc48343fcd5d498f47b1684311f) Thanks [@pdlug](https://github.com/pdlug)! - Evaluate the historical identity-class reconstruction once per query instead of
once per candidate edge
An identity-expanded traversal hop under a historical coordinate (`asOf`,
`asOfRecorded`, or a non-current `view()`) has no materialized closure to read,
so it rebuilds classes from the assertion ledger. That rebuild used to sit inside
the correlated edge predicate, where SQLite re-materialized it for every
candidate _(source row, edge)_ pair, and under `sameIdAcrossKinds: "fold"` each
rebuild also scanned the structural same-id relation across the graph — a
quadratic term that quadrupled per doubling of graph size.
The reconstruction is now a single materialized query-level relation of
`(seed_kind, seed_id, kind, id)` rows, seeded by the nodes that have identity
peers rather than by the frontier, so it depends on nothing a traversal step
carries and is built once for the whole statement. Each step widens its frontier
onto that relation with an outer join, which turns the candidate-edge lookup into
the same ordinary indexed equality a traversal without identity expansion uses.
On the narrow-edge fixture (SQLite, all _n_ nodes acting as source rows) the hop
drops from 122/486/1984/8261 ms at _n_ = 250/500/1000/2000 to 7/7/14/28 ms, and
grows linearly rather than quadratically.
Results are unchanged at every coordinate. Current-coordinate traversal still
reads the materialized closure through its existing correlated predicate.
- [#425](https://github.com/nicia-ai/typegraph/pull/425) [`92354bf`](https://github.com/nicia-ai/typegraph/commit/92354bfb559c255a03b4b5e91741b87b98de5777) Thanks [@pdlug](https://github.com/pdlug)! - Ask props bags whether they carry a key with `Object.hasOwn` rather than `in`.
A props bag is data: its keys come from a JSON column, so a schema may declare a
field named after an `Object.prototype` member — `toString`, `constructor`,
`valueOf` — and such a field is ordinary data that survives Zod validation and
the JSON round-trip untouched. `in` cannot answer "does this row carry this
property" for such a bag, because `"toString" in {}` is `true`: a row that does
not carry the key reads as though it does, and the read that follows yields the
inherited prototype member instead of stored data.
This is a lost-write fix, not only hardening. In a graph merge, a fork's bag is
its full intended state, so a base property absent from it was deleted by that
fork. Under `in` that deletion was never detected for a prototype-named field:
no deletion tombstone was written and the base value survived the merge, silently
discarding the fork's write. The same misclassification credited a branch that
does not carry such a property with the inherited prototype member as if it were
a stored value, letting an invented claim compete in conflict resolution and be
reported to the caller as that branch's value. A schema diff also reported a
removed prototype-named property as an incompatible schema change rather than a
removal, because the absent field resolved to a function that was then compared
as though it were the field's new schema.
The edge fold and the node cluster union were affected in the same way, and their
worst outcome was a committed function. For a property the SURVIVING row does not
carry, both ask that row's bag for the value to keep, so under `in` they took the
inherited `Object.prototype` member and wrote that function into the merged row
instead of the value a member actually carried. The cluster union additionally
routed such a property through the separate base-property-conflict policy on the
strength of a base member that does not carry it, so the wrong policy decided the
committed value. The edge fold's claim filter separately counted a member that says
NOTHING about such a field as having AUTHORED it; the shared value collector
discarded that phantom claim, so the two agreed only by one absorbing the other's
mistake.
Two guards were quietly weakened rather than corrupted. Graph-extension validation
accepted a unique constraint on an undeclared field named after a prototype
member — it answered "declared" against the prototype — and went on to index a
field that does not exist. The evolve guard that refuses re-adding a kind whose
data cleanup is still pending never counted such a kind as added, so it skipped
the refusal.
The convention now has one owner, `hasOwnKey`, applied across graph-merge node and
edge property resolution, schema-diff property classification, schema-removal
reconciliation, interchange unknown-property stripping, graph-extension document
validation, query and index schema-field validation, the evolve pending-removal
guard, edge `matchOn` composite-key and match-comparison reads, and
embedding/fulltext field extraction. `in` remains correct, and still in use,
when both the key and membership question are internal: a discriminated union's
tag, a capability probe, a brand check, and the deliberate `Object.prototype`
lookup in selective projection. A user-supplied field name is always checked as
an own key, even when the schema shape itself is statically known, so names such
as `__proto__` and `constructor` cannot masquerade as declared fields through
`Object.prototype`.
A plain `bag[field]` walks the prototype chain exactly as `field in bag` does, so
the same misreading reached two more read paths that never used `in` at all. An
edge's `matchOn` composite key and its per-field match comparison
(`getOrCreateByEndpoints`, edge upsert dedup) read a caller's stored and input
props by a schema-declared field name; a field named after a prototype member
that neither bag carries as an own key now reads as `undefined` on both sides
instead of the same inherited function, so a match or non-match decision is never
made on a phantom shared value. `syncEmbeddings` and `computeFulltextContent`
read a declared embedding or searchable field the same way, so an undeclared
row no longer surfaces a prototype function as if it were the field's stored
value there either.
`then` and `toJSON` complete the same class from the other end. They are the two
names JavaScript itself probes — the thenable check and the `JSON.stringify` hook
— so every proxy standing in for a row resolved them to `undefined` up front to
stay safe to await and to serialize. They are also legal schema field names, and
answering them by NAME before consulting the data made the read side lie: a
declared field called `toJSON` came back `undefined` through smart selection while
the full mapper returned the stored string, so the same query answered differently
depending on whether the optimizer engaged — exactly the equivalence selective
projection exists to preserve. The predicate builders (query, traversal,
collection, and index WHERE) made such a field unaddressable outright, and field
tracking dropped a declared `then`, so the projection could not have carried it.
**A declared `then` or `toJSON` field is now tracked, projected, readable through
smart selection, and usable in a predicate.** The rule is the one the surrounding
fixes already follow: ask the data question first — `hasOwnKey` for a
materialized row, `hasDeclaredField` for a proxy whose key set is the schema —
and fall back to the probe exemption only once the answer is "not data", which is
what keeps `await` and `JSON.stringify` working on a partially projected row.
Returning an own `then` is safe as well as correct: props decode from a JSON
column, so the value can never be callable, and the thenable check ignores a
non-callable `then` exactly as it does on the plain objects the full path returns.
`isInteropProbeKey` owns which names those are, and an ESLint rule bans the bare
name comparison that used to stand in for the decision.
`__proto__`, the case originally reported, is the NARROW variant. Every VALIDATED
write path blocks it: Zod drops an own `__proto__` key, and `bag["__proto__"] = value`
assigns a prototype rather than creating a key, so an assignment-built bag cannot
carry one either. It is still reachable through `trustedImportGraph`, which by
contract does not validate properties and writes a caller's bag verbatim — the
stored JSON parses back with `__proto__` as an own key on both dialects. Recorded
here so the two are not confused: a prototype-named field needs nothing unusual at
all, while `__proto__` needs the trusted path.
- [#381](https://github.com/nicia-ai/typegraph/pull/381) [`e6fb356`](https://github.com/nicia-ai/typegraph/commit/e6fb35669ed0bfeab1cfce64aafc72afaf5698a2) Thanks [@pdlug](https://github.com/pdlug)! - identity: answer current different-ness with one probe on the separation relation
`identity.areDifferent()` and the `assertSame` contradiction precheck resolved
both identity classes and then loaded every current `different` assertion
touching one of them, scanning in JS for one that spanned the pair. That scan
grew with class size and, past the backend's bind budget, took more than one
statement. Both now probe the derived separation relation on its primary key
`(graph_id, class_key_low, class_key_high)` instead: `areDifferent` reads the
assertion ledger not at all, and the precheck reads it only to name the
conflicting assertion in the typed error it is already about to throw.
Results and typed errors are unchanged. Reads at a valid-time `asOf` or a
recorded coordinate still reconstruct from the ledger, since the separation
relation projects current assertions onto current classes. A probe never
answers "not separated" when it could not read: a missing relation refuses with
`IDENTITY_STORAGE_MISSING`, and any other driver failure propagates unchanged so
transient conflicts stay classifiable.
- [#404](https://github.com/nicia-ai/typegraph/pull/404) [`479ca78`](https://github.com/nicia-ai/typegraph/commit/479ca783781d9449a7b20422446c68fd702f516b) Thanks [@pdlug](https://github.com/pdlug)! - Fix `bulkUpsertById` throwing on a repeated id whose row does not exist yet.
`bulkUpsertById` applies items in order, so a repeated id in one batch is
last-write-wins — but that only held for an id that already existed. The create
branch queued its create without registering the id in the batch-local pending
map, so a second copy of a **new** id queued a second create and the batch failed
with `Node already exists` / `Edge already exists` (a unique-constraint violation
on some paths). Callers feeding a batch straight from a stream or a changeset,
where a key can legitimately appear twice, hit this on first delivery of a key.
A queued create is now registered like a queued update: a later copy of the id
takes the update path over the queued create, which runs after the batch's
creates, so the final row is exactly what the equivalent sequence of `upsertById`
calls produces — the later copy's props merged over the created row, one version
bump per real write, and the created row's validity lower bound. With
`coalesceUnchangedUpserts` enabled, a value-identical second copy of a new id now
coalesces against the queued create instead of writing a second time. Nodes and
edges are both fixed; for edges, as for an id that already existed, a later
copy's `from` / `to` are ignored because an update never repoints an edge.
Two smaller consequences of routing every queued write through the same state: a
repeated id whose dirty check rejected an earlier item's props no longer reports
the wrong error, and no later copy can coalesce against a stale prefetched row
after an earlier item queued a write.
- [#362](https://github.com/nicia-ai/typegraph/pull/362) [`9982960`](https://github.com/nicia-ai/typegraph/commit/9982960e66343c6980b8d6e87e0cb159a981e72a) Thanks [@pdlug](https://github.com/pdlug)! - Translate PostgreSQL read-only and missing-`TEMP` failures during graph
analytics into `UnsupportedBackendCapabilityError`, preserving the driver error
as the cause. Both refusal points are covered: a standby that rejects the
read-write working-table transaction, and a role that cannot create the
temporary table inside it.
- [#420](https://github.com/nicia-ai/typegraph/pull/420) [`d82fdaf`](https://github.com/nicia-ai/typegraph/commit/d82fdaf369109fe791be00ffe330351fb8aa4d00) Thanks [@pdlug](https://github.com/pdlug)! - Harden two failures at the operations/backend boundary: a create the engine
refuses now reports the condition it actually hit, and the last UPDATE path that
could store an inverted valid-time window no longer can.
**A create refused by the engine reports "already exists", not a raw driver
error.** A create learns an id is taken either from its own existence probe or
from the engine refusing the INSERT, and the second used to escape as a
`DrizzleQueryError` whose `.message` is the raw INSERT text. One condition
therefore surfaced as a typed user error down one path and an opaque system error
down the other, and callers could not branch on it at all. The engine's report is
now classified structurally and both routes raise the same `ValidationError`, on
the single and batch create paths for nodes and edges alike.
Two things reach the engine's path. A NODE create probes first, but the probe and
the INSERT are two statements and PostgreSQL's default READ COMMITTED does not
serialize the two write transactions, so a concurrent create of the same new id
can commit in between — the issue's reproduction. An EDGE create has no existence
probe at all, so the engine's refusal is its only report of a taken id, on every
backend and with no race involved.
Classification is structural, never message text: SQLSTATE 23505 plus the
PostgreSQL protocol's own constraint and relation fields, and SQLite's extended
result code, which distinguishes a primary-key duplicate (1555) from any other
unique-index duplicate (2067) in the code itself.
Every such refusal, from either route, now carries the new exported issue code
`ENTITY_ALREADY_EXISTS`, so a caller can recognize it without matching on the
message. `details.entityType` and `details.kind` say what was refused;
`details.id` names the taken id, and is absent only when the refused statement
inserted more than one row, because the engine reports that the statement
collided without saying which row did. No race is needed to reach that: a bulk
create of edges, whose ids the caller supplied and which nothing probes, is
refused this way on every backend.
The classification is scoped to the primary key on purpose. A `unique: true`
index declaration materializes a UNIQUE INDEX on the same relation, and violating
that is a declared-uniqueness failure about the row's VALUES rather than a
duplicate identity — PostgreSQL reports it under the index's own name and SQLite
under a different extended code, so it never matches and is unaffected. Neither is
a declared `unique` constraint conflict, which still raises `UniquenessError`.
SQLite never reached the node race: `BEGIN IMMEDIATE` gives the writer slot to
one transaction at a time, so a second create cannot sit between its probe and its
INSERT while the first commits. Its probe is authoritative there, and the refusal
was already the typed error — it now carries the code too. A duplicate EDGE id on
SQLite did surface as a raw `SqliteError`, and now raises the same error as it
does on PostgreSQL.
**A node resurrection stores the bound its window guard measured against.** A
resurrection rewrites `valid_from` rather than retaining it, so the guard that
refuses inverted windows has no stored bound to check and used the write instant
instead — sampled in the operations layer, while the backend went on to stamp its
own, strictly later, sample. A `validTo` at the guard's instant passed as
zero-width and committed as NEGATIVE width a millisecond later, the exact shape
the previous release exists to refuse. The operations layer now passes the
instant it validated against explicitly, so the bound that is checked is the
bound that is stored. Stating both endpoints is unaffected; the only change to a
successful write is that a resurrection's `valid_from` is the operations layer's
instant rather than the backend's — sampled a moment earlier inside the same
locked write, before the uniqueness entries it re-checks and re-inserts.
Edge resurrection was never exposed: an edge RETAINS its stored `valid_from`
unless the write names a new one, so its guard measures against a value already
on disk and predicts nothing.
- [#403](https://github.com/nicia-ai/typegraph/pull/403) [`c0279fc`](https://github.com/nicia-ai/typegraph/commit/c0279fc9cafac91b7a6ec31c343bd8fbdd12d743) Thanks [@pdlug](https://github.com/pdlug)! - graph-merge: credit the branch that authored a merged row's end-of-validity
Ending a row's validity is authored state, but a branch whose only change to a
row was its window could contribute the instant the merge committed and still be
absent from the merge's provenance. An identity is staged once, so a branch that
merely moved an inherited edge's window had its staged copy skipped whenever
another branch's property edit already staged that edge — and the provenance for
edges is derived from the staged copies. Nodes were worse: a window change had no
provenance path at all, so a window-only node ending was credited to nobody even
when no other branch touched the row.
The credit now comes from the window resolution itself, which is the only phase
that knows whose claim was committed. It credits exactly the branches whose claim
IS the resolved end — a claim that lost the least-claim rule contributed nothing
to the committed row, and remains visible in `MergeReport.validityEnds` under
`claimedBy`. A branch that both edited a row's properties and moved its window
stays one contribution.
The staged copy that carries a window-only ending is no longer credited for
carrying it: that copy exists only to give the ending a row to write, its
properties are the base's, and the branch holding it is whichever sorted first —
possibly one whose claim the merge discarded. Which branch carries the row is
left exactly as it was, because that branch also labels the base's properties in
the repoint fold's property union, where a relabelled contribution can change
which value a fold commits. Merge outcomes are unchanged; only the provenance is.
## 0.45.0
### Minor Changes
- [#355](https://github.com/nicia-ai/typegraph/pull/355) [`2882b23`](https://github.com/nicia-ai/typegraph/commit/2882b23ef6daed20041ccea56e3cfa76a8435c7a) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.nodes..updateWhere()` for typed, transactional set-based node
updates selected by property and independent relationship predicates. The
operation validates complete after-images and atomically maintains uniqueness,
fulltext, vector, history, and revision state on SQLite and PostgreSQL. Its
cross-backend storage primitive returns every updated after-image and provides
bind-budgeted, graph- and concrete-kind-scoped uniqueness cleanup so rebuilding
reservations cannot clear same-id nodes of another kind.
- [#352](https://github.com/nicia-ai/typegraph/pull/352) [`872d196`](https://github.com/nicia-ai/typegraph/commit/872d1960b470352cdcf5412a1303e59b776ce1ed) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.repairContributions()`, a privileged, idempotent repair pass for
strategy-owned contribution storage. It re-audits declarations from the active
persisted graph, non-destructively retries `missing-marker` and
`failed-materialization` findings, reports `stale` and `orphaned-marker` as
`requires-rebuild`, and returns a fresh post-repair diagnostic result. Repair
targets remain backend-owned so callers do not need access to TypeGraph-managed
tables, physical names, or DDL.
- [#353](https://github.com/nicia-ai/typegraph/pull/353) [`c225605`](https://github.com/nicia-ai/typegraph/commit/c22560576e2e22f5808ead72e552ead4b7f8743c) Thanks [@pdlug](https://github.com/pdlug)! - Add Store-level heterogeneous bulk edge reads that keep database round trips independent of schema breadth.
### Patch Changes
- [#351](https://github.com/nicia-ai/typegraph/pull/351) [`ff8e428`](https://github.com/nicia-ai/typegraph/commit/ff8e4280456b21985033977ec9be553bad06d63c) Thanks [@pdlug](https://github.com/pdlug)! - Ensure the kind-removal status table before `evolve()` checks it, so databases
created before TypeGraph 0.44 can evolve without manual backend initialization.
Concurrent PostgreSQL focused-table ensures also retry the catalog uniqueness
race that `CREATE TABLE IF NOT EXISTS` can surface during replica startup.
- [#350](https://github.com/nicia-ai/typegraph/pull/350) [`f752543`](https://github.com/nicia-ai/typegraph/commit/f752543749ac6914592e5f765f6e153b69e72518) Thanks [@pdlug](https://github.com/pdlug)! - Clarify that schema-managed Stores are immutable schema snapshots. After
`evolve()` changes the schema, callers must use the returned Store or the
updated `StoreRef.current` for subsequent work; a previously captured Store is
not mutated and its managed writes are rejected by the schema-version fence.
Document how long-lived caches detect schema commits from other processes with
`getCommittedSchemaVersion()` and refresh through a verified Store open.
Correct the `StoreRef` contract to say that the replacement is installed before
a successful schema-changing call resolves, rather than claiming that the
in-memory ref update is atomic with the persisted schema commit.
## 0.44.0
### Minor Changes
- [#331](https://github.com/nicia-ai/typegraph/pull/331) [`a1f1fde`](https://github.com/nicia-ai/typegraph/commit/a1f1fdec1be98dcbf243db738fb43cf634cae278) Thanks [@pdlug](https://github.com/pdlug)! - Add batched multi-source edge reads: `bulkFindFrom` / `bulkFindTo`
`EdgeCollection` could only read the edges of ONE endpoint at a time, so
rendering a page of N nodes with their relationships cost N statements. The new
`store.edges..bulkFindFrom(froms, options?)` and `bulkFindTo(tos,
options?)` read a whole SET of endpoints in set-oriented statements per
endpoint kind and bind-budget chunk,
returning the edges grouped per input (index `i` holds the edges of input `i`,
empty array when an endpoint has none).
This widens the predicate rather than batching the calls: `from_id = ?` becomes
`from_id IN (...)`, the same prefix seek on the edge relation's system index.
Temporal semantics are identical to `findFrom` / `findTo` — same default mode,
same `temporalMode` / `asOf` options, same soft-delete filtering, same
per-endpoint ordering — and a `StoreView` exposes both methods pinned to its
coordinate. Pass `limitPerInput` to bound each endpoint's fan-out (applied in
SQL via `ROW_NUMBER()` where the backend supports window functions). Inputs
larger than the backend's bound-parameter budget are split across statements
transparently.
Backend authors: this adds a new **optional** `GraphBackend` operation,
`findEdgesByEndpointSet(params)`, with its own `FindEdgesByEndpointSetParams`.
`FindEdgesByKindParams` is unchanged.
It is a separate operation rather than optional fields on `findEdgesByKind` so
that a backend which does not implement it cannot degrade silently. Optional
params would have left an existing backend type-correct while it ignored the id
list and returned every edge of the kind — which the collection would rebucket
into a correct-looking answer at unbounded cost. Support is now detected by the
method's presence, before any read is issued, and `bulkFindFrom` / `bulkFindTo`
refuse with a typed `ConfigurationError` on a backend without it rather than
looping `findFrom` per input.
The parameter shape also makes the previously-validated illegal states
unrepresentable: one `side` instead of two id lists, no scalar `fromId` /
`toId` to disagree with a set, and no `limit` / `offset` / `after` to slice a
read the backend splits into bind-budget chunks.
- [#334](https://github.com/nicia-ai/typegraph/pull/334) [`7a2e16b`](https://github.com/nicia-ai/typegraph/commit/7a2e16b70685609b19cda6683dc1d545d5aa5f9a) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.verifyContributions()`, an owner-agnostic diagnostic that crosses
each contribution currently expected by the active graph and backend strategies
against its durable marker and the physical catalog. Nothing on the open path
probes the catalog — boot and the runtime asserts short-circuit on a per-instance
signature cache and then on the marker row alone — so a database whose
strategy-owned tables were dropped out of band opened completely clean and
failed at the first fulltext or vector read. The diagnostic reports detected
problems as `orphaned-marker` (marker records a success, table absent),
`missing-marker` (table present, nothing attests it), `failed-materialization`
(the marker records a failed attempt and no table was produced — marker and
catalog agree, and it is broken anyway), or `stale` (marker recorded at a
different shape), with the `owner` / `logicalName` /
`physicalName` and, for vector slots, the `kind` and `fieldPath` needed to route
to the state-specific repair without reconstructing internal marker strings.
For vector slots, `missing-marker` and `failed-materialization` use the
non-destructive forced ensure; only `orphaned-marker` and `stale` rebuild vector
storage with `store.reembedVectorField`. `lastError` carries the reason the
marker recorded, when it recorded one: `state` says which repair to run,
`lastError` says why it broke. A contribution with neither a marker nor a table
was never attempted and is omitted, as are retired markers and unsupported
vector slots, so an empty result is not proof of initialization. It is read-only
(one existence query per contribution table, no DDL, no writes) and deliberately
not a boot step; the fast-path caching stays the default. Backends that cannot
probe their own catalog throw `ConfigurationError` rather than reporting a clean
bill of health.
- [#335](https://github.com/nicia-ai/typegraph/pull/335) [`7950bb0`](https://github.com/nicia-ai/typegraph/commit/7950bb0de0b2cb93fd71c4cf6644df0b65126b6b) Thanks [@pdlug](https://github.com/pdlug)! - Support list-valued parameters in `in()` / `notIn()`
`field.in(param("ids"))` now binds the whole list at
`.prepare().execute({ ids: [...] })`, so the canonical "fetch these ids"
query can finally be prepared. The list rides on a single bound parameter
that the dialect unpacks (`json_each` on SQLite, `jsonb_array_elements_text`
on PostgreSQL), which keeps arity out of the SQL text: one compiled statement
serves every list length, and a list of any size costs one bound parameter
instead of one per element. An empty list is valid — `in([])` matches nothing,
`notIn([])` matches everything.
A `ParameterRef` passed among the _elements_ of a literal list
(`in(["a", param("b")])`) was previously coerced to a literal and silently
produced wrong results. It now throws `UnsupportedPredicateError` naming the
supported form. A name used both as a list and as a scalar in one query is
rejected at `prepare()`.
List elements are validated against the field's type before binding, so
`[1, "a"]` against a number field is rejected with a `ConfigurationError`
rather than failing on PostgreSQL and silently matching nothing on SQLite.
This matches the literal form, which already refuses a mixed list.
Non-finite numbers (`NaN`, `±Infinity`) are now rejected in any parameter
binding, list or scalar. `JSON.stringify` turns them into `null`, so a list
binding became SQL NULL — `notIn(param("x"))` with `[NaN]` filtered out every
row — and SQLite binds a scalar `NaN` as NULL, so `eq(param("x"))` with `NaN`
quietly matched nothing. Both now throw.
`DialectAdapter` gains two members, `inListParameter` and `packListValue`;
custom dialect adapters must implement them.
- [#329](https://github.com/nicia-ai/typegraph/pull/329) [`03e87bd`](https://github.com/nicia-ai/typegraph/commit/03e87bd1d370408af7dcb640729e9579508db29f) Thanks [@pdlug](https://github.com/pdlug)! - Export the committed-schema reads from the package root. `getActiveSchema`,
`isSchemaInitialized`, and the `SerializedSchema` type now sit next to
`getCommittedSchemaVersion` in `@nicia-ai/typegraph`, so answering "what kinds
does this database already have?" no longer requires finding the
`@nicia-ai/typegraph/schema` subpath or querying `typegraph_schema_versions`
by hand. `getActiveSchema` and `getCommittedSchemaVersion` now cross-reference
each other in their docstrings.
- [#332](https://github.com/nicia-ai/typegraph/pull/332) [`2fb8925`](https://github.com/nicia-ai/typegraph/commit/2fb89258b2047fe8ae6d2ce97cb7f219200285e5) Thanks [@pdlug](https://github.com/pdlug)! - Fix `migrateSchema()` silently dropping runtime-committed kinds
`migrateSchema(backend, graph, currentVersion)` committed `graph` verbatim. It
did not fold the persisted graph extension, so kinds committed at runtime by
`Store.evolve()` — which live in `schema_doc.extension`, not in the
compile-time graph — were erased from the active schema document while their
rows stayed in `typegraph_nodes` / `typegraph_edges`, reachable by nothing.
The persisted `deprecatedKinds` set was erased the same way.
This was reachable by following the library's own advice: the `MigrationError`
raised for a breaking change told callers to "use `getSchemaChanges()` to
review, then `migrateSchema()` to apply", and doing so with the graph they
passed to `createStoreWithSchema` destroyed every `evolve()`-committed kind.
Two changes:
- **The persisted graph extension (and deprecated-kind set) is now folded in**,
exactly as `createStoreWithSchema` and `getSchemaChanges` already did.
`migrateSchema` was the last commit path that did not. Callers pass the graph
they have; runtime-committed kinds survive.
- **A commit that would drop a kind still holding rows is refused** with a
`MigrationError` whose `details.reason` is the new `"kind-removal"`
discriminant and whose `details.droppedKinds` names them. Pass
`{ discardDroppedKindRows: true }` if losing those rows is the intent — the
name says what the flag does, because the next reconcile deletes them.
The guard fires on the actual harm — rows the next reconcile would delete —
not on kind removal as such. Dropping an _empty_ kind is unaffected, so the documented three-deploy
removal flow (stop writing → delete the rows → drop from `defineGraph()` and
migrate) still works exactly as written; Deploy 2 is now what makes Deploy 3
legal instead of being merely advisory. Live rows only, matching the
`excludeDeleted` default of the equivalent probe in `Store.evolve()`.
Breaking property changes — the documented reason to reach for
`migrateSchema()` — are unaffected.
`MaterializeRemovalsEntry` gains a `"skipped"` variant, carrying
`reason: "kind-is-live"`. `materializeRemovals()` returns it when a queued
removal names a kind the active schema declares again, so the decline is
reported rather than leaving the queue at a non-zero depth with nothing
explaining why. Consumers that switch exhaustively on `status` must handle it.
The type is now a discriminated union, so `"failed"` carries a required
`error` and `"skipped"` a required `reason`.
`Store.evolve()` refuses to re-add a kind whose data cleanup is still pending,
with a `ConfigurationError` naming the kind and pointing at
`materializeRemovals()`. Reads filter only by `(graph_id, kind)`, so re-adding
before cleanup made the previous incarnation's rows visible alongside the new
ones — and the cleanup was then declined because the kind was live, so they
were never reclaimed. The documented cycle (remove → `materializeRemovals` →
re-add) is unaffected.
Two further corrections found while reviewing the above:
- **A stale store can no longer resurrect a removed kind.** The fold now
strips the supplied graph's own extension slice before applying the
persisted one, so the committed document is a function of the database
alone. Previously `migrateSchema(backend, store.graph, v)` — `store.graph`
is public and returns the merged graph — unioned a stale slice back in and
silently undid `Store.removeKinds()`, leaving a kind the schema called live
while its `typegraph_kind_removals` row stayed queued for a later
hard-delete. `Store.#catchUpToStored` has stripped for this exact reason;
the schema layer now matches it.
- **`discardDroppedKindRows`'s documentation was wrong.** It claimed the dropped
kind's rows stay and that `materializeRemovals` "will never clean them up".
`materializeRemovals` re-derives removals by walking schema-version history,
so the next reconcile hard-deletes them regardless. The flag buys a
committed schema, not retained data; the docstring now says so and points
callers at copying the rows out first.
### Patch Changes
- [#328](https://github.com/nicia-ai/typegraph/pull/328) [`dc2a386`](https://github.com/nicia-ai/typegraph/commit/dc2a386fa2b6275b7d0f3d6d80e2959ea094365b) Thanks [@pdlug](https://github.com/pdlug)! - State `store.batch()`'s real cost where callers see it. `batch()` runs its
queries in sequence, keeping at most one in flight — at least one statement
each, and two for a query whose selective-field mapping falls back after its
statement has already executed. So it caps concurrency at best and will not fix
an N+1. It is also not a snapshot: PostgreSQL's default read-committed isolation
lets a later query in the batch observe a commit the earlier ones did not, and
there is no public way to get one across fluent queries, since a transaction
context exposes no query builder.
The docstrings for `batch()`, `BatchableQuery`, `executeOn`, and the edge
`batchFind*` methods now lead with that, and point at the set-oriented and
chunked alternatives, described by what they actually do: `.traverse()` compiles
a chain to one statement, `store.subgraph()` costs 2 statements on SQLite and 3
on PostgreSQL, `getByIds()` issues one statement per bind-limit chunk (falling
back to concurrent per-id lookups where the backend exposes no batch read), and
`bulkFindByIndex()` costs a probe plus that same chunked hydration.
The docs site is corrected to match, including claims that `batch()` "minimizes
round-trips for reads", that `batchFind*` collapses N reads into "a single
transactional round-trip", that `subgraph()` is a single statement, and that
`getByIds()` is a single query. Transaction support no longer implies a
transport shape anywhere: Durable Objects use an ambient transaction with no
framing statements, and the non-transactional path may still reuse one client.
The changelog entry that shipped `batch()` carries a correction note rather than
a silent rewrite.
Execution semantics are unchanged. One public diagnostic changes: the
`ConfigurationError` message for a batch endpoint read on a read-only
`StoreView` no longer calls `batch()` a "batch loader".
- [#344](https://github.com/nicia-ai/typegraph/pull/344) [`ea05d0d`](https://github.com/nicia-ai/typegraph/commit/ea05d0da9f4956062b895138be03ad3b36fa289b) Thanks [@pdlug](https://github.com/pdlug)! - Fence deferred kind cleanup against concurrent schema re-adds. Removal now
rechecks the active schema and atomically deletes live rows, recorded-time
intervals, vector storage, and contribution markers under the schema lock.
Custom backends that implement the optional `schemaWriteTransaction` capability
must expose transaction-bound statement execution, table-existence probing,
schema DDL, and vector-contribution marker deletion on its callback target.
- [#347](https://github.com/nicia-ai/typegraph/pull/347) [`1616e93`](https://github.com/nicia-ai/typegraph/commit/1616e9380de834afa4912c91b79e45ed8edcd122) Thanks [@pdlug](https://github.com/pdlug)! - Fence schema-version commits against concurrent schema-managed Store writes.
SQLite uses its immediate writer transaction; PostgreSQL locks the active schema
row in shared mode for managed writes and exclusive mode for schema commits.
Managed writes revalidate their Store schema version while holding the fence, so
stale queued writes fail instead of landing against a schema that no longer
accepts them. Snapshot-isolated PostgreSQL transactions may raise the database's
native serialization failure; callers retry the whole transaction, and graph
merge does so automatically. Schema-managed Stores on non-transactional or
custom backends without the fence now fail closed on writes. Raw `createStore()`
instances, direct backend writes, and Stores whose schema metadata was reset by
`clear()` remain outside the versioned guarantee.
- [#342](https://github.com/nicia-ai/typegraph/pull/342) [`d481054`](https://github.com/nicia-ai/typegraph/commit/d481054ebbac7fd6ec06b8d0b5cfd28313efab89) Thanks [@pdlug](https://github.com/pdlug)! - Make the documented store query hooks fire for query-builder statements,
including prepared queries, batched queries, and selective-projection retries.
Each submitted statement now reports its SQL, parameters, row count, duration,
and failures through the existing `StoreHooks` callbacks.
- [#343](https://github.com/nicia-ai/typegraph/pull/343) [`347d5e3`](https://github.com/nicia-ai/typegraph/commit/347d5e3ba1c634697b873044a9770436a119604b) Thanks [@pdlug](https://github.com/pdlug)! - Avoid repeated selective-projection fallback queries. Smart-select planning now
covers common high-value threshold branches, and prepared queries remember a
missing-field fallback so later executions fetch the full row directly.
## 0.43.0
### Minor Changes
- [#320](https://github.com/nicia-ai/typegraph/pull/320) [`010132a`](https://github.com/nicia-ai/typegraph/commit/010132a6ce2625b83f6256ef78bbc9bbd78867ee) Thanks [@pdlug](https://github.com/pdlug)! - Allow idempotent endpoint-based edge writes to set application-time validity.
`getOrCreateByEndpoints` now accepts `validFrom` and `validTo`, while
`bulkGetOrCreateByEndpoints` accepts them per item. Creation applies both
fields, updates and resurrections apply `validTo`, and pure found results leave
the existing window unchanged.
## 0.42.1
### Patch Changes
- [#317](https://github.com/nicia-ai/typegraph/pull/317) [`8024711`](https://github.com/nicia-ai/typegraph/commit/80247111bde9282dbbe1a9ef3c31ca66bd16ae39) Thanks [@pdlug](https://github.com/pdlug)! - Prevent large PGlite bulk writes from silently leaving the connection unable
to return rows. PGlite backends now advertise their safe 32,767-parameter
limit, PostgreSQL batch sizes follow the active backend capability, and
over-budget statements fail before driver dispatch.
## 0.42.0
### Minor Changes
- [#313](https://github.com/nicia-ai/typegraph/pull/313) [`a797a8b`](https://github.com/nicia-ai/typegraph/commit/a797a8b1ebe077869e42a334816b172314cb0132) Thanks [@pdlug](https://github.com/pdlug)! - Expose the schema-commit surface's decisions as data instead of prose, so
callers can pre-flight a proposal and classify a failure without matching
message text.
`MigrationError` now carries a stable `details.reason` discriminant —
`"schema-behind" | "breaking-change" | "no-active-version" |
"version-not-found"` (exported as `MIGRATION_FAILURE_REASONS`) — plus the
structured `details.diff` for the outcomes that computed one. Branch on
`details.diff.hasBreakingChanges` to tell an additive change from an
incompatible one, with no re-query and no substring matching. Note that
`MigrationErrorDetails.reason` is now required rather than an optional free-text
string.
For pre-flight, `classifySchemaChanges(diff)` reduces a diff to
`"identical" | "additive" | "incompatible"`, and the existing SELECT-only
`getSchemaChanges` is now reachable from a store handle: `store.schemaChanges()`
returns the diff and `store.requiresMigration()` answers the boolean predicate
(also `true` when nothing has been committed yet). A least-privilege runtime can
detect that it needs the privileged bootstrap instead of discovering the
migration wall partway through a request.
Documents two operational facts that were previously invisible at the call site:
kinds are scoped to the `graph_id` (a namespace _is_ a graph id — separate
declaration sites do not isolate kinds), and running many `graph_id`s with
divergent schemas in one database is a supported multi-tenant pattern, including
the one cross-graph coupling (SQL index names are database-global, so identical
kind+index shapes share a physical index and divergent shapes fail loudly).
Also fixes `getSchemaChanges` to fold in the persisted graph-extension before
diffing, matching what the commit path already does. Without it a compile-time
graph was compared against a stored schema that also contains runtime-committed
kinds, so those kinds read as removals and an unchanged schema was reported as
requiring a breaking migration.
### Patch Changes
- [#313](https://github.com/nicia-ai/typegraph/pull/313) [`a797a8b`](https://github.com/nicia-ai/typegraph/commit/a797a8b1ebe077869e42a334816b172314cb0132) Thanks [@pdlug](https://github.com/pdlug)! - Stop reporting a reordered declaration as a schema change. Restating a kind
with its properties, enum members, or edge endpoints listed in a different
order is a semantic no-op, but the diff compared those arrays positionally and
reported the kind as `modified` — forcing callers into a privileged migration
for a schema that had not actually changed. A reordered `enum` was even
classified `breaking`, i.e. a pure reordering demanded a destructive-migration
decision.
`required`, `enum`, and edge `fromKinds` / `toKinds` are now compared as the
sets they are, in both the modified-vs-unmodified decision and the
breaking-change severity classification. Genuine changes — added or removed
properties, newly required properties, changed enum members, different edge
endpoints — are detected exactly as before.
The normalization is deliberately scoped to diff comparison and is **not**
applied to the canonical form behind `computeSchemaHash`, so no schema hash
already committed to a database changes.
The normalization walks the document as JSON Schema rather than as plain JSON,
because a key's meaning depends on where it appears. Recursion is an
**allowlist** of known schema-valued keywords; everything else is preserved
verbatim:
- Instance data (`default`, `const`, `examples`) and unknown extension keys —
Zod's `.meta()` merges arbitrary keys straight into the generated schema — are
compared verbatim. Recursing into them would sort a nested key merely _named_
`required`, silently normalizing away a real change to a stored value.
- Keys under `properties`, `patternProperties`, `dependentSchemas`, `$defs`, and
`definitions` are user-chosen field names, not keywords, so a field _named_
`default` still has its subschema normalized like any other.
- `dependentRequired` maps a name to a set of names, so each set is
order-normalized.
The allowlist fails in the safe direction: an unrecognized schema-valued keyword
is left unsorted, so a reordering inside it reads as a change rather than being
hidden.
## 0.41.0
### Minor Changes
- [#311](https://github.com/nicia-ai/typegraph/pull/311) [`008fa20`](https://github.com/nicia-ai/typegraph/commit/008fa2008d199f997f45af967ba2f1a1fbad4970) Thanks [@pdlug](https://github.com/pdlug)! - Make verified adapter stores reusable across connections so serverless/edge
deployments that open a fresh database connection per request can verify once
per isolate instead of paying a schema-reconcile round-trip on every request.
`AdapterStore` now exposes `reconciledSchema`, an opaque snapshot of a store's
reconciled (compile-time + runtime-committed) graph and committed schema
version. Pass it to a synchronous `createAdapterStore(graph, backend,
{ reconciled })` — which issues **zero** database queries and still validates
reads and writes against runtime-committed kinds — or call
`store.withBackend(freshBackend)` to rebind an already-verified store onto a new
connection with no re-verify (the store's connection is captured immutably, so
this returns a new equivalent store rather than mutating in place). The new
`getCommittedSchemaVersion(backend, graphId)` reads the committed version with a
single indexed SELECT, the cheap cross-isolate probe for detecting when another
process committed a schema change and the cached snapshot must be refreshed.
## 0.40.0
### Minor Changes
- [#308](https://github.com/nicia-ai/typegraph/pull/308) [`db2dc31`](https://github.com/nicia-ai/typegraph/commit/db2dc31e2a66f3195c0d5d7e3df19864cb64672c) Thanks [@pdlug](https://github.com/pdlug)! - Replace timestamp-only `RecordedInstant` values with versioned anchors that
encode a strict per-graph logical revision alongside a non-decreasing physical
wall-time high-water mark. Recorded relations store numeric revisions while the
public anchor remains one durable string. Upgrade timestamp-only preview tables
with `migrateLegacyRecordedTime()` and remap external checkpoints with
`migrateRecordedAnchor()`. Driver timestamps are normalized without host-local
timezone parsing, migration integrity failures are typed, and the retained
anchor map can be dropped automatically after its final graph is cleaned up.
History-enabled async store factories now reject an unmigrated recorded schema
at open, including when the legacy tables are empty.
## 0.39.0
### Minor Changes
- [#306](https://github.com/nicia-ai/typegraph/pull/306) [`cd4e0eb`](https://github.com/nicia-ai/typegraph/commit/cd4e0ebf1fa8a51e8a965e667801941a5360e097) Thanks [@pdlug](https://github.com/pdlug)! - Add bounded, deterministic `scan()` pagination to recorded-time node and edge collections so adapters can reconstruct complete historical snapshots without retaining a separate identity inventory.
- [#305](https://github.com/nicia-ai/typegraph/pull/305) [`4349766`](https://github.com/nicia-ai/typegraph/commit/4349766fc9d5db7046f05cefab6e56f4dd4d655a) Thanks [@pdlug](https://github.com/pdlug)! - Harden adapter capability surfaces and document their migrations.
This is source-breaking for adapter code that reads `tx.sql` without first
narrowing `tx.sqlAvailability === "available"`: non-available union arms now
omit `sql` instead of exposing it as an optional `never`/`undefined` property.
The runtime history and revision-tracking guards remain fail-loud for JavaScript
and type-suppressed callers.
Add `openProvenanceStore(targetStore)` as the preferred graph-merge provenance
API while retaining `openProvenanceStore(backend, targetGraphId)` for standalone
inspection tools. On Cloudflare D1 and Durable Object SQLite, ignore only a
recognized `SQLITE_AUTH` rejection of the performance-only `analysis_limit`
PRAGMA and continue with scoped `ANALYZE`; unexpected maintenance failures stay
visible.
## 0.38.0
### Minor Changes
- [#297](https://github.com/nicia-ai/typegraph/pull/297) [`474afe6`](https://github.com/nicia-ai/typegraph/commit/474afe602f44ef60e313015d853c4705b30e8790) Thanks [@pdlug](https://github.com/pdlug)! - Add global and weighted multi-seed personalized PageRank with induced-subgraph,
temporal-view, direction, convergence-tolerance, and working-memory options.
- [#303](https://github.com/nicia-ai/typegraph/pull/303) [`58855f7`](https://github.com/nicia-ai/typegraph/commit/58855f7b9e4f808805aecb4d1169fadd8e6aaab1) Thanks [@pdlug](https://github.com/pdlug)! - Move query compilation behind TypeGraph-owned backend and SQL-fragment
abstractions so strict consumers no longer typecheck unused Drizzle dialect
declarations. Add a Drizzle-free `core` entrypoint and managed full-Store
entrypoints for local SQLite and PGlite, with packed TypeScript 5 and 6
regression coverage for both databases.
The portable `@nicia-ai/typegraph/indexes` entrypoint is also Drizzle-free.
Direct Drizzle index-builder helpers moved to
`@nicia-ai/typegraph/adapters/drizzle/indexes`.
Advanced adapter APIs now use TypeGraph's `SqlFragment` instead of Drizzle
`SQL`: this includes query `compile()` results, custom `GraphBackend`
implementations, and custom fulltext/vector strategies. Use `toSQL()` for a
dialect-rendered `{ sql, params }` result, or `renderSqlite()` /
`renderPostgres()` when rendering a fragment directly.
Custom backend and strategy authors can import the complete, Drizzle-free
contract vocabulary from `@nicia-ai/typegraph/backend`. This entrypoint names
every operation parameter, row, strategy payload, dialect port, SQL fragment
chunk, and supporting schema/index type referenced by those contracts. API
Extractor enforces zero forgotten exports for new entrypoints and fingerprints
the complete pre-existing debt set so added or removed leaks cannot pass
silently.
The default `Store` is now the portable TypeGraph surface. It keeps the full
graph API and graph-owned transactions while omitting adapter-native handles and
caller-owned transaction adoption. Drizzle integration entrypoints return
`AdapterStore` when precisely typed `tx.sql`,
`withTransaction`, or `withRecordedTransaction` interoperability is required.
`createStore`, `createStoreWithSchema`, and `createVerifiedStore` now return the
portable contract; use their `createAdapterStore`,
`createAdapterStoreWithSchema`, and `createVerifiedAdapterStore` counterparts
when the application deliberately needs adapter-native interoperability.
Migration map:
| 0.37 use | 0.38 replacement |
| ------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------- |
| `createStoreWithSchema(...)` followed by `tx.sql` | `createAdapterStoreWithSchema(...)` |
| `Store` | `AdapterStore` |
| `TransactionContext` | `AdapterTransactionContext` |
| `HistoryTransactionContext` | `TransactionContext` for portable history stores, or `AdapterHistoryTransactionContext` for adapter history stores |
| `AdoptedTransaction` | The adapter's concrete native-handle type, such as `AnySqliteDatabase` or `AnyPgTransaction`, passed as the adapter generic |
`HistoryTransactionContext` and `MeasurableHistoryTransactionContext` were
removed rather than retained as aliases. Portable stores have one SQL-free
transaction context across live and history modes. Adapter contexts expose
`sql` only after `sqlAvailability === "available"`; the other discriminated
union arms omit the property. Store evolution is now generic over the exact
Store flavor, so adapter/history/recorded-read surfaces and compatible
`StoreRef` values are preserved without downstream casts. Remove casts that
existed only to restore the old widened evolution result.
Requiring the `sqlAvailability` check is source-breaking for adapter code that
previously read `tx.sql` from the unnarrowed union. Narrow on the discriminant
before passing the handle to even an `unknown`-typed sink.
`GraphBackend` is now the portable TypeGraph backend port. Native transaction
adoption lives on `AdapterBackend`, so a capability-less
backend cannot be passed to an adapter-store factory. Portable transaction
contexts and adapter transaction contexts both expose the same runtime-enforced
read-only `TransactionReadBackend`. Adapter contexts add only the precisely
typed native `sql` handle; TypeGraph internals reach the full transaction
backend through a non-public, non-enumerable runtime port. Backend functions
are now receiver-free (`this: void`); custom backends must close over their
state instead of depending on method receivers.
PostgreSQL adapter stores now expose and accept `AnyPgTransaction` for native
transaction interoperability. A root Drizzle PostgreSQL database is rejected
at compile time and runtime; pass only the transaction handle received by a
caller-owned `db.transaction(...)`. SQLite adoption remains database-handle
based so its documented manual-`BEGIN` integration continues to work.
Public `TransactionOptions` now contains only caller-selectable isolation and
access modes. TypeGraph's temporary-write authorization is an internal,
globally branded capability and is no longer expressible through the public
transaction contract. Fulltext and vector strategy members are readonly
function properties, closing TypeScript's method-bivariance loophole for
third-party implementations. Dialect adapter members use the same receiver-free
function-property contract, and `TransactionOptions` is exported from the root
entrypoint for portable transaction consumers.
Managed SQLite and PGlite factories preserve the precise live, history, or
recorded-read Store flavor selected by their options, including when options
are widened before the call. This keeps unavailable write and native-adapter
capabilities unrepresentable instead of relying on runtime failures.
Every Store flavor exposes the safe, Drizzle-free `store.capabilities`
descriptor for runtime feature checks without exposing backend operations.
`AdapterHistoryStore.backend` exposes the narrower `HistoryStoreBackend`, which
omits raw SQL, native import, graph clearing, and nested backend transactions so
capture-bypassing writes are absent at both type and runtime levels.
Backend capability narrowing now uses an exhaustive runtime allowlist instead
of default-forwarding proxy overlays. New `GraphBackend` members must be
classified explicitly, preventing adapter capabilities from leaking through a
history wrapper. Store evolution also preserves each refined Store flavor and
accepts invariant `StoreRef` values for that exact replacement surface.
Add checked-in API Extractor reports derived from every package export. CI now
fails when the emitted public declaration surface changes without an
intentional report update.
Direct SQL fragment values now pass through the same dialect binding
normalization as placeholders and compiled queries. Runtime store,
transaction, schema, and recorded-read ports use versioned global symbols so
mixed ESM/CJS or duplicated bundle instances interoperate safely. Dialect
policies outside the compiler are exhaustive records or switches, so adding a
new SQL dialect cannot silently inherit SQLite behavior.
Remove the transitional `SQL`, `SqlRenderDialect`, and `AdoptedTransaction`
aliases. Import `SqlFragment` and `SqlDialect` directly. The constructors that
brand arbitrary fragments as executable SQL are now internal; public compiled
SQL values come from TypeGraph's query compiler. Managed local stores now live
at `/sqlite/local` and `/postgres/pglite`; bring-your-own-connection APIs live
under `/adapters/drizzle/sqlite...` and `/adapters/drizzle/postgres...`.
The old `@nicia-ai/typegraph/sqlite` and
`@nicia-ai/typegraph/postgres` entrypoints were removed. They are not
compatibility aliases because those names now distinguish the managed Store
API from bring-your-own-connection adapters. Move imports as follows:
| 0.37 entrypoint | 0.38 entrypoint |
| ------------------------------------------------- | ----------------------------------- |
| `/sqlite` | `/adapters/drizzle/sqlite` |
| `/sqlite/local` for `createLocalSqliteBackend` | `/adapters/drizzle/sqlite/local` |
| `/sqlite/libsql` | `/adapters/drizzle/sqlite/libsql` |
| `/postgres` | `/adapters/drizzle/postgres` |
| `/postgres/pglite` for `createLocalPgliteBackend` | `/adapters/drizzle/postgres/pglite` |
The managed `/sqlite/local` and `/postgres/pglite` entrypoints keep Drizzle
out of their public declarations, but their built-in database implementation
still uses Drizzle internally. `drizzle-orm` therefore remains a required peer
of the 0.38 package; declaration isolation does not imply installation
isolation.
- [#298](https://github.com/nicia-ai/typegraph/pull/298) [`f178663`](https://github.com/nicia-ai/typegraph/commit/f1786634e7867b892805349dc922269e155d1d65) Thanks [@pdlug](https://github.com/pdlug)! - Add deterministic synchronous label propagation with exact binary tie-breaking,
induced node-kind scope, temporal and recorded-time views, bind-independent
neighbor voting, early period-two oscillation detection, and an
`onMaxIterations` completion contract: `"throw"` (default) returns only a
converged labeling, while `"return"` yields the exact fixed-round Graphalytics
CDLP labeling.
## 0.37.1
### Patch Changes
- [#292](https://github.com/nicia-ai/typegraph/pull/292) [`0152c3b`](https://github.com/nicia-ai/typegraph/commit/0152c3bf4a0bd3c047931041d1505b69a25fa05a) Thanks [@pdlug](https://github.com/pdlug)! - Restore graph algorithms on Cloudflare Durable Objects SQLite. The
auto-detected `do-sqlite` profile now marks temporary-table graph analytics as
unsupported, routes shortest-path and reachability algorithms through their
inline fallback, and rejects temporary-table-only algorithms with the existing
typed capability error instead of leaking workerd's `SQLITE_AUTH` failure.
- [#293](https://github.com/nicia-ai/typegraph/pull/293) [`9309ec3`](https://github.com/nicia-ai/typegraph/commit/9309ec3f839b53474d389b9376f7851c87753e28) Thanks [@pdlug](https://github.com/pdlug)! - Speed up exact weakly connected components with indexed changed-label
frontiers, changed-row-only writes, and one fewer working-table join. Preserve
synchronous convergence across bind-limited edge-kind chunks, and align
shortest-path identity tie-breaks with portable binary ordering.
## 0.37.0
### Minor Changes
- [#269](https://github.com/nicia-ai/typegraph/pull/269) [`92479d4`](https://github.com/nicia-ai/typegraph/commit/92479d44ba7f0fc76985d51174cb9801c055fd4d) Thanks [@pdlug](https://github.com/pdlug)! - Vector storage now rides the [#135](https://github.com/nicia-ai/typegraph/issues/135) durable-contribution machinery, so the
runtime never issues DDL on the embedding hot path.
Previously every vector op (`upsertEmbedding` / `deleteEmbedding` /
`vectorSearch` / `createVectorIndex`) lazily ran `CREATE TABLE IF NOT EXISTS`
for its per-`(kind, field)` table on whatever connection it executed on. On a
least-privilege Postgres role (USAGE on `public`, full DML, but no `CREATE`)
this failed with `permission denied for schema public` (SQLSTATE 42501) — even
when the table already existed, because Postgres runs the schema aclcheck before
the `IF NOT EXISTS` short-circuit. The fulltext path already avoided this via
durable markers; vectors now do too.
What changed:
- **Boot (privileged):** `createStoreWithSchema` provisions every embedding
`(kind, field)` table + a durable contribution marker, enumerated from the
graph. `evolve()` provisions any embedding fields it introduces. A slot
already provisioned at a _different_ shape (the declared dimension changed)
is warned about and left untouched — boot stays reachable so
`store.reembedVectorField()` can recreate it; until then, writes to that
field fail with a `stale` `StoreNotInitializedError` that points at
`reembedVectorField`.
- **Runtime writes (DML-only):** `upsertEmbedding` (single and batch) and
`deleteEmbedding` assert the durable marker with a cached, signature-checked
SELECT and run DML — never DDL. `createVerifiedStore` verifies vector markers
at attach, alongside fulltext.
- **Vector reads are not marker-gated:** `store.search.vector`,
`store.search.hybrid`, and query-builder `.similarTo()` predicates compile to
SQL against the per-field table directly (searches may override the metric at
query time, so their slot legitimately differs from the provisioned shape);
against an un-provisioned database they surface the engine's missing-relation
error, which `createVerifiedStore` catches at attach.
- `reembedVectorField` re-stamps the marker after recreating storage at a new
dimension; vector-field reclaim (`materializeRemovals`) clears the marker when
it drops a table.
**Breaking:** vector ops now require a prior privileged `createStoreWithSchema`
(exactly as fulltext already does). A plain `createStore` + embedding write with
no provisioning step throws `StoreNotInitializedError` instead of lazily
creating the table.
**Migration:** after upgrading, run `createStoreWithSchema(graph, adminBackend)`
once under the schema-owner role. It creates the per-field vector tables +
markers; least-privilege runtimes then assert markers (SELECT) and run vector
DML with zero DDL — no `GRANT CREATE` required.
Consumers that boot manually (raw DDL + the sync `createStore` attach +
`backend.ensureRuntimeContributions`) provision vectors the same way: the new
`resolveGraphVectorSlots(graph)` export enumerates every embedding
`(kind, field)` slot, and `backend.ensureVectorSlotContribution(slot)`
materializes each — the exact step `createStoreWithSchema` performs. Batch
counterparts (`backend.ensureVectorSlotContributions(slots)` /
`backend.assertVectorSlotsInitialized(slots)`) resolve every slot's markers
with one graph-scoped query — what boot and verified attach use, and the
right choice for many embedding fields over a remote connection.
- [#284](https://github.com/nicia-ai/typegraph/pull/284) [`26f5b4a`](https://github.com/nicia-ai/typegraph/commit/26f5b4a3f129353e0ec92040497ddca685563c7f) Thanks [@pdlug](https://github.com/pdlug)! - TypeGraph's base-relation indexes are now **system-index declarations** — a
single declared list (`SYSTEM_INDEX_DECLARATIONS`) that both dialect schemas
derive from and that materializes onto already-initialized databases.
Previously the base indexes were hand-written twice (once per dialect schema)
and applied only by first-boot bootstrap DDL, so an index added in a newer
library version never reached an existing database without manual DDL (the
gap [#282](https://github.com/nicia-ai/typegraph/issues/282) exposed). Now:
- **Single source, parity by construction.** `createSqliteTables` /
`createPostgresTables` build their node/edge/recorded-relation indexes from
the same declarations, and a cross-dialect extraction test asserts the two
generated DDL scripts' full index sets stay identical.
- **Upgrade path.** `createStoreWithSchema` brings a database's system
indexes up to the running library version at boot — `CREATE INDEX
CONCURRENTLY` on PostgreSQL, riding the same status table, drift
signatures, invalid-leftover healing, and cross-caller claim protocol as
graph-declared indexes. A database whose indexes all exist settles from
three concurrent catalog/status reads (scoped to the session
`search_path`, so schema-per-tenant databases never observe each other's
indexes) with no index DDL and no status writes — the only DDL on that
warm path is the idempotent status-table `CREATE TABLE IF NOT EXISTS`
ensure step every materialize verb runs. A system index that is
physically absent or invalid is rebuilt
even when a stale success row survives (dump/restore, manual drop).
Failures — including status-table infrastructure errors — degrade to a
warning: indexes are a performance concern and the store still boots.
Deployments that must not run index builds inline at boot pass
`systemIndexes: "skip"` to `createStoreWithSchema` and materialize
out-of-band.
- **New API: `store.materializeSystemIndexes()`** for deployments that boot
without `createStoreWithSchema` (zero-DDL attach) — call once under a
DDL-capable role after upgrading. Strict where the boot path is lenient:
throws `ConfigurationError` on backends without DDL/status primitives.
- `IndexEntity` gains a `"system"` member; system status rows carry the
relation key (e.g. `"recordedNodes"`) in their `kind` column.
Generated DDL is unchanged for default and short custom table names — same
index names, columns, and order — so existing databases and drizzle-kit
migrations are unaffected. Names that would exceed PostgreSQL's 63-char
identifier bound (very long custom table names) are now deterministically
truncated + hash-suffixed instead of being silently truncated by the engine
into collisions. System index names are reserved: a graph-declared index
using one is rejected at table definition and by `materializeIndexes()`
(previously its `CREATE INDEX IF NOT EXISTS` silently no-opped against the
differently-shaped system index while recording success). Legacy databases
that predate the recorded relations skip those indexes cleanly instead of
attempting failing DDL at every boot.
- [#273](https://github.com/nicia-ai/typegraph/pull/273) [`42f6941`](https://github.com/nicia-ai/typegraph/commit/42f6941fdb601629c0f45ec54fa9f3e50bd028ae) Thanks [@pdlug](https://github.com/pdlug)! - Add `trustedImportGraph` and `trustedImportGraphStream` for atomic initial loads
into a fresh, dedicated database. The distinct trusted surface bypasses schema,
reference, cardinality, and conflict validation; uses prepared SQLite writes or
PostgreSQL `UNNEST` ingestion; defers rebuildable secondary indexes; refreshes
planner statistics; and rolls the complete stream back on any failure.
The first version rejects non-empty TypeGraph data tables, recorded history,
revision tracking, uniqueness constraints, searchable fields, vector fields,
and backends without the required native transactional path.
- [#279](https://github.com/nicia-ai/typegraph/pull/279) [`c44eeac`](https://github.com/nicia-ai/typegraph/commit/c44eeac36805be144b25123fafb5c52ae30b73a4) Thanks [@pdlug](https://github.com/pdlug)! - Add exact `store.algorithms.weaklyConnectedComponents()` for transactional
SQLite and PostgreSQL backends. Results include deterministic component
representatives and sizes, honor valid/recorded temporal views, and fail with a
typed convergence error instead of returning partial labels when the configured
iteration budget is exhausted. Callers can restrict WCC to a `nodeKinds`
induced subgraph, retaining isolated in-scope nodes without seeding unrelated
node kinds.
PostgreSQL iterative operations now refresh temporary-table planner statistics
after sufficiently large seeds and multiplicative growth, avoiding plans based
on the engine's initial one-row estimate. The policy also covers growing BFS
working tables and is a no-op on SQLite.
Set-based reachability now deduplicates edge targets before target-node
visibility checks and avoids computing unused predecessor paths. This reduces
dense-frontier work while preserving minimum-depth results and cross-backend
semantics.
- [#288](https://github.com/nicia-ai/typegraph/pull/288) [`17a3f83`](https://github.com/nicia-ai/typegraph/commit/17a3f83d806e0b77748c346f5c320d8d55080b91) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.algorithms.weightedShortestPath` — a minimum-total-weight path
search weighting each traversed edge by a numeric edge property (LDBC
Interactive IC14 shape). Runs frontier-based relaxation on the shared
iterative substrate with best-target pruning, works on both execution paths
(temporary working table and inline fallback), and honors valid-time and
recorded-time coordinates including pinned StoreViews. Edge weights are
audited up front: negative, non-numeric, out-of-range, or (without
`defaultWeight`) missing weights throw the new typed
`InvalidEdgeWeightError`. Weight arithmetic is IEEE 754 double precision on
both backends, so total weights are backend-identical; among
equal-total-weight paths the returned node sequence is too, except when the
`edges` list exceeds the backend's bind-parameter budget (hundreds of edge
kinds in one call).
### Patch Changes
- [#282](https://github.com/nicia-ai/typegraph/pull/282) [`923219d`](https://github.com/nicia-ai/typegraph/commit/923219d6854a0a97cc186c2f4f27e6564bb935be) Thanks [@pdlug](https://github.com/pdlug)! - Add a `(graph_id, id)` index to the live and recorded node tables so bare-id
lookups (a node's `id` without its `kind`) seek instead of scanning the
graph's node partition — the composite keys lead with `kind`, so they can't
serve that probe. `store.algorithms.degree()`'s node-kind subquery is the
main consumer: ~95 ms → sub-millisecond at LDBC SNB SF1 (3.16M nodes) on
SQLite, at the live and recorded coordinates alike.
New databases get both indexes at bootstrap. Existing databases adopt them
with a one-time `await backend.bootstrapTables()` — every statement is
`CREATE … IF NOT EXISTS`, so the call is idempotent and only creates what's
missing. On PostgreSQL this issues a plain `CREATE INDEX` (briefly locks
writes on large tables); schedule it, or apply the equivalent
`CREATE INDEX CONCURRENTLY` statements manually.
- [#265](https://github.com/nicia-ai/typegraph/pull/265) [`35ab2a0`](https://github.com/nicia-ai/typegraph/commit/35ab2a02af728df9059750518ddbdd12e489450e) Thanks [@pdlug](https://github.com/pdlug)! - Docs: scope the `coalesceUnchangedUpserts` benefit correctly. Coalescing
eliminates _re-delivery_ churn (an already-applied change delivered again,
value-identical to the live row). It does not make a full replay-from-zero
free when the stream supersedes values in place: re-applying an older value
over the live row is a genuine change, and restoring the current value
afterwards is another, so such a replay still writes — and leaves a spurious
back-and-forth band in the live store's recorded history. Churn-free rebuilds
replay into a fresh store instead. Clarified in the option's TSDoc and in the
"Materializing external event logs" guide; no behavior change.
- [#289](https://github.com/nicia-ai/typegraph/pull/289) [`199b33a`](https://github.com/nicia-ai/typegraph/commit/199b33aabffa304b6a11f34ad7ca9b0e5f449218) Thanks [@pdlug](https://github.com/pdlug)! - Cap SQLite-backed Durable Object statements at Cloudflare's 100-bound-parameter
limit. Structural client detection now makes platform identity authoritative
over stale execution hints, and capability overrides cannot raise the hard
ceiling. Recorded-history capture and every capability-driven SQLite batch path
chunk large writes before workerd rejects the query, while SQLite literal list
predicates use one JSON-bound parameter instead of one bind per element.
- [#285](https://github.com/nicia-ai/typegraph/pull/285) [`9949562`](https://github.com/nicia-ai/typegraph/commit/99495623057610e502671515709c29d8f5139ae2) Thanks [@pdlug](https://github.com/pdlug)! - Cut three overheads out of the iterative graph algorithms, root-caused with
`EXPLAIN (ANALYZE, BUFFERS)` against LDBC SNB SF1 on PostgreSQL.
Weakly connected components no longer re-validates node visibility per edge in
its propagate rounds. The working table is seeded through the same
graph/kind/temporal filters inside the same snapshot and both edge endpoints
are already joined against it, so membership is the visibility proof; the
per-edge `typegraph_nodes` index loops (hundreds of thousands per round on
SF1) added nothing. Results are byte-identical.
Traversal rounds now carry their own bookkeeping instead of issuing follow-up
statements: seeding returns the frontier through `INSERT … RETURNING`, and
bidirectional shortest-path rounds detect the frontier meeting inside the
expansion statement rather than with a separate probe per round. A
shortest-path traversal that used to issue two to three statements per round
now issues one, roughly halving round-trip latency on latency-bound
connections. The working-table `ANALYZE` policy is unchanged in its
thresholds but no longer runs when no further round will read the table.
When several equal-depth meetings exist, the tie now breaks by node id then
kind in code-unit order on both backends — previously the selection followed
the database collation, so a PostgreSQL cluster with a linguistic default
collation could pick a different (equally shortest) path.
New option: iterative algorithm calls (`reachable`, `shortestPath`,
`canReach`, `neighbors`, `weaklyConnectedComponents`) accept
`workingMemory?: string`, an opt-in, transaction-scoped override of the
session's `work_mem`, applied on PostgreSQL with `SET LOCAL` semantics via
parameterized `set_config`. By default (option omitted) operations inherit
the server's configured `work_mem` — nothing is overridden. `work_mem` is a
threshold each sort/hash operator (and each parallel worker) may allocate up
to, not a per-operation budget, and concurrent calls multiply it; set it
deliberately (e.g. `"64MB"`) for large single-tenant analytical runs where
the configured default spills whole-graph sorts to disk (measured ~106MB
external merges per WCC round on SF1). The override never touches the
session or server setting, is validated as `kB|MB|GB` within
PostgreSQL's accepted `work_mem` range (64kB–2147483647kB) with the same
typed error on both backends, and is ignored by SQLite.
- [#290](https://github.com/nicia-ai/typegraph/pull/290) [`247c1b7`](https://github.com/nicia-ai/typegraph/commit/247c1b77b8c30d2f03520d52214bfaebbf1a0e6c) Thanks [@pdlug](https://github.com/pdlug)! - Fix PostgreSQL pointer-level `pathIsNull()` / `pathIsNotNull()` predicates
misclassifying two stored value shapes. The previous text-comparison form
(`#>> path = 'null'`) went three-valued on a stored JSON `null` — so
`pathIsNull()` silently failed to match those rows on PostgreSQL while
matching them on SQLite — and misread the JSON _string_ `"null"` as null,
falsely matching it with `pathIsNull()` and excluding it from
`pathIsNotNull()`. Both predicates are now type-based (`jsonb_typeof`) and
never SQL NULL, converging on SQLite's (correct) semantics. Field-level
`isNull()` / `isNotNull()` predicates were already correct and are unchanged.
Behavior change on PostgreSQL for affected data: rows holding a JSON `null`
now match `pathIsNull()`, and rows holding the string `"null"` no longer do.
- [#283](https://github.com/nicia-ai/typegraph/pull/283) [`8306680`](https://github.com/nicia-ai/typegraph/commit/830668032328e5fdc031deccfaf0c452f466a15c) Thanks [@pdlug](https://github.com/pdlug)! - Selective `ORDER BY … LIMIT` queries now compile with late materialization:
the query sorts and limits a lean candidate set carrying only identity, sort
keys, and predicate columns, then re-fetches the deferred projection columns
by primary key for only the surviving rows — instead of extracting every
projected column for every candidate and discarding all but the `LIMIT`
survivors after the sort. At LDBC SNB SF1, IC9's top-20 over a 1.18M-comment
fan-out stops extracting `content` 1.18M times, ~30–37% faster on SQLite.
The transform fires only on the selective `.select()` path with `ORDER BY`
and a positive `LIMIT` at the live coordinate. Aggregates, vector/fulltext,
optional (LEFT JOIN) traversals, edge-field projections, non-selective
queries, and recorded-time reads keep the flat plan unchanged.
- [#274](https://github.com/nicia-ai/typegraph/pull/274) [`2a889aa`](https://github.com/nicia-ai/typegraph/commit/2a889aaea095a842fdd6f1b4a97feba5e6026d82) Thanks [@pdlug](https://github.com/pdlug)! - Replace path-enumerating recursive CTEs in `reachable`, `neighbors`,
`shortestPath`, and `canReach` with set-based breadth-first search.
Transactional SQLite and PostgreSQL backends now execute graph iterations
against a connection-local temporary working table, de-duplicated by node kind
and ID on every round. Non-transactional backends retain parity through a
bind-limit-aware inline frontier. Traversals run in one snapshot where the
backend supports transactions, preserve temporal filtering, and clean up
temporary state on success or failure.
## 0.36.0
### Minor Changes
- [#261](https://github.com/nicia-ai/typegraph/pull/261) [`5bc7b53`](https://github.com/nicia-ai/typegraph/commit/5bc7b5333d30392c31605161436441b3e8602447) Thanks [@pdlug](https://github.com/pdlug)! - Return a receipt from `store.withRecordedTransaction`, and add scoped write
measurement with `tx.measure`.
- **`store.withRecordedTransaction(externalTx, fn)` now returns
`Promise>`** instead of `Promise`. The adopted path
is the only way to get exactly-once cursors and graph writes atomically on a
history store, and it now surfaces the same receipt `transactionWithReceipt`
does: `receipt.writes` for dropped-change detection and `receipt.recorded` as
the per-transaction replay anchor (`undefined` for a read-only callback or a
non-history store).
**BREAKING:** the adopted path now returns the result under `.result`. Migrate
by destructuring:
```typescript
// Before
const x = await store.withRecordedTransaction(externalTx, fn);
// After
const { result: x } = await store.withRecordedTransaction(externalTx, fn);
```
- **Scoped receipts — `tx.measure((scoped) => ...)`.** On the receipt-enabled
contexts (`transactionWithReceipt`, `withRecordedTransaction`), `tx.measure`
runs its callback with a **scoped context** — a second view over the same
transaction — and returns a `TransactionOutcome` whose receipt counts exactly
the writes made **through that scoped context** (`scoped.nodes` /
`scoped.edges`). So a framework can attribute writes to user code it invoked
(e.g. a materializer measuring `project(scoped, change)` to detect a dropped
change) while its own bookkeeping — written through the outer `tx` — stays out
of the count. Attribution is by which context you write through, not by
timing, which makes overlapping and concurrent measures safe by construction
(two scopes racing under `Promise.all` never cross-count). Nesting composes;
measured writes still count in the outer receipt; a scoped receipt's
`recorded` is always `undefined`. Plain `store.transaction()` contexts have no
`measure` (that path runs no recorder and stays zero-overhead). New exported
types: `MeasurableTransactionContext`, `MeasurableHistoryTransactionContext`,
`ScopedMeasure`.
- **Adopted contexts seal on return.** A transaction context retained and
written through _after_ its `withRecordedTransaction` callback resolves now
fails loud on both paths — the history path's capture guard is checked
_before_ the live write (so a swallowed error can no longer commit an
uncaptured row), and the non-history path seals its receipt-tracked
collections (so a post-return write can't persist a row the already-returned
receipt never counted).
- [#262](https://github.com/nicia-ai/typegraph/pull/262) [`34468a0`](https://github.com/nicia-ai/typegraph/commit/34468a04f9cb6abff34177282d31f1240f1254d1) Thanks [@pdlug](https://github.com/pdlug)! - Add an opt-in `coalesceUnchangedUpserts` store option for at-least-once /
replay materializers.
Idempotent event-log projectors converge live state correctly, but every
re-delivery of a byte-identical value still performed a real write:
`upsertById` on an existing id called `updateNode` unconditionally, allocating
a fresh recorded instant and a new history row. A full replay of an N-event log
therefore rewrote every row and grew recorded history by N — the recovery /
rebuild workload inflates history the most.
With `createStore(graph, backend, { coalesceUnchangedUpserts: true })`, an
`upsertById` (or `bulkUpsertById` item) whose validated props are
value-identical to the existing **live** row performs **no write at all**: no
`updateNode`, no recorded-time capture, no history row, no revision-anchor
advance, and no `update` operation hooks. It resolves with the existing node.
The dirty-check compares the storage-normalized representation (props run
through the kind's Zod schema, key-order-independent), so it answers exactly
"would the persisted value differ?".
A write still happens (never coalesced) when the row is soft-deleted (an upsert
resurrects it), when an explicit `validFrom` / `validTo` is passed, or when any
prop differs. Default off, because some consumers want an audit row per
re-delivery. Covered symmetrically for edge `bulkUpsertById` (props only —
endpoints are the edge's identity).
Receipt semantics are unchanged and need no new signal: a coalesced upsert
still counts as one write intent (`writes.total`) but captures nothing
(`recorded` stays `undefined`) — the same two-signal shape as a no-op delete,
which at-least-once consumers already handle by carrying the prior anchor
forward.
- [#260](https://github.com/nicia-ai/typegraph/pull/260) [`35d03ae`](https://github.com/nicia-ai/typegraph/commit/35d03ae0fd4647286927d617210f20a7e47df4b6) Thanks [@pdlug](https://github.com/pdlug)! - Make the store transaction surface tell the truth about raw SQL and history
capture.
- **New `tx.sqlAvailability` discriminant.** Every transaction context now
carries a required `sqlAvailability: "available" | "history" |
"revisionTracking" | "unavailable"` field. Branch on it instead of
truthiness-testing `tx.sql`: under `history: true` / `revisionTracking: true`
the raw handle is present-but-throwing (so `if (tx.sql)` read truthy and then
threw), and it is `undefined` only on the non-transactional fallback. `"available"`
means `tx.sql` is a usable raw handle; `"history"` / `"revisionTracking"` mean
raw SQL is disabled here; `"unavailable"` means the backend has no transactions
(`tx.sql === undefined`, no atomicity).
- **`store.withTransaction()` on a history-enabled store is now a compile error.**
It always threw at runtime; the call site now rejects the argument with a
message pointing at `store.withRecordedTransaction()`. The runtime guard is
unchanged for suppressed calls.
- **Branchable recorded-capture guard codes.** The `ConfigurationError`s these
guards throw carry a stable `details.code`
(`RECORDED_CAPTURE_REQUIRES_CALLBACK_TRANSACTION`,
`RECORDED_CAPTURE_RAW_SQL_DISABLED`, `REVISION_TRACKING_RAW_SQL_DISABLED`), now
exported as `RECORDED_CAPTURE_GUARD_CODES` with a `RecordedCaptureGuardCode`
type and an `isRecordedCaptureGuardError(error, code?)` type guard — so a
portable caller can distinguish "history forbids raw SQL here" from "this
backend has no transactions" without substring-matching the message.
- **Fixed `withRecordedTransaction`'s JSDoc**, which incorrectly promised
`tx.sql`; on the adopted path you already hold the pinned connection, so write
your own relational tables through the external transaction handle you passed
in.
## 0.35.0
### Highlights
TypeGraph 0.35 improves bulk ingestion, search, and traversal performance across SQLite and PostgreSQL. Bulk creation and import batch their validation and side effects, large autocommit loads refresh planner statistics, and built-in hybrid search combines retrieval, fusion, and hydration in one SQL statement. Search gains property filters, pagination, subclass scope, and explicit approximate vector retrieval; exact retrieval remains the default.
Graph branches can now use durable revision origins and counters to validate their base without fingerprinting every live row. Streaming interchange supports larger graph copies, and `transactionWithReceipt()` exposes completed collection write intents and recorded commit coordinates. Index declarations gain GIN, trigram, and system-column keys.
Correctness fixes keep repeated and prepared queries bound to a fresh read instant, preserve temporal metadata through branch copies, and enforce exact vector results even when an approximate index exists. Current temporal reads now use the application clock consistently with writes.
### Upgrade notes
- Audit persisted schemas as well as source definitions for endpoint-incompatible `implies()` relations. These are now rejected when constructing or loading a registry.
- Backend rows expose `props` as `RowProps`, which may be JSON text or a parsed object. Direct backend consumers should use `rowPropsToObject()` or `rowPropsToJsonText()` instead of unconditionally calling `JSON.parse()`.
- Custom vector strategies must declare the new filtered approximate-search capability, including whether they guarantee a full result page.
- Existing databases retain their old edge traversal indexes until those indexes are explicitly rebuilt. The detailed entries below include the PostgreSQL and SQLite rebuild procedures; rerunning `CREATE INDEX IF NOT EXISTS` alone does not widen an existing index.
- New records without an explicit `validFrom` now begin at their creation instant. Existing open-left records retain their stored bounds, and interchange preserves those bounds through branch copies. Approximate vector retrieval requires `{ approximate: true }`; applications that previously received approximate results on the default path may see a different cost for the corrected exact search.
### Minor Changes
- [#231](https://github.com/nicia-ai/typegraph/pull/231) [`839f536`](https://github.com/nicia-ai/typegraph/commit/839f53621998d41704537e45408872d49452cf1c) Thanks [@pdlug](https://github.com/pdlug)! - Aggregate queries now support `.orderBy()`. Previously `ExecutableAggregateQuery`
exposed `limit()` but no way to order results, so `.aggregate({...}).limit(n)`
returned an arbitrary `n` groups rather than the top `n` — the most common
aggregate shape ("top N groups by count/sum") required fetching every group
and sorting in JS.
`.orderBy(key, direction?)` takes any output name from `.aggregate({...})` —
either a grouped field or an aggregate alias — and can be chained for
multi-key sorts:
```typescript
store
.query()
.from("Author", "a")
.traverse("wrote", "e")
.to("Book", "b")
.groupByNode("a")
.aggregate({ author: field("a", "name"), bookCount: count("b") })
.orderBy("bookCount", "desc")
.limit(2)
.execute();
```
Ordering resolves against the projected SELECT-list output alias rather than
recompiling the underlying expression, so it works uniformly for grouped
fields and aggregates on both SQLite and PostgreSQL with no dialect-specific
handling.
- [#212](https://github.com/nicia-ai/typegraph/pull/212) [`dcdd542`](https://github.com/nicia-ai/typegraph/commit/dcdd54246fef1e93839196d7029e4dbadbc72b42) Thanks [@pdlug](https://github.com/pdlug)! - Autocommit `bulkCreate` and `bulkInsert` calls (nodes and edges) now
refresh planner
statistics automatically when a single call writes 1,000 rows or more,
closing the stale-statistics window after bulk loads where the planner
keeps pre-load row estimates until ANALYZE runs (observed 25-200x
slowdowns on traversal and fulltext shapes). Tune the threshold or
disable with the new `autoRefreshStatistics` store option
(`createStore(graph, backend, { autoRefreshStatistics: 5000 })` or
`false`). Bulk writes inside a caller-provided transaction never
auto-refresh — statistics cannot see uncommitted rows — and a refresh
failure degrades to a warning without failing the committed write.
`importGraph()` keeps its existing built-in refresh.
- [#195](https://github.com/nicia-ai/typegraph/pull/195) [`e48dfa2`](https://github.com/nicia-ai/typegraph/commit/e48dfa2531148892ca7f5432a3ced6068b464807) Thanks [@pdlug](https://github.com/pdlug)! - `bulkCreate` now batches its round trips end to end instead of degenerating
into per-row statements around one multi-row INSERT.
- Validation probes: per-row existence checks collapse into one `getNodes`
per kind, and per-row uniqueness pre-checks into one `checkUniqueBatch`
per (constraint, kind) — the batch validation caches are primed up front,
so the per-row checks run against memory. Validation now runs as a
synchronous first pass, so a later row's validation error can surface
before an earlier row's constraint error (both fail the whole batch).
- Side effects: uniqueness entries write through a new `insertUniqueBatch`
(multi-row conditional upsert with the same per-entry `UniquenessError`
semantics), fulltext sync goes through the existing `upsertFulltextBatch`,
and embedding sync through a new `upsertEmbeddingBatch` per
(kind, field) — implemented for pgvector, sqlite-vec, and libSQL native
vectors via an optional `VectorStrategy.buildUpsertBatch` seam with a
per-row fallback for custom strategies.
Measured on the write bench (in-memory SQLite, 100-row batches of nodes
with searchable + embedding fields): ~1,600 → ~4,100 rows/s (~2.6×). The
win compounds on per-statement-networked engines (Turso, D1, Neon), where
each eliminated statement is a network round trip.
- [#194](https://github.com/nicia-ai/typegraph/pull/194) [`b3668c9`](https://github.com/nicia-ai/typegraph/commit/b3668c96db58127f983695fa6df8f39662ed761b) Thanks [@pdlug](https://github.com/pdlug)! - Default-path performance tuning for SQLite and bulk maintenance verbs.
- `createLocalSqliteBackend` now applies connection pragmas at open:
`journal_mode=WAL`, `synchronous=NORMAL`, and a 5s `busy_timeout`. On
file-backed databases this makes single-operation writes roughly 5×
faster than the better-sqlite3 driver defaults (rollback journal,
`synchronous=FULL`), because each write no longer pays a full-durability
fsync in journal mode. Override individual values via the new `pragmas`
option, or pass `pragmas: false` to keep driver defaults.
- The SQLite backend now detects the connection's real bound-parameter
budget instead of assuming the historic 999: better-sqlite3 compiles in
`SQLITE_MAX_VARIABLE_NUMBER=32766` (probed via `PRAGMA compile_options`,
with a `sqlite_version() >= 3.32` fallback), Cloudflare D1 is capped at
its documented 100, and undetectable async drivers keep the conservative
999 floor. Batch chunk math derives from the detected budget, so bulk
inserts on better-sqlite3 use ~33× fewer statements (111-row chunks →
3,640-row chunks), and batched writes on D1 no longer exceed its
per-statement limit. `capabilities.maxBindParameters` reports the
detected value and remains overridable.
- `importGraph()` now refreshes planner statistics (`ANALYZE`)
automatically after an import that created or updated rows, and
`store.materializeIndexes()` does the same on SQLite after creating
indexes. Stale statistics after bulk loads previously degraded
traversals ~10× on PostgreSQL and some FTS5 queries ~30× on SQLite until
the engine caught up on its own. Both verbs accept
`refreshStatistics: false` to opt out. On PostgreSQL,
`materializeIndexes()` builds with `CREATE INDEX CONCURRENTLY` and skips
the automatic refresh (concurrent same-index builds from two callers can
deadlock when a refresh shifts their timing) — call
`store.refreshStatistics()` after materializing.
- PostgreSQL `refreshStatistics()` now issues one `ANALYZE (SKIP_LOCKED)`
per table instead of a single multi-table `ANALYZE`. A multi-table
ANALYZE is one transaction acquiring several ShareUpdateExclusive locks
in sequence, and ANALYZE's lock class conflicts with in-flight
`CREATE INDEX CONCURRENTLY` builds — the old shape could deadlock
against concurrent index DDL; the new one can never join a lock-wait
cycle (a locked table is skipped and covered by the next refresh or
autovacuum).
- [#247](https://github.com/nicia-ai/typegraph/pull/247) [`191e877`](https://github.com/nicia-ai/typegraph/commit/191e877796fde30ad606993948decea7305fd367) Thanks [@pdlug](https://github.com/pdlug)! - Declare, as a typed capability, whether a backend's filtered approximate vector
search can silently return a short page.
Every approximate (ANN) search TypeGraph issues carries at least one row filter —
the liveness predicate that hides soft-deleted and out-of-validity rows — and a
`.where(...)` predicate narrows it further. Where the engine applies that filter
relative to the index traversal decides whether the page fills:
- **`sqlite-vec`** pushes the filter into the `vec0` KNN candidate set. Exact —
the only engine here that guarantees a full page.
- **`pgvector` ≥ 0.8** re-enters the index for more candidates
(`hnsw.iterative_scan` / `ivfflat.iterative_scan`, applied automatically).
Much better recall than a post-filter, but **not** a guarantee: the iterative
scan stops at `hnsw.max_scan_tuples` / `ivfflat.max_probes`, and on
**pgvector < 0.8** there is no iterative scan at all — the backend detects
that at runtime, warns once, and the search stays `ef_search`-bounded.
- **`libsql-native`** cannot do either: DiskANN's `vector_top_k` is a table
function with no filter pushdown. TypeGraph over-fetches `4 × (limit + offset)`
neighbors and post-filters, so once more than that headroom is filtered out the
search returns **fewer than `limit` rows even though more matches exist**.
Heavy tombstone drift — routine in a temporal store — is what makes this real
rather than theoretical.
That asymmetry was previously only a code comment. `VectorCapabilities` now
carries a required `filteredApproximateSearch: { mode, guaranteesFullPage }`.
**Read `guaranteesFullPage`, not `mode`** — `mode`
(`"filter-pushdown" | "iterative-scan" | "post-filter"`) names the mechanism the
strategy asks for, but only `guaranteesFullPage` reflects the runtime-dependent,
scan-bounded reality (it is `true` for `sqlite-vec` alone). It is documented in
the backend parity matrix, and boundary tests execute the difference against real
libSQL, sqlite-vec, and pgvector: the same 200-vector fixture, the same filter,
the same `limit`.
**Breaking for custom vector strategies only.** `VectorCapabilities` gained a
required field, so a hand-written `VectorStrategy` must now declare both its mode
and whether it guarantees a full page. That is deliberate: an omitted declaration
would inherit an engine promise the strategy may not keep.
- [#198](https://github.com/nicia-ai/typegraph/pull/198) [`a9477bb`](https://github.com/nicia-ai/typegraph/commit/a9477bb28ee887a1a93c103a64912e8563de9d76) Thanks [@pdlug](https://github.com/pdlug)! - Property filters that a btree can never serve now have a declarative index
story: `defineNodeIndex` / `defineEdgeIndex` accept
`method: "gin" | "trigram"` (default `"btree"`, unchanged).
- `method: "gin"` emits a PostgreSQL expression GIN (`jsonb_path_ops`) over
the field's jsonb extraction, serving the array containment predicates
(`contains` / `containsAll` / `containsAny` on array fields). Verified to
match TypeGraph's compiled `(props #> ARRAY[…]) @> $1` form under
parameterized prepared statements — note that a hand-written
whole-column `GIN (props)` never matches these expressions (the previous
docs guidance recommended one; corrected).
- `method: "trigram"` emits an expression GIN with `gin_trgm_ops` over the
field's text extraction, serving substring and case-insensitive matches
(`contains` / `startsWith` / `endsWith` / `like` / `ilike` on string
fields). `materializeIndexes()` installs `pg_trgm`
(`CREATE EXTENSION IF NOT EXISTS`) on first use.
Both are materialize-only (like vector ANN indexes) and PostgreSQL-only:
`materializeIndexes()` reports them as `skipped` on SQLite, whose
substring-search story is FTS5 fulltext. GIN-family declarations take
exactly one field and reject `unique`, `coveringFields`, and `where`;
`method: "btree"` is canonicalized by absence so existing stored schema
documents and materialization signatures are unchanged. `bulkFindByIndex`
rejects GIN-family indexes (it compiles equality probes, which only btree
declarations serve).
- [#204](https://github.com/nicia-ai/typegraph/pull/204) [`94eea90`](https://github.com/nicia-ai/typegraph/commit/94eea90ead38c69c0ac5b55bad34036f45578b87) Thanks [@pdlug](https://github.com/pdlug)! - perf: `store.search.hybrid` now runs as a single SQL statement on the built-in backends — both sources, weighted RRF fusion, liveness, and node hydration composed into one round trip (previously two search statements plus an id-hydration fetch, with fusion in JS). Results are identical to the previous path; the saving scales with per-statement cost (serverless drivers, D1/Durable Objects, remote databases). `GraphBackend` gains an optional `hybridSearch` member; backends without it (custom backends, capability profiles without window functions) keep the multi-statement fallback.
- [#223](https://github.com/nicia-ai/typegraph/pull/223) [`a161d70`](https://github.com/nicia-ai/typegraph/commit/a161d70895da5101706602ac13e4cca4b7fc6a62) Thanks [@pdlug](https://github.com/pdlug)! - Add `asNodeId` and `asEdgeId` constructors for branding persisted ids that
round-trip through untyped storage before being passed back to read, update, or
delete APIs.
- [#241](https://github.com/nicia-ai/typegraph/pull/241) [`8f3e772`](https://github.com/nicia-ai/typegraph/commit/8f3e7727dde0d46415b90e715138c0a9766cd2b5) Thanks [@pdlug](https://github.com/pdlug)! - Fixes `implies(edgeA, edgeB)` silently accepting endpoint-incompatible edge
pairs. Previously an ontology declaration like `implies(about, writes)` — where
`about` connects `Paper -> Topic` and `writes` connects `Author -> Paper` —
was accepted without complaint, and `expand: "implying"` query traversal
would then silently fold `about` rows into a `writes` traversal even though
the two edges connect entirely different node kinds.
`implies()` relations are now validated wherever a query-capable
`KindRegistry` is built — `createStore()`/`createStoreWithSchema()` for a
live graph definition, and `deserializeSchema(...).buildRegistry()` for a
persisted schema — including relations authored through
`store.evolve({ ontology })`. A relation is accepted when every kind the
implying edge allows on a side (`from`/`to`) is assignable — equal, or a
`subClassOf` descendant — to at least one kind the implied edge allows on
that same side; otherwise construction throws a `ConfigurationError`
describing the incompatible kinds and how to fix the declaration.
**Breaking change — two things to know before upgrading.**
_It breaks the load path, not just graph definition._ `deserializeSchema(...)`
runs the same endpoint check inside `buildRegistry()`, so a schema **already
persisted** under 0.34 that carries a now-rejected `implies()` relation throws
at the first `buildRegistry()` after the upgrade — no code change of yours
required to trigger it. Audit persisted schemas before rolling out, not only
the graph definitions in source.
_It rejects superset domains, not only disjoint ones._ A relation is accepted
only when every kind the implying edge allows on a side is assignable to at
least one kind the implied edge allows on that side. So `implies(a, b)` where
`a` is declared `from: [Person]` and `b` is declared `from: [Employee]` (with
`Employee subClassOf Person`) is **rejected**, even though every `a` row on
disk might in fact start at an `Employee`: `Person` is not assignable to
`Employee`. The declaration, not the data, is what the traversal folds on, and
a `Person`-rooted `a` row folded into a `b` traversal would be unsound. The
same rule is what makes the previously-silent disjoint case (`Paper -> Topic`
implying `Author -> Paper`) an error.
Fix such relations by narrowing the implying edge's endpoints, adding a
`subClassOf` relation to bridge the mismatch, or removing the `implies()`
declaration.
- [#195](https://github.com/nicia-ai/typegraph/pull/195) [`e48dfa2`](https://github.com/nicia-ai/typegraph/commit/e48dfa2531148892ca7f5432a3ced6068b464807) Thanks [@pdlug](https://github.com/pdlug)! - `importGraph` now processes each `batchSize` slice with batched round trips
instead of fully single-row statements. Nodes: one `getNodes` per kind for
existence, one `checkUniqueBatch` per (constraint, kind) for uniqueness
pre-checks, one multi-row insert, and one batched side-effect pass
(uniqueness entries, fulltext, embeddings) for the accepted creates.
Edges: one `getNodes` per endpoint kind for reference liveness, one
`getEdges` for existence, and one multi-row insert.
Per-row semantics are unchanged: conflicts route by `onConflict`, a
uniqueness conflict is recorded as a per-row error entry (the rest of the
import proceeds), reference validation still rejects missing or tombstoned
endpoints, and rows repeating an id within a slice fall back to the
per-row path so they observe the first occurrence's row exactly as before.
Measured on the write bench (in-memory SQLite, 500 nodes + 500 edges per
import): ~26k → ~96k entities/s (~4×). The win compounds on
per-statement-networked engines (Turso, D1, Neon), where the old path paid
one round trip per row and the new one pays a handful per slice.
- [#236](https://github.com/nicia-ai/typegraph/pull/236) [`31aee82`](https://github.com/nicia-ai/typegraph/commit/31aee82608518411e4e9f905c96c52348f7cf08f) Thanks [@pdlug](https://github.com/pdlug)! - `defineNodeIndex` accepts a new `keySystemColumns` option: system columns
(e.g. `"id"`) to include in the index key, positioned after the `scope`
prefix and before `fields`/`coveringFields`. `fields` is now optional (was
a required non-empty tuple) — an index must declare at least one of
`fields`, `coveringFields`, or `keySystemColumns`.
This closes a real gap: a covering index can only serve a query's join
index-only (avoiding a heap fetch per candidate row) if the index's key
matches the join's actual predicate. Queries that join on a system column
directly (e.g. TypeGraph's compiled `n.id = e.from_id` for a reverse
traversal) had no way to declare a matching index, since `fields`/
`coveringFields` only ever accept the node's own schema properties.
`keySystemColumns: ["id"]` (plus `coveringFields` for whatever the query
also projects) now lets that same join be served index-only.
Rejects edge-only system columns (`from_kind`/`from_id`/`to_kind`/
`to_id`) on a node index, and rejects any column already implied by
`scope`. Not supported with `method: "gin" | "trigram"` (same restriction
as `coveringFields`). Also rejects `unique: true` combined with
`keySystemColumns: ["id"]` — every node's `id` is already unique per
row, so a unique index keyed on `id` plus other columns can never
enforce a meaningful constraint across those other columns. Canonicalized
by absence, like `method`: indexes that don't use it produce byte-identical
names/hashes to before this field existed, so existing stored schema
documents and materialization signatures are unaffected.
- [#208](https://github.com/nicia-ai/typegraph/pull/208) [`586b2b0`](https://github.com/nicia-ai/typegraph/commit/586b2b05f3f501f3d53db1dbb2ec247e17a67294) Thanks [@pdlug](https://github.com/pdlug)! - fix: `materializeIndexes` serializes same-index builds across callers on PostgreSQL via a durable claim in the status table (two concurrent same-name expression-index `CREATE INDEX CONCURRENTLY` builds can deadlock — no safe-snapshot exemption). Losers wait and converge as `alreadyMaterialized`; a crashed builder's claim expires after a 15-minute lease and the takeover drops the INVALID index leftover before rebuilding (relational indexes now self-heal instead of requiring manual repair). With same-index builds serialized, the automatic post-create `ANALYZE` is re-enabled on PostgreSQL.
- [#201](https://github.com/nicia-ai/typegraph/pull/201) [`b52ae3b`](https://github.com/nicia-ai/typegraph/commit/b52ae3b3358435de9774f348fc94ab7140bdc7eb) Thanks [@pdlug](https://github.com/pdlug)! - perf: eliminate the PostgreSQL JSONB parse→stringify→parse round trip per row.
**Public backend row contract change:** rows returned by `GraphBackend` read methods now carry `props` as `RowProps = string | Readonly>` — JSON text on SQLite, the driver-parsed object on PostgreSQL. Code that consumed backend rows directly with `JSON.parse(row.props)` must switch to the new `rowPropsToObject(row.props)` (or `rowPropsToJsonText` when text is required); both helpers and the `RowProps` type are exported from the package root. Store-level APIs (`store.nodes.*`, `store.query()`, search, export) are unaffected — they already return parsed objects.
- [#249](https://github.com/nicia-ai/typegraph/pull/249) [`d2a6feb`](https://github.com/nicia-ai/typegraph/commit/d2a6feb8a99aaafa247c7bf97f9670c56608a870) Thanks [@pdlug](https://github.com/pdlug)! - Add revision-anchored graph branches and streaming interchange. Stores can opt
into `revisionTracking: true` (or use `history: true`) so branch and merge
validation read a durable per-graph origin and revision instead of
fingerprinting every live row or accepting a coincident revision from another
store. Physical branch clones now stream bounded interchange batches, enabling
large branch copies, exports, and imports without materializing the full graph
in memory. Direct backend writes remain outside the revision-tracking contract;
tracked stores fail loudly if `tx.sql` would bypass that contract.
- [#203](https://github.com/nicia-ai/typegraph/pull/203) [`801768d`](https://github.com/nicia-ai/typegraph/commit/801768d2e1a63a0d3bda9d40a46a7f03deddffbd) Thanks [@pdlug](https://github.com/pdlug)! - feat: facade search scoping — `store.search.{vector,fulltext,hybrid}` accept `where` (a property predicate compiled by the shared query compiler into the search statement's candidate set), `offset` (rank-relative pagination pushed into the engine), and `includeSubClasses` (search `subClassOf` descendants and merge into one ranking). Filters compile into the search statement's candidate set — exact on pgvector, sqlite-vec, tsvector, and FTS5, where a filtered search returns `limit` hits whenever enough matches exist; libSQL DiskANN post-filters a 4× over-fetched ANN set, so its recall against the filter is bounded by that headroom. Search now applies full current-read semantics (validity windows, not just tombstones), matching `find()`.
- [#205](https://github.com/nicia-ai/typegraph/pull/205) [`17bbe54`](https://github.com/nicia-ai/typegraph/commit/17bbe5419a246c95bbab9f6bc7da64f6691e159e) Thanks [@pdlug](https://github.com/pdlug)! - feat: `.similarTo(vector, k, { approximate: true })` — opt-in approximate retrieval for the inline vector predicate. Each declaring kind's relevance branch compiles to the engine's native ANN search form (vec0 `MATCH … k=`, libSQL `vector_top_k`, pgvector's index-eligible scan), scoped to the query's candidate nodes via the same pushdown the search facade uses, so composed predicates and traversals still constrain results. Never applied silently: the default remains the exact distance scan, and slots declared `indexType: "none"` keep it even with the opt-in.
- [#245](https://github.com/nicia-ai/typegraph/pull/245) [`ef6def6`](https://github.com/nicia-ai/typegraph/commit/ef6def6b67e306a9cdb40e78723dad6d36f89647) Thanks [@pdlug](https://github.com/pdlug)! - `createLocalSqliteBackend`'s `pragmas` option accepts two new fields:
`cacheSizeKib` (`PRAGMA cache_size`) and `mmapSizeBytes` (`PRAGMA
mmap_size`). Both default to `undefined`, leaving SQLite's own built-in
defaults (a 2MiB page cache, mmap disabled) untouched — existing callers
are unaffected.
SQLite's 2MiB default cache is fine for a small embedded database, but
once a database's working set exceeds it, every page a query touches past
that point pays a fresh disk read instead of a cache hit — including pages
an otherwise fully covering index would have served from cache alone. Set
`cacheSizeKib` (and optionally `mmapSizeBytes`) once a database's working
set is known to exceed the default, the same way you'd size a page cache
for any other embedded or server database engine.
- [#197](https://github.com/nicia-ai/typegraph/pull/197) [`f420a92`](https://github.com/nicia-ai/typegraph/commit/f420a922a1f168891ee4de54e91cc9ca1638deed) Thanks [@pdlug](https://github.com/pdlug)! - SQLite CRUD statements now reuse the prepared-statement cache. The
operation backend's read/write helpers previously executed through
drizzle's `db.all()` / `db.run()`, which re-prepares every statement on
every call — only the query engine's `backend.execute` path used the
prepared-statement LRU. On synchronous drivers (better-sqlite3,
bun:sqlite) CRUD statements and the per-write transaction frames
(`BEGIN IMMEDIATE` / `COMMIT` / `ROLLBACK`) now route through the
execution adapter's compiled path, so a repeated operation shape re-binds
parameters against a cached prepared statement. A warmed CRUD cycle
re-prepares nothing. Async drivers (remote libsql/Turso, D1) have no
statement cache and keep the existing execution path.
Measured on the write bench (in-memory SQLite, order-controlled A/B):
single-op creates ~18.3k → ~28.8k ops/s (~1.6×), transaction-batched
creates ~23.9k → ~36k ops/s (~1.5×).
- [#251](https://github.com/nicia-ai/typegraph/pull/251) [`f23f7a5`](https://github.com/nicia-ai/typegraph/commit/f23f7a5d15fd5fb59de3667c7d7b10e1975690d4) Thanks [@pdlug](https://github.com/pdlug)! - `createLocalSqliteBackend`'s `pragmas` option accepts a new field:
`walAutocheckpointPages` (`PRAGMA wal_autocheckpoint`). Defaults to
`undefined`, leaving SQLite's own built-in default (1,000 pages, ~4MiB)
untouched — existing callers are unaffected.
SQLite's default checkpoints WAL back into the main database file every
~4MiB. That's fine for a normal read/write mix, but a large bulk load pays
increasingly expensive checkpoints as the database file grows over the
course of the load — each checkpoint has to flush WAL frames into a B-tree
that's larger, and less page-cache-resident, than the one before it. A
local repro (real `bulkInsert()` calls, 100K/500K/2M synthetic rows)
confirmed this: raising `walAutocheckpointPages` cut a 2M-row bulk load's
wall-clock time by over 50% at the largest scale tested, with the effect
growing at larger row counts. Set `walAutocheckpointPages` for a
bulk-insert-heavy workload; `0` disables automatic checkpointing entirely
for callers that would rather run one explicit `PRAGMA wal_checkpoint`
after the load finishes.
- [#222](https://github.com/nicia-ai/typegraph/pull/222) [`7588634`](https://github.com/nicia-ai/typegraph/commit/758863402ec69b3724acb93f07a51eaf23132dc7) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.transactionWithReceipt()`, which runs a transaction and returns a
receipt summarizing completed collection write intents and, for
history-enabled stores, the recorded commit instant allocated by the
transaction.
- [#233](https://github.com/nicia-ai/typegraph/pull/233) [`e0e6304`](https://github.com/nicia-ai/typegraph/commit/e0e6304bc17b6d9c004376fd77ddf8cc3b0cc252) Thanks [@pdlug](https://github.com/pdlug)! - Every `TypeGraphError` subclass with a fixed-shape `details` payload now
declares a narrowed `readonly details` type (e.g.
`RestrictedDeleteError.details` is `RestrictedDeleteErrorDetails`, not the
base class's `Readonly>`), so reading structured
fields like `error.details.edgeCount` no longer requires a cast. The new
`XxxErrorDetails` types (`NodeNotFoundErrorDetails`,
`EdgeNotFoundErrorDetails`, `KindNotFoundErrorDetails`,
`NodeConstraintNotFoundErrorDetails`, `NodeIndexNotFoundErrorDetails`,
`EndpointNotFoundErrorDetails`, `EndpointErrorDetails`,
`UniquenessErrorDetails`, `CardinalityErrorDetails`, `DisjointErrorDetails`,
`RestrictedDeleteErrorDetails`, `VersionConflictErrorDetails`,
`SchemaMismatchErrorDetails`, `MigrationErrorDetails`,
`EagerMaterializationErrorDetails`, `StaleVersionErrorDetails`,
`SchemaContentConflictErrorDetails`, `StoreNotInitializedErrorDetails`,
`DatabaseOperationErrorDetails`, `EmbeddingDimensionChangedErrorDetails`) are
exported from the package root alongside the existing
`ValidationErrorDetails`. Classes with intentionally open, per-call-site
details (`ConfigurationError`, `UnsupportedPredicateError`,
`CompilerInvariantError`, `BackendDisposedError`) are unchanged.
- [#206](https://github.com/nicia-ai/typegraph/pull/206) [`995b964`](https://github.com/nicia-ai/typegraph/commit/995b9643927f498deabb157113bcb5ecb5883ca9) Thanks [@pdlug](https://github.com/pdlug)! - perf: cascade deletes batch their edge removals — new optional `GraphBackend.deleteEdgesBatch` / `hardDeleteEdgesBatch` members issue one statement per bind-budget chunk instead of one per connected edge (50-edge cascade on local PostgreSQL: 24.4ms → 3.6ms), with recorded-time capture preserved. `getOrCreate` variants no longer run the full Zod parse twice on the create leg.
### Patch Changes
- [#247](https://github.com/nicia-ai/typegraph/pull/247) [`191e877`](https://github.com/nicia-ai/typegraph/commit/191e877796fde30ad606993948decea7305fd367) Thanks [@pdlug](https://github.com/pdlug)! - Fix: the synthetic CTE column names that carry selectively-extracted `props`
fields are now bounded to PostgreSQL's identifier limit.
A selected top-level `props` field is extracted once inside the CTE that owns it,
under a generated column name encoding the query alias and the field name. The
encoding was unambiguous but unbounded, and PostgreSQL silently truncates
identifiers at 63 **bytes** — so two distinct `(alias, field)` pairs sharing a
long prefix could collapse onto one column name after truncation, yielding an
ambiguous-column error or the wrong value.
Long names are now truncated on a UTF-8 character boundary and disambiguated with
a hash of the full, untruncated pair — the same guard the sibling subgraph
projection path already used, now extracted into one shared helper. Names that
already fit are emitted unchanged, so compiled SQL for ordinary queries is
byte-for-byte what it was.
- [#247](https://github.com/nicia-ai/typegraph/pull/247) [`191e877`](https://github.com/nicia-ai/typegraph/commit/191e877796fde30ad606993948decea7305fd367) Thanks [@pdlug](https://github.com/pdlug)! - Document a semantic consequence of batched writes: within **one backend batch
call**, every row whose timestamp TypeGraph generates shares a single instant,
sampled once for that call — not once per row, and not once per bind-budget
chunk. `bulkCreate()` and `bulkInsert()` issue one such call, so all of their
rows tie.
Creating the same rows one at a time through `create()` gives each its own
timestamp, so `ORDER BY created_at` was a total order there and is only a
partial one after a bulk write. Two things it is **not** safe to conclude:
- **`importGraph()` is not one instant.** It slices nodes and edges into
`batchSize` batches and drives one backend call per slice, so each slice
samples its own timestamp. Rows that carry an explicit `validFrom` in the
import payload keep it verbatim; only generated defaults are affected.
- **Ids are not a sequence.** The default generator is a random NanoID, and
callers may supply arbitrary ids, so `ORDER BY id` is not insertion order.
`(created_at, id)` is a _deterministic_ tiebreak, not a chronology. If input
order matters, persist an explicit sequence column.
One instant per batch call is the intended semantics — it is what makes a bulk
write a single point in valid time rather than a smear — and it is the same
choice `valid_from` already made. Nothing changes in behavior; this note exists
because the batching work that landed this release moved several paths onto it.
- [#248](https://github.com/nicia-ai/typegraph/pull/248) [`c379045`](https://github.com/nicia-ai/typegraph/commit/c37904505ee0cf17a9a62f4f7e6769be61319670) Thanks [@pdlug](https://github.com/pdlug)! - Perf: cache compiled query SQL across executions again, without freezing the
read instant.
The read-freshness fix recompiled a query's full AST to SQL on every
`execute()` so a reused or prepared query would always see the latest rows.
That kept results fresh but made the recommended `.prepare()`-once-`.execute()`-
many pattern pay a full compile per call (a point lookup ~58µs, a three-hop
traversal ~450µs of pure JS compilation).
Only the bound "current" read instant varies between two compilations of the
same query; the SQL text is identical. So a query now compiles once into a
cached statement whose read instant is a reserved execution-time placeholder,
and each execution fills a fresh instant into it and runs the cached text
directly. Repeated point-query execution drops from ~47µs to ~2.4µs (near the
raw-execution floor) while staying just as fresh — a row created after
`prepare()` or the first `execute()` is still visible on the next call.
The cache applies to `ExecutableQuery`, prepared queries, aggregate queries,
and set operations, on backends that can compile and run raw SQL text
(synchronous SQLite and PostgreSQL backends); other backends — including async
SQLite profiles that do not expose `executeRaw` — fall back to per-call
recompilation unchanged. Statements whose execution depends on the compiled
SQL object — pgvector approximate-scan GUC tuning and parameter-blind-plan
avoidance — keep running through the standard execution path. `param()` now
rejects the reserved read-instant name, and aggregate queries (which have no
`.prepare()`) reject `param()` with clear guidance instead of a downstream
binding error.
- [#244](https://github.com/nicia-ai/typegraph/pull/244) [`b38a537`](https://github.com/nicia-ai/typegraph/commit/b38a537d1f0abcb5925a94b7e0845fb1184509ff) Thanks [@pdlug](https://github.com/pdlug)! - Fix: "current" temporal reads now evaluate validity against the application
clock, not the database clock — repairing a read-after-write consistency
violation on Postgres.
`valid_from` is stamped from the application clock (`Date.toISOString()`) on
write, but a "current" read compiled its validity filter against the database
clock (`valid_from <= NOW()` on Postgres). On any deployment where the
application-server clock runs ahead of the database-server clock — i.e. the
app and database on separate hosts, which is the norm — a freshly-created node
or edge could be missing from the very "current" read that immediately
followed its creation, until the database clock caught up. SQLite (a single
in-process clock) was never exposed.
The "current" read now binds the application clock (`nowIso()`) as a
parameter — the same clock `valid_from`, the facade search-currency filter,
and the recorded/logical clock already use — across every current-read path
(standard and recursive queries, subgraph extraction, graph algorithms, and
recorded-time reads). The temporal-visibility clock is now a single source.
Because the current-read instant is no longer dialect-specific, the internal
`DialectAdapter.currentTimestamp()` seam has been removed.
**Know the consistency model this buys you.** Reads and writes now share one
clock — _the clock of the process that issued them_. Read-after-write
consistency therefore holds **per application process**: a node you just
created is visible to the very next current read from that same process,
which is the guarantee the bug broke. It does **not** extend across processes.
Two application servers with skewed clocks, writing to one PostgreSQL
database, can still miss each other's fresh rows: a row stamped
`valid_from = T` by the server that runs ahead stays invisible to a current
read from the server that runs behind until its own clock passes `T`. The
window equals the skew between the two application hosts, not between an
application host and the database. If you need cross-process read-after-write
consistency, keep application clocks disciplined (NTP), or read at an explicit
`asOf` coordinate rather than `current`.
- [#247](https://github.com/nicia-ai/typegraph/pull/247) [`191e877`](https://github.com/nicia-ai/typegraph/commit/191e877796fde30ad606993948decea7305fd367) Thanks [@pdlug](https://github.com/pdlug)! - Fix: `store.algorithms.degree()` undercounted edges written before an endpoint
declaration changed.
To let the composite edge indexes seek — both lead with the endpoint kind
column, so a bare `from_id = ?` cannot — the direction filter supplied the
missing kind equality by enumerating the endpoint kinds the _graph declaration_
permits for the counted edge kinds. That enumeration is complete only for rows
written under the current declaration. Narrow `knows` from `from: [Person]` to
`from: [Employee]`, and every `Person`-rooted `knows` edge already on disk drops
out of the filter: `degree()` silently returns a number too small, with no error
and no warning.
The filter now derives the kind from the counted node itself, via an
uncorrelated scalar subquery. This is exact by construction: an edge row stores
the _actual_ kind of each endpoint node (the write path copies it off the
endpoint reference) and a node's kind is immutable for the life of its id, so
for any edge incident to a node, the endpoint kind on that node's side is that
node's kind and nothing else — however the declaration later evolves.
It is also a better filter. An equality on one kind replaces an `IN` list over
every declared endpoint and its `subClassOf` descendants, and both engines hoist
the uncorrelated subquery to a constant (a Postgres InitPlan, a SQLite one-shot
scalar subquery), so the seek is unchanged. `EXPLAIN QUERY PLAN` still shows
`typegraph_edges_from_idx` / `_to_idx` seeks with no partition scan.
`degree()` of an id that names no node is `0`, as before.
- [#200](https://github.com/nicia-ai/typegraph/pull/200) [`472ac1c`](https://github.com/nicia-ai/typegraph/commit/472ac1c20a6751a52121da8732f6c562fe5124c8) Thanks [@pdlug](https://github.com/pdlug)! - `degree()` direction filters are now shaped for the default edge indexes.
The filters previously compiled to bare `from_id = ?` / `to_id = ?`, which
neither composite edge index can seek (both lead with the endpoint kind
column) — so degree counts relied on engine-specific rescue: SQLite
skip-scan (only with fresh statistics) or PostgreSQL 18's new btree skip
scan, and degenerated to partition scans everywhere else (PostgreSQL ≤ 17,
SQLite with stale statistics).
The filters now enumerate the endpoint kinds the graph declaration permits
for the counted edge kinds, expanded through the subClassOf closure — the
same set edge writes validate against — making `edges_from_idx` /
`edges_to_idx` structurally seekable on every engine and version.
Measured on PostgreSQL 18 (where the old form was already skip-scan
rescued): 0.30ms → 0.06ms per call; on older PostgreSQL the old form
could not use these indexes at all. An edge set that declares no endpoint
kinds on the required side now returns 0 without a round trip.
Behavior note: because the counted set is now restricted to edges whose
stored endpoint kind falls within the declaration's `subClassOf` closure,
`degree()` no longer counts an edge whose stored `from_kind` / `to_kind`
lies _outside_ that closure — e.g. a row written before the endpoint
declaration was narrowed, or written directly through the backend bypassing
endpoint validation. This matches how typed traversals already treat such
rows (invisible to a schema-consistent read), but it is a change from the
previous "count every edge touching this node regardless of stored kind"
behavior.
- [#220](https://github.com/nicia-ai/typegraph/pull/220) [`7b48543`](https://github.com/nicia-ai/typegraph/commit/7b4854310fc042410e31f2e14abc19a9e61e44a2) Thanks [@pdlug](https://github.com/pdlug)! - Edge delete, edge hard delete, and node hard delete no longer re-read
the row inside the write transaction. The in-transaction preflight was
pure round-trip fat on these paths: nothing consumed the row, and the
writes are already concurrency-correct on their own — the tombstone
UPDATE is guarded by `deleted_at IS NULL` and the hard deletes are
id-keyed and idempotent, so a row deleted concurrently between the
outside gate and the write lock degrades to a 0-row no-op with
identical observable behavior (verified including recorded-time history
under a deliberately staled gate). One less statement per delete
(~20% of the per-op round trips on client/server engines). Node SOFT
delete keeps its preflight deliberately: its pipeline consumes the
pre-image for uniqueness-key cleanup, now documented in place.
- [#227](https://github.com/nicia-ai/typegraph/pull/227) [`09754a6`](https://github.com/nicia-ai/typegraph/commit/09754a6e4435425e8a55e9a0b991fcbd66daccbf) Thanks [@pdlug](https://github.com/pdlug)! - Batches edge creation's endpoint-existence checks in `bulkCreate`/`bulkInsert`
into one `getNodes` call per distinct (kind) referenced across the whole
batch, instead of an individual `getNode` probe per edge (mirroring the
batched existence/uniqueness pre-check node creation already had via
`primeBatchValidationCaches`). Found while investigating why a real
LDBC SNB SF1 bulk load (millions of nodes and edges) was far slower than
expected: a controlled 1M-row reproduction showed `bulkInsert` edge-batch
time growing from ~90ms to ~630ms per 2,000-row batch as the graph grew,
while an equivalent node-only batch (no edges) stayed roughly flat. The
edge batch path validated each edge's `from`/`to` endpoints with a
`getNode` call per edge — for a batch with mostly-unique endpoints, that's
thousands of individual round trips per batch instead of one batched
fetch per distinct node kind. With the fix, the same 1M-edge reproduction's
per-batch time drops to roughly ~90-160ms and its growth curve flattens
substantially (the residual growth matches the same mild index-maintenance
cost already seen on plain node inserts). No behavior change: this is a
pure internal optimization to `executeEdgeCreateNoReturnBatch`/
`executeEdgeCreateBatch`; callers observe identical results, just fewer
round trips.
- [#245](https://github.com/nicia-ai/typegraph/pull/245) [`ef6def6`](https://github.com/nicia-ai/typegraph/commit/ef6def6b67e306a9cdb40e78723dad6d36f89647) Thanks [@pdlug](https://github.com/pdlug)! - The default edge traversal indexes (`{table}_from_idx` / `{table}_to_idx`,
created for every graph on both SQLite and PostgreSQL) were missing two
things a traversal join needs to be served fully index-only:
- **`valid_from`** — one of the three system columns every compiled
query's soft-delete / temporal-validity predicate checks (`deleted_at`
and `valid_to` were already covered; `valid_from` wasn't).
- **The join's target-id column** — a compiled traversal reads `n.id =
e.to_id` for an outgoing traversal, or `n.id = e.from_id` for an
incoming one (`standard-builders.ts`), but neither index carried the
_other_ endpoint's id column, so the join to the target node still
required a heap-row fetch even once the predicate columns above were
covered.
Both gaps produce the same symptom: SQLite's plan reads `USING INDEX`,
never `USING COVERING INDEX`, so every candidate edge pays a heap-row
fetch. That fetch is free while the table fits in the page cache. Once it
doesn't — a real LDBC SNB benchmark run measured this at 10x data volume,
where the nodes table outgrew available cache — every one of those
fetches becomes a genuine random disk read, and with thousands of
candidates per traversal that alone produced a multi-second/minute
latency cliff on an otherwise sub-millisecond query shape. Both indexes
now carry all five columns beyond their existing seek prefix
(`deleted_at`, `valid_from`, `valid_to`, plus the other endpoint's id),
confirmed via `EXPLAIN QUERY PLAN` against the actual SQL `execute()`
sends (not `toSQL()`'s wider, unoptimized output) to flip to `USING
COVERING INDEX`.
**Existing databases get none of this until you rebuild the indexes.**
The widened indexes materialize on **fresh databases only**.
`generateSqliteMigrationSQL()` / `generatePostgresMigrationSQL()` emit
`CREATE INDEX IF NOT EXISTS` under the _same index name_, and that is a
no-op against an index that already exists — regardless of how the column
list changed. An upgraded deployment silently keeps its narrow index, and
keeps the latency cliff, until it runs the rebuild below. Upgrading the
package is not enough; there is no automatic migration.
```sql
-- SQLite: no CONCURRENTLY equivalent; drop and let the next migration
-- run (generateSqliteMigrationSQL(), or a createStoreWithSchema boot,
-- which re-issues idempotent DDL) recreate them.
DROP INDEX IF EXISTS typegraph_edges_from_idx;
DROP INDEX IF EXISTS typegraph_edges_to_idx;
-- PostgreSQL: CREATE INDEX CONCURRENTLY does not block writes, but it
-- cannot run inside a transaction and needs its own connection. Rename
-- the old index out of the way first so the new one can use the
-- production name without a window where neither exists.
ALTER INDEX typegraph_edges_from_idx RENAME TO typegraph_edges_from_idx_old;
CREATE INDEX CONCURRENTLY "typegraph_edges_from_idx" ON "typegraph_edges"
("graph_id", "from_kind", "from_id", "kind", "to_kind", "deleted_at", "valid_from", "valid_to", "to_id");
DROP INDEX CONCURRENTLY typegraph_edges_from_idx_old;
ALTER INDEX typegraph_edges_to_idx RENAME TO typegraph_edges_to_idx_old;
CREATE INDEX CONCURRENTLY "typegraph_edges_to_idx" ON "typegraph_edges"
("graph_id", "to_kind", "to_id", "kind", "from_kind", "deleted_at", "valid_from", "valid_to", "from_id");
DROP INDEX CONCURRENTLY typegraph_edges_to_idx_old;
```
- [#217](https://github.com/nicia-ai/typegraph/pull/217) [`fce0a0f`](https://github.com/nicia-ai/typegraph/commit/fce0a0f18b90e7b6f5b5d395681231865b21fb52) Thanks [@pdlug](https://github.com/pdlug)! - Non-approximate `.similarTo()` is now genuinely exact when an ANN index
exists. pgvector serves any `ORDER BY embedding <=> q LIMIT k` from a
matching HNSW/IVFFlat index, so after `materializeIndexes()` the
default (non-approximate) inline vector predicate silently returned
approximate results — measured recall 0.980 unfiltered and 0.000 under
a selective filter at 50k docs, where the index frontier starves at the
default ef_search and returns entirely wrong rows. The exact branch now
orders by `(distance + 0.0)`, which the index opclass cannot match,
forcing the true flat scan on every engine (numerically identity;
inert on SQLite/libSQL whose ANN forms are opt-in constructs).
Behavior change: exact queries that were silently index-served get
correct results and flat-scan latency (50k x 384 dims: ~39ms instead of
~23ms-but-wrong). The sanctioned fast path remains
`similarTo(..., { approximate: true })`, which is unchanged. The
`bench:vector` lane's `vector:exact-postindex-recall` and
`vector:exact-filtered-postindex-recall` rows now read 1.000.
- [#210](https://github.com/nicia-ai/typegraph/pull/210) [`76422c6`](https://github.com/nicia-ai/typegraph/commit/76422c64189baa2c83287a99a1fea6a13bbfe976) Thanks [@pdlug](https://github.com/pdlug)! - perf: PostgreSQL fulltext queries are now parsed with the kind's DECLARED language as a plan-time constant (the same winning-language rule the write path applies to rows), instead of referencing the per-row `language` column. The per-row form made every tsquery non-constant, so the GIN index on `tsv` could never serve a match and every search re-parsed the query per row — measured 12.9ms → 2.3ms at 5,000 docs for the parse elimination alone, with GIN service now possible as corpora grow. Applies to the facade and the inline `$fulltext` predicate; mixed-language subclass aliases and explicit per-query overrides behave as before.
- [#207](https://github.com/nicia-ai/typegraph/pull/207) [`5cbcb35`](https://github.com/nicia-ai/typegraph/commit/5cbcb35f9df972a6f36975b43adad2d7b110bfd1) Thanks [@pdlug](https://github.com/pdlug)! - perf: recorded-time capture acquires the PostgreSQL graph-write advisory lock once per transaction instead of once per captured write (`pg_advisory_xact_lock` is reentrant and held to transaction end, so the repeats were pure round trips). A 50-write recorded transaction drops from N+1 lock round trips to 1; measured 1.7× on the transaction shape.
- [#215](https://github.com/nicia-ai/typegraph/pull/215) [`0eb2fd8`](https://github.com/nicia-ai/typegraph/commit/0eb2fd8ba778fb6e9cf6469481805a1c8cd86647) Thanks [@pdlug](https://github.com/pdlug)! - The single-statement hybrid search now emits the candidates set
(liveness/currency filter, or the compiled `where` predicate query)
once, as a CTE shared by the vector and fulltext legs, instead of
embedding — and re-executing — a private copy inside each leg. The
duplicate evaluation was most expensive with a `where` filter, whose
compiled candidates query ran twice per search: measured on PostgreSQL,
filtered hybrid drops 26.5ms → 17.1ms at 5k docs (bench shape
11.8ms → 8.6ms; unfiltered 6.1ms → 4.9ms). This also removes a subtle
inconsistency where each leg stamped its own currency instant. SQLite
is unchanged within noise (in-process re-execution was cheap).
- [#247](https://github.com/nicia-ai/typegraph/pull/247) [`191e877`](https://github.com/nicia-ai/typegraph/commit/191e877796fde30ad606993948decea7305fd367) Thanks [@pdlug](https://github.com/pdlug)! - Fix: hybrid search's two execution paths agreed on scores but not on ties, and
neither was deterministic across PostgreSQL databases.
Relevance ranking breaks a score tie on `node_id`. Left bare, PostgreSQL sorts
that under the database's default text collation: an `en_US.UTF-8` database
orders `a, A, b, B` where byte order gives `A, B, a, b`. So the same query
returned different pages on two databases whose `datcollate` differed, and
disagreed with SQLite (whose `BINARY` collation is byte order) throughout.
Three seams had to move together, because a hybrid search's tiebreak decides the
page twice — once in the per-source ranks, and again in the fused ordering the
ranks produce:
- The single-statement hybrid search now renders `node_id COLLATE "C"` in both
per-source `ROW_NUMBER()` windows and in the final `ORDER BY`.
- The standalone fulltext search's `ORDER BY … , node_id` is C-collated too, so
the multi-statement fallback's fulltext ranks match.
- The fallback now re-ranks each leg's rows before assigning ranks, rather than
trusting the order the source SQL happened to return for a single kind. The
vector source breaks a distance tie arbitrarily — it carries no `node_id`
tiebreak, because a second sort key would cost pgvector its ordered index scan
— so its arrival order was never a sound basis for a rank. That re-rank sorts
with a new code-point comparator rather than JavaScript's UTF-16 code-unit
`<`, which disagrees with byte order for astral characters such as emoji.
All three orderings now coincide, and the single-statement and multi-statement
paths return identical hits, ranks, and scores even when every score ties.
Results only change where they were previously non-deterministic.
- [#213](https://github.com/nicia-ai/typegraph/pull/213) [`a243f3b`](https://github.com/nicia-ai/typegraph/commit/a243f3bc323f8d7377454f06c2349fb87386963c) Thanks [@pdlug](https://github.com/pdlug)! - `importGraph`'s default `batchSize` is now 1,000 (was 100), and the
default now actually applies: options are parsed through
`ImportOptionsSchema` at the function boundary, so direct calls that
omit fields with schema defaults (e.g. `{ onConflict: "error" }`)
resolve them instead of reading `undefined`. `ImportOptions` is now the
schema's input type — fields with defaults are optional for callers.
Each import batch pays fixed per-round-trip costs (existence probe,
unique pre-check, one multi-row insert), so the old default dominated
import time on client/server engines: a 20k-node + 5k-edge import on
PostgreSQL drops from 1,515ms to 781ms (16.5k → 32k entities/s).
SQLite imports are insensitive to the value (in-process, no round
trips). Explicit `batchSize` values are unaffected.
Fulltext batch upserts and deletes are now split by the driver's
bind-parameter budget in the backend wrappers, like node/edge/unique
inserts already were. Previously a searchable import slice emitted ONE
FTS5 (or tsvector) statement over every row — 6 binds per row, so a
1,000-row slice overflowed SQLite's 999-bind fallback ceiling and D1's
~100-bind cap, and 6,000-row slices overflowed even better-sqlite3's
32,766 budget ("too many SQL variables").
- [#221](https://github.com/nicia-ai/typegraph/pull/221) [`9b61809`](https://github.com/nicia-ai/typegraph/commit/9b618098b6c6f4917f79a23f4b1f0477428de0b3) Thanks [@pdlug](https://github.com/pdlug)! - Inline `.similarTo(..., { approximate: true })` now actually uses the
ANN index on PostgreSQL. Two defects compounded: the candidates
membership subquery carried a `DISTINCT` that kept the planner off the
ordered index scan entirely (even `enable_seqscan = off` could not
rescue it — duplicates are irrelevant to `IN` membership, so the
DISTINCT bought nothing), and the inline path never applied the
pgvector GUCs the search facade uses, so even an index-served filtered
scan would have starved at the default ef_search frontier. The compiler
now emits duplicate-tolerant membership candidates for the engine-form
branch and brands ANN-bearing statements; the PostgreSQL backend wraps
branded statements with the facade's GUC overrides
(`hnsw.iterative_scan = strict_order` / `ivfflat.iterative_scan =
relaxed_order` on transaction-capable drivers with pgvector >= 0.8;
the settings are transaction-scoped, so non-transactional backends
such as neon-http keep the plain bounded scan). Set operations merge
operand brands onto the combined statement, so a union with an
approximate operand is wrapped too. Measured at 50k x 384 dims:
unfiltered approximate 174ms -> 2.1ms (recall 0.995), filtered
approximate 3.8ms at recall 1.000 on filter-independent corpora. The
JOIN consumers of the scoped candidates (exact branch, fulltext CTE)
keep their DISTINCT — a join does multiply rows on duplicates — and the
non-approximate path's exactness guarantee is untouched.
- [#224](https://github.com/nicia-ai/typegraph/pull/224) [`b5886cd`](https://github.com/nicia-ai/typegraph/commit/b5886cdad183dcba80586344935278a79f9ed795) Thanks [@pdlug](https://github.com/pdlug)! - Document external event-log materialization patterns and verify the
export/import bulk-copy path into graph-merge branches.
- [#199](https://github.com/nicia-ai/typegraph/pull/199) [`d01d6c7`](https://github.com/nicia-ai/typegraph/commit/d01d6c76be56efb393a3cd5506e6a5690995c409) Thanks [@pdlug](https://github.com/pdlug)! - Subgraph extraction is ~4× faster on PostgreSQL. The final node/edge
fetches filtered ids with `IN (SELECT id FROM included_ids)`; PostgreSQL
pulls that form up into a join whose recursive-CTE row estimate (~10 rows
for a single-row seed) drives the planner into a nested-loop join filter —
measured at ~10 million discarded rows on the depth-3 benchmark shape.
PostgreSQL now evaluates membership against the materialized closure ids
with a parameterized `text[]` semi-join
(`EXISTS (SELECT 1 FROM unnest($ids) AS t(id) WHERE t.id = column)`) rather
than pulling the recursive CTE into that join; SQLite keeps `IN (subquery)`,
which it already evaluates optimally.
Measured (benchmark suite, 1,200 users / depth-3 stress shape): PostgreSQL
subgraph full hydration 322ms → 82ms, depth-2 11.5ms → 7.1ms; SQLite
unchanged.
- [#247](https://github.com/nicia-ai/typegraph/pull/247) [`191e877`](https://github.com/nicia-ai/typegraph/commit/191e877796fde30ad606993948decea7305fd367) Thanks [@pdlug](https://github.com/pdlug)! - Fix: serialize the statements TypeGraph issues on a transaction's pinned
Postgres connection, so its own graph writes never present two queries to one
connection at once.
A transaction pins one connection, and the PostgreSQL wire protocol carries one
statement at a time. node-postgres hid that behind an internal queue, deprecated
it in `pg@8.22` ("Calling client.query() when the client is already executing a
query is deprecated and will be removed in pg@9.0. Use async/await or an
external async flow control mechanism instead"), and removes the queue in
`pg@9`. TypeGraph overlapped statements on a pinned connection in two ways:
- **Always on, no user concurrency required.** The node write pipeline issues
`Promise.all([syncEmbeddings, syncFulltext])` for any schema that has both a
`searchable()` field and an `embedding()` field, so every single `create()`,
`update()`, or resurrect on such a schema put two statements on the wire.
- **User-driven.** `store.transaction(async (tx) => { await Promise.all([...]) })`
is a documented, recommended pattern.
Transaction-scoped backends now run every statement they issue through a
per-connection queue. Concurrency at the API surface is unchanged — a
`Promise.all` of graph writes still works, and on a pooled (non-transactional)
backend the statements still run genuinely concurrently. The queue serializes
only what already had to be serial. A multi-statement `SET LOCAL`-scoped vector
search (snapshot / set / select / restore) runs as one exclusive group, so two
concurrent searches can no longer interleave and apply each other's `efSearch`.
The transaction boundary also **drains and closes** the queue before the driver
emits `COMMIT` / `ROLLBACK`. Those control statements do not travel through the
queue, so without the drain a rollback could overlap a live statement. And a
callback that rejects out of a `Promise.all` leaves its siblings running: their
statements would otherwise land on the connection _after_ the pool had reclaimed
it, executing inside an unrelated transaction. Such a statement is now refused
with a new `TransactionClosedError` (normally invisible — `Promise.all` has
already rejected with the original failure and discards this one).
**Scope: the queue mediates only TypeGraph's own statements.** The raw Drizzle
handle exposed as `tx.sql` (for writing your own relational tables in the same
atomic boundary) bypasses it. Running a raw statement concurrently with a graph
write — or with another raw statement — still races on the one pinned
connection, and `drainAndClose` cannot wait for a raw statement it never saw.
Await each `tx.sql` statement before the next write; this is inherent to a
single-connection transaction, not something TypeGraph can enforce over a handle
it doesn't mediate. `adoptTransaction()` likewise serializes the statements it
issues but never closes the queue — the caller owns that transaction's end.
- [#219](https://github.com/nicia-ai/typegraph/pull/219) [`ee93b77`](https://github.com/nicia-ai/typegraph/commit/ee93b77581e6bcbddccf5256dbb2b321b827e361) Thanks [@pdlug](https://github.com/pdlug)! - Statements whose good plan depends on their parameter values (the
subgraph id-array fetches, marked internally with the custom-plan
brand) now opt out of statement preparation per call on the postgres-js
driver too, via `sql.unsafe(text, params, { prepare: false })`.
Previously postgres-js prepared them like everything else, so after
five executions PostgreSQL flipped them to a generic, parameter-blind
plan — the same cliff fixed for node-postgres in the subgraph
shared-traversal change (measured there: 21ms → 310ms on the edge
fetch). Scalar-parameter statements keep the driver's prepared default.
- [#246](https://github.com/nicia-ai/typegraph/pull/246) [`d5aafe8`](https://github.com/nicia-ai/typegraph/commit/d5aafe845f95a503070ac485994afb46b3a82cac) Thanks [@pdlug](https://github.com/pdlug)! - **Critical fix**: `.prepare()`d queries, and any `ExecutableQuery`/`UnionableQuery`/`ExecutableAggregateQuery` instance whose `.execute()` was called more than once, could silently miss rows created after the query was first compiled.
A "current" (live) temporal-validity read binds its read instant (`currentReadInstant()`) at SQL compile time. All four query-builder classes cached their compiled SQL text across calls — `.prepare()` compiled once and every subsequent `execute({...})` reused that same SQL text, and a reused `ExecutableQuery`/`UnionableQuery`/`ExecutableAggregateQuery` instance cached its first `.execute()`'s compilation the same way. Both patterns froze "now" at the moment of first compilation: any row created afterward had a `valid_from` later than the frozen instant, so `valid_from <= now` silently evaluated to false for it, for the query's entire remaining lifetime.
This is a regression introduced by the `current-read-app-clock` fix (the [#242](https://github.com/nicia-ai/typegraph/issues/242) clock-skew correction): the prior behavior (`NOW()` / `strftime('now')`, evaluated fresh by the database on every execution) did not have this problem. It is more severe than [#242](https://github.com/nicia-ai/typegraph/issues/242) — that bug required app/DB clock skew across separate hosts; this one reproduces unconditionally, in a single process, on the very next insert after a query is prepared or first executed. `.prepare()`-once-`.execute()`-many is this library's own documented, recommended pattern, so this affected the common case, not an edge case.
**Fix**: none of the four classes cache compiled SQL text across calls anymore — each `execute()`/`compile()`/`toSQL()` call recompiles fresh, so `currentReadInstant()` is re-evaluated every time. `.prepare()` still builds and structurally validates the query AST once (so a malformed query still fails fast, before the first `execute()`); only the SQL-text compilation moved from prepare-time to each execute-time call. `param()`-bound values are unaffected — those were already correctly re-bound per call.
- [#209](https://github.com/nicia-ai/typegraph/pull/209) [`5e24882`](https://github.com/nicia-ai/typegraph/commit/5e24882536a242d75a2ec9973bfb0301027da92c) Thanks [@pdlug](https://github.com/pdlug)! - perf: facade search candidate handling planned poorly at scale. The hybrid statement's fused CTE is now MATERIALIZED (PostgreSQL inlines single-use CTEs, re-executing the fusion subtree once per candidate node row under a nested-loop join), and unfiltered facade searches use a flat, parameter-bound current-read candidates subquery instead of a compiled builder query whose per-row SQL clock calls dominated on SQLite. Semantics are unchanged — validity windows and tombstones are still enforced, with the instant bound as a parameter. Only searches with a `where` predicate compile a builder query as candidates; `includeSubClasses` expands at the store level and each concrete kind uses the flat form.
- [#202](https://github.com/nicia-ai/typegraph/pull/202) [`b45cfc3`](https://github.com/nicia-ai/typegraph/commit/b45cfc3e6d141a6f037544572f862f00c27d5571) Thanks [@pdlug](https://github.com/pdlug)! - fix: facade search (`store.search.vector` / `fulltext` / `hybrid`) now computes top-k over live nodes in SQL. Previously the search statement ranked side-table rows alone and hydration dropped tombstoned ids afterward, silently returning fewer than `limit` hits under index drift. Liveness is pushed into the KNN/MATCH SQL on every engine — exact on pgvector ≥0.8 (HNSW via `hnsw.iterative_scan = strict_order`; IVFFlat via `ivfflat.iterative_scan = relaxed_order` with an in-statement re-sort), sqlite-vec (vec0 primary-key `IN` pushdown), tsvector, and FTS5; libSQL DiskANN over-fetches 4× and post-filters (documented recall bound).
- [#237](https://github.com/nicia-ai/typegraph/pull/237) [`48f324b`](https://github.com/nicia-ai/typegraph/commit/48f324b905c9d0e2aa52371780e3c443b596040a) Thanks [@pdlug](https://github.com/pdlug)! - Fixes `.select()` query projections losing the `NodeId` brand on node `id`
fields. Previously `ctx.alias.id` in a `.select()` callback was typed as plain
`string`, so feeding a projected node id back into `getById`/`getByIds`
required an unsafe cast (`as never` or worse). `SelectableNode.id` is now
typed `NodeId`, matching what `getById`/`getByIds` already require — no
runtime change, no cast needed.
Edge ids from `.select()` stay plain `string` on purpose: `traverse()`
defaults to `expand: "inverse"`, which can back an edge alias with a row of
the registered _inverse_ edge kind, so the alias's static edge type doesn't
reliably describe the row. Use `asEdgeId` to re-brand a projected edge id
before a point read.
- [#247](https://github.com/nicia-ai/typegraph/pull/247) [`191e877`](https://github.com/nicia-ai/typegraph/commit/191e877796fde30ad606993948decea7305fd367) Thanks [@pdlug](https://github.com/pdlug)! - Fix: a set operation now binds one "current" read instant across all of its
operands.
`UNION` / `INTERSECT` / `EXCEPT` compile each operand independently, and each
operand compiled its own temporal-validity filter from a fresh `nowIso()`
sample. A compound `SELECT` is evaluated against a single snapshot, so two
samples microseconds apart let the two halves of an `INTERSECT` or `EXCEPT`
disagree about whether a row created between them is current — a row could
satisfy the left operand's `valid_from <= now` and not the right's.
Compilation of a set operation (including nested ones) now runs under a single
pinned instant. Ordinary single-leaf queries were already consistent — they bind
one instant per compile — and are unaffected.
- [#226](https://github.com/nicia-ai/typegraph/pull/226) [`4cd6b4c`](https://github.com/nicia-ai/typegraph/commit/4cd6b4ca8275c2dad53d85c085347814528b3074) Thanks [@pdlug](https://github.com/pdlug)! - Fixes a scaling bug in the SQLite backend's `refreshStatistics()` (the
planner-statistics refresh `bulkCreate`/`bulkInsert` trigger automatically
after a large autocommit write — see the `autoRefreshStatistics` store
option). It ran a bare, unscoped `ANALYZE`, which does two things wrong on
SQLite: it re-analyzes every table in the database file (not just
TypeGraph's own tables — already fixed on the Postgres backend), and it
does a full, unbounded table/index scan per call (Postgres's `ANALYZE`
samples a fixed-size set of rows regardless of table size; SQLite's does
not unless bounded). A caller streaming a bulk load through repeated
`bulkInsert()` calls — the only practical way to load a multi-million-row
dataset without holding it all in memory — re-triggers this once each
batch's row count crosses the threshold; with unbounded per-call cost
growing with total table size, total load time integrated to O(n²)
instead of O(n) (observed: a 2M-row bulk load that never finished after
4.5+ hours). `refreshStatistics()` on SQLite now scopes ANALYZE to
TypeGraph's own tables and sets `PRAGMA analysis_limit` first, bounding
each call's cost the way Postgres's already was. A 100k-row reproduction
of the original shape now completes in ~8s with load time growing
log-ishly with table size (2x from first batch to last), not
quadratically.
- [#218](https://github.com/nicia-ai/typegraph/pull/218) [`b601484`](https://github.com/nicia-ai/typegraph/commit/b601484e95f11f61d4b086f493a95e2b0c4f9c18) Thanks [@pdlug](https://github.com/pdlug)! - Non-approximate `.similarTo()` on SQLite now routes through sqlite-vec's
vec0 KNN form. vec0's KNN is brute-force in C — exact by construction —
so the default path keeps identical results (pinned against
JS-computed ground truth) while dropping from the SQL distance scan to
engine speed: 489ms → 124ms for top-10 over 50k 384-dim embeddings.
Declared via a new `searchIsExact` flag on the vector-strategy
contract; pgvector and libSQL leave it unset (their engine forms are
approximate) and are unchanged. The metric gate still applies: an
explicit metric override that differs from the slot's declared metric
falls back to the SQL scan, which is correct for any metric.
- [#211](https://github.com/nicia-ai/typegraph/pull/211) [`a216569`](https://github.com/nicia-ai/typegraph/commit/a21656906eec3cfc532200b1709d6356e6047d71) Thanks [@pdlug](https://github.com/pdlug)! - Subgraph extraction on PostgreSQL now runs the recursive traversal once
instead of twice. The node and edge fetches previously each embedded the
full recursive CTE; the closure ids are now fetched in one statement and
passed to both fetches as a single `text[]` parameter, filtered via an
`EXISTS` semi-join over `unnest`. Those id-filtered fetches execute as
unnamed statements so PostgreSQL plans them against the actual array on
every call — a named prepared statement flips to a generic plan after
five executions, which mis-plans array-cardinality-dependent filters
(measured 21ms → 310ms on the edge fetch). Depth-3 stress subgraph
(1,109 nodes / 4,513 edges, wide payloads): 82.9ms → 30.9ms full
hydration, 72.3ms → 15.6ms with SQL projection. SQLite keeps its
existing single-statement-per-fetch form, which is already optimal for
an in-process engine.
- [#234](https://github.com/nicia-ai/typegraph/pull/234) [`d042a30`](https://github.com/nicia-ai/typegraph/commit/d042a304979ea32f5777480b2cd28a8a02b1f339) Thanks [@pdlug](https://github.com/pdlug)! - perf: push selected top-level `props` field extractions into the
start/traversal CTEs instead of carrying the whole raw `props` JSONB/JSON
column outward for later extraction at the final projection. Each
selected field is extracted once, inline, as its own typed CTE column
(named from a length-prefixed encoding of its alias and field, so
distinct alias/field pairs can never collide on the same column name);
the outer projection and any matching `ORDER BY` on the same field just
reference that column directly instead of re-extracting from a
carried-forward `_props` column.
Found while investigating why a covering index on a system column (see
`keySystemColumns`) still couldn't get Postgres to serve an indexed join
index-only: the compiled query was asking for the entire `props` column
in the join step even though the final `.select()` only needed one
extracted field, so the specific indexed expression was never actually
what got read from the table. No behavior change: compiled query results
are identical; this only changes which columns each CTE carries and
where field extraction happens.
- [#242](https://github.com/nicia-ai/typegraph/pull/242) [`6b884b6`](https://github.com/nicia-ai/typegraph/commit/6b884b66b3f642bfc2a65064f51c63ce317c4cc9) Thanks [@pdlug](https://github.com/pdlug)! - Fix: creating a node or edge without an explicit `validFrom` now stamps the
operation's own creation timestamp instead of storing SQL `NULL`.
`NULL` is interpreted by temporal filters as open-left validity ("valid
since forever"), so a record created without `validFrom` was visible at
_any_ historical `asOf` instant — including ones before the record existed.
This contradicted the documented contract ("omitted `validFrom` defaults to
now") and is fixed at the insert layer for every write path: `create`,
`createFromRecord`, `upsertById`/`upsertByIdFromRecord` (create branch),
`bulkCreate`, `bulkInsert`, `bulkUpsertById`, and get-or-create, for both
nodes and edges.
`branch()`'s working-copy clone now also exports with `includeTemporal:
true`, so a fork's `validFrom`/`validTo` exactly match the base's — without
this, the clone would re-stamp any implicit `validFrom` to the fork's own
(later) creation time, narrowing the fork's valid-time window relative to
the base it was cloned from. This includes rows that still have a `NULL`
`valid_from` (predating this fix, or written directly via the backend):
`exportGraph`/`importGraph` now round-trip a confirmed open-left window as
an explicit `null` rather than silently dropping it, so a legacy row's
"valid since forever" semantics survive a clone unchanged instead of being
narrowed to the clone's own creation time.
`exportGraph`/`importGraph` round trips still default `includeTemporal` to
`false`; without it, imported records get a fresh `validFrom` at import
time rather than the source's original value (see the Interchange docs).
Custom `GraphBackend` implementations that build their own inserts (rather
than reusing the bundled Drizzle operation builders) should apply the same
rule: an omitted `validFrom` defaults to the row's creation instant, and an
explicit `null` is preserved as SQL `NULL` (open-left).
- [#214](https://github.com/nicia-ai/typegraph/pull/214) [`583fbb3`](https://github.com/nicia-ai/typegraph/commit/583fbb3782d78b16e07f92082da37ab299c3d966) Thanks [@pdlug](https://github.com/pdlug)! - PostgreSQL ANN index builds (`materializeIndexes()` on pgvector
HNSW/IVFFlat) now retry serially when the parallel build exhausts
shared memory. Parallel builds stage the index graph in dynamic shared
memory, and resource-constrained hosts — e.g. containers with the 64MB
`/dev/shm` default — reject the allocation with SQLSTATE class 53
(observed: 53100 from `dsm_impl_posix` on a 50k x 384-dim HNSW build).
The retry drops the INVALID leftover from the failed CONCURRENTLY
build, pins the vector table to `parallel_workers = 0`, rebuilds in
local memory, and restores the setting. Non-resource failures still
surface as before. Serial builds are slower — raise `/dev/shm` and
`maintenance_work_mem` where you control the host — but a slow index
beats a silently missing one.
## 0.34.0
### Highlights
TypeGraph 0.34 adds provenance-backed source retraction through `@nicia-ai/typegraph/provenance`. Applications map their graph kinds to sources, justifications, facts, premises, and derivations; retracting a source then makes unsupported reachable facts non-current while preserving their edges and recorded history. `unRetract` reverses the belief transition without treating it as a domain delete.
The release also strengthens transaction and import behavior: operation hooks report durably committed writes, conflicting updates preserve uniqueness reservations, tombstones survive update-mode imports, and incremental merges refuse inherited-row changes that occurred after planning. SQLite transaction handling and deterministic keyset pagination receive correctness fixes.
### Upgrade notes
- Provenance retraction requires `history: true` and captures TypeGraph-managed writes. Out-of-band SQL remains outside recorded capture.
- `onOperationEnd` now runs after the enclosing transaction commits. Consumers using hooks for metrics, cache invalidation, or audit events should account for the later notification boundary.
- Incompatible property-schema changes, including narrowed enums and nested type changes, are now classified as breaking migrations.
- `importGraph(..., { onConflict: "update" })` skips soft-deleted target rows. Use `onUnknownProperty: "allow"` for fidelity-preserving imports and `"strip"` when schema normalization is intended.
- An incremental merge that detects a changed inherited target row returns a retryable `BaseVersionMismatchError`; recompute the merge against current state.
### Minor Changes
- [#188](https://github.com/nicia-ai/typegraph/pull/188) [`0b0f4ea`](https://github.com/nicia-ai/typegraph/commit/0b0f4ea23ee2310cc2c160d24385eb94ebfdc5a8) Thanks [@pdlug](https://github.com/pdlug)! - Add the `@nicia-ai/typegraph/provenance` subpath for provenance-backed source
retraction. The first slice maps user graph kinds to source, justification,
fact, premise, and derivation roles; supports multiple source node kinds and
terminal fact kinds; requires `{ history: true }`; applies TypeGraph-managed
belief transitions by making unsupported facts non-current; and keeps
recorded-time replay available before and after retraction. A transition only
touches facts reachable from the flipped sources, and closing a fact's currency
is a belief-status change rather than a domain delete — the fact's edges are
left untouched (no `restrict`/`cascade`/`disconnect` enforcement), so
`unRetract` is an exact inverse of `retract`. PostgreSQL transitions serialize
with TypeGraph-managed history writes on the same graph; out-of-band SQL
remains outside recorded capture.
### Patch Changes
- [#188](https://github.com/nicia-ai/typegraph/pull/188) [`0b0f4ea`](https://github.com/nicia-ai/typegraph/commit/0b0f4ea23ee2310cc2c160d24385eb94ebfdc5a8) Thanks [@pdlug](https://github.com/pdlug)! - Stop opening a write transaction on `getOrCreateByConstraint`'s found path.
The single-item node getOrCreate wrapped its whole body — probe included — in
a transaction, so the common "already exists" case paid for `BEGIN IMMEDIATE`
on SQLite (and, under history capture, the per-graph advisory lock on
Postgres), and the nested create's operation hooks fired inside that outer
transaction, reporting success before a COMMIT that could still fail. The
probe now runs as a pure read; the create and update/resurrect legs each open
their own (hooked) transaction, so `onOperationEnd` means durably committed. A
concurrent create that reserves the key between the probe and the insert
surfaces as a uniqueness conflict and is converged by a single re-probe. The
bulk variant keeps its one enclosing transaction (atomic batch, hooks skipped
by design). Edge `getOrCreateByEndpoints` gets the same probe-first shape.
- [#191](https://github.com/nicia-ai/typegraph/pull/191) [`2cad229`](https://github.com/nicia-ai/typegraph/commit/2cad2293f2d937aff7f53a1318525814eeb05533) Thanks [@pdlug](https://github.com/pdlug)! - Guard `mergeIncremental()` against inherited-row lost updates. The incremental
commit path re-checked new-row identity resolution and per-row resurrect/strip
hazards, but not whether a committed row the plan mutates still held the value
the plan merged against — so a concurrent write to an inherited row between
planning (reads taken outside the transaction) and commit was silently
discarded. The commit now re-reads, in-transaction, every committed target row
the plan will change and aborts with a retryable `BaseVersionMismatchError` if it
drifted, matching the snapshot merge path's TOCTOU contract. This covers all four
mutating paths: node writes and node deletions (checked by `version`), and edge
upserts and edge deletions (checked by a content signature over endpoints,
liveness, and canonical props, since edges carry no version column).
- [#188](https://github.com/nicia-ai/typegraph/pull/188) [`0b0f4ea`](https://github.com/nicia-ai/typegraph/commit/0b0f4ea23ee2310cc2c160d24385eb94ebfdc5a8) Thanks [@pdlug](https://github.com/pdlug)! - `importGraph(..., { onConflict: "update" })` now skips soft-deleted target rows
instead of failing. Import never resurrects a tombstone: a node or edge that
exists only as a tombstone counts as `skipped`, keeps its tombstone, and gets no
uniqueness/embedding/fulltext side effects (a uniqueness reservation held by a
tombstoned node would block live creates of the same value). Previously the
update path attempted a live-row update that threw and aborted the whole
import. `onUnknownProperty: "allow"` is also pinned as the fidelity-preserving
strategy: it validates known fields but persists the given properties
byte-for-byte — no transform re-application, no default injection — so an
export→import round trip cannot corrupt values whose schema transforms are not
idempotent; use `"strip"` for a normalizing import.
- [#188](https://github.com/nicia-ai/typegraph/pull/188) [`0b0f4ea`](https://github.com/nicia-ai/typegraph/commit/0b0f4ea23ee2310cc2c160d24385eb94ebfdc5a8) Thanks [@pdlug](https://github.com/pdlug)! - Fix a uniqueness-reservation corruption on a conflicting node update.
`updateUniquenessEntries` mutated one constraint's sidecar at a time — releasing
the old key before proving the new one free — so a caller that catches the
resulting `UniquenessError` and still commits the transaction (notably
`importGraph(..., { onConflict: "update" })`, which reports the conflict per row)
left the node's already-mutated sidecars in a corrupt state: an earlier
constraint's old key released (letting a later create silently duplicate it) or a
new key wrongly reserved, while the row itself stayed unchanged. The update now
runs in two passes — preflight every changed constraint's new key first, then
apply all sidecar deletes and inserts only after every key is proven free — so a
conflict throws with zero partial writes, for every caller of the shared
node-write pipeline and for nodes with any number of unique constraints.
- [#188](https://github.com/nicia-ai/typegraph/pull/188) [`0b0f4ea`](https://github.com/nicia-ai/typegraph/commit/0b0f4ea23ee2310cc2c160d24385eb94ebfdc5a8) Thanks [@pdlug](https://github.com/pdlug)! - Make in-memory libsql databases safe across transactions, and fail loud on
re-entrant root access. Local `@libsql/client` connections (`file:` paths and
`file::memory:`) now frame transactions with raw `BEGIN IMMEDIATE`/`COMMIT` on
the client's single stable connection instead of `client.transaction()`, which
permanently hands that connection to the transaction and lazily opens a fresh —
for `:memory:`, empty — database afterwards
(tursodatabase/libsql-client-ts#229). Remote Turso connections keep using the
driver's per-stream transactions. Separately, a store-level operation awaited
from inside a `store.transaction` callback on the same SQLite backend (root
store instead of the `tx` context) used to deadlock permanently — the open
transaction holds the backend's serialized execution slot — and is now rejected
with a `ConfigurationError` that points at the transaction-scoped context.
- [#189](https://github.com/nicia-ai/typegraph/pull/189) [`fe21158`](https://github.com/nicia-ai/typegraph/commit/fe2115836d084a86613ae94a4403651d8316713a) Thanks [@pdlug](https://github.com/pdlug)! - Classify incompatible property-schema changes as breaking schema migrations. The
migration diff previously compared only the top-level JSON-Schema token of each
property, so a changed property type (e.g. `string` → `number`), a changed array
item type (`string[]` → `number[]`), a narrowed enum, or a type change nested
inside an object all auto-migrated silently as a non-blocking warning, leaving
stored rows that no longer satisfy the declared schema; edge property changes
were unconditionally treated as safe. Node and edge property diffs now share one
recursive, conservative classifier: a change is `safe` only when it can be proven
non-breaking (a new optional property, a metadata-only edit, or an additive
optional field nested inside an object). Everything else — a removed property, a
newly required property, an in-place type change, a changed array item schema, an
enum/const/composition change, a same-type constraint change, or a breaking
change nested inside an object — is `breaking` and blocks auto-migration. The
`warning` severity is no longer emitted for property changes.
- [#190](https://github.com/nicia-ai/typegraph/pull/190) [`1bfa9c2`](https://github.com/nicia-ai/typegraph/commit/1bfa9c28d04f03b9f82e23bf0a97417aba544767) Thanks [@pdlug](https://github.com/pdlug)! - Fix two silent query-correctness bugs. Keyset pagination (`paginate`/`stream`)
now appends a unique `id` tiebreaker to the ORDER BY so a non-unique sort no
longer drops equal-key rows across pages. And every compiled `LIKE`/`ILIKE` now
emits `ESCAPE '\'` — including the case-sensitive `like` path, which previously
omitted it — so escaped `%`/`_`/`\` match literally on SQLite as they already
did on PostgreSQL, in both the auto-escaped operators
(`contains`/`startsWith`/`endsWith`) and raw `like`/`ilike` patterns, and
whether the pattern is a literal or a bound parameter (previously SQLite had no
default LIKE escape character, so the two backends — and the direct vs prepared
paths — diverged).
- [#188](https://github.com/nicia-ai/typegraph/pull/188) [`0b0f4ea`](https://github.com/nicia-ai/typegraph/commit/0b0f4ea23ee2310cc2c160d24385eb94ebfdc5a8) Thanks [@pdlug](https://github.com/pdlug)! - Fix a uniqueness-reservation loss on node resurrection. Resurrecting a
soft-deleted node through `getOrCreateByConstraint` (or any
`clearDeleted: true` upsert) ran the diff-based uniqueness maintenance, which
skips a key that did not change — but the soft delete had already removed the
node's uniqueness entries, so the resurrected node held NO reservation and a
later `create` with the same unique value silently succeeded, duplicating it.
A resurrecting update now re-checks and re-inserts the entries for its new
props, exactly as the provenance reopen path does.
- [#188](https://github.com/nicia-ai/typegraph/pull/188) [`0b0f4ea`](https://github.com/nicia-ai/typegraph/commit/0b0f4ea23ee2310cc2c160d24385eb94ebfdc5a8) Thanks [@pdlug](https://github.com/pdlug)! - Open SQLite business-write transactions with `BEGIN IMMEDIATE` on the sync
(better-sqlite3) path, matching schema writes and the async libsql/Drizzle path.
A deferred `BEGIN` acquired the reserved write lock only on the first write, so a
read-then-write inside a transaction could fail with "database is locked" against
a writer on another connection to the same file; taking the lock at the start of
the transaction lets SQLite's busy timeout wait for it instead. The per-backend
serialized write queue continues to order a single backend's own transactions.
- [#192](https://github.com/nicia-ai/typegraph/pull/192) [`2af3a06`](https://github.com/nicia-ai/typegraph/commit/2af3a065d9d54b0ac89c32dc27d637a4eedc58cf) Thanks [@pdlug](https://github.com/pdlug)! - Type-check the remaining StoreView read-name buckets. `CURRENT_ONLY_READ_NAMES`
and `EDGE_BATCH_READ_NAMES` were plain `as const` arrays while every sibling
bucket carried a `satisfies readonly (keyof Collection)[]` guard, so a renamed
or mistyped method in those two would have gone uncaught at compile time. All
six buckets are now checked against the live collection keys. Compile-time only.
- [#188](https://github.com/nicia-ai/typegraph/pull/188) [`0b0f4ea`](https://github.com/nicia-ai/typegraph/commit/0b0f4ea23ee2310cc2c160d24385eb94ebfdc5a8) Thanks [@pdlug](https://github.com/pdlug)! - Operation hooks now mean "durably committed" everywhere. `onOperationEnd`
previously fired when an operation completed, even when that operation ran
inside an enclosing transaction whose COMMIT later failed — so hook consumers
(metrics, cache invalidation, audit logs) were told a rolled-back write
succeeded. Operations inside `store.transaction` now defer their success
hooks until the transaction commits, and a failed transaction converts every
completed operation's pending success into `onError`. Edge
`getOrCreateByEndpoints` no longer wraps its write legs in an outer
transaction (each leg commits — and reports — on its own, with a
probe/create race converged by one retry), and provenance transitions route
their source-flip and per-fact hooks through the same deferred lifecycle.
Inside an adopted transaction (`withTransaction` /
`withRecordedTransaction`) the commit belongs to the caller and cannot be
observed; hooks there keep firing at operation completion, as documented.
## 0.33.0
### Highlights
TypeGraph 0.33 adds recorded/system time alongside valid time. Enabling `history: true` captures TypeGraph-managed node and edge writes into historical relations at a monotonic per-graph commit instant, making it possible to reconstruct values that were later corrected or deleted.
`store.asOfRecorded()` exposes read-only reconstruction through point reads, queries, subgraph extraction, and graph algorithms. Applications can pin recorded and valid time independently to distinguish when a fact was true from when TypeGraph knew it. External history producers can bind compatible recorded relations through `recordedRelation()` and `recordedRead`.
### Upgrade notes
- Capture is opt-in and does not backfill existing rows. Enable it on a fresh graph for complete history; existing entities are first captured on their next managed write. Capture requires a transactional backend with statement execution.
- Use typed collection writes for captured transactions. Raw `tx.sql` is unavailable under `history: true`; adopt caller-owned transactions through `withRecordedTransaction(externalTx, callback)` so capture flushes before commit.
- Use `recordedNow()` as a post-write reconstruction anchor after checking for `undefined`. Recorded instants can advance ahead of wall-clock time, and recorded coordinates must use canonical UTC ISO 8601.
- Recorded views expose only reconstructible reads. Current full-text/vector indexes and unsupported broad collection reads are refused; `asOfRecorded(T)` pins both time axes to `T` unless composed with an explicit valid-time view.
- Custom backend and SQL tooling must honor the new backend-role and row-versus-statement SQL intent brands. `recordedRead` descriptors must come from `recordedRelation()` and cannot be combined with `history: true`.
### Minor Changes
- [#186](https://github.com/nicia-ai/typegraph/pull/186) [`655407a`](https://github.com/nicia-ai/typegraph/commit/655407a9c225e8eca0aff5f636ed17ca99f3e382) Thanks [@pdlug](https://github.com/pdlug)! - Add recorded / system-time capture — TypeGraph's second temporal axis. Where valid time (`validFrom` / `validTo`, queried via `asOf` / `includeEnded`) records _when a fact was true in the world_, recorded time records _when TypeGraph captured a managed node/edge write_. Together they answer "what did TypeGraph reconstruct as true, as of a captured commit instant?" — surfacing values that were later corrected (à la SQL:2011 system-versioned tables).
Enable capture per store with `createStore(graph, backend, { history: true })`. TypeGraph collection writes through that store are then captured into recorded-time relations (`typegraph_recorded_nodes` / `typegraph_recorded_edges`), stamped with a per-graph monotonic commit instant from a `typegraph_recorded_clock` (serialized on PostgreSQL via a per-graph advisory lock). Capture is opt-in and has **no backfill** — enable it on a fresh graph, since an entity that already exists is first recorded the next time it is written. It requires a transactional backend with statement execution (the built-in SQLite / PostgreSQL backends).
Read at a recorded instant with `store.asOfRecorded(T)`, which returns a narrow read-only `RecordedStoreView`. Direct `store.asOfRecorded(T)` is diagonal bitemporal sugar (recorded _and_ valid axes both at `T`); chain `store.asOf(validT).asOfRecorded(recordedT)` to pin the two axes independently, or `store.view({ mode }).asOfRecorded(recordedT)` to compose recorded time with any valid-time mode (e.g. `includeTombstones`). `store.recordedNow()` returns the recorded high-water mark; after guarding the `undefined` case, passing that value to `store.asOfRecorded(...)` is a deterministic "as things stand now" anchor. Recorded instants are monotonic and can run briefly ahead of wall-clock time under bursty writes, so the wall clock is not a reliable anchor right after a write.
The recorded view is a **reconstructing** lens that exposes only reads which can be faithfully rebuilt from the history relations: point reads (`nodes..getById` / `getByIds` and the edge equivalents), a sealed `query()`, `subgraph()`, and the graph algorithms (`reachable` / `canReach` / `shortestPath` / `degree`). Broad collection reads (`find` / `count` / `findFrom`), `search`, and fulltext / vector predicates refuse with a `ConfigurationError` / `UnsupportedPredicateError` — those indexes reflect current state only. `T` must be a canonical UTC ISO-8601 timestamp (`YYYY-MM-DDTHH:mm:ss.sssZ`).
The public live-read and algorithm option types explicitly reject internal recorded coordinates, while recorded internals use a branded `RecordedInstant` so only validated canonical recorded instants can flow through the reconstructing paths.
Recorded read binding is now explicit without exposing TypeGraph's internal capture binding. `history: true` enables TypeGraph-managed capture and binds the built-in recorded relations internally, while the factory-branded `recordedRelation({ schema })` / `recordedRead` path is the external-read-source API for hosts that populate a row-compatible recorded relation outside TypeGraph's writer wrapper. The store validates that runtime `recordedRead` values come from `recordedRelation({ schema })`, rejects `recordedRead` combined with `history: true`, and factory-brands/freezes SQL schema and recorded-read descriptors so they cannot be structurally forged as plain objects. Store overloads reflect that split: history-enabled stores expose `HistoryStore`, read-bound live stores expose `RecordedReadStore`, and captured-history stores expose `HistorySafeBackend` / `HistoryTransactionContext` types that hide raw statement / DDL write seams from the typed `backend`, `transaction()`, and `withRecordedTransaction()` surfaces.
Writes under `history: true` flush capture at transaction commit, so they must go through the typed collections: raw `tx.sql` is disabled (it would bypass capture), and `store.withTransaction(externalTx)` is replaced by the callback form `store.withRecordedTransaction(externalTx, async (tx) => ...)`, which gives capture a flush point before the caller commits. `store.clear()` clears the recorded relations alongside the live tables.
Node creates now run atomically on transactional backends with uniqueness, vector, and fulltext finalization, and node delete cascades now run atomically even without `history: true`. A failed finalize step rolls back the node row instead of leaving a partially indexed row behind. Overlapping PostgreSQL cascades may hold locks longer, so callers should keep normal deadlock-retry handling around concurrent deletes.
Backend and SQL execution contracts are more explicit for maintainers and extension authors: backend role brands separate graph-write paths from raw/bulk paths, `execute` / `executeStatement` now require row-vs-statement SQL intent brands, transaction backends are composed from explicit backend facets instead of `Omit`, and backend wrappers use an exact overlay helper that preserves prototype/proxy backends while catching typoed override keys at compile time.
Exports `RecordedStoreView` and its collection types (`RecordedStoreViewNodeCollection` / `RecordedStoreViewNodeCollections`, `RecordedStoreViewEdgeCollection` / `RecordedStoreViewEdgeCollections`, `TypedRecordedStoreViewEdgeCollection`).
**Performance:** recorded reads reconstruct from the history relations rather than the live tables, so they are slower than current-state reads — most noticeably for full-graph `subgraph` / algorithm reconstructions on PostgreSQL. Use `asOfRecorded` for audit and point-in-time reconstruction, not hot-path reads.
## 0.32.0
### Minor Changes
- [#182](https://github.com/nicia-ai/typegraph/pull/182) [`0f0e771`](https://github.com/nicia-ai/typegraph/commit/0f0e77161d473b5c3b2d2e224d930c611eb4b123) Thanks [@pdlug](https://github.com/pdlug)! - Close the TOCTOU windows in graph-merge commits. A merge resolves its plan from reads taken before the commit transaction, so a write landing on the target in between could previously be committed over. Now, inside the commit transaction: `merge()` and `mergeAgainstBase()` re-validate the target's base@V content fingerprint, and `mergeIncremental()` re-runs its new-vs-base identity resolution (the unique-constraint and block-index probes). All three fail with `BaseVersionMismatchError` — instead of committing a stale plan or a duplicate entity — when the target changed in that window. Merge commits run at `SERIALIZABLE` isolation with bounded retry on serialization failures and deadlocks, making the guards race-free on multi-writer Postgres. `Store.transaction()` accepts optional `TransactionOptions` (isolation level) and `TransactionContext` exposes the transaction-scoped `backend`.
- [#185](https://github.com/nicia-ai/typegraph/pull/185) [`4e23be8`](https://github.com/nicia-ai/typegraph/commit/4e23be8d6af94b965bdcf90e911dc0e1c49d2bad) Thanks [@pdlug](https://github.com/pdlug)! - Add `StoreView`, a read-only `(mode, asOf)` lens over a `Store` that pins a temporal coordinate and routes every supported read through it (the as-of database value, à la Datomic `(d/as-of db t)` / SQL:2011 `FOR SYSTEM_TIME AS OF`). Construct one with `store.asOf(T)` (valid-time) or `store.view({ mode, asOf })` for the other public modes (`current` / `includeEnded` / `includeTombstones`). The view exposes pinned `nodes` / `edges` collections (`getById` / `getByIds` / `find` / `count`, edge `findFrom` / `findTo`), a pre-pinned `query()`, `subgraph()`, and the graph algorithms (`reachable` / `canReach` / `shortestPath` / `neighbors` / `degree`). It is read-only by construction — writes and temporally-unscoped reads refuse with a clear error — and `search` refuses on a non-`current` pin (the fulltext / vector index reflects current state only).
Internally every pinned surface injects a single opaque `ReadCoordinate` through one helper, so a future temporal axis (recorded / system time) lands on every surface at once instead of splitting per surface. The view's read surface is derived from a read/write split of the live collection types (`NodeTemporalReads` / `NodeCurrentReads` / `NodeWrites` and edge equivalents, now exported) with a `test-d` conformance check, so a new collection read cannot silently bypass the view's pinning decision.
- **`store.snapshot()`.** Sugar for `store.asOf(new Date().toISOString())` — a read-only view pinned to the current instant captured once at construction. Unlike `store.view({ mode: "current" })` (which tracks "now" live), a snapshot is a stable point-in-time value where every surface observes the same instant. Mirrors Datomic's `(d/db conn)`.
- **Sealed pinned query.** `view.query()` now returns a query builder whose temporal axis is sealed — calling `.temporal(...)` on it throws — so a pinned view cannot be silently re-coordinated per query.
- **Current-only reads.** Constraint / index lookups (`findByConstraint`, `bulkFindByConstraint`, `bulkFindByIndex`), which have no temporal axis, are now available on a `current` view (delegating to the live store) and refuse with a clear error on a temporal pin — instead of being unavailable on every view.
**Breaking — `find` / `count` signature:** `store.nodes..find(...)` / `count(...)` and `store.edges..find(...)` / `count(...)` now take the temporal coordinate as a **second** argument rather than inline in the filter object: `find(filter?, temporal?)` / `count(filter?, temporal?)`. For example, `nodes.Person.find({ where, temporalMode: "asOf", asOf })` becomes `nodes.Person.find({ where }, { temporalMode: "asOf", asOf })`, and `edges.worksAt.count({ temporalMode: "includeEnded" })` becomes `edges.worksAt.count(undefined, { temporalMode: "includeEnded" })`. Old call sites that inlined `temporalMode` / `asOf` are now type errors. `getById` / `getByIds` / `findFrom` / `findTo` / node `count` are unchanged (they already took a trailing temporal argument).
**Breaking — canonical `validFrom` / `validTo` on write:** `create` / `update` / `bulk*` now require canonical fixed-width UTC ISO timestamps (`YYYY-MM-DDTHH:mm:ss.sssZ`) for `validFrom` / `validTo`, rejecting date-only, zoned-offset, variable/missing-millisecond, and rollover values with a `ValidationError`. This makes the _stored_ values that temporal filters compare as text always sort chronologically — the same contract the `asOf` read coordinate already enforces, applied uniformly to every timestamp in the system. Convert non-canonical inputs with `new Date(value).toISOString()`. There is no migration: pre-existing non-canonical rows are left as-is (recreate them if affected) — acceptable pre-1.0.
**Behavior change:** `store.edges..findFrom(...)` / `findTo(...)` / `findByEndpoints(...)` (and their `batchFindFrom` / `batchFindTo` / `batchFindByEndpoints` variants) now honor the temporal model like `getById` / `find` instead of returning every non-soft-deleted edge. With no temporal argument, the graph's default `temporalMode` applies — so under the default `"current"` mode, edges outside their `validFrom` / `validTo` window are now excluded. Pass `temporalMode` / `asOf` to read at another coordinate (e.g. `temporalMode: "includeEnded"` to recover the previous "all non-deleted" behavior). `findByEndpoints` / `batchFindByEndpoints` gain a trailing `temporal?` argument and are now pinnable on a `StoreView` (no longer refused on a temporal pin). The internal `getOrCreate*ByEndpoints` identity lookup is unaffected — it deliberately matches against all edges regardless of validity window.
**Read coordinates:** `asOf`, `.temporal("asOf", T)`, algorithms, subgraph, and `StoreView` require canonical UTC ISO timestamps (`YYYY-MM-DDTHH:mm:ss.sssZ`) for the same lexicographic-comparison reason.
## 0.31.0
### Highlights
TypeGraph 0.31 introduces graph branching and semantic merge through `@nicia-ai/typegraph/graph-merge`. `branch()` creates an isolated working copy; `merge()` reconciles one or more branches using stable IDs, declared uniqueness constraints, blocking keys, and optional similarity scoring. Canonical survivors receive merged properties and repointed edges, with conflicts and source contributions recorded in the report.
Snapshot merges check the branch's base token, while `mergeIncremental()` supports targets that have advanced since the fork. Applications can configure property and delete/modify conflict policies, use ontology-aware type reconciliation, and optionally persist provenance in a separate sidecar graph.
### Upgrade notes
- Merge requires a transaction-capable backend and refuses non-atomic execution. Vector and hybrid entity-resolution strategies require a configured embedder.
- Choose snapshot or incremental merge according to whether the target may advance after the fork. A base mismatch requires a fresh branch or an appropriate incremental merge workflow.
- Optional `persistProvenance` runs after the graph commit. A persistence failure is a report warning and does not roll back the merged graph.
### Minor Changes
- [#178](https://github.com/nicia-ai/typegraph/pull/178) [`6b6e418`](https://github.com/nicia-ai/typegraph/commit/6b6e4186642c65d58c939250458b6521efbc40c7) Thanks [@pdlug](https://github.com/pdlug)! - Add `@nicia-ai/typegraph/graph-merge`, a TypeGraph-native branch and semantic merge subpath for deterministic entity-resolution merges across graph forks.
## 0.30.0
### Minor Changes
- [#171](https://github.com/nicia-ai/typegraph/pull/171) [`f5defd3`](https://github.com/nicia-ai/typegraph/commit/f5defd35b331e56f282d4eb501b98d3b9affe562) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.nodes..bulkFindByIndex(indexName, items, options?)` — batched candidate retrieval against declared node indexes, including non-unique ones. For each input record it returns the live nodes that share that record's declared index key, for import reconciliation, dedup-candidate discovery, and joining records against the graph by a composite key. Each input yields its own array (candidate retrieval, not a uniqueness guarantee); buckets preserve input order and are ordered by node id.
TypeGraph owns the index semantics: keys are computed from `index.fields` only (reusing the index's own extraction expressions), the partial `where` is applied to stored rows, and a missing/`undefined` indexed field matches a stored `NULL` via a new null-safe-equality dialect adapter. An optional `limitPerInput` caps each bucket — in SQL via `ROW_NUMBER()` when the backend supports window functions, otherwise capped in memory with the same result. Date-typed key fields are rejected with `ConfigurationError` because they can't compare identically across SQLite and PostgreSQL. Unknown index names throw `NodeIndexNotFoundError`.
`createLocalSqliteBackend` also gains a `capabilities` override for simulating engine capability gaps (e.g. `windowFunctions: false`) in tests.
- [#173](https://github.com/nicia-ai/typegraph/pull/173) [`bd96cfb`](https://github.com/nicia-ai/typegraph/commit/bd96cfbeadde11c6986fb667f9a86b0ba0b5b1bd) Thanks [@pdlug](https://github.com/pdlug)! - Add the `backend.capabilities.windowFunctions` capability and reject relevance-ranking queries before SQL generation when a custom backend profile disables SQL window functions.
## 0.29.0
### Highlights
TypeGraph 0.29 makes vector storage and hybrid search portable across pgvector, sqlite-vec, and libSQL/Turso through a pluggable `VectorStrategy`. Embeddings move into graph-scoped, fixed-dimension storage per field, with strategy-derived metric and index capabilities, legacy migration tooling, and field re-embedding after dimension changes.
PGlite gains first-class backend support and a local factory for running PostgreSQL and pgvector in process. SQLite set operations now compile their operands through the full query compiler, preserving traversal, search, nested ordering, and pagination behavior across backends.
### Upgrade notes
- Existing vector deployments must run `migrateLegacyEmbeddings()` after upgrading. Search no longer reads the shared legacy `typegraph_node_embeddings` table. Use `reembedVectorField()` when an embedding field's dimensions change.
- Install the optional PGlite peers when using the local PGlite backend. Configure `vector: false` when the engine has no vector extension.
- Custom capability consumers must remove the descriptive-only `jsonb`, `ginIndexes`, `partialIndexes`, `cte`, and `returning` flags. Query semantics are provided by the shared compiler and configured strategies.
- Interchange payloads with `source.type: "typegraph-cloud"` must be retagged as `"external"`; the removed source variant now fails validation.
### Minor Changes
- [#161](https://github.com/nicia-ai/typegraph/pull/161) [`9e86269`](https://github.com/nicia-ai/typegraph/commit/9e862695c6a3341af5d8acbd4f652738bd7727ca) Thanks [@pdlug](https://github.com/pdlug)! - Add cross-backend vector and hybrid search through a pluggable
`VectorStrategy`, closing [#157](https://github.com/nicia-ai/typegraph/issues/157). TypeGraph now has first-class vector storage and
search for libSQL/Turso, sqlite-vec, and pgvector behind the same semantic
search APIs.
Backend highlights:
- libSQL/Turso stores fixed-dimension embeddings in `F32_BLOB(N)` columns,
supports cosine/L2 search, and can use DiskANN through `libsql_vector_idx`
and `vector_top_k`.
- sqlite-vec uses `vec0` KNN tables instead of brute-force vector scans.
- pgvector uses graph-scoped, per-field `vector(N)` tables with HNSW/IVFFlat
materialization.
- Backends advertise vector metrics, index types, and dimension limits from the
active strategy, and `createSqliteBackend` / `createPostgresBackend` accept a
custom `vector?: VectorStrategy`.
The release also adds migration and lifecycle tooling for the new storage model:
- `migrateLegacyEmbeddings(...)` copies existing rows out of the legacy shared
`typegraph_node_embeddings` table.
- `store.reembedVectorField(kind, fieldPath, { embed? })` recreates a field's
storage after an embedding dimension change and can re-embed existing rows.
- `store.materializeRemovals()` reclaims vector tables for removed embedding
fields and reports them in `MaterializeRemovalsResult.reclaimedVectorFields`.
**Breaking storage change:** vector embeddings now live in graph-scoped,
fixed-dimension per-field storage instead of the shared
`typegraph_node_embeddings` table. Search no longer reads the legacy table.
Deployments with existing embeddings must run `migrateLegacyEmbeddings(...)`
once after upgrading; deployments without stored embeddings need no migration.
- [#165](https://github.com/nicia-ai/typegraph/pull/165) [`ae5bfdc`](https://github.com/nicia-ai/typegraph/commit/ae5bfdc55aae3531bcd75f0770cb2812ad9682d9) Thanks [@pdlug](https://github.com/pdlug)! - Reduce `BackendCapabilities` to the flags the library actually consumes:
`transactions`, `vector`, and `fulltext`.
The descriptive-only flags `jsonb`, `ginIndexes`, `partialIndexes`, `cte`, and
`returning` were never read anywhere to gate a query feature or pick an index
strategy. `jsonb`/`ginIndexes` additionally misrepresented SQLite, which has
native JSON (`json_extract`/`json_each`) and supports B-tree expression indexes
on scalar JSON properties at parity with PostgreSQL — the only real JSON
difference (GIN containment acceleration) is a Postgres performance
characteristic, not a gated capability.
If you were reading any of these removed flags, branch on
`backend.dialect === "postgres"` instead, or rely on the dialect layer
(JSON-path predicates, `WITH` queries, `RETURNING`, partial indexes, and
`defineNodeIndex`/`defineEdgeIndex` work the same on both backends).
- [#163](https://github.com/nicia-ai/typegraph/pull/163) [`0175a25`](https://github.com/nicia-ai/typegraph/commit/0175a2585029aa1b6ceabc9889074a72b8895d03) Thanks [@pdlug](https://github.com/pdlug)! - Add first-class support for [PGlite](https://pglite.dev/) (Postgres-in-WASM),
closing [#160](https://github.com/nicia-ai/typegraph/issues/160).
- **Execution fast-path fix.** `createPostgresBackend` now detects a PGlite
`db.$client` and routes it to the unnamed positional query wrapper. PGlite's
`.query` has no node-postgres named-statement config form — passing one
desyncs its single connection (`08P01`), so under the default
`prepareStatements: true` every query previously failed. PGlite works
unchanged with `createPostgresBackend(drizzle(pglite))` now.
- **`createLocalPgliteBackend`** — a batteries-included helper under the new
`@nicia-ai/typegraph/postgres/pglite` entry, the Postgres analog of
`createLocalSqliteBackend`. It constructs an in-process PGlite engine
(in-memory by default, or any `dataDir`), loads pgvector, runs the schema
DDL, and returns `{ backend, db, client }` whose `close()` disposes the
engine. Pass `vector: false` to skip the extension, or `vector: `
to bring your own pgvector build.
`@electric-sql/pglite` (and, for vector support, `@electric-sql/pglite-pgvector`
on PGlite ≥ 0.5) are optional peer dependencies. The biggest payoff: the
Postgres dialect and pgvector path can now be exercised in plain `pnpm test`
with zero Docker.
- [#162](https://github.com/nicia-ai/typegraph/pull/162) [`48a6ffc`](https://github.com/nicia-ai/typegraph/commit/48a6ffc3e63459e7a2535a936a8c9c3fbcd29a99) Thanks [@pdlug](https://github.com/pdlug)! - Add `vector: false` to `createPostgresBackend` to disable the vector stack.
The Postgres backend wires `pgvectorStrategy` by default, assuming a standalone
Postgres server has the pgvector extension installed. An in-process Postgres
(PGlite) built without that extension can't honor it — the default strategy's
`vector(N)` DDL hard-fails the moment an embedding is written or
`CREATE EXTENSION vector` runs. Passing `vector: false` turns the stack off:
the backend advertises no `capabilities.vector` and omits the
embedding/search methods, mirroring a SQLite connection without sqlite-vec, so
the store never routes vector work to it.
Real-Postgres behavior is unchanged — the default remains `pgvectorStrategy`.
- [#158](https://github.com/nicia-ai/typegraph/pull/158) [`bc07847`](https://github.com/nicia-ai/typegraph/commit/bc07847cbde20eedd01781062e0403856cb46079) Thanks [@pdlug](https://github.com/pdlug)! - Export the ontology transitive-closure utilities (`computeTransitiveClosure`, `invertClosure`, `isReachable`) from the package root. These were previously internal-only. Exposing them lets consumers reason over `subClassOf` / `equivalentTo` hierarchies — e.g. reconciling node types when merging graphs from independent sources.
- [#166](https://github.com/nicia-ai/typegraph/pull/166) [`a32d31f`](https://github.com/nicia-ai/typegraph/commit/a32d31f7bbe9fc4657eb956e86900eaf1c283ef9) Thanks [@pdlug](https://github.com/pdlug)! - Remove the `typegraph-cloud` source type from the interchange
`GraphDataSourceSchema`.
TypeGraph Cloud is not a publicly available product, so the `typegraph-cloud`
variant has been dropped from the graph-data source discriminated union, and the
corresponding interchange documentation has been removed. `GraphDataSource` now
accepts only `typegraph-export` and `external`.
**Breaking:** importing data whose `source.type` is `"typegraph-cloud"` now
fails schema validation. Re-tag such payloads as `"external"` before importing.
- [#165](https://github.com/nicia-ai/typegraph/pull/165) [`ae5bfdc`](https://github.com/nicia-ai/typegraph/commit/ae5bfdc55aae3531bcd75f0770cb2812ad9682d9) Thanks [@pdlug](https://github.com/pdlug)! - Support the full query feature set inside SQLite set operations
(`UNION`/`UNION ALL`/`INTERSECT`/`EXCEPT`).
Previously the SQLite set-operation compiler hand-rolled a thin subset of leaf
compilation and rejected leaves that used traversals, `EXISTS`/`IN` subqueries,
vector or fulltext predicates, `GROUP BY`/`HAVING`, or per-leaf
`ORDER BY`/`LIMIT`/`OFFSET` — throwing `UnsupportedPredicateError` at execution
time. PostgreSQL accepted all of these. The result was a portability cliff: a
combined query developed against PostgreSQL could throw the moment the backend
was switched to SQLite.
Both dialects now compile every leaf with the full query compiler and only
differ in how each operand is wrapped. SQLite forbids parenthesized compound
operands, but it does allow a `WITH` clause inside a FROM-subquery, so each
operand is emitted as `SELECT * FROM ()`. This keeps every leaf's CTEs
(traversal joins, recursive expansions, vector/fulltext relevance) scoped to its
own subquery and lets per-leaf `ORDER BY`/`LIMIT`/`OFFSET` live inside the wrap.
Nested set operations are wrapped the same way, preserving the AST's grouping
regardless of the dialect's native compound-operator associativity. As a
side effect, vector/fulltext predicates in set-operation leaves now use the
backend's configured relevance strategy instead of falling back to the dialect
default.
Note: `GROUP BY`/`HAVING` leaves are supported at the compiler level, but the
query builder still does not expose `.union()`/`.intersect()`/`.except()` on
aggregate queries — that builder gate is unchanged and applies equally to both
backends.
### Patch Changes
- [#165](https://github.com/nicia-ai/typegraph/pull/165) [`ae5bfdc`](https://github.com/nicia-ai/typegraph/commit/ae5bfdc55aae3531bcd75f0770cb2812ad9682d9) Thanks [@pdlug](https://github.com/pdlug)! - Fix `ORDER BY`/`LIMIT`/`OFFSET` being silently dropped on a nested set-operation
operand.
When a set operation was nested inside another — e.g.
`a.union(b).limit(10).intersect(c)` — the inner compound's suffix clauses were
applied only at the top level, so the inner `limit`/`offset` were ignored and
the outer operation ran over the full (unlimited) inner result. The compiler now
emits each nested compound's own `ORDER BY`/`LIMIT`/`OFFSET` inside its operand
subquery on both SQLite and PostgreSQL.
- [#165](https://github.com/nicia-ai/typegraph/pull/165) [`ae5bfdc`](https://github.com/nicia-ai/typegraph/commit/ae5bfdc55aae3531bcd75f0770cb2812ad9682d9) Thanks [@pdlug](https://github.com/pdlug)! - Validate set-operation leaf vector predicates against the configured vector
strategy rather than only the dialect's fallback metric list, so a custom
strategy's metric (e.g. `inner_product` on SQLite) is accepted inside
`UNION`/`INTERSECT`/`EXCEPT` leaves exactly as it is in a standalone query.
Reject a per-query fulltext `language` override on the query-builder path
(`.$fulltext.matches(..., { language })`) when the strategy's tokenizer is fixed
at table-create time (SQLite/FTS5), matching the store-level search guard
instead of silently ignoring the option.
## 0.28.1
### Patch Changes
- [#154](https://github.com/nicia-ai/typegraph/pull/154) [`6703c88`](https://github.com/nicia-ai/typegraph/commit/6703c880d3d9047149f91d1db4a27b414983c632) Thanks [@pdlug](https://github.com/pdlug)! - Fix `isMissingTableError` missing DrizzleQueryError-wrapped Postgres
"relation does not exist" errors, breaking fresh/partial Postgres boot ([#153](https://github.com/nicia-ai/typegraph/issues/153)).
`isMissingTableError` (the shared "relation not bootstrapped yet"
discriminant for `loadActiveSchemaWithBootstrap`, `readActiveSchemaPure`,
and the [#135](https://github.com/nicia-ai/typegraph/issues/135) durable-marker gate) classified failures by inspecting only
`error.message`. On Postgres, drizzle-orm wraps every query-builder call
(`db.select()`, `db.insert()`, …) in a `DrizzleQueryError` whose `.message`
is the failed SQL text; the real driver error — carrying both
`relation "…" does not exist` and SQLSTATE `42P01` — is preserved on
`error.cause`, which the helper never walked. So the helper returned
`false` and a benign "not bootstrapped yet" surfaced as a hard fault.
This regressed `createStoreWithSchema` after the [#149](https://github.com/nicia-ai/typegraph/issues/149)/[#152](https://github.com/nicia-ai/typegraph/issues/152) read-only
pre-check: `ensureRuntimeContributions` now calls `getMarker` (a
query-builder read) on the possibly-absent
`typegraph_contribution_materializations` table _before_ `ensureMarkerTable()`.
On Postgres that read throws a `DrizzleQueryError`, the helper missed it,
and the open rethrew instead of materializing — breaking seed, first boot,
and test global-setup on any fresh or partial Postgres database (base
tables present, marker table absent — e.g. drizzle-kit-managed schemas).
SQLite was unaffected because better-sqlite3 throws a raw error whose
`.message` literally contains `no such table`.
`isMissingTableError` now walks the `error.cause` chain (cycle-safe) and
additionally keys on the locale-independent SQLSTATE `42P01`, rather than
matching only the outermost `.message`. Existing message patterns are
retained, so all prior matches still hold; the fix applies uniformly to
all three call sites, including the latent slow-path blind spot in
`loadActiveSchemaWithBootstrap` / `readActiveSchemaPure`.
## 0.28.0
### Minor Changes
- [#150](https://github.com/nicia-ai/typegraph/pull/150) [`f9b1300`](https://github.com/nicia-ai/typegraph/commit/f9b1300a031eb758ae456fcd97ba8cbfdf93a2b8) Thanks [@pdlug](https://github.com/pdlug)! - Add a per-search `efSearch` knob for tuning pgvector HNSW recall ([#148](https://github.com/nicia-ai/typegraph/issues/148)).
`store.search.vector` and the vector half of `store.search.hybrid` now
accept an optional `efSearch` — the HNSW search frontier
(`hnsw.ef_search`, default 40). pgvector caps a single index scan at
`ef_search` candidates, so the hybrid over-fetch (`vectorK = 4 * limit`
by default) silently under-delivers once `vectorK` climbs past the
session default; the floor is `efSearch >= vectorK` and ~2–4× is the
high-recall target. Being per-search lets one connection pool serve both
a latency-sensitive interactive path and a recall-sensitive batch path.
The Postgres backend applies it transaction-locally
(`SET LOCAL hnsw.ef_search`) around the vector `SELECT`, so it never
leaks to the next query on a pooled connection — `SET LOCAL` issued in
autocommit would roll off with the statement and the next pooled query
would see the session default. Omitting `efSearch` opens no transaction
and preserves today's behavior exactly. Validated as a positive integer
≤ 1000 (pgvector's ceiling).
Scope: pgvector HNSW only. sqlite-vec has no equivalent frontier knob
and treats it as a no-op; transaction-less Postgres drivers
(`drizzle-orm/neon-http`) ignore it with a one-time warning. IVFFlat's
`ivfflat.probes` is a follow-up.
### Patch Changes
- [#152](https://github.com/nicia-ai/typegraph/pull/152) [`761c672`](https://github.com/nicia-ai/typegraph/commit/761c672a991ea75454e441a4baf5939792da9505) Thanks [@pdlug](https://github.com/pdlug)! - Fix `ensureRuntimeContributions` running marker-table DDL on every store
open ([#149](https://github.com/nicia-ai/typegraph/issues/149)).
`createStoreWithSchema` → `ensureRuntimeContributions` previously ran the
`typegraph_contribution_materializations` marker DDL
(`ensureMarkerTable()` → `CREATE TABLE IF NOT EXISTS …`) on **every** open
for any graph with runtime contributions (e.g. `searchable()` fields),
even when every contribution was already materialized. The per-materializer
`initializedGraphIds` cache is per-instance, so a deployment that builds a
fresh backend per request (the norm on serverless Postgres) got an empty
cache each time and re-ran the DDL on every open — which intermittently
fails on connections that can't run it (observed on Cloudflare Workers +
the Neon serverless driver) and surfaces as a wrapped `DrizzleQueryError`
rather than a clean `MigrationError`.
`ensureRuntimeContributions` now does a read-only pre-check first, mirroring
the SELECT-only `assertInitialized`: when every runtime contribution is
already materialized (marker present, signature matches, no recorded error)
it returns without `ensureMarkerTable()` / `materializeOne`. A missing
marker table, or any missing/stale/failed contribution, still falls through
to the unchanged privileged first-materialization path. Warm per-request
opens are now DDL-free.
Note: the canonical runtime attach for the least-privilege / per-request
deployment model remains `createVerifiedStore` (zero DDL by construction);
`createStoreWithSchema` also runs bootstrap and auto-migration DDL and is
still intended to run once under a privileged role. This change is
defense-in-depth for the marker DDL specifically.
## 0.27.0
### Minor Changes
- [#144](https://github.com/nicia-ai/typegraph/pull/144) [`30a1cfd`](https://github.com/nicia-ai/typegraph/commit/30a1cfdba6f55240f3251de1ebdb05d69a66ea4c) Thanks [@pdlug](https://github.com/pdlug)! - Add `createVerifiedStore` and `assertSchemaCurrent` — the runtime
counterparts of `createStoreWithSchema` for the least-privilege
deployment model.
`createStoreWithSchema()` runs DDL (bootstrap, safe auto-migrations,
durable contribution materialization) and must run under a role with
`CREATE` privileges. For applications that want their runtime under a
least-privilege, DML-only role, the previous options were `createStore`
(zero-DDL attach with no schema gate — drift goes undetected until a
hot-path operation trips) or hand-rolling a SELECT-only verification
dance from `getActiveSchema` + `getSchemaChanges`.
This release adds two cleanly named entrypoints that share the same
zero-DDL verification path:
- **`createVerifiedStore(graph, backend, options?)`** — a SELECT-only
attach (zero DDL) with a verification gate. Reads the active schema
row and contribution markers, folds the persisted graph extension,
and refuses to construct the Store unless the database is at the
same schema version as the code graph. Returns
`Promise<[Store, SchemaValidationResult]>` mirroring
`createStoreWithSchema`. Throws `MigrationError` on any drift (safe
or breaking — the least-privilege runtime cannot migrate),
`ConfigurationError` when no schema has been initialized, and
`StoreNotInitializedError` when the schema is current but
runtime-contribution markers (e.g. fulltext) are missing/stale.
- **`assertSchemaCurrent(backend, graph)`** — the same verification gate
exposed as a standalone predicate for readiness probes / healthchecks.
Returns the `SchemaValidationResult` or throws the same errors.
The recommended deployment shape is now:
1. **Migration step** (privileged role with DDL/`CREATE`): run
`createStoreWithSchema()` once at startup, or apply
`generatePostgresMigrationSQL` / `generateSqliteMigrationSQL` plus a
one-shot `createStoreWithSchema()` to materialize runtime
contributions.
2. **Runtime** (least-privilege, DML-only role): attach with
`createVerifiedStore()`. Zero DDL on the runtime path; schema drift
fails fast with a clean `MigrationError` instead of leaking into
hot-path operations or 500ing on a permission error.
Internal: factored a pure `mergeStoredGraphExtension` helper out of
`loadAndMergeGraphExtensionDocument` so the SELECT-only verifier reuses
the same parse + extension-merge + deprecated-kind logic without going
through the bootstrap-capable loader. No behavior change for the
existing schema entrypoints.
Documentation: "Database roles & least privilege" in `backend-setup.md`
now folds in `createVerifiedStore` as the canonical runtime attach;
`schema-management.md` covers Basic / Managed / Verified stores side by
side; `troubleshooting.md` adds entries for `MigrationError` from a
verifying attach and `ConfigurationError` on uninitialized databases.
### Patch Changes
- [#144](https://github.com/nicia-ai/typegraph/pull/144) [`30a1cfd`](https://github.com/nicia-ai/typegraph/commit/30a1cfdba6f55240f3251de1ebdb05d69a66ea4c) Thanks [@pdlug](https://github.com/pdlug)! - Surface `MigrationError` before runtime-contribution DDL on a pending
breaking migration ([#143](https://github.com/nicia-ai/typegraph/issues/143)).
`loadActiveSchemaWithBootstrap` ran `ensureRuntimeContributions` (fulltext
contribution DDL) **before** `ensureSchema` computed the schema diff and
threw `MigrationError`. Contribution DDL is derived from the current code
graph, so against a database still on the old schema version it was applied
to a stale table shape. On Postgres the first failing statement aborts the
surrounding transaction, and the error that escaped was the idempotent
marker-table `CREATE TABLE IF NOT EXISTS
"typegraph_contribution_materializations"` (collateral damage), not a clean
`MigrationError`. Consumers using the documented migrate-on-`MigrationError`
recovery pattern never saw a `MigrationError`, so the first request after
every breaking schema change 500'd until a concurrent boot won the migration
race.
`loadActiveSchemaWithBootstrap` no longer materializes runtime
contributions. `createStoreWithSchema` remains the single canonical
durable-marker writer and runs the materialization step **after**
`ensureSchema`, so the breaking-change gate is always reached first and a
pending breaking migration throws `MigrationError` on the first request —
making the migrate-then-retry recovery path work as documented. The pre-[#129](https://github.com/nicia-ai/typegraph/issues/129)
`ensureFulltextTable` fallback is preserved at the canonical writer. No API
changes.
## 0.26.0
### Highlights
TypeGraph 0.26 lets graph writes share a transaction with application-owned relational writes. `store.withTransaction(externalTx)` binds graph collections to the caller's connection, while `tx.sql` exposes the transaction handle when TypeGraph owns the boundary. Cloudflare Durable Objects SQLite gains asynchronous storage-transaction support, allowing graph and application writes to roll back together across awaited operations.
Strategy-owned tables now follow one `TableContribution` contract. Full-text initialization becomes a durable, signature-checked materialization fact, allowing runtime operations to verify readiness without issuing DDL inside a business transaction.
### Upgrade notes
- Initialize the parent Store through `createStoreWithSchema()` before adopting business transactions. Full-text operations refuse missing, stale, or failed materialization markers with `StoreNotInitializedError`.
- `withTransaction()` requires real rollback support and refuses transactionless backends such as D1 and Neon HTTP. On synchronous better-sqlite3 connections, use an explicit transaction boundary instead of an async Drizzle transaction callback; Durable Objects use the asynchronous storage transaction runner.
- Inside a TypeGraph-owned transaction, use its `tx.sql` handle for relational writes so they share the same connection. Using the outer database object on pooled backends can escape the transaction.
- Custom `FulltextStrategy` implementations must replace `generateDdl()` with `ownedTables()`, declaring each table and its supporting indexes. Custom backends should implement the contribution initialization and durable materialization contracts described below.
### Minor Changes
- [#139](https://github.com/nicia-ai/typegraph/pull/139) [`f1ea17c`](https://github.com/nicia-ai/typegraph/commit/f1ea17cafab281d61741b1d2ad0b26a769efaa5a) Thanks [@pdlug](https://github.com/pdlug)! - Cross-store atomicity: share one transaction across the TypeGraph store and an
external Drizzle connection ([#134](https://github.com/nicia-ai/typegraph/issues/134)).
Applications that persist into the same database through two layers — Drizzle
for relational rows and TypeGraph for graph nodes/edges — previously had no way
to make a write that spans both layers all-or-nothing. `store.transaction()`
and `db.transaction()` each opened a _separate_ transaction on a _separate_
connection, so a failure between the two writes left either a stray relational
row or a committed graph node with a dangling foreign reference.
**What ships (additive — no breaking changes):**
- New `Store.withTransaction(externalTx): TransactionContext`. The caller
owns the transaction; `store.withTransaction(sqlTx)` returns a
transaction-scoped `{ nodes, edges }` bound to that _exact_ connection, so
both layers commit or roll back together. It is driver-agnostic; how you
open the transaction is not.
Async drivers (node-postgres, `neon-serverless` Pool, libsql):
```ts
await db.transaction(async (sqlTx) => {
const connector = await createConnectorRow(sqlTx, input); // Drizzle
const txStore = store.withTransaction(sqlTx);
await txStore.nodes.ArtifactSource.create({
// TypeGraph
connectorId: connector.id,
});
}); // one COMMIT / ROLLBACK
```
Synchronous `better-sqlite3` cannot use `db.transaction(async …)` (its
driver rejects an `async` callback); open the transaction with explicit
`BEGIN`/`COMMIT`/`ROLLBACK` instead and pass the connection to
`withTransaction`. See the "Cross-Store Transactions" recipe for both
shapes.
- New optional `GraphBackend.adoptTransaction(externalTx)` member, implemented
by the Drizzle Postgres and SQLite backends, plus the new `AdoptedTransaction`
type.
**Guarantees.** The adopted context reuses the parent store's already-resolved
schema: it runs no `createStoreWithSchema` / `evolve` / `migrateSchema` and
emits **no DDL inside the caller's business transaction**. Building on [#135](https://github.com/nicia-ai/typegraph/issues/135),
fulltext operations assert the durable materialization marker (a cached
`SELECT`, never DDL) and throw `StoreNotInitializedError` on a
missing/stale/failed marker rather than migrating mid-transaction — so boot the
parent store via `createStoreWithSchema` once at startup. When the backend
cannot provide real rollback (`backend.capabilities.transactions === false`:
`drizzle-orm/neon-http`, Cloudflare D1, SQLite `transactionMode: "none"`),
`withTransaction` throws `ConfigurationError` rather than silently degrading —
a non-atomic fallback is safe for graph-only writes but dangerous for
cross-store flows, where the caller's relational write _would_ still commit.
- [#142](https://github.com/nicia-ai/typegraph/pull/142) [`02c98a9`](https://github.com/nicia-ai/typegraph/commit/02c98a9933c888fcd732053e8cb47991614d2ec9) Thanks [@pdlug](https://github.com/pdlug)! - Transactional writes for Cloudflare Durable Objects SQLite (`do-sqlite`)
([#140](https://github.com/nicia-ai/typegraph/issues/140)).
A store backed by `drizzle(ctx.storage)` previously fell back to
non-transactional behavior, so TypeGraph mutations could not be composed
atomically with a product's own relational ledger tables (e.g.
`document_versions`, `change_events`) inside a Durable Object.
**What ships (additive — no breaking changes):**
- New SQLite `transactionMode: "do-sqlite"`, **auto-detected** for
`drizzle(ctx.storage)`. Such backends now advertise
`capabilities.transactions: true`.
- `store.transaction(async (tx) => …)` and the caller-owned
`store.withTransaction(db)` shape both work on Durable Objects. TypeGraph
delegates to the async storage runner `ctx.storage.transaction(async …)`
(surfaced by Drizzle as `db.$client.transaction`), which rolls back SQL
writes across `await`. Drizzle's own `db.transaction()` on DO is
`ctx.storage.transactionSync` and cannot span an `await`, so it is
deliberately not used. There is no Drizzle transaction handle on DO — the
storage transaction is ambient on the object — so the tx-scoped backend
binds the outer `db`.
```ts
await ctx.storage.transaction(async () => {
const txStore = store.withTransaction(db);
await txStore.nodes.Document.update(documentId, props);
await db.insert(documentVersions).values(versionRow);
await db.insert(changeEvents).values(eventRow);
}); // one storage-transaction COMMIT / ROLLBACK across both layers
```
- A latent detection bug is fixed: drizzle's Durable Objects session class is
`SQLiteDOSession` (not the previously-checked `SQLiteDurableObjectSession`),
so a real `drizzle(ctx.storage)` store was misclassified.
- New `TransactionContext.sql` — the raw Drizzle handle bound to the same
transaction — for graph-owned cross-store writes across **all**
transactional backends (Postgres, libsql, better-sqlite3, do-sqlite):
```ts
await store.transaction(async (tx) => {
await tx.nodes.Document.update(documentId, props);
// tx.sql is the AdoptedTransaction union — cast to your concrete
// Drizzle database type at the call site.
const sqlTx = tx.sql as NodePgDatabase;
await sqlTx.insert(documentVersions).values(versionRow);
await sqlTx.insert(changeEvents).values(eventRow);
});
```
This is the graph-owned counterpart of `store.withTransaction` (where the
caller owns the boundary). On Postgres/libsql it is a correctness
requirement — the outer `db` would write on a different connection and
escape the transaction. `tx.sql` is `undefined` only on the
non-transactional fallback. Its static type is the `AdoptedTransaction`
union; cast to your concrete Drizzle database type at the call site.
**Guarantees.** Building on [#135](https://github.com/nicia-ai/typegraph/issues/135), no schema/bootstrap/fulltext DDL ever runs
inside the business transaction: `bootstrapTables` and the durable
materialization marker run outside any storage transaction, while the
schema-version commit uses the `do-sqlite` runner (data only). Boot the parent
store via `createStoreWithSchema` once at object startup.
**Out of scope.** Cloudflare D1 stays `transactionMode: "none"`:
`D1Database.batch(...)` is transactional but not an interactive runner. A
batch-only D1 mode is tracked separately.
- [#138](https://github.com/nicia-ai/typegraph/pull/138) [`bcf1e48`](https://github.com/nicia-ai/typegraph/commit/bcf1e4819754f1839a236d350d70bab9103607ce) Thanks [@pdlug](https://github.com/pdlug)! - Durable, enforced fulltext materialization ([#135](https://github.com/nicia-ai/typegraph/issues/135)).
Strategy-owned fulltext table/index DDL was materialized lazily, guarded by an
**in-memory, per-backend-instance boolean latch** (`fulltextEnsured`), and
interleaved into the read/write data path. That was correct only by accident
(idempotent DDL + a warm process) and at the wrong durability scope; it was
inconsistent with how vector indexes are tracked and it blocked cross-store
transaction adoption ([#134](https://github.com/nicia-ai/typegraph/issues/134)). "Is this graph's fulltext storage materialized?"
is now a **durable, queryable database fact** instead of a process boolean.
**Breaking (behavioral): fulltext now requires an explicit boot step.**
`createStore()` is a synchronous, zero-I/O _attach_ — it never creates tables,
repairs DDL, or writes materialization markers. The durable marker is written
exclusively by the async boot path, `createStoreWithSchema(graph, backend)`,
which must run once at application startup (outside request handlers and
adopted transactions). A fulltext read/write — or a transaction that touches
fulltext — against a database with no valid marker now throws the new
`StoreNotInitializedError` instead of lazily emitting DDL on the hot path.
Consumers already using `createStoreWithSchema` need no changes; consumers
relying on lazy fulltext creation via bare `createStore()` must add a
`createStoreWithSchema` call at boot.
**What ships:**
- New `@nicia-ai/typegraph` exports: `StoreNotInitializedError` and the
`StoreNotInitializedReason` (`"missing" | "stale" | "failed"`) it carries in
`details.reason`.
- New per-deployment table `typegraph_contribution_materializations`, a
sibling of `typegraph_index_materializations` (the declared-index status
table is deliberately left unchanged). Keyed by [#129](https://github.com/nicia-ai/typegraph/issues/129) contribution identity
`(graph_id, logical_name, owner, table_name)`; `signature` is a separate
content-hash column, so a same-identity row with a drifted signature is a
loud error, never a silent re-materialize. Failed re-attempts preserve the
prior success timestamp via the same COALESCE rule as index
materializations.
- New backend primitives (SQLite + Postgres):
`ensureContributionMaterializationsTable`, `getContributionMaterialization`,
`recordContributionMaterialization`, and
`assertRuntimeContributionsInitialized`. `ensureRuntimeContributions`
and `ensureFulltextTable` now take a `graphId` and
route through the durable-marker writer (short-circuiting when the recorded
signature already matches). `createStoreWithSchema` records the marker after
the schema version is resolved, covering the cold-initialize path.
- The six fulltext-touching methods (`upsertFulltext`, `deleteFulltext`,
`upsertFulltextBatch`, `deleteFulltextBatch`, `fulltextSearch`,
`hardDeleteNode`) stop ensuring and instead assert the durable marker
(resolved once per backend instance, cached). The transaction path performs
zero DDL: the tx-scoped backend's fulltext methods assert the cached marker
at point of use (a `SELECT`, never `CREATE`), so a transaction that never
touches fulltext requires no fulltext initialization and one that does runs
pure DML on the adopted transaction.
This makes [#134](https://github.com/nicia-ai/typegraph/issues/134) (cross-store transaction adoption) sound by construction: a
transaction-adopting primitive consults the durable fact and refuses with a
clear `StoreNotInitializedError` if the store was never initialized, instead
of emitting `CREATE INDEX` inside the caller's business transaction.
- [#136](https://github.com/nicia-ai/typegraph/pull/136) [`9aa2d31`](https://github.com/nicia-ai/typegraph/commit/9aa2d31b8beddbf8f0dea08c4d9435ab3255b580) Thanks [@pdlug](https://github.com/pdlug)! - Unified `TableContribution` contract for strategy-owned tables ([#129](https://github.com/nicia-ai/typegraph/issues/129)).
"What tables does TypeGraph own?" was previously split across four
uncoordinated surfaces (Drizzle named exports, tables-factory
recursion, strategy raw DDL, per-table `ensureXTable` methods). Adding
a new strategy- or backend-owned table without also wiring an
`ensureXTable` + bootstrap probe re-opened the gap [#128](https://github.com/nicia-ai/typegraph/issues/128) closed. This
refactor routes every owned table through one shape.
**Breaking (custom `FulltextStrategy` implementers only):**
`FulltextStrategy.generateDdl(tableName): string[]` is replaced by
`ownedTables(primaryTableName): readonly StrategyTableContribution[]`.
A strategy now _declares_ its tables, Drizzle-free, as already
authoritative contributions (`logicalName`, `owner`, resolved
`tableName`, idempotent `createDdl` for the table **and its supporting
indexes**, `runtimeEnsure`). The two shipped strategies
(`tsvectorStrategy`, `fts5Strategy`) and all internal callers are
migrated; consumers using only the shipped strategies need no changes.
**What ships:**
- New `@nicia-ai/typegraph` export: `TableContribution` and
`StrategyTableContribution` (its strategy-declaration alias). Each
contribution carries a stable, deployment-independent `logicalName`
plus the resolved physical `tableName` (distinct identity vs.
drift-signature inputs) — the prerequisite that lets [#135](https://github.com/nicia-ai/typegraph/issues/135) make
fulltext materialization a durable, decidable fact instead of an
in-memory per-backend latch.
- `postgresContributions()` / `sqliteContributions()` are the single
source of truth for DDL generation and the bootstrap ensure.
`generatePostgresDDL` / `generateSqliteDDL` iterate contributions;
the `table === tables.fulltext` reference-identity hack is gone from
DDL generation. drizzle-kit visibility for the default Postgres
strategy comes from the schema barrel exporting the matching
`tables.fulltext` object (one object, not two); a non-default
strategy exports its own.
- New backend method `ensureRuntimeContributions()`, which runs each
`runtimeEnsure` contribution's full idempotent `createDdl` (table +
supporting indexes) so a partial state (table present, index
missing) self-heals — not a probe-and-skip.
`loadActiveSchemaWithBootstrap` calls it scoped to `runtimeEnsure`
contributions only (the strategy-owned fulltext table today), so
startup does not regress into broad DDL/probing across every table.
`ensureFulltextTable` is retained as a thin back-compat wrapper.
DDL statement ordering changes from "all CREATE TABLE, then all CREATE
INDEX, then fulltext" to per-contribution "table then its own
indexes". Safe because TypeGraph's tables carry no cross-table foreign
keys; raw migration SQL byte output differs accordingly.
Prerequisite for [#135](https://github.com/nicia-ai/typegraph/issues/135) (durable fulltext materialization), which is in
turn the prerequisite for [#134](https://github.com/nicia-ai/typegraph/issues/134) (cross-store transaction adoption).
## 0.25.1
### Patch Changes
- [#130](https://github.com/nicia-ai/typegraph/pull/130) [`dbe52dc`](https://github.com/nicia-ai/typegraph/commit/dbe52dc5d1346543b5aab5b4380df85bdbf66750) Thanks [@pdlug](https://github.com/pdlug)! - Fix drizzle-kit-managed fulltext bootstrap gap on both Postgres and SQLite ([#128](https://github.com/nicia-ai/typegraph/issues/128)).
Consumers managing typegraph storage via `drizzle-kit push` /
`drizzle-kit generate` (`export * from "@nicia-ai/typegraph/postgres"`
or `…/sqlite"`) got every typegraph table EXCEPT
`typegraph_node_fulltext`. The fulltext table was strategy-owned raw
DDL — the schema modules exposed only `fulltextTableName: string`,
not a Drizzle table — so drizzle-kit silently skipped it. The
`bootstrapTables` fallback in `loadActiveSchemaWithBootstrap` only
fires on a missing-table error from `getActiveSchema`; once
drizzle-kit had created `typegraph_schema_versions`, that branch
stopped triggering and `searchable()` writes failed at runtime with
`relation/table "typegraph_node_fulltext" does not exist`.
Two fixes ship together:
- **`backend.ensureFulltextTable()` (both backends).** A focused
narrow-ensure that mirrors the existing
`ensureIndexMaterializationsTable` /
`ensureKindRemovalsTable` /
`ensureReconciliationMarkersTable` idiom — single-table
`CREATE … IF NOT EXISTS`, no Postgres SHARE-lock deadlock under
concurrent replica startup. The backend wraps every method that
emits fulltext SQL (`upsertFulltext` / `deleteFulltext` and their
batch variants, `fulltextSearch`, and `hardDeleteNode` whose
cascade unconditionally deletes from the fulltext table) to call
the ensure first. A per-backend latch makes the per-call cost a
single boolean check after the first invocation, so the wrapping
is safe on the hot path. `loadActiveSchemaWithBootstrap` also
calls the ensure as a belt-and-suspenders for the
`createStoreWithSchema` path. Together these cover both async
schema-aware boot AND the sync `createStore` path — the bare
bootstrap-load probe alone would miss the latter. This is the
canonical fix and the **only** viable one for SQLite (FTS5
virtual tables aren't drizzle-kit-modelable).
- **Typed Drizzle pg-core table for `tsvectorStrategy` (Postgres
only).** `createPostgresTables()` now returns
`tables.fulltext` — a typed `pgTable` for the default
`tsvector` + GIN stack — alongside `tables.fulltextTableName`.
The new `fulltext` named export is included in
`@nicia-ai/typegraph/postgres`, so `export *` lets drizzle-kit
generate migrations for the fulltext table the same way it does
for `nodes`/`edges`/etc. Custom `tsvector`/`regconfig` column
types are exported alongside the existing `vector` column.
`generatePostgresDDL` deliberately skips the typed Drizzle table
(the column-walker can't reproduce the `GENERATED ALWAYS AS (…)
STORED`clause) and continues to defer to
`tsvectorStrategy.generateDdl()` for the runtime DDL emit. The
two paths agree byte-for-byte; a drift sentinel test catches any
divergence.
Alternate Postgres fulltext strategies (pg_trgm, ParadeDB,
pgroonga) still own their own DDL via
`FulltextStrategy.generateDdl()` and the bootstrap probe runs it.
Drizzle-kit consumers using a non-default strategy must override
`tables.fulltext` in their schema barrel with their strategy's
own table.
Documented the SQLite FTS5 virtual-table caveat and the new
Postgres `tables.fulltext` export in
`apps/docs/src/content/docs/integration.md`.
## 0.25.0
### Highlights
0.25.0 is the runtime schema evolution release. It adds graph extensions,
unified index declarations and materialization, dynamic queries over
runtime-declared kinds, runtime access to compiled props schemas, and a safer
transactional schema-version commit path.
- Graph extensions let applications commit reviewed JSON schema proposals as
durable TypeGraph schema versions without redeploying application code.
- Compile-time, graph-extension, relational, and vector indexes now share one
canonical declaration channel and flow through `Store.materializeIndexes()`.
- Dynamic query builder methods let typed queries traverse runtime-declared node
and edge kinds while still validating kind names, endpoints, and field
predicates at query-build time.
- `Store` now exposes compiled Zod props schemas for compile-time and
graph-extension kinds through `getNodePropsSchema`, `getEdgePropsSchema`, and
their `OrThrow` variants.
- Node and edge definitions now accept JSON-serializable `annotations` for
consumer-owned metadata such as UI hints, audit policy, and provenance.
### Upgrade notes
- Existing deployments with manually managed schemas should add the one-active
schema-version partial unique index:
`typegraph_schema_versions_one_active_per_graph_idx` on `(graph_id)` where
`is_active` is true (`TRUE` on Postgres, `1` on SQLite).
- Manually managed schemas should also sync the generated DDL for the new
TypeGraph status tables, including `typegraph_index_materializations`,
`typegraph_kind_removals`, and `typegraph_reconciliation_markers`.
- Run schema migrations from a transactional backend. Edge or HTTP-only
non-transactional drivers can continue serving normal reads and writes after
the schema is established.
- Tests that deep-compare the full `SchemaValidationResult` object may need to
switch to partial matching because `initialized` and `migrated` now include
`committedRow`.
#### Custom backends and index consumers
These changes affect custom `GraphBackend` implementations and advanced index
consumers; ordinary `createStoreWithSchema`, query, and collection callers
should not need code changes.
- `insertSchema` and `setActiveSchema` were removed from `GraphBackend`.
Implement `commitSchemaVersion` and `setActiveVersion` instead.
- `commitSchemaVersion` and `setActiveVersion` require transactional behavior.
Non-transactional drivers such as Cloudflare D1, Durable Objects,
`drizzle-orm/neon-http`, and SQLite backends configured with
`transactionMode: "none"` refuse these primitives for schema commits.
- `createFulltextIndex` and `dropFulltextIndex` were removed from
`GraphBackend`; fulltext storage remains owned by the active backend fulltext
strategy.
- The old `NodeIndex`, `EdgeIndex`, and `TypeGraphIndex` types were removed from
`@nicia-ai/typegraph/indexes`. Use `NodeIndexDeclaration`,
`EdgeIndexDeclaration`, or `IndexDeclaration`.
- Custom backends should add the new optional materialization/removal primitives
when they want first-class support for index status loading, removal
reconciliation markers, and vector index materialization.
### Minor Changes
#### New APIs
- `defineGraphExtension(input)` and `validateGraphExtension(input, options?)`.
- `Store.evolve`, `Store.deprecateKinds`, `Store.undeprecateKinds`,
`Store.removeKinds`, `Store.materializeRemovals`, and dynamic collection
accessors for graph-extension kinds.
- `defineGraph({ indexes })`, `defineNodeIndex`, `defineEdgeIndex`, `andWhere`,
`orWhere`, `notWhere`, and the `@nicia-ai/typegraph/indexes` subpath for
advanced index tooling.
- `Store.materializeIndexes(options?)` plus `MaterializeIndexesResult` status
reporting.
- `embedding(dimensions, options?)` vector index options and exported vector
index declaration/configuration types.
- `fromDynamic`, `traverseDynamic`, `optionalTraverseDynamic`, and `toDynamic`
on the query builder.
- `SchemaValidationResult.initialized` and `.migrated` now include
`committedRow: SchemaVersionRow`.
- `SqlTableNames` now includes `uniques` so cleanup paths can honor custom
physical table names.
#### Performance and reliability
- Schema commits now use a transactional `commitSchemaVersion` backend primitive
instead of the old insert-then-activate sequence, fixing the orphan schema-row
crash window.
- `materializeIndexes` bulk-loads materialization status in one round trip and
records per-index drift/failure state in `typegraph_index_materializations`.
- `materializeRemovals` records a reconciliation watermark, honors custom table
names, and cleans secondary embedding/fulltext/unique rows for removed node
kinds.
- Schema hash and parsed-schema caches avoid repeated serialization, SHA-256,
and Zod parse work on no-change startup and repeated store creation.
- Graph-extension merge/compile paths share caches and fast paths for idempotent
or partially overlapping evolves.
- Postgres vector-index drops now run per-metric DDL concurrently.
#### Pull requests
- [#103](https://github.com/nicia-ai/typegraph/pull/103) - Add per-kind
`annotations`.
- [#106](https://github.com/nicia-ai/typegraph/pull/106) - Add atomic schema
version commits.
- [#107](https://github.com/nicia-ai/typegraph/pull/107) - Add compile-time
index declarations to graph definitions and serialized schemas.
- [#112](https://github.com/nicia-ai/typegraph/pull/112) - Add
`Store.materializeIndexes`.
- [#117](https://github.com/nicia-ai/typegraph/pull/117) - Unify vector indexes
with the index declaration channel.
- [#118](https://github.com/nicia-ai/typegraph/pull/118) - Add graph
extensions.
- [#125](https://github.com/nicia-ai/typegraph/pull/125) - Add dynamic query
traversal methods.
- [#126](https://github.com/nicia-ai/typegraph/pull/126) - Expose runtime Zod
props schemas.
- [#127](https://github.com/nicia-ai/typegraph/pull/127) - Pre-release cleanup
and performance pass.
## 0.24.1
### Patch Changes
- [#99](https://github.com/nicia-ai/typegraph/pull/99) [`755df5a`](https://github.com/nicia-ai/typegraph/commit/755df5a8d8114fbc72047f436132bfe105d02823) Thanks [@pdlug](https://github.com/pdlug)! - Internal: dependency bump pass (patch/minor only — TypeScript and `@types/node` held back as separate majors).
Notable runtime/peer-relevant moves: `nanoid` 5.1.9 → 5.1.11 (only published runtime dep); dev/peer `zod` 4.3.6 → 4.4.3, `@libsql/client` 0.17.2 → 0.17.3.
Also drops the `export` keyword on 14 types that were never reachable through any public entry point (`src/index.ts`, `./schema`, `./indexes`, `./sqlite`, `./postgres`, etc.) and had no internal importers. These were leaked-internal types surfaced by a sensitivity change in `knip` 6.11. No symbol on the documented API surface changed; consumers importing only via the package's declared `exports` paths are unaffected.
## 0.24.0
### Minor Changes
- [#97](https://github.com/nicia-ai/typegraph/pull/97) [`8747df8`](https://github.com/nicia-ai/typegraph/commit/8747df8c003589f985e86ca654cf796fa5230e34) Thanks [@pdlug](https://github.com/pdlug)! - SQLite: implement `backend.vectorSearch`, unblocking `store.search.hybrid()` on SQLite.
The hybrid retrieval facade has been Postgres-only since [#88](https://github.com/nicia-ai/typegraph/issues/88): SQLite shipped fulltext (`fulltextSearch`) and embedding persistence (`upsertEmbedding` / `deleteEmbedding`), but never the `vectorSearch` method that `executeHybridSearch` requires for RRF fusion. `.similarTo()` on SQLite still worked because the predicate path goes through the query compiler, not the backend facade — but anyone reaching for `store.search.hybrid()` on SQLite hit `ConfigurationError: Backend does not support vector search`.
This release wires up the SQLite half of that contract:
- `buildVectorSearchSqlite` issues `vec_distance_cosine` / `vec_distance_l2` against the embeddings BLOB column, mirroring the Postgres SQL shape (same WHERE / ORDER BY / score expression / minScore semantics).
- `createSqliteBackend` exposes `vectorSearch` on the backend object whenever `hasVectorEmbeddings` is true (parallel to the existing `upsertEmbedding` gate).
- `inner_product` is rejected — sqlite-vec has no `vec_distance_ip` function.
```typescript
import { createLocalSqliteBackend } from "@nicia-ai/typegraph/sqlite/local";
const { backend } = createLocalSqliteBackend(); // sqlite-vec auto-loaded
const store = createStore(graph, backend);
const ranked = await store.search.hybrid("Document", {
limit: 10,
vector: { fieldPath: "embedding", queryEmbedding },
fulltext: { query: "climate adaptation" },
});
```
**Performance.** On the standard search-shapes bench (500 docs, 384-dim), SQLite hybrid clocks in at **0.8ms** — about 3× faster than PostgreSQL's 2.5ms on the same shape. The bench harness now measures it on both backends; the previously-blank SQLite cell in the search comparison table is filled in.
## 0.23.0
### Minor Changes
- [#95](https://github.com/nicia-ai/typegraph/pull/95) [`6f3bf30`](https://github.com/nicia-ai/typegraph/commit/6f3bf30b4ac7c51a5528e1001dc97e05146801b7) Thanks [@pdlug](https://github.com/pdlug)! - PostgreSQL: official postgres-js / Neon support, server-side prepared statements on the fast path, and a `refreshStatistics()` API.
**Four drivers supported.** `createPostgresBackend` has always been driver-agnostic, but only `node-postgres` was covered in CI. This release adds:
- **`drizzle-orm/postgres-js`** — full adapter + integration suite coverage (~250 tests run against both `pg` and `postgres-js` against a real PostgreSQL).
- **`drizzle-orm/neon-serverless`** — `@neondatabase/serverless` Pool over WebSockets. Wiring smoke tests verify driver detection, fast-path routing, Date→string normalization, and capability surface; the shared code paths are exercised by the `pg` integration suite since this driver is pg-Pool-protocol-compatible.
- **`drizzle-orm/neon-http`** — `@neondatabase/serverless` `neon(url)` over HTTP. Auto-detected so `capabilities.transactions` is set to `false` (HTTP can't hold a session); single-statement reads, writes, and migrations work normally. Smoke tests verify the detection and capability override.
Same `createPostgresBackend(db)` entry point regardless of driver.
```typescript
// postgres-js
import postgres from "postgres";
import { drizzle } from "drizzle-orm/postgres-js";
const backend = createPostgresBackend(
drizzle(postgres(process.env.DATABASE_URL)),
);
// Neon serverless (edge runtimes)
import { Pool } from "@neondatabase/serverless";
import { drizzle } from "drizzle-orm/neon-serverless";
const backend = createPostgresBackend(
drizzle(new Pool({ connectionString: env.NEON_DATABASE_URL })),
);
```
**On Neon HTTP vs WebSockets:** both work. The HTTP driver (`drizzle-orm/neon-http`) is best for stateless edge workloads — TypeGraph auto-disables transactions since HTTP can't hold a session, and `store.transaction(...)` falls through to non-transactional sequential execution. Use the WebSocket driver (`drizzle-orm/neon-serverless`) when you need atomic multi-statement writes.
**~6× faster on multi-hop traversals via server-side prepared statements.** The execution adapter now uses `node-postgres`'s named prepared statements transparently — each unique compiled SQL string gets a stable counter-derived statement name (cached by SQL text), so PostgreSQL caches the plan after first execution. Combined with routing `execute()` through the fast path directly (skipping Drizzle's session wrapper), this drops the 3-hop benchmark from ~7.5ms to ~0.8ms median, putting TypeGraph-on-PostgreSQL at parity with Neo4j on every single-query and multi-hop shape we measure.
The change is invisible to callers; existing code keeps working. postgres-js is unchanged (it handles its own preparation internally).
**New `store.refreshStatistics()` / `backend.refreshStatistics()` API.** Call once after a large initial import or bulk backfill. Without fresh stats, the planner can pick suboptimal execution plans — on PostgreSQL this is the difference between a 0.5ms and 5ms forward traversal; on SQLite it's the difference between 0.9ms and 23ms fulltext search. Autovacuum / background statistics catch up eventually, but explicit invocation gives correct latencies immediately.
```typescript
for (const batch of batches) {
await store.nodes.Document.bulkCreate(batch);
}
await store.refreshStatistics();
```
Implementations: SQLite runs `ANALYZE`; PostgreSQL runs `ANALYZE` on TypeGraph-managed tables only. Costs ~20ms on SQLite, ~80ms on PostgreSQL at the sizes this library is designed for.
**Type surface changes:**
- `GraphBackend` now requires a `refreshStatistics(): Promise` method. `TransactionBackend` still excludes it (statistics refresh isn't meaningful inside a transaction). External `GraphBackend` implementations (uncommon) need to add a no-op or proper implementation.
- `PostgresBackendOptions` adds an optional `capabilities?: Partial` for users who need to override capability flags (e.g., for custom HTTP-style drivers).
- `PostgresBackendOptions` also adds `prepareStatements?: boolean` (default `true`) and `preparedStatementCacheMax?: number` (default `256`). The prepared-statement name cache is now LRU-bounded so high-cardinality SQL text doesn't grow unbounded in either the Node process or in PostgreSQL's per-session prepared-statement memory. Set `prepareStatements: false` when pooling through pgbouncer in transaction-pool mode.
See [`backend-setup`](https://typegraph.dev/backend-setup#choosing-a-postgresql-driver) for the runtime-to-driver matrix, per-driver setup snippets, and post-bulk-load guidance.
## 0.22.0
### Minor Changes
- [#93](https://github.com/nicia-ai/typegraph/pull/93) [`1e9ae18`](https://github.com/nicia-ai/typegraph/commit/1e9ae18c0219c8168f0584b65b41a9ec2c564b60) Thanks [@pdlug](https://github.com/pdlug)! - Add `countEdges(edgeAlias)` and `countDistinctEdges(edgeAlias)` — edge-count aggregators that skip the target-node join in the count aggregate fast path.
The default `count(targetAlias)` counts edges whose target node is currently live under the query's temporal mode, which requires joining the edges to the target node table on every aggregation. For the common "how many follow relationships does this user have?" question, that join is unnecessary work: you want to count edges, not reach through each edge to validate the target.
```typescript
import { count, countEdges, field } from "@nicia-ai/typegraph";
const result = await store
.query()
.from("User", "u")
.optionalTraverse("follows", "e", { expand: "none" })
.to("User", "target")
.groupByNode("u")
.aggregate({
name: field("u", "name"),
// Counts live edges, regardless of target-node validity.
// Skips the typegraph_nodes join entirely — ~1.7x faster on
// SQLite, ~1.35x on PostgreSQL at benchmark scale.
followCount: countEdges("e"),
// Counts edges to live targets. Keeps the target-node join
// so the target's temporal window is honored.
liveFollowCount: count("target"),
})
.execute();
```
**When to use which:**
- `count(targetAlias)` — when the semantic question is "how many of this user's follows point to a live user?" The target-node join enforces the target's `validTo` / `deleted_at` filters.
- `countEdges(edgeAlias)` — when the semantic question is "how many follow relationships does this user have?" The edge's own temporal and deletion filters are enforced; target validity is not consulted.
- `countDistinctEdges(edgeAlias)` — same semantics as `countEdges` but with `COUNT(DISTINCT ...)`. Useful under ontology-driven expansions where the same edge can appear multiple times in join output.
The two can be mixed in one aggregate. When present together, the compiler keeps the target-node join but switches it to a `LEFT JOIN` with node-side filters pushed into the `ON` clause so edge counts reflect all live edges while node counts only reflect edges to live targets.
No change to existing `count(...)` behavior. This is purely additive — code that currently uses `count("targetAlias")` continues to count live targets exactly as before.
### Patch Changes
- [#93](https://github.com/nicia-ai/typegraph/pull/93) [`1e9ae18`](https://github.com/nicia-ai/typegraph/commit/1e9ae18c0219c8168f0584b65b41a9ec2c564b60) Thanks [@pdlug](https://github.com/pdlug)! - Push `LIMIT` past `GROUP BY` in the count aggregate fast path when it's safe.
When `groupByNode(...).aggregate({ x: count(alias) })` is paired with an optional traversal and a `.limit(n)` that doesn't depend on the aggregate (no `ORDER BY`, or an `ORDER BY` restricted to group keys), the compiler now emits the `LIMIT` inside the start CTE. The `GROUP BY` runs over `n` rows instead of the full start set — `O(limit)` grouping work instead of `O(|start|)`. When `OFFSET` is also set, it rides along with the `LIMIT` into the start CTE and the outer `SELECT` drops its own `LIMIT`/`OFFSET` so neither clause is double-applied.
The fast path also picks `INNER JOIN` over `LEFT JOIN` for the target-node join whenever a `whereNode()` predicate applies to the target alias, so those predicates constrain every aggregate — including `countEdges(...)`. `LEFT JOIN` remains the strategy when only temporal/delete filters apply to the target, so `countEdges` and `count(target)` can coexist in one query with divergent semantics.
No change to query semantics — aggregate counts still reflect the same `count(target)` as before, including the target node's temporal and deletion filters. No change to aggregate queries without a `LIMIT`. No change on SQLite or PostgreSQL query shapes outside the fast path.
Measured impact: scopes down group-by work for "top-N by count"-style aggregate queries. No impact on the blog-post benchmark's full-graph aggregate (which measures the ungrouped 1,200-user case and intentionally runs without a `LIMIT`).
- [#93](https://github.com/nicia-ai/typegraph/pull/93) [`1e9ae18`](https://github.com/nicia-ai/typegraph/commit/1e9ae18c0219c8168f0584b65b41a9ec2c564b60) Thanks [@pdlug](https://github.com/pdlug)! - Fix `generateSqliteDDL` and `generatePostgresMigrationSQL` emitting `(unknown, unknown, ...)` for indexes threaded through `createSqliteTables({}, { indexes })` or `createPostgresTables({}, { indexes })`.
The DDL generator's SQL-chunk flattener didn't handle two cases that appear inside index expression keys: Drizzle column references nested inside a SQL stream (whose `.getSQL()` wraps the column back inside a self-referential SQL object, causing the previous logic to recurse and fall through to `"unknown"`), and `StringChunk` values stored as single-element arrays (`[""]`).
Expression indexes now emit correctly in both dialects, e.g.
```sql
CREATE INDEX IF NOT EXISTS "idx_tg_node_user_city_cov_name_…" ON "typegraph_nodes"
("graph_id", "kind", (json_extract("props", '$."city"')), (json_extract("props", '$."name"')));
```
Added a regression test in `tests/indexes.test.ts` asserting that DDL from `createSqliteTables`/`createPostgresTables` never contains `(unknown` and includes the expected column and `json_extract` / `ARRAY['…']` expressions.
- [#93](https://github.com/nicia-ai/typegraph/pull/93) [`1e9ae18`](https://github.com/nicia-ai/typegraph/commit/1e9ae18c0219c8168f0584b65b41a9ec2c564b60) Thanks [@pdlug](https://github.com/pdlug)! - Emit `NOT MATERIALIZED` on PostgreSQL traversal and start CTEs so the planner can inline them and see their inner row statistics.
PostgreSQL defaults to materializing any CTE referenced more than once. TypeGraph's traversal compilation references each CTE twice — once from the next hop's join, once from the final SELECT — which triggers materialization under the default rules. Materialized CTEs have opaque statistics to the planner, causing poor join orderings and wildly off row estimates on multi-hop queries over larger graphs.
Introduces a `emitNotMaterializedHint` dialect capability (`true` for PostgreSQL, `false` for SQLite, which ignores the hint entirely) and threads it through the start-CTE and traversal-CTE emitters. The hint matches what an expert would write by hand for the same query shape.
Impact on the TypeGraph benchmark suite:
- Multi-hop traversal plans no longer carry opaque materializations, so the planner picks index-scan orderings appropriate to the starting row's selectivity.
- No visible change on SQLite (the hint is not emitted).
- Guards against regressions on larger graphs where materialized CTE plans degenerate into cross-product-plus-filter.
- [#93](https://github.com/nicia-ai/typegraph/pull/93) [`1e9ae18`](https://github.com/nicia-ai/typegraph/commit/1e9ae18c0219c8168f0584b65b41a9ec2c564b60) Thanks [@pdlug](https://github.com/pdlug)! - Persist vector embeddings on the SQLite backend when sqlite-vec is loaded.
Previously, `store.nodes.X.create({ ..., embedding: [...] })` on SQLite validated the embedding and inserted the node, but the embedding itself was silently dropped — the SQLite backend didn't implement `upsertEmbedding`/`deleteEmbedding`, so the store's embedding-sync path quietly no-op'd. Vector predicates like `d.embedding.similarTo(q, 20, { metric: "cosine" })` then ran against an empty `typegraph_node_embeddings` table and returned zero rows without error.
This release wires up both methods on the SQLite backend. They encode embeddings to `vec_f32('[...]')` BLOBs on write and rely on sqlite-vec at query time — same storage shape the existing `.similarTo()` compilation already targets. Activation is opt-in via a new `hasVectorEmbeddings` option on `createSqliteBackend` so callers that haven't loaded sqlite-vec don't hit `no such function: vec_f32` at write time. `createLocalSqliteBackend` best-effort-loads sqlite-vec at startup and flips the option automatically, so the common local setup works without configuration.
```typescript
// Local backend: sqlite-vec is loaded automatically when installed.
const { backend } = createLocalSqliteBackend();
// BYO drizzle connection: pass hasVectorEmbeddings after loading sqlite-vec.
import sqliteVec from "sqlite-vec";
sqliteVec.load(sqlite);
const backend = createSqliteBackend(drizzle(sqlite), {
tables,
hasVectorEmbeddings: true,
});
```
`getEmbedding` and the hybrid-search facade (`store.search.hybrid(...)`) remain PostgreSQL-only — decoding the raw BLOB back to `number[]` via `vec_to_json` and exposing a hybrid-search backend method are tracked separately.
## 0.21.0
### Highlights
TypeGraph 0.21 adds full-text search and hybrid vector/text retrieval. Mark string fields with `searchable()` and TypeGraph maintains native PostgreSQL tsvector/GIN or SQLite FTS5 storage. The node-level `$fulltext.matches()` predicate composes with property filters, graph traversal, and vector similarity; `store.search` provides full-text and hybrid helpers with reciprocal-rank fusion and optional snippets.
Full-text behavior is supplied by a pluggable `FulltextStrategy`, with query-mode, language, prefix, and highlighting capabilities checked against the active strategy. Existing graph data can be indexed through the paginated `rebuildFulltext()` maintenance API.
### Upgrade notes
- Node and edge property names beginning with `$` are now reserved for query accessors. Rename those fields before upgrading; graph definition rejects them with `ConfigurationError`.
- Use `store.search.rebuildFulltext()` to index existing data after declaring searchable fields. It reports skipped invalid properties and is a maintenance operation rather than a transactionally frozen scan of the whole graph.
- Query options must be supported by the configured strategy. SQLite FTS5 cannot honor per-query language overrides; unsupported modes, overrides, and snippet requests are refused.
- `findNodesByKind` now breaks equal creation-time ties by ID. Callers should not rely on the previously unspecified order.
### Minor Changes
- [#88](https://github.com/nicia-ai/typegraph/pull/88) [`6f681d5`](https://github.com/nicia-ai/typegraph/commit/6f681d59f16ef7d7651627999cce6cada01d024e) Thanks [@pdlug](https://github.com/pdlug)! - Add fulltext search and hybrid (vector + fulltext) retrieval. Declare `searchable()` string fields on any node schema and TypeGraph keeps a native FTS index in sync — `tsvector` + GIN on PostgreSQL, FTS5 on SQLite. Query it through a node-level `n.$fulltext.matches()` predicate that composes with metadata filters, graph traversal, and vector similarity in one SQL statement.
```typescript
import { defineNode, searchable, embedding } from "@nicia-ai/typegraph";
const Document = defineNode("Document", {
schema: z.object({
title: searchable({ language: "english" }),
body: searchable({ language: "english" }),
tenantId: z.string(),
embedding: embedding(1536),
}),
});
// Fulltext + metadata filter in a single query
const results = await store
.query()
.from("Document", "d")
.whereNode("d", (d) =>
d.$fulltext.matches("climate change", 20).and(d.tenantId.eq(tenant)),
)
.select((ctx) => ctx.d)
.execute();
// Hybrid: vector + fulltext fused with Reciprocal Rank Fusion at the SQL layer
const hybrid = await store
.query()
.from("Document", "d")
.whereNode("d", (d) =>
d.$fulltext
.matches("climate", 50)
.and(d.embedding.similarTo(queryVector, 50))
.and(d.tenantId.eq(tenant)),
)
.select((ctx) => ctx.d)
.limit(10)
.execute();
// Store-level helper with tunable RRF weights and snippets
const tuned = await store.search.hybrid("Document", {
limit: 10,
vector: { fieldPath: "embedding", queryEmbedding: queryVector },
fulltext: { query: "climate change", includeSnippets: true },
fusion: { method: "rrf", k: 60, weights: { vector: 1, fulltext: 1.5 } },
});
```
Query modes cover `websearch` (Google-style syntax — default), `phrase`, `plain`, and `raw` (dialect-native tsquery / FTS5 MATCH). Highlighting via `ts_headline` / `snippet()` is opt-in per query. No extensions required: Postgres uses the built-in `tsvector` + GIN (works on every managed provider); SQLite uses FTS5 which is statically linked into the standard `better-sqlite3` / `libsql` / `bun:sqlite` distributions. See `/fulltext-search` for the full guide.
### Added
- `n.$fulltext` — node-level fulltext accessor; `.matches(query, k?, options?)` composes against the combined `searchable()` content. `$fulltext` is exposed on every `NodeAccessor`; a runtime guard throws a clear error if the node kind has no `searchable()` fields. `k` defaults to 50.
- `store.search` facade — `store.search.fulltext()`, `store.search.hybrid()`, and `store.search.rebuildFulltext()` grouped under one namespace. Lazy-initialized and cached on first access.
- `FulltextSearchHit`, `VectorSearchHit`, and `HybridSearchHit` are generic over the node type (`FulltextSearchHit`). `store.search.fulltext("Document", ...)` returns hits with `hit.node` narrowed to the Document node shape — no cast required.
- `backend.upsertFulltextBatch` + `backend.deleteFulltextBatch` — symmetric batched fulltext primitives. Homogeneous batch shape, duplicate-nodeId dedupe last-write-wins, per-row fallback when unset.
- `store.search.rebuildFulltext(nodeKind?, { pageSize?, maxSkippedIds? })` — rebuilds the fulltext index from existing node data using keyset pagination on `id` (stable under shared timestamps and light concurrent writes). Transacts per page; cleans stale rows for soft-deleted nodes; validates `pageSize` as a positive integer; counts corrupt / non-object props as `skipped` and surfaces offending IDs via `skippedIds` without aborting. `maxSkippedIds` (default 10,000) lets operators investigating systemic corruption collect the full list. Concurrent hard-deletes between pages may be missed — document as maintenance operation.
- Keyset pagination on `findNodesByKind` via new `{ orderBy, after }` params.
- `QueryBuilder.fuseWith({ k?, weights? })` — tunable RRF on the query-builder path. Flat `HybridFusionOptions` shape, identical to `store.search.hybrid`'s `fusion` option. Throws at compile time if the query lacks either a `.similarTo()` or `n.$fulltext.matches()`. Shares its validator with `store.search.hybrid({ fusion })` so `method`, `k`, and per-source weights are checked identically on both paths.
- `FulltextStrategy` — pluggable abstraction (exported from the top-level entry) that owns the **entire** SQL pipeline for a dialect's fulltext support: DDL, upsert (single + batch), delete (single + batch), MATCH condition, rank expression, and snippet expression. Ships `tsvectorStrategy` (Postgres built-in `tsvector`) and `fts5Strategy` (SQLite FTS5); dialect adapters expose `fulltext: FulltextStrategy | undefined`. Alternate Postgres stacks (pg_trgm, ParadeDB / pg_search, pgroonga) choose their own column layout, index type, and projection — TypeGraph's operation layer just delegates to the active strategy. Strategies declare prefix-query support explicitly via `FulltextStrategy.supportsPrefix`, so capability discovery stays correct for strategies that support prefix matching via dedicated syntax without advertising raw-mode pass-through.
- Backend-level fulltext strategy override: `createPostgresBackend(db, { fulltext })` and `createSqliteBackend(db, { fulltext })` accept a `FulltextStrategy` that takes precedence over the dialect default. Threaded through to compiler passes, backend-direct search SQL, all write SQL, DDL generation, and capability discovery — so a ParadeDB-backed Postgres `store.search.hybrid()` fuses the same way a tsvector-backed one does, without any call-site changes.
- Option validation: `store.search.fulltext` and `store.search.hybrid` validate caller options against the active `FulltextStrategy` (falling back to `BackendCapabilities.fulltext.{phraseQueries, highlighting, languages}` when no strategy is attached). A `mode` outside `strategy.supportedModes` throws, `includeSnippets: true` on a strategy whose `supportsSnippets` is false throws, and a per-query `language` override on a strategy whose `supportsLanguageOverride` is false (e.g. SQLite FTS5) throws. Advisory warning for unknown languages on strategies that honor overrides. `$fulltext.matches()` is validated against the dialect strategy's `supportedModes` at compile time.
- One-time `console.warn` when a node kind has multiple `searchable()` fields with conflicting `language` values. The first field's language wins on the stored row; the warning makes the silent collapse visible so users know to split multilingual content across dedicated node kinds.
- Snippet highlighting uses `…` consistently across both shipped strategies (`ts_headline` on Postgres, `snippet()` on SQLite). One stylesheet applies everywhere.
- `FulltextSearchResult.score` is always `number`. The Postgres adapter coerces `numeric`-as-string driver returns at the backend boundary so downstream code never sees a union type.
- Hybrid SQL emitter uses a deterministic `COALESCE(fulltext.node_id, embeddings.node_id) ASC` tiebreak, matching the JS-side `localeCompare(nodeId)` tiebreak used by `store.search.hybrid` — both hybrid paths produce identical top-k under RRF score ties.
- Postgres fulltext table schema: `language` is `regconfig` (not `TEXT`) and `tsv` is a `GENERATED ALWAYS AS (to_tsvector("language", "content")) STORED` column. Postgres owns the `content / language → tsv` invariant; the strategy's write SQL doesn't recompute `tsv` inline. The `content` column is populated verbatim, and the per-query `language` override path still accepts a text parameter (cast to `regconfig` at query time). SQLite's FTS5 virtual table is unchanged.
### Changed
- **`defineNode()` / `defineEdge()` reject `$`-prefixed property names.** The `$` namespace is reserved for node-level accessors (starting with `$fulltext`). A `ConfigurationError` is raised at graph-definition time instead of silently shadowing user fields at query time. Rename any such fields before upgrading.
- **`findNodesByKind` offset pagination now has a deterministic tiebreaker** (`ORDER BY created_at DESC, id DESC`). Row order was previously under-specified when `created_at` values collided; callers that happened to rely on an implementation-dependent order may see different tie-breaking.
## 0.20.0
### Minor Changes
- [#85](https://github.com/nicia-ai/typegraph/pull/85) [`12055d0`](https://github.com/nicia-ai/typegraph/commit/12055d053b22cfadd1439c9a667307fae77af6a2) Thanks [@pdlug](https://github.com/pdlug)! - Add Tier 1 graph algorithms on `store.algorithms.*`: `shortestPath`, `reachable`, `canReach`, `neighbors`, and `degree`.
```typescript
// Find the shortest path through a set of edge kinds
const path = await store.algorithms.shortestPath(alice, bob, {
edges: ["knows"],
maxHops: 6,
});
// Enumerate reachable nodes within a depth bound
const reachable = await store.algorithms.reachable(alice, {
edges: ["knows"],
maxHops: 3,
});
// Fast existence check
const connected = await store.algorithms.canReach(alice, bob, {
edges: ["knows"],
});
// k-hop neighborhood (source always excluded)
const twoHop = await store.algorithms.neighbors(alice, {
edges: ["knows"],
depth: 2,
});
// Count incident edges
const total = await store.algorithms.degree(alice, { edges: ["knows"] });
```
All traversal algorithms compile to a single recursive-CTE query and share the dialect primitives used by `.recursive()` and `store.subgraph()`, so SQLite and PostgreSQL yield identical semantics. Node arguments accept either a raw ID string or any object with an `id` field — `Node`, `NodeRef`, and the lightweight records returned by the algorithms themselves all work. See `/graph-algorithms` for the full reference.
- [#85](https://github.com/nicia-ai/typegraph/pull/85) [`12055d0`](https://github.com/nicia-ai/typegraph/commit/12055d053b22cfadd1439c9a667307fae77af6a2) Thanks [@pdlug](https://github.com/pdlug)! - Graph algorithms (`store.algorithms.*`) and `store.subgraph()` now honor the store's temporal model.
**New:** Every algorithm and `store.subgraph()` accept `temporalMode` and `asOf` options, matching the shape already used by `store.query()` and collection reads. When neither is supplied, the resolved mode falls back to `graph.defaults.temporalMode` (typically `"current"`).
```typescript
// Snapshot at a point in time
await store.algorithms.shortestPath(alice, bob, {
edges: ["knows"],
temporalMode: "asOf",
asOf: "2023-01-15T00:00:00.000Z",
});
await store.subgraph(rootId, {
edges: ["has_task"],
temporalMode: "includeEnded",
});
```
The filter applies to both nodes and edges along the traversal, is orthogonal to `cyclePolicy`, and is honored by the shortest-path self-path short-circuit.
**BREAKING:** `store.subgraph()` previously ignored graph temporal settings and filtered only by `deleted_at IS NULL` (equivalent to `"includeEnded"`). It now defaults to `graph.defaults.temporalMode`. Callers that relied on walking through validity-ended rows must pass `temporalMode: "includeEnded"` explicitly. Soft-delete filtering is unchanged under the default `"current"` mode, so most callers see no difference.
### Patch Changes
- [#87](https://github.com/nicia-ai/typegraph/pull/87) [`f52bba6`](https://github.com/nicia-ai/typegraph/commit/f52bba63befe8111d13d04cfb9659371f7061625) Thanks [@pdlug](https://github.com/pdlug)! - Fix SQLite temporal filter timestamp format in graph algorithms and subgraph.
`buildReachableCte`, `resolveTemporalFilter`, and `fetchSubgraphEdges` compiled
temporal filters without passing `dialect.currentTimestamp()`, so on SQLite they
fell back to raw `CURRENT_TIMESTAMP` (`YYYY-MM-DD HH:MM:SS`). Stored
`valid_from` / `valid_to` use ISO-8601 (`YYYY-MM-DDTHH:MM:SS.sssZ`), and because
`T` sorts above space, same-day ISO timestamps compare incorrectly against raw
`CURRENT_TIMESTAMP`. Under `temporalMode: "current"` this caused
`reachable` / `canReach` / `neighbors` / `shortestPath` / `degree` and the
`subgraph` edge hydration to misclassify rows whose `valid_from` or `valid_to`
fell on today's date, disagreeing with `store.query()` and collection reads.
All three call sites now inject the dialect-specific current timestamp
(`strftime('%Y-%m-%dT%H:%M:%fZ','now')` on SQLite, `NOW()` on PostgreSQL),
matching the query compiler.
## 0.19.0
### Minor Changes
- [#83](https://github.com/nicia-ai/typegraph/pull/83) [`206f464`](https://github.com/nicia-ai/typegraph/commit/206f46467342eee6a060c83e057bbf1befb31c1a) Thanks [@pdlug](https://github.com/pdlug)! - **BREAKING:** `store.subgraph()` now returns an indexed result instead of flat arrays.
The result shape changes from `{ nodes: Node[], edges: Edge[] }` to:
```typescript
{
root: Node | undefined;
nodes: ReadonlyMap;
adjacency: ReadonlyMap>;
reverseAdjacency: ReadonlyMap>;
}
```
This eliminates the indexing boilerplate every consumer had to write before traversing the subgraph. Nodes are keyed by ID for O(1) lookup, and edges are organized into forward/reverse adjacency maps keyed by `nodeId → edgeKind`.
Migration:
- `result.nodes` is now a `Map` — use `.size` instead of `.length`, `.values()` instead of direct iteration, `.has(id)` / `.get(id)` instead of `.find()`
- `result.edges` is removed — access edges via `result.adjacency.get(fromId)?.get(edgeKind)` or `result.reverseAdjacency.get(toId)?.get(edgeKind)`
- `result.root` provides the root node directly (no lookup needed)
## 0.18.0
### Minor Changes
- [#80](https://github.com/nicia-ai/typegraph/pull/80) [`0845fa9`](https://github.com/nicia-ai/typegraph/commit/0845fa92a653ed107057cf350414e13745fff8d8) Thanks [@pdlug](https://github.com/pdlug)! - Add first-class libsql backend at `@nicia-ai/typegraph/sqlite/libsql`
### New convenience export
`createLibsqlBackend(client, options?)` wraps `@libsql/client` with automatic DDL
execution and correct async execution profile. The caller retains ownership of the
client, enabling shared-driver setups. Works with local files, in-memory databases,
and remote Turso URLs.
```typescript
import { createClient } from "@libsql/client";
import { createLibsqlBackend } from "@nicia-ai/typegraph/sqlite/libsql";
const client = createClient({ url: "file:app.db" });
const { backend, db } = await createLibsqlBackend(client);
const store = createStore(graph, backend);
```
### Bug fixes for async SQLite drivers
- **`db.get()` crash on empty results** — switched to `db.all()[0]` to work around
Drizzle's `normalizeRow` crash when libsql returns no rows
([drizzle-team/drizzle-orm#1049](https://github.com/drizzle-team/drizzle-orm/issues/1049))
- **`instanceof Promise` check fails for Drizzle thenables** — all SQLite exec helpers
now use unconditional `await` since Drizzle returns `SQLiteRaw` objects that are
thenable but not `Promise` instances
([drizzle-team/drizzle-orm#2275](https://github.com/drizzle-team/drizzle-orm/issues/2275))
### Internal improvements
- Extracted `wrapWithManagedClose()` helper for idempotent backend close with teardown
- Shared adapter and integration test suites now accept async backend factories
- libsql backend runs the full shared test suite (214 tests)
## 0.17.0
### Minor Changes
- [#77](https://github.com/nicia-ai/typegraph/pull/77) [`b9fc057`](https://github.com/nicia-ai/typegraph/commit/b9fc057e0dd62bd0f059bb78a20d18d91b1b87be) Thanks [@pdlug](https://github.com/pdlug)! - feat: support orderBy on edge properties in query builder
The `orderBy` method now accepts edge aliases in addition to node aliases, allowing results to be ordered by properties on traversed edges. This eliminates the need to denormalize ordering fields onto nodes or sort in memory.
```typescript
store
.query()
.from("Person", "p")
.traverse("worksAt", "e")
.to("Company", "c")
.orderBy("e", "salary", "asc") // order by edge property
.select((ctx) => ({ name: ctx.p.name, salary: ctx.e.salary }))
.execute();
```
Also fixes CTE alias resolution for edge aliases in `groupBy` and vector order-by compilation paths.
Closes [#76](https://github.com/nicia-ai/typegraph/issues/76)
## 0.16.2
### Patch Changes
- [#73](https://github.com/nicia-ai/typegraph/pull/73) [`1c95d8e`](https://github.com/nicia-ai/typegraph/commit/1c95d8ec641442cecb38e00fab4c6d10eb162c2c) Thanks [@pdlug](https://github.com/pdlug)! - fix: dispose serialized execution queue on backend close to prevent unhandled rejections
When the SQLite backend's underlying database is destroyed while operations are still queued (e.g., during Cloudflare Workers test teardown), the serialized execution queue now properly disposes pending promises. Calling `backend.close()` signals the queue to suppress errors from in-flight tasks and reject new operations with `BackendDisposedError`.
Fixes [#72](https://github.com/nicia-ai/typegraph/issues/72)
## 0.16.1
### Patch Changes
- [#70](https://github.com/nicia-ai/typegraph/pull/70) [`cebf681`](https://github.com/nicia-ai/typegraph/commit/cebf681c76820db9d63c29f2eb64ed92b1eb3ad5) Thanks [@pdlug](https://github.com/pdlug)! - Widen ID parameters on `DynamicNodeCollection` and `DynamicEdgeCollection` to accept plain `string` instead of branded `NodeId`/`EdgeId` types, removing the need for casts when using the dynamic collection API with IDs from edge metadata, snapshots, or external input.
## 0.16.0
### Minor Changes
- [#66](https://github.com/nicia-ai/typegraph/pull/66) [`2f241a9`](https://github.com/nicia-ai/typegraph/commit/2f241a98fc6ec78702bcaa609e1fce9b5a1ae4f4) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.getNodeCollection(kind)` and `store.getEdgeCollection(kind)` methods for runtime string-keyed collection access. Returns the full collection API with widened generics (`DynamicNodeCollection` / `DynamicEdgeCollection`), or `undefined` if the kind is not registered. Eliminates the need for `Reflect.get(store.nodes, kind) as SomeType` patterns when iterating kinds, resolving nodes from edge metadata, or building generic graph tooling like snapshots and summaries.
## 0.15.0
### Minor Changes
- [#63](https://github.com/nicia-ai/typegraph/pull/63) [`546a7eb`](https://github.com/nicia-ai/typegraph/commit/546a7eb3693141fa8ad236c9aad3333abf635893) Thanks [@pdlug](https://github.com/pdlug)! - `createStoreWithSchema()` now auto-creates base tables on a fresh database. Previously, calling it against a database without pre-existing TypeGraph tables (e.g. a new Cloudflare Durable Object) would throw a raw "no such table" error. The function now detects missing tables and bootstraps them automatically via the new optional `bootstrapTables` method on `GraphBackend`. Both SQLite and PostgreSQL backends implement this method. `createStore()` remains unchanged for users who manage DDL manually.
- [#64](https://github.com/nicia-ai/typegraph/pull/64) [`6b84b42`](https://github.com/nicia-ai/typegraph/commit/6b84b42bd9e626ca01f48d8a5bd3c18c5bfee80d) Thanks [@pdlug](https://github.com/pdlug)! - Add `StoreProjection` utility type for typing reusable helpers that work across graphs sharing a common subgraph. The type projects a store's collection surface onto a subset of node and edge keys, with node constraint names erased so that graphs registering the same node types with different unique constraints remain cross-assignable. Both `Store` and `TransactionContext` are structurally assignable to any `StoreProjection` whose keys are a subset of `G`. Also exports `GraphNodeCollections` and `GraphEdgeCollections` shared mapped types.
### Patch Changes
- [#59](https://github.com/nicia-ai/typegraph/pull/59) [`36742a1`](https://github.com/nicia-ai/typegraph/commit/36742a11f47b2e1903c13ce6abce3e72285f0dbf) Thanks [@pdlug](https://github.com/pdlug)! - Reject empty `fields` arrays at the type level in `defineNodeIndex` and `defineEdgeIndex`. Previously, passing `fields: []` was accepted by TypeScript but threw at runtime. The `fields` property now requires a non-empty tuple, surfacing the error at compile time.
- [#60](https://github.com/nicia-ai/typegraph/pull/60) [`dca5aba`](https://github.com/nicia-ai/typegraph/commit/dca5abad98cdb4df0ca546796f89c6470bdcf680) Thanks [@pdlug](https://github.com/pdlug)! - Export `SchemaValidationResult` and `SchemaManagerOptions` types from the root package entry point so users can type the return value of `createStoreWithSchema()` without reaching into internal subpaths.
## 0.14.0
### Minor Changes
- [#54](https://github.com/nicia-ai/typegraph/pull/54) [`bf6997a`](https://github.com/nicia-ai/typegraph/commit/bf6997afd5889556961977f45bdc9c8d38021902) Thanks [@pdlug](https://github.com/pdlug)! - ### Breaking: default recursive traversal depth lowered from 100 to 10
Unbounded `.recursive()` traversals are now capped at 10 hops instead of 100. Graphs with branching factor _B_ produce O(_B_^depth) rows before cycle detection can prune them — the previous default of 100 made exponential blowup easy to trigger accidentally.
If your traversals relied on the implicit 100-hop cap, add an explicit `.maxHops(100)` call. The `MAX_EXPLICIT_RECURSIVE_DEPTH` ceiling (1000) is unchanged.
### Schema parse validation
Serialized schema documents read from the database are now validated against a Zod schema at the parse boundary. Malformed, truncated, or incompatible schema documents will throw a `DatabaseOperationError` with path-level detail instead of propagating silently. Enum fields (`temporalMode`, `cardinality`, `deleteBehavior`, etc.) are validated against the known literal unions.
### Type safety improvements
- Added `useUnknownInCatchVariables`, `noFallthroughCasesInSwitch`, and `noImplicitReturns` to tsconfig
- Drizzle row mappers now use runtime type checks (`asString`/`asNumber`) instead of unsafe `as` casts
- `NodeMeta` and `EdgeMeta` are now derived from row types via mapped types
- All non-null assertions (`!`) eliminated from source code
- Hardcoded constants extracted to shared `constants.ts`
- Duplicate `fnv1aBase36` function consolidated into `utils/hash.ts`
## 0.13.0
### Minor Changes
- [#52](https://github.com/nicia-ai/typegraph/pull/52) [`1e3da4a`](https://github.com/nicia-ai/typegraph/commit/1e3da4aa814f3baf67a0cb54c9c753508eecf0f0) Thanks [@pdlug](https://github.com/pdlug)! - Add `batchFindFrom`, `batchFindTo`, and `batchFindByEndpoints` to edge collections for use with `store.batch()`.
Edge collection lookup methods (`findFrom`, `findTo`, `findByEndpoints`) execute immediately and cannot participate in `store.batch()`. The new `batchFind*` variants return a `BatchableQuery` instead, enabling edge lookups to share a single transactional connection alongside fluent queries.
```typescript
const [skills, employer, colleague] = await store.batch(
store.edges.hasSkill.batchFindFrom(alice),
store.edges.worksAt.batchFindFrom(alice),
store.edges.knows.batchFindByEndpoints(alice, bob),
);
```
- **`batchFindFrom(from)`** — deferred variant of `findFrom`
- **`batchFindTo(to)`** — deferred variant of `findTo`
- **`batchFindByEndpoints(from, to, options?)`** — deferred variant of `findByEndpoints`, returns 0-or-1 element array
All three preserve the same endpoint type constraints as their immediate counterparts.
Closes [#51](https://github.com/nicia-ai/typegraph/issues/51).
## 0.12.0
### Minor Changes
- [#50](https://github.com/nicia-ai/typegraph/pull/50) [`a59416d`](https://github.com/nicia-ai/typegraph/commit/a59416d8cbc641fd7611ee5d5b0fb115aea59450) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.batch()` for executing multiple queries over a single connection with snapshot consistency.
- **Single connection**: Acquires one connection via an implicit transaction, eliminating pool pressure from parallel `Promise.all` patterns (N connections → 1).
- **Snapshot consistency**: All queries see the same database state — no interleaved writes between results.
- **Typed tuple results**: Returns a mapped tuple preserving each query's independent result type, projection, filtering, sorting, and pagination.
> **Correction (see #325).** The "snapshot consistency" bullet above was never
> accurate and is retained only as the historical record. `batch()` opens its
> implicit transaction without an isolation option, so PostgreSQL runs it at the
> default read-committed isolation and a later query in the batch _can_ observe a
> commit the earlier ones did not. The "single connection" bullet describes the
> transactional path; connection reuse is otherwise the adapter's business, not a
> consequence of `capabilities.transactions`. `batch()` also never pipelined,
> despite the original issue specifying it.
- **`BatchableQuery` interface**: Satisfied by both `ExecutableQuery` (from `.select()`) and `UnionableQuery` (from set operations like `.union()`, `.intersect()`). Exposes `executeOn()` for backend-delegated execution.
- **Minimum 2 queries**: Enforced at the type level — single queries should use `.execute()` directly.
```typescript
const [people, companies] = await store.batch(
store
.query()
.from("Person", "p")
.select((ctx) => ({ id: ctx.p.id, name: ctx.p.name })),
store
.query()
.from("Company", "c")
.select((ctx) => ({ id: ctx.c.id, name: ctx.c.name }))
.orderBy("c", "name", "asc")
.limit(5),
);
// people: readonly { id: string; name: string }[]
// companies: readonly { id: string; name: string }[]
```
Closes [#47](https://github.com/nicia-ai/typegraph/issues/47).
- [#48](https://github.com/nicia-ai/typegraph/pull/48) [`753d9eb`](https://github.com/nicia-ai/typegraph/commit/753d9ebc6aa02f0f01bc52abc1de255b2d1bbd91) Thanks [@pdlug](https://github.com/pdlug)! - Add field-level projection to `store.subgraph()` via a declarative `project` option.
- **Declarative field selection**: Specify which properties to keep per node/edge kind. Projected nodes always retain `kind` and `id`; projected edges always retain structural endpoint fields. Kinds omitted from `project` remain fully hydrated.
- **SQL-level extraction**: Projected property fields are extracted via `json_extract()` / JSONB path expressions directly in the query, avoiding full `props` blob transfer for projected kinds.
- **All-or-nothing metadata**: Include `"meta"` in the field list for the full metadata object, or omit it entirely. No partial metadata selection — the struct is small enough that subsetting adds complexity without meaningful savings.
- **`defineSubgraphProject()` helper**: Curried identity function that preserves literal types for reusable projection configs. Without it, storing a projection in a variable widens field arrays to `string[]`, defeating compile-time narrowing.
- **Type-safe results**: Result types narrow per-kind based on the projection — accessing omitted fields is a compile-time error. Works through both inline literals and `defineSubgraphProject()`.
```typescript
const result = await store.subgraph(rootId, {
edges: ["has_task", "uses_skill"],
maxDepth: 2,
project: {
nodes: {
Task: ["title", "meta"],
Skill: ["name"],
},
edges: {
uses_skill: ["priority"],
},
},
});
// result.nodes — Task has { kind, id, title, meta }; Skill has { kind, id, name }
// result.edges — uses_skill has { id, kind, fromKind, fromId, toKind, toId, priority }
```
Closes [#46](https://github.com/nicia-ai/typegraph/issues/46) (alternative implementation — declarative arrays instead of callbacks).
## 0.11.1
### Patch Changes
- [#41](https://github.com/nicia-ai/typegraph/pull/41) [`68d5432`](https://github.com/nicia-ai/typegraph/commit/68d5432f830978bc05f888134ed1a69644ed97b9) Thanks [@pdlug](https://github.com/pdlug)! - Fix `.paginate()` dropping `id` from selective query results and `orderBy()` mishandling system fields.
- **Fix silent data loss in `.paginate()` + `.select()`**: `FieldAccessTracker.record()` no longer allows a system field (`id`, `kind`) to be downgraded to a props field, which caused the SQL projection to extract from `props->>'id'` (nonexistent) instead of the `id` column.
- **Fix `orderBy()` for system fields**: `orderBy("alias", "id")` now emits `ORDER BY cte.alias_id` instead of `ORDER BY json_extract(cte.alias_props, '$.id')`.
- **Add `gt`/`gte`/`lt`/`lte` to `StringFieldAccessor`**: Enables keyset cursor pagination via `whereNode("a", (a) => a.id.lt(cursor))`.
Fixes [#40](https://github.com/nicia-ai/typegraph/issues/40).
## 0.11.0
### Minor Changes
- [#38](https://github.com/nicia-ai/typegraph/pull/38) [`e26e4a5`](https://github.com/nicia-ai/typegraph/commit/e26e4a5282d9e59ab517a68dede37c38bea2a1e9) Thanks [@pdlug](https://github.com/pdlug)! - Add `createFromRecord()` and `upsertByIdFromRecord()` to `NodeCollection`.
These methods accept `Record` instead of `z.input`, providing an escape hatch for dynamic-data scenarios (changesets, migrations, imports) where the data shape is determined at runtime. Runtime Zod validation is unchanged — only the compile-time type gate is relaxed. The return type remains fully typed as `Node`.
Closes [#37](https://github.com/nicia-ai/typegraph/issues/37).
## 0.10.0
### Minor Changes
- [#33](https://github.com/nicia-ai/typegraph/pull/33) [`da14806`](https://github.com/nicia-ai/typegraph/commit/da14806b665418c7761b5db37641b23eb2914304) Thanks [@pdlug](https://github.com/pdlug)! - Add `store.subgraph()` for typed BFS neighborhood extraction from a root node.
Given a root node ID, traverses specified edge kinds using a recursive CTE and returns all reachable nodes and connecting edges as fully typed discriminated unions.
**Options:**
- `edges` — edge kinds to traverse (required)
- `maxDepth` — maximum traversal depth (default: 10)
- `direction` — `"out"` (default) or `"both"` for undirected traversal
- `includeKinds` — filter returned nodes to specific kinds (traversal still follows all reachable nodes)
- `excludeRoot` — omit the root node from results
- `cyclePolicy` — cycle detection strategy (default: `"prevent"`)
**Type utilities exported:**
- `AnyNode` / `AnyEdge` — discriminated unions of all node/edge runtime types in a graph
- `SubsetNode` / `SubsetEdge` — narrowed unions for a subset of kinds
- `SubgraphOptions` / `SubgraphResult` — fully generic option and result types
- [#35](https://github.com/nicia-ai/typegraph/pull/35) [`0ebc59c`](https://github.com/nicia-ai/typegraph/commit/0ebc59cf1f8d714b0d63c0759d08ed88face022c) Thanks [@pdlug](https://github.com/pdlug)! - Add runtime discriminated union types: `AnyNode`, `AnyEdge`, `SubsetNode`, `SubsetEdge`.
These pure type-level utilities produce discriminated unions of runtime node/edge instances from a graph definition. Unlike `AllNodeTypes` (union of type _definitions_), `AnyNode` gives the union of runtime `Node` values — discriminated by `kind` for exhaustive `switch` narrowing. `SubsetNode` narrows the union to a specific set of kinds.
## 0.9.2
### Patch Changes
- [#27](https://github.com/nicia-ai/typegraph/pull/27) [`c2f0811`](https://github.com/nicia-ai/typegraph/commit/c2f0811863a61608c16901ce1fc61fdfbc26cb3f) Thanks [@pdlug](https://github.com/pdlug)! - Fix `count(alias, field)` and `countDistinct(alias, field)` ignoring the field argument in SQL compilation.
Both functions always compiled to `COUNT(alias_id)` / `COUNT(DISTINCT alias_id)` regardless of the field argument, because:
1. The aggregate emitters in `standard-builders.ts` and `set-operations.ts` hardcoded `_id` for count/countDistinct instead of calling `compileFieldValue()` like sum/avg/min/max do.
2. `collectRequiredColumnsByAlias` in `standard-pass-pipeline.ts` explicitly skipped marking the field as required for count/countDistinct, so the CTE wouldn't include the `_props` column even if the emitter were fixed.
Now `count("p", "email")` correctly compiles to `COUNT(json_extract(p_props, '$."email"'))` and `countDistinct("b", "genre")` compiles to `COUNT(DISTINCT json_extract(b_props, '$."genre"'))`.
## 0.9.1
### Patch Changes
- [#24](https://github.com/nicia-ai/typegraph/pull/24) [`733bf8a`](https://github.com/nicia-ai/typegraph/commit/733bf8abfd7b0fa9901a08ff67ce1c9343a2e961) Thanks [@pdlug](https://github.com/pdlug)! - Fix `checkUniqueBatch` exceeding SQL bind parameter limit on SQLite/D1/Durable Objects.
Bulk constraint operations (`bulkGetOrCreateByConstraint`, `bulkFindByConstraint`) passed all keys in a single `IN (...)` clause. With hundreds of unique keys, this exceeded SQLite's 999 bind parameter limit, causing `SQLITE_ERROR: too many SQL variables`.
The fix chunks the keys array in `checkUniqueBatch` using the same pattern already used by `getNodes`, `insertNodesBatch`, and other batch operations. SQLite chunks at 996 keys per query (999 max − 3 fixed params), PostgreSQL at 65,532.
## 0.9.0
### Minor Changes
- [#21](https://github.com/nicia-ai/typegraph/pull/21) [`88beee4`](https://github.com/nicia-ai/typegraph/commit/88beee42ce0ecfe2064b0b3889653e889b0c74aa) Thanks [@pdlug](https://github.com/pdlug)! - Add `transactionMode` to SQLite execution profile, fixing Cloudflare Durable Object compatibility.
`createSqliteBackend` previously used raw `BEGIN`/`COMMIT`/`ROLLBACK` SQL for all sync SQLite drivers. This crashes on Cloudflare Durable Object SQLite (via `drizzle-orm/durable-sqlite`) because the driver does not support raw transaction SQL through `db.run()`.
The new `transactionMode` option (`"sql"` | `"drizzle"` | `"none"`) controls how transactions are managed:
- `"sql"` — TypeGraph issues `BEGIN`/`COMMIT`/`ROLLBACK` directly (default for better-sqlite3, bun:sqlite)
- `"drizzle"` — delegates to Drizzle's `db.transaction()` (default for async drivers)
- `"none"` — transactions disabled (default for D1 and Durable Objects)
D1 and Durable Object sessions are auto-detected by Drizzle session name. Users can override via `executionProfile: { transactionMode: "..." }`.
**Breaking:** `isD1` removed from `SqliteExecutionProfileHints` and `SqliteExecutionProfile`. Use `transactionMode: "none"` instead. `D1_CAPABILITIES` removed — capabilities are now derived from `transactionMode`.
## 0.8.0
### Minor Changes
- [#19](https://github.com/nicia-ai/typegraph/pull/19) [`5b1dec6`](https://github.com/nicia-ai/typegraph/commit/5b1dec64f280a2ec638c69b6fa5a1bc08ba92e88) Thanks [@pdlug](https://github.com/pdlug)! - Support unconstrained edges in `defineGraph`.
Edges defined without `from`/`to` constraints (e.g., `defineEdge("sameAs")`) can now be passed directly to `defineGraph` without an `EdgeRegistration` wrapper. They are automatically allowed to connect any node type in the graph to any other.
- **`EdgeEntry` widened** — accepts any `EdgeType`, not just those with endpoints
- **`NormalizedEdges`** — falls back to all graph node types when `from`/`to` are undefined
- Constrained edges, `EdgeRegistration` wrappers, and narrowing validation are unchanged
## 0.7.0
### Minor Changes
- [#16](https://github.com/nicia-ai/typegraph/pull/16) [`0a2f08f`](https://github.com/nicia-ai/typegraph/commit/0a2f08fa7d755ee6adb59db4d34a26a3863c0c79) Thanks [@pdlug](https://github.com/pdlug)! - Tighten type safety across store and collection APIs.
**Breaking:** `TypedNodeRef` has been renamed to `NodeRef` and the old untyped `NodeRef` has been removed. Replace `TypedNodeRef` with `NodeRef` — the type is structurally identical. Unparameterized `NodeRef` (with the new default) covers the old untyped usage.
- **`EdgeId`** — branded edge ID type, mirroring `NodeId`. Prevents mixing IDs from different edge types at compile time.
- **`Edge`** — edge instances now carry endpoint node types. `edge.fromId` is `NodeId`, `edge.toId` is `NodeId`, and `edge.id` is `EdgeId`.
- **`getNodeKinds` / `getEdgeKinds`** — return `readonly (keyof G["nodes"] & string)[]` instead of `readonly string[]`.
- **`constraintName` literal unions** — `findByConstraint`, `getOrCreateByConstraint`, and their bulk variants now only accept constraint names that exist on the node registration, catching typos at compile time.
## 0.6.0
### Minor Changes
- [#14](https://github.com/nicia-ai/typegraph/pull/14) [`45624e0`](https://github.com/nicia-ai/typegraph/commit/45624e0ef5caf28c5a7bf8931f0ae96ce542c20d) Thanks [@pdlug](https://github.com/pdlug)! - Restructure SQLite/Postgres entry points to decouple DDL generation from native dependencies.
**Breaking changes:**
- `./drizzle`, `./drizzle/sqlite`, `./drizzle/postgres`, `./drizzle/schema/sqlite`, `./drizzle/schema/postgres` entry points are removed. Import backend factories, schema tables/factories, and DDL helpers from `./sqlite` and `./postgres`.
- `createLocalSqliteBackend` moves from `./sqlite` to `./sqlite/local`. The `./sqlite` entry point no longer depends on `better-sqlite3`.
- `getSqliteMigrationSQL` is renamed to `generateSqliteMigrationSQL`.
- `getPostgresMigrationSQL` is renamed to `generatePostgresMigrationSQL`.
- Individual table type aliases (`NodesTable`, `EdgesTable`, `UniquesTable`, `SchemaVersionsTable`, `EmbeddingsTable`) are removed from both schema modules. Use `SqliteTables["nodes"]` or `PostgresTables["edges"]` instead.
**Migration guide:**
| Before | After |
| ------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------- |
| `import { ... } from "@nicia-ai/typegraph/drizzle/sqlite"` | `import { ... } from "@nicia-ai/typegraph/sqlite"` |
| `import { ... } from "@nicia-ai/typegraph/drizzle/postgres"` | `import { ... } from "@nicia-ai/typegraph/postgres"` |
| `import { ... } from "@nicia-ai/typegraph/drizzle/schema/sqlite"` | `import { ... } from "@nicia-ai/typegraph/sqlite"` |
| `import { ... } from "@nicia-ai/typegraph/drizzle/schema/postgres"` | `import { ... } from "@nicia-ai/typegraph/postgres"` |
| `import { createLocalSqliteBackend } from "@nicia-ai/typegraph/sqlite"` | `import { createLocalSqliteBackend } from "@nicia-ai/typegraph/sqlite/local"` |
| `getSqliteMigrationSQL()` | `generateSqliteMigrationSQL()` |
| `getPostgresMigrationSQL()` | `generatePostgresMigrationSQL()` |
| `NodesTable`, `EdgesTable`, `UniquesTable`, `SchemaVersionsTable`, `EmbeddingsTable` | `SqliteTables["nodes"]` / `PostgresTables["nodes"]` (and corresponding table keys) |
## 0.5.0
### Minor Changes
- [#12](https://github.com/nicia-ai/typegraph/pull/12) [`c40b8a4`](https://github.com/nicia-ai/typegraph/commit/c40b8a4c99f5ccddaf1bceea8c927f6aeb0300f4) Thanks [@pdlug](https://github.com/pdlug)! - Add read-only lookup methods and store-level clear for graph data management.
**New APIs:**
- `findByConstraint` / `bulkFindByConstraint` — look up nodes by a named uniqueness constraint without creating. Returns `Node | undefined` (or `(Node | undefined)[]` for bulk). Soft-deleted nodes are excluded.
- `findByEndpoints` — look up an edge by `(from, to)` with optional `matchOn` property fields without creating. Returns `Edge | undefined`. Soft-deleted edges are excluded.
- `store.clear()` — hard-delete all data for the current graph (nodes, edges, uniques, embeddings, schema versions). Resets collection caches so the store is immediately reusable with raw, unversioned semantics; reopen it through a managed factory before relying on schema-version fencing.
## 0.4.0
### Minor Changes
- [#10](https://github.com/nicia-ai/typegraph/pull/10) [`550eec6`](https://github.com/nicia-ai/typegraph/commit/550eec6bbe34427be9095fe59571b55f75c68792) Thanks [@pdlug](https://github.com/pdlug)! - Add node and edge get-or-create operations with explicit API naming.
**New APIs:**
- `getOrCreateByConstraint` / `bulkGetOrCreateByConstraint` — deduplicate nodes by a named uniqueness constraint
- `getOrCreateByEndpoints` / `bulkGetOrCreateByEndpoints` — deduplicate edges by `(from, to)` with optional `matchOn` property fields
- `hardDelete` for node and edge collections
- `action: "created" | "found" | "updated" | "resurrected"` result discriminant
**Breaking changes:**
- `upsert` → `upsertById`, `bulkUpsert` → `bulkUpsertById`
- `onConflict: "skip" | "update"` → `ifExists: "return" | "update"`
- `ConstraintNotFoundError` → `NodeConstraintNotFoundError`
- Removed generic `FindOrCreate*` type exports in favor of explicit `NodeGetOrCreateByConstraint*` and `EdgeGetOrCreateByEndpoints*` types
## 0.3.1
### Patch Changes
- [#8](https://github.com/nicia-ai/typegraph/pull/8) [`4732792`](https://github.com/nicia-ai/typegraph/commit/4732792a9ff7ed665f55bb314029c06024f5b62e) Thanks [@pdlug](https://github.com/pdlug)! - Fix `AnyPgDatabase` type to accept standard Drizzle instances created without an explicit schema
## 0.3.0
### Minor Changes
- [#6](https://github.com/nicia-ai/typegraph/pull/6) [`4553aed`](https://github.com/nicia-ai/typegraph/commit/4553aedf3cd7390acb7509e1c321a42bed225f1e) Thanks [@pdlug](https://github.com/pdlug)! - Big performance increases, cleaner APIs, prepared queries, and batch collection
APIs.
### Breaking Changes
**Renamed APIs:**
- `selectAggregate()` is now `aggregate()`
- `EdgeTypeNames` / `NodeTypeNames` are now `EdgeKinds` / `NodeKinds` (including getter functions)
**Traversal expansion:** `includeImplyingEdges` replaced with `expand` option supporting four modes: `"none"`, `"implying"`, `"inverse"`, and `"all"` (default: `"inverse"`)
**Recursive traversal:** The chained methods `.maxHops()`, `.minHops()`, `.collectPath()`, and `.withDepth()` are consolidated into a single `recursive()` call with an options object:
```ts
// Before
.traverse("p", "knows", "friend").recursive().maxHops(5).collectPath()
// After
.traverse("p", "knows", "friend").recursive({ maxHops: 5, path: true })
```
New `cyclePolicy: "prevent" | "allow"` option (default: `"prevent"`). Unbounded recursion capped at depth 100; explicit `maxHops` validated up to 1,000.
**Store:** `Store` class is now a type-only export — use `createStore()`. `StoreConfig` replaced by `StoreOptions`.
**Moved to `@nicia-ai/typegraph/schema`:** All schema management APIs (`serializeSchema`, `deserializeSchema`, `initializeSchema`, `ensureSchema`, `migrateSchema`, `computeSchemaDiff`, `getMigrationActions`, `isBackwardsCompatible`, and related types) are now imported from the new `@nicia-ai/typegraph/schema` entry point.
**Removed from main entry:** `KindRegistry`, Result utilities (`ok`/`err`/`isOk`/`isErr`/`unwrap`/`unwrapOr`), date helpers (`encodeDate`/`decodeDate`), validation utilities, and compiler/profiler internals.
### New Features
**Prepared queries** — precompile queries once and execute repeatedly with different bindings at zero recompilation cost:
```ts
const prepared = store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.name.eq(param("name")))
.select((ctx) => ctx.p)
.prepare();
const alice = await prepared.execute({ name: "Alice" });
const bob = await prepared.execute({ name: "Bob" });
```
**Batch collection APIs:**
- `getByIds(ids)` — batched lookup preserving input order, returns `undefined` for missing IDs
- `bulkInsert` — void-returning fire-and-forget ingestion
- `bulkCreate` — multi-row `INSERT ... RETURNING` instead of per-item inserts
- `bulkUpsert` (edges) — batch lookup instead of N+1 sequential calls
**Node `find({ where })`** — filter nodes using the full query predicate system directly from collections.
### Performance
- SQL compiler restructured into plan/passes/emitter pipeline with predicate pre-indexing, column pruning, and single-hop recursive lowering
- Drizzle backend split into modular operations with dialect-driven strategy dispatch
- SQLite prepared statement caching with LRU eviction
- Compilation caching on immutable query builder instances
- Bind-limit-aware batch chunking (SQLite: 999 params, PostgreSQL: 65,535 params)
- Benchmark regression guardrails added to CI for both SQLite and PostgreSQL
## 0.2.0
### Minor Changes
- [`bdd5f34`](https://github.com/nicia-ai/typegraph/commit/bdd5f349453b19e9616f00d7591b436195feb925) Thanks [@pdlug](https://github.com/pdlug)! - Improve support for custom table names and use web crypto to support both node and edge runtimes.
## 0.1.1
### Patch Changes
- [`6f16bf9`](https://github.com/nicia-ai/typegraph/commit/6f16bf93ebd0811f386df63b80b8b80a3ee26c2f) Thanks [@pdlug](https://github.com/pdlug)! - Verify npmjs trusted publishing
## 0.1.0
### Minor Changes
- [`3d78324`](https://github.com/nicia-ai/typegraph/commit/3d78324472ac4cb4ac929b52c7501c08a5e7b6ca) Thanks [@pdlug](https://github.com/pdlug)! - Initial public release
# Data Sync Patterns
> Strategies for synchronizing external data with your TypeGraph store
When adding TypeGraph as a graph overlay to an existing application, you need
to keep your graph data in sync with your source of truth. This guide covers
practical patterns for syncing external data into TypeGraph using the bulk
operations API.
## Overview
Most applications adding TypeGraph will have existing data in relational tables,
external APIs, or document stores. Rather than migrating this data, TypeGraph
works alongside it as an overlay that provides:
- Graph traversals across your existing entities
- Semantic search via vector embeddings
- Relationship discovery and inference
The key challenge is keeping the graph in sync with your source data. We cover three approaches:
| Approach | Best For | Complexity |
|----------|----------|------------|
| [On-demand sync](#on-demand-sync) | Low-volume, real-time needs | Low |
| [Batch sync](#batch-sync) | Bulk imports, periodic refresh | Medium |
| [Event-driven sync](#event-driven-sync) | High-volume, near-real-time | Higher |
## Bulk Operations API
TypeGraph provides bulk operations for efficient sync workflows:
```typescript
// Create or update a single node
await store.nodes.Document.upsertById(id, props);
// Create many nodes at once
await store.nodes.Document.bulkCreate(items);
// Insert many nodes without returning results (dedicated fast path)
await store.nodes.Document.bulkInsert(items);
// Create or update many nodes at once
await store.nodes.Document.bulkUpsertById(items);
// Delete many nodes at once
await store.nodes.Document.bulkDelete(ids);
```
### upsertById
Creates a node if it doesn't exist, or updates it if it does. This includes
"un-deleting" soft-deleted nodes:
```typescript
// First call creates the node
const doc1 = await store.nodes.Document.upsertById("doc_123", {
title: "Original Title",
content: "...",
});
// Second call updates the existing node
const doc2 = await store.nodes.Document.upsertById("doc_123", {
title: "Updated Title",
content: "...",
});
// doc1.id === doc2.id - same node, updated in place
```
To reopen a previously ended fact without changing its identity, use the
explicit clear operation:
```typescript
await store.nodes.Document.upsertById("doc_123", currentProps, {
clearValidTo: true,
});
```
Omitting both end fields preserves the current end; `validTo` sets it;
`clearValidTo: true` removes it. With `coalesceUnchangedUpserts`, replaying a
clear against an already-open row skips the write, but capability validation
still runs first.
### bulkCreate
Efficiently creates multiple nodes in a single operation. Uses a single
multi-row INSERT with RETURNING when the backend supports it:
```typescript
const documents = await store.nodes.Document.bulkCreate([
{ props: { title: "Doc 1", content: "..." } },
{ props: { title: "Doc 2", content: "..." } },
{ props: { title: "Doc 3", content: "..." }, id: "custom_id" },
]);
```
If you only need write side effects and not created payloads, use `bulkInsert`.
### bulkInsert
Inserts multiple nodes without returning results. This is the dedicated fast path
for bulk ingestion — automatically wrapped in a transaction:
```typescript
await store.nodes.Document.bulkInsert([
{ props: { title: "Doc 1", content: "..." } },
{ props: { title: "Doc 2", content: "..." } },
{ props: { title: "Doc 3", content: "..." }, id: "custom_id" },
]);
```
Prefer `bulkInsert` over `bulkCreate` when you don't need results.
### bulkUpsertById
Creates or updates multiple nodes. Ideal for sync workflows where you don't
know which records already exist:
```typescript
// Sync a batch of external records
const externalRecords = await fetchExternalData();
const synced = await store.nodes.Document.bulkUpsertById(
externalRecords.map((record) => ({
id: record.id, // Use external ID as graph node ID
props: {
title: record.title,
content: record.body,
source: { table: "documents", id: record.id },
},
}))
);
```
A feed can deliver the same key twice in one page, so items are applied in
order: the first item for an id creates or updates the row, and every later copy
is an update over the value the earlier item wrote. The batch ends in the same
state the equivalent sequence of `upsertById` calls would leave, whether or not
the row existed before the batch. Because a later copy is an update, it merges
over the earlier value — a field the last copy omits keeps what an earlier copy
set. Edges follow the same rule, except that an update never repoints an edge:
the endpoints of the first write stand.
#### One batch cannot hand a unique value from one row to another
Item order settles which value each id ends up with, but the writes themselves
are grouped: every create in the batch runs before every update. A batch where
one item **releases** a `unique` constraint value and a later item **claims** it
therefore fails, even though the same operations applied one at a time succeed:
```typescript
// "alice@example.com" is currently held by person_a.
await store.nodes.Person.bulkUpsertById([
{ id: "person_a", props: { email: "moved@example.com" } }, // releases it
{ id: "person_b", props: { email: "alice@example.com" } }, // claims it
]);
// ✗ UniquenessError: person_b's create is checked while person_a still holds
// the value. On a backend with transactions nothing is written at all.
```
Edges have the same limitation for a `cardinality` slot: a batch that ends the
lone `oneActive` edge from a source while creating its replacement fails with a
`CardinalityError`, because the replacement's create is checked before the
update that frees the slot.
This is a stated property of the bulk APIs rather than a bug to work around
blindly. A batch describes the **set** of rows you want, not a script to reach
them by, and grouping the creates apart from the updates is what makes a batch a
couple of statements instead of one per item. The failure is loud and typed —
never a silently dropped write.
Two workarounds, both exact:
- **Split the handoff across two batches**: one carrying the items that release
the value, then one carrying the items that claim it. Reordering items
*within* a single batch does not help, since the grouping ignores that order.
This keeps bulk throughput for everything else.
- **Apply the conflicting items as sequential `upsertById` calls** (for edges,
as `update` then `create`), which frees the value before the claim is checked.
If a sync feed can legitimately swap unique values between records, the
two-batch shape is the reliable one: send one `bulkUpsertById` for the ids that
already exist, then a second for the new ones. The interchange importer
(`importGraph` with `onConflict: "update"`) is the other option that handles it
directly — it applies each row in document order regardless of batch size, at the
cost of the per-row validation and reporting that a bulk write skips.
### bulkReplaceById
Use node `bulkReplaceById` when each source record is authoritative and should
replace the stored document rather than merge with it:
```typescript
await store.nodes.Document.bulkReplaceById(
externalRecords.map((record) => ({
id: record.id,
props: {
title: record.title,
content: record.body,
},
}))
);
```
Omitted optional fields are removed. IDs must be distinct within the call;
replacement is a set of final documents, not an ordered patch stream. On an
eligible serverless backend this is the read-free sync path: TypeGraph submits
creation, replacement, resurrection, claims, and search sidecars as one atomic
program instead of reading every stored document before writing.
### bulkDelete
Deletes multiple nodes by ID. Silently ignores IDs that don't exist:
```typescript
// Remove nodes that no longer exist in source
const deletedIds = await findDeletedRecords();
await store.nodes.Document.bulkDelete(deletedIds);
```
### getOrCreate APIs
Use get-or-create methods when your dedupe key is not a direct ID:
```typescript
// Match by a named uniqueness constraint
const byEmail = await store.nodes.User.getOrCreateByConstraint(
"user_email",
{ email: "alice@example.com", name: "Alice" },
{ ifExists: "update" }
);
// byEmail.action: "created" | "found" | "updated" | "resurrected"
// Match edges by endpoints (+ optional matchOn fields)
const membership = await store.edges.memberOf.getOrCreateByEndpoints(
user,
org,
{ role: "admin", source: "sync" },
{
matchOn: ["role"],
ifExists: "update",
validFrom: sourceMembership.startedAt,
validTo: sourceMembership.endedAt,
onImmutableLowerBound: "preserve",
}
);
// membership.action: "created" | "found" | "updated" | "resurrected"
```
For endpoint writes, `validFrom` applies when a new edge is created and on the
`"resurrected"` branch, where it restates the revived row's whole window. With
`onImmutableLowerBound: "preserve"`, an `"updated"` live edge keeps its stored
lower bound while still applying props and `validTo`; the default `"refuse"`
policy instead refuses a different stated start. A `"found"` result performs no
write and preserves the existing validity window. To reopen an ended live edge,
pass `clearValidTo: true` together with `ifExists: "update"`; the default return
mode refuses that combination rather than ignoring the clear. An end that precedes the
row's effective start is refused — see
[Inverted validity windows](/errors/#inverted_validity_window).
### Edge Bulk Operations
Edges also support bulk operations:
```typescript
// Create many edges at once (returns created edges)
const edges = await store.edges.relatedTo.bulkCreate([
{ from: doc1, to: doc2, props: { confidence: 0.9 } },
{ from: doc1, to: doc3, props: { confidence: 0.7 } },
{ from: doc2, to: doc3, props: { confidence: 0.8 } },
]);
// Insert many edges without returning results (fast path)
await store.edges.relatedTo.bulkInsert([
{ from: doc1, to: doc2, props: { confidence: 0.9 } },
{ from: doc1, to: doc3, props: { confidence: 0.7 } },
{ from: doc2, to: doc3, props: { confidence: 0.8 } },
]);
// Delete many edges at once
await store.edges.relatedTo.bulkDelete(edgeIds);
```
## On-Demand Sync
Sync individual records when they're accessed or modified. Best for low-volume
scenarios where you want real-time consistency.
```typescript
import { type Store } from "@nicia-ai/typegraph";
import { db, documents } from "./drizzle-schema";
interface AppDocument {
id: string;
title: string;
content: string;
updatedAt: Date;
}
async function syncDocument(store: Store, doc: AppDocument) {
// Generate embedding for semantic search
const embedding = await generateEmbedding(doc.content);
// Upsert ensures we create or update as needed
return store.nodes.Document.upsertById(doc.id, {
title: doc.title,
content: doc.content,
embedding,
source: { table: "documents", id: doc.id },
});
}
// Sync on read - ensure graph is current before querying
async function getRelatedDocuments(documentId: string) {
// First, ensure the source document is synced
const appDoc = await db.select().from(documents).where(eq(documents.id, documentId)).get();
if (!appDoc) throw new Error("Document not found");
await syncDocument(store, appDoc);
// Now query the graph for relationships
return store
.query()
.from("Document", "d")
.whereNode("d", (d) => d.id.eq(documentId))
.traverse("relatedTo", "r")
.to("Document", "related")
.select((ctx) => ({
id: ctx.related.id,
title: ctx.related.title,
confidence: ctx.r.confidence,
}))
.execute();
}
// Sync on write - update graph when source changes
async function updateDocument(documentId: string, updates: Partial) {
// Update source of truth first
const [updated] = await db
.update(documents)
.set({ ...updates, updatedAt: new Date() })
.where(eq(documents.id, documentId))
.returning();
// Then sync to graph
await syncDocument(store, updated);
return updated;
}
```
## Batch Sync
Process records in batches for bulk imports or periodic refresh. Best for
large datasets or when you need to backfill data.
### Basic Batch Sync
```typescript
interface SyncOptions {
batchSize?: number;
onProgress?: (processed: number, total: number) => void;
}
async function syncAllDocuments(store: Store, options: SyncOptions = {}) {
const { batchSize = 100, onProgress } = options;
// Get total count for progress reporting
const [{ count }] = await db.select({ count: sql`count(*)` }).from(documents);
let processed = 0;
let offset = 0;
while (offset < count) {
// Fetch a batch from source
const batch = await db.select().from(documents).limit(batchSize).offset(offset);
if (batch.length === 0) break;
// Generate embeddings in parallel (respecting API rate limits)
const embeddings = await batchGenerateEmbeddings(batch.map((d) => d.content));
// Bulk upsert the batch
await store.nodes.Document.bulkUpsertById(
batch.map((doc, i) => ({
id: doc.id,
props: {
title: doc.title,
content: doc.content,
embedding: embeddings[i],
source: { table: "documents", id: doc.id },
},
}))
);
processed += batch.length;
offset += batchSize;
onProgress?.(processed, count);
}
return { processed, total: count };
}
// Usage
await syncAllDocuments(store, {
batchSize: 50,
onProgress: (processed, total) => {
console.log(`Synced ${processed}/${total} documents`);
},
});
```
### Incremental Sync
Only sync records that have changed since the last sync:
```typescript
interface SyncState {
lastSyncAt: Date;
}
async function incrementalSync(store: Store, state: SyncState): Promise {
const since = state.lastSyncAt;
const now = new Date();
// Fetch only changed records
const changed = await db
.select()
.from(documents)
.where(gt(documents.updatedAt, since))
.orderBy(documents.updatedAt);
if (changed.length > 0) {
const embeddings = await batchGenerateEmbeddings(changed.map((d) => d.content));
await store.nodes.Document.bulkUpsertById(
changed.map((doc, i) => ({
id: doc.id,
props: {
title: doc.title,
content: doc.content,
embedding: embeddings[i],
source: { table: "documents", id: doc.id },
},
}))
);
console.log(`Synced ${changed.length} changed documents`);
}
// Handle deletions (if your source tracks them)
const deleted = await db
.select({ id: documents.id })
.from(documents)
.where(and(gt(documents.deletedAt, since), isNotNull(documents.deletedAt)));
if (deleted.length > 0) {
await store.nodes.Document.bulkDelete(deleted.map((d) => d.id));
console.log(`Removed ${deleted.length} deleted documents`);
}
return { lastSyncAt: now };
}
```
### Scheduled Sync Job
Run incremental sync on a schedule:
```typescript
import { CronJob } from "cron";
// Store sync state (in production, persist this to a database)
let syncState: SyncState = { lastSyncAt: new Date(0) };
// Run every 5 minutes
const syncJob = new CronJob("*/5 * * * *", async () => {
try {
syncState = await incrementalSync(store, syncState);
} catch (error) {
console.error("Sync failed:", error);
// Alert, retry, etc.
}
});
syncJob.start();
```
## Event-Driven Sync
React to changes in your source data via events, webhooks, or database triggers.
Best for high-volume scenarios requiring near-real-time sync.
### Message Queue Pattern
```typescript
import { Queue, Worker } from "bullmq";
// Define sync job types
interface SyncJob {
type: "upsert" | "delete";
entityType: "Document" | "User";
entityId: string;
}
// Producer: Enqueue sync jobs when source data changes
const syncQueue = new Queue("sync");
async function onDocumentCreated(doc: AppDocument) {
await syncQueue.add("sync", {
type: "upsert",
entityType: "Document",
entityId: doc.id,
});
}
async function onDocumentUpdated(doc: AppDocument) {
await syncQueue.add("sync", {
type: "upsert",
entityType: "Document",
entityId: doc.id,
});
}
async function onDocumentDeleted(docId: string) {
await syncQueue.add("sync", {
type: "delete",
entityType: "Document",
entityId: docId,
});
}
// Consumer: Process sync jobs
const syncWorker = new Worker(
"sync",
async (job) => {
const { type, entityType, entityId } = job.data;
if (type === "delete") {
await store.nodes[entityType].delete(entityId);
return;
}
// Fetch current state from source
const record = await fetchEntity(entityType, entityId);
if (!record) {
// Record was deleted between enqueue and processing
await store.nodes[entityType].delete(entityId);
return;
}
// Generate embedding if needed
const embedding = await generateEmbedding(record.content);
// Upsert to graph
await store.nodes[entityType].upsertById(entityId, {
...record,
embedding,
source: { table: entityType.toLowerCase() + "s", id: entityId },
});
},
{
concurrency: 10,
connection: redis,
}
);
```
### Webhook Handler
Process webhooks from external systems:
```typescript
import { Hono } from "hono";
const app = new Hono();
app.post("/webhooks/documents", async (c) => {
const event = await c.req.json<{
type: "created" | "updated" | "deleted";
data: AppDocument;
}>();
switch (event.type) {
case "created":
case "updated": {
const embedding = await generateEmbedding(event.data.content);
await store.nodes.Document.upsertById(event.data.id, {
title: event.data.title,
content: event.data.content,
embedding,
source: { table: "documents", id: event.data.id },
});
break;
}
case "deleted": {
await store.nodes.Document.delete(event.data.id);
break;
}
}
return c.json({ ok: true });
});
```
### Database Triggers (PostgreSQL)
Use LISTEN/NOTIFY for real-time sync from PostgreSQL:
```typescript
import { Client } from "pg";
// Set up listener
const listener = new Client({ connectionString: process.env.DATABASE_URL });
await listener.connect();
await listener.query("LISTEN document_changes");
listener.on("notification", async (msg) => {
if (msg.channel !== "document_changes") return;
const payload = JSON.parse(msg.payload!);
const { operation, id } = payload;
if (operation === "DELETE") {
await store.nodes.Document.delete(id);
return;
}
// Fetch and sync the changed document
const doc = await db.select().from(documents).where(eq(documents.id, id)).get();
if (doc) {
const embedding = await generateEmbedding(doc.content);
await store.nodes.Document.upsertById(id, {
title: doc.title,
content: doc.content,
embedding,
source: { table: "documents", id },
});
}
});
```
Corresponding PostgreSQL trigger:
```sql
CREATE OR REPLACE FUNCTION notify_document_changes()
RETURNS TRIGGER AS $$
BEGIN
PERFORM pg_notify(
'document_changes',
json_build_object(
'operation', TG_OP,
'id', COALESCE(NEW.id, OLD.id)
)::text
);
RETURN COALESCE(NEW, OLD);
END;
$$ LANGUAGE plpgsql;
CREATE TRIGGER document_changes_trigger
AFTER INSERT OR UPDATE OR DELETE ON documents
FOR EACH ROW EXECUTE FUNCTION notify_document_changes();
```
## Syncing Relationships
When syncing data that includes relationships, sync nodes first, then edges:
```typescript
interface ExternalUser {
id: string;
name: string;
email: string;
managerId?: string;
}
async function syncUsers(users: ExternalUser[]) {
// Step 1: Sync all user nodes first
await store.nodes.User.bulkUpsertById(
users.map((u) => ({
id: u.id,
props: {
name: u.name,
email: u.email,
source: { table: "users", id: u.id },
},
}))
);
// Step 2: Sync manager relationships
// First, remove all existing manages edges (clean slate approach)
const existingEdges = await store.edges.manages.find();
if (existingEdges.length > 0) {
await store.edges.manages.bulkDelete(existingEdges.map((e) => e.id));
}
// Then create edges for users with managers
const usersWithManagers = users.filter((u) => u.managerId);
await store.edges.manages.bulkInsert(
usersWithManagers.map((u) => ({
from: { kind: "User" as const, id: u.managerId! },
to: { kind: "User" as const, id: u.id },
})),
);
}
```
## Handling Sync Failures
### Retry with Exponential Backoff
```typescript
async function syncWithRetry(
fn: () => Promise,
options: { maxRetries?: number; baseDelay?: number } = {}
): Promise {
const { maxRetries = 3, baseDelay = 1000 } = options;
let lastError: Error | undefined;
for (let attempt = 0; attempt <= maxRetries; attempt++) {
try {
return await fn();
} catch (error) {
lastError = error as Error;
if (attempt < maxRetries) {
const delay = baseDelay * Math.pow(2, attempt);
console.warn(`Sync attempt ${attempt + 1} failed, retrying in ${delay}ms`);
await new Promise((resolve) => setTimeout(resolve, delay));
}
}
}
throw lastError;
}
// Usage
await syncWithRetry(() => store.nodes.Document.bulkUpsertById(items));
```
### Dead Letter Queue
Track failed syncs for manual intervention:
```typescript
interface FailedSync {
entityType: string;
entityId: string;
error: string;
failedAt: Date;
attempts: number;
}
const failedSyncs: FailedSync[] = [];
async function syncWithDLQ(entityType: string, entityId: string, syncFn: () => Promise) {
try {
await syncWithRetry(syncFn);
} catch (error) {
failedSyncs.push({
entityType,
entityId,
error: (error as Error).message,
failedAt: new Date(),
attempts: 3,
});
console.error(`Sync failed after retries: ${entityType}:${entityId}`);
}
}
// Periodically retry or alert on failed syncs
async function processFailedSyncs() {
for (const failed of failedSyncs) {
console.log(`Failed sync: ${failed.entityType}:${failed.entityId} - ${failed.error}`);
// Retry, alert, or log for manual intervention
}
}
```
## Best Practices
### Use Consistent IDs
Map external IDs to graph node IDs consistently:
```typescript
// Good: Use external ID directly when it's unique and stable
await store.nodes.Document.upsertById(externalDoc.id, { ... });
// Good: Namespace if IDs might collide across sources
await store.nodes.Document.upsertById(`notion:${notionPage.id}`, { ... });
await store.nodes.Document.upsertById(`gdrive:${driveFile.id}`, { ... });
```
### Track Sync Metadata
Store sync information for debugging and auditing:
```typescript
const Document = defineNode("Document", {
schema: z.object({
title: z.string(),
content: z.string(),
embedding: embedding(1536).optional(),
source: externalRef("documents"),
// Sync metadata
lastSyncedAt: z.string().datetime().optional(),
syncVersion: z.number().optional(),
}),
});
await store.nodes.Document.upsertById(doc.id, {
...props,
lastSyncedAt: new Date().toISOString(),
syncVersion: (existingNode?.syncVersion ?? 0) + 1,
});
```
### Validate Before Sync
Validate external data before syncing to avoid corrupting your graph:
```typescript
const ExternalDocumentSchema = z.object({
id: z.string().min(1),
title: z.string().min(1),
content: z.string(),
});
async function syncDocument(rawDoc: unknown) {
const result = ExternalDocumentSchema.safeParse(rawDoc);
if (!result.success) {
console.error("Invalid document data:", result.error);
return;
}
await store.nodes.Document.upsertById(result.data.id, {
title: result.data.title,
content: result.data.content,
});
}
```
### Monitor Sync Health
Track sync metrics for observability:
```typescript
const syncMetrics = {
successful: 0,
failed: 0,
lastSyncDuration: 0,
lastSyncAt: null as Date | null,
};
async function monitoredSync(fn: () => Promise) {
const start = Date.now();
try {
await fn();
syncMetrics.successful++;
} catch (error) {
syncMetrics.failed++;
throw error;
} finally {
syncMetrics.lastSyncDuration = Date.now() - start;
syncMetrics.lastSyncAt = new Date();
}
}
// Expose metrics endpoint
app.get("/metrics/sync", (c) => c.json(syncMetrics));
```
## Next Steps
- [Integration Patterns](/integration) - Database setup and deployment patterns
- [Semantic Search](/semantic-search) - Add vector embeddings during sync
- [Query Builder](/queries/overview) - Query your synced graph data
# Fulltext Search
> BM25-style fulltext search with hybrid retrieval for RAG applications
TypeGraph supports fulltext search directly in your SQLite or PostgreSQL
database — no external search service required. Combine it with semantic
search to get **hybrid retrieval**: the gold-standard pattern for RAG
applications.
## Overview
Vector search is great at finding *semantically* similar content, but it
misses exact matches: proper nouns, SKU numbers, code identifiers, rare
technical terms. Fulltext search handles those. Running both and fusing
the results with Reciprocal Rank Fusion typically beats either approach
alone.
**Key capabilities:**
- Declare `searchable()` string fields in your Zod schema
- Native BM25 ranking (SQLite FTS5) and `ts_rank_cd` (PostgreSQL tsvector)
- Google-style query syntax: quoted phrases, `-excluded`, `OR`
- `n.$fulltext.matches()` predicate composes with metadata filters and graph traversal
- Hybrid search via `$fulltext.matches()` + `.similarTo()` in one query, fused with RRF
- Tunable RRF via `.fuseWith({ k, weights })` on the query builder, or `store.search.hybrid({ fusion })`
## Use Cases
### Hybrid RAG
Combine exact-match retrieval with semantic similarity:
```typescript
const hits = await store.search.hybrid("Document", {
limit: 10,
vector: {
fieldPath: "embedding",
queryEmbedding: await embed(question),
},
fulltext: { query: question },
});
const context = hits.map((h) => h.node.content).join("\n\n");
```
### Multi-tenant fulltext with metadata filters
The most important composition — `$fulltext.matches()` in the same query as
any other predicate:
```typescript
const results = await store
.query()
.from("Document", "d")
.whereNode("d", (d) =>
d.$fulltext.matches("climate change", 20).and(d.tenantId.eq(tenant.id))
)
.select((ctx) => ctx.d)
.execute();
```
### Authorised search via graph traversal
Only return documents the user is allowed to read:
```typescript
const results = await store
.query()
.from("User", "u")
.whereNode("u", (u) => u.id.eq(currentUserId))
.traverse("canRead", "e")
.to("Document", "d")
.whereNode("d", (d) => d.$fulltext.matches(userQuery, 10))
.select((ctx) => ctx.d)
.execute();
```
## Schema Design
### Declaring Searchable Fields
Use `searchable()` to mark string fields for fulltext indexing:
```typescript
import { defineNode, searchable } from "@nicia-ai/typegraph";
import { z } from "zod";
const Document = defineNode("Document", {
schema: z.object({
title: searchable({ language: "english" }),
body: searchable({ language: "english" }),
tenantId: z.string(),
published: z.boolean(),
}),
});
```
### How Indexing Works
TypeGraph stores one fulltext row per node. When you create or update a
node, the values of every `searchable()` field are concatenated and
indexed as a single document. This single-document-per-node design lets a
single query find matches that span multiple source fields — a title hit
plus a body hit both contribute to the same score.
- **PostgreSQL**: the `typegraph_node_fulltext` table carries a
`tsvector` column populated at INSERT time, with a GIN index.
- **SQLite**: the same shape is backed by an FTS5 virtual table with
BM25 ranking.
Sync is automatic — the fulltext index stays in sync with node data
through every `create`, `update`, `upsert`, and `delete` (soft and hard).
### Searchable Options
```typescript
searchable({
language: "english", // Postgres regconfig / SQLite FTS5 tokenizer
})
```
- **`language`**: Postgres uses this as the `regconfig` for stemming
(`english`, `spanish`, `french`, etc.). SQLite FTS5 tokenizer is fixed
at table creation time, so the language is stored but treated as
metadata.
### Adding `searchable()` to an Existing Graph
When you add `searchable()` to a field on a node kind that already has
rows in production, those pre-existing rows are not indexed until you
backfill the index:
```typescript
const stats = await store.search.rebuildFulltext();
console.log(
`Upserted ${stats.upserted}, cleared ${stats.cleared}, ` +
`skipped ${stats.skipped} across ${stats.kinds.length} kinds`,
);
if (stats.skippedIds && stats.skippedIds.length > 0) {
console.warn("Nodes with corrupt props were skipped:", stats.skippedIds);
}
// For systemic corruption, raise the cap to collect the full list:
const forensic = await store.search.rebuildFulltext(undefined, {
maxSkippedIds: 1_000_000,
});
```
`store.search.rebuildFulltext()` iterates nodes with keyset pagination on `id`
(stable under shared timestamps and light concurrent writes), transacts
per page, and cleans up stale fulltext rows for soft-deleted nodes.
Rebuild is a maintenance operation: concurrent hard-deletes between page
fetches can be missed by a single pass. Run during a maintenance window
for full consistency. Scope to a single kind with
`store.search.rebuildFulltext("Document")` to avoid scanning unrelated
data.
Also useful for:
- Recovering after a `DROP TABLE` / `TRUNCATE` of the fulltext table.
- Re-tokenizing after changing `language` on a `searchable()` field.
- Recovering from bulk inserts that bypassed the store layer.
### Checking Whether Search Is Ready
`store.search.rebuildFulltext()` fixes *content*. It cannot fix storage
that is missing, unattested, or provisioned at the wrong shape — and it
throws `StoreNotInitializedError` when it is, because the hot-path gate
refuses fulltext writes until both the deployment marker attests the shared
table and the graph-local activation marker admits this graph. To find out
which situation you are in without writing anything:
```typescript
const health = await store.probeContributions();
const fulltext = health.entries.find(
(entry) => entry.contribution === "fulltext",
);
if (fulltext?.state !== "ready") {
console.warn(`fulltext search is ${fulltext?.state}`, fulltext?.detail);
}
```
The probe writes nothing, so it is safe to call from a health check, on a
read path, or on a replica — which is the point: the alternative was to
run `store.repairContributions()`, a write with repair side effects, and
hope. When it reports `degraded`, escalate through the contribution health
ladder — probe, then `repairContributions()`, then
`rebuildContribution("fulltext")`, which is scoped to the calling graph: it
deletes and refills only that graph's rows in the shared fulltext table, and
drops and recreates the table itself only when no other graph has rows in
it. The three rungs, what each one writes, and when to stop are in
[Contribution health: probe, repair, rebuild](/troubleshooting#contribution-health-probe-repair-rebuild).
## Database Setup
### Initialization is required (boot via `createStoreWithSchema`)
Fulltext storage is **durably materialized once, at application boot**,
by `createStoreWithSchema`:
```typescript
// Run this once at startup — outside request handlers and transactions.
const [store] = await createStoreWithSchema(graph, backend);
```
`createStore(graph, backend)` is a synchronous, zero-I/O *attach*: it
does not create tables, repair DDL, or record that fulltext storage is
materialized. A fulltext read or write — a `searchable()` field write,
`store.search.fulltext()`, `store.search.hybrid()`,
`n.$fulltext.matches()`, `store.search.rebuildFulltext()`, or a
transaction that touches fulltext — against a database that was never
initialized throws `StoreNotInitializedError`. Use `createStore()` only
to attach to a database a prior `createStoreWithSchema` boot already
initialized. Graphs with no `searchable()` fields are unaffected.
### PostgreSQL
No extensions required. The built-in `tsvector` type and GIN indexes
work on every managed Postgres (RDS, Supabase, Neon, Cloud SQL, Aiven).
The fulltext table's DDL ships in `bootstrapTables()` and the migration
SQL. The first privileged `createStoreWithSchema` boot records one
deployment-scoped physical marker for that shared table plus a graph-local
activation marker. Later graphs reuse the physical attestation and need only
the DML activation write; fulltext operations require both markers (see above):
```typescript
import { generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
// Includes the fulltext table, tsvector column, GIN index, and pgvector
const migrationSQL = generatePostgresMigrationSQL();
```
### SQLite
No extensions required. FTS5 is compiled into the standard SQLite
distribution shipped with `better-sqlite3`, `libsql`, `bun:sqlite`, and
most other drivers. The FTS5 virtual table uses the
`porter unicode61 remove_diacritics 2` tokenizer.
## Querying
### `n.$fulltext.matches()` — The Query Predicate
`n.$fulltext.matches(query, k?, options?)` is a node-level fulltext
predicate. It's exposed on every `NodeAccessor`; at runtime it throws a
clear `UnsupportedPredicateError` if the node kind has no `searchable()`
fields, with a suggestion for how to fix the schema.
> **Visible in types, guarded at runtime.** `$fulltext` is present on every
> `NodeAccessor` at the TypeScript level for simplicity — a type-level
> brand would not survive modifiers like `.min(1).optional()`. The runtime
> check is the single source of truth: adding a `searchable()` field is
> what makes `.matches()` actually work. A query that type-checks can still
> throw `UnsupportedPredicateError` the first time it runs if no field
> was declared searchable.
```typescript
store
.query()
.from("Document", "d")
.whereNode("d", (d) => d.$fulltext.matches("climate change"))
.select((ctx) => ctx.d)
.execute();
```
It compiles to a JOIN against the fulltext index, adds an ORDER BY on
relevance rank, and applies the top-k limit — all in a single SQL
statement that composes with every other query-builder feature.
**`k` vs `limit`**: `k` (the second positional arg) is the top-k cap
applied **inside the fulltext CTE** — how many candidates to pull from
the index before outer filtering and fusion. It defaults to `50`, which
is fine for single-predicate use. `.limit()` on the query controls the
**final result count**. When feeding into RRF (`store.search.hybrid` or
`.fuseWith()`), pass a larger `k` per predicate (e.g. 200) so there are
enough candidates for the fused ranking to be meaningful.
Traversal happens after candidate generation. A candidate can therefore fan
out into several match rows. Use query-level `.where((ctx) => ...)` to filter
those completed rows, then apply an explicit `.orderBy()` and `.limit()` for
the final result. The final order does not change which nodes entered the
top-k candidate set, and the final limit counts match rows rather than distinct
source nodes.
### Query Modes
```typescript
d.$fulltext.matches("climate -warming", 10, { mode: "websearch" })
// Google-style: quoted phrases, -excluded terms, OR operator
d.$fulltext.matches("climate change", 10, { mode: "phrase" })
// Exact phrase match
d.$fulltext.matches("climate change", 10, { mode: "plain" })
// All terms must appear (implicit AND), no special syntax
d.$fulltext.matches("climate & !warming", 10, { mode: "raw" })
// Dialect-native syntax (tsquery on Postgres, FTS5 on SQLite)
```
**When to use each:**
| Mode | Best For | Example |
|------|----------|---------|
| `websearch` (default) | User-facing search boxes | `"climate change" -hoax OR warming` |
| `phrase` | Proper nouns, exact quotes | `"New York Times"` |
| `plain` | Programmatic queries | `climate change` |
| `raw` | Advanced users who know the dialect syntax | `climate<->change` |
### Composing with Filters
Fulltext is just another predicate — combine with `.and()`:
```typescript
store
.query()
.from("Document", "d")
.whereNode("d", (d) =>
d.$fulltext
.matches("machine learning", 20, { mode: "websearch" })
.and(d.published.eq(true))
.and(d.publishedAt.gte("2024-01-01"))
.and(d.tenantId.eq(tenant))
)
.select((ctx) => ctx.d)
.execute();
```
### Composing with Graph Traversal
`$fulltext.matches()` works inside any traversal:
```typescript
// Find documents matching "climate" that were written by someone I follow
const results = await store
.query()
.from("Person", "me")
.whereNode("me", (p) => p.id.eq(currentUserId))
.traverse("follows", "f")
.to("Person", "author")
.traverse("authored", "a", { direction: "in" })
.to("Document", "d")
.whereNode("d", (d) => d.$fulltext.matches("climate", 10))
.select((ctx) => ({
title: ctx.d.title,
author: ctx.author.name,
}))
.execute();
```
### Hybrid Search (Query Builder)
Use `$fulltext.matches()` and `.similarTo()` in the same query and
TypeGraph automatically fuses the two signals with Reciprocal Rank Fusion
at the SQL layer:
```typescript
const results = await store
.query()
.from("Document", "d")
.whereNode("d", (d) =>
d.$fulltext
.matches("renewable energy", 50)
.and(d.embedding.similarTo(queryVector, 50))
.and(d.tenantId.eq(tenant))
)
.select((ctx) => ctx.d)
.limit(10)
.execute();
```
The compiled SQL builds two CTEs (one for the vector side, one for the
fulltext side), orders each by relevance, and the outer query sorts by
`1/(60 + rank_vector) + 1/(60 + rank_fulltext)`. One round-trip, fully
composable with any other predicate.
### Tuning RRF
Defaults (k=60, equal weights) suit most workloads. Bias toward fulltext
for exact-match queries, toward vectors for conceptual queries:
```typescript
store
.query()
.from("Document", "d")
.whereNode("d", (d) =>
d.$fulltext
.matches("renewable energy", 50)
.and(d.embedding.similarTo(queryVector, 50))
.and(d.tenantId.eq(tenant))
)
.fuseWith({ k: 60, weights: { vector: 1.0, fulltext: 1.5 } })
.limit(10)
.execute();
```
`.fuseWith()` throws at compile time if the query lacks either a
`.similarTo()` or a `$fulltext.matches()`. Validation rejects non-finite
or negative `k`/weights. The same shape is accepted by
`store.search.hybrid({ fusion })` and validated by the same function on
both paths.
### Hybrid Search (Store API)
For tunable RRF parameters, use `store.search.hybrid()`. On the built-in
backends this runs as a **single SQL statement** — both sources, RRF
fusion, and node hydration composed together — so a hybrid query costs one
round trip instead of three. The saving scales with per-statement cost:
decisive on serverless HTTP drivers, Cloudflare D1 / Durable Objects, and
remote databases; on a local low-latency connection the two paths are
within a few milliseconds of each other. (Kind expansions via
`includeSubClasses`, and custom backends without the composed statement,
transparently use a multi-statement path with identical results.)
```typescript
const results = await store.search.hybrid("Document", {
limit: 10,
vector: {
fieldPath: "embedding",
queryEmbedding: await embed(question),
metric: "cosine",
k: 50, // Candidates to retrieve from vector
},
fulltext: {
query: question,
k: 50, // Candidates to retrieve from fulltext
includeSnippets: true,
},
fusion: {
method: "rrf",
k: 60, // RRF constant (classic default)
weights: {
vector: 1.0,
fulltext: 1.5, // Weight fulltext higher for exact-match workloads
},
},
});
```
Each hit carries sub-scores from both halves so you can debug ranking:
```typescript
for (const hit of results) {
console.log(hit.node.title, hit.score);
console.log(" vector rank:", hit.vector?.rank);
console.log(" fulltext rank:", hit.fulltext?.rank);
console.log(" snippet:", hit.fulltext?.snippet);
}
```
### Fulltext-Only Store API
For quick fulltext lookups that don't need the query builder:
```typescript
const hits = await store.search.fulltext("Document", {
query: "quarterly earnings",
limit: 10,
mode: "websearch",
includeSnippets: true,
});
for (const hit of hits) {
console.log(hit.node.title, hit.score, hit.snippet);
}
```
#### Options reference
| Option | Type | Default | Description |
|--------|------|---------|-------------|
| `query` | `string` | — *(required)* | User-supplied query string. Parsed according to `mode`. |
| `limit` | `number` | — *(required)* | Max rows to return. Must be a positive integer. |
| `mode` | `"websearch" \| "phrase" \| "plain" \| "raw"` | `"websearch"` | Parser for `query`. See [Query Modes](#query-modes). |
| `language` | `string` | the kind's declared language | Override the stemming/tokenization language for this query. By default the query is parsed with the language the kind's `searchable()` fields declare — a plan-time constant, which is what lets PostgreSQL serve the match from the `tsv` GIN index (a per-row language reference makes the tsquery non-constant and forces a scan). Postgres only — SQLite FTS5's tokenizer is fixed at table-create time and a per-query override throws. |
| `minScore` | `number` | *(none)* | Drop hits whose backend-native score is below this threshold. Score units depend on the strategy. |
| `includeSnippets` | `boolean` | `false` | Return a highlighted `…` snippet per hit. Noticeably slower than plain search — request only for final-page results. |
| `where` | `(accessor) => Predicate` | *(none)* | Property predicate compiled into the search statement's candidate set — the engine ranks only matching rows, so a filter never shrinks results below `limit` when enough matches exist (libSQL DiskANN: bounded by its 4× over-fetch). Same accessor and semantics as `store.nodes..find({ where })`. |
| `offset` | `number` | `0` | Rank-relative pagination: skip the first `offset` ranked hits. |
| `includeSubClasses` | `boolean` | `false` | Also search `subClassOf` descendant kinds and merge their scores into one ranking. |
The same three options are available on `store.search.vector` and
`store.search.hybrid` (where `where` and `includeSubClasses` apply to both
halves). Search always follows current-read semantics: tombstoned nodes and
nodes outside their validity window never rank.
Returned hits are `FulltextSearchHit>` with `node`, `score`
(higher = more relevant), `rank` (1-based), and `snippet` (when
requested).
## Reciprocal Rank Fusion
RRF is the de facto standard for combining ranked lists from multiple
retrievers. The formula:
```text
score(doc) = Σ weight_source / (k + rank_source)
```
Where `k` is the RRF constant (classic default: 60), `rank_source` is
the document's 1-based ordinal rank in each source, and `weight_source`
lets you bias toward one retriever.
**Why it works:** RRF is rank-based, not score-based. It doesn't care
that vector distances are in `[0, 2]` while BM25 scores are
unbounded — it only cares about ordinal position. That makes it robust
to heterogeneous score distributions across retrievers.
**Tuning tips:**
- Over-fetch from each side (default: 4× the requested limit). More
candidates per source = better recall.
- Bump `weights.fulltext` higher when exact matches matter (names, IDs,
proper nouns). Bump `weights.vector` for conceptual queries.
- Leave `k = 60` alone unless benchmarks show otherwise.
## Best Practices
### All Searchable Fields Share One Index
TypeGraph indexes all `searchable()` fields on a node as one document
(see [How Indexing Works](#how-indexing-works)). There's a single
`n.$fulltext` accessor per node — `searchable()` declarations on
individual fields are what bring it into existence and what determine
which text gets indexed.
### Use `includeSnippets` Sparingly
Highlighting (`ts_headline` on Postgres, `snippet()` on SQLite) is
noticeably slower than plain search. Request it only for final-page
results, not for large over-fetch pools.
### Pair with a Reranker for Top Quality
RRF is a strong baseline, but production RAG systems typically add a
cross-encoder reranker (Cohere Rerank, `bge-reranker`, etc.) as a final
stage. TypeGraph gives you the candidate set — the reranker picks the
winning order:
```typescript
const candidates = await store.search.hybrid("Document", {
limit: 50, // Over-fetch for reranker
vector: { fieldPath: "embedding", queryEmbedding },
fulltext: { query },
});
const reranked = await cohere.rerank({
query,
documents: candidates.map((c) => c.node.content),
top_n: 10,
});
```
### Filter Before You Fuse
Applying predicates via `.and()` shrinks the candidate pool before the
fusion ORDER BY, which improves both latency and ranking quality — there
are fewer irrelevant candidates competing for top positions:
```typescript
// Fast: tenant filter applied inside each CTE
.whereNode("d", (d) =>
d.$fulltext.matches(query, 50)
.and(d.embedding.similarTo(queryVec, 50))
.and(d.tenantId.eq(tenant))
)
// Slow: tenant filter applied AFTER fusion
.whereNode("d", (d) =>
d.$fulltext.matches(query, 5000)
.and(d.embedding.similarTo(queryVec, 5000))
)
// ...then filter results in JS
```
## Limitations
### One Fulltext Predicate Per Query
A single query can contain at most one `$fulltext.matches()` predicate.
This mirrors the constraint on `.similarTo()` and keeps the RRF fusion
model well-defined. If you need to search multiple terms, combine them
into one query string using websearch mode:
```typescript
// Good
d.$fulltext.matches("climate change OR global warming", 20)
// Rejected (at query-build time, not by the type checker)
d.$fulltext.matches("climate", 10).and(d.$fulltext.matches("warming", 10))
```
This invariant is enforced when the query is compiled
(`UnsupportedPredicateError`), not by TypeScript — so a surprising
second `.matches()` call surfaces as a runtime error the first time
the query runs.
### No `.matches()` Under OR or NOT
Fulltext predicates must appear at top level or inside AND groups. They
rewrite query structure (adding a CTE and ORDER BY) in a way that isn't
compatible with disjunction or negation semantics.
### Tokenizer Is Fixed on SQLite
FTS5 tokenizer options are set at CREATE VIRTUAL TABLE time. TypeGraph
ships with `porter unicode61 remove_diacritics 2` — a solid default for
English and accented Latin-script languages. For CJK or other tokenizers,
create the fulltext table manually with your preferred options.
### No Per-Field Weighting
All searchable fields on a node contribute equally to the combined
document. Postgres `setweight()`-style per-field bias is a planned
extension; today, structure your fields to put the most important text
first or split high-weight content into a dedicated kind. This
limitation applies even when you [swap in a custom
`FulltextStrategy`](#custom-fulltext-strategies) — TypeGraph
concatenates searchable fields into one `content` string before handing
it to the strategy.
## Custom Fulltext Strategies
`createPostgresBackend(db, { fulltext })` and
`createSqliteBackend(db, { fulltext })` accept a `FulltextStrategy`
that owns the **entire** fulltext pipeline — DDL, INSERT/UPSERT
(single + batch), DELETE (single + batch), MATCH condition, rank
expression, and snippet generation. The same strategy flows through
`store.search.fulltext()`, `store.search.hybrid()`,
`$fulltext.matches()` in the query builder,
`store.search.rebuildFulltext()`, and `bootstrapTables()` DDL.
Use this when the built-in `tsvector` isn't the right fit — for
example, BM25 inside Postgres (ParadeDB / `pg_search`), trigram
similarity (`pg_trgm`), or fulltext optimized for CJK languages
(`pgroonga`).
Most SQLite users should leave the default `fts5Strategy` in place.
### Constraints on alternate strategies
Before implementing a strategy, know what the abstraction does **not**
let you change today:
- **Side table is mandatory.** Every strategy writes one row per
`(graph_id, node_kind, node_id)` to a dedicated fulltext table. A
strategy cannot skip the side table and index a column on the main
nodes table directly (e.g. a GIN trigram index on
`typegraph_nodes.props`). Strategies *can* choose the column layout,
index type, and any computed projection inside that side table.
- **Content is pre-concatenated.** TypeGraph joins every
`searchable()` field value with `\n` before the strategy sees it —
`UpsertFulltextParams.content` is a single string. Per-field
indexing (`setweight`, per-column BM25 boosts, pgroonga
per-column weights) is not plumbed through today; a richer
per-field payload is planned but not yet part of the public
strategy contract.
- **One language per row.** When a node has multiple `searchable()`
fields with different `language` values, the first field's
language wins and is recorded on the row. TypeGraph emits a
one-time warning per conflicting schema; true multilingual
indexing needs a dedicated node kind per language.
### Strategy skeleton
The `FulltextStrategy` contract and every type referenced by it are exported
from the Drizzle-free backend-authoring entrypoint.
Fields below are the minimum surface; see
`src/query/dialect/fulltext-strategy.ts` in the TypeGraph source for
`tsvectorStrategy` and `fts5Strategy` as full references.
```typescript
import {
sql,
type FulltextStrategy,
type SqlFragment,
} from "@nicia-ai/typegraph/backend";
/**
* Example: a trigram-based strategy on top of pg_trgm. Illustrative —
* not production code. pg_trgm supports plain-term matching only, so
* `supportedModes` advertises `"plain"` and rejects everything else at
* compile time.
*/
export const pgTrgmStrategy: FulltextStrategy = {
name: "pg_trgm",
supportedModes: ["plain"],
supportsSnippets: false, // no native highlight; emit NULL snippet
supportsPrefix: false, // trigram similarity, not prefix
supportsLanguageOverride: false,
languages: ["simple"],
matchCondition(table, query) {
return sql`${sql.identifier(table)}."content" % ${query}`;
},
rankExpression(table, query) {
return sql`similarity(${sql.identifier(table)}."content", ${query})`;
},
snippetExpression() {
// `supportsSnippets: false` — callers get NULL and skip the field.
return sql`NULL`;
},
// Declares the table(s) this strategy owns as authoritative
// TableContributions. pg_trgm brings its own table (not the typed
// Drizzle `tables.fulltext`), so it is emitted verbatim from
// `createDdl` and is invisible to drizzle-kit unless you export your
// own table object. `runtimeEnsure: true` because no
// drizzle-kit-managed setup can create it.
ownedTables(primaryTableName) {
return [
{
logicalName: "fulltext",
owner: "pg_trgm",
tableName: primaryTableName,
createDdl: [
`CREATE EXTENSION IF NOT EXISTS pg_trgm;`,
`CREATE TABLE IF NOT EXISTS "${primaryTableName}" (
"graph_id" TEXT NOT NULL,
"node_kind" TEXT NOT NULL,
"node_id" TEXT NOT NULL,
"content" TEXT NOT NULL,
"language" TEXT NOT NULL,
"updated_at" TIMESTAMPTZ NOT NULL,
PRIMARY KEY ("graph_id", "node_kind", "node_id")
);`,
`CREATE INDEX IF NOT EXISTS "${primaryTableName}_trgm_idx"
ON "${primaryTableName}" USING GIN ("content" gin_trgm_ops);`,
],
runtimeEnsure: true,
},
];
},
buildUpsert(table, params, timestamp) {
return [
sql`
INSERT INTO ${sql.identifier(table)}
("graph_id", "node_kind", "node_id", "content", "language", "updated_at")
VALUES (${params.graphId}, ${params.nodeKind}, ${params.nodeId},
${params.content}, ${params.language}, ${timestamp})
ON CONFLICT ("graph_id", "node_kind", "node_id")
DO UPDATE SET
"content" = EXCLUDED."content",
"language" = EXCLUDED."language",
"updated_at" = EXCLUDED."updated_at"
`,
];
},
buildBatchUpsert(table, params, timestamp) {
if (params.rows.length === 0) return [];
// Dedup last-write-wins by nodeId, then emit a single multi-VALUES INSERT.
// Postgres ON CONFLICT rejects repeated conflict keys inside one statement.
// (The shipped helpers in fulltext-strategy.ts show this pattern.)
return [/* … */];
},
buildDelete(table, params) {
return [
sql`
DELETE FROM ${sql.identifier(table)}
WHERE "graph_id" = ${params.graphId}
AND "node_kind" = ${params.nodeKind}
AND "node_id" = ${params.nodeId}
`,
];
},
buildBatchDelete(table, params) {
if (params.nodeIds.length === 0) return [];
return [/* DELETE … WHERE node_id IN (…) */];
},
};
```
Wire it in at backend construction:
```typescript
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
const backend = createPostgresBackend(db, { fulltext: pgTrgmStrategy });
```
Capabilities (`phraseQueries`, `prefixQueries`, `highlighting`,
`languages`) are derived automatically from the strategy, so
`store.search.fulltext({ mode: "websearch" })` now throws
`ConfigurationError` before any SQL is generated — the strategy's
`supportedModes` is the source of truth.
## Troubleshooting
### `StoreNotInitializedError: fulltext storage … is not initialized`
The database was never booted through `createStoreWithSchema`, so the
deployment-scoped physical marker or this graph's activation marker is missing
(or it is `stale` — the strategy/DDL changed since it was recorded, or
`failed` — the last boot-time attempt errored). Bare `createStore()`
deliberately does **not** self-heal this on the hot path.
Fix: call `createStoreWithSchema(graph, backend)` once at application
startup — outside request handlers and adopted transactions — before any
fulltext operation. If you previously relied on fulltext tables being
created lazily on first write via `createStore()`, that path was removed;
move the initialization to an explicit boot step. A `stale` reason means
the recorded shape no longer matches the active strategy/DDL: migrate or
drop the fulltext table and re-run the boot, or restore the original
strategy.
`ContributionUnavailableError` with `state: "physical-storage-missing"` is
different: the physical fulltext table disappeared after initialization. Gated
fulltext searches and searchable node writes raise this typed error; compiled
query-builder predicates can still surface the engine's missing-relation error.
Run `store.rebuildContribution("fulltext")`; rerunning ordinary initialization
cannot reconstruct the missing indexed content. Use `probeContributions()` at
startup when an application must detect this out-of-band catalog damage before
the first dependent read or write.
### `Cannot call .$fulltext.matches() on alias "x"`
`$fulltext` is exposed on every node accessor at the type level, but
calling `.matches()` requires the node kind to have at least one
`searchable()` field — otherwise there's no indexed content to search.
The runtime guard throws a clear error pointing at the alias:
```text
Cannot call .$fulltext.matches() on alias "d" — its node kind has no
fields declared with searchable(). Add at least one:
`title: searchable({ language: "english" })`.
```
Fix by adding a searchable field to the schema:
```typescript
// Before:
title: z.string(),
// After:
title: searchable({ language: "english" }),
```
Refinements like `.min(1)` and `.trim()` are preserved — you can write
`searchable({ language: "english" }).min(1)` and the field is still
indexed.
### Empty fulltext results after bulk insert
TypeGraph syncs the fulltext index inline with each node write. If you
bulk-inserted via raw SQL that bypassed the store layer, the fulltext
table won't have entries. Re-run the inserts through
`store.nodes.X.create()` / `.bulkCreate()`, run
`store.search.rebuildFulltext()` to populate the index from existing rows,
or issue `backend.upsertFulltext()` / `backend.upsertFulltextBatch()`
calls directly.
### After adding `searchable()` to existing data
See [Adding `searchable()` to an Existing Graph](#adding-searchable-to-an-existing-graph)
above for the rebuild recipe and the caveats that apply to concurrent
workloads.
### `"Fulltext match predicates cannot be nested under OR or NOT"`
See [No `.matches()` Under OR or NOT](#no-matches-under-or-or-not)
above. Move the `$fulltext.matches()` to the top level or inside an
`.and()`.
### Hybrid results miss obvious matches
Increase the per-source `k` (over-fetch). The default is 4× the final
`limit`, which is tuned for small result pages. Large corpora benefit
from `k: 200` or higher on each side.
### Postgres: `text search configuration "xyz" does not exist`
The `language` you passed to `searchable({ language })` must be an
installed `regconfig` on your Postgres server. Every stock install ships
`simple`, `english`, `french`, `german`, `italian`, `portuguese`,
`russian`, `spanish`, and `swedish`; anything else requires an extension
(`zhparser` for Chinese, `pg_trgm` for trigram-based matching, or a
custom dictionary).
TypeGraph emits a `console.warn` at query time when you pass a language
outside the backend-advertised list, but a typo or missing extension
only fails when Postgres tries to build the `tsvector`. To diagnose:
```sql
SELECT cfgname FROM pg_ts_config;
```
Pick a name from that list, or install the extension that provides the
one you want. If you're running a managed Postgres (RDS, Supabase, Neon,
Cloud SQL, Aiven), check the provider's docs for which language extensions
are enabled — some require a restart or explicit enabling.
### Swapping to a custom fulltext strategy
See [Custom Fulltext Strategies](#custom-fulltext-strategies) for the
full interface, constraints, and a skeleton implementation.
## API Reference
- **Schema**: [`searchable()`](/queries/predicates#searchable)
- **Predicate**: [`n.$fulltext.matches()`](/queries/predicates#searchable)
- **Tunable fusion**: `QueryBuilder.fuseWith({ k, weights })`
- **Rebuild**: `store.search.rebuildFulltext(nodeKind?, { pageSize? })`
- **Store API**: `store.search.fulltext()` and `store.search.hybrid()` —
see the [Schemas & Stores reference](/schemas-stores).
See also:
- [Semantic Search](/semantic-search) — vector embeddings and
`.similarTo()`
- [Predicates reference](/queries/predicates) — complete predicate
catalog
- [Knowledge Graph for RAG](/examples/knowledge-graph-rag) — end-to-end
example combining fulltext, vector, and graph traversal
# Graph Algorithms
> Traversal, connectivity, label propagation, and PageRank on store.algorithms
Graph queries like "are Alice and Bob connected?" or "who is within two hops
of this node?" are common enough that writing them as recursive CTEs by hand
gets repetitive. TypeGraph exposes the high-utility algorithms as a small
facade on the store:
```typescript
store.algorithms.shortestPath(alice, bob, { edges: ["knows"] });
store.algorithms.weightedShortestPath(alice, bob, {
edges: ["knows"],
weightProperty: "strength",
});
store.algorithms.reachable(alice, { edges: ["knows"] });
store.algorithms.canReach(alice, bob, { edges: ["knows"] });
store.algorithms.neighbors(alice, { edges: ["knows"], depth: 2 });
store.algorithms.degree(alice, { edges: ["knows"] });
store.algorithms.weaklyConnectedComponents({
edges: ["knows"],
nodeKinds: ["Person"],
});
store.algorithms.labelPropagation({
edges: ["knows"],
nodeKinds: ["Person"],
});
store.algorithms.pageRank({ edges: ["knows"], nodeKinds: ["Person"] });
store.algorithms.personalizedPageRank({
edges: ["knows"],
seeds: [{ id: alice.id, kind: "Person" }],
});
```
Traversal calls use a set-based breadth-first frontier. Each level expands the
current `(id, kind)` set with a bind-limit-aware SQL query. Transactional
backends keep the visited set in a temporary working table; backends without a
pinned transaction use a chunked inline working relation. Each node is admitted
only at its minimum depth. Reachability rounds deduplicate edge targets before
checking target-node visibility and do not compute unused predecessors.
`shortestPath` and `canReach` search from both endpoints, retain predecessors for
path reconstruction, and stop when the frontiers meet.
`degree` remains a single `COUNT` query. Exact algorithms behave identically on
SQLite and PostgreSQL; PageRank follows the numerical-tolerance contract below.
The inline fallback is intentionally limited to bounded traversals. Whole-graph
iterative algorithms such as WCC, label propagation, and PageRank require a pinned transactional connection and
fail through their typed capability gate when one is unavailable; they never
silently switch to a second full algorithm implementation.
## When to Reach for Algorithms
| You want to... | Use |
|----------------|-----|
| Find the fewest-hop route between two nodes | `shortestPath` |
| Find the cheapest route by a numeric edge property | `weightedShortestPath` |
| List every node reachable from a source | `reachable` |
| Check whether a node is reachable at all | `canReach` |
| Get the k-hop neighborhood of a node | `neighbors` |
| Count incident edges (in, out, or both) | `degree` |
| Partition all visible nodes by undirected connectivity | `weaklyConnectedComponents` |
| Find deterministic communities by neighbor-label voting | `labelPropagation` |
| Rank nodes by global structural importance | `pageRank` |
| Rank nodes relative to weighted seed nodes | `personalizedPageRank` |
| Filter, sort, or project over traversal results | `.query().traverse()` / `.recursive()` |
| Hydrate an entity plus all its relationships | `store.subgraph()` |
Traversal algorithms return lightweight `{ id, kind, depth }` records rather
than fully hydrated nodes. WCC returns one
`{ id, kind, componentId, componentKind, size }` membership per visible node.
Label propagation returns `{ id, kind, labelId, labelKind }` memberships.
PageRank returns `{ id, kind, score }` records ordered by descending score.
Use `store.nodes..getByIds(...)` when you need the full node data.
## Shared Options
Every traversal algorithm takes the same base options:
| Option | Type | Default | Description |
|--------|------|---------|-------------|
| `edges` | `readonly EdgeKinds[]` | *(required)* | Edge kinds to follow |
| `maxHops` | `number` | `10` | Maximum traversal depth (1 – 1000) |
| `direction` | `"out" \| "in" \| "both"` | `"out"` | Edge direction |
| `cyclePolicy` | `"prevent" \| "allow"` | `"prevent"` | Compatibility option; both values use set-based node de-duplication |
| `temporalMode` | `TemporalMode` | `graph.defaults.temporalMode` | Filter applied to nodes and edges along the traversal — see [Temporal Behavior](#temporal-behavior) |
| `asOf` | `string` (ISO-8601) | *(none)* | Snapshot timestamp, required when `temporalMode: "asOf"` |
| `workingMemory` | `string` | *(inherits server `work_mem`)* | Opt-in, transaction-scoped `work_mem` override for iterative rounds on PostgreSQL (`SET LOCAL` semantics); validated as `kB\|MB\|GB` within 64kB–2147483647kB, ignored by SQLite |
`direction: "both"` treats edges as undirected. `cyclePolicy` remains accepted
for compatibility with recursive query-builder traversals, but it does not
change these algorithms: their node-set results and shortest paths never need
to revisit a node.
## shortestPath
Finds the fewest-hop path from `from` to `to`. Returns `undefined` when no
path exists within `maxHops`.
```typescript
const path = await store.algorithms.shortestPath(alice, bob, {
edges: ["knows"],
maxHops: 6,
});
if (path) {
console.log(`${path.depth} hops:`, path.nodes.map((n) => n.id));
}
```
The result contains the ordered sequence of nodes (endpoints inclusive) and
the hop count:
```typescript
type ShortestPathResult = Readonly<{
nodes: readonly Readonly<{ id: string; kind: string }>[];
depth: number;
}>;
```
Source equal to target returns a zero-length path containing just that
node. Endpoints that don't pass the resolved temporal filter return
`undefined` — see [Temporal Behavior](#temporal-behavior).
When several minimum-hop paths exist, predecessor and meeting-node ties use
the smallest `(id, kind)` identity under portable binary ordering, so the
selected path is identical across backends and execution strategies.
## weightedShortestPath
Finds the minimum-total-weight path from `from` to `to`, weighting each
traversed edge by a numeric property stored on it. This is the shape of
LDBC's "trusted connection paths" query (Interactive IC14): the cheapest
route where each hop's cost reflects, say, interaction strength — which is
usually not the fewest-hop route.
```typescript
const path = await store.algorithms.weightedShortestPath(alice, bob, {
edges: ["knows"],
weightProperty: "interactionCost",
direction: "both",
});
if (path) {
console.log(`weight ${path.totalWeight} over ${path.depth} hops`);
}
```
```typescript
type WeightedShortestPathResult = Readonly<{
nodes: readonly Readonly<{ id: string; kind: string }>[];
depth: number; // hop count, nodes.length - 1
totalWeight: number; // sum of traversed edge weights
}>;
```
Options differ from the unweighted traversals:
| Option | Type | Default | Description |
|--------|------|---------|-------------|
| `edges` | `readonly EdgeKinds[]` | *(required)* | Edge kinds to follow |
| `weightProperty` | `string` | *(required)* | Top-level edge property holding each edge's non-negative numeric weight |
| `defaultWeight` | `number` | *(none)* | Substituted for edges missing the property; must be non-negative and within the audit's upper bound (~9.7e289) |
| `direction` | `"out" \| "in" \| "both"` | `"out"` | Edge direction |
| `maxIterations` | `number` | `1000` | Relaxation-round backstop; exceeding it throws `GraphAlgorithmConvergenceError` |
| `temporalMode` / `asOf` / `workingMemory` | | | Same as the shared options above |
There is no `maxHops`: cost-ordered search does not settle nodes in hop
order, so a hop bound is not a natural stopping rule here. The algorithm
relaxes frontier nodes round by round (with parallel edges collapsing to
their cheapest member), prunes any candidate costing strictly more than a
known path to the target — equal-cost candidates stay in play so every
equal-cost route to the target is considered — and stops when no distance
improves.
**Weights are validated up front.** Before any traversal rounds run, every
visible edge of the selected kinds is audited; the call throws a typed
`InvalidEdgeWeightError` naming the offending edge when a weight is:
- **negative** — the pruning that makes the search terminate early assumes
non-negative weights, so they are rejected rather than silently mis-answered;
- **non-numeric** — a JSON string like `"5"` does not count; the property
must be stored as a JSON number (a JSON `null` counts as missing, not
non-numeric);
- **out of range** — a magnitude above ~9.7e289 (bounded so path sums can
never overflow the double range, on either backend) or a nonzero
magnitude below the smallest IEEE 754 double. One engine caveat:
SQLite's JSON parser rounds sub-denormal text like `1e-400` to `0`
before SQL can observe it, so only PostgreSQL can reject that case;
- **missing** without a configured `defaultWeight`.
The audit covers the selected edge kinds globally (not just edges the
traversal happens to reach), so a data problem fails deterministically no
matter which endpoints you query. Weight arithmetic uses IEEE 754 double
precision on both backends: total weights are always backend-identical,
and — unless a single call's `edges` list exceeds the backend's
bind-parameter budget (hundreds of kinds, where equal-weight predecessor
ties can resolve differently) — so is the returned node sequence.
Path extraction honors the backend's
[`recursiveTraversal`](/backend-setup#recursive-traversal-capability) declaration. A supported
backend reconstructs the result with one recursive statement. A backend declaring
`{ supported: false, reason }` still runs the full weighted search when it supports the temporary
working-table operations the algorithm requires; TypeGraph reconstructs the same result by walking
predecessors instead, using `path.depth + 1` extraction statements. This affects round trips, not
the selected path or its weight. The unweighted traversal algorithms do not emit a recursive CTE
and are unaffected by this capability.
## reachable
Returns every node reachable from `from` within `maxHops`, annotated with
the shortest depth at which it was discovered.
```typescript
const reachable = await store.algorithms.reachable(alice, {
edges: ["knows"],
maxHops: 3,
});
// [{ id, kind, depth: 0 }, { id, kind, depth: 1 }, ...]
```
Pass `excludeSource: true` to drop the zero-depth source entry. Results are
sorted by ascending depth.
## canReach
Fast boolean check that uses the same balanced bidirectional BFS as
`shortestPath` and stops as soon as the two visited sets meet.
```typescript
const connected = await store.algorithms.canReach(alice, bob, {
edges: ["knows"],
maxHops: 6,
});
```
Use this when you only care whether a path exists. It shares the same search
cost as `shortestPath` but does not return the reconstructed path.
## neighbors
The k-hop neighborhood of a node, with the source always excluded. `depth`
defaults to `1`, so `neighbors(alice)` returns Alice's immediate
connections.
```typescript
const immediate = await store.algorithms.neighbors(alice, {
edges: ["knows"],
});
const twoHop = await store.algorithms.neighbors(alice, {
edges: ["knows"],
depth: 2,
});
```
Semantically equivalent to `reachable({ maxHops: depth, excludeSource: true })`,
but the name reads more naturally for neighborhood queries.
## degree
Counts edges incident to a node under the resolved temporal filter.
Self-loops contribute once when `direction` is `"both"`.
```typescript
const friends = await store.algorithms.degree(alice, {
edges: ["knows"],
direction: "out",
});
const connections = await store.algorithms.degree(alice, {
edges: ["knows"],
// direction: "both" is the default
});
const everything = await store.algorithms.degree(alice);
// No `edges` option counts all edge kinds in the graph
```
| Option | Type | Default | Description |
|--------|------|---------|-------------|
| `edges` | `readonly EdgeKinds[]` | all kinds | Edge kinds to count (empty array returns 0) |
| `direction` | `"out" \| "in" \| "both"` | `"both"` | Count outgoing, incoming, or either |
| `temporalMode` | `TemporalMode` | `graph.defaults.temporalMode` | Filter applied to the counted edges |
| `asOf` | `string` (ISO-8601) | *(none)* | Snapshot timestamp, required when `temporalMode: "asOf"` |
`degree` runs a single `COUNT` query, not a recursive CTE, so it's
efficient even for hub nodes with thousands of edges.
## PageRank and Personalized PageRank
`pageRank` scores every visible node by the stationary probability of a random
walk over the selected edges. `personalizedPageRank` uses the same power
iteration but teleports to weighted seed nodes instead of uniformly across the
graph. Both methods operate on the visible induced graph; `nodeKinds` can narrow
that graph without allowing transitions through excluded nodes.
```typescript
const globalScores = await store.algorithms.pageRank({
edges: ["knows", "cites"],
nodeKinds: ["Person", "Paper"],
direction: "out",
topK: 20,
});
const relatedToAlice = await store.algorithms.personalizedPageRank({
edges: ["knows"],
direction: "both",
seeds: [
{ id: alice.id, kind: "Person", weight: 3 },
{ id: bob.id, kind: "Person" }, // weight defaults to 1
],
});
// [{ id, kind, score }, ...] — highest score first
```
| Option | Type | Default | Description |
|--------|------|---------|-------------|
| `edges` | `readonly EdgeKinds[]` | *(required)* | Edge kinds defining transitions |
| `nodeKinds` | `readonly NodeKinds[]` | all visible kinds | Induced node subgraph to rank |
| `direction` | `"out" \| "in" \| "both"` | `"out"` | Follow stored, reversed, or undirected transitions |
| `dampingFactor` | `number` in `[0, 1)` | `0.85` | Probability of following an edge rather than teleporting |
| `tolerance` | positive finite `number` | `1e-8` | Maximum per-node score change accepted as convergence |
| `maxIterations` | positive integer | `1000` | Power-iteration backstop; exceeding it throws |
| `topK` | positive safe integer | all scores | Return the first K scores after deterministic ordering |
| `workingMemory` | PostgreSQL memory string | inherited | Same transaction-scoped override described above |
Personalized seeds are qualified by the full `(kind, id)` identity so same-ID
nodes of different kinds remain distinct. Seed weights must be finite and
positive; duplicates are combined and the vector is normalized. A seed outside
the temporal/node-kind scope throws `ConfigurationError` instead of silently
losing teleport mass.
Dangling-node mass is redistributed through the same teleport vector, keeping
the scores normalized to approximately one. Parallel physical edges retain
their multiplicity. With `direction: "both"`, a physical self-loop contributes
once rather than once per expansion direction.
PageRank uses double-precision arithmetic on both backends. Repeated runs on
SQLite are deterministic; on PostgreSQL a plan change between runs can reorder
floating-point summation and shift scores by a few last bits, which can reorder
near-tied rows. SQLite and PostgreSQL scores are expected to agree within the
requested numerical tolerance rather than bit-for-bit. Exact score ties use
portable binary `(id, kind)` ordering. As with WCC, an exhausted iteration
budget throws `GraphAlgorithmConvergenceError`—partial scores are never
returned.
`topK` bounds only result extraction: TypeGraph applies the limit in SQL after
ordering by descending score and the deterministic tie-breaker, so rows beyond
K do not reach the driver. PageRank still initializes and iterates over the
entire visible induced graph; `topK` does not make the graph computation itself
partial.
Convergence needs roughly `ln(1/tolerance) / ln(1/dampingFactor)` rounds in the
worst case. With the default `tolerance` and `maxIterations`, damping factors
above roughly `0.985` cannot converge in time — every round still runs against
the full working table before the budget-exhaustion error is thrown — so raise
`maxIterations` (or loosen `tolerance`) alongside a high damping factor. The
tolerance is also absolute: typical scores are on the order of `1/N`, so
reliably ranking the low-score tail of a large graph calls for a
proportionally smaller tolerance.
## labelPropagation
Runs deterministic synchronous Community Detection using Label Propagation
(CDLP) over the undirected projection of selected edge kinds. Each visible node
starts with its own `(id, kind)` identity as its label. A round adopts the most
frequent label among the node's visible neighbors from the previous round;
equal vote counts resolve to the minimum label under portable binary ordering.
```typescript
const memberships = await store.algorithms.labelPropagation({
edges: ["knows", "worksAt"],
nodeKinds: ["Person", "Company"], // optional; all kinds by default
maxIterations: 1000, // default
onMaxIterations: "throw", // default; "return" accepts the fixed-round labeling
});
// [{ id, kind, labelId, labelKind }, ...]
```
The graph is a neighbor set for voting: parallel edges and repeated selected
edge kinds do not multiply a vote, and self-loops do not make a node its own
neighbor. An isolated in-scope node therefore retains its initial label.
Labels and vote counts are exact integers, and edge-kind chunks are accumulated
before a label is staged, so results do not depend on bind limits or chunk
order. Changed-node frontiers restrict each later round to nodes whose neighbor
labels may have changed.
Synchronous voting has no self-vote, so structures whose neighborhoods mirror
each other never converge — they alternate between two labelings forever. This
covers every tree-shaped component (an isolated edge pair, a path, a star, an
org-chart hierarchy), every even-length cycle, and complete bipartite blocks;
empirically most sparse random graphs contain at least one such component.
One oscillating component anywhere in the selection prevents global
convergence, and raising `maxIterations` cannot help. Dense neighborhoods
built on odd cycles — triangles, cliques, and communities of them — converge.
`onMaxIterations` selects the completion contract:
- `"throw"` (default) returns only a converged labeling. A detected
period-two oscillation throws `GraphAlgorithmConvergenceError` immediately
rather than burning the remaining budget, and exhausting `maxIterations`
throws the same typed error. Partial labelings are never returned.
- `"return"` yields the labeling after exactly `maxIterations` synchronous
rounds (or at convergence, whichever comes first) — the fixed-round
contract of the LDBC Graphalytics CDLP benchmark. Synchronous rounds are
deterministic and chunk-independent, so this labeling is exact and
identical on SQLite and PostgreSQL. Use it for tree-shaped or mixed data
where a converged labeling need not exist; a detected oscillation
fast-forwards to the parity-exact final labeling instead of running every
remaining round.
Like WCC and PageRank, label propagation requires
`backend.capabilities.graphAnalytics?.supported === true`, runs in one
repeatable snapshot, honors temporal and recorded-time views, and accepts the
transaction-scoped `workingMemory` option.
## weaklyConnectedComponents
Computes an exact partition over the undirected projection of the selected edge
kinds. By default every visible node is returned. `nodeKinds` restricts the
operation to the induced subgraph over those kinds; an in-scope node with no
selected incident edge forms a singleton component.
```typescript
const memberships =
await store.algorithms.weaklyConnectedComponents({
edges: ["knows", "worksAt"],
nodeKinds: ["Person", "Company"], // optional; all kinds by default
minComponentSize: 2, // optional; singleton components are omitted
maxIterations: 1000, // default
});
// [{ id, kind, componentId, componentKind, size }, ...]
```
The component identity is the smallest `(id, kind)` member under portable
binary ordering. Results and representatives are therefore deterministic on
SQLite and PostgreSQL even when the PostgreSQL database uses a linguistic
default collation. Edges are always treated as undirected—WCC has no
`direction` option.
`minComponentSize` is an inclusive positive safe-integer filter: a value of
`2` returns every member of components containing two or more nodes, while a
value of `1` preserves the default result. TypeGraph computes component sizes
and filters memberships in extraction SQL, before rows reach the driver. The
full visible induced graph is still processed to convergence, so the option
bounds result materialization rather than WCC computation.
WCC is iterative and runs multiple SQL rounds in one repeatable snapshot. It
requires `backend.capabilities.graphAnalytics?.supported === true`; built-in
SQLite and PostgreSQL connections that permit temporary tables advertise
support. Cloudflare D1, Durable Objects SQLite, `neon-http`, and other
restricted backends throw `UnsupportedBackendCapabilityError` before temporary
state is created. PostgreSQL is checked again at execution time, because a read
replica or a role without `TEMP` can reject the working-table transaction even
when the backend's static capability is true: a replica refuses the read-write
transaction itself, and a role without `TEMP` refuses the `CREATE TEMP TABLE`
inside it. Both are reported as `UnsupportedBackendCapabilityError`; its `cause`
preserves the original database error. If propagation has not converged after
`maxIterations`, TypeGraph throws `GraphAlgorithmConvergenceError` rather than
returning a partial partition.
Each round expands only the indexed frontier of nodes whose label changed in
the previous round. Candidate labels remain staged until every edge-kind chunk
has run, preserving synchronous and bind-limit-independent iteration semantics;
only changed rows are written back. Late convergence rounds therefore avoid
rescanning the full edge set and do not churn unchanged working rows.
PostgreSQL refreshes planner statistics for a sufficiently large temporary
working table and refreshes them again after multiplicative growth. This avoids
plans based on PostgreSQL's initial one-row estimate for a new temporary table.
The policy is automatic, applies to WCC, label propagation, PageRank, and growing traversal frontiers, and is
a no-op on SQLite.
Iterative operations (WCC, label propagation, PageRank, and the working-table traversals) accept an opt-in
`workingMemory` override of the session's `work_mem` for their rounds. When
set, it is applied with `SET LOCAL work_mem` semantics inside the operation's
own transaction — the session and server settings are never modified, and the
override ends with the transaction. When omitted (the default), operations
inherit the server's configured `work_mem`.
Note that `work_mem` is a threshold each sort/hash operator (and each parallel
worker) may allocate up to, **not** a per-operation budget: a single round can
allocate several multiples of it, and concurrent algorithm calls multiply
again. Raise it deliberately — for example `workingMemory: "64MB"` keeps
whole-graph rounds from spilling their sorts to disk on large single-tenant
analytical runs (such as LDBC SNB SF1-scale benchmarks) — rather than as a
blanket setting on a shared cluster. The value must be a plain integer with a
`kB`, `MB`, or `GB` suffix within PostgreSQL's accepted `work_mem` range
(64kB to 2147483647kB) — both backends reject malformed or out-of-range
values with the same error. SQLite validates and otherwise ignores it.
Without `nodeKinds`, WCC seeds every visible node, so unrelated nodes still
appear as singleton components. On heterogeneous graphs, pass the kinds that
define the graph being analyzed. For example, `{ edges: ["knows"], nodeKinds:
["Person"] }` avoids seeding posts and comments while retaining isolated people.
Temporal views expose the same facade with the coordinate sealed:
```typescript
const historical = await store
.asOf("2024-01-01T00:00:00.000Z")
.algorithms.weaklyConnectedComponents({ edges: ["knows"] });
```
## Passing Nodes or IDs
Every node-oriented algorithm accepts either a raw ID string or any object with
an `id: string` field — `Node`, `NodeRef`, the lightweight records returned by
traversals, and `store.subgraph()` results all work. WCC is a whole-graph
operation and does not take a node identifier. Global PageRank is also
whole-graph; personalized PageRank instead takes explicit `{ id, kind, weight? }`
seed identities.
```typescript
const alice = await store.nodes.Person.getById(aliceId);
// All equivalent
store.algorithms.canReach(alice, bobId, { edges: ["knows"] });
store.algorithms.canReach(alice.id, bobId, { edges: ["knows"] });
store.algorithms.canReach(aliceId, { kind: "Person", id: bobId }, {
edges: ["knows"],
});
```
## Direction and Cycles
`direction: "both"` lets you treat a directed edge kind as undirected for
reachability questions:
```typescript
// Did Dave ever know Alice, regardless of who "added" whom first?
const knew = await store.algorithms.canReach(dave, alice, {
edges: ["knows"],
direction: "both",
});
```
Cycles cannot multiply work: a node enters the visited set once, at its minimum
depth. This differs from `.recursive()`, whose path-returning semantics use the
per-path behavior described in
[Cycle Detection](/queries/recursive#cycle-detection). The algorithm
`cyclePolicy` option is therefore compatibility-only.
## Depth Limits
`maxHops` is capped at `MAX_EXPLICIT_RECURSIVE_DEPTH` (1000). The default of 10
covers typical connectivity questions on most graphs. Within that bound,
single-source traversal examines each reached node once and each incident edge
at most once per direction, rather than enumerating paths.
Each expanded level costs a database round trip (inline working relations are
split to respect the backend bind limit). Transactional backends keep those
statements in one snapshot and drop their temporary working table in a
`finally` cleanup; non-transactional backends use their normal best-effort
multi-statement behavior. The application-clock valid-time instant is pinned
once for the whole traversal, and recorded-time views remain pinned to their
requested recorded coordinate.
## Temporal Behavior
Algorithms honor the same temporal model as the rest of the store. The
default temporal mode is `graph.defaults.temporalMode` (typically
`"current"`), and every algorithm accepts per-call `temporalMode` and
`asOf` options.
```typescript
// Default: uses graph.defaults.temporalMode — typically "current".
await store.algorithms.shortestPath(alice, bob, { edges: ["knows"] });
// Snapshot at a specific point in time. Both nodes and edges must have
// been valid at that timestamp to participate in the traversal.
await store.algorithms.shortestPath(alice, bob, {
edges: ["knows"],
temporalMode: "asOf",
asOf: "2023-01-15T00:00:00.000Z",
});
// Include validity-ended (but not soft-deleted) rows — useful for
// historical traversal without needing a specific timestamp.
await store.algorithms.reachable(alice, {
edges: ["knows"],
temporalMode: "includeEnded",
});
// Include soft-deleted rows too. Traversal can cross through tombstones.
await store.algorithms.canReach(alice, ghost, {
edges: ["knows"],
temporalMode: "includeTombstones",
});
```
**Semantic rules:**
- The temporal filter applies to **both nodes and edges** along the
traversal. An edge can only be traversed if it passes the filter *and*
its endpoint node passes too.
- `asOf` is required when `temporalMode: "asOf"` and rejected (throws
`ValidationError`) in every other mode — pinning an instant is only
meaningful in `"asOf"` mode.
- Temporal filtering is orthogonal to `cyclePolicy` — cycle detection
operates on path membership, not on time. A node that was valid → ended
→ re-valid is not treated as "visited" just because it appears in two
validity periods.
- The shortest-path self-path short-circuit also respects the resolved
mode: calling `shortestPath(a, a, ...)` returns `undefined` if node `a`
does not pass the temporal filter, and a zero-hop result otherwise.
## End-to-End Example
The runnable example
[`examples/14-research-copilot.ts`](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/14-research-copilot.ts)
combines every algorithm with semantic search and ontology-expanded topic
matching over a corpus of landmark ML papers. It produces an
explainable literature-review digest in one run against a single SQLite
file — a good starting point for your own RAG + graph workloads.
## What's Not Included
These algorithms cover shortest path (weighted and unweighted),
reachability, neighborhoods, degree, weakly connected components,
deterministic label propagation, and global and personalized PageRank. They
do **not** cover:
- Strongly connected components
- Topological sort
- Centrality measures beyond degree (betweenness, closeness, eigenvector)
- Modularity-optimizing community detection such as Leiden or Louvain
For those, export edges via `.query().traverse()` or `store.subgraph()` and
use a specialized library such as
[graphology](https://graphology.github.io/) in memory. See
[Limitations](/limitations#graph-analytics-limits) for the full list
of excluded analytics.
# Graph Extensions
> Extend a TypeGraph schema at runtime — durable, multi-process safe, with full Zod validation and unique-constraint enforcement.
Graph extensions let your application declare new node and edge
kinds **at runtime** — durable across restarts, with semantic parity to
compile-time `defineNode` / `defineEdge`. The motivating use case:
**agent-driven schema induction**, where an LLM proposes a typed schema
from a corpus, an operator approves it, and the live graph immediately
ingests under the new schema with no code change or restart.
:::note[See it end-to-end]
For a runnable scenario with an operator-approved agent in TypeScript,
see [Agent-Driven Schema](/examples/agent-driven-schema). For the same
loop driven by an open-weight LLM against public-record clinical data —
with a repair loop and smoke-test pattern — see
[`pdlug/typegraph-clinical-demo`](https://github.com/pdlug/typegraph-clinical-demo).
:::
This guide covers the core verbs:
| Verb | Purpose |
| ------------------------------------------------ | ----------------------------------------------------------------- |
| `defineGraphExtension` | Build a typed extension (pure value, no I/O) |
| `store.evolve(extension)` | Atomically commit a new schema version with the extension applied |
| `store.introspect()` | Snapshot the merged schema, persisted extension, version, and hash |
| `store.materializeIndexes()` | Run declared `CREATE INDEX` DDL against the live database |
| `store.deprecateKinds(...)` / `undeprecateKinds` | Soft-deprecate kinds for codegen / lint signaling |
| `store.removeKinds(...)` | Remove graph-extension-declared kinds from the active schema |
| `store.materializeRemovals()` | Delete rows queued by graph-extension-kind removal |
For the schema-management primitives that graph extensions ride on top
of, see [Schema Migrations](/schema-management) and [Evolving
Schemas](/schema-evolution).
## When to use graph extensions
Use them when **the kind set is not known at code time**:
- Agent / LLM proposes a new typed schema from observed data.
- Multi-tenant deployments where each tenant defines their own kinds.
- ETL pipelines that ingest sources with shifting structure.
- Plugins / extensions that contribute kinds at install time.
For everything else — kinds you can declare in TypeScript at deploy time
— use the compile-time DSL (`defineNode`, `defineEdge`, `defineGraph`).
The compile-time path is type-safe end-to-end; graph extensions trade
some of that type-safety for the ability to evolve without redeploying.
## A complete example
```ts
import { z } from "zod";
import {
createStoreWithSchema,
defineGraph,
defineNode,
defineGraphExtension,
} from "@nicia-ai/typegraph";
import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local";
// 1. Boot with a compile-time kind.
const Document = defineNode("Document", {
schema: z.object({ title: z.string(), body: z.string() }),
});
const baseGraph = defineGraph({
id: "research_corpus",
nodes: { Document: { type: Document } },
edges: {},
});
const { backend } = createLocalSqliteBackend();
const [store] = await createStoreWithSchema(baseGraph, backend);
// 2. An agent proposes a new kind at runtime.
const proposal = defineGraphExtension({
nodes: {
Paper: {
description: "An academic paper inferred from the corpus",
properties: {
title: { type: "string", minLength: 1 },
doi: { type: "string", minLength: 1 },
year: { type: "number", int: true, min: 1900, max: 2100 },
},
unique: [{ name: "paper_doi_unique", fields: ["doi"] }],
},
},
indexes: [
{
entity: "node",
kind: "Paper",
name: "paper_by_doi",
fields: ["doi"],
unique: true,
},
],
});
// 3. Operator approves; commit atomically.
const evolved = await store.evolve(proposal);
// 4. Use the dynamic-collection accessor (the type system does not
// widen for extension kinds — see "Reaching extension kinds" below).
const papers = evolved.getNodeCollection("Paper")!;
await papers.create({
title: "Attention is all you need",
doi: "10.5555/3295222.3295349",
year: 2017,
});
```
A complete runnable version is in [`examples/16-graph-extensions.ts`](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/16-graph-extensions.ts).
## Instantiate a graph template
For many tenant graphs with the same final shape, register a verified schema
once and instantiate each target as schema version 1. The template document
stays in the database; instantiation sends only its id, the target graph id,
and a client-computed content hash.
```ts
const [source] = await createAdapterStoreWithSchema(baseGraph, adminBackend);
const template = await registerGraphTemplate(adminBackend, {
templateId: "research-corpus-v1",
reconciled: source.reconciledSchema,
});
const target = await instantiateGraph(adminBackend, {
template,
graphId: "tenant-123",
});
// The returned target snapshot can create a request-scoped Store without
// another schema read. `tenantGraph` is the same compile-time shape with id
// "tenant-123".
const store = createAdapterStore(tenantGraph, runtimeBackend, {
reconciled: target.reconciled,
});
```
Instantiation is idempotent for the same template and target. A target already
initialized with a different schema is refused. On PostgreSQL the clone
statement takes the same graph advisory-lock key as schema commits; SQLite's
schema and marker writes are serialized by its writer lock.
Templates clone the source graph's graph-local runtime-contribution activation
markers along with its schema row. Deployment-scoped physical attestations,
such as the shared fulltext table marker, already belong to the database and
are not duplicated per template target. A target can therefore be reopened with
`createVerifiedStore` or `createVerifiedAdapterStore` from a later serverless
isolate without another schema reconciliation or provisioning DDL step.
Registration and instantiation are DML-only. They assert the deployment-wide
base-schema marker and throw `BaseSchemaMigrationError` rather than attempting
DDL when adoption is missing, stale, or newer than the running library. Run the
normal privileged `createStoreWithSchema` boot or the published external
migration before handing runtime requests to a DML-only role.
Template eligibility follows the backend's capabilities. Embedding declarations
are allowed when the backend is configured with `vector: false`, because that
backend has no graph-scoped vector storage to clone. A vector-enabled backend
continues to refuse embedding-bearing schema-only templates.
## The graph extension
`defineGraphExtension` accepts a structured value describing the new
kinds. The extension is JSON-serializable — that's load-bearing for
durability (see [Restart parity](#restart-parity-the-load-bearing-invariant)).
### Document format versioning
Every document carries a `version` field (currently `1`). The validator
stamps the version automatically when you call
`defineGraphExtension`, so consumer code never has to set it
explicitly. Stored documents from before this field existed are
treated as `version: 1` (the legacy default).
The forward-compat policy:
- **Additive minor changes** (new optional property modifier, new
`format` value, new top-level slice within the same major) ride
forward without bumping `version`. The validator does not reject
unknown top-level keys, and the persistence-side zod is `.loose()`
on every nested object — an older runtime reading a newer extension
silently ignores unknown fields and continues working.
- **Breaking changes** bump `version` to a higher major. An older
runtime reading a higher-version extension fails with
`GRAPH_EXTENSION_VERSION_UNSUPPORTED` and an actionable error
pointing the operator at upgrading the library — there is no
automatic downgrade path. The current major is exported as
`CURRENT_GRAPH_EXTENSION_VERSION` for tooling that wants to
pre-flight check.
- **Legacy extensions** (committed before `version` existed) and
extensions that explicitly omit `version` are interpreted as
`LEGACY_GRAPH_EXTENSION_VERSION`, pinned permanently to `1`. This
is deliberately distinct from `CURRENT`: when a future v2 ships,
legacy v1 extensions continue parsing as v1 (so the version-mismatch
path can route them through migration) rather than being silently
re-classified as v2 by a default-equals-current rule.
```ts
import {
CURRENT_GRAPH_EXTENSION_VERSION,
LEGACY_GRAPH_EXTENSION_VERSION,
} from "@nicia-ai/typegraph";
console.log(CURRENT_GRAPH_EXTENSION_VERSION); // 1 (today; bumps with breaking changes)
console.log(LEGACY_GRAPH_EXTENSION_VERSION); // 1 (always; the pre-versioning default)
```
### Property types (the v1 subset)
The following types are supported. The set is deliberately small so that
LLM-induced schemas can be audited at a glance and so the persistence
layer never has to reconstruct opaque Zod refinements from JSON.
| Type | JSON shape |
| --------- | ------------------------------------------------------------------------- |
| `string` | `{ type: "string", minLength?, maxLength?, pattern?, format? }` |
| `number` | `{ type: "number", int?, min?, max? }` |
| `boolean` | `{ type: "boolean" }` |
| `enum` | `{ type: "enum", values: ["a", "b", ...] }` |
| `array` | `{ type: "array", items: }` |
| `object` | `{ type: "object", properties: { foo: , ... } }` (one nesting only) |
Supported string formats: `"datetime"`, `"uri"`, `"email"`, `"uuid"`,
`"date"`. These route to the corresponding Zod factories
(`z.iso.datetime()`, `z.url()`, `z.email()`, `z.uuid()`, `z.iso.date()`).
Modifiers available on every property:
- `optional?: true` — omits from the required set.
- `description?: string` — surfaces in tooling.
- `searchable?: SearchableModifier` — string only; routes through the
`searchable()` brand for fulltext indexing.
- `embedding?: { dimensions: number }` — array-of-number only; routes
through the `embedding()` brand for vector search.
### Unique constraints
Pass `unique: [{ name, fields, ... }]` per kind. The `name` is required
(used as the diffing identity key) and must be unique within the kind.
Supports `scope`, `collation`, and a restricted `where` clause limited
to `isNull` / `isNotNull` (the only operations round-trippable through
the persisted form).
### Relational indexes
Pass `indexes: [...]` at the document top level to declare relational
indexes for graph-extension or compile-time host kinds:
```ts
const proposal = defineGraphExtension({
nodes: {
Paper: {
properties: {
doi: { type: "string" },
title: { type: "string", searchable: { language: "english" } },
},
},
},
indexes: [
{
entity: "node",
kind: "Paper",
name: "paper_by_doi",
fields: ["doi"],
unique: true,
},
],
});
```
Index `name`s are unique across the merged graph. A graph-extension index that
reuses a compile-time index name, or a later extension that reuses an
earlier graph-extension index name for a different declaration, is rejected.
### Edges
Edges follow the same shape as nodes and declare endpoints using kind names.
An array-valued `to` permits every combination of `from` and `to` kinds.
A map-valued `to` restricts targets by source kind:
```ts
const proposal = defineGraphExtension({
edges: {
assignedTo: {
from: ["Employee", "Student"],
to: {
Employee: ["Department"],
Student: ["Course"],
},
properties: {},
},
},
});
const evolved = await store.evolve(proposal);
```
All endpoint kinds must resolve in the merged graph when `evolve()` runs.
Map keys must exactly cover `from`, and each target array must be nonempty.
The map is persisted and restored on restart, preserving the same runtime pair
validation as [compile-time declarations](/core-concepts#source-dependent-targets).
Adding allowed pairs broadens an extension edge. Removing a pair tightens it,
even if the overall source and target kind sets stay the same. Tightening
currently requires the **entire edge kind** to be empty, not just the removed pair.
### Ontology
Pass `ontology: [{ metaEdge, from, to }, ...]` to declare ontology
relations between kinds (subClassOf, partOf, etc.). The meta-edge name
must match a meta-edge known to the merged graph.
## `store.evolve(extension, options?)`
```ts
const evolved = await store.evolve(extension);
const evolved = await store.evolve(extension, { ref });
const evolved = await store.evolve(extension, { eager: {} });
```
`evolve` is the consumer-facing primitive that drives extension. It:
1. **Catches up to persisted state** — folds any persisted
extension and deprecation set into the local baseline so a
stale store doesn't trample another writer's progress.
2. **Merges** the new extension into the baseline graph. Re-declaring
an existing extension kind with the same shape is a no-op; with a
non-additive change against existing rows it throws
`IncompatibleChangeError` (code `INCOMPATIBLE_CHANGE`).
3. **Atomically commits** a new schema version via `commitSchemaVersion`
(CAS on the active version).
4. **Returns the resulting `Store`** carrying the extended graph. The type
parameter `G` does NOT widen — see [Reaching extension
kinds](#reaching-extension-kinds-from-the-type-system) below.
:::caution[Use the returned Store]
`Store` instances are immutable schema snapshots. After `evolve()` resolves,
use its returned Store for **every subsequent Store operation in the same
request**, not just operations involving newly added kinds. A Store captured
before the commit remains pinned to the previous schema version, so its next
managed write fails the schema-version fence.
:::
### Plan outside and apply inside a caller-owned transaction
`planEvolution(extension)` prepares a named schema change before the caller
opens its write transaction. It returns an immutable `"noop"` or `"change"`
plan with `graphId`, `baseline: { version, hash }`, and
`result: { version, hash }`. Change plans expose an ordered `requirements`
array whose entries name new-kind additions, empty-kind checks, vector slots,
and identity work. A `new-kind` entry describes a graph delta; it does not
indicate that a removal is queued or direct callers to run
`materializeRemovals()`. The plan is opaque and bound to the loaded TypeGraph
module: it cannot be serialized, cloned, or reconstructed. It can be passed between
compatible Stores for the same graph that use the same loaded module; apply
still checks the active graph and fenced baseline version/hash. The
default `{ source: "database" }` reloads the active schema. `{ source:
"cached" }` uses a previously loaded planning snapshot on the same Store; a
cached plan is not fresh database evidence. A stale baseline is refused during
apply, so retry by rolling back the whole caller transaction and replanning
outside it.
Use `withEvolvedTransaction(nativeTx, plan, callback, { waitBudgetMs })` when
the schema change, TypeGraph writes, and application SQL must share the
caller's commit. Enter this boundary before other TypeGraph callbacks on the
same native transaction. The callback receives an evolved transaction context,
including its reads, collections, and supported composition operations. It
does not receive a replacement root Store.
```ts
const cachedStore = store;
const ref = { current: cachedStore };
const plan = await cachedStore.planEvolution(proposal);
const writerStore = cachedStore.withBackend(writerBackend);
const provisional = await db.transaction(async (nativeTx) => {
const outcome = await writerStore.withEvolvedTransaction(
nativeTx,
plan,
async (tx) => {
const person = await tx.nodes.Person.create({ name: "Ada" });
return person.id;
},
plan.status === "change" ? { waitBudgetMs: 5000 } : undefined,
);
await nativeTx.insert(applicationEvents).values({
personId: outcome.result,
schemaVersion: outcome.receipt.schema.version,
});
return outcome;
});
const refreshed = await cachedStore.refreshSchema({
minVersion: provisional.receipt.schema.version,
ref,
});
```
The callback and receipt finish before the outer SQL transaction commits.
The receipt's schema version/hash and recorded anchor are provisional until
that commit succeeds. A callback failure must reject the outer transaction;
catching it and committing does not prove rollback. Callback contexts and
queries built from them expire when the callback returns, including after a
failure. Use `refreshSchema()` only after awaiting a successful outer commit.
When the cached Store already matches `minVersion`, refresh returns it
without a read; that shortcut does not check for a newer database version.
Otherwise refresh reads the active schema, accepts a newer committed version,
and refuses a missing or older one. It applies no extension or storage
provisioning.
The adopted apply path supports metadata-only changes, required-empty checks,
and transactional identity and vector provisioning. Configure the adapter with
`schemaProvisioning: "transactional"` on a privileged connection when the plan
names new vector slots or identity work. The adapter's default DML-only policy
refuses those requirements before the schema fence, DDL, callback, or mutation.
The privileged path rechecks storage on the caller's fenced session, provisions
the required relations and vector contribution markers there, and rolls them
back with the outer transaction. Missing bootstrap tables still refuse; run
the normal bootstrap before serving adopted evolution requests.
The apply path uses a bounded schema fence, baseline validation, and a
version CAS; it can issue multiple statements. `waitBudgetMs` bounds fence
acquisition for change plans; no-op plans refuse an explicit `waitBudgetMs`.
On timeout, roll back and retry the complete native transaction.
Do not pass `ref` or eager-index options to `withEvolvedTransaction`; they
cannot be honored before the outer commit and are refused.
Generic and concurrent eager indexes remain explicit maintenance after commit:
call `materializeIndexes()` on the refreshed Store when they are needed.
When a wiring pass produces a no-op, ordinary recorded adoption is sufficient
and never takes the exclusive evolution fence. Reconcile only if the Store is
behind the named snapshot; matching versions make this refresh a cached read:
```typescript
if (plan.status === "noop") {
const current = await store.refreshSchema({ minVersion: plan.baseline.version });
await db.transaction((nativeTx) =>
current.withRecordedTransaction(nativeTx, async (tx) => {
await tx.nodes.Person.create({ name: "Ada" });
}),
);
}
```
### The `ref` pattern
`Store` is immutable by construction — `evolve()` returns the Store for the
resulting schema, using a fresh instance when the schema advances. Long-lived
consumer code that holds the Store in a singleton needs a way to re-bind it
before the operation completes. Pass `options.ref: { current: store }` (a
`StoreRef>`):
```ts
const ref: StoreRef> = { current: store };
const evolved = await ref.current.evolve(extension, { ref });
// `ref.current === evolved`; use either reference from here on.
const papers = ref.current.getNodeCollectionOrThrow("Paper");
await papers.create({ title: "...", doi: "...", year: 2024 });
```
Capture `ref.current` once at request entry only when that request will not
change the schema. If it does call `evolve()`, switch to the returned Store or
dereference `ref.current` again after the call. `StoreRef` is structurally
just `{ current: T }`; the library doesn't provide a factory because the
consumer composes the handle themselves (it could be a Vue ref, MobX
observable, Zustand atom, etc.).
The ref covers schema changes made by calls that receive it. If another process
or isolate can advance the schema, probe the committed version before reusing a
cached Store and perform a verified open when it changes. See
[Per-request connections: cache the verified Store](/integration#per-request-connections-cache-the-verified-store)
for the complete `getCommittedSchemaVersion()` recipe.
### Eager materialization
Pass `eager: {}` to materialize indexes immediately after the schema
commit:
```ts
const evolved = await store.evolve(extension, { eager: {} });
```
Or pass options for finer control:
```ts
// Restrict to the extension kind whose index was declared in the
// proposal above.
const evolved = await store.evolve(extension, {
eager: { kinds: ["Paper"], stopOnError: true },
});
```
Omit `eager` to skip materialization and run `materializeIndexes()`
later.
Per-index failures throw `EagerMaterializationError` AFTER the new
`Store` is constructed and `ref.current` is updated, so the caller can
recover via the ref handle. The schema commit is **not** rolled back
if materialization fails — eager is convenience, not a transaction.
```ts
const ref = { current: store };
try {
await store.evolve(extension, { ref, eager: {} });
} catch (error) {
if (error instanceof EagerMaterializationError) {
// Schema is committed; ref.current is the new store.
log.warn(
{ failed: error.failedIndexNames },
"indexes did not materialize; will retry",
);
await ref.current.materializeIndexes();
} else {
throw error;
}
}
```
## Reaching extension kinds from the type system
TypeScript can't see kinds that don't exist at compile time. The
`Store` returned by `evolve()` keeps the same generic parameter as
the original — `evolved.nodes.Paper` would not type-check.
The escape hatch is `store.getNodeCollection(kind)` and
`store.getEdgeCollection(kind)`, which return a typed
`DynamicNodeCollection` / `DynamicEdgeCollection`:
```ts
const papers = evolved.getNodeCollection("Paper");
if (papers === undefined) {
throw new Error("Paper kind not registered on this store");
}
await papers.create({ title: "...", doi: "...", year: 2024 });
const all = await papers.find({});
```
The throwing variants `getNodeCollectionOrThrow(kind)` /
`getEdgeCollectionOrThrow(kind)` are the right call when the caller
already knows the kind has been evolved onto the store — they raise
`KindNotFoundError` with the offending `kindName`, `entity`, and host
`graphId` instead of returning `undefined`, so a typo fails loudly at
the call site rather than crashing later on `papers!.create(...)`.
`DynamicNodeCollection` exposes the same CRUD surface as
`store.nodes.X` — `create`, `getById`, `find`, `update`, `delete`,
etc. — but with `DynamicNode` element types since the specific Zod schema isn't
visible to TypeScript at the call site.
In TypeScript, nodes returned by a dynamic collection carry the nominal
`DynamicNode` type, preserving the requested kind literal. That proof lets
runtime kinds participate directly in Operational Identity without weakening
compile-time references to arbitrary string kinds:
```ts
const paper = await evolved
.getNodeCollectionOrThrow("Paper")
.create({ title: "Runtime schemas" });
await evolved.identity.assertSame(document, paper);
```
Identity reads can consequently return `IdentityNodeReference` values for
either compile-time or runtime kinds. See [Operational Identity](/identity).
For consumers that need the live Zod schema itself — MCP tool wrappers
that validate inputs before forwarding to `collection.create`, or
agent prompts that want richer JSON Schema than `introspect()`
exposes — `store.getNodePropsSchema(kind)` /
`getNodePropsSchemaOrThrow(kind)` (and the edge counterparts) return
the exact `z.ZodObject` the store uses internally. Identity holds:
`evolved.getNodePropsSchema("Paper")` is the same instance the store
parses against on `papers.create(...)`.
```ts
import { z } from "zod";
const schema = evolved.getNodePropsSchemaOrThrow("Paper");
const parsed = schema.parse(input); // same Zod issues as papers.create surfaces
const jsonSchema = z.toJSONSchema(schema); // for MCP tool descriptions
```
These accessors return only the props validator. Operation-level
checks — uniqueness, endpoint resolution, temporal validity, backend
constraints — still run only through `collection.create` / `update`.
See [Dynamic Props Schema Access](/schemas-stores#dynamic-props-schema-access)
for the full reference.
For codegen consumers, the kind set is reachable by iterating the
registry's `nodeKinds` and `edgeKinds` maps:
```ts
const allNodeKinds = [...store.registry.nodeKinds.keys()];
const allEdgeKinds = [...store.registry.edgeKinds.keys()];
const personType = store.registry.getNodeType("Person"); // NodeType | undefined
```
`KindRegistry` also exposes `hasNodeType(name)` / `hasEdgeType(name)`
for existence checks.
### Querying extension kinds
`store.query()` requires every `from` / `traverse` / `to` kind to be a
compile-time literal in `Store`. The string-keyed siblings
`fromDynamic` / `traverseDynamic` / `optionalTraverseDynamic` /
`toDynamic` admit kinds added via `evolve()` so an MCP server (or any
caller working from kind names in a string variable) can build typed
multi-hop traversals without `as any`:
```ts
const rows = await store.query()
.fromDynamic("Paper", "p")
.traverseDynamic("authoredBy", "a")
.toDynamic("Author", "u")
.whereNode("p", (p) => p.field("year").number().gte(2020))
.select((ctx) => ({ paper: ctx.p, author: ctx.u, edge: ctx.a }))
.execute();
```
Each method runtime-validates against the registry: a typo throws
`KindNotFoundError`, and a `toDynamic` target that isn't a valid
endpoint for the current edge / direction throws `EndpointError`.
Compile-time `from` / `traverse` / `to` are unchanged.
When the extension document is available in typed code, mint Store-bound
runtime-kind evidence from the exact persisted definition. The token narrows
collections and query aliases without asking callers to restate the schema in
Zod:
```ts
const tagKind = store.runtimeNodeKind("Tag", extension.nodes.Tag);
const taggedWithKind = store.runtimeEdgeKind(
"taggedWith",
extension.edges.taggedWith,
);
const tags = store.getNodeCollectionOrThrow(tagKind);
const rows = await store.query()
.fromDynamic(tagKind, "tag") // ctx.tag has the declared Tag fields
.traverseDynamic(taggedWithKind, "edge") // ctx.edge is narrowed too
.toDynamic("Document", "document")
.select((ctx) => ({ label: ctx.tag.label, weight: ctx.edge.weight }))
.execute();
```
Definitions, rather than hand-authored Zod schemas, are the type evidence:
TypeGraph compares the complete graph-extension declaration that TypeScript
infers against the definition persisted for that kind. This keeps refinements
such as enums, optionality, arrays, and numeric constraints on one authoritative
surface. Tokens are bound to the issuing Store and active schema hash; use a
fresh token after reopening or evolving a Store. Token lookups intentionally
use the throwing `getNodeCollectionOrThrow` / `getEdgeCollectionOrThrow`
variants because valid Store-issued evidence already proves that the kind
exists.
Predicate accessors on dynamic aliases use a `.field(name)`
discriminator:
- `BaseFieldAccessor` methods (`eq`, `isNull`, `in`, `notIn`) are
available directly on `field("name")`.
- Type-specific predicates sit behind a discriminator method that
asserts the field's type — `.string()` / `.number()` / `.date()` /
`.array()` / `.object()` / `.embedding()`. Each validates against the
registered Zod schema and throws `TypeError` on mismatch, so
`field("year").string()` against a number field is caught at
query-build time, not as a silent "method is undefined" later.
- `.field("missing")` throws when the property isn't on the schema.
#### Mixed typed and dynamic aliases
Typed and dynamic aliases interleave freely in one query. The
predicate accessor is resolved per alias — a typed alias keeps its
narrow `StringFieldAccessor` etc., while a dynamic alias gets `.field()`:
```ts
const rows = await store.query()
.from("Document", "d") // typed compile-time kind
.traverseDynamic("taggedWith", "e") // runtime edge
.toDynamic("Tag", "n") // runtime target
.whereNode("d", (d) => d.title.eq("the doc")) // typed: direct
.whereNode("n", (n) => n.field("label").string().eq("research")) // dynamic: discriminator
.select((ctx) => ({ doc: ctx.d, tag: ctx.n }))
.execute();
```
A typed `traverse("typedEdge", "e")` followed by `.toDynamic(target, "n")`
keeps the edge alias `e` typed — `e.role.eq(...)` works directly, no
discriminator needed. Only the dynamic-declared aliases use `.field()`.
#### Optional dynamic traversal
`optionalTraverseDynamic` is the LEFT-JOIN sibling — papers without
authors still surface, with the edge and target aliases as `undefined`:
```ts
const rows = await store.query()
.fromDynamic("Paper", "p")
.optionalTraverseDynamic("authoredBy", "a")
.toDynamic("Author", "u")
.select((ctx) => ({ paper: ctx.p, author: ctx.u, edge: ctx.a }))
.execute();
// row.author and row.edge are undefined for papers without an authoredBy edge.
```
### Search facade
The `store.search` facade — `fulltext`, `vector`, `hybrid`, and
`rebuildFulltext` — accepts any registered kind, compile-time or
runtime, with no type cast. The hit's `node` type narrows to the
concrete typed node only when the kind literal is statically known
in `Store`; extension kinds widen to the base `Node`. Misspelled
kind names throw `KindNotFoundError` at the call site instead of
returning empty results.
```ts
// Compile-time kind: hit.node.title is narrowed.
const compileTimeHits = await store.search.fulltext("Document", {
query: "climate",
limit: 10,
});
// Extension kind: same call shape, no cast. hit.node is the base
// `Node` shape since "Paper" isn't in the static `G`.
const runtimeHits = await store.search.fulltext("Paper", {
query: "attention transformer",
limit: 10,
});
```
## `store.introspect()`
`introspect()` returns a frozen snapshot of the merged schema and the
durable-state metadata the store has loaded so far. Its shape:
| Field | Type | Notes |
| ---------------------- | ------------------------------------- | ---------------------------------------------------------------------------------------------- |
| `graphId` | `string` | The graph's stable id. |
| `annotations` | `GraphAnnotations \| undefined` | Merged graph-scoped metadata from `defineGraph` and runtime extensions. |
| `kinds` | `readonly KindIntrospection[]` | Merged node kinds with `origin: "compile-time" \| "runtime"`, description, annotations, etc. |
| `edges` | `readonly EdgeIntrospection[]` | Merged edge kinds with the same origin discriminator and endpoint information. |
| `ontology` | `readonly OntologyIntrospection[]` | Ontology relations declared on either tier. |
| `deprecatedKinds` | `ReadonlySet` | Kinds flagged via `deprecateKinds(...)`. Informational, not a gate. |
| `extension` | `GraphExtension \| undefined` | The persisted graph-extension document, or `undefined` when no extensions have been committed. |
| `schemaVersion` | `number \| undefined` | Active schema version on the backend. `undefined` until the first commit. |
| `schemaHash` | `string \| undefined` | Hash of the active schema document. `undefined` under the same condition. |
```ts
const intro = store.introspect();
console.log(intro.schemaVersion); // e.g. 2
console.log(intro.annotations?.displayName);
console.log(intro.extension?.nodes?.Paper); // ExtensionNodeDef or undefined
console.log([...intro.deprecatedKinds]); // ["LegacyDocument"]
```
The `extension` field round-trips: passing it back through
`defineGraphExtension(intro.extension!)` and `evolve()` against an
empty graph reconstructs the same extension kinds.
For schema tooling that has an extension document but no Store, call
`introspectGraphExtension(extension)`. It compiles the extension through the
same TypeGraph compiler used by `evolve()` and returns `kinds` and `edges`
with JSON Schema `properties`, descriptions, annotations, and endpoint names.
The result describes only the supplied document; it has no graph ID, committed
schema version, or schema hash.
```ts
import { introspectGraphExtension } from "@nicia-ai/typegraph";
const declaration = introspectGraphExtension(extension);
console.log(declaration.kinds[0]?.properties);
```
Graph extensions may also carry graph-scoped annotations:
```ts
const extension = defineGraphExtension({
annotations: {
displayName: "Customer knowledge",
capabilities: { semanticSearch: true },
},
});
```
Annotation keys are shallow-merged. A later extension replaces the complete
value of each key it supplies; it does not recursively merge nested objects.
## Population statistics and stored-data validation
`await store.describe()` pairs the merged schema introspection with current
population statistics. It returns node and edge counts for every declared kind
plus present, explicit-null, and non-null counts (and non-null coverage) for
each directly addressable declared property:
```ts
const description = await store.describe();
const people = description.statistics.nodes.find(
(entry) => entry.kind === "Person",
);
console.log(people?.count);
console.log(
people?.properties.find((property) => property.path === "/email")?.coverage,
);
```
The schema coordinate includes the active schema version and hash when present
and a `schemaFence`. TypeGraph reads that coordinate before and after the
bounded, sequential SQL aggregate statements and refuses the result if it
changed. Node and edge properties are queried separately, and wide schemas are
split into fixed-width path batches. The database, rather than TypeGraph's
JavaScript process, computes all counts. Coverage follows ordinary nested JSON
Schema `properties`; TypeGraph intentionally does not invent population
semantics through `$ref`, unions, intersections, arrays, or conditionals.
`validateStore()` remains authoritative for those schemas. Concurrent writes
can affect different `describe()` path batches differently; the schema fence
detects schema changes, not data changes.
Use `validateStore()` to find rows that no longer satisfy a kind's current
declared Zod schema, for example after tightening a rule around existing data:
```ts
let cursor: string | undefined;
do {
const page = await store.validateStore({
entity: "node",
kind: "Person",
pageSize: 250,
...(cursor === undefined ? {} : { cursor }),
});
for (const failure of page.violations) {
console.log(failure.id, failure.path, failure.reason);
}
cursor = page.nextCursor;
} while (cursor !== undefined);
```
Undeclared properties are healthy semi-structured state and are never reported
as violations, including when the authored Zod object is strict. Each failure
names the record id, JSON-pointer path, top-level property when applicable,
Zod issue code, and reason.
`pageSize` is the number of records scanned, not a cap on violations: one record
can contribute several Zod issues. Each request performs a bounded SQL keyset
scan (`LIMIT pageSize + 1`) and reports `scannedCount`; it never materializes or
rescans the complete kind just to continue. Cursors bind the entity, kind,
schema fence, and last scanned id. A schema change throws
`StoreAnalysisCursorStaleError`. Data pages are deliberately live rather than a
claimed cross-request snapshot, so concurrent inserts, updates, and deletes can
affect later pages. Each page still reads the schema coordinate before and after
its data statement and refuses a concurrent schema flip.
Both analysis methods are current-only. They are absent from `StoreView`;
recorded/as-of population analysis is deferred until it can be backed by an
equally explicit temporal contract. Root-store calls use sequential SQL
statements plus schema bracketing, so they also work on non-interactive
transactional adapters.
Transaction callbacks expose the same methods through their pinned session.
Choose `repeatable_read` or `serializable` when every `describe()` aggregate or
every `validateStore()` page must observe one database snapshot, and consume
all validation pages before the callback returns:
```ts
await store.transaction(
async (tx) => {
const description = await tx.describe();
let cursor: string | undefined;
do {
const page = await tx.validateStore({
entity: "node",
kind: "Person",
...(cursor === undefined ? {} : { cursor }),
});
cursor = page.nextCursor;
} while (cursor !== undefined);
return description;
},
{ isolationLevel: "repeatable_read", accessMode: "read_only" },
);
```
Calling root `store.describe()` from inside a transaction callback still does
not join that transaction; use `tx.describe()` or `tx.validateStore()` for the
bound-session behavior.
## `store.materializeIndexes(options?)`
```ts
const result = await store.materializeIndexes();
// Restrict to specific compile-time or extension kinds.
const result = await store.materializeIndexes({ kinds: ["Paper"] });
const result = await store.materializeIndexes({ stopOnError: true });
```
`materializeIndexes` runs `CREATE INDEX` DDL for the indexes declared
on the merged graph and tracks per-deployment status in
`typegraph_index_materializations`. It's a separate verb from
`evolve()` because:
- DDL is **per-database**, not per-graph (two replicas of the same
`schema_doc` are still two databases — DDL has to run on each).
- Postgres uses `CREATE INDEX CONCURRENTLY` so live tables never take
an `AccessExclusiveLock`. CIC cannot run inside a transaction, which
is why `materializeIndexes` runs at the top-level backend, never
inside `transaction()`.
- Best-effort by default: per-index failures land in the result with
the captured `Error` and the loop continues. Pass
`stopOnError: true` to halt on the first failure.
The returned `MaterializeIndexesResult` has one entry per declared
index with `status: "created" | "alreadyMaterialized" | "failed" | "skipped"`.
The `skipped` status surfaces when the backend recognizes the
declaration but has no separate ANN index to build for it — e.g.
sqlite-vec (KNN lives in the `vec0` virtual table), SQLite without a
vector engine, or `embedding(dims, { indexType: "none" })` opting out
of automatic materialization.
Graph-extension-declared relational indexes use the same declaration shape as
compile-time `defineNodeIndex` / `defineEdgeIndex`, but in a
JSON-serializable form. They are persisted in `schema_doc.extension`,
re-derived on restart, and surface in `store.graph.indexes` with
`origin: "runtime"`.
### Vector indexes
Vector indexes are **auto-derived** from `embedding()` brands on both
compile-time and extension node kinds. Every top-level node field
declared with `embedding(dims, opts?)` produces one
`VectorIndexDeclaration` that flows through `materializeIndexes()`
like any relational index. No extra wiring required.
```ts
const Document = defineNode("Document", {
schema: z.object({
title: z.string(),
// Auto-derives a cosine HNSW vector index with pgvector
// defaults (m=16, ef_construction=64).
embedding: embedding(384),
}),
});
// Customize the auto-derived index by passing options at the brand.
const Image = defineNode("Image", {
schema: z.object({
embedding: embedding(512, { metric: "l2", m: 32, efConstruction: 100 }),
}),
});
// Opt out of automatic materialization while keeping the embedding.
const Manual = defineNode("Manual", {
schema: z.object({
embedding: embedding(384, { indexType: "none" }),
}),
});
```
Embeddings live in per-`(graphId, nodeKind, fieldPath)` typed tables
named `tg_vec___` (each carrying the field's
fixed dimension) — there is no single shared embeddings table. The
privileged migrator provisions each table plus a durable marker: at boot
via `createStoreWithSchema`, and for a field a runtime `evolve()`
introduces, by that `evolve()` call. The runtime hot path then asserts
the marker (never DDL), so a least-privilege role can read/write
embeddings.
On `materializeIndexes()`:
- Postgres with pgvector: emits `CREATE INDEX ... USING hnsw ...` (or
`ivfflat`) on the field's per-`(graphId, kind, field)` vector table
and reports `created`.
- SQLite with `sqlite-vec`: KNN lives in the `vec0` virtual table, so
there's no separate ANN index to build; declarations report
`skipped` (with a reason), not `failed`.
- libSQL / Turso: the DiskANN index is created via the strategy's own
DDL (`libsql_vector_idx` + `vector_top_k`).
- SQLite without a vector engine: declarations report `skipped` with a
reason indicating the backend lacks vector support.
The vector declaration's identity key within a single graph is
`(kind, fieldPath)` — v1 allows at most one vector index per
(kind, field) pair. The auto-derived deterministic declaration
name is `tg_vec_{kind}_{field}_{metric}` — clean and scannable for
inspection in `pg_indexes` and result entries. Changing the metric
requires a different declaration name and explicit
re-materialization.
Cross-graph disambiguation lives at the materialization boundary,
not in the declaration name. Vector status rows in
`typegraph_index_materializations` are keyed on the compound
`{graphId}::{declaration.name}` for both auto-derived and explicit
`VectorIndexDeclaration` entries — so two graphs reusing the same
declaration name (whether auto-derived from the same kind/field or
constructed explicitly via `defineGraph({ indexes: [...] })`) don't
collide in the status table. Each graph's `materializeIndexes()`
call creates its own physical pgvector index on that graph's
per-`(graphId, kind, field)` vector table and records its own status row.
### Fulltext indexes (out of scope for v1)
Fulltext indexes are NOT in the unified declaration channel for v1.
The fulltext table's canonical index (Postgres GIN on `tsv`, SQLite
FTS5 virtual table) is created with the table itself by
`bootstrapTables` per the active `FulltextStrategy`. Per-kind fulltext
indexes are an "advanced strategy" surface that doesn't fit the
relational-style declaration model and is reserved for future work.
### Caveats (Postgres)
- `IF NOT EXISTS` does not validate shape — only that something with
that name exists. Drift detection uses TypeGraph's recorded
signature, not PG metadata. Signature mismatch surfaces as
`failed` with a `different signature` message.
- Failed `CONCURRENTLY` builds leave invalid indexes
(`pg_index.indisvalid = false`). v1 surfaces this as a `failed`
result; the operator drops the invalid index manually before retry.
## `store.deprecateKinds(...)` / `undeprecateKinds(...)`
```ts
await store.deprecateKinds(["LegacyDocument"]);
console.log([...store.introspect().deprecatedKinds]); // ["LegacyDocument"]
await store.undeprecateKinds(["LegacyDocument"]);
```
Soft-deprecation surfaces in
`store.introspect().deprecatedKinds: ReadonlySet` for
introspection (codegen, UI tooling, lints) but does not gate reads,
writes, or queries. Bumps the schema version like any other change;
idempotent — re-deprecating an already-deprecated kind is a no-op.
Use cases:
- Codegen routes around deprecated kinds when generating new client
code.
- Lint rules flag new code that touches deprecated kinds.
- UI tooling hides deprecated kinds from picker menus.
## `store.removeKinds(...)` / `materializeRemovals()`
`removeKinds()` removes graph-extension-declared kinds from the active schema.
It is intentionally two-phase:
1. **Schema commit.** `removeKinds(names)` rewrites the persisted graph
extension without the named graph-extension kinds, cascades extension edges
and ontology relations that can no longer resolve, and commits a
new schema version with CAS.
2. **Data cleanup.** `materializeRemovals()` deletes rows for removed
node and edge kinds on the current deployment.
```ts
const withoutPaper = await evolved.removeKinds(["Paper"]);
await withoutPaper.materializeRemovals();
```
Pass `{ eager: {} }` to run cleanup inline after the schema commit:
```ts
const withoutPaper = await evolved.removeKinds(["Paper"], { eager: {} });
```
Removing an embedding field from a surviving kind orphans its
per-`(graphId, kind, field)` `tg_vec_*` table; `materializeRemovals()`
reclaims it and reports the count in
`MaterializeRemovalsResult.reclaimedVectorFields`.
For source-dependent edges, removing a node kind removes its source entry and
any target references to it. A source entry whose targets are exhausted is also
removed. The edge kind survives while another valid pair remains; it is cascaded
only when no pairs remain. For example, removing `Course` from the `assignedTo`
extension above preserves `Employee → Department`.
Removal only applies to graph-extension-declared kinds. Removing a compile-time
kind throws `RemoveCompileTimeKindError`; deploy new TypeScript code
for compile-time schema removal. Removing a graph-extension kind that is still
referenced by a compile-time edge or ontology relation throws
`KindHasReferentsError`, because TypeGraph cannot rewrite your
compiled graph for you.
## Restart parity (the load-bearing invariant)
The graph extension is the **durable source of truth**.
Every call to `evolve()` persists the merged document into
`schema_doc.extension`. On startup, `createStoreWithSchema()`
reads it back, runs the same compiler, and reconstructs identical
Zod-bearing `GraphDef`. Net: an extension kind defined via `evolve()` is
indistinguishable from a compile-time kind after restart.
Verify this in your own tests:
```ts
const [store] = await createStoreWithSchema(baseGraph, backend);
const evolved = await store.evolve(proposal);
await evolved.getNodeCollection("Paper")!.create({ title: "...", doi: "...", year: 2024 });
// Different process / different deployment / fresh store...
const [restored] = await createStoreWithSchema(baseGraph, backend);
expect(restored.registry.hasNodeType("Paper")).toBe(true);
const all = await restored.getNodeCollection("Paper")!.find({});
expect(all).toHaveLength(1);
```
## Multi-process safety
Concurrent writers compete on the `commitSchemaVersion` CAS. One wins;
the loser sees one of two errors with very different recovery semantics:
- **`StaleVersionError`** — the local view of the active version is out
of date. Routine race signal: refetch and retry.
- **`SchemaContentConflictError`** — a different writer wrote a row at
the same version with a different content hash. NOT a routine race.
Two writers tried to commit semantically different schemas at the
same version, which means one of them is operating on an inconsistent
view of the world. Surface to the operator; do not blindly retry.
Retry recipe (only catches `StaleVersionError`):
```ts
async function evolveWithRetry(
ref: StoreRef>,
extension: GraphExtension,
attempts = 3,
): Promise> {
for (let attempt = 0; attempt < attempts; attempt++) {
try {
return await ref.current.evolve(extension, { ref });
} catch (error) {
if (error instanceof StaleVersionError) {
// Refetch happens implicitly inside evolve()'s next call —
// catch-up auto-merges the persisted state into the local
// baseline, so the next attempt diffs against fresh state.
continue;
}
// SchemaContentConflictError, GraphExtensionValidationError,
// EagerMaterializationError, etc. all surface to the caller —
// they require operator intervention or different handling, not
// blind retry.
throw error;
}
}
throw new Error(`Failed to evolve after ${attempts} attempts`);
}
```
The internal `#catchUpToStored` step inside `evolve()` (and
`deprecateKinds`, `materializeIndexes`) folds the persisted graph-extension
document and deprecation set into the local baseline before computing
the next state, so a stale store applying an extension on top of an
out-of-date baseline doesn't trample another writer's progress.
## Trust boundary
When the graph extension originates from an **untrusted
source** — an LLM completion, user input, an external API — treat it
as untrusted data. Specifically:
- **Validation runs at the boundary.** `defineGraphExtension(doc)`
rejects any input that doesn't match the v1 subset
(`GraphExtensionValidationError` with per-issue paths). Don't skip
this step. If you want Result-style handling for untrusted JSON, call
`validateGraphExtension(raw, { strict: true })` and surface the
structured issues before calling `evolve()`.
- **Property types are deliberately small.** The supported set
excludes things like `bigint`, `Date`, custom Zod refinements, and
arbitrary functions. An LLM cannot inject executable code by
proposing an extension document.
- **Operator approval is your gate.** The library doesn't enforce
human-in-the-loop — your application does. Show the diff to a human
before calling `evolve()`.
- **Persisted documents are part of your data.** They're stored in
`schema_doc` along with every other schema artifact; back them up,
audit them, version-control them.
## Out of scope for v1
- **Fulltext index unification.** Vector indexes flow through the
unified channel (auto-derived from `embedding()` brands). Fulltext
is still per-strategy: the GIN / FTS5 index is created with the
fulltext table at `bootstrapTables` time. Per-kind fulltext indexes
are reserved for future work.
- **Multiple vector indexes per (kind, field).** v1 allows at most
one. To use a different metric for the same field, use a different
field name or wait for v2.
- **Hard-blocking reads/writes on deprecated kinds.** Deprecation is
informational. If you want strict enforcement, wrap collection
access yourself.
- **Auto drop+recreate on signature drift.** `materializeIndexes`
surfaces drift as a `failed` result; manual remediation is required
to avoid risky lock semantics.
## See also
- [Schema Migrations](/schema-management) — the lower-level primitives `evolve()` rides on.
- [Evolving Schemas](/schema-evolution) — recipes for compile-time schema changes.
- [Errors](/errors) — `EagerMaterializationError`, `GraphExtensionValidationError`, `StaleVersionError`, `SchemaContentConflictError`.
# Graph Merge
> Branch a TypeGraph store, let many writers edit it independently, and fold their work back into one canonical graph with deterministic entity resolution, conflict reporting, edge repointing, and provenance.
Graph Merge turns a TypeGraph store into something you can **fork, edit in
parallel, and reconcile** — the way you already fork, branch, and merge code.
Several writers (agents, importers, reviewers, background workers) each build
graph changes in isolation, and a single deterministic step folds them back into
one canonical graph: duplicate entities are resolved, edges are repointed onto
the survivors, disagreements are surfaced (never silently overwritten), and you
get a full report of what happened and who contributed it.
It ships as a core package subpath:
```typescript
import { branch, merge } from "@nicia-ai/typegraph/graph-merge";
```
Everything here is defined over ordinary TypeGraph stores, schemas, indexes,
backends, and ontology semantics — there is no separate service to run.
## What you can build
Graph Merge exists because "append everything" is the wrong default for graphs:
it produces duplicate entities and dangling relationships. With a real merge
primitive you can build:
- **Multi-agent knowledge-graph construction.** Run N extraction agents in
parallel, each on its own branch, then merge. The same real-world entity
discovered by three agents collapses to one canonical node; every agent's
edges follow it; disagreements come back as conflicts to adjudicate.
- **Parallel ETL / import reconciliation.** Ingest an EHR export, a claims
feed, and a lab feed as independent branches and reconcile them into one
patient-care graph — by exact identifier, blocking key, or fuzzy name match.
- **Master-data / entity dedup (CRM, FHIR, catalogs).** Use declared `unique`
constraints as definitional identity and similarity scoring for the rest.
- **Human-in-the-loop review queues.** `planMerge()` returns the exact proposed
write set, conflicts, and entity-resolution evidence without changing the
target. Persist that JSON artifact, review it in another process, and apply
the reviewed bytes later with `applyMergePlan()`.
- **Incremental ingestion against a live graph.** `mergeIncremental()` lets new
batches land on a target that has *advanced* since the branch was taken,
re-discovering already-committed entities instead of duplicating them.
- **Semantic deduplication.** Plug in an embedder for `vector` or `hybrid`
similarity to collapse near-duplicates that exact and trigram matching miss.
The throughline: **isolation while writing, determinism while merging, and a
report you can act on.**
## How it works
The mental model is a three-act lifecycle:
1. **`branch()`** stamps the base store's `base@V` and materializes an
isolated, independently-mutable working copy. With `revisionTracking: true`
(or `history: true`), `base@V` uses the store's durable revision anchor: a
per-graph random origin plus a monotonic clock. Validation therefore does
not fingerprint every live row or mistake a coincident revision in a
separately created store for the branch's base. Existing stores retain the
schema-and-content-fingerprint fallback. Writers edit the working copy with
the normal store API; the base is never touched.
2. Writers do whatever they want — create nodes/edges, modify inherited rows,
delete inherited rows.
3. **Plan, then apply.** `planMerge()` diffs every branch against the base and
runs a fixed planning pipeline. `applyMergePlan()` validates the serialized
artifact and its digest, checks its revision fence inside the write
transaction, then mechanically applies the already-resolved writes:
```text
stage (diff every branch)
→ generate candidates (exact unique · blocking key · similarity)
→ cluster (group nodes that are the same entity)
→ canonicalize (pick a survivor, union properties, resolve conflicts)
→ repoint + dedupe edges onto survivors
→ reconcile delete/modify and types
→ emit a revision-fenced JSON plan
→ validate + commit transactionally + build the report
```
`merge()` remains the one-call convenience wrapper over this same lifecycle;
it plans and immediately applies. If the target Store carries a reconciled
schema version, its commit acquires
and validates the normal schema-write fence before row DML. A raw target
remains outside that guarantee. PostgreSQL serialization failures are retried
automatically around the complete merge commit.
The pipeline is **deterministic by construction**: candidate sets are sorted,
clusters resolve by stable keys, and every conflict is decided on an explicit
`branchOrder` (or lexicographic branch id) — *never* wall-clock arrival. Merging
the same branches in any order yields the same committed graph and the same
normalized report. That property is what makes a merge safe to retry, cache, and
reason about.
## Quick start
Create a base store, fork one branch per writer, write to the branch stores,
then merge them back into the target.
```typescript
import { createStoreWithSchema } from "@nicia-ai/typegraph";
import { asBranchId, branch, isOk, merge, unwrap } from "@nicia-ai/typegraph/graph-merge";
const [base] = await createStoreWithSchema(graph, baseBackend, {
// Recommended for graphs that branch repeatedly or stay live while agents work.
revisionTracking: true,
});
// branch() is backend-agnostic: you supply a factory for each branch's backend.
const makeBranchBackend = async () => createFreshBackend();
const sourceA = unwrap(await branch(base, makeBranchBackend, { id: asBranchId("source-a") }));
const sourceB = unwrap(await branch(base, makeBranchBackend, { id: asBranchId("source-b") }));
await sourceA.store.nodes.Patient.create({ name: "Anna Rivera", birthDate: "1974-03-09", mrn: "MRN-001" });
await sourceB.store.nodes.Patient.create({ name: "Ana Rivera", birthDate: "1974-03-09", mrn: "MRN-001" });
const result = await merge(base, [sourceA, sourceB], {
resolve: {
Patient: {
block: (node) => node.mrn ?? node.birthDate,
similarity: { kind: "fulltext", fields: ["name"] },
threshold: 0.78,
},
},
onPropertyConflict: "flag",
branchOrder: [sourceA.id, sourceB.id],
});
if (!isOk(result)) throw result.error;
console.log(result.data.resolutions); // the two patients collapsed to one
console.log(result.data.conflicts); // the "Anna" vs "Ana" spelling disagreement
```
`branch()` returns a `Result`; `unwrap` throws on failure (or branch on
`isOk`). The default working-copy strategy clones the base through TypeGraph's
streaming interchange, so each branch gets a fresh backend from your factory
without building a graph-sized export document in memory.
## Reviewable plan/apply lifecycle
Use the two-step API when approval must happen before accepted graph truth
changes. The target must have `revisionTracking: true` or `history: true` so the
plan can carry a durable, store-specific revision fence.
```typescript
import {
applyMergePlan,
applyMergePlanInTransaction,
isOk,
planMerge,
} from "@nicia-ai/typegraph/graph-merge";
const planned = await planMerge(base, [sourceA, sourceB], {
resolve: {
Patient: {
block: (node) => node.mrn ?? node.birthDate,
similarity: { kind: "fulltext", fields: ["name"] },
threshold: 0.78,
},
},
});
if (!isOk(planned)) throw planned.error;
// Persist outside the target graph, or send to a separate review process.
const stored = JSON.stringify(planned.data);
const reviewed = JSON.parse(stored);
// Later, against the same unchanged target:
const applied = await applyMergePlan(base, reviewed);
if (!isOk(applied)) throw applied.error;
console.log(applied.data.merged); // actual committed effects
```
Planning does not mutate the target. The public plan contains only JSON-safe,
deterministically ordered data: its target/schema/revision fence, resolved write
set, review information, match evidence, and a stable content digest. It never
contains a Store, backend, `Map`, `Set`, callback, or embedder. Applying it does
not re-run blocking, candidate generation, similarity scoring, embeddings,
canonical selection, or conflict callbacks. The final report copies the
reviewed entity-resolution evidence unchanged.
The envelope is deliberately explicit: `formatVersion` selects the wire schema;
`digest` identifies its canonical content; `mode`, `target`, and `anchors` state
what was observed; `proposed` summarizes the review; `writes` is the complete
mechanical write set; and `review` holds the conflicts, resolutions, evidence,
diagnostics, warnings, and other report material known before apply.
### Apply a plan with application writes
Use `applyMergePlanInTransaction(target, tx, artifact)` when the merge, a graph
receipt or anchor, and application SQL must share one caller-owned commit. Build
`tx` by passing the native transaction to the **same target Store's**
`withRecordedTransaction()` callback. Apply the plan before any other write to
the target graph in that transaction; after it returns, the callback may make
more graph writes and the caller may run more SQL on the native handle.
```typescript
await db.transaction(async (nativeTx) => {
const { result: report, receipt } = await target.withRecordedTransaction(
nativeTx,
async (tx) => {
const applied = await applyMergePlanInTransaction(target, tx, reviewed);
await tx.nodes.MergeReceipt.create({
planDigest: reviewed.digest.value,
mergedNodes: applied.merged.nodes,
});
return applied;
},
);
await nativeTx.insert(mergeRuns).values({
planDigest: reviewed.digest.value,
recordedAt: receipt.recorded,
mergedNodes: report.merged.nodes,
});
}); // await this outer commit before reporting success
```
The adopted applier returns `Promise` and throws a typed
`MergeError` on refusal or failure. It does not open, commit, roll back, or retry
a transaction. Let the exception reject the outer callback so all merge and
application writes roll back together. Never catch it inside the transaction
and then commit. When the driver reports a retryable transaction failure, retry
the entire outer transaction, including the application writes; do not add an
inner retry or nested transaction around the merge.
PostgreSQL requires the transaction's observed isolation to be `READ COMMITTED`.
SQLite acquires its serialized writer slot before checking the plan fence. The
plan must explicitly have `persistProvenance: false`: atomic sidecar provenance
persistence is refused on this path. `includeInReport` remains supported, so
the returned report can still contain the in-memory provenance index.
On a history store, `receipt.recorded` is allocated after the
`withRecordedTransaction()` capture callback returns. The caller can persist
that anchor with application SQL on `nativeTx` before the outer commit, as the
example above does. Await the outer commit before treating the report or receipt
as durable.
The plan's `proposed` summary describes **proposed changes**. It deliberately
does not call them “merged”: `MergeReport.merged` is reserved for the actual
effects returned after a successful transaction. Coalescing and idempotent
identity operations can make actual counts differ from the proposal.
### Merge after schema evolution in one caller transaction
Prepare the evolution first, then call
`planMergeForEvolution(target, evolutionPlan, branches, options?)` outside the
write transaction. This route resolves writes against the graph produced by
the evolution plan while checking the current target's durable data and
revision fence. The serialized merge plan names the resulting schema
version/hash. If the target schema or revision changes during planning, the
planner refuses the artifact; replan outside the transaction.
Branches forked from the original baseline can merge existing kinds. To
include a newly added kind, call
`branchForEvolution(target, evolutionPlan, makeBackend)` before the caller
transaction (on PostgreSQL, pass the working-copy manager's `makeBackend`; see
[PostgreSQL table-backed working copies](#postgresql-table-backed-working-copies)),
then add data on that isolated branch. The planner accepts
branches from either one matching baseline; a mixed set of old-schema and
resulting-schema forks is refused.
Pass `{ revisionJournal: false }` as the fourth `branchForEvolution()` argument
when its working copy does not need journal-backed changed-key lineage. The
branch remains revision-tracked, and merge planning uses the portable diff
when no other lineage source is available.
```typescript
const evolutionPlan = await target.planEvolution(extension);
const futureBranch = unwrap(
await branchForEvolution(target, evolutionPlan, makeIsolatedBackend),
);
try {
await futureBranch.store.getNodeCollectionOrThrow("Tag").create({ label: "New" });
const mergePlan = unwrap(
await planMergeForEvolution(target, evolutionPlan, [futureBranch]),
);
await db.transaction(async (nativeTx) => {
const { result: report, receipt } = await target.withEvolvedTransaction(
nativeTx,
evolutionPlan,
(tx) => applyMergePlanInTransaction(target, tx, mergePlan),
);
await nativeTx.insert(mergeRuns).values({
mergedNodes: report.merged.nodes,
schemaVersion: receipt.schema.version,
});
});
} finally {
await futureBranch.close();
}
```
Apply the merge before other graph writes in the evolved callback. The
applier uses the evolved graph and checks the plan's resulting schema and
revision fences on the same caller session. Passing a merge plan for the old
schema refuses before merge mutation. Evolution's schema CAS is not treated
as a prior callback entity write. Roll back the entire native transaction on
any refusal; the schema change, merge, recorded capture, and application SQL
then roll back together. The report and receipt are provisional until the
outer commit succeeds. An adapter configured with
`schemaProvisioning: "transactional"` can provision required identity or vector
storage on the same native session before the merge callback. The default
DML-only policy refuses such requirements before the schema fence or merge
mutation. Bootstrap base storage before adopting either route; run generic
eager index maintenance separately after the outer commit.
### Candidate write sets for a planned schema
`planCandidateWriteSetForEvolution()` is the branch-free counterpart for a
bounded candidate batch. First use
`captureCandidateWriteSetTargetForEvolution(target, evolutionPlan)` when
authoring the JSON document; it records the evolution plan's resulting schema
identity rather than the currently active one. The planner stages the candidate
against that resulting graph and returns the same resulting-schema merge
artifact accepted by `withEvolvedTransaction()`.
Candidate resolution still includes the committed target as an accepted source.
Existing unique matches and property conflicts are therefore visible in the
reviewed plan before the evolution transaction begins, rather than surfacing as
late write-time failures.
```typescript
const evolutionPlan = await target.planEvolution(extension);
const writeSet = {
formatVersion: 1 as const,
sourceId: "import-batch-42",
target: captureCandidateWriteSetTargetForEvolution(target, evolutionPlan),
nodes: [{
kind: "Tag",
id: "import-batch-42:tag-1",
properties: { label: "Research" },
validFrom: "2026-01-01T00:00:00.000Z",
}],
edges: [],
};
const mergePlan = unwrap(await planCandidateWriteSetForEvolution({
target,
evolutionPlan,
makeBackend: makeIsolatedBackend,
writeSet,
}));
await db.transaction(async (nativeTx) =>
target.withEvolvedTransaction(nativeTx, evolutionPlan, (tx) =>
applyMergePlanInTransaction(target, tx, mergePlan),
),
);
```
The schema change and accepted candidate writes share the caller's one
transaction and recorded revision. If another writer changes the target while
planning, `MergePlanningStaleError` is an expected concurrency result: discard
the candidate plan, recapture the target for a new evolution plan, and replan.
For a frozen ancestor and a live destination, use the named incremental planner:
```typescript
const planned = await planMergeIncremental({
forkPoint,
target,
branches,
options,
});
if (!isOk(planned)) throw planned.error;
const applied = await applyMergePlan(target, planned.data);
```
When `target` records history, a durable branch can use its sealed recorded
fork point without keeping a second frozen Store:
```typescript
const forkPoint = created.branch.recordedForkPoint;
if (forkPoint === undefined) throw new Error("History was not captured at fork");
const planned = await planMergeIncremental({
forkPoint,
target,
branches: [created.branch],
options: { onBasePropertyConflict: "flag" },
});
```
`recordedForkPoint` is available when the source captured history at fork time;
it contains both the recorded instant and the branch's `base@V` token. The
planner reads ancestor rows from the target's recorded relations, validates the
origin, schema, and revision anchor, and enumerates only changed keys when
lineage can prove a complete delta. A missing or incompatible anchor is refused
before planning. The direct `mergeIncremental()` wrapper accepts the same fork
point. Keep the durable descriptor with the branch: reopening restores the
recorded fork point from the sealed origin.
The same target revision must still be current when the reviewed plan is
applied. If it moved during planning, planning returns
`MergePlanningStaleError` and no artifact. This is an expected retry-and-replan
outcome under concurrency: recapture the target, create a new plan, and review
its new digest before retrying. If it moved afterwards, `applyMergePlan()`
returns `StaleMergePlanError` before plan writes. Re-plan, review the new
digest and proposal, then apply the new artifact; never edit an
old plan or retry it as though it still represented the target. A successful
plan is single-use: a second or concurrent application is stale.
Persisting a plan or approval in the target graph also advances this revision.
For exact-plan approval, use external storage or a separate graph ID; writes to
that graph do not advance this target's revision. This does not provide atomic
writes across graphs, and any intervening target write still requires a fresh
plan. For candidate batches whose review records belong in the target itself,
use the durable review protocol below.
`merge()` and `mergeIncremental()` remain convenient compatibility wrappers.
They invoke the same planner and applier contiguously and return the same
`MergeReport` shape as before, now with match evidence on each resolution. Use
the wrappers when no external approval boundary is needed.
:::caution[Sensitive plans and trust]
A plan contains the complete resolved writes and may therefore contain personal,
regulated, or otherwise sensitive application data. Protect it like the source
graph: encrypt it where appropriate, restrict access, and avoid logging it. The
digest identifies the exact canonical artifact and detects accidental or
unrecorded changes. It is **not** a signature, proof of origin, authentication,
or authorization. Authenticate untrusted storage and authorize the caller before
passing a plan to `applyMergePlan()`.
:::
### Durable candidate review in the target graph
`planCandidateWriteSetReview()` separates immutable review evidence from a
revision-bound execution plan. Its `MergeReviewArtifact` retains the original
candidate write set, reviewed plan, normalized merge options, explicit policy
identity/context, and target baseline. You can persist this artifact and later
approval records in the target before calling
`revalidateCandidateWriteSetReview()` to compute a fresh execution plan.
Both review versions support candidate write sets only. They do not rebase arbitrary artifacts
from `planMerge()` or `planMergeIncremental()`.
Candidate planning on revision-tracked graphs reads existing candidate ids and
edge endpoints by key, then seeds only those rows in the transient working
copy. On identity-enabled graphs, it also follows live same-id peers and
current identity assertions from those references to a fixed point. The
planner reads peers of a candidate edge with `one` cardinality by source,
peers of a `unique` edge by its endpoint pair, and the active peer of a
`oneActive` edge by source. The active-only read checks an open `validTo` even
when `validFrom` is in the future, and does not return ended history. These reads
let the transient copy enforce the same cardinality rule as a complete clone.
On graphs with ontology relations, it also reads live nodes sharing each
candidate reference's id across kinds, so disjointness sees the same peers as a
complete clone. Ontology subtype relationships remain graph metadata.
The candidate diff and its target baseline are bounded to that dependency set and
any committed rows recalled by configured unique or index sources. Planning
still fences the target revision before and after these reads. With edge
match-identity constraints, a backend offering `findEdgesByMatchIdentity`
seeds the exact durable owners named by the candidate. A missing keyed read,
an owner excluded from the clone projection, or a target without revision
tracking uses the complete clone path. A custom backend lacking the optional
`findActiveEdgesBySourceV1` read also uses that path for `oneActive` graphs.
On the complete clone path, when the copy and target really share one serialized connection, clone export
is materialized before import, but its snapshot still holds the connection's
exclusive stream lease while it is collected. Concurrent review calls on that
resource can therefore return a merge error caused by a `ConfigurationError`
with `details.code: INTERCHANGE_SHARED_SERIALIZED_BACKEND_SNAPSHOT` or
`INTERCHANGE_SERIALIZED_IMPORT_IN_PROGRESS`. Await the whole review call before
starting another on the same serialized resource. A `pg.Pool` with more than one
connection is not one serialized resource; do not declare its pool object as
`{ mode: "shared" }` just because the working copies use that pool. See
[Serialized connections](/backend-setup#serialized-connections).
The following continues the [candidate write set example](#constraint-aware-ingestion-branches).
`Artifact`, `Decision`, and `evidence` are application-defined node/edge kinds;
`proposal` is an existing node. The target enables `history` or `revisionTracking`.
```typescript
import {
applyMergePlan,
planCandidateWriteSetReview,
revalidateCandidateWriteSetReview,
unwrap,
} from "@nicia-ai/typegraph/graph-merge";
const policy = {
id: "acceptance-policy-v1",
context: { requiredApprovals: 1, resolverVersion: "2026-09-01" },
};
const review = unwrap(
await planCandidateWriteSetReview({
target: store,
makeBackend,
writeSet,
policy,
}),
);
const artifact = await store.nodes.Artifact.create(
{ content: JSON.stringify(review) },
{ id: review.digest.value },
);
await store.edges.evidence.create(proposal, artifact, { note: "review" });
// After an authenticated reviewer approves under the application's policy:
const decision = await store.nodes.Decision.create({
approved: true,
reviewDigest: review.digest.value,
});
await store.edges.evidence.create(decision, artifact, { note: "approval" });
// Later: authenticate the stored artifact and decision, then check current
// authorization, approval validity, and policy before reusing that approval.
const persisted = await store.nodes.Artifact.getById(artifact.id);
if (persisted === undefined) throw new Error("Missing review artifact");
const checked = unwrap(
await revalidateCandidateWriteSetReview({
target: store,
makeBackend,
review: JSON.parse(persisted.content),
policy,
}),
);
if (checked.status !== "compatible") {
console.log(checked.differences);
throw new Error("Create a new review and obtain a new approval");
}
if (checked.reviewDigest.value !== decision.reviewDigest) {
throw new Error("Approval does not identify the validated review");
}
// Keep this fresh execution plan ephemeral: another target write makes it stale.
const applied = unwrap(await applyMergePlan(store, checked.plan));
// A separate commit AFTER successful apply; see the recovery boundary below.
await store.nodes.Artifact.create({
content: JSON.stringify({
reviewDigest: checked.reviewDigest,
approvalId: decision.id,
executionPlanDigest: checked.plan.digest,
executionTarget: checked.plan.target,
report: applied,
}),
});
```
All approval records and links must be committed before final revalidation.
Do not persist each replacement execution plan in the target: that repeats the
staleness cycle. Retain the original review, and use the returned `reviewDigest`
plus the fresh plan's `digest` and `target` fence to relate approval to execution.
Both review APIs return `Result<..., MergeError>`. Revalidation accepts the
persisted artifact as `unknown` and replans its retained candidate input once
target, policy, and baseline checks pass.
Supply current merge `options` and `policy` again; callbacks are never restored
from serialized data.
| Revalidation status | Meaning and next step |
| --- | --- |
| `compatible` | Includes a fresh `plan` and the original `reviewDigest`. Application policy may reuse approval; authorize the action and apply promptly. Compatibility itself grants no permission. |
| `changed` | `differences` identify changed policy/options, baseline entities/identity, or plan fields. Obtain a new review and approval. A `plan` is included only when fresh planning completed. |
| `incompatible` | The graph ID, schema identity, or revision origin differs. Approval cannot be reused for this target; resolve the mismatch and create a new review. |
Malformed/unsupported artifacts, mismatched digests, and missing required
evidence return `MergeReviewError` (`GRAPH_MERGE_REVIEW`). Existing typed planning
and constraint errors remain errors rather than compatibility statuses. A target
change during evidence capture/planning returns `MergePlanningStaleError`.
The V1 baseline is deliberately conservative:
- Every original node and edge row, including tombstones and validity metadata,
must remain unchanged. Editing an old audit record requires a new review even
when the candidate's resolved writes would be identical.
- Expected absences for candidate/write/guard references must remain absent.
Same-ID nodes of other kinds are also guarded, because they can change implicit
identity membership. Complete archival identity evidence must remain unchanged.
- Newly added rows can coexist with approval only when fresh planning produces
identical resolved writes, guards, conflicts, evidence, provenance, and other
plan content. Candidate-derived anchors and the execution digest/fence are
regenerated. There is no exemption for an “audit” kind.
For an eligible revision-tracked graph, pass
`reviewScope: "candidate"` to `planCandidateWriteSetReview()` to emit V2
candidate-scoped evidence. V2 fingerprints the candidate's node and edge ids,
edge endpoints, resolved writes, and plan guards, including expected absences
across kinds. On Operational Identity graphs it also records the reachable
identity assertion and same-id peer closure, plus assertion-ID collision
evidence. Revalidation expands that retained identity scope, rereads the
referenced rows, and replans the candidate under a new target fence. An unrelated original row may change
without invalidating V2 when it cannot affect the fresh resolved plan; V1
would report that row change. Applications whose approval policy needs the
V1 whole-graph rule should omit `reviewScope`. The review artifact records
its version and scope, so revalidation applies the rule originally reviewed.
Candidate-scoped review refuses graphs outside those eligibility rules.
On a `oneActive` graph, a custom backend must expose
`findActiveEdgesBySourceV1` for candidate-scoped review; the complete-clone
candidate planner and V1 review remain available when it does not.
On an Operational Identity graph, a custom Store runtime must also expose
endpoint-scoped and assertion-ID-scoped identity reads. Without both reads,
ordinary candidate planning uses the complete working-copy clone and V1 review
remains available; an explicit V2 candidate-scoped review request is refused.
Applicable store constraints still run during atomic application. Compatibility
does not promise that apply will succeed: new rows may introduce constraint
conflicts, and any write between revalidation and apply causes
`StaleMergePlanError`. A failed application commits no partial candidate node,
edge, or identity writes. Revalidate again after a stale refusal; require reapproval if
the result changes.
`policy.id` identifies your policy implementation; `policy.context` explicitly
records every opaque dependency that can change its decision. Include callback
and resolver versions, model/prompt versions, external configuration or data
versions, and any application state used to authorize approval reuse. Use an
empty context only when no such dependencies exist. TypeGraph captures callback
presence and serializable options, but cannot discover callback code, closure
state, external reads, or hidden application policy dependencies.
The producer must supply complete evidence, and the application must authenticate
the entire stored review and its approval. Content addressing and SHA-256 detect
content changes; anyone able to replace evidence can recompute a digest. A valid
digest is neither proof that the baseline was complete nor authorization to
reuse approval. Enforce artifact immutability and access control in your storage
or application. The review contains candidate data and an entire reviewed plan,
so protect it with the same care as graph data.
V1 review capture and revalidation read and fingerprint the complete target
graph and archival identity ledger. The artifact stores one fingerprint per
original row plus expected absences. Budget graph-sized reads and artifact
storage for V1. V2 candidate-scoped review uses bounded point and identity
closure reads for its baseline on eligible graphs.
The execution receipt above is a separate commit. If its write fails or the
process stops after apply, the merge may already be committed without a receipt.
Retain the original review and approval, and reconcile committed history and
application operation identity before repairing the receipt. Do not treat a
missing receipt as permission to replay the candidate; applying its old execution
plan is stale, and replanning is not a duplicate-execution check. To commit the
receipt atomically with the merge, create it in an `afterApply` callback as
described in [Composing application checks and writes](#composing-application-checks-and-writes).
Review revalidation alone does not add that guarantee.
See the runnable [durable merge review example](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/27-durable-merge-review.ts)
for the complete schema and lifecycle.
## Composing application checks and writes
Pass execution callbacks to `applyMergePlan()` when a reviewed candidate and
related application records must commit together:
```typescript
const result = await applyMergePlan(target, reviewedPlan, {
beforeApply: async (reads) => {
const resource = await reads.nodes.Resource.getById(resourceId);
if (resource?.owner !== "unclaimed") {
throw new Error("Resource is already claimed");
}
},
afterApply: async (tx, applied) => {
await tx.nodes.Resource.update(resourceId, { owner: "accepted" });
const decision = await tx.nodes.Decision.create({
status: "accepted",
changedNodes: applied.merged.nodes,
});
await tx.edges.decides.create(decision.id, resourceId, {});
},
});
if (!isOk(result)) throw result.error;
// The plan and application writes have now committed together.
```
`Resource`, `Decision`, and `decides` stand for types registered in your graph.
Import `MergePlanApplyOptions`, `MergePlanReadContext`, and `MergePlanApplied`
from `@nicia-ai/typegraph/graph-merge` to type reusable helpers. Callbacks are
execution options: they are not stored in the artifact or covered by its digest.
The transaction acquires the schema fence before the graph write lock, validates
schema, revision origin, and revision, and then calls `beforeApply`. This context
exposes node and edge collection reads and, for identity-enabled graphs, identity
reads. Write methods, native SQL, and a root Store are absent. Application writes
before plan application are deliberately unsupported: an uncommitted write can
change the reviewed state without changing its durable revision yet.
After the precheck, TypeGraph performs the existing plan preflight, identity
checks, and writes. `afterApply` receives a `TransactionContext` whose reads see
those writes. Its `applied.merged` contains only the plan's provisional counts;
callback writes do not contribute to the final report's merge counts. Use the
supplied contexts for every graph operation. Calling the original Store inside a
callback does not enlist it in this transaction. Do not retain a context for later
work, and await every operation before returning.
Callbacks must resolve without a value. Throw or reject to abort; returning a
value, including an `Err`, is refused with `InvalidMergeOptionsError`. A callback
rejection, stale plan, merge failure, capture-flush failure, or commit failure
rolls back the combined graph operation. Errors are converted to the outer
`Result` after rollback; ordinary application errors are retained in the cause
chain. Existing typed merge errors and constraint translation remain intact.
Only the successful outer result confirms commit.
Transaction conflicts (PostgreSQL serialization failures or deadlocks) retry the
whole transaction up to **three attempts**, including both callbacks. Every
attempt checks the fence again; an intervening committed write makes the plan
stale rather than silently rebasing it. Keep callbacks safe to repeat. Do not send
messages, call external services with side effects, or publish an outcome inside
a callback. Perform those effects after successful completion, or write an
application outbox record through `tx` for later delivery. Returned contexts and
provisional outcomes are not durable notifications.
Protection covers the target graph's transactional state and participating
TypeGraph writers using its graph fence. It does not make an application policy a
declarative constraint: every writer changing that policy's state must enforce
it, for example through its own conditional operation. It does not cover other
graphs, arbitrary SQL, or external systems. SQLite uses its writer transaction;
Composed PostgreSQL applications use read-committed isolation and the graph
write lock with or without history. The lock statement records the effective
session isolation; incompatible or unknown isolation is refused before callbacks.
Standalone revision-tracking-only applications retain serializable isolation. Unsupported
transaction capabilities are refused before callbacks. Existing session-bound
fence and recorded-capture isolation checks still apply.
With history enabled, plan and application writes share the transaction's
recorded capture and flush, producing one per-graph recorded revision. Without
history, revision tracking likewise advances for the combined transaction.
Failure leaves no live changes or recorded revision from the failed attempt.
Existing valid-time bounds, including open bounds, retain their semantics.
Optional persisted merge provenance remains separate from recorded history:
provenance records are persisted only after successful graph commit, and a
persistence failure remains a report warning. Callbacks do not receive a
post-commit provenance result. Previously committed review records in the target
still invalidate a plan's revision fence; this API does not relax plan staleness.
For a durable candidate review, revalidate the stored review first, then pass
the compatible result's fresh `plan` and these callbacks to `applyMergePlan()`.
## Scaling branches and interchange
`revisionTracking: true` is the recommended mode for long-lived, repeatedly
branched graphs. It advances one durable revision anchor inside each successful
Store write transaction. The anchor combines a per-graph random origin with the
monotonic commit clock, so a branch can only match the store that created it —
not an independent database whose clock happens to share the same timestamp. A
branch and its merge precondition then read that constant-size anchor instead of
hashing every live node and edge. Stores created with `history: true` already
have the same guarantee through their recorded-time commit clock.
On PostgreSQL, the guarantee serializes writes to the same graph with a
transaction-scoped advisory lock. That is the correct trade-off for a live graph
whose branch merges must fail closed, but it can reduce throughput and increase
write latency for a high-concurrency, single-graph workload. Partition that
workload across graphs or leave revision tracking off when the content-fingerprint
fallback is acceptable.
Turning revision tracking off does **not** turn off all serialization.
*Constrained* writes now take the same per-graph mutual exclusion regardless of
`revisionTracking` or `history`, because their check-then-write is only sound if
nothing else writes the graph in between: edge cardinality (`one`, `unique`,
`oneActive`, and the `getOrCreateByEndpoints` create and resurrect legs),
node-kind disjointness on create, and a `kindWithSubClasses` uniqueness
constraint that actually expands to more than one kind — a scope covering a
single kind probes exactly the row the uniques table's primary key then
reserves, so that key is already its fence. Everything else — an unconstrained
create, a delete, a cardinality-`many` edge — pays nothing, so the cost is
proportional to the constraints you actually declared. On PostgreSQL that
exclusion is the same transaction-scoped advisory lock; on SQLite it is the
`BEGIN IMMEDIATE` writer slot the backend already takes. A backend running
without transactions (D1, `neon-http`, or `transactionMode: "none"`) has neither
and cannot be fenced.
This unlocks:
- Many concurrent agent, importer, or review branches without base-version
validation growing with the graph.
- Large graph copies, backup/export, and transfer pipelines that keep only one
interchange batch resident at a time via `exportGraphStream()` and
`importGraphStream()`.
- A safe fast path for a live base: a branch is rejected if any tracked base
write lands before its merge commits, rather than silently merging a stale
plan.
Streaming removes the graph-sized heap spike, but a physical working copy still
copies `O(graph)` rows and snapshot merge staging still compares branch state to
the base. Bundled backends page those comparisons across declared kinds, so
unused kinds do not each cost a database statement; custom backends without the
cross-kind read retain per-kind keyset pagination. Disposable candidate clones
also skip statistics refresh. Copy-on-write logical branches and delta-only
staging remain the next larger architectural step.
Revision tracking covers writes through the Store API. Direct backend writes and
raw graph-table writes through `tx.sql` bypass the anchor, so applications using
either escape hatch must avoid them for a branchable graph or retain the default
content-fingerprint validation. On transactional backends, streaming export holds
one read-only repeatable-read transaction across nodes, edges, and identity
assertions, so every chunk belongs to one committed snapshot. A snapshot stream
cannot be piped directly into a target that writes through the same serialized
connection: the same SQLite backend, distinct wrappers sharing one better-sqlite3
handle or one local (`file:`/`:memory:`) libSQL client, a bare `pg`/neon
`Client` (a checked-out `PoolClient` included), a `pg` `Pool` capped at one
connection (`{ max: 1 }`, and equally the uncoerced string forms `{ max: "1" }`
and the legacy `{ poolSize: "1" }` that `max: process.env.PG_MAX` produces), a
postgres-js client capped at one connection (`{ max: 1 }`, `?max=1` in the URL,
or `PGMAX=1`), distinct PGlite backend wrappers sharing one in-process
connection, or Cloudflare Durable Object storage, whose transaction frame is
ambient on the storage object — materialize it first or import it into an
independent backend. Pooled connections, HTTP drivers, remote libSQL, and
separate handles on one database are deliberately not treated as serialized:
each statement gets an independent connection there,
so refusing would refuse work that succeeds. The exclusion is one **exclusive** lease
per serialized connection, not a one-time check and not a cross-kind-only rule:
at most one long-lived interchange stream of any kind holds a given connection,
so all four pairings are refused — import behind export snapshot (even through a
user-wrapped stream that no longer identifies its source backend), export
snapshot behind streaming import, export behind export, and import behind
import. Whichever long-lived stream starts second gets a typed
`ConfigurationError` instead of both hanging; its `details.code` names the
condition holding the connection and `details.requested` / `details.heldBy` name
the pairing that was refused (see
[Interchange serialized-connection guard codes](/errors#interchange-serialized-connection-guard-codes)).
Every long-lived import claims that lease, not only the chunk-streaming one:
`importGraph` holds it for the whole call and `trustedImportGraph` /
`trustedImportGraphStream` for the whole trusted session, so those APIs can throw
this `ConfigurationError` too — new in 0.46 for trusted import, which previously
threw only `TrustedImportError`. TypeGraph's branch cloner detects
the shared-client case and materializes its snapshot before importing it.
Non-transactional backends can export identity-disabled graphs without this
snapshot guarantee. Identity-enabled stores already require a transactional
backend at construction, so every identity export has the snapshot guarantee.
## Entity resolution
Resolution is configured **per node kind** in `resolve`. A kind that is omitted
merges *by id only*: its new nodes and edges are copied through, but no fuzzy
matching runs. Each configured kind composes up to three candidate sources, all
feeding one shared scorer:
| Source | What it matches | Configured by |
| ------------ | ---------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------- |
| Exact unique | Two staged nodes sharing all of a declared `unique` constraint's values — a *definitional* match that bypasses scoring | the graph's `unique` constraints |
| Blocking key | Cheap pre-grouping so similarity only compares plausibly-related nodes | `block` (staged) / `blockIndex` (vs. committed base) |
| Similarity | Fuzzy scoring of candidate pairs against a `threshold` | `similarity` + `threshold` |
```typescript
resolve: {
Patient: {
block: (node) => node.mrn ?? node.birthDate, // cheap candidate grouping
similarity: { kind: "fulltext", fields: ["name"] },
threshold: 0.78, // pairs scoring >= 0.78 merge
},
}
```
### Blocking: `block` vs `blockIndex`
Blocking bounds the otherwise-`O(n²)` pairwise comparison by only comparing
nodes that share a cheap key.
- **`block(node) => string | undefined`** is an arbitrary function over staged
nodes — a normalized email, a tenant id, a birth date, a `soundex(name)`.
Returning `undefined` puts the node in the shared *unblocked* bucket.
- **`blockIndex`** names a declared `defineNodeIndex` and is the **new-vs-base**
block key: it lets the merge query *already-committed* nodes that share a
staged node's index key and propose them as candidates. It powers incremental
ingestion (see [Snapshot vs incremental](#snapshot-vs-incremental)) and is
ignored on the snapshot `merge()` path.
```typescript
import { defineNodeIndex } from "@nicia-ai/typegraph";
const patientCohort = defineNodeIndex(Patient, { name: "patient_cohort_idx", fields: ["cohort"] });
const graph = defineGraph({ /* ... */ indexes: [patientCohort] });
// In resolve, recall committed patients in the same cohort:
resolve: { Patient: { blockIndex: "patient_cohort_idx", similarity: { kind: "fulltext", fields: ["name"] }, threshold: 0.85 } }
```
### Keyless windows
A node with no block key and no unique signature lands in the *unblocked*
bucket, which is otherwise compared all-vs-all. For large unblocked sets, set
`keyless` to switch to bounded single-pass **sorted-neighbourhood**: nodes are
sorted by their similarity text and each is compared only to its next `window`
neighbours — `O(n·window)` instead of `O(n²)`, still fully deterministic.
```typescript
resolve: {
Article: {
similarity: { kind: "fulltext", fields: ["title"] },
threshold: 0.8,
keyless: { window: 20 }, // compare each unblocked article to its 20 nearest neighbours
},
}
```
### Similarity strategies
Four strategies cover the spectrum from zero-dependency to embedding-powered:
| Strategy | Needs embedder? | Use case |
| ---------- | --------------- | ---------------------------------------------------------------------------------------------------------------- |
| `fulltext` | No | Portable in-memory Sørensen–Dice trigram score over one or more fields (e.g. `name`). The cross-backend default. |
| `custom` | No | Your own deterministic `score(a, b) => number` — domain rules, weighted field blends, edit distance. |
| `vector` | Yes | Cosine similarity over one field's embedding. Catches semantic near-duplicates. |
| `hybrid` | Yes | Blend `vector` and `fulltext` by `weights` (default 0.5 / 0.5). |
The `fulltext` scorer runs **in memory** over the staged candidate text — it
deliberately does not consult database fulltext indexes, because branch
candidates are staged working-copy rows, not indexed search results. That keeps
scoring deterministic and identical across SQLite and Postgres.
For `vector` / `hybrid`, supply an `embedder` (batched, async, deterministic —
the same text must always map to the same vector):
```typescript
const result = await merge(base, branches, {
embedder: async (texts) => texts.map((text) => embedModel(text)), // text[] -> Float32Array[]
resolve: {
Article: {
similarity: { kind: "hybrid", fields: ["title", "summary"], weights: { vector: 0.7, fulltext: 0.3 } },
threshold: 0.84,
},
},
});
```
A `vector`/`hybrid` strategy with no embedder configured fails with a typed
`SimilarityUnavailableError`, never a silent no-op.
## Conflicts
When merged contributors disagree on a property value, Graph Merge **resolves by
an explicit, deterministic policy and records what it did** — it never lets
arrival order decide.
### Property conflicts
| Policy | Behavior |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `flag` (default) | Commit the deterministic survivor value (or the committed base value, for base-vs-branch) and record a `PropertyConflict` for review. The graph still gets a value; the disagreement is surfaced rather than resolved toward another branch. |
| `lastWriteWins` | Pick the value from the highest-priority branch (earliest in `branchOrder`) — *logical* order, never wall-clock. |
| `provenanceWeighted` | Pick the value from the highest-weight branch (see `provenanceWeights`). Ties fall back to branch order. |
| function | Delegate: `(conflict) => JsonValue` lets application code decide per conflict. |
There are **two** property-conflict knobs, deliberately separate so a fuzzy
branch match can never silently overwrite committed data:
- `onPropertyConflict` — staged branch vs. staged branch.
- `onBasePropertyConflict` — committed base vs. a branch (new-vs-base merges).
Defaults to `flag` independently, and does **not** inherit `onPropertyConflict`.
`provenanceWeighted` reads per-branch trust weights you supply:
```typescript
const result = await merge(base, branches, {
onPropertyConflict: "provenanceWeighted",
provenanceWeights: new Map([
[authoritativeFeed.id, 1.0], // the system of record wins ties of value
[bestEffortAgent.id, 0.2],
]),
});
```
### Delete / modify conflicts
An inherited node or edge that one branch **deletes** while another **modifies**
is neither a pure delete nor a pure modify. `onDeleteModifyConflict` governs it
for both nodes and edges:
| Policy | Behavior |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `flag` (default) | The modification survives **and** an unresolved `DeleteModifyConflict` is recorded — a merge must never silently destroy the only branch still carrying data. |
| `deleteWins` | Honor the delete; discard the modification; record the conflict. |
| `modifyWins` | Resurrect the row; keep the modification; record the conflict. |
Independent edits to the *same* inherited row by different branches are
**three-way merged against the base**: a field only one branch changed takes
that change with no conflict; only fields multiple branches changed to differing
values become conflicts. This holds for node *and* edge properties, so disjoint
edits compose instead of clobbering each other.
## Edges follow their entities
When nodes collapse, their edges must too. After clustering, Graph Merge:
1. **Repoints** every edge endpoint onto its cluster's canonical survivor.
2. **Drops** any edge whose endpoint was finally deleted (recorded in `dropped`).
3. **Dedupes** edges that repointing brought together, as a pure set keyed by
`(from, type, to, props)` — so `x → a` and `x → b` both landing on `x → c*`
collapse to one edge.
4. **Reconciles** edges that collapse that way but disagree on properties, via
the same conflict policy as nodes — over the properties each side actually
*changed*, so an inherited row's untouched value never competes with (or
outvotes) a value some branch authored.
Steps 3 and 4 are scoped to collisions **repointing caused**: edges are grouped by
the endpoint pair they named *before* repointing, and one row per pair collapses. A
TypeGraph store is a multigraph — nothing enforces uniqueness on `(from, kind, to)`,
`create()` makes a parallel edge, and `getOrCreateByEndpoints()` is the opt-in
set-semantics accessor — so a branch that created a parallel edge merges as a
parallel edge, and a window claim lands on the row its author touched. A repointed
edge landing on endpoints that already have several parallel rows merges into one of
them; the rest keep their own properties and windows. What makes two staged edges
"the same row" is their **edge id**, not equal properties: one inherited edge
staged by several branches folds into a single write, while a branch-created edge
is a new row even when its properties happen to match an existing one's.
When such a collapse mixes an **inherited** edge with a branch-created one, the
inherited row is the one kept: a collapse rewrites the row it keeps and does not end
the rows folded into it, so writing onto the row the target already holds is what
keeps a committed edge from being left beside the row that replaced it. This mirrors
the node rule below, and it is also what the surviving edge id in `PropertyConflict`,
window resolutions, and provenance names. A collapse of branch-created edges alone
keeps the lexicographically-minimal edge id.
Inherited edges that a branch **deleted** are removed from the target, and
inherited edges **modified** by multiple branches go through the same base-aware
three-way merge as nodes — so an edge's `since` edited by one branch and `note`
edited by another keep *both* edits. The collapse in step 4 is base-aware for the
same reason: a staged copy of an inherited row carries that row's whole property
bag, and only the values it *changed* count as claims. The clearest case is a row
staged solely to carry an end-of-validity — it authored no property, so it
contributes no claim and raises no conflict, whatever its branch's rank.
## Ontology type reconciliation
With `reconcileTypes: "ontology"`, two staged nodes that share an id but carry
subtype-compatible kinds (via the graph's `subClassOf` closure) are collapsed to
the **most-specific** common type, recorded as a `TypeReconciliation`. A base
`Doctor` and a branch `SpecialistDoctor` reconcile to `SpecialistDoctor` instead
of being dropped as incompatible. The default `"off"` keeps identity strictly
`(kind, id)`.
```typescript
const graph = defineGraph({ /* ... */ ontology: [subClassOf(SpecialistDoctor, Doctor)] });
const result = await merge(base, branches, { reconcileTypes: "ontology" });
```
## Choosing the survivor
By default a cluster's canonical survivor is the member with the
lexicographically-minimal id. A committed member always wins instead, so its
committed identity and the edges already attached to it stay stable: on
new-vs-base merges that is a committed base member, and on incremental merges
it is also a node the live target committed after the fork point, such as one
an earlier branch's merge added. Override the staged-vs-staged choice with
`canonical`:
```typescript
const result = await merge(base, branches, {
canonical: (cluster) => preferGoldenSource(cluster.members), // pick which id survives
});
```
## Scaling & safety
Two guards keep a merge bounded and predictable on large or pathological inputs:
- **`maxComparisonsPerKind`** caps fuzzy comparisons per kind. On overflow,
`onComparisonCeiling` decides: `"error"` (default) fails with a typed error,
or `"mergeByIdOnly"` skips similarity for that kind (still honoring exact
unique matches) and records a warning. Tighten your `block` to shrink buckets
rather than raising the ceiling blindly.
- **`clusterMaxDiameter`** optionally splits over-broad clusters: if a cluster's
single-link diameter exceeds the bound, the weakest edges are dropped
deterministically until every sub-cluster fits. This stops a chain of
near-matches (`a~b~c~…`) from fusing genuinely-distinct entities.
```typescript
const result = await merge(base, branches, {
maxComparisonsPerKind: 50_000,
onComparisonCeiling: "mergeByIdOnly",
clusterMaxDiameter: 2,
});
```
## The merge report
`merge()` returns `Result`. The report is the
**application boundary** — show conflicts to an operator, write a review record,
persist provenance, or feed a downstream step.
```typescript
type MergeReport = {
merged: {
nodes: number;
edges: number;
identity: { asserted: number; retracted: number }; // ledger effects
};
resolutions: EntityResolution[]; // collapse membership + decisive match evidence
conflicts: PropertyConflict[]; // per-property disagreements + how they resolved
deleteModifyConflicts: DeleteModifyConflict[]; // node/edge delete-vs-modify cases
typeReconciliations: TypeReconciliation[]; // ontology kind collapses
// Node drops (deleted endpoints, incompatible members), edge drops, identity
// drops (identity:duplicate-assertion, identity:endpoints-collapsed,
// identity:retraction-target-mismatch, identity:deletion-overruled), and
// lower-bound deltas the commit cannot apply (window-not-applicable)
dropped: DroppedItem[];
// Inherited rows whose end-of-validity the merge resolved. Each entry carries
// validTo for a set/move or clearValidTo: true for a reopening.
validityEnds: ValidityEndResolution[];
baseAmbiguities: BaseAmbiguity[]; // new-vs-base matches that spanned >= 2 committed entities
provenance: ProvenanceIndex; // byBranch(id) -> { nodeIds, edgeIds }
warnings: string[]; // non-fatal advisories (ceiling skips, provenance-persist failures)
candidateDiagnostics?: CandidateDiagnostics; // bounded, opt-in scored comparisons
provenancePersisted?: { graphId: string; count: number }; // when persistProvenance ran
};
```
A typical operator loop: auto-apply when `conflicts` and
`deleteModifyConflicts` are empty; otherwise enqueue them for review alongside
`resolutions` so the reviewer sees what merged and why.
### Why two entities matched
Every multi-member `EntityResolution` has `decisiveEdges`: a deterministic
minimal connectivity witness. A resolution over N distinct `(kind, id)`
identities normally has N−1 edges. Endpoints retain both kind and id, so
same-id nodes of different kinds remain distinguishable during ontology
reconciliation.
A same-id ontology retype remains a `TypeReconciliation`, rather than creating
an id-merge resolution. Its optional `decisiveEdges` carries the accepted retype
witness without changing the meaning of the existing resolution collection.
Each edge records every candidate source that proposed the pair in stable order.
Definitional evidence names the trusted rule, such as a unique constraint, and
does not pretend the internal forced match was a perfect similarity score.
Scored evidence records the strategy descriptor, actual score, and threshold
used by the shared scorer:
```typescript
type MatchEvidence =
| {
a: { kind: string; id: string };
b: { kind: string; id: string };
sources: MatchSource[];
decision: "definitional";
}
| {
a: { kind: string; id: string };
b: { kind: string; id: string };
sources: MatchSource[];
decision: "scored";
strategy: MatchStrategy;
score: number;
threshold: number;
};
```
Built-in source metadata distinguishes block, unique, base-unique, base-index,
keyless, and ontology-retype proposals. Several sources proposing the same pair
are all retained after deduplication. Strategy metadata describes `fulltext`,
`vector`, `hybrid`, or `custom` configuration, never custom function source.
Default evidence excludes the raw compared values and rejected pairs because
those may contain PII and can make reports enormous.
Candidate diagnostics are explicit and bounded:
```typescript
const planned = await planMerge(base, branches, {
...options,
candidateDiagnostics: { limit: 1_000 },
});
```
When enabled, the report and reviewable plan include accepted and rejected
scored comparisons in canonical order. A definitional edge removed by the base
ambiguity or diameter guard is also retained with its exclusion reason, so the
final partition remains explainable. The collection also carries `total`,
`limit`, and `truncated`.
The limit is deterministic: the same candidate set produces the same retained
prefix regardless of branch, source, or backend enumeration order. Diagnostics
still omit raw compared values; join their `(kind, id)` references to application
data only in an appropriately protected evaluation environment.
## Provenance
Provenance answers *which branch contributed each merged node and edge*. A
contribution is anything a branch authored into the committed row — the
properties it staged, the modification that survived, or the end-of-validity the
merge applied.
- **Report-only (default, `provenance: true`)** — `report.provenance.byBranch(id)`
returns the `{ nodeIds, edgeIds }` that branch contributed. In-memory; it
evaporates after the call.
- **Durable (`persistProvenance: true`)** — one `{branch, sourceId} → canonical`
row per contribution is upserted into a *sidecar* graph on the target's
backend (its own namespaced tables; your domain schema is untouched). The
sidecar is opened and claimed **before** the merge commits, so a sidecar graph
id TypeGraph cannot claim refuses the whole merge and leaves the target
unmodified; only the row write itself is post-commit and best-effort, where a
transient failure surfaces as a `warnings` entry rather than a failed merge.
Re-running the same merge upserts (deterministic ids), never duplicates.
`openProvenanceStore` only ever opens a sidecar graph id it can prove it owns,
and ownership is **marker-first**: a durable `ProvenanceOwner` marker row is the
sidecar's first write of any kind, committed inside the schema fence *before*
the sidecar schema is registered. A never-seen id is free to claim only when it
holds no row in **any** per-graph table — nodes and edges, but equally
recorded-time history, the revision clock and origins, identity assertions and
their derived closure and separation, fulltext, and unique keys — because a
plain `createStore` writes rows without registering a schema, so an unregistered
id is not by itself evidence of a free namespace. Ownership is then the marker
alone, checked independently of the schema hash, because an application is free
to define the same `Provenance` shape at an unrelated id. Because the marker
comes first, the resumable interrupted state is **marker without schema** (or a
marker beside a pre-marker legacy schema): that resumes by registering or
migrating the schema. The opposite state — the exact current sidecar schema with
no marker — is one TypeGraph cannot produce, and is refused unconditionally
whatever the graph contains, empty and provenance-shaped included, since
contents an application could have written are not evidence of authorship.
**What a claim costs, on PostgreSQL.** One writer class takes neither the
per-graph fence nor the graph's active schema row: a schema-less raw
`createStore` writer, or a direct `backend.insertNode` / `insertEdge` call. At
READ COMMITTED its insert could commit between the claim's re-inspection and the
claim's own commit, leaving the marker on an id an application had just made its
own. To close that, the claim issues
`LOCK TABLE , IN SHARE ROW EXCLUSIVE MODE` inside the fence and
before the re-inspection. That mode excludes every `INSERT` / `UPDATE` /
`DELETE` on those two tables **for every graph on the database** — they are
shared tables — while still admitting readers. So while a claim runs, every node
and edge write database-wide waits.
The bound is what makes it acceptable: the lock is taken **only inside a claim**,
which happens when a sidecar is created, upgraded from the pre-marker schema, or
resumed after a crash — never on the common path, where an already-owned sidecar
opens with no fence at all. Its duration is the re-inspection's probes plus one
`INSERT`, with no caller code and no caller I/O inside it. The mode is
`SHARE ROW EXCLUSIVE` rather than plain `SHARE` because it must be
self-exclusive: two concurrent claims on different sidecar ids hold different
advisory locks, so under `SHARE` both would acquire it and then both request
`ROW EXCLUSIVE` for their own marker insert — a lock-upgrade deadlock PostgreSQL
resolves by aborting one of them. SQLite takes no such lock; `BEGIN IMMEDIATE`
already owns the engine's single writer slot.
Refusals carry the code `GRAPH_MERGE_PROVENANCE_ID_COLLISION` and one of five
`details.reason` values — `application-graph`, `empty-legacy-sidecar`,
`unupgradeable-legacy-sidecar`, `unowned-exact-schema-graph`, or
`corrupt-ownership-marker` — so the remediation matches what is actually there
instead of generic advice; a backend with no transactional schema fence refuses
an unclaimed sidecar with `GRAPH_MERGE_PROVENANCE_CLAIM_UNFENCED` (an
already-owned sidecar still opens there). Under `persistProvenance: true` both
of those arrive as a typed `InvalidMergeOptionsError` naming
`details.option: "persistProvenance"`, with the originating `ConfigurationError`
as its `cause` — see
[Merge provenance sidecar codes](/errors#merge-provenance-sidecar-codes).
Query persisted provenance back later:
```typescript
import { openProvenanceStore, readProvenance } from "@nicia-ai/typegraph/graph-merge";
const store = await openProvenanceStore(target);
const fromAgentA = await readProvenance(store, { branchId: "agent-a" }); // what did agent A contribute?
const whoMadeX = await readProvenance(store, { canonicalId: "patient-123" }); // who contributed node X?
```
Inspection tools that have a backend and graph id but not the target's
`GraphDef` can use the standalone overload:
```typescript
const store = await openProvenanceStore(backend, targetGraphId);
```
## Snapshot vs incremental
A branch is forked from a `base@V` — a token combining the base's schema hash
with the store's durable revision anchor when `revisionTracking: true` or
`history: true` is on, or a complete live-content fingerprint otherwise.
The revision anchor is namespaced by a durable per-graph origin, which
`Store.clear()` rotates. A lineage-capable untracked store whose backend
supports that origin relation also carries it beside its content fingerprint.
The two merge entry points differ in how they
treat that token.
The token is printable text, so it can be stored anywhere an application
keeps descriptors, plans, and fork points, including PostgreSQL `text` and
`jsonb` columns. Treat it as opaque: compare it whole and never parse it.
Tokens minted by releases before this format, which separated components
with a NUL character, are refused with a `BaseVersionMismatchError` whose
`details.reason` is `"legacy-token-format"`. Re-branch or re-plan from the
current target. Earlier `engine:` anchors and untracked content tokens without
the active schema version also require re-branching; they cannot match the
current target's token.
The token is printable text, so it can be stored anywhere an application
keeps descriptors, plans, and fork points, including PostgreSQL `text` and
`jsonb` columns. Treat it as opaque: compare it whole and never parse it.
Tokens minted by releases before this format, which separated components
with a NUL character, are refused with a `BaseVersionMismatchError` whose
`details.reason` is `"legacy-token-format"`. Re-branch or re-plan from the
current target.
**`merge()` is a snapshot merge.** Every branch must have forked from the
target's *current* `base@V`. If the target advanced since the branch was taken,
`merge()` returns a `BaseVersionMismatchError` rather than risk clobbering newer
data. This is the right model for "fork, do work, merge back" within one round.
**`mergeIncremental()` is a fork-point merge into a live target.** It merges
branches that forked from a frozen `forkPoint` into a `target` that may have
*moved on*. Additions are re-discovered against already-committed entities (via
`blockIndex` / unique constraints) so a re-seen entity updates the committed row
instead of duplicating it. Inherited node and edge modifications/deletions are
also propagated through the same three-way planner, with the live target kept
authoritative when it changed concurrently.
```typescript
import { mergeIncremental } from "@nicia-ai/typegraph/graph-merge";
const result = await mergeIncremental({
forkPoint, // the frozen ancestor the branches forked from
target, // the live committed graph (may have advanced)
branches,
options: {
resolve: { Patient: { blockIndex: "patient_cohort_idx", similarity: { kind: "fulltext", fields: ["name"] }, threshold: 0.85 } },
onBasePropertyConflict: "flag", // required: never overwrite a newer committed value
},
});
```
`mergeIncremental()` requires `onBasePropertyConflict: "flag"` — any other value
is refused with `InvalidMergeOptionsError` — so a stale branch value can never
overwrite a newer committed value during new-vs-base recall.
The `forkPoint` must stay **frozen for the duration of the call**: every branch
diff is computed against it, and the commit transaction re-reads its `base@V`
before applying anything, so a write landing on the fork point mid-merge is
refused with `BaseVersionMismatchError` instead of committing diffs against an
ancestor that no longer exists. Only the `target` may advance while the merge
runs.
If both the branch and the live target changed the same inherited row, the target
value/deletion wins and the conflict is reported. Both `merge()` and
`mergeIncremental()` commit **transactionally** and require a
transaction-capable target backend. Managed targets also acquire the
schema-version write fence; raw targets remain outside schema fencing. On
PostgreSQL, serialization failures from either the target-content guard or the
schema fence are retried automatically around the complete commit.
### Lineage and pruned diffs
A backend may declare a `lineage` capability: an opaque, whole-database
`revision(session)` it can report and compare, plus `changesSince(session,
revision, graphId)`, which names every node and edge of one graph that
changed (inserted, updated, deleted, or resurrected) after that revision — or
admits `{ kind: "unbounded" }` when it cannot bound the answer. Bundled SQLite
and PostgreSQL stores with `revisionTracking: true` also provide bounded
lineage through a DML journal when history capture is disabled. `lineageRevisionNow()`
mints a public anchor and `changesSince(anchor)` returns changed node and edge
keys. Use that anchor API rather than `revisionNow()`, which returns a clock
value without the graph's origin identity. The journal is installed when the
store is provisioned through `createStoreWithSchema()`, or explicitly with
`installRevisionChangesJournal(backend)` from `@nicia-ai/typegraph/schema`
under a schema owner role. Existing installations must first adopt base schema
version 4 through a privileged schema open or generated base-schema migration.
Runtime lineage checks the journal and its triggers without issuing DDL; a
revision-tracked store without history fails with `REVISION_JOURNAL_NOT_READY`
when the journal is not ready. Short-lived clones that do not need this bounded
lineage can set `revisionJournal: false`. Writes before the first anchor are
outside that anchor's range.
Node and edge inserts, updates, and deletes are recorded by database triggers.
Identity-only revisions and revisions whose write provenance is incomplete
produce `{ kind: "unbounded" }` rather than an incomplete key list. Custom
backends must provide their own lineage capability to get bounded results.
Each trigger is attached to a whole physical node, edge, or identity table; it
records every write to that table and uses `graph_id` to identify the affected
graph. On shared tables this captures writes from every graph, not only graphs
whose stores enabled the journal. Journal rows are retained per revision and
never cleaned up automatically; applications should avoid installing triggers
on shared tables unless cross-graph capture is intended, and should plan an
external retention policy that preserves every revision still used as a branch
anchor. `resolveLineage(store)`
selects backend lineage first, then captured history, then the first-party
revision journal. A lineage source is consulted only to avoid rework; it never
changes what a merge decides.
`revision()` reports `:`, never the bare clock value alone:
the durable, random per-graph revision-origin nonce
(`typegraph_revision_origins`) plus the recorded-time clock. Two
independently created stores that share a `graphId`, or the SAME store
across a `Store.clear()` boundary, can mint numerically comparable clock
values, and the origin is what keeps `changesSince` from mistaking one for
the other — a revision whose origin no longer matches the graph's LIVE
origin row is `unbounded`, regardless of what its numeric clock value is.
The recorded-relations derivation's delta is trustworthy only when EVERY
writer to the graph goes through a store that captures history — a precondition
it can partially, but not fully, enforce itself. `changesSince` proves
completeness directly rather than inferring it from a high-water mark: every
integer revision between the requested one and the graph's current clock
must carry direct evidence — a `recorded_from` or a non-sentinel
`recorded_to` — in one of the three recorded relations (nodes, edges,
identity assertions). This catches an incomplete record wherever the hole
falls, including a `revisionTracking`-only `Store` (no `history`) that
advanced the shared clock without inserting a row and was later FOLLOWED by
a capturing commit — a later capturing commit cannot retroactively supply
the missing evidence, so the gap is caught regardless of what comes after
it. What it CANNOT detect: a non-capturing writer bypassing every `Store`
entirely (a raw `GraphBackend` write, or an engine-side mutation outside
TypeGraph), which leaves no evidence to be short of. Route every writer
through a capturing `Store` if a `"keys"` delta from this source must be
exhaustive.
`session` is the connection the caller's decision is bound to — a
session-less bag could never be pinned to anything, so this one always
carries one. A caller planning outside any transaction (`branch()`'s
fork-revision capture, the pruning below) passes the root backend it holds;
a caller re-validating a content fingerprint inside an open commit transaction
reads through that transaction's own handle, so the fingerprint observes the
transaction's snapshot and establishes dependencies on the rows it covers.
**Untracked stores use a complete fingerprint.** An engine-wide revision and
node/edge-only `changesSince` result cannot fence an identity-only write. It
also cannot establish read dependencies on the graph state used in planning.
For this reason, a store without TypeGraph revision tracking fingerprints live
nodes, edges, and current identity assertions even if its backend exposes
`lineage`. Where supported, the token also carries the durable graph origin.
The commit transaction checks the origin and recomputes the fingerprint before
applying its writes. Previously minted `engine:` base tokens are retired; re-branch
from the current store rather than applying an old merge.
**`Store.clear()` rotates the revision origin.** For revision-tracked stores,
`clear()` deletes and re-mints the per-graph origin in the same transaction.
A branch forked before that clear cannot merge into the post-clear store even
when its revision clock has the same numeric value.
The origin row is also read fresh on every mint (`computeBaseVersion`,
`Store.revisionOriginNow()`), never cached on a `Store` instance. Two live
`Store` objects can legitimately observe the same graph — nothing requires
that only one `Store` ever exists per database — and only one of them runs
`clear()` at a time; a stale per-instance cache on the other would keep
minting anchors from the origin that existed before the clear, so a branch
it forks would fail every merge at commit until that `Store` happened to be
recreated. Reading fresh means a second `Store` over a graph another `Store`
just cleared sees the rotation immediately, with nothing to recreate.
**Pruning the diff.** `branch()` also records a `forkRevision` on the
returned `GraphBranch` — the fork's own `lineage.revision(session)`, read
right after the working copy is created and before any write reaches it,
with the working copy's own root backend as the session (this runs strictly
outside any transaction). For the recorded-relations source this is
origin-bearing like any other reading, so clearing and repopulating the
FORK itself to the same revision count `forkRevision` held is caught the
same way a cleared BASE store already is — there is no separate guard for
the fork side to add, because the token itself now carries the check. When
staging a branch for merge, its diff against the base is restricted to the
union of two deltas: what changed on the *fork* since `forkRevision`, and
what changed on the *base* since the anchor in its own `base@V` — instead of
enumerating every live row on both sides. A key absent from both deltas
cannot have changed since the fork point, so narrowing the read to their
union cannot miss anything the full diff would have found; it only fetches
fewer rows to compare. Pruning is a pure optimization with one rule:
whenever either side cannot supply a bounded delta, the merge falls back to
comparing every live row, exactly as it always has. That covers no
`forkRevision` (a hand-built branch, or one whose store resolved no
`lineage`); either side's `changesSince` answering `unbounded` or
REJECTING (a transient engine error never fails a merge the full diff would
have completed); and the base's own anchor failing to resolve against the
base store's lineage at all — an origin mismatch between a revision-anchored
`base` and the base store's live revision row, a revision anchor minted
before the base store's first tracked write, or an old engine anchor that
must be re-branched. Nothing about *what* a merge decides depends on
whether its diff was pruned.
## Working copies
`branch()` is backend-agnostic. The default `cloneWorkingCopyStrategy` exports
the base through TypeGraph's interchange and imports it into a fresh store on a
backend your factory provides — so it works identically across SQLite, Postgres,
and in-process PGlite, and needs no schema changes. The import is
fidelity-preserving: undeclared properties that `validateStore()` treats as
healthy semi-structured data are carried through. Stripping them would make a
later merge invent deletions against the original base.
```typescript
// Each branch gets its own in-memory SQLite backend:
import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local";
const makeBackend = async () => createLocalSqliteBackend().backend;
const fork = unwrap(await branch(base, makeBackend, { id: asBranchId("worker-1") }));
```
For a custom isolation mechanism (e.g. a future copy-on-write namespace), pass a
`WorkingCopyStrategy` as the fourth argument to `branch()` — its single `create`
method receives the base store and the `BaseVersion` `branch()` already
stamped off it, and returns an independently-mutable store over the same
graph definition.
**A branch is a data fork.** `branch()` records the clone's committed schema
`(version, hash)` at fork time, and the merge refuses (typed, as
`BaseVersionMismatchError`) any branch whose store ran a schema operation
afterwards — `evolve()`, `migrateSchema()`, or `removeKinds()` — even a
round-trip migration that restores the original document hash. Those
operations mutate rows through their own preflights, and projecting the side
effects into a merge would detach them from the schema change that caused
them. Apply schema changes to the target first (or re-fork), then merge.
### PostgreSQL table-backed working copies
`createPostgresWorkingCopyManager` allocates a private set of TypeGraph tables
in the source PostgreSQL database. It derives the table inventory and base
schema marker from TypeGraph's PostgreSQL schema contributions, copies the
source graph with fenced `INSERT ... SELECT` statements, and records ownership
in `typegraph_working_copy_allocations`. The control backend, source backend,
and backends returned by `connect` must all reach the same database, and
`control` and `connect` must run as the same role
([One database role](#one-database-role)). TypeGraph checks the allocation's
private ownership token through each connection.
The control backend must execute DDL inside its PostgreSQL transactions;
its root `executeDdl` port is not required.
```typescript
import { drizzle } from "drizzle-orm/node-postgres";
import {
createPostgresBackend,
createPostgresTables,
} from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import { createPostgresWorkingCopyManager } from "@nicia-ai/typegraph/adapters/drizzle/postgres/working-copy";
import {
asBranchId,
branchDurable,
destroyDurableBranch,
reopenDurableBranch,
unwrap,
} from "@nicia-ai/typegraph/graph-merge";
const control = createPostgresBackend(drizzle(pool));
const copies = createPostgresWorkingCopyManager({
control,
connect: (names, allocation) =>
Promise.resolve(
createPostgresBackend(drizzle(pool), {
tables: createPostgresTables(names),
...(allocation === undefined ?
{}
: { vector: allocation.vectorStrategy }),
}),
),
});
const { branch: copy, descriptor } = unwrap(
await branchDurable(sourceStore, copies.durable, {
id: asBranchId("candidate-42"),
allocationId: "candidate-allocation-42",
}),
);
await copy.close(); // Releases the connection; the tables remain.
const reopened = unwrap(
await reopenDurableBranch(graph, descriptor, copies.durable),
);
await reopened.close();
unwrap(await destroyDurableBranch(descriptor, copies.durable));
```
The same manager exposes `ephemeral` for `branch()`; closing that branch drops
its tables. `listUnsealedAllocations({ after, limit })` pages through durable
allocations awaiting seal and ephemeral allocations. These rows may still have
active owners; the ledger alone cannot identify a crashed process. After
confirming that no live branch or allocation uses a row, call
`abortAllocation(id)` to remove it. A durable branch's descriptor contains only
the allocation ID, not connection credentials. Pass `sourceTableNames` when the source backend uses
custom status table names; pass `reopenOptions` to restore process-local hooks
or query options on a later process. An external `recordedRead` binding is
refused because its relation is outside the owned table inventory. Reopen
options cannot replace the allocation's schema, recorded-read binding,
history mode, or revision-tracking mode. Pass `operations` to let
`durable.operations` commit host mutations atomically with immutable evidence;
see [Atomic operations and immutable evidence](#atomic-operations-and-immutable-evidence).
Each durable allocation also owns an evidence table (`op_evidence`) in
the allocation's schema, addressed through that schema rather than the
connection's `search_path`; destroy refuses to drop it while undelivered
evidence remains.
The source backend and every backend returned by `connect` must expose the
complete PostgreSQL `tableNames` inventory, including history, identity, and
status relations. The manager refuses missing or mismatched bindings with a
`BranchError` before cloning or opening a Store. For `ephemeral` and `durable`,
`connect` runs after the allocation tables are created, so custom callbacks may
inspect those tables; on binding failure, the manager removes the new tables and
ledger row. `makeBackend` connects earlier, before it provisions anything.
The table-backed strategy supports bundled tsvector fulltext, declared
PostgreSQL B-tree, GIN, and trigram graph indexes, and pgvector sidecars. It
builds each declared graph index on private tables under stable
allocation-scoped physical names while keeping logical index names and schema
hashes unchanged. `materializeIndexes()` can retry or repair indexes after
reopen; destroy removes their owned tables and indexes. When `connect` receives
an allocation vector strategy, pass it to `createPostgresBackend`; the strategy
assigns stable table and index names from the ledger-reserved physical prefix.
Allocation claims and all initial table and vector DDL commit together, so a
colliding or failed provision leaves no partly owned sidecars. Source vector sidecars
are copied under the same transaction locks as TypeGraph relations. The ledger
stores every relation name declared by each slot's `ownedTables()` contribution,
so destroy can remove them in reverse declaration order without a graph object.
Reopening requires the graph's vector slots and owned-relation inventory to
match the persisted allocation manifest. Older ledger rows that stored only
`tableName()` remain readable as single-relation slots. A declared vector slot
whose source sidecar is absent is refused because its contents cannot be
snapshotted exactly.
The `ephemeral` and `durable` copies have a fixed schema: `evolve`, kind
removal, and deprecation refuse before mutation. Use `makeBackend`, below, when
the working copy's schema must change. Custom fulltext strategies still need a host-level database
fork. The source and every copy connection, including durable reopen, must use
the bundled `tsvectorStrategy`: a custom strategy may own additional physical
tables whose rows cannot be copied safely from the generic contribution
inventory. A connection with fulltext disabled is refused for the same reason.
System index maintenance remains available. Source table locks cover the
entire TypeGraph relation set and vector sidecars while the SQL clone runs, so a
large clone briefly blocks writes to other graphs in the same database.
#### One database role
The manager supports one deployment shape: the `control` backend and every
session `connect` returns run as the **same PostgreSQL role**. TypeGraph reads
`current_user` on both sessions and refuses a difference with a
`ConfigurationError` whose `details.code` is `WORKING_COPY_ROLE_MISMATCH`, and
the refused allocation is not left behind.
The reason is ownership. A `control` session provisions and removes every
allocation, but the Store that opens on a connected backend issues its own DDL:
runtime-contribution markers, the revision journal and its triggers, system and
declared indexes, and vector tables an evolved graph introduces. Only a table's
owner (or a member of the owning role, or a superuser) can drop it, and the
comparison is by role name, so a `connect` role that is merely a member of
`control`'s role is refused rather than trusted. A different role would leave
the tables it creates behind on close and `abortAllocation`. The shared role
therefore needs `CREATE` on the schema.
`makeBackend` calls `connect` before it writes the ledger row or any DDL and
refuses a mismatch there, so nothing is allocated. `ephemeral` and `durable`
call `connect` after their allocation tables exist, so they refuse right after
it, before cloning or opening a Store, and remove the new allocation; a durable
reopen refuses the same way and leaves the sealed allocation untouched.
#### One schema per allocation
Every allocation lives in one schema: the `control` session's current schema when
the allocation is made, recorded in the ledger's `schema_name` column. No
`search_path` decides where an allocation's relations are created or dropped, so
a `connect` pool whose connections lead with different schemas cannot strand
tables that removal never finds.
- **Provisioning** fixes its transaction's search path to that schema before it
claims the ledger row, so the tables it creates land there whichever pooled
connection runs it, and the claim records the schema the statement itself
observed.
- **The connected backend** receives table names that carry the schema. A
backend built with `createPostgresTables(names)` over that object runs the DDL
it issues lazily (bundled tables a Store ensures on first use, fulltext and
contribution storage, schema-write transactions) with the schema leading its
search path, and the allocation's pgvector strategy names its tables and
indexes through the schema. `CREATE INDEX CONCURRENTLY` cannot run in a
transaction; it creates the index in the schema of the table it names, which
is already the allocation's. The backend's catalog probes (table, index, and
column lookups, including the recorded-time compatibility check a
`history: true` Store runs) read the allocation's schema, not the session's
current one. Extensions are database-global and create no
allocation relation, but their DDL still runs through the same DDL runner
wherever the write fence takes no lock: there, a backend built over a caller's
own transaction is subject to the same session check as any other lazy DDL
(below). Under a lock fence, a pooled backend installs the extension in its own
transaction, as before; a backend built over a caller's own transaction runs it
as a savepoint inside that transaction and makes no session check, because the
extension creates no allocation relation.
- **Refusals.** A connection whose backend was built over a *copy* of `names`
(which carries no schema) is refused with a `BranchError`. A `connect` driver
that cannot hold an interactive transaction (`drizzle-orm/neon-http`) is
refused with a `ConfigurationError`
(`ALLOCATION_SCHEMA_REQUIRES_INTERACTIVE_TRANSACTIONS`), because it cannot run
its DDL under a fixed schema. A backend built over a caller's own transaction
runs its lazy DDL and schema writes, and adopts that transaction for a schema
write, only when that session's current schema is the allocation's; otherwise
it is refused with a `ConfigurationError`
(`ALLOCATION_SCHEMA_SESSION_MISMATCH`). The caller owns that session's search
path, so it is checked rather than rewritten.
- **Removal** (`close`, `abort`, `destroy`, `abortAllocation`) searches the
catalog across every schema for relations named with the allocation's reserved
prefixes. It drops those in the recorded schema, schema-qualified in one
statement, and deletes the ledger row in the same transaction. If a drop fails
(a view that depends on an allocation table, for example) the transaction rolls
back, the row stays, and the allocation remains in `listUnsealedAllocations()`
for `abortAllocation()` once the dependency is gone. If any such relation sits
in a different schema, removal refuses with a `BranchError` that names the
schemas found and keeps the row, because deleting the row would discard the
only pointer to them. Three cases are worded differently. When the recorded
schema holds none of them, the schema was renamed or the tables moved (the
message says the relations are "not in its schema"; move the tables back or
correct the row's `schema_name` and remove again). When every relation found
elsewhere has a same-named relation in the recorded schema, it is a stale copy
left in another schema, such as a backup or restore schema (the message says
the allocation "also has relations" there; drop the copy and remove again,
since the copy blocks removal until it is gone). When some relations moved and
others stayed, for example one table moved to a backup schema while the rest
remain, the allocation is split and the relations elsewhere may be the only
copy (the message says the allocation "is split across schemas"; the
suggestion drops nothing, so move the relations back or correct the row's
`schema_name`). `details` carries `allocationId`, `schema`, `foundIn`, and
`schemas`, and `suggestion` names the recovery step. If the
allocation's relations exist nowhere (its tables were dropped entirely) there
is nothing to recover, and removal deletes the ledger row, so a crashed owner's
allocation cannot stay listed forever.
- **Ledger rows from before the schema was recorded** (written by 0.72.0) carry
no schema. They resolve through the session that removes them and reopen
without binding, and follow the same removal rule: relations found in a schema
other than the removing session's refuse removal and name that schema. `control` adds the column to an
existing ledger the first time it runs.
The connection must still be able to *resolve* the allocation's tables, so its
`search_path` must include the schema, typically `public`. A per-role `"$user"`
schema ahead of it is fine. The ledger itself lives where `control`'s session
creates it, so run `control` with one consistent `search_path`.
#### `makeBackend` for branches, candidate planning, and evolution previews
`copies.makeBackend` is a `MakeBackend`, so PostgreSQL callers no longer
hand-roll table prefixes, DDL, and cleanup. It fits every API that takes one:
`branch`, `ingestionBranch`, `planCandidateWriteSet`,
`planCandidateWriteSetReview` (including sparse staging), `branchForEvolution`,
and `planCandidateWriteSetForEvolution`.
```typescript
import { branch, branchForEvolution } from "@nicia-ai/typegraph/graph-merge";
const fork = unwrap(await branch(sourceStore, copies.makeBackend));
const preview = unwrap(
await branchForEvolution(sourceStore, evolutionPlan, copies.makeBackend),
);
```
Each call allocates a fresh allocation in the same ledger, in the `ephemeral`
state, and returns an **empty, schema-mutable** backend: the caller (or the
branch API) seeds it and may commit new kinds and fields, which the fixed-schema
`ephemeral` and `durable` copies refuse. Closing the backend drops the
allocation. While it is live it appears in `listUnsealedAllocations()`, and if
its owner crashes without closing it, `abortAllocation(id)` removes everything
it owns.
Because the graph is unknown when the backend is allocated:
- **Vector tables.** A graph that declares embeddings creates its per-field
pgvector tables after allocation, so the ledger manifest cannot list them.
Dropping an allocation therefore also removes every table in its schema whose
name starts with the allocation's reserved vector prefix. That prefix is
fixed-length and never truncated, so it cannot match another allocation's
tables. `connect` always receives the allocation vector strategy for
`makeBackend`; bind it with
`createPostgresBackend({ vector: allocation.vectorStrategy })`. A connection
that binds any other vector strategy is refused with a `BranchError`, because
it could create tables the allocation does not own. Pass `vector: false` to
opt out of vector support.
- **Graph indexes.** PostgreSQL index names are database-global, so a declared
index cannot reuse its logical name on a private table. `makeBackend` scopes
each declaration to the allocation (`gix_`) the first time the
Store's `materializeIndexes()` sees it, leaving logical names and schema
hashes unchanged and never touching the source's or another allocation's
indexes. A backend you derive from the returned one with `deriveBackend`
inherits the scoping; one you build by copying its members does not.
- **Fulltext.** The same bundled `tsvectorStrategy` requirement applies as for
the cloned copies.
`control` and `connect` must run as the same role
([One database role](#one-database-role)). Both must also use the allocation's
schema ([One schema per allocation](#one-schema-per-allocation)); a pooled
connection's own `search_path` does not decide where anything is created.
### Forked working copies
A second bundled strategy, `forkedWorkingCopyStrategy({ fork, connect })`,
targets a fork-capable host instead of a streamed-interchange clone: `fork`
asks the host itself to produce a complete, independent copy of the database
`baseStore` is on, and `connect` opens a backend on that copy.
```typescript
import {
asBranchId,
branch,
forkedWorkingCopyStrategy,
unwrap,
type ForkHandle,
} from "@nicia-ai/typegraph/graph-merge";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import { decorateBackend } from "@nicia-ai/typegraph/backend";
import { drizzle } from "drizzle-orm/node-postgres";
import { Pool } from "pg";
// A host whose fork call returns a new connection string for the branch —
// this is the shape of the copy-on-write branching APIs some Postgres hosts
// offer (Neon and Supabase branches, for example), without either SDK.
type HostBranch = ForkHandle & Readonly<{ connectionString: string }>;
const strategy = forkedWorkingCopyStrategy({
fork: async () => {
const created = await hostBranchApi.createBranch(baseDatabaseId);
return {
connectionString: created.connectionString,
dispose: async () => hostBranchApi.deleteBranch(created.id),
};
},
connect: async (fork) => {
// `createPostgresBackend` takes a Drizzle database, not a pool — open
// one here. Its `close()` deliberately does not end a caller-owned pool
// (Drizzle leaves connection lifecycle to the caller), so compose the
// pool's own shutdown into this fork's `close` through the public
// `decorateBackend` (never a spread) — `branch()`'s composed close then
// ends the pool along with releasing the fork.
const pool = new Pool({ connectionString: fork.connectionString });
const backend = createPostgresBackend(drizzle(pool));
return decorateBackend(backend, {
close: async () => {
await backend.close();
await pool.end();
},
});
},
});
// `makeBackend` is ignored once an explicit strategy is supplied — pass a
// factory whose only job is to reject if it is ever called by mistake.
const rejectMakeBackend = () =>
Promise.reject(new Error("makeBackend must not be called"));
const worker = unwrap(
await branch(
base,
rejectMakeBackend,
{ id: asBranchId("worker-1") },
strategy,
),
);
// ... write on worker.store, plan and apply the merge ...
await worker.close();
```
`TFork` must extend `ForkHandle` (`{ dispose?: () => Promise }`).
`forkedWorkingCopyStrategy` supplies ephemeral copies only. Its base-version
comparison checks the graph's schema and revision or live-content anchor;
the host fork must preserve the full physical database, including TypeGraph
sidecars and extensions. A durable host strategy must persist its branch ID
and attest the sealed origin when reopening it.
For a hosted PostgreSQL branch such as [Neon](https://neon.com/docs/get-started-with-neon/workflow-primer),
`connect` must use that branch's connection string and compute endpoint for
every pooled checkout and transaction. Reusing the source pool can appear to
pass a base-version check while writing to the source. Doltgres can pin a
connection through a [database revision specifier](https://www.doltgres.com/docs/reference/version-control/branches/);
avoid session-level branch switching on a pool whose checkouts may retain
different branch state. Doltgres exposes native branch and merge commands,
but TypeGraph continues to use its own merge planner and apply path; native
merge and Doltgres backend support require separate conformance testing.
`create()` calls `fork(baseStore)`, then `connect(fork)`; the connected
backend's `close` is composed with the fork's `dispose` through `deriveBackend`
(never a spread), so `worker.close()` — the branch's public release call —
releases both the connection and the fork. A `connect` failure disposes the fork before
rethrowing, leaving nothing open and the base untouched.
A fork inherits the base's WHOLE construction option set — hooks, upsert
coalescing, the SQL schema (custom table names), the auto-refresh-statistics
threshold, query defaults, and an externally-bound recorded-read relation —
read once through `Store.workingCopyOptions`, plus `history`/
`revisionTracking`, matched to the base's own `historyEnabled`/
`revisionTrackingEnabled`. This is safe precisely because a fork is the SAME
physical database as the base: a custom `schema` names relations the fork
carries too, and an external `recordedRead` binding points at one. The clone
strategy inherits only `revisionTracking` — its fresh backend is a distinct,
empty database, so a schema naming the base's tables or a `recordedRead`
binding populated nowhere on the clone would misdirect it.
Because the fork's store reads and writes through the base's table names,
`connect()`'s backend must bind those SAME names. `create()` compares the
connected backend's own table bindings against the base's own resolved SQL
schema (`Store.revisionSchema` — the base's explicit `schema` option, or its
backend's own `tableNames` otherwise), and refuses with a `BranchError`,
closing the backend first, when they disagree: a backend bound to different
(often just the default) table names would read and write through tables the
fork's rows were never written to.
**A fork preserves what a clone drops, and that is why it is safe to merge.**
The clone strategy above streams the base through public interchange with
`includeDeleted: false`, so it omits every soft-deleted row entirely: the
interchange `meta` schema has no `deletedAt` field, so a tombstoned row would
otherwise round-trip as LIVE and read as a spurious resurrection on the
clone's diff. It also regenerates `created_at`/`updated_at` on import — safe
only because the merge's state diff always compares against the *original*
base store, never the clone. A fork is never rebuilt through
`exportGraphStream`/`importGraphStream`, so none of that applies: tombstones,
`created_at`/`updated_at`, and the `version` column carry over unchanged, and
— with `history: true` — the fork physically carries the base's recorded
relations, so `store.asOfRecorded()` answers from
that history. A clone-based branch never enables history, so the same call on
it refuses outright.
`create()` asserts `computeBaseVersion(forkStore) === base` right after
attaching the store, where `base` is the token `branch()` already stamped off
the ORIGINAL base store before invoking the strategy — cheap when the base has
revision tracking (an O(1) anchor compare), an O(graph) content fingerprint
otherwise, and computed exactly once either way. This proves base-token
equality at the instant the fork was taken, not byte-for-byte physical
identity: the untracked fingerprint deliberately omits tombstones,
`created_at`/`updated_at`, the `version` column, and recorded history (the "A
fork preserves what a clone drops" paragraph above) — providing those
unchanged is the FORK MECHANISM's job, not something this assertion re-verifies
on every branch. That is still the right fence: the merge's lost-update guard
reads `version` and the diff reads tombstones/timestamps straight off the
fork, so a `fork` that is not a true physical copy breaks them regardless of
what the content fingerprint agrees on. A mismatch closes the backend first
and refuses with a `BranchError` carrying `forkVersion`/`baseVersion` in
`error.details`; `branch()` catches it and returns that `BranchError` as the
`cause` of the outer `BranchError` it resolves with. Only a base-token
mismatch is refused here — a fork taken while the base was mid-write, or a
`fork` that returns a different graph; divergence confined to the physical
state the token omits (tombstones, timestamps, row versions, recorded
history) passes the fence, and keeping that state faithful remains the fork
mechanism's contract.
`create()` also refuses BEFORE ever attaching a store when `connect()`'s
backend aliases the base's own backend: the same backend object, one derived
from the other through `deriveBackend`, or two wrappers sharing one underlying
connection. Without this check, a `connect()` that mistakenly hands back the
base's own backend (a cached factory keyed by database name, say) would pass
every fence below trivially — every write on the "fork" would actually mutate
the base, and closing the working copy would close the base's own backend. The
refusal disposes only the fork (never the aliased backend, which the base
still owns) and throws a `BranchError` naming `connect()`. This cannot detect
every aliasing shape: a fresh backend built over the base's own connection
pool is indistinguishable from a real fork's connection when that pool audits
as independent (the normal case for a default-size `pg.Pool`) — a pooled
checkout genuinely is a different connection from the pool's perspective.
`ingestionBranch()` stays clone-based. Its strategy derives a working-copy
schema with node uniqueness deferred so an untrusted batch's repeated keys can
reach entity resolution before validation; a host-level fork carries the
base's schema exactly, uniqueness included, with no hook to relax it.
:::caution[Suspend hazard]
A fork-capable host that suspends idle compute to reclaim it between requests
drops that compute's in-process state, including anything memoized against a
particular connection or session. TypeGraph's own locking already assumes
this rather than trusting a lock survives idle time: the recorded-write lock
memo (`RecordedGraphLockMemo`, populated by
`memoizeAcquiredRecordedGraphWriteLock`) and the schema-fence lease
(`memoizeLeasedSchemaFence`) are both keyed weakly by the transaction-scoped
backend object, so they hold for exactly one transaction's lifetime and
re-acquire on the next one, and the write fence itself (see
[Write fence declaration](/backend-setup#write-fence-declaration-writefence))
is resolved and its lock taken fresh per transaction, never cached across
one. An ordinary sequence of separate `store` calls — each its own
transaction — therefore tolerates a suspend between any two of them.
What does NOT tolerate a suspend is a single `store.transaction` callback:
every read and write the callback issues, and the lock it holds, runs on one
native database transaction over one connection, so a suspend partway
through drops that connection out from under the callback and aborts
whatever was in flight. Keep a `store.transaction` callback's wall-clock
duration short and free of anything that could let the host suspend
underneath it — an external API call, a human approval step, a long queue
wait — and commit a long-running workflow across multiple `store.transaction`
calls instead of holding one open across such a wait.
:::
### Durable host-native branches
`branchDurable()` is the persistent counterpart to `branch()`. A
`DurableWorkingCopyStrategy` allocates a host branch, opens a Store on it, and
returns a non-secret JSON locator. TypeGraph seals the immutable fork origin
beside that allocation and returns a `DurableBranchDescriptor` that can cross a
queue, process, deployment, or machine boundary.
For a remote host, persist a chosen `{ id, allocationId }` before calling
`branchDurable(base, strategy, { id, allocationId })`. `create()` receives both
and must refuse an allocation ID that may already exist. If the host allocates
a branch but its response is lost, use host tooling to inspect the ID and
recover or remove the allocation before retrying. A failed create reports both
IDs for that reconciliation. The host must never allocate a second physical
copy for the same ID or return a sealed copy as though it were new.
```typescript
import {
applyDurableMergePlan,
branchDurable,
destroyDurableBranch,
planMerge,
reopenDurableBranch,
unwrap,
} from "@nicia-ai/typegraph/graph-merge";
const created = unwrap(await branchDurable(base, durableStrategy));
await created.branch.store.nodes.Person.create({ name: "Ada" });
// Releases this process's connection and writer lease. The host branch stays.
await created.branch.close();
await queue.put(JSON.stringify(created.descriptor));
// A later process reconstructs the ordinary GraphBranch used by planning.
const descriptor = JSON.parse(await queue.get()) as typeof created.descriptor;
const reopened = unwrap(
await reopenDurableBranch(graph, descriptor, durableStrategy),
);
const plan = unwrap(await planMerge(base, [reopened]));
// Applies the complete TypeGraph plan inside the target transaction.
const report = unwrap(
await applyDurableMergePlan({
target: base,
branch: reopened,
descriptor,
strategy: durableStrategy,
plan,
}),
);
await reopened.close();
unwrap(await destroyDurableBranch(descriptor, durableStrategy));
```
Closing and destroying are deliberately separate. `GraphBranch.close()` closes
the backend and releases its access lease, but leaves the persistent allocation
reopenable. `destroyDurableBranch()` asks the strategy to attest the complete
origin and delete or archive that allocation atomically. A descriptor is
untrusted input: TypeGraph checks its allocation id, graph definition, branch
id, base token, schema anchor, and engine revision against the origin the host
sealed. The allocation id is independent of the caller's branch id, so
swapping or relabeling a locator cannot authorize deletion of another copy
even when two copies were given the same branch id.
Strategies write new locators using `version` and may list older supported
locator versions in `readableVersions`. Every method must understand each
listed version, including destroy and evidence access.
The strategy locator must be JSON-safe and **must not contain secrets**. Use a
branch id, database id, or other lookup key, then resolve credentials from
strategy-owned configuration. TypeGraph returns the locator to application code
so a connection URL, password, or bearer token placed there can escape through
ordinary descriptor storage. Framework cleanup errors deliberately omit the
locator and raw host cleanup error from diagnostic details.
#### Exact forks and access leases
After `strategy.create()` returns, TypeGraph recomputes `base@V` from the source.
A source write racing allocation therefore refuses and aborts the working copy
instead of sealing a branch from the wrong ancestor. TypeGraph then accepts an
exact matching working-copy token as the fast path. When a strategy creates an
equivalent persistent copy with an independent revision namespace, TypeGraph
instead verifies that its complete merge-visible graph state has no delta from
the source, fencing the source again after enumeration. The host remains
responsible for physical fidelity outside TypeGraph's graph semantics.
To enable lineage-pruned merge diffs, `create()` may return `forkRevision`
captured atomically with the physical fork. When it cannot prove that cut, omit
the revision and TypeGraph compares the complete graph state; reading a later
revision after the copy was opened could miss an intervening branch write.
Every `create()` and `reopen()` also returns a `DurableWorkingCopyAccess`:
- `engine-fenced` says the database provides sound cross-client isolation and
change fencing for the full Store planning/apply access pattern, across every
connection and process that could mutate the working copy.
- `exclusive` carries an allocation-wide writer lease. The strategy must acquire
it before returning and exclude every other process and backend instance.
TypeGraph closes the backend first, then releases the lease; a failed release
is retried by the next `close()` call.
Do not use `engine-fenced` merely because one backend object serializes its own
calls. A `caller-serialized` backend owns one in-memory queue per backend
instance, so two reopened pools or two processes still race. Such an engine must
use a host-wide `exclusive` lease, and a concurrent reopen must wait or refuse.
Merge planning also assumes the working copy is quiescent while it is diffed.
#### Native database branches
A strategy may allocate a working copy using a database-native branch, but
`applyDurableMergePlan()` always applies the approved TypeGraph plan through
the target Store transaction. The former native-merge callback was removed:
it could commit outside the transaction that checked the target revision.
A future native merge capability needs a host-native compare-and-swap on the
actual target, plus proof that the full physical diff equals the approved
TypeGraph writes, including schema, history, identity, and sidecars.
For a Doltgres strategy, pin each Store connection to the intended database
branch. [Doltgres revision specifiers](https://www.doltgres.com/docs/reference/version-control/branches/)
provide that connection-level selection. Its
[`DOLT_BRANCH()` and `DOLT_MERGE()` functions](https://www.doltgres.com/docs/reference/version-control/dolt-sql-functions/)
implicitly commit the current transaction, so a fence checked before those
functions cannot by itself protect their target.
#### Atomic operations and immutable evidence
A `DurableWorkingCopyStrategy` may also expose an optional `operations`
capability (`DurableOperationCapability`). It lets a durable host combine one
opaque graph mutation with its immutable operation evidence in a **single host
transaction**.
TypeGraph owns descriptor validation, sealed-origin attestation, request
canonicalization, and evidence validation; the host owns the database mechanics.
```typescript
import {
durableBranchHasUndeliveredEvidence,
getDurableOperation,
markDurableOperationDelivered,
operateDurableBranch,
scanDurableOperations,
unwrap,
} from "@nicia-ai/typegraph/graph-merge";
const request = {
idempotencyKey: "statement-42",
// Host-defined, JSON-safe description of the graph change to apply.
mutation: { kind: "statement", op: "upsert", payload: { subject: "s-1" } },
// Host evidence, retained verbatim. TypeGraph never interprets either field.
metadata: { source: "etl", schemaVersion: 3 },
};
const outcome = unwrap(
await operateDurableBranch(descriptor, durableStrategy, request),
);
if (outcome.outcome === "unsupported") {
// The strategy applied no mutation and wrote no evidence; TypeGraph refuses
// rather than emulating atomicity with best effort or callbacks that run
// outside the evidence transaction.
throw new Error(`Missing capabilities: ${outcome.dimensions.join(", ")}`);
}
console.log(outcome.outcome); // "applied" | "replayed"
// Newly applied evidence is always false. A replay returns the current
// committed delivery state, which may already be true.
console.log(outcome.evidence.delivered);
```
Both `mutation` and `metadata` are **JSON-safe host values**. TypeGraph never
interprets their application fields; it canonicalizes `metadata` plus `mutation`
into the `operationDigest` and otherwise carries them through untouched. The
digest covers the complete request except the idempotency key, so reusing a key
with a different mutation *or* different metadata conflicts. Non-JSON content is
refused before any host call.
`metadata` is retained as evidence; `mutation` is the host's own description of
the graph change it must apply atomically with the evidence row. The strategy
attests the caller's `expectedOrigin` against the allocation the descriptor
names, exactly as reopen and destroy do. Every committed
operation returns `before`/`after` coordinates — the merge-visible `base`
fingerprint and, when the working copy resolves lineage, the engine `revision`.
TypeGraph validates that the returned evidence echoes the canonical request and
digest; a host cannot forge a different digest, echo a different request, or
return non-JSON metadata (`DurableOperationEvidenceError`).
**Idempotency.** The strategy treats `idempotencyKey` as its unique key:
- Identical key **and** digest: returns the previously committed evidence
(`outcome: "replayed"`) and re-applies nothing. Because delivery marking is
monotonic, a replay after delivery legitimately returns `delivered: true`.
- Identical key with a **different** digest: refuses with
`DurableOperationConflictError` and mutates nothing.
A first application (`outcome: "applied"`) must return `delivered: false`.
TypeGraph rejects `applied` evidence that is already delivered, so a host cannot
bypass downstream delivery or the destroy fence. It also validates the complete
host outcome envelope: malformed outcomes and empty, duplicate, or unknown
`unsupported` dimensions return `DurableOperationEvidenceError`.
**Evidence access and delivery.**
- `getDurableOperation(descriptor, strategy, idempotencyKey)` reads one
operation's evidence, or `undefined` when it was never committed.
- `scanDurableOperations(descriptor, strategy, { after?, limit? })` returns
`{ operations, cursor, hasMore }` in monotonic commit order, with ties broken
deterministically. Pass the opaque `cursor` back as `after` to resume, even
after `hasMore: false`; later commits must sort after that cursor. An empty
page echoes `after`, and only an empty initial scan omits `cursor`. `limit` defaults to
`DURABLE_OPERATION_SCAN_DEFAULT_LIMIT` (100) and may not exceed
`DURABLE_OPERATION_SCAN_MAX_LIMIT` (1000); a larger page is refused.
- `markDurableOperationDelivered(descriptor, strategy, idempotencyKey)` marks
one operation delivered, idempotently: marking an already-delivered operation
returns the same evidence and writes nothing, and an unknown key returns
`undefined`.
- `durableBranchHasUndeliveredEvidence(descriptor, strategy)` reports whether
any committed evidence is still undelivered — the queryable half of the
destroy fence below.
`operateDurableBranch()` is the only orchestrator that tolerates a missing
capability: a strategy with no `operations` returns the explicit `unsupported`
outcome (`dimensions: ["atomicMutation"]`) having executed no host call. `get`,
`scan`, `markDelivered`, and `hasUndelivered` instead refuse with a typed
`DurableOperationUnsupportedError`. TypeGraph never emulates the atomic
guarantee: a callback that runs inside the strategy's own evidence transaction
(as `apply` does in the bundled PostgreSQL manager below) is the host's atomic
mutation, while best effort or a callback outside that transaction is refused.
**Destroy fence.** A strategy with `operations` MUST refuse destruction while
undelivered evidence remains, throwing `DurableEvidenceUndeliveredError`;
`destroyDurableBranch()` preserves that typed refusal instead of flattening it
into a generic branch failure, so the caller can still recover the evidence.
Deliver (or archive) the outstanding evidence before destroying the branch.
Concurrent `operate` and `destroy` are serialized by the host's own transaction:
either the operation commits first (destroy then observes undelivered evidence
and refuses) or destroy commits first (the operation fails against the removed
allocation). No partial state is ever observable.
##### Bundled PostgreSQL manager
`createPostgresWorkingCopyManager` implements the capability when given an
`operations` option. `apply` is how the host's opaque mutation reaches the
graph; TypeGraph still never interprets `mutation`.
```typescript
const copies = createPostgresWorkingCopyManager({
control,
connect,
operations: {
graph,
// Runs inside the transaction that commits the evidence row. A throw rolls
// back both the mutation and the evidence.
apply: async (transaction, mutation) => {
await applyHostMutation(transaction, mutation);
},
},
});
const outcome = unwrap(
await operateDurableBranch(descriptor, copies.durable, request),
);
```
`operations.graph` is required because a capability member receives only the
descriptor, so the manager must reopen the allocation from the graph the host
names. Before any connection or transaction opens, every member checks that
graph against the sealed allocation's attested origin: its graph id and its
version-blind definition hash must equal the ones the branch was forked with, so
a graph that reuses the id with a different definition is refused. `apply`
receives the transaction-scoped context of the allocation's fixed-schema Store,
the same context `store.transaction` provides, so the allocation's fixed schema
applies. Without the option, `copies.durable.operations` is undefined and
`operateDurableBranch()` returns `unsupported` (`atomicMutation`).
Each durable allocation owns one evidence relation under its ledger-reserved
physical prefix, created in the provisioning transaction and dropped by destroy.
The ledger records whether an allocation has one (`operation_evidence`).
`operate` takes the allocation lock on the allocation's own transaction session,
attests the sealed origin, resolves idempotency, takes the graph write lock,
computes the `before` coordinates, calls `apply`, computes the `after`
coordinates once the transaction's revision bookkeeping has run, and inserts
undelivered evidence, all in one transaction. The allocation lock is a
transaction-scoped advisory lock keyed on the allocation id, in a namespace of
its own so it can never collide with a graph's write lock. It serializes
operations per allocation, so the evidence sequence that backs the opaque scan
cursor is commit order, and each operation's `before` equals the previous
operation's `after` whenever every writer to the allocation goes through
`operate` or takes the graph write lock. Ordinary writes take that lock on an
allocation that tracks history or revisions, so a direct write cannot commit
between `before` and `apply`; on an allocation that tracks neither, a direct
write is not fenced and the evidence's `before`/`after` pair may include it.
The graph write lock is graph-wide. While `apply` runs, tracked writes to the
source graph and to every sibling working copy of it wait on that lock, so keep
`apply` short and do not wait on other graph writers inside it.
Coordinates always carry `base`. They also carry `revision`, the engine
revision, when the allocation resolves lineage, which is when it tracks history
or revisions; both are read on the transaction's own session so they describe
one state. An allocation that tracks neither reports no `revision`, and its
`base` values are content fingerprints, which read the whole graph twice per
operation.
**Isolation is observed, not assumed.** `operate`, `markDelivered`, and destroy
each request READ COMMITTED, and the statement that takes the allocation lock
also reports the isolation level its session actually runs at. Any other level
is refused before anything is read or written, with a `ConfigurationError` whose
`details.code` is `WORKING_COPY_ISOLATION_UNSUPPORTED`, because the request is
honored only where a backend supports it and a role or server default of
REPEATABLE READ would otherwise give the fence and the idempotency lookup a
snapshot older than the lock wait. A `control` or `connect` wrapper must
therefore forward the transaction `isolationLevel` option. The refusal only
fires when a wrapper drops the requested option and the session's default is
not READ COMMITTED.
The same check runs everywhere the manager drops an allocation, not only in
destroy and `abortAllocation`: closing an ephemeral working-copy store, closing
a `makeBackend` backend, and the cleanup after a failed allocation. The first
two surface the refusal from `close()`. The cleanup swallows it so the
allocation's original failure reaches the caller, which leaves the allocation
behind. Every such orphan is discoverable with `listUnsealedAllocations` and is
removed by `abortAllocation` once `control` forwards the option.
**Destroy fence.** Destroy (and `abortAllocation`) takes the same allocation
lock. `destroyDurableBranch()` refuses with `DurableEvidenceUndeliveredError`
while undelivered evidence exists, even from a manager built without
`operations`; delivering the evidence requires a manager built with
`operations`. An in-flight `operate` and a destroy on one allocation serialize
on the lock: whichever commits first decides the other's outcome. The destroy
waits at most `cleanupLockTimeoutMs` (5000 ms by default); one that outwaits a
long `apply` fails with the database's lock timeout having committed nothing,
and can be retried after the operation settles. `get`, `scan`, and
`hasUndelivered` take no allocation lock, so they never wait behind an `apply`.
A destroy that commits after any member has attested the sealed row but before
that member holds the allocation (before its connection is attested, before
`operate` mints the revision origin, or before `get`, `scan`, or `hasUndelivered`
reads the evidence relation) fails the member with one `BranchError`
(`changed owner or was destroyed during the operation`), the same error
`operate` and `markDelivered` raise against a removed allocation. It is never a
raw missing-relation error or the "connection is not bound to the allocation
database" refusal, which is reserved for a connection that reaches a different
database than the one the ledger names.
The fence follows the manager's [removal rule](#one-schema-per-allocation). It
reads the evidence relation only in the allocation's own schema. Evidence
relations that sit in another schema refuse removal before the fence runs and
are kept. The fence runs only when the evidence relation is among the
relations removal drops. An allocation whose evidence relation is gone has no
evidence left to deliver and nothing to recover, so destroy removes its
remaining relations and its ledger row, exactly as it does for an allocation
whose relations exist nowhere.
**Mixed-version deployments.** Only managers on this version take the allocation
lock and honor the destroy fence. A manager from an earlier release that shares
the ledger destroys an allocation without consulting its evidence, so
undelivered evidence is lost with the allocation, and it does not drop the
evidence relation, so a later `allocate` with the same id refuses because
`op_evidence` exists without a ledger row. Upgrade every process that
shares a working-copy ledger before any of them creates or destroys a durable
allocation. To recover an orphaned evidence relation, read its undelivered rows
(`WHERE NOT delivered`) and deliver them, then drop the relation the refusal
names and retry. TypeGraph never drops it for you, because it may hold the only
copy of undelivered evidence.
An allocation provisioned by an earlier release has no evidence relation, and
its ledger row says so without any statement that changes the database.
`operate` returns `unsupported` with `dimensions: ["evidenceStore"]`. The only
statement it runs is one read-only ledger `SELECT` through `control`; it runs no
DDL, takes no lock, calls no `connect`, and applies and writes nothing. The read
members report no evidence: `get` and `markDelivered` return `undefined`, `scan`
returns an empty page (echoing `after`), and `hasUndelivered` returns `false`.
Re-fork the branch to gain evidence.
### Constraint-aware ingestion branches
For a bounded candidate batch, `planCandidateWriteSet()` hides the transient
branch lifecycle completely. It accepts a validated, versioned JSON document,
stages it through the same constraint-aware ingestion implementation, delegates
to incremental merge planning, and closes the working copy on every outcome.
The result is the ordinary `MergePlanArtifact`, so review and application use
the same APIs as every other merge plan.
On eligible revision-tracked graphs, planning
seeds existing candidate rows, edge
endpoints, cardinality peers, live same-id ontology peers, and any reachable
current identity component into the disposable working copy.
The resolver still queries the live target for declared unique and index peers,
and the plan retains its ordinary provenance, conflicts, digest, and commit-time
fences. Existing undeclared target properties survive staging; extra candidate
properties are refused. A custom backend without the active-only source read
uses the complete clone path for `oneActive` graphs. A custom backend without
the keyed match-identity owner read, or a candidate whose owner is excluded
from the clone projection, also uses that path. Other ineligible graphs use
the complete clone path so staging
still checks constraints that can depend on rows beyond the candidate's ids.
```typescript
import {
captureCandidateWriteSetTarget,
planCandidateWriteSet,
unwrap,
} from "@nicia-ai/typegraph/graph-merge";
const writeSet = {
formatVersion: 1,
sourceId: "provider-a",
target: await captureCandidateWriteSetTarget(store),
nodes: [
{
kind: "Patient",
id: "provider-a:123",
properties: { name: "Ana", mrn: "123" },
validFrom: "2026-01-01T00:00:00.000Z",
},
],
edges: [],
} as const;
const plan = unwrap(
await planCandidateWriteSet({
target: store,
makeBackend,
writeSet: JSON.parse(JSON.stringify(writeSet)),
options: {
resolve: {
Patient: {
blockIndex: "patient_mrn_candidates",
similarity: { kind: "fulltext", fields: ["name"] },
threshold: 0.9,
},
},
},
}),
);
```
`sourceId` is the stable attribution carried into conflicts, resolutions, and
provenance; node and edge ids remain the contribution source ids. The target
schema identity prevents a document authored against one graph contract from
being staged against another. `validFrom` is required (and may be `null`) so
replaying identical JSON cannot acquire a new import-time timestamp and change
the plan digest.
This adapter applies TypeGraph's existing entity/property merge semantics. Two
distinct records that both validate do not conflict merely because an
application interprets their subject, predicate, time, source, or value fields
as disagreement. Domain-specific acceptance and Statement semantics remain in
the consuming application.
Use `ingestionBranch()` when an untrusted ingestion batch may contain aliases
that deliberately repeat a canonical node's unique key. An ordinary `branch()`
keeps the complete graph schema and rejects the duplicate during staging,
before entity resolution can review and collapse it. An ingestion branch
materializes an honest working-copy schema with only node uniqueness deferred;
schema validation, edge endpoint checks, disjointness, and edge cardinality
still apply immediately.
```typescript
import { asNodeId } from "@nicia-ai/typegraph";
import {
applyMergePlan,
asBranchId,
ingestionBranch,
planMergeIncremental,
unwrap,
} from "@nicia-ai/typegraph/graph-merge";
import { importGraph } from "@nicia-ai/typegraph/interchange";
const incoming = unwrap(
await ingestionBranch(base, makeBackend, {
id: asBranchId("provider-a"),
}),
);
const imported = await importGraph(incoming, providerDocument, {
onConflict: "error",
onUnknownProperty: "error",
});
if (!imported.success) throw new Error("Provider import was rejected");
const alias = await incoming.nodes.Patient.getById(
asNodeId("incoming-patient"),
);
if (alias === undefined) throw new Error("Imported patient was not found");
// `canonicalPatient` is an existing Patient read from the base before forking.
// The repeated MRN and its identity evidence can be staged together.
await incoming.identity.assertSame(canonicalPatient, alias);
const plan = unwrap(
await planMergeIncremental({
forkPoint: base,
target: base,
branches: [incoming],
options: {
onBasePropertyConflict: "flag",
resolve: {
Patient: {
blockIndex: "patient_mrn_candidates",
similarity: { kind: "fulltext", fields: ["name"] },
threshold: 0.9,
},
},
},
}),
);
const applied = unwrap(await applyMergePlan(base, plan));
await incoming.close();
```
Both `importGraph()` and `importGraphStream()` accept the returned handle, so an
interchange document can be staged without a hand-written collection copy loop.
Import remains the single owner of node-first ordering, validity windows, edge
endpoint order and reference validation.
On an identity-enabled graph, the handle also exposes an assertion-only
`IdentityAssertionWriteFacade` as `identity`: `assertSame`, `assertDifferent`,
`bulkAssertSame`, and `bulkAssertDifferent`. This lets a batch stage aliases
that repeat unique keys and the explicit identity evidence needed to reconcile
them before merge-time constraint validation. Assertion contradictions and
invalid endpoints are still refused while staging; only node uniqueness is
deferred.
The returned handle exposes those ingestion collections and identity assertion
writes, not the branch's underlying `Store`. Identity reads and retractions,
schema operations, transactions, and runtime internals remain unavailable, so
callers cannot bypass the deferred-constraint contract. As with `Store`, the
`identity` property is absent at the type level when the graph does not enable
Operational Identity.
The original graph definition remains the merge contract:
`applyMergePlan()` validates node uniqueness against the entire resolved write
set in the target transaction. Valid key handoffs and swaps are accepted as one
set. If reviewed resolution leaves two live owners of the same unique key, the
merge returns `MergeConstraintConflictError` and commits no graph or provenance
writes.
The derived schema is persisted on the working-copy backend, so the relaxed
contract is auditable and an explicit reattachment with an equivalent graph
definition verifies the same constraint behavior. `ingestionBranch()` does not
expose a general reopen/resume API. Deferral is not an in-memory flag and does
not disable database constraints ad hoc. Ingestion branches require a backend
with the batch uniqueness operations needed for atomic final validation.
Unsupported backends are refused rather than falling back to sequential checks.
## Valid-time windows
**A new row's window travels with the merge.** An explicitly open-left row stays
open-left through snapshot and incremental merges, including edge repointing.
Reviewable plans serialize that lower bound as `validFrom: null`; an omitted
plan field means the write states no lower-bound change. JSON export/import and
plan application preserve the distinction.
A branch-authored node or edge window — including a deliberately ended one on a resurrection — is written
as-is by the commit rather than reset to merge time. When the incremental
target itself also created the surviving row, the target's committed window
wins.
**An inherited row's end-of-validity is merged.** Both
`update(id, {}, { validTo })` and `update(id, {}, { clearValidTo: true })` on a
branch are ordinary writes. The merge carries the set, move, or reopening to the
target even when the row's properties are untouched:
```typescript
await fork.store.nodes.Patient.update(asNodeId("pat-1"), {}, { validTo: "2030-06-01T00:00:00.000Z" });
const report = unwrap(await merge(base, [fork]));
// base now holds pat-1 with valid_to = 2030-06-01, and:
report.validityEnds;
// [{ entity: "node", kind: "Patient", id: "pat-1",
// validTo: "2030-06-01T00:00:00.000Z", claimedBy: ["worker-1"] }]
await fork.store.nodes.Patient.update(asNodeId("pat-1"), {}, { clearValidTo: true });
// A merge now reopens pat-1 and reports:
// [{ entity: "node", kind: "Patient", id: "pat-1",
// clearValidTo: true, claimedBy: ["worker-1"] }]
```
An ending is treated as a **sibling of deletion**, not as a property, because it
makes the same kind of statement: *this stopped being true*. That single choice
explains the whole contract:
| Situation | Outcome |
| --------- | ------- |
| One branch ends the row | That end is written — including a *later* end, which extends the window. |
| One branch reopens the row | The end is cleared with `clearValidTo: true`. |
| Several branches end it differently | No conflict. The **earliest** end wins, and `report.validityEnds` names every claiming branch. |
| Sibling branches end and reopen it | The end wins as the stronger monotone claim; every claimant remains visible in `report.validityEnds`. |
| The incremental target already ended it | The target's end stands. A branch never re-windows a row the target itself windowed, and the row is left out of the merge's writes entirely — but the discarded claims are still reported, as an entry carrying `precedence: "target"` and the target's own instant. |
| One branch ends it, another deletes it | Deleted, with **no** `DeleteModifyConflict` — the stronger statement absorbs the weaker one. |
| A branch re-states the end the target holds | No write at all — nothing is staged, so there is no version bump or history row even with `coalesceUnchangedUpserts` off. |
| No branch touched the window | Untouched. A properties-only edit never passes a window, so the committed one stands. |
The earliest-end rule is fixed, not a policy knob: it is commutative and
associative, so the merge stays order-independent, and `onPropertyConflict`
never sees a property your schema does not have.
**The branch that authored the committed end is credited.** An ending is
authored state, so its author is a contributor to that row in
`report.provenance` and in the durable sidecar — even when moving the window is
the only thing that branch changed. Credit follows the *committed* end: when
several branches end a row differently, only the branches whose claim equals the
written instant are credited, while `validityEnds[].claimedBy` still names every
claimant, winning or not. An ending a deletion absorbed commits nothing, so it
credits nobody, and neither does an entry marked `precedence: "target"` — the
merge committed none of that end.
**Every claim the merge observed is visible in `validityEnds`, applied or not.**
An entry with no `precedence` is one the merge *decided*: `validTo` is the
instant it wrote, or `clearValidTo: true` says it reopened the row. An entry
with `precedence: "target"` is one it did **not** — the incremental target had
already changed that end, so the entry describes the target's set or clear,
`claimedBy` names the branch claims that were thrown away, and nothing was
written or credited for the row. A row no branch claimed at all produces no
entry, since there was nothing to discard.
`validityEnds` reports claims about rows inherited from the fork point. If the
fork point is empty, every branch row is branch-created and the array is always
empty. A demo or topology that needs to exercise this report must seed the row
before branching, then end that inherited row on one or more branches.
Because an ending is not a modification, `onDeleteModifyConflict` never sees
one: a row whose *only* change is its window loses to a concurrent deletion even
under `"prefer-modify"`, since there is no modification to prefer. A row with a
properties edit *and* an ending keeps the usual delete/modify behavior on the
properties, and its ending rides along only if that modification survives.
**What is still NOT merged, and why.** On a row that is live in both the base
and the branch, `validTo` is the only window field a branch can author *and* the
commit can apply. A row's lower bound is immutable outside resurrection —
`validFrom` is written only when a soft-deleted row is brought back — so that
lower-bound delta remains observable in a fork but unapplicable:
| Observed delta | Reachable how | Merged? |
| -------------- | ------------- | ------- |
| `validTo` set or moved | `update(id, {}, { validTo })` | **Yes** |
| `validTo` cleared back to open | `update(id, {}, { clearValidTo: true })` | **Yes** |
| `validFrom` changed | soft-delete + resurrect inside the fork | No |
Rather than silently ignore it, the merge reports the lower-bound change in
`report.dropped` with reason `"window-not-applicable"`. Reconciling a value the
commit would then drop is worse than not merging it: the report would claim a
change that never happened.
Delete+resurrect can also make an ended base row appear open because resurrection
creates a fresh window. When `validFrom` changed, that open end is part of the
same non-applicable resurrection artifact; it is not treated as a branch-authored
`clearValidTo`, and an incremental target artifact does not outrank another
branch's explicit end claim.
Full interval reconciliation (intersecting `[validFrom, validTo]` across
branches) is deliberately out of scope — it needs a write path that moves a live
row's lower bound, which contradicts the temporal model, and it would silently
discard a branch's extension.
## Forking one graph namespace
`forkGraphNamespace(sourceStore, privateBackend, operationKey)` copies one
history-enabled graph into an independently allocated PostgreSQL database. It
copies the graph's committed schema, current rows, tombstones, recorded-time
relations, revision clock and journal, identity relations, and TypeGraph
materialization records. It checks a repeatable-read source snapshot against a
pre-cut `base@V` token, compares every copied row before target commit, and
returns `{ store, proof, abort }`. One source transaction holds that snapshot
for the entire copy, from its first source read through the target copy and
digest checks. The source can accept writes after the snapshot cut, while the
long-lived snapshot remains open until copying finishes; `proof.sourceBase`
identifies the copied cut.
```typescript
import {
forkGraphNamespace,
prepareNamespaceForkTarget,
} from "@nicia-ai/typegraph/graph-merge";
// Run with the schema owner role before the runtime fork.
await prepareNamespaceForkTarget(sourceStore, privateBackend);
const fork = await forkGraphNamespace(sourceStore, privateBackend, "restore-42");
// Owner role again: builds IVFFlat indexes over the copied rows.
await fork.store.materializeIndexes();
const historical = await fork.store
.asOfRecorded(receipt.recorded)
.nodes.Item.getById(receipt.itemId);
// Publish the private database through your own placement registry only after
// checking the fork and any application-specific restore invariants.
// Before publication, await fork.abort() to discard an unchanged copy.
```
The caller provisions and owns `privateBackend`. It may contain other graph
namespaces, but it must contain no rows for the source graph. TypeGraph refuses
a connection to the source database, including an aliased backend object.
`prepareNamespaceForkTarget()` is the owner-side step, and the fork itself
issues no DDL. It installs the retry ledger, creates the graph's per-field
pgvector tables, and builds every index the source has materialized for the
graph with the DDL the source used. It writes no graph rows and no
materialization records, so it can run before the target is empty-checked,
and running it again is harmless. Indexes whose build never completed on the
source are neither built nor required. IVFFlat indexes are the exception:
IVFFlat clusters the rows present when it is built, so building one on an
empty table gives poor recall. They are not built by preparation and their
materialization records are not copied; run `fork.store.materializeIndexes()`
after the fork to build them over the copied rows. Every other index the fork
carried is already recorded, so that call only builds the IVFFlat ones. An
IVFFlat index left on the target by an aborted fork has no record, so the
next fork's `materializeIndexes()` drops and rebuilds it over the new rows. The target stays private
until the caller changes its own placement pointer;
TypeGraph does not publish it. `abort()` atomically removes the copied graph
and operation marker while preserving unrelated namespaces, and refuses if the
target has changed. A retry with the same operation key returns the same proof
after checking the target digest and base token; a different key cannot reuse
the populated target.
This first-party copy supports the bundled PostgreSQL table layout, bundled
`pgvector` embedding storage, and default `tsvector` fulltext storage.
Embeddings are copied, digested, and verified like every other graph relation,
and `abort()` removes them. A graph with embedding fields forks only between
backends with the same vector storage: pgvector on both sides, or
`vector: false` on both, where embeddings live only in node properties. A
vector-disabled source never wrote the vector tables a pgvector target would
search, so that pair is refused. The fork refuses custom table mappings, custom vector or fulltext
strategies, and contribution-owned tables it cannot copy and validate. The
current copy buffers one relation at a time and
inserts rows in bounded batches, so operators should size the private copy
process for its largest graph relation. It does not use interchange, whose payload lacks
recorded history and tombstones.
## Determinism
Graph Merge is built to be reproducible, which is what lets you retry, cache,
diff, and test a merge with confidence:
- Candidate sets are sorted before clustering; clusters resolve by stable keys.
- Conflict resolution consults only the captured `branchOrder` (or lexicographic
branch id) — never wall-clock.
- The committed graph and the normalized report are a pure function of the
*unordered* branch set.
Use `branchOrder` to make preference explicit wherever a policy needs ordering:
```typescript
const branchOrder = [systemOfRecord.id, agentA.id, agentB.id];
const result = await merge(base, [agentB, systemOfRecord, agentA], {
branchOrder,
onPropertyConflict: "lastWriteWins", // systemOfRecord wins, regardless of input order
});
```
## Errors
Most entry points return a `Result`; the error arm is a typed `TypeGraphError`
subclass you can branch on. `applyMergePlanInTransaction()` instead throws a
typed `MergeError` so a caller-owned transaction callback cannot resolve and
commit after a partially applied failure:
| Error | When |
| ---------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `BranchError` | `branch()` or `ingestionBranch()` could not materialize a working copy. |
| `BaseVersionMismatchError` | A branch forked from a different `base@V` than the target now has (snapshot `merge()`). Also the typed replan error `mergeIncremental()`'s in-transaction guards raise, and the by-ID freshness check both commit modes run, when the target moved in the plan→commit window. |
| `IdentityMergeConflictError` | Code `GRAPH_MERGE_IDENTITY_CONFLICT`. Thrown by both `merge()` and `mergeIncremental()` for identity contradictions, assertion-ID collisions, and retract/reassert races. See the [identity guide](/identity/#interchange-and-branch-merge). |
| `MergeConstraintConflictError` | Code `GRAPH_MERGE_CONSTRAINT_CONFLICT`. The resolved plan would violate a deterministic store constraint, such as edge cardinality or node uniqueness. Its category is `constraint`, its `cause` is the original typed store error, and its details expose the original constraint fields. No graph or provenance writes commit. |
| `InvalidMergeOptionsError` | Code `GRAPH_MERGE_INVALID_OPTIONS`. The supplied option combination is invalid, `mergeIncremental()` was given the snapshot-only `target` option instead of silently ignoring it, or `mergeIncremental()`'s `onBasePropertyConflict` is not `"flag"`. |
| `SimilarityUnavailableError` | A `vector`/`hybrid` strategy was requested with no `embedder`. |
| `MergeConflictError` | A conflict could not be resolved under the configured policy. |
| `MergePlanCapabilityError` | Public planning was requested for a target without the durable revision guarantee needed across processes and time. Enable `revisionTracking` or `history`. |
| `MergePlanningStaleError` | The target moved while planning was reading it. This is an expected retry-and-replan outcome under concurrent writers: no plan was returned, so recapture the target and create a new plan before retrying. |
| `StaleMergePlanError` | The target revision changed after planning, or this plan was already applied. Review a newly-created plan. |
| `InvalidMergePlanError` | The input is not a valid plan artifact. More specific subclasses distinguish unsupported versions, digest changes, and target/schema/origin mismatches. |
| `CandidateSourceError` | A built-in candidate source failed; details identify its source id, entity kind, and operation. |
| `CandidateWriteSetError` | Code `GRAPH_MERGE_CANDIDATE_WRITE_SET`. Candidate JSON is malformed, targets another graph schema, cannot be staged, or violates the active graph contract. The accepted graph is unchanged. |
| `MergeReviewError` | Code `GRAPH_MERGE_REVIEW`. Durable review evidence is malformed, unsupported, incomplete, or inconsistent, or review options cannot be represented safely. |
| `DurableOperationError` | Code `GRAPH_MERGE_OPERATION`. System-category failure while calling a durable-operation host, including transport and strategy failures. |
| `DurableOperationRequestError` | Code `GRAPH_MERGE_OPERATION_REQUEST`. User-category refusal for an invalid durable-operation request, descriptor, or scan option. |
| `DurableOperationConflictError` | Code `GRAPH_MERGE_OPERATION_CONFLICT`. Constraint-category refusal when an idempotency key is reused with a different operation digest. The previously committed operation is returned untouched; nothing new is written. |
| `DurableOperationUnsupportedError` | Code `GRAPH_MERGE_OPERATION_UNSUPPORTED`. The strategy's `operations` capability lacks a requested member; TypeGraph refuses rather than emulating the atomic guarantee. |
| `DurableOperationEvidenceError` | Code `GRAPH_MERGE_OPERATION_EVIDENCE`. System-category failure because a host returned malformed or request-inconsistent operation evidence. |
| `DurableEvidenceUndeliveredError` | Code `GRAPH_MERGE_OPERATION_UNDELIVERED`. `destroyDurableBranch()` was refused because committed operation evidence is still undelivered. Deliver or archive it first; the typed refusal is preserved so the evidence stays recoverable. |
| `MatchEvidenceError` | Evidence could not be constructed safely, including a custom scorer returning `NaN` or infinity. |
| `MergeError` | Any other merge failure (e.g. comparison-ceiling `"error"`, a non-transactional target). `MERGE_ERROR_CODES` enumerates the codes. |
## Example
See [FHIR Graph Merge](/examples/fhir-graph-merge) for a complete runnable
snapshot merge that reconciles two independently-extracted patient-care branches,
and [Incremental Merge](/examples/incremental-merge) for live-target ingestion
against an advancing base with persisted, queryable provenance.
# Operational Identity
> Assert, retract, query, and historize identity between graph nodes
The TypeGraph Identity Profile records identity facts between **individual
nodes**. It is deliberately smaller than OWL: `same` is symmetric and
transitive, `different` is symmetric and class-lifted, and neither relation
substitutes properties or automatically expands every graph query.
## Enable the profile
Identity is graph-level and opt-in:
```typescript
const graph = defineGraph({
id: "knowledge",
nodes: { Person: { type: Person }, Author: { type: Author } },
edges: {},
identity: { sameIdAcrossKinds: "fold" },
});
```
The option is serialized with the schema. Enabled graph types expose the full
facade as `store.identity` and `tx.identity`, a read-only facade as
`StoreView.identity`, and the identity traversal option. These surfaces use
conditional **presence**: on an identity-disabled graph type, `identity` does
not exist on `Store`, `TransactionContext`, or the read-only views at all —
reaching for it is a compile error, not a `never`-typed property. A runtime
`ConfigurationError` with details code `IDENTITY_NOT_ENABLED` backs those
getters too, for widened or `any`-typed handles TypeScript can't check (a
JavaScript caller, or a store handle that lost its precise graph type).
Constraint-aware `IngestionBranch` handles follow the same conditional-presence
contract, but expose only `assertSame`, `assertDifferent`, and their bulk forms.
Reads and retractions stay unavailable so untrusted batches can stage identity
evidence without gaining the full operational surface. See
[Constraint-aware ingestion branches](/graph-merge/#constraint-aware-ingestion-branches)
for the staging and merge workflow.
At runtime, a disabled graph does no identity work: no identity locks, probes,
closure computation, or identity SQL run. That guarantee is scoped to runtime
behavior — a bundled backend still provisions the identity tables' schema (not
work) when it bootstraps a fresh database, independent of whether the specific
graph passed to `createStore`/`createStoreWithSchema` declares `identity`.
`sameIdAcrossKinds: "fold"` preserves TypeGraph's structural ID rule: live
nodes of different kinds with the same ID belong to one identity class. No
assertion row is manufactured for that implicit membership. Use
`sameIdAcrossKinds: "ignore"` to enable the assertion ledger without joining
equal IDs across kinds; only explicit `same` assertions then join classes.
## Write and read identity
```typescript
const alice = await store.nodes.Person.create(
{ name: "Alice" },
{ id: "person-alice" },
);
const author = await store.nodes.Author.create(
{ penName: "A. Example" },
{ id: "author-alice" },
);
const result = await store.identity.assertSame(alice, author);
// result.action is "created" or "existing"; result.assertion is durable truth
await store.identity.membersOf(alice);
// [{ kind: "Author", id: "author-alice" },
// { kind: "Person", id: "person-alice" }]
await store.identity.representativeOf(alice);
await store.identity.nodesOf(alice); // hydrated, kind-discriminated nodes
await store.identity.areSame(alice, author);
await store.identity.assertionsOf(alice);
await store.identity.explainSame(alice, author);
const ended = await store.identity.retractAssertion(result.assertion.id);
// ended?.validTo is the exact assertion end instant
```
The complete write surface is:
- `assertSame(a, b)` and `assertDifferent(a, b)`
- `bulkAssertSame(pairs)` and `bulkAssertDifferent(pairs)`
- `retractAssertion(id)`
- `retractSameAssertion(a, b)` and `retractDifferentAssertion(a, b)`
- `bulkRetractAssertions(ids)`
Bulk methods are eager and, on PostgreSQL, run under one graph identity lock
(see [Operational notes](#operational-notes) — SQLite serializes through its
single-writer lock instead). `bulkAssertSame` and `bulkAssertDifferent`
preserve input order and return exactly one result per input pair. Reasserting
a current semantic pair is idempotent; assertion results distinguish
`action: "created"` from `action: "existing"`. Retraction methods return the
ended assertion (or `undefined` for a missing current assertion).
`bulkRetractAssertions` does **not** share that one-result-per-input shape: it
dedupes the input ids and returns only the assertions that were actually open,
in dense, first-occurrence input order — so the result does not align
index-by-index with the input array. Self-assertions are rejected.
Assertion IDs use the exported private-symbol-branded
`IdentityAssertionId` type so unrelated strings cannot be passed accidentally.
When you hold a plain assertion-ID string that came from persistence or an
interchange document, re-enter the branded type with the `asIdentityAssertionId(value)`
caster rather than a `as` assertion.
Assertions may state an explicit half-open validity window. Scalar methods take
the window as their third argument; bulk methods carry one window per pair:
```typescript
await store.identity.assertSame(alice, legacyAlice, {
validFrom: "2020-01-01T00:00:00.000Z",
validTo: "2022-01-01T00:00:00.000Z",
});
await store.identity.bulkAssertDifferent([
{ a: alice, b: bob, validFrom: "2023-01-01T00:00:00.000Z" },
{ a: alice, b: carol }, // ordinary current assertion semantics
]);
```
A past-ended window affects historical reads only. An open window beginning in
the past affects both historical and current reads. Repeating the exact
relation, pair, and window is idempotent. A second open window for an already
current semantic pair is refused rather than silently collapsed onto a
different `validFrom`. Empty objects retain the ordinary unwindowed semantics,
including inside a mixed bulk call.
Runtime-evolved nodes carry a nominal dynamic-node type, so they flow through
the same identity surface without a cast:
```typescript
const evolved = await store.evolve(extension);
const person = await evolved.nodes.Person.create({ name: "Alice" });
const tag = await evolved
.getNodeCollectionOrThrow("Tag")
.create({ label: "author" });
await evolved.identity.assertSame(person, tag);
await evolved.identity.membersOf(tag);
```
Reference reads return `IdentityNodeReference` values covering both
compile-time graph kinds and registered runtime kinds. This widening is
necessary even when a read starts from `person`, because its class can contain
`tag`. Their IDs retain the appropriate nominal brand, and `nodesOf` hydrates
the class into static kind-discriminated members or `DynamicNode` values for
runtime members. A plain `{ kind: string, id: string }` does not prove that the
kind came through the evolved Store; pass the dynamic node or a nominal dynamic
reference returned by an identity read. Unknown and removed kinds still fail
at runtime with `KindNotFoundError`. A missing, deleted, or
coordinate-invisible input returns `undefined`, `[]`, or `false` according to
the method. A visible singleton returns itself from `membersOf` and
`representativeOf`, and `areSame(ref, ref)` is true. `areDifferent` lifts an
explicit different assertion across both identity classes and also reflects
ontology `disjointWith` constraints. Representatives are deterministic: the
code-point-smallest `(kind, id)` visible member wins.
`explainSame(a, b)` returns a shortest path of persisted `same` assertions and
implicit same-ID folds connecting two visible references. Each step names its
endpoints and either the assertion or `type: "same-id-fold"`. It returns `[]`
for one visible reference and `undefined` when the references are distinct or
not visible at the read coordinate. Use `store.asOf(instant).identity` for a
historical explanation.
Historical identity reads and identity-expanded traversals use the kinds
registered on the current Store. Assertions involving a removed kind remain
in recorded history but no longer connect active classes.
`classes({ limit, kinds?, cursor? })` lists visible classes, including
singletons, in representative order. A kind filter selects classes containing
at least one visible member of the requested kinds; each result still includes
all of that class's visible members. Pass `nextCursor` to the next call until
it is absent. The cursor is exclusive and applies to the same graph, read
coordinate, and kind filter. When `kinds` is omitted, the scan uses the
registered runtime kinds present when each page is requested; adding a runtime
kind during that scan changes the filter and invalidates its cursor. At current
coordinates, the database finds visible representatives
for the page and expands members only for those classes; discovering
representatives still examines the visible node set. Historical coordinates
reconstruct all visible classes before applying the page boundary. For paging
across writes, use a recorded-time coordinate when recorded history is enabled:
valid-time `asOf` reads still observe later changes to the live tables.
## Integrity and lifecycle
Ordinary unwindowed assertions require live endpoints. Explicit windows require
both endpoint rows to cover the assertion's whole half-open interval; an ended
or late-starting endpoint raises `IdentityEndpointValidityError`. Future bounds
and inverted windows raise `IdentityValidityWindowError`. Zero-width windows
are accepted as empty history. Contradictions are checked throughout every
overlapping segment, including transitive `same` paths; adjacent half-open
windows do not overlap.
`assertSame` fails when a current
`different` assertion spans the two classes or when any member kinds are
ontology-disjoint. `assertDifferent` fails when both endpoints are already in
one class. These checks, folding, node deletion, import, schema-transition
validation, and closure rebuild share one per-graph lock and one mutation
coordinator.
Soft-deleting a node ends its current assertions. Hard-deleting it removes
every current and ended assertion touching the node from the live assertion
ledger; when recorded history is enabled, earlier recorded coordinates remain
queryable. On every graph, a `create()` or `upsertById()` for a soft-deleted
same-`(kind, id)` row **resurrects** that row rather than erroring:
its properties are replaced and its validity window is reset, so `validFrom`
becomes the resurrection instant — unless the write carries an explicit
window, which is honored as given (this is how merge preserves
branch-authored windows). A resurrecting node write that supplies only a
historical `validTo` takes the same **born-already-ended** exception a create
takes: no lower bound is stored ("ended at T, start unknown") rather than a
start after its own end, so the row reads back at every `asOf` before that end
and `meta.validFrom` is `undefined`. One stated window reaches one stored shape
whichever node path resets it — `create()` on a fresh id, `create()` on a
tombstone, or a resurrecting `upsertById()`. (Edge resurrection instead keeps
its stored lower bound, so `getOrCreateByEndpoints` can resurrect an edge
directly into the ended state — but the end it names is held to that retained
bound, so reviving an edge into a window that closed before the edge began is
refused as a `ValidationError`, and means passing both bounds.) This graph-wide
rule does not depend on the
identity profile. Resurrection does not revive ended assertions, but folding
runs again over the resurrected node when configured. Kind removal
cascades assertion and closure rows for the removed kinds. Tightening ontology
disjointness is rejected when it would make a persisted class contradictory.
`rebuildIdentityClosure(store)` repairs the derived current closure from live
nodes and current assertions. It validates integrity and never advances the
content revision. Schema-managed rebuilds, including automatic startup repair
of derived identity relations, pin the schema version used by the rebuild.
If a concurrent migration advances that version first, repair refuses with
`StaleVersionError` without overwriting the newer closure. Reopen using the
current graph definition before retrying.
### The database-level backstop
The checks above are code deciding whether a write is legal, and code can be
wrong. Underneath them TypeGraph maintains a second derived relation — the
**separation relation** — that holds one row per pair of identity classes a
current `different` assertion keeps apart, keyed by the two class keys under a
`CHECK (class_key_low < class_key_high)` constraint.
Every transaction that fuses two identity classes relabels the affected
separation rows in the same statement batch. Fusing two classes that were
separated relabels both sides of their shared row to one key, the constraint
rejects it, and the transaction aborts — in the engine, with no application
code in the way. A write that reaches the ledger through a path that skipped
identity validation therefore still cannot commit a contradictory graph; it
fails with an `IdentitySeparationViolationError` naming the `different`
assertion it contradicts.
Nothing about the identity API changes. The relation is derived and
maintained wherever the closure is, `rebuildIdentityClosure(store)` recomputes
it from the ledger, and store-open validation checks it against that
recomputation the same way it checks the closure.
## Temporal identity
Integrity is **structural**; reads are **coordinate-visible**.
Current reads use a materialized closure and then filter members through the
same visibility predicate ordinary node reads use. `store.identity` and
`store.asOf(now).identity` therefore agree.
Non-current valid-time and recorded-time views reconstruct one fixed point over
both explicit `same` assertions and same-ID folding edges. A structurally
existing but coordinate-invisible bridge can conduct identity without being
returned as a member. Recorded assertions are captured in the same commit as
the truth-bearing write.
The assertion's validity window and the commit that recorded it are independent
coordinates. A retrospective assertion is therefore invisible before its
recorded-time commit even when its valid-time window reaches farther into the
past. Archival export includes the endpoint temporal bounds needed to validate
those windows on import, and graph merge carries branch-authored bounded
assertions without turning them into current truth.
Identity profile and ontology rules are schema-level interpretation, not a
third temporal dimension. Historical views apply the Store's pinned
`sameIdAcrossKinds` mode and ontology to the assertions and nodes visible at the
requested coordinate. Changing those schema rules can therefore reinterpret
older coordinates; it does not rewrite the recorded assertion ledger.
```typescript
const before = await store.recordedNow();
const historical = store.asOfRecorded(before!);
await historical.identity.membersOf(alice);
```
### Folds and time
Implicit same-id folds (`sameIdAcrossKinds: "fold"`) conduct based on a node's
**lifecycle** — whether it currently exists and is not soft-deleted — not its
valid-time window. A node created today with a backdated `validFrom` is
valid-time visible in the past (an ordinary node read at that past coordinate
returns it), but it does not conduct a fold there: the fold only takes effect
once the node actually exists. Symmetrically, a node with a future `validFrom`
does not suppress its folds today — it already exists and is live, so it
folds now even though it is not yet valid-time visible. Explicit `same` and
`different` assertions are unaffected by this: they carry their own validity
windows and conduct exactly when they are current. This keeps the fold
computation tied to write events rather than to valid-time windows, so the
materialized closure used by current reads and by `asOf(now)` reads is
identical — a fixed-point reconstruction of "current" never needs to
special-case valid-time skew on the folding edge itself.
## Identity-expanded traversal
Traversal expansion is per hop and defaults off:
```typescript
const results = await store
.query()
.from("Person", "person")
.traverse("authored", "edge", { includeIdentityMembers: true })
.to("Document", "document")
.select((ctx) => ({ edge: ctx.edge, document: ctx.document }))
.execute();
```
The hop considers coordinate-visible members of the source class, returns the
physical edge and target rows, preserves their provenance, and deduplicates a
physical edge within the step — with one legitimate exception: a self-inverse
edge (`inverseOf(edgeKind, edgeKind)`) traversed with `expand` between two
identity-folded peers can yield the same physical edge twice, once per
direction/target it matches through the fold. That is not a dedup bug; the
edge genuinely satisfies the traversal from both of its endpoints. Recursive
traversal supports the same option.
TypeGraph does not perform automatic graph-wide expansion and collection reads
such as `getById` have no identity option.
Both coordinates reach the candidate edge the same way — an ordinary indexed
equality on the class member, never a membership test evaluated per candidate
edge. How each one reaches the class differs, because what a class costs to
compute differs.
At the **current** coordinate the maintained closure already *is* the class
relation, so each traversal step seeks into it from its own frontier rows: the
frontier row's class through the closure's primary key, that class's members
through the class index, each member's node for its visibility. Cost is
proportional to the frontier and the size of its classes — never to how many
identity classes the graph holds. Measured on SQLite with *n* Person nodes, each
folded with a Company and an Alias peer sharing its id (a three-member class per
source), all *n* acting as source rows and every edge leaving the Company peer:
| source rows | fan-out | matching edges | before | after |
| --- | --- | --- | --- | --- |
| 250 | 1 | 250 | 67 ms | 6 ms |
| 1000 | 1 | 1000 | 1077 ms | 9 ms |
| 2000 | 1 | 2000 | 4616 ms | 19 ms |
| 1000 | 8 | 8000 | 8611 ms | 13 ms |
| 500 | 200 | 100,000 | 51,602 ms | 77 ms |
Growth is linear in graph size where it used to quadruple per doubling: the hop
no longer evaluates membership per candidate *(source row, edge)* pair. The
number to plan around is the last row — a hundred thousand matching edges over a
five-hundred-row frontier is where the old per-source rescan dominated.
A **historical** hop — one under `asOf`, `asOfRecorded`, or a non-current
`view()` — cannot use the materialized closure, because the closure represents
only the present. Its rows come from a reconstruction of identity classes out of
the assertion ledger, and under `sameIdAcrossKinds: "fold"` that reconstruction
also has to consider the structural same-id relation, which is proportional to
the number of live nodes in the graph. No frontier row narrows that fixed point,
so it is built once per statement into a materialized relation every traversal
step joins. Measured on the narrow-edge fixture that isolates the term (SQLite,
*n* Person nodes each folded with a Company peer, all *n* acting as source rows,
fan-out 1):
| *n* | before | after |
| --- | --- | --- |
| 250 | 122 ms | 7 ms |
| 500 | 486 ms | 7 ms |
| 1000 | 1984 ms | 14 ms |
| 2000 | 8261 ms | 28 ms |
Growth is linear in graph size where it used to quadruple per doubling.
The caveat that remains is the historical one, and it is worth planning around: a
past-coordinate hop rebuilds the whole graph's classes even when you asked about
one node, so its floor is a pass over the identity population regardless of how
narrow the frontier is. A **current** hop has no such floor — a single-start-row
hop over 50,000 folded triples measures 1 ms on SQLite against 387 ms when the
class relation was still built graph-wide, and nine unrelated 501-member classes
cost it nothing at all (0.5 ms on SQLite, 2.4 ms on PostgreSQL, against 564 ms
and 568 ms). Pick the coordinate you actually need: reading the present is the
cheaper question by a wide margin.
## Interchange and branch merge
Interchange format `2.0` optionally carries an identity section. State export
(the default) includes current assertions. Import into a populated target is
target-oriented: an existing current semantic pair keeps its target assertion
ID and `validFrom`. Working-copy branch cloning imports into an empty target and
preserves source IDs and `validFrom` exactly.
```typescript
const state = await exportGraph(store, { includeTemporal: true });
const archive = await exportGraph(store, {
identityMode: "archival",
includeDeleted: true,
});
```
Identity-enabled exports default `includeTemporal` to `true`, because importing
identity truth must prove that both endpoints existed throughout each assertion
window. Explicitly setting `includeTemporal: false` on an identity-enabled graph
is refused.
Archival mode also includes ended assertions. Those rows are restored after
shape validation and do not affect current closure. An ending a node deletion
caused carries that node as `endedBy`, so a round-trip preserves why each
assertion ended and not merely that it did; import rejects an `endedBy` on an
open assertion, or one naming a node that is not an endpoint of the assertion
it ends. Ended assertions can
reference soft-deleted nodes, and by default (`includeDeleted: false`) export
joins every assertion against its endpoints' live rows — an assertion with a
soft-deleted endpoint is silently **dropped from the export entirely**, not
carried with a dangling reference. Pair `identityMode: "archival"` with
`includeDeleted: true` to keep those assertions in the archive. Interchange
documents carry no `deletedAt` field, so a node exported only because of
`includeDeleted: true` re-imports as **live** — an `includeDeleted` archive
resurrects its soft-deleted nodes on import rather than restoring them as
deleted. Weigh that trade-off deliberately for a backup: without
`includeDeleted`, soft-deleted endpoints and the assertions that reference them
are silently absent; with it, those nodes come back alive. Recorded side
tables are not part of interchange.
Graph merge includes identity truth in staleness fingerprints and diffs.
Duplicate current assertions use the earliest `validFrom`, then the
code-point-smallest assertion ID — unless one candidate is already committed
on the target with the exact staged truth, which always wins: the applier is
idempotent per semantic pair, so a challenger could never actually be
written. A node deletion cascades into ending the assertions touching it, at
the node's own deletion instant, and records the deleted node on every row it
ends — so the diff reads which endings that deletion caused and stages each
one with its cause, however close in time the branch's own retractions
fell. When a delete/modify
conflict resolution keeps the node, an ending is dropped along with the
overruled deletion that caused it (reported as
`identity:deletion-overruled`), while a retraction a branch made itself
survives the deletion being overruled — including one the deleting branch
made before deleting the node, even in the deletion's own millisecond. A hard
delete removes the assertion rows outright, taking the recorded cause with
them and leaving nothing to separate cause from intent, so those endings count
as cascades. `merge()` detects identity conflicts at plan
time and returns them as a typed `IdentityMergeConflictError` — direct
opposing relations on one endpoint pair, transitive contradictions reached
through a chain of `same` assertions no single branch wrote, retract/reassert
races, and an assertion over a node another branch deleted. A branch that
retracts a pair and also reasserts it itself (convergent, not racing) merges
cleanly. This is mechanical truth propagation, not semantic entity
reconciliation. Plan time is the early surface, not the only one: any
identity refusal that still escapes to the applier inside the commit
transaction is translated into the same typed `IdentityMergeConflictError`,
with the original error preserved as its cause (identity environment and
storage-corruption codes pass through untranslated — they are not statements
about merge truth). See
[`IdentityMergeConflictError`](/errors/#identitymergeconflicterror)
for the exact `merge()` signature and how to catch it.
### Independent targets and assertion IDs
`mergeIncremental()` accepts a target that has moved on from the branches'
fork point, so a branch's assertion IDs can meet a ledger that assigned those
IDs independently. Snapshot `merge()` still requires its target to match the
branches' base@V exactly, but the same by-ID contract governs the divergence
a branch can create within its own lineage (hard-delete/recreate replacement)
and the plan→commit window. The contract is by ID, on complete truth:
- **One assertion ID, one complete truth.** A planned assertion whose ID the
target's ledger — ended rows included — already binds to a different
complete truth (relation, endpoints, validity) refuses at plan time as
`IdentityMergeConflictError`. An exact match is applied idempotently.
- **Retractions carry the truth they retract.** A branch retraction ends the
target's current row for its ID only when that row *is* the truth the
branch retracted. When the target reuses the ID for different truth, the
retraction is skipped and reported in `MergeReport.dropped` as
`identity:retraction-target-mismatch` — the branch's own assertion is
already absent from the target, and ending the target's unrelated row
would delete truth the branch never saw.
- **Truth replacement is a conflict, not a silent keep.** Within one lineage
a branch can legally rebind an assertion ID by hard-deleting an endpoint
(which physically removes the row) and importing the ID for different
truth. The diff stages that replacement as a retraction plus a new
assertion; because the target's ledger still holds the ID's prior truth in
an ended row, the plan-time one-ID-one-truth check refuses it typed rather
than silently keeping either side's truth.
- **The commit re-verifies IDs.** Both commit modes re-read every planned
assertion and retraction ID inside the commit transaction and refuse
plan→commit drift as `BaseVersionMismatchError` — retrying recomputes the
plan from current state. One deliberate exception: a planned retraction
whose row another writer already ended is accepted as a no-op, not drift.
`MergeReport.merged.identity` reports the rows the applier actually
created and ended; idempotent skips are excluded.
- **The commit proves the result, not the plan.** After its identity writes,
and still inside the same transaction, a merge re-derives the identity
classes it touched from the written state and refuses a contradiction there
as `IdentityMergeConflictError` — so a plan validated against state that has
since moved cannot leave a contradictory ledger behind. The whole merge rolls
back; there is no partial commit. If the derived classes disagree with the
materialized closure, the closure is rebuilt inside the same transaction and
the check re-runs, which repairs a lagging closure atomically with a merge
that is otherwise sound.
## Operational notes
On PostgreSQL, every identity-affecting node write on an identity-enabled graph
serializes on a per-graph advisory transaction lock: at most one writer per
graph proceeds at a time. This is a correctness guarantee for the assertion
ledger and closure, and it is also a throughput ceiling — concurrent writers to
the same graph queue behind the lock. Writes to other graphs, and all reads,
are unaffected.
First-time enablement is heavier than steady state. It takes a `SHARE` lock on
the shared nodes table, which briefly blocks writes for **every** graph in that
database, and it loads the whole graph to build the initial identity closure.
Plan enablement for a quiet window on large databases. `evolve()` on an
identity-enabled graph re-runs the same closure rebuild, so schema evolution
carries a comparable one-time cost proportional to graph size.
Changing `sameIdAcrossKinds` is a **breaking** schema change — a `fold`↔`ignore`
flip rewrites the materialized identity closure and changes every
`areSame`/`membersOf`/`includeIdentityMembers` answer against existing data —
so it requires the same explicit `migrateSchema()` opt-in as any other
breaking change; it never auto-migrates silently. Identity-relevant ontology
changes (`disjointWith`, `equivalentTo`/deprecated `sameAs`, or `subClassOf`)
are likewise persisted semantic migrations, not a local runtime toggle.
`createStoreWithSchema` and explicit `migrateSchema()` both rebuild and
validate the closure atomically with the schema commit that carries the
change. While the flip is unapplied, store construction refuses with
`ConfigurationError` details code `IDENTITY_PROFILE_MIGRATION_PENDING`
whenever the identity change is the only breaking one in the diff; a
migration that also breaks other schema surfaces raises the generic
`MigrationError` enumerating everything. First-time identity *enablement* (`autoMigrate: false` on a graph
newly declaring `identity: { ... }`) is a safe, additive change, and
`createStoreWithSchema` refuses to return a Store while it is pending with
`ConfigurationError` details code `IDENTITY_ENABLEMENT_PENDING`. The very
first schema commit of an identity-enabled graph is an enablement too: a
legacy database populated through an unmanaged `createStore` gets the same
atomic fold scan, contradiction validation, and closure build during
initialization — an empty database just makes them cheap no-ops.
## Migrating from type-level factories
The ontology factories `sameAs(A, B)` and `differentFrom(A, B)` are deprecated:
they relate **types**, not individual rows, and `differentFrom` never enforced
instance identity. To migrate:
1. Add `identity: { sameIdAcrossKinds: "fold" }` to the graph.
2. Open it with `createStoreWithSchema` so the capability is persisted and
existing cross-kind same-ID groups are validated and materialized.
3. Replace type-level facts with `store.identity` assertions between concrete
node references.
4. Use `equivalentTo` or `disjointWith` when the intended relation is genuinely
between kinds.
On PostgreSQL, first-time enablement waits for in-flight node writes before it
builds the initial identity closure. Quiesce or restart any store instances
that were opened with the identity-disabled schema before allowing writes to
resume; stale instances do not participate in identity locking.
Identity requires interactive atomic transactions. Bundled SQLite and
PostgreSQL drivers support it; Cloudflare D1 and `drizzle-orm/neon-http` reject
an enabled graph with `ConfigurationError` details code
`IDENTITY_REQUIRES_ATOMIC_BACKEND`. Identity-disabled graphs continue to work
on those drivers.
Durable entity handles, identity-group IDs, semantic reconciliation, automatic
OWL property substitution, and graph-wide identity expansion are reserved
future capabilities and are not implied by this profile.
# Integration Patterns
> Strategies for integrating TypeGraph into your application architecture
This guide covers common integration patterns for adding TypeGraph to existing
applications, from simple setups to production deployment strategies.
## Direct Drizzle Integration (Shared Database)
If you're already using Drizzle ORM, TypeGraph can share your existing database
connection. TypeGraph tables coexist alongside your application tables.
```typescript
import { drizzle } from "drizzle-orm/node-postgres";
import { Pool } from "pg";
import { createPostgresBackend, generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import { createStore } from "@nicia-ai/typegraph";
// Your existing Drizzle setup
const pool = new Pool({ connectionString: process.env.DATABASE_URL });
const db = drizzle(pool);
// Add TypeGraph tables to your existing database
await pool.query(generatePostgresMigrationSQL());
// Create TypeGraph backend using the same connection
const backend = createPostgresBackend(db);
const store = createStore(graph, backend);
// For pure TypeGraph operations, use store.transaction()
await store.transaction(async (tx) => {
const person = await tx.nodes.Person.create({ name: "Alice" });
const company = await tx.nodes.Company.create({ name: "Acme" });
await tx.edges.worksAt.create(person, company, { role: "Engineer" });
});
```
### Mixed Drizzle + TypeGraph Transactions
When combining TypeGraph operations with direct Drizzle queries in the same atomic transaction,
create a temporary backend from the Drizzle transaction:
```typescript
await db.transaction(async (tx) => {
// Direct Drizzle operations
await tx.insert(auditLog).values({ action: "user_created" });
// TypeGraph operations in the same transaction
const txBackend = createPostgresBackend(tx);
const txStore = createStore(graph, txBackend);
await txStore.nodes.Person.create({ name: "Alice" });
});
```
This pattern is only needed when you must combine both in one atomic transaction.
**When to use:**
- You want a single database to manage
- Your graph data relates to existing tables
- You need cross-cutting transactions
**Considerations:**
- TypeGraph tables use the `typegraph_` prefix to avoid collisions
- Run TypeGraph migrations alongside your application migrations
- Connection pool is shared, so size accordingly
## Drizzle-Kit Managed Migrations (Recommended)
If you use `drizzle-kit` to manage migrations, you can import TypeGraph's table
definitions directly into your schema file. This lets drizzle-kit generate
migrations for all tables—both yours and TypeGraph's—in one place.
### Setup
**1. Import TypeGraph tables into your schema:**
```typescript
// schema.ts
import { sqliteTable, text, integer } from "drizzle-orm/sqlite-core";
// Import TypeGraph tables (these are standard Drizzle table definitions)
export * from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
// Or for PostgreSQL:
// export * from "@nicia-ai/typegraph/adapters/drizzle/postgres";
// Your application tables
export const users = sqliteTable("users", {
id: text("id").primaryKey(),
name: text("name").notNull(),
email: text("email").notNull(),
});
```
**2. Generate migrations normally:**
```bash
npx drizzle-kit generate
```
Drizzle-kit will now see all tables—TypeGraph's and yours—and generate migrations
for them.
**3. Apply migrations:**
```bash
npx drizzle-kit migrate
# Or for Cloudflare D1:
wrangler d1 migrations apply your-database
```
**4. Create the backend:**
```typescript
import { drizzle } from "drizzle-orm/better-sqlite3";
import Database from "better-sqlite3";
import { createSqliteBackend, tables } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
import { createStore } from "@nicia-ai/typegraph";
const sqlite = new Database("app.db");
const db = drizzle(sqlite);
// Use the same tables that drizzle-kit manages
const backend = createSqliteBackend(db, { tables });
const store = createStore(graph, backend);
```
### Custom Table Names
To avoid conflicts or match your naming conventions, use the factory function:
```typescript
// schema.ts
import { createSqliteTables } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
// Create tables with custom names
export const typegraphTables = createSqliteTables({
nodes: "myapp_graph_nodes",
edges: "myapp_graph_edges",
uniques: "myapp_graph_uniques",
schemaVersions: "myapp_graph_schema_versions",
embeddings: "myapp_graph_embeddings",
fulltext: "myapp_graph_fulltext",
indexMaterializations: "myapp_graph_index_materializations",
kindRemovals: "myapp_graph_kind_removals",
reconciliationMarkers: "myapp_graph_reconciliation_markers",
});
// Export individual tables for drizzle-kit
export const { nodes: myappGraphNodes, edges: myappGraphEdges, uniques: myappGraphUniques, schemaVersions: myappGraphSchemaVersions, embeddings: myappGraphEmbeddings, indexMaterializations: myappGraphIndexMaterializations, kindRemovals: myappGraphKindRemovals, reconciliationMarkers: myappGraphReconciliationMarkers } = typegraphTables;
// SQLite fulltext is an FTS5 virtual table — drizzle-kit can't model
// virtual tables, so this name is exposed as a string. The backend
// creates the FTS5 table on first store boot via a focused
// `ensureFulltextTable()` ensure (idempotent CREATE VIRTUAL TABLE
// IF NOT EXISTS), so drizzle-kit-managed setups work without an
// extra manual step.
export const myappGraphFulltextTableName = typegraphTables.fulltextTableName;
```
For PostgreSQL with the default `tsvectorStrategy`, the factory
**does** return a typed Drizzle table — `tables.fulltext` — alongside
the others, so drizzle-kit-managed setups pick up the fulltext table
automatically:
```typescript
// schema.ts (PostgreSQL)
import { createPostgresTables } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
export const typegraphTables = createPostgresTables({
// …same names as above…
});
export const {
nodes: myappGraphNodes,
edges: myappGraphEdges,
// …
fulltext: myappGraphFulltext,
indexMaterializations: myappGraphIndexMaterializations,
// …
} = typegraphTables;
```
If you swap in an alternate Postgres fulltext strategy (pg_trgm,
ParadeDB / pg_search, pgroonga), the typed `tsvector`-shaped table
won't match what your strategy needs. Override `tables.fulltext` in
your schema barrel with your strategy's own Drizzle table, or skip
the typed export and rely on the backend's runtime
`ensureFulltextTable()` ensure to bootstrap your strategy's DDL.
Then pass the same tables to the backend:
```typescript
import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
import { typegraphTables } from "./schema";
const backend = createSqliteBackend(db, { tables: typegraphTables });
```
### Adding TypeGraph Indexes
The table factory functions also accept `indexes`, which drizzle-kit will include in migrations:
```ts
// schema.ts
import { createSqliteTables } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
import { defineNodeIndex } from "@nicia-ai/typegraph/indexes";
import { Person } from "./graph";
const personEmail = defineNodeIndex(Person, { fields: ["email"] });
export const typegraphTables = createSqliteTables({}, { indexes: [personEmail] });
```
For PostgreSQL, use `createPostgresTables` from `@nicia-ai/typegraph/adapters/drizzle/postgres`.
See [Indexes](/performance/indexes) for covering fields, partial indexes, and profiler integration.
Beyond accelerating queries, a declared index powers `store.nodes..bulkFindByIndex(indexName, items)` — a
batched lookup that returns, for each incoming record, the live nodes sharing its index key. This is the primitive for
**import reconciliation** and **dedup-candidate discovery**: probe an entire import batch against the graph in one
query to decide create-vs-merge per record (the key may be non-unique, so each record yields its own candidate list).
See the [batched index lookup reference](/performance/indexes#batched-index-lookup-bulkfindbyindex) and the runnable
[`examples/17-bulk-find-by-index.ts`](https://github.com/nicia-ai/typegraph/blob/main/packages/typegraph/examples/17-bulk-find-by-index.ts).
If you only need PostgreSQL adapter exports, import from `@nicia-ai/typegraph/adapters/drizzle/postgres`:
```typescript
import { createPostgresBackend, tables } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
```
### PostgreSQL with pgvector
For PostgreSQL with vector search, ensure the pgvector extension is enabled
before running migrations:
```sql
CREATE EXTENSION IF NOT EXISTS vector;
```
When multiple allocations share one PostgreSQL database, give each backend a
stable namespace so its pgvector tables and indexes remain physically isolated:
```typescript
import { createPgvectorStrategy } from "@nicia-ai/typegraph";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
const backend = createPostgresBackend(pool, {
vector: createPgvectorStrategy("tenant-a"),
});
```
Keep the namespace stable for the lifetime of the allocation. The default
`pgvectorStrategy` continues to use the existing `tg_vec` / `tg_vecidx` names.
Then in your schema:
```typescript
// schema.ts
export * from "@nicia-ai/typegraph/adapters/drizzle/postgres";
export const users = pgTable("users", { ... });
```
**When to use:**
- You already use drizzle-kit for migrations
- You want a single migration workflow for all tables
- You need Cloudflare D1 or other platforms that require drizzle-kit migrations
**Advantages over raw SQL migrations:**
- Single source of truth for schema
- Type-safe schema in TypeScript
- Drizzle-kit handles migration diffs automatically
- Works with all drizzle-kit supported platforms
## Separate Database
Use a dedicated database when you want isolation between your application data
and graph data.
```typescript
import { Pool } from "pg";
import { drizzle } from "drizzle-orm/node-postgres";
import { createPostgresBackend, generatePostgresMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
// Application database (your existing setup)
const appPool = new Pool({ connectionString: process.env.APP_DATABASE_URL });
const appDb = drizzle(appPool);
// Dedicated TypeGraph database
const graphPool = new Pool({ connectionString: process.env.GRAPH_DATABASE_URL });
const graphDb = drizzle(graphPool);
await graphPool.query(generatePostgresMigrationSQL());
const backend = createPostgresBackend(graphDb);
const store = createStore(graph, backend);
```
**When to use:**
- Your primary database doesn't support required features (e.g., pgvector)
- You want independent scaling for graph operations
- Compliance requires data separation
- You're adding graph capabilities to a legacy system
**Considerations:**
- No cross-database transactions (use eventual consistency patterns)
- Sync data between databases via application logic or events
- Separate backup/restore procedures
## In-Memory (Ephemeral Graphs)
Use in-memory SQLite for temporary graphs, caching, or computation.
```typescript
import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local";
function createEphemeralStore(graph: GraphDef) {
const { backend } = createLocalSqliteBackend();
return createStore(graph, backend);
}
// Use case: Build a temporary graph for computation
async function computeRecommendations(userId: string): Promise {
const tempStore = createEphemeralStore(recommendationGraph);
// Load relevant data into temporary graph
const userData = await fetchUserData(userId);
await populateGraph(tempStore, userData);
// Run graph algorithms
const results = await tempStore
.query()
.from("User", "u")
.whereNode("u", (u) => u.id.eq(userId))
.traverse("similar", "s")
.to("Product", "p")
.select((ctx) => ctx.p)
.execute();
return results;
}
```
**When to use:**
- Temporary computation graphs
- Request-scoped graph state
- Graph-based caching with expiration
- Isolated test fixtures
**Considerations:**
- Data lost on process termination
- Memory usage scales with graph size
- No persistence—rebuild on restart
## Hybrid Overlay (Graph on Existing Data)
Add graph relationships on top of existing relational data without migrating
your data model. Your existing tables remain the source of truth; TypeGraph
stores only the relationships and graph-specific metadata.
Use the `externalRef()` helper to create type-safe references to external tables:
```typescript
import { createExternalRef, defineEdge, defineGraph, defineNode, embedding, externalRef } from "@nicia-ai/typegraph";
import { z } from "zod";
// Define nodes that reference your existing tables
const User = defineNode("User", {
schema: z.object({
// Type-safe reference to your existing users table
source: externalRef("users"),
// Denormalized fields for graph queries (optional)
displayName: z.string().optional(),
}),
});
const Document = defineNode("Document", {
schema: z.object({
source: externalRef("documents"),
embedding: embedding(1536).optional(),
}),
});
// Graph-only relationships not in your relational schema
const relatedTo = defineEdge("relatedTo", {
schema: z.object({
relationship: z.enum(["cites", "extends", "contradicts"]),
confidence: z.number().min(0).max(1),
}),
});
const authored = defineEdge("authored");
const graph = defineGraph({
id: "document_graph",
nodes: { User, Document },
edges: {
relatedTo: { type: relatedTo, from: [Document], to: [Document] },
authored: { type: authored, from: [User], to: [Document] },
},
});
```
The `externalRef()` helper validates that references include both the table name
and ID, catching errors at insert time:
```typescript
// Valid: includes table and id
await store.nodes.Document.create({
source: { table: "documents", id: "doc_123" },
});
// Error: wrong table name (caught by TypeScript and runtime validation)
await store.nodes.Document.create({
source: { table: "users", id: "doc_123" }, // Type error!
});
// Use createExternalRef() for a cleaner API
const docRef = createExternalRef("documents");
await store.nodes.Document.create({
source: docRef("doc_456"),
});
```
**Syncing with external data:**
```typescript
// Sync helper: Create or update graph node from app data
async function syncDocument(store: Store, appDocument: AppDocument) {
const existing = await store
.query()
.from("Document", "d")
.whereNode("d", (d) => d.source.get("id").eq(appDocument.id))
.select((ctx) => ctx.d)
.first();
if (existing) {
await store.nodes.Document.update(existing.id, {
embedding: await generateEmbedding(appDocument.content),
});
return existing;
}
return store.nodes.Document.create({
source: { table: "documents", id: appDocument.id },
embedding: await generateEmbedding(appDocument.content),
});
}
// Query combining graph traversal with app data hydration
async function findRelatedDocuments(documentId: string) {
// Get graph relationships
const related = await store
.query()
.from("Document", "d")
.whereNode("d", (d) => d.source.get("id").eq(documentId))
.traverse("relatedTo", "r")
.to("Document", "related")
.select((ctx) => ({
source: ctx.related.source,
relationship: ctx.r.relationship,
confidence: ctx.r.confidence,
}))
.execute();
// Hydrate with full data from app database
const externalIds = related.map((r) => r.source.id);
const fullDocuments = await appDb.select().from(documents).where(inArray(documents.id, externalIds));
return related.map((r) => ({
...r,
document: fullDocuments.find((d) => d.id === r.source.id),
}));
}
```
**When to use:**
- Adding graph capabilities to an existing application
- Semantic search over existing content
- Relationship discovery without schema changes
- Gradual migration from relational to graph thinking
**Considerations:**
- Maintain sync between app data and graph nodes
- Decide what to denormalize (tradeoff: query speed vs. sync complexity)
- The `table` field in `externalRef` enables referencing multiple external sources
## Background Embedding Workers
Decouple embedding generation from request handling using background jobs.
```typescript
// job-queue.ts - Define the embedding job
interface EmbeddingJob {
nodeType: string;
nodeId: string;
content: string;
}
// worker.ts - Process embedding jobs
import { createStore } from "@nicia-ai/typegraph";
async function processEmbeddingJob(job: EmbeddingJob) {
const { nodeType, nodeId, content } = job;
// Generate embedding (expensive operation)
const embedding = await openai.embeddings.create({
model: "text-embedding-ada-002",
input: content,
});
// Update the node
const collection = store.nodes[nodeType as keyof typeof store.nodes];
await collection.update(nodeId, {
embedding: embedding.data[0].embedding,
});
}
// api-handler.ts - Enqueue jobs on create/update
async function createDocument(data: DocumentInput) {
// Create node without embedding (fast)
const doc = await store.nodes.Document.create({
title: data.title,
content: data.content,
// embedding: undefined - will be populated by worker
});
// Enqueue embedding job (non-blocking)
await jobQueue.add("generate-embedding", {
nodeType: "Document",
nodeId: doc.id,
content: data.content,
});
return doc;
}
```
**Batch processing for bulk imports:**
```typescript
async function backfillEmbeddings(batchSize = 100) {
let processed = 0;
while (true) {
// Find nodes missing embeddings
const nodes = await store
.query()
.from("Document", "d")
.whereNode("d", (d) => d.embedding.isNull())
.select((ctx) => ({
id: ctx.d.id,
content: ctx.d.content,
}))
.limit(batchSize)
.execute();
if (nodes.length === 0) break;
// Batch embed
const embeddings = await openai.embeddings.create({
model: "text-embedding-ada-002",
input: nodes.map((n) => n.content),
});
// Batch update
await store.transaction(async (tx) => {
for (const [i, node] of nodes.entries()) {
await tx.nodes.Document.update(node.id, {
embedding: embeddings.data[i].embedding,
});
}
});
processed += nodes.length;
console.log(`Processed ${processed} documents`);
}
}
```
**When to use:**
- Embedding generation is slow (100-500ms per call)
- You want fast API response times
- Bulk importing existing content
- Retry logic for API failures
**Considerations:**
- Handle job failures and retries
- Consider rate limits on embedding APIs
- Queries on `embedding` should handle null values during population
## Testing
For test setup patterns, seed data strategies, and profiler-based index coverage checks,
see the dedicated [Testing](/testing) guide.
## Deployment Patterns
### Edge and Serverless
Deploy TypeGraph at the edge using SQLite-compatible runtimes.
> **Note:** Edge environments cannot use `@nicia-ai/typegraph/adapters/drizzle/sqlite/local`
> because it depends on `better-sqlite3`, a native Node.js addon. Instead, use
> `@nicia-ai/typegraph/adapters/drizzle/sqlite` which is driver-agnostic.
**Cloudflare Durable Objects (SQLite) — transactional:**
```typescript
import { drizzle } from "drizzle-orm/durable-sqlite";
import { createAdapterStoreWithSchema } from "@nicia-ai/typegraph";
import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
export class GraphObject {
constructor(private ctx: DurableObjectState) {}
async fetch(): Promise {
const db = drizzle(this.ctx.storage);
const backend = createSqliteBackend(db); // auto-detects "do-sqlite"
const [store] = await createAdapterStoreWithSchema(graph, backend);
// Atomic across TypeGraph + the caller's own relational tables:
await store.transaction(async (tx) => {
await tx.nodes.Document.update(documentId, props);
if (tx.sqlAvailability !== "available") {
throw new Error(`Native transaction unavailable: ${tx.sqlAvailability}`);
}
const sqlTx = tx.sql;
await sqlTx.insert(documentVersions).values(versionRow);
});
return new Response("ok");
}
}
```
Unlike D1, Durable Objects expose an interactive storage transaction runner,
so `store.transaction()` / `store.withTransaction()` are fully atomic
(`capabilities.execution.interactiveTransactions: true`). See
[Backend Setup](/backend-setup#cloudflare-durable-objects-sqlite) and the
[Cross-Store Transactions recipe](/recipes#cross-store-transactions-drizzle--typegraph).
**Cloudflare Workers with D1:**
```typescript
// worker.ts
import { drizzle } from "drizzle-orm/d1";
import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
export default {
async fetch(request: Request, env: Env): Promise {
const db = drizzle(env.DB);
const backend = createSqliteBackend(db);
const store = createStore(graph, backend);
// Handle request with graph queries
const results = await store
.query()
.from("Document", "d")
.whereNode("d", (d) => d.embedding.similarTo(queryEmbedding, 5))
.select((ctx) => ctx.d)
.execute();
return Response.json(results);
},
};
```
**Turso (libSQL):**
```typescript
import { createClient } from "@libsql/client";
import { drizzle } from "drizzle-orm/libsql";
import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
const client = createClient({
url: process.env.TURSO_DATABASE_URL!,
authToken: process.env.TURSO_AUTH_TOKEN,
});
const db = drizzle(client);
const backend = createSqliteBackend(db);
const store = createStore(graph, backend);
```
> For Turso and D1, use [drizzle-kit managed migrations](#drizzle-kit-managed-migrations-recommended)
> to set up the schema.
**Bun with built-in SQLite:**
Bun runs locally, so you can use the Node.js-compatible path with better-sqlite3, or
use bun:sqlite with drizzle-kit managed migrations:
```typescript
import { Database } from "bun:sqlite";
import { drizzle } from "drizzle-orm/bun-sqlite";
import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
const sqlite = new Database("app.db");
const db = drizzle(sqlite);
const backend = createSqliteBackend(db);
const store = createStore(graph, backend);
```
> Use [drizzle-kit managed migrations](#drizzle-kit-managed-migrations-recommended)
> to set up the schema with bun:sqlite.
**When to use:**
- Low-latency requirements (data close to users)
- Serverless functions with graph queries
- Read-heavy workloads
**Considerations:**
- SQLite limitations (single-writer, no pgvector)
- Cold start times include DB initialization
- Vector search (cosine/L2): sqlite-vec on the local better-sqlite3 backend;
libSQL's built-in vectors on the libSQL / Turso backend
### Per-Request Connections (Cache the Verified Store)
Some serverless Postgres setups — Cloudflare Workers behind Hyperdrive, or any
platform that pools connections for you — want a **fresh connection per
request**. `createVerifiedAdapterStore` reconciles the committed schema and
checks index materialization at open time (a few `SELECT`s), so re-opening a
verified store on every request adds that cost to every graph-backed route.
Verify **once per isolate**, then build a zero-query store per request from the
cached reconciled schema:
```typescript
import { createAdapterStore, createVerifiedAdapterStore, getCommittedSchemaVersion, type GraphBackend, type ReconciledSchema } from "@nicia-ai/typegraph";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
type Cached = {
reconciled: ReconciledSchema;
version: number | undefined;
};
// Per-isolate cache, plus the in-flight reconciliation. Memoizing the promise
// collapses concurrent cold (or stale) requests onto ONE verify instead of each
// running its own — otherwise a burst of first requests reproduces the fan-out
// stampede this pattern exists to avoid.
let cached: Cached | undefined;
let inFlight: Promise | undefined;
function reconcileOnce(verifyBackend: GraphBackend): Promise {
inFlight ??= (async () => {
const [store, result] = await createVerifiedAdapterStore(graph, verifyBackend);
cached = {
reconciled: store.reconciledSchema,
version: result.status === "unchanged" ? result.version : store.reconciledSchema.version,
};
return cached;
})().finally(() => {
inFlight = undefined;
});
return inFlight;
}
export default {
async fetch(request: Request, env: Env): Promise {
const backend = createPostgresBackend(newPoolForThisRequest(env));
// Cold start: concurrent first requests all await the same reconciliation.
let snapshot = cached ?? (await reconcileOnce(backend));
// A one-row probe detects a schema commit from another isolate; a moved
// version refreshes through the same single-flight path.
const committed = await getCommittedSchemaVersion(backend, graph.id);
if (committed !== snapshot.version) snapshot = await reconcileOnce(backend);
// Zero database round-trips. Reads and writes still validate against
// runtime-committed kinds carried by the reconciled snapshot.
const store = createAdapterStore(graph, backend, { reconciled: snapshot.reconciled });
const results = await store
.query()
.from("Document", "d")
.select((ctx) => ctx.d)
.execute();
return Response.json(results);
},
};
```
`store.reconciledSchema` is an opaque snapshot of the reconciled graph
(compile-time kinds folded with any runtime-committed kinds) plus the committed
version it reflects. `createAdapterStore(graph, backend, { reconciled })` issues
**no** queries and validates writes against that snapshot, so kinds committed at
runtime remain writable without re-verifying. If you already hold a verified
store and only need to swap the connection, `store.withBackend(freshBackend)`
returns an equivalent store bound to the new connection with no re-verify.
The `getCommittedSchemaVersion` probe is your read-your-writes seam: one
round-trip, far cheaper than the full verified open (which also reconciles the
schema and checks index materialization), and re-verify only fires when the
version actually moves. Skip the probe only if your schema changes exclusively
during a deployment that also clears the cache. Otherwise reads may use the
stale schema snapshot, and the write fence rejects managed writes until the
cache is refreshed.
### Read Replica Separation
Route heavy graph queries to read replicas while writes go to primary.
```typescript
import { Pool } from "pg";
import { drizzle } from "drizzle-orm/node-postgres";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
// Primary for writes
const primaryPool = new Pool({
connectionString: process.env.PRIMARY_DATABASE_URL,
max: 10,
});
const primaryDb = drizzle(primaryPool);
const primaryBackend = createPostgresBackend(primaryDb);
const primaryStore = createStore(graph, primaryBackend);
// Replica for reads
const replicaPool = new Pool({
connectionString: process.env.REPLICA_DATABASE_URL,
max: 50, // Higher pool for read-heavy workloads
});
const replicaDb = drizzle(replicaPool);
const replicaBackend = createPostgresBackend(replicaDb);
const replicaStore = createStore(graph, replicaBackend);
// Route based on operation
export const stores = {
write: primaryStore,
read: replicaStore,
};
// Usage
async function searchDocuments(query: string) {
// Read from replica
return stores.read
.query()
.from("Document", "d")
.whereNode("d", (d) => d.embedding.similarTo(queryEmbedding, 10))
.select((ctx) => ctx.d)
.execute();
}
async function createDocument(data: DocumentInput) {
// Write to primary
return stores.write.nodes.Document.create(data);
}
```
**When to use:**
- Heavy read workloads (semantic search, graph traversals)
- Write/read ratio is heavily skewed toward reads
- Need to scale read capacity independently
**Considerations:**
- Replication lag means reads may be slightly stale
- Don't use replica for read-after-write scenarios
- Monitor replication lag in production
### Multi-Tenant Architecture
Four approaches for multi-tenant deployments, each with different tradeoffs.
#### Option 1: Shared tables with tenant isolation (simplest)
```typescript
import { defineNode, defineGraph } from "@nicia-ai/typegraph";
// Include tenantId in your node schemas
const Document = defineNode("Document", {
schema: z.object({
tenantId: z.string(),
title: z.string(),
content: z.string(),
}),
});
// Always filter by tenant in queries
function createTenantQuery(store: Store, tenantId: string) {
return {
searchDocuments: (query: string) =>
store
.query()
.from("Document", "d")
.whereNode("d", (d) => d.tenantId.eq(tenantId).and(d.embedding.similarTo(queryEmbedding, 10)))
.select((ctx) => ctx.d)
.execute(),
createDocument: (data: Omit) => store.nodes.Document.create({ ...data, tenantId }),
};
}
// Middleware extracts tenant and creates scoped API
function withTenant(req: Request) {
const tenantId = req.headers.get("x-tenant-id")!;
return createTenantQuery(store, tenantId);
}
```
#### Option 2: Separate `graph_id` per tenant (divergent schemas, one database)
Every TypeGraph row is keyed by `graph_id`, and so is the committed schema. Two
graphs with different `id`s coexist in one database with **independent schemas** —
tenant A can declare kinds tenant B has never heard of, and neither sees the
other's nodes, edges, or kind namespace.
```typescript
function tenantGraph(tenantId: string) {
return defineGraph({
id: `tenant_${tenantId}`, // the isolation boundary
nodes: { Document: { type: Document } },
edges: {},
});
}
// Each tenant commits — and evolves — its own schema, in the same database.
const [store] = await createStoreWithSchema(tenantGraph("acme"), backend);
```
What `graph_id` isolates:
- **Kinds** — the kind namespace is per `graph_id`. Declaring `Invoice` in one
graph does not create it in another.
- **Data** — nodes and edges are filtered by `graph_id` on every read and write.
- **Schema version and evolution** — each graph owns its committed schema
document and version, so tenants migrate independently.
This is the cheap alternative to N physical databases when you want divergent
per-tenant schemas without N connections. See
[Graph identity and the kind namespace](/schemas-stores#graph-identity-and-the-kind-namespace).
**Caveat 1 — index names are database-global.** Materialized SQL index names
are derived from `(kind, fields, shape)` and are **not** namespaced by `graph_id`,
because a SQL index name is a database-global identifier. Two graphs declaring
the same kind name *and* the same index therefore resolve to one physical index:
- **Same shape** — the second graph reuses the first graph's index. For a
graph-scoped index this is safe: the index is keyed by `graph_id` (or
`graph_id, kind`), so each graph still gets its own region of the index.
- **Different shape** (say one `unique`, one not) — materialization fails loudly
with a signature-drift error instead of silently sharing a mismatched index.
Rename the declaration, or drop the existing index and retry.
**Caveat 2 — a unique index must stay graph-scoped.** `scope` decides which
TypeGraph system columns prefix the index key:
| `scope` | Key prefix | Unique constraint applies |
|---------|-----------|---------------------------|
| `"graphAndKind"` (default) | `(graph_id, kind)` | per kind, per graph |
| `"graph"` | `(graph_id)` | per graph |
| `"none"` | *(none)* | **across every graph in the table** |
A `unique` index declared with `scope: "none"` omits `graph_id` from the key, so
the database enforces that value as unique across **all** graphs sharing the
table — one tenant's row will block another tenant's insert. That is a real
cross-tenant effect, and it holds whether or not two graphs share the physical
index. Keep unique indexes on the default `"graphAndKind"` (or `"graph"`) scope
in a multi-graph database; reserve `scope: "none"` for non-unique indexes where
you deliberately want one index spanning every graph.
Subject to those two rules, per-`graph_id` isolation holds: reads and writes
stay filtered by `graph_id`, and the coupling is confined to physical index
reuse.
#### Option 3: Schema per tenant (PostgreSQL)
```typescript
import { sql } from "drizzle-orm";
async function createTenantStore(tenantId: string) {
const schemaName = `tenant_${tenantId}`;
// Create schema if not exists
await pool.query(`CREATE SCHEMA IF NOT EXISTS ${schemaName}`);
// Run migrations in tenant schema
await pool.query(`SET search_path TO ${schemaName}`);
await pool.query(generatePostgresMigrationSQL());
await pool.query(`SET search_path TO public`);
// Create Drizzle instance with schema
const db = drizzle(pool, { schema: { schemaName } });
const backend = createPostgresBackend(db);
return createStore(graph, backend);
}
// Cache tenant stores
const tenantStores = new Map();
async function getTenantStore(tenantId: string): Promise {
if (!tenantStores.has(tenantId)) {
tenantStores.set(tenantId, await createTenantStore(tenantId));
}
return tenantStores.get(tenantId)!;
}
```
#### Option 4: Database per tenant (strongest isolation)
```typescript
interface TenantConfig {
id: string;
databaseUrl: string;
}
async function createTenantStore(config: TenantConfig) {
const pool = new Pool({ connectionString: config.databaseUrl });
await pool.query(generatePostgresMigrationSQL());
const db = drizzle(pool);
const backend = createPostgresBackend(db);
return {
store: createStore(graph, backend),
close: () => pool.end(),
};
}
// Connection manager with LRU eviction
class TenantConnectionManager {
private stores = new Map Promise }>();
private maxConnections = 100;
async getStore(tenantId: string): Promise {
if (!this.stores.has(tenantId)) {
if (this.stores.size >= this.maxConnections) {
await this.evictOldest();
}
const config = await fetchTenantConfig(tenantId);
this.stores.set(tenantId, await createTenantStore(config));
}
return this.stores.get(tenantId)!.store;
}
private async evictOldest() {
const [oldestId, oldest] = this.stores.entries().next().value;
await oldest.close();
this.stores.delete(oldestId);
}
}
```
**Comparison:**
| Approach | Isolation | Complexity | Scaling | Cost |
| ------------------- | --------------- | ---------- | --------------------------- | ------- |
| Shared tables | Low (row-level) | Low | Single DB | Lowest |
| Schema per tenant | Medium | Medium | Single DB, separate schemas | Low |
| Database per tenant | High | High | Independent DBs | Highest |
**When to use each:**
- **Shared tables**: SaaS with many small tenants, cost-sensitive
- **Schema per tenant**: Moderate isolation needs, PostgreSQL only
- **Database per tenant**: Enterprise customers requiring data isolation, compliance requirements
## Next Steps
- [Quick Start](/getting-started) - Basic setup and first graph
- [Semantic Search](/semantic-search) - Vector embeddings and similarity
- [Performance](/performance/overview) - Optimization strategies
# Graph Interchange
> Import and export graph data for backups, migrations, and external integrations
TypeGraph provides a standardized interchange format for importing and exporting
graph data. Use it for:
- Backing up and restoring graph data
- Migrating data between environments
- Exchanging data with external systems
## Quick Start
```typescript
import { importGraph, exportGraph, GraphDataSchema } from "@nicia-ai/typegraph/interchange";
// Export your graph
const backup = await exportGraph(store);
// Import into another store
const result = await importGraph(targetStore, backup, {
onConflict: "update",
onUnknownProperty: "strip",
});
console.log(`Imported ${result.nodes.created} nodes, ${result.edges.created} edges`);
```
## Interchange Format
The interchange format is a JSON structure validated by Zod schemas. You can use
`GraphDataSchema` to validate data before import, or export the schema as JSON
Schema for API documentation.
```typescript
import { GraphDataSchema } from "@nicia-ai/typegraph/interchange";
// Validate incoming data
const validated = GraphDataSchema.parse(jsonData);
// Export as JSON Schema for API docs
import { toJSONSchema } from "zod";
const jsonSchema = toJSONSchema(GraphDataSchema);
```
### Format Structure
```typescript
interface GraphData {
formatVersion: "2.0";
exportedAt: string; // ISO datetime
source: {
type: "typegraph-export" | "external";
// Additional source-specific fields
};
nodes: Array<{
kind: string;
id: string;
properties: Record;
validFrom?: string | null;
validTo?: string;
meta?: {
version?: number;
createdAt?: string;
updatedAt?: string;
};
}>;
edges: Array<{
kind: string;
id: string;
from: { kind: string; id: string };
to: { kind: string; id: string };
properties: Record;
validFrom?: string | null;
validTo?: string;
meta?: {
createdAt?: string;
updatedAt?: string;
};
}>;
identity?: {
profile: "typegraph-identity-v1";
mode: "state" | "archival";
assertions: Array<{
id: string;
relation: "same" | "different";
a: { kind: string; id: string };
b: { kind: string; id: string };
validFrom: string;
validTo?: string;
}>;
};
}
```
`validFrom` has three states: the key **absent** means it wasn't requested
(`includeTemporal: false`, the default) — import defaults it to the
import's own creation timestamp, unless the record also states a `validTo`
at or before that instant, in which case it is imported with no lower bound
("ended at T, start unknown") rather than one past its own end. An
**explicit `null`** means the source row is confirmed to have no lower bound
(open-left validity) — import preserves that instead of re-stamping it. A
**string** is an explicit value, carried through unchanged.
### Format Version Compatibility
Exports always write `formatVersion: "2.0"`. The read side — both
`importGraph`/`importGraphStream` and `GraphDataSchema.parse` — additionally
accepts `"1.0"`. A 1.0 document is structurally a valid 2.0 document: the only
2.0 change is the additive optional `identity` section, so pre-existing 1.0
exports validate and import unchanged. You never need to rewrite the version
field of an older backup; validation and import handle both.
## Exporting Data
Use `exportGraph` to serialize your graph data:
```typescript
import { exportGraph } from "@nicia-ai/typegraph/interchange";
// Export everything
const fullExport = await exportGraph(store);
// Export specific node kinds
const peopleOnly = await exportGraph(store, {
nodeKinds: ["Person", "Organization"],
});
// Export specific edge kinds
const relationshipsOnly = await exportGraph(store, {
edgeKinds: ["worksAt", "knows"],
});
// Include metadata (version, timestamps)
const withMeta = await exportGraph(store, {
includeMeta: true,
});
// Include temporal fields (validFrom, validTo)
const withTemporal = await exportGraph(store, {
includeTemporal: true,
});
// Include soft-deleted records
const withDeleted = await exportGraph(store, {
includeDeleted: true,
});
// Identity-enabled graphs export current assertions by default.
// Include ended assertion history explicitly:
const archival = await exportGraph(store, {
identityMode: "archival",
});
// A self-contained archive pairs archival identity with includeDeleted:
const selfContainedArchive = await exportGraph(store, {
identityMode: "archival",
includeDeleted: true,
});
```
**Archival identity and soft-deleted endpoints:** `identityMode: "archival"`
also exports *ended* assertions, and an ended assertion can reference an
endpoint that was later soft-deleted. A default export (`includeDeleted:
false`) joins every assertion against its endpoints' live rows, so an
assertion touching a soft-deleted endpoint is silently **dropped from the
export** — not carried with a dangling reference. This is silent archive
loss, not a dangling-endpoint problem. When the archive must stand alone
(backup, cold storage), pair it with `includeDeleted: true` so those
assertions and their endpoints travel with it.
That pairing has its own honest trade-off: the interchange format has no
`deletedAt` field, so a node included only because of `includeDeleted: true`
carries no record that it was deleted. Re-importing that archive resurrects
the node as **live**. Choose deliberately: without `includeDeleted`, a
backup silently loses soft-deleted endpoints and the assertions referencing
them; with it, those nodes come back alive on restore.
On import, every ended assertion's endpoints must exist as node rows in the
target (soft-deleted rows qualify) — historical reads conduct identity
through ended assertions, so an endpoint that never existed would become a
phantom bridge joining real nodes at past coordinates. The store's own
exports satisfy this by construction; a hand-built document that fails it is
recorded as an `entityType: "identity"` entry in `result.errors`.
### Export Options
| Option | Type | Default | Description |
|--------|------|---------|-------------|
| `nodeKinds` | `string[]` | all | Filter to specific node types |
| `edgeKinds` | `string[]` | all | Filter to specific edge types |
| `includeMeta` | `boolean` | `false` | Include version and timestamps |
| `includeTemporal` | `boolean` | `false` | Include validFrom/validTo fields |
| `includeDeleted` | `boolean` | `false` | Include soft-deleted records |
| `identityMode` | `"state" \| "archival"` | `"state"` | Export current identity assertions, or current plus ended assertions |
| `signal` | `AbortSignal` | none | Cancel the export: roll its snapshot transaction back and release the connection. See [Cancelling an export](#cancelling-an-export) |
`exportGraphStream` also accepts `idleTimeoutMs`, a positive integer with no
default. It bounds how long a delivered chunk may remain unacknowledged before
the stream settles itself. The clock stops as soon as the consumer requests the
next chunk, so a slow database read does not count as consumer idleness.
**Round-trip caveat:** with the default `includeTemporal: false`, exported
records carry no `validFrom`/`validTo`. On import, an omitted `validFrom`
defaults to the *import's own* creation timestamp (a born-already-ended record
is the exception noted above, and keeps no lower bound) — so a plain
`exportGraph` + `importGraph` round trip does **not** reproduce the
source's original valid-time window; every imported record becomes valid
from import time forward. Pass `includeTemporal: true` on export when the
clone needs to match the source's `asOf` behavior exactly (this is what
`branch()` does internally).
Identity-enabled graphs are the exception: their exports default temporal
fields on, because assertion windows cannot be validated against endpoints
without the endpoint bounds. Explicit `includeTemporal: false` is refused for
those graphs.
**Repair a legacy graph before exporting it.** A row an older library version
stored with a backwards window (`valid_from > valid_to`) exports as it is stored
and is then refused **per row** on re-import, because the import validates the
stated pair — so an unrepaired graph does not round-trip. Run
[`repairInvertedValidityWindows`](/schema-management#repairing-inverted-validity-windows)
first; it normalizes those rows to the open-left shape import accepts.
**`includeTemporal: true` with `onConflict: "update"`:** an update leg sends the
document's `validTo` and never its `validFrom`, because a live row's lower bound
is history. A document whose `validFrom` names a different instant than the
target row holds is therefore stating a bound the import will not apply, and that
row is reported as a per-row error carrying
[`IMMUTABLE_VALIDITY_LOWER_BOUND`](/errors/#immutable_validity_lower_bound)
rather than updated under a bound it ignored. This is reachable whenever a
temporal export is replayed over rows that were created separately — the same
document imported into a fresh graph creates those rows with their stated bounds
and is unaffected. To update props over existing rows from a temporal export,
either omit `validFrom` from the update document, export with
`includeTemporal: false`, or import into a fresh graph and swap it in.
### Cancelling an export
On a backend reporting `capabilities.execution.interactiveTransactions`, an export holds one
repeatable-read snapshot transaction for its whole life, and on a
single-connection backend it holds that connection's exclusive
interchange-stream lease with it. (A backend without transactions — SQLite
`transactionMode: "none"`, the session-less HTTP Postgres drivers — opens
neither: its export paginates statement by statement, so a write committed
mid-stream can appear in the pages that follow. That is a declared capability
gap, not something the stream papers over.) Every *cooperative* exit gives both back,
because each one runs the stream's `finally`: `break` or `throw` out of a `for
await`, and an explicit `iterator.return()`.
A consumer that pulls `next()` and then simply **drops the iterator** has no
cooperative exit. Async-generator `finally` blocks do not run on garbage
collection, so that snapshot transaction stays open for the life of the process
— and on a serialized connection every later export and every later import is
then refused for a stream nobody is reading. If you might abandon an iterator,
pass a `signal` or configure an idle timeout:
```typescript
const controller = new AbortController();
const iterator = exportGraphStream(store, {
batchSize: 1000,
signal: controller.signal,
idleTimeoutMs: 30_000,
})[Symbol.asyncIterator]();
try {
for (;;) {
const next = await Promise.race([
iterator.next(),
deadline(30_000), // resolves to a sentinel, leaving the pull in flight
]);
if (next === TIMED_OUT) {
// Do NOT just walk away: this is the leak. Aborting rolls the snapshot
// back and frees the connection.
controller.abort(new Error("export deadline exceeded"));
break;
}
if (next.done === true) break;
await write(next.value);
}
} finally {
controller.abort();
}
```
Aborting rejects the pull that is in flight — and any later pull from a consumer
that walked away and came back — with
[`ExportStreamCancelledError`](/errors#exportstreamcancellederror), carrying the
signal's own reason as `cause`, so a cancelled export is never mistaken for a
complete one. The message states what was actually settled: a snapshot rolled
back and a connection released on a transactional backend, or merely abandoned
reads on one that never held either. Aborting a signal *before* the first pull refuses the export
outright: no transaction is opened and no lease claimed. Aborting one that has
already finished does nothing, so a single controller can safely span a whole
job. `exportGraph` accepts `signal` too — there it simply makes the call reject
instead of running to completion.
When `idleTimeoutMs` expires, a later pull rejects with
[`ExportStreamIdleTimeoutError`](/errors#exportstreamidletimeouterror), with code
`"INTERCHANGE_EXPORT_STREAM_IDLE_TIMEOUT"`. Its `details.graphId` and
`details.idleTimeoutMs` identify the stream and configured bound. The timeout is
stream-only: `exportGraph` owns and promptly advances its internal consumer, so
it accepts `signal` but not `idleTimeoutMs`. The option has no default because a
stream may intentionally spend an unbounded amount of time processing a chunk;
callers that cannot choose a safe idle bound should retain an `AbortController`
and abort on their own job deadline instead.
There is deliberately no garbage-collection fallback. A `FinalizationRegistry`
cannot close this gap: any cleanup state able to settle an abandoned stream has
to reach the stream's internals, and a registry holds its state strongly, so
doing so would keep the abandoned stream reachable and the finalizer would never
run. Explicit cancellation and the idle timeout are the mechanisms.
## Importing Data
Use `importGraph` to load data into a store:
```typescript
import { importGraph } from "@nicia-ai/typegraph/interchange";
const result = await importGraph(store, data, {
onConflict: "update",
onUnknownProperty: "strip",
validateReferences: true,
batchSize: 1000,
});
if (result.success) {
console.log(`Created: ${result.nodes.created} nodes, ${result.edges.created} edges`);
console.log(`Updated: ${result.nodes.updated} nodes, ${result.edges.updated} edges`);
console.log(`Skipped: ${result.nodes.skipped} nodes, ${result.edges.skipped} edges`);
console.log(`Identity: ${result.identity.created} created, ${result.identity.skipped} skipped`);
} else {
console.error("Import had errors:", result.errors);
}
```
### Durable edge match identities during import
The target graph declaration, not the interchange document, determines an
edge's durable `matchIdentity`. After validating and normalizing an incoming
edge's properties, normal import builds the same canonical endpoint/property
key used by collection creates and writes it with the edge row. An import
therefore cannot bypass convergence by choosing a new edge id.
An incoming create whose durable identity is already owned by a different row,
or was claimed by an earlier edge in the same import slice, is recorded as a
per-edge error using `EdgeMatchIdentityConflictError`. This decision is
separate from `onConflict`, which handles an existing row with the incoming
edge's own id; `skip` and `update` do not authorize taking another row's durable
identity.
Bundled backends discover existing owners with set-oriented exact endpoint-pair
reads and insert claimless durable slices with bind-budgeted,
conflict-arbitrated batch statements. A slice that also needs cardinality
claims remains atomic. If an exceptional batch refusal cannot identify the
losing row, TypeGraph rolls the slice back to a savepoint and retries its rows
individually so `result.edges.created` and `result.errors` remain honest. A
custom or non-transactional backend that cannot prove that rollback refuses
the import with `IMPORT_EDGE_BATCH_RETRY_REQUIRES_SAVEPOINT`; retrying a
possibly committed prefix could otherwise double-count or misattribute rows.
When the document carries an `identity` section, `result.identity` reports
`{ created, skipped }` counts for imported assertions (skipped covers an
exact re-import of an assertion that already exists under the same id).
A rejected assertion — an unknown endpoint, a contradiction against the
target's existing identity truth, or a reused assertion id that names
different truth — is recorded in `result.errors` with `entityType:
"identity"`, `kind` set to the assertion's relation (`"same"` or
`"different"`), and `id` set to the assertion id, mirroring how node/edge
errors carry `kind`/`id`.
**Partial-commit caveat:** identity assertions are applied one at a time, and
a mid-batch failure (a contradiction or id conflict partway through the
`identity.assertions` array) stops the identity import but does not roll back
the assertions already applied before it — they remain committed. They are
**not** reflected in `result.identity.created`, since that count is only
reported on success; the count under-reports rather than invents a number for
committed-but-unaccounted work. The failure that stopped the batch is the one
error entry you see in `result.errors`.
### Import Options
| Option | Type | Default | Description |
|--------|------|---------|-------------|
| `onConflict` | `"skip" \| "update" \| "error"` | required | How to handle existing entities |
| `onUnknownProperty` | `"error" \| "strip" \| "allow"` | `"error"` | How to handle extra properties |
| `validateReferences` | `boolean` | `true` | Verify edge endpoints exist |
| `batchSize` | `number` | `1000` | Batch size for database operations. Each batch pays fixed per-round-trip costs, so undersized batches slow client/server imports; inserts are still split by the driver bind budget internally. |
### Trusted initial import
`trustedImportGraph` and `trustedImportGraphStream` are a separate,
intentionally trusted path for loading a fresh dedicated database. They do not
turn off validation on `importGraph`; they bypass the normal store write
pipeline entirely.
```typescript
import {
trustedImportGraphStream,
type GraphInterchangeChunk,
} from "@nicia-ai/typegraph/interchange";
async function* chunks(): AsyncIterable {
yield { type: "header", header };
for await (const nodes of readNodeBatches()) {
yield { type: "nodes", nodes };
}
for await (const edges of readEdgeBatches()) {
yield { type: "edges", edges };
}
}
const result = await trustedImportGraphStream(store, chunks());
console.log(result); // { nodes: 1000000, edges: 5000000 }
```
The contract is deliberately narrow:
- The TypeGraph node and edge tables must be globally empty. A different graph
in the same database also makes the database non-empty.
- The caller guarantees property shapes, endpoint existence, edge endpoint
types, cardinality, duplicate-free IDs, and duplicate-free durable edge match
identities. Only stream ordering and known kind names are checked. Trusted
import still derives each declared edge identity and stores it with the row,
so a collision reaches the database arbiter and rolls back the complete
trusted-import transaction.
- Recorded-time history, revision tracking, node uniqueness constraints,
`searchable()` fields, and `embedding()` fields are rejected in this first
version because their sidecar writes would otherwise be skipped.
- Operational Identity-enabled target stores are rejected with
`details.reason === "identity_unsupported"`; identity-bearing input is
rejected with `details.reason === "invalid_stream"`. The trusted session
writes only the node and edge relations, so it cannot persist assertions or
materialize the derived closure — refusing both cases keeps identity truth
from being silently dropped. Use `importGraphStream` for an export that
carries Operational Identity assertions. This restriction does not apply to
an edge registration's `matchIdentity`, which trusted import materializes in
the edge relation itself.
- Nodes must precede edges. The `meta` timestamps and node version in an
interchange row are not restored; the import creates new storage metadata.
- The complete stream is one transaction. Data insertion, temporary secondary
index removal, index rebuilding, and planner statistics either all commit or
all roll back.
- A schema-managed Store acquires and validates its schema-write fence inside
that transaction before loading rows. A stale managed import fails before
row DML; a raw Store remains explicitly outside the schema-fencing guarantee.
Supported native paths are synchronous prepared-statement SQLite
(`better-sqlite3` and Bun SQLite) and transaction-capable PostgreSQL adapters
with raw execution support (including node-postgres, postgres.js, and PGlite).
Remote libSQL/Turso, D1, and HTTP-only PostgreSQL adapters reject the call with
`TrustedImportError` and `details.reason === "backend_unsupported"`.
Use `importGraph`/`importGraphStream` for external or uncertain data, conflict
handling, incremental loads, and any graph with the unsupported features above.
Use collection `bulkInsert` when the data is trusted but the database is not a
fresh dedicated target.
### Conflict Strategies
**`skip`** - Keep existing data, ignore incoming:
```typescript
// Useful for incremental imports where you don't want to overwrite
await importGraph(store, data, { onConflict: "skip" });
```
**`update`** - Merge incoming data into existing:
```typescript
// Useful for syncing updates from an external source
await importGraph(store, data, { onConflict: "update" });
```
**`error`** - Fail if any entity already exists:
```typescript
// Useful for initial imports where duplicates indicate a problem
await importGraph(store, data, { onConflict: "error" });
```
#### An edge id held by a different edge
Edge ids are unique per graph, but the import's existence probe (`getEdge` /
`getEdges`) is keyed on `(graph_id, id)` alone. So a document edge whose id is
already held by a row with a different **immutable identity** — its `kind` or
either of its endpoints — finds that row. That question — *is this the same
edge?* — is prior to *what do we do about the same edge?*, so it is answered
**before** the conflict strategy and all three strategies answer alike: the row
is reported as a per-row entry in `result.errors`, whose `error` message is
prefixed `INTERCHANGE_EDGE_KIND_CONFLICT` and names each component that differs
alongside the value the document stated. The stored row is left untouched.
```typescript
const result = await importGraph(store, data, { onConflict: "update" });
const identityConflicts = result.errors.filter((entry) =>
entry.error.startsWith("INTERCHANGE_EDGE_KIND_CONFLICT"),
);
```
One prefix covers the whole class rather than a second one for endpoint
mismatches: the condition is a single fact and the recovery is a single action,
and a caller that had to match two prefixes to catch one condition would
eventually match only one. The token still reads `…_KIND_CONFLICT` because it is
the published, branchable string; it now covers every identity component.
`ImportError` carries no `code` field, so the message prefix is the branchable
token — the same `CODE: message` idiom the validity-window import refusals use.
Give the incoming edge a distinct id, or import it under the identity the stored
row already carries.
Previously both non-`error` strategies were silent about this: `update` wrote
the incoming edge's properties onto the *other* row with nothing in
`result.errors`, and `skip` counted the document's edge as already present when
no matching edge existed anywhere — so it was never created and never reported.
Comparing `kind` alone closed only half of it: because endpoints are immutable,
a document naming the incumbent's kind and id but different endpoints still read
as the same edge, so `update` overwrote the incumbent's properties and silently
retained its old endpoints. The update is additionally issued with all five
identity components in the statement's own `WHERE`, so the check cannot be raced
by a concurrent hard-delete-and-recreate; an update that consequently matches no
row is reported as the same per-row error rather than aborting the import.
Nodes were never affected: their probe is `getNode(graphId, kind, id)`, which is
kind-scoped, so a cross-kind id collision simply reads as absent.
#### An update target that changed under the import
`onConflict: "update"` is a read-then-write pair: the import probes the stored
row, validates the document's validity window against that row's `valid_from`,
and then writes. Every part of that verdict is restated in the UPDATE's own
`WHERE` — for edges the five identity components above, and for **both** nodes
and edges the effective validity lower bound, whenever the window check actually
read it. A concurrent hard-delete-and-recreate between the probe and the write
therefore matches no row instead of landing a decision computed for a row that
is gone (which would have ignored a `validFrom` the document stated, or
persisted a `validTo` below the new row's `validFrom`).
The bound is read — and so restated — when the document states a `validFrom` to
compare against it, or a lone `validTo` to check for an inverted window. A
document that states **neither** makes no claim about the row's window, so its
properties update is not fenced on the bound and a concurrent recreate that only
moved the bound does not refuse it. This matches `store.nodes.*.update` exactly:
a write asserts what its decision read, and nothing more.
A write that matches no row is reported per row, so an import whose earlier
rows are already written is not aborted for it:
```typescript
const result = await importGraph(store, data, { onConflict: "update" });
const raced = result.errors.filter(
(entry) =>
entry.error.startsWith("INTERCHANGE_NODE_UPDATE_TARGET_CHANGED") ||
entry.error.startsWith("INTERCHANGE_EDGE_KIND_CONFLICT"),
);
```
`INTERCHANGE_NODE_UPDATE_TARGET_CHANGED` is the node-side prefix;
edges reuse `INTERCHANGE_EDGE_KIND_CONFLICT`, whose message now also names the
validity lower bound. Re-export the source and retry.
A node update refused this way leaves no partial trace, and neither does one
refused for a uniqueness conflict. The row write and the uniqueness transition
are one unit: the new keys are claimed first (the claim is what decides the
conflict), the row write follows, and the old keys are released only once it
lands — with the claims given back if it does not. Fulltext and embedding
sidecars are written only after the row update reports a match. So a row that
`result.errors` reports is a row the import did not change, even though the
transaction around it commits.
### Unknown Property Handling
When importing data that has properties not defined in your schema:
**`error`** - Reject the import (default, safest):
```typescript
await importGraph(store, data, { onUnknownProperty: "error" });
// Throws if data has { name: "Alice", unknownField: "value" }
```
**`strip`** - Remove unknown properties silently:
```typescript
await importGraph(store, data, { onUnknownProperty: "strip" });
// { name: "Alice", unknownField: "value" } becomes { name: "Alice" }
```
**`allow`** - Pass through to storage:
```typescript
await importGraph(store, data, { onUnknownProperty: "allow" });
// Behavior depends on your database and schema strictness
```
## Backup and Restore
### Creating Backups
```typescript
import { exportGraph } from "@nicia-ai/typegraph/interchange";
import fs from "fs/promises";
async function createBackup(store: Store, backupDir: string) {
const timestamp = new Date().toISOString().replace(/[:.]/g, "-");
const filename = `backup-${timestamp}.json`;
const data = await exportGraph(store, {
includeMeta: true,
includeTemporal: true,
});
await fs.writeFile(
`${backupDir}/${filename}`,
JSON.stringify(data, null, 2)
);
return filename;
}
```
### Restoring from Backup
```typescript
import { importGraph, GraphDataSchema } from "@nicia-ai/typegraph/interchange";
import fs from "fs/promises";
async function restoreBackup(store: Store, backupPath: string) {
const json = await fs.readFile(backupPath, "utf-8");
const data = GraphDataSchema.parse(JSON.parse(json));
const result = await importGraph(store, data, {
onConflict: "update", // or "error" for clean restore
onUnknownProperty: "error",
});
if (!result.success) {
throw new Error(`Restore failed: ${result.errors.map(e => e.error).join(", ")}`);
}
return result;
}
```
## Migration Between Environments
Move data from development to staging, or staging to production:
```typescript
import { createStore } from "@nicia-ai/typegraph";
import { exportGraph, importGraph } from "@nicia-ai/typegraph/interchange";
import { graph } from "./schema";
async function migrateData(
sourceBackend: GraphBackend,
targetBackend: GraphBackend,
) {
const sourceStore = createStore(graph, sourceBackend);
const targetStore = createStore(graph, targetBackend);
// Export from source
const data = await exportGraph(sourceStore);
// Import to target
const result = await importGraph(targetStore, data, {
onConflict: "error", // Ensure clean migration
onUnknownProperty: "error",
validateReferences: true,
});
return result;
}
```
## Building Custom Import Pipelines
For complex import scenarios, you can build pipelines using the Zod schemas:
```typescript
import {
GraphDataSchema,
InterchangeNodeSchema,
InterchangeEdgeSchema,
type GraphData,
} from "@nicia-ai/typegraph/interchange";
// Transform external data to interchange format
function transformExternalData(externalRecords: ExternalRecord[]): GraphData {
const nodes = externalRecords.map((record) => ({
kind: "Document",
id: record.externalId,
properties: {
title: record.name,
content: record.body,
source: { system: "external", id: record.externalId },
},
}));
// Validate each node
const validatedNodes = nodes.map((node) => InterchangeNodeSchema.parse(node));
return {
formatVersion: "2.0",
exportedAt: new Date().toISOString(),
source: {
type: "external",
description: "Imported from external CMS",
},
nodes: validatedNodes,
edges: [],
};
}
```
## Error Handling
Import returns detailed error information for partial failures:
```typescript
const result = await importGraph(store, data, { onConflict: "error" });
if (!result.success) {
for (const error of result.errors) {
console.error(
`Failed to import ${error.entityType} ${error.kind}:${error.id}: ${error.error}`
);
}
// Decide how to handle partial import
if (result.nodes.created > 0 || result.edges.created > 0) {
console.log("Partial import completed, some entities were created");
}
}
```
### Serialized-connection refusals
Row-level failures are reported in `result.errors`, but one class of failure is
thrown instead: two long-lived interchange streams cannot share a single
serialized database connection, so whichever starts second is refused with a
typed `ConfigurationError`. The lease is exclusive — one stream of any kind per
connection — and every long-lived import claims it, so `importGraph`,
`importGraphStream`, `trustedImportGraph`, and `trustedImportGraphStream` can all
throw it, as can `exportGraphStream` when an import already holds the connection
— there, on the stream's first pull, since the claim begins when the snapshot
transaction opens rather than when the iterable is constructed. The codes (`INTERCHANGE_SHARED_SERIALIZED_BACKEND_SNAPSHOT`,
`INTERCHANGE_SAME_SQLITE_BACKEND_SNAPSHOT`,
`INTERCHANGE_SERIALIZED_IMPORT_IN_PROGRESS`) and the `details.requested` /
`details.heldBy` pairing they carry are documented in
[Interchange serialized-connection guard codes](/errors#interchange-serialized-connection-guard-codes).
Which connections count as serialized is read off the driver, so a driver
TypeGraph cannot identify is left unmarked and a stream pair on it can still
wedge. `createSqliteBackend` and `createPostgresBackend` take a
`serializedResource` declaration for both directions of that gap —
`{ mode: "shared", resource: client }` marks a connection TypeGraph cannot see,
`{ mode: "independent" }` escapes a detection that is wrong for your topology.
The escape hatch lifts the shared-resource refusal between two distinct backends
only: one SQLite backend exporting into **itself** stays refused with
`INTERCHANGE_SAME_SQLITE_BACKEND_SNAPSHOT`, because that is one handle holding
the one snapshot transaction its own import needs. That surviving refusal is
SQLite-only, so on PostgreSQL a backend declared independent may export into
itself. See
[Serialized connections](/backend-setup#serialized-connections).
## Best Practices
### Validate Before Import
Always validate external data before importing:
```typescript
import { GraphDataSchema } from "@nicia-ai/typegraph/interchange";
const result = GraphDataSchema.safeParse(untrustedData);
if (!result.success) {
console.error("Invalid data:", result.error.format());
return;
}
await importGraph(store, result.data, options);
```
### Use Transactions for Consistency
Import operations use transactions when the backend supports them. When the
Store carries a reconciled schema version, every import batch acquires and
validates the same schema-write fence as collection writes; a stale managed
Store fails before row DML. A raw Store remains outside that guarantee. On a raw
Store without transaction support, consider smaller batch sizes to minimize
partial-failure impact; a managed Store on that backend fails closed on its
first write.
### Test with `onConflict: "error"` First
When setting up a new import pipeline, use `onConflict: "error"` to catch
unexpected duplicates early:
```typescript
// Development/testing
await importGraph(store, data, { onConflict: "error" });
// Production (after validation)
await importGraph(store, data, { onConflict: "update" });
```
### Monitor Import Results
Log import statistics for observability:
```typescript
const result = await importGraph(store, data, options);
logger.info("Import completed", {
success: result.success,
nodesCreated: result.nodes.created,
nodesUpdated: result.nodes.updated,
nodesSkipped: result.nodes.skipped,
edgesCreated: result.edges.created,
edgesUpdated: result.edges.updated,
edgesSkipped: result.edges.skipped,
identityCreated: result.identity.created,
identitySkipped: result.identity.skipped,
errorCount: result.errors.length,
});
```
## Next Steps
- [Data Sync](/data-sync) - Patterns for keeping external data in sync
- [Schema Migrations](/schema-management) - Managing schema changes over time
- [Integration Patterns](/integration) - Database setup and deployment
# Limitations
> Known constraints and backend-specific limitations
This page documents TypeGraph's known limitations and constraints.
## Backends Without Atomic Transactions
Some runtimes cannot hold a multi-statement database session and therefore
cannot offer atomic transactions:
- **Cloudflare D1** — the D1 binding has no interactive transaction
primitive (`D1Database.batch(...)` is transactional but batch-only).
- **`drizzle-orm/neon-http`** — Neon's HTTP driver issues each statement as
an independent request; there is no session to bind a transaction to.
Cloudflare **Durable Objects** SQLite is *not* in this list: a store backed
by `drizzle(ctx.storage)` is auto-detected as `transactionMode: "do-sqlite"`,
reports `capabilities.execution.interactiveTransactions: true`, and is fully atomic. An
`AdapterStore` created from that backend also exposes the adapter-only
`store.withTransaction` and `tx.sql` surfaces. See
[Backend Setup](/backend-setup#cloudflare-durable-objects-sqlite).
These backends report `capabilities.execution.interactiveTransactions: false`. Read-only
`store.batch(...)` still runs, but each query may use an independent connection
and observe a different database snapshot. (Whether the queries nonetheless
reuse one connection is up to the adapter — the no-transaction path hands each
query the same backend object.) Note this is a difference of degree, not of
kind: on PostgreSQL, `batch()`'s implicit transaction runs at the default
read-committed isolation, so queries there can also observe interleaved commits.
Write behavior depends on how the Store was constructed. A schema-managed Store
fuses its schema fence into a write's own statement when the write fuses, and
fails closed for writes that need the transaction-scoped schema or constraint
fence otherwise — see
[The guard every fused write shares](#the-guard-every-fused-write-shares)
below for which writes fuse and which refuse. A raw `createStore()` /
`createAdapterStore()` without a reconciled snapshot still has no interactive
transaction boundary. `store.transaction(fn)` refuses with a typed capability
error rather than pretending to provide rollback; direct backend writes remain
raw. Eligible operations that use a certified atomic SQL program can still be
available on these roots, but that transport guarantee is separate from the
interactive transaction capability.
These backends cannot honor the `isolationLevel` option on
`store.transaction(...)`; the method refuses before invoking its callback, so
the collection-read snapshot recipe documented elsewhere does not apply here.
```typescript
// On a raw D1 / neon-http Store, this refuses before the callback runs.
await store.transaction(async (tx) => {
await tx.nodes.Person.create({ name: "Alice" });
});
```
**If you require atomicity or schema-version fencing, branch on the capability:**
```typescript
if (store.capabilities.execution.interactiveTransactions) {
await store.transaction(async (tx) => {
/* atomic */
});
} else {
// Use independent operations, or a supported certified atomic operation.
const person = await store.nodes.Person.create({ name: "Alice" });
const company = await store.nodes.Company.create({ name: "Acme" });
await store.edges.worksAt.create(person, company, { role: "Engineer" });
}
```
If you need atomic writes from an edge runtime, use
`drizzle-orm/neon-serverless` (WebSocket-backed Pool) instead of
`drizzle-orm/neon-http`.
### Four kinds of write atomicity
TypeGraph distinguishes an interactive transaction, a static adapter batch, a
certified atomic SQL program, and an authoritative one-statement command. `store.transaction(...)` is the
interactive Store API: it pins a session and groups the callback's operations.
A static batch is adapter-internal (such as D1 `batch()` or a multi-row
insert); it is not a public Store transaction and cannot make arbitrary Store
calls atomic. A certified atomic SQL program is a closed ordered statement
sequence whose transport preserves result slots and parameters and rolls back
primary and sidecar writes when a later statement fails. An authoritative command is a single `commands.execute` write
whose database statement returns the decision it made. It can provide a safe
transactionless create/found path only when the backend has a durable arbiter.
Operational Identity, single-edge claim/cardinality enforcement, and any
undeclared dynamic `matchOn` convergence that may write still require an
interactive transaction and fail closed on a backend that cannot provide one.
Outside the native durable-convergence envelope, an all-live
`ifExists: "return"` endpoint batch is read-only and can return from its
set-oriented root read without a transaction. Inside the native envelope, the
authoritative upsert program runs before the Store knows every identity is
live. It preserves the logical `"found"` result in one exchange, but may take
incumbent-row locks and produce write amplification. Eligible direct edge
batches on bundled roots are a separate exception: their closed native program
carries the claim sidecars inside one atomic exchange. A
declared edge `matchIdentity` persists
a canonical endpoint/property key and has a unique database arbiter; eligible
root `getOrCreateByEndpoints` calls can therefore use the authoritative
one-statement command. The durable identity does not make unrelated Store
operations, claims, or history/revision side effects transactionless.
### The guard every fused write shares
Every static batch and every certified atomic program asserts the active
schema version inside the very statement that writes, never as a preceding
check — the fused create's `WHERE … is_active` predicate, or the program's
leading `schema_fence` CTE. A stale version makes that statement match zero
rows, so the write commits nothing, and the store re-reads and reports
`StaleVersionError` instead of writing against a version that already moved
on. This is what lets a `"batch"`-tier backend (`capabilities.execution.unitOfWork === "batch"`
— Cloudflare D1's `batch()`, Neon HTTP's `transaction(queries)`, which fix
every statement before the first one runs and commit them together with no
session in between) run schema-managed creates, updates and deletes, and
bulk writes at all: the fence travels inside the one exchange it can hold,
instead of needing a session to hold it separately. A singleton node update,
`upsertById`, or delete fuses the same way as a create, through a one-entry
certified atomic program, whenever its kind carries no declared unique
constraint — except a node delete, which fuses even when the kind DOES
carry one, because the atomic delete program releases that claim in the
same statement. A singleton edge update or delete fuses the same way
(`EdgeCollection` has no `upsertById`).
A write that needs more than that one guarded statement — because it must
read a value it wrote earlier in the same write, hold an interactive
callback open across round trips, maintain Operational Identity's closure,
hold history's per-graph lock across a whole write cascade, or hold one
transaction across a schema commit's compare-and-swap — refuses on a
`"batch"`-tier backend with `BATCH_WRITE_UNSUPPORTED`, naming which of those
it needed:
| `reason` | What it needs |
| --- | --- |
| `interactive-callback` | Hold an interactive callback transaction open across several round trips (`store.transaction(fn)`). |
| `constraint-needs-probe` | Read a value it wrote earlier in the same write before deciding what to write next (a declared constraint's probe-then-write). |
| `identity` | Read and write Operational Identity's closure across several round trips inside one held transaction. |
| `history` | Hold the per-graph write lock and clock open across a whole write cascade (`history: true` / `revisionTracking: true`). |
| `schema-commit` | Hold one transaction across its compare-and-swap read and its activating write (`commitSchemaVersion` / `setActiveVersion`). |
A write that simply cannot fuse — an ineligible write kind, a singleton
create/update/`upsertById` on a kind with a declared unique constraint, a
tombstone-resurrection write a supplied id falls through to, or a derived
backend — refuses with `SCHEMA_WRITE_FENCE_UNSUPPORTED` instead and carries
no `batchRefusal` reason: that gate has no proven need to name, only its own
plain limitation.
See [`BATCH_WRITE_UNSUPPORTED`](/errors#batch_write_unsupported) for where
each reason surfaces in an error's `details`.
## libsql Single-Connection Transactions
For local `@libsql/client` connections (`file:` paths and `file::memory:`),
`createLibsqlBackend` frames transactions with raw `BEGIN IMMEDIATE`/`COMMIT`
statements on the client's single stable connection. It deliberately avoids
`client.transaction()`, which hands the client's connection to the transaction
and lazily opens a new one afterwards — for an in-memory database that new
connection is a fresh, empty database
([tursodatabase/libsql-client-ts#229](https://github.com/tursodatabase/libsql-client-ts/issues/229)).
In-memory databases therefore work for all operations, including transactions.
Remote Turso connections (`libsql://`, `http(s)://`) run each transaction on
its own stream via the driver.
The trade-off of a single connection: a store-level operation awaited from
**inside** a `store.transaction` callback (on the root store, rather than the
`tx` context) can never run — the open transaction occupies the backend's
serialized execution slot until it completes — so the backend rejects it with
a `ConfigurationError` instead of deadlocking.
```typescript
// ✅ In-memory works, including transactions
const client = createClient({ url: "file::memory:" });
// ❌ Root-store access inside a transaction callback throws
await store.transaction(async (tx) => {
await store.nodes.Person.find(); // ConfigurationError — use tx.nodes
await tx.nodes.Person.find(); // ✅ transaction-scoped access
});
```
## Recursive Traversal Depth
Variable-length traversals use two depth caps and an explicit cycle policy:
1. Unbounded traversals (no `maxHops` option) are capped at 10 hops.
2. Explicit `maxHops` values are validated up to 1000 hops (`maxHops: >1000` throws).
3. Cycle prevention is on by default. To skip cycle checks for speed, opt into
`cyclePolicy: "allow"` (which may revisit nodes across hops).
This prevents runaway queries while still supporting deep, intentionally bounded traversals.
```typescript
// Implicitly limited to 10 hops
store
.query()
.from("Person", "p")
.traverse("reportsTo", "e")
.recursive()
.to("Person", "manager");
// Explicit limits up to 1000 are honored
store
.query()
.from("Person", "p")
.traverse("reportsTo", "e")
.recursive({ maxHops: 200 }) // honored
.to("Person", "manager");
// Explicit limits above 1000 throw
store
.query()
.from("Person", "p")
.traverse("reportsTo", "e")
.recursive({ maxHops: 2000 }) // throws
.to("Person", "manager");
```
The unbounded-traversal limit is defined as `MAX_RECURSIVE_DEPTH`:
```typescript
import { MAX_RECURSIVE_DEPTH } from "@nicia-ai/typegraph";
// MAX_RECURSIVE_DEPTH = 10
```
## Connection Management
Managed Store factories own their local SQLite or PGlite connection, and their
`store.close()` method releases it. The local backend factories
`createLocalSqliteBackend` and `createLocalPgliteBackend` likewise expose an
owned backend whose `close()` releases its resources.
Bring-your-own adapter factories leave connection ownership with you. For
`createSqliteBackend`, `createPostgresBackend`, and `createLibsqlBackend`, you
are responsible for:
1. **Creating and configuring** the database connection
2. **Implementing connection pooling** for production use
3. **Closing connections** when done
```typescript
import Database from "better-sqlite3";
import { drizzle } from "drizzle-orm/better-sqlite3";
import { createSqliteBackend, generateSqliteMigrationSQL } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
// You manage the connection
const sqlite = new Database("app.db");
sqlite.exec(generateSqliteMigrationSQL());
const db = drizzle(sqlite);
const backend = createSqliteBackend(db);
const store = createStore(graph, backend);
// You close the connection
sqlite.close();
```
For production deployments, use connection pooling:
```typescript
import { Pool } from "pg";
import { drizzle } from "drizzle-orm/node-postgres";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
const pool = new Pool({
connectionString: process.env.DATABASE_URL,
max: 20, // Maximum connections
});
const db = drizzle(pool);
const backend = createPostgresBackend(db);
```
In the bring-your-own example above, `store.close()` leaves the supplied driver
open. Close that driver or pool through its own API.
## Predicate Serialization
Where predicates in unique constraints cannot be serialized. If you use schema
serialization for versioning or migration, predicates are stored as
`"[predicate]"` and cannot be reconstructed.
```typescript
// This predicate works at runtime...
unique({
name: "email_unique_when_active",
fields: ["email"],
where: (props) => props.status.isNotNull(),
});
// ...but serializes as:
// { "where": "[predicate]" }
```
**Workaround:** For full schema serialization support, avoid predicates in unique constraints.
Use application-level validation instead.
## Vector Search Backend Requirements
Vector and hybrid search work across all primary backends via a pluggable
`VectorStrategy`. Each backend advertises its capabilities through
`backend.capabilities.vector` (`{ supported, metrics, indexTypes, maxDimensions }`):
| Backend | Requirement | Metrics |
|---------|-------------|---------|
| PostgreSQL | pgvector extension (HNSW / IVFFlat) | cosine, l2, inner_product |
| SQLite | sqlite-vec extension (`vec0` KNN) | cosine, l2 |
| libSQL / Turso | built-in native engine (DiskANN); nothing to load | cosine, l2 |
| D1 | Not supported | — |
Note that `inner_product` is PostgreSQL-only — sqlite-vec and libSQL support
cosine and l2 only.
Using vector predicates on unsupported backends throws `UnsupportedPredicateError`:
```typescript
try {
await store
.query()
.from("Document", "d")
.whereNode("d", (d) => d.embedding.similarTo(queryVector, 10))
.select((ctx) => ctx.d)
.execute();
} catch (error) {
if (error instanceof UnsupportedPredicateError) {
// Vector search not available on this backend
}
}
```
## Query Builder Type Inference
Complex query chains may occasionally require explicit type annotations when TypeScript cannot
infer the full type. This is rare but can occur with deeply nested selects or unions.
```typescript
// If type inference fails, add explicit type
const results = await store
.query()
.from("Person", "p")
.select((ctx) => ({
name: ctx.p.name as string, // Explicit annotation
}))
.execute();
```
## Bulk Operation Limits
Bulk operations (`bulkCreate`, `bulkInsert`, `bulkUpsertById`, `bulkDelete`) have practical limits based on your database:
| Database | Recommended Batch Size |
|----------|----------------------|
| SQLite | 500-1000 items |
| PostgreSQL | 1000-5000 items |
For larger datasets, batch your operations:
```typescript
const BATCH_SIZE = 1000;
for (let i = 0; i < items.length; i += BATCH_SIZE) {
const batch = items.slice(i, i + BATCH_SIZE);
await store.nodes.Person.bulkCreate(batch);
}
```
### Native node bulk eligibility
Bundled PostgreSQL roots using a recognized session-capable driver, Neon HTTP,
Cloudflare D1, and libSQL roots can use one schema-fenced native atomic program
for schema-managed `nodes.bulkInsert()` and `nodes.bulkCreate()` calls when the
node has no Operational Identity, history, or revision work. The program can
compose fulltext/vector projections with the complete supported uniqueness and
disjointness claim set for every member. Session-capable
PostgreSQL executes that program on one pinned
transaction; Neon HTTP, D1, and libSQL submit one transport batch. Advertised
same-kind or hierarchy-wide uniqueness claims, disjointness claims, and mixed
families are acquired in canonical order; compatibility reads preserve rows
written under legacy claim axes. Claim-free members may participate alongside
claimed members. IDs may be generated, caller-supplied, or mixed, and
`bulkCreate()` returns rows in input order. This is an internal optimization,
not a general Store batch API. Identity-enabled nodes, history/revision
tracking, a member beyond the executor's declared claim-input budget, missing
schema-fence support, and other unsupported shapes fail closed to the existing
transaction or fallback behavior.
The transport inventory for the supported libSQL root records one client
`batch` submission and zero client `execute` calls for both generated-ID
claim-free batches, multiple-claim batches, cross-scope claims, and
claim-plus-projection batches. This is a measured submission count, not a
wall-clock RTT benchmark;
fallback paths are intentionally not assigned a latency claim.
On D1, claim work is chunked inside the same submission rather than imposing a
batch-wide ceiling. Each member has 87 claim-input binds after its row and
fence: a canonical claim costs six, each legacy hierarchy-wide uniqueness
probe costs nine, and each legacy disjointness probe costs six. Custom
executors should call the exported `atomicNodeClaimInputCost()` owner rather
than reproduce this formula. A member beyond that bound retains the portable
behavior.
Direct `edges.bulkInsert()` and `edges.bulkCreate()` calls on those same roots
use one schema-fenced native program when history and revision capture are
disabled. Declared durable match identities and `one`, `unique`, or
`oneActive` cardinality are maintained inside that exchange; any endpoint,
identity, or cardinality refusal rolls the whole call back. Transaction-scoped
stores, derived backends, custom backends without the corresponding exact-root
semantic registration, and dynamic get-or-create convergence retain the
interactive path.
Direct edge `bulkDelete()` calls use the same exact-root exchange and refuse a
foreign-kind ID atomically. Restricted node `bulkDelete()` also releases every
unique or disjoint claim owned by rows it tombstones in the same program, while
enforcing live connected edges in SQL. Identity, projection, history,
revision, cascade, and disconnect shapes retain their transaction path. `bulkUpsertById()`
remains a resolved mutation set because it must read and schema-validate a
database preimage before its writes are known. Bundled serverless roots can
submit an eligible distinct-ID, live-row resolved set as one native exchange
after that read. Bundled session-capable PostgreSQL can bind the same program
to the exact collection-opened, caller-supplied, or adopted transaction; this
is a bounded statement sequence on the pinned session, not one network
exchange. Update-only sets use a guarded update; sets containing both
fresh creates and updates include a terminal database assertion that rolls the
whole exchange back when any guarded postimage is absent. Repeated IDs,
resurrections, temporal changes, claims, edge sidecars (including durable edge
match identity), history/revision capture, ordinary derived backends, and
unregistered sessions use the interactive path. On D1's 100-parameter budget,
each native statement carries at most 17 node mutations or 6 edge mutations.
Larger eligible sets are chunked inside the same atomic transport submission;
each chunk has its own terminal postimage assertion, so one refusal rolls every
sibling chunk back rather than weakening the set contract. A D1 submission is
bounded to 512 node members or 187 edge members; larger sets fail closed to the
portable path instead of building an unbounded request. Other backends derive
their statement width from their declared bind budget and retain an absolute
512-member submission ceiling. The
operation returns an explicit `unsupported` verdict before issuing program SQL;
the Store never infers fallback safety from a missing result. Once a session
program starts, a savepoint preserves the surrounding transaction for typed
refusal diagnosis.
Node `bulkReplaceById()` avoids that structural preimage read by accepting only
complete replacement documents and distinct IDs. On an eligible bundled root,
the complete call—including claim ownership changes and fulltext/vector
sidecars—uses one atomic transport submission. Live rows retain their stored
validity windows; tombstones receive a freshly stamped window. Operational
Identity and history/revision capture use the portable path. Custom backends
must register and semantically certify the independent `replaceNodes` family;
transport registration or another node family is not evidence for replacement.
Eligible singleton `update()` and `delete()` calls reuse those same registered
families. Plain or projected node updates, unconstrained non-durable-identity edge updates,
all direct edge deletes, and plain restricted node deletes remain two-exchange
operations—one authoritative read/gate and one atomic mutation—because
TypeGraph must validate merged update properties and must preserve the rule
that a missing delete fires no operation hooks. This removes explicit
transaction transport from the eligible shape; it does not turn claims, edge
sidecars, temporal, captured, derived-backend, or caller-transaction writes into
autocommit operations.
That singleton update path uses optimistic convergence: the mutation asserts
the row preimage it read and retries a moved preimage up to four times. Under
sustained same-row contention it can throw `DatabaseOperationError` where an
interactive transaction would have waited to serialize the writers. This
applies to eligible `update()` calls and the live-row leg of `upsertById()` on
registered exact-root atomic transports. Caller transactions and other
ineligible shapes continue to use the serialized transaction path. Applications
using an atomic root should retry the operation when sustained contention can
move the row throughout all four attempts.
### One `bulkUpsertById` batch cannot hand a constrained value between rows
`bulkUpsertById` applies items in order for the purpose of deciding each row's
final props, but it groups the writes: every create in the batch runs before
every update. A batch where one item **releases** a constrained value and a later
item **claims** it therefore fails, where the same operations applied one at a
time succeed.
- Nodes: releasing and re-claiming a `unique` constraint value in one batch
throws `UniquenessError` — the claiming create is checked while the releasing
row still reserves the value.
- Edges: ending the lone `oneActive` edge from a source while creating its
replacement throws `CardinalityError`, for the same reason.
Bulk semantics are set-like, not scripted — a batch states the rows you want,
not an order to reach them in — so this is a stated limitation rather than a
pending fix. It always surfaces as a typed error, never as a dropped write.
Split the handoff across two batches (release, then claim), or apply the
conflicting items one at a time — as sequential `upsertById` calls for nodes,
and as `update` then `create` for edges, which have no single-item upsert. See
[Data Sync](/data-sync#one-batch-cannot-hand-a-unique-value-from-one-row-to-another)
for the worked example.
## Graph Analytics Limits
TypeGraph ships focused algorithms on `store.algorithms.*` — shortest path
(weighted and unweighted), reachability, k-hop neighborhoods, degree, exact
weakly connected components, deterministic label propagation, and
global/personalized PageRank. See
[Graph Algorithms](/graph-algorithms) for the full API.
The following heavier analytics are **not** provided:
- Modularity-optimizing community detection such as Leiden or Louvain
- Centrality measures beyond degree (betweenness, closeness, eigenvector)
- Strongly connected components
- Topological sort
- Graph partitioning
For these use cases, export your data via `.query().traverse()` or
`store.subgraph()` and use a specialized library such as
[graphology](https://graphology.github.io/) in memory, or move to a
dedicated graph database.
## Single Database Deployment
TypeGraph is designed for single-database deployments. It does not support:
- Distributed storage across multiple databases
- Sharding
- Cross-database queries
- Replication coordination
For distributed graph workloads, consider a dedicated graph database.
## Temporal Query Limitations
Temporal queries (`asOf`, `includeEnded`) work correctly but have some constraints:
- Point-in-time queries cannot be combined with streaming (`.stream()`)
- `validFrom` defaults to the record's own creation timestamp when omitted, so `asOf` queries
work out of the box; an end boundary still requires an explicit `validTo` — an open `validTo`
means "still valid". A record written with a `validTo` at or before its own creation instant
is "born already ended" and stores no lower bound instead, so it reads back at every `asOf`
before that end
- Rows an **older library version** stored with a backwards window
(`valid_from > valid_to`) are readable at no coordinate, and upgrading does not rewrite
them. Making them observable is an explicit operator action: run
`repairInvertedValidityWindows({ relations: "live-and-recorded", mode: "apply" })` while
writers are stopped, then re-baseline any outstanding merge branches. See
[Repairing inverted validity windows](/schema-management#repairing-inverted-validity-windows)
- Clock skew between application servers can affect temporal accuracy
### Recorded / system time (`history: true`)
Recorded-time capture (`createStore(graph, backend, { history: true })`) and
`store.asOfRecorded(T)` add a second temporal axis with these constraints. Use
`createAdapterStore(..., { history: true })` instead when the application must
adopt a caller-owned transaction:
- **Opt-in, no backfill.** Capture only sees changes committed after it is
enabled; an entity that already exists is first recorded the next time it is
written. Enable it on a fresh graph for complete history.
- **TypeGraph-write capture.** Built-in capture records TypeGraph collection
writes only. Out-of-band database writes and row-returning raw SQL paths are
not captured into the recorded relations.
- **Reconstructing reads only.** A recorded view exposes point reads
(`getById` / `getByIds`), bounded deterministic `scan()` pages, `query()`,
`subgraph()`, and the graph algorithms. Broad filtered collection reads
(`find` / `count` / `findFrom`), `search`, and fulltext / vector predicates are
refused — those indexes reflect current state and cannot answer a
recorded-time query.
- **Transactional backend required.** Capture needs a backend with atomic
transactions and statement execution — the built-in SQLite / PostgreSQL
backends qualify. A custom backend must implement `executeStatement` (optional
on the `GraphBackend` interface, but required once `history: true` is set) or
enabling capture throws a `ConfigurationError` at write time. On an
`AdapterHistoryStore`, raw `tx.sql` is disabled under `history: true`; adopt
external transactions with `store.withRecordedTransaction(...)` instead of
`store.withTransaction(...)` (which is a compile error on a history store).
- **Reconstruction cost.** Recorded reads rebuild from the history relations and
are slower than live reads, most noticeably for full-graph subgraph /
algorithm reconstructions on PostgreSQL.
- **PostgreSQL capture requires `READ COMMITTED`.** Every captured commit
advances a single recorded-clock row for the graph. TypeGraph refuses
PostgreSQL `REPEATABLE READ` / `SERIALIZABLE` history-capture transactions
because snapshot isolation cannot safely allocate that per-graph recorded
clock inside the captured transaction. Omit the transaction isolation option,
or set it to `read_committed`.
- **Recorded anchors are per graph.** Each captured transaction advances a
fixed-width logical revision and pairs it with a non-decreasing physical
wall-time high-water mark. TypeGraph does not provide a cross-graph recorded
anchor. See
[Logical revision and physical time](/queries/temporal#logical-revision-and-physical-time).
- **The preview schema needs an offline migration.** Timestamp-only anchors and
PostgreSQL recorded relations using `timestamptz` predate numeric recorded
revisions and the `r1::` API encoding. Run
`migrateLegacyRecordedTime()` while writers are stopped, then use
`migrateRecordedAnchor()` for checkpoints held outside TypeGraph. See
[Migrating preview recorded time](/schema-management#migrating-preview-recorded-time).
## Schema Migration Constraints
Automatic migrations (`createStoreWithSchema`) only handle additive changes:
| Change Type | Auto-Migrated |
|-------------|---------------|
| Add new node type | Yes |
| Add new edge type | Yes |
| Add optional property | Yes |
| Add required property | No |
| Remove property | No |
| Rename type | No |
| Change property type | No |
Breaking changes throw `MigrationError` and require manual migration.
# LLM Support
> Machine-readable documentation for AI assistants and coding tools
TypeGraph documentation is available in formats optimized for Large Language
Models (LLMs) following the [llms.txt specification](https://llmstxt.org/). All
files are generated from the same source docs as the website.
## Recommended Retrieval Order
For coding agents, use these files progressively:
1. Start with [`/llms-small.txt`](/llms-small.txt) for implementation and debugging tasks.
2. Use [`/llms-full.txt`](/llms-full.txt) only when you need deep reference content.
3. Load [`/_llms-txt/examples.txt`](/_llms-txt/examples.txt) only when you need full end-to-end patterns.
## Available Files
| File | Purpose | Size |
|------|---------|------|
| [`/llms.txt`](/llms.txt) | Index with page titles, descriptions, and links | Small |
| [`/llms-small.txt`](/llms-small.txt) | Core docs for implementation and debugging tasks | Medium |
| [`/llms-full.txt`](/llms-full.txt) | Complete documentation in a single file | Large |
| [`/_llms-txt/examples.txt`](/_llms-txt/examples.txt) | Complete application examples | Medium |
## Copy-Paste Agent Instructions
Use this in repository-level agent instruction files (`AGENTS.md`,
`CLAUDE.md`, etc.):
```md
TypeGraph (`@nicia-ai/typegraph`) is a TypeScript-first embedded knowledge graph
library with typed nodes, edges, queries, and schema management over SQLite and
PostgreSQL backends.
When working with TypeGraph code (graph definitions, node/edge schemas, store
operations, query builder, backend setup, or migrations):
1. Load https://typegraph.dev/llms-small.txt first.
2. Use https://typegraph.dev/llms-full.txt only for deep API/reference lookup.
3. Load https://typegraph.dev/_llms-txt/examples.txt only for end-to-end implementation patterns.
4. Prefer current API docs over inferred behavior from old snippets.
```
# Materializing External Event Logs
> How to project at-least-once event streams into TypeGraph without making TypeGraph an event-log product
External logs are the transport. TypeGraph is the typed, entity-resolved
materialization and merge layer.
Use this pattern when agents or integration runtimes already run on an event log
or stream: Electric Durable Streams, database changefeeds, message queues, or a
custom append-only feed. The log owns delivery, ordering, replay, and offsets.
TypeGraph owns the current graph, valid-time facts, recorded-time history, and
mergeable working copies. The sibling
[`agent-stream-graph`](https://github.com/nicia-ai/agent-stream-graph) package is
the reference implementation of this posture.
## The Shape of the Problem
External log consumers usually have three properties:
- **At-least-once delivery.** A change can be delivered more than once, especially
after a crash or reconnect.
- **Resume from a cursor.** The consumer persists the last source offset it has
safely processed.
- **Replay.** Reprocessing old events is normal: for recovery, backfills, or
rebuilding a derived graph.
That means a projector must be idempotent. Re-delivering the same source change
should converge on the same graph state, not create duplicates.
## Idempotent Projectors
Use stable source ids as TypeGraph ids whenever the source has them. For nodes,
that usually means `upsertById`. For edges, prefer `getOrCreateByEndpoints`.
Avoid `create` in a log projector unless the source event itself carries a
unique id you pass as the TypeGraph id.
```typescript
async function projectChange(
tx: TransactionContext,
change: Change,
) {
const issue = await tx.nodes.Issue.upsertById(
change.issueId,
{
title: change.title,
state: change.state,
},
{
validFrom: change.issueValidFrom,
onImmutableLowerBound: "preserve",
},
);
const actor = await tx.nodes.Actor.upsertById(change.actorId, {
name: change.actorName,
});
await tx.edges.changedBy.getOrCreateByEndpoints(
issue,
actor,
{ action: change.action },
{
ifExists: "update",
validFrom: change.relationshipValidFrom,
validTo: change.relationshipValidTo,
onImmutableLowerBound: "preserve",
},
);
}
```
The important rule is that the second delivery of the same change takes the same
code path and reaches the same row identities.
The `"preserve"` policy makes `validFrom` create/resurrection-only input for
both node and edge writes: a later revision updates props and `validTo` without
trying to rewrite the live row's start. Without it, the default `"refuse"`
policy raises `IMMUTABLE_VALIDITY_LOWER_BOUND` when a revision states a
different start. The edge also explicitly selects `ifExists: "update"`; the
default is `"return"`, which is right for create-once relationships but writes
neither revised props nor a closing `validTo` when the edge already exists.
### `matchOn` widens the identity key — don't reach for it by default
`getOrCreateByEndpoints` matches on the endpoints `(from, to)` alone unless you
pass `matchOn`. Endpoints-only is the **more** idempotent choice and is right for
most projectors: a re-delivered edge between the same two nodes converges on the
one existing edge regardless of how its properties drifted between deliveries.
`matchOn` adds the named property fields to the match key, so it *widens*
identity — two edges between the same endpoints are now distinct if they differ
on a matched field. Use it only when the relationship model genuinely allows
several parallel edges between one pair (say, one `changedBy` edge per distinct
`action`), and know the footgun: if a re-delivered change carries a **changed**
value in a matched field, it no longer matches the earlier edge and you get a
**second** edge instead of convergence. Reach for `matchOn` when the domain
needs the extra edges, not as a reflex.
Validity timestamps do not become part of this identity key. If the same
endpoints can have multiple application-time periods, include a stable period
or source-event identifier in the edge schema and in `matchOn`. This keeps a
re-delivery of one period convergent without collapsing a later period into the
same edge.
## Cursor Bookkeeping
A cursor is application state: the last source offset you have safely processed.
It should advance only at a source offset boundary, after every change in that
batch has been projected. Where the cursor lives — a row in your own relational
table, or a node in the graph — decides which guarantees you can get.
### Exactly-once with an adopted transaction
To commit the projected batch **and** the cursor as one unit, let the caller own
the transaction and adopt it with
[`store.withRecordedTransaction(externalTx, fn)`](/schemas-stores/#transaction-receipts).
The graph writes and your own cursor write land on the same connection inside the
same commit: either both persist or neither does, so the cursor can never advance
past a batch the graph did not durably record.
Two constraints make this the *only* sanctioned transactional recipe on a store
created with `createAdapterStore` or `createAdapterStoreWithSchema` and
`{ history: true }` — which the Transaction Receipts and Bitemporal sections
below both require:
- **Write your own tables through the external handle you passed in**, never
through `tx.sql`. Under history capture the typed transaction context omits
`sql` (raw SQL would bypass recorded-time capture); suppressed access reaches
a runtime guard and raises a
[`ConfigurationError`](/errors/#recorded-capture-guard-codes). The external
handle *is* the pinned connection, so writing your cursor row through it keeps
both layers in the one transaction.
- **`store.withTransaction()` — the non-recorded sibling — is a compile error on
a history store** (its `externalTx` argument is rejected against a message
type), and its runtime guard throws
`RECORDED_CAPTURE_REQUIRES_CALLBACK_TRANSACTION`. It has no flush point before
the caller commits, so recorded-time capture could not seal. Use
`withRecordedTransaction` instead.
**Async drivers (Postgres / libsql)** open the boundary with `db.transaction`:
```typescript
const receipt = await db.transaction(async (dbTx) => {
const outcome = await store.withRecordedTransaction(dbTx, async (tx) => {
for (const change of batch.changes) {
await projectChange(tx, change);
}
});
// The cursor row goes through the external handle, in the same transaction.
await dbTx
.insert(streamCursors)
.values({ sourceId: batch.sourceId, offset: batch.endOffset })
.onConflictDoUpdate({
target: streamCursors.sourceId,
set: { offset: batch.endOffset },
});
return outcome.receipt;
}); // one COMMIT / ROLLBACK across both layers
```
**Synchronous `better-sqlite3`** cannot adopt an `async` transaction callback
(its driver rejects a promise-returning `db.transaction`), so the caller frames
the boundary by hand with `BEGIN IMMEDIATE` / `COMMIT` / `ROLLBACK` on the single
connection:
```typescript
await db.run(sql`BEGIN IMMEDIATE`);
try {
const { receipt } = await store.withRecordedTransaction(db, async (tx) => {
for (const change of batch.changes) {
await projectChange(tx, change);
}
await db.run(sql`
INSERT INTO stream_cursor (source_id, offset)
VALUES (${batch.sourceId}, ${batch.endOffset})
ON CONFLICT (source_id) DO UPDATE SET offset = excluded.offset
`);
});
await db.run(sql`COMMIT`);
// persist receipt.recorded as the offset's replay anchor — see below
} catch (error) {
await db.run(sql`ROLLBACK`); // graph writes and cursor roll back together
throw error;
}
```
The graph writes and your own statements share the caller's one pinned
connection. TypeGraph serializes the statements its collections issue; sequence
your own raw statements yourself (don't `Promise.all` them with graph writes) so
two queries never race on that connection.
For an adapter-backed materializer that runs against several backends, branch
on capability rather than message-matching: use
[`tx.sqlAvailability`](/recipes/#cross-store-transactions-drizzle--typegraph) to
decide whether raw SQL is usable inside `store.transaction`, and
[`isRecordedCaptureGuardError(error, code?)`](/errors/#recorded-capture-guard-codes)
to recognize a history-store guard when you catch one.
### At-least-once with a separate cursor store
When the runtime already owns checkpointing, or the backend cannot provide atomic
transactions (`backend.capabilities.execution.interactiveTransactions === false` — Cloudflare D1,
`drizzle-orm/neon-http`), keep the cursor outside the graph transaction. The
pattern is at-least-once plus idempotence: a crash after the graph writes but
before the cursor write replays the batch, which is safe precisely because the
projector converges.
This fallback requires a raw Store. A schema-managed Store refuses writes on a
non-transactional backend because it cannot hold the schema-version fence. Use a
transactional driver, or deliberately construct a raw Store and own schema/write
coordination yourself.
```typescript
for (const change of batch.changes) {
// Each successful projection may commit before a later projection or cursor write fails.
await projectChange(store, change);
}
await cursorStore.save({
sourceId: batch.sourceId,
offset: batch.endOffset,
});
```
This at-least-once path plus an idempotent projector is the workload TypeGraph is
built for. It is also the one that churns recorded history the hardest: every
re-delivery of a byte-identical change rewrites its row, allocating a fresh
recorded instant and a new history row per delivery. Enable
[`coalesceUnchangedUpserts: true`](/schemas-stores/#createstoregraph-backend-options)
on the store to suppress that. A node `upsertById` and an edge endpoint
get-or-create update perform no write, history row, or revision advance when
its validated props and requested window already equal the live row. Their bulk
forms have the same behavior. See
[Transaction Receipts](#transaction-receipts) for how a coalesced upsert reads on
a receipt.
Every captured transaction receives one versioned recorded instant: a strict
per-graph logical revision paired with a non-decreasing physical wall-time
high-water mark. High commit rates consume revisions without pushing the
timestamp beyond observed wall time. A backward clock correction holds the
physical component at its prior value until the clock catches up, preserving
cumulative diagonal checkpoint replay. Group changes by their durable
replay/checkpoint boundary so one addressable source position consumes one
recorded instant where practical. Cap transaction size independently: a source
may expose one coarse checkpoint for a very large initial sync, but that does
not make an unbounded transaction safe.
Recorded clocks are independent per graph, and there is no cross-graph
`recordedNow()` snapshot. See
[Logical revision and physical time](/queries/temporal#logical-revision-and-physical-time)
for the anchor encoding and replay semantics.
**Coalescing eliminates *re-delivery* churn, not replay cost.** The win is
scoped to re-delivery of the current value — the realistic at-least-once case,
where a change that was already applied arrives again (a crash-window replay, a
duplicate) and is value-identical to the live row. A full **replay-from-zero**
over the current state is different: if the stream contains in-place updates,
replaying `insert a=1 … update a=2` re-applies `a=1` over the live `a=2` — a
genuine backward change that writes — and then `a=2` restores it. Both writes
are correct (the replay faithfully re-walks each historical state), but
"coalescing makes replay free" holds only for streams whose rows never
supersede each other. It also leaves a spurious `a=2 → a=1 → a=2` band in the
live store's recorded history, stamped at replay time. To rebuild without
either cost, replay into a **fresh store** and publish it, rather than
re-applying the log over the current state.
**Historical ends need a historical start on the creating event.** When a fresh
store creates a row without a stated `validFrom`, TypeGraph uses the ingest
instant. A later replayed event whose historical `validTo` precedes that ingest
instant is therefore an `INVERTED_VALIDITY_WINDOW`, even if the source timeline
itself was ordered. An event-time decoder must emit `validFrom` on the event
that first creates each node or edge; `onImmutableLowerBound: "preserve"` then
lets later revisions carry their source bound without trying to move the stored
start.
### In-graph cursors and the receipt
A cursor can also live inside the graph as an ordinary node — convenient, and it
travels with the graph. But if you also use the transaction receipt (next
section) to detect a projector that dropped a change, an in-graph cursor
**corrupts that signal**: the receipt counts writes per transaction with no
attribution, so the cursor's own upsert is indistinguishable from the projector's
writes. A projector that drops a change in a transaction that also checkpoints an
in-graph cursor produces `writes.total === 1` from the cursor alone —
`writes.total > 0` no longer means "the projector wrote," the drop goes
undetected, and the cursor advances past the lost event.
Two ways out:
- **Scope the projector with `tx.measure`.** On a receipt-enabled context
(`transactionWithReceipt` or `withRecordedTransaction`),
[`tx.measure((scopedTx) => …)`](/schemas-stores/#scoped-receipts-txmeasure)
hands your callback a **scoped context** and returns a sub-receipt that counts
exactly the writes made **through that scoped context** (`scopedTx.nodes` /
`scopedTx.edges`). Run the projector through `scopedTx`; write the cursor
through the outer `tx` (or your own table). Attribution is by which context you
write through, not by timing, so the cursor's write counts only in the outer
receipt and the scope reflects the projector alone. This is what makes an
in-graph cursor and drop-detection composable; see the full loop below.
- **Keep the cursor in your own relational table.** The
[exactly-once recipe](#exactly-once-with-an-adopted-transaction) makes that
atomic anyway, and it keeps the belief graph pristine: an in-graph cursor node
still lands in every `asOfRecorded` reconstruction of the graph, so a consumer
that wants recorded-time reads to show only projected facts should keep the
cursor out of the graph entirely.
## Transaction Receipts
When you need to know what a projector did, use `store.transactionWithReceipt`
(TypeGraph owns the boundary) or `store.withRecordedTransaction` (you adopt an
open transaction); both return a `TransactionOutcome` with a `receipt`.
The receipt carries **two signals that deliberately disagree**, and a
materializer needs both. `receipt.writes.total` counts completed write intents at
the collection surface; `receipt.recorded` (on a `{ history: true }` store) is
the recorded commit instant this transaction allocated, or `undefined` when
nothing was captured or explicitly requested. The common, load-bearing case is
where they diverge:
| case | `writes.total` | `recorded` |
| --------------------------------------------- | -------------- | --------------- |
| projector wrote | `> 0` | defined |
| no-op delete of an absent key (a real intent) | `1` | **`undefined`** |
| coalesced upsert (value-identical, opt-in) | `1` | **`undefined`** |
| projector dropped the change | `0` | `undefined` |
| explicit recorded revision request | `0` | defined |
A no-op delete completes a write *intent* but captures nothing; a
[coalesced upsert](/schemas-stores/#createstoregraph-backend-options) is the same
shape by design. In both, `writes.total` counts (the method resolved) but
`recorded` is `undefined`. **An offset whose transaction reports
`recorded === undefined` must carry the prior anchor forward** — otherwise
replay-by-offset breaks at exactly the offsets where nothing changed.
Call `requestRecordedRevision()` when that offset must instead receive its own
anchor despite making no entity change.
Two counting rules bite materializers specifically, both worth internalizing
before you read `writes.total` as "the projector did work":
- **Bulk methods count by input length**, so `bulkCreate([])` contributes `0`. A
projector that filters a batch down to nothing and issues an empty bulk call
must not read as a writer.
- **A method that rejects counts `0`** — even on SQLite, where a failed statement
does **not** abort the surrounding transaction. A projector that swallows a
write error and commits can persist rows the receipt never counted, so do not
read the receipt as rows-affected in that scenario.
### The full materializer loop
Putting the pieces together: an adopted transaction for exactly-once cursors, a
`tx.measure`-scoped projector so a single dropped change is caught within a
multi-change batch — the outer receipt only tells you the whole batch wrote
nothing, whereas a `measure` scope attributes writes per change by having the
projector write through the scoped context it receives (any cursor written
through the outer `tx` stays out of that count) — `receipt.recorded` as the
per-offset replay anchor, and `writes.total === 0` on a non-delete change as the
drop signal.
`withRecordedTransaction` flushes recorded-time capture and resolves **before**
the caller's commit, so `outcome.receipt.recorded` is already known inside the
`db.transaction` callback. Write the cursor advance **and** its replay anchor
through `dbTx` there, in the same commit as the graph writes. Persisting the
anchor after the commit — as a separate step — would reopen the exactly-once gap
the adopted transaction exists to close: a crash between the commit and the
anchor write leaves the cursor advanced with no anchor, and that offset can never
be replayed.
If an accepted source position must be addressable even when the projector makes
no entity changes, request an explicit revision in the callback. Await the
`withRecordedTransaction` outcome, then insert the application-owned cursor and
returned anchor through the still-open native transaction before its commit:
```typescript
await db.transaction(async (dbTx) => {
const { receipt } = await store.withRecordedTransaction(dbTx, async (tx) => {
tx.requestRecordedRevision();
await projectBatch(tx, batch);
});
await dbTx.insert(cursors).values({
source: batch.source,
offset: batch.offset,
recorded: receipt.recorded,
});
});
```
```typescript
let lastAnchor: RecordedInstant | undefined = await loadLastAnchor(); // on resume
lastAnchor = await db.transaction(async (dbTx) => {
const outcome = await store.withRecordedTransaction(dbTx, async (tx) => {
for (const change of batch.changes) {
// The projector writes through the scoped context, so `projected`
// counts its writes alone — nothing else in the transaction.
const projected = await tx.measure((scopedTx) =>
projectChange(scopedTx, change),
);
// A non-delete change that wrote nothing was silently dropped.
if (
projected.receipt.writes.total === 0 &&
change.operation !== "delete"
) {
throw new DroppedChangeError(change); // rolls the whole batch back
}
}
});
// `recorded` is undefined when the batch captured nothing (all drops, no-op
// deletes, or coalesced upserts) — carry the prior anchor forward so replay
// by offset still resolves. Anchor comes from the receipt, never from a
// post-commit store.recordedNow() (see below).
const anchor = outcome.receipt.recorded ?? lastAnchor;
// Cursor and anchor commit atomically with the graph writes: no window where
// the cursor has advanced past an offset whose anchor was never persisted.
await dbTx
.insert(offsetAnchors)
.values({ sourceId: batch.sourceId, offset: batch.endOffset, recorded: anchor });
await dbTx
.insert(streamCursors)
.values({ sourceId: batch.sourceId, offset: batch.endOffset })
.onConflictDoUpdate({
target: streamCursors.sourceId,
set: { offset: batch.endOffset },
});
return anchor; // updates lastAnchor only once the transaction commits
});
```
**Take the replay anchor from `receipt.recorded`, never from a post-commit
`store.recordedNow()`.** `recordedNow()` is the graph-global recorded
high-water mark, advanced by **any** writer to the graph. Between your commit and
your read of it, a concurrent writer can advance it, and `asOfRecorded(that)`
then reconstructs a belief your stream never produced. The receipt hands you the
instant *this* transaction allocated; that is the only anchor that reconstructs
exactly what this offset materialized.
## Bitemporal Mapping
External streams usually carry domain time and delivery time. Keep those
separate:
- **Event time belongs in valid time.** If a source change says a fact became
true on January 1, pass that timestamp as `validFrom`; if it ended on January
31, pass `validTo`.
- **Ingest time is recorded time.** TypeGraph records when the graph committed
the write. Recorded time is allocated by the backend and cannot be backdated.
- **Backfills collapse recorded instants to now.** Replaying historical events
today writes historical valid-time facts with today's recorded-time anchors.
That is correct SQL:2011 bitemporal behavior, not a bug.
To replay by source offset, load the anchor you saved for that offset and read a
recorded-time view. The `receipt.recorded` you persisted is a branded
`RecordedInstant`, but round-tripping through your cursor table stores it as a
plain string — re-brand it with `asRecordedInstant` on the way back before
passing it to `asOfRecorded`:
```typescript
import { asRecordedInstant } from "@nicia-ai/typegraph";
const stored = await offsetAnchors.anchorFor(offset); // plain string from storage
const anchor = asRecordedInstant(stored); // validates + re-brands
const graphAtOffset = store.asOfRecorded(anchor);
const issue = await graphAtOffset.nodes.Issue.getById(issueId);
```
If the cursor table contains timestamp-only anchors from the recorded-time
preview, migrate the TypeGraph relations first and remap those cursor values
with `migrateRecordedAnchor({ backend, graphId, anchor: stored })`. See
[Migrating preview recorded time](/schema-management#migrating-preview-recorded-time).
That answers "what did the materialized graph know after offset X?" even if
later corrections changed or deleted rows. See
[Recorded time](/queries/temporal/#recorded-time-bitemporal) for the full view
surface.
### Refresh planner statistics after a large replay
A **custom** replay or backfill loop — one built from the projector recipes above
— runs its writes **inside a caller-provided transaction**, which never
auto-refreshes the query planner's table statistics: `ANALYZE` from another
connection cannot see rows that are still uncommitted, so the store deliberately
skips the automatic refresh it does after large autocommit bulk writes. Left
alone, the planner keeps pre-load row estimates and can pick an
order-of-magnitude-slower plan. After a large custom replay, refresh once:
```typescript
await replayEverything();
await store.refreshStatistics(); // once, after the bulk replay commits
```
The interchange path handles this for you: `importGraph` and `importGraphStream`
call `refreshStatistics()` once after the import commits (see
[Bulk Copy Between Stores](#bulk-copy-between-stores)), so a bulk copy needs no
manual refresh.
## Bulk Copy Between Stores
To copy a materialized graph into another store — most often a graph-merge
working copy — stream interchange directly from source to target with
`exportGraphStream` / `importGraphStream`. This is the same path
[graph-merge](/interchange/) uses internally, so a copy produces byte-identical
merge results, conflicts, and provenance to a native branch:
```typescript
import {
exportGraphStream,
importGraphStream,
} from "@nicia-ai/typegraph/interchange";
const result = await importGraphStream(
branch.store,
exportGraphStream(beliefStore, {
nodeKinds: ["Belief", "Claim"],
edgeKinds: ["supports"],
includeTemporal: true,
}),
{ onConflict: "update" },
);
if (!result.success) {
throw new Error(`copy failed with ${result.errors.length} import errors`);
}
```
Two option defaults are exactly right here and worth stating because they are not
obvious:
- **`includeDeleted` defaults to `false`, and the copy clones live state — it
does not synchronize deletions.** The exporter simply omits soft-deleted rows;
it cannot round-trip `deletedAt` at all (the wire format carries no deletion
flag). So a fact deleted on the source is merely *absent* from the stream: on a
fresh target it never appears, but on a populated target an existing live row
**stays live** — the copy never deletes it. If the target must reflect
deletions, apply them through your projector, not the bulk copy.
- **`includeTemporal` must be set to `true`** (it defaults to `false`). It is
what carries each fact's original `validFrom` / `validTo` across the copy;
without it the import re-stamps every fact with the *copy's* wall clock,
destroying valid-time fidelity in the merged branch.
`importGraphStream` preserves ids, routes existing rows through normal
`onConflict` handling, validates edge endpoints (`validateReferences` defaults to
`true`), and refreshes planner statistics once after the import commits.
## Cursor-Based Resumption and Electric
The examples above assume a per-change offset. **Electric does not provide one** —
every change in a `ShapeStream` catch-up batch shares the stream's
`lastOffset`. A cursor keyed on Electric's offset can therefore only advance at a
**batch boundary**, after the whole batch is projected. Advancing mid-batch is
unsafe: Electric's `read(after)` is strictly-after, so resuming from a
mid-batch offset permanently skips that batch's remaining changes. Project the
whole batch, then checkpoint the cursor once at its boundary.
# Multiple Graphs
> Using separate graph definitions for different domains in the same application
TypeGraph supports multiple graphs for applications that have distinct data domains that benefit from separate graph definitions.
## When to Use Multiple Graphs
Use separate graphs when you have:
- **Distinct domains**: A RAG system for documents and a business network for suppliers have different node types,
edge semantics, and query patterns
- **Independent lifecycles**: One graph might evolve rapidly while another is stable
- **Team ownership**: Different teams own different graphs, with separate schema review processes
- **Different retention policies**: Document chunks might be ephemeral while business relationships are long-lived
**Don't use multiple graphs** when:
- You need cross-graph queries or traversals (use a single graph with ontology relations instead)
- The domains are closely related (e.g., Users and Documents that Users author)
- You're trying to solve multi-tenancy (use tenant isolation patterns instead)
## Example: Documents and Business Network
A company needs two graphs:
1. **Documents graph**: Powers semantic search over internal documents
2. **Organization graph**: Tracks suppliers, partners, and contracts
### Defining the Graphs
```typescript
// graphs/documents.ts
import { z } from "zod";
import { defineNode, defineEdge, defineGraph, embedding } from "@nicia-ai/typegraph";
const Document = defineNode("Document", {
schema: z.object({
title: z.string(),
source: z.string(),
createdAt: z.string().datetime(),
}),
});
const Chunk = defineNode("Chunk", {
schema: z.object({
content: z.string(),
embedding: embedding(1536),
position: z.number().int(),
}),
});
const hasChunk = defineEdge("hasChunk");
export const documentsGraph = defineGraph({
id: "documents",
nodes: {
Document: { type: Document },
Chunk: { type: Chunk },
},
edges: {
hasChunk: { type: hasChunk, from: [Document], to: [Chunk] },
},
});
```
```typescript
// graphs/organization.ts
import { z } from "zod";
import { defineNode, defineEdge, defineGraph, subClassOf } from "@nicia-ai/typegraph";
const Organization = defineNode("Organization", {
schema: z.object({
name: z.string(),
domain: z.string().optional(),
}),
});
const Supplier = defineNode("Supplier", {
schema: z.object({
name: z.string(),
domain: z.string().optional(),
category: z.enum(["materials", "services", "logistics"]),
}),
});
const Partner = defineNode("Partner", {
schema: z.object({
name: z.string(),
domain: z.string().optional(),
partnershipLevel: z.enum(["bronze", "silver", "gold"]),
}),
});
const Contract = defineNode("Contract", {
schema: z.object({
title: z.string(),
value: z.number(),
startDate: z.string().datetime(),
endDate: z.string().datetime().optional(),
status: z.enum(["draft", "active", "expired"]).default("draft"),
}),
});
const supplies = defineEdge("supplies");
const hasContract = defineEdge("hasContract");
export const organizationGraph = defineGraph({
id: "organization",
nodes: {
Organization: { type: Organization },
Supplier: { type: Supplier },
Partner: { type: Partner },
Contract: { type: Contract },
},
edges: {
supplies: { type: supplies, from: [Supplier], to: [Organization] },
hasContract: { type: hasContract, from: [Organization], to: [Contract] },
},
ontology: [
subClassOf(Supplier, Organization),
subClassOf(Partner, Organization),
],
});
```
### Creating Stores
Both graphs can share the same database backend. Each graph's data is isolated by its `id`.
```typescript
// stores.ts
import { createStore } from "@nicia-ai/typegraph";
import { createPostgresBackend } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import { drizzle } from "drizzle-orm/node-postgres";
import { Pool } from "pg";
import { documentsGraph } from "./graphs/documents";
import { organizationGraph } from "./graphs/organization";
const pool = new Pool({ connectionString: process.env.DATABASE_URL });
const db = drizzle(pool);
const backend = createPostgresBackend(db);
// Same backend, different stores
export const documentsStore = createStore(documentsGraph, backend);
export const organizationStore = createStore(organizationGraph, backend);
```
### Using the Stores
Each store is fully independent with its own typed API:
```typescript
// Semantic search in documents
async function searchDocuments(query: string, embedding: number[]) {
return documentsStore
.query()
.from("Chunk", "c")
.whereNode("c", (c) => c.embedding.similarTo(embedding, 10))
.select((ctx) => ({
content: ctx.c.content,
position: ctx.c.position,
}))
.execute();
}
// Business queries in organization
async function getActiveSuppliers(category: string) {
return organizationStore
.query()
.from("Supplier", "s")
.whereNode("s", (s) => s.category.eq(category))
.traverse("hasContract", "e")
.to("Contract", "c")
.whereNode("c", (c) => c.status.eq("active"))
.select((ctx) => ({
supplier: ctx.s.name,
contract: ctx.c.title,
value: ctx.c.value,
}))
.execute();
}
```
## Coordinating Across Graphs
Since cross-graph queries aren't supported, coordinate at the application level.
### Shared Identifiers
Use consistent IDs when entities relate across graphs:
```typescript
// When ingesting a supplier's documents, use the supplier ID as a reference
async function ingestSupplierDocument(
supplierId: string,
title: string,
content: string,
embedding: number[]
) {
// Store document with supplier reference in metadata
const doc = await documentsStore.nodes.Document.create({
title,
source: `supplier:${supplierId}`,
createdAt: new Date().toISOString(),
});
const chunk = await documentsStore.nodes.Chunk.create({
content,
embedding,
position: 0,
});
await documentsStore.edges.hasChunk.create(doc, chunk, {});
return doc;
}
// Later, find documents for a supplier
async function getSupplierDocuments(supplierId: string) {
return documentsStore
.query()
.from("Document", "d")
.whereNode("d", (d) => d.source.eq(`supplier:${supplierId}`))
.select((ctx) => ctx.d)
.execute();
}
```
### Application-Level Joins
Combine results from multiple graphs in your application:
```typescript
interface SupplierWithDocuments {
supplier: { name: string; category: string };
documents: Array<{ title: string }>;
}
async function getSupplierOverview(
supplierId: string
): Promise {
// Parallel queries to both graphs
const [supplier, documents] = await Promise.all([
organizationStore.nodes.Supplier.getById(supplierId),
getSupplierDocuments(supplierId),
]);
return {
supplier: {
name: supplier.name,
category: supplier.category,
},
documents: documents.map((d) => ({ title: d.title })),
};
}
```
### Event-Driven Sync
For loose coupling, use events to keep graphs in sync:
```typescript
// When a supplier is created, set up document ingestion
eventBus.on("supplier.created", async (event) => {
const { supplierId, name } = event.payload;
// Create a placeholder document node for future ingestion
await documentsStore.nodes.Document.create({
title: `${name} - Supplier Profile`,
source: `supplier:${supplierId}`,
createdAt: new Date().toISOString(),
});
});
// When a supplier is deleted, clean up related documents
eventBus.on("supplier.deleted", async (event) => {
const { supplierId } = event.payload;
const docs = await documentsStore
.query()
.from("Document", "d")
.whereNode("d", (d) => d.source.eq(`supplier:${supplierId}`))
.select((ctx) => ctx.d.id)
.execute();
for (const docId of docs) {
await documentsStore.nodes.Document.delete(docId);
}
});
```
## Separate Backends
For stronger isolation, use separate database connections:
```typescript
// Documents in PostgreSQL with pgvector for embeddings
const documentsPool = new Pool({
connectionString: process.env.DOCUMENTS_DATABASE_URL,
});
const documentsBackend = createPostgresBackend(drizzle(documentsPool));
export const documentsStore = createStore(documentsGraph, documentsBackend);
// Organization data in a separate database
const orgPool = new Pool({
connectionString: process.env.ORG_DATABASE_URL,
});
const orgBackend = createPostgresBackend(drizzle(orgPool));
export const organizationStore = createStore(organizationGraph, orgBackend);
```
**When to separate backends:**
- Different performance profiles (vector search vs. relational queries)
- Compliance requirements (PII in one database, analytics in another)
- Independent scaling needs
- Different backup/retention policies
## Schema Management
Each graph has independent schema versioning:
```typescript
import { createStoreWithSchema } from "@nicia-ai/typegraph";
// Each graph tracks its own schema version
const [documentsStore, docsSchemaResult] = await createStoreWithSchema(
documentsGraph,
backend
);
const [orgStore, orgSchemaResult] = await createStoreWithSchema(
organizationGraph,
backend
);
// Check migration status independently
if (docsSchemaResult.status === "migrated") {
console.log("Documents schema was migrated");
}
if (orgSchemaResult.status === "migrated") {
console.log("Organization schema was migrated");
}
```
## Inspecting What a Database Holds
Graphs sharing a backend are separated by `graph_id` inside TypeGraph's tables. Two reads answer the
questions an operator asks about that layout without depending on it: which graphs live in this
database, and how many rows one graph holds.
### `listGraphIds(backend, options?)`
Lists the graph ids that hold data, one bounded page at a time:
```typescript
import { listGraphIds } from "@nicia-ai/typegraph";
let after: string | undefined;
for (;;) {
const page = await listGraphIds(backend, { prefix: "tenant-", after, limit: 100 });
if (page.length === 0) break;
for (const graphId of page) console.log(graphId);
after = page.at(-1);
}
```
| Option | Meaning |
| -------- | ----------------------------------------------------------------------------------- |
| `prefix` | Only ids starting with this exact, case-sensitive text. `%` and `_` are not wildcards. |
| `after` | Exclusive cursor: only ids ordered after this one. Pass the last id of the previous page. |
| `limit` | Page size from 1 to 1000. Defaults to 100. Anything else throws `ConfigurationError`. |
Ids come back in byte order (UTF-8 code point order) on every backend, so `Tenant-x` sorts before
`tenant-a` on SQLite and PostgreSQL alike and a cursor resumes exactly where the last page ended,
whatever the database collation. The reserved deployment marker id that TypeGraph uses for
deployment-scoped contribution markers is never listed. A graph appears while it has nodes, edges or a
committed schema version, which are exactly the relations a default `store.clear()` empties. A
cleared graph therefore stops being listed even though `store.clear()` keeps its contribution
markers unless you pass `preserveContributionMaterializations: false`, and even when a
revision-tracked store reseeded its `recordedClock` row during the clear.
Each page walks graph ids by index seek, one seek per graph per relation, instead of reading every
row. The walk starts at the cursor or prefix and stops after the page, so a page costs about `limit`
seeks wherever it sits, however many graphs the database holds and however many rows they contain.
SQLite serves the seeks from the `graph_id`-leading primary keys, which are already in byte order.
PostgreSQL orders ordinary text indexes by the database collation, so it serves them from a
byte-ordered (`COLLATE "C"`) `graph_id` index that base-schema version 5 adds to `nodes`, `edges` and
`schema_versions`; see [Base-schema version 5](/backend-setup#base-schema-version-5-byte-ordered-graph_id-indexes-postgresql)
for what it costs and how to build it ahead of an upgrade. On a 20,000-graph, 50-rows-per-graph
PostgreSQL 18 database a page takes about 3 ms, where the same read took about 400 ms before the
index; at 200 graphs of 5,000 rows it is about 3 ms either way. A database without the index (its
base schema not adopted yet, or DDL managed by hand) lists the same ids by reading and de-duplicating
every row of those relations for each page, measured at 47 to 105 ms a page at these sizes. A backend that
declares no recursive traversal does the same. Use the listing for operator tooling, not on a request
path. Rows that exist only outside those relations, such as orphaned recorded history or contribution
markers, do not make a graph appear; `inspectGraphStorage` counts every relation.
The read runs in one read-only transaction where the backend supports it. It needs the backend's
catalog probes to tell a table that was never provisioned from an empty one, and throws
`ConfigurationError` on a custom backend that has none.
### `inspectGraphStorage(store)`
Counts one graph's rows in every relation that can hold them:
```typescript
import { inspectGraphStorage } from "@nicia-ai/typegraph";
await store.clear();
const { graphId, relations, totalRows } = await inspectGraphStorage(store);
const leftovers = relations.filter((relation) => relation.rows > 0);
// [{ relation: "contributionMaterializations", table: "typegraph_contribution_materializations", rows: 2 }]
```
`relations` lists every graph-scoped relation under its logical key (`nodes`, `edges`, `uniques`,
`edgeClaims`, `identityAssertions`, `recordedNodes`, `fulltext`, `schemaVersions`, and so on) with
the physical `table` it resolved to on this backend, so custom table names are reported as
configured. The graph's per-field vector tables come from its vector slots and the active vector
strategy and are reported as `vector:.`. A relation whose table the database never
provisioned counts as `0` rather than failing.
Use it to verify that `store.clear()` left nothing behind. Two relations can legitimately hold a
row after a clear, by design:
- `contributionMaterializations` is preserved unless you pass
`preserveContributionMaterializations: false`.
- `recordedClock` is reseeded inside the clear transaction on a store with live revision tracking
(without history).
Every other relation reads `0` after a clear, and other graphs in the same database are untouched.
#### Consistency of the counts
Each relation is counted by its own statement, so `relations` and `totalRows` describe one state of
the graph only when every statement read the same snapshot. The result carries a `consistency`
field that says whether they did:
| `consistency` | Meaning |
| --- | --- |
| `"snapshot"` | Every count came from one snapshot: `relations` and `totalRows` describe a state the graph was in. |
| `"per-statement"` | Each relation was counted independently. A write between two counts can leave the result describing a state that never existed together, for example rows in `nodes` beside an empty `schemaVersions`. |
The read asks for a read-only `repeatable read` transaction, but it does not trust the request: a
transaction wrapper can drop the isolation option, and a role or database can default the level.
The effective level is read on the counting session itself, inside the first count statement, so
the answer costs no extra round trip. What that gives on each backend:
- **SQLite (better-sqlite3, libSQL, and other drivers with interactive transactions):**
always `"snapshot"`. A SQLite transaction reads one snapshot whatever level was requested.
- **PostgreSQL (`pg`, `postgres-js`, PGlite):** `"snapshot"` when the session was observed at
`repeatable read` or `serializable`, which is what the request produces. `"per-statement"` when
it ran at `read committed`, which happens when a wrapper around `backend.transaction` does not
forward its options and the role or database defaults to `read committed`; the same wrapper
under a `repeatable read` default still reports `"snapshot"`, because the level is observed, not
requested. A backend that declares no session isolation read cannot be observed and reports
`"per-statement"`.
- **Backends without interactive transactions (Cloudflare D1, `neon-http`):** `"per-statement"`,
because there is no transaction to share a snapshot. The exception is a graph with at most one
provisioned relation, which is one statement and so trivially consistent.
The evidence proves the isolation of the session that ran the first count. A backend wrapper that
violates the transaction contract by handing the root pool through as its transaction backend can
run later counts on other sessions, which no observation on the first one can detect.
The read never refuses on a weaker level: it is a diagnostic. Treat `"per-statement"` counts as an
approximation. To verify a clear with them, make sure nothing else writes the graph while you read,
or read twice and compare.
## Shared Subgraph Helpers
When multiple graphs share a common set of node and edge types, you can write reusable
helpers that accept any store containing that shared subgraph. The `StoreProjection` utility
type makes this type-safe without coupling to a specific graph definition.
### Defining shared types and graphs
Start with the shared node and edge types, then define the graphs that use them:
```typescript
import {
createStore,
defineNode,
defineEdge,
defineGraph,
type Node,
type StoreProjection,
} from "@nicia-ai/typegraph";
const Document = defineNode("Document", {
schema: z.object({ title: z.string() }),
});
const Chunk = defineNode("Chunk", {
schema: z.object({ text: z.string() }),
});
const Comment = defineNode("Comment", {
schema: z.object({ text: z.string() }),
});
const hasChunk = defineEdge("hasChunk", { from: [Document], to: [Chunk] });
const aboutChunk = defineEdge("aboutChunk", { from: [Comment], to: [Chunk] });
const reviewGraph = defineGraph({
id: "review",
nodes: {
Document: { type: Document },
Chunk: { type: Chunk },
Comment: { type: Comment },
Label: { type: Label },
},
edges: { hasChunk, aboutChunk, hasLabel },
});
const catalogGraph = defineGraph({
id: "catalog",
nodes: {
Document: {
type: Document,
unique: [{ name: "title_unique", fields: ["title"], scope: "kind", collation: "binary" }],
},
Chunk: { type: Chunk },
Comment: { type: Comment },
Category: { type: Category },
},
edges: { hasChunk, aboutChunk, inCategory },
});
```
### Projecting a shared subgraph
Define a projection against either graph — it picks only the shared keys:
```typescript
type CoreStore = StoreProjection<
typeof reviewGraph,
"Document" | "Chunk" | "Comment",
"hasChunk" | "aboutChunk"
>;
```
### Writing a reusable helper
```typescript
async function addComment(
store: CoreStore,
chunk: Node,
text: string,
) {
const comment = await store.nodes.Comment.create({ text });
await store.edges.aboutChunk.create(comment, chunk);
return comment;
}
```
### Using across different graphs
The same `addComment` function works with any store whose graph includes the projected
nodes and edges — even if the graphs diverge on other types or unique constraints:
```typescript
const reviewStore = createStore(reviewGraph, backend);
const catalogStore = createStore(catalogGraph, backend);
await addComment(reviewStore, chunk, "needs revision");
await addComment(catalogStore, chunk, "good categorization");
```
The projection also works inside transactions — `TransactionContext` is structurally
assignable to `StoreProjection` for the same keys:
```typescript
await reviewStore.transaction(async (tx) => {
await addComment(tx, chunk, "transactional comment");
});
```
### What the projection strips
`StoreProjection` erases node constraint names, making constraint-based methods like
`findByConstraint` uncallable through the projection. This is intentional: unique
constraints are graph-registration-level details that typically differ between graphs
sharing the same node types. If you need constraint access, type the helper against a
specific `Store` instead.
## Caveats
**No cross-graph queries**: You cannot traverse from a node in one graph to a node in another. If you need this, consider:
- Merging the graphs into one with clear ontology separation
- Using application-level joins as shown above
**Separate ontology closures**: Each graph computes its own `subClassOf`, `implies`, etc. closures. Ontology relations
don't span graphs.
**Independent transactions**: A transaction in one store doesn't include the other. For cross-graph consistency, use
sagas or eventual consistency patterns.
**Shared tables**: When using the same backend, both graphs write to the same `typegraph_nodes` and `typegraph_edges`
tables, differentiated by `graph_id`. This is fine for most cases but means a database-level issue affects both
graphs.
## Next Steps
- [Multi-Tenant SaaS](./examples/multi-tenant) - Isolating data by tenant within a single graph
- [Schema Migrations](./schema-management) - Versioning and migrations
- [Integration Patterns](./integration) - More deployment strategies
# Ontology & Reasoning
> Semantic relationships, type hierarchies, and inference
## When Do You Need an Ontology?
An ontology captures **meaning** about your data—relationships that exist at the type level, not just instance
level. You need ontology when:
- **Type hierarchies**: "A Podcast is a type of Media" (query for Media, get Podcasts too)
- **Concept relationships**: "Machine Learning is narrower than AI" (topic navigation)
- **Constraints**: "A Person cannot also be an Organization" (prevent invalid data)
- **Edge implications**: Query `knows` through more-specific `marriedTo` rows when explicitly requested
- **Bidirectional queries**: "manages and managedBy are inverses" (traverse in either direction)
Without ontology, you'd implement these manually—if statements scattered throughout your code, hand-rolled
validation, duplicate queries. Ontology centralizes this logic in your schema.
## How It Works
TypeGraph treats semantic relationships between types as **meta-edges**—edges at the type level rather than instance level:
```typescript
// Instance edges: relationships between INSTANCES
// "Alice knows Bob"
const knows = defineEdge("knows");
// Meta-edges: relationships between TYPES
// "Employee subClassOf Person"
subClassOf(Employee, Person);
```
When you define an ontology, TypeGraph:
1. **Precomputes closures** at store initialization (not query time)
2. **Expands only the query operations that explicitly opt in** (except inverse
traversal, whose store default is `"inverse"` and can be changed)
3. **Enforces the documented constraints** when building a registry or writing data
It does not run a general reasoner, materialize implied edges, substitute
properties between types, or automatically expand every query.
## Verified Support Matrix
| Relation / feature | Runtime contract |
| --- | --- |
| `subClassOf` | Transitive registry closure, write-path endpoint assignability, and opt-in node-query expansion with `includeSubClasses` |
| `disjointWith` | Same-ID collision enforcement, propagated through interleaved `subClassOf` and `equivalentTo` closure (`sameAs` remains a deprecated equivalence alias) |
| `implies` | Transitive registry closure and opt-in traversal expansion with `expand: "implying"`; endpoints are validated |
| `inverseOf` | Single inverse partner, endpoint reversal validation, and traversal expansion with `expand: "inverse"` (the default store setting) |
| `equivalentTo` | Registry lookups and graph-merge type reconciliation; no automatic query or property behavior. `sameAs` is folded in as a full alias — the merge type reconciler and the registry treat a `sameAs` declaration identically to `equivalentTo` |
| `broader` / `narrower` | Transitive registry introspection only |
| `partOf` / `hasPart` | Transitive registry introspection only |
| `relatedTo` | Symmetric direct registry introspection through `getRelatedKinds` only |
| Type-level `sameAs` | Deprecated name for `equivalentTo` (see above); prefer calling `equivalentTo` directly |
| Type-level `differentFrom` | Deprecated and decorative — never enforced instance identity; migrate to the graph-level TypeGraph Identity Profile |
| Custom `metaEdge()` properties | Serialized introspection metadata only; custom transitivity, symmetry, inverse, and inference settings are not executed |
## Core Meta-Edges
TypeGraph provides a standard set of meta-edges:
```typescript
import { subClassOf, broader, narrower, equivalentTo, sameAs, differentFrom, disjointWith, partOf, hasPart, relatedTo, inverseOf, implies } from "@nicia-ai/typegraph";
```
### Subsumption (Type Inheritance)
**`subClassOf`**: Defines type inheritance where instances of the child are also instances of the parent.
```typescript
subClassOf(Podcast, Media);
subClassOf(Article, Media);
subClassOf(Company, Organization);
```
**Query Behavior:**
Subclass expansion is **opt-in** via `includeSubClasses: true`:
```typescript
// Without expansion: returns only nodes with kind="Media"
const mediaOnly = await store
.query()
.from("Media", "m")
.select((ctx) => ctx.m)
.execute();
// With expansion: returns Media, Podcast, AND Article nodes
const allMedia = await store
.query()
.from("Media", "m", { includeSubClasses: true })
.select((ctx) => ctx.m)
.execute();
// Results include nodes of kind "Media", "Podcast", and "Article"
```
This is a fundamental difference from traditional ORM inheritance—TypeGraph stores the concrete type
(`kind: "Podcast"`) in the database, and expands at query time when requested.
### Hierarchical (Concept Hierarchy)
**`broader`** and **`narrower`**: Define conceptual hierarchy without identity.
```typescript
broader(MachineLearning, ArtificialIntelligence);
broader(DeepLearning, MachineLearning);
broader(ArtificialIntelligence, Technology);
```
**Important**: This is different from `subClassOf`. A topic instance of "ML" is related to "AI",
but is **not** an instance of "AI".
```typescript
// Get all topics narrower than Technology
const narrowerTopics = registry.expandNarrower("Technology");
// ["ArtificialIntelligence", "MachineLearning", "DeepLearning", ...]
```
### Equivalence
**`equivalentTo`**: Defines semantic equivalence between types or external IRIs.
```typescript
equivalentTo(Person, "https://schema.org/Person");
equivalentTo(Organization, "https://schema.org/Organization");
```
**`sameAs`** and **`differentFrom`** are deprecated type-level factories.
`sameAs` is currently a type-equivalence alias; `differentFrom` is decorative.
For durable individual identity, enable the graph-level TypeGraph Identity
Profile and use `store.identity`. That ledger deliberately does not provide OWL
property substitution or automatic graph-wide query expansion.
### Constraints
**`disjointWith`**: Declares that two types cannot share the same ID.
```typescript
disjointWith(Person, Organization);
disjointWith(Podcast, Article);
```
Disjointness is inherited by subclasses. If `Company subClassOf Organization`,
then `disjointWith(Person, Organization)` also makes `Person` and `Company`
disjoint.
**Effect**: Attempting to create a node that violates disjointness throws `DisjointError`:
```typescript
// Create a Person with ID "entity-1"
await store.nodes.Person.create({ name: "Alice" }, { id: "entity-1" });
// Throws DisjointError: Person and Organization are disjoint
await store.nodes.Organization.create({ name: "Acme" }, { id: "entity-1" });
```
**Coherence rules**: `disjointWith` cannot contradict the rest of the ontology.
A kind disjoint with itself, a kind disjoint with one of its own subclass
ancestors, a common subclass of two disjoint parents, and a kind declared both
`equivalentTo` and `disjointWith` another are all rejected, including overlaps
reached through mixed equivalence/subclass paths. These checks run
both when you construct a graph and when a persisted schema is reloaded, so a
document written by an older, more permissive version can fail validation on
load with a `ConfigurationError` whose details code is
`ONTOLOGY_DISJOINT_CONFLICT`. To recover, fix the graph definition and, for a
persisted schema, correct the stored document before upgrading (or rewrite it
through the previous minor version, which still accepts it). The same
construction-and-reload rule applies to the other ontology coherence checks
(duplicate relations, hierarchical self-loops and cycles, and inverse-partner
uniqueness).
### Composition
**`partOf`** and **`hasPart`**: Define compositional relationships.
```typescript
partOf(Chapter, Book);
hasPart(Book, Chapter);
partOf(Episode, Podcast);
hasPart(Podcast, Episode);
```
### Edge Relationships
**`inverseOf`**: Declares two edge kinds as inverses of each other.
```typescript
inverseOf(manages, managedBy);
inverseOf(cites, citedBy);
inverseOf(follows, followedBy);
```
**Effect**: You can query in either direction using the registry:
```typescript
const inverse = registry.getInverseEdge("manages"); // "managedBy"
```
You can also expand traversals to include inverse edge kinds at query time:
```typescript
const relationships = await store
.query()
.from("Person", "p")
.traverse("manages", "e", { expand: "inverse" })
.to("Person", "other")
.select((ctx) => ({
other: ctx.other.name,
via: ctx.e.kind,
}))
.execute();
```
For symmetric relationships, declare an edge as its own inverse:
```typescript
inverseOf(collaboratesWith, collaboratesWith);
```
An edge may have only one distinct inverse partner. Every allowed pair must be
compatible with a reversed pair in its partner, in both traversal directions,
using equal kinds or `subClassOf` assignability. Matching the independent source
and target unions is insufficient for source-dependent edges. A self-inverse
edge must satisfy the same reversed-pair check against itself.
**`implies`**: Declares that one edge kind implies another exists.
```typescript
implies(marriedTo, knows);
implies(bestFriends, friends);
implies(friends, knows);
```
**Effect**: Query for `knows` can include `marriedTo`, `bestFriends`, and `friends` edges:
```typescript
const connections = await store
.query()
.from("Person", "p")
.traverse("knows", "e", { expand: "implying" })
.to("Person", "other")
.select((ctx) => ctx.other)
.execute();
```
**Endpoint compatibility is required.** `implies(edgeA, edgeB)` only makes
sense if every node kind `edgeA` can connect could also, in principle,
satisfy `edgeB`'s own domain/range — otherwise `expand: "implying"` would
traverse rows whose kinds don't match what the traversal actually asked for.
Every allowed pair in `edgeA` must match a single allowed pair in `edgeB`:
both endpoints must be assignable — equal, or a `subClassOf` descendant —
to their corresponding endpoint in that pair. For
[source-dependent targets](/core-concepts#source-dependent-targets), finding
the source in one entry and the target in another does not suffice.
An incompatible pair (say, `Author -> Paper` implying
`Paper -> Topic`) throws `ConfigurationError` wherever the graph is built
into a store or committed as a schema version (`createStore`,
`createStoreWithSchema`, `store.evolve({ ontology })`) — including relations
authored through a graph extension, not just `implies()` calls in code.
## Using the Ontology
### In Graph Definition
```typescript
const graph = defineGraph({
id: "knowledge_base",
nodes: { ... },
edges: { ... },
ontology: [
// Type hierarchy
subClassOf(Podcast, Media),
subClassOf(Article, Media),
subClassOf(Company, Organization),
// Concept hierarchy
broader(MachineLearning, ArtificialIntelligence),
broader(DeepLearning, MachineLearning),
// Constraints
disjointWith(Person, Organization),
disjointWith(Media, Person),
// Composition
partOf(Episode, Podcast),
// Edge relationships
inverseOf(cites, citedBy),
implies(marriedTo, knows),
],
});
```
### Registry Lookups
The type registry (accessed via `store.registry`) provides methods to query the ontology:
```typescript
const registry = store.registry;
// Subsumption
registry.isSubClassOf("Podcast", "Media"); // true
registry.expandSubClasses("Media"); // ["Media", "Podcast", "Article"]
// Hierarchy
registry.expandNarrower("Technology"); // ["AI", "ML", "DL", ...]
registry.expandBroader("DeepLearning"); // ["ML", "AI", "Technology"]
// Constraints
registry.areDisjoint("Person", "Organization"); // true
registry.getDisjointKinds("Person"); // ["Organization", "Media", ...]
// Edge relationships
registry.getInverseEdge("cites"); // "citedBy"
registry.getImpliedEdges("marriedTo"); // ["knows"]
registry.getImplyingEdges("knows"); // ["marriedTo", "bestFriends", "friends"]
registry.getRelatedKinds("MachineLearning"); // ["DataScience", ...]
```
## Custom Meta-Edges
Define domain-specific meta-edges for serialized introspection metadata:
```typescript
import { metaEdge } from "@nicia-ai/typegraph";
// Custom meta-edge for prerequisite relationships
const prerequisiteOf = metaEdge("prerequisiteOf", {
transitive: true,
inference: "hierarchy",
description: "Learning prerequisite (Calculus prerequisiteOf LinearAlgebra)",
});
// Custom meta-edge for superseding relationships
const supersedes = metaEdge("supersedes", {
transitive: true,
inference: "substitution",
description: "Replacement relationship (v2 supersedes v1)",
});
```
### Meta-Edge Properties
Each custom meta-edge can carry these properties as metadata. In the current
release they do **not** make the registry compute a custom closure or make the
query builder execute custom inference. Only the built-in relations in the
support matrix have runtime behavior.
| Property | Type | Description |
| ------------ | --------------- | ------------------------- |
| `transitive` | `boolean` | A→B, B→C implies A→C |
| `symmetric` | `boolean` | A→B implies B→A |
| `reflexive` | `boolean` | A→A is always true |
| `inverse` | `string` | Name of inverse meta-edge |
| `inference` | `InferenceType` | How this affects queries |
### Inference Types
For custom meta-edges, `inference` is descriptive metadata for consumers:
| Type | Description |
| ---------------- | -------------------------------------------- |
| `"subsumption"` | Query for X includes instances of subclasses |
| `"hierarchy"` | Enables broader/narrower traversal |
| `"substitution"` | Can substitute equivalent types |
| `"constraint"` | Validation rules |
| `"composition"` | Part-whole navigation |
| `"association"` | Discovery/recommendation |
| `"none"` | No automatic inference |
## Closure Computation
TypeGraph precomputes transitive closures at store initialization:
```typescript
// subClassOf closure
// If: Podcast subClassOf Media, Episode subClassOf Media
// Then: expandSubClasses("Media") = ["Media", "Podcast", "Episode"]
// implies closure
// If: marriedTo implies partneredWith, partneredWith implies knows
// Then: getImpliedEdges("marriedTo") = ["partneredWith", "knows"]
```
This makes queries efficient—expansion happens at query compilation time, not execution time.
## Best Practices
### Separate `subClassOf` from `broader`
These have different semantics:
- `subClassOf`: Type membership (a Podcast instance is also a Media instance)
- `broader`: Conceptual relation (ML **relates to** AI, but ML instance ≠ AI instance)
```typescript
// CORRECT: Type hierarchy
subClassOf(Podcast, Media);
// CORRECT: Concept hierarchy
broader(MachineLearning, ArtificialIntelligence);
// WRONG: Don't mix them
// subClassOf(MachineLearning, ArtificialIntelligence);
```
### Use Disjoint Constraints
Prevent impossible combinations:
```typescript
// Good: Prevent ID conflicts
disjointWith(Person, Organization);
disjointWith(Person, Product);
disjointWith(Organization, Product);
```
### Model Edge Hierarchies with Implies
```typescript
// Relationship hierarchy: specific → general
implies(marriedTo, partneredWith);
implies(partneredWith, knows);
implies(parentOf, relatedTo);
implies(siblingOf, relatedTo);
implies(relatedTo, knows);
```
### Use InverseOf for Bidirectional Queries
```typescript
inverseOf(manages, managedBy);
inverseOf(follows, followedBy);
inverseOf(cites, citedBy);
```
This lets you query efficiently in either direction without duplicating edges.
## API Reference
### Ontology Functions
#### `subClassOf(child, parent)`
Declares type inheritance.
```typescript
function subClassOf(child: NodeType, parent: NodeType): OntologyRelation;
```
#### `broader(narrower, broader)`
Declares hierarchical relationship (narrower concept to broader concept).
```typescript
function broader(narrower: NodeType, broader: NodeType): OntologyRelation;
```
#### `narrower(broader, narrower)`
Declares hierarchical relationship (broader concept to narrower concept).
```typescript
function narrower(broader: NodeType, narrower: NodeType): OntologyRelation;
```
#### `equivalentTo(a, b)`
Declares semantic equivalence between types or with external IRIs.
```typescript
function equivalentTo(
a: NodeType | string,
b: NodeType | string
): OntologyRelation;
```
#### `sameAs(kindA, kindBOrIri)`
Deprecated type-level alias of `equivalentTo`, including the equivalence with
external IRIs. Migrate to the graph-level TypeGraph Identity Profile for
individual identity.
```typescript
function sameAs(kindA: NodeType, kindBOrIri: NodeType | string): OntologyRelation;
```
#### `differentFrom(a, b)`
Deprecated decorative type-level relation. Migrate to the graph-level TypeGraph
Identity Profile for individual identity.
```typescript
function differentFrom(a: NodeType, b: NodeType): OntologyRelation;
```
#### `disjointWith(a, b)`
Declares mutual exclusion (types cannot share the same ID).
```typescript
function disjointWith(a: NodeType, b: NodeType): OntologyRelation;
```
#### `partOf(part, whole)`
Declares compositional relationship (part to whole).
```typescript
function partOf(part: NodeType, whole: NodeType): OntologyRelation;
```
#### `hasPart(whole, part)`
Declares compositional relationship (whole to part).
```typescript
function hasPart(whole: NodeType, part: NodeType): OntologyRelation;
```
#### `relatedTo(a, b)`
Declares a symmetric association available through
`registry.getRelatedKinds(kind)`. It has no query behavior.
```typescript
function relatedTo(a: NodeType, b: NodeType): OntologyRelation;
```
#### `inverseOf(edgeA, edgeB)`
Declares edge types as inverses of each other.
```typescript
function inverseOf(edgeA: AnyEdgeType, edgeB: AnyEdgeType): OntologyRelation;
```
#### `implies(edgeA, edgeB)`
Declares that one edge type implies another exists.
```typescript
function implies(edgeA: AnyEdgeType, edgeB: AnyEdgeType): OntologyRelation;
```
Each allowed pair in `edgeA` must be assignable to one allowed pair in `edgeB`
(equal, or a `subClassOf` descendant, on both endpoints). Throws
`ConfigurationError` when the graph is built into a store or committed as a
schema version if they aren't — see [Edge Relationships](#edge-relationships)
above.
#### `metaEdge(name, options?)`
Creates a custom meta-edge for domain-specific relationships.
```typescript
function metaEdge(
name: string,
options?: {
transitive?: boolean;
symmetric?: boolean;
reflexive?: boolean;
inverse?: string;
inference?: InferenceType;
description?: string;
},
): MetaEdge;
```
### Type Registry API
The type registry is available via `store.registry` and provides methods to query the ontology at runtime.
#### `isSubClassOf(child, parent)`
Checks if a type is a subclass of another.
```typescript
registry.isSubClassOf(child: string, parent: string): boolean;
registry.isSubClassOf("Podcast", "Media"); // true
```
#### `expandSubClasses(type)`
Returns a type and all its subclasses.
```typescript
registry.expandSubClasses(type: string): readonly string[];
registry.expandSubClasses("Media"); // ["Media", "Podcast", "Article"]
```
#### `areDisjoint(a, b)`
Checks if two types are disjoint.
```typescript
registry.areDisjoint(a: string, b: string): boolean;
registry.areDisjoint("Person", "Organization"); // true
```
#### `getDisjointKinds(type)`
Returns all types disjoint with the given type.
```typescript
registry.getDisjointKinds(type: string): readonly string[];
registry.getDisjointKinds("Person"); // ["Organization", "Media", ...]
```
#### `expandNarrower(type)`
Returns all types narrower than the given type (via `broader` relationships).
```typescript
registry.expandNarrower(type: string): readonly string[];
registry.expandNarrower("Technology"); // ["AI", "ML", "DeepLearning", ...]
```
#### `expandBroader(type)`
Returns all types broader than the given type.
```typescript
registry.expandBroader(type: string): readonly string[];
registry.expandBroader("DeepLearning"); // ["MachineLearning", "AI", "Technology"]
```
#### `getInverseEdge(edgeType)`
Returns the inverse of an edge type.
```typescript
registry.getInverseEdge(edgeType: string): string | undefined;
registry.getInverseEdge("manages"); // "managedBy"
```
#### `getImpliedEdges(edgeType)`
Returns edges implied by an edge type.
```typescript
registry.getImpliedEdges(edgeType: string): readonly string[];
registry.getImpliedEdges("marriedTo"); // ["knows"]
```
#### `getImplyingEdges(edgeType)`
Returns edges that imply an edge type.
```typescript
registry.getImplyingEdges(edgeType: string): readonly string[];
registry.getImplyingEdges("knows"); // ["marriedTo", "bestFriends", "friends"]
```
#### `expandImplyingEdges(edgeType)`
Returns an edge type and all edges that imply it.
```typescript
registry.expandImplyingEdges(edgeType: string): readonly string[];
registry.expandImplyingEdges("knows"); // ["knows", "marriedTo", "bestFriends", "friends"]
```
# Indexes
> Define and create indexes for TypeGraph queries
TypeGraph stores node and edge properties in a JSON `props` column. When you filter or order by
JSON properties at scale, you typically need **expression indexes** on those JSON paths.
TypeGraph includes built-in indexes for common access patterns (lookups by ID, edge traversals,
temporal filtering), but application-specific indexes are up to you.
The `@nicia-ai/typegraph/indexes` entrypoint provides:
- **Type-safe index definitions** for node and edge schemas
- **Dialect-specific DDL generation** for PostgreSQL and SQLite
- **Drizzle schema integration** so drizzle-kit can generate migrations
- **Profiler integration** so recommendations account for indexes you already have
## Quick Start (Drizzle / drizzle-kit)
Define your indexes once and pass them into the Drizzle schema factories:
```ts
import { defineEdge, defineNode } from "@nicia-ai/typegraph";
import { createPostgresTables } from "@nicia-ai/typegraph/adapters/drizzle/postgres";
import { andWhere, defineEdgeIndex, defineNodeIndex } from "@nicia-ai/typegraph/indexes";
import { z } from "zod";
const Person = defineNode("Person", {
schema: z.object({
email: z.string().email(),
name: z.string(),
createdAt: z.date(),
isActive: z.boolean().optional(),
}),
});
const worksAt = defineEdge("worksAt", {
schema: z.object({
role: z.string(),
}),
});
export const personEmail = defineNodeIndex(Person, {
fields: ["email"],
unique: true,
coveringFields: ["name"],
where: (w) => andWhere(w.deletedAt.isNull(), w.isActive.eq(true)),
});
export const worksAtRoleOut = defineEdgeIndex(worksAt, {
fields: ["role"],
direction: "out",
where: (w) => w.deletedAt.isNull(),
});
// drizzle-kit will include these indexes in generated migrations
export const typegraphTables = createPostgresTables(
{},
{
indexes: [personEmail, worksAtRoleOut],
},
);
```
For SQLite, use `createSqliteTables`:
```ts
import { createSqliteTables } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
export const typegraphTables = createSqliteTables(
{},
{
indexes: [personEmail, worksAtRoleOut],
},
);
```
## Node Indexes
`defineNodeIndex(nodeType, config)` creates an index definition for node properties and TypeGraph
system columns.
**Key options:**
- `keys`: ordered node B-tree keys with an explicit direction. Property and system-column keys can
be interleaved to match a complete query order. This node-only option is mutually exclusive with
`fields` and `keySystemColumns`, and does not define a `bulkFindByIndex` lookup key.
- `fields`: JSON property paths used for filtering/ordering (B-tree expression keys). Optional if
`keys`, `coveringFields`, or `keySystemColumns` supply the key instead.
- `coveringFields`: additional properties frequently selected with the same filters. These become
additional index keys to enable index-only reads when combined with smart select.
- `keySystemColumns`: system columns (e.g. `"id"`) to include in the key, after the `scope` prefix
and before `fields`. See [Keying on system columns](#keying-on-system-columns-keysystemcolumns).
Rejects `"id"` combined with `unique: true` — every node's `id` is already unique per row, so a
unique index keyed on `id` plus other columns can never enforce a meaningful constraint across
those other columns.
- `unique`: create a unique index.
- `scope`: prefixes index keys with TypeGraph system columns (default is `"graphAndKind"`).
- `where`: partial index predicate (portable DSL, compiled per dialect).
:::note[Covering fields vs PostgreSQL INCLUDE]
TypeGraph properties live inside `props`, so indexes are built on expressions. PostgreSQL `INCLUDE`
does not support expressions, so `coveringFields` are implemented as additional index keys rather
than an `INCLUDE (...)` clause.
Because `coveringFields` become index keys, they:
- Increase index size (more key data)
- Affect index ordering (can help `ORDER BY`, but changes sort/range behavior)
- Must be maintained on writes like any other key
:::
Field ordering emits the database's native `NULLS FIRST` / `NULLS LAST` suffix. This keeps the
primary sort expression aligned with a matching B-tree expression index. PostgreSQL and SQLite
still choose plans from their statistics and the full query shape, so confirm important queries
with `EXPLAIN`; an applicable index does not guarantee that every data distribution will use it.
### Directed node keys
Use `keys` when the direction and interleaving of the B-tree keys must match a query's complete
ordering. Scope columns remain first, keys retain declaration order, and `coveringFields` remain
last:
```ts
const newestPerson = defineNodeIndex(Person, {
keys: [
{ field: "createdAt", direction: "desc" },
{ system: "id", direction: "asc" },
],
coveringFields: ["name"],
});
const page = await store
.query()
.from("Person", "person")
.select((fields) => ({ name: fields.person.name }))
.orderBy("person", "createdAt", "desc")
.orderBy("person", "id", "asc")
.paginate({ first: 50 });
```
The resulting key order is `(graph_id, kind, createdAt DESC, id ASC, name)`. In the first version,
`keys` is supported for node B-tree indexes only and refuses
`unique: true`. It intentionally does not participate in `bulkFindByIndex`: that operation probes
the equality lookup fields declared through `fields`, while `keys` describes physical ordering.
Use a separate legacy `fields` index when the same graph also needs keyed candidate lookup.
### Nested JSON Paths
For top-level properties, use the field name:
```ts
defineNodeIndex(Person, { fields: ["email"] });
```
For nested properties inside `props`, use a JSON pointer:
```ts
defineNodeIndex(Person, { fields: ["/metadata/priority"] });
```
You can also pass pointer segments:
```ts
defineNodeIndex(Person, { fields: [["metadata", "priority"] as const] });
```
### Index Scope
Index `scope` controls which TypeGraph system columns are prefixed ahead of your JSON keys:
- `"graphAndKind"` (default): prefixes with `(graph_id, kind)` to match most TypeGraph queries.
- `"graph"`: prefixes with `graph_id` only (rare; useful for cross-kind queries within a graph).
- `"none"`: no system prefix (rare; usually only correct for global queries).
### Keying on system columns (`keySystemColumns`)
`fields` and `coveringFields` only ever reference your schema's own properties, and `scope` only
ever prefixes `graph_id`/`kind` — neither can put a system column like `id` into the index **key**.
Some query shapes need that. A reverse traversal — `.traverse("hasCreator", { direction: "in"
}).to("Post", "post")` — compiles a join on the target node's own `id` (`n.id = e.from_id`). If you
also filter, sort, or select a prop on that same node (say `creationDate`), a covering index needs
`id` in its key to match that join — `graph_id`/`kind` and `creationDate` alone aren't enough,
because the index can't be chosen for an `id`-equality join it doesn't cover.
```ts
const postRecent = defineNodeIndex(Post, {
keySystemColumns: ["id"],
coveringFields: ["creationDate"],
});
// -> CREATE INDEX ... ON typegraph_nodes (graph_id, kind, id, (props #>> ARRAY['creationDate']))
```
**Validation:**
- Node indexes only. Rejects edge-only system columns (`fromKind` / `fromId` / `toKind` / `toId`).
- Rejects a column already implied by `scope` (e.g. `graph_id` when `scope: "graphAndKind"`).
- Not supported with `method: "gin" | "trigram"` (same restriction as `coveringFields`).
## Edge Indexes
`defineEdgeIndex(edgeType, config)` works the same way as node indexes, with one extra option:
- `direction`: `"out" | "in" | "none"` (default `"none"`). When set, the index keys are prefixed
with the join key used by traversal queries (`from_id` for `"out"`, `to_id` for `"in"`).
This makes it easy to create indexes that match `.traverse()` patterns.
**When to use `direction`:**
- `"out"`: optimize outbound traversals that join on `from_id` (start node → edges).
- `"in"`: optimize inbound traversals that join on `to_id` (end node → edges).
- `"none"`: for edge queries not anchored by a traversal join key (less common).
## Partial Indexes (WHERE)
Use `where` to create partial indexes with a small, typed predicate DSL.
System columns are available (e.g. `deletedAt`, `createdAt`, `fromId`), as well as your schema
properties (e.g. `email`, `role`).
```ts
import { andWhere, defineNodeIndex } from "@nicia-ai/typegraph/indexes";
const activeEmail = defineNodeIndex(Person, {
fields: ["email"],
where: (w) => andWhere(w.deletedAt.isNull(), w.isActive.eq(true)),
});
```
## Covering Indexes
To maximize the benefit of [smart select optimization](/performance/overview#smart-select),
create indexes that include both the filter columns and selected columns. This enables index-only
scans where the database satisfies the entire query from the index.
```ts
// Index covers email filter AND name selection
const personEmailWithName = defineNodeIndex(Person, {
fields: ["email"],
coveringFields: ["name"],
where: (w) => w.deletedAt.isNull(),
});
```
**Generated PostgreSQL:**
```sql
CREATE INDEX idx_person_email_name ON typegraph_nodes
(graph_id, kind, ((props #>> ARRAY['email'])), ((props #>> ARRAY['name'])))
WHERE deleted_at IS NULL;
```
**Generated SQLite:**
```sql
CREATE INDEX idx_person_email_name ON typegraph_nodes
(graph_id, kind, json_extract(props, '$.email'), json_extract(props, '$.name'))
WHERE deleted_at IS NULL;
```
:::caution[PostgreSQL: JSONB expression indexes don't get a true Index Only Scan]
A covering index like the one above lets PostgreSQL avoid a full **table scan**, but not
necessarily a **heap fetch** per matching row. PostgreSQL's `Index Only Scan` optimization — skip
the heap entirely when the index already has everything the query needs — does not extend to
JSONB extraction expressions (`props #>> ARRAY[...]`), only to real stored columns. Even a
correctly-shaped `coveringFields` index shows as a plain `Index Scan` in `EXPLAIN`, not
`Index Only Scan`, and PostgreSQL still visits the heap row for every match.
This matters most for high-fan-in ranked reads — e.g. "each of N friends' most recent post, top 10
overall" — where a query visits many candidate rows and discards most of them after sorting. If
`EXPLAIN (ANALYZE, BUFFERS)` shows most of a query's cost coming from a repeated `Index Scan` with
high `Buffers: shared hit` relative to the rows actually returned, this is likely why — confirmed
against a real workload at real scale, not a theoretical concern.
TypeGraph doesn't currently offer a built-in way to materialize a hot prop as a real column (this
is an active area of investigation). Until then, if this is a genuine hot-path bottleneck on
PostgreSQL, the workaround is maintaining your own
[stored generated column](https://www.postgresql.org/docs/current/ddl-generated-columns.html)
alongside `props` with a plain index over it — a real column does get `Index Only Scan`, confirmed
via `EXPLAIN (ANALYZE, BUFFERS)` showing `Heap Fetches: 0` (run `VACUUM ANALYZE`, not just
`ANALYZE`, after backfilling — the visibility map needs to be current before PostgreSQL will prove
it immediately).
We haven't specifically verified whether SQLite's JSON-extraction covering indexes have the same
limitation for this query shape.
:::
## Batched Index Lookup (`bulkFindByIndex`)
`store.nodes..bulkFindByIndex(indexName, items, options?)` takes many in-memory records and
returns the live nodes that share each record's declared **index key** — batched candidate retrieval
for import reconciliation, dedup-candidate discovery, and joining incoming records against the graph
by a declared composite key.
```ts
const candidates = await store.nodes.Person.bulkFindByIndex("person_active_name", [
{ props: { isActive: true, name: "Ana" } },
{ props: { name: "Bo" } }, // missing isActive → matches stored null
]);
// readonly Node[][] — one bucket per input, ordered by node id
```
Semantics:
- **One bucket per input**, in input order; empty input returns `[]`. The index may be non-unique, so
each bucket is a (possibly empty) array — this is candidate retrieval, not a uniqueness guarantee.
For unique lookups prefer `bulkFindByConstraint` (backed by the uniqueness side-table).
- TypeGraph computes the lookup key from **`index.fields` only** (JSON-pointer extraction, reusing the
index's own extraction expressions). `keys`, `coveringFields`, and `keySystemColumns` are not part
of the probe key. An index declared without `fields` has nothing to probe by and throws
`ConfigurationError`.
- The index's partial `where` is applied in SQL to **stored** rows only; probes carry index-field
values, nothing else. Only the indexed fields are validated — full records are not required.
- A missing/`undefined` indexed field matches stored `NULL` (null-safe equality). Live,
non-soft-deleted nodes only.
- `options.limitPerInput` caps each bucket (ordered by node id); unbounded by default — no silent
truncation. A non-positive value throws `ValidationError`. An unknown index name throws
`NodeIndexNotFoundError`. A non-scalar probe value throws `ValidationError`. On backends that
support SQL window functions the cap is applied in-database (`ROW_NUMBER()`); on backends without
them (`capabilities.windowFunctions: false`) it degrades to an in-memory cap after fetching the
matching ids — same result, but it transfers all matching ids for low-selectivity keys.
- **Key field types:** string, number, and boolean keys are supported. **Date-typed key fields are
not** — they throw `ConfigurationError`, because SQLite compares stored ISO text byte-wise while
PostgreSQL compares `timestamptz` instants, so the same instant in different ISO forms would match
on one backend but not the other. Use a string-encoded key, or `store.query(...).where(...)` for
date predicates.
The lookup is correct whether or not the physical index has been materialized; materialize it (see
below) for the query planner to actually use it. Null-safe predicates may be less reliably
index-accelerated than plain equality.
## Choosing the Right Index Type
TypeGraph's `defineNodeIndex` / `defineEdgeIndex` generate **B-tree expression indexes** — the right
choice for scalar equality, range, and ordering queries. But JSON properties can also hold arrays
and objects, which need different index strategies.
| Data shape | Query pattern | Index type | TypeGraph utility? |
| -------------------------------------- | ---------------------------------------------- | ------------------------ | ------------------------------------ |
| Scalar (`string`, `number`, `boolean`) | `eq()`, `gt()`, `in()`, `orderBy()` | B-tree expression | Yes — `defineNodeIndex` |
| Array of scalars | `contains()`, `containsAll()`, `containsAny()` | GIN (PostgreSQL) | No — use raw SQL |
| Nested object | `hasKey()`, `pathEquals()`, `pathContains()` | GIN or B-tree expression | Partially — B-tree on specific paths |
### B-tree expression indexes (scalar properties)
Best for equality, range, sorting, and prefix matching on individual JSON fields. This is what
`defineNodeIndex` and `defineEdgeIndex` generate.
```ts
// Good for: .whereNode("p", (p) => p.email.eq("..."))
defineNodeIndex(Person, { fields: ["email"] });
// Good for: .orderBy("p", "createdScore", "desc")
defineNodeIndex(Person, { fields: ["createdScore"] });
```
### GIN indexes (array containment — PostgreSQL only)
TypeGraph compiles array predicates to PostgreSQL's JSONB containment operator over the field's
extraction expression — `(props #> ARRAY['tags']) @> $1`. Declare a containment index with
`method: "gin"` and TypeGraph emits the matching **expression GIN** (`jsonb_path_ops`):
```typescript
const personTags = defineNodeIndex(Person, {
fields: ["tags"],
method: "gin",
});
// materialized via store.materializeIndexes():
// CREATE INDEX ... USING GIN (("props" #> ARRAY['tags']) jsonb_path_ops);
```
The index accelerates all containment predicates on that field:
```typescript
// contains: does the tags array include "typescript"?
.whereNode("p", (p) => p.tags.contains("typescript"))
// containsAll: does it include BOTH "typescript" AND "graphql"?
.whereNode("p", (p) => p.tags.containsAll(["typescript", "graphql"]))
// containsAny: does it include "typescript" OR "graphql"?
.whereNode("p", (p) => p.tags.containsAny(["typescript", "graphql"]))
```
:::caution[Whole-column `GIN (props)` does not work]
PostgreSQL matches expression indexes structurally. A whole-column
`CREATE INDEX ... USING GIN (props)` serves `props @> …` — **not** the
per-field `(props #> ARRAY['tags']) @> …` expressions TypeGraph compiles,
so such an index is never used. Declare `method: "gin"` per field (or
hand-write the same expression form) instead.
:::
:::note[SQLite]
SQLite has no GIN equivalent; `materializeIndexes()` reports gin/trigram
declarations as `skipped` there. Array containment on SQLite uses
`json_each()` scans, which can't be indexed.
:::
### Trigram indexes (substring and case-insensitive matching — PostgreSQL only)
`contains` / `startsWith` / `endsWith` / `ilike` on string fields compile to
`ILIKE` on PostgreSQL, which a B-tree can never serve for infix patterns.
Declare `method: "trigram"` and TypeGraph emits an expression GIN with
`gin_trgm_ops` (installing the `pg_trgm` extension on first
materialization):
```typescript
const personName = defineNodeIndex(Person, {
fields: ["name"],
method: "trigram",
});
// CREATE INDEX ... USING GIN (("props" #>> ARRAY['name']) gin_trgm_ops);
.whereNode("p", (p) => p.name.contains("smith")) // served by the index
.whereNode("p", (p) => p.name.ilike("%SMITH%")) // also served
```
On SQLite these declarations are `skipped` — SQLite's substring-search
story is [FTS5 fulltext](/fulltext-search) via `searchable()` fields.
GIN-family methods take exactly one field and don't support `unique`,
`coveringFields`, or `where`; the query's `graph_id` / `kind` filters apply
as residual conditions over the index's candidate rows.
### Combining B-tree and GIN
For kinds where you filter on both scalar fields (equality, range) and array or substring
predicates, declare both index types:
```typescript
const personEmail = defineNodeIndex(Person, { fields: ["email"] });
const personTags = defineNodeIndex(Person, { fields: ["tags"], method: "gin" });
```
PostgreSQL's query planner can use both indexes together via a BitmapAnd scan when a query filters
on both a scalar field and an array field.
## Generating SQL (No drizzle-kit)
If you manage migrations yourself, generate DDL snippets:
```ts
import { generateIndexDDL } from "@nicia-ai/typegraph/indexes";
const sql = generateIndexDDL(personEmail, "postgres");
// → CREATE INDEX ...;
```
## Verifying Index Usage
Use `EXPLAIN ANALYZE` to verify your indexes are being used:
```sql
-- PostgreSQL
EXPLAIN ANALYZE SELECT props #>> ARRAY['email'], props #>> ARRAY['name']
FROM typegraph_nodes
WHERE graph_id = 'my_graph'
AND kind = 'Person'
AND deleted_at IS NULL
AND (props #>> ARRAY['email']) = 'alice@example.com';
-- SQLite
EXPLAIN QUERY PLAN SELECT json_extract(props, '$.email'), json_extract(props, '$.name')
FROM typegraph_nodes
WHERE graph_id = 'my_graph'
AND kind = 'Person'
AND deleted_at IS NULL
AND json_extract(props, '$.email') = 'alice@example.com';
```
Look for "Index Scan" or "Index Only Scan" (PostgreSQL) or "USING INDEX" (SQLite) in the output.
## Profiler Integration
Pass your existing indexes to the [Query Profiler](/performance/profiler) so recommendations
focus on what you *don't* have:
```ts
import { QueryProfiler } from "@nicia-ai/typegraph/profiler";
import { toDeclaredIndexes } from "@nicia-ai/typegraph/indexes";
const profiler = new QueryProfiler({
declaredIndexes: toDeclaredIndexes([personEmail, worksAtRoleOut]),
});
```
## Limitations
- The default `defineNodeIndex` / `defineEdgeIndex` method generates B-tree expression indexes for
**scalar** properties (`string`, `number`, `boolean`, `Date`). Array containment and substring
matching are served by [`method: "gin"`](#gin-indexes-array-containment--postgresql-only) and
[`method: "trigram"`](#trigram-indexes-substring-and-case-insensitive-matching--postgresql-only)
declarations.
- GIN-family methods are PostgreSQL-only; `materializeIndexes()` reports them as `skipped` on
SQLite, which has no equivalent for JSON containment acceleration (substring search on SQLite is
served by FTS5 fulltext).
- Embedding fields live in per-`(graphId, kind, field)` vector tables (`tg_vec_*`) and are indexed
through `store.materializeIndexes()` (pgvector builds an HNSW / IVFFlat ANN index; sqlite-vec and
libSQL report `skipped`/build their own). See [Semantic Search](/semantic-search).
## System Indexes
TypeGraph ships a set of **system indexes** on its own relations (nodes, edges, and the recorded
history tables) — the traversal, listing, temporal-validity, and bare-id access paths every
compiled query relies on. They are declared once (`SYSTEM_INDEX_DECLARATIONS`, exported for
inspection), and both dialects' schemas derive from that single list, so SQLite and PostgreSQL
always carry the same set.
You normally never manage them: fresh databases get them at bootstrap, and
`createStoreWithSchema()` brings an already-initialized database up to the running library
version's set on boot (with `CREATE INDEX CONCURRENTLY` on PostgreSQL). Deployments that boot
without `createStoreWithSchema` can run `store.materializeSystemIndexes()` once after a library
upgrade; it shares `materializeIndexes()`'s status tracking, drift signatures, and concurrent-build
claim protocol, and settles with no DDL when everything already exists.
## Next Steps
- [Performance Overview](/performance/overview) — Best practices, N+1 prevention, batch patterns
- [Query Profiler](/performance/profiler) — Automatic index recommendations
# Performance Overview
> Understanding the performance characteristics of TypeGraph
TypeGraph is designed to be a high-performance, low-overhead layer on top of
your relational database. By leveraging the power of modern SQL engines (SQLite
and PostgreSQL) and precomputing complex relationships, TypeGraph ensures that
your knowledge graph scales with your application.
## Performance Philosophy
1. **One Fluent Query, One Statement**: Every fluent query — including multi-hop traversals —
compiles to a single SQL statement, so its statement count never grows with the size of the
graph. This prevents compiler-generated N+1 work inside that query; application code can still
create an N+1 by issuing separate reads in a loop. (Compilation, not execution: a query whose
selective-field mapping falls back re-runs as a full fetch, costing a second statement. See
[Batch reads](#batch-reads).)
2. **Precomputed Ontology**: Transitive closures, subclass hierarchies, and edge implications are
computed once at schema initialization, not during every query.
3. **Batching & Transactions**: Bulk collection APIs minimize round-trips for writes. On the read
side that job belongs to the query compiler — `store.batch()` only caps concurrency at one query
in flight, it does not reduce round trips and it is not a snapshot.
4. **Zero-Cost Abstractions**: Type safety and ontological reasoning add no measurable runtime overhead.
## N+1 Prevention
A common performance problem in ORMs is the N+1 query: you fetch N entities, then issue one
query per entity to load related data. TypeGraph's fluent query compiler eliminates that pattern
inside one graph-shaped query; it cannot eliminate separate collection reads issued by application
code.
Every query — regardless of how many traversals it chains — compiles to a **single SQL statement**
using Common Table Expressions (CTEs). Each traversal step becomes a CTE that joins against the
previous one:
```typescript
// This compiles to ONE SQL statement, not 3 separate queries
const results = await store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.name.eq("Alice"))
.traverse("worksAt", "employment")
.to("Company", "c")
.traverse("locatedIn", "location")
.to("City", "city")
.select((ctx) => ({
person: ctx.p.name,
company: ctx.c.name,
city: ctx.city.name,
}))
.execute();
```
The generated SQL looks like:
```sql
WITH cte_p AS (
SELECT ... FROM typegraph_nodes
WHERE graph_id = ? AND kind IN ('Person') AND ...
),
cte_employment AS (
SELECT ... FROM typegraph_edges e
JOIN typegraph_nodes n ON ...
WHERE e.graph_id = ? AND ...
),
cte_location AS (
SELECT ... FROM typegraph_edges e
JOIN typegraph_nodes n ON ...
WHERE e.graph_id = ? AND ...
)
SELECT ... FROM cte_p
JOIN cte_employment ON ...
JOIN cte_location ON ...
```
This holds for all query types:
- Multi-hop traversals (N CTEs, 1 statement)
- [Recursive traversals](/queries/recursive) (WITH RECURSIVE, 1 statement)
- Aggregations with traversals (CTEs + GROUP BY, 1 statement)
- [Set operations](/queries/combine) (UNION/INTERSECT/EXCEPT of CTEs, 1 statement)
The fluent query needs no dataloader for that joined read because the database handles its entire
join graph in one execution. Separate reads can still form an N+1; use a traversal, `batchOnce()`,
`neighbors()` / `countNeighbors()`, or `subgraph()`. Inside `batchOnce()`, its batch-scoped `read` builder
creates composable versions of the set-oriented reads when unlike result shapes must share one
statement. Chunked collection reads remain useful for homogeneous ID and endpoint sets.
## Batch Write Patterns
### Remote edge convergence
For a latency-sensitive `getOrCreateByEndpoints()` path, declare the canonical
identity on the edge registration instead of supplying an ad hoc `matchOn` list
at each call:
```typescript
const graph = defineGraph({
id: "work",
nodes: { Person: { type: Person }, Company: { type: Company } },
edges: {
worksAt: {
type: worksAt,
from: [Person],
to: [Company],
cardinality: "many",
matchIdentity: { name: "employment", fields: ["role"] },
},
},
});
await store.edges.worksAt.getOrCreateByEndpoints(alice, acme, {
role: "engineer",
});
```
On a schema-managed bundled root backend (including D1 and neon-http), with
claims, sidecars, history, and revision work absent, the default
`ifExists: "return"` path combines the
schema fence, endpoint validation, unique arbitration, and created/found result
into one statement. A typical Neon WebSocket miss therefore falls from roughly
five sequential requests to one. The found path is also one request, but the
PostgreSQL implementation performs a no-op conflict update: it takes a row lock
and can create write amplification, so it is not a substitute for a hot read
cache.
Dynamic call-level `matchOn`, constrained single-edge writes, history/revision
stores, caller-owned transactions, and custom backends without the matching
semantic program retain the transactional path required by their additional
contracts. In particular, writes made through `store.transaction()` remain on
the interactive path and do not receive the root-path exchange-count reduction.
Eligible durable bulk endpoint
convergence now submits one closed native atomic exchange: the durable identity
arbiter, endpoint validation, and ordered created/found
results are all resolved by the program. This removes the outside probe, the
transaction open/commit, and the per-item write legs for the eligible shape.
The fallback bulk path still discovers exact directed endpoint pairs in
set-oriented bind-budget chunks and retains its transactional contract. See
[`getOrCreateByEndpoints`](/schemas-stores#getorcreatebyendpointsfrom-to-props-options)
for field restrictions, migration rules, and PostgreSQL retry guidance.
The one-request path is an authoritative command, not a general Store batch:
the backend statement owns endpoint validation, durable-key arbitration, and
the created/found result. Static adapter batches (including multi-row inserts)
are separate internal optimizations and do not turn a sequence of public Store
calls into one atomic operation. Use `store.transaction(...)` when several
operations—including claims, Operational Identity, history, or revision
sidecars—must commit together. Undeclared dynamic `matchOn` convergence keeps
that interactive-transaction requirement; only a schema-declared durable
`matchIdentity` can qualify for the one-statement root command.
Bulk endpoint convergence has a narrower native envelope than direct edge
inserts. A schema-declared durable `matchIdentity` with `cardinality: "many"`,
the declaration's match fields, default `ifExists: "return"`, and no temporal
mutation qualifies on an exact bundled root. The libSQL transport inventory
records one client `batch` submission and zero client `execute` calls for a
multi-item eligible call; this is a submission-count measurement, not a
wall-clock benchmark. Dynamic `matchOn`, `ifExists: "update"`, constrained
cardinality, temporal options, caller transactions, derived backends, custom
backends without a registered durable-convergence family, and history/revision
stores intentionally retain the fallback path.
Outside the native envelope, an all-live `ifExists: "return"` batch is the
read-only exception: every backend may return that result from its single
set-oriented root read without opening a confirmation transaction. Inside the
native envelope, the authoritative upsert program runs first. An all-live call
still returns `"found"` in one exchange, but the conflict-update mechanism may
take incumbent-row locks and produce write amplification. Any batch outside
that envelope that may create, resurrect, or update requires the complete
transactional fallback.
If an otherwise eligible batch resolves a tombstoned identity, the native
attempt rolls back and transactionless convergence refuses with the typed
`CONSTRAINT_WRITE_FENCE_UNSUPPORTED` (`edgeMatchKeyConvergence`) error. Use a
transaction-capable backend when resurrection must merge partial properties
through the graph's Zod update schema.
Cloudflare D1's 100-parameter budget admits at most seven **unique durable
identities** in the native convergence program; duplicate inputs reuse their
first identity and do not consume another program entry. Above that ceiling,
an all-live return completes through the read-only one-read path described
above. A
batch that needs a write uses the portable fallback on a transaction-capable
backend and refuses on a transactionless D1 root rather than splitting one
atomic convergence contract across multiple submissions.
Bundled PostgreSQL roots using a recognized session-capable driver, Neon HTTP,
Cloudflare D1, and libSQL also expose native write programs for eligible
ingestion calls. Schema-managed `nodes.bulkInsert(items)` and
`nodes.bulkCreate(items)` run one schema-fenced atomic program when the node has
no Operational Identity, history, or revision work. The program composes each
member's complete advertised uniqueness/disjointness claim set with its
fulltext/vector projection transitions. Same-kind and hierarchy-wide
uniqueness, generated and caller IDs, and mixed claim families share this one
program boundary.
Neon HTTP, D1, and libSQL submit that program as one transport batch. Session-capable
PostgreSQL runs its statements on one pinned Drizzle transaction, with one SQL
statement per bind-budget chunk. The batch may use generated IDs,
caller-supplied IDs, or a mixture of both; `bulkCreate()` restores its rows to
input order. Claim work is chunked by member inside the same atomic submission,
so Cloudflare D1 no longer has a batch-wide claimed-member ceiling. Its
100-parameter budget leaves 87 claim-input binds per member after the row and
fence: a canonical claim costs six, each legacy hierarchy-wide uniqueness
probe costs nine, and each legacy disjointness probe costs six. Custom
executors should call the exported `atomicNodeClaimInputCost()` owner rather
than reproduce this formula. A member beyond that complexity, identity-enabled and
history/revision-tracked shapes, and other unsupported node work retain the
existing transaction or fallback path. This per-member budget is distinct from
the seven unique durable edge identities admitted by the convergence program
above. A successful program needs no diagnostic reads. A refused claim first
rolls the entire native batch back, then uses committed-state fence and claim
reads to recover the same typed error as the portable path. Bundled backends
diagnose the complete refused input with set-oriented node and uniqueness reads
rather than one probe per member. Custom backends without those batch reads
advance through 32-member concurrency windows until the earliest refusal can be
selected in input order; when no claim explains the rollback, every input is
covered before the honest terminal error. These failure-only reads are not part
of the successful-write RTT count.
Both `edges.bulkInsert(items)` and `edges.bulkCreate(items)` use the same
schema-fenced atomic program when the store has no history or revision capture.
The program validates live endpoints, arbitrates declared durable
`matchIdentity`, and maintains `one`, `unique`, and `oneActive` cardinality
claims at the write boundary. It rolls back the whole call when any
bind-budget chunk or constraint sidecar fails and restores `bulkCreate()`
results to input order.
Eligible `bulkDelete()` calls use the same mutation-program boundary. Direct
edge batches on bundled roots submit one schema-fenced atomic program; the
statement refuses an ID owned by another edge collection and rolls back every
chunk. Restricted node batches also submit one exchange and release uniqueness
and disjointness claims owned by the tombstoned rows in that program. The node
statement rechecks connected live edges at the write boundary, so a restricted
delete cannot race an earlier application-side probe. Cascade, disconnect,
projection, identity, captured, derived-backend, unregistered custom-backend, and
caller-transaction shapes keep the interactive path.
The same exact-root programs serve eligible singleton `update()` and
`delete()` calls without changing their per-operation hook contract. On an
interactive PostgreSQL root, the guarded mutation owns a short transaction;
on the single-submission transports it remains one batch. A node
update with no unique/identity sidecars (disjointness has no update-side
transition) may carry its fulltext/vector replacements in the same program. A
`cardinality: "many"` edge update with no durable match identity, performs one
authoritative preimage read, validates and merges properties in TypeGraph, then
submits one guarded atomic update. Every direct edge delete, and a restricted
node delete with supported claim cleanup, performs its existing live-row gate
and then submits one guarded atomic delete. Missing or tombstoned deletes remain
hook-free read-only no-ops. Temporal mutations, node cascade/disconnect,
history/revision capture, ordinary derived backends, and unregistered custom
families retain the complete portable transaction path.
The guarded update converges optimistically rather than holding a transaction
lock across its read and write. A one-row update gets four attempts; sustained
same-row contention can still end in `DatabaseOperationError`, while larger
resolved batches retain their two-attempt budget to bound retry cost.
These operations are registered through an exact-resource mutation execution
profile. Create and delete are **closed programs**: validation and arbitration
can be expressed by the submitted SQL itself. Node `bulkReplaceById()` is also
a closed program: every item supplies a complete document, so eligible bundled
roots submit missing-row creation, live replacement, tombstone resurrection,
claim release/acquisition, and fulltext/vector transitions in one read-free
atomic exchange. Live rows preserve their validity windows; resurrected rows
receive a fresh stamped window. Operational Identity and history/revision
capture retain the portable path.
`bulkUpsertById()` is deliberately different. It first reads authoritative
stored properties, then merges and validates them before its write set is
known. On bundled serverless roots, a
distinct-ID batch with no claims, Operational Identity, durable
edge match identity, temporal mutation, history, or revision capture submits
its resolved mutation set, including node fulltext/vector replacements, as one
atomic exchange. An eligible set on bundled
session-capable PostgreSQL stays inside its exact open transaction and
dispatches the same reviewed program through a separate registration bound to
that pinned transaction. Its `applied | unsupported` result is explicit:
`unsupported` is returned only before any program SQL runs, after which the
collection enters the complete portable path. Update-only sets use one
guarded set update. A set containing both fresh creates and live updates carries
both legs plus a zero-write terminal postimage assertion in the same native
batch; an incomplete preimage deliberately aborts the batch before any create
can commit. Repeated IDs, resurrections, temporal changes, claims, edge
sidecars, and
unregistered transaction sessions retain the consolidated interactive path.
The exact session may be collection-opened, supplied by `store.transaction()`,
or adopted from the caller. A
session program uses a savepoint so a deliberate database refusal can be rolled
back and diagnosed without poisoning the caller's surrounding PostgreSQL
transaction.
Measured at the libSQL transport boundary, eligible plain node
`bulkInsert()` and `bulkCreate()` calls each submit one exchange for generated,
caller-supplied, and mixed ID batches. A one-chunk unconstrained edge
batch remains 1 exchange, a durable-match batch drops from 6 transport
submissions to 1 atomic exchange, and a cardinality-constrained batch drops
from 8 transport submissions to 1 atomic exchange. Eligible one-chunk node and
edge `bulkDelete()` calls likewise submit 1 atomic exchange instead of a
transaction plus per-row probes/writes. On the portable edge path, one batched
authoritative read plus one set-based soft delete replaces the former two
statements per input. Eligible `nodes.bulkReplaceById()` calls submit 1 atomic
exchange with no preimage read, including claim and projection sidecars.
Eligible update-only and mixed create/update node and edge
`bulkUpsertById()` calls whose preimages fit one bind-budget read submit 2
exchanges: one batched preimage read and one atomic mutation exchange. Mixed
sets previously required separate create and update submissions after the read,
so they fall from 3 exchanges to 2 (33%). D1's 100-parameter budget admits 17
node mutations or 6 edge mutations per mutation statement. One D1 native
submission accepts at most 512 node members or 187 edge members; larger sets
fail closed to the portable path rather than constructing an unbounded
transport request. Within that ceiling, the mutation program chunks statements
inside one atomic batch, and a terminal postimage assertion for every chunk
rolls the complete submission back if any guarded member moved. When the
preimage read also exceeds its bind budget, it costs one read exchange per read
chunk plus the single atomic mutation submission. Other backends derive their
per-statement chunk size from their declared parameter budget and retain an
absolute 512-member submission ceiling. The native exchange still contains the
SQL statements needed for inserts and node projection sidecars; it groups
fulltext and per-vector-slot transitions into set statements and submits them
as one transaction so they do not each pay network latency. The exact previous
count varies by driver and endpoint shape. The program also proves the exact
durable contribution-marker identity and strategy signature in that
submission. A newly constructed backend therefore pays no separate cold marker
read before an eligible projected write. Missing, stale, failed, or
unmaterialized evidence aborts the whole submission; the failure path then
reads committed marker state to recover the existing typed contribution
diagnostic. The marker proof is an additional SQL statement inside that atomic
submission, so the optimization removes a network exchange rather than all
server-side proof work. Schemas with more marker identities may require more
than one proof statement within the same submission.
On session-capable PostgreSQL, the preimage read and mutation program remain in
one collection-owned transaction. The program has a bounded number of
statements independent of row count within its bind ceiling; it is not described
as a single network exchange because wire-protocol drivers execute those
statements on the pinned session. The gain is removal of the portable
per-family/per-member write-plan fan-out while retaining whole-call rollback.
These are internal execution optimizations, not a public Store batch API.
History/revision capture, ordinary derived or custom backends, dynamic
get-or-create convergence, and other unsupported shapes retain their
transaction or fallback behavior. Eligible mixed sets inside a bundled
PostgreSQL transaction are the narrow session-bound exception.
### Single vs bulk operations
For small numbers of writes, individual `create()` calls inside a transaction are fine. For larger
volumes, use the bulk collection APIs — they use multi-row INSERTs and handle parameter chunking
internally.
| Method | Returns results | Use case |
| ----------------------------------------- | --------------- | ---------------------------------------------------- |
| `bulkCreate(items)` | Yes | Need created nodes back |
| `bulkInsert(items)` | No | Maximum throughput ingestion |
| `bulkUpsertById(items)` | Yes | Idempotent import (create or update by ID) |
| `bulkReplaceById(items)` | Yes | Idempotent complete-document replacement by ID |
| `bulkDelete(ids)` | No | Mass soft-delete |
| `trustedImportGraphStream(store, chunks)` | No | Fastest initial load into a fresh dedicated database |
The collection APIs remain the default: they validate data and maintain every
configured constraint and sidecar. For a one-time initial load whose producer
already guarantees those invariants, the distinct
[`trustedImportGraphStream`](/interchange#trusted-initial-import) surface uses a
single transaction, engine-native inserts, and deferred secondary-index builds.
It intentionally rejects non-empty databases and graph features it cannot yet
maintain.
### PostgreSQL parameter limits
PostgreSQL's protocol can encode 65,535 bind parameters, while TypeGraph uses a portable
65,533-parameter budget across its bundled drivers. Bulk operations are automatically chunked to
stay within that budget:
- Node inserts: ~7,200 per chunk (9 params per node)
- Edge inserts: ~4,680 per chunk (budgeted at 14 params per durable edge)
You don't need to chunk manually — pass arrays of any size and TypeGraph handles the rest.
### Transaction wrapping
On a transaction-capable backend, each bulk method call is atomic across all of its bind-budget
chunks. Eligible plain `nodes.bulkInsert()` and `nodes.bulkCreate()` calls,
eligible plain node `bulkDelete()` calls, and direct edge
`bulkInsert()` / `bulkCreate()` / `bulkDelete()` calls also provide
whole-call atomicity on bundled transactionless roots through one native
atomic exchange, including durable-match and cardinality-constrained edge
batches. Other
bulk shapes on a transactionless root either refuse when their contract requires a fence or use
their documented non-atomic path. A certified atomic SQL program is available only to operations
whose closed statement contract has been proven by the backend conformance runner; it does not
make arbitrary Store calls atomic.
`store.transaction()` refuses before
invoking its callback on a transactionless root; it never presents sequential writes as atomic.
To commit several bulk calls as one unit on a transaction-capable backend, wrap them in a
transaction:
```typescript
// Atomic: all-or-nothing for the entire import
await store.transaction(async (tx) => {
await tx.nodes.Person.bulkCreate(people);
await tx.nodes.Company.bulkCreate(companies);
await tx.edges.worksAt.bulkCreate(employments);
});
```
Without the wrapping transaction, a failure in a later bulk call leaves earlier calls committed.
### Choosing the right pattern
```typescript
// Small batch (< 100 items): individual creates in a transaction are fine
await store.transaction(async (tx) => {
for (const person of people) {
await tx.nodes.Person.create(person);
}
});
// Medium batch (100–10,000 items): bulkCreate
const created = await store.nodes.Person.bulkCreate(people);
// Large batch (10,000+ items): bulkInsert (no result allocation)
await store.nodes.Person.bulkInsert(people);
// Idempotent import: bulkUpsertById (creates or updates by ID)
await store.nodes.Person.bulkUpsertById(itemsWithIds);
// Fresh dedicated database + already-validated producer:
await trustedImportGraphStream(store, interchangeChunks);
```
### Batch sizing for large multi-call imports
For a dataset too large for a single `bulkInsert`/`bulkCreate` call (e.g., streaming rows from a
file in a loop), the *size* of each call matters, not just the total row count. Each call is its
own transaction, and — per the default
[`autoRefreshStatistics`](/backend-setup#refreshing-planner-statistics-after-bulk-loads) — can
trigger a planner-statistics refresh on its own. In a large-scale bulk-load benchmark, batches of
~2,000 rows per call were consistently ~25-30% slower per row than batches of ~20,000+: fewer,
larger calls amortize both the per-call transaction commit and the statistics refresh across more
rows. Prefer batch sizes in the tens of thousands when looping over many calls for a large import,
and consider `autoRefreshStatistics: false` plus one `store.refreshStatistics()` call after the
loop if per-call refreshes still dominate.
### Batch reads
`getByIds()` on node and edge collections uses `SELECT ... WHERE id IN (...)` — one statement per
bind-limit chunk, so a single statement for id counts under the limit — instead of N individual
queries. Results are returned in input order with `undefined` for missing entries.
```typescript
const [alice, bob] = await store.nodes.Person.getByIds([aliceId, bobId]);
```
For multiple independent embeddable reads with different shapes and filters, use
[`store.batchOnce()`](/schemas-stores#batch-query-execution) to execute exactly one statement:
```typescript
const [activeUsers, recentOrders] = await store.batchOnce(() => [
store
.query()
.from("User", "u")
.whereNode("u", (u) => u.status.eq("active"))
.select((ctx) => ({ id: ctx.u.id, name: ctx.u.name })),
store
.query()
.from("Order", "o")
.select((ctx) => ({ id: ctx.o.id, total: ctx.o.total }))
.orderBy("o", "createdAt", "desc")
.limit(20),
]);
```
Runtime arrays of compatible `read.subgraph()` calls can opt into shared traversal and hydration
with `store.batchOnce(build, { shareSubgraphs: true })`. This helps overlapping, payload-heavy
neighborhoods; it adds overhead for membership and per-request reconstruction, so the independent
one-statement plan remains the default. Measure the real root overlap and projection rather than
enabling sharing universally.
Compare the [concrete sharing examples](#choosing-shared-subgraphs) before enabling the option.
The callback's batch-scoped builder composes set-oriented reads in the same call:
```typescript
const [latest, versionCount, detail] = await store.batchOnce((read) => [
read.neighbors(document, {
edges: ["hasVersion"],
orderBy: { by: "node", field: "sequence", direction: "desc" },
limit: 1,
}),
read.countNeighbors(document, { edges: ["hasVersion"] }),
read.subgraph(document.id, { edges: ["hasSection"], maxDepth: 2 }),
]);
```
The returned collection may be a runtime-sized array, including `roots.map(...)`. Empty arrays
execute zero statements; every nonempty array, including a singleton, executes exactly one or is
refused before execution. The portable ceiling is 500 reads and the final statement must fit the
backend's bind-parameter budget. Since all member rows are returned in materialized JSON envelopes,
use bounded projections and limits; `batchOnce()` does not stream, predict payload size, or impose a
response-byte cap.
When a request needs several independent neighborhoods, prefer one runtime batch over awaiting
`store.subgraph()` in a loop, especially when the database is remote:
```typescript
const neighborhoods = await store.batchOnce((read) =>
roots.map((root) =>
read.subgraph(root.id, {
edges: ["knows", "worksAt"],
maxDepth: 2,
project: {
nodes: { Person: ["name"], Company: ["name"] },
edges: { knows: [], worksAt: ["role"] },
},
}),
),
);
```
This changes several database round trips into one. Each subgraph still has its own recursive CTE and
hydration work: By default, `batchOnce()` does not merge roots, share traversal, or guarantee less database CPU.
For one large closure, the direct backend-tuned `store.subgraph()` path can be faster. Use
`batchOnce()` when round-trip latency across several independent, bounded subgraphs is the cost to
remove, and measure both forms when database work dominates.
When the response needs a list of matching child records per parent, use a relation grouped by the
parent key and [`expr.collect()` with named scalar fields](/queries/relations#ordered-collections).
That aggregates the selected fields into ordered records in one query. Apply
[`topPerPartition()`](/queries/relations#top-n-per-parent) before grouping when each parent needs
only its highest-priority or most recent N children. The database chooses winners before returning
results; it may still scan and sort all candidates. For several independent,
bounded subgraphs, keep using the `batchOnce(read => roots.map(root => read.subgraph(...)))` pattern
above; record collection does not replace subgraph hydration.
Use `store.batch()` when queued edge collection reads must participate. It runs them in sequence.
On a transactional backend it still issues at least one statement per query plus
`begin`/`commit`, so N queries are N+2 round trips at best; without transactions there is no
framing. It buys a connection profile that never peaks at N — not lower latency, and not a snapshot
(PostgreSQL's default read-committed isolation lets a later query see a newer commit).
Edge collection `batchFind*` methods (`batchFindFrom`, `batchFindTo`, `batchFindByEndpoints`) also
participate in `store.batch()`. On a transactional backend they move N `findFrom`/`findTo` calls
into one transaction — the statement count is unchanged either way. If the round trips are what
hurt, replace the calls with `store.neighbors()` / `store.countNeighbors()` or a traversal (one
statement), or compose the batch-scoped `read.neighbors()`, `read.countNeighbors()`, and
`read.subgraph()` forms in `batchOnce()`.
The same read family is available on `TransactionContext`. Use `tx.neighbors()`,
`tx.countNeighbors()`, or `tx.subgraph()` for a direct read that must see earlier writes in the
callback. Use `tx.batchOnce()` to combine independent transaction-bound reads into exactly one
statement on the held connection; fluent items in that batch start from `tx.query()`. Direct
`tx.subgraph()` also uses its one-statement plan so it never submits concurrent statements to the
held transaction connection.
Direct `store.subgraph()` and batch-scoped `read.subgraph()` share one semantic planner and produce the
same result, but intentionally use different physical plans. The direct form uses 2 statements
on SQLite and 3 on PostgreSQL so each backend can hydrate a closure efficiently. The scoped form
uses 1 statement everywhere to make cross-shape composition possible. On PostgreSQL, prefer the
direct form for a standalone large closure; use the batch-scoped form when eliminating network round
trips across several independent reads matters more than optimizing that closure in isolation.
To read the edges of a *set* of endpoints, prefer `bulkFindFrom` / `bulkFindTo` (see
[Edge Collections](/schemas-stores#edge-collections)).
Where `store.batch()` runs N singleton reads over one connection, these widen the endpoint predicate
itself to `from_id IN (...)` — one set-oriented statement per endpoint kind and bind-budget chunk,
on the same index prefix seek the singleton read uses — and return the edges grouped per input:
```typescript
const people = await store.nodes.Person.find({ limit: 50 });
const jobsPerPerson = await store.edges.worksAt.bulkFindFrom(people);
// jobsPerPerson[i] holds the worksAt edges of people[i]
```
This is the fix for the "list view with relationship counts" N+1: statement count grows with
endpoint kinds and bind-budget chunks instead of with every item on the page. Pass `limitPerInput`
to bound each endpoint's fan-out.
If a view spans several source kinds and edge kinds, use the Store-level
`bulkFindEdgesFrom` operation instead of calling each licensed edge collection separately. It
accepts heterogeneous source groups and edge kinds, then executes one set-oriented statement per
bind-budget chunk. Round trips therefore grow with input size, not with the number of licensed
`(source kind, edge kind)` combinations:
```typescript
const edgesBySource = await store.bulkFindEdgesFrom({
sources: [
{ kind: "Company", ids: companyIds },
{ kind: "Person", ids: personIds },
],
edgeKinds: ["employs", "owns", "dependsOn"],
});
// edgesBySource[i] identifies its source and contains that source's matching edges
```
:::note[Operation hooks]
Bulk operations (`bulkCreate`, `bulkInsert`, `bulkUpsertById`, `bulkDelete`) skip per-item operation hooks for
throughput, and the bulk hooks (`onBulkOperationStart` / `onBulkOperationEnd`) do not stand in for them — those fire
only for node `updateWhere`, so a bulk method emits no hook events at all, neither per-item nor bulk. To observe every
individual write, call the single-item method instead. Query hooks still fire normally. See
[Schemas & Stores](/schemas-stores#observability-hooks) for details.
:::
#### Choosing shared subgraphs
Consider a `Person` graph with `name` and a long `biography` property, connected by
outgoing `knows` edges. The following examples use the same Store and bounded depth.
Sharing preserves a separate subgraph for each input root; it changes the database
plan and response encoding, not the returned neighborhoods.
**Good candidate: overlapping neighborhoods with substantial properties.** Ada and
Bea both know Cara, and Cara knows Dev:
```text
Ada ──knows──▶ Cara ──knows──▶ Dev
Bea ──knows──▶ Cara
```
Both depth-two subgraphs contain Cara and Dev. If their biographies are several
kilobytes each, independent plans repeat that property data. Enable sharing so the
compatible reads hydrate those shared entities once:
```typescript
const overlapping = await store.batchOnce(
(read) =>
[ada.id, bea.id].map((rootId) =>
read.subgraph(rootId, {
edges: ["knows"],
maxDepth: 2,
project: { nodes: { Person: ["name", "biography"] } },
}),
),
{ shareSubgraphs: true },
);
// overlapping[0]: Ada, Cara, Dev
// overlapping[1]: Bea, Cara, Dev
// Each result owns independent projected values, including nested objects.
```
**Keep the default: disjoint neighborhoods.** Suppose Erin knows Finn, Finn knows
Gia, Hana knows Ivan, and Ivan knows Jules, with no connections between the groups:
```text
Erin ──knows──▶ Finn ──knows──▶ Gia
Hana ──knows──▶ Ivan ──knows──▶ Jules
```
No entities are reused across roots. Sharing adds membership information without
removing duplicate biographies. Default `batchOnce()` still combines the two reads
into one statement:
```typescript
const disjoint = await store.batchOnce((read) =>
[erin.id, hana.id].map((rootId) =>
read.subgraph(rootId, {
edges: ["knows"],
maxDepth: 2,
project: { nodes: { Person: ["name", "biography"] } },
}),
),
);
```
**Keep the default initially: overlapping roots, identity-only results.** A graph
preview might need only connectivity, even when the stored biographies are large.
Use empty property selections to retain identities and edge endpoints without
transferring biographies:
```typescript
const connectivity = await store.batchOnce((read) =>
[ada.id, bea.id].map((rootId) =>
read.subgraph(rootId, {
edges: ["knows"],
maxDepth: 2,
project: { nodes: { Person: [] }, edges: { knows: [] } },
}),
),
);
```
Overlap alone does not make sharing a payload optimization here: little repeated
property data remains to remove. Sharing may still improve latency, so measure
both response size and elapsed time before choosing it.
| Eight-root benchmark shape | Shared vs default batch encoded bytes | Starting choice |
| --- | --- | --- |
| 75% overlap, full 2,048-byte payload | About 27–29% fewer | Try sharing |
| Disjoint roots, full 256-byte payload | About 21–25% more | Default batching |
| 75% overlap, identities only | About 10–18% more | Default batching; measure latency |
These response-size differences appeared in small local SQLite and PostgreSQL runs
and a same-region remote Neon PostgreSQL run; they are not universal thresholds.
Bytes measure JSON encoding at the backend boundary, not protocol traffic. In the
remote run, shared batches were faster at the median in all three shapes, including
the disjoint and identity-only shapes whose encoded responses grew. Default
`batchOnce()` was faster at the median than concurrent direct subgraph calls in all
three shapes. The remote client egress was observed in Bend, Oregon, and the pooled
database endpoint was in Oregon; a simple pooled `SELECT 1` round trip measured
24 ms at the median. Tail latency varied, so compare elapsed time and response
size on the deployment route that matters to your application.
See the [SQLite report](https://github.com/nicia-ai/typegraph/blob/6196354c/packages/benchmarks/reports/subgraph-batch-sqlite-2026-09-14.md)
and [PostgreSQL report](https://github.com/nicia-ai/typegraph/blob/6196354c/packages/benchmarks/reports/subgraph-batch-postgres-2026-09-14.md)
for local timings and methodology, and the [remote Neon report](https://github.com/nicia-ai/typegraph/blob/cffd0082906bffd5f5e993dfeede1e01e6e6f300/packages/benchmarks/reports/subgraph-batch-neon-oregon-2026-09-15.md)
for 40 raw samples per shape and mode, placement, and reproduction commands. The
older PostgreSQL report also includes a separately labeled delay simulation; the
Neon measurements used actual network calls without injected delay.
Sharing also requires compatible options. Different edge sets, depths, temporal
coordinates, projections, or edge windows can keep reads in separate groups even
when their results overlap. For example, a social `knows` neighborhood and an
employment `worksAt` neighborhood still fit in one `batchOnce()` call, but enabling
sharing does not fuse those incompatible plans.
## Connection Management
Managed local Store and backend factories own and close their SQLite or PGlite
resources. Bring-your-own adapter integrations leave the supplied connection or
pool under application control. See [Backend Setup](/backend-setup#connection-management)
for the ownership matrix and shutdown examples.
### PostgreSQL pooling
Always use a connection pool in production. An individual query holds a connection only while each
statement runs. Most queries issue a single statement; a query whose selective-field mapping falls
back issues a second. `store.transaction()` holds one connection for the whole callback, and
`store.batch()` does the same for its implicit transaction.
```typescript
import { Pool } from "pg";
const pool = new Pool({
connectionString: process.env.DATABASE_URL,
max: 20, // Size based on your concurrency needs
idleTimeoutMillis: 30_000,
connectionTimeoutMillis: 2_000,
});
pool.on("error", (err) => {
console.error("Unexpected pool error", err);
});
```
**Sizing guidance:** Each concurrent query holds one connection for as long as its statement runs.
A pool of 10–20 connections handles most workloads. If you're running bulk imports in parallel,
size up accordingly.
**Reducing pool pressure with `batch()`:** When loading multiple independent queries (e.g., a
detail page with several relationship types), `Promise.all` can acquire up to N connections
simultaneously — fewer if the pool is undersized or saturated, in which case it queues instead.
[`store.batch()`](/schemas-stores#batch-query-execution) keeps at most one query in flight, so peak
connection use is 1 — on a transactional backend that is literally one checked-out connection for
the implicit transaction; elsewhere it is one at a time, and whether the adapter reuses the same
client is its own business. It does not reduce the statement count, and read-committed isolation
means it is not a snapshot.
### SQLite concurrency
SQLite is single-writer. For best throughput:
- Use WAL with `synchronous=NORMAL`. `createLocalSqliteBackend` applies both
(plus a 5s `busy_timeout`) automatically; on a bring-your-own connection set
them yourself: `sqlite.pragma("journal_mode = WAL")`,
`sqlite.pragma("synchronous = NORMAL")`. On file databases this makes
single-operation writes roughly 5× faster than the driver defaults.
- Batch writes in transactions rather than issuing many small commits. (One
nuance: pure bulk appends of fresh pages can run marginally faster under the
rollback journal than WAL, since WAL writes pages twice — the per-commit wins
dominate everywhere else.)
- For read-heavy workloads, SQLite performs well without pooling since `better-sqlite3` is synchronous
### Transaction isolation
PostgreSQL transactions accept an optional isolation level:
```typescript
await store.transaction(
async (tx) => {
// Serializable isolation for strict consistency
const snapshot = await tx.nodes.Account.getById(accountId);
// ...
},
{ isolationLevel: "serializable" },
);
```
Available levels: `read_uncommitted`, `read_committed` (default), `repeatable_read`, `serializable`.
Schema-managed Stores fence writes against concurrent schema-version commits.
That includes Stores opened by `createStoreWithSchema`,
`createAdapterStoreWithSchema`, `createVerifiedStore`, or
`createVerifiedAdapterStore`; an adapter Store constructed with a cached
`{ reconciled }` snapshot; and Stores returned by `evolve()` or rebound from
one of those Stores. `store.introspect().schemaVersion !== undefined` is the
runtime test.
PostgreSQL reacquires and validates the active-schema row lock at every managed
write. The lock is normally reentrant and remains held to transaction end, but
the repeated check is required because rolling back to a caller-created
savepoint releases row locks acquired after that savepoint. At
`repeatable_read` or `serializable`, a concurrent schema commit can raise
PostgreSQL's normal serialization failure; retry the whole transaction.
Graph-merge commits already retry those failures automatically. Raw
`createStore` / `createAdapterStore` instances without a reconciled snapshot,
and writes issued directly through a backend, do not carry schema metadata and
remain outside this guarantee. `store.clear()` also resets the cleared Store to
that raw state.
SQLite always operates at `serializable` isolation.
## Query Optimization Features
### Precomputed Closures
When you define an ontology (e.g., `subClassOf`, `implies`), TypeGraph precomputes the full
transitive closure at store initialization. Queries like
`.from("Parent", "p", { includeSubClasses: true })` use a pre-calculated list of kinds rather than
recursive lookups at runtime.
### Smart Select
For explicit SQL field selection, prefer [`project()`](/queries/expressions#projection-and-mapping).
Use `map()` afterward for JavaScript transformations. This avoids legacy selector probing
and makes the database projection explicit.
TypeGraph automatically optimizes queries based on which fields your `select()` callback accesses.
When you select specific fields, TypeGraph generates SQL that only extracts those fields using
`json_extract()` (SQLite) or JSONB path extraction (PostgreSQL), rather than fetching the entire
`props` blob.
```typescript
// Optimized: Only fetches email and name from the database
const results = await store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.email.eq("alice@example.com"))
.select((ctx) => ({
email: ctx.p.email,
name: ctx.p.name,
}))
.execute();
// SQL: SELECT json_extract(props, '$.email'), json_extract(props, '$.name') ...
```
This optimization pairs well with [covering indexes](/performance/indexes#covering-indexes): if
your index contains both the filter keys and the selected keys, the database can serve the query
straight from the index instead of scanning the whole table — though on PostgreSQL specifically,
this stops short of a true `Index Only Scan` for JSONB-extracted fields; see the
[covering indexes](/performance/indexes#covering-indexes) section for the concrete limitation and
a workaround.
**When optimization applies:**
| Pattern | Optimized? | Reason |
| ----------------------------------------------- | ---------- | ---------------------------------- |
| `ctx => ({ email: ctx.p.email })` | Yes | Simple field extraction |
| `ctx => [ctx.p.id, ctx.p.name]` | Yes | Multiple fields in array |
| `ctx => ctx.p` | No | Whole node returned |
| `ctx => ({ upper: ctx.p.email.toUpperCase() })` | Yes | Field extracted; method runs in JS |
| `ctx => ({ ...ctx.p })` | No | Spread requires full node |
The optimization is transparent — if your callback can't be optimized, TypeGraph automatically
falls back to fetching the full node data.
For data-dependent callbacks, TypeGraph first plans with representative values, including a
high-value pass that covers common numeric threshold branches. If an unobserved branch accesses an
additional field at execution time, the first miss may require a second statement that fetches the
full row. Prepared queries remember that missing-field failure and use the full-row plan directly on
later executions. Comparisons against arbitrary string values can still take an unobserved branch;
the high-value pass does not guarantee that every possible callback path is planned in advance.
:::note[Select callback purity]
Smart select applies to `.execute()`, `.paginate()`, and `.stream()`. The `select()` callback may be evaluated
multiple times during planning/optimization, so it should be pure (no side effects).
:::
:::note[Known limitations]
Smart select is not currently applied to queries that include variable-length traversals (recursive CTEs),
even when the select callback is otherwise optimizable.
:::
### Built-in Indexes
The default TypeGraph schema includes optimized indexes for the most common access patterns:
- **Graph + Kind + ID**: Primary key for node lookups
- **Graph + From/To ID**: Optimized for edge traversals
- **Temporal columns**: Indexes on `valid_from`, `valid_to`, and `deleted_at`
For application-specific indexes on JSON properties, see [Indexes](/performance/indexes).
### SQL Compilation
Each builder method (`.where()`, `.limit()`, `.orderBy()`, etc.) returns a new immutable instance.
A reused query instance compiles **once**. The first `.execute()` builds a cached template and every
later call reuses it — for standard queries, aggregate queries, set-operation queries (`union`,
`intersect`, `except`), and prepared queries alike. Explicit `.toSQL()` / `.compile()` calls are the
exception: they compile on demand every time, because producing the statement is the thing the
caller asked for.
The subtlety a cache like that has to survive is freshness. A "current" (live) read filters on
temporal validity as of the instant it runs, so a template with a concrete "now" baked into it would
freeze that instant for the query instance's whole lifetime, hiding every row created afterward. The
template therefore reserves the read instant as a **placeholder** rather than a value, and each
execution fills it with a fresh instant alongside that call's bindings. Nothing in the statement's
text depends on either, so reuse costs no freshness.
```typescript
const activeUsers = store
.query()
.from("User", "u")
.whereNode("u", (u) => u.status.eq("active"))
.select((ctx) => ctx.u);
// One compilation, two executions. The read instant is bound per call, so a
// user created between these two is visible to the second one.
await activeUsers.execute();
await activeUsers.execute();
```
Two things fall back to compiling on every call:
- **Backends that cannot execute pre-compiled SQL text** — a custom or async backend, i.e. one
without `executeRaw`.
- **Statements whose execution semantics ride on the compiled SQL object rather than its text**,
even on PostgreSQL with `executeRaw` fully available. Two query shapes do: **approximate vector
search** (`similarTo(..., { approximate: true })`, which carries the pgvector / `sqlite-vec`
iterative-scan wrapper) and **`store.subgraph()` on PostgreSQL**, whose id-array fetches are
marked to force a custom plan so the planner sizes them against the actual array rather than
reusing a generic one. Flattening either to cacheable text would silently drop the behavior it
depends on, so they are excluded deliberately — the trade is a template hit against correct
execution, and correctness wins.
Compilation is pure, in-memory string-building with no I/O, so both fallbacks are cheap; the query's
database round-trip dominates either way. Worth knowing if you are profiling a vector query and
expecting the compile-once behavior described above — that is the one shape where it does not apply.
### Prepared Queries
For hot paths that execute the same query shape with different values, `.prepare()` builds and
structurally validates the query AST once — a malformed query fails fast, before the first
`.execute()`, instead of on first use — and compiles the statement once into a cached template. Each
`.execute(bindings)` fills that template's placeholders (a fresh read instant plus the call's own
parameter values) and runs the cached text directly through `executeRaw`.
Because arity never reaches the SQL text, a list-valued parameter reuses the same template no matter
how long the list is:
```typescript
const byIds = store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.id.in(param("ids")))
.select((ctx) => ctx.p)
.prepare();
await byIds.execute({ ids: ["a", "b", "c"] });
await byIds.execute({ ids: ["d"] }); // same compiled statement
```
Best for: validating a query shape once, then reusing it with different parameter values. The saved
compilation is real but small — the database round-trip still dominates.
See [Prepared Queries](/queries/execute#prepared-queries) for usage details.
### Subgraph extraction
For the "load entity with all relationships" pattern, [`store.subgraph()`](/schemas-stores#subgraph-extraction)
is a backend-tuned option for a single bounded neighborhood. It compiles to a recursive CTE
that fans out across all specified edge
types in a fixed 2 statements on SQLite and 3 on PostgreSQL — no matter how many relationship kinds
are involved, or how much it returns. See
[Choosing a query strategy](/schemas-stores#choosing-a-query-strategy) for guidance on when to use
`subgraph()` vs the fluent query builder vs manual `findFrom` calls.
The [`project` option](/schemas-stores#subgraph-projection) further reduces overhead by extracting
only the specified fields per kind at the SQL level via `json_extract()` / JSONB paths, skipping
full `props` blob transfer and metadata columns for projected kinds.
## Best Practices
### Filter early
Use `.whereNode()` and `.whereEdge()` for match constraints that should restrict expansion.
The compiler applies them at the matching stage regardless of their position in the chain.
Use scoped `.where()` when the condition must filter completed rows; moving a completed-row
condition into an optional match or recursive hop can change its meaning.
### Select specific fields
When you only need certain fields, use `project()` to make the SQL projection explicit.
Legacy `select()` also supports [smart select optimization](#smart-select). Smaller projections
reduce transferred data and may benefit from covering indexes, subject to the engine limitations
described above.
```typescript
// Preferred: Only fetches what you need
.project((e) => ({ name: e.p.name, email: e.p.email }))
// Avoid when possible: Fetches entire props blob
.select((ctx) => ctx.p)
```
### Use specific kinds
Unless you specifically need to query across a hierarchy, avoid `includeSubClasses: true`. Being
specific about the node kind allows the SQL engine to use more restrictive index scans.
### Use cursor pagination
For large datasets, prefer `.paginate()` over `.limit()` and `.offset()`. Keyset pagination
(using cursors) avoids the `O(N)` cost of skipping rows in standard SQL offsets.
### Index your filter and sort properties
TypeGraph's built-in indexes cover structural lookups (by ID, by edge endpoints). Properties you
filter or sort on in `whereNode()`, `whereEdge()`, and `orderBy()` need application-specific
[expression indexes](/performance/indexes). Use the [Query Profiler](/performance/profiler) to
identify which properties need coverage.
## Profile Your Queries
Use the [Query Profiler](/performance/profiler) to identify missing indexes and understand
query patterns in your application. The profiler captures property access patterns and generates
prioritized index recommendations.
```typescript
import { QueryProfiler } from "@nicia-ai/typegraph/profiler";
const profiler = new QueryProfiler();
const profiledStore = profiler.attachToStore(store);
// Run your application or test suite...
const report = profiler.getReport();
console.log(report.recommendations);
```
## Benchmarks
TypeGraph uses a deterministic performance sanity suite as its benchmark and regression gate.
The suite seeds a realistic graph shape and measures end-to-end query latency across:
- forward and reverse traversals
- inverse/symmetric traversal (`expand: "inverse"` / `expand: "all"`)
- 2-hop and 3-hop traversals
- aggregate queries
- cached execute vs prepared execute
- deep traversals (`10`/`100`/`1000` hop recursive with `cyclePolicy: "allow"`)
Guardrail thresholds enforce expected behavior in CI (for example, traversal latency caps and
ratio checks such as reverse/forward and deep-hop scaling).
Deep-recursive benchmark probes explicitly set `cyclePolicy: "allow"` to isolate recursive CTE
expansion cost; the default `cyclePolicy: "prevent"` prioritizes cycle-safe semantics and is
expected to be slower on long traversals.
*Note: Real-world performance varies by hardware, database driver, network latency (for PostgreSQL),
and schema/data shape.*
Benchmark configuration and guardrails
Current suite configuration:
| Setting | Value |
| ----------------------------------- | ----- |
| Seed users | 1200 |
| Follows per user | 10 |
| Posts per user | 5 |
| Batch size | 250 |
| Warmup iterations | 2 |
| Sample iterations (median reported) | 15 |
Default guardrails:
| Check | Threshold |
| ------------------------------------------ | --------- |
| reverse/forward ratio | <= 6x |
| inverse traversal latency | <= 500ms |
| inverse/forward ratio | <= 10x |
| 3-hop latency | <= 500ms |
| 3-hop/2-hop ratio | <= 8x |
| aggregate latency | <= 500ms |
| aggregate distinct latency | <= 700ms |
| aggregateDistinct/aggregate ratio | <= 4x |
| cached execute latency | <= 500ms |
| prepared execute latency | <= 500ms |
| prepared/cached ratio | <= 2x |
| 10-hop recursive latency | <= 250ms |
| 100-hop recursive latency | <= 1000ms |
| 100-hop-recursive/10-hop-recursive ratio | <= 30x |
| 1000-hop recursive latency | <= 5000ms |
| 1000-hop-recursive/100-hop-recursive ratio | <= 20x |
Backend-specific overrides:
| Backend | Check | Threshold |
| ---------- | -------------------------- | --------- |
| SQLite | 1000-hop recursive latency | <= 7000ms |
| PostgreSQL | inverse traversal latency | <= 1000ms |
| PostgreSQL | inverse/forward ratio | <= 30x |
| PostgreSQL | 3-hop latency | <= 1000ms |
| PostgreSQL | aggregate distinct latency | <= 1200ms |
| PostgreSQL | prepared execute latency | <= 700ms |
### Real-world workload validation
Beyond the synthetic guardrail suite above, TypeGraph is also exercised against the
[LDBC Social Network Benchmark (SNB) Interactive](https://github.com/ldbc/ldbc_snb_interactive_v1)
workload — a standard, independently-defined graph benchmark, not a TypeGraph-specific one — at
SF1 scale (~10k persons, ~1M posts, ~2M comments). This surfaced and fixed two real scaling bugs in
the library: an unbounded `ANALYZE` cost on bulk SQLite loads, and an N+1 endpoint-existence check
in batched edge creation. It also directly produced the `keySystemColumns` guidance and the
PostgreSQL index-only-scan caveat in [Indexes](/performance/indexes#covering-indexes). The
benchmark source lives in `packages/benchmarks/src/real/` in the repository.
### Running benchmarks locally
```bash
pnpm bench
```
For guardrail mode (fails on regression thresholds):
```bash
pnpm --filter @nicia-ai/typegraph-benchmarks perf:check
```
Run the same guardrailed suite against PostgreSQL:
```bash
POSTGRES_URL=postgresql://typegraph:typegraph@127.0.0.1:5432/typegraph_test \
pnpm --filter @nicia-ai/typegraph-benchmarks perf:check:postgres
```
By default the SQLite suite runs against an in-memory database, which
measures engine and compile cost but not WAL/fsync behavior. Add
`--storage=file` (or use the `perf:file` / `perf:check:file` scripts) to run
against a temporary on-disk database — the lane that reflects real local
deployments.
A separate write-throughput bench measures single-op creates,
transaction-amortized creates, `bulkCreate`, search-indexed creates
(fulltext + vector sync), and `importGraph`, normalized to milliseconds per
operation:
```bash
pnpm --filter @nicia-ai/typegraph-benchmarks bench:write # sqlite, in-memory
pnpm --filter @nicia-ai/typegraph-benchmarks bench:write:file # sqlite, on-disk
POSTGRES_URL=... pnpm --filter @nicia-ai/typegraph-benchmarks bench:write:postgres
```
The write bench is report-only (no guardrails): write latency is dominated
by fsync behavior on the file lane and needs per-machine calibration.
The benchmark source code is located in `packages/benchmarks/src/`.
## Next Steps
- [Indexes](/performance/indexes) — Define custom indexes for your schema
- [Query Profiler](/performance/profiler) — Identify missing indexes automatically
- [Backend Setup](/backend-setup) — Connection setup, pooling, and lifecycle
# Query Profiler
> Capture query patterns and generate index recommendations
The Query Profiler captures property access patterns from your queries and generates
index recommendations. Use it during development or in test suites to identify missing indexes.
## Quick Start
```typescript
import { QueryProfiler } from "@nicia-ai/typegraph/profiler";
// Create a profiler and attach it to your store
const profiler = new QueryProfiler();
const profiledStore = profiler.attachToStore(store);
// Run queries as normal - they're automatically tracked
await profiledStore
.query()
.from("Person", "p")
.whereNode("p", (p) => p.email.eq("alice@example.com"))
.select((ctx) => ({ name: ctx.p.name }))
.execute();
// Get recommendations
const report = profiler.getReport();
for (const rec of report.recommendations) {
console.log(
`[${rec.priority}] ${rec.entityType}:${rec.kind} ${rec.fields.join(", ")}`,
);
console.log(` ${rec.reason}`);
}
```
## How It Works
The profiler uses JavaScript Proxy to transparently wrap your store and query builders. When
queries execute, it extracts property access patterns from the query AST:
- **Filter patterns**: Properties used in `.whereNode()` and `.whereEdge()` predicates
- **Sort patterns**: Properties used in `.orderBy()`
- **Select patterns**: Properties accessed in `.select()` callbacks
- **Group patterns**: Properties used in `.groupBy()`
The profiler then compares these patterns against your declared indexes and generates
recommendations for missing coverage.
## Kinds and `includeSubClasses`
When you query with `includeSubClasses: true`, a single alias can represent multiple kinds.
When the profiler is attached to a store, it uses the graph schema to attribute a property access
only to kinds where that JSON path exists. This avoids recommending indexes for unrelated subclasses.
## Attaching to a Store
```typescript
const profiler = new QueryProfiler();
const profiledStore = profiler.attachToStore(store);
// The profiled store behaves exactly like the original
await profiledStore.nodes.Person.create({ email: "bob@example.com", name: "Bob" });
// Queries are tracked automatically
await profiledStore.query().from("Person", "p").select((ctx) => ctx.p).execute();
// Access the profiler from the store
profiledStore.profiler.getReport();
```
The profiled store exposes a `profiler` property for convenient access.
## Declaring Existing Indexes
Pass your existing indexes so the profiler doesn't recommend indexes you already have:
```typescript
import { QueryProfiler } from "@nicia-ai/typegraph/profiler";
import { toDeclaredIndexes } from "@nicia-ai/typegraph/indexes";
import { personEmail, worksAtRole } from "./indexes";
const profiler = new QueryProfiler({
declaredIndexes: toDeclaredIndexes([personEmail, worksAtRole]),
});
```
You can also declare indexes manually:
```typescript
const profiler = new QueryProfiler({
declaredIndexes: [
{
entityType: "node",
kind: "Person",
fields: ["/email"],
unique: true,
name: "idx_person_email",
},
{
entityType: "node",
kind: "Person",
fields: ["/name"],
unique: false,
name: "idx_person_name",
},
],
});
```
## Understanding the Report
```typescript
const report = profiler.getReport();
```
The report contains:
### `recommendations`
Prioritized index recommendations sorted by importance:
```typescript
for (const rec of report.recommendations) {
console.log(
`[${rec.priority}] ${rec.entityType}:${rec.kind} ${rec.fields.join(", ")}`,
);
console.log(` Reason: ${rec.reason}`);
console.log(` Frequency: ${rec.frequency}`);
}
```
**Priority levels:**
- `high`: Property accessed 10+ times in filters/sorts (configurable)
- `medium`: Property accessed 5-9 times (configurable)
- `low`: Property accessed 3-4 times (configurable)
### `unindexedFilters`
Properties used in filter predicates that lack index coverage:
```typescript
for (const path of report.unindexedFilters) {
const target =
path.target.__type === "prop" ? path.target.pointer : path.target.field;
console.log(`Unindexed filter: ${path.entityType}:${path.kind} ${target}`);
}
```
### `patterns`
Raw property access statistics:
```typescript
for (const [key, stats] of report.patterns) {
console.log(`${key}: ${stats.count} accesses`);
console.log(` Contexts: ${[...stats.contexts].join(", ")}`);
console.log(` Predicates: ${[...stats.predicateTypes].join(", ")}`);
}
```
### `summary`
Session statistics:
```typescript
console.log(`Total queries: ${report.summary.totalQueries}`);
console.log(`Unique patterns: ${report.summary.uniquePatterns}`);
console.log(`Duration: ${report.summary.durationMs}ms`);
```
## Test Assertions
Use `assertIndexCoverage()` to fail tests when queries filter on unindexed properties:
```typescript
import { describe, it, beforeAll, afterAll } from "vitest";
import { QueryProfiler } from "@nicia-ai/typegraph/profiler";
describe("Query Performance", () => {
let profiler: QueryProfiler;
let profiledStore: ProfiledStore;
beforeAll(() => {
profiler = new QueryProfiler({
declaredIndexes: toDeclaredIndexes([personEmail, personName]),
});
profiledStore = profiler.attachToStore(store);
});
// Run your test suite against profiledStore...
it("all filtered properties should be indexed", () => {
// Throws if any filter property lacks an index
profiler.assertIndexCoverage();
});
});
```
## Configuration
```typescript
const profiler = new QueryProfiler({
// Indexes you already have
declaredIndexes: [...],
// Minimum frequency to generate a recommendation (default: 3)
minFrequencyForRecommendation: 5,
// Optional priority thresholds (defaults: 5 and 10)
mediumFrequencyThreshold: 8,
highFrequencyThreshold: 20,
});
```
## Lifecycle Methods
```typescript
// Reset collected data (keeps configuration)
profiler.reset();
// Detach from store (allows reattachment)
profiler.detach();
// Check attachment status
if (profiler.isAttached) {
console.log("Profiler is attached to a store");
}
```
## Manual Recording
For custom integrations, record queries directly from their AST:
```typescript
const query = store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.email.eq("test@example.com"))
.select((ctx) => ctx.p);
// Record without executing
profiler.recordQuery(query.toAst());
```
## Composite Index Detection
The profiler understands composite index prefix matching. If you have an index on `["email", "name"]`,
queries filtering on just `email` are considered covered:
```typescript
const profiler = new QueryProfiler({
declaredIndexes: [
{
entityType: "node",
kind: "Person",
fields: ["/email", "/name"],
unique: false,
name: "idx_email_name",
},
],
});
// This query IS covered (uses the email prefix of the composite index)
await profiledStore
.query()
.from("Person", "p")
.whereNode("p", (p) => p.email.eq("test@example.com"))
.execute();
// No recommendation generated for email
```
## Best Practices
1. **Profile realistic workloads**: Run your actual queries or test suite, not synthetic benchmarks.
2. **Profile before optimizing**: Don't guess which indexes you need - let the profiler tell you.
3. **Use in CI**: Add `assertIndexCoverage()` to your test suite to catch regressions.
4. **Declare all indexes**: Pass your existing indexes so recommendations are accurate.
5. **Review frequency**: High-frequency patterns are most important to index.
## Next Steps
- [Indexes](/performance/indexes) - Create the indexes the profiler recommends
- [Performance Overview](/performance/overview) - Best practices and smart select
# Project Structure
> Recommended patterns for organizing TypeGraph in your codebase
How you organize your TypeGraph code depends on your project's size and complexity.
This guide covers recommended patterns from simple single-file setups to large multi-domain graphs.
## Small Projects
For projects with a handful of node and edge types, keep everything in two files:
```text
src/
graph.ts # Node/edge definitions + graph
graph-store.ts # Store instantiation
```
### graph.ts
Contains all definitions and exports the graph:
```typescript
import { z } from "zod";
import { defineNode, defineEdge, defineGraph, disjointWith } from "@nicia-ai/typegraph";
// Node definitions
export const Person = defineNode("Person", {
schema: z.object({
name: z.string(),
email: z.string().email().optional(),
}),
});
export const Company = defineNode("Company", {
schema: z.object({
name: z.string(),
industry: z.string().optional(),
}),
});
// Edge definitions
export const worksAt = defineEdge("worksAt", {
schema: z.object({
role: z.string().optional(),
since: z.string().optional(),
}),
});
// Graph definition
export const graph = defineGraph({
id: "my_app",
nodes: {
Person: { type: Person },
Company: { type: Company },
},
edges: {
worksAt: { type: worksAt, from: [Person], to: [Company] },
},
ontology: [disjointWith(Person, Company)],
});
```
### graph-store.ts
Instantiates and exports the store:
```typescript
import { createStore } from "@nicia-ai/typegraph";
import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local";
import { graph } from "./graph";
const { backend } = createLocalSqliteBackend({ path: "./data.db" });
export const store = createStore(graph, backend);
```
This separation keeps the schema definition (which is static) separate from
store instantiation (which involves runtime configuration like database paths).
## Medium Projects
When your graph grows to 10+ node types or you want better organization, split definitions into separate files:
```text
src/graph/
index.ts # Re-exports + defineGraph
nodes.ts # All node definitions
edges.ts # All edge definitions
ontology.ts # Ontological relations
store.ts # Store instantiation
```
### nodes.ts
```typescript
import { z } from "zod";
import { defineNode } from "@nicia-ai/typegraph";
export const Person = defineNode("Person", {
schema: z.object({
name: z.string(),
email: z.string().email().optional(),
role: z.string().optional(),
}),
});
export const Company = defineNode("Company", {
schema: z.object({
name: z.string(),
industry: z.string().optional(),
founded: z.number().optional(),
}),
});
export const Project = defineNode("Project", {
schema: z.object({
name: z.string(),
status: z.enum(["planning", "active", "completed"]),
}),
});
// ... more node definitions
```
### edges.ts
```typescript
import { z } from "zod";
import { defineEdge } from "@nicia-ai/typegraph";
export const worksAt = defineEdge("worksAt", {
schema: z.object({
role: z.string().optional(),
since: z.string().optional(),
}),
});
export const manages = defineEdge("manages");
export const assignedTo = defineEdge("assignedTo", {
schema: z.object({
assignedAt: z.string().optional(),
}),
});
// ... more edge definitions
```
### ontology.ts
```typescript
import { subClassOf, disjointWith, inverseOf } from "@nicia-ai/typegraph";
import { Person, Company, Project } from "./nodes";
import { manages } from "./edges";
export const ontology = [
disjointWith(Person, Company),
disjointWith(Person, Project),
disjointWith(Company, Project),
// inverseOf(manages, reportsTo),
];
```
### index.ts
Combines everything into the graph definition:
```typescript
import { defineGraph } from "@nicia-ai/typegraph";
import { Person, Company, Project } from "./nodes";
import { worksAt, manages, assignedTo } from "./edges";
import { ontology } from "./ontology";
export const graph = defineGraph({
id: "my_app",
nodes: {
Person: { type: Person },
Company: { type: Company },
Project: { type: Project },
},
edges: {
worksAt: { type: worksAt, from: [Person], to: [Company] },
manages: { type: manages, from: [Person], to: [Person] },
assignedTo: { type: assignedTo, from: [Project], to: [Person] },
},
ontology,
});
// Re-export for convenience
export * from "./nodes";
export * from "./edges";
export { store } from "./store";
```
### store.ts
```typescript
import { createStore } from "@nicia-ai/typegraph";
import { createLocalSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite/local";
import { graph } from "./index";
const { backend } = createLocalSqliteBackend({ path: "./data.db" });
export const store = createStore(graph, backend);
```
## Large Projects
For large graphs with distinct domains, group related nodes and edges together:
```text
src/graph/
index.ts # Combines all domains
store.ts # Store instantiation
domains/
users.ts # User, Profile, Team + related edges
content.ts # Document, Comment, Tag + related edges
projects.ts # Project, Task, Milestone + related edges
```
### domains/users.ts
```typescript
import { z } from "zod";
import { defineNode, defineEdge, subClassOf, disjointWith } from "@nicia-ai/typegraph";
// Nodes
export const User = defineNode("User", {
schema: z.object({
email: z.string().email(),
name: z.string(),
role: z.enum(["admin", "member", "guest"]),
}),
});
export const Profile = defineNode("Profile", {
schema: z.object({
bio: z.string().optional(),
avatarUrl: z.string().optional(),
}),
});
export const Team = defineNode("Team", {
schema: z.object({
name: z.string(),
description: z.string().optional(),
}),
});
// Edges
export const hasProfile = defineEdge("hasProfile");
export const memberOf = defineEdge("memberOf", {
schema: z.object({ joinedAt: z.string().optional() }),
});
export const leads = defineEdge("leads");
// Domain-specific ontology
export const usersOntology = [
disjointWith(User, Team),
disjointWith(User, Profile),
];
// Export for graph assembly
export const usersNodes = {
User: { type: User },
Profile: { type: Profile },
Team: { type: Team },
};
export const usersEdges = {
hasProfile: { type: hasProfile, from: [User], to: [Profile] },
memberOf: { type: memberOf, from: [User], to: [Team] },
leads: { type: leads, from: [User], to: [Team] },
};
```
### index.ts
Assembles domains into the final graph:
```typescript
import { defineGraph } from "@nicia-ai/typegraph";
import { usersNodes, usersEdges, usersOntology } from "./domains/users";
import { contentNodes, contentEdges, contentOntology } from "./domains/content";
import { projectsNodes, projectsEdges, projectsOntology } from "./domains/projects";
export const graph = defineGraph({
id: "my_app",
nodes: {
...usersNodes,
...contentNodes,
...projectsNodes,
},
edges: {
...usersEdges,
...contentEdges,
...projectsEdges,
},
ontology: [
...usersOntology,
...contentOntology,
...projectsOntology,
],
});
// Re-export types for convenience
export * from "./domains/users";
export * from "./domains/content";
export * from "./domains/projects";
export { store } from "./store";
```
## Cross-Domain Edges
When edges connect nodes from different domains, define them at the graph level:
```typescript
// index.ts
import { defineEdge } from "@nicia-ai/typegraph";
import { User } from "./domains/users";
import { Document } from "./domains/content";
import { Project } from "./domains/projects";
// Cross-domain edges
const authored = defineEdge("authored");
const assignedTo = defineEdge("assignedTo");
export const graph = defineGraph({
// ...
edges: {
...usersEdges,
...contentEdges,
...projectsEdges,
// Cross-domain
authored: { type: authored, from: [User], to: [Document] },
assignedTo: { type: assignedTo, from: [User], to: [Project] },
},
});
```
## Naming Conventions
| Element | Convention | Example |
|---------|------------|---------|
| Node definitions | PascalCase | `Person`, `Company` |
| Edge definitions | camelCase | `worksAt`, `hasAuthor` |
| Graph IDs | snake_case | `my_app`, `content_graph` |
| Files | kebab-case | `graph-store.ts`, `project-structure.ts` |
| Query aliases | short lowercase | `p`, `c`, `e1` |
## Type Exports
Export types alongside definitions for use in your application:
```typescript
// graph/nodes.ts
import { type Node, type NodeProps, type NodeId } from "@nicia-ai/typegraph";
export const Person = defineNode("Person", { /* ... */ });
// Convenience type exports
export type PersonNode = Node;
export type PersonProps = NodeProps;
export type PersonId = NodeId;
```
This lets consumers import types directly:
```typescript
import { type PersonNode, type PersonProps } from "./graph";
function displayPerson(person: PersonNode) {
console.log(person.name);
}
function validatePersonInput(data: unknown): PersonProps {
return Person.schema.parse(data);
}
```
## Framework Integration
### Next.js / React Server Components
Keep the store in a server-only module:
```text
src/
graph/
index.ts
store.server.ts # Server-only store
```
```typescript
// store.server.ts
import "server-only";
import { createStore } from "@nicia-ai/typegraph";
import { graph } from "./index";
// ...
```
### Edge Runtimes (Cloudflare Workers, Vercel Edge)
Use the Drizzle backend with edge-compatible drivers:
```typescript
// graph/store.ts
import { createStore } from "@nicia-ai/typegraph";
import { createSqliteBackend } from "@nicia-ai/typegraph/adapters/drizzle/sqlite";
import { drizzle } from "drizzle-orm/d1";
import { graph } from "./index";
export function createGraphStore(env: { DB: D1Database }) {
const db = drizzle(env.DB);
const backend = createSqliteBackend(db);
return createStore(graph, backend);
}
```
## Next Steps
- [Getting Started](/getting-started) - Build your first graph
- [Schemas & Types](/core-concepts) - Deep dive into node and edge definitions
- [Integration](/integration) - Database setup and Drizzle integration
# Provenance and Retraction
> Track source lineage for derived facts, retract bad sources, and use recorded time to replay what the graph believed before and after the transition.
Provenance and Retraction is the TypeGraph subpath for source lineage and
belief transitions. It maps your ordinary graph kinds onto four roles:
- one or more retractable source node kinds with a boolean `retracted` flag
- a justification node that represents an AND support rule
- one or more derived fact node kinds
- two typed edges: premises point to justifications, and justifications derive facts
The API lives at `@nicia-ai/typegraph/provenance`:
```typescript
import { createRetractionCapability } from "@nicia-ai/typegraph/provenance";
const provenance = createRetractionCapability(store, {
source: { kind: "Source" },
justification: { kind: "Justification" },
fact: { kinds: ["Fact"] },
premiseOf: { kind: "premiseOf" },
derives: { kind: "derives" },
});
```
Use `source: { kinds: [...] }` when different source node kinds share the same
boolean retraction field:
```typescript
const provenance = createRetractionCapability(store, {
source: { kinds: ["ScannerSource", "VendorSource"] },
justification: { kind: "Justification" },
fact: { kinds: ["Vulnerability", "DeployDecision"] },
premiseOf: { kind: "premiseOf" },
derives: { kind: "derives" },
});
```
`store` must be created with `{ history: true }`. Retraction mutates graph row
currency, so TypeGraph-managed recorded capture is required:
```typescript
const [store] = await createStoreWithSchema(graph, backend, {
history: true,
});
```
For a complete runnable version, see
[Provenance Retraction](/examples/provenance-retraction).
## Graph shape
Define the roles as normal TypeGraph nodes and edges.
```typescript
const Source = defineNode("Source", {
schema: z.object({
label: z.string(),
retracted: z.boolean().default(false),
}),
});
const Fact = defineNode("Fact", {
schema: z.object({ label: z.string() }),
});
const TerminalFact = defineNode("TerminalFact", {
schema: z.object({ label: z.string() }),
});
const Justification = defineNode("Justification", {
schema: z.object({ label: z.string() }),
});
const premiseOf = defineEdge("premiseOf");
const derives = defineEdge("derives");
const graph = defineGraph({
id: "claims",
nodes: {
Source: { type: Source },
Fact: { type: Fact },
TerminalFact: { type: TerminalFact },
Justification: { type: Justification },
},
edges: {
premiseOf: { type: premiseOf, from: [Source, Fact], to: [Justification] },
derives: { type: derives, from: [Justification], to: [Fact, TerminalFact] },
},
});
```
A justification fires when all of its premise nodes are in the well-founded
support set. Sources are in support unless their `retracted` flag is true. Facts
enter support when at least one firing justification derives them.
Fact kinds only need to appear in `premiseOf.from` if they can support another
justification. Terminal facts can be listed in `fact.kinds` and `derives.to`
without being valid premise endpoints.
## Retraction
`retract(source)` sets the source flag, recomputes support from the current
provenance graph, and makes unsupported facts non-current. A transition only
touches facts reachable from the flipped sources, and closing a fact is a
belief-status change, not a domain delete: none of the fact's edges are
deleted (its `onDelete` behavior is not enforced), so `unRetract` restores the
fact exactly as it was.
```typescript
const before = await store.recordedNow();
const report = await provenance.retract({ kind: "Source", id: sourceId });
const after = await store.recordedNow();
const previous = before ? store.asOfRecorded(before) : undefined;
const current = after ? store.asOfRecorded(after) : undefined;
```
The report partitions facts relative to the retracted source:
- `died`: facts that were believed before and lost grounded support
- `survivedVia`: affected facts that still have a firing justification
- `unaffected`: previously believed facts outside the source's provenance
`unRetract(source)` clears the source flag, recomputes support, and reopens
facts that regain support.
Use `retractMany(sources)` or `unRetractMany(sources)` to change several source
flags in one recorded transaction:
```typescript
const report = await provenance.retractMany([
{ kind: "ScannerSource", id: scannerId },
{ kind: "VendorSource", id: vendorId },
]);
```
## Recorded time
Retraction uses TypeGraph-managed writes, so before and after states are visible
through recorded-time reads. On PostgreSQL, provenance transitions serialize
with TypeGraph-managed history writes on the same graph before computing and
applying fact currency. Capture is scoped to TypeGraph-managed writes; it does
not claim to observe out-of-band database mutations.
```typescript
const factBefore = before ? await store.asOfRecorded(before).nodes.Fact.getById(factId) : undefined;
const factAfter = after ? await store.asOfRecorded(after).nodes.Fact.getById(factId) : undefined;
```
Use `holding()` when you only need the current well-founded believed facts:
```typescript
const facts = await provenance.holding();
```
# Subqueries
> EXISTS, IN, and correlated subqueries for complex filtering
Subqueries let you filter based on conditions that depend on related data—check if related records
exist, or if values appear in another query's results.
## EXISTS
Check if related records exist:
```typescript
import { exists, fieldRef } from "@nicia-ai/typegraph";
// Find people who have authored at least one PR
const authors = await store
.query()
.from("Person", "p")
.whereNode("p", () =>
exists(
store
.query()
.from("PullRequest", "pr")
.traverse("author", "e", { direction: "in" })
.to("Person", "author")
.whereNode("author", (a) => a.id.eq(fieldRef("p", ["id"])))
.select((ctx) => ({ id: ctx.pr.id }))
.toAst()
)
)
.select((ctx) => ctx.p)
.execute();
```
## NOT EXISTS
Find records without related records:
```typescript
import { notExists, fieldRef } from "@nicia-ai/typegraph";
// Find people with no pull requests
const nonContributors = await store
.query()
.from("Person", "p")
.whereNode("p", () =>
notExists(
store
.query()
.from("PullRequest", "pr")
.traverse("author", "e", { direction: "in" })
.to("Person", "author")
.whereNode("author", (a) => a.id.eq(fieldRef("p", ["id"])))
.select((ctx) => ({ id: ctx.pr.id }))
.toAst()
)
)
.select((ctx) => ctx.p)
.execute();
```
## IN
Check if a value is in a subquery result set:
```typescript
import { inSubquery, fieldRef } from "@nicia-ai/typegraph";
// Find people who work at tech companies
const techWorkers = await store
.query()
.from("Person", "p")
.whereNode("p", () =>
inSubquery(
fieldRef("p", ["companyId"]),
store
.query()
.from("Company", "c")
.whereNode("c", (c) => c.industry.eq("Technology"))
.aggregate({
id: fieldRef("c", ["id"], { valueType: "string" }),
})
.toAst()
)
)
.select((ctx) => ctx.p)
.execute();
```
## NOT IN
Exclude values that appear in a subquery:
```typescript
import { notInSubquery, fieldRef } from "@nicia-ai/typegraph";
// Find people not in the blocklist
const allowedUsers = await store
.query()
.from("Person", "p")
.whereNode("p", () =>
notInSubquery(
fieldRef("p", ["id"]),
store
.query()
.from("BlockedUser", "b")
.aggregate({
userId: fieldRef("b", ["props", "userId"], { valueType: "string" }),
})
.toAst()
)
)
.select((ctx) => ctx.p)
.execute();
```
## fieldRef()
The `fieldRef()` function creates a reference to a field in the outer query for use in subquery predicates:
```typescript
import { fieldRef } from "@nicia-ai/typegraph";
fieldRef("alias", ["field"]) // Reference a single field
fieldRef("alias", ["nested", "path"]) // Reference a nested field
```
**Parameters:**
| Parameter | Type | Description |
|-----------|------|-------------|
| `alias` | `string` | The alias of the node/edge in the outer query |
| `path` | `string[]` | Path to the field (array for nested access) |
## Helpers Reference
| Function | Description |
|----------|-------------|
| `exists(subqueryAst)` | True if subquery returns any rows |
| `notExists(subqueryAst)` | True if subquery returns no rows |
| `inSubquery(fieldRef, subqueryAst)` | True if field value is in subquery results |
| `notInSubquery(fieldRef, subqueryAst)` | True if field value is not in subquery results |
For `inSubquery()` and `notInSubquery()`, the subquery must project exactly one
scalar column. Prefer `aggregate({ ... })` with a single field.
## Real-World Examples
### Users with Recent Activity
```typescript
// Find users who logged in within the last 7 days
const activeUsers = await store
.query()
.from("User", "u")
.whereNode("u", () =>
exists(
store
.query()
.from("LoginEvent", "e")
.whereNode("e", (e) =>
e.userId.eq(fieldRef("u", ["id"]))
.and(e.timestamp.gte(sevenDaysAgo))
)
.select((ctx) => ({ id: ctx.e.id }))
.toAst()
)
)
.select((ctx) => ctx.u)
.execute();
```
### Products Not in Any Cart
```typescript
// Find products that haven't been added to any cart
const unpopularProducts = await store
.query()
.from("Product", "p")
.whereNode("p", () =>
notExists(
store
.query()
.from("CartItem", "ci")
.whereNode("ci", (ci) => ci.productId.eq(fieldRef("p", ["id"])))
.select((ctx) => ({ id: ctx.ci.id }))
.toAst()
)
)
.select((ctx) => ctx.p)
.execute();
```
### Users in Specific Teams
```typescript
// Find users who are members of either the engineering or design team
const targetTeamIds = ["team-eng", "team-design"];
const teamMembers = await store
.query()
.from("User", "u")
.whereNode("u", () =>
inSubquery(
fieldRef("u", ["id"]),
store
.query()
.from("TeamMembership", "tm")
.whereNode("tm", (tm) => tm.teamId.in(targetTeamIds))
.aggregate({
userId: fieldRef("tm", ["props", "userId"], {
valueType: "string",
}),
})
.toAst()
)
)
.select((ctx) => ctx.u)
.execute();
```
## Query Debugging
For debugging or advanced use cases, you can inspect the query AST or generated SQL.
### View the AST
```typescript
const query = store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.status.eq("active"))
.select((ctx) => ctx.p);
const ast = query.toAst();
console.log(JSON.stringify(ast, null, 2));
```
### View Generated SQL
`toSQL()` returns the SQL text and bound parameters for the current backend dialect:
```typescript
const { sql, params } = query.toSQL();
console.log("SQL:", sql);
console.log("Parameters:", params);
```
This is useful for:
- Debugging query behavior
- Understanding performance characteristics
- Logging queries in production
- Running the query with a custom executor
## Next Steps
- [Filter](/queries/filter) - Basic filtering with predicates
- [Combine](/queries/combine) - Set operations
- [Execute](/queries/execute) - Running queries
# Aggregate
> GROUP BY, aggregate functions, and HAVING clauses
TypeGraph supports SQL-style aggregations for analytics and reporting. Group nodes by properties,
compute aggregates like COUNT and SUM, and filter groups with HAVING clauses.
Call [`asRelation()`](/queries/relations/) on an aggregate query to filter its completed output or
aggregate those results again. Both expression aggregates and compatibility aggregates support
prepared execution and one-statement batching through the shared relation API.
Typed expression callbacks are recommended for new queries. They provide schema-checked operands,
computed aggregate arguments, and inferred nullable result types:
```typescript
import { expr } from "@nicia-ai/typegraph";
const companySizes = await store
.query()
.from("Person", "p")
.groupBy((e) => [e.p.department])
.having((e) => expr.gt(expr.count(e.p.id), expr.literal(5)))
.aggregate((e) => ({
department: e.p.department,
employees: expr.count(e.p.id),
payroll: expr.sum(e.p.salary),
}))
.execute();
```
The string helpers below remain supported as compatibility adapters. See
[Database Expressions](/queries/expressions) for arithmetic, conditions, projection, and scope
safety.
## When to Use Aggregations
Aggregations are useful for:
- **Analytics dashboards**: Employee counts by department, revenue by region
- **Reporting**: Average order value, total sales by product category
- **Data exploration**: Find groups meeting certain criteria
- **Metrics**: Count active users, sum transaction amounts
## Basic Aggregation
Use `groupBy()` and `aggregate()` with aggregate helper functions:
```typescript
import { count, field } from "@nicia-ai/typegraph";
const companySizes = await store
.query()
.from("Person", "p")
.traverse("worksAt", "e")
.to("Company", "c")
.groupBy("c", "name") // Group by company name
.aggregate({
companyName: field("c", "name"), // Include the grouped field
employeeCount: count("p"), // Count people in each group
})
.execute();
// Result: [{ companyName: "Acme Corp", employeeCount: 42 }, ...]
```
## Aggregate Functions
Import aggregate functions from `@nicia-ai/typegraph`:
```typescript
import { count, countDistinct, sum, avg, min, max, field } from "@nicia-ai/typegraph";
```
### count
Count rows in each group:
```typescript
count("p") // COUNT(p.id) - count all nodes
count("p", "department") // COUNT(p.props.department) - count non-null values
```
### countDistinct
Count unique values:
```typescript
countDistinct("p") // COUNT(DISTINCT p.id)
countDistinct("p", "department") // COUNT(DISTINCT p.props.department)
```
Distinct counts support string, number, Boolean, and date fields. Structured JSON, array,
embedding, and unresolved dynamic fields are refused so SQLite and PostgreSQL cannot disagree about
value equality.
### sum
Sum numeric values:
```typescript
sum("p", "salary") // SUM(p.props.salary)
```
### avg
Average of numeric values:
```typescript
avg("p", "age") // AVG(p.props.age)
```
### min / max
Minimum and maximum values:
```typescript
min("p", "hireDate") // MIN(p.props.hireDate)
max("p", "salary") // MAX(p.props.salary)
```
Aggregate results follow JavaScript value conventions. `count()` and
`countDistinct()` always return a number, including `0` for an empty input.
`sum()`, `avg()`, `min()`, and `max()` return `undefined` when SQL produces
`NULL`, such as an aggregate over an empty input. Minimum and maximum preserve
the schema field's scalar type, so string results remain strings and date
results are decoded as `Date` values. Minimum and maximum support string, number,
and date fields; schema-known boolean and structured operands are refused.
Sum and average require numeric fields.
### field
Include a grouped field in the output:
```typescript
field("p", "department") // The grouped field value
field("c", "id") // Node ID
field("c", "name") // Property value
```
## Multiple Aggregations
Combine multiple aggregates in one query:
```typescript
import { count, countDistinct, sum, avg, min, max, field } from "@nicia-ai/typegraph";
const departmentStats = await store
.query()
.from("Employee", "e")
.groupBy("e", "department")
.aggregate({
department: field("e", "department"),
headcount: count("e"),
uniqueRoles: countDistinct("e", "role"),
avgSalary: avg("e", "salary"),
minSalary: min("e", "salary"),
maxSalary: max("e", "salary"),
totalPayroll: sum("e", "salary"),
})
.execute();
```
## Grouping by Multiple Fields
Chain `groupBy()` calls for multi-column grouping:
```typescript
const breakdown = await store
.query()
.from("Employee", "e")
.groupBy("e", "department")
.groupBy("e", "level")
.aggregate({
department: field("e", "department"),
level: field("e", "level"),
count: count("e"),
avgSalary: avg("e", "salary"),
})
.execute();
// Result: [
// { department: "Engineering", level: "Senior", count: 15, avgSalary: 150000 },
// { department: "Engineering", level: "Junior", count: 8, avgSalary: 80000 },
// { department: "Sales", level: "Senior", count: 5, avgSalary: 120000 },
// ...
// ]
```
## Grouping by Node
Use `groupByNode()` to group by unique nodes (by ID):
```typescript
const projectContributions = await store
.query()
.from("Commit", "c")
.traverse("author", "e")
.to("Developer", "d")
.groupByNode("d") // Group by developer node
.aggregate({
developerId: field("d", "id"),
developerName: field("d", "name"),
commitCount: count("c"),
})
.execute();
```
## Filtering Groups with HAVING
Use `having()` to filter groups based on aggregate values (SQL's HAVING clause):
```typescript
import { count, havingGt } from "@nicia-ai/typegraph";
// Only departments with more than 5 employees
const largeDepartments = await store
.query()
.from("Employee", "e")
.groupBy("e", "department")
.having(havingGt(count("e"), 5)) // HAVING COUNT(e) > 5
.aggregate({
department: field("e", "department"),
headcount: count("e"),
})
.execute();
```
### Available HAVING Helpers
```typescript
import {
having,
havingGt,
havingGte,
havingLt,
havingLte,
havingEq,
} from "@nicia-ai/typegraph";
// Comparison helpers
havingGt(aggregate, value) // >
havingGte(aggregate, value) // >=
havingLt(aggregate, value) // <
havingLte(aggregate, value) // <=
havingEq(aggregate, value) // =
// Generic comparison (for custom operators)
having(aggregate, "gt", value)
```
### Multiple HAVING Conditions
Chain multiple having conditions:
```typescript
const qualifiedDepartments = await store
.query()
.from("Employee", "e")
.groupBy("e", "department")
.having(havingGte(count("e"), 5)) // At least 5 employees
.having(havingGte(avg("e", "salary"), 100000)) // Average salary >= 100k
.aggregate({
department: field("e", "department"),
headcount: count("e"),
avgSalary: avg("e", "salary"),
})
.execute();
```
## Aggregations with Traversals
Combine graph traversals with aggregations:
```typescript
const topContributors = await store
.query()
.from("PullRequest", "pr")
.whereNode("pr", (pr) => pr.state.eq("merged"))
.traverse("targetsRepo", "e1")
.to("Repository", "repo")
.traverse("author", "e2", { direction: "in" })
.to("Developer", "dev")
.groupBy("repo", "name")
.groupBy("dev", "name")
.aggregate({
repository: field("repo", "name"),
developer: field("dev", "name"),
prCount: count("pr"),
linesChanged: sum("pr", "linesAdded"),
})
.limit(50)
.execute();
```
## Ordering Aggregated Results
Aggregate queries have their own `orderBy(key, direction?)`, called after
`.aggregate({...})`. `key` is any output name from the fields object — a
grouped field or an aggregate alias — so `limit()` finally means "top N",
not "an arbitrary N":
```typescript
const topDepartments = await store
.query()
.from("Employee", "e")
.groupBy("e", "department")
.aggregate({
department: field("e", "department"),
headcount: count("e"),
totalSalary: sum("e", "salary"),
})
.orderBy("totalSalary", "desc")
.limit(10)
.execute();
```
## Real-World Example: Team Analytics
```typescript
import { count, countDistinct, sum, avg, field, havingGt } from "@nicia-ai/typegraph";
// 1. Productivity by department
const departmentMetrics = await store
.query()
.from("Developer", "dev")
.traverse("authored", "e")
.to("PullRequest", "pr")
.whereNode("pr", (pr) => pr.state.eq("merged"))
.groupBy("dev", "department")
.aggregate({
department: field("dev", "department"),
developerCount: countDistinct("dev"),
totalPRs: count("pr"),
totalLinesAdded: sum("pr", "linesAdded"),
avgLinesPerPR: avg("pr", "linesAdded"),
})
.execute();
// 2. Active reviewers (reviewed > 10 PRs)
const activeReviewers = await store
.query()
.from("Developer", "d")
.traverse("reviewed", "r")
.to("PullRequest", "pr")
.groupByNode("d")
.having(havingGt(count("pr"), 10))
.aggregate({
developer: field("d", "name"),
reviewCount: count("pr"),
})
.orderBy("reviewCount", "desc")
.execute();
// 3. Repository health
const repoHealth = await store
.query()
.from("Repository", "r")
.traverse("contains", "e")
.to("PullRequest", "pr")
.groupByNode("r")
.aggregate({
repo: field("r", "name"),
openPRs: count("pr"),
avgAge: avg("pr", "daysOpen"),
})
.execute();
```
## Next Steps
- [Shape](/queries/shape) - Output transformation with `select()`
- [Order](/queries/order) - Ordering and limiting results
- [Traverse](/queries/traverse) - Graph traversals
# Combine
> Set operations with union(), intersect(), and except()
Combine operations merge results from multiple queries using set operations. Use `union()` to combine
results, `intersect()` to find common results, and `except()` to exclude results.
For new SQL projections, use [`project().asRelation()`](/queries/relations/) to combine visible output
columns, then order, filter, prepare, or batch the combined relation. The `select()` examples below
describe the compatibility API; its JavaScript result mapper does not define SQL row equality.
## Set Operations Overview
| Operation | Description | Duplicates |
|-----------|-------------|------------|
| `union()` | Combine results from both queries | Removed |
| `unionAll()` | Combine results from both queries | Kept |
| `intersect()` | Results that appear in both queries | Removed |
| `except()` | Results in first query but not second | Removed |
## union()
Combine results from multiple queries, removing duplicates:
```typescript
const activeOrAdmin = await store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.status.eq("active"))
.select((ctx) => ({ id: ctx.p.id, name: ctx.p.name }))
.union(
store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.role.eq("admin"))
.select((ctx) => ({ id: ctx.p.id, name: ctx.p.name }))
)
.execute();
```
This returns all active users PLUS all admins, with duplicates removed (active admins appear once).
### Selection Shape Must Match
Both queries must have the same selection shape:
```typescript
// Valid: Same shape
query1.select((ctx) => ({ id: ctx.p.id, name: ctx.p.name }))
.union(
query2.select((ctx) => ({ id: ctx.p.id, name: ctx.p.name }))
)
// Invalid: Different shapes - will cause an error
query1.select((ctx) => ({ id: ctx.p.id }))
.union(
query2.select((ctx) => ({ id: ctx.p.id, name: ctx.p.name }))
)
```
## unionAll()
Combine results keeping duplicates:
```typescript
const allMentions = await store
.query()
.from("Comment", "c")
.whereNode("c", (c) => c.mentions.contains(userId))
.select((ctx) => ({ id: ctx.c.id, text: ctx.c.text }))
.unionAll(
store
.query()
.from("Post", "p")
.whereNode("p", (p) => p.mentions.contains(userId))
.select((ctx) => ({ id: ctx.p.id, text: ctx.p.content }))
)
.execute();
```
Use `unionAll()` when:
- You want to preserve duplicates
- Performance matters (no deduplication overhead)
- You're counting occurrences
## intersect()
Find results that appear in both queries:
```typescript
const activeAdmins = await store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.status.eq("active"))
.select((ctx) => ({ id: ctx.p.id }))
.intersect(
store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.role.eq("admin"))
.select((ctx) => ({ id: ctx.p.id }))
)
.execute();
```
This returns only users who are BOTH active AND admins.
### Equivalent to AND
`intersect()` can often be replaced with combined predicates:
```typescript
// Using intersect
query1.intersect(query2)
// Often equivalent to
.whereNode("p", (p) =>
p.status.eq("active").and(p.role.eq("admin"))
)
```
Use `intersect()` when the queries are complex or involve different traversal paths.
## except()
Find results in the first query but not the second (set difference):
```typescript
const nonAdminActive = await store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.status.eq("active"))
.select((ctx) => ({ id: ctx.p.id }))
.except(
store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.role.eq("admin"))
.select((ctx) => ({ id: ctx.p.id }))
)
.execute();
```
This returns active users who are NOT admins.
### Order Matters
Unlike `union()` and `intersect()`, the order of queries in `except()` matters:
```typescript
// Active users who are NOT admins
activeUsers.except(admins)
// Admins who are NOT active (different result!)
admins.except(activeUsers)
```
## Chaining Set Operations
Chain multiple set operations:
```typescript
const complexSet = await store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.status.eq("active"))
.select((ctx) => ({ id: ctx.p.id }))
.union(
store.query()
.from("Person", "p")
.whereNode("p", (p) => p.role.eq("admin"))
.select((ctx) => ({ id: ctx.p.id }))
)
.except(
store.query()
.from("Person", "p")
.whereNode("p", (p) => p.suspended.eq(true))
.select((ctx) => ({ id: ctx.p.id }))
)
.execute();
// (active OR admin) AND NOT suspended
```
## Ordering and Limiting Combined Results
Apply ordering and limits after set operations:
```typescript
const results = await query1
.union(query2)
.orderBy("name", "asc")
.limit(100)
.execute();
```
## Real-World Examples
### Multi-Source Search
Search across different node types:
```typescript
import { expr } from "@nicia-ai/typegraph";
async function globalSearch(term: string) {
const people = store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.name.ilike(`%${term}%`))
.project((fields) => ({
id: fields.p.id,
type: expr.literal("person" as string),
title: fields.p.name,
}))
.asRelation();
const companies = store
.query()
.from("Company", "c")
.whereNode("c", (c) => c.name.ilike(`%${term}%`))
.project((fields) => ({
id: fields.c.id,
type: expr.literal("company" as string),
title: fields.c.name,
}))
.asRelation();
return people
.union(companies)
.limit(20)
.execute();
}
```
### Exclude Blocklist
```typescript
const eligibleUsers = await store
.query()
.from("User", "u")
.whereNode("u", (u) => u.status.eq("active"))
.select((ctx) => ({ id: ctx.u.id, email: ctx.u.email }))
.except(
store
.query()
.from("BlockedUser", "b")
.traverse("blockedUser", "e")
.to("User", "u")
.select((ctx) => ({ id: ctx.u.id, email: ctx.u.email }))
)
.execute();
```
### Find Common Connections
```typescript
async function mutualFriends(userId1: string, userId2: string) {
const user1Friends = store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.id.eq(userId1))
.traverse("follows", "e")
.to("Person", "friend")
.select((ctx) => ({ id: ctx.friend.id, name: ctx.friend.name }));
const user2Friends = store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.id.eq(userId2))
.traverse("follows", "e")
.to("Person", "friend")
.select((ctx) => ({ id: ctx.friend.id, name: ctx.friend.name }));
return user1Friends
.intersect(user2Friends)
.execute();
}
```
### Deduplicate Recursive Results
Remove duplicate nodes from recursive traversals by projecting proven node identities:
```typescript
// Get unique reachable nodes (recursive may return duplicates via different paths)
const uniqueNodes = await store
.query()
.from("Node", "start")
.traverse("linkedTo", "e")
.recursive()
.to("Node", "reachable")
.project((fields) => ({
kind: fields.reachable.kind,
id: fields.reachable.id,
}))
.asRelation()
.distinctNodes({ kind: "kind", id: "id" })
.execute();
```
`distinctNodes()` accepts only the `kind` and `id` columns from one proven node binding. This avoids
choosing an arbitrary edge, path, or payload value when several matches reach the same node.
## Using Set Operations with batch()
Set operation queries implement the `BatchableQuery` interface, so you can include them
in `store.batch()` alongside regular queries:
```typescript
const [adminOrOwner, companies] = await store.batch(
store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.role.eq("admin"))
.select((ctx) => ({ id: ctx.p.id, name: ctx.p.name }))
.union(
store
.query()
.from("Person", "p")
.whereNode("p", (p) => p.role.eq("owner"))
.select((ctx) => ({ id: ctx.p.id, name: ctx.p.name })),
),
store
.query()
.from("Company", "c")
.select((ctx) => ({ id: ctx.c.id, name: ctx.c.name })),
);
```
See [Batch Query Execution](/schemas-stores#batch-query-execution) for details.
## Next Steps
- [Advanced](/queries/advanced) - Subqueries with `exists()` and `inSubquery()`
- [Execute](/queries/execute) - Running queries
- [Compose](/queries/compose) - Reusable query fragments
# Compose
> Reusable query transformations with pipe() and fragment composition
Compose operations let you create reusable query transformations. Use `pipe()` to apply
transformations and `createFragment()` to build typed, composable query parts.
## The pipe() Method
Apply a transformation function to a query builder:
```typescript
const results = await store
.query()
.from("User", "u")
.pipe((q) => q.whereNode("u", ({ status }) => status.eq("active")))
.pipe((q) => q.orderBy("u", "createdAt", "desc"))
.select((ctx) => ctx.u)
.execute();
```
Each `pipe()` receives the current builder and returns a modified builder, enabling chained transformations.
## Defining Reusable Fragments
Extract common patterns into reusable functions:
```typescript
// Define reusable fragments
const activeOnly = (q) =>
q.whereNode("u", ({ status }) => status.eq("active"));
const recentFirst = (q) =>
q.orderBy("u", "createdAt", "desc");
const first10 = (q) =>
q.limit(10);
// Use in queries
const results = await store
.query()
.from("User", "u")
.pipe(activeOnly)
.pipe(recentFirst)
.pipe(first10)
.select((ctx) => ctx.u)
.execute();
```
## Typed Fragments with createFragment()
For full type safety, use the `createFragment()` factory:
```typescript
import { createFragment } from "@nicia-ai/typegraph";
// Create a typed fragment factory for your graph
const fragment = createFragment();
// Define typed fragments
const activeUsers = fragment((q) =>
q.whereNode("u", ({ status }) => status.eq("active"))
);
const withRecentPosts = fragment((q) =>
q.traverse("authored", "a")
.to("Post", "p")
.whereNode("p", ({ createdAt }) => createdAt.gte("2024-01-01"))
);
// Compose into queries
const results = await store
.query()
.from("User", "u")
.pipe(activeUsers)
.pipe(withRecentPosts)
.select((ctx) => ({
user: ctx.u,
post: ctx.p,
}))
.execute();
```
## Composing Fragments
Use `composeFragments()` to combine multiple fragments into one:
```typescript
import { composeFragments, limitFragment, orderByFragment } from "@nicia-ai/typegraph";
// Compose multiple fragments into one
const paginatedActiveUsers = composeFragments(
(q) => q.whereNode("u", ({ status }) => status.eq("active")),
(q) => q.orderBy("u", "createdAt", "desc"),
(q) => q.limit(20)
);
// Apply as a single transformation
const results = await store
.query()
.from("User", "u")
.pipe(paginatedActiveUsers)
.select((ctx) => ctx.u)
.execute();
```
## Helper Fragments
TypeGraph provides pre-built helper fragments:
```typescript
import {
limitFragment,
offsetFragment,
orderByFragment,
composeFragments
} from "@nicia-ai/typegraph";
// Pre-built fragments
const paginated = composeFragments(
orderByFragment("u", "createdAt", "desc"),
limitFragment(20),
offsetFragment(40)
);
const results = await store
.query()
.from("User", "u")
.pipe(paginated)
.select((ctx) => ctx.u)
.execute();
```
### Available Helpers
| Helper | Description |
|--------|-------------|
| `limitFragment(n)` | Limits results to n rows |
| `offsetFragment(n)` | Skips the first n rows |
| `orderByFragment(alias, field, direction)` | Orders by a field |
## Fragments with Traversals
Fragments can include traversals:
```typescript
// Fragment that adds a manager traversal
const withManager = fragment((q) =>
q.traverse("reportsTo", "r").to("User", "manager")
);
// Fragment that adds department info
const withDepartment = fragment((q) =>
q.traverse("belongsTo", "b").to("Department", "dept")
);
// Compose for a complete employee view
const employeeDetails = composeFragments(withManager, withDepartment);
const results = await store
.query()
.from("User", "u")
.pipe(employeeDetails)
.select((ctx) => ({
employee: ctx.u,
manager: ctx.manager,
department: ctx.dept,
}))
.execute();
```
## Post-Select Fragments
`pipe()` is also available on `ExecutableQuery`:
```typescript
// Define a pagination fragment for executable queries
const paginate = (q) =>
q.orderBy("u", "name", "asc").limit(10).offset(20);
const results = await store
.query()
.from("User", "u")
.select((ctx) => ({ name: ctx.u.name, email: ctx.u.email }))
.pipe(paginate)
.execute();
```
## Real-World Patterns
### Search with Conditional Filters
```typescript
function searchUsers(filters: {
status?: string;
role?: string;
search?: string;
}) {
let query = store.query().from("User", "u");
// Apply filters conditionally using pipe
if (filters.status) {
query = query.pipe((q) =>
q.whereNode("u", ({ status }) => status.eq(filters.status))
);
}
if (filters.role) {
query = query.pipe((q) =>
q.whereNode("u", ({ role }) => role.eq(filters.role))
);
}
if (filters.search) {
query = query.pipe((q) =>
q.whereNode("u", ({ name }) => name.ilike(`%${filters.search}%`))
);
}
return query.select((ctx) => ctx.u).execute();
}
```
### Configurable Pagination
```typescript
function createPaginationFragment(options: {
sortField: string;
sortDir: "asc" | "desc";
page: number;
pageSize: number;
}) {
return composeFragments(
orderByFragment("u", options.sortField, options.sortDir),
limitFragment(options.pageSize),
offsetFragment((options.page - 1) * options.pageSize)
);
}
// Use with any query
const pagination = createPaginationFragment({
sortField: "createdAt",
sortDir: "desc",
page: 2,
pageSize: 25,
});
const results = await store
.query()
.from("User", "u")
.pipe(pagination)
.select((ctx) => ctx.u)
.execute();
```
### Domain-Specific Query Helpers
```typescript
// Create domain-specific query helpers
const userQueries = {
active: (q) => q.whereNode("u", ({ status }) => status.eq("active")),
verified: (q) => q.whereNode("u", ({ emailVerified }) => emailVerified.eq(true)),
withRole: (role: string) => (q) =>
q.whereNode("u", ({ role: r }) => r.eq(role)),
withPosts: (q) => q.traverse("authored", "a").to("Post", "p"),
recentlyActive: (q) => q.whereNode("u", ({ lastLogin }) =>
lastLogin.gte(new Date(Date.now() - 7 * 24 * 60 * 60 * 1000).toISOString())
),
};
// Compose for specific use cases
const activeAdmins = await store
.query()
.from("User", "u")
.pipe(userQueries.active)
.pipe(userQueries.verified)
.pipe(userQueries.withRole("admin"))
.select((ctx) => ctx.u)
.execute();
```
## Type Definitions
For advanced use cases, TypeGraph exports fragment type definitions:
```typescript
import type {
QueryFragment,
FlexibleQueryFragment,
TraversalFragment
} from "@nicia-ai/typegraph";
```
- **`QueryFragment`** - A typed fragment transformation
- **`FlexibleQueryFragment