Skip to content

GRAPH INSERT EDGE leaves the collection descriptor inconsistent with the replicated metadata log; next restart permanently wedges all DDL #188

Description

@mkhairi

Version / build tested against

origin/main @ eea86b279

Deployment mode

Origin — single node (local)

Engine(s) involved

Graph (overlay), Document (strict)

Summary

A single GRAPH INSERT EDGE against a document_strict collection leaves that collection's on-disk descriptor version out of step with the replicated metadata log (local prior 2, replicated 1). Nothing fails while the daemon stays up — but on the next restart, metadata replay hits a descriptor version anomaly for exactly the graph-touched collection and the metadata applier stops advancing its watermark, retrying the same rejected entry forever. From that point every DDL statement on the daemon times out, descriptor lease releases time out, and filtered count(*) hangs; the state survives further restarts and the only recovery is discarding the data directory. This may share a root cause with the edge double-count (one insert registers edge_count 2 / node_count 4 in SHOW GRAPH STATS): a doubled apply would also explain a double descriptor bump.

Steps to reproduce

-- fresh data directory, psql on 6432
CREATE COLLECTION g_min (id TEXT PRIMARY KEY, name TEXT) WITH (engine='document_strict');
-- CREATE COLLECTION
GRAPH INSERT EDGE IN g_min FROM 'a' TO 'b' TYPE 'knows' PROPERTIES '{}';
-- INSERT EDGE

-- restart the daemon, then:
CREATE COLLECTION after_restart (id TEXT PRIMARY KEY) WITH (engine='document_strict');
-- ERROR:  metadata propose: configuration error: metadata propose timed out after 5s waiting for log index 7 (current: 1)

-- control: the same fresh-datadir + restart WITHOUT the edge insert boots and accepts DDL normally

Expected behavior

Restart after a graph edge insert replays the metadata log cleanly; the applier accepts or reconciles the replayed descriptor version, and DDL keeps working.

Actual behavior

The applier rejects the entry (descriptor version anomaly for 'g_min': replicated version 1 is inconsistent with local prior 2 (expected 2 or prior+1)) and retries forever — 59 identical rejections over 10 minutes with zero watermark progress. All DDL times out daemon-wide, filtered count(*) hangs; point reads and inserts on existing collections keep working. No self-heal, survives restarts; only a data-directory rebuild recovers.

What actually happened? (check all that are true)

  • The server crashed, hung, or failed to start (the metadata plane hangs permanently; the process itself stays up)
  • Core functionality is broken with no acceptable workaround

Proposed severity

SEV-1 — Critical: the data directory is irreversibly damaged — the persisted metadata state can never be replayed again, and the only recovery is discarding the directory. If the process-stays-up detail reads as SEV-2 to you, no objection.

Reproducibility

Always — every attempt (two fresh data directories in a row).

Last known-good version / commit (if a regression)

(blank — restart-wedge symptoms observed on earlier 0.4.0 heads too; never known good)

Environment & logs

Linux x86_64, release build.

WARN metadata apply: durable host-side effect failed; not advancing watermark — Raft will re-deliver and retry
     index=12 last_applied=11
     error=descriptor version anomaly for 'g_min': replicated version 1 is inconsistent with local prior 2 (expected 2 or prior+1)
WARN QueryLeaseScope drop: background release failed error=configuration error: descriptor lease release did not apply within 5s

Before submitting

  • I searched existing issues and this is not a duplicate.
  • I reproduced this on a released tag or a current main build (not a stale local branch).
  • This is not a security vulnerability (those go to a private advisory).

Activity

  1. added
    status:needs-triageAwaiting maintainer triage (severity + priority)
    type:bugA defect — broken, incorrect, or lost data
    sev:1-criticalData loss, corruption, security, or crash; no workaround
    priority:P0Drop everything — fix now
    area:cluster-raftRaft, replication, consensus safety
    and removed
    status:needs-triageAwaiting maintainer triage (severity + priority)
    on Jul 20, 2026
  2. farhan-syah commented on Jul 20, 2026

    @farhan-syah
    Member

    Maintainer triage: SEV-1 / P0, confirmed.

    A deterministic GRAPH INSERT EDGE followed by restart leaves the durable catalog unreplayable and wedges daemon-wide DDL until the data directory is discarded. That meets the repository's corruption/no-workaround threshold, even though the process remains alive. The primary ownership is metadata replication/apply (area:cluster-raft) with the graph insert path as the trigger.

    Keeping this separate from #190 for now: they share a trigger, but descriptor replay and graph-stat accounting are distinct invariants until diagnosis proves one double-apply mechanism causes both.

  3. self-assigned this
    on Jul 21, 2026
  4. added a commit that references this issue on Jul 21, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:cluster-raftRaft, replication, consensus safetyengine:graphGraph overlaysev:1-criticalData loss, corruption, security, or crash; no workaroundstatus:confirmedReproduced by a maintainertype:bugA defect — broken, incorrect, or lost data

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions