Skip to content

Cluster epoch: isolate GarnetEpoch and replace the busy-spin barrier with a jittered self-poll - #2203

Merged
Vasileios Zois (vazois) merged 15 commits into
mainfrom
vazois/garnet-epoch-improv
Oct 7, 2026
Merged

Vasileios Zois (vazois) merged 15 commits into
mainfrom
vazois/garnet-epoch-improv

Conversation

@vazois

@vazois Vasileios Zois (vazois) commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Problem

Adding a primary to a live cluster can wedge the donor: MIGRATE ... SLOTSRANGE pins CPU (~3598% observed) and makes no progress for an hour. The donor stops answering PING and logs connection failures while recipients still own zero slots. Root cause (see investigation): the config-epoch barrier that a bump waits on is an uncapped cooperative busy-spin (await Task.Yield() over a full session scan) with no backoff, so a bump that cannot immediately converge burns whole cores indefinitely.

This PR isolates the epoch barrier and replaces the busy-spin with a jittered, capped self-poll that idles instead of pinning cores, and wires production to it.

What this PR does

1. Isolate the epoch into its own instance. Extracts epoch state out of ClusterProvider into a dedicated GarnetEpoch<TEpochObserver> (libs/cluster/Server/Epochs/). Quiescence is determined through a new IEpochObserver abstraction, with ServerEpochObserverSource wrapping the live StoreWrapper scan. This isolates and clarifies the current API and makes the barrier testable without a live server. TEpochObserver is constrained to struct so the JIT specializes the generic and the quiescence check devirtualizes/inlines.

2. Replace the busy-spin with a self-poll barrier (enabled in production). BumpAndWaitForEpochTransitionAsync increments the epoch, then:

  • Fast path: a short adaptive SpinWait (40 iterations) covers the common case where sessions drain immediately.
  • Slow path: a jittered, capped exponential-backoff self-poll (Task.Delay) of the same IEpochObserver quiescence predicate.

There is no waker. A releasing session does nothing — the session hot path is untouched (zero locking/allocation on release). Each parked bump re-scans on its own jittered schedule, so concurrent bumps need no coordination (correctness rests on quiescence monotonicity: once all sessions cross the target epoch, every later re-scan still observes quiescence), and a genuinely stuck bump idles at the capped poll rate (maxParkDelay, default 50 ms) instead of burning a core. The bounded slice is also the liveness backstop: a session that quiesces is always seen by the next re-scan. Per-bump jitter decorrelates concurrent waiters (no thundering herd).

Production (ClusterProvider.BumpAndWaitForEpochTransitionAsync) now calls this path directly. The legacy busy-spin is retained as BumpAndSpinWaitForEpochTransitionAsync only as an A/B baseline for the benchmark below; it shares the exact same IEpochObserver predicate, so the two differ only in how they wait, not in what they wait for.

3. Add a reusable backoff primitive. ExponentialBackoff (libs/common/) — a lock-free, allocation-free struct implementing capped exponential backoff with equal-jitter in the upper portion of each slice (Random.Shared). The epoch self-poll derives each park slice from it.

4. Prove the problem and the fix (BDN + custom diagnoser).

  • CpuDiagnoser — a custom BenchmarkDotNet IDiagnoser reporting the per-op CPU the wall-clock mean cannot see, split into KernelMode CPU / UserMode CPU / Total CPU columns (kernel = PrivilegedProcessorTime, i.e. syscalls/scheduler). CPU is normalized over the actual workload iterations only.
  • EpochBumpContention — an A/B benchmark that runs the identical contention under both strategies (WaitMode = Spin | Poll), measuring the cost the bumping thread (the waiter) pays to complete one config-epoch bump and wait for every session to cross the new epoch, while AcquiringThreads background sessions continuously acquire/release their epoch.

5. Component tests. GarnetEpochBarrierTests (7 tests) covering the fast path, idle/ahead sessions, poll completion on blocking-session release, re-acquire at/above target, cancellation, concurrent bumps each converging, and a release-race soak.

Benchmark evidence

Environment: BenchmarkDotNet v0.15.8, net10.0, Windows 11, AMD Ryzen 7 PRO 7840U (16 logical / 8 physical cores), IterationCount=12, WarmupCount=3. The CpuDiagnoser columns are whole-process CPU per operation, summed across all threads, so a value larger than the wall-clock mean means more than one core was busy.

Waiter penalty and CPU burn (EpochBumpContention)

What it measures: the cost paid by the bumping thread (the waiter) to complete one config-epoch bump and wait for every session to cross the new epoch, while AcquiringThreads background sessions continuously acquire and release their epoch (randomized short holds). The WaitMode axis runs the identical contention under both strategies. Waiter penalty is the mean wall-clock time of one bump-and-wait; the CPU columns are whole-process CPU consumed per bump.

AcquiringThreads WaitMode Waiter penalty KernelMode CPU UserMode CPU Total CPU Allocated
2 Spin 14.33 ms 26.1 ms 64.2 ms 90.4 ms 141 B
2 Poll 23.71 ms 0.0 ms 0.1 ms 0.2 ms –
4 Spin 19.78 ms 20.1 ms 86.3 ms 106.4 ms 165 B
4 Poll 26.55 ms 0.0 ms 0.2 ms 0.2 ms –
8 Spin 17.55 ms 35.4 ms 85.9 ms 121.2 ms 146 B
8 Poll 27.15 ms 0.2 ms 0.2 ms 0.4 ms –

Reading it: the legacy Spin barrier burns ~90–121 ms of CPU per bump — several cores for a ~15–20 ms wait, and the user-mode burn grows with thread count (90 → 106 → 121 ms) — and allocates 140–165 B per bump (Task.Yield continuation boxing). This is the production ~3598%-CPU saturation in miniature. The Poll barrier cuts bump-wait CPU to ~0.2–0.4 ms (≈250–600× less) and allocates nothing, at the cost of a modest, bounded ~9 ms of extra wall-clock per bump (the jittered park quantum replacing frantic spinning).

Why this trade-off is right: the config-epoch barrier fires rarely and sits off the data hot path. Trading ~9 ms of wall-clock per bump to eliminate a full-core CPU burn is exactly what keeps the donor answering PING and lets the migration scan make progress. CPU is the axis that matters on a live server — it is the difference between serving traffic during migration and starving.

Notes / follow-ups

  • The self-poll design intentionally has no waker, signal, or queue: with nothing to release waiters, a FIFO/wake mechanism has no role, and quiescence monotonicity makes concurrent bumps safe without shared coordination. This is a deliberate simplification over the earlier wake-based prototype.
  • End-to-end migrate/failover test reproducing the CPU symptom is a natural follow-up.

Copilot AI balanced review requested due to automatic review settings October 5, 2026 19:24

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The solution references a missing project, and the CPU diagnoser uses an incompatible operation count that understates its metrics.

Review effort: Balanced
Findings: 3 Medium severity · 2 Low severity

Open (5)
What changed in this PR

This PR isolates cluster epoch tracking, adds an experimental wake-based barrier while preserving production spin behavior, and introduces tests and benchmarks.

Changes:

  • Extracts epoch tracking and observer logic into dedicated types.
  • Adds the wake signal, barrier implementation, and component tests.
  • Adds CPU/allocation benchmarks and performance-gate configuration.
File Description
Garnet.slnx Registers an epoch litmus project.
libs/​common/​Synchronization/​AsyncManualResetSignal.cs Adds a reusable async signal.
libs/​cluster/​Session/​ClusterSession.cs Reads epochs through GarnetEpoch.
libs/​cluster/​Server/​Epochs/​GarnetEpoch.cs Implements spin and wake barriers.
libs/​cluster/​Server/​Epochs/​EpochObserver.cs Abstracts session quiescence scans.
libs/​cluster/​Server/​ClusterProvider.cs Delegates epoch management.
test/​cluster/​Garnet.test.cluster/​GarnetEpochWakeBarrierTests.cs Tests wake-barrier behavior.
benchmark/​BDN.benchmark/​Diagnostics/​CpuDiagnoser.cs Adds process CPU metrics.
benchmark/​BDN.benchmark/​Cluster/​EpochWakeBarrier.cs Benchmarks release-path overhead.
benchmark/​BDN.benchmark/​Cluster/​EpochBumpContention.cs Compares spin and wake contention.
benchmark/​BDN.benchmark/​Cluster/​EpochBenchParams.cs Defines composite benchmark parameters.
test/​BDNPerfTests/​BDN_Benchmark_Config.json Adds allocation expectations.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread benchmark/BDN.benchmark/Diagnostics/CpuDiagnoser.cs
Comment thread libs/common/Synchronization/AsyncManualResetSignal.cs Outdated
Comment thread test/BDNPerfTests/BDN_Benchmark_Config.json Outdated
Comment thread libs/cluster/Server/Epochs/GarnetEpoch.cs Outdated
Comment thread libs/cluster/Server/Epochs/GarnetEpoch.cs Outdated
@vazois Vasileios Zois (vazois) changed the title Cluster epoch: isolate GarnetEpoch and add wake-based bump barrier (with CPU-burn proof harness) Cluster epoch: isolate GarnetEpoch and replace the busy-spin barrier with a jittered self-poll Oct 5, 2026

@kevin-montrose kevin-montrose left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Couple notes for efficiency

Comment thread libs/cluster/Server/Epochs/GarnetEpoch.cs Outdated
Comment thread libs/cluster/Server/Epochs/GarnetEpoch.cs Outdated
Comment thread libs/cluster/Server/Epochs/GarnetEpoch.cs Outdated
Comment thread libs/cluster/Session/ClusterSession.cs Outdated
Comment thread libs/cluster/Server/Epochs/EpochObserver.cs
Comment thread libs/cluster/Server/Epochs/EpochObserver.cs
@kevin-montrose
kevin-montrose self-requested a review October 7, 2026 17:18
@vazois
Vasileios Zois (vazois) merged commit b0255a3 into main Oct 7, 2026
169 checks passed
@vazois
Vasileios Zois (vazois) deleted the vazois/garnet-epoch-improv branch October 7, 2026 21:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants