Skip to content

Add reproducible shared-worktree performance baseline - #177

Merged
Shengyu Fu (shengyfu) merged 7 commits into
mainfrom
shengyfu-shared-index-benchmarks
Oct 6, 2026
Merged

Shengyu Fu (shengyfu) merged 7 commits into
mainfrom
shengyfu-shared-index-benchmarks

Conversation

@shengyfu

@shengyfu Shengyu Fu (shengyfu) commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Outcome

37 selected paired cases / 74 mode runs, backed by 43 validated raw pairs / 86 runs. Lossless JSON samples, backend diagnostics, provenance, counters, hashes and cleanup evidence are committed.

Full baseline, immutable artifacts, methodology and reproduction commands.

Merge prerequisite: #175 must land first. Its root-script discovery at e7b62505f104f72b5cffaa3c816dd06b000d58f7 runs this suite in the Ubuntu/macOS/Windows matrix. Standalone #177's preexisting hosted CI does not yet execute these Python tests.

Not a universal speedup: LF memory/storage improve, but startup and end-to-end CLI queries are slower. Clean CRLF reverses the memory/storage benefit. Keep shared mode opt-in and evaluate intended workloads before runtime rollout.

Fresh workload Ready time ordinary / shared Endpoint PSS ordinary / shared Persistent storage ordinary / shared
LF, 32 views, 2 MiB/tree 2.05 / 10.19 s 272.0 / 19.1 MiB 58.75 / 1.84 MiB
CRLF, 32 views, 2 MiB/tree 1.93 / 17.20 s 266.9 / 316.5 MiB 58.84 / 148.54 MiB
LF, 32 views, 32 MiB/tree supplement 18.74 / 25.73 s 2129.6 / 107.9 MiB 927.81 / 28.37 MiB
CRLF, 4 views, 32 MiB/tree supplement 2.28 / 23.10 s 274.7 / 758.4 MiB 116.16 / 321.74 MiB

Primary LF32 query p95: 11.45 / 35.48 ms, 960 samples/mode. Scale rows have only three samples/view, not robust tail estimates. RSS sums do not prove physical mmap deduplication; Linux PSS is additionally reported. Ordinary restart can serve an existing index while reconciling; shared readiness includes initial verification/checkpoint publication.

LF32's default120s idle pass completed32 extra reconciliations: 8,320 reads, zero extractions, 3.56 server-CPU seconds. One lookup at120.107s reported ready:false during reconciliation; no idle queries were issued, and polling cannot establish its duration. CRLF32 needed8,256 initial private extractions versus LF32's32 gitfile extractions.

Provenance and corrections

  • Release binary source: 2b2d497cc945c33ba4674b71c04f2ffd65221aff, the separately owned JSON-offset fix found by strict smoke. No production fix included here.
  • Frozen measured harness: 5c528fdd75e84a8e781db8c1c75fb94c31a160df; SHA-256 7165905b20edf2765a1dea46d3428ac2de670e664f9185e3aefe9857ad228650.
  • Primary: four workloads ×1/4/16/32 ×fresh/restart;256×8192-byte files;30samples/view;12 bounded churn rounds;130s idle observation. Supplement:4096×8192-byte files, LF1/4/16/32 andCRLF4 fresh; at most roughly1GiB working-tree text.
  • A timestamp audit found a coordinator Windows smoke overlapping the primary's first51s. Original bytes remain unchanged. LF1/4/16 fresh/restart table rows use a separately retained 07:08:01–07:09:18 UTC corrective run after a confirmed quiet window, with identical source/workload/sample settings. Six replacements plus originals produce43 retained pairs and37 selected pairs; no averaging or silent rewriting.
  • Post-measurement successors harden optional provenance, evidence validation and setup safety only. Successful timing commands, metric sampling and equality gates are unchanged. All43 historical pairs pass current validation with original hashes; no timed baseline was rerun or relabeled for these changes.

Safety and review follow-ups

Stdlib Python; deterministic real Git worktrees; isolated temporary fixtures/external storage; bounded deadlines and owned cleanup. Every indexed query requires positive --stats backend proof and strict JSON match/context/offset/span parity, with filename checks and timed-sample digests. No hardware-sensitive thresholds or global cache purges.

Windows children start suspended before Job assignment; teardown waits for zero active processes and handles exited parents and malformed detach output. Optional Rust probe failures become null with reasons; containment/reap failures remain fatal. Required metadata/argument failures occur before report construction and exit diagnostically; measurement/validation/cleanup failures afterward emit partial failed reports.

Latest safety follow-up (71ffa86): cap generated files at4,096, preserving the measured scale. Before population, query actual filesystem allocation geometry and reserve rounded CRLF-expanded file sizes, per-entry metadata allowance, auxiliary entries, Git/index/publication headroom and a fixed256MiB margin. Failed geometry probes fail closed. This is a preflight heuristic, not a quota or a guarantee against concurrent disk use. One shared parameter checker now enforces parser/report types and bounds, including churn_interval, file size/count, threads, idle and timeout. Focused tests cover the previously allowed131,072×256-byte case, reserve thresholds and missing/out-of-range report parameters.

Validation requires all documented evidence and cross-checks process inventories, samples, aggregates/deltas and cleanup. Storage groups are separate sequential walks, not an atomic partition: later checkpoint growth can exceed an earlier total. Storage validation checks types/completeness rather than imposing a false simultaneous-accounting inequality. Unsupported metrics are null with reasons; Linux I/O is not direct physical-device accounting.

Validation

  • Current root discovery: 27 tests pass Windows;26 plus one Windows-only skip Linux.
  • Current-source real LF/fresh smokes passed on Windows and native-ext4 Linux: one view, eight256-byte files, three queries/view. Allocation preflight, pre-finalization/finalized report validation and owned cleanup pass. Functional timings are excluded from the baseline.
  • Frozen source: strict all-scenario1/4 Windows fresh and native-ext4 Linux fresh/restart smokes passed. Coordinator independently audited combined execution and immutable artifacts.
  • All four artifacts decompress losslessly; raw/compressed hashes match;43 historical pairs remain valid.
  • Hosted checks cover format, Clippy, existing three-OS tests, CodeQL and CLA. Live check state is authoritative; pre-Add recurring shared release qualification #175 CI does not run the new Python suite.

Standalone branch from maine9d55dbbf232f0e228695f15a640cf207d4b3348. No production/runtime/lifecycle-test/workflow/cache/GC/base-migration changes. No merge requested. Interrupted, pre-backend-proof and restricted-output experiments are excluded from accepted evidence.

Shengyu Fu (shengyfu) and others added 2 commits October 5, 2026 23:27
Compare owned ordinary servers with a shared daemon using deterministic Git worktrees, strict JSON and backend gates, portable resource accounting, and process-tree cleanup tests.

Co-authored-by: Copilot App <[email protected]>
Retain lossless raw samples, metadata, backend proof and cleanup records with artifact hashes. Document LF savings alongside higher startup/query costs, CRLF private-overlay costs, readiness-contract differences and idle reconciliation work.

Co-authored-by: Copilot App <[email protected]>
@shengyfu
Shengyu Fu (shengyfu) marked this pull request as ready for review October 6, 2026 07:05
Copilot AI balanced review requested due to automatic review settings October 6, 2026 07:05
…rection

Retain the original primary JSON unchanged, disclose coordinator smoke overlap, and select six matching LF1/4/16 corrective cases. No harness or binary changes; preserve raw replacement samples and hashes.

Co-authored-by: Copilot App <[email protected]>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Optional tool probes can abort the harness, and the new test suite is not executed by CI.

Review effort: Balanced
Findings: 2 Medium severity

Open (2)
What changed in this PR

Adds a reproducible benchmark comparing ordinary per-worktree indexes with the shared daemon, without changing production behavior.

Changes:

  • Adds a strict cross-platform benchmark harness and tests.
  • Documents methodology, results, limitations, and reproduction steps.
  • Retains validated raw benchmark artifacts.
File Description
README.md Links and summarizes the shared-worktree baseline.
SHARED_INDEX_BENCHMARKS.md Documents results, methodology, safety, and provenance.
scripts/​benchmark_shared.py Implements benchmark execution, validation, metrics, and cleanup.
scripts/​test_benchmark_shared.py Tests validation and process cleanup behavior.
scripts/​benchmark-results/​2026-10-06-linux-primary.json.gz Stores primary benchmark data.
scripts/​benchmark-results/​2026-10-06-linux-scale-lf.json.gz Stores LF scale data.
scripts/​benchmark-results/​2026-10-06-linux-scale-crlf.json.gz Stores CRLF scale data.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread scripts/benchmark_shared.py Outdated
Comment thread scripts/test_benchmark_shared.py
Copilot AI balanced review requested due to automatic review settings October 6, 2026 07:12

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Report validation, failure reporting, and CI coverage have unresolved correctness gaps.

Review effort: Balanced
Findings: 3 Medium severity

Open (3)

Comment thread scripts/benchmark_shared.py
…failures

Record unavailable optional Rust probes with bounded deadlines while retaining fatal containment/reap failures. Add focused regressions and explicitly document the measured-source identity and companion CI merge prerequisite.

Co-authored-by: Copilot App <[email protected]>
Copilot AI balanced review requested due to automatic review settings October 6, 2026 07:20

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

Validate startup, counter/status, resource, storage, refresh/churn/idle and per-process cleanup evidence, with full-report and mutation regressions. Preserve main's pre-finalization ordering; Windows/native-Linux eight-file real smokes and all 43 immutable historical pairs pass. Document validation-only provenance and the observed idle readiness transition.

Co-authored-by: Copilot App <[email protected]>
Copilot AI balanced review requested due to automatic review settings October 6, 2026 07:40

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Storage validation accepts contradictory evidence, and cross-platform test discovery still depends on unmerged PR #175.

Review effort: Balanced
Findings: 1 Medium severity · 2 Low severity

Open (3)
Resolved since last review (1)

Comment thread scripts/benchmark_shared.py
Comment thread SHARED_INDEX_BENCHMARKS.md Outdated
Comment thread SHARED_INDEX_BENCHMARKS.md Outdated
Do not claim sequential directory walks form an atomic partition. Document the scope of storage checks and the required provenance failures that occur before a report exists; retain all measured data and executable code unchanged.

Co-authored-by: Copilot App <[email protected]>
Copilot AI balanced review requested due to automatic review settings October 6, 2026 07:50

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Artifact validation is incomplete, and the disk guard can underreserve high-file-count fixtures.

Review effort: Balanced
Findings: None

Resolved since last review (3)
Previously missed (2)

In code that hasn't changed since last review

Medium severity Account for filesystem overhead in high-file-count storage guard

scripts/​benchmark_shared.py:470

This reserve counts logical payload bytes, so it substantially underestimates the allowed high-file-count case. For example, 131072 256-byte files across 32 worktrees consume about 16 GiB on a typical 4 KiB-allocation filesystem before Git/index overhead, while this check reserves only about 4.4 GiB. The guard can therefore pass and then exhaust the temp volume; account for filesystem allocation/inode overhead or impose a separate file-count cap.

Medium severity Validate churn interval and enforce parser safety bounds

scripts/​benchmark_shared.py:1124

validate() omits churn_interval entirely and does not mirror the parser's safety bounds for file size/count, threads, idle duration, or timeout. As a result, the documented audit command can qualify a report with a missing churn pause or impossible run parameters, even though these values materially affect reproducibility. Require every measurement parameter and enforce the same constraints as parse_args().

Cap generated files at4096 while preserving measured cases. Account for filesystem allocation units, CRLF expansion and metadata before population. Use one type/range validator for CLI and reports, including churn_interval. Add boundary and guard regressions; all43 frozen pairs and real Windows/native-Linux setup smokes remain valid.

Co-authored-by: Copilot App <[email protected]>
Copilot AI balanced review requested due to automatic review settings October 6, 2026 08:15

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The cross-platform process-containment harness and large evidence set warrant human review, while CI discovery still depends on unmerged PR #175.

Review effort: Balanced
Findings: None

@shengyfu
Shengyu Fu (shengyfu) merged commit 677a4ba into main Oct 6, 2026
12 checks passed
@shengyfu
Shengyu Fu (shengyfu) deleted the shengyfu-shared-index-benchmarks branch October 6, 2026 15:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants