You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Not a universal speedup: LF memory/storage improve, but startup and end-to-end CLI queries are slower. Clean CRLF reverses the memory/storage benefit. Keep shared mode opt-in and evaluate intended workloads before runtime rollout.
Fresh workload
Ready time ordinary / shared
Endpoint PSS ordinary / shared
Persistent storage ordinary / shared
LF, 32 views, 2 MiB/tree
2.05 / 10.19 s
272.0 / 19.1 MiB
58.75 / 1.84 MiB
CRLF, 32 views, 2 MiB/tree
1.93 / 17.20 s
266.9 / 316.5 MiB
58.84 / 148.54 MiB
LF, 32 views, 32 MiB/tree supplement
18.74 / 25.73 s
2129.6 / 107.9 MiB
927.81 / 28.37 MiB
CRLF, 4 views, 32 MiB/tree supplement
2.28 / 23.10 s
274.7 / 758.4 MiB
116.16 / 321.74 MiB
Primary LF32 query p95: 11.45 / 35.48 ms, 960 samples/mode. Scale rows have only three samples/view, not robust tail estimates. RSS sums do not prove physical mmap deduplication; Linux PSS is additionally reported. Ordinary restart can serve an existing index while reconciling; shared readiness includes initial verification/checkpoint publication.
LF32's default120s idle pass completed32 extra reconciliations: 8,320 reads, zero extractions, 3.56 server-CPU seconds. One lookup at120.107s reported ready:false during reconciliation; no idle queries were issued, and polling cannot establish its duration. CRLF32 needed8,256 initial private extractions versus LF32's32 gitfile extractions.
Provenance and corrections
Release binary source: 2b2d497cc945c33ba4674b71c04f2ffd65221aff, the separately owned JSON-offset fix found by strict smoke. No production fix included here.
Primary: four workloads ×1/4/16/32 ×fresh/restart;256×8192-byte files;30samples/view;12 bounded churn rounds;130s idle observation. Supplement:4096×8192-byte files, LF1/4/16/32 andCRLF4 fresh; at most roughly1GiB working-tree text.
A timestamp audit found a coordinator Windows smoke overlapping the primary's first51s. Original bytes remain unchanged. LF1/4/16 fresh/restart table rows use a separately retained 07:08:01–07:09:18 UTC corrective run after a confirmed quiet window, with identical source/workload/sample settings. Six replacements plus originals produce43 retained pairs and37 selected pairs; no averaging or silent rewriting.
Post-measurement successors harden optional provenance, evidence validation and setup safety only. Successful timing commands, metric sampling and equality gates are unchanged. All43 historical pairs pass current validation with original hashes; no timed baseline was rerun or relabeled for these changes.
Safety and review follow-ups
Stdlib Python; deterministic real Git worktrees; isolated temporary fixtures/external storage; bounded deadlines and owned cleanup. Every indexed query requires positive --stats backend proof and strict JSON match/context/offset/span parity, with filename checks and timed-sample digests. No hardware-sensitive thresholds or global cache purges.
Windows children start suspended before Job assignment; teardown waits for zero active processes and handles exited parents and malformed detach output. Optional Rust probe failures become null with reasons; containment/reap failures remain fatal. Required metadata/argument failures occur before report construction and exit diagnostically; measurement/validation/cleanup failures afterward emit partial failed reports.
Latest safety follow-up (71ffa86): cap generated files at4,096, preserving the measured scale. Before population, query actual filesystem allocation geometry and reserve rounded CRLF-expanded file sizes, per-entry metadata allowance, auxiliary entries, Git/index/publication headroom and a fixed256MiB margin. Failed geometry probes fail closed. This is a preflight heuristic, not a quota or a guarantee against concurrent disk use. One shared parameter checker now enforces parser/report types and bounds, including churn_interval, file size/count, threads, idle and timeout. Focused tests cover the previously allowed131,072×256-byte case, reserve thresholds and missing/out-of-range report parameters.
Validation requires all documented evidence and cross-checks process inventories, samples, aggregates/deltas and cleanup. Storage groups are separate sequential walks, not an atomic partition: later checkpoint growth can exceed an earlier total. Storage validation checks types/completeness rather than imposing a false simultaneous-accounting inequality. Unsupported metrics are null with reasons; Linux I/O is not direct physical-device accounting.
Validation
Current root discovery: 27 tests pass Windows;26 plus one Windows-only skip Linux.
Current-source real LF/fresh smokes passed on Windows and native-ext4 Linux: one view, eight256-byte files, three queries/view. Allocation preflight, pre-finalization/finalized report validation and owned cleanup pass. Functional timings are excluded from the baseline.
Frozen source: strict all-scenario1/4 Windows fresh and native-ext4 Linux fresh/restart smokes passed. Coordinator independently audited combined execution and immutable artifacts.
All four artifacts decompress losslessly; raw/compressed hashes match;43 historical pairs remain valid.
Hosted checks cover format, Clippy, existing three-OS tests, CodeQL and CLA. Live check state is authoritative; pre-Add recurring shared release qualification #175 CI does not run the new Python suite.
Standalone branch from maine9d55dbbf232f0e228695f15a640cf207d4b3348. No production/runtime/lifecycle-test/workflow/cache/GC/base-migration changes. No merge requested. Interrupted, pre-backend-proof and restricted-output experiments are excluded from accepted evidence.
…rection
Retain the original primary JSON unchanged, disclose coordinator smoke overlap, and select six matching LF1/4/16 corrective cases. No harness or binary changes; preserve raw replacement samples and hashes.
Co-authored-by: Copilot App <[email protected]>
…failures
Record unavailable optional Rust probes with bounded deadlines while retaining fatal containment/reap failures. Add focused regressions and explicitly document the measured-source identity and companion CI merge prerequisite.
Co-authored-by: Copilot App <[email protected]>
Do not claim sequential directory walks form an atomic partition. Document the scope of storage checks and the required provenance failures that occur before a report exists; retain all measured data and executable code unchanged.
Co-authored-by: Copilot App <[email protected]>
Account for filesystem overhead in high-file-count storage guard
scripts/benchmark_shared.py:470
This reserve counts logical payload bytes, so it substantially underestimates the allowed high-file-count case. For example, 131072 256-byte files across 32 worktrees consume about 16 GiB on a typical 4 KiB-allocation filesystem before Git/index overhead, while this check reserves only about 4.4 GiB. The guard can therefore pass and then exhaust the temp volume; account for filesystem allocation/inode overhead or impose a separate file-count cap.
Validate churn interval and enforce parser safety bounds
scripts/benchmark_shared.py:1124
validate() omits churn_interval entirely and does not mirror the parser's safety bounds for file size/count, threads, idle duration, or timeout. As a result, the documented audit command can qualify a report with a missing churn pause or impossible run parameters, even though these values materially affect reproducibility. Require every measurement parameter and enforce the same constraints as parse_args().
Cap generated files at4096 while preserving measured cases. Account for filesystem allocation units, CRLF expansion and metadata before population. Use one type/range validator for CLI and reports, including churn_interval. Add boundary and guard regressions; all43 frozen pairs and real Windows/native-Linux setup smokes remain valid.
Co-authored-by: Copilot App <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Outcome
37 selected paired cases / 74 mode runs, backed by 43 validated raw pairs / 86 runs. Lossless JSON samples, backend diagnostics, provenance, counters, hashes and cleanup evidence are committed.
Full baseline, immutable artifacts, methodology and reproduction commands.
Merge prerequisite: #175 must land first. Its root-script discovery at e7b62505f104f72b5cffaa3c816dd06b000d58f7 runs this suite in the Ubuntu/macOS/Windows matrix. Standalone #177's preexisting hosted CI does not yet execute these Python tests.
Not a universal speedup: LF memory/storage improve, but startup and end-to-end CLI queries are slower. Clean CRLF reverses the memory/storage benefit. Keep shared mode opt-in and evaluate intended workloads before runtime rollout.
Primary LF32 query p95: 11.45 / 35.48 ms, 960 samples/mode. Scale rows have only three samples/view, not robust tail estimates. RSS sums do not prove physical mmap deduplication; Linux PSS is additionally reported. Ordinary restart can serve an existing index while reconciling; shared readiness includes initial verification/checkpoint publication.
LF32's default120s idle pass completed32 extra reconciliations: 8,320 reads, zero extractions, 3.56 server-CPU seconds. One lookup at120.107s reported
ready:falseduring reconciliation; no idle queries were issued, and polling cannot establish its duration. CRLF32 needed8,256 initial private extractions versus LF32's32 gitfile extractions.Provenance and corrections
2b2d497cc945c33ba4674b71c04f2ffd65221aff, the separately owned JSON-offset fix found by strict smoke. No production fix included here.5c528fdd75e84a8e781db8c1c75fb94c31a160df; SHA-2567165905b20edf2765a1dea46d3428ac2de670e664f9185e3aefe9857ad228650.Safety and review follow-ups
Stdlib Python; deterministic real Git worktrees; isolated temporary fixtures/external storage; bounded deadlines and owned cleanup. Every indexed query requires positive
--statsbackend proof and strict JSON match/context/offset/span parity, with filename checks and timed-sample digests. No hardware-sensitive thresholds or global cache purges.Windows children start suspended before Job assignment; teardown waits for zero active processes and handles exited parents and malformed detach output. Optional Rust probe failures become null with reasons; containment/reap failures remain fatal. Required metadata/argument failures occur before report construction and exit diagnostically; measurement/validation/cleanup failures afterward emit partial failed reports.
Latest safety follow-up (
71ffa86): cap generated files at4,096, preserving the measured scale. Before population, query actual filesystem allocation geometry and reserve rounded CRLF-expanded file sizes, per-entry metadata allowance, auxiliary entries, Git/index/publication headroom and a fixed256MiB margin. Failed geometry probes fail closed. This is a preflight heuristic, not a quota or a guarantee against concurrent disk use. One shared parameter checker now enforces parser/report types and bounds, includingchurn_interval, file size/count, threads, idle and timeout. Focused tests cover the previously allowed131,072×256-byte case, reserve thresholds and missing/out-of-range report parameters.Validation requires all documented evidence and cross-checks process inventories, samples, aggregates/deltas and cleanup. Storage groups are separate sequential walks, not an atomic partition: later checkpoint growth can exceed an earlier total. Storage validation checks types/completeness rather than imposing a false simultaneous-accounting inequality. Unsupported metrics are null with reasons; Linux I/O is not direct physical-device accounting.
Validation
Standalone branch from main
e9d55dbbf232f0e228695f15a640cf207d4b3348. No production/runtime/lifecycle-test/workflow/cache/GC/base-migration changes. No merge requested. Interrupted, pre-backend-proof and restricted-output experiments are excluded from accepted evidence.