Skip to content

[Tsavorite] Default to NativeStorageDevice on x64 Linux (+ segment-boundary flush and write-error propagation fixes) - #1991

Merged
Badrish Chandramouli (badrishc) merged 19 commits into
mainfrom
badrishc/switch-to-native-device-linux
Aug 1, 2026
Merged

Badrish Chandramouli (badrishc) merged 19 commits into
mainfrom
badrishc/switch-to-native-device-linux

Conversation

@badrishc

@badrishc Badrish Chandramouli (badrishc) commented Jul 29, 2026 •

Copy link
Copy Markdown
Collaborator

Description of Change

Makes DeviceType.Native (the NativeStorageDevice, backed by libaio / io_uring) the default storage device on x64 Linux. Previously Linux defaulted to the managed RandomAccessLocalStorageDevice; Windows already defaulted to Native. This exercises the higher-performance native backend by default and, in doing so, surfaced (and fixes) latent storage bugs that the managed device had been silently masking.

Switching the default is a one-line change in Devices.GetDefaultDeviceType(); the rest of this PR fixes the real bugs that the stricter native device exposed.

Key changes:

  • Default device (Devices.cs, DeviceType.cs): GetDefaultDeviceType() now returns Native on Windows and on x64 Linux, and RandomAccess everywhere else. The Linux native library is only shipped for linux-x64 (runtimes/linux-x64), so non-x64 Linux (e.g. arm64) and platforms without a Native implementation (e.g. macOS) fall back to the managed device instead of failing to load the native library on first IO. All unit tests that use the default device type (Garnet server tests and the Tsavorite test projects) now run on the native device on x64 Linux CI.

  • Recovery flush past the page/segment boundary (AllocatorBase.cs, ObjectAllocatorImpl.cs) — real bug fix: A recovery flush starts at scanFromAddress, which for the first record on a page is one PageHeader (64 bytes) past the page start. AsyncFlushPagesForRecovery left that flush marked non-partial while the object allocator hardcodes numBytesToWrite = PageSize, so the flush re-included the header and computed a PageSize + PageHeader-byte write. When a segment holds a single page (SegmentSize == aligned page size, as in the manySegments/4k cluster tests) that write crosses into the next segment. The managed RandomAccess device silently tolerated this by growing the segment file; NativeStorageDevice correctly rejects any write whose end exceeds segment_size. Fixed at the source by marking a mid-page recovery flush as partial (fromAddress > page start). The existing partial-flush path then derives the write end from untilAddress (the page end) instead of the hardcoded full page, and the first-record header-include still writes the whole page — so the flush writes exactly one page. The recovery-flush assertion in ObjectAllocatorImpl.WriteAsync is updated to reflect that a recovery flush may be front-partial but always extends to the page end. The now-dead partial recompute of numBytesToWrite in the WritePage path is removed (the write uses alignedBufferSize and numBytesToWrite is not read afterward). The snapshot/hybrid-log boundary that selects which records copy their objects is carried separately via formerFlushedUntilAddress, so record selection is unchanged. (An earlier revision clamped the write length in ObjectAllocatorImpl.WriteAsync; that treated the symptom and has been reverted in favor of this root-cause fix.)

  • Write-error propagation through flush / checkpoint (PageAsyncResultTypes.cs, CircularDiskWriteBuffer.cs, AllocatorBase.cs, IndexCheckpoint.cs, MallocFixedPageSize.cs, DeviceLogCommitCheckpointManager.cs) — durability fix: The object-allocator flush completion (CountdownCallbackAndContext.Decrement) hardcoded errorCode: 0 when invoking the upper-layer callback, so AllocatorBase.AsyncFlushPageCallback (which is already error-aware and records failures in errorList) treated a failed object-store flush as success and advanced FlushedUntilAddress past unwritten data. Likewise the snapshot, main-index, overflow-bucket, and checkpoint-metadata completions logged the write error but still signaled success, allowing a checkpoint to commit after a failed write. These now retain/forward the first non-zero error and fault the corresponding completion. All of these changes are error-path only; success paths are unchanged.

  • Native misalignment guard exception type (NativeStorageDevice.cs, DeviceTests.cs): ThrowMisaligned now throws TsavoriteException instead of IOException, consistent with the class's other precondition/validation guards (segment/sector-size checks). IOException remains reserved for actual kernel IO completion failures (the Read/Write completion callbacks are unchanged). Three DeviceTests assertions updated to match.

Key Technical Details

Affected types/interfaces:

  • Devices.GetDefaultDeviceType() — platform/architecture-gated default device selection.
  • AllocatorBase.AsyncFlushPagesForRecovery — a mid-page recovery flush is marked partial, so the write length is derived from untilAddress (the page end) rather than the object allocator's hardcoded full page.
  • CountdownCallbackAndContext / DiskWriteCallbackContext.Release(errorCode) — object-log flush error threading.
  • FlushCompletionTracker.SetException, TaskCompletionSource fault paths in IndexCheckpoint / MallocFixedPageSize / DeviceLogCommitCheckpointManager.
  • NativeStorageDevice.ThrowMisaligned — TsavoriteException for precondition failures.

What NOT to Do (for future agents)

  • ❌ Don't leave a mid-page flush marked non-partial while passing a full-PageSize length. Starting past the PageHeader with a full-page length overshoots the page (and, with single-page segments, the segment). A flush that starts mid-page must be marked partial so its length is derived from untilAddress (the page end).
  • ❌ Don't issue a device read/write whose end exceeds the segment size. A record starting in segment N lives entirely in segment N's file; every IO must fit within a single segment. Managed devices silently tolerate oversized segment files, but NativeStorageDevice (now the x64 Linux default) rejects such writes with an error — so this is enforced in practice.
  • ❌ Don't report a flush/checkpoint completion as successful when the underlying device write failed — advancing FlushedUntilAddress or committing a checkpoint past unwritten data causes silent data loss on recovery.

Edge Cases

Scenario Risk Mitigation
Non-x64 Linux (e.g. arm64) Medium Falls back to RandomAccess; native lib is x64-only
macOS / other non-Windows, non-Linux Low Falls back to RandomAccess (no Native impl)
SegmentSize == aligned page size (4k segments) High Mid-page recovery flush marked partial; write length == page size
Genuine device write error during flush/checkpoint High Error propagated; flush/checkpoint faults instead of silently succeeding

Testing

  • Full Tsavorite suites on the native default (x64 Linux): recovery (197 passed / 0 failed), hlog (incl. NativeStorageDevice category), recordops, session, session.context, main.
  • Garnet checkpoint/recover (SeSaveRecover*), disk-spill (low-memory), AOF recovery, and admin command suites.
  • Cluster replication (incl. the previously-failing ClusterSRPrimaryCheckpointRetrieve disk-based checkpoint-sync test, 8/8) and multilog — the suites that regressed on the native default before the recovery-flush fix.
  • CI runs the full Linux + Windows × net8.0/net10.0 × Debug/Release matrix on the native default.

Native device on all shipped platforms + CI to build the prebuilts

Making the native device the Linux default requires it to actually work on every platform Garnet ships. Previously only glibc linux-x64 (and win-x64) prebuilts existed, so Alpine/musl broke: the shipped libnative_device.so is a glibc build whose DT_NEEDED is libaio.so.1t64, which musl's loader (Alpine ships libaio.so.1) cannot satisfy — the first storage-tier/AOF write threw DllNotFoundException.

Key changes:

  • Ship a linux-musl-x64 prebuilt (libnative_device.so + libnative_device_libaio.so) built against musl; it links the portable libaio.so.1/liburing.so.2 SONAMEs. Alpine now runs the native device by default (verified: the musl .so is mmap'd in the server process, full write/SAVE/recover works).
  • RID-aware loader: NativeStorageDevice resolves runtimes/<rid>/native/ from OS + architecture + libc (linux-x64, linux-musl-x64, linux-arm64, linux-musl-arm64, win-x64, win-arm64, osx-*) instead of hardcoding linux-x64/win-x64.
  • GetDefaultDeviceType returns Native on musl x64 too (a real musl prebuilt is shipped); non-x64 and other RIDs without a shipped prebuilt still fall back to RandomAccess.
  • Glob packaging: Tsavorite.core.csproj ships Device/runtimes/** so a new RID folder committed by CI is packaged automatically (no csproj edit).
  • C++17 filesystem portability: the native C++ used <experimental/filesystem>, which the newest MSVC (VS 2026 / 17.10+) removed, breaking the Windows build. Added filesystem_compat.h (prefers C++17 <filesystem>, falls back to experimental) and compile the library as C++17. Regenerated the linux-x64/linux-musl-x64 prebuilts from the C++17 source.
  • .github/workflows/native-build.yml — the mechanism to build the canonical prebuilts: a CMake matrix builds every RID (Linux glibc/musl × x64/arm64 via containers + QEMU; Windows x64/arm64 via MSVC), verifies the exported C ABI, and — on manual dispatch with update_repo=true — opens a PR overwriting the checked-in runtimes/<rid>/native/ binaries. It also runs on any libs/storage/Tsavorite/cc/** change to verify all platforms still build. A reusable build-native.sh is shared by the workflow and developers.
  • macOS is deferred: the native C++ (file_linux.cc) is hard-wired to libaio/io_uring/O_DIRECT, none of which exist on macOS — it will not compile there. macOS keeps the managed RandomAccess device until a file_darwin.cc backend (POSIX AIO / kqueue + F_NOCACHE) is written; the workflow documents this exclusion.

Validated: full docker image validation (basic, Lua, and default + native device persistence across all 5 images incl. Alpine musl) 67/0; build-native.sh builds + verifies locally for linux-x64, linux-musl-x64, and linux-musl-arm64 (QEMU); the native-build workflow's four Linux jobs are green on GitHub runners.

Not included (deliberately scoped out)

  • A network-session teardown robustness change (RespServerSession.TryConsumeMessages) was considered but reverted to keep this PR focused. TryConsumeMessages is intentionally an exception firewall for the IO-completion/receive machinery; rethrowing from it is a separate, hot-path contract change that warrants its own PR with dedicated TLS + Lua stress testing.

Issues Fixed

No tracking issue was found for this work. Happy to open one (Enhancement) and link it if desired.

Co-authored-by: Copilot [email protected]

Make DeviceType.Native the default on Linux (previously RandomAccess);
Windows already defaulted to Native. Platforms without a Native
implementation (e.g. macOS) continue to use RandomAccess. This cascades
to all unit tests that use the default device type (Garnet server tests
and the Tsavorite test projects).

Also make NativeStorageDevice's misaligned-I/O guard throw
TsavoriteException instead of IOException, consistent with the class's
other precondition/validation guards (segment/sector-size checks);
IOException remains reserved for actual kernel I/O completion failures.
Update the corresponding DeviceTests assertions to match.

Co-authored-by: Copilot <[email protected]>
Copilot-Session: acafd58e-b5ce-420c-aec8-0e6a492d69fd
…dary

A full-page recovery flush whose fromAddress points just past the PageHeader
double-counted the header: endOffset was computed as startOffset + PageSize,
then startOffset was reset to 0 to include the header while keeping the inflated
end, producing a PageSize + PageHeader byte write. When a segment holds a single
page (SegmentSize == aligned page size), that write crosses into the next segment.

Managed devices (RandomAccess) silently tolerated this by growing the segment
file; NativeStorageDevice correctly rejects a write whose end exceeds the segment.
Clamp the page-relative end to PageSize so a full-page flush writes exactly the
page and never past it. Records never span pages, so this only removes the header
double-count.

Co-authored-by: Copilot <[email protected]>
Copilot-Session: acafd58e-b5ce-420c-aec8-0e6a492d69fd
…fault to x64 Linux

Two production-readiness fixes from code review:

- ObjectAllocatorImpl: make the page-boundary clamp unconditional instead of
  only inside the isFirstRecordOnPage branch, so no flush path (mid-page start or
  partial snapshot, in addition to full-page recovery starting past the
  PageHeader) can emit a write whose end crosses the page/segment boundary.

- Devices.GetDefaultDeviceType: default to Native on Linux only when the process
  architecture is x64, matching the shipped prebuilt native library
  (runtimes/linux-x64). Other Linux architectures (e.g. arm64) fall back to the
  managed RandomAccess device instead of failing to load the native library on
  the first storage IO.

Co-authored-by: Copilot <[email protected]>
Copilot-Session: acafd58e-b5ce-420c-aec8-0e6a492d69fd
…ardown

Two pre-existing robustness fixes surfaced by the code review of the native
device switch (the native device surfaces real IO errors that managed devices
previously masked). Both are error-path only; success paths are unchanged.

Durability: on a device write error, the object-allocator flush completion
(CountdownCallbackAndContext, reached via CircularDiskWriteBuffer) hardcoded
errorCode 0 when invoking the upper-layer callback, so AllocatorBase's
AsyncFlushPageCallback (which is already error-aware and records failures in
errorList) treated a failed flush as success and advanced FlushedUntilAddress
past unwritten data. Retain and forward the first non-zero error instead.
Likewise, the snapshot, main-index, overflow-bucket, and commit-metadata
checkpoint completions logged the error but signaled success; they now fault
their completion (TrySetException / throw) so a checkpoint cannot commit after
a failed write.

Session: RespServerSession.TryConsumeMessages caught a fatal exception, disposed
the sender, and returned 0, letting the network receive loop re-enter on the
disposed sender (null response-object pointers -> NullReferenceException). It now
rethrows; both the TLS and non-TLS receive loops already tear the connection
down cleanly on exception.

Co-authored-by: Copilot <[email protected]>
Copilot-Session: acafd58e-b5ce-420c-aec8-0e6a492d69fd
Copilot AI review requested due to automatic review settings July 29, 2026 06:16

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR switches Tsavorite’s default log device to the native backend on Windows and x64 Linux, and hardens multiple flush/checkpoint paths so that native-device strictness (segment-boundary enforcement + I/O error reporting) cannot be silently bypassed.

Changes:

  • Default Devices.GetDefaultDeviceType() to DeviceType.Native on Windows and x64 Linux, with fallback to RandomAccess elsewhere.
  • Clamp object-allocator flush lengths to the page boundary to prevent segment-crossing writes.
  • Propagate device write errors through flush/checkpoint completion paths so failures fault the checkpoint instead of committing past unwritten data; also rethrow fatal exceptions in RespServerSession.TryConsumeMessages.

Reviewed changes

Copilot reviewed 12 out of 12 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
libs/storage/Tsavorite/cs/test/test.hlog/DeviceTests.cs Updates misalignment tests to expect TsavoriteException.
libs/storage/Tsavorite/cs/src/core/Utilities/PageAsyncResultTypes.cs Threads device write error codes through multi-write completion callbacks.
libs/storage/Tsavorite/cs/src/core/Index/Recovery/IndexCheckpoint.cs Faults main-index checkpoint completion on flush error.
libs/storage/Tsavorite/cs/src/core/Index/CheckpointManagement/DeviceLogCommitCheckpointManager.cs Converts checkpoint-metadata write errors into a thrown TsavoriteException.
libs/storage/Tsavorite/cs/src/core/Device/NativeStorageDevice.cs Uses TsavoriteException for misalignment precondition failures.
libs/storage/Tsavorite/cs/src/core/Device/DeviceType.cs Updates DeviceType documentation around defaults.
libs/storage/Tsavorite/cs/src/core/Device/Devices.cs Implements x64-Linux gating for defaulting to Native.
libs/storage/Tsavorite/cs/src/core/Allocator/ObjectSerialization/CircularDiskWriteBuffer.cs Passes error codes into write completion release path.
libs/storage/Tsavorite/cs/src/core/Allocator/ObjectAllocatorImpl.cs Clamps flush length to stay within a page/segment boundary.
libs/storage/Tsavorite/cs/src/core/Allocator/MallocFixedPageSize.cs Faults overflow-bucket checkpoint completion on flush error.
libs/storage/Tsavorite/cs/src/core/Allocator/AllocatorBase.cs Faults snapshot flush completion on page flush errors.
libs/server/Resp/RespServerSession.cs Rethrows fatal exceptions so connection teardown is deterministic.
Comments suppressed due to low confidence (1)

libs/storage/Tsavorite/cs/src/core/Utilities/PageAsyncResultTypes.cs:275

  • firstErrorCode is never reset when a new callback/context is installed via Set(). Since CountdownCallbackAndContext is reused across partial flush ranges, a prior non-zero error will be propagated to later successful flushes unless firstErrorCode is cleared when starting a new batch.
        public void Set(DeviceIOCompletionCallback callback, object context, uint numBytes)
        {
            this.callback = callback;
            this.context = context;
            this.numBytes = numBytes;

Comment thread libs/storage/Tsavorite/cs/src/core/Utilities/PageAsyncResultTypes.cs Outdated
Comment thread libs/storage/Tsavorite/cs/src/core/Device/DeviceType.cs Outdated
…firewall

Reverts the earlier change that made RespServerSession.TryConsumeMessages
rethrow from its catch-all. TryConsumeMessages is intentionally an exception
firewall: it is invoked from the network IO-completion/receive machinery, and
the broad catch guarantees no unexpected exception escapes into that machinery.
On error it logs, disposes the sender (closing the connection), and returns so
the server degrades gracefully and stays stable.

The rethrow was not required for the native-device switch (the failing cluster
tests were fixed by the ObjectAllocator recovery-flush fix), and it changes a
long-standing network hot-path contract with subtle cross-path implications
(TLS async readers, Lua redis.call). Any teardown-cleanliness improvement
belongs in a separate, dedicated change with TLS + Lua stress coverage.

Co-authored-by: Copilot <[email protected]>
Copilot-Session: acafd58e-b5ce-420c-aec8-0e6a492d69fd
… during migration

The vector-set migration stress test hammers VADD/VEMB while slots migrate
back and forth. Its retry loop already treats MOVED/timeout/connection as
retryable migration transients, but a raw Execute can also transiently return
a nil/null RedisResult mid-migration. Casting that nil to int threw a
NullReferenceException that escaped the retry filter (and the VEMB read cast a
nil to a null string[] that would NRE on .Length), intermittently failing the
test. Switching the default Linux device to NativeStorageDevice shifts IO
timing during migration and made this pre-existing flake surface more often.

Treat a nil/null result as one more retryable transient in both the write and
read loops. Data integrity is unaffected: only writes that return 1 are
recorded, and the end-of-test VEMB validation independently verifies every
recorded element survived all migrations.

Co-authored-by: Copilot <[email protected]>
Copilot-Session: acafd58e-b5ce-420c-aec8-0e6a492d69fd
…tErrorCode note

- CountdownCallbackAndContext.ToString() now null-checks `context` (it was
  guarded by the `callback` null check but dereferenced `context`, which can be
  null).
- DeviceType.RandomAccess doc now reflects the x64 gating: the Linux native
  library ships only for x64, so non-x64 Linux (e.g. arm64) and macOS default to
  RandomAccess; Windows and x64 Linux default to Native.
- Documented why firstErrorCode is intentionally not reset in
  CountdownCallbackAndContext.Set(): a fresh instance is created per partial
  flush (OnBeginPartialFlush -> new()), and a device write can RecordError before
  OnPartialFlushComplete installs the callback via Set(); resetting there would
  discard a pre-Set error and let a failed flush report success.

Co-authored-by: Copilot <[email protected]>
Copilot-Session: acafd58e-b5ce-420c-aec8-0e6a492d69fd
@badrishc
Badrish Chandramouli (badrishc) force-pushed the badrishc/switch-to-native-device-linux branch from 2779805 to c05fb65 Compare July 29, 2026 20:10
A recovery flush starts at scanFromAddress, which for the first record on a page
is one PageHeader (64 bytes) past the page start. AsyncFlushPagesForRecovery left
that flush marked non-partial while the object allocator hardcodes
numBytesToWrite = PageSize, so the flush re-included the header and computed a
PageSize + PageHeader-byte write. That write overshoots the page and, when a
segment holds a single page, the segment boundary: managed devices silently grow
the segment file, NativeStorageDevice rejects it.

Mark the mid-page recovery flush as partial (fromAddress > page start). The
existing partial path then derives the write end from untilAddress (the page end),
and the first-record header-include still writes the whole page, so the write is
exactly one page. Record selection is unchanged: the snapshot/hybrid-log boundary
is carried separately via formerFlushedUntilAddress.

Updates the recovery-flush assertion in ObjectAllocatorImpl to reflect that a
recovery flush may be front-partial but always extends to the page end. Removes
the now-dead partial recompute of numBytesToWrite in the WritePage path (the write
uses alignedBufferSize and numBytesToWrite is not read afterwards). Also reverts
the earlier defensive clamp, which treated the symptom.

Validated: ClusterSRPrimaryCheckpointRetrieve 8/8; Tsavorite recovery suite 197/0;
full cluster replication suite 107/0.

Co-authored-by: Copilot <[email protected]>
Copilot-Session: acafd58e-b5ce-420c-aec8-0e6a492d69fd
@badrishc
Badrish Chandramouli (badrishc) force-pushed the badrishc/switch-to-native-device-linux branch from c05fb65 to 0c2a887 Compare July 29, 2026 20:23
… for all platforms

Switching the Linux default to the native device surfaced that Alpine (musl) had
no working native library: the shipped prebuilt is a glibc build whose DT_NEEDED is
libaio.so.1t64, which musl's loader (and Alpine's libaio.so.1) cannot satisfy, so
the first storage-tier/AOF write threw DllNotFoundException. The native C++ itself
builds cleanly on musl; we simply were not shipping a musl binary.

Changes:
- Ship a linux-musl-x64 prebuilt (libnative_device.so + libnative_device_libaio.so)
  built against musl; it links the portable libaio.so.1 / liburing.so.2 SONAMEs.
- Make the NativeStorageDevice loader RID-aware: it resolves runtimes/<rid>/native/
  from OS + architecture + libc (linux-x64, linux-musl-x64, linux-arm64,
  linux-musl-arm64, win-x64, win-arm64, osx-*) instead of hardcoding linux-x64/win-x64.
- Package prebuilts with a recursive glob over Device/runtimes/** so a new RID folder
  is shipped automatically without a csproj edit.
- Devices.GetDefaultDeviceType now returns Native on musl x64 too (a real musl prebuilt
  is shipped); non-x64 and other unshipped RIDs still fall back to RandomAccess.
- Add .github/workflows/native-build.yml: builds native_device via CMake for every RID
  (linux glibc/musl x64+arm64 via containers/QEMU; win x64+arm64 via MSVC), verifies the
  exported C ABI, and on manual dispatch opens a PR updating the checked-in prebuilts.
  macOS is intentionally excluded until a file_darwin.cc backend exists (the current C++
  is hard-wired to libaio/io_uring/O_DIRECT and will not compile on macOS).
- Add libs/storage/Tsavorite/cc/build-native.sh: reproducible per-platform build helper
  used by both the workflow and developers.
- Update the io_uring init-failure diagnostic (container seccomp blocking io_uring_setup
  is the real cause; musl is now supported), the native README, .gitattributes (mark
  native libs binary), and the docker validation script (Alpine now tests the native
  device too).

Validated: full docker image validation (default + native device persistence across all
5 images incl. Alpine musl) 67/0; native-build.sh recipe built + verified locally for
linux-x64 glibc, linux-musl-x64, and linux-musl-arm64 (QEMU); Tsavorite device tests 89/0;
full solution build clean.

Co-authored-by: Copilot <[email protected]>
Copilot-Session: acafd58e-b5ce-420c-aec8-0e6a492d69fd
…n (fix Windows jobs)

The hosted Windows runner ships Visual Studio 18 (2026), so the hardcoded
'Visual Studio 17 2022' generator failed with 'could not find any instance of
Visual Studio'. Let CMake auto-detect the installed VS generator (-A selects the
target arch), drop the now-redundant Spectre-libs install step (the image already
includes the Spectre-mitigated CRT and ARM64 toolset), and locate dumpbin via
vswhere instead of a hardcoded 2022 path.

Co-authored-by: Copilot <[email protected]>
Copilot-Session: acafd58e-b5ce-420c-aec8-0e6a492d69fd
…r Windows/modern toolsets

Adding the Windows jobs to the native-build workflow surfaced that the C++ used
<experimental/filesystem>, which the newest MSVC toolset (Visual Studio 2026 /
17.10+) has removed, so the win-x64 and win-arm64 builds failed with C1083
"Cannot open include file: 'experimental/filesystem'".

Introduce filesystem_compat.h, which prefers C++17 <filesystem> and only falls
back to <experimental/filesystem> where <filesystem> is unavailable, exposing a
single `tsv_fs` alias. Replace the hard-coded std::experimental::filesystem uses
in file_system_disk.h and native_device.h with it, and compile the native library
as C++17 (bump CMAKE_CXX_STANDARD to 17; add /std:c++17 for MSVC and -std=c++17
for gcc/clang). No C++17-removed features are used (no std::auto_ptr etc.).

Regenerated the checked-in linux-x64 and linux-musl-x64 prebuilts from the C++17
source. The win-x64/win-arm64 and arm64 Linux prebuilts are produced by the
native-build workflow (run it with update_repo=true to refresh all of them).

Validated: Tsavorite device tests 89/0 on the regenerated glibc binary; docker
image validation (default + alpine, native + default device persistence) 28/0;
glibc and musl builds compile clean under C++17 locally.

Co-authored-by: Copilot <[email protected]>
Copilot-Session: acafd58e-b5ce-420c-aec8-0e6a492d69fd
The build succeeds but CMAKE_RUNTIME_OUTPUT_DIRECTORY is overridden, so
native_device.dll is not at build/src/Release/. Find the dll/pdb recursively under
build/ instead of assuming a generator-specific path.

Co-authored-by: Copilot <[email protected]>
Copilot-Session: acafd58e-b5ce-420c-aec8-0e6a492d69fd
/CETCOMPAT (CET shadow stack) is x86/x64-only; MSVC rejects it for ARM64 with
LNK1246, breaking the win-arm64 native build. Guard the flag on
CMAKE_GENERATOR_PLATFORM != ARM64 so win-x64 is unchanged and win-arm64 links.

Co-authored-by: Copilot <[email protected]>
Copilot-Session: acafd58e-b5ce-420c-aec8-0e6a492d69fd
…-build CI

Replaces the locally-built linux-x64/linux-musl-x64 prebuilts and the stale
(C++14) win-x64 dll with the canonical binaries produced by the native-build
workflow, and adds the previously-missing linux-arm64, linux-musl-arm64, and
win-arm64 prebuilts. All six were built from the C++17 source by
.github/workflows/native-build.yml (run 30671785943, all jobs green) and had
their exported C ABI verified in that workflow.

Runtime-validated locally: linux-x64 (Tsavorite device tests 89/0) and
linux-musl-x64 (Alpine native device active + storage-tier persistence). The
arm64 and Windows binaries are build- and symbol-verified by CI; enabling Native
by default on arm64 is deferred pending validation on real arm64 hardware
(GetDefaultDeviceType keeps arm64 on the managed device; the prebuilts allow
explicit --device-type Native there).

Co-authored-by: Copilot <[email protected]>
Copilot-Session: acafd58e-b5ce-420c-aec8-0e6a492d69fd
…ched PR branch

Reworks the workflow's publish step so a developer working on a PR that changes
the native device can get the rebuilt binaries into their own PR:

- Dispatching native-build on a feature branch with update_repo=true now commits
  the refreshed runtimes/<rid>/native/ binaries directly onto that branch (rebasing
  onto the latest tip first), so the open PR updates in place and ci.yml then tests
  the C# against the new native code. Dispatching on main/dev opens a review PR
  instead (protected branches are never pushed to directly).
- Uses an optional NATIVE_BINARIES_PAT secret for the push so it can re-trigger
  ci.yml; falls back to GITHUB_TOKEN (push lands, ci.yml runs on the next push or a
  manual re-run) when the secret is absent. Documented in the workflow header and README.
- Adds a per-branch concurrency group so two dispatches can't race on the push.
- Tightens the build-only trigger to exclude cc/README.md so a docs-only edit does
  not spin up the 6-platform build; the filter remains scoped to the native sources
  under libs/storage/Tsavorite/cc/** (a change under cs/** — including the checked-in
  binaries the publish step commits — never triggers native-build, avoiding loops).

Co-authored-by: Copilot <[email protected]>
Copilot-Session: acafd58e-b5ce-420c-aec8-0e6a492d69fd
…re mitigation

The Windows build requires the Spectre-mitigated CRT libraries (CMakeLists uses
/Qspectre /guard:cf /sdl; a missing component fails the link with MSB8040, per
cc/README.md). The workflow previously relied implicitly on the hosted image
shipping them. Make it explicit and verified:

- Add an "Ensure MSVC Spectre-mitigated CRT libraries" step that detects the
  lib\spectre\<arch> component (and the ARM64 VC toolset for the arm64 target) and
  installs it on demand, so the build no longer silently depends on image contents.
- After the build, verify the produced native_device.dll actually carries the
  security mitigations by checking the PE load config for Control Flow Guard
  (/guard:cf) and the stack Security Cookie (/GS,/sdl) — both set alongside
  /Qspectre — failing the job if they are absent.

Co-authored-by: Copilot <[email protected]>
Copilot-Session: acafd58e-b5ce-420c-aec8-0e6a492d69fd
…vice docs

macOS is not a shipped native platform here; drop the macOS/darwin explanatory
commentary from the native-build workflow header and the GetDefaultDeviceType
doc comment to keep the comments focused on what is actually built and shipped.

Co-authored-by: Copilot <[email protected]>
Copilot-Session: acafd58e-b5ce-420c-aec8-0e6a492d69fd
Regenerated by the Build Native Device workflow (run 30677804058) from the native sources on 'badrishc/switch-to-native-device-linux'.
@badrishc
Badrish Chandramouli (badrishc) merged commit d81227e into main Aug 1, 2026
218 checks passed
@badrishc
Badrish Chandramouli (badrishc) deleted the badrishc/switch-to-native-device-linux branch August 1, 2026 06:25
@github-actions github-actions Bot locked and limited conversation to collaborators Sep 30, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants