Repository navigation
Conversation
Introduce a transport-agnostic multi-rail parallel read layer: - multi_rail.rs: RailReader trait, RailManager, MultiRailReader with stripe-to-rail planning (locality + round-robin), scoped-thread concurrent reads, completion aggregation, per-stripe checksum verification, and RailReadStats bottleneck attribution. - mock_rail.rs: Mock transport with configurable per-rail bandwidth and fault injection (late_write_after_cancel + epoch guard), so the scheduling / state-machine / buffer-lifecycle path is exercisable on machines without RDMA hardware. - tests/multi_rail_mock.rs: byte-exact multi-rail aggregation, single-rail pass-through, A/B demonstration of the late-write hazard (without cancel: corruption reproduced; with cancel: epoch guard blocks it), and round-robin stripe planning. - cs-mock-bench: rail-count sweep capacity model reporting aggregate bandwidth, speedup and bottleneck classification. Invariants: the on-disk stripe layout and the upper read interface (LookupObject / ReadByDescriptor) are untouched; single-rail deployments behave exactly as before (pass-through).
- rdma.rs: implement the multi-rail RailReader trait for RdmaClient. Registration reuses the cached MR path (register_raw_buffer_cached); stripe reads go through get_descriptor_stripes_sge (wire tag 15), so each rail RDMA-WRITEs its stripe subset straight into disjoint regions of the caller's destination buffer. Works for Soft-RoCE and physical NICs alike. - examples/softroce_dual_rail.rs: wire two RdmaClients (rxe0 / rxe1) into a RailManager for dual-rail Soft-RoCE validation on commodity Ethernet.
…ards, and lint fixes
mock_rail: corrupt_stripe/stall/fail/short_read injection; multi_rail: ReadOptions{timeout,expected_generation,max_inflight_bytes} with epoch/version/inflight guards + internal Arc timeout buffer (no UAF); 7 new robustness tests (11/11 pass); Chinese comments translated to English per CLAUDE.md; clippy fixes; .gitignore excludes .workbuddy/.codegraph; Cargo.toml adds twox-hash 1.6 + indexmap std feature. Covers contest robustness matrix, keeps single-rail pass-through and no upstream interface change.
- examples/softroce_dual_rail.rs: from wiring-only to full e2e: PUT via gRPC control plane, lookup_object for a fresh descriptor, locally derived per-stripe xxh3-64 checksums, single-rail baseline read, dual-rail read, byte-exact + checksum verification, and a speedup / bottleneck report. Endpoints default to 127.0.0.1:50053/50054 (RDMA control channels) to avoid the gRPC :50051 port; GID index is configurable (CS_RAIL_GID, default 1 for RXE). - configs/server-wsl2-softroce.toml: single-machine demo config with 4 MiB striping threshold/chunk so a 64 MiB object becomes 16 stripes (production default of 256 MiB would not stripe it).
Full design doc for the multi-rail parallel read proposal (DaoCloud#33): background, goals/non-goals, codebase generalization analysis, architecture, stripe-to-rail scheduling, late-RDMA-WRITE memory-safety defences, v1 failure semantics matrix, resource governance and backpressure, compatibility invariants, observability / bottleneck attribution, the 11-test plan, labelled capacity-model numbers, known limitations, and reproduction steps.
Explicit coverage of the two remaining proposal-doc items: the rail tuple (id, local device/port/GID, remote endpoint, QP/CQ, PD+MRs), static-config discovery with per-rail identity/config/fault injection, and the QP/MR/WR/CQE/buffer lifecycle contract.
RDMA-CM resolves the device from the destination address; 127.0.0.1 would resolve to loopback where no rxe device exists. Default endpoints now match setup-wsl2-rxe.sh (rxe0->veth0=192.168.96.110, rxe1->veth1=192.168.96.111).
Idempotent helper that creates the veth pair, assigns the 192.168.96.110/111 addresses, enables local delivery, and binds rxe0 / rxe1. Part of the reproduction material for the multi-rail proposal (requires a CONFIG_RDMA_RXE=y kernel).
With CONFIG_RDMA_RXE=y the driver is built-in and modprobe reports 'not found' even when RXE is fully functional. Fall back to a create/delete probe link so the check passes for built-in configs.
…ment mapping Server (tag-15 stripe-subset path): - Fix boundary check in serve_get_stripes: stripe end is an object-space offset and must be compared against descriptor size, not the client window (req.max_size). Previously every stripe beyond index 0 missed silently when the object exceeded the advertised window. Client: - multi_rail: surface per-rail index and underlying error on failure instead of a bare 'one or more rails failed'. - RdmaClient::read_stripes: translate the per-stripe segment table into a single segment spanning the whole registered object window, matching the server-side map_range_to_segments contract (segments are treated as contiguous object-space coverage from 0). Guard the translation (segment count, single registration, natural stripe offsets). - Add BufferView::len()/is_empty() accessors. Example (softroce_dual_rail): - Full e2e: PUT 64MiB via gRPC, lookup descriptor, single-rail baseline vs dual-rail read, per-stripe xxh3-64 verification plus byte equality, speedup report. Result on WSL2 Soft-RoCE (2 veth pairs, rxe0/rxe1): 73.8 MB/s single-rail -> 217.6 MB/s dual-rail = 2.95x. Config (server-wsl2-softroce.toml): - Complete mandatory [gc] fields (GcConfig has no serde defaults). - Switch io_executor to tier_b (tier_a does not implement read_aligned_into_ptr_batch used by the stripe-subset path). Upstream build fixes under --features rdma: - rdma.rs test: pass ttl_seconds to build_put_request. - rdma_bench.rs: add missing Args fields in test initializer. Verified: lib(rdma) 4/4, multi_rail_mock 11/11.
fresshman
marked this pull request as ready for review
October 7, 2026 02:42
…o materials The multi-rail read path registered destination buffers through a cache keyed by (base_ptr, length). A single client performing consecutive reads into a per-iteration Vec silently got zeros back: once the first Vec is dropped the allocator frequently reuses the same virtual address, the cache hits on that key, and the reused MR still pins the *old* physical pages while the CPU reads through the *new* mapping — so the NIC's RDMA WRITE lands where the caller never looks. The trait contract already said "pins the region for the read's duration", so the caching implementation was the deviation. Fix (option A+C): - Default path: RailReader::register now registers afresh per read via register_raw_buffer; the owning MR is held in the new `in_flight_mr` field, giving it exactly the register -> read_stripes lifetime the trait documents. - Pooled path: renamed to register_raw_buffer_pooled, documented as requiring long-lived buffer-pool semantics, plus invalidate_mr_cache() for callers that recycle buffers at a reused address. No public signature changed; the wire protocol, disk layout and upper read API are untouched. Tests / evidence: - tests/multi_rail_mr_lifetime.rs: consecutive reads with address reuse, object switching on one reader, and a compile-time check of the pooled API (rdma feature gated). - Soft-RoCE (real verbs): 20 measurement points across 32/64/128/256/512/1024 MiB and concurrency 1/2, every one verify_ok=true; e2e dual-rail stays verify_ok=true. New tooling (contest deliverables e/f/g): - examples/softroce_concurrency.rs: object-size x concurrency sweep reporting aggregate goodput, p50/p95 latency and process resources (RSS, registered memory, inflight bytes). - build-and-test.cmd: one-shot build + test + mock-bench driver (no RDMA HW). - docs/wsl2-softroce-setup.md: WSL2 RXE two-rail bring-up guide. Docs: design doc gains 8.1 (MR caching vs buffer lifetime case study), the measured concurrency/resource tables in 12.3 (dual-rail speedup peaks at 3.70x at 128 MiB, decays past 256 MiB as RXE's single-QP CPU serialization binds), and a new limitation row in 13.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #33
Summary
Generalizes the existing per-node concurrent stripe read
(
get_descriptor_stripes_sge, wire tag 15) to per-rail concurrency: oneclient worker transfers disjoint stripe subsets of the same object over N
independent RDMA paths concurrently, aggregating NIC bandwidth.
Hard invariants: on-disk stripe layout untouched; upper read interface
(
LookupObject/ReadByDescriptor) untouched; single-NIC deployments are apass-through (byte-for-byte identical behavior); v1 fails the whole request
safely if any required rail fails — no partial data, no in-request retry.
Changes
multi_rail.rs—RailReadertrait,RailManager,MultiRailReader:stripe→rail planning (locality + round-robin), scoped-thread concurrent
reads, completion aggregation, per-stripe checksum verification, and
RailReadStatswith a bottleneck classifier (network / storage / software).mock_rail.rs— mock transport with configurable per-rail bandwidth andfault injection (disconnect, under-delivery,
late_write_after_cancel+epoch guard), making the full scheduling / completion / buffer-lifecycle
path testable without RDMA hardware.
rdma.rs—RailReaderimplemented forRdmaClient(real verbs path;the default read path registers the destination buffer afresh per read and
holds the MR only across
register→read_stripes, reads go throughget_descriptor_stripes_sge). Works for Soft-RoCE and physical NICs.tests/multi_rail_mock.rs— 11 hardware-free tests (see below).tests/multi_rail_mr_lifetime.rs— 3 regression guards for MR lifetime(consecutive reads with address reuse, object switching on one reader, and a
compile-time check of the pooled registration API).
cs-mock-bench— rail-count sweep capacity model with bottleneckattribution.
examples/softroce_dual_rail.rs— full dual-rail e2e on Soft-RoCE:gRPC PUT → descriptor lookup → single-rail baseline vs dual-rail read →
per-stripe xxh3-64 verification + byte equality → speedup report
(rxe0/rxe1 against a server listening via
CS_RDMA_DEVICES).examples/softroce_concurrency.rs— object-size x concurrency sweep on thereal verbs path, reporting aggregate goodput, p50/p95 latency and process
resources (RSS, registered memory, inflight bytes) — contest deliverable e.
setup-wsl2-rxe.sh+kv-service/configs/server-wsl2-softroce.toml+docs/wsl2-softroce-setup.md+build-and-test.cmd— reproducible two-railRXE bring-up, demo server config (4 MiB striping threshold, tier_b io_uring
executor), setup guide and a one-shot build/test/bench driver for WSL2.
Memory safety (late RDMA WRITE hazard)
rail's registration until join; cancel flips a logical flag only, MRs are
not deregistered early.
live == epoch;bumping
liveon cancel makes the guard skip them.aggregation; a cancelled request never reaches commit.
Test plan
cargo test -p contextstore-client-rs --test multi_rail_mock— 11/11 pass(default features, no RDMA hardware needed):
epoch guard blocks it);
cargo test -p contextstore-client-rs --test multi_rail_mr_lifetime— 3/3(the pooled-API check is
rdma-feature gated):register afresh every read;
invalidate_mr_cacheexists with the documented signature.cargo test -p contextstore-client-rs --features rdma --lib— 4/4.Capacity model (Mock, WSL2; labelled accordingly)
64 MiB object, 4 MiB chunks, 16 stripes, 125 MB/s per rail:
Soft-RoCE end-to-end (WSL2, rxe0/rxe1 over two veth pairs)
Setup: custom WSL2 kernel 6.6.87.2 with
CONFIG_RDMA_RXE=y, two RXE devicesbound to two veth pairs (192.168.96.110/111), server started with
CS_RDMA_DEVICES=rxe0:0.0.0.0:50053:1,rxe1:0.0.0.0:50054:1, thenexamples/softroce_dual_rail.rsPUTs a 64 MiB object (16 stripes x 4 MiB)via gRPC and reads it back single-rail vs dual-rail with per-stripe xxh3-64
verification plus full byte equality:
(Single runs; run-to-run variance on WSL2 is noticeable, e.g. single-rail
64 MiB measured 73.8–110.1 MB/s across runs. The dual-rail advantage was
consistent in every run.)
The e2e pass surfaced and fixed two real bugs in the existing tag-15 path:
serve_get_stripescompared a stripe's object-space end offsetagainst the client window (
req.max_size), silently missing every stripebeyond index 0 whenever the object exceeded the advertised window. Now
compared against the descriptor size.
contiguous object-space coverage from 0 (
map_range_to_segments), so thesparse per-stripe table shifted writes to wrong offsets.
RdmaClientnow advertises one segment spanning the registered object window (the
RailReadertrait keeps the per-stripe table, which the mock asserts on;translation happens in the RDMA implementation with guards).
Also fixes two pre-existing
--features rdmabuild breaks (build_put_requesttest missing
ttl_seconds;rdma_benchtestArgsinitializer missingfields), and completes the mandatory
[gc]fields in the demo config.The concurrency sweep surfaced and fixed a third, more subtle bug:
register_raw_buffer_cachedkeyed MRs by(base_ptr, length). A client doing consecutive reads into a per-iterationVecsilently read back zeros — after the firstVecis freed the allocatoroften returns the same virtual address, the cache hits, and the reused MR
still pins the old physical pages while the CPU reads the new mapping, so
the NIC's RDMA WRITE lands where the caller never looks. The trait already
documented "pins the region for the read's duration", so the cache was the
deviation. The default path now registers afresh per read (MR held in
in_flight_mracrossregister→read_stripes); the pooled path survivesas explicitly-named
register_raw_buffer_pooledwithinvalidate_mr_cache.No public signature changed; wire protocol, disk layout and upper read API
untouched. Guarded by
multi_rail_mr_lifetime.rs.Soft-RoCE concurrency / resource sweep (deliverable e)
examples/softroce_concurrency.rs, real verbs, same object / layout / server /client across rows. Every one of the 20 points verified (
verify_ok=true):Dual-rail speedup vs object size (conc=1): 32 MiB 1.54x, 64 MiB 2.79x,
128 MiB 3.70x (peak), 256 MiB 1.48x, 512 MiB 1.11x, 1024 MiB 1.22x.
Reading it honestly: multi-rail pays off in the 64–128 MiB band; below it
the per-read connect/first-WR cost dominates; above 256 MiB the speedup
collapses and absolute goodput falls to ~112 MB/s because on Soft-RoCE every
rail is emulated in the same host CPU and long chains of 4 MiB WRITEs serialize
through one QP per rail — the CPU, not the NIC, is the ceiling, and a second
rail competes for it. Concurrency helps single-rail more than dual-rail in the
sweet spot (single up to 4.46x vs dual 1.20–2.49x), and at 256 MiB dual
concurrency 1→2 regresses to 0.81x — exactly the ceiling the inflight-budget
guard exists to bound. Resources stay predictable: RSS tracks live buffers
(70 → 2054 MiB), registered memory is
size × workers × rails, inflight issize × workers.Environment honesty
Soft-RoCE (RXE) is a software RoCE implementation: both rails share the host
CPU. The >2x speedup on 2 rails is therefore not NIC bandwidth aggregation —
the single-rail baseline is capped by CPU serialization on one QP/softirq
path, and a second rail engages additional cores. The claim is the relative
single-vs-dual-rail speedup on identical code paths; no claim of physical
multi-NIC aggregate bandwidth; no GPU involved. Absolute numbers on real
hardware will differ (mock model predicts rail-bound scaling up to the
software verification ceiling).