Skip to content

Multi-rail parallel read: aggregate multiple RDMA paths for single-worker stripe reads - #34

Open
fresshman wants to merge 12 commits into
DaoCloud:mainfrom
fresshman:feat/multi-rail-read
Open

fresshman wants to merge 12 commits into
DaoCloud:mainfrom
fresshman:feat/multi-rail-read

Conversation

@fresshman

@fresshman fresshman commented Oct 6, 2026 •

Copy link
Copy Markdown

Refs #33

Summary

Generalizes the existing per-node concurrent stripe read
(get_descriptor_stripes_sge, wire tag 15) to per-rail concurrency: one
client worker transfers disjoint stripe subsets of the same object over N
independent RDMA paths concurrently, aggregating NIC bandwidth.

Hard invariants: on-disk stripe layout untouched; upper read interface
(LookupObject / ReadByDescriptor) untouched; single-NIC deployments are a
pass-through (byte-for-byte identical behavior); v1 fails the whole request
safely if any required rail fails — no partial data, no in-request retry.

Changes

  • multi_rail.rs — RailReader trait, RailManager, MultiRailReader:
    stripe→rail planning (locality + round-robin), scoped-thread concurrent
    reads, completion aggregation, per-stripe checksum verification, and
    RailReadStats with a bottleneck classifier (network / storage / software).
  • mock_rail.rs — mock transport with configurable per-rail bandwidth and
    fault injection (disconnect, under-delivery, late_write_after_cancel +
    epoch guard), making the full scheduling / completion / buffer-lifecycle
    path testable without RDMA hardware.
  • rdma.rs — RailReader implemented for RdmaClient (real verbs path;
    the default read path registers the destination buffer afresh per read and
    holds the MR only across register → read_stripes, reads go through
    get_descriptor_stripes_sge). Works for Soft-RoCE and physical NICs.
  • tests/multi_rail_mock.rs — 11 hardware-free tests (see below).
  • tests/multi_rail_mr_lifetime.rs — 3 regression guards for MR lifetime
    (consecutive reads with address reuse, object switching on one reader, and a
    compile-time check of the pooled registration API).
  • cs-mock-bench — rail-count sweep capacity model with bottleneck
    attribution.
  • examples/softroce_dual_rail.rs — full dual-rail e2e on Soft-RoCE:
    gRPC PUT → descriptor lookup → single-rail baseline vs dual-rail read →
    per-stripe xxh3-64 verification + byte equality → speedup report
    (rxe0/rxe1 against a server listening via CS_RDMA_DEVICES).
  • examples/softroce_concurrency.rs — object-size x concurrency sweep on the
    real verbs path, reporting aggregate goodput, p50/p95 latency and process
    resources (RSS, registered memory, inflight bytes) — contest deliverable e.
  • setup-wsl2-rxe.sh + kv-service/configs/server-wsl2-softroce.toml +
    docs/wsl2-softroce-setup.md + build-and-test.cmd — reproducible two-rail
    RXE bring-up, demo server config (4 MiB striping threshold, tier_b io_uring
    executor), setup guide and a one-shot build/test/bench driver for WSL2.

Memory safety (late RDMA WRITE hazard)

  1. pin-until-completion: the read holds the destination buffer and every
    rail's registration until join; cancel flips a logical flag only, MRs are
    not deregistered early.
  2. epoch buffer guard: late writes land only while live == epoch;
    bumping live on cancel makes the guard skip them.
  3. two-phase commit + per-stripe checksum: stripes are verified after
    aggregation; a cancelled request never reaches commit.

Test plan

cargo test -p contextstore-client-rs --test multi_rail_mock — 11/11 pass
(default features, no RDMA hardware needed):

  • byte-exact multi-rail aggregation; single-rail pass-through;
  • late-write hazard A/B (without cancel: corruption reproduced; with cancel:
    epoch guard blocks it);
  • stripe→rail round-robin planning;
  • per-stripe checksum failure detected; single-rail disconnect safe-fail;
  • partial completion / under-delivery detected;
  • read timeout: error surfaces, caller's buffer untouched;
  • inflight-budget backpressure rejection;
  • object version-change (generation) rejection.

cargo test -p contextstore-client-rs --test multi_rail_mr_lifetime — 3/3
(the pooled-API check is rdma-feature gated):

  • consecutive reads into per-iteration buffers (address reuse) stay correct and
    register afresh every read;
  • switching objects on one reader never serves a stale registration;
  • invalidate_mr_cache exists with the documented signature.

cargo test -p contextstore-client-rs --features rdma --lib — 4/4.

Capacity model (Mock, WSL2; labelled accordingly)

64 MiB object, 4 MiB chunks, 16 stripes, 125 MB/s per rail:

rails  agg_MB/s  bal_MB/s  theo_MB/s  speedup  bottleneck
1      123.5     125.0     125.0      0.99x    network: slowest rail caps the object
2      244.9     250.0     250.0      1.96x    software: post-read stripe verification (client CPU)
3      327.4     333.3     375.0      2.62x    software: post-read stripe verification
4      484.1     500.0     500.0      3.87x    software: post-read stripe verification

Soft-RoCE end-to-end (WSL2, rxe0/rxe1 over two veth pairs)

Setup: custom WSL2 kernel 6.6.87.2 with CONFIG_RDMA_RXE=y, two RXE devices
bound to two veth pairs (192.168.96.110/111), server started with
CS_RDMA_DEVICES=rxe0:0.0.0.0:50053:1,rxe1:0.0.0.0:50054:1, then
examples/softroce_dual_rail.rs PUTs a 64 MiB object (16 stripes x 4 MiB)
via gRPC and reads it back single-rail vs dual-rail with per-stripe xxh3-64
verification plus full byte equality:

object   single-rail      dual-rail        speedup  verify
64 MiB   73.8 MB/s        217.6 MB/s       2.95x    xxh3-64 per stripe + byte equality
128 MiB  93.3 MB/s        384.1 MB/s       4.11x    xxh3-64 per stripe + byte equality

(Single runs; run-to-run variance on WSL2 is noticeable, e.g. single-rail
64 MiB measured 73.8–110.1 MB/s across runs. The dual-rail advantage was
consistent in every run.)

The e2e pass surfaced and fixed two real bugs in the existing tag-15 path:

  • server: serve_get_stripes compared a stripe's object-space end offset
    against the client window (req.max_size), silently missing every stripe
    beyond index 0 whenever the object exceeded the advertised window. Now
    compared against the descriptor size.
  • client: the server maps stripe offsets onto the segment table as
    contiguous object-space coverage from 0 (map_range_to_segments), so the
    sparse per-stripe table shifted writes to wrong offsets. RdmaClient
    now advertises one segment spanning the registered object window (the
    RailReader trait keeps the per-stripe table, which the mock asserts on;
    translation happens in the RDMA implementation with guards).

Also fixes two pre-existing --features rdma build breaks (build_put_request
test missing ttl_seconds; rdma_bench test Args initializer missing
fields), and completes the mandatory [gc] fields in the demo config.

The concurrency sweep surfaced and fixed a third, more subtle bug:

  • client, memory safety: register_raw_buffer_cached keyed MRs by
    (base_ptr, length). A client doing consecutive reads into a per-iteration
    Vec silently read back zeros — after the first Vec is freed the allocator
    often returns the same virtual address, the cache hits, and the reused MR
    still pins the old physical pages while the CPU reads the new mapping, so
    the NIC's RDMA WRITE lands where the caller never looks. The trait already
    documented "pins the region for the read's duration", so the cache was the
    deviation. The default path now registers afresh per read (MR held in
    in_flight_mr across register → read_stripes); the pooled path survives
    as explicitly-named register_raw_buffer_pooled with invalidate_mr_cache.
    No public signature changed; wire protocol, disk layout and upper read API
    untouched. Guarded by multi_rail_mr_lifetime.rs.

Soft-RoCE concurrency / resource sweep (deliverable e)

examples/softroce_concurrency.rs, real verbs, same object / layout / server /
client across rows. Every one of the 20 points verified (verify_ok=true):

size   mode    conc  agg_MB/s  p50_ms    RSS_MiB  reg_MiB  infl_MiB
32MiB  single  1      128.4     296.3        70       32       32
32MiB  dual    1      197.1     216.9        70       64       32
32MiB  single  2      489.8     118.1        70       64       64
32MiB  dual    2      490.0     118.2        70      128       64
64MiB  single  1      121.8     657.2       134       64       64
64MiB  dual    1      339.9     213.4       134      128       64
64MiB  single  2      498.6     235.1       134      128      128
64MiB  dual    2      492.9     236.4       134      256      128
128MiB single  1      111.1    1365.9       262      128      128
128MiB dual    1      411.3     305.6       262      256      128
128MiB single  2      495.9     474.0       262      256      256
128MiB dual    2      495.2     471.9       262      512      256
256MiB single  1      116.0    2485.6       518      256      256
256MiB dual    1      171.3    1617.4       518      512      256
256MiB single  2      133.3    3723.7       518      512      512
256MiB dual    2      138.1    3889.5       518     1024      512
512MiB single  1      101.3    5550.9      1030      512      512
512MiB dual    1      112.1    4558.4      1030     1024      512
1024MiB single 1       92.6   12551.7      2054     1024     1024
1024MiB dual   1      112.5    9004.3      2054     2048     1024

Dual-rail speedup vs object size (conc=1): 32 MiB 1.54x, 64 MiB 2.79x,
128 MiB 3.70x (peak), 256 MiB 1.48x, 512 MiB 1.11x, 1024 MiB 1.22x.

Reading it honestly: multi-rail pays off in the 64–128 MiB band; below it
the per-read connect/first-WR cost dominates; above 256 MiB the speedup
collapses and absolute goodput falls to ~112 MB/s
because on Soft-RoCE every
rail is emulated in the same host CPU and long chains of 4 MiB WRITEs serialize
through one QP per rail — the CPU, not the NIC, is the ceiling, and a second
rail competes for it. Concurrency helps single-rail more than dual-rail in the
sweet spot (single up to 4.46x vs dual 1.20–2.49x), and at 256 MiB dual
concurrency 1→2 regresses to 0.81x — exactly the ceiling the inflight-budget
guard exists to bound. Resources stay predictable: RSS tracks live buffers
(70 → 2054 MiB), registered memory is size × workers × rails, inflight is
size × workers.

Environment honesty

Soft-RoCE (RXE) is a software RoCE implementation: both rails share the host
CPU. The >2x speedup on 2 rails is therefore not NIC bandwidth aggregation —
the single-rail baseline is capped by CPU serialization on one QP/softirq
path, and a second rail engages additional cores. The claim is the relative
single-vs-dual-rail speedup on identical code paths; no claim of physical
multi-NIC aggregate bandwidth; no GPU involved. Absolute numbers on real
hardware will differ (mock model predicts rail-bound scaling up to the
software verification ceiling).

Introduce a transport-agnostic multi-rail parallel read layer:

- multi_rail.rs: RailReader trait, RailManager, MultiRailReader with
  stripe-to-rail planning (locality + round-robin), scoped-thread
  concurrent reads, completion aggregation, per-stripe checksum
  verification, and RailReadStats bottleneck attribution.
- mock_rail.rs: Mock transport with configurable per-rail bandwidth
  and fault injection (late_write_after_cancel + epoch guard), so the
  scheduling / state-machine / buffer-lifecycle path is exercisable on
  machines without RDMA hardware.
- tests/multi_rail_mock.rs: byte-exact multi-rail aggregation,
  single-rail pass-through, A/B demonstration of the late-write hazard
  (without cancel: corruption reproduced; with cancel: epoch guard
  blocks it), and round-robin stripe planning.
- cs-mock-bench: rail-count sweep capacity model reporting aggregate
  bandwidth, speedup and bottleneck classification.

Invariants: the on-disk stripe layout and the upper read interface
(LookupObject / ReadByDescriptor) are untouched; single-rail
deployments behave exactly as before (pass-through).
- rdma.rs: implement the multi-rail RailReader trait for RdmaClient.
  Registration reuses the cached MR path (register_raw_buffer_cached);
  stripe reads go through get_descriptor_stripes_sge (wire tag 15), so
  each rail RDMA-WRITEs its stripe subset straight into disjoint
  regions of the caller's destination buffer. Works for Soft-RoCE and
  physical NICs alike.
- examples/softroce_dual_rail.rs: wire two RdmaClients (rxe0 / rxe1)
  into a RailManager for dual-rail Soft-RoCE validation on commodity
  Ethernet.
…ards, and lint fixes

mock_rail: corrupt_stripe/stall/fail/short_read injection; multi_rail: ReadOptions{timeout,expected_generation,max_inflight_bytes} with epoch/version/inflight guards + internal Arc timeout buffer (no UAF); 7 new robustness tests (11/11 pass); Chinese comments translated to English per CLAUDE.md; clippy fixes; .gitignore excludes .workbuddy/.codegraph; Cargo.toml adds twox-hash 1.6 + indexmap std feature. Covers contest robustness matrix, keeps single-rail pass-through and no upstream interface change.
- examples/softroce_dual_rail.rs: from wiring-only to full e2e: PUT via
  gRPC control plane, lookup_object for a fresh descriptor, locally
  derived per-stripe xxh3-64 checksums, single-rail baseline read,
  dual-rail read, byte-exact + checksum verification, and a speedup /
  bottleneck report. Endpoints default to 127.0.0.1:50053/50054 (RDMA
  control channels) to avoid the gRPC :50051 port; GID index is
  configurable (CS_RAIL_GID, default 1 for RXE).
- configs/server-wsl2-softroce.toml: single-machine demo config with
  4 MiB striping threshold/chunk so a 64 MiB object becomes 16 stripes
  (production default of 256 MiB would not stripe it).
Full design doc for the multi-rail parallel read proposal (DaoCloud#33):
background, goals/non-goals, codebase generalization analysis,
architecture, stripe-to-rail scheduling, late-RDMA-WRITE memory-safety
defences, v1 failure semantics matrix, resource governance and
backpressure, compatibility invariants, observability / bottleneck
attribution, the 11-test plan, labelled capacity-model numbers, known
limitations, and reproduction steps.
Explicit coverage of the two remaining proposal-doc items: the rail
tuple (id, local device/port/GID, remote endpoint, QP/CQ, PD+MRs),
static-config discovery with per-rail identity/config/fault injection,
and the QP/MR/WR/CQE/buffer lifecycle contract.
RDMA-CM resolves the device from the destination address; 127.0.0.1
would resolve to loopback where no rxe device exists. Default endpoints
now match setup-wsl2-rxe.sh (rxe0->veth0=192.168.96.110,
rxe1->veth1=192.168.96.111).
Idempotent helper that creates the veth pair, assigns the
192.168.96.110/111 addresses, enables local delivery, and binds rxe0 /
rxe1. Part of the reproduction material for the multi-rail proposal
(requires a CONFIG_RDMA_RXE=y kernel).
With CONFIG_RDMA_RXE=y the driver is built-in and modprobe reports
'not found' even when RXE is fully functional. Fall back to a
create/delete probe link so the check passes for built-in configs.
…ment mapping

Server (tag-15 stripe-subset path):
- Fix boundary check in serve_get_stripes: stripe end is an object-space
  offset and must be compared against descriptor size, not the client
  window (req.max_size). Previously every stripe beyond index 0 missed
  silently when the object exceeded the advertised window.

Client:
- multi_rail: surface per-rail index and underlying error on failure
  instead of a bare 'one or more rails failed'.
- RdmaClient::read_stripes: translate the per-stripe segment table into
  a single segment spanning the whole registered object window, matching
  the server-side map_range_to_segments contract (segments are treated
  as contiguous object-space coverage from 0). Guard the translation
  (segment count, single registration, natural stripe offsets).
- Add BufferView::len()/is_empty() accessors.

Example (softroce_dual_rail):
- Full e2e: PUT 64MiB via gRPC, lookup descriptor, single-rail baseline
  vs dual-rail read, per-stripe xxh3-64 verification plus byte equality,
  speedup report. Result on WSL2 Soft-RoCE (2 veth pairs, rxe0/rxe1):
  73.8 MB/s single-rail -> 217.6 MB/s dual-rail = 2.95x.

Config (server-wsl2-softroce.toml):
- Complete mandatory [gc] fields (GcConfig has no serde defaults).
- Switch io_executor to tier_b (tier_a does not implement
  read_aligned_into_ptr_batch used by the stripe-subset path).

Upstream build fixes under --features rdma:
- rdma.rs test: pass ttl_seconds to build_put_request.
- rdma_bench.rs: add missing Args fields in test initializer.

Verified: lib(rdma) 4/4, multi_rail_mock 11/11.
@fresshman
fresshman marked this pull request as ready for review October 7, 2026 02:42
…o materials

The multi-rail read path registered destination buffers through a cache keyed
by (base_ptr, length). A single client performing consecutive reads into a
per-iteration Vec silently got zeros back: once the first Vec is dropped the
allocator frequently reuses the same virtual address, the cache hits on that
key, and the reused MR still pins the *old* physical pages while the CPU reads
through the *new* mapping — so the NIC's RDMA WRITE lands where the caller
never looks. The trait contract already said "pins the region for the read's
duration", so the caching implementation was the deviation.

Fix (option A+C):
- Default path: RailReader::register now registers afresh per read via
  register_raw_buffer; the owning MR is held in the new `in_flight_mr` field,
  giving it exactly the register -> read_stripes lifetime the trait documents.
- Pooled path: renamed to register_raw_buffer_pooled, documented as requiring
  long-lived buffer-pool semantics, plus invalidate_mr_cache() for callers that
  recycle buffers at a reused address. No public signature changed; the wire
  protocol, disk layout and upper read API are untouched.

Tests / evidence:
- tests/multi_rail_mr_lifetime.rs: consecutive reads with address reuse,
  object switching on one reader, and a compile-time check of the pooled API
  (rdma feature gated).
- Soft-RoCE (real verbs): 20 measurement points across 32/64/128/256/512/1024
  MiB and concurrency 1/2, every one verify_ok=true; e2e dual-rail stays
  verify_ok=true.

New tooling (contest deliverables e/f/g):
- examples/softroce_concurrency.rs: object-size x concurrency sweep reporting
  aggregate goodput, p50/p95 latency and process resources (RSS, registered
  memory, inflight bytes).
- build-and-test.cmd: one-shot build + test + mock-bench driver (no RDMA HW).
- docs/wsl2-softroce-setup.md: WSL2 RXE two-rail bring-up guide.

Docs: design doc gains 8.1 (MR caching vs buffer lifetime case study), the
measured concurrency/resource tables in 12.3 (dual-rail speedup peaks at 3.70x
at 128 MiB, decays past 256 MiB as RXE's single-QP CPU serialization binds),
and a new limitation row in 13.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant