Skip to content

Add independent multi-rail RDMA descriptor reads - #32

Draft
yanchaomei wants to merge 20 commits into
DaoCloud:mainfrom
yanchaomei:feat/independent-multirail-read
Draft

yanchaomei wants to merge 20 commits into
DaoCloud:mainfrom
yanchaomei:feat/independent-multirail-read

Conversation

@yanchaomei

@yanchaomei yanchaomei commented Oct 1, 2026 •

Copy link
Copy Markdown

Summary

LookupObject → validate descriptor/placement and stripe coverage
  → optionally discover fabric/listener capabilities and match configured local devices
  → schedule stripes across independent RDMA paths
  → one QP/CQ/MR and compact receive buffer per rail
  → join, verify, reassemble, recheck object version → publish atomically

Independent implementation from DaoCloud main at b5c6451; PR #28 is not its base. Current head is c1c1de1. The on-disk stripe layout and existing read APIs are unchanged. The opt-in Rust SDK entry point is KvClient::read_multi_rail_into; one configured rail remains usable. LookupObject may include additive, ephemeral fabric/listener capabilities; RailReader::discover_from_placement matches them to trusted local paths. Manual RailRoute remains available for older servers, and old clients ignore the new protobuf field. Global/per-rail task, staging, registered-memory and in-flight budgets are enforced, with a fixed 32-route/QP cap per read. Linux additionally caps registration at 80% of a finite process RLIMIT_MEMLOCK.

Review guides: design, lifecycle, failure and upgrade semantics, rail discovery contract and compatibility, and five-act competition story with evidence boundaries.

The server honors tag-15 scatter destinations in slab and registered-buffer fallback paths. Tier A now emits pointer-stream completions; missing bytes cannot be acknowledged as success. Uncertain CQ completion retains source memory and retires the server QP/CQ connection. This now also covers the older complete-object GET's cache-hit, cache-miss and per-chunk fallback sources; a cache-miss CQ error cannot fall back on the uncertain QP/CQ. A client GET control error likewise destroys its QP before the caller can release its MR. A test-only, bounded delay can target one listener, and stripe-subset/legacy GET CQ waiting has a configurable 100–30,000 ms limit. The ARM64 O_DIRECT correction fixes the baseline flag selecting O_DIRECTORY.

Evidence

  • Red → green safety regression: A real RXE test reproduced a pooled client's stale reply being accepted after GET timeout. The fixed client retires its connection; the fixed server also closes uncertain-CQ connections. Expected red, green and server logs.
  • Legacy GET safety regression: On isolated RXE, the client completed its QP handshake, then applied a 1 ms GET deadline; after failure it rejected QP reuse and a recycled destination retained its sentinel for six seconds. The server logged a 2 s CQ timeout with 0/64 completions and retired the connection without entering per-chunk fallback. Client and server receipt hashes.
  • Build/tests: On final x86_64 code, server RDMA release library 102 passed, client RDMA library 44 passed, 1 intentional Mock benchmark ignored, CLI unit tests and strict client Clippy passed; rdma,io-uring,metrics and no-RDMA builds passed. The isolated Redis/two-node E2E suite passed 4/4. Earlier ARM64 checks included 70 default server tests, 96 feature tests and Python 116 passed/11 skipped; the later CQ/discovery changes were verified on x86_64.
  • Physical HCA, one rail: SKV node1→node2 ConnectX-6 Dx restored the same checksummed 64 MiB/16-stripe object in slab and fallback modes. Dead listener, in-flight cancellation, checksum corruption/restore, timeout followed by late WRITE, buffer reuse and recovery passed. Five reads returned 335,544,320 B payload; server HCA transmit and client HCA receive counters increased by 354,554,880 and 355,865,600 counter-derived bytes. Raw samples and topology. Both stripe directories were on one SATA SSD; this is not multi-disk or dual-HCA evidence.
  • Two Soft-RoCE rails, real Verbs: Two isolated Ubuntu KVM guests use two separate virtio/tap/bridge networks, four RXE devices and two server listeners. One Worker reconstructed the same object with 32 MiB per rail. Independent Rail-1 shutdown/recovery, unreachable listener, old generation, corruption of a Rail-1 stripe, dual in-flight cancellation, partial completion then late WRITE, slab-disabled fallback and connection retirement were exercised. Topology/counters, test manifest, VM deployment scripts.
  • Advertised rail discovery and rolling upgrades: With only fabric-a/rxe_c0 and fabric-b/rxe_c1 configured locally, one Worker discovered both remote listeners through Placement and restored the same 64 MiB object at 32 MiB per rail. An actual second-link shutdown left the first rail successful, the second failed, and the caller buffer unchanged; recovery passed. Old client/new server, new client/old server manual mode, and explicit discovery rejection on the old server all passed. Fourteen full raw receipts and source hashes. These are functional single reads, not a new paired performance matrix.
  • Latest submission brief and four-minute demonstration: The 2026-10-03 release contains the updated three-page Chinese PDF, verified ZIP, full discovery/upgrade raw JSON and the original four-minute video. The video is an edited, narrated rendering of live manual-route Soft-RoCE commands from commit c439c59; the later discovery feature is verified by separate real-Verbs transcripts and is not represented as a scene in that video. The ZIP contains only allowlisted public material and per-file checksums; it is not a competition submission receipt. The original baseline release remains available unchanged.
  • Fair software-RoCE comparison: Same object/layout/guests/server/concurrency within each cell, one warmup, 3 trials×5 reads with rail-order alternation. Median one→two-rail latency: 64 MiB 352.8→329.4 ms (1.071×), 128 MiB 703.3→647.8 ms (1.086×), 256 MiB 1424.5→1279.4 ms (1.113×). 90 request samples, summary, analysis. CPU/read and RSS rose. With four concurrent 64 MiB reads in one process/shared RailReader, aggregate gain fell to 1.8%, retained as a negative result. Concurrent raw samples. Finite guest memlock rejects the fourth 128 MiB read before dispatch; raising only that isolated guest's limit allows all four.

Limits: RXE/QEMU, both guests and one virtual disk share a physical host. These measurements prove real Verbs paths and software scaling, not HCA offload, separate PCIe/NUMA rails, independent NVMe bandwidth or physical multi-NIC aggregation. All available SKV nodes had one online HCA port; host RXE modules conflict with installed OFED. This PR currently has no GitHub Actions checks or maintainer review, and the competition receipt remains pending. The original four-minute video is an edited terminal record of the baseline manual-route implementation; the later discovery change is supported by separate full raw Verbs receipts.

Merge Danger

Door: Two-way. The feature is opt-in; stored object layout is unchanged.

Blast Radius: RDMA data path and an additive PlacementDescriptor protobuf field. Tier A pointer streaming, CQ deadlines, QP retirement and fallback source lifetimes touch existing descriptor reads. Keep this PR in draft for maintainer review and CI approval; physical two-HCA aggregation remains to be measured when two live ports are available.

Build a rail planner and bounded read state machine directly on ContextStore main. Keep each rail's incoming stripes in a compact registered buffer using the existing tag-15 SGE protocol, then validate stripe coverage, acknowledgements, checksums, and post-read object identity before publishing bytes.

Support explicit per-node listener routes, weighted byte scheduling, runtime enablement and cooldown, cancellation, resource budgets, per-rail metrics and sysfs topology. Add a labeled benchmark CLI, Mock fault and overhead checks, hardware-gated E2E tests, CI coverage, and compatibility documentation.
Use one architecture-aware flag for Tier A and Tier B. The previous x86 literal maps to O_DIRECTORY on arm64 and made striped reads fail with ENOTDIR; regression tests compare the chosen value with native libc on each Linux build.
Retain active-read and staging reservations across the post-transfer metadata lookup, serialize cancellation with final publication, and route slab-fallback writes through tag-15 destination segments. Keep uncertain server WRITE sources registered until their QP is destroyed. Add focused regression tests and document checksum configuration limits.
Add a parameterized collector for paired one- and two-rail Mock reads, and commit the raw 180 latency samples, run summaries, and environment manifest. Label the data as software-only so it cannot be mistaken for HCA throughput.
Reserve per-rail task slots and bytes atomically across concurrent requests, release the transfer credits when workers finish, and retain the final-object budget through publication. Add red-green regression tests for both per-rail limits.
Rerun the paired 64, 256, and 512 MiB software-only matrix on the exact per-rail budget implementation. Keep the 180 raw samples and environment manifest with LF CSV output for clean review.
Implement Tier A pointer streaming for slab-backed striped reads, reject missing completion bytes, and retain uncertain slab allocations through QP teardown. Count failed CQEs when harvesting the queue to avoid an unnecessary final poll timeout. Add isolated fault-injection hooks and hardware-gated tests for single-rail success, disconnect, cancellation, checksum corruption, and late WRITEs, plus a validation config.
Record exact 64 MiB sample latencies for slab and fallback modes, HCA port-counter snapshots, fault-injection outcomes, and the shared-SATA storage boundary. Keep the hardware data separate from the software-only two-rail Mock matrix.
Run the same 64/256/512 MiB one- and two-rail Mock matrix on SKV x86_64, preserving 180 per-read samples, run summaries, and an environment manifest separately from physical HCA measurements.
Add a fault hook scoped to one RDMA listener and a dual-rail late-WRITE regression using real Verbs. Make stripe-subset CQ polling deadline configurable, preserving uncertain source memory until QP teardown. Add a paired real-Verbs benchmark collector and document Soft-RoCE limitations.
Reverse one- and two-rail execution order on even trials to reduce a systematic warm-cache or host-load bias in the reported software RoCE matrix.
Add an opt-in CLI concurrency mode that keeps full reads in one process, joins every request, checks object hashes, and reports batch throughput, CPU, RSS, and per-rail bytes. Add a paired real-Verbs collector for concurrency cells above one.
Reserve 20 percent of a finite Linux RLIMIT_MEMLOCK for runtime headroom and reject excess object reads before QP dispatch. Verify that a 128 MiB by four request on the isolated RXE guest now reports a rail resource limit instead of a late ibv_reg_mr ENOMEM; add real two-rail checksum and cancellation tests.
Stop the server handler after a stripe-subset CQ timeout or failed completion so late CQEs cannot satisfy a later request on the same QP. On client GET control errors, destroy the QP before the caller can deregister or reuse its MR, rejecting later operations on that connection. Reproduce stale response reuse on old binaries and verify the fixed behavior over two RXE rails.
Commit paired 64/128/256 MiB raw Verbs samples, concurrent 2/4 request batches, topology and per-rail virtual-link counters, the bounded memlock/fault test receipts, an analysis script, and isolated KVM deployment scripts. Explicitly label all results as software RXE on one host and one virtual disk.
The red RXE test exercised an old client against a server with CQ retirement; it showed that client-side control timeout also requires QP retirement. Correct the public test receipt wording without changing the recorded output or hashes.
Pin slab and per-chunk registered sources until QP destruction on post or CQ errors, preventing cache-miss fallback from reusing an uncertain queue. Add a post-handshake GET timeout regression verified on two isolated RXE rails and record the raw receipt hashes.
Add a reviewable lifecycle and failure design note plus the five-act story, and link both from the main README. Keep Soft-RoCE, Mock, and physical HCA claims separate.
Add optional fabric/listener capabilities to LookupObject without changing stored stripe layout. Match them to explicit local Verbs paths, reject stale or ambiguous ownership, bound QP tasks, preserve manual routes and old clients, and verify two RXE rails plus rolling-upgrade behavior.
Publish full command outputs, source-file hashes, and raw receipt SHA-256 values for discovered dual read, second-link failure/recovery, old-client/new-server and new-client/old-server compatibility, release builds, and two-node regression.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant