Repository navigation
Conversation
postSend() posts every send buffer by reference, so the NIC has to read a small RPC message from host memory before sending it. Request 236 bytes of inline data in qpCreate(), falling back to none if the device refuses, and set IBV_SEND_INLINE when a send fits the granted size. 236 bytes is the largest Send whose WQE fits the 256-byte BlueFlame buffer of current mlx5 NICs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
alxrxs
force-pushed
the
ib-inline-small-sends
branch
from
October 2, 2026 12:59
e556ccc to
125bcf1
Compare
The 236-byte limit keeps an inlined Send within the 256-byte BlueFlame buffer of mlx5 NICs. Apply it only on Mellanox/NVIDIA devices (vendor ID 0x02c9); on other devices the inline size the QP was granted is the limit. If QP creation with 236 bytes of inline data fails, try 64 bytes before falling back to none. irdma (Intel E810) rejects requests above 101 bytes, so on those NICs nothing was inlined. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
After a refused 236-byte inline request, QP creation was retried with 64 bytes, then with none. Devices take less than 236 by different amounts: Intel's irdma refuses more than 216, 101 or 48 bytes depending on the generation (101 on an E810, 48 on an X722), Alibaba's erdma more than 96. With 64 as the only middle step, an E810 or erdma QP got 64 bytes, a newer irdma QP 64 instead of 216, and an X722 QP none. The verbs API cannot report the limit, so step down to it. A device whose kernel driver has a fixed limit, found by the vendor ID IBDevice already holds, starts there: 216, then 101, then 48 on Intel, 96 on Alibaba, each capped at 236. Other devices start at 236. Each refused size is followed by the next one 16 bytes smaller, down to 0, so a failure unrelated to inline data still ends as before, after a few more attempts at QP creation. Devices that accept 236 (mlx5, rxe) are unchanged. libfabric's verbs provider also probes the limit by trial QP creation (vrb_find_max_inline()). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A refused inline size makes ibv_create_qp() fail with EINVAL. Any other failure (ENOMEM, for example) is now reported at once with its own errno, not retried at smaller inline sizes. Each step down is logged at DBG. Each IBDevice also remembers the inline size its last QP was created with, and new QPs on that device start there. A device that refuses the first size (an Intel device limited to 48 bytes goes 216, 101, 48) then steps down once rather than on every connection. If the remembered size is refused, the step-down continues from it. The size is kept per device because a host can have NICs from different vendors. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
alxrxs
marked this pull request as ready for review
October 6, 2026 13:39
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
IBSocket::postSend() posts every send buffer by reference and the QP is created with max_inline_data = 0, so for a small RPC request or response the NIC first has to read the message from host memory. In a separate ping-pong microbenchmark that sends the same small RDMA write with and without IBV_SEND_INLINE, inlining cut one-way latency by about 0.43 us on Xeon 8480C hosts with ConnectX-7 and 0.89 us on EPYC 7742 hosts with ConnectX-6.
This requests 236 bytes of inline data in qpCreate(), less if the device refuses (see below), and sets IBV_SEND_INLINE in postSend() when the message fits the granted size. 236 bytes is the largest Send whose inlined WQE fits the 256-byte BlueFlame buffer of current mlx5 NICs. The 236-byte limit applies only on Mellanox/NVIDIA NICs (vendor ID 0x02c9); other devices inline up to the size they grant. The verbs API cannot report the inline limit, so devices whose driver has a fixed one try it first, by vendor ID (Intel 216, 101 or 48 depending on the irdma generation, 101 on the E810 and 48 on the X722; Alibaba eRDMA 96), and a device that refuses a size is asked for 16 bytes less each time until it accepts, down to none. Only EINVAL counts as a refusal; any other failure is reported as before. Each device remembers the size its last QP was created with, so later connections start there instead of stepping down again. With the default max_sge of 16 the send WQE stays the same size. IBSocket already posts its empty liveness-check and close writes with IBV_SEND_INLINE. Nothing on the wire changes, so old and new peers still work together.
I haven't measured 3FS itself. I compiled IBConnect.cc, IBSocket.cc, IBDevice.cc and RDMABuf.cc (clang 17, -Wall -Wextra, no warnings) but could not build the whole tree or run it.
Generated with Claude Code