Skip to content

IBSocket: post small sends inline - #430

Open
alxrxs wants to merge 4 commits into
deepseek-ai:mainfrom
alxrxs:ib-inline-small-sends
Open

alxrxs wants to merge 4 commits into
deepseek-ai:mainfrom
alxrxs:ib-inline-small-sends

Conversation

@alxrxs

@alxrxs alxrxs commented Sep 28, 2026 •

Copy link
Copy Markdown

IBSocket::postSend() posts every send buffer by reference and the QP is created with max_inline_data = 0, so for a small RPC request or response the NIC first has to read the message from host memory. In a separate ping-pong microbenchmark that sends the same small RDMA write with and without IBV_SEND_INLINE, inlining cut one-way latency by about 0.43 us on Xeon 8480C hosts with ConnectX-7 and 0.89 us on EPYC 7742 hosts with ConnectX-6.

This requests 236 bytes of inline data in qpCreate(), less if the device refuses (see below), and sets IBV_SEND_INLINE in postSend() when the message fits the granted size. 236 bytes is the largest Send whose inlined WQE fits the 256-byte BlueFlame buffer of current mlx5 NICs. The 236-byte limit applies only on Mellanox/NVIDIA NICs (vendor ID 0x02c9); other devices inline up to the size they grant. The verbs API cannot report the inline limit, so devices whose driver has a fixed one try it first, by vendor ID (Intel 216, 101 or 48 depending on the irdma generation, 101 on the E810 and 48 on the X722; Alibaba eRDMA 96), and a device that refuses a size is asked for 16 bytes less each time until it accepts, down to none. Only EINVAL counts as a refusal; any other failure is reported as before. Each device remembers the size its last QP was created with, so later connections start there instead of stepping down again. With the default max_sge of 16 the send WQE stays the same size. IBSocket already posts its empty liveness-check and close writes with IBV_SEND_INLINE. Nothing on the wire changes, so old and new peers still work together.

I haven't measured 3FS itself. I compiled IBConnect.cc, IBSocket.cc, IBDevice.cc and RDMABuf.cc (clang 17, -Wall -Wextra, no warnings) but could not build the whole tree or run it.

Generated with Claude Code

postSend() posts every send buffer by reference, so the NIC has to read
a small RPC message from host memory before sending it. Request 236
bytes of inline data in qpCreate(), falling back to none if the device
refuses, and set IBV_SEND_INLINE when a send fits the granted size. 236
bytes is the largest Send whose WQE fits the 256-byte BlueFlame buffer
of current mlx5 NICs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@alxrxs
alxrxs force-pushed the ib-inline-small-sends branch from e556ccc to 125bcf1 Compare October 2, 2026 12:59
alxrxs and others added 3 commits October 4, 2026 00:30
The 236-byte limit keeps an inlined Send within the 256-byte BlueFlame
buffer of mlx5 NICs. Apply it only on Mellanox/NVIDIA devices (vendor
ID 0x02c9); on other devices the inline size the QP was granted is the
limit.

If QP creation with 236 bytes of inline data fails, try 64 bytes before
falling back to none. irdma (Intel E810) rejects requests above 101
bytes, so on those NICs nothing was inlined.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
After a refused 236-byte inline request, QP creation was retried with 64
bytes, then with none. Devices take less than 236 by different amounts:
Intel's irdma refuses more than 216, 101 or 48 bytes depending on the
generation (101 on an E810, 48 on an X722), Alibaba's erdma more than
96. With 64 as the only middle step, an E810 or erdma QP got 64 bytes, a
newer irdma QP 64 instead of 216, and an X722 QP none.

The verbs API cannot report the limit, so step down to it. A device
whose kernel driver has a fixed limit, found by the vendor ID IBDevice
already holds, starts there: 216, then 101, then 48 on Intel, 96 on
Alibaba, each capped at 236. Other devices start at 236. Each refused
size is followed by the next one 16 bytes smaller, down to 0, so a
failure unrelated to inline data still ends as before, after a few more
attempts at QP creation. Devices that accept 236 (mlx5, rxe) are
unchanged. libfabric's verbs provider also probes the limit by trial QP
creation (vrb_find_max_inline()).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A refused inline size makes ibv_create_qp() fail with EINVAL. Any other
failure (ENOMEM, for example) is now reported at once with its own
errno, not retried at smaller inline sizes. Each step down is logged at
DBG.

Each IBDevice also remembers the inline size its last QP was created
with, and new QPs on that device start there. A device that refuses the
first size (an Intel device limited to 48 bytes goes 216, 101, 48) then
steps down once rather than on every connection. If the remembered size
is refused, the step-down continues from it. The size is kept per device
because a host can have NICs from different vendors.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@alxrxs
alxrxs marked this pull request as ready for review October 6, 2026 13:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant