Skip to content

Add a singular accessor for acquiring one datagram frame - #202

Open
rnro wants to merge 1 commit into
apple:mainfrom
rnro:alloc/singular-datagram-accessor
Open

rnro wants to merge 1 commit into
apple:mainfrom
rnro:alloc/singular-datagram-accessor

Conversation

@rnro

@rnro rnro commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

QUIC builds one packet at a time when its prefetched batch runs dry, and the
only way down the stack was getDatagramsToSend, which built a FrameArray to
carry a single frame. Added getDatagramToSend beside it through the datagram
linkage, the base linkage's dispatch and the path-labelled QUIC helper, with
defaults that forward to a batch of one so existing linkages and protocols are
unaffected. UDP, IP and the in-process bridge override it, and
buildSinglePacketForKeyState now takes that route. The batched path is
untouched. The bridge's two paths share the check that withholds datagrams when
a test blocks packet generation.

Measured on my Mac against main with the package's QUIC benchmark tools.
Allocation counts come from full malloc stack logging, which records every
allocation:

QUICTransfer -size 1200, per message      24.37 -> 24.35
QUICTransfer, per 500 KB transfer       6,187.6 -> 6,179.5
QUICHandshake, per connection           1,947.4 -> 1,945.8
QUICStreamLoad, per stream                111.8 -> 109.9

The package benchmarks build most packets in batches and run on a fast
allocator, so they show little: only the stream load, which builds packets one
at a time, moves much. The saving is larger where a bottom protocol hands out
its own buffers one frame at a time, where traffic is dominated by packets built
singly, such as acknowledgment-driven or bursty sends, and where heap allocation
is costly.

Wall-clock time comes from running the two builds alternately for 9 rounds and
comparing the two runs within each round. Nothing changed beyond noise.

QUIC builds one packet at a time when its prefetched batch runs dry, and the
only way down the stack was `getDatagramsToSend`, which built a `FrameArray` to
carry a single frame. Added `getDatagramToSend` beside it through the datagram
linkage, the base linkage's dispatch and the path-labelled QUIC helper, with
defaults that forward to a batch of one so existing linkages and protocols are
unaffected. UDP, IP and the in-process bridge override it, and
`buildSinglePacketForKeyState` now takes that route. The batched path is
untouched. The bridge's two paths share the check that withholds datagrams when
a test blocks packet generation.

Measured on my Mac against `main` with the package's QUIC benchmark tools.
Allocation counts come from full malloc stack logging, which records every
allocation:

  QUICTransfer -size 1200, per message      24.37 -> 24.35
  QUICTransfer, per 500 KB transfer       6,187.6 -> 6,179.5
  QUICHandshake, per connection           1,947.4 -> 1,945.8
  QUICStreamLoad, per stream                111.8 -> 109.9

The package benchmarks build most packets in batches and run on a fast
allocator, so they show little: only the stream load, which builds packets one
at a time, moves much. The saving is larger where a bottom protocol hands out
its own buffers one frame at a time, where traffic is dominated by packets built
singly, such as acknowledgment-driven or bursty sends, and where heap allocation
is costly.

Wall-clock time comes from running the two builds alternately for 9 rounds and
comparing the two runs within each round. Nothing changed beyond noise.

@tfpauly tfpauly left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have some overall concerns with this approach. I get that for perf, we want to optimize a common case for single frames, but this approach is inconsistently applied (doesn't handle sending datagrams, or sending or receiving stream data), and also creates an explosion of specific functions. I'd rather take an approach of making the backing for FrameArrays cheaper for the "small" case

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants