Conversation
QUIC builds one packet at a time when its prefetched batch runs dry, and the only way down the stack was `getDatagramsToSend`, which built a `FrameArray` to carry a single frame. Added `getDatagramToSend` beside it through the datagram linkage, the base linkage's dispatch and the path-labelled QUIC helper, with defaults that forward to a batch of one so existing linkages and protocols are unaffected. UDP, IP and the in-process bridge override it, and `buildSinglePacketForKeyState` now takes that route. The batched path is untouched. The bridge's two paths share the check that withholds datagrams when a test blocks packet generation. Measured on my Mac against `main` with the package's QUIC benchmark tools. Allocation counts come from full malloc stack logging, which records every allocation: QUICTransfer -size 1200, per message 24.37 -> 24.35 QUICTransfer, per 500 KB transfer 6,187.6 -> 6,179.5 QUICHandshake, per connection 1,947.4 -> 1,945.8 QUICStreamLoad, per stream 111.8 -> 109.9 The package benchmarks build most packets in batches and run on a fast allocator, so they show little: only the stream load, which builds packets one at a time, moves much. The saving is larger where a bottom protocol hands out its own buffers one frame at a time, where traffic is dominated by packets built singly, such as acknowledgment-driven or bursty sends, and where heap allocation is costly. Wall-clock time comes from running the two builds alternately for 9 rounds and comparing the two runs within each round. Nothing changed beyond noise.
rnro
requested review from
agnosticdev,
ekinnear,
kkuk24,
rpaulo and
tfpauly
as code owners
October 1, 2026 19:54
tfpauly
reviewed
Oct 1, 2026
tfpauly
left a comment
Collaborator
There was a problem hiding this comment.
I have some overall concerns with this approach. I get that for perf, we want to optimize a common case for single frames, but this approach is inconsistently applied (doesn't handle sending datagrams, or sending or receiving stream data), and also creates an explosion of specific functions. I'd rather take an approach of making the backing for FrameArrays cheaper for the "small" case
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
QUIC builds one packet at a time when its prefetched batch runs dry, and the
only way down the stack was
getDatagramsToSend, which built aFrameArraytocarry a single frame. Added
getDatagramToSendbeside it through the datagramlinkage, the base linkage's dispatch and the path-labelled QUIC helper, with
defaults that forward to a batch of one so existing linkages and protocols are
unaffected. UDP, IP and the in-process bridge override it, and
buildSinglePacketForKeyStatenow takes that route. The batched path isuntouched. The bridge's two paths share the check that withholds datagrams when
a test blocks packet generation.
Measured on my Mac against
mainwith the package's QUIC benchmark tools.Allocation counts come from full malloc stack logging, which records every
allocation:
The package benchmarks build most packets in batches and run on a fast
allocator, so they show little: only the stream load, which builds packets one
at a time, moves much. The saving is larger where a bottom protocol hands out
its own buffers one frame at a time, where traffic is dominated by packets built
singly, such as acknowledgment-driven or bursty sends, and where heap allocation
is costly.
Wall-clock time comes from running the two builds alternately for 9 rounds and
comparing the two runs within each round. Nothing changed beyond noise.