Conversation
`SecFramerAESGCM.headerProtection` computed every mask with one-shot `CCCrypt`, which creates and releases a cryptor per call. That is three allocations each time a packet is sealed or opened, six per packet for a connection that both sends and receives, and the largest single source of allocations on the QUIC data path. Header protection uses AES-ECB, which carries no state from one block to the next, so `SecFramerKeys` now creates a `HeaderProtectionCryptor` when the keys are installed and every packet reuses it. Creating the cryptors costs a connection pair 48 allocations up front, against the 39 its handshake packets used to spend, so a connection that does nothing but handshake pays about 9 more. Measured on my Mac against `main` with the package's QUIC benchmark tools. Allocation counts come from full malloc stack logging, which records every allocation: QUICTransfer -size 1200, per message 24.37 -> 18.09 QUICTransfer, per 500 KB transfer 6,187.6 -> 3,597.6 QUICHandshake, per connection 1,947.4 -> 1,955.1 QUICStreamLoad, per stream 111.8 -> 100.2 Wall-clock time comes from running `main`, this change and the other changes measured alongside it in a rotating order for 9 rounds, and comparing each run with `main`'s in the same round. Changes moved paths they do not touch by up to about 1.3%, so differences that size count as noise. The 500 KB transfers took 2.1% less time and the stream load 4.8% less, faster in all 9 rounds each; the 1,200-byte messages and the handshakes did not change beyond noise.
rnro
requested review from
agnosticdev,
kkuk24,
rpaulo and
tfpauly
as code owners
October 1, 2026 19:54
agnosticdev
approved these changes
Oct 2, 2026
| packetBuffer.baseAddress! + packet.sampleRange.lowerBound, | ||
| into: maskBuffer.baseAddress! | ||
| ) | ||
| } |
Collaborator
There was a problem hiding this comment.
Yep, for Darwin users this is a nice win!
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
SecFramerAESGCM.headerProtectioncomputed every mask with one-shotCCCrypt,which creates and releases a cryptor per call. That is three allocations each
time a packet is sealed or opened, six per packet for a connection that both
sends and receives, and the largest single source of allocations on the QUIC
data path. Header protection uses AES-ECB, which carries no state from one
block to the next, so
SecFramerKeysnow creates aHeaderProtectionCryptorwhen the keys are installed and every packet reuses it.
Creating the cryptors costs a connection pair 48 allocations up front, against
the 39 its handshake packets used to spend, so a connection that does nothing
but handshake pays about 9 more.
Measured on my Mac against
mainwith the package's QUIC benchmark tools.Allocation counts come from full malloc stack logging, which records every
allocation:
Wall-clock time comes from running
main, this change and the other changesmeasured alongside it in a rotating order for 9 rounds, and comparing each run
with
main's in the same round. Changes moved paths they do not touch by up toabout 1.3%, so differences that size count as noise. The 500 KB transfers took
2.1% less time and the stream load 4.8% less, faster in all 9 rounds each; the
1,200-byte messages and the handshakes did not change beyond noise.