Skip to content

Create the header protection cryptor once per key - #198

Open
rnro wants to merge 1 commit into
apple:mainfrom
rnro:alloc/header-protection-cryptor
Open

rnro wants to merge 1 commit into
apple:mainfrom
rnro:alloc/header-protection-cryptor

Conversation

@rnro

@rnro rnro commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

SecFramerAESGCM.headerProtection computed every mask with one-shot CCCrypt,
which creates and releases a cryptor per call. That is three allocations each
time a packet is sealed or opened, six per packet for a connection that both
sends and receives, and the largest single source of allocations on the QUIC
data path. Header protection uses AES-ECB, which carries no state from one
block to the next, so SecFramerKeys now creates a HeaderProtectionCryptor
when the keys are installed and every packet reuses it.

Creating the cryptors costs a connection pair 48 allocations up front, against
the 39 its handshake packets used to spend, so a connection that does nothing
but handshake pays about 9 more.

Measured on my Mac against main with the package's QUIC benchmark tools.
Allocation counts come from full malloc stack logging, which records every
allocation:

QUICTransfer -size 1200, per message      24.37 -> 18.09
QUICTransfer, per 500 KB transfer       6,187.6 -> 3,597.6
QUICHandshake, per connection           1,947.4 -> 1,955.1
QUICStreamLoad, per stream                111.8 -> 100.2

Wall-clock time comes from running main, this change and the other changes
measured alongside it in a rotating order for 9 rounds, and comparing each run
with main's in the same round. Changes moved paths they do not touch by up to
about 1.3%, so differences that size count as noise. The 500 KB transfers took
2.1% less time and the stream load 4.8% less, faster in all 9 rounds each; the
1,200-byte messages and the handshakes did not change beyond noise.

`SecFramerAESGCM.headerProtection` computed every mask with one-shot `CCCrypt`,
which creates and releases a cryptor per call. That is three allocations each
time a packet is sealed or opened, six per packet for a connection that both
sends and receives, and the largest single source of allocations on the QUIC
data path. Header protection uses AES-ECB, which carries no state from one
block to the next, so `SecFramerKeys` now creates a `HeaderProtectionCryptor`
when the keys are installed and every packet reuses it.

Creating the cryptors costs a connection pair 48 allocations up front, against
the 39 its handshake packets used to spend, so a connection that does nothing
but handshake pays about 9 more.

Measured on my Mac against `main` with the package's QUIC benchmark tools.
Allocation counts come from full malloc stack logging, which records every
allocation:

  QUICTransfer -size 1200, per message      24.37 -> 18.09
  QUICTransfer, per 500 KB transfer       6,187.6 -> 3,597.6
  QUICHandshake, per connection           1,947.4 -> 1,955.1
  QUICStreamLoad, per stream                111.8 -> 100.2

Wall-clock time comes from running `main`, this change and the other changes
measured alongside it in a rotating order for 9 rounds, and comparing each run
with `main`'s in the same round. Changes moved paths they do not touch by up to
about 1.3%, so differences that size count as noise. The 500 KB transfers took
2.1% less time and the stream load 4.8% less, faster in all 9 rounds each; the
1,200-byte messages and the handshakes did not change beyond noise.

@agnosticdev agnosticdev left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you Rick!

packetBuffer.baseAddress! + packet.sampleRange.lowerBound,
into: maskBuffer.baseAddress!
)
}

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep, for Darwin users this is a nice win!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants