Skip to content

Seal and open packets through CryptoKit's span API where the OS has it - #197

Open
rnro wants to merge 1 commit into
apple:mainfrom
rnro:alloc/gcm-in-place
Open

rnro wants to merge 1 commit into
apple:mainfrom
rnro:alloc/gcm-in-place

Conversation

@rnro

@rnro rnro commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

The package supports OS 26, where CryptoKit lacks span-based seal(inPlace:)
and open(inPlace:), so it ships copying wrappers with the same names. Because
they share CryptoKit's signatures, they shadow its API throughout the module,
and a client supporting OS 26 takes the copying path even on OS 27.
DISABLE_SHIM_CRYPTO_SPAN_APIS only avoids that by requiring an OS 27
deployment target. That costs three allocations per packet, one to seal and two
to open.

Added sealInPlace and openInPlace entry points that call CryptoKit's span
API when the running OS has it, and otherwise fall back to the wrappers, now
named sealInPlaceThroughSealedBox and openInPlaceThroughSealedBox after how
they work. With DISABLE_SHIM_CRYPTO_SPAN_APIS they call CryptoKit directly, as
before.

Measured on my Mac against main with the package's QUIC benchmark tools.
Allocation counts come from full malloc stack logging, which records every
allocation:

QUICTransfer -size 1200, per message      24.37 -> 19.17
QUICTransfer, per 500 KB transfer       6,187.6 -> 4,027.2
QUICHandshake, per connection           1,947.4 -> 1,919.3
QUICStreamLoad, per stream                111.8 -> 102.3

Wall-clock time comes from running main, this change and the other changes
measured alongside it in a rotating order for 9 rounds, and comparing each run
with main's in the same round. Changes moved paths they do not touch by up to
about 1.3%, so differences that size count as noise. The 500 KB transfers took
8.3% less time, the handshakes 1.8% less and the stream load 16.2% less, faster
in all 9 rounds each; the 1,200-byte messages did not change beyond noise.

The package supports OS 26, where CryptoKit lacks span-based `seal(inPlace:)`
and `open(inPlace:)`, so it ships copying wrappers with the same names. Because
they share CryptoKit's signatures, they shadow its API throughout the module,
and a client supporting OS 26 takes the copying path even on OS 27.
`DISABLE_SHIM_CRYPTO_SPAN_APIS` only avoids that by requiring an OS 27
deployment target. That costs three allocations per packet, one to seal and two
to open.

Added `sealInPlace` and `openInPlace` entry points that call CryptoKit's span
API when the running OS has it, and otherwise fall back to the wrappers, now
named `sealInPlaceThroughSealedBox` and `openInPlaceThroughSealedBox` after how
they work. With `DISABLE_SHIM_CRYPTO_SPAN_APIS` they call CryptoKit directly, as
before.

Measured on my Mac against `main` with the package's QUIC benchmark tools.
Allocation counts come from full malloc stack logging, which records every
allocation:

  QUICTransfer -size 1200, per message      24.37 -> 19.17
  QUICTransfer, per 500 KB transfer       6,187.6 -> 4,027.2
  QUICHandshake, per connection           1,947.4 -> 1,919.3
  QUICStreamLoad, per stream                111.8 -> 102.3

Wall-clock time comes from running `main`, this change and the other changes
measured alongside it in a rotating order for 9 rounds, and comparing each run
with `main`'s in the same round. Changes moved paths they do not touch by up to
about 1.3%, so differences that size count as noise. The 500 KB transfers took
8.3% less time, the handshakes 1.8% less and the stream load 16.2% less, faster
in all 9 rounds each; the 1,200-byte messages did not change beyond noise.

@agnosticdev agnosticdev left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So you're saying that passing through the wrappers was the part that is taking extra allocations, not actually when DISABLE_SHIM_CRYPTO_SPAN_APIS is enabled, correct?

@rnro

rnro commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

So you're saying that passing through the wrappers was the part that is taking extra allocations, not actually when DISABLE_SHIM_CRYPTO_SPAN_APIS is enabled, correct?

Yes, that's correct. This change means that even if DISABLE_SHIM_CRYPTO_SPAN_APIS isn't enabled, but the fast path is available, it's taken.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants