Skip to content

Make FramePadding.countLeadingZeroBytes faster - #190

Open
MahdiBM wants to merge 3 commits into
apple:mainfrom
MahdiBM:mmbm-count-zeros
Open

MahdiBM wants to merge 3 commits into
apple:mainfrom
MahdiBM:mmbm-count-zeros

Conversation

@MahdiBM

@MahdiBM MahdiBM commented Oct 1, 2026 •

Copy link
Copy Markdown

#126 caught my eye. It can be made faster.

Result on M6 processors (in ns):

byte-count old new speedup
8 1.3 1.5 0.9x
32 2.1 1.5 1.5x
48 2.6 1.9 1.4x
96 4.1 1.7 2.5x
128 5.2 1.9 2.8x
256 9.8 2.9 3.3x
512 18.9 4.8 3.9x
768 28.6 7.3 3.9x
1200 46.2 10.3 4.5x
random 8–1200 23.4 7.0 3.3x

private static func countLeadingZeroBytes(_ bytes: RawSpan) -> Int {
let count = bytes.byteCount
let wordSize = MemoryLayout<UInt64>.size
let vectorSize = 32

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried 16/64 as well, 32 seems to be the sweet spot.

var orResult: UInt8 = 0
/// This loop is auto-vectorized.
for idx in 0..<vectorSize {
orResult |= bytes.unsafeLoadUnaligned(fromUncheckedByteOffset: offset &+ idx, as: UInt8.self)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(Doesn't really matter if we load bytes or words in auto-vectorized code.)

@tfpauly tfpauly added the 🔨 semver/patch No public API change. label Oct 1, 2026
@MahdiBM

MahdiBM commented Oct 1, 2026

Copy link
Copy Markdown
Author

(Test failures look to be unrelated to this PR)

@rpaulo

rpaulo commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Thanks for this. Can you change the tests to test > 32 byte buffers?

Comment thread Sources/SwiftNetwork/QUIC/QUICFrame.swift Outdated
Comment thread Sources/SwiftNetwork/QUIC/QUICFrame.swift Outdated
/// 1. Speed through the zeros via vectorization.
while offset &+ vectorSize <= count {
var orResult: UInt8 = 0
/// This loop is auto-vectorized.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing really guarantees autovec. We could unroll it and get the same perf benefit without relying on autovec:

while offset &+ vectorSize <= count {
    let a = bytes.unsafeLoadUnaligned(fromUncheckedByteOffset: offset,      as: UInt64.self)
    let b = bytes.unsafeLoadUnaligned(fromUncheckedByteOffset: offset &+ 8,  as: UInt64.self)
    let c = bytes.unsafeLoadUnaligned(fromUncheckedByteOffset: offset &+ 16, as: UInt64.self)
    let d = bytes.unsafeLoadUnaligned(fromUncheckedByteOffset: offset &+ 24, as: UInt64.self)
    if (a | b | c | d) != 0 {
        break 
    }
    offset &+= 32
}

@MahdiBM MahdiBM Oct 1, 2026 •

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I disagree. Swift and LLVM are pretty consistent in auto-vectorizing code. They can auto-vectorize much more complicated loops as well. There are some general conditions and all but the mechanism is consistent and in this specific case it's pretty much the simplest case. Generally speaking some of the conditions are:

  • The loop is branchless.
    • e.g. no byte-load precondition that can't be optimized out.
    • And of course, only minimal condition blocks that are reduced to e.g. conditional-select (e.g. CSEL) instructions.
  • The code is inlined (no external subroutine calls).
  • The length of the loop is known (not necessarily compile-time known).
  • There are no data dependencies to a previous iteration of the loop.
  • For LLVM specifically, It either never accepts table lookups or hardly ever does (I've never seen it to, and I haven't tried reading through LLVM code; I might at some point).
    • Generally table lookups are limited even when manually writing SIMD code.

In this specific code it's just 1 OR + byte-load so you don't really need to go verify all those statements either. Just that I'm trying to convince you that auto-vectorization is more trustable and consistent than what I feel like you think.

@MahdiBM MahdiBM Oct 1, 2026 •

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also the code you provided is auto-vectorized as well and that's why the 2 of the codes don't make much of a difference. If the code was scalar it would perform worse than the auto-vectorized code (General rule of thumb, but not true 100% of the time specially if there are a low amount of loop iterations, say 8/16 for example).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fair enough

@rpaulo rpaulo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved pending test modifications

@agnosticdev agnosticdev left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, approved pending testing.

@MahdiBM

MahdiBM commented Oct 2, 2026

Copy link
Copy Markdown
Author

@rpaulo @agnosticdev please re-run CI (See 7301108).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

🔨 semver/patch No public API change.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants