Conversation
| private static func countLeadingZeroBytes(_ bytes: RawSpan) -> Int { | ||
| let count = bytes.byteCount | ||
| let wordSize = MemoryLayout<UInt64>.size | ||
| let vectorSize = 32 |
There was a problem hiding this comment.
I tried 16/64 as well, 32 seems to be the sweet spot.
| var orResult: UInt8 = 0 | ||
| /// This loop is auto-vectorized. | ||
| for idx in 0..<vectorSize { | ||
| orResult |= bytes.unsafeLoadUnaligned(fromUncheckedByteOffset: offset &+ idx, as: UInt8.self) |
There was a problem hiding this comment.
(Doesn't really matter if we load bytes or words in auto-vectorized code.)
|
(Test failures look to be unrelated to this PR) |
|
Thanks for this. Can you change the tests to test > 32 byte buffers? |
| /// 1. Speed through the zeros via vectorization. | ||
| while offset &+ vectorSize <= count { | ||
| var orResult: UInt8 = 0 | ||
| /// This loop is auto-vectorized. |
There was a problem hiding this comment.
Nothing really guarantees autovec. We could unroll it and get the same perf benefit without relying on autovec:
while offset &+ vectorSize <= count {
let a = bytes.unsafeLoadUnaligned(fromUncheckedByteOffset: offset, as: UInt64.self)
let b = bytes.unsafeLoadUnaligned(fromUncheckedByteOffset: offset &+ 8, as: UInt64.self)
let c = bytes.unsafeLoadUnaligned(fromUncheckedByteOffset: offset &+ 16, as: UInt64.self)
let d = bytes.unsafeLoadUnaligned(fromUncheckedByteOffset: offset &+ 24, as: UInt64.self)
if (a | b | c | d) != 0 {
break
}
offset &+= 32
}
There was a problem hiding this comment.
I disagree. Swift and LLVM are pretty consistent in auto-vectorizing code. They can auto-vectorize much more complicated loops as well. There are some general conditions and all but the mechanism is consistent and in this specific case it's pretty much the simplest case. Generally speaking some of the conditions are:
- The loop is branchless.
- e.g. no byte-load precondition that can't be optimized out.
- And of course, only minimal condition blocks that are reduced to e.g. conditional-select (e.g. CSEL) instructions.
- The code is inlined (no external subroutine calls).
- The length of the loop is known (not necessarily compile-time known).
- There are no data dependencies to a previous iteration of the loop.
- For LLVM specifically, It either never accepts table lookups or hardly ever does (I've never seen it to, and I haven't tried reading through LLVM code; I might at some point).
- Generally table lookups are limited even when manually writing SIMD code.
In this specific code it's just 1 OR + byte-load so you don't really need to go verify all those statements either. Just that I'm trying to convince you that auto-vectorization is more trustable and consistent than what I feel like you think.
There was a problem hiding this comment.
Also the code you provided is auto-vectorized as well and that's why the 2 of the codes don't make much of a difference. If the code was scalar it would perform worse than the auto-vectorized code (General rule of thumb, but not true 100% of the time specially if there are a low amount of loop iterations, say 8/16 for example).
rpaulo
left a comment
There was a problem hiding this comment.
Approved pending test modifications
agnosticdev
left a comment
There was a problem hiding this comment.
Thanks, approved pending testing.
|
@rpaulo @agnosticdev please re-run CI (See 7301108). |
#126 caught my eye. It can be made faster.
Result on M6 processors (in ns):