bitshuffle_hperf's vector code is x86-only, so aarch64 decoded with the classic NEON path, whose
bit stage emulates x86 movemask (8 and + 8 cmeq + 2 x 9-instruction pairwise-add reductions per
16 bytes, then eight scattered 16-bit stores): ~29 instructions per 8x8 bit block plus a byte
transpose pass with scattered 8-byte stores, and a scalar path for 1-byte elements.
bitshuffle_neon.c follows hperf's structure (per byte plane bit-untranspose, then byte interleave)
but does the 8x8 bit transpose NEON-natively: the eight rows stay in eight registers, bits are
exchanged between them with shift + bit-select in three levels, and one zip level plus two st4
write the 128 contiguous output bytes. GCC 13 -O3 gives ~75 instructions per 16 blocks (~4.7 per
block) in the bit stage; the byte interleave is one st2/st4 per 16 elements. Other element sizes
keep the classic stages with the caller's scratch. x86 is unchanged (the file compiles empty).
Byte-identical under qemu-aarch64 (GCC 13, -O3 and -O0) to the classic decoder, hperf and the
x86 hperf/AVX2 path: element sizes 1/2/3/4/8, every block of 8..1040 elements plus 1536..16384,
seven data patterns, and real EIGER2/PILATUS4 bslz4 chunks (whose decoded image hash also
matches an hdf5plugin read). Same licence and component as bitshuffle_hperf, so licenses/ and
THIRD_PARTY_NOTICES.md are unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C