CPU pixel loops: ring sums held in registers, vectorised vertical window, shared packed mask

AdaptiveSpotFinderCPU::AccumulateRings: consecutive pixels mostly share a ring, so the ring's
sum/sum2/count and the fused azint sums are held in locals while they do and stored when the ring
changes - the same additions in the same order (az_sum2 is still contracted to the same FMA), without
a store-and-reload chain through memory on every pixel.

ImageSpotFinderCPU::DetectPass: the vertical-sum update (add the entering row, take out the leaving
one) is one branch-free loop over the raw image that GCC vectorises (int64 lanes). A pixel strong in
the previous pass used to be substituted per pixel through a bit test, which kept the loop scalar;
it is now added with its row and taken out again from the few set bits of prev_strong. Integer sums,
so the same totals. (A first, fully branch-free version that kept the per-pixel bit test did not
vectorise on the prev_strong path and was measured slower; this is its replacement.)

ImagePreprocessorCPU: the per-engine std::vector<bool> built bit by bit from the 32-bit mask
(~10 core-s per cytc run, one per worker per pass) is replaced by 32-pixel mask words that PixelMask
derives once beside its binary mask; each engine copies 2 MB. A branch-free rewrite of the Analyze
loop was measured and dropped: the loop is bound by reading the decompressed image (330 vs 328
core-s on cytc), so only the mask test changed.

Measured (perf, 499 Hz, CPU-only build, cytc, first version of this change): AccumulateRings
591 -> 539 core-s. Byte-identical p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz on myob, cytc, lyso,
sparse (CPU) and myob, lyso (GPU).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
This commit is contained in:
2026-09-28 02:19:31 +02:00
co-authored by Claude Opus 5.5
parent f4f9157cc0
commit c8692e320c
6 changed files with 107 additions and 46 deletions
@@ -5,10 +5,7 @@
ImagePreprocessorCPU::ImagePreprocessorCPU(const DiffractionExperiment &experiment, const PixelMask &mask)
: ImagePreprocessor(experiment),
mask_1bit(npixels, false) {
for (int i = 0; i < npixels; i++)
mask_1bit[i] = (mask.GetMask().at(i) != 0);
}
mask_bits(mask.GetPackedMask()) {}
ImageStatistics ImagePreprocessorCPU::Analyze(ImagePreprocessorBuffer &processed_image, const uint8_t *image_ptr, CompressedImageMode image_mode) {
switch (image_mode) {
@@ -43,7 +40,7 @@ ImageStatistics ImagePreprocessorCPU::Analyze(ImagePreprocessorBuffer &processed
sat_pixel_val = static_cast<T>(saturation_limit);
for (int i = 0; i < npixels; i++) {
if (mask_1bit[i] != 0) {
if ((mask_bits[i / 32] >> (i % 32)) & 1U) {
processed_image[i] = INT32_MIN;
++ret.masked_pixel_count;
} else if (image[i] == err_pixel_val) {