Two fixes found in the post-merge profile, where the previous commit's loops did not do what they
were written to do:
- AccumulateRingsBlock: the ring_overflow.emplace_back call inside the loop left no callee-saved
registers for the locals, so GCC kept sum/sum2/count on the stack and the store-reload chain was
still there (29% of the function's samples on one stack add). The azint sums get a loop of their own
over the block, and the out-of-histogram values are listed by a second loop run only when the block
has any, in the same pixel order; now the sums live in registers. Same additions in the same order.
- DetectPass: inlined into DetectPass's main loop the vertical slide was not vectorised (the
vectorised copies GCC made were for the other call sites). It is now a free function,
SlideVertical, which GCC vectorises on its own.
Measured (perf, 75 s of myob, CPU-only, relative to the unchanged FlagRow/AnalyzeBlock):
AccumulateRingsBlock -4..-11%, DetectPass + SlideVertical -7..-14%. Byte-identical p.hkl, p.mtz,
p_P1.mtz, p_unmerged.mtz on myob, cytc, lyso, sparse (CPU) and myob (GPU); ImageSpotFinderCPU*,
AdaptiveSpotFinder, SpotFinding, AzimuthalIntegration and portable tests pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C