AdaptiveSpotFinderCPU::AccumulateRingsBlock reads each pixel once for both the
ring histogram and the fused azimuthal profile (was two loops), and no longer
keeps the per-ring integer sums: they are taken from the histogram, as the two
sigma-clip passes already were (ClipRings(INFINITY)). Integer sums, so the same
totals; the profile's float sums keep their pixel order.
FlagRow is branch-free and works a 32-pixel word at a time, so it vectorises;
pixels outside every ring meet a +inf threshold in an extra ring_thr entry.
BraggPredictionRot::Calc takes A*h, A*h + B*k, C*l and 4*S0*S0 out of the inner
loops; p0 is the same ((A*h) + (B*k)) + (C*l) as before.
Measured on cytc (CPU build, both binaries run concurrently on a loaded box):
AccumulateRingsBlock -35%, FlagRow -47%, Calc + Coord ops -30% cycles; whole run
-5% cycles. p.mtz md5 unchanged on myob/cytc/thau, CPU and GPU builds.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB