The standalone GPU azimuthal integration did three shared-memory atomics per
pixel on the same few ring addresses, which is what it was limited by. It now
reads four pixels per thread as vector loads and keeps a running total per ring,
flushed when the ring changes - the scheme the adaptive finder's ring pass
(reduce_rings_shared) already uses. The npix % 4 leftovers are done one at a time.
Used wherever the fused adaptive engine is not (fixed-threshold spot finding, the
broker's non-adaptive path). Measured on a 16 Mpx sweep (1800 frames,
--no-adaptive-spots, RTX 5080): 843 -> 295 us per call (min 621 -> 196 us).
Not bit-identical, and the old kernel was not either: float atomics arrive in any
order, so two runs of the OLD kernel already differ by up to 1.7e-6 relative in
the per-frame profile; new vs old differs by up to 1.9e-6, the same order. Per-ring
pixel counts are identical. Default rugnux runs do not reach this kernel (p.mtz
md5 unchanged on three sets); on the fixed-threshold path p_unmerged.mtz is
md5-identical to the old kernel's. New test: GPU vs CPU engine on a pixel count
that is not a multiple of four, with masked and saturated pixels.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB