The fused adaptive spot finder's azimuthal profile (and the standalone AzIntEngineGPU) summed the
corrected pixel values - pixel times a float correction - with float atomicAdd, in shared memory
and then across blocks. Float addition is not associative, so the per-ring sums depended on the
order the threads and blocks happened to arrive in, and on the grid size chosen for the GPU. The
profile, and the per-image background estimate read from it, moved in their last bits between two
runs of the same command: on a thaumatin JUNGFRAU 4M sweep one bkg value in _plot.txt differed in
the 4th digit between runs. The raw ring sums that drive detection were already exact integers;
only the reported profile was affected.
Each corrected value (and its square) is now quantised to 2^-12 on its own and summed as a signed
64-bit integer, the convention the raw ring sums and the Bragg-integration profile already use.
Integer addition is associative, so the sums are the same whatever the order or grid.
The repeat test of the fused finder now also requires the profile and its std to repeat bit for
bit (it failed on the old kernel at the first repeat); the AzIntEngineGPU test re-runs 20 times.
Fused finder benchmark: 0.313 -> 0.324 ms/frame.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi