03dee2dda7d4e84759188a9708710457afc8218c
The standalone GPU azimuthal integration did three shared-memory atomics per pixel on the same few ring addresses, which is what it was limited by. It now reads four pixels per thread as vector loads and keeps a running total per ring, flushed when the ring changes - the scheme the adaptive finder's ring pass (reduce_rings_shared) already uses. The npix % 4 leftovers are done one at a time. Used wherever the fused adaptive engine is not (fixed-threshold spot finding, the broker's non-adaptive path). Measured on a 16 Mpx sweep (1800 frames, --no-adaptive-spots, RTX 5080): 843 -> 295 us per call (min 621 -> 196 us). Not bit-identical, and the old kernel was not either: float atomics arrive in any order, so two runs of the OLD kernel already differ by up to 1.7e-6 relative in the per-frame profile; new vs old differs by up to 1.9e-6, the same order. Per-ring pixel counts are identical. Default rugnux runs do not reach this kernel (p.mtz md5 unchanged on three sets); on the fixed-threshold path p_unmerged.mtz is md5-identical to the old kernel's. New test: GPU vs CPU engine on a pixel count that is not a multiple of four, with masked and saturated pixels. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
Jungfraujoch
Application to receive data from the PSI JUNGFRAU and EIGER detectors.
All documentation is now placed in docs/ subdirectory and for the current version hosted on Jungfraujoch Read The Docs page.
Languages
C++
77.9%
HTML
6%
C
4.7%
TypeScript
3.4%
Cuda
3%
Other
4.9%