The GPU finder read the whole image (int32 values + pixel_to_bin) three times per frame for its
per-ring background: a plain pass and two 3-sigma clip passes. As the CPU finder already does
(87cfc879b), the plain pass now also records every valid pixel in a per-ring value histogram
(HIST_VALUES = 1024 bins per ring, global memory, integer atomics aggregated per warp by
(ring, value)), and pixels outside [0, 1024) on an overflow list. Each clip pass is then taken from
the histogram and the list (clip_rings_from_hist, one warp per ring) - integer sums of the same
values, so the ring statistics, thresholds and spots are exactly the image passes'. If the list
overflows its capacity (npix / 32 entries), the clip passes read the image as before; the choice is
made on the device, so there is no extra host sync.
The clip bounds are now written as explicit fused multiply-adds (clip_keep). ptxas already fused
`mean -/+ clip_k * sigma` into FFMA in the image-pass kernel (checked in the sm_120 SASS); spelling it
out keeps the histogram clip and the image clip on exactly the same line.
Memory: nbins * 4 KiB histogram + 8 B * npix / 32 overflow list per engine, only on the
shared-memory path (at most ~1750 rings, so <= 7 MB of histogram); the list is 4.5 MB for a 16M EIGER2.
Exactness: a temporary side-by-side check (histogram clips vs forced image clips on every frame)
gave 0 differing ring sums/counts on myob, cytc, lyso and sparse (overflow list at most 308
entries); p.hkl, p.mtz, p_P1.mtz and p_unmerged.mtz are byte-identical to rc173 on all four.
New test AdaptiveSpotFinderGPU_HistogramClipMatchesImagePasses covers rings above/straddling the
histogram range and the overflow fallback.
Timing (RTX 5080, nsys, GPU locked; myob 16M EIGER2): reduce_rings 230 us x 3 launches/frame
(3.20 s) -> plain pass 308 us + 2 x 4 us early-exit launches + 2 x 12 us clip_rings_from_hist
(1.58 s total); summed kernel time 11.28 -> 9.18 s; image loops 5.89/5.09 -> 5.49/4.69 s. cytc:
kernels 12.86 -> 10.75 s. Small detectors gain less (lyso 0.40 -> 0.32 s, sparse 1.19 -> 1.01 s of
ring reduction), as the histogram costs relatively more in the plain pass there.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C