The two pre-scan steps that were still CPU-bound in a GPU build now run where the
projection already is.
- FindBeamCenterFromBackground: the per-iteration binning pass and the two clip rounds
run on the device (BeamCenterBackgroundGPU); the fit itself stays on the host. Each
cell is summed in the host's order (pixel order within the host's row blocks, blocks
in order), and the per-pixel cell/derivative formula is shared (BackgroundBand.h).
The angles come from BackgroundAtan2 (IEEE ops only) instead of atan2f, and both
translation units are compiled without FMA contraction, so host and device give the
same bits: 0 of 6.5 M pixels in a different cell, identical walks on the three
in-house rotation sets. With glibc/CUDA atan2f and default contraction ~30 pixels per
16 Mpx sweep changed cell and the fitted centre moved by up to 0.05 px.
- ShadowFinder::GetMask: the whole mask (pooling, ring medians, components, morphology,
hole fill, arm search) runs on the device from ShadowAccumulatorGPU's projection
(ShadowMaskGPU), so the 360 MB projection no longer comes back; the mean projection is
divided on the device too (same bits). The two small fits over rings and sectors
(BlockedOutTo, HarmonicFit) are shared with the host path in ShadowFinderInternal.h.
Integers, comparisons, sorts and components are exact; the polarization trig, the
Poisson log and the arm-search azimuth are not, so a pixel at a threshold can differ.
The one-time change against the previous CPU arithmetic (BackgroundAtan2, no
contraction), measured on the myoglobin, cytochrome C and thaumatin rotation sets:
ring centre moves 0.002-0.045 px (fit sigma 0.75-1.2 px), beam-centre capture
0.01-0.04 px; beam-stop mask differs on 31 / 144 / 53 pixels of 259k / 144k / 198k
(25 of the myoglobin ones are GPU-vs-CPU arithmetic in the mask, the rest follow the
centre); hot-pixel mask identical. Spot width, integration radii, bandwidth, beam-centre
arbitration, indexing, space group, cell, resolution and the merged statistics table
are identical; only the error model moves in its 4th digit. CPU build: the same
centres and decisions.
Timing (GPU, box at load 30-38): ring walk 0.54 -> 0.23-0.27 s, mask 1.24-1.44 ->
0.18-0.22 s, beam-centre capture walk 1.1-1.3 -> 0.31-0.35 s.
Tests: ShadowFinder_DeviceMaskMatchesHost, BeamCenterFromBackground_DeviceMatchesHost
(bit-exact), plus [ShadowFinder], [BeamCenter], [HotPixelFinder].
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB