MaskDefectivePixels re-read its 60 sample frames with 8 CPU workers - host
bitshuffle/LZ4 decode, CPU preprocessing, then per frame a scatter by ring-sector,
nth_element medians and per-pixel sums over ~22 B/px of state - and was bound by
memory bandwidth: 2.9 s wall on an EIGER2 16M sweep, the largest single step left in
the pre-scan.
With a GPU each worker now decodes and preprocesses the frame on the device
(ImagePreprocessorGPU::AnalyzeCompressed, as the image loops do), and
HotPixelFinderGPU takes the order statistics there: one block per ring-sector and
per ring, radix selection eight bits at a time, which is exact for any values, so
the histogram-plus-fallback of the host path is not needed. Only the per-key
statistics (count, sector median, ring median and MAD) come back; levels and
thresholds are computed on the host by the same function the CPU path now calls
(AddLevels), uploaded, and a per-pixel kernel does AddImage's loop line for line,
including the float comparison. The per-pixel sums stay on the device and are
downloaded once when the mask is read. The CPU path is unchanged. The constructor's
~470 MB of per-pixel arrays are no longer zeroed on the calling thread: they are
allocated unwritten and first written by the existing parallel key pass.
Exactness: a temporary check that ran both finders side by side on the same
frames (CPU decode + preprocess against GPU decode + preprocess) found every
per-pixel and per-key sum identical on myob, cytc, lyso and sparse; the new test
HotPixelFinder_DeviceMatchesHost compares the masks on synthetic frames that take
every branch (histogram fallback, negatives, saturated/error/masked pixels,
134 masked). p.hkl, p.mtz, p_P1.mtz and p_unmerged.mtz are byte-identical to the
previous GPU build on myob and cytc; the masked counts (3/2/0) are unchanged.
myob GPU, quiet box: MaskDefectivePixels 2.90 s -> 0.42 s, total 20.4 -> 19.4 s,
peak RSS 3.52 -> 3.60 GB.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Unmasked persistent hot pixels are integrated into whichever reflection's box they fall in; under
rotation one pixel collects a different reflection on every frame that reaches it, and the merge
carries intensities hundreds to thousands of times their shell mean on one or two observations.
HotPixelFinder (rugnux/HotPixels.{h,cpp}) reads the pre-scan sample a second time, once the
beam-stop projection has measured where the background puts the beam, and on each frame calls a
pixel lit when it exceeds its 2 px iso-2theta ring level (max of the ring and 1/16-sector medians)
by 3.3 sigma (sqrt(level) or the ring's MAD) + 2. A pixel lit on at least max(k1, kB) frames is
persistent: k1 = 1 + ceil((osc + 5 deg)/|zeta| / frame spacing) is more than one reflection can
light, kB the binomial bound (0.01 family-wise over the detector) from the ring's own lit rate.
A persistent pixel is masked, as the new PixelMask bit 10, only if it stands alone (component of
persistent pixels <= 2), reads on average >= 10x its ring and its mean excess is above the Poisson
bound; pixels holding the error value on most frames are masked with them. Counting sensors (thickness > 0)
and rotation data only; a CCD is left alone. One log line reports the counts.
Drawing the rings about the file's centre, as a first version did, masked pixels along the
background fall-off on a sweep whose file centre is 171 px from the background's and cost it 14%
ISa; about the measured centre that sweep is within 1%. Masking the detector's outermost row and
column unconditionally was tried and dropped: the persistence test already catches the hot pixels
there, and the whole lines bought nothing measurable.
Numbers below are from the looser first criterion (no isolation / 10x gate), against rc173-final
on the same base: merged reflections > 30x their shell mean gone on the
sets with proven hot pixels (7brr 21 -> 0, 9ih9 15 -> 0, 8xte 10 -> 0, 6z8o worst 1706x -> 40x);
6z8o CC1/2 0.50 -> 0.995, ISa 8.6 -> 12.8, CC to model 0.82 -> 0.90; 8xte ISa 8.0 -> 10.8; 6u7g
ISa 9.8 -> 12.0; controls (lyso_x06da_ref, marCCD) unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C