Files
Jungfraujoch/image_analysis/beam_stop/ShadowMaskGPU.h
T
leonarski_fandClaude Opus 5.5 f849e2d1be Pre-scan on the GPU: background beam-centre walk and beam-stop mask
The two pre-scan steps that were still CPU-bound in a GPU build now run where the
projection already is.

- FindBeamCenterFromBackground: the per-iteration binning pass and the two clip rounds
  run on the device (BeamCenterBackgroundGPU); the fit itself stays on the host. Each
  cell is summed in the host's order (pixel order within the host's row blocks, blocks
  in order), and the per-pixel cell/derivative formula is shared (BackgroundBand.h).
  The angles come from BackgroundAtan2 (IEEE ops only) instead of atan2f, and both
  translation units are compiled without FMA contraction, so host and device give the
  same bits: 0 of 6.5 M pixels in a different cell, identical walks on the three
  in-house rotation sets. With glibc/CUDA atan2f and default contraction ~30 pixels per
  16 Mpx sweep changed cell and the fitted centre moved by up to 0.05 px.
- ShadowFinder::GetMask: the whole mask (pooling, ring medians, components, morphology,
  hole fill, arm search) runs on the device from ShadowAccumulatorGPU's projection
  (ShadowMaskGPU), so the 360 MB projection no longer comes back; the mean projection is
  divided on the device too (same bits). The two small fits over rings and sectors
  (BlockedOutTo, HarmonicFit) are shared with the host path in ShadowFinderInternal.h.
  Integers, comparisons, sorts and components are exact; the polarization trig, the
  Poisson log and the arm-search azimuth are not, so a pixel at a threshold can differ.

The one-time change against the previous CPU arithmetic (BackgroundAtan2, no
contraction), measured on the myoglobin, cytochrome C and thaumatin rotation sets:
ring centre moves 0.002-0.045 px (fit sigma 0.75-1.2 px), beam-centre capture
0.01-0.04 px; beam-stop mask differs on 31 / 144 / 53 pixels of 259k / 144k / 198k
(25 of the myoglobin ones are GPU-vs-CPU arithmetic in the mask, the rest follow the
centre); hot-pixel mask identical. Spot width, integration radii, bandwidth, beam-centre
arbitration, indexing, space group, cell, resolution and the merged statistics table
are identical; only the error model moves in its 4th digit. CPU build: the same
centres and decisions.

Timing (GPU, box at load 30-38): ring walk 0.54 -> 0.23-0.27 s, mask 1.24-1.44 ->
0.18-0.22 s, beam-centre capture walk 1.1-1.3 -> 0.31-0.35 s.

Tests: ShadowFinder_DeviceMaskMatchesHost, BeamCenterFromBackground_DeviceMatchesHost
(bit-exact), plus [ShadowFinder], [BeamCenter], [HotPixelFinder].

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-03 12:19:05 +02:00

40 lines
2.0 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
// Included only under JFJOCH_USE_CUDA.
#include <cstdint>
#include <vector>
#include <cuda_runtime.h>
// What the mask is drawn about: the detector, the centre the rings are drawn about, and the geometry
// the polarization factor is read off (ShadowFinder keeps it as a DiffractionGeometry; the device
// takes it as numbers).
struct ShadowMaskSetup {
int width = 0, height = 0;
float beam_x = 0.0f, beam_y = 0.0f;
float det_matrix[9] = {}; // row major
float pixel_size_mm = 0.0f;
float distance_mm = 0.0f;
bool has_polarization = false;
float polarization = 0.0f;
};
// ShadowFinder::GetMask on the device, from the projection ShadowAccumulatorGPU holds there. Step for
// step the host's algorithm - the same pooling, ring medians, components, morphology and arm search -
// and the same answer wherever the arithmetic is exact: every integer, comparison, sort and component
// is. What is not is the floating point the two compilers evaluate differently - the polarization
// factor's trigonometry, the Poisson test's logarithm and the azimuth of the arm search - so a pixel
// within a rounding of one of those thresholds can come out the other way.
std::vector<uint32_t> ShadowMaskOnDevice(const ShadowMaskSetup &setup, const std::vector<uint32_t> &pixel_mask,
const int64_t *max_value, const int64_t *sum_value,
const uint32_t *valid_count, uint32_t frames, cudaStream_t stream);
// The mean projection the host's GetMeanProjection makes, computed where the sums are: the same
// division, so the same bits, and a quarter of the bytes to bring back.
std::vector<float> MeanProjectionOnDevice(const std::vector<uint32_t> &pixel_mask, const int64_t *sum_value,
const uint32_t *valid_count, size_t npixels, cudaStream_t stream);