The two pre-scan steps that were still CPU-bound in a GPU build now run where the projection already is. - FindBeamCenterFromBackground: the per-iteration binning pass and the two clip rounds run on the device (BeamCenterBackgroundGPU); the fit itself stays on the host. Each cell is summed in the host's order (pixel order within the host's row blocks, blocks in order), and the per-pixel cell/derivative formula is shared (BackgroundBand.h). The angles come from BackgroundAtan2 (IEEE ops only) instead of atan2f, and both translation units are compiled without FMA contraction, so host and device give the same bits: 0 of 6.5 M pixels in a different cell, identical walks on the three in-house rotation sets. With glibc/CUDA atan2f and default contraction ~30 pixels per 16 Mpx sweep changed cell and the fitted centre moved by up to 0.05 px. - ShadowFinder::GetMask: the whole mask (pooling, ring medians, components, morphology, hole fill, arm search) runs on the device from ShadowAccumulatorGPU's projection (ShadowMaskGPU), so the 360 MB projection no longer comes back; the mean projection is divided on the device too (same bits). The two small fits over rings and sectors (BlockedOutTo, HarmonicFit) are shared with the host path in ShadowFinderInternal.h. Integers, comparisons, sorts and components are exact; the polarization trig, the Poisson log and the arm-search azimuth are not, so a pixel at a threshold can differ. The one-time change against the previous CPU arithmetic (BackgroundAtan2, no contraction), measured on the myoglobin, cytochrome C and thaumatin rotation sets: ring centre moves 0.002-0.045 px (fit sigma 0.75-1.2 px), beam-centre capture 0.01-0.04 px; beam-stop mask differs on 31 / 144 / 53 pixels of 259k / 144k / 198k (25 of the myoglobin ones are GPU-vs-CPU arithmetic in the mask, the rest follow the centre); hot-pixel mask identical. Spot width, integration radii, bandwidth, beam-centre arbitration, indexing, space group, cell, resolution and the merged statistics table are identical; only the error model moves in its 4th digit. CPU build: the same centres and decisions. Timing (GPU, box at load 30-38): ring walk 0.54 -> 0.23-0.27 s, mask 1.24-1.44 -> 0.18-0.22 s, beam-centre capture walk 1.1-1.3 -> 0.31-0.35 s. Tests: ShadowFinder_DeviceMaskMatchesHost, BeamCenterFromBackground_DeviceMatchesHost (bit-exact), plus [ShadowFinder], [BeamCenter], [HotPixelFinder]. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
40 lines
2.0 KiB
C++
40 lines
2.0 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
// Included only under JFJOCH_USE_CUDA.
|
|
|
|
#include <cstdint>
|
|
#include <vector>
|
|
|
|
#include <cuda_runtime.h>
|
|
|
|
// What the mask is drawn about: the detector, the centre the rings are drawn about, and the geometry
|
|
// the polarization factor is read off (ShadowFinder keeps it as a DiffractionGeometry; the device
|
|
// takes it as numbers).
|
|
struct ShadowMaskSetup {
|
|
int width = 0, height = 0;
|
|
float beam_x = 0.0f, beam_y = 0.0f;
|
|
float det_matrix[9] = {}; // row major
|
|
float pixel_size_mm = 0.0f;
|
|
float distance_mm = 0.0f;
|
|
bool has_polarization = false;
|
|
float polarization = 0.0f;
|
|
};
|
|
|
|
// ShadowFinder::GetMask on the device, from the projection ShadowAccumulatorGPU holds there. Step for
|
|
// step the host's algorithm - the same pooling, ring medians, components, morphology and arm search -
|
|
// and the same answer wherever the arithmetic is exact: every integer, comparison, sort and component
|
|
// is. What is not is the floating point the two compilers evaluate differently - the polarization
|
|
// factor's trigonometry, the Poisson test's logarithm and the azimuth of the arm search - so a pixel
|
|
// within a rounding of one of those thresholds can come out the other way.
|
|
std::vector<uint32_t> ShadowMaskOnDevice(const ShadowMaskSetup &setup, const std::vector<uint32_t> &pixel_mask,
|
|
const int64_t *max_value, const int64_t *sum_value,
|
|
const uint32_t *valid_count, uint32_t frames, cudaStream_t stream);
|
|
|
|
// The mean projection the host's GetMeanProjection makes, computed where the sums are: the same
|
|
// division, so the same bits, and a quarter of the bytes to bring back.
|
|
std::vector<float> MeanProjectionOnDevice(const std::vector<uint32_t> &pixel_mask, const int64_t *sum_value,
|
|
const uint32_t *valid_count, size_t npixels, cudaStream_t stream);
|