The two pre-scan steps that were still CPU-bound in a GPU build now run where the projection already is. - FindBeamCenterFromBackground: the per-iteration binning pass and the two clip rounds run on the device (BeamCenterBackgroundGPU); the fit itself stays on the host. Each cell is summed in the host's order (pixel order within the host's row blocks, blocks in order), and the per-pixel cell/derivative formula is shared (BackgroundBand.h). The angles come from BackgroundAtan2 (IEEE ops only) instead of atan2f, and both translation units are compiled without FMA contraction, so host and device give the same bits: 0 of 6.5 M pixels in a different cell, identical walks on the three in-house rotation sets. With glibc/CUDA atan2f and default contraction ~30 pixels per 16 Mpx sweep changed cell and the fitted centre moved by up to 0.05 px. - ShadowFinder::GetMask: the whole mask (pooling, ring medians, components, morphology, hole fill, arm search) runs on the device from ShadowAccumulatorGPU's projection (ShadowMaskGPU), so the 360 MB projection no longer comes back; the mean projection is divided on the device too (same bits). The two small fits over rings and sectors (BlockedOutTo, HarmonicFit) are shared with the host path in ShadowFinderInternal.h. Integers, comparisons, sorts and components are exact; the polarization trig, the Poisson log and the arm-search azimuth are not, so a pixel at a threshold can differ. The one-time change against the previous CPU arithmetic (BackgroundAtan2, no contraction), measured on the myoglobin, cytochrome C and thaumatin rotation sets: ring centre moves 0.002-0.045 px (fit sigma 0.75-1.2 px), beam-centre capture 0.01-0.04 px; beam-stop mask differs on 31 / 144 / 53 pixels of 259k / 144k / 198k (25 of the myoglobin ones are GPU-vs-CPU arithmetic in the mask, the rest follow the centre); hot-pixel mask identical. Spot width, integration radii, bandwidth, beam-centre arbitration, indexing, space group, cell, resolution and the merged statistics table are identical; only the error model moves in its 4th digit. CPU build: the same centres and decisions. Timing (GPU, box at load 30-38): ring walk 0.54 -> 0.23-0.27 s, mask 1.24-1.44 -> 0.18-0.22 s, beam-centre capture walk 1.1-1.3 -> 0.31-0.35 s. Tests: ShadowFinder_DeviceMaskMatchesHost, BeamCenterFromBackground_DeviceMatchesHost (bit-exact), plus [ShadowFinder], [BeamCenter], [HotPixelFinder]. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
83 lines
4.9 KiB
C++
83 lines
4.9 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <optional>
|
|
#include <utility>
|
|
#include <vector>
|
|
|
|
#include "BeamCenterFFT.h"
|
|
#include "../../common/DiffractionExperiment.h"
|
|
#include "../../common/PixelMask.h"
|
|
|
|
struct BeamCenterEstimate {
|
|
float beam_x_pxl = 0.0f;
|
|
float beam_y_pxl = 0.0f;
|
|
float sigma_pxl = 0.0f; // 1 sigma on the fitted shift, the larger of the two axes
|
|
};
|
|
|
|
// Beam centre from the isotropy of the scattered background, before anything is indexed.
|
|
//
|
|
// The solvent and air scatter is isotropic in 2-theta about the beam, so a centre that is off
|
|
// shifts each azimuthal sector's radial profile by a different amount. Sector k's profile is
|
|
// m_k * g(2theta + d_k), with the shift d_k = Jx_k*dx + Jy_k*dy and m_k an amplitude that
|
|
// absorbs anything multiplicative and azimuthal - a holder arm, a cryostream shadow, a
|
|
// flat-field gradient. Fitting the amplitude alongside the shift is what makes this usable:
|
|
// a 50% shadow over one sextant otherwise reads as several tens of pixels of centre error.
|
|
//
|
|
// The leverage comes from the CURVATURE of the radial profile - the water ring - because for a
|
|
// pure exponential decay g' is proportional to g and shift and amplitude are indistinguishable.
|
|
// The sigma measures that leverage, so it grows as the curvature weakens, and the caller's gate on
|
|
// it is what keeps an ill-determined centre out. It is a precision and not an accuracy: on a
|
|
// background with NO curvature at all there is nothing to separate the two parameters, the fit
|
|
// follows the noise in g' instead, and it does so confidently.
|
|
//
|
|
// `mean` is a per-pixel projection over a few tens of frames, NAN where no frame contributed.
|
|
// nthreads = 0 asks for all hardware threads. The pixels are split into a fixed number of row
|
|
// blocks whatever that count is, so the answer does not depend on it.
|
|
//
|
|
// `start` is where the walk begins; the centre in the file when it is not given. The walk advances
|
|
// by a bounded distance per iteration, so where it starts decides how much of its budget is spent
|
|
// travelling and - on a surface with more than one basin - which fixed point it can reach at all.
|
|
//
|
|
// With a GPU the passes over the pixels run on it (BeamCenterBackgroundGPU); allow_device = false
|
|
// keeps them on the host, which is what the parity test compares against.
|
|
std::optional<BeamCenterEstimate>
|
|
FindBeamCenterFromBackground(const DiffractionExperiment &experiment, const PixelMask &mask,
|
|
const std::vector<float> &mean, size_t nthreads = 0,
|
|
std::optional<std::pair<float, float>> start = {},
|
|
bool allow_device = true);
|
|
|
|
// The precision of a centre that is the FFT capture alone, with no walk behind it. The capture is
|
|
// a half-pixel grid position read off a surface, measured over 75 rotation datasets at a median
|
|
// 3.2 px and a 90th percentile of 11 px from the truth, so this is what it knows the centre to -
|
|
// a capture precision and not a fit precision. It is deliberately far above the ceiling the
|
|
// callers adopt a centre on: a capture is evidence about where the beam is, not a measurement of
|
|
// where it is.
|
|
constexpr float BEAM_CENTER_CAPTURE_SIGMA_PXL = 5.0f;
|
|
|
|
// The beam centre from the background, captured globally and then refined locally.
|
|
//
|
|
// The walk in FindBeamCenterFromBackground is a good local refiner and a poor searcher: it moves
|
|
// one or two pixels per iteration, it is seeded at the centre in the file, and on a background
|
|
// whose isotropy is broken it has a second basin to fall into. The FFT score is the opposite -
|
|
// it evaluates EVERY candidate centre on the detector in one transform set, at a cost that does
|
|
// not depend on how wrong the file is, but it returns a half-pixel grid position and no sigma.
|
|
// Composing them takes the reach from one and the precision from the other: the capture chooses
|
|
// the basin, the walk finishes inside it and reports what it knows the answer to.
|
|
//
|
|
// The shadow of the beam stop is blanked out of the image the capture scores. A one-sided opaque
|
|
// region imposes a centrosymmetry of its own that can beat the background's - measured, an umbra
|
|
// over 9.6 % of the detector put the capture 48 px out, and masking it put it back to 1.1 px,
|
|
// while a random mask of the same area changed nothing.
|
|
//
|
|
// Where the walk declines at the capture the walk is asked again from the centre in the file, and
|
|
// only where neither start gives it something to fit does the capture stand alone, at
|
|
// BEAM_CENTER_CAPTURE_SIGMA_PXL - which is what lets a caller that only needs a hypothesis to test
|
|
// still get one. Returns nothing only where the capture has no candidate either.
|
|
std::optional<BeamCenterEstimate>
|
|
FindBeamCenter(const DiffractionExperiment &experiment, const PixelMask &mask,
|
|
const std::vector<float> &mean, size_t nthreads = 0,
|
|
BeamCenterFFTResult *capture = nullptr);
|