The per-image ice score was read off the PLAIN azimuthal profile. That profile is a per-ring mean, so a few strong Bragg reflections landing in a ring's q bin lift it exactly as ice would. Measured over 37 rotation crystals, that did not merely add noise - it INVERTED the metric: the two highest-scoring crystals had no ice at all (4.23 and 4.06), while a clean control read 1.57. A decoy null - the identical statistic evaluated at q positions where hexagonal ice cannot be - reaches 1.51 at its 99th percentile and 2.70 at its maximum, so that metric cannot support any absolute threshold whatsoever. The adaptive spot finder already computes the right input for its own threshold: a sigma-clipped per-resolution-ring background, in the same bins. A powder ring is azimuthally smooth and survives the clip; Bragg peaks do not. On the clipped profile the clean population tightens to 1.00-1.22 and the crystals with confirmed ice sit at 2.08-2.37, against a decoy null that never exceeds 1.29. That channel is blind to one thing: ice in large crystallites diffracts as DISCRETE spots and leaves the radial profile flat. So a second channel counts found spots on the rings against the same q width of ice-free flanks beside them. The two barely overlap - the smooth-ice crystals read 2.1-2.4 / ~1.0 and the textured ones ~1.1 / 3.8-17.6, while a clean crystal reads 1.04 on both. Both are then used as a GATE (--ice-min-score 1.5, --ice-min-spot-ratio 2.0, both calibrated on the battery, 0 disables): the eleven fixed hexagonal bands cover 16-26 % of the unique reflections at typical resolutions whether or not the crystal has ice, so flagging, the exclusion from the scale fit and the merge-time CC1/2 ring mask are now all skipped when neither channel sees any. The gate is applied in the full pipeline and in --scale, which reads the stored per-image values back out of the _process.h5. Also fixes the merge-time mask's control: the shoulder now excludes reflections that are themselves on an ice ring. The rings are not evenly spaced - 1.947/1.916/1.882 A sit 0.05-0.06 apart in q - so for those three the [w,3w) shoulder landed squarely on the neighbours and the test compared ice against ice. Measured, that is the only thing this changes: it removes firings on those three rings and leaves every other firing's CC pair identical to three decimals. And the online ice half-width, which was 0.02 in the API against 0.03 offline, so the same data got a narrower band online than the measured ~0.06 ring FWHM justifies. Battery (37 rotation crystals, against the previous behaviour): space groups 34/37 in both and NO crystal's space group changes; 6 crystals gain unique reflections, 1 loses. Best of them gains 7082 unique reflections with R_meas 16.0 -> 14.3, CC1/2 95.9 -> 97.3 and ISa 13.7 -> 19.0; another goes R_meas 54.9 -> 42.9, CC1/2 84.0 -> 90.4, ISa 3.9 -> 5.5; a third reaches CC1/2 99.4 from 95.7 at an unchanged reflection count. The one crystal that loses reflections improves on both R_meas and CC1/2. Not done here: the ScanResult/API/plot-type/frontend/viewer layers for the new spot_count_ice_control (they need the OpenAPI regeneration). Message, CBOR, HDF5 write/read and the receiver plots are. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
112 lines
6.4 KiB
C++
112 lines
6.4 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
// GPU adaptive spot finder that FUSES azimuthal integration and spot finding into one image pass.
|
|
//
|
|
// The CPU adaptive finder (AdaptiveSpotFinderCPU) and the azimuthal integrator both bin every pixel
|
|
// into resolution rings and reduce (sum / sum^2 / count). Today azint runs on the GPU while the
|
|
// adaptive finder re-does the identical per-ring reduction on the HOST - a wasted second pass over a
|
|
// ~10 MP image. This engine does the ring reduction on the GPU and drives BOTH products from it:
|
|
// - the azimuthal-integration profile (mean intensity per ring, in flat-field-corrected space), and
|
|
// - the per-ring background (mean, sigma, peak-excluded via two sigma-clip passes) that sets the
|
|
// self-calibrating spot-detection threshold (in raw photon counts).
|
|
// It then flags strong pixels (value >= ring threshold) into a packed bit buffer and hands that
|
|
// buffer - still on the device - to SpotExtractorGPU, which builds the spots there.
|
|
//
|
|
// Numerically it reproduces AdaptiveSpotFinderCPU: the same three-pass robust background, the same
|
|
// per-ring threshold formula (shared via AdaptiveThreshold.h, computed on the host once per frame),
|
|
// and the same raw-count detection test. The only differences from the CPU are those inherent to a
|
|
// GPU reduction (float per-ring accumulation in atomic order vs the CPU's serial double sums), which
|
|
// shift a handful of borderline pixels at most. The corrected sums for the azint profile are
|
|
// accumulated in the SAME plain first pass, so one reduction feeds both products.
|
|
|
|
#include <memory>
|
|
#include <vector>
|
|
|
|
#include "ImageSpotFinder.h"
|
|
#include "SpotExtractorGPU.h"
|
|
#include "SpotFindingSettings.h"
|
|
#include "../../common/AzimuthalIntegrationProfile.h"
|
|
#include "../../common/AzimuthalIntegrationMapping.h"
|
|
#include "../indexing/CUDAMemHelpers.h"
|
|
#include "../indexing/CudaSharedTables.h"
|
|
|
|
class AdaptiveSpotFinderGPU : public ImageSpotFinder {
|
|
const AzimuthalIntegrationMapping &mapping;
|
|
std::shared_ptr<CudaStream> stream;
|
|
|
|
const int nbins;
|
|
const size_t npix;
|
|
|
|
int reduce_threads = 128;
|
|
int reduce_blocks = 0;
|
|
int flag_threads = 256;
|
|
int flag_blocks = 0;
|
|
size_t shared_plain = 0; // per-block shared bytes for the plain pass (raw + corrected rings)
|
|
size_t shared_clip = 0; // per-block shared bytes for a sigma-clip pass (raw rings only)
|
|
bool use_shared = true; // false -> nbins too large for shared memory, use the global-atomics kernel
|
|
|
|
// Static mapping inputs: geometry-only, so one copy per GPU shared with every other engine on it
|
|
// (see CudaSharedTables.h) rather than one copy per worker thread.
|
|
std::shared_ptr<CudaDevicePtr<uint16_t>> gpu_pixel_to_bin;
|
|
std::shared_ptr<CudaDevicePtr<float>> gpu_corrections;
|
|
|
|
// Raw per-ring accumulators (re-zeroed each pass) + derived stats used to clip and threshold.
|
|
// double, like the CPU engine's ring accumulators: the ring sigma is the cancelling difference
|
|
// sum2/n - m^2, and the block atomics that fill these arrive in an arbitrary order.
|
|
CudaDevicePtr<unsigned long long> gpu_sum;
|
|
CudaDevicePtr<unsigned long long> gpu_sum2;
|
|
CudaDevicePtr<uint32_t> gpu_count;
|
|
CudaDevicePtr<float> gpu_mean; // per-ring raw mean (clip predicate)
|
|
CudaDevicePtr<float> gpu_sigma; // per-ring raw sigma (clip predicate)
|
|
|
|
// Corrected per-ring accumulators (plain first pass only) -> azimuthal-integration profile.
|
|
CudaDevicePtr<float> gpu_sum_corr;
|
|
CudaDevicePtr<float> gpu_sum2_corr;
|
|
|
|
// Per-ring detection threshold (host-computed, uploaded) and the strong-pixel bit buffer.
|
|
CudaDevicePtr<float> gpu_thr;
|
|
CudaDevicePtr<uint32_t> gpu_strong;
|
|
|
|
// Host mirrors of the small per-ring transfers.
|
|
std::vector<unsigned long long> host_sum; // clipped raw sum } input to the host threshold computation
|
|
std::vector<unsigned long long> host_sum2; // clipped raw sum^2 } (exact integers - see the kernel)
|
|
std::vector<uint32_t> host_count; // clipped raw count }
|
|
std::vector<float> host_thr; // per-ring threshold (empty -> frame had no valid pixels)
|
|
std::vector<float> host_bkg; // clipped per-ring mean, NaN where the ring is too sparse to trust
|
|
std::vector<float> prof_sum; // plain corrected sum } azimuthal-integration profile
|
|
std::vector<float> prof_sum2; // plain corrected sum^2 }
|
|
std::vector<uint32_t> prof_count; // plain pixel count }
|
|
|
|
SpotExtractorGPU extractor; // builds the spots from gpu_strong without it leaving the device
|
|
|
|
AzimuthalIntegrationProfile last_profile; // filled every Run(), retrievable via GetProfile()
|
|
|
|
// One reduction pass over the image into the raw accumulators. clip_k <= 0 -> plain pass (all
|
|
// valid pixels); clip_k > 0 -> keep only pixels within clip_k sigma of the current gpu_mean.
|
|
// accumulate_corrected additionally fills gpu_sum_corr/gpu_sum2_corr for the profile (plain pass).
|
|
void ReducePass(const ImagePreprocessorBuffer &image, float clip_k, bool accumulate_corrected);
|
|
// Finalize gpu_mean/gpu_sigma from the current raw accumulators (per ring).
|
|
void FinalizeStats();
|
|
// Host: per-ring threshold from the clipped raw stats and the single knob E (false pixels/frame).
|
|
void ComputeThresholds(const SpotFindingSettings &settings);
|
|
|
|
public:
|
|
AdaptiveSpotFinderGPU(const AzimuthalIntegrationMapping &mapping, std::shared_ptr<CudaStream> stream);
|
|
~AdaptiveSpotFinderGPU() override = default;
|
|
AdaptiveSpotFinderGPU(const AdaptiveSpotFinderGPU &) = delete;
|
|
AdaptiveSpotFinderGPU &operator=(const AdaptiveSpotFinderGPU &) = delete;
|
|
|
|
void Detect(const ImagePreprocessorBuffer &image, const SpotFindingSettings &settings) override;
|
|
void SetResolutionMask(const std::vector<bool> &mask) override;
|
|
const std::vector<DiffractionSpot> &ExtractComponents(const ImagePreprocessorBuffer &image,
|
|
const SpotFindingSettings &settings) override;
|
|
|
|
// The azimuthal profile computed as a byproduct of the last Detect() - lets this engine replace the
|
|
// separate azint pass in the analysis pipeline.
|
|
[[nodiscard]] const AzimuthalIntegrationProfile &GetProfile() const { return last_profile; }
|
|
[[nodiscard]] const std::vector<float> &GetRingBackground() const override { return host_bkg; }
|
|
};
|