Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 8m41s
Build Packages / build:windows:nocuda (push) Successful in 16m50s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 18m22s
Build Packages / build:windows:cuda (push) Successful in 19m40s
Build Packages / build:viewer-tgz:cpu (push) Successful in 21m2s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m43s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m21s
Build Packages / build:rugnux:windows (push) Successful in 10m45s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 27m37s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 27m50s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 19m52s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 22m6s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 25m47s
Build Packages / build:rpm (rocky9) (push) Successful in 23m52s
Build Packages / build:rpm (rocky8) (push) Successful in 26m33s
Build Packages / Generate python client (push) Successful in 45s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2404) (push) Successful in 23m57s
Build Packages / DIALS test (push) Successful in 25m7s
Build Packages / XDS test (durin plugin) (push) Successful in 11m18s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 27m24s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m58s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m43s
Build Packages / Unit tests (push) Successful in 1h22m26s
The self-calibrating finder was meant to replace the classic finder's FIXED PHOTON FLOOR with a per-resolution-ring threshold read off the image's own noise. As written it replaced the local-box SNR test as well, and that is the defect: a whole-ring threshold is an ABSOLUTE contour with no feedback from a pixel's own surroundings, so the area a spot puts above it grows as sigma^2*ln(peak/threshold) and never saturates. Measured on a strongly diffracting rotation set, the detected footprint grows by +8.05 pixels per e-fold of peak, so the brightest reflections came out as 100-500 pixel blobs and were then discarded for exceeding the size bound - every one of the ten strongest on an image. Intersecting with the local box gives -0.24 pixels per e-fold, the classic finder's own number to two decimals. WHY the local box is the right partner, rather than merely the incumbent: it is a prominence rule whose reference level is a 961-pixel mean. A spot inflates the box's own variance and the peak divides out of the acceptance test, so it cuts at a fixed FRACTION of the spot's own height. Referring that level to fewer pixels makes it inherit their shot noise - at FIXED footprint, estimating the level from 961 pixels, from 25, and from the single maximum gives centroid residuals of 0.524, 0.539 and 0.656 - so flat growth and a stable centroid turn out to be two ends of one dial. A contour on the bare maximum has the flattest growth of anything tried (+0.1) and merges worst. The two arms bind in different regimes, which is why intersecting beats choosing: on serial stills the ring threshold is 0.6x the classic floor, on this rotation sweep 2.3-6.0x. Stills are a strict no-op - 175 components against 175, identical per frame - so the +40% in stills indexing that the adaptive threshold was introduced for is untouched. What it buys, stated as one fact rather than two. Across five geometry pins spanning 1.1 mm it indexes the most frames of any arm tried, 0.831 against 0.803, and integrates 3.04 to 5.76% more observations - but those are the SAME number: regressing observation count on indexing rate over four arms leaves residuals of +/-0.7 percentage points against swings of -7 to +4.5%, so the extra observations ARE the extra indexed frames, not better data per frame. CC1/2, the only statistic here carrying per-observation quality, is +0.66 at one pin and -0.06 at the other: not harmed, not improved. <I/sigma>, ISa and R_meas cannot arbitrate on this data - across those pins each crosses zero as a monotone function of the pin. WHY an absolute contour indexes fewer frames, when its spot list is equal or better on every axis measured - recall, top-1000 recall, centroid, ice fraction, component count - is the interesting part, and it is not a detection effect at all: ITS OWN SIZE BOUND DELETES THE BRIGHTEST REFLECTIONS ON THE FRAME. A component is discarded because it grew past 200 px, and it grew past 200 px because it was bright, so the deletions are drawn from the head of the indexing budget rather than uniformly from it: they are 11x enriched in the top 250 of the thousand spots handed to the indexer, and the bound's own real deletions sit at MEDIAN RANK 12. Turning the bound off recovers 66% and 50% of the deficit at the two pins, against a bar registered at 33% before the run. Three of us dismissed this for most of a day on the grounds that the gates delete only ~4% of what is detected. That arithmetic was right and the denominator was wrong - a rate is not an impact when the thing being lost is selected for the property that makes it matter. Reworking the bound instead was measured and rejected: it recovers half the deficit, and it cannot be done without re-admitting what the bound is for - 68 components past 200 px, of which 8 are real and 60 are junk, where the intersect gets the 8 without the 60. The residual once the bound is off, +1.08%/+1.70%, is the contour itself. Component merging is ruled out separately: geometrically impossible here, 33.9 px minimum reflection separation against components spanning 10 px. So is a ranking effect - the intersect's lead runs +0.06% at --max-spots 250, +3.46% at 1000 and +14.26% at 2000, which is backwards for a selection artefact. Costs 0.48 ms per image in the finder, and 0.044 px of bright-spot centroid precision - measured convention-free, by fitting a line to a reflection's own centroid across five frames, after an XDS-referenced figure proved to be four fifths aperture convention. It also makes the compactness gate above it safe. On the absolute contour that gate is net damage, deleting 37 genuine reflections per ten frames; once the footprint stops growing nothing reaches its threshold at all. Also fixes a real but unexercised defect in PoissonThreshold, where the exact tail handed over to a normal approximation with a step. It changes nothing here: the clipped ring sigma is over-dispersed 1.2-4.9x against sqrt(mu) because it still contains diffraction, so the Gaussian arm wins every ring above mu=50 and none of the 522 thresholds move. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FBumeJVx4oeXxiBRpkrE5H
124 lines
7.2 KiB
C++
124 lines
7.2 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
// GPU adaptive spot finder that FUSES azimuthal integration and spot finding into one image pass.
|
|
//
|
|
// The CPU adaptive finder (AdaptiveSpotFinderCPU) and the azimuthal integrator both bin every pixel
|
|
// into resolution rings and reduce (sum / sum^2 / count). Today azint runs on the GPU while the
|
|
// adaptive finder re-does the identical per-ring reduction on the HOST - a wasted second pass over a
|
|
// ~10 MP image. This engine does the ring reduction on the GPU and drives BOTH products from it:
|
|
// - the azimuthal-integration profile (mean intensity per ring, in flat-field-corrected space), and
|
|
// - the per-ring background (mean, sigma, peak-excluded via two sigma-clip passes) that sets the
|
|
// self-calibrating spot-detection threshold (in raw photon counts).
|
|
// The ring threshold is the photon-count FLOOR and nothing more: this engine derives from the classic
|
|
// ImageSpotFinderGPU and intersects the ring mask with that engine's local-box SNR mask, exactly as
|
|
// AdaptiveSpotFinderCPU does on the host (the reason the local test is not optional is written out
|
|
// there). The result is left in the inherited bit buffer, still on the device, and the inherited
|
|
// SpotExtractorGPU builds the spots from it.
|
|
//
|
|
// Numerically it reproduces AdaptiveSpotFinderCPU: the same three-pass robust background, the same
|
|
// per-ring threshold formula (shared via AdaptiveThreshold.h, computed on the host once per frame),
|
|
// and the same intersection with the local test. The only differences from the CPU are those inherent
|
|
// to a GPU reduction (float per-ring accumulation in atomic order vs the CPU's serial double sums),
|
|
// which shift a handful of borderline pixels at most. The corrected sums for the azint profile are
|
|
// accumulated in the SAME plain first pass, so one reduction feeds both products.
|
|
|
|
#include <memory>
|
|
#include <vector>
|
|
|
|
#include "ImageSpotFinderGPU.h"
|
|
#include "SpotExtractorGPU.h"
|
|
#include "SpotFindingSettings.h"
|
|
#include "../../common/AzimuthalIntegrationProfile.h"
|
|
#include "../../common/AzimuthalIntegrationMapping.h"
|
|
#include "../indexing/CUDAMemHelpers.h"
|
|
#include "../indexing/CudaSharedTables.h"
|
|
|
|
class AdaptiveSpotFinderGPU : public ImageSpotFinderGPU {
|
|
const AzimuthalIntegrationMapping &mapping;
|
|
|
|
const int nbins;
|
|
const size_t npix;
|
|
|
|
int reduce_threads = 256;
|
|
int reduce_blocks = 0; // global-atomics fallback
|
|
int reduce_blocks_plain = 0; // as many blocks as actually fit, per shared-memory footprint
|
|
int reduce_blocks_clip = 0;
|
|
int flag_threads = 256;
|
|
int flag_blocks = 0;
|
|
size_t shared_plain = 0; // per-block shared bytes for the plain pass (raw + corrected rings)
|
|
size_t shared_clip = 0; // per-block shared bytes for a sigma-clip pass (raw rings only)
|
|
bool use_shared = true; // false -> nbins too large for shared memory, use the global-atomics kernel
|
|
|
|
// Static mapping inputs: geometry-only, so one copy per GPU shared with every other engine on it
|
|
// (see CudaSharedTables.h) rather than one copy per worker thread.
|
|
std::shared_ptr<CudaDevicePtr<uint16_t>> gpu_pixel_to_bin;
|
|
std::shared_ptr<CudaDevicePtr<float>> gpu_corrections;
|
|
|
|
// Raw per-ring accumulators (re-zeroed each pass) + derived stats used to clip and threshold.
|
|
// double, like the CPU engine's ring accumulators: the ring sigma is the cancelling difference
|
|
// sum2/n - m^2, and the block atomics that fill these arrive in an arbitrary order.
|
|
CudaDevicePtr<unsigned long long> gpu_sum;
|
|
CudaDevicePtr<unsigned long long> gpu_sum2;
|
|
CudaDevicePtr<uint32_t> gpu_count;
|
|
CudaDevicePtr<float> gpu_mean; // per-ring raw mean (clip predicate)
|
|
CudaDevicePtr<float> gpu_sigma; // per-ring raw sigma (clip predicate)
|
|
|
|
// Corrected per-ring accumulators (plain first pass only) -> azimuthal-integration profile.
|
|
CudaDevicePtr<float> gpu_sum_corr;
|
|
CudaDevicePtr<float> gpu_sum2_corr;
|
|
|
|
// Per-ring detection threshold (host-computed, uploaded) and the ring-threshold mask that is
|
|
// intersected into the inherited bit buffer.
|
|
CudaDevicePtr<float> gpu_thr;
|
|
CudaDevicePtr<uint32_t> gpu_ring;
|
|
|
|
// Host mirrors of the small per-ring transfers.
|
|
std::vector<unsigned long long> host_sum; // clipped raw sum } input to the host threshold computation
|
|
std::vector<unsigned long long> host_sum2; // clipped raw sum^2 } (exact integers - see the kernel)
|
|
std::vector<uint32_t> host_count; // clipped raw count }
|
|
std::vector<float> host_thr; // per-ring threshold (empty -> frame had no valid pixels)
|
|
std::vector<float> host_bkg; // clipped per-ring mean, NaN where the ring is too sparse to trust
|
|
std::vector<float> prof_sum; // plain corrected sum } azimuthal-integration profile
|
|
std::vector<float> prof_sum2; // plain corrected sum^2 }
|
|
std::vector<uint32_t> prof_count; // plain pixel count }
|
|
|
|
// Every per-ring array above is a device-to-host copy once per frame. A D2H copy into PAGEABLE
|
|
// memory blocks the host until it completes, whatever stream it was issued on - which would stall
|
|
// Detect() between the plain pass and the clip passes, with the device then idle while the host
|
|
// enqueues them. Pinning the destinations makes the copies genuinely asynchronous, as the
|
|
// azimuthal-integration engine already does with its own.
|
|
CudaRegisteredVector<unsigned long long> host_sum_reg;
|
|
CudaRegisteredVector<unsigned long long> host_sum2_reg;
|
|
CudaRegisteredVector<uint32_t> host_count_reg;
|
|
CudaRegisteredVector<float> prof_sum_reg;
|
|
CudaRegisteredVector<float> prof_sum2_reg;
|
|
CudaRegisteredVector<uint32_t> prof_count_reg;
|
|
|
|
AzimuthalIntegrationProfile last_profile; // filled every Run(), retrievable via GetProfile()
|
|
|
|
// One reduction pass over the image into the raw accumulators. clip_k <= 0 -> plain pass (all
|
|
// valid pixels); clip_k > 0 -> keep only pixels within clip_k sigma of the current gpu_mean.
|
|
// accumulate_corrected additionally fills gpu_sum_corr/gpu_sum2_corr for the profile (plain pass).
|
|
void ReducePass(const ImagePreprocessorBuffer &image, float clip_k, bool accumulate_corrected);
|
|
// Finalize gpu_mean/gpu_sigma from the current raw accumulators (per ring).
|
|
void FinalizeStats();
|
|
// Host: per-ring threshold from the clipped raw stats and the single knob E (false pixels/frame).
|
|
void ComputeThresholds(const SpotFindingSettings &settings);
|
|
|
|
public:
|
|
AdaptiveSpotFinderGPU(const AzimuthalIntegrationMapping &mapping, std::shared_ptr<CudaStream> stream);
|
|
~AdaptiveSpotFinderGPU() override = default;
|
|
AdaptiveSpotFinderGPU(const AdaptiveSpotFinderGPU &) = delete;
|
|
AdaptiveSpotFinderGPU &operator=(const AdaptiveSpotFinderGPU &) = delete;
|
|
|
|
void Detect(const ImagePreprocessorBuffer &image, const SpotFindingSettings &settings) override;
|
|
|
|
// The azimuthal profile computed as a byproduct of the last Detect() - lets this engine replace the
|
|
// separate azint pass in the analysis pipeline.
|
|
[[nodiscard]] const AzimuthalIntegrationProfile &GetProfile() const { return last_profile; }
|
|
[[nodiscard]] const std::vector<float> &GetRingBackground() const override { return host_bkg; }
|
|
};
|