Build Packages / Create release (push) Successful in 24s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m29s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m43s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 8m27s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 9m53s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 9m58s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 11m22s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 13m39s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 18m37s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 16m32s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 24m11s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m30s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 19m3s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 20m23s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 19m41s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 21m0s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m38s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 14m33s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 17m55s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 20m50s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 18m38s
Build Packages / Unit tests (push) Successful in 1h46m14s
* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports. * jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls. * Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results. * Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable. * Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate. * Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do. * Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence. * Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags. * Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check. * Rugnux: Clear error messages when a data set needs more GPU or host memory than is available. Reviewed-on: #83 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
135 lines
7.8 KiB
C++
135 lines
7.8 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
// GPU adaptive spot finder that FUSES azimuthal integration and spot finding into one image pass.
|
|
//
|
|
// The CPU adaptive finder (AdaptiveSpotFinderCPU) and the azimuthal integrator both bin every pixel
|
|
// into resolution rings and reduce (sum / sum^2 / count). Today azint runs on the GPU while the
|
|
// adaptive finder re-does the identical per-ring reduction on the HOST - a wasted second pass over a
|
|
// ~10 MP image. This engine does the ring reduction on the GPU and drives BOTH products from it:
|
|
// - the azimuthal-integration profile (mean intensity per ring, in flat-field-corrected space), and
|
|
// - the per-ring background (mean, sigma, peak-excluded via two sigma-clip passes) that sets the
|
|
// self-calibrating spot-detection threshold (in raw photon counts).
|
|
// The ring threshold is the photon-count FLOOR and nothing more: this engine derives from the classic
|
|
// ImageSpotFinderGPU and intersects the ring mask with that engine's local-box SNR mask, exactly as
|
|
// AdaptiveSpotFinderCPU does on the host (the reason the local test is not optional is written out
|
|
// there). The result is left in the inherited bit buffer, still on the device, and the inherited
|
|
// SpotExtractorGPU builds the spots from it.
|
|
//
|
|
// Numerically it reproduces AdaptiveSpotFinderCPU: the same three-pass robust background, the same
|
|
// per-ring threshold formula (shared via AdaptiveThreshold.h, computed on the host once per frame),
|
|
// and the same intersection with the local test. The only differences from the CPU are those inherent
|
|
// to a GPU reduction (float per-ring accumulation in atomic order vs the CPU's serial double sums),
|
|
// which shift a handful of borderline pixels at most. The corrected sums for the azint profile are
|
|
// accumulated in the SAME plain first pass, so one reduction feeds both products.
|
|
|
|
#include <memory>
|
|
#include <vector>
|
|
|
|
#include "ImageSpotFinderGPU.h"
|
|
#include "SpotExtractorGPU.h"
|
|
#include "SpotFindingSettings.h"
|
|
#include "../../common/AzimuthalIntegrationProfile.h"
|
|
#include "../../common/AzimuthalIntegrationMapping.h"
|
|
#include "../indexing/CUDAMemHelpers.h"
|
|
#include "../indexing/CudaSharedTables.h"
|
|
|
|
class AdaptiveSpotFinderGPU : public ImageSpotFinderGPU {
|
|
const AzimuthalIntegrationMapping &mapping;
|
|
|
|
const int nbins;
|
|
const size_t npix;
|
|
|
|
int reduce_threads = 256;
|
|
int reduce_blocks = 0; // global-atomics fallback
|
|
int reduce_blocks_plain = 0; // as many blocks as actually fit, per shared-memory footprint
|
|
int reduce_blocks_clip = 0;
|
|
int flag_threads = 256;
|
|
int flag_blocks = 0;
|
|
size_t shared_plain = 0; // per-block shared bytes for the plain pass (raw + corrected rings)
|
|
size_t shared_clip = 0; // per-block shared bytes for a sigma-clip pass (raw rings only)
|
|
bool use_shared = true; // false -> nbins too large for shared memory, use the global-atomics kernel
|
|
|
|
// Static mapping inputs: geometry-only, so one copy per GPU shared with every other engine on it
|
|
// (see CudaSharedTables.h) rather than one copy per worker thread.
|
|
std::shared_ptr<CudaDevicePtr<uint16_t>> gpu_pixel_to_bin;
|
|
std::shared_ptr<CudaDevicePtr<float>> gpu_corrections;
|
|
|
|
// Raw per-ring accumulators (re-zeroed each pass) + derived stats used to clip and threshold.
|
|
// double, like the CPU engine's ring accumulators: the ring sigma is the cancelling difference
|
|
// sum2/n - m^2, and the block atomics that fill these arrive in an arbitrary order.
|
|
CudaDevicePtr<unsigned long long> gpu_sum;
|
|
CudaDevicePtr<unsigned long long> gpu_sum2;
|
|
CudaDevicePtr<uint32_t> gpu_count;
|
|
CudaDevicePtr<float> gpu_mean; // per-ring raw mean (clip predicate)
|
|
CudaDevicePtr<float> gpu_sigma; // per-ring raw sigma (clip predicate)
|
|
|
|
// The plain pass's valid pixels as a per-ring histogram of their values, and the (ring, value) of
|
|
// those outside [0, HIST_VALUES) - so the sigma-clip passes need not re-read the image. A list
|
|
// longer than overflow_cap is only counted, and the clip passes then read the image instead.
|
|
// Allocated only on the shared-memory path (few rings).
|
|
CudaDevicePtr<uint32_t> gpu_hist;
|
|
CudaDevicePtr<int2> gpu_overflow;
|
|
CudaDevicePtr<uint32_t> gpu_overflow_n;
|
|
uint32_t overflow_cap = 0;
|
|
|
|
// Corrected per-ring accumulators (plain first pass only) -> azimuthal-integration profile.
|
|
CudaDevicePtr<float> gpu_sum_corr;
|
|
CudaDevicePtr<float> gpu_sum2_corr;
|
|
|
|
// Per-ring detection threshold (host-computed, uploaded) and the ring-threshold mask that is
|
|
// intersected into the inherited bit buffer.
|
|
CudaDevicePtr<float> gpu_thr;
|
|
CudaDevicePtr<uint32_t> gpu_ring;
|
|
|
|
// Host mirrors of the small per-ring transfers.
|
|
std::vector<unsigned long long> host_sum; // clipped raw sum } input to the host threshold computation
|
|
std::vector<unsigned long long> host_sum2; // clipped raw sum^2 } (exact integers - see the kernel)
|
|
std::vector<uint32_t> host_count; // clipped raw count }
|
|
std::vector<float> host_thr; // per-ring threshold (empty -> frame had no valid pixels)
|
|
std::vector<float> host_bkg; // clipped per-ring mean, NaN where the ring is too sparse to trust
|
|
std::vector<float> prof_sum; // plain corrected sum } azimuthal-integration profile
|
|
std::vector<float> prof_sum2; // plain corrected sum^2 }
|
|
std::vector<uint32_t> prof_count; // plain pixel count }
|
|
|
|
// Every per-ring array above is a device-to-host copy once per frame. A D2H copy into PAGEABLE
|
|
// memory blocks the host until it completes, whatever stream it was issued on - which would stall
|
|
// Detect() between the plain pass and the clip passes, with the device then idle while the host
|
|
// enqueues them. Pinning the destinations makes the copies genuinely asynchronous, as the
|
|
// azimuthal-integration engine already does with its own.
|
|
CudaRegisteredVector<unsigned long long> host_sum_reg;
|
|
CudaRegisteredVector<unsigned long long> host_sum2_reg;
|
|
CudaRegisteredVector<uint32_t> host_count_reg;
|
|
CudaRegisteredVector<float> prof_sum_reg;
|
|
CudaRegisteredVector<float> prof_sum2_reg;
|
|
CudaRegisteredVector<uint32_t> prof_count_reg;
|
|
|
|
AzimuthalIntegrationProfile last_profile; // filled every Run(), retrievable via GetProfile()
|
|
|
|
// One reduction pass over the image into the raw accumulators. clip_k <= 0 -> plain pass (all
|
|
// valid pixels); clip_k > 0 -> keep only pixels within clip_k sigma of the current gpu_mean.
|
|
// accumulate_corrected additionally fills gpu_sum_corr/gpu_sum2_corr for the profile (plain pass).
|
|
void ReducePass(const ImagePreprocessorBuffer &image, float clip_k, bool accumulate_corrected);
|
|
// Finalize gpu_mean/gpu_sigma from the current raw accumulators (per ring).
|
|
void FinalizeStats();
|
|
// Host: per-ring threshold from the clipped raw stats and the single knob E (false pixels/frame).
|
|
void ComputeThresholds(const SpotFindingSettings &settings);
|
|
|
|
public:
|
|
static constexpr int32_t HIST_VALUES = 1024; // as AdaptiveSpotFinderCPU
|
|
|
|
AdaptiveSpotFinderGPU(const AzimuthalIntegrationMapping &mapping, std::shared_ptr<CudaStream> stream);
|
|
~AdaptiveSpotFinderGPU() override = default;
|
|
AdaptiveSpotFinderGPU(const AdaptiveSpotFinderGPU &) = delete;
|
|
AdaptiveSpotFinderGPU &operator=(const AdaptiveSpotFinderGPU &) = delete;
|
|
|
|
void Detect(const ImagePreprocessorBuffer &image, const SpotFindingSettings &settings) override;
|
|
|
|
// The azimuthal profile computed as a byproduct of the last Detect() - lets this engine replace the
|
|
// separate azint pass in the analysis pipeline.
|
|
[[nodiscard]] const AzimuthalIntegrationProfile &GetProfile() const { return last_profile; }
|
|
[[nodiscard]] const std::vector<float> &GetRingBackground() const override { return host_bkg; }
|
|
};
|