One changeset, developed together in response to a review of this branch, so the files carry several of the changes at once. Full test suite passes (733 cases). Spot finding - Split ImageSpotFinder into Detect() (flag strong pixels - the expensive per-pixel pass) and ExtractSpots() (CCL + min/max-pix + resolution mask), with Run() = both. The per-image min-pix escalation now detects ONCE and repeats only the cheap extraction, instead of re-running the whole finder four times per frame as it did on the default path. It also keeps the winning attempt's spot list rather than re-extracting it, so the frame that is integrated is exactly the frame that was scored - which a GPU re-extract could not guarantee (float atomic ordering). - spot_finding_time_s no longer swallows indexing time, and indexing_time_s now sums every escalation call instead of reporting only the last. Detection limits follow the detector - The azimuthal-integration upper q and the spot-finding high-resolution limit are now std::optional, in the C++ structs AND in the OpenAPI schema, and resolve to the detector's own maximum (DiffractionExperiment::GetDetectorMaxQ_ recipA). Adaptive detection reads a pixel's ring from the azimuthal bins, so a pixel outside that q range could never be strong - the integration range silently bounded what detection could see, regardless of the requested resolution limit. Regenerated the C++ and TypeScript clients; the viewer and the web frontend each gained a "to detector edge" switch. Detection defaults are now per workflow (measured, not assumed) - Stills: adaptive detection, min-pix chosen per image, no resolution clipping. - Rotation: fixed-threshold finder, min-pix 2, 1.5 A limit. On a 33-crystal rotation battery, adaptive detection helped four hard crystals but deterministically broke three (a lost space group, a halved indexing rate, a collapsed merge), and the detector-edge limit cost indexing on a strong rotation set (100.0 -> 96.8%). Each is still overridable by its flag, and --no-adaptive-spots is new. Indexer seed escalation - Stop escalating once a seed's lattice explains >= 90% of the seed spots. Previously any frame with >= 80 spots always paid three indexer calls, online broker included. Merge-consistency filter - --min-image-cc gated on a per-image CC computed BEFORE the stills partiality post-refinement and never refreshed; the refiner now recomputes it, so the reported CC describes the data that are actually merged. - Replaced the per-call cc_mask argument with one MergeOnTheFly flag, so the merge, the error model and MergeStats can no longer disagree about which images are in (the --scale path merged unfiltered while its statistics were filtered). Per-image B-factor refinement (-B) removed - Measured on four serial-stills datasets: it is a no-op where the per-image fit is well conditioned and actively harmful where it is not (CC1/2 -8.1, R_meas +23.2 on the weakest large-cell set, whose fits hit their [-50, 200] bounds on 14-25% of images). It had also been silently DISCARDED since the partiality post-refinement landed - reported but not applied. Rather than fix and keep a knob with no demonstrated benefit, the flag and the whole image_scale_b_factor chain are gone: setting, scaling fit, message field, CBOR, HDF5 write and read-back, per-image plot, OpenAPI enum, viewer column and checkbox, docs. ScaleOnTheFly no longer needs Ceres at all - the fit is a linear IRLS. (The Wilson per-image b_factor is a different quantity and stays.) Stills partiality width now fits both of its components - sigma^2 = gamma0^2 + (gamma_e*d*)^2 instead of a purely angular gamma_e*d* with gamma0 pinned to 0. Fitted per crystal by least squares of dist_ewald^2 on d*^2. The angular-only width is fitted over a d*^2-dense population, so it was pinned by the high-resolution edge and collapsed at low d*: median partiality 0.008 beyond 13 A for reflections that were plainly recorded, 55% of them under the merge's partiality floor, and the survivors divided by those values - which inflated the merged low-resolution intensity scale 3.6x (~ +9 A^2 of apparent B). Measured on 5000 stills: the ramp flattens to 0.89x, no observation is dropped any more (701750 -> 716811), shell-mean CC1/2 and R-free improve slightly. Note CC1/2, R_meas, completeness and a B-refining R-free are all blind to that ramp, which is why it survived earlier validation; the cost is high-resolution R_meas (98.5 -> 101.9 shell-averaged). Removed dead code from add-then-remove churn - Prediction-time "still partiality" (unreachable: no setter), the phantom IndexingSettings::min_indexed_spot_fraction knob (getter, no setter - now the constant it always was), StillsPartialityRefine's caller-less Settings constructor and its reference to a long-gone env var, ProcessImage's unread bool return, an unused include, and a dead viewer overlay hook. Also - Viewer: the magnifier compared a QImage with itself, so its scene rect was set once ever and it could not pan into a larger dataset; the hover tail timer could fire after leaveEvent and resurrect the resolution readout outside the image. - update_version.sh regenerated the frontend lock file BEFORE bumping the version (every release shipped an off-by-one lock), and did git rm/git add on a path that has not existed since the client moved to src/client - with no set -e, both failed silently. - fpga/pcie_driver/postinstall.sh tested "[ ! occurrences > 0 ]", which is a redirect, not a test, so dkms add never ran. - Unit tests for the adaptive-threshold host functions, which had none. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
102 lines
5.4 KiB
C++
102 lines
5.4 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
// GPU adaptive spot finder that FUSES azimuthal integration and spot finding into one image pass.
|
|
//
|
|
// The CPU adaptive finder (AdaptiveSpotFinderCPU) and the azimuthal integrator both bin every pixel
|
|
// into resolution rings and reduce (sum / sum^2 / count). Today azint runs on the GPU while the
|
|
// adaptive finder re-does the identical per-ring reduction on the HOST - a wasted second pass over a
|
|
// ~10 MP image. This engine does the ring reduction on the GPU and drives BOTH products from it:
|
|
// - the azimuthal-integration profile (mean intensity per ring, in flat-field-corrected space), and
|
|
// - the per-ring background (mean, sigma, peak-excluded via two sigma-clip passes) that sets the
|
|
// self-calibrating spot-detection threshold (in raw photon counts).
|
|
// It then flags strong pixels (value >= ring threshold) into a packed bit buffer and hands it to the
|
|
// shared host connected-component extractor (ImageSpotFinder::ExtractSpots).
|
|
//
|
|
// Numerically it reproduces AdaptiveSpotFinderCPU: the same three-pass robust background, the same
|
|
// per-ring threshold formula (shared via AdaptiveThreshold.h, computed on the host once per frame),
|
|
// and the same raw-count detection test. The only differences from the CPU are those inherent to a
|
|
// GPU reduction (float per-ring accumulation in atomic order vs the CPU's serial double sums), which
|
|
// shift a handful of borderline pixels at most. The corrected sums for the azint profile are
|
|
// accumulated in the SAME plain first pass, so one reduction feeds both products.
|
|
|
|
#include <memory>
|
|
#include <vector>
|
|
|
|
#include "ImageSpotFinder.h"
|
|
#include "SpotFindingSettings.h"
|
|
#include "../../common/AzimuthalIntegrationProfile.h"
|
|
#include "../../common/AzimuthalIntegrationMapping.h"
|
|
#include "../indexing/CUDAMemHelpers.h"
|
|
|
|
class AdaptiveSpotFinderGPU : public ImageSpotFinder {
|
|
const AzimuthalIntegrationMapping &mapping;
|
|
std::shared_ptr<CudaStream> stream;
|
|
|
|
const int nbins;
|
|
const size_t npix;
|
|
|
|
int reduce_threads = 128;
|
|
int reduce_blocks = 0;
|
|
int flag_threads = 256;
|
|
int flag_blocks = 0;
|
|
size_t shared_plain = 0; // per-block shared bytes for the plain pass (raw + corrected rings)
|
|
size_t shared_clip = 0; // per-block shared bytes for a sigma-clip pass (raw rings only)
|
|
bool use_shared = true; // false -> nbins too large for shared memory, use the global-atomics kernel
|
|
|
|
// Static mapping inputs (uploaded once).
|
|
CudaDevicePtr<uint16_t> gpu_pixel_to_bin;
|
|
CudaDevicePtr<float> gpu_corrections;
|
|
|
|
// Raw per-ring accumulators (re-zeroed each pass) + derived stats used to clip and threshold.
|
|
CudaDevicePtr<float> gpu_sum;
|
|
CudaDevicePtr<float> gpu_sum2;
|
|
CudaDevicePtr<uint32_t> gpu_count;
|
|
CudaDevicePtr<float> gpu_mean; // per-ring raw mean (clip predicate)
|
|
CudaDevicePtr<float> gpu_sigma; // per-ring raw sigma (clip predicate)
|
|
|
|
// Corrected per-ring accumulators (plain first pass only) -> azimuthal-integration profile.
|
|
CudaDevicePtr<float> gpu_sum_corr;
|
|
CudaDevicePtr<float> gpu_sum2_corr;
|
|
|
|
// Per-ring detection threshold (host-computed, uploaded) and the strong-pixel bit buffer.
|
|
CudaDevicePtr<float> gpu_thr;
|
|
CudaDevicePtr<uint32_t> gpu_strong;
|
|
|
|
// Host mirrors of the small per-ring transfers.
|
|
std::vector<float> host_sum; // clipped raw sum } input to the host threshold computation
|
|
std::vector<float> host_sum2; // clipped raw sum^2 }
|
|
std::vector<uint32_t> host_count; // clipped raw count }
|
|
std::vector<float> host_thr; // per-ring threshold (empty -> frame had no valid pixels)
|
|
std::vector<float> prof_sum; // plain corrected sum } azimuthal-integration profile
|
|
std::vector<float> prof_sum2; // plain corrected sum^2 }
|
|
std::vector<uint32_t> prof_count; // plain pixel count }
|
|
|
|
CudaRegisteredVector<uint32_t> output_buffer_reg; // pins the base-class bit buffer for fast D2H
|
|
|
|
AzimuthalIntegrationProfile last_profile; // filled every Run(), retrievable via GetProfile()
|
|
|
|
// One reduction pass over the image into the raw accumulators. clip_k <= 0 -> plain pass (all
|
|
// valid pixels); clip_k > 0 -> keep only pixels within clip_k sigma of the current gpu_mean.
|
|
// accumulate_corrected additionally fills gpu_sum_corr/gpu_sum2_corr for the profile (plain pass).
|
|
void ReducePass(const ImagePreprocessorBuffer &image, float clip_k, bool accumulate_corrected);
|
|
// Finalize gpu_mean/gpu_sigma from the current raw accumulators (per ring).
|
|
void FinalizeStats();
|
|
// Host: per-ring threshold from the clipped raw stats and the single knob E (false pixels/frame).
|
|
void ComputeThresholds(const SpotFindingSettings &settings);
|
|
|
|
public:
|
|
AdaptiveSpotFinderGPU(const AzimuthalIntegrationMapping &mapping, std::shared_ptr<CudaStream> stream);
|
|
~AdaptiveSpotFinderGPU() override = default;
|
|
AdaptiveSpotFinderGPU(const AdaptiveSpotFinderGPU &) = delete;
|
|
AdaptiveSpotFinderGPU &operator=(const AdaptiveSpotFinderGPU &) = delete;
|
|
|
|
void Detect(const ImagePreprocessorBuffer &image, const SpotFindingSettings &settings) override;
|
|
|
|
// The azimuthal profile computed as a byproduct of the last Detect() - lets this engine replace the
|
|
// separate azint pass in the analysis pipeline.
|
|
[[nodiscard]] const AzimuthalIntegrationProfile &GetProfile() const { return last_profile; }
|
|
};
|