Build Packages / Create release (push) Successful in 24s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m29s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m43s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 8m27s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 9m53s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 9m58s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 11m22s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 13m39s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 18m37s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 16m32s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 24m11s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m30s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 19m3s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 20m23s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 19m41s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 21m0s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m38s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 14m33s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 17m55s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 20m50s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 18m38s
Build Packages / Unit tests (push) Successful in 1h46m14s
* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports. * jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls. * Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results. * Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable. * Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate. * Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do. * Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence. * Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags. * Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check. * Rugnux: Clear error messages when a data set needs more GPU or host memory than is available. Reviewed-on: #83 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
111 lines
6.4 KiB
C++
111 lines
6.4 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <vector>
|
|
|
|
#include "ImageSpotFinderCPU.h"
|
|
#include "SpotFindingSettings.h"
|
|
#include "../../common/AzimuthalIntegrationMapping.h"
|
|
#include "../../common/AzimuthalIntegrationProfile.h"
|
|
|
|
// Self-calibrating strong-pixel detector for the offline (rugnux/viewer) path.
|
|
//
|
|
// The classic finder (ImageSpotFinderCPU) marks a pixel strong when it clears a *fixed* photon
|
|
// count AND a local-box SNR. The fixed photon floor is what forces per-dataset tuning: it must sit
|
|
// above the background (wants high) yet not bury weak spots (wants low), and the background level
|
|
// differs per dataset, so the sweet spot is narrow (~12 photons on one serial-stills set, ~5 on a
|
|
// weaker one).
|
|
//
|
|
// Here the floor is replaced by a per-resolution-ring threshold derived from a single portable
|
|
// number: E = the expected count of noise pixels tolerated per frame (default ~100). For a ring
|
|
// whose (peak-excluded) background mean is mu, the threshold is the smallest count whose Poisson
|
|
// upper tail is <= p = E / N_pixels, max'd with a Gaussian arm mu + z*sigma to absorb read/flat-field
|
|
// excess. Because it is set from the image's own noise, the SAME E lands at ~12 photons on the first
|
|
// set and ~5 on the weaker one with no user input.
|
|
//
|
|
// The ring threshold replaces the floor and ONLY the floor: the classic finder's local-box SNR test
|
|
// still has to pass, which is why this engine runs it (ImageSpotFinderCPU) and intersects the two
|
|
// masks.
|
|
//
|
|
// A whole-ring threshold is an ABSOLUTE contour with no feedback from the pixel's own surroundings,
|
|
// so the area a spot puts above it grows as sigma^2 * ln(peak/threshold) and never saturates: on a
|
|
// strongly diffracting rotation set the detected footprint grows by 8 pixels per e-fold of peak, so
|
|
// the brightest reflections came out as 100-500 pixel blobs. The local box has no such contour. The
|
|
// spot inflates the box's own variance, and the peak divides out of the acceptance test, so the box
|
|
// cuts every spot at roughly a fixed FRACTION of its own height - a peak-relative contour. Measured
|
|
// on the same set, the footprint then grows by -0.2 pixels per e-fold, i.e. not at all, and lands on
|
|
// the classic finder's own number to two decimals.
|
|
//
|
|
// The two arms bind in different regimes, which is the point of intersecting rather than choosing.
|
|
// On serial stills the ring background is a fraction of a count and the ring threshold lands BELOW
|
|
// the fixed floor the classic finder would use, so the ring arm decides and the local box passes
|
|
// everything - which is the whole reason this engine exists. On a bright rotation set the ring
|
|
// background is tens of counts, the ring threshold lands several times ABOVE that floor, and the
|
|
// local box decides. The engine is therefore never worse than the classic finder on footprint, and
|
|
// never worse than a fixed floor on a weak background.
|
|
class AdaptiveSpotFinderCPU : public ImageSpotFinderCPU {
|
|
const AzimuthalIntegrationMapping &mapping;
|
|
|
|
// per-ring scratch, sized to the mapping's bin count
|
|
// Exact integers: a preprocessed pixel is an int32 and the sentinels are skipped, so v and v*v
|
|
// are exact in 64 bits. That is what lets the GPU engine reproduce these bit for bit - integer
|
|
// addition is associative, so its block atomics can arrive in any order.
|
|
std::vector<int64_t> ring_sum;
|
|
std::vector<uint64_t> ring_sum2;
|
|
std::vector<int64_t> ring_cnt;
|
|
std::vector<float> ring_mean;
|
|
std::vector<float> ring_sigma;
|
|
std::vector<float> ring_thr;
|
|
// ring_mean of the last Detect(), NaN where the ring holds too few pixels to be its own background.
|
|
// Kept separately because ring_mean carries the previous frame's value for an empty ring.
|
|
std::vector<float> ring_bkg;
|
|
// Pixels at or above their ring's threshold, packed like output_buffer. Intersected with the
|
|
// local-box mask that ImageSpotFinderCPU::Detect leaves in output_buffer.
|
|
std::vector<uint32_t> ring_bits;
|
|
// The plain pass's valid pixels as a per-ring histogram of their values (HIST_VALUES bins per
|
|
// ring) plus a list of the values outside it, so the two sigma-clip passes sum over distinct
|
|
// values instead of over the image again. Integer sums, so the same totals.
|
|
static constexpr int32_t HIST_VALUES = 1024;
|
|
std::vector<uint32_t> ring_hist;
|
|
std::vector<std::pair<uint16_t, int32_t>> ring_overflow; // (ring, value)
|
|
|
|
// The azimuthal-integration profile, taken in the plain ring pass when asked for: the
|
|
// same pixels in the same order and the same arithmetic as AzIntEngineCPU, so the same sums, and
|
|
// one pass over the image less (the CPU twin of the fused GPU engine).
|
|
bool fuse_azint = false;
|
|
std::vector<float> azint_sum;
|
|
std::vector<float> azint_sum2;
|
|
std::vector<uint32_t> azint_count;
|
|
|
|
// Set by BeginRings(): the plain ring pass of the next Detect() is being accumulated block by block.
|
|
bool rings_from_blocks = false;
|
|
|
|
// Zero the sums of the plain ring pass (and of the fused profile).
|
|
void ResetRings();
|
|
// One sigma-clip pass over the plain pass's values.
|
|
void ClipRings(float clip_k);
|
|
// ring_mean / ring_sigma from the current sums.
|
|
void UpdateRingStatistics();
|
|
// Set the ring_bits of the pixels of one row from the thresholds of the current frame. ring_bits
|
|
// is zeroed before the first row.
|
|
void FlagRow(const ImagePreprocessorBuffer &image, int32_t row);
|
|
|
|
public:
|
|
explicit AdaptiveSpotFinderCPU(const AzimuthalIntegrationMapping &mapping);
|
|
void Detect(const ImagePreprocessorBuffer &image, const SpotFindingSettings &settings) override;
|
|
[[nodiscard]] const std::vector<float> &GetRingBackground() const override { return ring_bkg; }
|
|
|
|
// The plain ring pass taken while the image is being preprocessed, so the pixels are read while
|
|
// still in cache: BeginRings(), then AccumulateRingsBlock() over every block of the image in pixel
|
|
// order - the order keeps the profile's float sums the same - and the next Detect() of that image
|
|
// starts from these sums instead of passing over the image for them.
|
|
void BeginRings();
|
|
void AccumulateRingsBlock(const ImagePreprocessorBuffer &image, size_t first, size_t n);
|
|
|
|
void FuseAzimuthalIntegration(bool enable) { fuse_azint = enable; }
|
|
// The profile of the last Detect(), when fused.
|
|
void GetProfile(AzimuthalIntegrationProfile &profile) const;
|
|
};
|