Files
Jungfraujoch/rugnux/HotPixels.h
T
leonarski_f 84228bf8be
Build Packages / Create release (push) Successful in 24s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m29s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m43s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 8m27s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 9m53s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 9m58s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 11m22s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 13m39s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 18m37s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 16m32s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 24m11s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m30s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 19m3s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 20m23s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 19m41s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 21m0s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m38s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 14m33s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 17m55s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 20m50s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 18m38s
Build Packages / Unit tests (push) Successful in 1h46m14s
v1.0.0-rc.173 (#83)
* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports.
* jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls.
* Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results.
* Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable.
* Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate.
* Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do.
* Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence.
* Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags.
* Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check.
* Rugnux: Clear error messages when a data set needs more GPU or host memory than is available.

Reviewed-on: #83
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-09-29 15:57:32 +02:00

146 lines
8.1 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
// =============================================================================
// HotPixelFinder - pixels the pre-scan's own frames show to be defective
// =============================================================================
//
// A pixel that reads high on frame after frame, whatever the crystal is doing, is not recording
// diffraction. Unmasked, it is integrated as part of whichever reflection's box it falls in, and
// under rotation that is one pixel at one resolution collecting a different reflection on every
// frame that reaches it - a merged intensity hundreds or thousands of times its shell mean,
// supported by one or two observations the merge's outlier test cannot judge. The file's own
// mask misses such pixels on many detectors, and which pixels are hot moves with the threshold
// setting, so they are measured per dataset, on the pre-scan's sample of frames.
//
// On each frame a pixel is LIT when it is Poisson-significantly above the level of its own
// resolution ring - the larger of the median of the 2 px iso-2theta ring and of the ring's
// sixteenth in azimuth, so an ice or powder ring, the polarization dip and a partly shadowed ring
// all set their own level. It is PERSISTENT when it is lit on more of the sampled frames than
// either explanation that is not a defect allows:
//
// * one Bragg reflection: it stays on a pixel for (oscillation + rocking width) / |zeta| of
// rotation, i.e. on at most 1 + ceil(that / frame spacing) of frames sampled that far apart;
// * chance: independent reflections and noise light a pixel of this ring on a fraction q of the
// frames (measured, on the ring's pixels that are not candidates), and the number of lit frames
// that fewer than 0.01 of all the detector's pixels would reach by chance is the binomial bound.
//
// A persistent pixel is masked only if it stands alone (by itself or in a pair - a larger patch is a
// feature of the scattering, not a defect), reads on average at least ten times its ring, and its
// mean excess is above the Poisson bound: a weaker one cannot make an outlier. A pixel holding the
// detector's error value on most frames is masked with it: every frame already treats it as invalid,
// but the static mask did not know. The detector's outermost row and column get no rule of their
// own - a hot pixel there is caught like any other.
//
// None of this holds for a CCD, whose pixels carry a read-out offset and are not Poisson; the caller
// runs it for counting sensors only.
// =============================================================================
#include <atomic>
#include <cstdint>
#include <memory>
#include <mutex>
#include <vector>
#include "../common/DiffractionExperiment.h"
#include "../common/PixelMask.h"
#ifdef JFJOCH_USE_CUDA
#include "HotPixelsGPU.h"
#endif
class HotPixelFinder {
public:
// Sixteen sectors of the ring; the level is the larger of the two medians.
static constexpr int SECTORS = 16;
// Ring width in 2theta: this many pixels where the beam meets the detector, wider further out.
static constexpr float RING_WIDTH_PX = 2.0f;
// A ring needs this many valid pixels to have a level, a sector this many to have one of its own.
static constexpr int MIN_RING_PIXELS = 16;
static constexpr int MIN_SECTOR_PIXELS = 8;
// Lit: above level + LIT_NSIGMA * noise + LIT_OFFSET (noise = the larger of sqrt(level) and the
// ring's robust spread; the offset keeps a pixel in a ring with no background from being lit by
// one or two photons).
static constexpr float LIT_NSIGMA = 3.3f;
static constexpr float LIT_OFFSET = 2.0f;
// How long one reflection can stay on a pixel, degrees of rotation beyond the oscillation: +-3
// sigma of a rocking curve 2 degrees FWHM, wider than the crystals this is meant for.
static constexpr float ROCKING_WIDTH_DEG = 5.0f;
// Pixels expected to be called persistent by chance on the whole detector.
static constexpr double FAMILY_WISE_RATE = 0.01;
// A persistent pixel is masked only where it reads this many times its ring on average: the
// strength at which the pixels behind the merged outliers were found, and well clear of a pixel
// that is merely a little over-responding.
static constexpr double STRONG_RATIO = 10.0;
// The most frames a finder can be given: its per-pixel counters are 16-bit.
static constexpr int MAX_SAMPLED_FRAMES = UINT16_MAX;
struct Result {
std::vector<uint32_t> mask; // non-zero = masked (1 hot, 2 error value), converted geometry
size_t hot = 0; // persistent above the ring
size_t error = 0; // holding the error value on most frames
uint32_t frames = 0;
};
HotPixelFinder(const DiffractionExperiment &experiment, const PixelMask &mask, size_t nthreads);
// One preprocessed frame (INT32_MIN = masked or error value, INT32_MAX = saturated). Thread
// safe; `scratch` is the caller's own, reused across its calls.
void AddImage(const int32_t *image, std::vector<int32_t> &scratch);
#ifdef JFJOCH_USE_CUDA
// The same for a frame preprocessed on the device, on `frame`'s stream - the caller's own, as
// `scratch` above. The per-pixel sums are then kept on the device and replace the host's when the
// mask is read, so a finder is fed one way or the other, not both. Thread safe.
void AddDeviceImage(const int32_t *device_image, HotPixelFinderGPU::Frame &frame);
#endif
// The mask, from the frames added so far. oscillation_deg is the rotation per image and
// spacing_deg the rotation between two consecutive sampled frames.
[[nodiscard]] Result GetMask(double oscillation_deg, double spacing_deg, size_t nthreads);
private:
const DiffractionGeometry geometry;
const Coord axis; // rotation axis; zero where there is none
const size_t width, height;
int nrings = 0;
// Ring * SECTORS + sector of every pixel, -1 where the pixel is already masked. The per-pixel
// arrays are hundreds of megabytes on a large detector, so they are left unwritten when allocated
// and first written by the constructor's parallel pass rather than zeroed on one thread.
std::unique_ptr<int32_t[]> key;
// Where each ring-sector's pixels start in a scratch buffer laid out by key.
std::vector<uint32_t> key_begin;
size_t unmasked = 0;
// The per-pixel sums are guarded by bands of rows, so workers adding frames at the same time
// meet only when they reach the same band.
static constexpr size_t BANDS = 64;
mutable std::mutex band_mutex[BANDS];
std::atomic<size_t> next_band{0};
mutable std::mutex m;
uint32_t frames = 0;
std::unique_ptr<uint16_t[]> n_lit, n_error;
std::unique_ptr<int64_t[]> sum_value;
// A pixel's valid frames and the sum of its ring-sector level over them, kept per ring-sector - the
// frames its ring had a level on, and the sum of the level over those - less what the pixel's own
// error frames among them took out. Integers, so exactly the per-pixel sums, at a fraction of the
// memory traffic per frame.
std::vector<uint16_t> key_frames;
std::vector<int64_t> key_level_sum;
std::unique_ptr<uint16_t[]> n_error_ring_ok;
std::unique_ptr<int64_t[]> error_level_sum;
#ifdef JFJOCH_USE_CUDA
std::unique_ptr<HotPixelFinderGPU> gpu; // built by the first device frame
#endif
// Each ring-sector's level and lit threshold, from the frame's order statistics, and the frame's
// share of the per-key sums. The host and the device path both come through here.
void AddLevels(const std::vector<int32_t> &sector_level, const std::vector<int32_t> &ring_level,
const std::vector<float> &ring_spread, const std::vector<char> &ring_ok,
std::vector<int32_t> &level, std::vector<float> &threshold);
[[nodiscard]] int n_valid(size_t i) const { return key_frames[key[i]] - n_error_ring_ok[i]; }
[[nodiscard]] int64_t sum_level(size_t i) const { return key_level_sum[key[i]] - error_level_sum[i]; }
};