Exact: p.mtz and the pre-scan products (shadow mask, mean projection, defective-pixel mask, ring and capture centres, compared as hashes and hex floats) are bit-identical to rc174 on three in-house rotation sets, GPU and CPU builds. - ShadowFinder::GetMask: the serial parts run in parallel - connected components by row band joined with union-find (both the shadow and the transmitting-arm searches, and the hole fill), ring binning and the harmonic sector gather by blocks, gap bridging by line; ring pixel counts read off the ring offsets. Mean projection filled in parallel. - ShadowFinder host accumulation: one band-locked projection instead of a 20 B/px shard per pre-scan worker (2.7 GB zeroed and folded on a 16M detector); SetShardCount and the shard argument are gone. - FindBeamCenterFromBackground: the usable-pixel test is made once, the in-band pixels are kept in pixel order so the clipping rounds no longer sweep the whole detector, the 67 MB cell map is gone and the per-iteration block fold runs in parallel - same sums, same order. - HotPixelFinder::GetMask: the chance-rate counts in parallel (integers). Measured on a loaded box (load ~25 from other jobs), pre-scan window: GPU 5.9-6.5 s -> 3.2-3.4 s, CPU 8.4-9.0 s -> 6.1-7.4 s. The GPU-build pre-scan now ends with its background spot measurement (CPU spot finder on ~120 frames, ~13 core-s on 8 workers). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
155 lines
8.2 KiB
C++
155 lines
8.2 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <cstdint>
|
|
#include <atomic>
|
|
#include <future>
|
|
#include <memory>
|
|
#include <mutex>
|
|
#include <optional>
|
|
#include <vector>
|
|
|
|
#include "../../common/CompressedImage.h"
|
|
#include "../../common/DiffractionExperiment.h"
|
|
#include "../../common/DiffractionGeometry.h"
|
|
#include "../../common/JFJochMessages.h"
|
|
#include "../../common/PixelMask.h"
|
|
#ifdef JFJOCH_USE_CUDA
|
|
#include "ShadowAccumulatorGPU.h"
|
|
#endif
|
|
|
|
// Finds the beam-stop shadow - the central disk and the holder arm - from a set of images,
|
|
// mirroring the accumulate-then-finalize shape of DarkMaskAnalysis: feed frames with
|
|
// AddImage(), then read the mask once with GetMask(). The mask is in converted geometry
|
|
// and is 1 where the beam stop shadows the detector.
|
|
//
|
|
// The shadow is a place where the background is missing, so it is found by comparing each
|
|
// pixel's mean against the typical background at the same radius - the median over its ring,
|
|
// taken over the pixels not already known to be shadowed. That comparison holds wherever the
|
|
// ring still has unshadowed pixels to measure. Where it does not - a ring lying wholly inside
|
|
// the stop - there is nothing to compare against, and such a ring is shadow in its entirety.
|
|
//
|
|
// The background belongs to the beam and the shadow to the stop, and the two are not concentric:
|
|
// the stop sits off the beam by a sizeable fraction of its own radius. Only the per-ring
|
|
// comparison is used, so nothing here assumes they share a centre.
|
|
//
|
|
// Frames are chosen by the caller; the detection needs enough of them that the background
|
|
// is counted rather than guessed (see MIN_EXPECTED_COUNTS in the .cpp).
|
|
// Thread-safe: workers call AddImage concurrently. The projection is split into bands of rows, each
|
|
// with a lock of its own, and a worker adding a frame starts at a different band from the one before
|
|
// it, so workers meet only when they reach the same band.
|
|
class ShadowFinder {
|
|
mutable std::mutex m;
|
|
|
|
const int width;
|
|
const int height;
|
|
float beam_x;
|
|
float beam_y;
|
|
|
|
// What the background owes to the source rather than to the hardware. The scattered background
|
|
// is not flat around a ring: a polarized source suppresses it in its own plane by a factor that
|
|
// reaches three at the 2 theta a short detector distance puts in a corner - several times the
|
|
// dip this class is looking for - so the comparison divides it out before it compares. Of the
|
|
// corrections a ring carries this is the only one that varies along it; solid angle, detector
|
|
// and air absorption are all functions of 2 theta alone and the ring's own median absorbs them.
|
|
// The geometry is kept whole rather than reduced to a distance and a pixel size because the
|
|
// azimuth is the whole point: detector tilt, a quarter-turned image and an in-plane rotation
|
|
// all move the polarization plane across the stored image, and the geometry already knows where
|
|
// it lies. Only the centre is replaced, by the one the caller measured.
|
|
const DiffractionGeometry geometry;
|
|
const std::optional<float> polarization; // unset leaves the background as it was measured
|
|
|
|
std::vector<uint32_t> pixel_mask; // pixels already masked carry no background to test
|
|
|
|
// Per-pixel projection over the frames added so far (converted geometry). The sums and counts are
|
|
// integers and the maximum is a maximum, so the result does not depend on the order the frames
|
|
// arrive in.
|
|
struct Projection {
|
|
std::vector<int64_t> max_value;
|
|
std::vector<int64_t> sum_value;
|
|
std::vector<uint32_t> valid_count;
|
|
uint32_t frames = 0;
|
|
};
|
|
// The frames added on the host. Allocated by the first of them: with a GPU there are usually none.
|
|
Projection host;
|
|
std::mutex host_mutex; // guards the allocation and the frame count, not the sums
|
|
static constexpr size_t BANDS = 64;
|
|
std::mutex band_mutex[BANDS];
|
|
std::atomic<size_t> next_band{0};
|
|
|
|
#ifdef JFJOCH_USE_CUDA
|
|
// Present when a GPU is available. Frames it can decode are accumulated there instead of on the
|
|
// host - only the compressed chunk crosses PCIe - and its projection is folded in with the
|
|
// host projection when the mask is read. Frames it cannot take (anything but bitshuffle+LZ4) still
|
|
// go to the host, so a run mixing compressions is handled without a second code path.
|
|
// Built on a thread of its own: it allocates and clears several hundred megabytes of device
|
|
// memory, and cudaMalloc synchronises the whole device, so doing it in the constructor would
|
|
// stall the caller before it has read its first frame. The first AddImage waits for it, by
|
|
// which time the reads have been running for a while.
|
|
mutable std::future<std::unique_ptr<ShadowAccumulatorGPU>> gpu_pending;
|
|
mutable std::unique_ptr<ShadowAccumulatorGPU> gpu;
|
|
mutable std::mutex gpu_mutex;
|
|
|
|
// The accumulator once its construction has finished, or null if there is none.
|
|
[[nodiscard]] ShadowAccumulatorGPU *Gpu() const;
|
|
#endif
|
|
|
|
// Add the pixels [begin, end) of one frame to the host projection.
|
|
template<class T> void Add(const T *ptr, size_t begin, size_t end);
|
|
|
|
#ifdef JFJOCH_USE_CUDA
|
|
// The device's projection with the host's folded in. max_value is only taken from a projection
|
|
// that actually counted the pixel - one that never saw it holds 0, which would beat a genuinely
|
|
// negative maximum.
|
|
[[nodiscard]] Projection Reduce() const;
|
|
|
|
// That projection, made on its first read and kept. The ring-centre fit, the mask and the
|
|
// beam-centre capture all read the same one, and on a 16 Mpx detector each is 360 MB brought back
|
|
// from the device into fresh memory.
|
|
mutable std::optional<Projection> reduced;
|
|
#endif
|
|
// The projection the frames added so far make: the host's, or the one above where the device
|
|
// took frames. Called with `m` held.
|
|
[[nodiscard]] const Projection &Reduced() const;
|
|
|
|
public:
|
|
static constexpr uint32_t SHADOW = 1;
|
|
static constexpr uint32_t TRANSMITTING = 2;
|
|
|
|
ShadowFinder(const DiffractionExperiment &experiment, const PixelMask &mask);
|
|
|
|
// The centre the rings are drawn about. It starts as the file's, which is the only one there
|
|
// is when the finder is built; a caller that has measured one replaces it before reading the
|
|
// mask. A centre far from the truth draws the rings across the background's own radial
|
|
// fall-off instead of along it, and the comparison then describes the fall-off rather than the
|
|
// hardware. The projection is not centred on anything, so this may be set after the frames.
|
|
void BeamCenter(float x, float y);
|
|
|
|
// Accumulate one full converted-geometry image. Gap / masked pixels (the pixel type's sentinel
|
|
// extreme) are skipped. `buffer` is scratch space for decompression, reused across the calls of
|
|
// one worker.
|
|
void AddImage(const DataMessage &data, std::vector<uint8_t> &buffer);
|
|
|
|
// Compute the shadow mask (SHADOW, TRANSMITTING or 0 = keep), of the converted pixel count.
|
|
// TRANSMITTING marks the pieces of hardware that let part of the beam through, added after the
|
|
// shadow proper; both are masked, and a consumer that must not see those pieces can tell them apart.
|
|
// Recomputed on each call from the projection, which is put together on the first read of it
|
|
// (GetMask or GetMeanProjection): frames added after that are not seen.
|
|
// nthreads = 0 asks for all hardware threads. The per-pixel passes over a 16M-pixel detector
|
|
// dominate this, and they are all exactly parallel.
|
|
[[nodiscard]] std::vector<uint32_t> GetMask(size_t nthreads = 0) const;
|
|
|
|
// Mean counts per pixel over the frames added, NAN where nothing was counted. This is the
|
|
// projection GetMask() tests, so anything else that wants the background before indexing
|
|
// gets it without reading the frames a second time.
|
|
[[nodiscard]] std::vector<float> GetMeanProjection() const;
|
|
|
|
// Let go of the projection the two above read, once the caller has what it wants of it: on a
|
|
// 16 Mpx detector it is 360 MB. A later read puts it together again.
|
|
void ReleaseProjection();
|
|
|
|
[[nodiscard]] uint32_t GetFrameCount() const;
|
|
};
|