Files
Jungfraujoch/image_analysis/bragg_integration/BraggIntegrationEngineGPU.h
T
leonarski_fandClaude Opus 5.5 37a8c8e24e Rotation merge: drop rocking events with an overloaded pixel; capture uncertainty in the merge variance
A saturated pixel in a spot means the brightest part of the reflection was
not measured. The integration used to drop the peak frame's partial (its
peak pixel is unreadable) and keep the flanks, so the combine extrapolated
the event from its tails by the partiality model: on a strongly
diffracting small-molecule crystal the strongest low-order reflections
read 2-3x low and were the largest SHELXL misfits. XDS drops such a
reflection (OVERLOAD); so does rugnux now.

- Integration (CPU + GPU engines): a reflection is `overloaded` when a
  signal-disk pixel is saturated, or unreadable on this frame but not in
  the run's pixel mask - EIGER/PILATUS write their error value for a
  pixel they could not count, which the preprocessor turns into a masked
  pixel like a gap's. The engines now receive the PixelMask to tell the
  two apart (an earlier attempt that re-classified the marker as
  saturation in the preprocessor broke a dataset whose gaps are not in
  the file's mask). An overloaded reflection is kept with its box sum,
  unfitted, only so its event can be recognised.
- Rotation combine (CPU + GPU): an event with any overloaded partial is
  dropped whole; counted in the log and the report
  (OBSERVATIONS_REJECTED_OVERLOAD=). The unmerged MTZ export drops it too.
- Everything else that reads reflections leaves an overloaded one out:
  AcceptReflection (stills merge, per-image scaling), the post-refinement
  gather, the axial-row sums.
- Capture uncertainty: the merge rebuilds each full's variance at the
  reflection's mean (counting_variance / ModelSigma) and dropped the
  capture term the combine had put into sigma, so a full extrapolated
  from part of its rocking curve merged at the weight of a whole one.
  Fulls now carry it (Obs::capture) and the rebuilt variance adds
  (capture * <I>)^2, host and device.

SHELXL R1 on rugnux's own integration (harness), median fix -> this:
citric acid .0648 -> .0420 (XDS .051; 221 events dropped, EXTI 1.02 -> 0.29),
HEPES .0396 -> .0381 (184), aspirin 20 keV .0387 -> .0385 (6),
aspirin 25 keV .0376 -> .0375 (5); metformin/nidppe/dnba/lalanine/cytidine
no overloads, unchanged. YAG .116 -> .128 (87 dropped; its scale loop does
not settle either way). Proteins and private subset: see the branch report.
Tests: BraggIntegrationEngineCPU_SaturatedPeakIsFlaggedNotDropped (new),
BraggIntegrationEngineGPU_MatchesCPU (overloaded flag compared),
AcceptReflection_ResolutionLimits, [write_reflections], [large].

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-04 21:01:40 +02:00

98 lines
5.6 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <cstdint>
#include <memory>
#include <vector>
#include "BraggIntegrationEngine.h"
#include "../indexing/CUDAMemHelpers.h"
#include "../../common/PixelMask.h"
// CUDA engine: reproduces BraggIntegrationEngineCPU up to floating-point precision. Each stage is a
// kernel with one CUDA block per reflection cooperating over the small window via shared-memory
// reductions (the natural mapping for thousands of independent, tiny per-spot integrations).
//
// Pipeline (profile modes): reset -> mark_mask -> boxsum -> learn_profile -> build_profiles -> fit
// (the resolution shell is computed inline, so there is no separate shell pass). BoxSum mode stops
// after boxsum (that pass is the BraggIntegrate2D box integrator and the seed of the profile fit).
// The preprocessed image already lives on the device (ImagePreprocessorBufferGPU::getGPUBuffer());
// only the per-frame predicted centres are uploaded.
class BraggIntegrationEngineGPU : public BraggIntegrationEngine {
std::shared_ptr<CudaStream> stream;
int threads;
size_t fit_shared_bytes;
int rad_w = 0; // radial-background window of boxsum, in bins of one pixel
size_t boxsum_shared_bytes = 0;
size_t capacity = 0; // per-reflection device/host arrays hold at least this many reflections
// Whether d_mask / d_owner may still carry the marks of an earlier image. Run() clears what it
// marked before it returns when that is cheaper than clearing the frame, and this is then false in
// the steady state; it is true before the first call, after one that threw part-way through, and
// whenever the marks covered enough of the frame that clearing all of it was the cheaper choice.
bool dirty = true;
size_t mask_box_px = 0; // pixels one reflection's mark_mask box covers, at the widest aperture
// --- per-reflection device arrays (grown by EnsureCapacity) ---
CudaDevicePtr<float> d_px_x, d_px_y, d_d;
CudaDevicePtr<uint8_t> d_mark; // the reflection marks its signal region in d_mask
CudaDevicePtr<int> d_cx, d_cy;
CudaDevicePtr<float> d_I, d_sigma, d_bkg, d_bkg_var, d_var_bkg, d_obs_x, d_obs_y;
CudaDevicePtr<float> d_isum; // box-sum raw sum, for the radial correction
CudaDevicePtr<int> d_ninner, d_rbin, d_kbin;
// The run's pixel mask, one byte per pixel (PixelMask::GetBinaryMask), shared with the
// preprocessor's copy. See BraggIntegrationEngineCPU::static_mask.
std::shared_ptr<CudaDevicePtr<uint8_t>> d_static_mask;
CudaDevicePtr<uint8_t> d_ok, d_strong, d_has_obs, d_overloaded;
// --- radial background curvature correction (see BraggIntegrationEngine) ---
int n_rad = 0; // radial bins, 0 when the correction is off
CudaDevicePtr<unsigned long long> d_rad_sum; // integer pixel sums, see boxsum
CudaDevicePtr<float> d_k_diff;
CudaDevicePtr<int> d_rad_cnt;
// --- fixed-size device arrays ---
// The learning/fit math is single precision: FP64 is heavily throttled on consumer GPUs and the
// extraction is Poisson-noise limited, so float reproduces the double CPU path to ~1e-4.
CudaDevicePtr<uint8_t> d_mask; // per-pixel inner-stencil reflection mask
// Per-pixel (distance, reflection) key naming the nearest predicted centre; allocated only when
// an overlap treatment is on, so the default path costs no extra device memory.
CudaDevicePtr<uint32_t> d_owner;
// Fixed-point (see PROFILE_FIXED): a float atomicAdd here made the profile depend on the order
// the blocks arrived in, and with it every intensity fitted through it.
CudaDevicePtr<unsigned long long> d_shell_grid, d_global_grid; // learned profile accumulators (N_SHELL*GG, GG)
CudaDevicePtr<float> d_shell_P, d_global_P; // normalised profiles (empirical mode)
CudaDevicePtr<unsigned long long> d_mom; // learned 2nd moments, 3 per shell + global
CudaDevicePtr<float> d_sigma2_r, d_sigma2_t; // radial/tangential widths, N_SHELL + global
CudaDevicePtr<int> d_shell_n, d_global_n;
CudaDevicePtr<unsigned long long> d_invd2; // [min,max] inv-d^2 as monotonic bit patterns
// BraggIntegrationCounts, accumulated on the device so an image costs no transfer; brought back
// only when Counts() is asked for.
CudaDevicePtr<unsigned long long> d_counts;
// --- host staging (copied back once per frame) ---
// Pinned, like every other engine's staging: a copy out of pageable memory does not return until the
// driver has staged it through a bounce buffer, so eight of them in a row are eight serialised
// round-trips rather than eight queued transfers.
CudaHostPtr<float> h_px_x, h_px_y, h_d;
CudaHostPtr<uint8_t> h_mark;
CudaHostPtr<float> h_I, h_sigma, h_bkg, h_var_bkg, h_obs_x, h_obs_y;
CudaHostPtr<uint8_t> h_ok, h_has_obs, h_overloaded;
void EnsureCapacity(size_t n);
public:
BraggIntegrationEngineGPU(const DiffractionExperiment &experiment, std::shared_ptr<CudaStream> stream,
const PixelMask &mask);
std::vector<Reflection> Run(const ImagePreprocessorBuffer &image,
const std::vector<Reflection> &predicted, size_t npredicted,
int64_t image_number) override;
// Brings the two device counters back before answering. Synchronises the stream, so ask once a
// pass rather than once an image.
[[nodiscard]] BraggIntegrationCounts Counts() const override;
};