Files
Jungfraujoch/image_analysis/bragg_integration/BraggIntegrationEngineGPU.h
T
leonarski_fandClaude Opus 5.5 5803dee2ce Bragg integration: keep the background ring clear only of neighbours with flux on the frame
The r2 regions set aside from a reflection's r2..r3 background ring were
those of every prediction in the +-4 sigma rocking window. On a finely
sliced dense pattern most of them are the tails of reflections recorded on
the frames either side, which fill every ring while the frame shows nothing
there. 9ac2ca677 kept the reflections those rings starved by taking the ring
whole, neighbour pixels included, and relying on the high-side clip.

A prediction now masks the ring only where it puts at least 5% of its flux
on the frame (partiality >= 0.05). A ring still starved by those neighbours
has real flux in it, and the reflection is dropped, as before 9ac2ca677:
taking it whole let the neighbours' wings into the background. The CPU mask
holds two levels (tail, flux); the GPU writes them in two launches so the
flux mark wins, with no atomics.

The count the widened-radius guard reads keeps measuring the pattern's
density over every prediction, tails included, so the guard decides on
the quantity its 1.13% bound was read off. Read on the gated mask it let a
cubic set keep r1=6, and R_meas went 36.9 -> 43.2%. Masked but
unreadable ring pixels are no longer counted as the neighbours' doing.

Fixed radius, same code base, --model, the 0.05 deg / 7200-frame set:
                      pre-9ac2ca677  9ac2ca677  this
  rings starved            88.9%       88.9%    21.5%
  partials ingested         7.4M        67M      53M
  completeness             17.9%       79.7%    85.1%
  ISa                      12.8        14.3     14.7
  R_meas                    6.8%        8.4%     7.6%
  R_free                   0.139       0.176    0.171
  radial misfit            0.11        0.185    0.066
  R vs model, common hkl to 0.69 A (61k):
                           0.133       0.130    0.129
  (the R_free rise over pre-9ac2ca677 is composition: 63k -> 319k
  reflections to 0.69 A, the added ones weaker.) Gate at 0.2 instead:
  0.1% starved but radial misfit 0.35, R_free 0.180.

A crystal whose header-geometry pass indexes a 7x supercell:
9ac2ca677's whole rings raised that pass's I/sigma >= 2 count past the
refined pass's by more than 10%. RefinedPassIsWorse then sent the run
back to the supercell. Now: the true cell and space group as before
9ac2ca677, same ISa, 104M partials in the header pass against 219M.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-28 17:13:16 +02:00

93 lines
5.3 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <cstdint>
#include <memory>
#include <vector>
#include "BraggIntegrationEngine.h"
#include "../indexing/CUDAMemHelpers.h"
// CUDA engine: reproduces BraggIntegrationEngineCPU up to floating-point precision. Each stage is a
// kernel with one CUDA block per reflection cooperating over the small window via shared-memory
// reductions (the natural mapping for thousands of independent, tiny per-spot integrations).
//
// Pipeline (profile modes): reset -> mark_mask -> boxsum -> learn_profile -> build_profiles -> fit
// (the resolution shell is computed inline, so there is no separate shell pass). BoxSum mode stops
// after boxsum (that pass is the BraggIntegrate2D box integrator and the seed of the profile fit).
// The preprocessed image already lives on the device (ImagePreprocessorBufferGPU::getGPUBuffer());
// only the per-frame predicted centres are uploaded.
class BraggIntegrationEngineGPU : public BraggIntegrationEngine {
std::shared_ptr<CudaStream> stream;
int threads;
size_t fit_shared_bytes;
int rad_w = 0; // radial-background window of boxsum, in bins of one pixel
size_t boxsum_shared_bytes = 0;
size_t capacity = 0; // per-reflection device/host arrays hold at least this many reflections
// Whether d_mask / d_owner may still carry the marks of an earlier image. Run() clears what it
// marked before it returns when that is cheaper than clearing the frame, and this is then false in
// the steady state; it is true before the first call, after one that threw part-way through, and
// whenever the marks covered enough of the frame that clearing all of it was the cheaper choice.
bool dirty = true;
size_t mask_box_px = 0; // pixels one reflection's mark_mask box covers, at the widest aperture
// --- per-reflection device arrays (grown by EnsureCapacity) ---
CudaDevicePtr<float> d_px_x, d_px_y, d_d;
CudaDevicePtr<uint8_t> d_mark; // the reflection marks its signal region in d_mask
CudaDevicePtr<int> d_cx, d_cy;
CudaDevicePtr<float> d_I, d_sigma, d_bkg, d_bkg_var, d_var_bkg, d_obs_x, d_obs_y;
CudaDevicePtr<float> d_isum; // box-sum raw sum, for the radial correction
CudaDevicePtr<int> d_ninner, d_rbin, d_kbin;
CudaDevicePtr<uint8_t> d_ok, d_strong, d_has_obs;
// --- radial background curvature correction (see BraggIntegrationEngine) ---
int n_rad = 0; // radial bins, 0 when the correction is off
CudaDevicePtr<unsigned long long> d_rad_sum; // integer pixel sums, see boxsum
CudaDevicePtr<float> d_k_diff;
CudaDevicePtr<int> d_rad_cnt;
// --- fixed-size device arrays ---
// The learning/fit math is single precision: FP64 is heavily throttled on consumer GPUs and the
// extraction is Poisson-noise limited, so float reproduces the double CPU path to ~1e-4.
CudaDevicePtr<uint8_t> d_mask; // per-pixel inner-stencil reflection mask
// Per-pixel (distance, reflection) key naming the nearest predicted centre; allocated only when
// an overlap treatment is on, so the default path costs no extra device memory.
CudaDevicePtr<uint32_t> d_owner;
// Fixed-point (see PROFILE_FIXED): a float atomicAdd here made the profile depend on the order
// the blocks arrived in, and with it every intensity fitted through it.
CudaDevicePtr<unsigned long long> d_shell_grid, d_global_grid; // learned profile accumulators (N_SHELL*GG, GG)
CudaDevicePtr<float> d_shell_P, d_global_P; // normalised profiles (empirical mode)
CudaDevicePtr<unsigned long long> d_mom; // learned 2nd moments, 3 per shell + global
CudaDevicePtr<float> d_sigma2_r, d_sigma2_t; // radial/tangential widths, N_SHELL + global
CudaDevicePtr<int> d_shell_n, d_global_n;
CudaDevicePtr<unsigned long long> d_invd2; // [min,max] inv-d^2 as monotonic bit patterns
// BraggIntegrationCounts, accumulated on the device so an image costs no transfer; brought back
// only when Counts() is asked for.
CudaDevicePtr<unsigned long long> d_counts;
// --- host staging (copied back once per frame) ---
// Pinned, like every other engine's staging: a copy out of pageable memory does not return until the
// driver has staged it through a bounce buffer, so eight of them in a row are eight serialised
// round-trips rather than eight queued transfers.
CudaHostPtr<float> h_px_x, h_px_y, h_d;
CudaHostPtr<uint8_t> h_mark;
CudaHostPtr<float> h_I, h_sigma, h_bkg, h_var_bkg, h_obs_x, h_obs_y;
CudaHostPtr<uint8_t> h_ok, h_has_obs;
void EnsureCapacity(size_t n);
public:
BraggIntegrationEngineGPU(const DiffractionExperiment &experiment, std::shared_ptr<CudaStream> stream);
std::vector<Reflection> Run(const ImagePreprocessorBuffer &image,
const std::vector<Reflection> &predicted, size_t npredicted,
int64_t image_number) override;
// Brings the two device counters back before answering. Synchronises the stream, so ask once a
// pass rather than once an image.
[[nodiscard]] BraggIntegrationCounts Counts() const override;
};