Files
Jungfraujoch/image_analysis/bragg_integration/BraggIntegrationEngineGPU.h
T
leonarski_fandClaude Opus 5 0109676bdf Keep the constants out of the dual numbers, and clear only the boxes
Two costs in the per-image loop, each measured before it was touched.

Geometry refinement is the largest item in that loop - about half to two thirds of
its processor time on the datasets where the loop matters, and all of it on the host.
It runs three solves per image, and each one spends four fifths of itself inside the
solver at barely two iterations: the cost is not convergence, it is what every
residual evaluation does. The residual carried the blocks it does not refine as dual
numbers, so each evaluation recomputed the two detector rotations, the whole
orthogonalisation matrix, three cross products and the cell volume - all of them
constant for the image - through the derivative machinery, several million times per
run. Split the observed and predicted sides so the un-refined blocks pass as plain
doubles, evaluate the cell side once when the functor is built, and let the rotator
take a point whose type differs from the angle's. A dual number times a double is a
dual number times a dual number whose derivatives are zero, so the arithmetic is the
same one with the zeros removed.

Integration cleared the owner and mask images for the whole frame before every image.
On a large detector that is more than three hundred megabytes of writes to reset
pixels of which about one in twenty-five is ever marked, and it cost most of what the
integration kernels themselves cost. The marking kernel gained an unmarking mode - one
kernel, so the two cannot drift apart - and the engine clears whichever way is cheaper
for the frame in front of it, with a flag to force the full clear the first time and
after anything threw. The size test is not decoration: without it, clearing box by box
is slower than the memset on a small detector with many predictions, which is what the
measurement said before it was added.

Faster on thirteen of thirteen matched pairs: refinement by a quarter to a third,
whole-run wall by one to eight per cent depending on how much of the run is the loop.
The two changes pay in opposite regimes - refinement where the loop is processor-bound,
the clear where the detector is large enough for the card to be the constraint.

Every reflection file over seven datasets is byte-identical, and the solver did not
merely land in the same place: it took the same path, agreeing digit for digit on
iteration, residual and Jacobian evaluation counts. A new test runs two mismatched
frames through one engine and compares against a fresh one, which is what a mark left
behind would break.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016NNnL26LAvruQ9eLUUWvrJ
2026-08-24 21:49:57 +02:00

81 lines
4.5 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <cstdint>
#include <memory>
#include <vector>
#include "BraggIntegrationEngine.h"
#include "../indexing/CUDAMemHelpers.h"
// CUDA engine: reproduces BraggIntegrationEngineCPU up to floating-point precision. Each stage is a
// kernel with one CUDA block per reflection cooperating over the small window via shared-memory
// reductions (the natural mapping for thousands of independent, tiny per-spot integrations).
//
// Pipeline (profile modes): reset -> mark_mask -> boxsum -> learn_profile -> build_profiles -> fit
// (the resolution shell is computed inline, so there is no separate shell pass). BoxSum mode stops
// after boxsum (that pass is the BraggIntegrate2D box integrator and the seed of the profile fit).
// The preprocessed image already lives on the device (ImagePreprocessorBufferGPU::getGPUBuffer());
// only the per-frame predicted centres are uploaded.
class BraggIntegrationEngineGPU : public BraggIntegrationEngine {
std::shared_ptr<CudaStream> stream;
int threads;
size_t fit_shared_bytes;
int rad_w = 0; // radial-background window of boxsum, in bins of one pixel
size_t boxsum_shared_bytes = 0;
size_t capacity = 0; // per-reflection device/host arrays hold at least this many reflections
// Whether d_mask / d_owner may still carry the marks of an earlier image. Run() clears what it
// marked before it returns when that is cheaper than clearing the frame, and this is then false in
// the steady state; it is true before the first call, after one that threw part-way through, and
// whenever the marks covered enough of the frame that clearing all of it was the cheaper choice.
bool dirty = true;
size_t mask_box_px = 0; // pixels one reflection's mark_mask box covers, at the widest aperture
// --- per-reflection device arrays (grown by EnsureCapacity) ---
CudaDevicePtr<float> d_px_x, d_px_y, d_d;
CudaDevicePtr<int> d_cx, d_cy;
CudaDevicePtr<float> d_I, d_sigma, d_bkg, d_bkg_var, d_var_bkg, d_obs_x, d_obs_y;
CudaDevicePtr<float> d_isum; // box-sum raw sum, for the radial correction
CudaDevicePtr<int> d_ninner, d_rbin, d_kbin;
CudaDevicePtr<uint8_t> d_ok, d_strong, d_has_obs;
// --- radial background curvature correction (see BraggIntegrationEngine) ---
int n_rad = 0; // radial bins, 0 when the correction is off
CudaDevicePtr<unsigned long long> d_rad_sum; // integer pixel sums, see boxsum
CudaDevicePtr<float> d_k_diff;
CudaDevicePtr<int> d_rad_cnt;
// --- fixed-size device arrays ---
// The learning/fit math is single precision: FP64 is heavily throttled on consumer GPUs and the
// extraction is Poisson-noise limited, so float reproduces the double CPU path to ~1e-4.
CudaDevicePtr<uint8_t> d_mask; // per-pixel inner-stencil reflection mask
// Per-pixel (distance, reflection) key naming the nearest predicted centre; allocated only when
// an overlap treatment is on, so the default path costs no extra device memory.
CudaDevicePtr<uint32_t> d_owner;
// Fixed-point (see PROFILE_FIXED): a float atomicAdd here made the profile depend on the order
// the blocks arrived in, and with it every intensity fitted through it.
CudaDevicePtr<unsigned long long> d_shell_grid, d_global_grid; // learned profile accumulators (N_SHELL*GG, GG)
CudaDevicePtr<float> d_shell_P, d_global_P; // normalised profiles (empirical mode)
CudaDevicePtr<unsigned long long> d_mom; // learned 2nd moments, 3 per shell + global
CudaDevicePtr<float> d_sigma2_r, d_sigma2_t; // radial/tangential widths, N_SHELL + global
CudaDevicePtr<int> d_shell_n, d_global_n;
CudaDevicePtr<unsigned long long> d_invd2; // [min,max] inv-d^2 as monotonic bit patterns
// --- host staging (copied back once per frame) ---
std::vector<float> h_px_x, h_px_y, h_d;
std::vector<float> h_I, h_sigma, h_bkg, h_var_bkg, h_obs_x, h_obs_y;
std::vector<uint8_t> h_ok, h_has_obs;
void EnsureCapacity(size_t n);
public:
BraggIntegrationEngineGPU(const DiffractionExperiment &experiment, std::shared_ptr<CudaStream> stream);
std::vector<Reflection> Run(const ImagePreprocessorBuffer &image,
const std::vector<Reflection> &predicted, size_t npredicted,
int64_t image_number) override;
};