Files
Jungfraujoch/image_analysis/image_preprocessing/ImagePreprocessorBufferGPU.h
jungfrauandClaude Opus 5 9f49e5abb7
Build Packages / Unit tests (push) Failing after 5m38s
Build Packages / build:windows:nocuda (push) Successful in 19m41s
Build Packages / build:viewer-tgz:cpu (push) Successful in 21m1s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m33s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 24m7s
Build Packages / build:rpm (rocky9_sls9) (push) Failing after 18m49s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 24m30s
Build Packages / build:rpm (rocky8_sls9) (push) Failing after 25m37s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 27m57s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m11s
Build Packages / XDS test (durin plugin) (push) Successful in 10m24s
Build Packages / Generate python client (push) Successful in 34s
Build Packages / build:rpm (rocky9) (push) Failing after 15m21s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m12s
Build Packages / build:rpm (ubuntu2404) (push) Failing after 15m10s
Build Packages / build:rpm (rocky8) (push) Failing after 18m41s
Build Packages / XDS test (neggia plugin) (push) Successful in 11m24s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 11m50s
Build Packages / build:rpm (ubuntu2204) (push) Failing after 17m24s
Build Packages / DIALS test (push) Successful in 17m33s
Build Packages / build:windows:cuda (push) Successful in 23m37s
Keep 16-bit images 16-bit through the GPU pipeline
A detector reading out 16 bits had its frame widened to int32 the moment it was
decoded, and every per-pixel pass over that frame then moved four bytes a pixel to
carry two. Those passes - the ring statistics three times over, the strong-pixel
search, spot extraction, the azimuthal and ROI integrators, Bragg integration - are
the bulk of the image loop's device traffic, and 16 bits is the mode a fast
acquisition runs in, which is exactly where throughput matters.

The preprocessed image now keeps the width of its source. Two codes at the top of the
16-bit range carry the two special states, and they cannot collide with a real value:

  0xFFFF          masked, or the source's own bad-pixel marker.
  saturation      a pixel at or above the saturation limit. 0xFFFE where the limit
  code            leaves room - a 16-bit EIGER declares a count-rate limit of a few
                  thousand, so there is room to spare - and 0xFFFF where the limit is
                  the whole range, in which case the "is error" test has already
                  claimed 0xFFFF, nothing can be saturated, and 0xFFFE stays a real
                  value.

Either way a real value is strictly below the saturation limit and so below both
codes. Nothing is clipped and nothing is lost, and which code is in force is carried
with the image rather than assumed.

No engine learns a second convention. PixelView widens on load, so a masked pixel
still reads as INT32_MIN and a saturated one as INT32_MAX, and every existing
`v != INT32_MIN && v != INT32_MAX` test keeps its meaning. One code path, not two
instantiations that can drift apart; the branch is on a pointer that is the same for
every thread of every block, on kernels whose time is the loads it selects between.
The vector loads are kept - four pixels still arrive in one transaction, 16 bytes wide
or 8, whichever the image is.

The wide path is unchanged, and is still taken for anything that is not a 16-bit
source, and for any caller that wants the preprocessed image copied back to the host -
that mirror is int32 and the CPU engines know only that convention.

Measured on the one 16-bit dataset in the rotation test set, which is also the
smallest detector in it (2.5M pixels, where per-pixel work is a small part of the
loop): image loop 1.025 s -> 1.005 s at one GPU, whole run 5.64 s -> 5.52 s. The gain
scales with the frame, so a 16M-pixel detector - where six full-frame passes are 86 %
of the loop's GPU time - has much more to gain, and nothing here can measure that:
every other dataset in the test set is stored 32-bit.

Correctness on that dataset is exact where it can be: indexing rate, first-pass
validation score and the integrated partial count are identical to the wide path, and
its whole battery row - reflections, observations, space group, R_meas, CC1/2, ISa,
mosaicity - is unchanged. Battery 6m17s -> 6m16s, 21/24 space groups, no failures.

Also logs, once per run, the width the images are stored in, since it decides how much
of the frame moves through every pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 01:52:27 -04:00

46 lines
2.2 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <memory>
#include "ImagePreprocessorBuffer.h"
#include "../indexing/CUDAMemHelpers.h"
#include "PreprocessedPixel.h"
class ImagePreprocessorBufferGPU : public ImagePreprocessorBuffer {
CudaDevicePtr<int32_t> gpu_image;
CudaRegisteredVector<int32_t> buffer_reg;
// Set per image by the preprocessor: true once it has written this frame in its source's 16-bit
// width rather than widening it, with the code that frame uses for a saturated pixel.
bool narrow = false;
uint16_t narrow_sat_code = 0;
// Staging for Gather(). Its only caller is ImageSpotFinder::ExtractSpots, which gives up on a frame
// with UINT16_MAX or more strong pixels (the connected-component search rejects it anyway), so that
// is the largest gather that can be asked for.
static constexpr size_t MAX_GATHER = UINT16_MAX;
CudaDevicePtr<uint32_t> gpu_gather_index;
CudaDevicePtr<int32_t> gpu_gather_value;
// Own stream: every analysis engine synchronises its own stream before it returns, so the device
// image is final by the time a gather is asked for. The NULL stream would serialise all workers.
CudaStream gather_stream;
public:
// host_mirror = false skips the host copy of the preprocessed image entirely (and with it the
// page-locking): pass it when every engine reading this buffer runs on the device.
explicit ImagePreprocessorBufferGPU(size_t npixel, bool host_mirror = true);
int32_t *getGPUBuffer() override;
const int32_t *getGPUBuffer() const override;
// The device image is one allocation either way - the narrow form is the same memory read two
// bytes at a time, so a run that switches width allocates nothing and frees nothing.
bool IsNarrow() const override { return narrow; }
const uint16_t *getGPUBufferNarrow() const override;
uint16_t NarrowSatCode() const override { return narrow_sat_code; }
void SetNarrow(bool v, uint16_t sat_code) override { narrow = v; narrow_sat_code = sat_code; }
void Gather(const std::vector<uint32_t> &npixel, std::vector<int32_t> &values) const override;
};