Files
Jungfraujoch/image_analysis/spot_finding/ImageSpotFinderGPU.cu
T
leonarski_fandjungfrau 4dc2534dbf
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 18m57s
Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 16m55s
Build Packages / build:windows:cuda (push) Successful in 18m48s
Build Packages / build:viewer-tgz:cpu (push) Successful in 13m10s
Build Packages / build:viewer-tgz:cuda (push) Successful in 14m45s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 22m23s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 20m12s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 23m7s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 20m43s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 23m9s
Build Packages / XDS test (durin plugin) (push) Successful in 12m26s
Build Packages / build:rpm (rocky9) (push) Successful in 24m58s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 23m20s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (JFJoch plugin) (push) Successful in 12m37s
Build Packages / build:rpm (rocky8) (push) Successful in 27m58s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m38s
Build Packages / Build documentation (push) Successful in 59s
Build Packages / DIALS test (push) Successful in 23m16s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m38s
v1.0.0.rc-162 (#72)
**Files written by Jungfraujoch now import correctly in DIALS, XDS and pyFAI.** A tilted detector, a grid scan, a still recorded at a goniometer position, and saturated or unreadable pixels were each described in a way that a third-party program acted on wrongly. If you process Jungfraujoch data outside Jungfraujoch, prefer this release to any earlier one.

* HDF5: the detector tilt (`rot1`/`rot2`/`rot3`) is exported correctly in the NXmx transformation chain; untilted geometries are unaffected.
* HDF5: a still recorded at a goniometer position is no longer read back as a single image, and a grid scan records a stationary spindle so a program that requires a rotation axis can open it.
* HDF5: the sample transformation chain is written in mounting order, with a Smargon head position told apart from the spindle, one entry per image, `module_offset` as a float unit vector, and `offset_units` on every offset.
* HDF5: saturated, underloaded and unreadable pixels are described so a downstream program masks them - `saturation_value`, `underload_value`, `error_value` and `bit_depth_readout` are written correctly, and a data file missing next to a VDS master reads as the error marker rather than as zero counts.
* HDF5: the rotation axis is read back under whatever name it carries, and `mirror_y` records whether the assembled image is mirrored in Y relative to the detector's raw readout.
* A grid scan and a goniometer axis can both be set; they are no longer alternatives.
* `images_per_file` is chosen from the acquisition when it is not given: a rotation sweep of at most 20000 images goes into a single data file, a grid scan splits on whole fast-axis rows, and stills and serial keep 1000.
* The writer refuses a stream whose start message declares a different pixel format than its images carry, and a DECTRIS detector sending signed images is no longer declared unsigned.
* The image stream can carry the sample transformation chain (`transformations`, in the END message); a producer that does not send it gets the same chain built by the writer.
* rugnux: fixing the space group with `-S` no longer prevents the lattice from being found - a lattice indexed in a different setting is reindexed into that group's own setting, and a run whose crystal does not have that group's lattice stops and names the cell it indexed as, rather than reporting statistics that cannot describe it.
* rugnux: the per-image resolution estimate now predicts the resolution the merged data reach rather than the highest-resolution spot found, and is reported as `SPOT_RESOLUTION_ESTIMATE`.
* rugnux: two runs of the same command on the same images produce the same merged intensities; the azimuthal profile written alongside them is not yet reproducible in the same way.
* rugnux: the offline lattice refinement is bounded by iterations rather than by a wall clock, so a loaded machine can no longer refine to a different lattice; a live acquisition keeps its real-time bound.
* rugnux: the detector-frame modulation correction is fitted on a grid spanning the detector, so whether it is applied no longer depends on how far integration reached.
* rugnux: the geometry pre-pass no longer writes `<prefix>_01.mtz`, `_01.cif`, `_01.hkl` and `_01_image.dat`; the refined second pass writes those files under `<prefix>`, and that is the result to use.
* rugnux: `_process.h5` describes the pixel format of the images it links to, and is written on a thread of its own.
* rugnux: the detector geometry is also logged in XDS's convention (`ORGX`/`ORGY`, detector axis vectors, rotation axis), so it can be compared with an XDS refinement.
* rugnux: an image integrated in pyFAI through the `.poni` file written by `--mode calibration` comes out with the correct azimuth, and the file declares pyFAI's `orientation`, which needs pyFAI 2024.01 or newer. Radial integration is unchanged.
* rugnux: a rotation run is substantially faster throughout - beam-stop detection, first-pass indexing, geometry refinement, integration, scaling and merging - and observations outside the scaling resolution range are dropped as they are ingested. The refined geometry, the space group chosen and the merged statistics are unchanged.
* Faster spot finding and indexing, on the broker as well as in rugnux; the spots found and the lattices indexed are unchanged.
* A run reserves substantially less GPU memory: nothing is allocated for buffers that are never read, and a worker builds only the engines it uses.
* rugnux: with `-N` left at its default the per-image loop of `--mode mx` uses at most 16 workers per GPU, rather than one per hardware thread; an explicit `-N` is obeyed as given.
* CUDA 12 builds now contain device code for Volta, so the RHEL 8 packages and the portable Linux `.tgz` run on a V100; the CUDA 13 artefacts (RHEL 9, Ubuntu, Windows) remain Turing and newer.
* The build resolves a single Eigen for the whole project, and refuses to configure if Ceres picks up a different one; a build that mixed two Eigen versions was undefined behaviour and crashed at -O2.
* Documentation: a security page, and the supported GPU generations and minimum NVIDIA driver version of every released artefact.

**Breaking change to OpenAPI** - regenerate the client (`jfjoch-client` 1.0.0-rc.162, `frontend/src/client`):
* `dataset_settings.images_per_file` is no longer `default: 1000` and no longer accepts `0`; it is optional, and its minimum is 1. A client sending `0` (previously "one file for the whole run") is now rejected - omit the field instead, which for a rotation sweep gives the same single file.
* `file_writer_format` now defaults to `NXmxVDS`, matching the server's own default and the layout recommended for DIALS, XDS and CrystFEL. A generated client that fills in schema defaults and does not set the format explicitly will write VDS masters where it previously wrote legacy ones; set `NXmxLegacy` explicitly to keep them.

---------

Co-authored-by: jungfrau <jungfrau@mx-aare-test.psi.ch>
Reviewed-on: #72
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-08-25 08:21:39 +02:00

308 lines
16 KiB
Plaintext

// SPDX-FileCopyrightText: 2025 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
// GPU Spot finding developed by Hans-Christian Stadler (PSI)
// Copyright (2019-2023) Paul Scherrer Institute
#include "ImageSpotFinderGPU.h"
#include "../../common/JFJochException.h"
struct spot_parameters {
int32_t width;
int32_t height;
float strong_pixel_threshold2;
int32_t count_threshold;
};
// input X x Y pixels array
// output X x Y bit array
static constexpr int WARP_SIZE = 32; // assume warp size of 32 cuda threads per warp
inline void cuda_err(cudaError_t val) {
if (val != cudaSuccess)
throw JFJochException(JFJochExceptionCategory::GPUCUDAError, cudaGetErrorString(val));
}
// Write pixel results to bit array
// params: spot finding parameters
// out: pixel result bit array
// pixel: flat pixel index = bit index into bit array
// val: pixel result
// **NOTE**: assumes sizeof(*out) * 8 == WARP_SIZE
__device__ __forceinline__ void write_result(const spot_parameters& params, uint32_t* out, int32_t pixel, uint8_t val)
{
static_assert(sizeof(*out) * 8 == WARP_SIZE, "Violation of essential implementation assumption: WARP_SIZE must match output array element type bit size!");
static constexpr unsigned ALL_THREADS = UINT32_MAX;
const int32_t laneid = threadIdx.x & (WARP_SIZE - 1);
unsigned result = __ballot_sync(ALL_THREADS, val);
const int32_t idx = pixel / WARP_SIZE; // global uint32_t index
const int32_t bit = pixel % WARP_SIZE; // local bit index
if ((bit >= laneid) && (laneid == 0)) { // write to upper part of uint32_t
result <<= bit;
if (result)
atomicOr(&out[idx], result);
} else if ((bit < laneid) && (bit == 0)) { // write to lower part of uint32_t
result >>= laneid;
if (result)
atomicOr(&out[idx], result);
}
}
// Determine if pixel could be a spot
// params: spot finding parameters
// val: pixel value
// sum: window sum
// sum2: window sum of squares
// count: window valid pixels count
// return the pixel result: 0-no spot / 1-spot candidate
__device__ __forceinline__ uint8_t pixel_result(const spot_parameters& params, const int64_t val, int64_t sum, int64_t sum2, int64_t count)
{
sum -= val;
sum2 -= val * val;
count -= 1;
const int64_t var = count * sum2 - (sum * sum); // This should be divided by ((2*NBX+1) * (2*NBY+1)-1)*((2*NBX+1) * (2*NBY+1))
const int64_t in_minus_mean = val * count - sum; // Should be divided by ((2*NBX+1) * (2*NBY+1));
const int64_t tmp1 = in_minus_mean * in_minus_mean;
const float tmp2 = var * params.strong_pixel_threshold2;
bool snr_criterion;
if (params.strong_pixel_threshold2 == 0.0f)
snr_criterion = true;
else
snr_criterion = (count > ImageSpotFinder::MIN_VALID_PIXELS) && (in_minus_mean > 0) && (tmp1 > tmp2);
bool count_criterion = (params.count_threshold == 0.0f) || (val > params.count_threshold);
bool strong_pixel = snr_criterion && count_criterion;
if (val == INT32_MAX)
strong_pixel = true;
else if (val == INT32_MIN)
strong_pixel = false;
return strong_pixel ? 1 : 0;
}
// Find pixels that could be spots
// in: image input values
// out: pixel result bit array, 1 bit per pixel (0:no/1:candidate spot)
// params: spot finding parameters
//
// The algorithm uses multiple waves (blockDim.y) that run over sections of rows.
// Each wave will write output at the back row and read input at the front row.
// Each wave is split into column output sections (blockDim.x)
// A wave section (block) is responsible for a particular row/column section and
// maintains sum/sum2/count values per column for the output row.
// Every cuda thread is associated with a particular column. The thread maintains
// the sum/sum2/count values in shared memory for it's column. To do this, the input
// pixel values for the hight of the aggregation window are saved in shared memory.
__global__ void analyze_pixel(const int32_t *in, uint32_t *prev_out, uint32_t *out, const spot_parameters params)
{
// assumption: 2 * params.nby + 1 <= params.rows and 2 * params.nbx + 1 <= params.width
const int32_t window = 2 * (int)ImageSpotFinder::NBX + 1; // vertical window
const int32_t writeSize = blockDim.x - 2 * ImageSpotFinder::NBX; // output columns per block
const int32_t cmin = blockIdx.x * writeSize; // lowest output column
const int32_t cmax = min(cmin + writeSize, static_cast<int32_t>(params.width)); // past highest output column
const int32_t col = cmin + threadIdx.x - ImageSpotFinder::NBX; // thread -> column mapping
const bool data_col = (col >= 0) && (col < static_cast<int32_t>(params.width)); // read global mem
const bool result_col = (col >= cmin) && (col < cmax); // write result
const int32_t nWaves = gridDim.y; // number of waves
const int32_t rowsPerWave = (params.height + nWaves - 1) / nWaves; // rows per wave
const int32_t rmin = blockIdx.y * rowsPerWave; // lowest result row for this wave
const int32_t rmax = min(rmin + rowsPerWave, static_cast<int32_t>(params.height)); // past highest result row for this wave
// rowsPerWave is rounded up, so the last waves can start at or past the last row and have no
// rows to write. rmin depends only on blockIdx.y, so the whole block leaves together and the
// __syncthreads/__ballot_sync below stay collective.
if (rmin >= static_cast<int32_t>(params.height))
return;
const int32_t left = max(static_cast<int32_t>(threadIdx.x) - static_cast<int32_t>(ImageSpotFinder::NBX), 0); // leftmost column touched by this thread
const int32_t right = min(static_cast<int32_t>(threadIdx.x) + static_cast<int32_t>(ImageSpotFinder::NBX) + 1, static_cast<int32_t>(params.width)); // past rightmost column touched by this thread
int32_t back = rmin; // back of wave for writing
int32_t front = max(back - static_cast<int32_t>(ImageSpotFinder::NBX), 0); // front of wave for reading (needs to overtake back initially)
extern __shared__ int64_t shared_mem[];
int64_t* shared_sum = shared_mem; // shared buffer [blockDim.x]
int64_t* shared_sum2 = &shared_sum[blockDim.x]; // shared buffer [blockDim.x]
auto shared_count = reinterpret_cast<int16_t*>(&shared_sum2[blockDim.x]); // shared buffer [blockDim.x]
auto shared_val = reinterpret_cast<int32_t *>(&shared_count[blockDim.x]); // shared cyclic buffer [window, blockDim.x]
int64_t total_sum; // totals
int64_t total_sum2;
int32_t total_count;
// initialize sum, sum2, count, val buffers
shared_sum[threadIdx.x] = 0; // shared values without effect on totals
shared_sum2[threadIdx.x] = 0;
shared_count[threadIdx.x] = 0;
for (int i=0; i<window; i++)
shared_val[i * blockDim.x + threadIdx.x] = INT32_MIN; // value that is NOT counted
// wave front up to rmin + nby + 1
do {
// Rows past the last one contribute nothing; shared_val keeps its INT32_MIN initialiser.
if (data_col && (front < static_cast<int32_t>(params.height))) { // read at the front end of the wave
const int32_t npixel = front * params.width + col;
const bool sat = ((prev_out[npixel / 32] & (1U << (npixel % 32))) != 0);
const int32_t val = sat ? INT32_MAX : in[npixel];
shared_val[(front % window) * blockDim.x + threadIdx.x] = val;
if (val != INT32_MAX && val != INT32_MIN) {
shared_sum[threadIdx.x] += val;
shared_sum2[threadIdx.x] += static_cast<int64_t>(val) * val; // see the main loop
shared_count[threadIdx.x] += 1;
}
}
front++;
} while (front < rmin + static_cast<int32_t>(ImageSpotFinder::NBX) + 1);
// wave front up to rmax
do {
__syncthreads(); // make others see the shared values
uint8_t val = 0;
if (result_col) { // write at the back end of the wave
total_sum = total_sum2 = total_count = 0;
for (auto j = left; j < right; j++) {
total_sum += shared_sum[j];
total_sum2 += shared_sum2[j];
total_count += shared_count[j];
}
val = pixel_result(params, shared_val[(back % window) * blockDim.x + threadIdx.x], total_sum, total_sum2, total_count);
}
write_result(params, out, back * params.width + col, val);
back++;
__syncthreads(); // keep shared values until others have seen them
if (data_col) { // read at the front end of the wave
int16_t cnt = 0;
int32_t old = shared_val[(front % window) * blockDim.x + threadIdx.x];
if (old == INT32_MAX || old == INT32_MIN) {
old = 0; // no effect value
cnt = 1; // bring count to normal
}
int32_t val = INT32_MIN; // past the last row nothing enters the window
if (front < static_cast<int32_t>(params.height)) {
// A pixel the previous pass found strong reads as INT32_MAX, exactly as the priming
// and drain loops above and below do, and as the CPU finder's value_at() does on
// every read. Without it the second pass counts pass-1 spot pixels as background
// over the whole middle of the image, which is where nearly all rows are.
const int32_t npixel = front * params.width + col;
const bool sat = ((prev_out[npixel / 32] & (1U << (npixel % 32))) != 0);
val = sat ? INT32_MAX : in[npixel];
}
shared_val[(front % window) * blockDim.x + threadIdx.x] = val;
if (val == INT32_MAX || val == INT32_MIN) {
val = 0; // no effect value
cnt -= 1; // count diff from normal
}
shared_sum[threadIdx.x] += val - old;
// 64-bit squares: the accumulator is int64, but val*val in int32 wraps above 46340 and
// the detector saturates far higher (overload ~1e6), which corrupted the variance for
// every window containing a bright pixel.
shared_sum2[threadIdx.x] += static_cast<int64_t>(val) * val
- static_cast<int64_t>(old) * old;
shared_count[threadIdx.x] += cnt;
}
front++;
} while (front < rmax);
// wave back up to rmax
do {
__syncthreads(); // make others see the shared values
uint8_t val = 0;
if (result_col) { // write at the back end of the wave
total_sum = total_sum2 = total_count = 0;
for (auto j = left; j < right; j++) {
total_sum += shared_sum[j];
total_sum2 += shared_sum2[j];
total_count += shared_count[j];
}
val = pixel_result(params, shared_val[(back % window) * blockDim.x + threadIdx.x], total_sum, total_sum2, total_count);
}
write_result(params, out, back * params.width + col, val);
back++;
__syncthreads(); // keep shared values until others have seen them
if (data_col) { // read at the front end of the wave if possible
int16_t cnt = -1; // normal count diff
int32_t old = shared_val[(front % window) * blockDim.x + threadIdx.x];
if (old == INT32_MAX || old == INT32_MIN) {
old = 0; // no effect value
cnt += 1; // bring count to normal
}
int32_t val = 0;
if (front < params.height) {
const int32_t npixel = front * params.width + col;
const bool sat = ((prev_out[npixel / 32] & (1U << (npixel % 32))) != 0);
val = sat ? INT32_MAX : in[npixel];
if (val == INT32_MAX || val == INT32_MIN)
val = 0; // no effect value
else
cnt += 1; // count diff from normal
}
shared_sum[threadIdx.x] += val - old;
shared_sum2[threadIdx.x] += static_cast<int64_t>(val) * val
- static_cast<int64_t>(old) * old; // see the main loop
shared_count[threadIdx.x] += cnt;
}
front++;
} while (back < rmax);
}
ImageSpotFinderGPU::ImageSpotFinderGPU(int32_t in_width, int32_t in_height,
std::shared_ptr<CudaStream> in_stream)
: ImageSpotFinder(in_width, in_height, false),
stream(in_stream),
extractor(in_width, in_height, std::move(in_stream)) {
gpu_out_0 = CudaDevicePtr<uint32_t>(OutputSize());
gpu_out_1 = CudaDevicePtr<uint32_t>(OutputSize());
}
void ImageSpotFinderGPU::SetResolutionMaskBits(const std::vector<uint32_t> &packed_mask) {
ImageSpotFinder::SetResolutionMaskBits(packed_mask);
extractor.SetResolutionMask(res_mask_bits);
}
const std::vector<DiffractionSpot> &ImageSpotFinderGPU::ExtractComponents(const ImagePreprocessorBuffer &image,
const SpotFindingSettings &settings) {
extractor.Extract(gpu_out_1, image.getGPUBuffer(), settings, components);
return components;
}
void ImageSpotFinderGPU::Detect(const ImagePreprocessorBuffer &image, const SpotFindingSettings &settings) {
spot_parameters spot_params{};
spot_params.height = height;
spot_params.width = width;
spot_params.strong_pixel_threshold2 = settings.signal_to_noise_threshold * settings.signal_to_noise_threshold;
spot_params.count_threshold = settings.photon_count_threshold;
if (2 * NBX + 1 > windowSizeLimit)
throw JFJochException(JFJochExceptionCategory::SpotFinderError, "nbx exceeds window size limit");
if (2 * NBX + 1 > windowSizeLimit)
throw JFJochException(JFJochExceptionCategory::SpotFinderError, "nby exceeds window size limit");
if (windowSizeLimit > numberOfCudaThreads)
throw JFJochException(JFJochExceptionCategory::SpotFinderError, "window size limit exceeds number of cuda threads");
if (windowSizeLimit > spot_params.width)
throw JFJochException(JFJochExceptionCategory::SpotFinderError, "window size limit exceeds number of columns");
if (windowSizeLimit > spot_params.height)
throw JFJochException(JFJochExceptionCategory::SpotFinderError, "window size limit exceeds number of height");
const auto nWriters = numberOfCudaThreads - 2 * NBX;
const auto nBlocks = (spot_params.width + nWriters - 1) / nWriters;
const auto window = 2 * NBX + 1;
const auto sharedSize = (2 * sizeof(int64_t) + // sum, sum2
window * sizeof(int32_t) + // val
1 * sizeof(int16_t) // count
) * numberOfCudaThreads;
const dim3 blocks(nBlocks, numberOfWaves);
cuda_err(cudaMemsetAsync(gpu_out_0, 0, OutputSize() * sizeof(uint32_t), *stream));
cuda_err(cudaMemsetAsync(gpu_out_1, 0, OutputSize() * sizeof(uint32_t), *stream));
analyze_pixel<<<blocks, numberOfCudaThreads, sharedSize, *stream>>>
(image.getGPUBuffer(), gpu_out_1, gpu_out_0, spot_params);
analyze_pixel<<<blocks, numberOfCudaThreads, sharedSize, *stream>>>
(image.getGPUBuffer(), gpu_out_0, gpu_out_1, spot_params);
// The bit buffer stays on the device - ExtractComponents reads it there.
cuda_err(cudaStreamSynchronize(*stream));
}