Files
Jungfraujoch/image_analysis/spot_finding/ImageSpotFinderGPU.cu
T
leonarski_fandClaude Opus 5 16bf3408f0 Address code-review findings; make detection limits detector-driven
One changeset, developed together in response to a review of this branch, so the
files carry several of the changes at once. Full test suite passes (733 cases).

Spot finding
- Split ImageSpotFinder into Detect() (flag strong pixels - the expensive
  per-pixel pass) and ExtractSpots() (CCL + min/max-pix + resolution mask), with
  Run() = both. The per-image min-pix escalation now detects ONCE and repeats
  only the cheap extraction, instead of re-running the whole finder four times
  per frame as it did on the default path. It also keeps the winning attempt's
  spot list rather than re-extracting it, so the frame that is integrated is
  exactly the frame that was scored - which a GPU re-extract could not guarantee
  (float atomic ordering).
- spot_finding_time_s no longer swallows indexing time, and indexing_time_s now
  sums every escalation call instead of reporting only the last.

Detection limits follow the detector
- The azimuthal-integration upper q and the spot-finding high-resolution limit
  are now std::optional, in the C++ structs AND in the OpenAPI schema, and
  resolve to the detector's own maximum (DiffractionExperiment::GetDetectorMaxQ_
  recipA). Adaptive detection reads a pixel's ring from the azimuthal bins, so a
  pixel outside that q range could never be strong - the integration range
  silently bounded what detection could see, regardless of the requested
  resolution limit. Regenerated the C++ and TypeScript clients; the viewer and
  the web frontend each gained a "to detector edge" switch.

Detection defaults are now per workflow (measured, not assumed)
- Stills: adaptive detection, min-pix chosen per image, no resolution clipping.
- Rotation: fixed-threshold finder, min-pix 2, 1.5 A limit.
  On a 33-crystal rotation battery, adaptive detection helped four hard crystals
  but deterministically broke three (a lost space group, a halved indexing rate,
  a collapsed merge), and the detector-edge limit cost indexing on a strong
  rotation set (100.0 -> 96.8%). Each is still overridable by its flag, and
  --no-adaptive-spots is new.

Indexer seed escalation
- Stop escalating once a seed's lattice explains >= 90% of the seed spots.
  Previously any frame with >= 80 spots always paid three indexer calls, online
  broker included.

Merge-consistency filter
- --min-image-cc gated on a per-image CC computed BEFORE the stills partiality
  post-refinement and never refreshed; the refiner now recomputes it, so the
  reported CC describes the data that are actually merged.
- Replaced the per-call cc_mask argument with one MergeOnTheFly flag, so the
  merge, the error model and MergeStats can no longer disagree about which
  images are in (the --scale path merged unfiltered while its statistics were
  filtered).

Per-image B-factor refinement (-B) removed
- Measured on four serial-stills datasets: it is a no-op where the per-image fit
  is well conditioned and actively harmful where it is not (CC1/2 -8.1, R_meas
  +23.2 on the weakest large-cell set, whose fits hit their [-50, 200] bounds on
  14-25% of images). It had also been silently DISCARDED since the partiality
  post-refinement landed - reported but not applied. Rather than fix and keep a
  knob with no demonstrated benefit, the flag and the whole image_scale_b_factor
  chain are gone: setting, scaling fit, message field, CBOR, HDF5 write and
  read-back, per-image plot, OpenAPI enum, viewer column and checkbox, docs.
  ScaleOnTheFly no longer needs Ceres at all - the fit is a linear IRLS.
  (The Wilson per-image b_factor is a different quantity and stays.)

Stills partiality width now fits both of its components
- sigma^2 = gamma0^2 + (gamma_e*d*)^2 instead of a purely angular gamma_e*d*
  with gamma0 pinned to 0. Fitted per crystal by least squares of dist_ewald^2
  on d*^2. The angular-only width is fitted over a d*^2-dense population, so it
  was pinned by the high-resolution edge and collapsed at low d*: median
  partiality 0.008 beyond 13 A for reflections that were plainly recorded, 55%
  of them under the merge's partiality floor, and the survivors divided by those
  values - which inflated the merged low-resolution intensity scale 3.6x
  (~ +9 A^2 of apparent B). Measured on 5000 stills: the ramp flattens to 0.89x,
  no observation is dropped any more (701750 -> 716811), shell-mean CC1/2 and
  R-free improve slightly. Note CC1/2, R_meas, completeness and a B-refining
  R-free are all blind to that ramp, which is why it survived earlier validation;
  the cost is high-resolution R_meas (98.5 -> 101.9 shell-averaged).

Removed dead code from add-then-remove churn
- Prediction-time "still partiality" (unreachable: no setter), the phantom
  IndexingSettings::min_indexed_spot_fraction knob (getter, no setter - now the
  constant it always was), StillsPartialityRefine's caller-less Settings
  constructor and its reference to a long-gone env var, ProcessImage's unread
  bool return, an unused include, and a dead viewer overlay hook.

Also
- Viewer: the magnifier compared a QImage with itself, so its scene rect was set
  once ever and it could not pan into a larger dataset; the hover tail timer
  could fire after leaveEvent and resurrect the resolution readout outside the
  image.
- update_version.sh regenerated the frontend lock file BEFORE bumping the
  version (every release shipped an off-by-one lock), and did git rm/git add on
  a path that has not existed since the client moved to src/client - with no
  set -e, both failed silently.
- fpga/pcie_driver/postinstall.sh tested "[ ! occurrences > 0 ]", which is a
  redirect, not a test, so dkms add never ran.
- Unit tests for the adaptive-threshold host functions, which had none.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 09:07:00 +02:00

277 lines
14 KiB
Plaintext

// SPDX-FileCopyrightText: 2025 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
// GPU Spot finding developed by Hans-Christian Stadler (PSI)
// Copyright (2019-2023) Paul Scherrer Institute
#include "ImageSpotFinderGPU.h"
#include "../../common/JFJochException.h"
struct spot_parameters {
int32_t width;
int32_t height;
float strong_pixel_threshold2;
int32_t count_threshold;
};
// input X x Y pixels array
// output X x Y bit array
static constexpr int WARP_SIZE = 32; // assume warp size of 32 cuda threads per warp
inline void cuda_err(cudaError_t val) {
if (val != cudaSuccess)
throw JFJochException(JFJochExceptionCategory::GPUCUDAError, cudaGetErrorString(val));
}
// Write pixel results to bit array
// params: spot finding parameters
// out: pixel result bit array
// pixel: flat pixel index = bit index into bit array
// val: pixel result
// **NOTE**: assumes sizeof(*out) * 8 == WARP_SIZE
__device__ __forceinline__ void write_result(const spot_parameters& params, uint32_t* out, int32_t pixel, uint8_t val)
{
static_assert(sizeof(*out) * 8 == WARP_SIZE, "Violation of essential implementation assumption: WARP_SIZE must match output array element type bit size!");
static constexpr unsigned ALL_THREADS = UINT32_MAX;
const int32_t laneid = threadIdx.x & (WARP_SIZE - 1);
unsigned result = __ballot_sync(ALL_THREADS, val);
const int32_t idx = pixel / WARP_SIZE; // global uint32_t index
const int32_t bit = pixel % WARP_SIZE; // local bit index
if ((bit >= laneid) && (laneid == 0)) { // write to upper part of uint32_t
result <<= bit;
if (result)
atomicOr(&out[idx], result);
} else if ((bit < laneid) && (bit == 0)) { // write to lower part of uint32_t
result >>= laneid;
if (result)
atomicOr(&out[idx], result);
}
}
// Determine if pixel could be a spot
// params: spot finding parameters
// val: pixel value
// sum: window sum
// sum2: window sum of squares
// count: window valid pixels count
// return the pixel result: 0-no spot / 1-spot candidate
__device__ __forceinline__ uint8_t pixel_result(const spot_parameters& params, const int64_t val, int64_t sum, int64_t sum2, int64_t count)
{
sum -= val;
sum2 -= val * val;
count -= 1;
const int64_t var = count * sum2 - (sum * sum); // This should be divided by ((2*NBX+1) * (2*NBY+1)-1)*((2*NBX+1) * (2*NBY+1))
const int64_t in_minus_mean = val * count - sum; // Should be divided by ((2*NBX+1) * (2*NBY+1));
const int64_t tmp1 = in_minus_mean * in_minus_mean;
const float tmp2 = var * params.strong_pixel_threshold2;
bool snr_criterion;
if (params.strong_pixel_threshold2 == 0.0f)
snr_criterion = true;
else
snr_criterion = (count > ImageSpotFinder::MIN_VALID_PIXELS) && (in_minus_mean > 0) && (tmp1 > tmp2);
bool count_criterion = (params.count_threshold == 0.0f) || (val > params.count_threshold);
bool strong_pixel = snr_criterion && count_criterion;
if (val == INT32_MAX)
strong_pixel = true;
else if (val == INT32_MIN)
strong_pixel = false;
return strong_pixel ? 1 : 0;
}
// Find pixels that could be spots
// in: image input values
// out: pixel result bit array, 1 bit per pixel (0:no/1:candidate spot)
// params: spot finding parameters
//
// The algorithm uses multiple waves (blockDim.y) that run over sections of rows.
// Each wave will write output at the back row and read input at the front row.
// Each wave is split into column output sections (blockDim.x)
// A wave section (block) is responsible for a particular row/column section and
// maintains sum/sum2/count values per column for the output row.
// Every cuda thread is associated with a particular column. The thread maintains
// the sum/sum2/count values in shared memory for it's column. To do this, the input
// pixel values for the hight of the aggregation window are saved in shared memory.
__global__ void analyze_pixel(const int32_t *in, uint32_t *prev_out, uint32_t *out, const spot_parameters params)
{
// assumption: 2 * params.nby + 1 <= params.rows and 2 * params.nbx + 1 <= params.width
const int32_t window = 2 * (int)ImageSpotFinder::NBX + 1; // vertical window
const int32_t writeSize = blockDim.x - 2 * ImageSpotFinder::NBX; // output columns per block
const int32_t cmin = blockIdx.x * writeSize; // lowest output column
const int32_t cmax = min(cmin + writeSize, static_cast<int32_t>(params.width)); // past highest output column
const int32_t col = cmin + threadIdx.x - ImageSpotFinder::NBX; // thread -> column mapping
const bool data_col = (col >= 0) && (col < static_cast<int32_t>(params.width)); // read global mem
const bool result_col = (col >= cmin) && (col < cmax); // write result
const int32_t nWaves = gridDim.y; // number of waves
const int32_t rowsPerWave = (params.height + nWaves - 1) / nWaves; // rows per wave
const int32_t rmin = blockIdx.y * rowsPerWave; // lowest result row for this wave
const int32_t rmax = min(rmin + rowsPerWave, static_cast<int32_t>(params.height)); // past highest result row for this wave
const int32_t left = max(static_cast<int32_t>(threadIdx.x) - static_cast<int32_t>(ImageSpotFinder::NBX), 0); // leftmost column touched by this thread
const int32_t right = min(static_cast<int32_t>(threadIdx.x) + static_cast<int32_t>(ImageSpotFinder::NBX) + 1, static_cast<int32_t>(params.width)); // past rightmost column touched by this thread
int32_t back = rmin; // back of wave for writing
int32_t front = max(back - static_cast<int32_t>(ImageSpotFinder::NBX), 0); // front of wave for reading (needs to overtake back initially)
extern __shared__ int64_t shared_mem[];
int64_t* shared_sum = shared_mem; // shared buffer [blockDim.x]
int64_t* shared_sum2 = &shared_sum[blockDim.x]; // shared buffer [blockDim.x]
auto shared_count = reinterpret_cast<int16_t*>(&shared_sum2[blockDim.x]); // shared buffer [blockDim.x]
auto shared_val = reinterpret_cast<int32_t *>(&shared_count[blockDim.x]); // shared cyclic buffer [window, blockDim.x]
int64_t total_sum; // totals
int64_t total_sum2;
int32_t total_count;
// initialize sum, sum2, count, val buffers
shared_sum[threadIdx.x] = 0; // shared values without effect on totals
shared_sum2[threadIdx.x] = 0;
shared_count[threadIdx.x] = 0;
for (int i=0; i<window; i++)
shared_val[i * blockDim.x + threadIdx.x] = INT32_MIN; // value that is NOT counted
// wave front up to rmin + nby + 1
do {
if (data_col) { // read at the front end of the wave
const int32_t npixel = front * params.width + col;
const bool sat = ((prev_out[npixel / 32] & (1U << (npixel % 32))) != 0);
const int32_t val = sat ? INT32_MAX : in[npixel];
shared_val[(front % window) * blockDim.x + threadIdx.x] = val;
if (val != INT32_MAX && val != INT32_MIN) {
shared_sum[threadIdx.x] += val;
shared_sum2[threadIdx.x] += val * val;
shared_count[threadIdx.x] += 1;
}
}
front++;
} while (front < rmin + static_cast<int32_t>(ImageSpotFinder::NBX) + 1);
// wave front up to rmax
do {
__syncthreads(); // make others see the shared values
uint8_t val = 0;
if (result_col) { // write at the back end of the wave
total_sum = total_sum2 = total_count = 0;
for (auto j = left; j < right; j++) {
total_sum += shared_sum[j];
total_sum2 += shared_sum2[j];
total_count += shared_count[j];
}
val = pixel_result(params, shared_val[(back % window) * blockDim.x + threadIdx.x], total_sum, total_sum2, total_count);
}
write_result(params, out, back * params.width + col, val);
back++;
__syncthreads(); // keep shared values until others have seen them
if (data_col) { // read at the front end of the wave
int16_t cnt = 0;
int32_t old = shared_val[(front % window) * blockDim.x + threadIdx.x];
if (old == INT32_MAX || old == INT32_MIN) {
old = 0; // no effect value
cnt = 1; // bring count to normal
}
int32_t val = in[front * params.width + col];
shared_val[(front % window) * blockDim.x + threadIdx.x] = val;
if (val == INT32_MAX || val == INT32_MIN) {
val = 0; // no effect value
cnt -= 1; // count diff from normal
}
shared_sum[threadIdx.x] += val - old;
shared_sum2[threadIdx.x] += val * val - old * old;
shared_count[threadIdx.x] += cnt;
}
front++;
} while (front < rmax);
// wave back up to rmax
do {
__syncthreads(); // make others see the shared values
uint8_t val = 0;
if (result_col) { // write at the back end of the wave
total_sum = total_sum2 = total_count = 0;
for (auto j = left; j < right; j++) {
total_sum += shared_sum[j];
total_sum2 += shared_sum2[j];
total_count += shared_count[j];
}
val = pixel_result(params, shared_val[(back % window) * blockDim.x + threadIdx.x], total_sum, total_sum2, total_count);
}
write_result(params, out, back * params.width + col, val);
back++;
__syncthreads(); // keep shared values until others have seen them
if (data_col) { // read at the front end of the wave if possible
int16_t cnt = -1; // normal count diff
int32_t old = shared_val[(front % window) * blockDim.x + threadIdx.x];
if (old == INT32_MAX || old == INT32_MIN) {
old = 0; // no effect value
cnt += 1; // bring count to normal
}
int32_t val = 0;
if (front < params.height) {
const int32_t npixel = front * params.width + col;
const bool sat = ((prev_out[npixel / 32] & (1U << (npixel % 32))) != 0);
val = sat ? INT32_MAX : in[npixel];
if (val == INT32_MAX || val == INT32_MIN)
val = 0; // no effect value
else
cnt += 1; // count diff from normal
}
shared_sum[threadIdx.x] += val - old;
shared_sum2[threadIdx.x] += val * val - old * old;
shared_count[threadIdx.x] += cnt;
}
front++;
} while (back < rmax);
}
ImageSpotFinderGPU::ImageSpotFinderGPU(int32_t in_width, int32_t in_height,
std::shared_ptr<CudaStream> stream)
: ImageSpotFinder(in_width, in_height),
stream(std::move(stream)),
output_buffer_reg(output_buffer) {
gpu_out_0 = CudaDevicePtr<uint32_t>(OutputSize());
gpu_out_1 = CudaDevicePtr<uint32_t>(OutputSize());
}
void ImageSpotFinderGPU::Detect(const ImagePreprocessorBuffer &image, const SpotFindingSettings &settings) {
spot_parameters spot_params{};
spot_params.height = height;
spot_params.width = width;
spot_params.strong_pixel_threshold2 = settings.signal_to_noise_threshold * settings.signal_to_noise_threshold;
spot_params.count_threshold = settings.photon_count_threshold;
if (2 * NBX + 1 > windowSizeLimit)
throw JFJochException(JFJochExceptionCategory::SpotFinderError, "nbx exceeds window size limit");
if (2 * NBX + 1 > windowSizeLimit)
throw JFJochException(JFJochExceptionCategory::SpotFinderError, "nby exceeds window size limit");
if (windowSizeLimit > numberOfCudaThreads)
throw JFJochException(JFJochExceptionCategory::SpotFinderError, "window size limit exceeds number of cuda threads");
if (windowSizeLimit > spot_params.width)
throw JFJochException(JFJochExceptionCategory::SpotFinderError, "window size limit exceeds number of columns");
if (windowSizeLimit > spot_params.height)
throw JFJochException(JFJochExceptionCategory::SpotFinderError, "window size limit exceeds number of height");
const auto nWriters = numberOfCudaThreads - 2 * NBX;
const auto nBlocks = (spot_params.width + nWriters - 1) / nWriters;
const auto window = 2 * NBX + 1;
const auto sharedSize = (2 * sizeof(int64_t) + // sum, sum2
window * sizeof(int32_t) + // val
1 * sizeof(int16_t) // count
) * numberOfCudaThreads;
const dim3 blocks(nBlocks, numberOfWaves);
cuda_err(cudaMemsetAsync(gpu_out_0, 0, OutputSize() * sizeof(uint32_t), *stream));
cuda_err(cudaMemsetAsync(gpu_out_1, 0, OutputSize() * sizeof(uint32_t), *stream));
analyze_pixel<<<blocks, numberOfCudaThreads, sharedSize, *stream>>>
(image.getGPUBuffer(), gpu_out_1, gpu_out_0, spot_params);
analyze_pixel<<<blocks, numberOfCudaThreads, sharedSize, *stream>>>
(image.getGPUBuffer(), gpu_out_0, gpu_out_1, spot_params);
cuda_err(cudaMemcpyAsync(output_buffer.data(), gpu_out_1, OutputSize() * sizeof(uint32_t), cudaMemcpyDeviceToHost, *stream));
cuda_err(cudaStreamSynchronize(*stream));
}