Build Packages / Create release (push) Successful in 40s
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 6m58s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 8m35s
Build Packages / build:viewer-tgz:cpu (push) Successful in 9m50s
Build Packages / build:viewer-tgz:cuda (push) Successful in 11m18s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 12m38s
Build Packages / build:windows:nocuda (push) Successful in 17m14s
Build Packages / build:windows:cuda (push) Successful in 19m46s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 22m17s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 17m8s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 16m25s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 17m37s
Build Packages / build:rugnux:windows (push) Successful in 10m36s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 16m38s
Build Packages / Generate python client (push) Successful in 17s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 16m37s
Build Packages / Build documentation (push) Successful in 1m7s
Build Packages / build:rpm (rocky8) (push) Successful in 17m11s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 16m55s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 17m31s
Build Packages / build:rpm (rocky9) (push) Successful in 19m38s
Build Packages / Unit tests (push) Successful in 1h41m0s
From the five-agent review of rc168-rc170 and the ticking-bomb hunt: - Every CUDA kernel launch is followed by cuda_err(cudaGetLastError()) (or the file's own check idiom) - 53 launches in 11 files ran unchecked, so a non-sticky launch failure (out-of-resources on a shared GPU, a zero grid) silently handed stale device buffers downstream as good data. The FFT indexer got this check in bfe95b4ed; this is the same gap everywhere else. BeamCenterFFTGPU already checked every launch through CheckLastKernel. - ShadowFinder: a non-finite or absurd beam centre is refused before it can become a negative ring index (an out-of-bounds write) or an arbitrarily large per-ring table; a pixel whose polarization correction is not strictly positive is not usable - divided by zero it put an inf into the pooled means, which the running box sums turn into NaN for a whole row. - FileWriter: the network-supplied image number is bounded by the collection's declared number_of_images - unbounded it sized per-image vectors, a huge value was a fatal allocation and a wrapping product an out-of-bounds heap write. - ROICircle/ROIAzimuthal: parameters must be finite, not merely positive - NaN passes every <= test, inf passes > 0, and both reached the preview drawing where a non-finite loop bound hangs the rendering thread. - Reader + viewer: documented that SWMR / growing HDF5 files are not supported - a file we open is expected to be final, which is why re-opening the currently open path deliberately does not re-read it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
52 lines
2.0 KiB
Plaintext
52 lines
2.0 KiB
Plaintext
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#include "ImagePreprocessorBufferGPU.h"
|
|
|
|
static inline void cuda_err(cudaError_t val) {
|
|
if (val != cudaSuccess)
|
|
throw JFJochException(JFJochExceptionCategory::GPUCUDAError, cudaGetErrorString(val));
|
|
}
|
|
|
|
__global__ void gather_kernel(const int32_t *__restrict__ image,
|
|
const uint32_t *__restrict__ npixel,
|
|
int32_t *__restrict__ values,
|
|
int count) {
|
|
for (int i = blockIdx.x * blockDim.x + threadIdx.x; i < count; i += blockDim.x * gridDim.x)
|
|
values[i] = image[npixel[i]];
|
|
}
|
|
|
|
ImagePreprocessorBufferGPU::ImagePreprocessorBufferGPU(size_t npixel, bool host_mirror)
|
|
: ImagePreprocessorBuffer(npixel, host_mirror),
|
|
gpu_image(npixel),
|
|
// A no-op when the mirror was not allocated: CudaRegisteredVector skips an empty vector.
|
|
buffer_reg(buffer),
|
|
max_gather(StrongPixelLimit(npixel)),
|
|
gpu_gather_index(max_gather),
|
|
gpu_gather_value(max_gather) {
|
|
}
|
|
|
|
int32_t *ImagePreprocessorBufferGPU::getGPUBuffer() {
|
|
return gpu_image;
|
|
}
|
|
|
|
const int32_t *ImagePreprocessorBufferGPU::getGPUBuffer() const {
|
|
return gpu_image;
|
|
}
|
|
|
|
void ImagePreprocessorBufferGPU::Gather(const std::vector<uint32_t> &npixel, std::vector<int32_t> &values) const {
|
|
values.resize(npixel.size());
|
|
if (npixel.empty())
|
|
return;
|
|
|
|
const int count = static_cast<int>(npixel.size());
|
|
cudaMemcpyAsync(gpu_gather_index.get(), npixel.data(), count * sizeof(uint32_t),
|
|
cudaMemcpyHostToDevice, gather_stream);
|
|
gather_kernel<<<(count + 255) / 256, 256, 0, gather_stream>>>(
|
|
gpu_image.get(), gpu_gather_index.get(), gpu_gather_value.get(), count);
|
|
cuda_err(cudaGetLastError());
|
|
cudaMemcpyAsync(values.data(), gpu_gather_value.get(), count * sizeof(int32_t),
|
|
cudaMemcpyDeviceToHost, gather_stream);
|
|
cudaStreamSynchronize(gather_stream);
|
|
}
|