Build Packages / Create release (push) Successful in 24s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m29s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m43s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 8m27s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 9m53s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 9m58s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 11m22s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 13m39s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 18m37s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 16m32s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 24m11s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m30s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 19m3s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 20m23s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 19m41s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 21m0s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m38s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 14m33s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 17m55s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 20m50s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 18m38s
Build Packages / Unit tests (push) Successful in 1h46m14s
* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports. * jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls. * Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results. * Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable. * Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate. * Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do. * Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence. * Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags. * Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check. * Rugnux: Clear error messages when a data set needs more GPU or host memory than is available. Reviewed-on: #83 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
67 lines
3.3 KiB
C++
67 lines
3.3 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <cstdint>
|
|
#include <memory>
|
|
#include <mutex>
|
|
#include <vector>
|
|
|
|
#include "../image_analysis/indexing/CUDAMemHelpers.h"
|
|
|
|
// The device half of HotPixelFinder, for frames already preprocessed on the GPU: the per-frame order
|
|
// statistics (each ring-sector's median, each ring's median and median absolute deviation) and the
|
|
// per-pixel sums run where the image already is, and only the per-key statistics - tens of thousands
|
|
// of numbers - come to the host, which turns them into levels and thresholds with the very code the
|
|
// host path uses. Every statistic is an exact order statistic of integers and every sum an integer,
|
|
// so the sums, and with them the mask, are identical to what HotPixelFinder::AddImage produces.
|
|
class HotPixelFinderGPU {
|
|
const size_t npixels;
|
|
const size_t nkeys;
|
|
const int nrings;
|
|
const int sectors;
|
|
|
|
CudaDevicePtr<int32_t> key; // ring * sectors + sector of each pixel, -1 masked
|
|
CudaDevicePtr<uint32_t> pixels_by_key; // the unmasked pixels, grouped by key
|
|
CudaDevicePtr<uint32_t> key_begin; // where each key's pixels start in pixels_by_key
|
|
|
|
// The per-pixel sums, exactly those of HotPixelFinder.
|
|
CudaDevicePtr<uint16_t> n_lit, n_error, n_error_ring_ok;
|
|
CudaDevicePtr<int64_t> sum_value, error_level_sum;
|
|
// Frames are selected on their workers' streams in parallel, but each pixel's sums are plain
|
|
// read-modify-writes, so one frame at a time adds to them.
|
|
std::mutex accumulate_mutex;
|
|
|
|
public:
|
|
// One worker's buffers, on the stream its frames are preprocessed on.
|
|
struct Frame {
|
|
explicit Frame(std::shared_ptr<CudaStream> stream) : stream(std::move(stream)) {}
|
|
std::shared_ptr<CudaStream> stream;
|
|
CudaDevicePtr<uint32_t> count; // valid pixels per key
|
|
CudaDevicePtr<int32_t> sector_median; // per key
|
|
CudaDevicePtr<int32_t> ring_median, ring_mad;
|
|
CudaDevicePtr<int32_t> level;
|
|
CudaDevicePtr<float> threshold;
|
|
CudaDevicePtr<char> ring_ok;
|
|
};
|
|
|
|
HotPixelFinderGPU(const int32_t *key, size_t npixels, const std::vector<uint32_t> &key_begin, int nrings,
|
|
int sectors);
|
|
|
|
// The lower median of the valid values of each key (count[k] of them, 0 where there are none), and
|
|
// of each ring the median and the lower median of the absolute deviations from it.
|
|
void Statistics(const int32_t *device_image, Frame &frame, std::vector<uint32_t> &count,
|
|
std::vector<int32_t> §or_median, std::vector<int32_t> &ring_median,
|
|
std::vector<int32_t> &ring_mad);
|
|
|
|
// Add the frame to the per-pixel sums, with each key's level and lit threshold and each ring's
|
|
// verdict on whether it has a level at all - as HotPixelFinder::AddImage does.
|
|
void Accumulate(const int32_t *device_image, Frame &frame, const std::vector<int32_t> &level,
|
|
const std::vector<float> &threshold, const std::vector<char> &ring_ok);
|
|
|
|
// The per-pixel sums, npixels each.
|
|
void Download(uint16_t *n_lit, uint16_t *n_error, int64_t *sum_value, uint16_t *n_error_ring_ok,
|
|
int64_t *error_level_sum);
|
|
};
|