Build Packages / Create release (push) Successful in 24s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m29s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m43s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 8m27s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 9m53s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 9m58s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 11m22s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 13m39s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 18m37s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 16m32s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 24m11s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m30s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 19m3s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 20m23s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 19m41s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 21m0s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m38s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 14m33s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 17m55s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 20m50s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 18m38s
Build Packages / Unit tests (push) Successful in 1h46m14s
* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports. * jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls. * Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results. * Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable. * Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate. * Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do. * Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence. * Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags. * Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check. * Rugnux: Clear error messages when a data set needs more GPU or host memory than is available. Reviewed-on: #83 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
51 lines
2.5 KiB
C++
51 lines
2.5 KiB
C++
// SPDX-FileCopyrightText: 2025 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <vector>
|
|
|
|
#include "SpotFindingSettings.h"
|
|
#include "ImageSpotFinder.h"
|
|
#include "SpotExtractorGPU.h"
|
|
#include "../indexing/CUDAMemHelpers.h"
|
|
|
|
class ImageSpotFinderGPU : public ImageSpotFinder {
|
|
protected:
|
|
// Protected rather than private because AdaptiveSpotFinderGPU derives from this engine: it is
|
|
// this same local-box detection with the fixed photon floor replaced by a per-resolution-ring
|
|
// one, so it reuses the stream, the bit buffers and the extractor rather than owning a second
|
|
// set of them.
|
|
std::shared_ptr<CudaStream> stream;
|
|
|
|
CudaDevicePtr<uint32_t> gpu_out_0;
|
|
CudaDevicePtr<uint32_t> gpu_out_1; // holds the strong-pixel bits after Detect()
|
|
SpotExtractorGPU extractor;
|
|
|
|
private:
|
|
const int numberOfCudaThreads = 128; // #threads per block of analyze_pixel (one warp per 32 columns)
|
|
const int numberOfWaves = 32; // #row bands of analyze_pixel
|
|
const int windowSizeLimit = 32; // limit on the window size (2nby+1, 2nbx+1): a warp holds 2 x 32 columns
|
|
int candidateBlocks = 0; // grid of analyze_candidates (256 threads per block)
|
|
|
|
void RunDetect(const ImagePreprocessorBuffer &image, const SpotFindingSettings &settings,
|
|
const uint32_t *gpu_candidates);
|
|
protected:
|
|
// Detect, but with the result wanted only at the candidate pixels (device bit buffer, the layout
|
|
// of the output): gpu_out_1 ends up holding (Detect's result & candidates), and the second pass is
|
|
// evaluated at the candidates alone. Asynchronous - it does not wait for the stream.
|
|
void DetectAt(const ImagePreprocessorBuffer &image, const SpotFindingSettings &settings,
|
|
const uint32_t *gpu_candidates);
|
|
public:
|
|
ImageSpotFinderGPU(int32_t width, int32_t height, std::shared_ptr<CudaStream> stream);
|
|
~ImageSpotFinderGPU() override = default;
|
|
|
|
void Detect(const ImagePreprocessorBuffer &image, const SpotFindingSettings &settings) override;
|
|
void SetResolutionMaskBits(const std::vector<uint32_t> &packed_mask) override;
|
|
[[nodiscard]] uint32_t StrongPixelCount() const override { return extractor.StrongPixelCount(); }
|
|
const std::vector<DiffractionSpot> &ExtractComponents(const ImagePreprocessorBuffer &image,
|
|
const SpotFindingSettings &settings) override;
|
|
};
|
|
|
|
|