Build Packages / Unit tests (push) Successful in 1h22m15s
Build Packages / build:windows:nocuda (push) Successful in 18m0s
Build Packages / build:windows:cuda (push) Successful in 20m30s
Build Packages / build:viewer-tgz:cpu (push) Successful in 10m32s
Build Packages / build:viewer-tgz:cuda (push) Successful in 11m39s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 8m55s
Build Packages / build:rugnux:windows (push) Successful in 11m25s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 20m6s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 16m27s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 20m19s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 15m34s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 20m25s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 19m36s
Build Packages / build:rpm (rocky8) (push) Successful in 17m43s
Build Packages / build:rpm (rocky9) (push) Successful in 13m34s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 21m28s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 18m19s
Build Packages / DIALS test (push) Successful in 12m36s
Build Packages / XDS test (durin plugin) (push) Successful in 6m56s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 6m48s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m7s
Build Packages / Generate python client (push) Successful in 11s
Build Packages / Build documentation (push) Successful in 36s
Build Packages / Create release (push) Skipped
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 5m11s
* `rugnux --mode calibration` writes `<prefix>.json` beside the `.poni`, whose `dataset_settings` member is a `jfjoch_broker` `dataset_settings` body as it stands. * `rugnux` and `jfjoch_viewer` read PILATUS miniCBF sweeps natively, without conversion. * Masters written by other facilities open, including Eiger 1.x and third-party NXmx variants. * `rugnux` measures the beam centre on every run, and indexes with it when the file's value indexes nothing. * A detector swung out on a 2theta arm is placed where the file says it stands, and the calibration can hold the tilt fixed. * `rugnux` writes the unmerged MTZ by default, and a P1 merge beside it, so a wrong space group can be re-merged without reprocessing. * Significant improvements to symmetry handling in `rugnux`: the lattice, the point group, the setting and the systematic absences. * The `rugnux` report gives the resolution the CC1/2 fit reached, beside the range the reflections were written to. * The `rugnux` report gives the twinning statistics measured before the space group was decided, beside the ones measured after. * The `rugnux` report gives the strong-direction diffraction limit, and warns when CC1/2 is not monotone with resolution. * `rugnux` ranks screw axes on the evidence their absences carry, rather than on how many control reflections a candidate happens to have. * Twinning is no longer reported when the L-test contradicts it. * The `rugnux` report gives the detector tilt, the measured tilt and the direct beam beside the beam centre, and a post-refined beam centre is judged against the run's own measurement rather than the file's. * `--no-refine-tilt` holds the detector tilt at the value in the file, instead of zeroing it, when the calibration starts from the spots. * The `jfjoch_viewer` grid scan view draws the cells in the proportion of the scan steps, so the map has the shape of the scanned area. Reviewed-on: #76 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
101 lines
5.0 KiB
C++
101 lines
5.0 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
// Device-side connected-component extraction for the GPU spot finders.
|
|
//
|
|
// The GPU finders flag strong pixels into a packed bit buffer ON THE DEVICE. Reading spots out of it
|
|
// used to mean copying that whole buffer back (2.26 MB per frame at 18 MP) and scanning it bit by bit
|
|
// on the host. This does the whole extraction where the data already is, so nothing about the image
|
|
// comes back - only the finished spot list, a few hundred entries.
|
|
//
|
|
// The algorithm is the sparse formulation the ACTS/traccc project settled on for the same problem
|
|
// (sparse silicon-detector hits): the strong pixels are compacted into a list that is sorted by flat
|
|
// index, each pixel finds its at most FOUR backward 8-neighbours by binary search in that list, and
|
|
// the resulting graph is labelled with a lock-free union-find. A dense image-wide labelling
|
|
// (Playne-equivalence, BUF/BKE, nppiLabelMarkers, cv::cuda::connectedComponents) would label 18
|
|
// million pixels to find five hundred.
|
|
//
|
|
// It reproduces the host StrongPixelSet::sparseccl EXACTLY, not just equivalently:
|
|
// * both make a component's root its lowest list index, so both find the same roots;
|
|
// * labels are handed out by a prefix sum over the roots in ascending order, which is the order the
|
|
// host's second scan hands them out in, so the SPOT ORDER is identical;
|
|
// * the centroid sums are accumulated per component in ascending list order, in integers, term for
|
|
// term as DiffractionSpot::AddPixel does them, so there is no rounding for the two compilers to
|
|
// disagree about.
|
|
// tests/SpotExtractorGPUParityTest.cpp holds the two to each other on realistic, occupancy-swept and
|
|
// pathological frames, and checks that repeating a frame gives byte-identical output.
|
|
|
|
#include <cstdint>
|
|
#include <memory>
|
|
#include <vector>
|
|
|
|
#include "../../common/DiffractionSpot.h"
|
|
#include "../indexing/CUDAMemHelpers.h"
|
|
#include "SpotFindingSettings.h"
|
|
|
|
// Per-component sums, in exactly the form DiffractionSpot holds them: x and y are sum(col*photons)
|
|
// and sum(line*photons), not a centroid.
|
|
struct SpotExtractorGPUSpot {
|
|
int64_t x;
|
|
int64_t y;
|
|
int64_t photons;
|
|
int64_t max_photons;
|
|
int32_t pixel_count;
|
|
// Longer side of the component's bounding box, for SpotShapeAccepted. It fits in what used to be
|
|
// padding, so carrying it costs nothing.
|
|
int32_t bbox_side;
|
|
};
|
|
|
|
class SpotExtractorGPU {
|
|
std::shared_ptr<CudaStream> stream;
|
|
const int32_t width;
|
|
const size_t nwords;
|
|
|
|
// Strong pixels this engine's buffers hold, and above which the extraction gives up on the frame -
|
|
// StrongPixelLimit, so it follows the detector rather than standing at a constant. 104 bytes of
|
|
// device memory apiece, 29 MB on an 18-megapixel detector.
|
|
const uint32_t max_strong;
|
|
// Spots copied back together with their count in one transfer. A frame with more than this many
|
|
// surviving spots - far past anything indexable - simply takes a second copy.
|
|
static constexpr uint32_t SPOT_PREFIX = 4096;
|
|
|
|
int compact_blocks = 0;
|
|
|
|
CudaDevicePtr<uint32_t> gpu_res_mask; // packed, bit set = pixel excluded
|
|
CudaDevicePtr<uint32_t> gpu_block_count;
|
|
CudaDevicePtr<uint32_t> gpu_block_offset;
|
|
CudaDevicePtr<uint32_t> gpu_nstrong;
|
|
CudaDevicePtr<uint32_t> gpu_index; // strong pixels, sorted by flat index
|
|
CudaDevicePtr<int32_t> gpu_value;
|
|
CudaDevicePtr<uint32_t> gpu_parent; // union-find parent
|
|
CudaDevicePtr<uint32_t> gpu_root;
|
|
CudaDevicePtr<uint32_t> gpu_label; // compact label, indexed by root
|
|
CudaDevicePtr<int32_t> gpu_count; // pixels per component
|
|
CudaDevicePtr<SpotExtractorGPUSpot> gpu_spot;
|
|
CudaDevicePtr<SpotExtractorGPUSpot> gpu_spot_out;
|
|
CudaDevicePtr<uint32_t> gpu_nspot;
|
|
|
|
CudaHostPtr<uint32_t> host_nstrong;
|
|
CudaHostPtr<uint32_t> host_nspot;
|
|
CudaHostPtr<SpotExtractorGPUSpot> host_spot; // SPOT_PREFIX entries, pinned
|
|
std::vector<SpotExtractorGPUSpot> overflow_spot; // only for a frame with more spots than that
|
|
|
|
public:
|
|
SpotExtractorGPU(int32_t width, int32_t height, std::shared_ptr<CudaStream> stream);
|
|
|
|
void SetResolutionMask(const std::vector<uint32_t> &packed_mask);
|
|
|
|
// gpu_strong is the finder's device bit buffer, gpu_image the preprocessed image it was built
|
|
// from. Fills spots with every component of at most max-pix pixels, in the same order the host
|
|
// extractor would.
|
|
void Extract(const uint32_t *gpu_strong, const int32_t *gpu_image,
|
|
const SpotFindingSettings &settings, std::vector<DiffractionSpot> &spots);
|
|
|
|
// Strong pixels the last Extract() saw, after the resolution mask. Reported whether or not the
|
|
// frame was given up on, which is the point of it: a frame at or above max_strong yields no spots
|
|
// at all, and this is what says so.
|
|
[[nodiscard]] uint32_t StrongPixelCount() const { return *host_nstrong.get(); }
|
|
};
|