Build Packages / Create release (push) Successful in 24s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m29s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m43s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 8m27s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 9m53s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 9m58s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 11m22s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 13m39s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 18m37s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 16m32s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 24m11s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m30s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 19m3s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 20m23s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 19m41s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 21m0s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m38s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 14m33s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 17m55s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 20m50s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 18m38s
Build Packages / Unit tests (push) Successful in 1h46m14s
* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports. * jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls. * Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results. * Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable. * Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate. * Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do. * Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence. * Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags. * Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check. * Rugnux: Clear error messages when a data set needs more GPU or host memory than is available. Reviewed-on: #83 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
113 lines
5.7 KiB
C++
113 lines
5.7 KiB
C++
// SPDX-FileCopyrightText: 2025 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <memory>
|
|
#include <mutex>
|
|
#include <optional>
|
|
#include "DiffractionExperiment.h"
|
|
#include "PixelMask.h"
|
|
|
|
// The per-pixel geometry of the last mapping built through it, before the mask: every pixel's
|
|
// resolution, correction and bin. None of these depends on the mask - a masked pixel is only left
|
|
// out - so a mapping for the same geometry under another mask takes them from here and blanks its own
|
|
// masked pixels, instead of evaluating the geometry of every pixel again. Building a mapping is that
|
|
// evaluation, and on a 16 Mpx detector it is a tenth of a second on every core. One entry; a mapping
|
|
// being built holds it, so two built at once for the same geometry evaluate it once.
|
|
class AzimuthalIntegrationGeometryCache {
|
|
friend class AzimuthalIntegrationMapping;
|
|
std::mutex m;
|
|
std::vector<uint32_t> key; // bit patterns of everything SetupPixel reads
|
|
std::vector<uint16_t> pixel_to_bin;
|
|
std::vector<float> pixel_resolution;
|
|
std::vector<float> corrections;
|
|
};
|
|
|
|
class AzimuthalIntegrationMapping {
|
|
protected:
|
|
const AzimuthalIntegrationSettings settings;
|
|
const float wavelength;
|
|
const size_t width, height;
|
|
|
|
std::vector<float> bin_to_q;
|
|
std::vector<float> bin_to_2theta;
|
|
std::vector<float> bin_to_d;
|
|
std::vector<float> bin_to_phi;
|
|
std::vector<uint16_t> pixel_to_bin;
|
|
|
|
std::vector<float> pixel_resolution;
|
|
std::vector<float> corrections;
|
|
|
|
std::optional<float> polarization_factor;
|
|
|
|
// Checksums of the two tables the GPU engines upload, taken once here. They are part of the key
|
|
// the shared device-table cache looks them up by (CudaSharedTables.h), so an engine that hands
|
|
// its own checksum in does not have to hash tens of megabytes on its way to a cache hit - and
|
|
// there is one engine per worker per pass. Both vectors are written in the constructor and never
|
|
// touched again, so a checksum taken there stays true for the mapping's whole life.
|
|
uint64_t pixel_to_bin_checksum = 0;
|
|
uint64_t corrections_checksum = 0;
|
|
|
|
// Bit-packed "this pixel is outside the resolution limits" mask, memoised for the limits it was
|
|
// last built for. Every worker's spot finder wants the identical mask and building it walks every
|
|
// pixel, so it is built once and shared. The mapping is const and read from all the worker threads
|
|
// at once, hence the mutex; and the mask is handed out as a shared_ptr, so a caller keeps the one
|
|
// it was given even if another thread later replaces the cached one.
|
|
mutable std::mutex res_mask_mutex;
|
|
mutable std::optional<float> res_mask_high, res_mask_low;
|
|
mutable std::shared_ptr<const std::vector<uint32_t>> res_mask_bits;
|
|
|
|
size_t nthreads;
|
|
|
|
void UpdateMaxBinNumber();
|
|
|
|
// A null mask sets up every pixel.
|
|
void SetupRawGeom(const DiffractionExperiment& experiment, const std::vector<uint32_t> &mask);
|
|
void SetupConvGeomRows(const DiffractionGeometry &geom, const std::vector<uint32_t> *mask, size_t row0, size_t row_end);
|
|
void SetupConvGeom(const DiffractionGeometry &geom, const std::vector<uint32_t> *mask);
|
|
void SetupConvGeomCached(const DiffractionGeometry &geom, const std::vector<uint32_t> &mask,
|
|
AzimuthalIntegrationGeometryCache &cache);
|
|
|
|
void SetupPixel(const DiffractionGeometry &geom, const std::vector<uint32_t> *mask,
|
|
uint32_t pxl, uint32_t col, uint32_t row);
|
|
|
|
AzimuthalIntegrationMapping(const DiffractionExperiment& experiment,
|
|
const PixelMask& mask,
|
|
AzimuthalIntegrationGeometryCache *cache,
|
|
size_t nthreads);
|
|
public:
|
|
|
|
AzimuthalIntegrationMapping(const DiffractionExperiment& experiment,
|
|
const PixelMask& mask,
|
|
size_t nthreads = 0);
|
|
// The same mapping, with the geometry of the pixels taken from `cache` when it holds this geometry
|
|
// (AzimuthalIntegrationGeometryCache), and left there for the next one when it does not.
|
|
AzimuthalIntegrationMapping(const DiffractionExperiment& experiment,
|
|
const PixelMask& mask,
|
|
AzimuthalIntegrationGeometryCache &cache,
|
|
size_t nthreads = 0);
|
|
|
|
[[nodiscard]] uint16_t GetBinNumber() const;
|
|
[[nodiscard]] const std::vector<uint16_t>& GetPixelToBin() const;
|
|
[[nodiscard]] const std::vector<float> &GetBinToQ() const;
|
|
[[nodiscard]] const std::vector<float> &GetBinToD() const;
|
|
[[nodiscard]] const std::vector<float> &GetBinToTwoTheta() const;
|
|
[[nodiscard]] const std::vector<float> &GetBinToPhi() const;
|
|
[[nodiscard]] uint16_t QToBin(float q) const;
|
|
[[nodiscard]] const std::vector<float> &Corrections() const;
|
|
[[nodiscard]] const std::vector<float> &Resolution() const;
|
|
[[nodiscard]] uint64_t GetPixelToBinChecksum() const;
|
|
[[nodiscard]] uint64_t GetCorrectionsChecksum() const;
|
|
// Pixels the spot finders must ignore because their resolution falls outside the limits, packed
|
|
// 32 pixels to a word (bit set = ignore), in the finders' own layout.
|
|
[[nodiscard]] std::shared_ptr<const std::vector<uint32_t>>
|
|
ResolutionMaskBits(std::optional<float> high_res, std::optional<float> low_res) const;
|
|
[[nodiscard]] size_t GetWidth() const;
|
|
[[nodiscard]] size_t GetHeight() const;
|
|
[[nodiscard]] const AzimuthalIntegrationSettings& Settings() const;
|
|
[[nodiscard]] int32_t GetAzimuthalBinCount() const;
|
|
[[nodiscard]] int32_t GetQBinCount() const;
|
|
[[nodiscard]] size_t GetNThreads() const;
|
|
};
|