Files
Jungfraujoch/common/AzimuthalIntegrationMapping.h
T
leonarski_fandClaude Opus 5.5 4616921087 AzimuthalIntegrationMapping: reuse the per-pixel geometry of the last mapping; pre-scan's off the critical path
Building a mapping evaluates the geometry of every pixel (resolution, azimuth, solid-angle and
polarization corrections, bin) - about 2 CPU-s and 0.1-0.25 s of wall on a 16 Mpx detector - and a
rotation run built seven, several for a geometry it had already evaluated: the first pass repeats the
pre-scan's, and the canonical pass and the second of two concurrent probes repeat the probes'. None of
it depends on the mask beyond skipping masked pixels, so AzimuthalIntegrationGeometryCache keeps the
unmasked tables of the last geometry (keyed bit for bit on everything SetupPixel reads), and a
mapping for the same geometry copies them and blanks its own masked pixels to what SetupPixel leaves
there. One entry, shared by the run's copies; on myob four of the seven builds now evaluate, three
copy (~40-120 ms against ~240+ ms on a loaded box).

The pre-scan's mapping is only read by the spot measurement, so it is built at the start of that
background task rather than on the main thread ahead of the projection read. It also fills the cache
for the first pass.

Output identical on the three GPU reference sweeps; a new test checks cached against fresh mappings
under different masks and centres.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-27 09:12:25 +02:00

113 lines
5.7 KiB
C++

// SPDX-FileCopyrightText: 2025 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <memory>
#include <mutex>
#include <optional>
#include "DiffractionExperiment.h"
#include "PixelMask.h"
// The per-pixel geometry of the last mapping built through it, before the mask: every pixel's
// resolution, correction and bin. None of these depends on the mask - a masked pixel is only left
// out - so a mapping for the same geometry under another mask takes them from here and blanks its own
// masked pixels, instead of evaluating the geometry of every pixel again. Building a mapping is that
// evaluation, and on a 16 Mpx detector it is a tenth of a second on every core. One entry; a mapping
// being built holds it, so two built at once for the same geometry evaluate it once.
class AzimuthalIntegrationGeometryCache {
friend class AzimuthalIntegrationMapping;
std::mutex m;
std::vector<uint32_t> key; // bit patterns of everything SetupPixel reads
std::vector<uint16_t> pixel_to_bin;
std::vector<float> pixel_resolution;
std::vector<float> corrections;
};
class AzimuthalIntegrationMapping {
protected:
const AzimuthalIntegrationSettings settings;
const float wavelength;
const size_t width, height;
std::vector<float> bin_to_q;
std::vector<float> bin_to_2theta;
std::vector<float> bin_to_d;
std::vector<float> bin_to_phi;
std::vector<uint16_t> pixel_to_bin;
std::vector<float> pixel_resolution;
std::vector<float> corrections;
std::optional<float> polarization_factor;
// Checksums of the two tables the GPU engines upload, taken once here. They are part of the key
// the shared device-table cache looks them up by (CudaSharedTables.h), so an engine that hands
// its own checksum in does not have to hash tens of megabytes on its way to a cache hit - and
// there is one engine per worker per pass. Both vectors are written in the constructor and never
// touched again, so a checksum taken there stays true for the mapping's whole life.
uint64_t pixel_to_bin_checksum = 0;
uint64_t corrections_checksum = 0;
// Bit-packed "this pixel is outside the resolution limits" mask, memoised for the limits it was
// last built for. Every worker's spot finder wants the identical mask and building it walks every
// pixel, so it is built once and shared. The mapping is const and read from all the worker threads
// at once, hence the mutex; and the mask is handed out as a shared_ptr, so a caller keeps the one
// it was given even if another thread later replaces the cached one.
mutable std::mutex res_mask_mutex;
mutable std::optional<float> res_mask_high, res_mask_low;
mutable std::shared_ptr<const std::vector<uint32_t>> res_mask_bits;
size_t nthreads;
void UpdateMaxBinNumber();
// A null mask sets up every pixel.
void SetupRawGeom(const DiffractionExperiment& experiment, const std::vector<uint32_t> &mask);
void SetupConvGeomRows(const DiffractionGeometry &geom, const std::vector<uint32_t> *mask, size_t row0, size_t row_end);
void SetupConvGeom(const DiffractionGeometry &geom, const std::vector<uint32_t> *mask);
void SetupConvGeomCached(const DiffractionGeometry &geom, const std::vector<uint32_t> &mask,
AzimuthalIntegrationGeometryCache &cache);
void SetupPixel(const DiffractionGeometry &geom, const std::vector<uint32_t> *mask,
uint32_t pxl, uint32_t col, uint32_t row);
AzimuthalIntegrationMapping(const DiffractionExperiment& experiment,
const PixelMask& mask,
AzimuthalIntegrationGeometryCache *cache,
size_t nthreads);
public:
AzimuthalIntegrationMapping(const DiffractionExperiment& experiment,
const PixelMask& mask,
size_t nthreads = 0);
// The same mapping, with the geometry of the pixels taken from `cache` when it holds this geometry
// (AzimuthalIntegrationGeometryCache), and left there for the next one when it does not.
AzimuthalIntegrationMapping(const DiffractionExperiment& experiment,
const PixelMask& mask,
AzimuthalIntegrationGeometryCache &cache,
size_t nthreads = 0);
[[nodiscard]] uint16_t GetBinNumber() const;
[[nodiscard]] const std::vector<uint16_t>& GetPixelToBin() const;
[[nodiscard]] const std::vector<float> &GetBinToQ() const;
[[nodiscard]] const std::vector<float> &GetBinToD() const;
[[nodiscard]] const std::vector<float> &GetBinToTwoTheta() const;
[[nodiscard]] const std::vector<float> &GetBinToPhi() const;
[[nodiscard]] uint16_t QToBin(float q) const;
[[nodiscard]] const std::vector<float> &Corrections() const;
[[nodiscard]] const std::vector<float> &Resolution() const;
[[nodiscard]] uint64_t GetPixelToBinChecksum() const;
[[nodiscard]] uint64_t GetCorrectionsChecksum() const;
// Pixels the spot finders must ignore because their resolution falls outside the limits, packed
// 32 pixels to a word (bit set = ignore), in the finders' own layout.
[[nodiscard]] std::shared_ptr<const std::vector<uint32_t>>
ResolutionMaskBits(std::optional<float> high_res, std::optional<float> low_res) const;
[[nodiscard]] size_t GetWidth() const;
[[nodiscard]] size_t GetHeight() const;
[[nodiscard]] const AzimuthalIntegrationSettings& Settings() const;
[[nodiscard]] int32_t GetAzimuthalBinCount() const;
[[nodiscard]] int32_t GetQBinCount() const;
[[nodiscard]] size_t GetNThreads() const;
};