Build Packages / Create release (push) Successful in 24s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m29s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m43s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 8m27s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 9m53s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 9m58s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 11m22s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 13m39s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 18m37s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 16m32s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 24m11s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m30s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 19m3s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 20m23s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 19m41s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 21m0s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m38s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 14m33s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 17m55s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 20m50s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 18m38s
Build Packages / Unit tests (push) Successful in 1h46m14s
* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports. * jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls. * Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results. * Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable. * Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate. * Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do. * Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence. * Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags. * Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check. * Rugnux: Clear error messages when a data set needs more GPU or host memory than is available. Reviewed-on: #83 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
76 lines
3.6 KiB
C++
76 lines
3.6 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include "BraggIntegrationEngine.h"
|
|
|
|
class CompressedImage;
|
|
|
|
// Plain-C++ reference/fallback engine: a faithful serial re-expression of BraggIntegrate2D (box
|
|
// sum) and ProfileIntegrate2D (Kabsch profile fit) reading the preprocessed int32 image. Also the
|
|
// numeric oracle the CUDA engine is checked against.
|
|
class BraggIntegrationEngineCPU : public BraggIntegrationEngine {
|
|
// Core integrator, templated on a pixel sampler so it reads either the preprocessed int32 buffer
|
|
// or a raw CompressedImage of any pixel type - both presented per-pixel in the INT32_MIN(masked)/
|
|
// INT32_MAX(saturated) convention - without ever materialising a second full-image copy.
|
|
// Full-frame scratch of RunImpl: the reflection mask and the signal-region owner map. A call writes
|
|
// only around its reflections, so these keep only the 16x16-pixel tiles written since the last
|
|
// Clear(): a frame-sized array was mostly never read, yet over a sweep every page of it got
|
|
// touched, in every worker. Reading a tile nothing wrote gives `empty`.
|
|
template <class T>
|
|
class TiledFrame {
|
|
static constexpr int TILE = 16;
|
|
int tiles_x;
|
|
T empty;
|
|
std::vector<int32_t> tile_start; // per tile: where it starts in `pixels`, -1 = not written
|
|
std::vector<int32_t> written; // the tiles written, for Clear()
|
|
std::vector<T> pixels;
|
|
public:
|
|
TiledFrame(int width, int height, T empty)
|
|
: tiles_x((width + TILE - 1) / TILE), empty(empty),
|
|
tile_start(static_cast<size_t>(tiles_x) * ((height + TILE - 1) / TILE), -1) {}
|
|
T Get(int x, int y) const {
|
|
const int32_t start = tile_start[(y / TILE) * tiles_x + x / TILE];
|
|
return start < 0 ? empty : pixels[start + (y % TILE) * TILE + x % TILE];
|
|
}
|
|
T &At(int x, int y) {
|
|
const int t = (y / TILE) * tiles_x + x / TILE;
|
|
if (tile_start[t] < 0) {
|
|
tile_start[t] = static_cast<int32_t>(pixels.size());
|
|
pixels.resize(pixels.size() + TILE * TILE, empty);
|
|
written.push_back(t);
|
|
}
|
|
return pixels[tile_start[t] + (y % TILE) * TILE + x % TILE];
|
|
}
|
|
void Clear() {
|
|
for (int t : written)
|
|
tile_start[t] = -1;
|
|
written.clear();
|
|
pixels.clear();
|
|
}
|
|
};
|
|
TiledFrame<uint8_t> refl_mask;
|
|
TiledFrame<uint32_t> owner;
|
|
|
|
template <class Sampler>
|
|
std::vector<Reflection> RunImpl(const Sampler &img, const std::vector<Reflection> &predicted,
|
|
size_t npredicted, int64_t image_number);
|
|
|
|
public:
|
|
explicit BraggIntegrationEngineCPU(const DiffractionExperiment &experiment);
|
|
|
|
using BraggIntegrationEngine::Run; // keep the preprocessed-buffer overload visible
|
|
|
|
std::vector<Reflection> Run(const ImagePreprocessorBuffer &image,
|
|
const std::vector<Reflection> &predicted, size_t npredicted,
|
|
int64_t image_number) override;
|
|
|
|
// FPGA workflow: integrate straight off the assembled detector image, reading only the pixels
|
|
// inside each reflection disk (no whole-image conversion - the FPGA host cannot afford one at its
|
|
// frame rate). Masked pixels carry the type minimum and saturated the type maximum.
|
|
std::vector<Reflection> Run(const CompressedImage &image,
|
|
const std::vector<Reflection> &predicted, size_t npredicted,
|
|
int64_t image_number);
|
|
};
|