Build Packages / Create release (push) Successful in 17s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m22s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m37s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 9m33s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 10m39s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 11m4s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 13m19s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 17m37s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 18m49s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 19m10s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 24m26s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m31s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 18m54s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 20m45s
Build Packages / Generate python client (push) Successful in 37s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 20m20s
Build Packages / Build documentation (push) Successful in 1m32s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 14m37s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 21m6s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 19m49s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 20m29s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 17m2s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 14m27s
Build Packages / Unit tests (push) Successful in 1h18m12s
* Rugnux: Performance improvements on GPU and CPU (more of the pre-scan and of scaling on the GPU, faster CPU spot finding and crystal refinement), with unchanged results. * Rugnux: More robust processing - patches of persistently hot pixels are masked, an inconsistent merge triggers a retry at the measured beam centre, and builds targeting different CPU levels give the same results. * Rugnux: Improved scaling and merging - reflections with an overloaded pixel are dropped, as in XDS, sparse rotation sweeps are scaled more reliably, and French-Wilson amplitudes use an anisotropic Wilson prior. * Rugnux: Improved space-group determination - glide planes in groups without a centre of symmetry, screw axes from short or weak axial rows kept when a higher group is adopted, and more reliable decisions on twinned and pseudo-symmetric crystals. * Rugnux: Improved small-molecule processing - spots that grow wider than the integration disk and split spots are integrated over their measured footprint, sparse lattices are integrated on every frame, and the `.hkl` file holds unmerged scaled reflections (SHELX HKLF 4). * Rugnux: Reads Rigaku d*TREK SMV images (Saturn CCD), including detector 2theta and encoded pixel overflows; home-source (rotating-anode) datasets were added to the validation battery. * jfjoch_viewer: Fixed processing failing at the end with "Wrong JPEG library version" on Linux; the merge window shows the space group with proper subscripts and a checklist of crystal pathologies. Reviewed-on: #84 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
80 lines
3.9 KiB
C++
80 lines
3.9 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include "BraggIntegrationEngine.h"
|
|
#include "../../common/PixelMask.h"
|
|
|
|
class CompressedImage;
|
|
|
|
// Plain-C++ reference/fallback engine: a faithful serial re-expression of BraggIntegrate2D (box
|
|
// sum) and ProfileIntegrate2D (Kabsch profile fit) reading the preprocessed int32 image. Also the
|
|
// numeric oracle the CUDA engine is checked against.
|
|
class BraggIntegrationEngineCPU : public BraggIntegrationEngine {
|
|
// Core integrator, templated on a pixel sampler so it reads either the preprocessed int32 buffer
|
|
// or a raw CompressedImage of any pixel type - both presented per-pixel in the INT32_MIN(masked)/
|
|
// INT32_MAX(saturated) convention - without ever materialising a second full-image copy.
|
|
// Full-frame scratch of RunImpl: the reflection mask and the signal-region owner map. A call writes
|
|
// only around its reflections, so these keep only the 16x16-pixel tiles written since the last
|
|
// Clear(): a frame-sized array was mostly never read, yet over a sweep every page of it got
|
|
// touched, in every worker. Reading a tile nothing wrote gives `empty`.
|
|
template <class T>
|
|
class TiledFrame {
|
|
static constexpr int TILE = 16;
|
|
int tiles_x;
|
|
T empty;
|
|
std::vector<int32_t> tile_start; // per tile: where it starts in `pixels`, -1 = not written
|
|
std::vector<int32_t> written; // the tiles written, for Clear()
|
|
std::vector<T> pixels;
|
|
public:
|
|
TiledFrame(int width, int height, T empty)
|
|
: tiles_x((width + TILE - 1) / TILE), empty(empty),
|
|
tile_start(static_cast<size_t>(tiles_x) * ((height + TILE - 1) / TILE), -1) {}
|
|
T Get(int x, int y) const {
|
|
const int32_t start = tile_start[(y / TILE) * tiles_x + x / TILE];
|
|
return start < 0 ? empty : pixels[start + (y % TILE) * TILE + x % TILE];
|
|
}
|
|
T &At(int x, int y) {
|
|
const int t = (y / TILE) * tiles_x + x / TILE;
|
|
if (tile_start[t] < 0) {
|
|
tile_start[t] = static_cast<int32_t>(pixels.size());
|
|
pixels.resize(pixels.size() + TILE * TILE, empty);
|
|
written.push_back(t);
|
|
}
|
|
return pixels[tile_start[t] + (y % TILE) * TILE + x % TILE];
|
|
}
|
|
void Clear() {
|
|
for (int t : written)
|
|
tile_start[t] = -1;
|
|
written.clear();
|
|
pixels.clear();
|
|
}
|
|
};
|
|
TiledFrame<uint8_t> refl_mask;
|
|
TiledFrame<uint32_t> owner;
|
|
// The run's pixel mask, packed 32 pixels to a word (PixelMask::GetPackedMask). An unreadable pixel
|
|
// it does not explain was unreadable on this frame only - an overload (see Reflection::overloaded).
|
|
std::vector<uint32_t> static_mask;
|
|
|
|
template <class Sampler>
|
|
std::vector<Reflection> RunImpl(const Sampler &img, const std::vector<Reflection> &predicted,
|
|
size_t npredicted, int64_t image_number);
|
|
|
|
public:
|
|
BraggIntegrationEngineCPU(const DiffractionExperiment &experiment, const PixelMask &mask);
|
|
|
|
using BraggIntegrationEngine::Run; // keep the preprocessed-buffer overload visible
|
|
|
|
std::vector<Reflection> Run(const ImagePreprocessorBuffer &image,
|
|
const std::vector<Reflection> &predicted, size_t npredicted,
|
|
int64_t image_number) override;
|
|
|
|
// FPGA workflow: integrate straight off the assembled detector image, reading only the pixels
|
|
// inside each reflection disk (no whole-image conversion - the FPGA host cannot afford one at its
|
|
// frame rate). Masked pixels carry the type minimum and saturated the type maximum.
|
|
std::vector<Reflection> Run(const CompressedImage &image,
|
|
const std::vector<Reflection> &predicted, size_t npredicted,
|
|
int64_t image_number);
|
|
};
|