Files
Jungfraujoch/image_analysis/bragg_integration/BraggIntegrationEngineCPU.h
T
leonarski_fandClaude Opus 5.5 67ec3649a4 CPU Bragg integration: keep the full-frame scratch between images
RunImpl allocated the reflection mask (1 byte/pixel) and the owner map (4
bytes/pixel) afresh for every image - 90 MB at 18 Mpixel, above the malloc
mmap threshold, so each call paid an mmap, a zero-fill page fault per 4 kB and
an munmap (with its TLB shootdown across the other workers). On a CPU-only
rotation run of a 16M set that was 79 M page faults and 794 s of system time.

Both are now members of the engine, and each call clears the rectangles the
previous one wrote before it starts. Contents at every read are unchanged, so
the output is byte-identical; CPU-only wall 297 -> 262 s, system time
794 -> 214 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 17:23:12 +02:00

47 lines
2.5 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include "BraggIntegrationEngine.h"
class CompressedImage;
// Plain-C++ reference/fallback engine: a faithful serial re-expression of BraggIntegrate2D (box
// sum) and ProfileIntegrate2D (Kabsch profile fit) reading the preprocessed int32 image. Also the
// numeric oracle the CUDA engine is checked against.
class BraggIntegrationEngineCPU : public BraggIntegrationEngine {
// Core integrator, templated on a pixel sampler so it reads either the preprocessed int32 buffer
// or a raw CompressedImage of any pixel type - both presented per-pixel in the INT32_MIN(masked)/
// INT32_MAX(saturated) convention - without ever materialising a second full-image copy.
// Full-frame scratch of RunImpl: the reflection mask and the signal-region owner map. Kept from
// call to call rather than allocated per image - at 18 Mpixel that was 90 MB, a fresh mmap and a
// page fault per 4 kB every call - and put back to their empty state one rectangle at a time: the
// rectangles a call writes are recorded, and the next call clears exactly those before it starts.
struct Rect { int x0, x1, y0, y1; };
std::vector<uint8_t> refl_mask;
std::vector<uint32_t> owner;
std::vector<Rect> refl_mask_dirty;
std::vector<Rect> owner_dirty;
template <class Sampler>
std::vector<Reflection> RunImpl(const Sampler &img, const std::vector<Reflection> &predicted,
size_t npredicted, int64_t image_number);
public:
explicit BraggIntegrationEngineCPU(const DiffractionExperiment &experiment);
using BraggIntegrationEngine::Run; // keep the preprocessed-buffer overload visible
std::vector<Reflection> Run(const ImagePreprocessorBuffer &image,
const std::vector<Reflection> &predicted, size_t npredicted,
int64_t image_number) override;
// FPGA workflow: integrate straight off the assembled detector image, reading only the pixels
// inside each reflection disk (no whole-image conversion - the FPGA host cannot afford one at its
// frame rate). Masked pixels carry the type minimum and saturated the type maximum.
std::vector<Reflection> Run(const CompressedImage &image,
const std::vector<Reflection> &predicted, size_t npredicted,
int64_t image_number);
};