Two whole-image passes per frame out of the CPU spot finder (the pre-scan's finder on every build, and every image on the CPU-only build): - AccumulateRings ran three passes over the frame - the plain ring statistics and two sigma clips. The plain pass now also counts each ring's valid values in a histogram (0..1023, the rest in a short list), and the clip passes sum over the distinct values: each meets the same float test its pixels would, and the sums are integers, so the totals are the same. - The local test's first pass is read by DetectAt only inside a candidate's window. It now marks the row/32-column blocks those windows reach and keeps its sliding sums everywhere but skips the per-pixel test elsewhere; the bits it leaves unset are never read. md5-identical output on four sets (GPU) and three (CPU-only). CPU-only 16M: 203 -> 177 s and 303 -> 263 s; GPU 16M 23.8 -> 22.8 s (the pre-scan's finder). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
59 lines
2.6 KiB
C++
59 lines
2.6 KiB
C++
// SPDX-FileCopyrightText: 2025 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
|
|
#include <cstddef>
|
|
#include <vector>
|
|
|
|
#include "ImageSpotFinder.h"
|
|
#include "SpotFindingSettings.h"
|
|
#include "../../common/AzimuthalIntegrationMapping.h"
|
|
#include "../../common/DiffractionSpot.h"
|
|
|
|
// This is "slow" spot finder for image-based analysis
|
|
// To complement "fast" PSI detector module based spot finder
|
|
// This one is expected to be used in cases, where images are already assembled
|
|
// and it aims for 100 ms execution
|
|
|
|
class ImageSpotFinderCPU : public ImageSpotFinder {
|
|
// Output of the first pass, and the exclusion mask for the second.
|
|
std::vector<uint32_t> first_pass_buffer;
|
|
|
|
// One detection pass. prev_strong is the previous pass's bitmap, or nullptr for the first pass;
|
|
// a pixel set in it is treated as INT32_MAX, i.e. kept out of every local background it would
|
|
// fall into, and reported strong.
|
|
//
|
|
// With candidates set (SNR test only), the window sums of every candidate pixel are also recorded
|
|
// in candidate_windows, in pixel order.
|
|
void DetectPass(const ImagePreprocessorBuffer &image, const SpotFindingSettings &settings,
|
|
const uint32_t *prev_strong, std::vector<uint32_t> &out,
|
|
const uint32_t *candidates = nullptr);
|
|
|
|
// Local-box sums (centre pixel included) of one candidate pixel, as the first pass saw them.
|
|
struct CandidateWindow {
|
|
int32_t pxl;
|
|
int64_t sum, sum2, valid;
|
|
};
|
|
std::vector<CandidateWindow> candidate_windows;
|
|
// Where the first pass's own result is read by DetectAt - within NBX of a candidate - per row and
|
|
// 32-column block. Elsewhere that pass keeps its sums but skips the test.
|
|
std::vector<uint8_t> first_pass_needed;
|
|
|
|
protected:
|
|
// Detect, but with the result wanted only at the candidate pixels (packed like output_buffer):
|
|
// output_buffer ends up holding (Detect's result & candidates). The second pass is not run over
|
|
// the image: a candidate's second-pass window is its first-pass window with the pixels the first
|
|
// pass found strong taken out, and those are few, so they are subtracted from the sums the first
|
|
// pass recorded. Integer sums and the same test, so the bits are exactly Detect's.
|
|
void DetectAt(const ImagePreprocessorBuffer &image, const SpotFindingSettings &settings,
|
|
const std::vector<uint32_t> &candidates);
|
|
|
|
public:
|
|
ImageSpotFinderCPU(int32_t width, int32_t height);
|
|
void Detect(const ImagePreprocessorBuffer &image, const SpotFindingSettings &settings) override;
|
|
};
|
|
|
|
|