Every lattice decision counts spots or frames, and the sub-lattice's strong reflections win every count, so a crystal whose cell is doubled by a weak superstructure class (9min, 6z9g) is adopted at the half cell. This asks the question in intensities instead. On 60 frames spread over the sweep, after each frame's own integration, the 2a x 2b x 2c supercell of its primitive lattice is predicted to 3 A with the frame's own refined orientation and geometry and integrated on the same engine; the reflections are summed per parity class in two shells (20-5, 5-3 A), with a fit of intensity against partiality (I = a + b p) that separates what rocks like a Bragg reflection from what sits at the node whatever the rocking. Only the tested frames are predicted and nothing is retained, so memory is bounded (the previous prototype predicted the whole run through the merge and ran out of GPU memory). The probe's integrations are kept out of the engine's own counts, which the two-pass stencil guard reads. Results are bit-identical with and without it. REPORT ONLY - it decides nothing, because on the battery it does not yet separate a weak real class from what sits at the half-integer nodes of crystals whose cell is right. Real classes: 6z9g class 101 at 24 % of the lattice's intensity (29 % rocking), 9min 100 at 19 % (4 % rocking - its real class does not rock like the lattice either). On correct cells the largest classes reach 9-12 % raw (7n2s, 9i0a, 7os3) and 3-4 % rocking (7dkp, 7os3), and 7mzt reads 40 % / 19 % on a class the deposition does not have. The log line is the population a decision has to be calibrated on. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
106 lines
5.4 KiB
C++
106 lines
5.4 KiB
C++
// SPDX-FileCopyrightText: 2024 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <mutex>
|
|
|
|
#include "../common/JFJochMessages.h"
|
|
#include "../common/DiffractionExperiment.h"
|
|
#include "../common/AzimuthalIntegrationMapping.h"
|
|
#include "../common/PixelMask.h"
|
|
#include "../common/AzimuthalIntegrationProfile.h"
|
|
#include "bragg_prediction/BraggPrediction.h"
|
|
#include "bragg_integration/BraggIntegrationEngine.h"
|
|
#include "spot_finding/ImageSpotFinder.h"
|
|
#include "spot_finding/AdaptiveSpotFinderCPU.h"
|
|
#include "indexing/IndexerThreadPool.h"
|
|
#include "azint/AzIntEngine.h"
|
|
#include "roi/ROIIntegration.h"
|
|
#include "IndexAndRefine.h"
|
|
#include "image_preprocessing/ImagePreprocessor.h"
|
|
#include "image_preprocessing/ImagePreprocessorBuffer.h"
|
|
|
|
class CudaStream;
|
|
class AdaptiveSpotFinderGPU;
|
|
|
|
// MXAnalysisWithoutFPGA is not thread safe - it has to owned by a single thread
|
|
class MXAnalysisWithoutFPGA {
|
|
const DiffractionExperiment &experiment;
|
|
const AzimuthalIntegrationMapping &integration;
|
|
|
|
std::vector<uint8_t> decompression_buffer;
|
|
|
|
std::unique_ptr<ImagePreprocessor> preprocessor;
|
|
|
|
size_t npixels;
|
|
size_t xpixels;
|
|
|
|
// Built on first use: the fused adaptive finder produces the azimuthal profile as a by-product,
|
|
// so on the rugnux path this engine is constructed and then never run.
|
|
std::unique_ptr<AzIntEngine> azint;
|
|
AzIntEngine &AzInt();
|
|
std::unique_ptr<ROIIntegration> roi;
|
|
// Built on first use. Which finder an image takes arrives with its SpotFindingSettings, and
|
|
// with adaptive detection on - the default everywhere but the broker - this one is never asked
|
|
// for; on the GPU it is ~14 MB and 15 device allocations per worker.
|
|
std::unique_ptr<ImageSpotFinder> spotFinder;
|
|
ImageSpotFinder &FixedThresholdFinder();
|
|
// Self-calibrating finder, used when spot settings request adaptive detection. Kept alongside the
|
|
// default finder because the choice arrives with the per-image settings, not at construction. It is
|
|
// an AdaptiveSpotFinderCPU by default; on the GPU path, when the fused engine is enabled (rugnux
|
|
// offline only), it is instead an AdaptiveSpotFinderGPU that also computes the azimuthal profile,
|
|
// aliased through fused_adaptive so Analyze() can take that profile and skip the separate azint pass.
|
|
std::unique_ptr<ImageSpotFinder> adaptiveSpotFinder;
|
|
AdaptiveSpotFinderGPU *fused_adaptive = nullptr;
|
|
const bool enable_fused_adaptive_gpu;
|
|
IndexAndRefine &indexer;
|
|
std::unique_ptr<BraggPrediction> prediction;
|
|
std::unique_ptr<BraggIntegrationEngine> bragg_engine;
|
|
// What the supercell probe's integrations added to bragg_engine's counts (see BraggCounts).
|
|
BraggIntegrationCounts probe_counts;
|
|
std::unique_ptr<ImagePreprocessorBuffer> preprocessor_buffer;
|
|
const PixelMask &mask;
|
|
|
|
// Decompress the image into decompression_buffer (or read it straight from the message, when it is
|
|
// not compressed) and return where it landed.
|
|
const uint8_t *Decompress(const CompressedImage &image);
|
|
|
|
// Pixels outside the resolution limits, bit-packed. Built by the integration mapping, which is
|
|
// shared by every worker's engine and hands out the same mask to all of them.
|
|
std::shared_ptr<const std::vector<uint32_t>> mask_resolution;
|
|
// The limits mask_resolution was built for. Kept as the OPTIONAL the caller passed, so an unset
|
|
// high-resolution limit compares equal to itself and the mask is not rebuilt on every image.
|
|
std::optional<float> mask_high_res;
|
|
std::optional<float> mask_low_res;
|
|
void UpdateMaskResolution(const SpotFindingSettings& settings);
|
|
#ifdef JFJOCH_USE_CUDA
|
|
std::shared_ptr<CudaStream> stream; // kept so RebuildROI() can recreate the GPU ROI engine
|
|
#endif
|
|
public:
|
|
// enable_fused_adaptive_gpu turns on the fused GPU azint+adaptive spot finder (only takes effect on
|
|
// the GPU path with adaptive detection). The rugnux offline path and the interactive viewer enable
|
|
// it by default, as does the online receiver. It only changes performance - the fused engine
|
|
// reproduces the CPU finder's spots. Note it also decides whether the preprocessed image is copied
|
|
// back to the host each frame: that copy exists only for a CPU engine to read, and with the flag on
|
|
// no CPU engine is built, so the copy is skipped.
|
|
MXAnalysisWithoutFPGA(const DiffractionExperiment &experiment, const AzimuthalIntegrationMapping &integration,
|
|
const PixelMask &mask, IndexAndRefine &indexer, bool enable_fused_adaptive_gpu = false);
|
|
void Analyze(DataMessage &output, AzimuthalIntegrationProfile &profile, const SpotFindingSettings &spot_finding_settings);
|
|
|
|
// Surgical ROI-only paths used when a full re-analysis is not wanted: rebuild the
|
|
// ROI engine after the ROI set changes, recompute ROIs after preprocessing a new
|
|
// image (reanalyze off), or just rerun ROIs on the current preprocessed image (an
|
|
// interactive ROI move). A full Analyze() already computes ROIs, so needs nothing.
|
|
void RebuildROI();
|
|
void AnalyzeROIOnly(DataMessage &output);
|
|
void RunROIOnly(DataMessage &output);
|
|
|
|
// What this worker's Bragg integrator counted (BraggIntegrationCounts). Each worker builds its own
|
|
// analysis, so a caller that wants the run's totals sums this over the workers it started.
|
|
[[nodiscard]] BraggIntegrationCounts BraggCounts() const;
|
|
};
|
|
|
|
|
|
|