The profile is the MEAN of each bin, so a few strong reflections landing in a bin lift it exactly as a smooth powder ring does. That is the wrong quantity whenever the profile is wanted as a background rather than as a measurement of what is in the bin - the ice score being the case in point, where reading a plain profile INVERTED the metric: over 37 rotation crystals the two highest-scoring crystals had no ice at all. The adaptive spot finder already computes the right thing, a sigma-clipped per-resolution-ring background, as a byproduct of its own threshold. Where it runs, the ice score uses that. Where it does not - --no-adaptive-spots, --azint-only, and anything reading the profile the broker wrote - there was no way to get it. This adds one: azim_int_settings.sigma_clip (rugnux --azim-sigma-clip), 0 = off, minimum 2 because a tighter clip rejects a large part of a clean Gaussian bin and biases the estimate low rather than removing outliers. Two clip passes follow the plain one, matching the finder's recipe - the first pass's standard deviation is itself inflated by the peaks being removed, so one pass leaves a threshold that is still too generous. A bin with fewer than eight pixels is left alone: at the detector edge and behind the beam stop there is no spread to clip on. Both engines do it. On the GPU the accept range is computed by a small kernel and stays resident, so a clip pass is one more read of the same pixels and no round trip; the two accumulation kernels take the range as a pointer that is null on the plain pass. Measured on a JUNGFRAU rotation dataset, non-adaptive path: azimuthal integration 0.02 -> 0.06 ms per image, exactly the 3x the extra passes predict, against a 0.34 ms per-image total. Note what the result IS: the smooth background under the peaks, not the bin mean. It should not be switched on where a ring's integrated intensity is wanted - the powder-ring geometry fit reads ring peaks, and those are what a clip is designed to remove. Off by default, so nothing changes unless it is asked for. Not exposed over the REST API - that needs the generated model regenerated, which is a separate step. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
35 lines
1.3 KiB
C++
35 lines
1.3 KiB
C++
// SPDX-FileCopyrightText: 2025 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include "AzIntEngine.h"
|
|
#include "../indexing/CUDAMemHelpers.h"
|
|
#include "../indexing/CudaSharedTables.h"
|
|
|
|
class AzIntEngineGPU : public AzIntEngine {
|
|
std::shared_ptr<CudaStream> stream;
|
|
int threads;
|
|
int blocks;
|
|
size_t shared_needed;
|
|
size_t shared_size;
|
|
|
|
// Geometry-only tables: one copy per GPU, shared with every other engine on it (see
|
|
// CudaSharedTables.h) instead of one copy per worker thread.
|
|
std::shared_ptr<CudaDevicePtr<float>> gpu_azint_correction;
|
|
std::shared_ptr<CudaDevicePtr<uint16_t>> gpu_pixel_to_bin;
|
|
|
|
CudaDevicePtr<float> gpu_sum;
|
|
CudaDevicePtr<float> gpu_sum2;
|
|
CudaDevicePtr<uint32_t> gpu_count;
|
|
// Per-bin accept range for a sigma-clip pass; empty when clipping is off.
|
|
CudaDevicePtr<float> gpu_clip_lo;
|
|
CudaDevicePtr<float> gpu_clip_hi;
|
|
CudaRegisteredVector<float> cpu_sum_reg;
|
|
CudaRegisteredVector<float> cpu_sum2_reg;
|
|
CudaRegisteredVector<uint32_t> cpu_count_reg;
|
|
public:
|
|
AzIntEngineGPU(const AzimuthalIntegrationMapping& integration, std::shared_ptr<CudaStream> stream);
|
|
void Run(const ImagePreprocessorBuffer &image, AzimuthalIntegrationProfile &profile) override;
|
|
};
|