Files
Jungfraujoch/image_analysis/spot_finding/AdaptiveSpotFinderCPU.h
T
leonarski_fandClaude Opus 5 9f2696c1ac
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 8m41s
Build Packages / build:windows:nocuda (push) Successful in 16m50s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 18m22s
Build Packages / build:windows:cuda (push) Successful in 19m40s
Build Packages / build:viewer-tgz:cpu (push) Successful in 21m2s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m43s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m21s
Build Packages / build:rugnux:windows (push) Successful in 10m45s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 27m37s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 27m50s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 19m52s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 22m6s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 25m47s
Build Packages / build:rpm (rocky9) (push) Successful in 23m52s
Build Packages / build:rpm (rocky8) (push) Successful in 26m33s
Build Packages / Generate python client (push) Successful in 45s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2404) (push) Successful in 23m57s
Build Packages / DIALS test (push) Successful in 25m7s
Build Packages / XDS test (durin plugin) (push) Successful in 11m18s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 27m24s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m58s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m43s
Build Packages / Unit tests (push) Successful in 1h22m26s
Adaptive spot finding: the ring threshold replaces the floor, not the local test
The self-calibrating finder was meant to replace the classic finder's FIXED
PHOTON FLOOR with a per-resolution-ring threshold read off the image's own
noise. As written it replaced the local-box SNR test as well, and that is the
defect: a whole-ring threshold is an ABSOLUTE contour with no feedback from a
pixel's own surroundings, so the area a spot puts above it grows as
sigma^2*ln(peak/threshold) and never saturates.

Measured on a strongly diffracting rotation set, the detected footprint grows by
+8.05 pixels per e-fold of peak, so the brightest reflections came out as
100-500 pixel blobs and were then discarded for exceeding the size bound - every
one of the ten strongest on an image. Intersecting with the local box gives
-0.24 pixels per e-fold, the classic finder's own number to two decimals.

WHY the local box is the right partner, rather than merely the incumbent: it is
a prominence rule whose reference level is a 961-pixel mean. A spot inflates the
box's own variance and the peak divides out of the acceptance test, so it cuts
at a fixed FRACTION of the spot's own height. Referring that level to fewer
pixels makes it inherit their shot noise - at FIXED footprint, estimating the
level from 961 pixels, from 25, and from the single maximum gives centroid
residuals of 0.524, 0.539 and 0.656 - so flat growth and a stable centroid turn
out to be two ends of one dial. A contour on the bare maximum has the flattest
growth of anything tried (+0.1) and merges worst.

The two arms bind in different regimes, which is why intersecting beats choosing:
on serial stills the ring threshold is 0.6x the classic floor, on this rotation
sweep 2.3-6.0x. Stills are a strict no-op - 175 components against 175,
identical per frame - so the +40% in stills indexing that the adaptive threshold
was introduced for is untouched.

What it buys, stated as one fact rather than two. Across five geometry pins
spanning 1.1 mm it indexes the most frames of any arm tried, 0.831 against
0.803, and integrates 3.04 to 5.76% more observations - but those are the SAME
number: regressing observation count on indexing rate over four arms leaves
residuals of +/-0.7 percentage points against swings of -7 to +4.5%, so the
extra observations ARE the extra indexed frames, not better data per frame.
CC1/2, the only statistic here carrying per-observation quality, is +0.66 at one
pin and -0.06 at the other: not harmed, not improved. <I/sigma>, ISa and R_meas
cannot arbitrate on this data - across those pins each crosses zero as a
monotone function of the pin.

WHY an absolute contour indexes fewer frames, when its spot list is equal or
better on every axis measured - recall, top-1000 recall, centroid, ice fraction,
component count - is the interesting part, and it is not a detection effect at
all: ITS OWN SIZE BOUND DELETES THE BRIGHTEST REFLECTIONS ON THE FRAME. A
component is discarded because it grew past 200 px, and it grew past 200 px
because it was bright, so the deletions are drawn from the head of the indexing
budget rather than uniformly from it: they are 11x enriched in the top 250 of
the thousand spots handed to the indexer, and the bound's own real deletions sit
at MEDIAN RANK 12. Turning the bound off recovers 66% and 50% of the deficit at
the two pins, against a bar registered at 33% before the run.

Three of us dismissed this for most of a day on the grounds that the gates
delete only ~4% of what is detected. That arithmetic was right and the
denominator was wrong - a rate is not an impact when the thing being lost is
selected for the property that makes it matter.

Reworking the bound instead was measured and rejected: it recovers half the
deficit, and it cannot be done without re-admitting what the bound is for - 68
components past 200 px, of which 8 are real and 60 are junk, where the intersect
gets the 8 without the 60. The residual once the bound is off, +1.08%/+1.70%,
is the contour itself.

Component merging is ruled out separately: geometrically impossible here, 33.9
px minimum reflection separation against components spanning 10 px. So is a
ranking effect - the intersect's lead runs +0.06% at --max-spots 250, +3.46% at
1000 and +14.26% at 2000, which is backwards for a selection artefact.

Costs 0.48 ms per image in the finder, and 0.044 px of bright-spot centroid
precision - measured convention-free, by fitting a line to a reflection's own
centroid across five frames, after an XDS-referenced figure proved to be four
fifths aperture convention.

It also makes the compactness gate above it safe. On the absolute contour that
gate is net damage, deleting 37 genuine reflections per ten frames; once the
footprint stops growing nothing reaches its threshold at all.

Also fixes a real but unexercised defect in PoissonThreshold, where the exact
tail handed over to a normal approximation with a step. It changes nothing here:
the clipped ring sigma is over-dispersed 1.2-4.9x against sqrt(mu) because it
still contains diffraction, so the Gaussian arm wins every ring above mu=50 and
none of the 522 thresholds move.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FBumeJVx4oeXxiBRpkrE5H
2026-08-28 15:30:39 +02:00

76 lines
4.4 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <vector>
#include "ImageSpotFinderCPU.h"
#include "SpotFindingSettings.h"
#include "../../common/AzimuthalIntegrationMapping.h"
// Self-calibrating strong-pixel detector for the offline (rugnux/viewer) path.
//
// The classic finder (ImageSpotFinderCPU) marks a pixel strong when it clears a *fixed* photon
// count AND a local-box SNR. The fixed photon floor is what forces per-dataset tuning: it must sit
// above the background (wants high) yet not bury weak spots (wants low), and the background level
// differs per dataset, so the sweet spot is narrow (~12 photons on one serial-stills set, ~5 on a
// weaker one).
//
// Here the floor is replaced by a per-resolution-ring threshold derived from a single portable
// number: E = the expected count of noise pixels tolerated per frame (default ~100). For a ring
// whose (peak-excluded) background mean is mu, the threshold is the smallest count whose Poisson
// upper tail is <= p = E / N_pixels, max'd with a Gaussian arm mu + z*sigma to absorb read/flat-field
// excess. Because it is set from the image's own noise, the SAME E lands at ~12 photons on the first
// set and ~5 on the weaker one with no user input.
//
// The ring threshold replaces the floor and ONLY the floor: the classic finder's local-box SNR test
// still has to pass, which is why this engine runs it (ImageSpotFinderCPU) and intersects the two
// masks.
//
// A whole-ring threshold is an ABSOLUTE contour with no feedback from the pixel's own surroundings,
// so the area a spot puts above it grows as sigma^2 * ln(peak/threshold) and never saturates: on a
// strongly diffracting rotation set the detected footprint grows by 8 pixels per e-fold of peak, so
// the brightest reflections came out as 100-500 pixel blobs. The local box has no such contour. The
// spot inflates the box's own variance, and the peak divides out of the acceptance test, so the box
// cuts every spot at roughly a fixed FRACTION of its own height - a peak-relative contour. Measured
// on the same set, the footprint then grows by -0.2 pixels per e-fold, i.e. not at all, and lands on
// the classic finder's own number to two decimals.
//
// The two arms bind in different regimes, which is the point of intersecting rather than choosing.
// On serial stills the ring background is a fraction of a count and the ring threshold lands BELOW
// the fixed floor the classic finder would use, so the ring arm decides and the local box passes
// everything - which is the whole reason this engine exists. On a bright rotation set the ring
// background is tens of counts, the ring threshold lands several times ABOVE that floor, and the
// local box decides. The engine is therefore never worse than the classic finder on footprint, and
// never worse than a fixed floor on a weak background.
class AdaptiveSpotFinderCPU : public ImageSpotFinderCPU {
const AzimuthalIntegrationMapping &mapping;
// per-ring scratch, sized to the mapping's bin count
// Exact integers: a preprocessed pixel is an int32 and the sentinels are skipped, so v and v*v
// are exact in 64 bits. That is what lets the GPU engine reproduce these bit for bit - integer
// addition is associative, so its block atomics can arrive in any order.
std::vector<int64_t> ring_sum;
std::vector<uint64_t> ring_sum2;
std::vector<int64_t> ring_cnt;
std::vector<float> ring_mean;
std::vector<float> ring_sigma;
std::vector<float> ring_thr;
// ring_mean of the last Detect(), NaN where the ring holds too few pixels to be its own background.
// Kept separately because ring_mean carries the previous frame's value for an empty ring.
std::vector<float> ring_bkg;
// Pixels at or above their ring's threshold, packed like output_buffer. Intersected with the
// local-box mask that ImageSpotFinderCPU::Detect leaves in output_buffer.
std::vector<uint32_t> ring_bits;
void AccumulateRings(const ImagePreprocessorBuffer &image, float clip_k);
// Fill ring_bits from the thresholds of the current frame.
void FlagRings(const ImagePreprocessorBuffer &image);
public:
explicit AdaptiveSpotFinderCPU(const AzimuthalIntegrationMapping &mapping);
void Detect(const ImagePreprocessorBuffer &image, const SpotFindingSettings &settings) override;
[[nodiscard]] const std::vector<float> &GetRingBackground() const override { return ring_bkg; }
};