Pre-scan: parallel beam-stop mask, leaner background beam-centre fit

Exact: p.mtz and the pre-scan products (shadow mask, mean projection,
defective-pixel mask, ring and capture centres, compared as hashes and
hex floats) are bit-identical to rc174 on three in-house rotation sets,
GPU and CPU builds.

- ShadowFinder::GetMask: the serial parts run in parallel - connected
  components by row band joined with union-find (both the shadow and
  the transmitting-arm searches, and the hole fill), ring binning and
  the harmonic sector gather by blocks, gap bridging by line; ring pixel
  counts read off the ring offsets. Mean projection filled in parallel.
- ShadowFinder host accumulation: one band-locked projection instead of
  a 20 B/px shard per pre-scan worker (2.7 GB zeroed and folded on a
  16M detector); SetShardCount and the shard argument are gone.
- FindBeamCenterFromBackground: the usable-pixel test is made once, the
  in-band pixels are kept in pixel order so the clipping rounds no
  longer sweep the whole detector, the 67 MB cell map is gone and the
  per-iteration block fold runs in parallel - same sums, same order.
- HotPixelFinder::GetMask: the chance-rate counts in parallel (integers).

Measured on a loaded box (load ~25 from other jobs), pre-scan window:
GPU 5.9-6.5 s -> 3.2-3.4 s, CPU 8.4-9.0 s -> 6.1-7.4 s. The GPU-build
pre-scan now ends with its background spot measurement (CPU spot finder
on ~120 frames, ~13 core-s on 8 workers).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
This commit is contained in:
2026-10-03 10:46:49 +02:00
co-authored by Claude Opus 5.5
parent 2c97e654ed
commit 24e36ae740
6 changed files with 427 additions and 326 deletions
+18 -4
View File
@@ -281,11 +281,25 @@ HotPixelFinder::Result HotPixelFinder::GetMask(double oscillation_deg, double sp
// The chance rate per ring, from the pixels lit on no more than half of their frames: whatever
// lights those - reflections, zingers, noise above the bound - lights a defect-free pixel too.
// Counted in integers by blocks of rows in parallel, so the totals do not depend on the split.
std::vector<std::vector<int64_t>> block_lit(BANDS), block_seen(BANDS);
const size_t rows_per_band = (height + BANDS - 1) / BANDS;
ParallelFor(static_cast<int>(BANDS), nthreads, [&](int b) {
block_lit[b].assign(nrings, 0);
block_seen[b].assign(nrings, 0);
const size_t begin = std::min(width * height, b * rows_per_band * width);
const size_t end = std::min(width * height, (b + 1) * rows_per_band * width);
for (size_t i = begin; i < end; i++)
if (key[i] >= 0 && n_valid(i) > 0 && 2 * n_lit[i] <= n_valid(i)) {
block_lit[b][key[i] / SECTORS] += n_lit[i];
block_seen[b][key[i] / SECTORS] += n_valid(i);
}
});
std::vector<double> lit(nrings, 0.0), seen(nrings, 0.0);
for (size_t i = 0; i < width * height; i++)
if (key[i] >= 0 && n_valid(i) > 0 && 2 * n_lit[i] <= n_valid(i)) {
lit[key[i] / SECTORS] += n_lit[i];
seen[key[i] / SECTORS] += n_valid(i);
for (int r = 0; r < nrings; r++)
for (size_t b = 0; b < BANDS; b++) {
lit[r] += static_cast<double>(block_lit[b][r]);
seen[r] += static_cast<double>(block_seen[b][r]);
}
std::vector<int> k_chance(nrings, n + 1);
for (int r = 0; r < nrings; r++)