Conflicts in the three CPU pixel loops the fused pipeline split per block, resolved by carrying this
branch's loop bodies into the new functions with the per-pixel arithmetic unchanged:
- AdaptiveSpotFinderCPU: the register-held ring sums now live in AccumulateRingsBlock (stored at the
end of each block, so the additions stay in pixel order).
- ImagePreprocessorCPU::AnalyzeBlock reads the PixelMask-derived 32-pixel mask words at first + i.
- ImageSpotFinderCPU::DetectPass: rc173's new_row() marking/fill_row calls kept in place, the
vertical update replaced by the vectorised slide with the prev_strong fix-up.
Byte-identical p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz vs rc173 references on myob, cytc, lyso,
sparse (CPU-only build) and myob (GPU build); targeted tests (ImageSpotFinderCPU*, AdaptiveSpotFinder,
SpotFinding, PixelMask, Bragg*, RotationScale, AzimuthalIntegration, portable) pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
AdaptiveSpotFinderCPU::AccumulateRings: consecutive pixels mostly share a ring, so the ring's
sum/sum2/count and the fused azint sums are held in locals while they do and stored when the ring
changes - the same additions in the same order (az_sum2 is still contracted to the same FMA), without
a store-and-reload chain through memory on every pixel.
ImageSpotFinderCPU::DetectPass: the vertical-sum update (add the entering row, take out the leaving
one) is one branch-free loop over the raw image that GCC vectorises (int64 lanes). A pixel strong in
the previous pass used to be substituted per pixel through a bit test, which kept the loop scalar;
it is now added with its row and taken out again from the few set bits of prev_strong. Integer sums,
so the same totals. (A first, fully branch-free version that kept the per-pixel bit test did not
vectorise on the prev_strong path and was measured slower; this is its replacement.)
ImagePreprocessorCPU: the per-engine std::vector<bool> built bit by bit from the 32-bit mask
(~10 core-s per cytc run, one per worker per pass) is replaced by 32-pixel mask words that PixelMask
derives once beside its binary mask; each engine copies 2 MB. A branch-free rewrite of the Analyze
loop was measured and dropped: the loop is bound by reading the decompressed image (330 vs 328
core-s on cytc), so only the mask test changed.
Measured (perf, 499 Hz, CPU-only build, cytc, first version of this change): AccumulateRings
591 -> 539 core-s. Byte-identical p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz on myob, cytc, lyso,
sparse (CPU) and myob, lyso (GPU).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The CPU-only image loop is DRAM-bound on 16M frames: decode, preprocess, the adaptive finder's plain
ring pass and FlagRings each streamed the whole frame through memory.
- JFJochDecompressHperfBlocks hands each decoded bitshuffle block to a callback; with no output
buffer the block is unshuffled into a reused block-sized scratch (JFJochDecompressBlocks).
- MXAnalysisWithoutFPGA::PreprocessCPU preprocesses each block into the int32 buffer
(ImagePreprocessorCPU::AnalyzeBlock) and, when the fused CPU finder runs, puts it through the plain
ring pass + fused azint (AdaptiveSpotFinderCPU::AccumulateRingsBlock) while it is in cache. Detect()
then starts from those sums. The per-worker decompression buffer is no longer allocated for
bitshuffled data.
- FlagRings becomes FlagRow, called by DetectAt's first pass for row y+NBX just before that row enters
the vertical sums; first_pass_needed is marked from each row's candidates at the same point.
Exact: blocks arrive in pixel order, so the float azint sums see the same pixels in the same order;
the per-pixel expressions are unchanged; everything else is integer. p.hkl, p.mtz, p_P1.mtz and
p_unmerged.mtz byte-identical to rc173 on myob, cytc, lyso, sparse (CPU-only build), GPU myob
identical (the GPU path does not take this route).
CPU-only, 32 workers, under the gpulock on a shared (loaded) machine, base -> fused, two rounds
(second in reversed order):
myob 155.9 -> 97.7 s, 139.9 -> 88.1 s (loop 46.1 -> 26.8 s/pass; user 3690 -> 2221 s)
cytc 220.8 -> 156.6 s, 216.6 -> 155.1 s (user 5160 -> 4128 s)
lyso 134.5 -> 123.3 s, 69.5 -> 64.8 s
peak RSS myob 15.2 -> 10.8 GB, cytc 12.6 -> 10.4 GB, lyso 7.0 -> 6.7 GB
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
This is an UNSTABLE release. The release has significant modifications and bug fixes, if things go wrong, it is better to revert to 1.0.0-rc.132.
* jfjoch_broker: Better track time for each operation in the processing stack
* jfjoch_broker: Rewrite preprocessing of diffraction images in the non-FPGA workflow to better use GPUs (work in progress)
* jfjoch_broker: Remove ROI calculation in the non-FPGA workflow (work in progress)
* jfjoch_viewer: Toolbar displays image number starting from 1 (instead of 0)
Reviewed-on: #46