Commit Graph
2 Commits
Author SHA1 Message Date
leonarski_fandClaude Opus 5.5 9c5141c7c0 CPU pixel pipeline: decode, preprocess and ring pass per bitshuffle block; flag rings in the local sweep
The CPU-only image loop is DRAM-bound on 16M frames: decode, preprocess, the adaptive finder's plain
ring pass and FlagRings each streamed the whole frame through memory.

- JFJochDecompressHperfBlocks hands each decoded bitshuffle block to a callback; with no output
  buffer the block is unshuffled into a reused block-sized scratch (JFJochDecompressBlocks).
- MXAnalysisWithoutFPGA::PreprocessCPU preprocesses each block into the int32 buffer
  (ImagePreprocessorCPU::AnalyzeBlock) and, when the fused CPU finder runs, puts it through the plain
  ring pass + fused azint (AdaptiveSpotFinderCPU::AccumulateRingsBlock) while it is in cache. Detect()
  then starts from those sums. The per-worker decompression buffer is no longer allocated for
  bitshuffled data.
- FlagRings becomes FlagRow, called by DetectAt's first pass for row y+NBX just before that row enters
  the vertical sums; first_pass_needed is marked from each row's candidates at the same point.

Exact: blocks arrive in pixel order, so the float azint sums see the same pixels in the same order;
the per-pixel expressions are unchanged; everything else is integer. p.hkl, p.mtz, p_P1.mtz and
p_unmerged.mtz byte-identical to rc173 on myob, cytc, lyso, sparse (CPU-only build), GPU myob
identical (the GPU path does not take this route).

CPU-only, 32 workers, under the gpulock on a shared (loaded) machine, base -> fused, two rounds
(second in reversed order):
  myob  155.9 -> 97.7 s, 139.9 -> 88.1 s   (loop 46.1 -> 26.8 s/pass; user 3690 -> 2221 s)
  cytc  220.8 -> 156.6 s, 216.6 -> 155.1 s (user 5160 -> 4128 s)
  lyso  134.5 -> 123.3 s, 69.5 -> 64.8 s
  peak RSS myob 15.2 -> 10.8 GB, cytc 12.6 -> 10.4 GB, lyso 7.0 -> 6.7 GB

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-28 02:00:50 +02:00
leonarski_f c981e1b91c v1.0.0-rc.137 (#46)
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 10m7s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 10m35s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 11m8s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 9m24s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m29s
Build Packages / build:rpm (rocky8) (push) Successful in 10m27s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 11m41s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 11m1s
Build Packages / Generate python client (push) Successful in 45s
Build Packages / Unit tests (push) Has been skipped
Build Packages / Create release (push) Has been skipped
Build Packages / build:rpm (rocky9) (push) Successful in 12m48s
Build Packages / Build documentation (push) Successful in 1m3s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m10s
Build Packages / XDS test (durin plugin) (push) Successful in 8m59s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m32s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m39s
Build Packages / DIALS test (push) Successful in 13m13s
This is an UNSTABLE release. The release has significant modifications and bug fixes, if things go wrong, it is better to revert to 1.0.0-rc.132.

* jfjoch_broker: Better track time for each operation in the processing stack
* jfjoch_broker: Rewrite preprocessing of diffraction images in the non-FPGA workflow to better use GPUs (work in progress)
* jfjoch_broker: Remove ROI calculation in the non-FPGA workflow (work in progress)
* jfjoch_viewer: Toolbar displays image number starting from 1 (instead of 0)

Reviewed-on: #46
2026-04-25 19:59:21 +02:00