c8692e320c968680535a0209c1f664b766742ab4
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c8692e320c |
CPU pixel loops: ring sums held in registers, vectorised vertical window, shared packed mask
AdaptiveSpotFinderCPU::AccumulateRings: consecutive pixels mostly share a ring, so the ring's sum/sum2/count and the fused azint sums are held in locals while they do and stored when the ring changes - the same additions in the same order (az_sum2 is still contracted to the same FMA), without a store-and-reload chain through memory on every pixel. ImageSpotFinderCPU::DetectPass: the vertical-sum update (add the entering row, take out the leaving one) is one branch-free loop over the raw image that GCC vectorises (int64 lanes). A pixel strong in the previous pass used to be substituted per pixel through a bit test, which kept the loop scalar; it is now added with its row and taken out again from the few set bits of prev_strong. Integer sums, so the same totals. (A first, fully branch-free version that kept the per-pixel bit test did not vectorise on the prev_strong path and was measured slower; this is its replacement.) ImagePreprocessorCPU: the per-engine std::vector<bool> built bit by bit from the 32-bit mask (~10 core-s per cytc run, one per worker per pass) is replaced by 32-pixel mask words that PixelMask derives once beside its binary mask; each engine copies 2 MB. A branch-free rewrite of the Analyze loop was measured and dropped: the loop is bound by reading the decompressed image (330 vs 328 core-s on cytc), so only the mask test changed. Measured (perf, 499 Hz, CPU-only build, cytc, first version of this change): AccumulateRings 591 -> 539 core-s. Byte-identical p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz on myob, cytc, lyso, sparse (CPU) and myob, lyso (GPU). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C |
||
|
|
c981e1b91c |
v1.0.0-rc.137 (#46)
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 10m7s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 10m35s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 11m8s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 9m24s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m29s
Build Packages / build:rpm (rocky8) (push) Successful in 10m27s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 11m41s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 11m1s
Build Packages / Generate python client (push) Successful in 45s
Build Packages / Unit tests (push) Has been skipped
Build Packages / Create release (push) Has been skipped
Build Packages / build:rpm (rocky9) (push) Successful in 12m48s
Build Packages / Build documentation (push) Successful in 1m3s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m10s
Build Packages / XDS test (durin plugin) (push) Successful in 8m59s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m32s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m39s
Build Packages / DIALS test (push) Successful in 13m13s
This is an UNSTABLE release. The release has significant modifications and bug fixes, if things go wrong, it is better to revert to 1.0.0-rc.132. * jfjoch_broker: Better track time for each operation in the processing stack * jfjoch_broker: Rewrite preprocessing of diffraction images in the non-FPGA workflow to better use GPUs (work in progress) * jfjoch_broker: Remove ROI calculation in the non-FPGA workflow (work in progress) * jfjoch_viewer: Toolbar displays image number starting from 1 (instead of 0) Reviewed-on: #46 |