82e6583e007004ffa5f358202f97f3cc69df8dc5
5
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a277b8c458 |
Merge rc173 (fused CPU pixel pipeline) into cpu-loops-tail
Conflicts in the three CPU pixel loops the fused pipeline split per block, resolved by carrying this branch's loop bodies into the new functions with the per-pixel arithmetic unchanged: - AdaptiveSpotFinderCPU: the register-held ring sums now live in AccumulateRingsBlock (stored at the end of each block, so the additions stay in pixel order). - ImagePreprocessorCPU::AnalyzeBlock reads the PixelMask-derived 32-pixel mask words at first + i. - ImageSpotFinderCPU::DetectPass: rc173's new_row() marking/fill_row calls kept in place, the vertical update replaced by the vectorised slide with the prev_strong fix-up. Byte-identical p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz vs rc173 references on myob, cytc, lyso, sparse (CPU-only build) and myob (GPU build); targeted tests (ImageSpotFinderCPU*, AdaptiveSpotFinder, SpotFinding, PixelMask, Bragg*, RotationScale, AzimuthalIntegration, portable) pass. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C |
||
|
|
c8692e320c |
CPU pixel loops: ring sums held in registers, vectorised vertical window, shared packed mask
AdaptiveSpotFinderCPU::AccumulateRings: consecutive pixels mostly share a ring, so the ring's sum/sum2/count and the fused azint sums are held in locals while they do and stored when the ring changes - the same additions in the same order (az_sum2 is still contracted to the same FMA), without a store-and-reload chain through memory on every pixel. ImageSpotFinderCPU::DetectPass: the vertical-sum update (add the entering row, take out the leaving one) is one branch-free loop over the raw image that GCC vectorises (int64 lanes). A pixel strong in the previous pass used to be substituted per pixel through a bit test, which kept the loop scalar; it is now added with its row and taken out again from the few set bits of prev_strong. Integer sums, so the same totals. (A first, fully branch-free version that kept the per-pixel bit test did not vectorise on the prev_strong path and was measured slower; this is its replacement.) ImagePreprocessorCPU: the per-engine std::vector<bool> built bit by bit from the 32-bit mask (~10 core-s per cytc run, one per worker per pass) is replaced by 32-pixel mask words that PixelMask derives once beside its binary mask; each engine copies 2 MB. A branch-free rewrite of the Analyze loop was measured and dropped: the loop is bound by reading the decompressed image (330 vs 328 core-s on cytc), so only the mask test changed. Measured (perf, 499 Hz, CPU-only build, cytc, first version of this change): AccumulateRings 591 -> 539 core-s. Byte-identical p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz on myob, cytc, lyso, sparse (CPU) and myob, lyso (GPU). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C |
||
|
|
9c5141c7c0 |
CPU pixel pipeline: decode, preprocess and ring pass per bitshuffle block; flag rings in the local sweep
The CPU-only image loop is DRAM-bound on 16M frames: decode, preprocess, the adaptive finder's plain ring pass and FlagRings each streamed the whole frame through memory. - JFJochDecompressHperfBlocks hands each decoded bitshuffle block to a callback; with no output buffer the block is unshuffled into a reused block-sized scratch (JFJochDecompressBlocks). - MXAnalysisWithoutFPGA::PreprocessCPU preprocesses each block into the int32 buffer (ImagePreprocessorCPU::AnalyzeBlock) and, when the fused CPU finder runs, puts it through the plain ring pass + fused azint (AdaptiveSpotFinderCPU::AccumulateRingsBlock) while it is in cache. Detect() then starts from those sums. The per-worker decompression buffer is no longer allocated for bitshuffled data. - FlagRings becomes FlagRow, called by DetectAt's first pass for row y+NBX just before that row enters the vertical sums; first_pass_needed is marked from each row's candidates at the same point. Exact: blocks arrive in pixel order, so the float azint sums see the same pixels in the same order; the per-pixel expressions are unchanged; everything else is integer. p.hkl, p.mtz, p_P1.mtz and p_unmerged.mtz byte-identical to rc173 on myob, cytc, lyso, sparse (CPU-only build), GPU myob identical (the GPU path does not take this route). CPU-only, 32 workers, under the gpulock on a shared (loaded) machine, base -> fused, two rounds (second in reversed order): myob 155.9 -> 97.7 s, 139.9 -> 88.1 s (loop 46.1 -> 26.8 s/pass; user 3690 -> 2221 s) cytc 220.8 -> 156.6 s, 216.6 -> 155.1 s (user 5160 -> 4128 s) lyso 134.5 -> 123.3 s, 69.5 -> 64.8 s peak RSS myob 15.2 -> 10.8 GB, cytc 12.6 -> 10.4 GB, lyso 7.0 -> 6.7 GB Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C |
||
|
|
54c0100e8e |
v1.0.0-rc.157 (#67)
Build Packages / Unit tests (push) Successful in 1h28m28s
Build Packages / build:windows:nocuda (push) Successful in 14m45s
Build Packages / build:windows:cuda (push) Successful in 13m13s
Build Packages / build:viewer-tgz:cpu (push) Successful in 6m47s
Build Packages / build:viewer-tgz:cuda (push) Successful in 7m22s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 13m52s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 14m16s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m19s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 12m50s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 14m40s
Build Packages / build:rpm (rocky8) (push) Successful in 11m18s
Build Packages / build:rpm (rocky9) (push) Successful in 12m4s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 11m55s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m22s
Build Packages / DIALS test (push) Successful in 13m37s
Build Packages / XDS test (durin plugin) (push) Successful in 8m47s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m4s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m45s
Build Packages / Generate python client (push) Successful in 34s
Build Packages / Build documentation (push) Successful in 1m4s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 7m16s
This is an UNSTABLE release. It includes many experimental features, as well as many AI generated fixes. We recommend using rc.152 for production use. * rugnux: Rebrand the offline data-processing subsystem as `rugnux` and consolidate all offline analysis into the single `rugnux` binary - `jfjoch_process` is now `rugnux`, the former `jfjoch_azint` is now `rugnux --azint-only`, and `jfjoch_scale` is now `rugnux --scale` (see the new docs/NAMING.md and docs/RUGNUX.md). Scaling and merging are on by default for rotation and stills (`--no-merge` disables them), replacing the previous opt-in `-M, --scale-merge`. * rugnux: CLI fixes - default `-N` to all hardware threads, parse numeric option arguments strictly (reject non-numeric or trailing input instead of silently yielding 0), require `--wavelength > 0`, and correct the reproduced command line and `--scale` reference-cell handling. * rugnux: De-novo space-group improvements - recover genuine high symmetry and centred Bravais lattices from intensities, add an automatic CC1/2 high-resolution cutoff, and report L-test twinning statistics. * rugnux: Index weakly-diffracting low-resolution rotation data that previously failed (e.g. F-cubic crystals that diffract only to ~4 A on a detector reaching ~1.5 A). The per-frame indexing gate now measures the indexed fraction only within the resolution range the lattice actually diffracts to, so the many sub-diffraction ice/noise spots no longer make the fraction floor unreachable; the two-pass first pass tries several image-sampling schemes (spread across the whole rotation vs a consecutive wedge whose native stride keeps a reflection's rocking curve continuous, letting the FFT resolve a long axis) and keeps the one that indexes the most frames; and the de-novo space-group search no longer discards all reflections (and crashes) when every resolution shell falls below <I/sigma> = 1. * rugnux: Lower the low-resolution R-meas for strongly-diffracting rotation data - drop edge-of-sweep truncated fulls whose rocking curve was captured below `--min-captured-fraction` (default 0.7 for rotation), and report R-meas only over the observations kept by outlier rejection (matching XDS). The 0.7 default also strips the partiality-extrapolated fulls that dominate the intensity second moment on weakly-diffracting crystals, so the de-novo space-group search is no longer starved by the error-model I/sigma floor and recovers the correct symmetry (e.g. the F-cubic Benas crystals: Benas_3 -> F432, Benas_7 -> P6122, instead of P4/P1); on the reference battery every other crystal keeps its space group. * rugnux: Write the refined geometry (beam, tilt, axis) to _process.h5 and place non-standard mmCIF items under a reserved `jfjoch` prefix. * jfjoch_broker: Ordinary acquisition failures (receiver/writer/analysis problems, missed packets, writer disconnect) now return to the Idle state with an Error-severity message, so a run can be retried without an expensive re-initialisation; only failures that leave the detector in an undefined state (new JFJochCriticalException, e.g. PCIe/FPGA faults) go to the Error state and force re-initialisation. * jfjoch_broker: A synchronous /start now reports its failure to the HTTP caller instead of returning HTTP 200, and an incomplete or truncated dataset (missing packets, writer disconnect) is reported as an error rather than a "reduce frame rate" warning. * jfjoch_broker: Drop uncollected placeholder rows (number = -1) from the scan_result REST endpoint. * jfjoch_broker: Fix the inverted per-image compression ratio reported by the Lite receiver (was compressed/uncompressed instead of uncompressed/compressed). * jfjoch_broker: Bragg integration adds a quantization-noise variance floor with a box-sum fallback, and treats the type-maximum marker as an invalid pixel for unsigned image types. * jfjoch_writer: Detect file-overwrite conflicts at start for back-channel transports, and reset the writer when end-of-collection finalisation fails. * jfjoch_viewer: Preview overlays follow the geometry (resolution/ROI arcs, true beam centre, predictions, coral secondary-lattice spots, legend), add save-as-JPEG, and fix an HTTP live-follow memory leak. * Frontend: Improved aesthetics and usability, and added in-browser pixel-mask and JUNGFRAU-pedestal visualisation. * CI: Name the Windows installer jfjoch-viewer-* instead of jfjoch-*.Reviewed-on: #67 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch> |
||
|
|
c981e1b91c |
v1.0.0-rc.137 (#46)
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 10m7s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 10m35s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 11m8s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 9m24s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m29s
Build Packages / build:rpm (rocky8) (push) Successful in 10m27s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 11m41s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 11m1s
Build Packages / Generate python client (push) Successful in 45s
Build Packages / Unit tests (push) Has been skipped
Build Packages / Create release (push) Has been skipped
Build Packages / build:rpm (rocky9) (push) Successful in 12m48s
Build Packages / Build documentation (push) Successful in 1m3s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m10s
Build Packages / XDS test (durin plugin) (push) Successful in 8m59s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m32s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m39s
Build Packages / DIALS test (push) Successful in 13m13s
This is an UNSTABLE release. The release has significant modifications and bug fixes, if things go wrong, it is better to revert to 1.0.0-rc.132. * jfjoch_broker: Better track time for each operation in the processing stack * jfjoch_broker: Rewrite preprocessing of diffraction images in the non-FPGA workflow to better use GPUs (work in progress) * jfjoch_broker: Remove ROI calculation in the non-FPGA workflow (work in progress) * jfjoch_viewer: Toolbar displays image number starting from 1 (instead of 0) Reviewed-on: #46 |