docs: tighten the changelog and the decoding section for release
Build Packages / build:viewer-tgz:cpu (push) Successful in 11m6s
Build Packages / build:viewer-tgz:cuda (push) Successful in 15m41s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 19m45s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 20m58s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 24m16s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 24m47s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 22m49s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 17m45s
Build Packages / build:rpm (rocky9) (push) Successful in 18m59s
Build Packages / XDS test (durin plugin) (push) Successful in 11m18s
Build Packages / build:rpm (rocky8) (push) Successful in 25m1s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 19m31s
Build Packages / Generate python client (push) Successful in 49s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m30s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 23m30s
Build Packages / DIALS test (push) Successful in 17m23s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m53s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m21s
Build Packages / Unit tests (push) Successful in 1h56m58s
Build Packages / build:windows:nocuda (push) Successful in 18m13s
Build Packages / build:windows:cuda (push) Successful in 22m29s

The rc.161 changelog had grown entries several hundred words long and listed the
same area three or four times over. Collapse them by subject - spot finding,
resolution limits, space-group search, scaling, performance, correctness - and
hold each to one line, keeping the actionable detail in the breaking API entry.
Add the performance work that had not been written up: device-side image
decoding and the parallel first-pass candidate-cell refinement.

Section 0 of the CPU analysis document was the longest thing in it after two
core algorithm sections, and most of that was a profiling narrative and the
measurements that motivated the change rather than a description of what runs.
Cut it to the two kernels, the host-side block scan and the fallback rule. The
attribution stays; it is also in the reference list.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-03 07:44:55 +02:00
co-authored by Claude Opus 5
parent 7e47afe47f
commit 737cbde3ff
2 changed files with 32 additions and 61 deletions
+18 -27
View File
@@ -3,34 +3,25 @@
### 1.0.0-rc.161
This is an UNSTABLE release. It includes many experimental features, as well as many AI generated fixes. We recommend using rc.152 for production use.
* Spot finding: Self-calibrating **adaptive detection** added - no per-dataset threshold tuning.
* Spot finding: Fused GPU engine (azimuthal integration + detection in one image pass).
* rugnux: Adaptive detection is now the **default** for rotation data as well as stills.
* rugnux: Serial stills pick `--min-pix-per-spot` per image.
* rugnux: Stills **partiality post-refinement** added, on by default (`--simple-stills` disables).
* rugnux: De-novo **space-group search** made substantially more robust.
* rugnux: Rotation **scaling** hardened against a collapsed per-frame scale.
* rugnux: Per-image **geometry refinement** weights spots by confidence.
* rugnux: `--min-image-cc` now works for rotation data (opt-in).
* rugnux: New `--search-min-zeta` (rotation default 0.85).
* Bragg integration: the prediction/integration resolution limit now defaults to **as far as the detector reaches**, instead of a hardcoded 1.0 Å that nothing could change. Any experiment whose detector reached past 1.0 Å was silently losing everything beyond it (5 of 33 rotation test datasets do, down to 0.981 Å); the reflections were found by spot finding and then never integrated. rugnux `--integration-high-resolution <num>` sets a limit by hand (0 = no limit).
* Bragg prediction: how far the predictor walks the lattice is now a setting (`bragg_integration_settings.max_hkl`) rather than a hardcoded 100, for both rotation and stills. Offline (rugnux, viewer) it is **derived per crystal** from the refined cell as `ceil(max(a,b,c)/d_min) + 1` — an exact bound, so it recovers the outermost reflections that a fixed 100 was quietly truncating on a long-axis or high-resolution dataset, while costing less than before on a small cell. Measured: up to +17% observations and high-shell CC1/2 15.1% → 25.8%, with small cells bit-identical. `rugnux --max-hkl <n>` overrides it. Online the broker bootstraps 100, so a live acquisition keeps a predictable per-image cost whatever crystal is mounted; it is settable via the API and the frontend.
* rugnux: Reports how close a symmetry axis lies to the spindle.
* rugnux: Azimuthal-integration and spot-finding resolution limits default to the detector — including **rotation** data, which no longer keeps a 1.5 Å spot-finding limit (`--spot-high-resolution <num>` still sets one).
* **API (breaking for older clients)**: the `image_scale_b` value is gone from the `plot_type` enum, so a client that requests that plot now gets an error rather than a curve. Also, `azim_int_settings.high_q_recipA` and `spot_finding_settings.high_resolution_limit` are no longer `required`. Both now mean "as far as the detector reaches" when unset, and are **omitted** from the response instead of carrying a placeholder value. A client generated from an rc.160-or-earlier spec that assumes the field is always present (the generated C++ `from_json` does `j.at(...)`) raises when it is missing. Because the azimuthal-integration limit now defaults to unset, a stock rc.161 broker omits `high_q_recipA` from `GET /config/azim_int` with no operator action. Regenerate the client (`jfjoch-client` 1.0.0-rc.161, `frontend/src/client`) or read both fields as optional.
* rugnux: Several non-helping stills scaling/detection knobs removed.
* Reader: A stored dataset is read as **int32**, so the saturation cap comes from the file's own `saturation_value` instead of a container depth that no longer matched the pixels handed out (32-bit EIGER2 files were capping at 32767 and dropping their strongest reflections; 16- and 8-bit files were capping at half their range).
* jfjoch_viewer: Image rendering and interaction performance improved.
* rugnux: The space-group search takes **systematic absences from the merge of all observations** in every case. `--search-min-zeta` rescues the point group only, as documented - when the filtered merge was the one that found more symmetry, its screws and centering were being adopted too.
* rugnux: A search pass no longer alters the **production merge**. The observations it drops (`--search-min-zeta`, `--min-image-cc`) are restored for the next pass, and on GPU builds the search no longer ran on unscaled intensities.
* rugnux: The merged-sigma **systematic floor** is kept when ISa is too degenerate to report, instead of being switched off with it.
* rugnux: A **screw axis** is only claimed against an axial row with at least three control reflections, so a single strong reflection cannot invent one.
* rugnux: A **collapsed per-image scale** leaves that image unscaled on the stills path, as it already did for rotation; a stills partiality refinement whose solve failed no longer adopts its tilt.
* rugnux: Bad command-line input is **reported** (exit 1) instead of terminating the process; `--scaling-high-resolution` and `--integration-radius` are parsed like every other numeric option.
* rugnux: A run in which **nothing indexes** writes a `_process.h5` with the per-image analysis rather than finishing with no output at all.
* GPU: The fused adaptive spot finder sums its ring statistics in double, so it reproduces the CPU finder and itself run to run (no measurable cost). GPU merging no longer subtracts a negative intensity's Poisson term, and GPU Bragg prediction covers the full hkl range.
* Spot finding: Self-calibrating **adaptive detection** added, now the **default** for rotation as well as stills - no per-dataset threshold tuning; a fused GPU engine runs it together with azimuthal integration in one image pass.
* Spot finding: Serial stills pick `--min-pix-per-spot` per image; connected components run on the GPU; detection is configurable over the API, where a `high_resolution_limit` of 0 means "no limit" instead of throwing.
* Resolution limits: Bragg integration, azimuthal integration and spot finding all default to **as far as the detector reaches**, replacing a fixed 1.0 Å integration limit and a 1.5 Å rotation spot-finding limit that silently discarded everything beyond them; `--integration-high-resolution` and `--spot-high-resolution` still set one by hand.
* Bragg prediction: How far the predictor walks the lattice is a setting (`bragg_integration_settings.max_hkl`) instead of a fixed 100, derived per crystal offline from the refined cell (`--max-hkl` overrides); the broker keeps a fixed bootstrap so a live acquisition has a predictable per-image cost.
* rugnux: De-novo **space-group search** substantially more robust - centering ranked by net absences and judged on absent-class strength, merohedral-twin over-promotion vetoed, and genuine high-symmetry groups recovered on weak data.
* rugnux: The space-group search takes systematic absences from the merge of all observations, needs at least three control reflections on an axial row to claim a **screw axis**, and no longer alters the production merge.
* rugnux: Stills **partiality post-refinement** added, on by default (`--simple-stills` disables); several non-helping stills scaling/detection knobs removed.
* rugnux: **Scaling** hardened against a collapsed per-frame scale on the stills path as well as rotation, and the merged-sigma systematic floor is kept when ISa is too degenerate to report.
* rugnux: `--min-image-cc` now works for rotation data (opt-in); new `--search-min-zeta` (rotation default 0.85); reports how close a symmetry axis lies to the spindle.
* Performance: bitshuffle+LZ4 images are **decoded on the GPU** rather than on the host (1.8x on a large rotation run), and first-pass rotation indexing refines candidate cells in parallel instead of on one thread per scheme.
* Performance: Assorted work removed from the hot paths - GPU lookup tables shared per device, fewer device allocations, cheaper candidate filtering and merged-file writing; jfjoch_viewer image rendering and interaction improved.
* Reader: A stored dataset is read as **int32**, so the saturation cap comes from the file's own `saturation_value` (32-bit files were capping at 32767 and dropping their strongest reflections; 16- and 8-bit files at half their range).
* GPU: Spot finding, merging and Bragg prediction now match the CPU path and reproduce run to run.
* rugnux: **Geometry refinement** no longer commits a solve that did not converge, and per-image refinement weights spots by confidence.
* Lattice search: Corrected the reindex matrices for two Niggli characters.
* Azimuthal integration: A flat ring no longer yields a NaN standard deviation, and a low-Q limit above the maximum is rejected instead of producing undefined behaviour.
* Spot finding: A `high_resolution_limit` of 0 means "no limit" over the API too, instead of throwing on every image.
* rugnux: Bad command-line input is **reported** (exit 1) instead of terminating the process, and a run in which nothing indexes still writes a `_process.h5`.
* Broker / writer / image pusher: Image-slot lifetime, locking and deactivation fixes.
* **API (breaking for older clients)**: `image_scale_b` is gone from the `plot_type` enum, and `azim_int_settings.high_q_recipA` and `spot_finding_settings.high_resolution_limit` are no longer `required` - both mean "as far as the detector reaches" when unset and are omitted from the response, which raises in a client generated from an rc.160-or-earlier spec. Regenerate the client (`jfjoch-client` 1.0.0-rc.161, `frontend/src/client`) or read both fields as optional.
### 1.0.0-rc.160
This is an UNSTABLE release. It includes many experimental features, as well as many AI generated fixes. We recommend using rc.152 for production use.
+14 -34
View File
@@ -30,53 +30,33 @@ The methods are inspired and reuising solutions implemented in:
- S. French & K. Wilson, "On the treatment of negative intensity observations", *Acta Cryst.* **A34** (1978), 517-525 (Bayesian amplitude estimation from intensities).
- A. T. Brünger, "Free R value: a novel statistical quantity for assessing the accuracy of crystal structures", *Nature* **355** (1992), 472-475 (R-free cross-validation).
- M. Wojdyr, "GEMMI: A library for structural biology", *J. Open Source Softw.* **7** (2022), 4200 (model / structure-factor / map machinery used in §14).
- J. P. Wright, "Experiences with GPU decompression for bitshuffle + LZ4 data", HDF5 User Group meeting (2021), and [github.com/jonwright/bslz4decoders](https://github.com/jonwright/bslz4decoders) decoding bitshuffle+LZ4 on the GPU instead of the host, which is the idea behind §0.
- J. P. Wright, "Experiences with GPU decompression for bitshuffle + LZ4 data", HDF5 User Group meeting (2021), and [github.com/jonwright/bslz4decoders](https://github.com/jonwright/bslz4decoders) (device-side decoding of bitshuffle+LZ4 images, §0).
(list is not exhaustive)
## 0. Getting the image onto the GPU: device-side bitshuffle+LZ4 decoding
Images arrive bitshuffle+LZ4 compressed (HDF5 filter 32008), and everything from §1 onwards runs on
the GPU when one is present. The obvious arrangement — decompress on the host, upload the image
turns out to be the most expensive part of the whole per-image loop. Profiling an 18 Mpx rotation
dataset showed the host-to-device copy occupying **78% of the loop** against **39% for all kernels
combined**: 3600 copies of 72.4 MB each, and running at only 12.5 GB/s of an available 2728
because the host-side decompression was itself saturating host memory bandwidth.
the GPU when one is present. Instead of decompressing on the host and uploading the image, the
compressed chunk is uploaded — a few MB rather than tens of MB — and decoded on the device. The
approach follows Jon Wright (ESRF); the kernels are Jungfraujoch's own.
So the compressed chunk is uploaded instead — about 4 MB rather than 72 MB, an 18× reduction — and
decoded on the device. This removes the transfer *and* the host decompression that was throttling
it. Measured on that dataset the run went from 49.4 s to 28.0 s with byte-identical output.
**This approach is Jon Wright's (ESRF)** — see his HDF5 User Group talk "Experiences with GPU
decompression for bitshuffle + LZ4 data" (2021) and
[github.com/jonwright/bslz4decoders](https://github.com/jonwright/bslz4decoders). The kernels in
`image_analysis/image_preprocessing/BSLZ4DecoderGPU.cu` are our own, but the idea, and the
demonstration that it is worth doing at all, are his.
Two kernels do the work, mirroring the CPU decoder exactly:
Two kernels mirror the CPU decoder:
1. **LZ4, one warp per bitshuffle block.** Blocks are independent, so the parallelism is across
them; within a warp every lane runs the same sequence parser over the same bytes (a broadcast
read, no divergence) and the literal and match copies are split across the 32 lanes so the
stores coalesce. An overlapping match is treated as a pattern of period `offset` sourced from
bytes that already precede the write position, which keeps it parallel rather than a serial byte
loop — the long zero runs of sparse detector data arrive here with `offset == 1`. One thread per
block instead, with each thread streaming its own 8 kB region, measured 13× slower.
them; within a warp every lane runs the same sequence parser over the same bytes, and the
literal and match copies are split across the 32 lanes so the stores coalesce. An overlapping
match is treated as a pattern of period `offset` sourced from bytes that already precede the
write position, which keeps it parallel rather than a serial byte loop.
2. **The bitshuffle inverse**, one CUDA block per bitshuffle block: bit un-transpose of each
byte-plane into shared memory, then interleave the planes back into elements so the final store
is coalesced. For 8-bit images (a real DECTRIS mode) there is a single plane and the interleave
degenerates to a copy, matching `bitshuf_decode_block`'s `elem_size == 1` branch.
is coalesced. For 8-bit images there is a single plane and the interleave degenerates to a copy.
The block offsets inside the container can only be discovered by reading the block lengths in
order, so that scan stays on the host; it is a few hundred microseconds per frame against the
milliseconds it saves.
order, so that scan stays on the host.
**Only BSHUF_LZ4 is decoded on the device.** The zstd variants (`BSHUF_ZSTD`, `BSHUF_ZSTD_RLE`,
`BSHUF_ZSTD_RLE_HUFF`) have no device decoder, and neither does an uncompressed or float image;
for those `BSLZ4DecoderGPU::Supports()` returns false and the pipeline decompresses on the host and
uploads as before. The fallback is explicit rather than implicit, so an algorithm we cannot decode
on the device is a slower path and never a wrong answer. `tests/BSLZ4DecoderGPUTest.cpp` holds the
device decoder against the CPU one byte for byte, on data produced by the production compressor,
for every element size the detectors emit.
Only `BSHUF_LZ4` is decoded on the device. For the zstd variants (`BSHUF_ZSTD`, `BSHUF_ZSTD_RLE`,
`BSHUF_ZSTD_RLE_HUFF`), and for uncompressed or float images, `BSLZ4DecoderGPU::Supports()` returns
false and the pipeline decompresses on the host and uploads as before.
## 1. Geometry, reciprocal-space mapping, and basic quantities