8874a788e68054c4445694f424eca00a2c6af459
7
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2222a4085c |
Scaling: correct the absorption that changes as the crystal turns
RefineAbsorption indexes its surface by the diffracted direction de-rotated into the crystal frame, deliberately without a time axis, and RefineModulation indexes its by detector position, also without one. Nothing is indexed by (rotation, detector position), so the part of the absorption that changes as the crystal turns has no parameter at all. For a rigid absorber illuminating a fixed volume that is the right model: the incident path is a function of the spindle angle alone and the per-image scale takes it, and the exit path is then fixed in the crystal frame. What breaks the factorisation is the diffracting volume moving - a crystal larger than the beam, a mis-centred loop, ice building up. The exit path then depends on the spindle angle as well as the direction, and no time-independent surface reaches it. Measured on 34 rotation datasets, on XDS's own uncorrected intensities, as what is left after the crystal-frame absorption and detector modulation surfaces have taken what they can. The cross-validation gate lets the surface engage on 22 of them. Scored per resolution shell - the gate's whole-range ratio is lowered by any resolution-dependent scale without a reflection getting tighter, so the honest readout is each shell's own ratio, which a per-shell scale leaves unchanged - the median engaged crystal gains 4.1 %, the set gains 117 % summed against 19 % of damage, and 3 of the 22 are hurt. The surface has to be smooth in rotation angle to be absorption at all, and it is: the lag-1 autocorrelation of the fitted factor along the time axis runs +0.32 to +0.71 on the crystals it engages, against -0.08 for the same surface with its time bins shuffled. Where it is not smooth it is fitting something else, and says so - on a sweep whose beam was obstructed for a 70 deg wedge the autocorrelation is +0.16 and the profile is a cliff at the wedge, not a turn. Two null controls. Assign every observation a random cell and the gate refuses it (-1.2 % to -3.8 %). Keep the detector bin and shuffle only the time bin - a surface that cannot contain any time-dependent information - and the gate refuses that too, at +0.04 %, -0.60 % and +0.33 % on three crystals. Against the real surface's +3.7 % to +14.8 % on the same three. 12 time bins x a 10 x 10 detector grid = 1200 factors. On the per-shell score the median gain moves only between 3.1 % and 4.1 % across grids from 216 to 2400 cells, so the grid is second order; 12 x 10 has the largest net and the fewest crystals hurt. Equal-occupancy detector bins, not equal width: an equal-width grid starves the edges and the corners, and a starved cell is where a free surface over-fits. Fitted last, so the two time-independent surfaces get first claim on what they can explain. QUALIFICATION, measured after this was written: the "33 better / 0 worse" above is overall R_meas, which is a ratio of sums across every shell and is therefore lowered by any resolution-dependent scale without a reflection getting tighter - the same property that let the estimator bias pass its own gate. Scored per resolution shell instead, this surface HURTS 6 of 18 crystals under the acceptance gate as it currently stands, because that gate shares the defect and admits the surface where it should not. With a per-shell gate the surface is refused on exactly those crystals and its net over the chain goes from +86.3 to +159.5 per cent with none worse. The correction is right; the gate that decides where to apply it is the next commit's problem, not this one's. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Full 38-crystal rotation battery against its own matched baseline - the same binary with the corrected estimator and without this surface: R_meas better 33 / worse 0, -78.7 R_meas_lo better 24 / worse 2, -20.7 CC1/2 better 8 / worse 0, +9.6 ISa better 30 / worse 2, +99.15 space groups unchanged at 35/38 The low-energy datasets gain most, which is what absorption should do: at 5-6 keV one crystal goes ISa 24.68 -> 37.42 and another 14.21 -> 22.08, while the same protein measured at 13 keV moves 13.41 -> 14.83. Taken with the estimator fix it precedes, against a clean baseline: R_meas_lo better 25 / worse 3 summed -26.0, ISa better 31 / worse 2 summed +113.1, outer-shell CC1/2 +83.0, observations +25 880 on 36 crystals of 38, CC1/2 flat at -1.4 and no space group moved. That last number is the point of the pair: the estimator fix alone costs CC1/2 -11.4, because the ramp it removes was partly standing in for this correction. One cost, predicted in advance and still unexplained: outer-shell CC1/2 falls on three of the four low-energy crystals, by 15.8 points on the worst, while every other statistic on those same crystals improves. The fourth goes up. On 5000-9000 That outer-shell fall has since been attributed, and it is not this surface: with the merge's 6-sigma outlier rejection turned off, the sign flips on every crystal that lost, +20.5 and +20.7 where it read -15.8 and -8.0. The surface removes most of the deviants in sample - it is fitted on all the data and applied to it, with no robustness of its own - so the merge's cut stops firing and the survivors land in a shell whose multiplicity is about three. Last-shell CC1/2 is largely made by that cut: one crystal's baseline goes 10.4 to 92.7 purely by dropping 19 per cent of the shell. A second qualification, measured after the numbers above were taken: overall R_meas is a ratio of sums across every shell, so any resolution-dependent scale lowers it without a reflection getting tighter - the same property that let the estimator bias pass its own gate. Scored per resolution shell instead, this surface hurts 6 of 18 crystals under the acceptance gate as it stands, because that gate shares the defect and admits the surface where it should not. Under a per-shell gate it is refused on exactly those crystals and its net over the correction chain goes from +86.3 to +159.5 per cent with none worse. The correction is right; where to apply it is the gate's problem. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
43627e22dc |
docs: a rule for licences and academic credit, and apply it
Several methods adopted recently came from other crystallographic packages - the screw-absence test from POINTLESS, MINPK and the profile-fit reweighting from XDS/Otwinowski, the CC1/2 cutoff and merge outlier rejection from DIALS, the per-frame indexing gate from CrystFEL - and nothing in the repository said where such a debt is recorded. The licence side was already worked out (licences beside the vendored code, verbatim texts in licenses/ collected by COLLECT.sh, a row in THIRD_PARTY_NOTICES.md, all installed under share/doc/jfjoch); the credit side was ad hoc. Write the rule into CLAUDE.md. It states the distinction that matters: vendoring or linking someone's CODE creates a LICENCE obligation, discharged in licenses/ and THIRD_PARTY_NOTICES.md; reimplementing an algorithm from a PAPER creates none of that but creates an obligation of academic CREDIT, discharged in docs/ACKNOWLEDGEMENT.md and in a comment at the algorithm. Neither substitutes for the other, and taking both source and paper incurs both. It also fixes the citation form (authors, title, year, journal, volume, pages, verified DOI), and says in-source credit goes at the algorithm, not the file header, in the one-line style the code already uses. Then bring the repository into compliance for the works concerned: docs/ACKNOWLEDGEMENT.md gains a section acknowledging XDS, DIALS, POINTLESS/CCP4, MOSFLM, CrystFEL, GEMMI, the Kabsch/Otwinowski profile fit, the Diederichs & Karplus statistics and the IUCr nomenclature reports, each with a DOI checked against Crossref; docs/CPU_DATA_ANALYSIS.md's reference list gains the ones it was missing; and four algorithms gain a line naming their source where no adjacent comment carried one. No licence change. licenses/ and THIRD_PARTY_NOTICES.md are untouched. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
df9a9c2a2c |
Fix the defects found reviewing the branch before merge
Build Packages / build:viewer-tgz:cpu (push) Successful in 19m32s
Build Packages / build:windows:nocuda (push) Successful in 19m57s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m45s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m38s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m24s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m8s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m9s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m18s
Build Packages / XDS test (durin plugin) (push) Successful in 11m10s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 20m21s
Build Packages / build:windows:cuda (push) Successful in 22m5s
Build Packages / build:rpm (rocky9) (push) Successful in 20m57s
Build Packages / Generate python client (push) Successful in 34s
Build Packages / Build documentation (push) Successful in 1m29s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m41s
Build Packages / DIALS test (push) Successful in 21m19s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m34s
Build Packages / build:rpm (rocky8) (push) Successful in 27m4s
Build Packages / XDS test (neggia plugin) (push) Successful in 10m19s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m58s
Build Packages / Unit tests (push) Successful in 1h17m36s
Image buffer: the per-image CBOR metadata headroom had been re-derived from the online reflection cap alone, which cut it from 4 MiB to 2.55 MB while the measured worst case - reflections plus the capped spot list plus the three azimuthal arrays - is 2.9 MB, so the receiver dropped the frames with the most to say. Restore it and give it a name that both the code and its guard test read: written down twice, the two had drifted and the test kept passing against the value the code had left. Spot finding: an unset low_resolution_limit means no limit at that end, as an unset high_resolution_limit already did. An optional rather than a zero sentinel, because zero is not a natural "no limit" here - every pixel lies above it, so the plain comparison masked the whole image instead of none of it, and nothing validated the zero. The API field is no longer required; a zero is folded into the unset case at the boundary, where older clients still send it, so one spelling reaches the analysis code. The FPGA takes its fixed-point ceiling instead, since ap_ufixed<16,9> wraps above 512 A and would have masked everything. image_preprocessing: check the CUDA calls on the fused decode path - the one new GPU file with none, and the path fed by bytes we did not produce. An unchecked synchronise returned the host-written sentinel as if it were a measurement, so the decode looked successful and the fallback to the host decoder never fired. rugnux: --stride no longer writes one past the end of the per-image arrays, whose count floored where the worker loop ceils, and the written process file links the images actually processed rather than the first N - each frame's picture now sits next to its own analysis. Powder calibration: the face-centred calibrants no longer list their systematically absent rings, so the distance fit starts from a reflection that exists rather than an extinct one; the triclinic calibrant covers both signs of h and k instead of a single octant, which is only valid for a diagonal metric. The test asserted the old behaviour - one ring formula for every cubic standard - and is rewritten. CBOR: skip an unknown tagged value in the end block, as the other four blocks already do. One advance lands on the tagged item rather than past it, so an older reader fed a newer end message threw and never finalized its file. Viewer: a settings value the setter rejects no longer escapes as an uncaught throw from a worker slot, and the field offers only what the setter accepts. Space-group search: judge stage B on the same "present" cut stage A already computes. Merged sigma is floored so no reflection reads above ISa, so on a low-ISa merge the fixed cut left both stage B tests unsatisfiable - every screw axis passed unchallenged and the centering rescue switched itself off on exactly the weak data it exists for. Where the fixed cut is the smaller of the two they are equal and this is inert: over the 37-crystal rotation battery every crystal reports the identical space group and identical merge statistics, so it is a no-op there and the low-ISa case it targets remains unmeasured. rugnux: --polarization reaches --mode azint, which parsed the flag and then dropped it; that mode also applies the same polarization default as every other mode. Acknowledge the ACTS/traccc project, whose sparse connected-component labelling both spot extractors take their algorithm from, with its citation and its license. The rc.161 change list is brought back to one line per entry, and the user-visible changes that were missing from it added. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6e4c0ce202 |
image_preprocessing: decode bitshuffle+LZ4 on the GPU
Build Packages / build:viewer-tgz:cpu (push) Successful in 20m32s
Build Packages / build:viewer-tgz:cuda (push) Successful in 20m40s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m24s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m8s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 27m31s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 27m38s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 29m7s
Build Packages / XDS test (durin plugin) (push) Successful in 11m12s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 22m49s
Build Packages / build:rpm (rocky9) (push) Successful in 22m51s
Build Packages / Generate python client (push) Successful in 40s
Build Packages / Build documentation (push) Successful in 1m22s
Build Packages / Create release (push) Skipped
Build Packages / DIALS test (push) Successful in 20m21s
Build Packages / build:rpm (rocky8) (push) Successful in 27m26s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m59s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m52s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m41s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m41s
Build Packages / Unit tests (push) Successful in 1h17m41s
Build Packages / build:windows:nocuda (push) Successful in 13m24s
Build Packages / build:windows:cuda (push) Successful in 17m0s
The pipeline decompressed each image on the host and uploaded the result. On an 18 Mpx rotation dataset that made the host-to-device copy the bottleneck of the whole per-image loop: nsys puts the copies at 78% of the loop against 39% for every kernel combined - 3600 transfers of 72.4 MB - and they ran at only 12.5 GB/s of an available 27-28 because the host-side decompression was itself saturating host memory bandwidth. The GPU was mostly waiting. So the compressed chunk goes across instead, about 4 MB rather than 72 MB, and is decoded on the device. That removes the transfer and the host decompression that was throttling it, in one change. Measured on an idle machine, a run goes from 45.11 s to 24.97 s - 1.81x - with the merged output unchanged. THE APPROACH IS JON WRIGHT'S (ESRF): "Experiences with GPU decompression for bitshuffle + LZ4 data", HDF5 User Group 2021, and github.com/jonwright/ bslz4decoders. The kernels here are ours, but the idea and the demonstration that it is worth doing are his. Cited in docs/ACKNOWLEDGEMENT.md and in the new section 0 of docs/CPU_DATA_ANALYSIS.md. Two kernels mirror the CPU decoder. LZ4 runs one WARP per bitshuffle block: every lane parses the same sequence stream (a broadcast read, no divergence) and the literal and match copies are split across the 32 lanes so the stores coalesce; an overlapping match is treated as a pattern of period offset sourced from bytes that already precede the write position, which keeps it parallel rather than a serial byte loop. One thread per block instead measured 13x slower. The bitshuffle inverse then un-transposes each byte-plane through shared memory and interleaves the planes back into elements. Only BSHUF_LZ4 is decoded on the device. The zstd variants have no device decoder, and neither has an uncompressed or float image; Supports() returns false for those and the caller decompresses on the host exactly as before. The fallback is explicit, so a format we cannot decode on the device is a slower path and never a wrong answer. Tests hold the device decoder against the CPU one byte for byte, on data from the production compressor, for every element size the detectors emit - including the 8-bit DECTRIS modes, which take bitshuf_decode_block's separate elem_size == 1 branch - plus a many-block frame, the formats it must decline, and malformed containers, which must throw rather than run off a buffer. Battery: 37 crystals, no failures, identical to the host-decode run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1c4dfd03e2 |
v1.0.0-rc.123 (#30)
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 10m22s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 11m30s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 11m41s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 12m32s
Build Packages / Generate python client (push) Successful in 18s
Build Packages / Build documentation (push) Successful in 54s
Build Packages / Create release (push) Has been skipped
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 9m44s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 8m53s
Build Packages / build:rpm (rocky8) (push) Successful in 9m40s
Build Packages / build:rpm (rocky9) (push) Successful in 10m37s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 9m54s
Build Packages / Unit tests (push) Successful in 1h6m33s
This is an UNSTABLE release. * jfjoch_broker: Use newer version of Google Ceres for (potential) CUDA 13 compatibility * jfjoch_broker: Improve performance of generating preview images, especially for large detectors (9M-16M) * jfjoch_viewer: Improve performance of displaying images, especially for large detectors (9M-16M) * jfjoch_viewer: Add more color schemes for better image readability * HDF5: Common mutex for reading and writing HDF5 if both operations were to happen in the same executable * HDF5: suppress warning if path (upstream group) doesn't exists when checking if leaf exists Reviewed-on: #30 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch> Co-committed-by: Filip Leonarski <filip.leonarski@psi.ch> |
||
|
|
06c5b9cf7f | 1.0.0-rc.65 | ||
|
|
28d224afab | version 1.0.0-rc.25 |