Files
Jungfraujoch/image_analysis/scale_merge/Merge.h
T
leonarski_f 749db470ca
Build Packages / build:rpm (rocky9) (push) Successful in 19m56s
Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 16m57s
Build Packages / build:windows:cuda (push) Successful in 19m18s
Build Packages / build:viewer-tgz:cpu (push) Successful in 14m48s
Build Packages / build:viewer-tgz:cuda (push) Successful in 16m18s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 14m19s
Build Packages / build:rugnux:windows (push) Successful in 10m34s
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 8m49s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 20m55s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 17m4s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 20m48s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 19m15s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 24m26s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 20m32s
Build Packages / build:rpm (rocky8) (push) Successful in 23m39s
Build Packages / Generate python client (push) Successful in 46s
Build Packages / Build documentation (push) Successful in 1m45s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (durin plugin) (push) Successful in 11m3s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 11m30s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m10s
Build Packages / XDS test (neggia plugin) (push) Successful in 10m17s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 23m12s
Build Packages / DIALS test (push) Successful in 20m12s
v1.0.0-rc.164 (#74)
* rugnux now tells you whether a crystal diffracts anisotropically and how far it reaches in each direction, without a second program: a new `9. DIFFRACTION ANISOTROPY` section in `<prefix>_report.txt` and matching `_reflns.pdbx_aniso_B_tensor_*` / `_reflns.jfjoch_aniso_*` items in the merged mmCIF report the anisotropic deltaB, the diffraction limit along each principal direction, and a `NOT DETECTED` / `DETECTED` / `CANNOT DETERMINE` verdict measured against the data set's own systematic error. It is a description only - no intensity is corrected, no reflection is removed, and the merged data do not depend on direction.
* rugnux can hand its integrated observations to another scaling program: `--export-unmerged` writes `<prefix>_unmerged.mtz`, an unmerged MTZ readable by aimless, pointless, careless and `iotbx.merging_statistics`, in `--mode mx` and `--mode scale` alike. Each rotation reflection's partials are summed into one full; `--export-unmerged-partials` writes one row per image instead. Intensities carry the Lorentz-polarization factor and nothing else, since those programs scale the data themselves. Lattice-centring absences are not written; screw and glide absences are.
* rugnux integrates crystals with broad spots better - where it changes anything, per-shell mean I/sigma improves by up to 31% and R_meas by up to 24% - because on rotation data the integration signal radius is now taken from the crystal's own measured spot width instead of a fixed 4 px. `--adaptive-integration-radius=off` restores the fixed radius and an explicit `--integration-radius` still overrides both. The widened radius applies to the final integration pass only, and a pattern too dense for it is re-integrated at 4 px with a note in the log.
* rugnux discards fewer stills reflections for want of a background ring, improving per-shell R_meas over most of the signal-bearing range: the stills background ring now runs to 14 px instead of 12. The gain reverses in shells below a mean I/sigma of about 4.
* rugnux determines the space group with thresholds that mean the same thing on a weak crystal as on a strong one: symmetry operators are scored on resolution-normalised intensities (E squared) instead of raw merged intensities, and a reflection counts as genuinely present on its counting significance instead of on the merged I/sigma, which saturates at the merge's own ISa. The search resolution cut is no longer able to move the answer, and the twin-law H bound moves from 1.70 to 1.85, which stops one class of correct high-symmetry assignment being refused as twinning.
* rugnux says what the space-group search tested and what it could not: the twin-law disagreement H is printed for every operator together with the adopted point group's H ratio and its bound; alternatives that are not on the reported lattice are named with how their cell differs; and a lattice centring the data could not test - the crystal having been integrated on the primitive sub-cell, so the reflections it extinguishes were never measured - is marked `UNTESTED` and warned about where it is adopted, as coming from the lattice metric rather than from the intensities.
* rugnux `--mode scale` re-merges a `_process.h5` in the right symmetry without being told it: the file now records the space group on every run - a two-pass rotation run wrote none before, so re-merging defaulted to P1 - together with the change of basis under `/entry/MX/reindexMatrix` where the lattice was re-seated, and `--mode scale` also reports the Wilson B-factor estimate instead of `WILSON_B= nan`. A file written before this stops with a message naming the two cells and the override to use, instead of failing inside the merge. A third-party reader of a `_process.h5` must apply `reindexMatrix` where it is present.
* rugnux installs on its own, as a package called `rugnux` - `dnf install rugnux` or `apt install rugnux` - instead of arriving inside `jfjoch-viewer`. It pulls in none of the acquisition stack, so a machine that only processes data no longer has to carry the broker, the detector libraries or Qt to get it. Installing it over a `jfjoch-viewer` from rc.163 or earlier, which still owns `/usr/bin/rugnux`, upgrades cleanly rather than failing on the duplicate file.
* rugnux is also a standalone download, built for arm64 as well as x86_64: `rugnux-<version>-linux-{x86_64|aarch64}-cuda<major>.tgz` and `rugnux-<version>-win64-cuda<major>.zip` on the release page, for machines that are not managed by a package manager. The aarch64 build targets GH200 and DGX Spark, and is untested on hardware.
* Every portable Linux binary is now a single self-contained file: cuFFT is linked statically instead of being shipped beside the executable and found through an rpath, so `rugnux` and `jfjoch_viewer` need nothing but an NVIDIA driver, and only to use the GPU. The `.rpm`/`.deb` continue to take cuFFT from the distribution. The developer utilities `jfjoch_extract_hkl` and `jfjoch_recompress` are no longer packaged anywhere.
* Jungfraujoch needs six fewer shared libraries on the machine - libopenblas and libmetis, and libgfortran, libquadmath, libgomp and libz behind them - because the Ceres LAPACK, METIS and SuiteSparse back-ends are no longer built. Nothing in the code ever selected them, and results are unchanged.
* The PCIe driver DKMS package builds for the kernel it is being installed for instead of the running one, so a module built while a kernel update is being applied loads after the reboot.
* The PCIe driver builds on RHEL 9.5 and later, and on their CentOS Stream, Rocky and AlmaLinux equivalents, where the `vm_flags` kernel interface was backported into the 5.14 kernel.
* A data collection started with `async_start` that fails to start - a writer refusing to overwrite an existing file, for instance - is reported as an error by `/wait_until_running` and `/wait_till_done` instead of as a timeout and a successful collection respectively. The error message is the one the writer gave.
* A calibration that is cancelled or that fails to collect its pedestals is no longer reported as a successful one. The broker goes to `Inactive` with an error message and has to be initialized again, instead of sitting in `Idle` looking ready to measure while holding partial pedestals - data collected in that state was silently mis-converted.
* A failed `/initialize` is reported to `/wait_until_running` and `/wait_till_done` as soon as it happens, instead of when their timeout expires.
* `space_group_number` accepts space groups up to 230 in the API schema, so cubic space groups can be recorded. The broker always accepted them; the generated clients rejected them before the request was sent.
* The results report's `REPORT_VERSION` is 3, two sections having been added. Existing key names and table columns are unchanged.
* The merged statistics table has **9** resolution shells instead of 10, which is what XDS reports. The bins were already XDS's - equal steps in 1/d^2 between the lowest- and the highest-resolution reflection the merge kept - so at the same resolution limits the two tables now have the same shell boundaries and can be read row for row. `--resolution-shells` sets a different count.
* `rugnux --model` now settles the frame the merged reflections are written in, not only the frame the R-factors and the maps are computed in: the `.mtz`/`.cif`/`.hkl` come out in the model's indexing, and where the data were merged in the model's enantiomorph they take the model's hand and space group - which on anomalous data puts I(+) and I(-) the right way round. The indexing choice is logged with the winning R-free and the runner-up, so a decision made within noise is visible.
* `rugnux --model` can resolve the indexing ambiguity of a **serial stills** run, which a model could not do before: structure factors computed from the model become the per-image reference, the same role a reference MTZ plays. It needs the cell and space group up front (`-C` / `-S`). Without one or the other, a merohedral serial run still merges both hands together and says so.
* The rugnux documentation opens with a quick start - the default run, and runs with a reference MTZ, with a model, or with the space group and cell pinned - and explains the indexing ambiguity: what it costs on rotation and on serial data, and which of `-z` / `--model` resolves it in each case. The long reference pages now carry a table of contents.

Reviewed-on: #74
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-08-26 22:47:00 +02:00

255 lines
13 KiB
C++

// SPDX-FileCopyrightText: 2025 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <algorithm>
#include <cmath>
#include <map>
#include <unordered_map>
#include <vector>
#include "../../common/Logger.h"
#include "../../common/DiffractionExperiment.h"
#include "../../common/Reflection.h"
#include "../IntegrationOutcome.h"
#include "AnisotropyAnalysis.h"
#include "HKLKey.h"
struct MergeStatisticsShell {
float d_min = 0.0f;
float d_max = 0.0f;
float mean_one_over_d2 = 0;
int total_observations = 0;
int unique_reflections = 0;
int possible_unique_reflections = 0;
double mean_i_over_sigma = 0.0;
double cc_half = 0.0f;
double cc_ref = NAN;
// Redundancy-independent merging R-factor (Diederichs & Karplus 1997), computed over the
// observations that enter the merge: R_meas = sum_hkl sqrt(n/(n-1)) sum_i|I_i-<I>| / sum I_i.
double r_meas = NAN;
// Anomalous signal-to-noise (XDS "SigAno" / mmCIF pdbx_absDiff_over_sigma_anomalous):
// <|I(+)-I(-)|> / <sigma(I(+)-I(-))> over acentric reflections measured in both hands. NaN when
// there is no anomalous split (e.g. Friedel-merged with no mates, or the stills path).
double abs_diff_over_sigma_anomalous = NAN;
};
// Why a stretch of the sweep came out much weaker than the rest of the run. Report-only: nothing is
// excluded on the strength of it. The codes are the field's own words - "crystal rotating out of the
// beam" (HKL-2000 manual), "loss of centring during crystal rotation" (autoPROC).
enum class SweepQualityReason {
NoDiffraction, // the range recorded essentially no diffraction from the indexed lattice
CrystalOutOfBeam, // frames were lost: the range gets a per-image scale far less often than the run
WeakDiffraction, // the frames all still index, but with much less intensity - cause not determined
LossOfCentring, // one cycle of modulation per revolution: the crystal is off the rotation axis
RadiationDamage // the range runs to the end of a sweep whose quality was already decaying
};
const char *SweepQualityReasonCode(SweepQualityReason reason); // machine-readable, e.g. "out_of_beam"
const char *SweepQualityReasonText(SweepQualityReason reason); // for a sentence, e.g. "out of beam"
struct SweepQualityRange {
// Inclusive, in processed-image ordinals - the numbering of <prefix>_image.dat and of every other
// per-image array rugnux writes. With -s/--stride the source image is start + ordinal * stride.
int first_image = 0;
int last_image = 0;
SweepQualityReason reason = SweepQualityReason::WeakDiffraction;
// Fraction of the run's typical diffracting power missing over the range: 0 = as good as the run,
// 1 = nothing at all. The rest are supporting numbers, relative to the run unless stated.
float severity = 0.0f;
float rotation_deg = 0.0f; // width of the range
float mean_relative_scale = 1.0f; // <per-image scale> / run median
float mean_relative_cc = 1.0f; // <per-image CC to merge> / run median
float indexed_fraction = 1.0f; // frames in the range that got a per-image scale at all
float relative_b = NAN; // mean of the radiation-damage monitor's per-batch B over the range
};
// Sweep-quality diagnostic (rotation only). `measured` separates "the run is clean" from "this did not
// run": the range list is empty in the first case and in the second alike.
struct SweepQuality {
bool measured = false;
float sweep_deg = 0.0f;
float flux_peak_to_trough = 1.0f; // the incident-flux proxy, over the whole run
float modulation_peak_to_trough = 1.0f; // depth of a DIAGNOSED once-per-revolution modulation (1 = none)
std::vector<SweepQualityRange> ranges;
};
struct MergeStatistics {
std::vector<MergeStatisticsShell> shells;
MergeStatisticsShell overall;
// Dataset-wide isotropic Wilson B-factor estimate (A^2) from the log-linear fit of the shell-mean
// merged intensity against 1/d^2 (CalcGlobalWilsonB) - the analogue of XDS's "WILSON LINE ... B=".
// Diagnostic only; not used in scaling. NaN when not determined.
double wilson_b = NAN;
double wilson_b_correlation = NAN;
// Radiation-damage monitor (rotation only): the relative Debye-Waller B change measured from the first
// to the last frame of the run (A^2; positive = high-resolution intensity fades with dose = damage) and
// the per-batch relative-B curve it was derived from. Measured before any decay/relative-B correction is
// applied, so it reports how much radiation damage was present. Diagnostic; NaN / empty for stills or
// when not determined. batch_deg is the rotation width per batch of the curve. A batch the data cannot
// measure is NaN in the curve, and delta_b is NaN when the curve is not a trend a single number
// summarises - damage is progressive, so a curve that is not is telling the user about something else.
double radiation_damage_delta_b = NAN;
std::vector<float> radiation_damage_b_batch;
double radiation_damage_batch_deg = 0.0;
// Stretches of the sweep over which the crystal delivered much less than the rest of the run
// (MeasureSweepQuality). Report-only - no observation is dropped because of it.
SweepQuality sweep_quality;
// Diffraction anisotropy (AnalyzeAnisotropy): the anisotropy tensor, the diffraction limit along
// each of its principal directions, and whether either is established above this dataset's own
// systematic error. Report-only - nothing is corrected and no reflection is removed. Empty
// (n_reflections = 0) when the diagnostic did not run.
AnisotropyResult anisotropy;
};
std::ostream &operator<<(std::ostream &output, const MergeStatisticsShell &in);
std::ostream &operator<<(std::ostream &output, const MergeStatistics &in);
struct MergeAccum {
int32_t h = 0;
int32_t k = 0;
int32_t l = 0;
float d = NAN;
double sum_wI = 0.0;
double sum_w = 0.0;
double sum_wI_half[2] = {0.0, 0.0};
double sum_w_half[2] = {0.0, 0.0};
size_t n_half[2] = {0, 0};
};
// XDS's error-model convention (Diederichs, Acta Cryst. D66 (2010) 733) is sigma^2 = a*(sigma0^2 +
// b*I^2), reported with ISa = 1/sqrt(a*b). Jungfraujoch fits sigma^2 = a*sigma0^2 + (b*<I>)^2 - the
// same `a`, but a `b` that is a FRACTION of the intensity - so b_xds = b^2/a, and the two ISa
// expressions are the same number: 1/sqrt(a * b^2/a) = 1/b. Only `b` needs converting, and only
// where it is reported: doing it here rather than in the fit leaves every merge weight untouched.
//
// NOTE this is NOT the `b` of SearchSpaceGroup's merge_systematic_b, which is a third, unrelated
// quantity (a fraction of I, fitted with no `a` at all) whose gate constants are calibrated in that
// convention. Do not "make them consistent".
struct XdsErrorModel {
double a = 1.0;
double b = 0.0;
double isa = 0.0;
};
inline XdsErrorModel ToXdsErrorModel(double a, double b) {
if (!(a > 0.0) || !(b > 0.0))
return {a, 0.0, 0.0};
return {a, b * b / a, 1.0 / b};
}
class MergeOnTheFly {
mutable std::mutex merged_mutex;
const int space_group_number = 1;
ScalingSettings scaling_settings;
IndexingSettings indexing_settings;
std::optional<UnitCell> reference_cell;
std::optional<double> high_resolution_limit;
std::optional<double> low_resolution_limit;
std::optional<double> image_cc_limit;
// Apply image_cc_limit in Mask(). One flag for the whole engine, not a per-call argument, so the
// merge, the error model and MergeStats can never disagree about which images are in.
bool filter_by_image_cc = false;
double min_partiality = 0.02;
// When set, ice-ring-flagged reflections are left out of this merge. Used for the P1 pass whose
// merged intensities feed the space-group search and the error model - those model fits must not
// see the ice-contaminated intensities. The final in-symmetry merge keeps them (for completeness).
bool exclude_ice_rings = false;
HKLKeyGenerator generator;
std::map<uint64_t, MergeAccum> accumulator;
// Global error model (XDS form): sigma_corr^2 = a*sigma^2 + (b*<I>)^2. a rescales the
// (under-estimated) counting variance; the (b*<I>)^2 term adds the intensity-
// proportional systematic error that counting statistics miss, so strong reflections
// are no longer over-weighted. ISa = 1/b is the asymptotic I/sigma. Refined from the
// scatter of symmetry equivalents (RefineErrorModel); identity until then.
bool error_model_active = false;
double error_model_a = 1.0;
double error_model_b = 0.0;
double error_model_chi2 = 0.0; // achieved median reduced chi^2 (~1.0 = honestly calibrated sigmas)
// The (b*I)^2 term uses the reflection's *mean* intensity (constant over its
// observations), so it inflates sigma without biasing the inverse-variance weights -
// using the per-observation I_i instead would over-weight down-fluctuated points.
std::unordered_map<uint64_t, float> error_model_mean_I;
[[nodiscard]] float CorrectedSigma(float I_corr, float sigma_corr, float image_scale_corr,
float var_bkg,
uint64_t hkl_key) const;
// Optional per-observation outlier rejection: drop observations whose corrected
// intensity lies more than reject_nsigma error-model sigmas from the reflection's
// *median* (a robust centre). The error-model sigma already captures the genuine
// (e.g. partiality) scatter, so this removes only the tail beyond it - zingers,
// overlaps, mis-indexed frames - not good partials. Populated by RefineErrorModel.
bool reject_outliers = false;
double reject_nsigma = 6.0;
std::unordered_map<uint64_t, float> reject_median_I;
size_t reject_count = 0;
bool Mask(const IntegrationOutcome &outcome);
public:
MergeOnTheFly(const DiffractionExperiment &x);
MergeOnTheFly& ReferenceCell(const std::optional<UnitCell> &cell);
MergeOnTheFly& ExcludeIceRings(bool input) { exclude_ice_rings = input; return *this; }
MergeOnTheFly& FilterByImageCC(bool input) { filter_by_image_cc = input; return *this; }
// Fit the global error model from the spread of symmetry-equivalent observations.
// Call once before merging; AddImage then applies it.
void RefineErrorModel(const std::vector<IntegrationOutcome> &outcomes);
[[nodiscard]] bool ErrorModelActive() const { return error_model_active; }
[[nodiscard]] double ErrorModelA() const { return error_model_a; }
[[nodiscard]] double ErrorModelB() const { return error_model_b; }
[[nodiscard]] double ErrorModelChi2() const { return error_model_chi2; }
// Outlier rejection (driven by ScalingSettings::GetOutlierRejectNsigma) reports its count.
[[nodiscard]] size_t RejectedCount() const { return reject_count; }
// image_id is the image's stable identity (its index in the outcomes vector). The CC1/2 half-set
// is a deterministic hash of it, so the split is reproducible run-to-run and independent of the
// order (or threading) of AddImage calls - not a draw from a shared RNG in call order.
void AddImage(const IntegrationOutcome& outcome, int64_t image_id);
// d_min_override, when set, is the effective high-resolution limit for the shell table (used for
// the automatic resolution cutoff computed by the caller); otherwise the manual
// ScalingSettings high-resolution limit stands. The number of shells is ScalingSettings::ReportShellCount.
MergeStatistics MergeStats(const std::vector<MergedReflection> &merged,
const std::vector<IntegrationOutcome> &reflections,
const std::vector<MergedReflection> &reference = {},
std::optional<double> d_min_override = std::nullopt);
std::vector<MergedReflection> ExportReflections();
};
std::vector<MergedReflection> MergeAll(const DiffractionExperiment &x,
const std::vector<IntegrationOutcome> &reflections);
// Pearson CC between one image's corrected intensities (I * image_scale_corr) and a reference set of
// full intensities, over the reflections that would enter the merge (non-ice, within the resolution
// limit, partiality above the floor, finite). {NAN, n} when fewer than 20 reflections qualify.
// This is the per-image image_scale_cc: ScaleOnTheFly sets it, and StillsPartialityRefine recomputes it
// after refining the partiality model, so the reported CC always describes the corrections that will be
// merged - which matters because --min-image-cc drops images by it.
std::pair<double, size_t> ImageReferenceCC(const std::vector<Reflection> &reflections,
const std::map<HKLKey, double> &reference,
const HKLKeyGenerator &generator,
std::optional<double> d_min_limit,
std::optional<double> d_max_limit,
double min_partiality);