Build Packages / build:rpm (rocky9) (push) Successful in 19m56s
Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 16m57s
Build Packages / build:windows:cuda (push) Successful in 19m18s
Build Packages / build:viewer-tgz:cpu (push) Successful in 14m48s
Build Packages / build:viewer-tgz:cuda (push) Successful in 16m18s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 14m19s
Build Packages / build:rugnux:windows (push) Successful in 10m34s
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 8m49s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 20m55s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 17m4s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 20m48s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 19m15s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 24m26s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 20m32s
Build Packages / build:rpm (rocky8) (push) Successful in 23m39s
Build Packages / Generate python client (push) Successful in 46s
Build Packages / Build documentation (push) Successful in 1m45s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (durin plugin) (push) Successful in 11m3s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 11m30s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m10s
Build Packages / XDS test (neggia plugin) (push) Successful in 10m17s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 23m12s
Build Packages / DIALS test (push) Successful in 20m12s
* rugnux now tells you whether a crystal diffracts anisotropically and how far it reaches in each direction, without a second program: a new `9. DIFFRACTION ANISOTROPY` section in `<prefix>_report.txt` and matching `_reflns.pdbx_aniso_B_tensor_*` / `_reflns.jfjoch_aniso_*` items in the merged mmCIF report the anisotropic deltaB, the diffraction limit along each principal direction, and a `NOT DETECTED` / `DETECTED` / `CANNOT DETERMINE` verdict measured against the data set's own systematic error. It is a description only - no intensity is corrected, no reflection is removed, and the merged data do not depend on direction.
* rugnux can hand its integrated observations to another scaling program: `--export-unmerged` writes `<prefix>_unmerged.mtz`, an unmerged MTZ readable by aimless, pointless, careless and `iotbx.merging_statistics`, in `--mode mx` and `--mode scale` alike. Each rotation reflection's partials are summed into one full; `--export-unmerged-partials` writes one row per image instead. Intensities carry the Lorentz-polarization factor and nothing else, since those programs scale the data themselves. Lattice-centring absences are not written; screw and glide absences are.
* rugnux integrates crystals with broad spots better - where it changes anything, per-shell mean I/sigma improves by up to 31% and R_meas by up to 24% - because on rotation data the integration signal radius is now taken from the crystal's own measured spot width instead of a fixed 4 px. `--adaptive-integration-radius=off` restores the fixed radius and an explicit `--integration-radius` still overrides both. The widened radius applies to the final integration pass only, and a pattern too dense for it is re-integrated at 4 px with a note in the log.
* rugnux discards fewer stills reflections for want of a background ring, improving per-shell R_meas over most of the signal-bearing range: the stills background ring now runs to 14 px instead of 12. The gain reverses in shells below a mean I/sigma of about 4.
* rugnux determines the space group with thresholds that mean the same thing on a weak crystal as on a strong one: symmetry operators are scored on resolution-normalised intensities (E squared) instead of raw merged intensities, and a reflection counts as genuinely present on its counting significance instead of on the merged I/sigma, which saturates at the merge's own ISa. The search resolution cut is no longer able to move the answer, and the twin-law H bound moves from 1.70 to 1.85, which stops one class of correct high-symmetry assignment being refused as twinning.
* rugnux says what the space-group search tested and what it could not: the twin-law disagreement H is printed for every operator together with the adopted point group's H ratio and its bound; alternatives that are not on the reported lattice are named with how their cell differs; and a lattice centring the data could not test - the crystal having been integrated on the primitive sub-cell, so the reflections it extinguishes were never measured - is marked `UNTESTED` and warned about where it is adopted, as coming from the lattice metric rather than from the intensities.
* rugnux `--mode scale` re-merges a `_process.h5` in the right symmetry without being told it: the file now records the space group on every run - a two-pass rotation run wrote none before, so re-merging defaulted to P1 - together with the change of basis under `/entry/MX/reindexMatrix` where the lattice was re-seated, and `--mode scale` also reports the Wilson B-factor estimate instead of `WILSON_B= nan`. A file written before this stops with a message naming the two cells and the override to use, instead of failing inside the merge. A third-party reader of a `_process.h5` must apply `reindexMatrix` where it is present.
* rugnux installs on its own, as a package called `rugnux` - `dnf install rugnux` or `apt install rugnux` - instead of arriving inside `jfjoch-viewer`. It pulls in none of the acquisition stack, so a machine that only processes data no longer has to carry the broker, the detector libraries or Qt to get it. Installing it over a `jfjoch-viewer` from rc.163 or earlier, which still owns `/usr/bin/rugnux`, upgrades cleanly rather than failing on the duplicate file.
* rugnux is also a standalone download, built for arm64 as well as x86_64: `rugnux-<version>-linux-{x86_64|aarch64}-cuda<major>.tgz` and `rugnux-<version>-win64-cuda<major>.zip` on the release page, for machines that are not managed by a package manager. The aarch64 build targets GH200 and DGX Spark, and is untested on hardware.
* Every portable Linux binary is now a single self-contained file: cuFFT is linked statically instead of being shipped beside the executable and found through an rpath, so `rugnux` and `jfjoch_viewer` need nothing but an NVIDIA driver, and only to use the GPU. The `.rpm`/`.deb` continue to take cuFFT from the distribution. The developer utilities `jfjoch_extract_hkl` and `jfjoch_recompress` are no longer packaged anywhere.
* Jungfraujoch needs six fewer shared libraries on the machine - libopenblas and libmetis, and libgfortran, libquadmath, libgomp and libz behind them - because the Ceres LAPACK, METIS and SuiteSparse back-ends are no longer built. Nothing in the code ever selected them, and results are unchanged.
* The PCIe driver DKMS package builds for the kernel it is being installed for instead of the running one, so a module built while a kernel update is being applied loads after the reboot.
* The PCIe driver builds on RHEL 9.5 and later, and on their CentOS Stream, Rocky and AlmaLinux equivalents, where the `vm_flags` kernel interface was backported into the 5.14 kernel.
* A data collection started with `async_start` that fails to start - a writer refusing to overwrite an existing file, for instance - is reported as an error by `/wait_until_running` and `/wait_till_done` instead of as a timeout and a successful collection respectively. The error message is the one the writer gave.
* A calibration that is cancelled or that fails to collect its pedestals is no longer reported as a successful one. The broker goes to `Inactive` with an error message and has to be initialized again, instead of sitting in `Idle` looking ready to measure while holding partial pedestals - data collected in that state was silently mis-converted.
* A failed `/initialize` is reported to `/wait_until_running` and `/wait_till_done` as soon as it happens, instead of when their timeout expires.
* `space_group_number` accepts space groups up to 230 in the API schema, so cubic space groups can be recorded. The broker always accepted them; the generated clients rejected them before the request was sent.
* The results report's `REPORT_VERSION` is 3, two sections having been added. Existing key names and table columns are unchanged.
* The merged statistics table has **9** resolution shells instead of 10, which is what XDS reports. The bins were already XDS's - equal steps in 1/d^2 between the lowest- and the highest-resolution reflection the merge kept - so at the same resolution limits the two tables now have the same shell boundaries and can be read row for row. `--resolution-shells` sets a different count.
* `rugnux --model` now settles the frame the merged reflections are written in, not only the frame the R-factors and the maps are computed in: the `.mtz`/`.cif`/`.hkl` come out in the model's indexing, and where the data were merged in the model's enantiomorph they take the model's hand and space group - which on anomalous data puts I(+) and I(-) the right way round. The indexing choice is logged with the winning R-free and the runner-up, so a decision made within noise is visible.
* `rugnux --model` can resolve the indexing ambiguity of a **serial stills** run, which a model could not do before: structure factors computed from the model become the per-image reference, the same role a reference MTZ plays. It needs the cell and space group up front (`-C` / `-S`). Without one or the other, a merohedral serial run still merges both hands together and says so.
* The rugnux documentation opens with a quick start - the default run, and runs with a reference MTZ, with a model, or with the space group and cell pinned - and explains the indexing ambiguity: what it costs on rotation and on serial data, and which of `-z` / `--model` resolves it in each case. The long reference pages now carry a table of contents.
Reviewed-on: #74
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
434 lines
29 KiB
C++
434 lines
29 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <cmath>
|
|
#include <cstdint>
|
|
#include <limits>
|
|
#include <optional>
|
|
#include <vector>
|
|
|
|
#include "../../common/DiffractionExperiment.h"
|
|
#include "../../common/Logger.h"
|
|
#include "../../common/Reflection.h"
|
|
#include "../../common/UnitCell.h"
|
|
#include "../IntegrationOutcome.h"
|
|
|
|
#include "Merge.h" // MergedReflection, MergeStatistics
|
|
#ifdef JFJOCH_USE_CUDA
|
|
#include <memory>
|
|
#include "RotationScaleMergeGPU.h"
|
|
#endif
|
|
|
|
// Dedicated, allocate-once scale+combine+merge for rotation data (the -P rot3d path): recompute the
|
|
// per-frame partiality from the (smoothed) mosaicity, robustly fit a per-image scale G, 3D-combine each
|
|
// rocking event's partials into fulls, refit a per-frame scale on the fulls (XDS order), and merge with
|
|
// a global error model.
|
|
//
|
|
// The per-frame partial observations are ingested ONCE into flat vectors; the hkl->ASU grouping is
|
|
// computed once per space group (by a sort, not a map) and reused across all scaling iterations; every
|
|
// hot step is a flat loop over those vectors, so the whole pipeline maps onto CUDA kernels (segmented
|
|
// reduction + per-frame solve) and runs GPU-resident when a GPU is present, with the CPU loops as the
|
|
// bit-parity fallback. CC1/2 and the per-image CC are computed once at the end, not every iteration.
|
|
//
|
|
// Used only for the self-scaling rotation case with per-image G (Rotation partiality, a fixed/forced
|
|
// mosaicity is honoured by the recompute). Post-scale-fulls correction stages (on by default via
|
|
// ScalingSettings::CorrectionSurfaces): a global Debye-Waller decay and a goniometer-frame absorption
|
|
// surface, both fitted on the host and pushed back to the resident (GPU) fulls before the merge.
|
|
// External-reference scaling, the stills B-factor and wedge refinement are unsupported (caller rejects).
|
|
// Stills use the per-image ScaleOnTheFly (fixed partiality) instead.
|
|
class RotationScaleMerge {
|
|
public:
|
|
struct Result {
|
|
std::vector<MergedReflection> merged;
|
|
MergeStatistics statistics;
|
|
// Two tiers, and they are different quantities. `isa` is the whole-range 1/sqrt(a*b) - which
|
|
// in this parameterisation is 1/b - and is what XDS's ISa means, so it is the one exported.
|
|
// `isa_asymptotic` is the strong-reflection tier, which XDS has no equivalent of and which can
|
|
// only ever be the more optimistic of the two. Both 0 if the model stayed at identity.
|
|
double isa = 0.0;
|
|
double isa_asymptotic = 0.0;
|
|
double error_model_a = 0.0; // XDS convention: sigma^2 = a*(sigma0^2 + b*I^2)
|
|
double error_model_b = 0.0;
|
|
// The overall CC1/2 as it stood BEFORE the correction surfaces were folded in. That is what the
|
|
// two-pass quality guard compares one pass against the other with: a pass whose intensities are
|
|
// discarded fits no surfaces (see full_stats), so judging the pass that does fit them by its
|
|
// corrected CC1/2 would set two different measurements against each other. Equal to
|
|
// statistics.overall.cc_half whenever no surface was fitted, and NaN when the caller did not
|
|
// ask for it (see measure_cc_before_corrections) - it costs a merge, so it is not measured on
|
|
// spec, and a caller that did not ask must not be handed a number that looks measured.
|
|
double cc_half_before_corrections = std::numeric_limits<double>::quiet_NaN();
|
|
};
|
|
|
|
// experiment: read live (its space group is changed by the caller between Run() calls).
|
|
// partial_outcomes: the per-frame partials; the final per-frame scale (G, CC, mosaicity) is written
|
|
// back onto them so the offline per-image scaling table is still exported.
|
|
// reference_cell: the consensus cell (for the completeness count and the cell-consistency mask).
|
|
RotationScaleMerge(const DiffractionExperiment &experiment,
|
|
std::vector<IntegrationOutcome> &partial_outcomes,
|
|
std::optional<UnitCell> reference_cell,
|
|
int scaling_iterations,
|
|
size_t nthreads,
|
|
Logger &logger,
|
|
std::string observation_dump_path = {});
|
|
|
|
// Copy the per-frame partials into the flat buffers. Call once before the first Run().
|
|
void Ingest();
|
|
|
|
// Scale (per-frame G) -> smooth G -> 3D combine -> scale fulls -> merge -> error model -> statistics
|
|
// for the space group currently set on the experiment, reusing the ingested buffers.
|
|
// for_search: the de-novo P1 pass whose merged intensities feed the space-group search - ice-ring
|
|
// reflections are dropped from the merge and the error model (kept otherwise, for completeness).
|
|
// full_stats: these merged intensities are an OUTPUT. False on the rotation two-pass geometry
|
|
// pre-pass, whose merge exists only to choose the space group and post-refine the geometry and
|
|
// whose reflections are never written: the correction surfaces, the report-only diagnostics, the
|
|
// R_meas re-walk, the anomalous split, the R-free flags and the French-Wilson amplitudes are then
|
|
// all skipped, because computing them fills in fields nothing reads. What the pre-pass IS read for
|
|
// - the merged intensities themselves, the error model, and the completeness / CC1/2 the second
|
|
// pass is judged against - is computed either way.
|
|
// measure_cc_before_corrections: also merge ONCE MORE, just before the correction surfaces, and
|
|
// report that merge's overall CC1/2 in the result. Only a caller comparing two passes of the same
|
|
// data needs it (the pre-pass fits no surfaces, so only the uncorrected number is the same
|
|
// measurement on both sides); it is a whole extra merge, so the offline re-scale path, which
|
|
// compares nothing, asks for it to be left out.
|
|
Result Run(bool for_search, bool full_stats, bool measure_cc_before_corrections);
|
|
|
|
// Override the high-resolution cut for the next Run() - used to gate the de-novo P1 search pass at
|
|
// <I/sigma> >= 1 without cutting the final in-symmetry merge. Reset to the manual limit afterwards.
|
|
void SetDMinLimit(std::optional<double> d_min_A) { d_min_limit = d_min_A; }
|
|
|
|
// Toggle the search-pass Lorentz filter (see search_min_zeta) between Run() calls, so the caller can
|
|
// produce both a filtered and an unfiltered search merge from the same ingested partials.
|
|
void SetSearchMinZeta(double zeta) { search_min_zeta = zeta; }
|
|
[[nodiscard]] double GetSearchMinZeta() const { return search_min_zeta; }
|
|
|
|
private:
|
|
// One integrated observation - a per-frame partial during scaling/combine, or a combined full during
|
|
// scale-fulls/merge. Flat (not nested per image); a POD so the arrays translate straight to CUDA.
|
|
struct Obs {
|
|
int32_t h, k, l;
|
|
float I, sigma, d, rlp, partiality, zeta, delta_phi, bkg, var_bkg;
|
|
// Fulls only, written by the combine: the full's variance as a function of intensity,
|
|
// var(I) = var_bkg + var_per_I * I. The merge rebuilds it at the reflection's mean.
|
|
float var_per_I = 0.0f;
|
|
float px = NAN, py = NAN; // predicted detector position (for the absorption surface; CPU path only)
|
|
float image_number; // fractional frame position (for 3D-combine contiguity)
|
|
int32_t frame; // index of the outcome whose per-frame scale G applies to this obs
|
|
uint8_t on_ice;
|
|
float corr; // image_scale_corr (working; updated by scaling)
|
|
int32_t group; // dense ASU-group id for the current space group; <0 = never mergeable
|
|
};
|
|
|
|
// One leverage-corrected error-model sample per usable full: its raw variance, its group's mean
|
|
// intensity, its squared deviation from that mean - and the resolution it sits at, because the fit is
|
|
// re-run over the samples that survive the automatic resolution cutoff. See MergeAndStats.
|
|
struct Sample { double s2, I2, dev2; float d; };
|
|
|
|
// The narrow per-observation record the ingest sort orders: the raw hkl the runs are cut on, the
|
|
// frame position that breaks a tie inside one, the observation's own index (which makes the order
|
|
// total - see the .cpp), and the resolution the range test reads. Twenty-four bytes against the
|
|
// Obs's eighty, and it is all the ingest needs before it knows which observations survive.
|
|
struct SortKey {
|
|
int32_t h, k, l;
|
|
float image_number;
|
|
int32_t idx;
|
|
float d;
|
|
};
|
|
|
|
const DiffractionExperiment &x;
|
|
std::vector<IntegrationOutcome> &partials_out; // written back at the end of scaling
|
|
std::optional<UnitCell> reference_cell;
|
|
size_t nthreads;
|
|
Logger &logger;
|
|
std::string observation_dump_path;
|
|
|
|
// Fixed settings snapshot (read once in the ctor).
|
|
int n_frames = 0;
|
|
double min_partiality = 0.02;
|
|
std::optional<double> d_min_limit;
|
|
std::optional<double> d_max_limit;
|
|
bool merge_friedel = true;
|
|
double capture_uncertainty_coeff = 0.0;
|
|
double min_captured_fraction = 0.0;
|
|
|
|
// Drop a frame's observations entirely when the frame disagrees with the merged reference below this
|
|
// correlation (--min-image-cc). A mis-centred or off-crystal frame still produces spots, still
|
|
// indexes and still integrates - it just measures something that is not the crystal's diffraction,
|
|
// and nothing downstream removes it. 0 = off.
|
|
double min_cc_for_image = 0.0;
|
|
|
|
// Exclude observations with |zeta| below this from the DE-NOVO SEARCH merge only (the final merge
|
|
// keeps everything). zeta is the sine of the angle between a reflection's rocking path and the
|
|
// spindle: near 0 it crosses the Ewald sphere almost tangentially, spends many frames in
|
|
// diffracting position and is measured badly. The symmetry search compares how equal an operator's
|
|
// paired intensities are, so it is answered by whichever reflections are worst measured - and when
|
|
// the spindle lies in a lattice plane, an operator permuting the two in-plane axes samples a
|
|
// different mixture of qualities than one that only flips signs, which is not a fair comparison.
|
|
// Unlike a bound on I/sigma this is pure geometry, identical in meaning on every dataset. 0 = off.
|
|
double search_min_zeta = 0.0;
|
|
double reject_nsigma = 0.0;
|
|
bool reject_outliers = false;
|
|
double rfree_fraction = 0.0;
|
|
int scaling_iter = 3;
|
|
bool scale_fulls = true;
|
|
bool refine_decay_b = false; // per-time-block Debye-Waller decay correction (radiation damage)
|
|
int absorption_iter = 0; // >0: fit a goniometer-frame absorption surface over this many iterations
|
|
int modulation_iter = 0; // >0: fit a detector-plane modulation (flat-field) surface, this many iterations
|
|
double relative_b_deg = 0.0; // >0: fit a per-batch relative-B (batch width in deg); 0 = off
|
|
double mosaicity_deg = 0.1;
|
|
// Automatic high-resolution cutoff for the written reflections + reported shells (post-merge; the
|
|
// scaling, combine and error model always run over the full range). Manual d_min_limit wins.
|
|
ResolutionCutoffMethod resolution_cutoff_method = ResolutionCutoffMethod::Off;
|
|
double resolution_cc_target = 0.30;
|
|
int report_shell_count = 9;
|
|
|
|
// Flat buffers, allocated once by Ingest() and reused across Run() calls.
|
|
std::vector<Obs> partials; // all per-frame partials, grouped by frame
|
|
std::vector<int32_t> frame_start, frame_count; // CSR ranges of `partials` per frame
|
|
std::vector<uint8_t> frame_cell_ok; // per-frame cell-consistency mask (1 = kept)
|
|
std::vector<uint8_t> finite_ok; // per-obs AcceptReflection finiteness (immutable; 1 = kept)
|
|
std::vector<double> g_partial; // per-frame partial scale G (the RESIDUAL after the flux)
|
|
std::vector<double> frame_flux; // per-frame incident flux, run median = 1 (see the .cpp)
|
|
// corr as it stood before the current pass's own filters (--search-min-zeta, --min-image-cc) zeroed
|
|
// observations out of ITS merge. Restored at the start of the next Run, because zeroing corr is
|
|
// permanent otherwise: the only thing that rewrites it is the scaling loop, and that skips any frame
|
|
// it cannot fit. Empty when there is nothing to put back.
|
|
std::vector<float> corr_before_pass_filters;
|
|
|
|
// Raw-hkl ordering, built ONCE by Ingest and reused: `perm` lists partial indices sorted by
|
|
// (raw h,k,l, image_number); each distinct raw hkl is a contiguous run [rawrun_start, +count) of it.
|
|
// The expensive sort happens once here, so per-pass combine (event split) and ASU grouping are linear.
|
|
std::vector<int32_t> perm;
|
|
std::vector<int32_t> rawrun_start, rawrun_count;
|
|
std::vector<int32_t> rawrun_h, rawrun_k, rawrun_l;
|
|
std::vector<float> rawrun_d; // representative resolution per raw hkl
|
|
std::vector<int32_t> rawrun_group; // dense ASU-group id per raw hkl (<0 = absent/out of range)
|
|
|
|
std::vector<Obs> fulls; // combined fulls (rebuilt each Run), sorted by frame
|
|
std::vector<int32_t> fulls_frame_start, fulls_frame_count; // CSR ranges of `fulls` per frame
|
|
std::vector<double> g_full; // per-frame scale on the fulls
|
|
|
|
// The five fields the host-side walks over the fulls actually read, pulled out of the 80-byte record
|
|
// once per merge (see MergeAndStats). `group` carries the usability decision: -1 means the full is
|
|
// not in this merge, which is what a negative group already meant. Members rather than locals for the
|
|
// same reason as FullsStaging below - the merge runs several times per run and this is tens of
|
|
// megabytes each time.
|
|
struct MergeFields {
|
|
std::vector<int32_t> group;
|
|
std::vector<float> I, sigma, corr, d;
|
|
void Resize(int n) { group.resize(n); I.resize(n); sigma.resize(n); corr.resize(n); d.resize(n); }
|
|
};
|
|
MergeFields merge_fields;
|
|
|
|
// One host array per field for the fulls download: the device hands back an array per field and the
|
|
// host gathers them into `fulls`. Members rather than locals in Run() because the whole
|
|
// scale->combine->merge chain runs several times per run and these are a few hundred megabytes
|
|
// between them, so as locals every chain allocates, faults in and zeroes the lot again.
|
|
struct FullsStaging {
|
|
std::vector<int32_t> h, k, l, frame, group;
|
|
std::vector<float> I, sigma, d, image_number, corr, px, py, var_bkg, var_per_I;
|
|
std::vector<uint8_t> on_ice;
|
|
void Resize(int n) {
|
|
h.resize(n); k.resize(n); l.resize(n); frame.resize(n); group.resize(n);
|
|
I.resize(n); sigma.resize(n); d.resize(n); image_number.resize(n); corr.resize(n);
|
|
px.resize(n); py.resize(n); var_bkg.resize(n); var_per_I.resize(n);
|
|
on_ice.resize(n);
|
|
}
|
|
};
|
|
FullsStaging fulls_staging;
|
|
|
|
// The merge accumulators (see MergeAndStats' run_merge): one entry per ASU group, plus the arrays
|
|
// the device kernel fills that the host unpacks into them. Members for the same reason as
|
|
// FullsStaging - a merge runs several times per run and this is a hundred megabytes between them.
|
|
struct Accum { double swI = 0, sw = 0, swIh[2] = {0, 0}, swh[2] = {0, 0}; size_t nh[2] = {0, 0}; float d = NAN; };
|
|
std::vector<Accum> merge_acc;
|
|
struct MergeAccumStaging {
|
|
std::vector<double> swI, sw, swIh0, swIh1, swh0, swh1, d;
|
|
std::vector<int32_t> nh0, nh1, rej;
|
|
void Resize(int n) {
|
|
swI.resize(n); sw.resize(n); swIh0.resize(n); swIh1.resize(n); swh0.resize(n); swh1.resize(n);
|
|
d.resize(n); nh0.resize(n); nh1.resize(n); rej.resize(n);
|
|
}
|
|
};
|
|
MergeAccumStaging merge_accum;
|
|
|
|
// Per-group scatter for the strong-reflection ISa asymptote (see MergeAndStats). A member for the
|
|
// same reason: 32 bytes a group, once per merge.
|
|
struct GroupScatter { double sum = 0, sum_sq = 0, sum_var = 0; int n = 0; };
|
|
std::vector<GroupScatter> asymptote_scatter;
|
|
|
|
// The error model's working pools (see MergeAndStats): the samples themselves, the scratch copy each
|
|
// fit partitions, the misfit-free subset the refit uses, the subset inside the resolution cutoff, and
|
|
// the per-sample chi2 the reported number is the median of. Members for the same reason as FullsStaging - one is 32 bytes per full and
|
|
// there are two dozen fits per run, so as locals this is gigabytes of pages faulted in and handed
|
|
// straight back. Every one of them is cleared and refilled before it is read.
|
|
std::vector<Sample> em_samples, em_fit_pool, em_refit_pool, em_cut_pool;
|
|
std::vector<double> em_chi2;
|
|
|
|
// Set by FitPerFrameG: which frames were fitted this call (so corr/G is updated only there).
|
|
std::vector<uint8_t> frame_scaled_scratch;
|
|
|
|
// Per-frame mosaicity smoothed in frame order (deterministic); used to recompute partiality and
|
|
// written back for the per-image scaling table. Empty if there is no per-frame mosaicity.
|
|
std::vector<float> mos_smooth;
|
|
|
|
// Radiation-damage monitor (measured by MeasureRadiationDamageB on the scaled fulls before any decay
|
|
// correction; report-only, copied into the result statistics by MergeAndStats). NaN / empty until set.
|
|
double rad_damage_delta_b = std::numeric_limits<double>::quiet_NaN(); // relative-B first->last (A^2)
|
|
std::vector<float> rad_damage_b_batch; // per-batch relative-B curve (A^2)
|
|
double rad_damage_batch_deg = 0.0; // rotation width per batch (deg)
|
|
|
|
// Sweep-quality diagnostic (MeasureSweepQuality; report-only, copied into the result statistics by
|
|
// MergeAndStats). Empty and not measured until it runs.
|
|
SweepQuality sweep_quality;
|
|
|
|
// Working per-group arrays (sized to the current group count; reused).
|
|
std::vector<int32_t> group_h, group_k, group_l;
|
|
|
|
#ifdef JFJOCH_USE_CUDA
|
|
// GPU engine: the whole hot path (scaling, combine, scale-fulls, per-frame CC, smooth-G, merge +
|
|
// error model) runs on the device, resident, when a GPU is present. Null / inactive otherwise, with
|
|
// the CPU loops as the bit-parity fallback. Built in Ingest.
|
|
std::unique_ptr<RotationScaleMergeGPU> gpu_;
|
|
bool gpu_active_ = false;
|
|
#endif
|
|
|
|
// --- helpers (each a flat pass; see the .cpp) ---
|
|
// Turn the per-frame mean background under the reflections (accumulated by the ingest fill loop) into
|
|
// the per-frame incident flux, which the finiteness pass then folds into rlp so that
|
|
// corr = rlp / (partiality * G) divides it out and G fits only the residual. See the .cpp for why a
|
|
// background is a usable flux meter and what it costs when it is not.
|
|
void MeasureIncidentFlux(const std::vector<double> &mean_bkg);
|
|
|
|
// Build the flat `partials` array (and the per-frame CSR, the finiteness mask and `perm`) from the
|
|
// source reflections, skipping the observations whose resolution can never be in range:
|
|
// --scaling-high/low-resolution are the coarsest limits any Run() uses (the space-group search only
|
|
// ever RAISES d_min), and an out-of-range raw hkl gets group -1 in every pass, which keeps it out of
|
|
// the scaling reference, the per-frame fit, the combine, the merge and the error model alike. On a
|
|
// crystal that integrates to the detector corner and merges well short of it that is most of the
|
|
// array, and an eighty-byte record built for it is eighty bytes written and then thrown away. Whole
|
|
// raw-hkl RUNS are skipped, on the same per-hkl resolution ComputeAsuGroups tests, so what survives -
|
|
// and the order of every sum formed over it - is exactly what it would have been had the whole array
|
|
// been built and then filtered. Everything is built when no manual limit was given.
|
|
void BuildInRangeObservations(const std::vector<SortKey> &keys);
|
|
|
|
// Compute the dense ASU-group id for the current space group by grouping the (pre-sorted) raw-hkl
|
|
// runs by their ASU key - one gemmi ASU reduction per distinct raw hkl, not per observation. Fills
|
|
// rawrun_group, the group_h/k/l representative tables, and partials[].group; returns the group count.
|
|
int ComputeAsuGroups(const HKLKeyGenerator &key_generator);
|
|
|
|
// Inverse-variance per-group mean of I*corr over `obs` (the merge reference).
|
|
void ReduceGroupMeans(const std::vector<Obs> &obs, int n_groups, std::vector<double> &out_mean) const;
|
|
|
|
// Robust per-frame G fit (IRLS, Cauchy k=3), unity=false uses the rotation partiality, unity=true the
|
|
// scale-fulls (partiality already folded in). Reads out_mean[group] as the reference intensity.
|
|
void FitPerFrameG(std::vector<Obs> &obs, const std::vector<int32_t> &fstart,
|
|
const std::vector<int32_t> &fcount, const std::vector<double> &group_mean_in,
|
|
bool unity, std::vector<double> &g);
|
|
|
|
// corr = rlp / (partiality * G[frame]); leaves corr unchanged for frames that could not be fit.
|
|
void UpdateCorr(std::vector<Obs> &obs, const std::vector<double> &g,
|
|
const std::vector<uint8_t> &frame_scaled) const;
|
|
|
|
void SmoothG(std::vector<Obs> &obs, std::vector<double> &g, int window) const;
|
|
|
|
// The windowed geometric mean of G over frames (the shared first half of SmoothG); the GPU path
|
|
// applies the resulting ratio to the resident corr in a kernel instead of the host obs loop.
|
|
void ComputeSmoothGWindow(const std::vector<double> &g, int window,
|
|
std::vector<double> &g_smooth) const;
|
|
|
|
// Drop the observations of any frame whose fitted per-frame scale collapsed far below the run
|
|
// median, reporting the per-frame corr factor the caller has to apply. See the .cpp for why nothing
|
|
// downstream can catch a collapsed scale on its own.
|
|
bool DropCollapsedScales(const std::vector<uint8_t> &fitted_mask, std::vector<double> &g,
|
|
std::vector<uint8_t> &apply, std::vector<double> &ratio) const;
|
|
|
|
// Smooth per-frame mosaicity in frame order and recompute each partial's partiality from it, so the
|
|
// per-frame partials of one rocking event tile the curve consistently (they sum toward 1) before the
|
|
// 3D combine. Deterministic (frame order); replaces the old arrival-order mosaicity moving average
|
|
// that prediction applied. SG-independent, so done once in Ingest.
|
|
void SmoothGeometry();
|
|
void SmoothMosaicityAndPartiality();
|
|
|
|
void Combine(); // partials -> fulls (CPU)
|
|
|
|
// Drop the fulls of any frame whose scale collapsed toward zero. The fulls are scaled with the Unity
|
|
// model, so their corr IS 1/G and a collapsed G multiplies every intensity on that frame without
|
|
// bound. Covers the CPU and GPU scaling paths alike; `from_staging` says the fulls' frame and corr
|
|
// are still in fulls_staging, which is where the scan reads them from when they are. Returns true if
|
|
// anything was dropped (the caller then has to push the corrected corr back to the device).
|
|
bool DropCollapsedFullScales(bool from_staging);
|
|
|
|
// Post-scale-fulls correction surfaces, each an alternating multiplicative fit of the host fulls' corr
|
|
// against the merged reference (cheap host loops; the corrected corr is re-uploaded to the resident
|
|
// fulls afterwards). Each is cross-validated (fit even frames, keep only if held-out odd equivalents
|
|
// improve) so it is a no-op when its systematic is absent. RefineDecay fits a global Debye-Waller B
|
|
// (resolution x time - radiation damage the resolution-flat per-frame G cannot capture; also gated on a
|
|
// physical total-dB floor). RefineAbsorption fits a smooth factor over the diffracted-beam direction in
|
|
// the goniometer frame (path-length / absorption; negligible at hard X-rays, matters at low energy).
|
|
void RefineDecay(int n_groups);
|
|
// Solve a smooth per-batch relative-B from the per-batch normal equations for b (num_c, den_c):
|
|
// data-fidelity + a second-difference (curvature) penalty, by Gauss-Seidel, each batch clamped to
|
|
// +-b_max. Returns the un-anchored curve; the caller sets the gauge and the clamp it can live with.
|
|
// Shared by the correction and the radiation-damage monitor.
|
|
std::vector<double> SolveCurvatureSmoothedB(const std::vector<double> &num,
|
|
const std::vector<double> &den, double b_max) const;
|
|
// Fit a smoothed per-batch relative-B curve (A^2 per batch) on the fulls over the ASU-group subset
|
|
// {group&1==gparity} (gparity<0 = all): the weighted s^2 slope of ln(Iref/Iobs) per batch against a
|
|
// subset-global reference, smoothed and zero-mean-anchored. Drives the per-batch correction.
|
|
std::vector<double> FitRelativeBCurve(int n_groups, int n_batch, int frames_per_batch, int gparity) const;
|
|
// Radiation-damage MONITOR (report-only): measure the per-batch relative-B on the scaled fulls before
|
|
// any decay correction and store the first->last relative-B change + the per-batch curve on this object
|
|
// (copied into the result statistics by MergeAndStats, then printed / logged / written to the mmCIF).
|
|
void MeasureRadiationDamageB(int n_groups);
|
|
// Sweep-quality diagnostic (report-only): find the contiguous stretches of the sweep over which the
|
|
// crystal delivered much less than the rest of the run, and say what each one looks like. Reads the
|
|
// per-frame scale (with the incident flux already divided out) and the per-frame CC to merge.
|
|
void MeasureSweepQuality(const std::vector<uint8_t> &partial_scaled, const std::vector<double> &cc,
|
|
const std::vector<int64_t> &cc_n);
|
|
// Per-batch relative-B, applied after RefineDecay: the single decay slope removes the average
|
|
// radiation-damage falloff, but the relative scattering power drifts NON-monotonically across a run
|
|
// (absorption path, crystal slippage, dose bursts). Refine one relative Debye-Waller B per batch
|
|
// (FitRelativeBCurve), anchored to zero mean (the constant part is a global Wilson-B, degenerate with
|
|
// overall scale). Guarded by a physical peak-to-peak floor and cross-validated by ASU-GROUP parity (a
|
|
// per-batch parameter cannot be scored on a held-out FRAME the batch owns; splitting the equivalents
|
|
// tests whether a batch's B generalises to reflections it was not fit on). Opt-in (--relative-b).
|
|
void RefineRelativeB(int n_groups);
|
|
void RefineAbsorption(int n_iter, int n_groups);
|
|
// Time-dependent absorption: the same cross-validated surface, indexed by (rotation, detector position)
|
|
// instead of by the crystal-frame direction alone. RefineAbsorption's parameterisation is the whole
|
|
// model only while the illuminated volume stays put; once the crystal drifts through the beam the exit
|
|
// path becomes a function of the spindle angle too, and nothing time-independent reaches it.
|
|
void RefineAbsorptionTime(int n_iter, int n_groups);
|
|
// Detector-plane modulation (flat-field): the same cross-validated surface fit as absorption, but the
|
|
// cell is the predicted detector position (px, py) instead of the goniometer-frame direction. Corrects
|
|
// detector-response / geometric systematics that vary with where a reflection lands; because it lives
|
|
// in the detector frame (not tied to the rotation) the same correction concept applies to stills.
|
|
void RefineModulation(int n_iter, int n_groups);
|
|
// Shared engine for the correction surfaces: given a per-full cell assignment (cell[i] in [0,ncell), or
|
|
// <0 to skip), fit a Tikhonov-regularised multiplicative factor per cell against the merged reference,
|
|
// cross-validate on even/odd frames, and fold it into corr only if the held-out equivalents improve.
|
|
void ApplyCellSurface(const std::vector<int32_t> &cell, int ncell, int n_iter, int n_groups,
|
|
const char *name);
|
|
|
|
// Sort `fulls` by peak frame and (re)build fulls_frame_start/count (the per-frame CSR the scale-fulls
|
|
// step slices). Shared by the CPU Combine tail and the GPU combine path.
|
|
void SortFullsByFrame();
|
|
|
|
// Per-frame CC vs the partial merge reference (CPU; the GPU equivalent is gpu_->ComputePartialCC).
|
|
void ComputePerFrameCC(const std::vector<double> &partial_group_mean,
|
|
std::vector<double> &cc, std::vector<int64_t> &cc_n) const;
|
|
|
|
// Write G/CC/mosaicity back onto the partials (once, at the end of partial scaling) from the given
|
|
// per-frame cc/cc_n, so the offline per-image scaling table is still exported.
|
|
void FinalizePerFrameScale(const std::vector<double> &cc, const std::vector<int64_t> &cc_n,
|
|
const std::vector<uint8_t> &frame_scaled);
|
|
|
|
// Error model + merge + statistics over the fulls (the last stage). n_groups is the fulls group count.
|
|
// fulls_resident: the (scaled) fulls + their group CSR are still on the GPU, so the em-stats / samples
|
|
// / merge-accumulate / R_meas reductions run there (only per-group + samples come back).
|
|
// full_stats: see Run().
|
|
Result MergeAndStats(int n_groups, bool for_search, bool fulls_resident, bool full_stats);
|
|
};
|