Files
Jungfraujoch/image_analysis/scale_merge/RotationScaleMerge.h
T
jungfrauandClaude Opus 5 b380da2a91 Judge the refined pass against the pre-pass on the same footing
The two-pass guard rolls back to the header geometry when the refined pass looks worse than the
pre-pass, and one of its three tests is a drop of more than 0.05 in CC1/2. Since the pre-pass stopped
fitting its correction surfaces - it exists to pick a space group and post-refine the geometry, and
its intensities are discarded - the two sides of that test were no longer measuring the same thing:
the pre-pass's CC1/2 came out uncorrected and the refined pass's corrected. On one crystal here that
flattered the refined pass by 0.008, and it is the wrong direction to be careless in, because it
makes the guard slower to fire on a pass that really is bad.

So measure the refined pass's CC1/2 before its surfaces are applied as well, and compare that. It
cannot be had from the half-set accumulate alone, which was the cheap thing to hope for: CC1/2
correlates half-set means built on the error model's sigmas over the reflections the automatic
resolution cutoff kept, so an accumulate on its own is a different quantity - and one biased low,
which would make the guard fire too eagerly. It takes the same merge the pre-pass now does, without
the statistics tail, before the surfaces run.

That is one extra merge on the one pass that has surfaces, so a caller asks for it explicitly rather
than paying for it by default: the two-pass driver does, --mode scale does not, because there is no
other pass to compare against. Where the caller wants it but nothing was corrected anyway, the merge
that already ran IS the uncorrected one and is reported as such; where nobody asked, the field stays
absent rather than being filled in from the statistic that reads almost the same and is not.

Also make the final in-symmetry merge unconditional, with P1 when no group was determined. The
comment there has always said P1 stands in that case and the condition did the opposite, which
would have written a search merge - zeta-filtered, ice-excluded, uncorrected - as the result. It
turns out to be unreachable: with any reflections at all the search returns a group, because no
symmetry leaves the identity point group whose representative is P1, a group with no screws or
centering leaves a symmorphic candidate that has no absences to contradict and so is always
eligible, and an empty merge throws in both engines before the search sees it. The two lines keep
that promise here instead of resting on eligibility gates in another file that a later change could
tighten without noticing what leaned on them; the reasoning is written at the site.

Battery unchanged on all 24 crystals - same space group, reflection count and R_meas as the run
before it - and the merged output is byte-identical on three crystals spanning the regimes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU
2026-08-23 14:53:39 -04:00

387 lines
26 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <cstdint>
#include <limits>
#include <optional>
#include <vector>
#include "../../common/DiffractionExperiment.h"
#include "../../common/Logger.h"
#include "../../common/Reflection.h"
#include "../../common/UnitCell.h"
#include "../IntegrationOutcome.h"
#include "Merge.h" // MergedReflection, MergeStatistics
#ifdef JFJOCH_USE_CUDA
#include <memory>
#include "RotationScaleMergeGPU.h"
#endif
// Dedicated, allocate-once scale+combine+merge for rotation data (the -P rot3d path): recompute the
// per-frame partiality from the (smoothed) mosaicity, robustly fit a per-image scale G, 3D-combine each
// rocking event's partials into fulls, refit a per-frame scale on the fulls (XDS order), and merge with
// a global error model.
//
// The per-frame partial observations are ingested ONCE into flat vectors; the hkl->ASU grouping is
// computed once per space group (by a sort, not a map) and reused across all scaling iterations; every
// hot step is a flat loop over those vectors, so the whole pipeline maps onto CUDA kernels (segmented
// reduction + per-frame solve) and runs GPU-resident when a GPU is present, with the CPU loops as the
// bit-parity fallback. CC1/2 and the per-image CC are computed once at the end, not every iteration.
//
// Used only for the self-scaling rotation case with per-image G (Rotation partiality, a fixed/forced
// mosaicity is honoured by the recompute). Post-scale-fulls correction stages (on by default via
// ScalingSettings::CorrectionSurfaces): a global Debye-Waller decay and a goniometer-frame absorption
// surface, both fitted on the host and pushed back to the resident (GPU) fulls before the merge.
// External-reference scaling, the stills B-factor and wedge refinement are unsupported (caller rejects).
// Stills use the per-image ScaleOnTheFly (fixed partiality) instead.
class RotationScaleMerge {
public:
struct Result {
std::vector<MergedReflection> merged;
MergeStatistics statistics;
// Two tiers, and they are different quantities. `isa` is the whole-range 1/sqrt(a*b) - which
// in this parameterisation is 1/b - and is what XDS's ISa means, so it is the one exported.
// `isa_asymptotic` is the strong-reflection tier, which XDS has no equivalent of and which can
// only ever be the more optimistic of the two. Both 0 if the model stayed at identity.
double isa = 0.0;
double isa_asymptotic = 0.0;
double error_model_a = 0.0; // XDS convention: sigma^2 = a*(sigma0^2 + b*I^2)
double error_model_b = 0.0;
// The overall CC1/2 as it stood BEFORE the correction surfaces were folded in. That is what the
// two-pass quality guard compares one pass against the other with: a pass whose intensities are
// discarded fits no surfaces (see full_stats), so judging the pass that does fit them by its
// corrected CC1/2 would set two different measurements against each other. Equal to
// statistics.overall.cc_half whenever no surface was fitted, and NaN when the caller did not
// ask for it (see measure_cc_before_corrections) - it costs a merge, so it is not measured on
// spec, and a caller that did not ask must not be handed a number that looks measured.
double cc_half_before_corrections = std::numeric_limits<double>::quiet_NaN();
};
// experiment: read live (its space group is changed by the caller between Run() calls).
// partial_outcomes: the per-frame partials; the final per-frame scale (G, CC, mosaicity) is written
// back onto them so the offline per-image scaling table is still exported.
// reference_cell: the consensus cell (for the completeness count and the cell-consistency mask).
RotationScaleMerge(const DiffractionExperiment &experiment,
std::vector<IntegrationOutcome> &partial_outcomes,
std::optional<UnitCell> reference_cell,
int scaling_iterations,
size_t nthreads,
Logger &logger,
std::string observation_dump_path = {});
// Copy the per-frame partials into the flat buffers. Call once before the first Run().
void Ingest();
// Scale (per-frame G) -> smooth G -> 3D combine -> scale fulls -> merge -> error model -> statistics
// for the space group currently set on the experiment, reusing the ingested buffers.
// for_search: the de-novo P1 pass whose merged intensities feed the space-group search - ice-ring
// reflections are dropped from the merge and the error model (kept otherwise, for completeness).
// full_stats: these merged intensities are an OUTPUT. False on the rotation two-pass geometry
// pre-pass, whose merge exists only to choose the space group and post-refine the geometry and
// whose reflections are never written: the correction surfaces, the report-only diagnostics, the
// R_meas re-walk, the anomalous split, the R-free flags and the French-Wilson amplitudes are then
// all skipped, because computing them fills in fields nothing reads. What the pre-pass IS read for
// - the merged intensities themselves, the error model, and the completeness / CC1/2 the second
// pass is judged against - is computed either way.
// measure_cc_before_corrections: also merge ONCE MORE, just before the correction surfaces, and
// report that merge's overall CC1/2 in the result. Only a caller comparing two passes of the same
// data needs it (the pre-pass fits no surfaces, so only the uncorrected number is the same
// measurement on both sides); it is a whole extra merge, so the offline re-scale path, which
// compares nothing, asks for it to be left out.
Result Run(bool for_search, bool full_stats, bool measure_cc_before_corrections);
// Override the high-resolution cut for the next Run() - used to gate the de-novo P1 search pass at
// <I/sigma> >= 1 without cutting the final in-symmetry merge. Reset to the manual limit afterwards.
void SetDMinLimit(std::optional<double> d_min_A) { d_min_limit = d_min_A; }
// Toggle the search-pass Lorentz filter (see search_min_zeta) between Run() calls, so the caller can
// produce both a filtered and an unfiltered search merge from the same ingested partials.
void SetSearchMinZeta(double zeta) { search_min_zeta = zeta; }
[[nodiscard]] double GetSearchMinZeta() const { return search_min_zeta; }
private:
// One integrated observation - a per-frame partial during scaling/combine, or a combined full during
// scale-fulls/merge. Flat (not nested per image); a POD so the arrays translate straight to CUDA.
struct Obs {
int32_t h, k, l;
float I, sigma, d, rlp, partiality, zeta, delta_phi, bkg, var_bkg;
// Fulls only, written by the combine: the full's variance as a function of intensity,
// var(I) = var_bkg + var_per_I * I. The merge rebuilds it at the reflection's mean.
float var_per_I = 0.0f;
float px = NAN, py = NAN; // predicted detector position (for the absorption surface; CPU path only)
float image_number; // fractional frame position (for 3D-combine contiguity)
int32_t frame; // index of the outcome whose per-frame scale G applies to this obs
uint8_t on_ice;
float corr; // image_scale_corr (working; updated by scaling)
int32_t group; // dense ASU-group id for the current space group; <0 = never mergeable
};
// The narrow per-observation record the ingest sort orders: the raw hkl the runs are cut on, the
// frame position that breaks a tie inside one, the observation's own index (which makes the order
// total - see the .cpp), and the resolution the range test reads. Twenty-four bytes against the
// Obs's eighty, and it is all the ingest needs before it knows which observations survive.
struct SortKey {
int32_t h, k, l;
float image_number;
int32_t idx;
float d;
};
const DiffractionExperiment &x;
std::vector<IntegrationOutcome> &partials_out; // written back at the end of scaling
std::optional<UnitCell> reference_cell;
size_t nthreads;
Logger &logger;
std::string observation_dump_path;
// Fixed settings snapshot (read once in the ctor).
int n_frames = 0;
double min_partiality = 0.02;
std::optional<double> d_min_limit;
std::optional<double> d_max_limit;
bool merge_friedel = true;
double capture_uncertainty_coeff = 0.0;
double min_captured_fraction = 0.0;
// Drop a frame's observations entirely when the frame disagrees with the merged reference below this
// correlation (--min-image-cc). A mis-centred or off-crystal frame still produces spots, still
// indexes and still integrates - it just measures something that is not the crystal's diffraction,
// and nothing downstream removes it. 0 = off.
double min_cc_for_image = 0.0;
// Exclude observations with |zeta| below this from the DE-NOVO SEARCH merge only (the final merge
// keeps everything). zeta is the sine of the angle between a reflection's rocking path and the
// spindle: near 0 it crosses the Ewald sphere almost tangentially, spends many frames in
// diffracting position and is measured badly. The symmetry search compares how equal an operator's
// paired intensities are, so it is answered by whichever reflections are worst measured - and when
// the spindle lies in a lattice plane, an operator permuting the two in-plane axes samples a
// different mixture of qualities than one that only flips signs, which is not a fair comparison.
// Unlike a bound on I/sigma this is pure geometry, identical in meaning on every dataset. 0 = off.
double search_min_zeta = 0.0;
double reject_nsigma = 0.0;
bool reject_outliers = false;
double rfree_fraction = 0.0;
int scaling_iter = 3;
bool scale_fulls = true;
bool refine_decay_b = false; // per-time-block Debye-Waller decay correction (radiation damage)
int absorption_iter = 0; // >0: fit a goniometer-frame absorption surface over this many iterations
int modulation_iter = 0; // >0: fit a detector-plane modulation (flat-field) surface, this many iterations
double relative_b_deg = 0.0; // >0: fit a per-batch relative-B (batch width in deg); 0 = off
double mosaicity_deg = 0.1;
// Automatic high-resolution cutoff for the written reflections + reported shells (post-merge; the
// scaling, combine and error model always run over the full range). Manual d_min_limit wins.
ResolutionCutoffMethod resolution_cutoff_method = ResolutionCutoffMethod::Off;
double resolution_cc_target = 0.30;
int report_shell_count = 10;
// Flat buffers, allocated once by Ingest() and reused across Run() calls.
std::vector<Obs> partials; // all per-frame partials, grouped by frame
std::vector<int32_t> frame_start, frame_count; // CSR ranges of `partials` per frame
std::vector<uint8_t> frame_cell_ok; // per-frame cell-consistency mask (1 = kept)
std::vector<uint8_t> finite_ok; // per-obs AcceptReflection finiteness (immutable; 1 = kept)
std::vector<double> g_partial; // per-frame partial scale G (the RESIDUAL after the flux)
std::vector<double> frame_flux; // per-frame incident flux, run median = 1 (see the .cpp)
// corr as it stood before the current pass's own filters (--search-min-zeta, --min-image-cc) zeroed
// observations out of ITS merge. Restored at the start of the next Run, because zeroing corr is
// permanent otherwise: the only thing that rewrites it is the scaling loop, and that skips any frame
// it cannot fit. Empty when there is nothing to put back.
std::vector<float> corr_before_pass_filters;
// Raw-hkl ordering, built ONCE by Ingest and reused: `perm` lists partial indices sorted by
// (raw h,k,l, image_number); each distinct raw hkl is a contiguous run [rawrun_start, +count) of it.
// The expensive sort happens once here, so per-pass combine (event split) and ASU grouping are linear.
std::vector<int32_t> perm;
std::vector<int32_t> rawrun_start, rawrun_count;
std::vector<int32_t> rawrun_h, rawrun_k, rawrun_l;
std::vector<float> rawrun_d; // representative resolution per raw hkl
std::vector<int32_t> rawrun_group; // dense ASU-group id per raw hkl (<0 = absent/out of range)
std::vector<Obs> fulls; // combined fulls (rebuilt each Run), sorted by frame
std::vector<int32_t> fulls_frame_start, fulls_frame_count; // CSR ranges of `fulls` per frame
std::vector<double> g_full; // per-frame scale on the fulls
// Set by FitPerFrameG: which frames were fitted this call (so corr/G is updated only there).
std::vector<uint8_t> frame_scaled_scratch;
// Per-frame mosaicity smoothed in frame order (deterministic); used to recompute partiality and
// written back for the per-image scaling table. Empty if there is no per-frame mosaicity.
std::vector<float> mos_smooth;
// Radiation-damage monitor (measured by MeasureRadiationDamageB on the scaled fulls before any decay
// correction; report-only, copied into the result statistics by MergeAndStats). NaN / empty until set.
double rad_damage_delta_b = std::numeric_limits<double>::quiet_NaN(); // relative-B first->last (A^2)
std::vector<float> rad_damage_b_batch; // per-batch relative-B curve (A^2)
double rad_damage_batch_deg = 0.0; // rotation width per batch (deg)
// Sweep-quality diagnostic (MeasureSweepQuality; report-only, copied into the result statistics by
// MergeAndStats). Empty and not measured until it runs.
SweepQuality sweep_quality;
// Working per-group arrays (sized to the current group count; reused).
std::vector<int32_t> group_h, group_k, group_l;
#ifdef JFJOCH_USE_CUDA
// GPU engine: the whole hot path (scaling, combine, scale-fulls, per-frame CC, smooth-G, merge +
// error model) runs on the device, resident, when a GPU is present. Null / inactive otherwise, with
// the CPU loops as the bit-parity fallback. Built in Ingest.
std::unique_ptr<RotationScaleMergeGPU> gpu_;
bool gpu_active_ = false;
// One host array per field for the fulls download: the device hands back an array per field and the
// host gathers them into `fulls`. Members rather than locals in Run() because the whole
// scale->combine->merge chain runs several times per run and these are a few hundred megabytes
// between them, so as locals every chain allocates, faults in and zeroes the lot again.
struct FullsStaging {
std::vector<int32_t> h, k, l, frame, group;
std::vector<float> I, sigma, d, image_number, corr, px, py, var_bkg, var_per_I;
std::vector<uint8_t> on_ice;
void Resize(int n) {
h.resize(n); k.resize(n); l.resize(n); frame.resize(n); group.resize(n);
I.resize(n); sigma.resize(n); d.resize(n); image_number.resize(n); corr.resize(n);
px.resize(n); py.resize(n); var_bkg.resize(n); var_per_I.resize(n);
on_ice.resize(n);
}
};
FullsStaging fulls_staging;
#endif
// --- helpers (each a flat pass; see the .cpp) ---
// Turn the per-frame mean background under the reflections (accumulated by the ingest fill loop) into
// the per-frame incident flux, which the finiteness pass then folds into rlp so that
// corr = rlp / (partiality * G) divides it out and G fits only the residual. See the .cpp for why a
// background is a usable flux meter and what it costs when it is not.
void MeasureIncidentFlux(const std::vector<double> &mean_bkg);
// Build the flat `partials` array (and the per-frame CSR, the finiteness mask and `perm`) from the
// source reflections, skipping the observations whose resolution can never be in range:
// --scaling-high/low-resolution are the coarsest limits any Run() uses (the space-group search only
// ever RAISES d_min), and an out-of-range raw hkl gets group -1 in every pass, which keeps it out of
// the scaling reference, the per-frame fit, the combine, the merge and the error model alike. On a
// crystal that integrates to the detector corner and merges well short of it that is most of the
// array, and an eighty-byte record built for it is eighty bytes written and then thrown away. Whole
// raw-hkl RUNS are skipped, on the same per-hkl resolution ComputeAsuGroups tests, so what survives -
// and the order of every sum formed over it - is exactly what it would have been had the whole array
// been built and then filtered. Everything is built when no manual limit was given.
void BuildInRangeObservations(const std::vector<SortKey> &keys);
// Compute the dense ASU-group id for the current space group by grouping the (pre-sorted) raw-hkl
// runs by their ASU key - one gemmi ASU reduction per distinct raw hkl, not per observation. Fills
// rawrun_group, the group_h/k/l representative tables, and partials[].group; returns the group count.
int ComputeAsuGroups(const HKLKeyGenerator &key_generator);
// Inverse-variance per-group mean of I*corr over `obs` (the merge reference).
void ReduceGroupMeans(const std::vector<Obs> &obs, int n_groups, std::vector<double> &out_mean) const;
// Robust per-frame G fit (IRLS, Cauchy k=3), unity=false uses the rotation partiality, unity=true the
// scale-fulls (partiality already folded in). Reads out_mean[group] as the reference intensity.
void FitPerFrameG(std::vector<Obs> &obs, const std::vector<int32_t> &fstart,
const std::vector<int32_t> &fcount, const std::vector<double> &group_mean_in,
bool unity, std::vector<double> &g);
// corr = rlp / (partiality * G[frame]); leaves corr unchanged for frames that could not be fit.
void UpdateCorr(std::vector<Obs> &obs, const std::vector<double> &g,
const std::vector<uint8_t> &frame_scaled) const;
void SmoothG(std::vector<Obs> &obs, std::vector<double> &g, int window) const;
// The windowed geometric mean of G over frames (the shared first half of SmoothG); the GPU path
// applies the resulting ratio to the resident corr in a kernel instead of the host obs loop.
void ComputeSmoothGWindow(const std::vector<double> &g, int window,
std::vector<double> &g_smooth) const;
// Drop the observations of any frame whose fitted per-frame scale collapsed far below the run
// median, reporting the per-frame corr factor the caller has to apply. See the .cpp for why nothing
// downstream can catch a collapsed scale on its own.
bool DropCollapsedScales(const std::vector<uint8_t> &fitted_mask, std::vector<double> &g,
std::vector<uint8_t> &apply, std::vector<double> &ratio) const;
// Smooth per-frame mosaicity in frame order and recompute each partial's partiality from it, so the
// per-frame partials of one rocking event tile the curve consistently (they sum toward 1) before the
// 3D combine. Deterministic (frame order); replaces the old arrival-order mosaicity moving average
// that prediction applied. SG-independent, so done once in Ingest.
void SmoothGeometry();
void SmoothMosaicityAndPartiality();
void Combine(); // partials -> fulls (CPU)
// Drop the fulls of any frame whose scale collapsed toward zero. The fulls are scaled with the Unity
// model, so their corr IS 1/G and a collapsed G multiplies every intensity on that frame without
// bound. Reads the host fulls, so it covers the CPU and GPU scaling paths alike. Returns true if
// anything was dropped (the caller then has to push the corrected corr back to the device).
bool DropCollapsedFullScales();
// Post-scale-fulls correction surfaces, each an alternating multiplicative fit of the host fulls' corr
// against the merged reference (cheap host loops; the corrected corr is re-uploaded to the resident
// fulls afterwards). Each is cross-validated (fit even frames, keep only if held-out odd equivalents
// improve) so it is a no-op when its systematic is absent. RefineDecay fits a global Debye-Waller B
// (resolution x time - radiation damage the resolution-flat per-frame G cannot capture; also gated on a
// physical total-dB floor). RefineAbsorption fits a smooth factor over the diffracted-beam direction in
// the goniometer frame (path-length / absorption; negligible at hard X-rays, matters at low energy).
void RefineDecay(int n_groups);
// Solve a smooth per-batch relative-B from the per-batch normal equations for b (num_c, den_c):
// data-fidelity + a second-difference (curvature) penalty, by Gauss-Seidel, each batch clamped to
// +-b_max. Returns the un-anchored curve; the caller sets the gauge and the clamp it can live with.
// Shared by the correction and the radiation-damage monitor.
std::vector<double> SolveCurvatureSmoothedB(const std::vector<double> &num,
const std::vector<double> &den, double b_max) const;
// Fit a smoothed per-batch relative-B curve (A^2 per batch) on the fulls over the ASU-group subset
// {group&1==gparity} (gparity<0 = all): the weighted s^2 slope of ln(Iref/Iobs) per batch against a
// subset-global reference, smoothed and zero-mean-anchored. Drives the per-batch correction.
std::vector<double> FitRelativeBCurve(int n_groups, int n_batch, int frames_per_batch, int gparity) const;
// Radiation-damage MONITOR (report-only): measure the per-batch relative-B on the scaled fulls before
// any decay correction and store the first->last relative-B change + the per-batch curve on this object
// (copied into the result statistics by MergeAndStats, then printed / logged / written to the mmCIF).
void MeasureRadiationDamageB(int n_groups);
// Sweep-quality diagnostic (report-only): find the contiguous stretches of the sweep over which the
// crystal delivered much less than the rest of the run, and say what each one looks like. Reads the
// per-frame scale (with the incident flux already divided out) and the per-frame CC to merge.
void MeasureSweepQuality(const std::vector<uint8_t> &partial_scaled, const std::vector<double> &cc,
const std::vector<int64_t> &cc_n);
// Per-batch relative-B, applied after RefineDecay: the single decay slope removes the average
// radiation-damage falloff, but the relative scattering power drifts NON-monotonically across a run
// (absorption path, crystal slippage, dose bursts). Refine one relative Debye-Waller B per batch
// (FitRelativeBCurve), anchored to zero mean (the constant part is a global Wilson-B, degenerate with
// overall scale). Guarded by a physical peak-to-peak floor and cross-validated by ASU-GROUP parity (a
// per-batch parameter cannot be scored on a held-out FRAME the batch owns; splitting the equivalents
// tests whether a batch's B generalises to reflections it was not fit on). Opt-in (--relative-b).
void RefineRelativeB(int n_groups);
void RefineAbsorption(int n_iter, int n_groups);
// Time-dependent absorption: the same cross-validated surface, indexed by (rotation, detector position)
// instead of by the crystal-frame direction alone. RefineAbsorption's parameterisation is the whole
// model only while the illuminated volume stays put; once the crystal drifts through the beam the exit
// path becomes a function of the spindle angle too, and nothing time-independent reaches it.
void RefineAbsorptionTime(int n_iter, int n_groups);
// Detector-plane modulation (flat-field): the same cross-validated surface fit as absorption, but the
// cell is the predicted detector position (px, py) instead of the goniometer-frame direction. Corrects
// detector-response / geometric systematics that vary with where a reflection lands; because it lives
// in the detector frame (not tied to the rotation) the same correction concept applies to stills.
void RefineModulation(int n_iter, int n_groups);
// Shared engine for the correction surfaces: given a per-full cell assignment (cell[i] in [0,ncell), or
// <0 to skip), fit a Tikhonov-regularised multiplicative factor per cell against the merged reference,
// cross-validate on even/odd frames, and fold it into corr only if the held-out equivalents improve.
void ApplyCellSurface(const std::vector<int32_t> &cell, int ncell, int n_iter, int n_groups,
const char *name);
// Sort `fulls` by peak frame and (re)build fulls_frame_start/count (the per-frame CSR the scale-fulls
// step slices). Shared by the CPU Combine tail and the GPU combine path.
void SortFullsByFrame();
// Per-frame CC vs the partial merge reference (CPU; the GPU equivalent is gpu_->ComputePartialCC).
void ComputePerFrameCC(const std::vector<double> &partial_group_mean,
std::vector<double> &cc, std::vector<int64_t> &cc_n) const;
// Write G/CC/mosaicity back onto the partials (once, at the end of partial scaling) from the given
// per-frame cc/cc_n, so the offline per-image scaling table is still exported.
void FinalizePerFrameScale(const std::vector<double> &cc, const std::vector<int64_t> &cc_n,
const std::vector<uint8_t> &frame_scaled);
// Error model + merge + statistics over the fulls (the last stage). n_groups is the fulls group count.
// fulls_resident: the (scaled) fulls + their group CSR are still on the GPU, so the em-stats / samples
// / merge-accumulate / R_meas reductions run there (only per-group + samples come back).
// full_stats: see Run().
Result MergeAndStats(int n_groups, bool for_search, bool fulls_resident, bool full_stats);
};