Files
Jungfraujoch/image_analysis/scale_merge/SearchSpaceGroup.h
T
leonarski_f 6dfe065365
Build Packages / Create release (push) Successful in 16s
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 8m27s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 9m15s
Build Packages / build:viewer-tgz:cpu (push) Successful in 10m11s
Build Packages / build:viewer-tgz:cuda (push) Successful in 12m6s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 15m44s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 16m1s
Build Packages / build:windows:nocuda (push) Successful in 17m29s
Build Packages / build:windows:cuda (push) Successful in 19m58s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 24m7s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 19m8s
Build Packages / build:rugnux:windows (push) Successful in 10m58s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 20m46s
Build Packages / Generate python client (push) Successful in 53s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 20m13s
Build Packages / Build documentation (push) Successful in 1m36s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 19m57s
Build Packages / build:rpm (rocky8) (push) Successful in 18m7s
Build Packages / build:rpm (rocky9) (push) Successful in 18m54s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 19m32s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 17m30s
Build Packages / Unit tests (push) Successful in 1h39m2s
v1.0.0-rc.172 (#82)
* Fixed `jfjoch_broker` cancelling every data collection with a CUDA "out of memory" error after long operation: GPU memory no longer leaks with each collection.
* Rugnux scales a rotation sweep until the per-frame scales settle instead of for a fixed three rounds, and says so when they did not - merged intensities, and the space group, resolution cut and frame rejection read off them, change accordingly; `--scaling-iterations` is now the cap on that loop (default 100).
* Rugnux places every frame of a marCCD, SMV or miniCBF series at the spindle angle its own header states, so a series with missing frames, or with angles written modulo 360, is no longer read at the wrong geometry or refused.
* Every rotation run writes two diagnostic files beside its reflections: `<prefix>_detector.jpg`, the detector projection with the pixel mask and the detected beam-stop shadow drawn on it, and `<prefix>_plot.txt`, one row per image.

Reviewed-on: #82
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-09-22 06:48:37 +02:00

874 lines
66 KiB
C++

// SPDX-FileCopyrightText: 2025 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <array>
#include <limits>
#include <optional>
#include <string>
#include <vector>
#include "../../common/Reflection.h"
#include "gemmi/symmetry.hpp"
#include "gemmi/unitcell.hpp"
// Determine the likely space group of a dataset from its P1-merged intensities, in the spirit
// of POINTLESS (Evans 2006):
//
// Stage A - point group (Laue) symmetry. Every candidate rotation operator is scored once by
// the correlation of E^2(h) with E^2(Rh), on resolution-normalised intensities as
// POINTLESS does. The chosen point group is the largest one all
// of whose operators are confirmed (high CC). A wrong operator scores ~0, so this is
// self-pruning - the unit cell is not needed.
//
// Stage B - space group within that point group. Each Sohncke space group of the point group
// predicts a set of systematically absent reflections (centering + screw axes). The
// chosen space group is the one that explains the MOST absences while every reflection
// it predicts absent is in fact weak. The symmorphic group (no absences) is the
// fallback when no screw/centering is supported. Enantiomorphic pairs (e.g. P4_1 vs
// P4_3) are indistinguishable from intensities and are reported as a pair.
struct SpaceGroupOperatorScore {
std::string op_triplet_hkl; // reciprocal-space triplet of the rotation, e.g. "-h,-k,l"
double cc = 0.0; // correlation of E^2(h) with E^2(Rh) - resolution-normalised
int n_pairs = 0; // independent reflection pairs the CC was computed from
bool present = false; // operator confirmed as a real symmetry of the intensities
// MEDIAN |I1-I2|/(I1+I2) over this operator's pairs - the disagreement the operator implies, with no
// sigma in it. Unlike chi^2 and the systematic-b, which are ratios to a merge error model that
// drifts with multiplicity, this is a property of the intensities alone; compared against an
// operator already confirmed on the same reflections it is what separates a real symmetry from a
// twin law (see max_operator_h_ratio). The median rather than the mean because a merohedral twin
// perturbs EVERY pair while a badly integrated minority perturbs only the tail.
double h_stat = 0.0;
// INTENSITY-WEIGHTED disagreement across this operator's pairs: sum|I1-I2| / sum(I1+I2), an
// R-factor form. Unlike h_stat (a median) it is dominated by the STRONG reflections, which is where
// a false operator relating unequal intensities shows the largest contrast and where a genuine one
// agrees best - the same reflections a median under-weights. Sigma-free, so it does not saturate
// with the merge's ISa the way the chi^2 and systematic-b gates do on a starved search merge. Used
// as the full-resolution merge-degradation gate (see max_operator_r_over_floor).
double r_stat = 0.0;
};
struct SpaceGroupCandidateScore {
gemmi::SpaceGroup space_group;
int absent_observed = 0; // observed reflections this SG predicts systematically absent
int absent_violations = 0; // of those, how many are nonetheless strongly present
int centering_absent = 0; // of absent_observed, the ones extinguished by the CENTERING
int screw_absent = 0; // ... and the ones extinguished by a SCREW (i.e. on axial rows)
// How much likelier the screw-absent class is if the screw exists than if it does not, in nats -
// the quantity that decides whether a screw may be claimed at all (see min_screw_absence_evidence,
// and ScrewZoneEvidence for the statistic). 0 when there are no screw absences, or when no axial
// row carried a control class to judge them against.
double screw_absence_evidence = 0.0;
// How unlikely the CENTERING-absent class would be without the centering, judged against the
// present class as its control (see AbsenceEvidence). 0 for a primitive group, and for a centred
// one whose class this merge never sampled. This is what ranks the candidates, in place of a net
// absence COUNT.
double centering_absence_evidence = 0.0;
// One entry per axial ROW (zone) the group's screws extinguish, each scored on its own evidence.
// A zone whose row carries no control reflections in this merge cannot be judged: it ABSTAINS,
// neither confirming the screw nor refusing it, rather than being folded into one pooled verdict
// where an unmeasured row can outweigh a confirmed one.
struct ScrewZone {
std::array<int, 3> row{}; // the row's direction, e.g. {0,1,0} for 0k0
int n_absent = 0;
int n_control = 0; // 0 = undetermined: this row has no control class in the merge
double evidence = 0.0; // nats; meaningless when n_control is 0
};
std::vector<ScrewZone> screw_zones;
// Of the three axial rows, how many this group's screws extinguish while the merge holds no
// control class to judge them - claims the data can neither confirm nor refuse. Zero unless at
// least one screw zone WAS judged, so the tie-break it feeds (the sort in SearchSpaceGroup) is
// only ever spent by a crystal that has already shown a screw; centring-driven axial conditions
// are not counted.
int unjudged_screw_claims = 0;
// A GLIDE plane extinguishes a two-dimensional ZONE (h0l with l odd for a c-glide normal to b),
// where a screw extinguishes a one-dimensional row. The zone is keyed by the mirror direction -
// by the ROTATION part of the improper operator, so a glide and its centring partner (the c and
// the n of C2/c, same plane, different translation) are ONE zone and not two, which would score
// the same reflections twice.
//
// Every member here is zero for a Sohncke group, and by construction rather than by measurement:
// a Sohncke group has no improper operator at all, so the loop that fills these never has a zone
// to fill. That is what makes glide detection inert on chiral (protein) data.
struct GlideZone {
std::string label; // the zone as a crystallographer names it: "h0l", "hk0", "hhl", ...
int n_absent = 0;
int n_control = 0; // the rest of the same zone; 0 = the zone is not measurable here
double evidence = 0.0; // nats, ScrewZoneEvidence on the zone's own control mean
// evidence / n_absent. THIS is what a glide is judged on, not `evidence`. The statistic is
// linear in the number of absences at fixed deadness, and a zone is a PLANE - hundreds to
// thousands of reflections where a screw row holds tens - so a whole-zone sum reaches
// hundreds of nats on a class that is merely a few times weak. Measured over 140 protein
// datasets: the largest false zone SUM is 331 nats, against 422 for the smallest genuine
// glide in the corpus - a factor of 1.3, no separation. Per reflection the same two are
// 0.65 and 5.95 - a factor of 9. The rate is the scale the two populations separate on.
double evidence_per_reflection = 0.0;
// The absent class in units of its zone's control mean - what the evidence is computed from,
// and the number a crystallographer reads: 0.005 means the extinguished class sits at half a
// percent of the rest of its own plane. Reported because the evidence is a log and a report
// that only prints nats cannot be sanity-checked against the diffraction pattern.
double mean_u = 0.0;
};
std::vector<GlideZone> glide_zones;
int glide_absent = 0; // observed reflections extinguished by a glide (not a screw)
// The WEAKEST zone's evidence per reflection - the minimum, not the maximum a screw takes. A
// screw's zones are separate conditions and the best one may speak for the group; a candidate
// that claims two glides claims BOTH planes are dead, so the one that is not dead refuses it.
// The safe direction: a false glide is an over-call, and over-calls are the dangerous ones.
double glide_absence_evidence = 0.0;
// A zone this merge cannot judge - too few absences, or no control class of its own. For a screw
// an unjudgeable zone ABSTAINS; for a glide it REFUSES. A screw is being weighed against its own
// absence in a group that otherwise fits; a glide is an extra claim on top of a group that
// already fits without it, and an extra claim with no evidence behind it is not a tie, it is a no.
bool glide_unmeasurable = false;
double absent_mean_i_over_sigma = 0.0;
double present_mean_i_over_sigma = 0.0;
bool consistent = false; // absent class confirmed weak (few violations)
bool selected = false; // chosen result (or its enantiomorph)
// The centering could not be TESTED: the group is centred, but this merge holds not one
// reflection of the class its centering extinguishes. That happens when the data were indexed
// and integrated on the primitive sub-cell rather than on the conventional centred one - the
// interstitial nodes were never predicted, so there is nothing here to confirm or disprove.
// The candidate then scores zero absences, which TIES it with the primitive group rather than
// supporting it, and a centering adopted in that state comes from the lattice METRIC and not
// from these intensities. Reported so that case is visible instead of silent.
bool centering_untested = false;
};
// How unlikely a predicted-absent class would be if the condition producing it did not exist, in
// nats: the -log lower tail of Beta(n_absent, n_control) at T = sum_u / (sum_u + n_control), where
// sum_u is the absent intensities expressed in units of their control class's mean. Used for a SCREW
// against the rest of its own axial row and for a CENTRING against the present class, so the two are
// in the same units and add. The control's own strength cancels, which is the property a count of
// absences does not have. Declared here because it is the statistic the candidate ranking is built
// on, and the comments below name it.
double AbsenceEvidence(double sum_u, int n_absent, int n_control);
// The same for a SCREW zone, with the control COUNT dropped (the b -> infinity limit), which makes it
// a likelihood ratio for the absent class rather than a tail probability. A screw's control class is
// the complement of its absent class on one axial row, so a candidate that predicts more of the row
// absent shrinks its own control and was charged for it - an equally dead SUPERSET of another
// candidate's absences could score lower. See SearchSpaceGroup.cpp.
double ScrewZoneEvidence(double sum_u, int n_absent);
// A screw zone's sum with `max_u` dropped and rescaled to estimate the same quantity (dividing by
// n - H_n, the expected sum of the other n-1 under the null, and multiplying by n). ScrewZoneEvidence
// is a sum, so it is dominated by the largest one or two members of a class that holds only a handful
// of reflections, and a zone's verdict could hang on one reflection whose merged value moved with the
// scaling. max_u is the largest TRIMMABLE member, not simply the largest: the caller withholds a
// member that both violates the absence and stands at or above its row's own present mean, because
// that is a reflection that is there rather than a measurement that moved, and trimming it removed
// the single datum refuting the claim. The sum a withheld member leaves is therefore higher than a
// true trim would give, which UNDER-states the evidence: this estimator only ever refuses a screw,
// never asserts one. See SearchSpaceGroup.cpp.
double TrimmedZoneSum(double sum_u, double max_u, int n_absent);
struct SearchSpaceGroupOptions {
// Lattice (metric) symmetry from LatticeSearch. When set, the point-group search is limited to
// the Sohncke subgroups of this system's holohedry - the metric is an upper bound on the
// intensity symmetry, so e.g. a tetragonal metric tests {1,2,222,4,422} and never 3/6/23. All
// subgroups are still tested (down to P1), so a pseudo-symmetric metric never forces a higher
// symmetry than the intensities support. Unset = search every Sohncke system.
std::optional<gemmi::CrystalSystem> lattice_system;
// Centering is NOT taken from the lattice metric: an indexer that returns the conventional cell
// (e.g. cubic 'P'-looking axes for a body-centered lattice) hides the centering, which lives only
// in the systematic absences. Stage B therefore tests every centering allowed by the point group
// and confirms it from the data (h+k+l etc. absent), rather than trusting a geometric hint.
// The conventional cell the merged reflections are indexed on. Only needed with the two
// enumeration options below: a setting names the symmetry directions by AXIS, so a candidate in a
// setting the cell does not support (a c-unique monoclinic group against a b-unique cell) is
// refused before it is scored. Unset with either option on means no non-reference setting is
// offered at all - the check is a precondition, not an optional extra.
std::optional<gemmi::UnitCell> cell;
// Enumerate the non-reference SETTINGS of the chosen point group in Stage B - P 2_1 2 2 and
// P 2 2_1 2 beside P 2 2 2_1, A 2 2 2 beside C 2 2 2, I 1 2 1 beside C 1 2 1. These are
// alternative namings at the same group ORDER, so this cannot promote the point group; what it
// adds is the ability to put a screw or a centering on the axis the data actually show it on.
// Without it a crystal whose screws are on b and c is reported as P 2 2 2_1 - the wrong group,
// not a lower one, because the only candidate that predicts the 00l absences and nothing else
// wins on zero extra evidence.
bool enumerate_all_settings = false;
// Enumerate the rotation sets that only a non-reference setting carries in Stage A - the a-unique
// and c-unique monoclinic 2-folds. Unlike the option above this DOES add promotions: a crystal
// whose only 2-fold lies on a or c currently has no rung to stand on and falls to P1. Requires a
// cell, and requires enumerate_all_settings for those point groups to be nameable in Stage B
// (no reference setting carries them), which this option turns on for them by itself.
bool enumerate_all_rotation_sets = false;
// Friedel mates are treated as equivalent when matching HKLs (i.e. anomalous signal ignored).
bool merge_friedel = true;
// Ignore reflections beyond this resolution (smaller d = higher resolution). 0 disables.
double d_min_limit_A = 0.0;
// Drop reflections weaker than this from the correlation stage only (the absence stage must
// keep weak reflections - that is where the screw-axis signal lives).
double min_i_over_sigma = 0.0;
// Drop reflections STRONGER than this resolution-normalised E^2 = I/<I>(shell) from the correlation
// stage only. A second lattice deposits intensity on one reciprocal position but not its symmetry
// mate, so an overlap-contaminated reflection is a one-sided E^2 outlier that poisons an operator's
// I(h)/I(Rh) correlation (it flipped a pseudo-merohedral P2_1 case to P1 under -A: excluding the E^2>9 tail, ~0.4%
// of reflections, lifted the 2-fold CC 0.33->0.53 back over the gate - values on the raw-I scale the
// operator CC used before it was normalised). Clean Wilson-distributed data
// almost never reaches E^2=9 (P(E^2>9) ~ 0.01-0.3%), so this removes essentially nothing there and
// only trims the overlap tail. 0 disables. Absences are unaffected (they need the weak tail).
// This E^2 is the one normalised over pass_absence, NOT the one the correlation is scored on: the
// cut helps DEFINE pass_cc, so normalising it over pass_cc would be circular.
double max_e_squared_for_cc = 9.0;
// --- Stage A: point group ---
// A rotation is accepted as a real symmetry when its E^2(h)/E^2(Rh) correlation reaches this over
// at least min_pairs_per_operator independent pairs.
//
// The correlation is on RESOLUTION-NORMALISED intensity (see SearchSpaceGroup.cpp), which is what
// sets this value. On raw I the two arms of every pair share the whole resolution fall-off, so a
// completely false operator still scores a large positive CC - measured over the rotation battery
// by shell-matched random pairing, that noise floor runs 0.05 to 0.53 across crystals, a SPREAD
// (0.46) wider than the entire true/false gap (0.38). One absolute number therefore meant a
// different test on every crystal: the old bound of 0.5 on raw I corresponds, crystal by crystal,
// to a normalised threshold anywhere from -0.03 to 0.47. On the normalised statistic the floor has
// a median of 0.015 and never exceeds 0.06, so a single number finally means the same thing on
// every crystal.
//
// 0.30 is the midpoint of the gap the battery leaves: the weakest genuine conjugacy-class mean is
// 0.35 and the strongest false candidate 0.23 (both on the pessimistic arm, which carries rugnux's
// own search merge onto the normalised scale). Scanning the constant against the battery: 0.20 lets
// a cubic over-promotion in, 0.35 loses a genuine monoclinic 2-fold, and 0.25-0.30 changes nothing.
// It is also the median of the per-crystal threshold the old 0.5 already imposed (0.309), so it is
// the behaviour-preserving choice at the median crystal and only redistributes strictness at the
// tails - away from the strong crystals, which were being handed margin out of their own noise
// floor, and towards the weak ones, which had none.
double min_operator_cc = 0.30;
// KNOWN-WEAK, DELIBERATELY LEFT AT 20 - do not raise it without reading this.
//
// 20 pairs is not a test. The null of the normalised CC - shell-matched random pairing over
// pass_cc with symmetry-related pairs vetoed, i.e. what a metrically-allowed but FALSE operator
// looks like - has a mean of -0.02 to +0.05 and a standard deviation C/sqrt(n) with C = 1.08
// (0.98 to 1.23 across crystals), flat from 20 to 5000 pairs. Its upper tail is far heavier than
// normal theory, because E^2 is exponentially distributed: measured over 142 (crystal, n) points
// the exceedance is log10 P = -1.03 z + 0.69 with z = (min_operator_cc - null mean)/sd, i.e. ONE
// DECADE PER STANDARD DEVIATION where the normal law would give five. At 20 pairs a false
// operator therefore clears 0.30 on 12.4% of draws (and cleared the old bound of 0.5 on 3.2%).
// A bound that held that to 1e-3 on the worst merge measured would be 251 pairs.
//
// It is left at 20 because a pair-count FLOOR is the wrong shape for this problem, and 251 was
// measured to do harm. Nothing in the rotation battery is confirmed on noise: over its 314
// candidate operators the sparsest carries 511 pairs and the weakest confirmed one stands 9.1
// sigma above its own null, so there is no live case to fix. Meanwhile SearchSpaceGroup is run
// twice per data set (see Rugnux.cpp) - once on all observations, which decides, and once on a
// Lorentz-filtered merge that is deliberately starved and only reports - and on that second arm
// genuine operators reach 98 pairs. At 251 the sparsest crystal's filtered arm loses 19 of its
// 23 operators and the run prints "the Lorentz-filtered one supports 1" for an arm whose own
// correlations are 0.91-0.97, every one of them 9 sigma clear. A floor discards an operator on
// pair count alone however decisive its CC is, and that is what makes it the wrong instrument.
//
// The right fix, when it is worth its own calibration and battery, is the pair-count-aware bound
// this null hands over directly: require cc >= max(min_operator_cc, mu + C*z/sqrt(n)). That
// rejects the 20-pair noise the floor is aimed at while keeping every operator above, because it
// asks how large the correlation is and not only how many pairs it came from.
int min_pairs_per_operator = 20;
// Per-operator CC alone cannot tell a real weak operator from a false strong one (a noisy crystal's
// genuine 2-fold can score below a pseudo-symmetric crystal's near-perfect false one). So a point
// group is also required to be SELF-CONSISTENT: merging the intensities under it must not inflate the
// reduced chi^2 (within-orbit scatter / sigma^2) beyond this factor times the most-consistent
// candidate. A false operator forces non-equivalent reflections together so they disagree by many
// sigma and chi^2 blows up; a real one leaves it ~flat even when the operator CC is only moderate.
// Calibrated on the rotation-test battery: every correct point group stays within ~1.7x the best
// subgroup even on weak / badly-integrated data (worst real case a P41212 at 1.71), while a twin
// law or pseudo-symmetry lands clearly higher (a merohedral R3->R32 twin 2-fold at 2.01). 1.85 sits
// between the two, so a partial merohedral twin is kept in its true lower symmetry (R3), not
// over-promoted to the holohedral R32.
double max_merge_chi2_ratio = 1.85;
// A genuine high-symmetry merge can drift just past max_merge_chi2_ratio when its data are only
// imperfectly scaled: each real symmetry step then adds a little systematic scatter, so a weakly-
// scaled cubic case lands at ratio ~2.0 - right where a merohedral twin (an R3->R32 case at ~1.95)
// also lands, so the chi^2 ratio alone cannot separate them. A candidate whose chi^2 is
// only this far past the best subgroup is therefore rescued if the SYSTEMATIC part of the extra
// scatter stayed small (max_systematic_b_ratio): merging under a genuine operator gains multiplicity
// without intensity-proportional disagreement, so the merge error model's b barely moves, whereas a
// twin forces non-equivalent reflections together and b balloons. Both bounds sit in the gap; the
// rescue only ever promotes, and only in this narrow chi^2 band.
double max_merge_chi2_rescue = 2.30;
// 1.78, and this bound is load-bearing in a way it was not when it was first set. Where a genuine
// high-symmetry step and a merohedral twin law BOTH fail the chi^2 ratio, the rescue is the only
// thing still choosing between them, and on a correctly-scaled search merge it is the only one of
// the four Stage-A statistics that can. Measured over the rotation battery on merges whose scaling
// loop runs the iterations it was asked for, a weak tetragonal 4 -> 422 step that is genuine and a
// trigonal 3 -> 32 step that is a twin law read
//
// chi^2 ratio 2.05 genuine vs 1.99 twin H ratio 1.82 genuine vs 1.67 twin
// b ratio 1.71 genuine vs 1.85 twin R_meas 1.31 genuine vs 1.36 twin
//
// - chi^2 and H INVERTED, the genuine step looking worse than the twin on both. That is what b is
// for and why it survives here: a twin's disagreement is proportional to I, and b is the statistic
// that isolates the intensity-proportional part, where chi^2 is normalised by a sigma the weak
// genuine case has too little of and H is a per-operator median that case's angular coverage
// inflates. 1.78 is the midpoint of the gap b leaves.
//
// It is a 4% margin on each side, narrower than the shift a change in how the search merge is
// scaled has already been measured to produce on these same four numbers. Re-measure it whenever
// that scaling changes rather than trusting it through such a change.
//
// 1.44, not the 1.78 those four numbers were read in. This bound is a bound on a RATIO of two b
// values, so its scale is set by how merge_systematic_b normalises its reduced chi^2 - and that
// normalisation changed, from the observation count N to the fit's degrees of freedom N - G.
// Every ratio in the corpus fell with it: a measured -23.7% median over three disjoint corpora,
// lower in 42 of 42 promotions, because the low-multiplicity parent loses a larger fraction of
// its dof than the candidate and so gains more b. Measured on
// one trigonal crystal, same merge, only the divisor differing: the parent's b 0.126 -> 0.181
// (+43%) against the candidate's 0.219 -> 0.254 (+16%), and the ratio 1.74 -> 1.40.
//
// Leaving the bound at 1.78 through that was a change of units without a conversion, and it
// loosened the RESCUE below onto a merohedral twin the old convention refused. Rescaling by the
// same factor restores every verdict exactly - b_new > 0.805 x 1.78 is the same test as
// b_old > 1.78 - which is what the divisor change intended and did not deliver. Derived three
// ways that agree: 1.78 x 0.808; the midpoint of the calibration pair above re-read in the
// corrected convention (genuine 1.71 -> 1.38, twin 1.85 -> 1.49); and a population-median
// derivation that argued 1.38.
double max_systematic_b_ratio = 1.44;
// Promotion gate on the operator disagreement H = median|I1-I2|/(I1+I2), taken as the ratio of the
// operators the promotion ADDS to the operators of the parent group already confirmed on the same
// reflections. A real symmetry operator relates equal intensities, so its H matches the parent's
// (ratio ~1); a merohedral twin law relates DIFFERENT reflections mixed in proportion alpha, so its
// H is systematically larger. Unlike the chi^2 and systematic-b ratios (genuine 1.00-3.47 /
// 1.09-3.89 vs twin 1.35-3.32 / 1.77-4.76, fully interleaved) it does not drift with multiplicity,
// because there is no sigma in it and the parent normalisation cancels data quality.
//
// A MEDIAN, not a mean. A merohedral twin mixes every reflection with its twin mate, so it shifts the
// whole distribution of |I1-I2|/(I1+I2); a minority of badly measured reflections shifts only the
// tail. Measured on the same real crystals, moving from the mean to the median leaves genuine
// promotions where they are (1.016 -> 1.013, 1.051 -> 1.067, 1.238 -> 1.231) and pushes every real
// twin UP (1.280 -> 1.622, 1.427 -> 2.010, 1.272 -> 1.447, 1.441 -> 1.522), widening the margin
// around this bound from 2.7% to 17.5%. Battery-neutral: 33 crystals, no point group changed.
//
// The parent normalisation is what makes this work, and it is not optional. Symmetry-related
// reflections never agree exactly on real data - absorption, illumination and partiality differ
// between them - and that systematic floor varies by crystal AND by operator (a cubic 3-fold permutes
// axes, relating far-apart parts of reciprocal space, so it disagrees more than the 2-folds of its
// parent even when the symmetry is perfectly real). Measuring an added operator against the parent's
// own operators, on the same reflections, is what divides that floor out. An ABSOLUTE bound on any
// per-operator agreement statistic cannot: measured absolute values for genuine symmetry span the
// whole range from 0.99 on strong data down to 0.71 on weak, straddling every twin.
//
// Its known limit is angular coverage, not data quality. What the parent normalisation cannot divide
// out is the part of the systematic floor that differs BETWEEN the added operators and the parent's,
// and that part grows as the measurement gets less uniform: on a full sweep a real tetragonal crystal
// reads the same H on all seven operators of 422 (0.051-0.057, ratio 1.03), but on a quarter of the
// frames of the same crystal the seven spread over 0.082-0.168 and the ratio reaches 1.45, the two
// diagonal 2-folds - left with the fewest surviving pairs - carrying the whole excess. So a refusal
// from a merge built on a small or lopsided part of the sweep says as much about the coverage as about
// the symmetry; the fix for that is upstream, in getting the sweep indexed.
//
// The bound is set from the gap between the two populations, measured over the rotation battery:
// genuine promotions read 0.85-1.57 (and 2.48 on a genuine orthorhombic step in a pair-starved arm),
// real merohedral twins 1.82 and 4.01. 1.25 sat INSIDE the genuine range - four genuine promotions
// already exceeded it and survived only because the two-arm rule happened to cover them, and one
// cubic case with no such cover was refused outright by a margin of 0.4%.
//
// Re-measured over ~780 reports (three batteries and four de-novo integration-radius arms), H is the
// only Stage-A guard that ever fires: 14 refusals on three crystals, no chi-squared-ratio refusal and
// no b veto anywhere. Two of the three are the genuine trigonal twins, reading 1.90/1.91 on every arm
// and 3.61-4.80; the third is a weak tetragonal crystal reading 1.73-1.80 whose 422 is confirmed by
// every operator correlation and which three independent tests - our own L test on the merge it was
// demoted into, xtriage on the same merge ("most likely due to an NCS axis parallel to the twin
// axis"), and the arithmetic that a twin at the implied fraction caps pair correlation at 0.55 while
// the added operators score 0.84-0.90 - agree is not twinned. Its excess is the coverage limit named
// above: the added 422 operators exchange h and k and carry a systematic floor the parent's
// sign-flipping 2-folds do not. So the populations collide at 1.80 and 1.90 rather than at 1.70, and
// the bound sits at the midpoint. Degrading the data drives the ratio towards 1, never up, and the
// estimator is precise (bootstrap standard error 0.2-3.7% of the ratio), so the collision is a real
// overlap and not sampling noise - which is also why scaling the bound by that error cannot work.
//
// Those two populations were measured on search merges whose scaling loop had been over-iterated,
// and on correctly-scaled ones they SWAP: the weak tetragonal case rises to 1.82 and the trigonal
// twin falls to 1.67, so no value of this bound keeps the one and refuses the other. H is kept as a
// necessary condition because it still refuses the twins that read far above either (3.4 and up),
// but it no longer decides the close case - max_systematic_b_ratio does.
double max_operator_h_ratio = 1.85;
// The H test needs at least this many pairs on both sides to mean anything.
int min_pairs_for_h = 200;
// Full-resolution merge-degradation gate. The operators a promotion ADDS are scored by their
// intensity-weighted R (r_stat: a SUM, so the strong low-resolution reflections that carry the
// contrast dominate it, where the median H under-weights them), and that R is judged against a
// CLEAN reference. This is the one Stage-A gate whose reference cannot be contaminated by false
// operators: the H ratio normalises against the PARENT group, and a parent can itself be a false
// promotion (a pseudo-tetragonal 2 -> 222 -> 422 cascade confirms 222 by pooling the one real
// 2-fold with two false ones, so the 422 step is then measured against an already-contaminated
// 222, and every ratio reads ~1.4 and passes). It also acts on the first step out of P1, where
// there is no parent to normalise against at all and the operator CC floor stood alone. Sigma-free, so unlike the chi^2 and b gates it does not saturate with the
// merge's ISa on a starved search merge.
//
// Two references, each scale-free so the bound means the same on every crystal (the property an
// absolute R bound lacks):
//
// - min_operator_r_contrast: when the data confirm more than one operator, the added operators'
// mean R placed on the scale the crystal itself defines, between the R of UNRELATED
// reflections (random_pairing_r, where a false operator sits) and the R of the best-agreeing
// operator anywhere (global_best_operator_r, where a genuine one sits):
//
// contrast = (random_pairing_r - r_added) / (random_pairing_r - global_best_operator_r)
//
// 1 means "agrees as well as the cleanest operator in the crystal", 0 "agrees no better than
// unrelated reflections". Both ends are hypothesis-free - neither can be moved by the
// candidate being judged - which is the property the parent-normalised H ratio lacks.
//
// This replaced a plain ratio r_added/global_best_operator_r, bound 2.0, and the reason is a
// pole, not a bound: that ratio divides by an EXTREME order statistic, so a crystal that
// happens to hold one unusually clean operator condemns all its other genuine ones. It is not
// rare. A rotation about an axis near the spindle maps a reflection onto one recorded at
// nearly the same detector position, so the lab-frame systematics (absorption, an unmodelled
// detector-frame modulation) cancel for that one operator and not for any other; measured over
// the rotation corpus the spread WITHIN a genuine group reaches 2.0-2.6x, i.e. the whole width
// of the old bound. A weak cubic crystal whose near-spindle 3-fold read 0.074 while its ten
// other genuine operators read 0.16-0.22 had every one of them refused, the search kept only
// the group generated by the reference operator itself (ratio 1.00 by construction), and the
// same 222 hypothesis on the same merge read 2.45 under a cubic enumeration and 1.02 under an
// orthorhombic one - a statistic that moves with which OTHER operators were enumerated.
// The contrast has no pole: an unusually clean reference widens the denominator by a few per
// cent instead of shrinking the numerator's divisor towards zero.
//
// Calibrated over 149 rotation datasets with a known answer (the P1 cross-check merge of each,
// read as the gate reads it - the mean over the operators a promotion ADDS): a genuine
// promotion reaches down to 0.77 and the worst false one on the corpus reads 0.695. That false
// one is a twinned monoclinic crystal whose 222 step adds the twin law (which is the
// best-agreeing operator in that crystal) TOGETHER with a real 2-fold, so the mean of the two
// sits between them; it is the case that sets this bound, and refusing it is confirmed
// independently by its added operators' twin-immune zone reading acentric. 0.72 sits in that
// gap, with ~0.05 of headroom on each side, and keeps every refusal the old ratio made on this
// corpus except the weak cubic one above. Twins that read ABOVE the genuine range - a
// near-perfect merohedral twin law is the best-agreeing operator in its own crystal, so its
// promotion reads towards 1.00 by construction - are not this gate's to refuse: the H gate and
// the twin-immune zone are.
double min_operator_r_contrast = 0.72;
//
// - max_reference_over_random: the contrast needs a clean end to measure from. Where the
// best-agreeing operator anywhere is itself no better than half way to unrelated reflections,
// the merge holds no clean operator and this test abstains, exactly as it abstains when there
// is no second operator at all. Over the same 149 datasets every macromolecular merge sits at
// 0.05-0.33 and only two small-molecule ones (0.61, 0.81) trip this.
double max_reference_over_random = 0.5;
//
// - max_operator_r_over_floor: the first step out of P1 confirms only the one operator, so there
// is no other to be the best - fall back to the merge's own random-noise R floor (merge_r_floor,
// from the half-dataset merges). The floor is a genuine-symmetry disagreement (two halves of
// the SAME reflections) at the same multiplicity, so it too is a clean scale-free reference.
// Measured: a genuine first 2-fold sits at ~1.3-2.3x the floor, a false one (a pseudo-C-centred
// metric coincidence) at ~4.7. 3.5 sits in that gap. Wider than the best-operator bound because
// the floor is a noisier reference than a confirmed operator (it can also be inflated by an
// indexing ambiguity that mixes hands), and the safe direction here is to stay in P1.
double max_operator_r_over_floor = 3.5;
//
// - max_operator_r_spread: the two tests above read a MEAN over the operators a promotion adds,
// so they are blind to an added set that MIXES real operators with false ones - the mean sits
// between the two populations and passes. That is what a pseudo-symmetric metric produces: a
// monoclinic crystal whose lattice is exactly hexagonal (a = c, beta = 120) carries two genuine
// horizontal 2-folds among the six a 622 promotion adds, and their R at the noise floor buys
// the other four a mean well inside the contrast bound. The mixture is invisible to any summary
// of the added set and obvious the moment the candidate's operators are compared with EACH
// OTHER: a point group is the claim that all of its operators are symmetries of one crystal, so
// they must agree with the merge about equally well.
//
// Measured over the 129 rotation datasets of one battery run that adopted a group of order 3 or
// more, as worst-agreeing over best-agreeing operator WITHIN the adopted group: every correct
// adoption sits at 1.0-2.64, the highest being the weak cubic crystal whose near-spindle 3-fold
// is unusually clean (the same crystal that motivated the contrast above), then 2.30 for a
// fine-sliced tetragonal one. The one over-called point group on that run reads 4.87. 3.2 sits
// in that gap with 0.56 of headroom below and 1.67 above - deliberately well clear of the
// 2.0-2.6 spread a genuine group reaches, which is exactly where the old ratio-to-best gate's
// bound of 2.0 sat and why that gate had to go.
double max_operator_r_spread = 3.2;
// Above this reduced chi^2 for the best subgroup (chi2_ref), the merged error model is treated as
// badly miscalibrated (weak, low-resolution data whose sigmas are far too small): the fixed-sigma
// chi^2 ratio then grows with point-group order for genuine high symmetry too and can no longer
// arbitrate, so a promotion is confirmed on the systematic-b test alone (which re-fits its own error
// and stays valid). A well-calibrated merge sits near 1; 3.0 (sigmas ~1.7x too small) marks the point
// where the ratio stops being trustworthy. The H and added-operator-R gates below still guard
// against a twin: both are necessary conditions no rescue can lift.
double chi2_ref_reliable = 3.0;
// The other end of the same statistic, where it also stops meaning anything. A merged reduced
// chi^2 decomposes as chi2 = a^2 + b^2 <(I/sigma)^2>: a measurement part, which a fitted error
// model pins at ~1 for every crystal, plus a systematic part b^2<(I/sigma)^2>. The ratio
// chi2_cand / chi2_ref is meant to compare the SYSTEMATIC parts - how much intensity-proportional
// disagreement the added operators bring, against how much the reference merge already carries -
// but it divides by the total. Expanding the bound, a promotion is allowed added variance
// (ratio - 1) x (a^2 + b_ref^2<(I/sigma)^2>): the second term is the intended relative allowance,
// the first is a fixed one that has nothing to do with the crystal. Where the reference merge is
// cleanly scaled its systematic part goes to zero, the fixed term is all that is left, and the
// "ratio" is really an absolute bound on b_added^2<(I/sigma)^2> - so the same real degradation
// reads larger the better the data are measured, and a crystal is refused for the cleanliness of
// its own reference merge. Measured on one crystal collected twice: the added operators cost the
// same absolute chi^2 both times (+1.16 and +1.07) and the same systematic b (0.080 and 0.074),
// but the cleaner reference (chi2_ref 1.21 against 1.44) turned that into 1.96x against 1.74x, and
// only the cleaner merge was refused. So below this chi2_ref the ratio is not read at all. The
// value is where the two parts of chi2_ref are equal - systematic content matching measurement
// content, a^2 ~ 1 plus as much again - not a level tuned to any dataset. The b test is NOT the
// fallback here as it is under chi2_ref_reliable: it divides by the reference's own b and fails in
// exactly the same way. What replaces the ratio is the ABSOLUTE form of the same question,
// against chi2_ref_reliable: the merge under the candidate must itself look calibrated. The H and
// added-operator-R gates still apply on top, as necessary conditions no rescue can lift.
double chi2_ref_informative = 2.0;
// Adopt this space group's point group instead of the one Stage A would choose, and go straight to
// Stage B with it. Used when one merge decides the point group and a different merge has to decide
// the absences: a caller cannot pass the point group by name, because gemmi reports both P321 and
// P312 as "32", so the group itself is passed and its rotations are taken from it. Stage A still
// runs (its operator scores and its refusal report are still wanted); only the choice is overridden.
std::optional<gemmi::SpaceGroup> fixed_point_group;
// --- Stage B: space group (screw axes / centering) ---
bool determine_space_group = true; // false: stop at the symmorphic representative
// Signed I/sigma above which a reflection counts as genuinely present. Signed, so negative
// noise is never mistaken for a real reflection. Used both to spot reflections that violate a
// wrongly-assumed absence, and to keep the correlation stage from pairing near-zero reflections.
// Read against the reflection's COUNTING sigma, not against the merged one - see merge_isa, which
// is what converts it into a cut on the quantity the merge actually exports.
double present_i_over_sigma = 3.0;
// ISa = 1/b of the error model the merge handed in was scaled with: the I/sigma one observation
// of an arbitrarily strong reflection can reach. 0 = not known, or b = 0 (no systematic term at
// all), and then present_i_over_sigma is used as it stands.
//
// It is needed because a merged I/sigma is not a measure of how strong a reflection is. The merge
// carries the error model's intensity-proportional term, sigma^2 = a*sigma0^2 + (b*I)^2, so
// I/sigma saturates: at ISa*sqrt(n) for a reflection observed n times, and at ISa exactly for one
// observed once. Above that knee the quantity stops rising with the intensity - measured over the
// rotation battery's space-group search merges, the median I/sigma of the top E^2 decile sits at
// 0.97 of ISa on the weakest merge, and its log-log slope against E^2 over the strong half of the
// merge is 0.008, i.e. flat. An absolute cut on it is therefore a different test on every crystal,
// and on a merge whose ISa is near the cut it is close to unreachable: 3% of reflections clear 3.0
// there, against a battery median of 53%.
//
// So the cut is converted to the counting scale once, using the reflection's own significance.
// For a reflection observed once sigma_counting^2 = sigma^2 - (b*I)^2, hence
// I/sigma_counting = t/sqrt(1-(t/ISa)^2) with t = I/sigma, and "I/sigma_counting >= T" is exactly
// "t >= T/sqrt(1+(T/ISa)^2)". That converted cut lies strictly below ISa for every ISa, so it is
// always reachable, and it is within 1% of T on any merge with ISa >= 21, so it leaves a healthy
// merge where it was. Multiplicity is taken as 1 deliberately and not estimated: the cut has to be
// reachable by the reflection the ceiling binds hardest, which is the one measured once.
double merge_isa = 0.0;
// Extra intensity gate for the systematic-absence test: a reflection also has to reach this
// resolution-normalised intensity E^2 = I / <I>(shell) to count as violating a predicted absence.
// The error model can under-estimate sigma on weak axial reflections and fake a high I/sigma, so a
// reflection at a few percent of the shell-mean intensity is judged absent regardless of its sigma
// (e.g. the monoclinic 2_1: 0k0-odd at ~1% of 0k0-even). 0 disables the gate (I/sigma only).
//
// For a SCREW the threshold is this fraction of the median E^2 of the axial row the screw
// constrains (floored at the plain value), not of an average reflection at that resolution - an
// axial row is often far stronger than the shell mean, and only the row is a fair comparison.
double present_e_squared = 0.3;
// A candidate's SCREW/glide absence conditions are accepted when at most this fraction of the
// reflections it predicts absent are in fact strongly present.
double max_absent_violation_fraction = 0.10;
// A CENTERING is accepted when its systematically-absent class is this much weaker than the present
// class - mean signed I/sigma of the centering-absent reflections <= this fraction of the present
// mean. A real centering cancels structure factors so its absent class sits near zero (ratio ~0-0.3
// across the test battery) even on noisy or obverse/reverse-twinned data; a false centering leaves it
// as strong as the present class (ratio ~1.0). 0.5 separates the two with wide margin. This strength
// test replaces a per-reflection violation-count gate for centering, which was brittle when noise
// pushed genuinely-absent reflections over I/sigma>3 (a true R3 was lost at 13.5% violations).
double max_absent_present_ratio = 0.5;
// Need at least this many observed reflections in the CENTERING-absent class before a centering is
// claimed (guards against deciding from a handful of reflections). A real centering extinguishes a
// third to a half of every reflection in the data set, so this is never the binding constraint on a
// centering that exists - it only refuses one the data barely sampled.
//
// It does NOT apply to screws. A screw's absent class is one row of reciprocal space by
// construction, and that row is often the one a rotation sweep records least: it lies near the
// spindle, where the blind cusp maps onto itself and symmetry cannot fill it in. Counting it
// measures the geometry of the sweep, not the strength of the evidence. Applied to screws this
// bound cost a monoclinic crystal its 2_1 - six 0k0-odd reflections, every one of them measured at
// |E^2| <= 0.013 with zero violations, against a 0k0 row averaging 1.44x the shell mean, refused
// for being six rather than eight. Screws answer to min_screw_absence_evidence below instead.
int min_absent_observed = 8;
// Evidence, in nats, that a screw's predicted-absent class really is absent, required before the
// screw may be claimed. In the spirit of the POINTLESS zone test (Evans, Acta Cryst D67, 282-292
// (2011), Appendix A3), which likewise scores an absence against the rest of its own axial row
// rather than against a global mean or a fixed cut, and likewise lets the confidence fall away with
// the number of axial reflections instead of refusing outright below a count.
//
// The statistic (see ScrewZoneEvidence) is the likelihood ratio of the absent class against its
// row's control class - the -log Beta tail AbsenceEvidence uses for a centring, with the control
// COUNT dropped so that candidates judged on the same row stay comparable. Two properties are
// what a violation count lacks: the row's own strength cancels, so a uniformly weak axial row
// decides nothing rather than reading as "absent"; and the scale is set by the number of
// reflections, so few-but-decisive and many-but-marginal are told apart. It is sigma-free by
// design - the merged sigma carries the error model's intensity-proportional term and so shrinks
// with I, reading much the same on an absent reflection as on a present one (the <I/s> columns of
// the candidate table show this directly).
//
// The zone's sum is taken with its largest member trimmed (TrimmedZoneSum), because a sum over
// half a dozen reflections is otherwise decided by its largest one and moved tens of nats by any
// upstream change that moves that one reflection.
//
// Measured over the probe crystals, an adopted screw reads 24-600 nats and the false ones (the
// 4_1/4_3 conditions of a crystal that has no screw at all, whose predicted-absent class is as
// strong as its control row) read -3.8 to -38.5. The bound sits in that gap - loose enough that
// three well-measured dead axial reflections clear it, tight enough that two do not.
//
// That last sentence is no longer true, and the reason is worth recording rather than re-tuning
// blind. MIN_U_PER_ABSENT_REFLECTION floors sum_u at 1e-3 per absent reflection, which caps a
// zone's evidence at -a*log(a*1e-3) + log(a!): 13.1 nats at two absences, 19.2 at three, 25.3 at
// four. A zone of three therefore CANNOT clear 20 however dead its reflections are, so the gate
// is in practice "four absences in one zone" - a count, which is the measure the whole statistic
// exists to avoid. A thin monoclinic sweep reaching only three 0k0-odd reflections loses its
// 2(1) to this. Lowering the floor to 7.7e-4 would let three through; that is a recalibration
// and needs its own battery.
double min_screw_absence_evidence = 20.0;
// ---- Glide planes (small-molecule space groups) -------------------------------------------
// Offer the NON-SOHNCKE space groups of the same proper-rotation set as candidates, so a glide
// plane can be named. Protein crystals are built from L-amino acids and are therefore chiral:
// their groups are Sohncke and contain rotations and screws only, so this can never change a
// protein answer except through a false positive - which is why it is measured rather than
// assumed (see min_glide_evidence_per_reflection). Small molecules are where it earns its place:
// three corpus datasets get the right cell and the right Sohncke subgroup and stop one c-glide
// short of the deposited group.
//
// Only a group whose ABSENCES differ from a Sohncke candidate's is offered. That is not an
// optimisation, it is the honest limit of the measurement: P2/m predicts exactly what P2
// predicts, so the inversion centre buys no candidate and is never claimed. See the note on
// Friedel's law at the top of SearchSpaceGroup.cpp.
bool enumerate_non_sohncke = true;
// How dead a glide zone must be, per reflection, in nats, before the glide may be claimed.
// ScrewZoneEvidence is asymptotically n * (log(1/ubar) - 1) where ubar is the absent class in
// units of its zone's control mean, so this is a bound on ubar and not on a sum: 2.0 nats is
// ubar <= exp(-3) = 0.050, i.e. the zone's absent class must be 20x weaker than the rest of its
// own plane. Measured over 140 protein datasets and 6 small-molecule ones: the largest protein
// zone reads 0.65 nats/reflection (ubar 0.19) and the next 0.12, while genuine glides read
// 4.24-5.95 (ubar 0.0003-0.0054). The bar sits between the two populations with about a factor
// of 4 in ubar to the false side and 10 to the true side.
double min_glide_evidence_per_reflection = 2.0;
// A zone with fewer absences than this, or fewer than half as many control reflections, cannot
// be judged and refuses the candidate that needs it. Small numbers are where a RATE is noisy:
// the corpus's one Sohncke small molecule shows an 8-reflection zone at 1.9 nats/reflection, on
// a crystal whose true group has no glide at all.
int min_glide_absent = 20;
// Workers for the operator-correlation stage, which is the bulk of the search: one pass over the
// whole merge per candidate rotation, and the rotations are independent of each other. 1 = serial.
size_t nthreads = 1;
};
// One row of the FINALIST LEDGER: what the intensities say about ONE point-group hypothesis, on the
// merge this search was handed. Every candidate the operator CCs confirmed gets a row, adopted or
// refused, so a run can show what the rival was and how far behind it came - the current flow keeps
// only the winner and the single highest refusal.
//
// Every number here is PAIRED or a RATIO TO A REFERENCE THE HYPOTHESIS CANNOT MOVE, never an absolute
// per-operator statistic: r_contrast places the added operators between unrelated reflections and the
// best-agreeing operator anywhere in the data, r_over_floor divides by the merge's own random-noise
// floor, chi2_over_best and b_over_parent by the same
// quantity under a subgroup of this very candidate. An absolute score would be the joint-likelihood
// design that was already refuted; a ratio to a clean within-crystal reference is the form thread1_A
// measured to separate genuine operators (1.06-1.15x) from false ones (2.4-4.66x).
//
// REPORT-ONLY. Nothing in the search reads these back; the adoption is made exactly as before.
struct PointGroupLedgerEntry {
std::string point_group_hm;
int order = 0;
double min_class_cc = 0.0; // weakest operator CC among the group's own operators
double chi2 = std::numeric_limits<double>::quiet_NaN(); // merge chi^2 folding under it
double chi2_over_best = std::numeric_limits<double>::quiet_NaN(); // ... over the most consistent subgroup's
double b_extra = 0.0; // systematic-b refit under it
double b_over_parent = std::numeric_limits<double>::quiet_NaN(); // ... over its largest confirmed subgroup's
double h_ratio = std::numeric_limits<double>::quiet_NaN();
double r_added = std::numeric_limits<double>::quiet_NaN(); // added operators' mean intensity-weighted R
double r_over_best = std::numeric_limits<double>::quiet_NaN(); // ... over the best-agreeing operator anywhere
double r_contrast = std::numeric_limits<double>::quiet_NaN(); // ... 0 = unrelated reflections, 1 = that best operator
double r_over_floor = std::numeric_limits<double>::quiet_NaN(); // ... over the merge's random-noise R floor
double r_spread = std::numeric_limits<double>::quiet_NaN(); // this group's own worst operator R over its best
bool adopted = false; // this is the group the search took
bool eligible = false; // it passed every consistency test (so it COULD have been taken)
std::string refused_reason; // why it did not, when it was the highest refusal
};
struct SearchSpaceGroupResult {
std::optional<gemmi::SpaceGroup> best_space_group;
// Other space groups that fit the data equally well (same systematic absences): enantiomorphic
// partners (P4_1 vs P4_3), origin-ambiguous pairs (I222 vs I2_12_12_1), or groups left
// undetermined by incomplete data. The data cannot choose between best_space_group and these.
std::vector<gemmi::SpaceGroup> alternatives;
std::string point_group_hm; // chosen point group, e.g. "422"
// The symmorphic space group representing that point group. This, not the name, identifies it:
// gemmi reports both P321 and P312 as "32", so two searches that disagree about which 2-folds are
// real look identical by name. Also what a caller passes back as fixed_point_group.
std::optional<gemmi::SpaceGroup> point_group_representative;
// Order of that point group (its number of rotations). Reported separately because Stage B can
// leave best_space_group unset - no candidate had enough absences to be eligible - while Stage A
// has confirmed the point group perfectly well, and a caller comparing two searches has to see
// the symmetry that was found either way. 0 only when no point group was chosen at all.
int point_group_order = 0;
std::vector<SpaceGroupOperatorScore> operator_scores; // Stage A, all distinct operators tested
std::vector<SpaceGroupCandidateScore> candidates; // Stage B, ranked
// The twin-law H ratio of the point-group promotion that was ADOPTED, and the bound it was
// judged against (max_operator_h_ratio). NaN when there is no ratio to form: P1, the first step
// out of it (no parent group to normalise against), or too few pairs on either side.
//
// Reported on every run, not only when the bound refuses a promotion, which is all it used to
// be. The two populations that bound sits between are 1.80 and 1.90 apart, so where a CONFIRMED
// promotion falls inside that window is exactly what a recalibration needs - and printing the
// number only on refusal made it unreadable from every run that came out right.
double h_ratio = std::numeric_limits<double>::quiet_NaN();
double h_ratio_bound = 0.0;
// The merge's own RANDOM-noise R floor: sum|Ihalf0-Ihalf1| / sum(Ihalf0+Ihalf1) over the present
// reflections, from the two half-dataset merges. This is the intensity-weighted disagreement two
// GENUINELY equivalent groups of observations of the SAME reflection show - i.e. the value the
// r_stat of a real symmetry operator would tend to (plus a modest systematic floor), measured on
// this very crystal at this very multiplicity. It is the clean, uncontaminated reference the
// operator R is judged against (see max_operator_r_over_floor), the one a false operator cannot
// move because it is built without applying any candidate symmetry. NaN if the merge carried no
// half-dataset intensities. Reported so a run shows the margin the gate had.
double merge_r_floor = std::numeric_limits<double>::quiet_NaN();
// The point group the CONFIRMED OPERATORS GENERATE, when that group is not itself one of the
// candidates the operator bar admitted. A candidate is admitted only if EVERY one of its
// operators clears min_operator_cc, so a confirmed set whose products are not all confirmed
// leaves the search standing on a subgroup while holding evidence for something larger, and
// nothing says so. Where that happens the generated group - or, when the basis the merge is
// indexed on names no such group, the smallest candidate containing it - is offered as a
// candidate and judged by the same consistency tests as every other. Empty on the usual run,
// where the confirmed operators already form one of the candidates.
std::string generated_point_group_hm;
bool generated_point_group_adopted = false;
// The half-integer pseudo-translation a screw zone was judged against, per axial row, in halves
// ({1,1,1} is (1/2,1/2,1/2)), with the depth of the modulation it puts on that row. Empty when
// the merge carries none, which is the usual case. A translation like this suppresses one parity
// class of every reflection, the axial rows included, so an absence test that expects every
// reflection to be equally strong reads the suppressed class as extinct and claims a screw axis
// that is not there.
struct RowPseudoTranslation {
std::array<int, 3> row{};
std::array<int, 3> halves{};
double ratio = 0.0; // <E^2>(suppressed)/<E^2>(control) at the lowest resolution measurable
};
std::vector<RowPseudoTranslation> pseudo_translations;
// The intensity-weighted operator R of the promotion that was ADOPTED (mean over the operators it
// adds), that R over merge_r_floor, that R over the globally best-agreeing operator, and where it
// falls on the scale between unrelated reflections and that best operator - the number the
// merge-degradation gate acts on once two operators are confirmed (the floor ratio decides the
// first step out of P1). Reported so a run shows the margin the gate had.
double r_added = std::numeric_limits<double>::quiet_NaN();
double r_over_floor = std::numeric_limits<double>::quiet_NaN();
double r_over_best = std::numeric_limits<double>::quiet_NaN();
double r_contrast = std::numeric_limits<double>::quiet_NaN();
double global_best_operator_r = std::numeric_limits<double>::quiet_NaN();
// The intensity-weighted R of UNRELATED reflections on this merge - shell-matched pairs of
// reflections no symmetry relates. The far end of the contrast scale above: where a false
// operator's R sits. NaN when the merge carried too few present reflections to form it.
double random_pairing_r = std::numeric_limits<double>::quiet_NaN();
// A HIGHER point group whose operators the intensities confirmed (Stage A) but whose promotion the
// consistency tests refused, with the reason. Processing continues in the lower group, which is the
// safe direction: merging a twinned crystal in the twin's holohedry averages non-equivalent
// reflections into each other and is unrecoverable from the output (and makes the run report that no
// twin law exists), whereas keeping the subgroup costs only redundancy and can be promoted later.
// Empty when nothing was refused. Surfaced to the user - a silent demotion is how a twin gets missed.
std::string refused_point_group_hm;
std::string refused_reason;
// Of the refused candidate: the twin fraction its added operators' own disagreement implies
// (0.5 - H, a lower bound; NaN where H was not measured) and whether the refusal came from a twin
// gate (the H ratio or the added-operator R) rather than from the merge chi^2.
double refused_twin_fraction = std::numeric_limits<double>::quiet_NaN();
bool refused_by_twin_gate = false;
// The symmorphic representative of that group - the same thing point_group_representative is, and
// for the same reason (a name cannot tell P321 from P312). It is what a caller hands back as
// fixed_point_group to ask what merging under the refused group would give. Unset when nothing
// was refused.
std::optional<gemmi::SpaceGroup> refused_point_group_representative;
// Every operator-confirmed point-group hypothesis with its full evidence vector, ranked by order.
// Report-only (see PointGroupLedgerEntry); the adoption above is unchanged by it. Empty on a search
// that confirmed no operator at all, which the report renders as "no finalists" rather than as an
// empty table - a blank one reads as "nothing was wrong" when it means "nothing was asked".
std::vector<PointGroupLedgerEntry> point_group_ledger;
// The refusal bounds the ledger's own channels were judged against, copied from the options so the
// report can print the MARGIN the adopted row had rather than only its value. The accept-side
// headroom is thin (measured: the lowest genuine adopted contrast on the corpus is 0.77 against a
// bound of 0.72 - see min_operator_r_contrast, which is where these numbers are kept), and a
// margin is the only form in which that is visible per run.
double r_contrast_bound = 0.0;
double r_over_floor_bound = 0.0;
// An axial row whose SCREW these data could not decide: the selected group and at least one of
// its `alternatives` disagree about whether that row carries screw absences, and nothing in this
// merge separated them. The adopted group is then the lowest-numbered member of that set - the
// symmorphic one, by convention and not by measurement - so a report printing only
// best_space_group reads as a determination where the honest answer is "P2 or P2_1, no answer
// possible". The written files still carry one group, because a reflection file cannot hold
// "maybe a screw"; this is what lets the report, the log and a scorer say which axis is open.
// Empty on the usual run, where every screw the data could show was judged.
struct UndeterminedScrew {
std::array<int, 3> row{}; // sign-canonicalised as everywhere else, e.g. {0,0,-1} for 00l
std::string row_label; // the row as a crystallographer names it: "h00", "0k0", "00l"
char axis = ' '; // the axis a screw would run along: 'a', 'b' or 'c'
int n_observed = 0; // reflections of the row in this merge at all - 0 = never recorded
int n_absent = 0; // of those, the ones a claiming candidate predicts absent
int n_control = 0; // ...and the rest of the row, which an absence is judged against
};
std::vector<UndeterminedScrew> undetermined_screws;
// The best SOHNCKE candidate, always, even when a glide plane was confirmed and best_space_group
// therefore names a non-Sohncke group. Both are reported on every run: a chiral (protein) crystal
// cannot have a glide, so a user who knows their sample is a protein must be able to see the
// Sohncke answer without re-running, and a user working on a small molecule must be able to see
// what the glide bought. Unset only when no Sohncke candidate was eligible.
std::optional<gemmi::SpaceGroup> sohncke_space_group;
// The best NON-SOHNCKE candidate - the "small molecule" answer, the group named with its glide
// planes. Unset when no glide zone in this merge cleared min_glide_evidence_per_reflection,
// which is the expected and measured outcome on every protein dataset.
std::optional<gemmi::SpaceGroup> glide_space_group;
// That candidate's glide zones, so a run shows WHICH plane was found dead and how dead it was.
std::vector<SpaceGroupCandidateScore::GlideZone> glide_zones;
};
SearchSpaceGroupResult SearchSpaceGroup(
const std::vector<MergedReflection>& merged,
const SearchSpaceGroupOptions& opt = {});
// The finalist ledger as a ranked table, highest point group first. Report-only; printed behind
// --finalist-ledger so the numbers cannot be read as a decision before they have been calibrated.
std::string FinalistLedgerToText(const SearchSpaceGroupResult& result);
// Stage A's operator statistic for ONE rotation, given by its Miller-index matrix (h -> M h) in the
// basis the merge is indexed on: the correlation of resolution-normalised E^2(h) with E^2(Mh), over
// the same reflection population and with the same Friedel-folded pairing Stage A uses. For a caller
// that wants to ask about a single operator without enumerating a point group - a metric two-fold, say.
// Nothing when the operator makes too few pairs to judge (min_pairs_per_operator); the returned
// `present` flag is the same min_operator_cc test Stage A applies.
std::optional<SpaceGroupOperatorScore> OperatorCorrelation(const std::vector<MergedReflection>& merged,
const gemmi::Mat33& hkl_matrix,
const SearchSpaceGroupOptions& opt = {});
std::string SearchSpaceGroupResultToText(
const SearchSpaceGroupResult& result,
size_t max_candidates_to_print = 20);