v1.0.0-rc.160 (#70)
Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:cuda (push) Successful in 18m44s
Build Packages / build:viewer-tgz:cpu (push) Successful in 6m11s
Build Packages / build:viewer-tgz:cuda (push) Successful in 6m54s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 9m40s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 10m41s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 10m10s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 10m4s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m5s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 12m23s
Build Packages / build:rpm (rocky8) (push) Successful in 11m30s
Build Packages / build:rpm (rocky9) (push) Successful in 12m51s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m8s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m21s
Build Packages / DIALS test (push) Successful in 13m22s
Build Packages / XDS test (durin plugin) (push) Successful in 9m2s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 7m55s
Build Packages / XDS test (neggia plugin) (push) Successful in 5m57s
Build Packages / Generate python client (push) Successful in 23s
Build Packages / Build documentation (push) Successful in 57s
Build Packages / Create release (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 10m24s
Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:cuda (push) Successful in 18m44s
Build Packages / build:viewer-tgz:cpu (push) Successful in 6m11s
Build Packages / build:viewer-tgz:cuda (push) Successful in 6m54s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 9m40s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 10m41s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 10m10s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 10m4s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m5s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 12m23s
Build Packages / build:rpm (rocky8) (push) Successful in 11m30s
Build Packages / build:rpm (rocky9) (push) Successful in 12m51s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m8s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m21s
Build Packages / DIALS test (push) Successful in 13m22s
Build Packages / XDS test (durin plugin) (push) Successful in 9m2s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 7m55s
Build Packages / XDS test (neggia plugin) (push) Successful in 5m57s
Build Packages / Generate python client (push) Successful in 23s
Build Packages / Build documentation (push) Successful in 57s
Build Packages / Create release (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 10m24s
This is an UNSTABLE release. It includes many experimental features, as well as many AI generated fixes. We recommend using rc.152 for production use. * rugnux: Add `--model model.pdb` - score the merged data against an atomic model and compute initial maps. It reports R-work/R-free (scaling the model to the observed amplitudes with an overall scale, an anisotropic B and a flat bulk solvent - the standard few-parameter model, so a batch of maps stays directly comparable) and writes 2Fo-Fc / Fo-Fc electron-density maps (CCP4) plus a map-coefficient MTZ. The structure itself is not refined; the model is only re-fractionalised into the data cell. * rugnux: The merged reflection output now carries French-Wilson amplitudes (|F| and its sigma) next to the intensities - MTZ `F`/`SIGF`, mmCIF `_refln.F_meas_au`, and the text HKL - computed with the correct centric/acentric Wilson prior and epsilon multiplicity, so a downstream program (e.g. phenix.refine) can refine against amplitudes. The intensity columns are unchanged. * rugnux: R-free test-set flags are now assigned deterministically and consistently across symmetry - a Bijvoet pair I(+)/I(-) is never split between the work and free sets, and the assignment is a reproducible per-hkl hash that depends only on the reflection index, so every dataset of one crystal form gets the same ~5% free set (what a multi-dataset campaign such as PanDDA needs). On small data the fraction is floored so the test set stays large enough for a stable R-free (~500 reflections, capped at 10%); it stays flat at 5% on ordinary data. When a reference MTZ carries a `FreeR_flag` column its test set is imported instead, letting a whole campaign inherit one shared free set. * rugnux: A reference MTZ (`--reference-mtz`) can now fix the space group and cell for rotation data too (previously rejected), without being used to scale - the rotation merge stays self-consistent. When the crystal has an indexing (merohedral) ambiguity - a lattice symmetry higher than its Laue symmetry, e.g. P3/P4/P6/C2 - the reference also resolves it: each candidate reindexing (identity plus the twin-law cosets of the metric symmetry) is scored by its intensity correlation against the reference and the data are re-merged in the best-correlating one. This is a metric-preserving relabelling of hkl (the cell is unchanged) and a no-op for a holohedral crystal such as lysozyme. * rugnux: `--model` validation now aligns the data to the model before scoring - the observed reflections are reindexed into the model's enantiomorph when the two differ only by hand (indistinguishable from merged intensities). A merohedral indexing ambiguity is resolved against the reference MTZ when one is given (so a whole campaign shares one indexing convention); only with a model and no reference does validation fall back to fitting each candidate reindexing and keeping the lowest R-free. * rugnux: De-novo symmetry - recover a genuine high-symmetry group whose data are imperfectly scaled. Such a merge's within-orbit chi² lands just past the self-consistency bound (each real symmetry step adds a little systematic scatter), right where a merohedral twin also lands, so the chi² ratio alone cannot separate them. The candidate is now rescued when the extra intensity-proportional systematic error it invokes stays small relative to the confirmed subgroup - a genuine symmetry step gains multiplicity without inflating the merge error model's b, whereas a twin forces non-equivalent reflections together and b balloons. Fixes cubic insulin (I23 instead of I222) with no change to any other crystal in the test battery, including the twins that must stay in their lower symmetry. * Docs: Document the French-Wilson amplitude estimation, R-free flagging, reference-based space-group/ambiguity resolution, and model-based validation/maps in CPU_DATA_ANALYSIS.md. * Frontend: The status-bar pill now shows a progress bar during detector calibration (previously only during measurement), and the calibration state and its button are labelled "Calibration"/"CALIBRATE" (the internal `Pedestal` state name is unchanged for back-compatibility).Reviewed-on: #70 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
This commit was merged in pull request #70.
This commit is contained in:
@@ -222,6 +222,14 @@ SearchSpaceGroupResult SearchSpaceGroup(
|
||||
}
|
||||
}
|
||||
|
||||
// Overlap guard (Stage A / correlation only): drop the extreme resolution-normalised-E tail, which
|
||||
// on a two-lattice crystal is the one-sided overlap contamination that poisons the operator CC.
|
||||
// See SearchSpaceGroupOptions::max_e_squared_for_cc. Absences (pass_absence) keep the full range.
|
||||
if (opt.max_e_squared_for_cc > 0.0)
|
||||
for (size_t i = 0; i < n; ++i)
|
||||
if (pass_cc[i] && Esq[i] > opt.max_e_squared_for_cc)
|
||||
pass_cc[i] = false;
|
||||
|
||||
std::unordered_map<HKLKey, int, HKLKeyHash> key_to_index;
|
||||
key_to_index.reserve(n * 2);
|
||||
for (size_t i = 0; i < n; ++i)
|
||||
@@ -374,7 +382,7 @@ SearchSpaceGroupResult SearchSpaceGroup(
|
||||
// has to grow to swallow it (mirroring the merge error model's b / ISa collapse). This isolates
|
||||
// the systematic part of the scatter, which the fixed-sigma chi^2 ratio cannot: a genuine but
|
||||
// imperfectly-scaled high-symmetry merge and a twin can share a chi^2 ratio (~2) yet differ
|
||||
// sharply here (cubic Ins_I_3 b x1.04 vs twin Ins_H_2 b x1.6).
|
||||
// sharply here (a genuine cubic step b x1.04 vs a merohedral twin b x1.6).
|
||||
auto merge_systematic_b = [&](const std::vector<gemmi::Op>& rotations) -> double {
|
||||
struct Acc { double sw = 0.0, swI = 0.0; int n = 0; };
|
||||
std::unordered_map<HKLKey, Acc, HKLKeyHash> grp;
|
||||
@@ -447,24 +455,24 @@ SearchSpaceGroupResult SearchSpaceGroup(
|
||||
for (const auto& c : pg_cands) {
|
||||
// A genuine symmetry operator merges equivalent reflections, so it barely changes the reduced
|
||||
// chi^2 relative to the best subgroup - across the whole rotation-test battery every correct
|
||||
// point group stays within ~1.7x, even on weak or badly-integrated data (F432 chi2_ref 8.3 ->
|
||||
// 1.15; Thau P41212 -> 1.71). A twin law or pseudo-symmetry forces non-equivalent reflections
|
||||
// together, so its ratio is markedly higher (Ins_H_2's twin 2-fold: R3 3.02 -> R32 6.07, ratio
|
||||
// 2.01). max_merge_chi2_ratio sits between the two. (An earlier log10(chi2_ref) widening
|
||||
// point group stays within ~1.7x, even on weak or badly-integrated data (a cubic F432 chi2_ref
|
||||
// 8.3 -> 1.15; a tetragonal P41212 -> 1.71). A twin law or pseudo-symmetry forces non-equivalent
|
||||
// reflections together, so its ratio is markedly higher (a merohedral twin 2-fold: R3 3.02 ->
|
||||
// R32 6.07, ratio 2.01). max_merge_chi2_ratio sits between the two. (An earlier log10(chi2_ref) widening
|
||||
// compensated for an under-calibrated error model that inflated real-symmetry ratios with data
|
||||
// weakness; the variance-floor fix removed that inflation, and the widening now only let the
|
||||
// twin through, so it is gone.)
|
||||
bool consistent = c.pg->rotations.empty() || !std::isfinite(c.chi2) ||
|
||||
!std::isfinite(chi2_ref) || c.chi2 <= chi2_ref * opt.max_merge_chi2_ratio;
|
||||
|
||||
// Rescue a genuine high-symmetry merge whose chi^2 lands just past the ratio bound because its
|
||||
// data are imperfectly scaled (see max_merge_chi2_rescue). The systematic-error test tells it
|
||||
// apart from a twin: promote only if the extra intensity-proportional error b, relative to the
|
||||
// largest confirmed subgroup (by rotation-set inclusion), stayed within max_systematic_b_ratio -
|
||||
// a genuine step barely moves it, a twin's balloons.
|
||||
if (!consistent && std::isfinite(c.chi2) && std::isfinite(chi2_ref)
|
||||
&& c.chi2 <= chi2_ref * opt.max_merge_chi2_rescue) {
|
||||
double parent_b = -1.0;
|
||||
// Systematic-error test vs the largest confirmed subgroup (by rotation-set inclusion): merging
|
||||
// under a genuine operator gains multiplicity without intensity-proportional disagreement, so the
|
||||
// merge error model's b barely moves; a merohedral twin forces non-equivalent reflections together
|
||||
// and b balloons. It both RESCUES a genuine step whose chi^2 drifts just past the ratio bound
|
||||
// (imperfectly scaled data) and VETOES a twin whose chi^2 now looks self-consistent but whose b
|
||||
// balloons - the chi^2 ratio alone no longer separates them.
|
||||
double parent_b = -1.0;
|
||||
if (!c.pg->rotations.empty()) {
|
||||
int parent_order = 0;
|
||||
for (const auto& s : pg_cands)
|
||||
if (s.order < c.order && s.order > parent_order
|
||||
@@ -473,9 +481,29 @@ SearchSpaceGroupResult SearchSpaceGroup(
|
||||
parent_order = s.order;
|
||||
parent_b = s.b_extra;
|
||||
}
|
||||
if (parent_b > 1e-4 && c.b_extra <= parent_b * opt.max_systematic_b_ratio)
|
||||
consistent = true;
|
||||
}
|
||||
// The chi^2 ratio is only trustworthy when the error model is calibrated. When even the best
|
||||
// subgroup's reduced chi^2 (chi2_ref) is far above 1 - weak, low-resolution data whose merged
|
||||
// sigmas are badly under-estimated - the ratio grows with point-group order for genuine high
|
||||
// symmetry too and wrongly rejects it (a true weak F432 reaches ratio ~14). The systematic-b test
|
||||
// re-fits its own error, so it stays valid under a broken sigma model: a genuine step's b barely
|
||||
// moves (b-ratio ~1) while a twin's balloons. So once chi2_ref shows the error model is unreliable,
|
||||
// a promotion is rescued on the b-test alone (subject to the balloon veto below); otherwise the
|
||||
// rescue is confined to the narrow chi^2 band just past the ratio bound.
|
||||
const bool miscalibrated = std::isfinite(chi2_ref) && chi2_ref > opt.chi2_ref_reliable;
|
||||
if (!consistent && parent_b > 1e-4 && c.b_extra <= parent_b * opt.max_systematic_b_ratio
|
||||
&& (miscalibrated || (std::isfinite(c.chi2) && std::isfinite(chi2_ref)
|
||||
&& c.chi2 <= chi2_ref * opt.max_merge_chi2_rescue)))
|
||||
consistent = true;
|
||||
// Veto a chi^2-passing promotion whose b clearly ballooned (above the largest genuine step, below a
|
||||
// twin); a genuine but imperfectly-scaled high-symmetry merge stays under the bound and is untouched.
|
||||
// The parent b is floored (min_systematic_b_for_veto) so a near-zero parent on excellent data cannot
|
||||
// fabricate a huge ratio out of a still-tiny absolute b (a genuine 422 at b=0.05 over a 222 parent at
|
||||
// b=0.008 is not a twin - a real twin drives b to ~0.19 regardless).
|
||||
if (consistent && parent_b > 1e-4
|
||||
&& c.b_extra > std::max(parent_b, opt.min_systematic_b_for_veto) * opt.max_systematic_b_veto)
|
||||
consistent = false;
|
||||
|
||||
if (!consistent)
|
||||
continue;
|
||||
if (c.order > best_pg_order || (c.order == best_pg_order && c.min_class_cc > best_pg_min_cc)) {
|
||||
@@ -517,7 +545,9 @@ SearchSpaceGroupResult SearchSpaceGroup(
|
||||
// a large, correct centering-absent set hide a few strong screw violations and over-claim
|
||||
// screw axes (e.g. I4_132 on I432 data).
|
||||
int centering_absent = 0, centering_violations = 0;
|
||||
double centering_absent_sum = 0;
|
||||
int screw_absent = 0, screw_violations = 0;
|
||||
int present_strong = 0;
|
||||
|
||||
for (size_t i = 0; i < n; ++i) {
|
||||
if (!pass_absence[i])
|
||||
@@ -534,6 +564,7 @@ SearchSpaceGroupResult SearchSpaceGroup(
|
||||
s.absent_observed += 1;
|
||||
absent_sum += IoverSigma[i];
|
||||
centering_absent += 1;
|
||||
centering_absent_sum += IoverSigma[i];
|
||||
if (present) { s.absent_violations += 1; centering_violations += 1; }
|
||||
} else if (gops.is_systematically_absent(hkl)) {
|
||||
s.absent_observed += 1;
|
||||
@@ -543,6 +574,7 @@ SearchSpaceGroupResult SearchSpaceGroup(
|
||||
} else {
|
||||
present_n += 1;
|
||||
present_sum += IoverSigma[i];
|
||||
if (present) present_strong += 1;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -551,8 +583,37 @@ SearchSpaceGroupResult SearchSpaceGroup(
|
||||
if (present_n > 0)
|
||||
s.present_mean_i_over_sigma = present_sum / present_n;
|
||||
|
||||
const bool centering_ok = centering_absent == 0 ||
|
||||
centering_violations <= opt.max_absent_violation_fraction * centering_absent;
|
||||
// Centering is judged by class STRENGTH, not a per-reflection violation count. A real centering
|
||||
// cancels structure factors, so its absent class is systematically weak - its mean signed
|
||||
// I/sigma sits well below the present class - regardless of noise or obverse/reverse twinning;
|
||||
// a false centering leaves the "absent" class as strong as the present one (mean ratio ~1). The
|
||||
// count-of-strong-violations gate is brittle on noisy/twinned data, where enough genuinely-absent
|
||||
// reflections randomly clear I/sigma>3 to trip the 10% bound though the class is 3-4x weaker (a
|
||||
// true R3 at 13.5% violations, absent 1.7 vs present 6.0). The mean is well-determined here
|
||||
// because a centering-absent class holds a third-to-half of all reflections. Screws keep the
|
||||
// count gate: their predicted-absent class is a handful of axial reflections, too few to average.
|
||||
const double present_mean = present_n > 0 ? present_sum / present_n : 0.0;
|
||||
const double centering_absent_mean =
|
||||
centering_absent > 0 ? centering_absent_sum / centering_absent : 0.0;
|
||||
// The centering-absent class proves itself weak in either of two floor-independent ways; a
|
||||
// FALSE centering (absent as strong as present) fails both:
|
||||
// (1) mean signed I/sigma well below the present class, OR
|
||||
// (2) its strong-reflection RATE well below the present class's own strong rate.
|
||||
// (2) is needed because weak / low-energy data carry a positive intensity floor (background /
|
||||
// profile leakage) that lifts <I/s>abs to ~1.5-2.3 even for genuinely extinct reflections; when
|
||||
// the present class is itself weak (small present_mean) that additive floor inflates the mean
|
||||
// ratio past the bound and hides a real centering - e.g. an I-centred cubic crystal at low
|
||||
// energy, whose true I-centering sat at ratio ~0.57. Normalising the violation count by the
|
||||
// present class's own strong rate cancels the shared floor and stays reliable on weak data
|
||||
// (both rates shrink together).
|
||||
const double present_strong_rate =
|
||||
present_n > 0 ? static_cast<double>(present_strong) / present_n : 0.0;
|
||||
const double centering_violation_rate =
|
||||
centering_absent > 0 ? static_cast<double>(centering_violations) / centering_absent : 0.0;
|
||||
const bool centering_ok = centering_absent == 0
|
||||
|| (present_n > 0 && centering_absent_mean <= opt.max_absent_present_ratio * present_mean)
|
||||
|| (present_strong_rate > 0.0
|
||||
&& centering_violation_rate <= opt.max_absent_present_ratio * present_strong_rate);
|
||||
const bool screw_ok = screw_absent == 0 ||
|
||||
screw_violations <= opt.max_absent_violation_fraction * screw_absent;
|
||||
s.consistent = centering_ok && screw_ok;
|
||||
@@ -561,8 +622,17 @@ SearchSpaceGroupResult SearchSpaceGroup(
|
||||
|
||||
// A candidate is eligible when its absences are confirmed and there are enough of them to
|
||||
// trust (the symmorphic group, with no absences, is always eligible as the fallback). Rank
|
||||
// eligible candidates by how many absences they explain - the screw/centering content that is
|
||||
// both real and maximal wins, instead of defaulting to the symmorphic group.
|
||||
// eligible candidates by how many absences they GENUINELY explain - absent_observed minus the
|
||||
// violations, not the gross count. A false super-centering over-claims: F222 on a C222 crystal
|
||||
// predicts every C absence (all genuinely weak) PLUS a block of C-present reflections it wrongly
|
||||
// calls absent, so its gross count is larger yet its net count only equals C222's. Its diluted
|
||||
// absent class (many true zeros + a strong block) also slips under the strength/rate gate, so the
|
||||
// gate cannot veto it alone; netting the violations puts the two level, and the fewer-violations
|
||||
// and lower-number tie-breaks then keep the honest, less-centred C222. The ranking is symmetric:
|
||||
// on a genuine F222 crystal F explains strictly more weak absences and still wins.
|
||||
auto net_absent = [](const SpaceGroupCandidateScore& s) {
|
||||
return s.absent_observed - s.absent_violations;
|
||||
};
|
||||
auto eligible = [&](const SpaceGroupCandidateScore& s) {
|
||||
return s.consistent && (s.absent_observed == 0 || s.absent_observed >= opt.min_absent_observed);
|
||||
};
|
||||
@@ -570,16 +640,24 @@ SearchSpaceGroupResult SearchSpaceGroup(
|
||||
[&](const SpaceGroupCandidateScore& a, const SpaceGroupCandidateScore& b) {
|
||||
if (eligible(a) != eligible(b))
|
||||
return eligible(a);
|
||||
if (a.absent_observed != b.absent_observed)
|
||||
return a.absent_observed > b.absent_observed;
|
||||
// Tie (e.g. I23 vs I2_13, indistinguishable by absences): lower space-group number.
|
||||
if (net_absent(a) != net_absent(b))
|
||||
return net_absent(a) > net_absent(b);
|
||||
if (a.absent_violations != b.absent_violations)
|
||||
return a.absent_violations < b.absent_violations; // prefer the honest, less over-claiming group
|
||||
// Genuinely indistinguishable (e.g. I23 vs I2_13, or an enantiomorphic pair): lower
|
||||
// space-group number is the representative.
|
||||
return a.space_group.number < b.space_group.number;
|
||||
});
|
||||
|
||||
if (!result.candidates.empty() && eligible(result.candidates.front())) {
|
||||
const int best_absent = result.candidates.front().absent_observed;
|
||||
// Alternatives are only the candidates with the SAME absence signature - identical absent AND
|
||||
// violation counts - as the winner: the enantiomorphic / origin-ambiguous partners the data
|
||||
// truly cannot separate. A super-centering that nets the same count but over-claims differs in
|
||||
// its violation count and is therefore not reported as an equal alternative.
|
||||
const int sel_absent = result.candidates.front().absent_observed;
|
||||
const int sel_violations = result.candidates.front().absent_violations;
|
||||
for (auto& s : result.candidates) {
|
||||
if (!eligible(s) || s.absent_observed != best_absent)
|
||||
if (!eligible(s) || s.absent_observed != sel_absent || s.absent_violations != sel_violations)
|
||||
continue;
|
||||
s.selected = true;
|
||||
if (!result.best_space_group.has_value())
|
||||
|
||||
Reference in New Issue
Block a user