Files
Jungfraujoch/tests/SearchSpaceGroupTest.cpp
T
leonarski_f 67dca388bd
Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:cuda (push) Successful in 18m44s
Build Packages / build:viewer-tgz:cpu (push) Successful in 6m11s
Build Packages / build:viewer-tgz:cuda (push) Successful in 6m54s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 9m40s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 10m41s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 10m10s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 10m4s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m5s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 12m23s
Build Packages / build:rpm (rocky8) (push) Successful in 11m30s
Build Packages / build:rpm (rocky9) (push) Successful in 12m51s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m8s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m21s
Build Packages / DIALS test (push) Successful in 13m22s
Build Packages / XDS test (durin plugin) (push) Successful in 9m2s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 7m55s
Build Packages / XDS test (neggia plugin) (push) Successful in 5m57s
Build Packages / Generate python client (push) Successful in 23s
Build Packages / Build documentation (push) Successful in 57s
Build Packages / Create release (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 10m24s
v1.0.0-rc.160 (#70)
This is an UNSTABLE release. It includes many experimental features, as well as many AI generated fixes. We recommend using rc.152 for production use.

* rugnux: Add `--model model.pdb` - score the merged data against an atomic model and compute initial maps. It reports R-work/R-free (scaling the model to the observed amplitudes with an overall scale, an anisotropic B and a flat bulk solvent - the standard few-parameter model, so a batch of maps stays directly comparable) and writes 2Fo-Fc / Fo-Fc electron-density maps (CCP4) plus a map-coefficient MTZ. The structure itself is not refined; the model is only re-fractionalised into the data cell.
* rugnux: The merged reflection output now carries French-Wilson amplitudes (|F| and its sigma) next to the intensities - MTZ `F`/`SIGF`, mmCIF `_refln.F_meas_au`, and the text HKL - computed with the correct centric/acentric Wilson prior and epsilon multiplicity, so a downstream program (e.g. phenix.refine) can refine against amplitudes. The intensity columns are unchanged.
* rugnux: R-free test-set flags are now assigned deterministically and consistently across symmetry - a Bijvoet pair I(+)/I(-) is never split between the work and free sets, and the assignment is a reproducible per-hkl hash that depends only on the reflection index, so every dataset of one crystal form gets the same ~5% free set (what a multi-dataset campaign such as PanDDA needs). On small data the fraction is floored so the test set stays large enough for a stable R-free (~500 reflections, capped at 10%); it stays flat at 5% on ordinary data. When a reference MTZ carries a `FreeR_flag` column its test set is imported instead, letting a whole campaign inherit one shared free set.
* rugnux: A reference MTZ (`--reference-mtz`) can now fix the space group and cell for rotation data too (previously rejected), without being used to scale - the rotation merge stays self-consistent. When the crystal has an indexing (merohedral) ambiguity - a lattice symmetry higher than its Laue symmetry, e.g. P3/P4/P6/C2 - the reference also resolves it: each candidate reindexing (identity plus the twin-law cosets of the metric symmetry) is scored by its intensity correlation against the reference and the data are re-merged in the best-correlating one. This is a metric-preserving relabelling of hkl (the cell is unchanged) and a no-op for a holohedral crystal such as lysozyme.
* rugnux: `--model` validation now aligns the data to the model before scoring - the observed reflections are reindexed into the model's enantiomorph when the two differ only by hand (indistinguishable from merged intensities). A merohedral indexing ambiguity is resolved against the reference MTZ when one is given (so a whole campaign shares one indexing convention); only with a model and no reference does validation fall back to fitting each candidate reindexing and keeping the lowest R-free.
* rugnux: De-novo symmetry - recover a genuine high-symmetry group whose data are imperfectly scaled. Such a merge's within-orbit chi² lands just past the self-consistency bound (each real symmetry step adds a little systematic scatter), right where a merohedral twin also lands, so the chi² ratio alone cannot separate them. The candidate is now rescued when the extra intensity-proportional systematic error it invokes stays small relative to the confirmed subgroup - a genuine symmetry step gains multiplicity without inflating the merge error model's b, whereas a twin forces non-equivalent reflections together and b balloons. Fixes cubic insulin (I23 instead of I222) with no change to any other crystal in the test battery, including the twins that must stay in their lower symmetry.
* Docs: Document the French-Wilson amplitude estimation, R-free flagging, reference-based space-group/ambiguity resolution, and model-based validation/maps in CPU_DATA_ANALYSIS.md.
* Frontend: The status-bar pill now shows a progress bar during detector calibration (previously only during measurement), and the calibration state and its button are labelled "Calibration"/"CALIBRATE" (the internal `Pedestal` state name is unchanged for back-compatibility).Reviewed-on: #70

Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-07-19 09:39:28 +02:00

193 lines
7.0 KiB
C++

#include <catch2/catch_all.hpp>
#include "../image_analysis/scale_merge/SearchSpaceGroup.h"
#include "gemmi/symmetry.hpp"
#include <algorithm>
#include <cmath>
#include <cstdint>
#include <string>
#include <tuple>
#include <unordered_set>
#include <vector>
namespace {
struct HKL {
int h = 0;
int k = 0;
int l = 0;
bool operator==(const HKL& o) const noexcept {
return h == o.h && k == o.k && l == o.l;
}
};
struct HKLHash {
size_t operator()(const HKL& x) const noexcept {
auto mix = [](uint64_t v) {
v ^= v >> 33;
v *= 0xff51afd7ed558ccdULL;
v ^= v >> 33;
v *= 0xc4ceb9fe1a85ec53ULL;
v ^= v >> 33;
return v;
};
return static_cast<size_t>(
mix(static_cast<uint64_t>(x.h)) ^
(mix(static_cast<uint64_t>(x.k)) << 1) ^
(mix(static_cast<uint64_t>(x.l)) << 2));
}
};
double CalcSyntheticD(int h, int k, int l) {
const double q2 = static_cast<double>(h * h + k * k + l * l);
return 40.0 / std::sqrt(q2 + 1.0);
}
double SyntheticIntensityFromAsu(const gemmi::Op::Miller& asu) {
uint64_t x = static_cast<uint64_t>((asu[0] + 31) * 73856093u) ^
static_cast<uint64_t>((asu[1] + 37) * 19349663u) ^
static_cast<uint64_t>((asu[2] + 41) * 83492791u);
x ^= x >> 13;
x *= 0x9e3779b97f4a7c15ULL;
x ^= x >> 17;
return 100.0 + static_cast<double>(x % 500);
}
std::vector<MergedReflection> GenerateMergedReflectionsForSpaceGroup(
const gemmi::SpaceGroup& sg,
int hmax = 8) {
std::vector<MergedReflection> merged;
std::unordered_set<HKL, HKLHash> added;
const gemmi::GroupOps gops = sg.operations();
const gemmi::ReciprocalAsu rasu(&sg);
for (int h = -hmax; h <= hmax; ++h) {
for (int k = -hmax; k <= hmax; ++k) {
for (int l = -hmax; l <= hmax; ++l) {
if (h == 0 && k == 0 && l == 0)
continue;
bool absent = false;
gemmi::Op::Miller hkl{{h, k, l}};
if (gops.is_systematically_absent(hkl))
absent = true;
const auto [asu, sign_plus] = rasu.to_asu_sign(hkl, gops);
if (!sign_plus)
continue;
const HKL key{h, k, l};
if (added.find(key) != added.end())
continue;
added.insert(key);
merged.push_back(MergedReflection{
.h = h,
.k = k,
.l = l,
.I = absent ? 0.0 : SyntheticIntensityFromAsu(asu),
.sigma = 1.0,
.d = CalcSyntheticD(h, k, l)
});
}
}
}
return merged;
}
}
TEST_CASE("SearchSpaceGroup detects synthetic space groups") {
struct Case {
std::string input_name;
std::string expected_short_name;
};
const std::vector<Case> cases = {
{"P 1", "P1"},
{"P 1 2 1", "P2"},
{"P 3 2 1", "P321"},
{"P 4 2 2", "P422"},
{"P 4 3 2", "P432"},
{"P 43 21 2", "P43212"},
{"P 6 2 2", "P622"},
{"C 1 2 1", "C2"},
{"C 2 2 2", "C222"},
{"I 4 3 2", "I432"},
{"I 21 21 21", "I212121"},
{"I 2 1 3", "I213"},
};
for (const auto& tc : cases) {
DYNAMIC_SECTION(tc.expected_short_name) {
const gemmi::SpaceGroup& sg = gemmi::get_spacegroup_by_name(tc.input_name);
const auto merged = GenerateMergedReflectionsForSpaceGroup(sg);
SearchSpaceGroupOptions opt;
opt.merge_friedel = true;
const auto result = SearchSpaceGroup(merged, opt);
// Several inputs cannot be told apart from intensities alone: enantiomorphic partners
// (P4_3 vs P4_1) and origin-ambiguous pairs (I2_12_12_1 vs I222, I2_13 vs I2_3) share
// the same systematic absences. The search reports those as alternatives, so the
// expected group must appear among the best group and its alternatives.
std::vector<std::string> accepted;
if (result.best_space_group.has_value())
accepted.push_back(result.best_space_group->short_name());
for (const auto& alt : result.alternatives)
accepted.push_back(alt.short_name());
INFO(SearchSpaceGroupResultToText(result));
REQUIRE(result.best_space_group.has_value());
CHECK(std::find(accepted.begin(), accepted.end(), tc.expected_short_name) != accepted.end());
}
}
}
// Regression: a real screw axis whose systematically-absent reflections carry a genuinely weak
// intensity but an UNDER-estimated sigma (so their I/sigma clears the "present" cut) must still be
// found. Reproduces a monoclinic 2_1 miss on weakly-diffracting monoclinic data, where the merged sigmas on
// the 0k0-odd reflections were ~2x too small and faked screw-axis violations. The E^2 intensity gate
// (present_e_squared) is what keeps those reflections classified absent.
TEST_CASE("SearchSpaceGroup finds a screw axis despite under-estimated sigmas on absent reflections") {
const gemmi::SpaceGroup& sg = gemmi::get_spacegroup_by_name("P 1 21 1");
auto merged = GenerateMergedReflectionsForSpaceGroup(sg, 18);
// Every systematically-absent (0k0, k odd) reflection: small-but-nonzero intensity (~2% of a
// normal reflection) with a far-too-small sigma, so I/sigma ~ 27 fakes a "present" reflection.
const gemmi::GroupOps gops = sg.operations();
int absent_count = 0;
for (auto& r : merged) {
const gemmi::Op::Miller hkl{{r.h, r.k, r.l}};
if (gops.is_systematically_absent(hkl)) {
r.I = 8.0f;
r.sigma = 0.3f;
++absent_count;
}
}
REQUIRE(absent_count >= 8); // enough predicted-absent reflections to be trusted
SearchSpaceGroupOptions opt;
opt.merge_friedel = true;
SECTION("intensity gate on (default): screw recovered") {
const auto result = SearchSpaceGroup(merged, opt);
INFO(SearchSpaceGroupResultToText(result));
REQUIRE(result.best_space_group.has_value());
CHECK(result.best_space_group->short_name() == "P21");
}
SECTION("intensity gate off (I/sigma only): the screw is missed") {
// Documents the failure the gate fixes: with I/sigma alone the too-small sigmas fake
// violations and the search falls back to the symmorphic group.
opt.present_e_squared = 0.0;
const auto result = SearchSpaceGroup(merged, opt);
INFO(SearchSpaceGroupResultToText(result));
REQUIRE(result.best_space_group.has_value());
CHECK(result.best_space_group->short_name() == "P2");
}
}