Files
Jungfraujoch/gemmi_gph/gemmi/binner.hpp
T
leonarski_f dd0bffb283
Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 11m6s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 10m27s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 10m54s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 9m25s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 10m5s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m33s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 11m19s
Build Packages / build:rpm (rocky8) (push) Successful in 12m23s
Build Packages / build:rpm (rocky9) (push) Successful in 13m21s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m30s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m55s
Build Packages / DIALS test (push) Successful in 13m42s
Build Packages / XDS test (durin plugin) (push) Successful in 9m26s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 6m41s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m12s
Build Packages / Generate python client (push) Successful in 19s
Build Packages / Build documentation (push) Successful in 52s
Build Packages / Create release (push) Skipped
Build Packages / build:viewer-tgz:cpu (push) Successful in 5m29s
Build Packages / build:viewer-tgz:cuda (push) Successful in 6m12s
Build Packages / build:windows:cuda (push) Successful in 18m36s
v1.0.0-rc.159 (#69)
This is an UNSTABLE release. It includes many experimental features, as well as many AI generated fixes. We recommend using rc.152 for production use.

* rugnux: Add `--model model.pdb` - score the merged data against an atomic model and compute initial maps. It reports R-work/R-free (scaling the model to the observed amplitudes with an overall scale, an anisotropic B and a flat bulk solvent - the standard few-parameter model, so a batch of maps stays directly comparable) and writes 2Fo-Fc / Fo-Fc electron-density maps (CCP4) plus a map-coefficient MTZ. The structure itself is not refined; the model is only re-fractionalised into the data cell.
* rugnux: The merged reflection output now carries French-Wilson amplitudes (|F| and its sigma) next to the intensities - MTZ `F`/`SIGF`, mmCIF `_refln.F_meas_au`, and the text HKL - computed with the correct centric/acentric Wilson prior and epsilon multiplicity, so a downstream program (e.g. phenix.refine) can refine against amplitudes. The intensity columns are unchanged.
* rugnux: R-free test-set flags are now assigned deterministically and consistently across symmetry - a Bijvoet pair I(+)/I(-) is never split between the work and free sets, and the assignment is a reproducible per-hkl hash that depends only on the reflection index, so every dataset of one crystal form gets the same ~5% free set (what a multi-dataset campaign such as PanDDA needs). On small data the fraction is floored so the test set stays large enough for a stable R-free (~500 reflections, capped at 10%); it stays flat at 5% on ordinary data. When a reference MTZ carries a `FreeR_flag` column its test set is imported instead, letting a whole campaign inherit one shared free set.
* rugnux: A reference MTZ (`--reference-mtz`) can now fix the space group and cell for rotation data too (previously rejected), without being used to scale - the rotation merge stays self-consistent. When the crystal has an indexing (merohedral) ambiguity - a lattice symmetry higher than its Laue symmetry, e.g. P3/P4/P6/C2 - the reference also resolves it: each candidate reindexing (identity plus the twin-law cosets of the metric symmetry) is scored by its intensity correlation against the reference and the data are re-merged in the best-correlating one. This is a metric-preserving relabelling of hkl (the cell is unchanged) and a no-op for a holohedral crystal such as lysozyme.
* rugnux: `--model` validation now aligns the data to the model before scoring - the observed reflections are reindexed into the model's enantiomorph when the two differ only by hand (indistinguishable from merged intensities). A merohedral indexing ambiguity is resolved against the reference MTZ when one is given (so a whole campaign shares one indexing convention); only with a model and no reference does validation fall back to fitting each candidate reindexing and keeping the lowest R-free.
* rugnux: De-novo symmetry - recover a genuine high-symmetry group whose data are imperfectly scaled. Such a merge's within-orbit chi² lands just past the self-consistency bound (each real symmetry step adds a little systematic scatter), right where a merohedral twin also lands, so the chi² ratio alone cannot separate them. The candidate is now rescued when the extra intensity-proportional systematic error it invokes stays small relative to the confirmed subgroup - a genuine symmetry step gains multiplicity without inflating the merge error model's b, whereas a twin forces non-equivalent reflections together and b balloons. Fixes cubic insulin (I23 instead of I222) with no change to any other crystal in the test battery, including the twins that must stay in their lower symmetry.
* Docs: Document the French-Wilson amplitude estimation, R-free flagging, reference-based space-group/ambiguity resolution, and model-based validation/maps in CPU_DATA_ANALYSIS.md.
* Frontend: The status-bar pill now shows a progress bar during detector calibration (previously only during measurement), and the calibration state and its button are labelled "Calibration"/"CALIBRATE" (the internal `Pedestal` state name is unchanged for back-compatibility).Reviewed-on: #69

Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-07-13 13:54:03 +02:00

179 lines
5.4 KiB
C++

// Copyright 2022 Global Phasing Ltd.
//
// Binning - resolution shells for reflections.
#ifndef GEMMI_BINNER_HPP_
#define GEMMI_BINNER_HPP_
#include <vector>
#include <limits> // for numeric_limits
#include "unitcell.hpp" // for UnitCell
namespace gemmi {
struct Binner {
enum class Method {
EqualCount,
Dstar,
Dstar2,
Dstar3,
};
void setup_from_1_d2(int nbins, Method method, std::vector<double>&& inv_d2,
const UnitCell* cell_) {
if (nbins < 1)
fail("Binner: nbins argument must be positive");
if (inv_d2.empty())
fail("Binner: no data");
if (cell_)
cell = *cell_;
if (!cell.is_crystal())
fail("Binner: unknown unit cell");
// first setup 2N bins to get both bin limits and middle points
limits.resize(2 * nbins);
if (method == Method::EqualCount) {
std::sort(inv_d2.begin(), inv_d2.end());
min_1_d2 = inv_d2.front();
max_1_d2 = inv_d2.back();
} else {
min_1_d2 = max_1_d2 = inv_d2.front();
for (double x : inv_d2) {
if (x < min_1_d2)
min_1_d2 = x;
if (x > max_1_d2)
max_1_d2 = x;
}
}
switch (method) {
case Method::EqualCount: {
double avg_count = double(inv_d2.size()) / limits.size();
for (size_t i = 1; i < limits.size(); ++i)
limits[i-1] = inv_d2[int(avg_count * i)];
break;
}
case Method::Dstar2: {
double step = (max_1_d2 - min_1_d2) / limits.size();
for (size_t i = 1; i < limits.size(); ++i)
limits[i-1] = min_1_d2 + i * step;
break;
}
case Method::Dstar: {
double min_1_d = std::sqrt(min_1_d2);
double max_1_d = std::sqrt(max_1_d2);
double step = (max_1_d - min_1_d) / limits.size();
for (size_t i = 1; i < limits.size(); ++i)
limits[i-1] = sq(min_1_d + i * step);
break;
}
case Method::Dstar3: {
double min_1_d3 = min_1_d2 * std::sqrt(min_1_d2);
double max_1_d3 = max_1_d2 * std::sqrt(max_1_d2);
double step = (max_1_d3 - min_1_d3) / limits.size();
for (size_t i = 1; i < limits.size(); ++i)
limits[i-1] = sq(std::cbrt(min_1_d3 + i * step));
break;
}
}
limits.back() = std::numeric_limits<double>::infinity();
mids.resize(nbins);
for (int i = 0; i < nbins; ++i) {
mids[i] = limits[2*i];
limits[i] = limits[2*i+1];
}
limits.resize(nbins);
}
template<typename DataProxy>
void setup(int nbins, Method method, const DataProxy& proxy,
const UnitCell* cell_=nullptr, size_t col_idx=0) {
if (col_idx >= proxy.stride())
fail("wrong col_idx in Binner::setup()");
cell = cell_ ? *cell_ : proxy.unit_cell();
std::vector<double> inv_d2;
inv_d2.reserve(proxy.size() / proxy.stride());
for (size_t offset = 0; offset < proxy.size(); offset += proxy.stride())
if (col_idx == 0 || !std::isnan(proxy.get_num(offset + col_idx)))
inv_d2.push_back(cell.calculate_1_d2(proxy.get_hkl(offset)));
setup_from_1_d2(nbins, method, std::move(inv_d2), nullptr);
}
void ensure_limits_are_set() const {
if (limits.empty())
fail("Binner not set up");
}
// Generic. Method-specific versions could be faster.
int get_bin_from_1_d2(double inv_d2) {
ensure_limits_are_set();
auto it = std::lower_bound(limits.begin(), limits.end(), inv_d2);
// it can't be limits.end() b/c limits.back() is +inf
return int(it - limits.begin());
}
int get_bin(const Miller& hkl) {
double inv_d2 = cell.calculate_1_d2(hkl);
return get_bin_from_1_d2(inv_d2);
}
// We assume that bins are seeked mostly for sorted reflections,
// so it's usually either the same bin as previously, or the next one.
int get_bin_from_1_d2_hinted(double inv_d2, int& hint) const {
if (inv_d2 <= limits[hint]) {
while (hint != 0 && limits[hint-1] > inv_d2)
--hint;
} else {
// limits.back() is +inf, so we won't overrun
while (limits[hint] < inv_d2)
++hint;
}
return hint;
}
int get_bin_hinted(const Miller& hkl, int& hint) const {
double inv_d2 = cell.calculate_1_d2(hkl);
return get_bin_from_1_d2_hinted(inv_d2, hint);
}
template<typename DataProxy>
std::vector<int> get_bins(const DataProxy& proxy) const {
ensure_limits_are_set();
int hint = 0;
std::vector<int> nums(proxy.size() / proxy.stride());
for (size_t i = 0, offset = 0; i < nums.size(); ++i, offset += proxy.stride())
nums[i] = get_bin_hinted(proxy.get_hkl(offset), hint);
return nums;
}
std::vector<int> get_bins_from_1_d2(const double* inv_d2, size_t size) const {
ensure_limits_are_set();
int hint = 0;
std::vector<int> nums(size);
for (size_t i = 0; i < size; ++i)
nums[i] = get_bin_from_1_d2_hinted(inv_d2[i], hint);
return nums;
}
std::vector<int> get_bins_from_1_d2(const std::vector<double>& inv_d2) const {
return get_bins_from_1_d2(inv_d2.data(), inv_d2.size());
}
double dmin_of_bin(int n) const {
return 1. / std::sqrt(n == (int) size() - 1 ? max_1_d2 : limits.at(n));
}
double dmax_of_bin(int n) const {
return 1. / std::sqrt(n == 0 ? min_1_d2 : limits.at(n-1));
}
size_t size() const { return limits.size(); }
UnitCell cell;
double min_1_d2;
double max_1_d2;
std::vector<double> limits; // upper limit of each bin
std::vector<double> mids; // the middle of each bin
};
} // namespace gemmi
#endif