Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 11m6s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 10m27s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 10m54s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 9m25s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 10m5s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m33s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 11m19s
Build Packages / build:rpm (rocky8) (push) Successful in 12m23s
Build Packages / build:rpm (rocky9) (push) Successful in 13m21s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m30s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m55s
Build Packages / DIALS test (push) Successful in 13m42s
Build Packages / XDS test (durin plugin) (push) Successful in 9m26s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 6m41s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m12s
Build Packages / Generate python client (push) Successful in 19s
Build Packages / Build documentation (push) Successful in 52s
Build Packages / Create release (push) Skipped
Build Packages / build:viewer-tgz:cpu (push) Successful in 5m29s
Build Packages / build:viewer-tgz:cuda (push) Successful in 6m12s
Build Packages / build:windows:cuda (push) Successful in 18m36s
This is an UNSTABLE release. It includes many experimental features, as well as many AI generated fixes. We recommend using rc.152 for production use. * rugnux: Add `--model model.pdb` - score the merged data against an atomic model and compute initial maps. It reports R-work/R-free (scaling the model to the observed amplitudes with an overall scale, an anisotropic B and a flat bulk solvent - the standard few-parameter model, so a batch of maps stays directly comparable) and writes 2Fo-Fc / Fo-Fc electron-density maps (CCP4) plus a map-coefficient MTZ. The structure itself is not refined; the model is only re-fractionalised into the data cell. * rugnux: The merged reflection output now carries French-Wilson amplitudes (|F| and its sigma) next to the intensities - MTZ `F`/`SIGF`, mmCIF `_refln.F_meas_au`, and the text HKL - computed with the correct centric/acentric Wilson prior and epsilon multiplicity, so a downstream program (e.g. phenix.refine) can refine against amplitudes. The intensity columns are unchanged. * rugnux: R-free test-set flags are now assigned deterministically and consistently across symmetry - a Bijvoet pair I(+)/I(-) is never split between the work and free sets, and the assignment is a reproducible per-hkl hash that depends only on the reflection index, so every dataset of one crystal form gets the same ~5% free set (what a multi-dataset campaign such as PanDDA needs). On small data the fraction is floored so the test set stays large enough for a stable R-free (~500 reflections, capped at 10%); it stays flat at 5% on ordinary data. When a reference MTZ carries a `FreeR_flag` column its test set is imported instead, letting a whole campaign inherit one shared free set. * rugnux: A reference MTZ (`--reference-mtz`) can now fix the space group and cell for rotation data too (previously rejected), without being used to scale - the rotation merge stays self-consistent. When the crystal has an indexing (merohedral) ambiguity - a lattice symmetry higher than its Laue symmetry, e.g. P3/P4/P6/C2 - the reference also resolves it: each candidate reindexing (identity plus the twin-law cosets of the metric symmetry) is scored by its intensity correlation against the reference and the data are re-merged in the best-correlating one. This is a metric-preserving relabelling of hkl (the cell is unchanged) and a no-op for a holohedral crystal such as lysozyme. * rugnux: `--model` validation now aligns the data to the model before scoring - the observed reflections are reindexed into the model's enantiomorph when the two differ only by hand (indistinguishable from merged intensities). A merohedral indexing ambiguity is resolved against the reference MTZ when one is given (so a whole campaign shares one indexing convention); only with a model and no reference does validation fall back to fitting each candidate reindexing and keeping the lowest R-free. * rugnux: De-novo symmetry - recover a genuine high-symmetry group whose data are imperfectly scaled. Such a merge's within-orbit chi² lands just past the self-consistency bound (each real symmetry step adds a little systematic scatter), right where a merohedral twin also lands, so the chi² ratio alone cannot separate them. The candidate is now rescued when the extra intensity-proportional systematic error it invokes stays small relative to the confirmed subgroup - a genuine symmetry step gains multiplicity without inflating the merge error model's b, whereas a twin forces non-equivalent reflections together and b balloons. Fixes cubic insulin (I23 instead of I222) with no change to any other crystal in the test battery, including the twins that must stay in their lower symmetry. * Docs: Document the French-Wilson amplitude estimation, R-free flagging, reference-based space-group/ambiguity resolution, and model-based validation/maps in CPU_DATA_ANALYSIS.md. * Frontend: The status-bar pill now shows a progress bar during detector calibration (previously only during measurement), and the calibration state and its button are labelled "Calibration"/"CALIBRATE" (the internal `Pedestal` state name is unchanged for back-compatibility).Reviewed-on: #69 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
130 lines
4.4 KiB
C++
130 lines
4.4 KiB
C++
// Copyright 2020 Global Phasing Ltd.
|
|
//
|
|
// Generating biological assemblies by applying operations
|
|
// from struct Assembly to a Model.
|
|
// Includes chain (re)naming utilities.
|
|
|
|
#ifndef GEMMI_ASSEMBLY_HPP_
|
|
#define GEMMI_ASSEMBLY_HPP_
|
|
|
|
#include "model.hpp" // for Model
|
|
#include "util.hpp" // for in_vector
|
|
#include "logger.hpp" // for Logger
|
|
|
|
namespace gemmi {
|
|
|
|
enum class HowToNameCopiedChain { Short, AddNumber, Dup };
|
|
|
|
struct ChainNameGenerator {
|
|
HowToNameCopiedChain how;
|
|
std::vector<std::string> used_names;
|
|
|
|
ChainNameGenerator(HowToNameCopiedChain how_) : how(how_) {}
|
|
ChainNameGenerator(const Model& model, HowToNameCopiedChain how_) : how(how_) {
|
|
if (how != HowToNameCopiedChain::Dup)
|
|
for (const Chain& chain : model.chains)
|
|
used_names.push_back(chain.name);
|
|
}
|
|
bool try_add(const std::string& name) {
|
|
if (in_vector(name, used_names))
|
|
return false;
|
|
used_names.push_back(name);
|
|
return true;
|
|
}
|
|
std::string make_short_name(const std::string& preferred) {
|
|
static const char symbols[] = {
|
|
'A','B','C','D','E','F','G','H','I','J','K','L','M',
|
|
'N','O','P','Q','R','S','T','U','V','W','X','Y','Z',
|
|
'a','b','c','d','e','f','g','h','i','j','k','l','m',
|
|
'n','o','p','q','r','s','t','u','v','w','x','y','z',
|
|
'0','1','2','3','4','5','6','7','8','9'
|
|
};
|
|
if (try_add(preferred))
|
|
return preferred;
|
|
std::string name(1, 'A');
|
|
for (char symbol : symbols) {
|
|
name[0] = symbol;
|
|
if (try_add(name))
|
|
return name;
|
|
}
|
|
name += 'A';
|
|
for (char symbol1 : symbols) {
|
|
name[0] = symbol1;
|
|
for (char symbol2 : symbols) {
|
|
name[1] = symbol2;
|
|
if (try_add(name))
|
|
return name;
|
|
}
|
|
}
|
|
fail("run out of 1- and 2-letter chain names");
|
|
}
|
|
|
|
std::string make_name_with_numeric_postfix(const std::string& base, int n) {
|
|
std::string name = base;
|
|
name += std::to_string(n);
|
|
while (!try_add(name)) {
|
|
name.resize(base.size());
|
|
name += std::to_string(++n);
|
|
}
|
|
return name;
|
|
}
|
|
|
|
std::string make_new_name(const std::string& old, int n) {
|
|
switch (how) {
|
|
case HowToNameCopiedChain::Short: return make_short_name(old);
|
|
case HowToNameCopiedChain::AddNumber: return make_name_with_numeric_postfix(old, n);
|
|
case HowToNameCopiedChain::Dup: return old;
|
|
}
|
|
unreachable();
|
|
}
|
|
};
|
|
|
|
inline void ensure_unique_chain_name(const Model& model, Chain& chain) {
|
|
ChainNameGenerator namegen(HowToNameCopiedChain::Short);
|
|
for (const Chain& ch : model.chains)
|
|
if (&ch != &chain)
|
|
namegen.try_add(ch.name);
|
|
chain.name = namegen.make_short_name(chain.name);
|
|
}
|
|
|
|
GEMMI_DLL Model make_assembly(const Assembly& assembly, const Model& model,
|
|
HowToNameCopiedChain how, const Logger& logging);
|
|
|
|
inline Assembly pseudo_assembly_for_unit_cell(const UnitCell& cell) {
|
|
Assembly assembly("unit_cell");
|
|
std::vector<Assembly::Operator> operators(cell.images.size() + 1);
|
|
// operators[0] stays as identity
|
|
for (size_t i = 1; i != operators.size(); ++i) {
|
|
const FTransform& op = cell.images[i-1];
|
|
operators[i].transform = cell.orth.combine(op.combine(cell.frac));
|
|
}
|
|
assembly.generators.push_back({{"(all)"}, {}, operators});
|
|
return assembly;
|
|
}
|
|
|
|
/// If called with assembly_name="unit_cell" changes structure to unit cell (P1).
|
|
/// \par keep_spacegroup preserves space group and unit cell - is it needed?
|
|
GEMMI_DLL void transform_to_assembly(Structure& st, const std::string& assembly_name,
|
|
HowToNameCopiedChain how, const Logger& logging,
|
|
bool keep_spacegroup=false, double merge_dist=0.2);
|
|
|
|
|
|
GEMMI_DLL Model expand_ncs_model(const Model& model, const std::vector<NcsOp>& ncs,
|
|
HowToNameCopiedChain how);
|
|
|
|
/// Searches and merges overlapping equivalent atoms from different chains.
|
|
/// To be used after expand_ncs() and make_assembly().
|
|
GEMMI_DLL void merge_atoms_in_expanded_model(Model& model, const UnitCell& cell,
|
|
double max_dist=0.2, bool compare_serial=true);
|
|
|
|
|
|
GEMMI_DLL void shorten_chain_names(Structure& st);
|
|
|
|
GEMMI_DLL void expand_ncs(Structure& st, HowToNameCopiedChain how, double merge_dist=0.2);
|
|
|
|
/// HowToNameCopiedChain::Dup adds segment name to chain name
|
|
GEMMI_DLL void split_chains_by_segments(Model& model, HowToNameCopiedChain how);
|
|
|
|
} // namespace gemmi
|
|
#endif
|