Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 11m6s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 10m27s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 10m54s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 9m25s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 10m5s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m33s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 11m19s
Build Packages / build:rpm (rocky8) (push) Successful in 12m23s
Build Packages / build:rpm (rocky9) (push) Successful in 13m21s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m30s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m55s
Build Packages / DIALS test (push) Successful in 13m42s
Build Packages / XDS test (durin plugin) (push) Successful in 9m26s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 6m41s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m12s
Build Packages / Generate python client (push) Successful in 19s
Build Packages / Build documentation (push) Successful in 52s
Build Packages / Create release (push) Skipped
Build Packages / build:viewer-tgz:cpu (push) Successful in 5m29s
Build Packages / build:viewer-tgz:cuda (push) Successful in 6m12s
Build Packages / build:windows:cuda (push) Successful in 18m36s
This is an UNSTABLE release. It includes many experimental features, as well as many AI generated fixes. We recommend using rc.152 for production use. * rugnux: Add `--model model.pdb` - score the merged data against an atomic model and compute initial maps. It reports R-work/R-free (scaling the model to the observed amplitudes with an overall scale, an anisotropic B and a flat bulk solvent - the standard few-parameter model, so a batch of maps stays directly comparable) and writes 2Fo-Fc / Fo-Fc electron-density maps (CCP4) plus a map-coefficient MTZ. The structure itself is not refined; the model is only re-fractionalised into the data cell. * rugnux: The merged reflection output now carries French-Wilson amplitudes (|F| and its sigma) next to the intensities - MTZ `F`/`SIGF`, mmCIF `_refln.F_meas_au`, and the text HKL - computed with the correct centric/acentric Wilson prior and epsilon multiplicity, so a downstream program (e.g. phenix.refine) can refine against amplitudes. The intensity columns are unchanged. * rugnux: R-free test-set flags are now assigned deterministically and consistently across symmetry - a Bijvoet pair I(+)/I(-) is never split between the work and free sets, and the assignment is a reproducible per-hkl hash that depends only on the reflection index, so every dataset of one crystal form gets the same ~5% free set (what a multi-dataset campaign such as PanDDA needs). On small data the fraction is floored so the test set stays large enough for a stable R-free (~500 reflections, capped at 10%); it stays flat at 5% on ordinary data. When a reference MTZ carries a `FreeR_flag` column its test set is imported instead, letting a whole campaign inherit one shared free set. * rugnux: A reference MTZ (`--reference-mtz`) can now fix the space group and cell for rotation data too (previously rejected), without being used to scale - the rotation merge stays self-consistent. When the crystal has an indexing (merohedral) ambiguity - a lattice symmetry higher than its Laue symmetry, e.g. P3/P4/P6/C2 - the reference also resolves it: each candidate reindexing (identity plus the twin-law cosets of the metric symmetry) is scored by its intensity correlation against the reference and the data are re-merged in the best-correlating one. This is a metric-preserving relabelling of hkl (the cell is unchanged) and a no-op for a holohedral crystal such as lysozyme. * rugnux: `--model` validation now aligns the data to the model before scoring - the observed reflections are reindexed into the model's enantiomorph when the two differ only by hand (indistinguishable from merged intensities). A merohedral indexing ambiguity is resolved against the reference MTZ when one is given (so a whole campaign shares one indexing convention); only with a model and no reference does validation fall back to fitting each candidate reindexing and keeping the lowest R-free. * rugnux: De-novo symmetry - recover a genuine high-symmetry group whose data are imperfectly scaled. Such a merge's within-orbit chi² lands just past the self-consistency bound (each real symmetry step adds a little systematic scatter), right where a merohedral twin also lands, so the chi² ratio alone cannot separate them. The candidate is now rescued when the extra intensity-proportional systematic error it invokes stays small relative to the confirmed subgroup - a genuine symmetry step gains multiplicity without inflating the merge error model's b, whereas a twin forces non-equivalent reflections together and b balloons. Fixes cubic insulin (I23 instead of I222) with no change to any other crystal in the test battery, including the twins that must stay in their lower symmetry. * Docs: Document the French-Wilson amplitude estimation, R-free flagging, reference-based space-group/ambiguity resolution, and model-based validation/maps in CPU_DATA_ANALYSIS.md. * Frontend: The status-bar pill now shows a progress bar during detector calibration (previously only during measurement), and the calibration state and its button are labelled "Calibration"/"CALIBRATE" (the internal `Pedestal` state name is unchanged for back-compatibility).Reviewed-on: #69 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
180 lines
6.6 KiB
C++
180 lines
6.6 KiB
C++
// Copyright 2019 Global Phasing Ltd.
|
|
//
|
|
// Searching for links based on the _chem_link table from monomer dictionary.
|
|
|
|
#ifndef GEMMI_LINKHUNT_HPP_
|
|
#define GEMMI_LINKHUNT_HPP_
|
|
|
|
#include <map>
|
|
#include "elem.hpp"
|
|
#include "model.hpp"
|
|
#include "monlib.hpp"
|
|
#include "neighbor.hpp"
|
|
#include "contact.hpp"
|
|
|
|
namespace gemmi {
|
|
|
|
struct LinkHunt {
|
|
struct Match {
|
|
const ChemLink* chem_link = nullptr;
|
|
int chem_link_count = 0;
|
|
int score = -1000;
|
|
CRA cra1;
|
|
CRA cra2;
|
|
bool same_image;
|
|
double bond_length = 0;
|
|
Connection* conn = nullptr;
|
|
};
|
|
|
|
double global_max_dist = 2.34; // ZN-CYS
|
|
const MonLib* monlib_ptr = nullptr;
|
|
std::multimap<std::string, const ChemLink*> links;
|
|
|
|
void index_chem_links(const MonLib& monlib, bool use_alias=true) {
|
|
std::map<ChemComp::Group, std::map<std::string, std::vector<std::string>>> aliases;
|
|
if (use_alias)
|
|
for (const auto& iter : monlib.monomers)
|
|
for (const ChemComp::Aliasing& a : iter.second.aliases)
|
|
for (const std::pair<std::string, std::string>& r : a.related) {
|
|
const ChemComp::Group& gr = ChemComp::is_nucleotide_group(a.group) ? ChemComp::Group::DnaRna : a.group;
|
|
aliases[gr][r.second].push_back(r.first);
|
|
}
|
|
|
|
for (const auto& iter : monlib.links) {
|
|
const ChemLink& link = iter.second;
|
|
if (link.rt.bonds.empty())
|
|
continue;
|
|
if (link.rt.bonds.size() > 1)
|
|
fprintf(stderr, "Note: considering only the first bond in %s\n",
|
|
link.id.c_str());
|
|
if (link.side1.comp.empty() && link.side2.comp.empty())
|
|
if (link.side1.group == ChemComp::Group::Null ||
|
|
link.side2.group == ChemComp::Group::Null ||
|
|
link.id == "SS")
|
|
continue;
|
|
const Restraints::Bond& bond = link.rt.bonds[0];
|
|
if (bond.value > global_max_dist)
|
|
global_max_dist = bond.value;
|
|
links.emplace(bond.lexicographic_str(), &link);
|
|
|
|
if (!use_alias || (!link.side1.comp.empty() && !link.side2.comp.empty()))
|
|
continue;
|
|
std::vector<std::string> *names1 = nullptr, *names2 = nullptr;
|
|
if (link.side1.comp.empty()) {
|
|
auto i = aliases.find(link.side1.group);
|
|
if (i != aliases.end()) {
|
|
auto j = i->second.find(bond.id1.atom);
|
|
if (j != i->second.end())
|
|
names1 = &j->second;
|
|
}
|
|
}
|
|
if (link.side2.comp.empty()) {
|
|
auto i = aliases.find(link.side2.group);
|
|
if (i != aliases.end()) {
|
|
auto j = i->second.find(bond.id2.atom);
|
|
if (j != i->second.end())
|
|
names2 = &j->second;
|
|
}
|
|
}
|
|
if (names1 && names2)
|
|
for (const std::string& n1 : *names1)
|
|
for (const std::string& n2 : *names2)
|
|
links.emplace(Restraints::lexicographic_str(n1, n2), &link);
|
|
else if (names1 || names2) {
|
|
const std::string& n1 = names1 ? bond.id2.atom : bond.id1.atom;
|
|
for (const std::string& n2 : (names1 ? *names1 : *names2))
|
|
links.emplace(Restraints::lexicographic_str(n1, n2), &link);
|
|
}
|
|
}
|
|
monlib_ptr = &monlib;
|
|
}
|
|
|
|
std::vector<Match> find_possible_links(Structure& st,
|
|
double bond_margin,
|
|
double radius_margin,
|
|
ContactSearch::Ignore ignore) {
|
|
std::vector<Match> results;
|
|
Model& model = st.first_model();
|
|
double search_radius = std::max(global_max_dist * bond_margin,
|
|
/*max r1+r2 ~=*/3.0 * radius_margin);
|
|
NeighborSearch ns(model, st.cell, std::max(5.0, search_radius));
|
|
ns.populate();
|
|
|
|
ContactSearch contacts((float) search_radius);
|
|
contacts.ignore = ignore;
|
|
contacts.for_each_contact(ns, [&](const CRA& cra1, const CRA& cra2,
|
|
int image_idx, double dist_sq) {
|
|
Match match;
|
|
|
|
// search for a match in chem_links
|
|
if (bond_margin > 0) {
|
|
auto range = links.equal_range(Restraints::lexicographic_str(
|
|
cra1.atom->name, cra2.atom->name));
|
|
// similar to MonLib::match_link()
|
|
for (auto iter = range.first; iter != range.second; ++iter) {
|
|
const ChemLink& link = *iter->second;
|
|
const Restraints::Bond& bond = link.rt.bonds[0];
|
|
if (dist_sq > sq(bond.value * bond_margin))
|
|
continue;
|
|
const ChemComp::Aliasing* aliasing1 = nullptr;
|
|
const ChemComp::Aliasing* aliasing2 = nullptr;
|
|
bool order1;
|
|
if (monlib_ptr->link_side_matches_residue(link.side1, cra1.residue->name, &aliasing1) &&
|
|
monlib_ptr->link_side_matches_residue(link.side2, cra2.residue->name, &aliasing2) &&
|
|
atom_match_with_alias(bond.id1.atom, cra1.atom->name, aliasing1))
|
|
order1 = true;
|
|
else if (monlib_ptr->link_side_matches_residue(link.side2, cra1.residue->name, &aliasing1) &&
|
|
monlib_ptr->link_side_matches_residue(link.side1, cra2.residue->name, &aliasing2) &&
|
|
atom_match_with_alias(bond.id2.atom, cra1.atom->name, aliasing1))
|
|
order1 = false;
|
|
else
|
|
continue;
|
|
int link_score = link.calculate_score(
|
|
order1 ? *cra1.residue : *cra2.residue,
|
|
order1 ? cra2.residue : cra1.residue,
|
|
order1 ? cra1.atom->altloc : cra2.atom->altloc,
|
|
order1 ? cra2.atom->altloc : cra1.atom->altloc,
|
|
order1 ? aliasing1 : aliasing2,
|
|
order1 ? aliasing2 : aliasing1);
|
|
match.chem_link_count++;
|
|
if (link_score > match.score) {
|
|
match.chem_link = &link;
|
|
match.score = link_score;
|
|
if (order1) {
|
|
match.cra1 = cra1;
|
|
match.cra2 = cra2;
|
|
} else {
|
|
match.cra1 = cra2;
|
|
match.cra2 = cra1;
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
// potential other links according to covalent radii
|
|
if (!match.chem_link) {
|
|
float r1 = cra1.atom->element.covalent_r();
|
|
float r2 = cra2.atom->element.covalent_r();
|
|
if (dist_sq > sq((r1 + r2) * radius_margin))
|
|
return;
|
|
match.cra1 = cra1;
|
|
match.cra2 = cra2;
|
|
}
|
|
|
|
// finalize
|
|
match.same_image = !image_idx;
|
|
match.bond_length = std::sqrt(dist_sq);
|
|
results.push_back(match);
|
|
});
|
|
|
|
// add references to st.connections
|
|
for (Match& match : results)
|
|
match.conn = st.find_connection_by_cra(match.cra1, match.cra2);
|
|
|
|
return results;
|
|
}
|
|
};
|
|
|
|
} // namespace gemmi
|
|
#endif
|