Files
Jungfraujoch/gemmi_gph/gemmi/metadata.hpp
T
leonarski_f dd0bffb283
Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 11m6s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 10m27s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 10m54s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 9m25s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 10m5s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m33s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 11m19s
Build Packages / build:rpm (rocky8) (push) Successful in 12m23s
Build Packages / build:rpm (rocky9) (push) Successful in 13m21s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m30s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m55s
Build Packages / DIALS test (push) Successful in 13m42s
Build Packages / XDS test (durin plugin) (push) Successful in 9m26s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 6m41s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m12s
Build Packages / Generate python client (push) Successful in 19s
Build Packages / Build documentation (push) Successful in 52s
Build Packages / Create release (push) Skipped
Build Packages / build:viewer-tgz:cpu (push) Successful in 5m29s
Build Packages / build:viewer-tgz:cuda (push) Successful in 6m12s
Build Packages / build:windows:cuda (push) Successful in 18m36s
v1.0.0-rc.159 (#69)
This is an UNSTABLE release. It includes many experimental features, as well as many AI generated fixes. We recommend using rc.152 for production use.

* rugnux: Add `--model model.pdb` - score the merged data against an atomic model and compute initial maps. It reports R-work/R-free (scaling the model to the observed amplitudes with an overall scale, an anisotropic B and a flat bulk solvent - the standard few-parameter model, so a batch of maps stays directly comparable) and writes 2Fo-Fc / Fo-Fc electron-density maps (CCP4) plus a map-coefficient MTZ. The structure itself is not refined; the model is only re-fractionalised into the data cell.
* rugnux: The merged reflection output now carries French-Wilson amplitudes (|F| and its sigma) next to the intensities - MTZ `F`/`SIGF`, mmCIF `_refln.F_meas_au`, and the text HKL - computed with the correct centric/acentric Wilson prior and epsilon multiplicity, so a downstream program (e.g. phenix.refine) can refine against amplitudes. The intensity columns are unchanged.
* rugnux: R-free test-set flags are now assigned deterministically and consistently across symmetry - a Bijvoet pair I(+)/I(-) is never split between the work and free sets, and the assignment is a reproducible per-hkl hash that depends only on the reflection index, so every dataset of one crystal form gets the same ~5% free set (what a multi-dataset campaign such as PanDDA needs). On small data the fraction is floored so the test set stays large enough for a stable R-free (~500 reflections, capped at 10%); it stays flat at 5% on ordinary data. When a reference MTZ carries a `FreeR_flag` column its test set is imported instead, letting a whole campaign inherit one shared free set.
* rugnux: A reference MTZ (`--reference-mtz`) can now fix the space group and cell for rotation data too (previously rejected), without being used to scale - the rotation merge stays self-consistent. When the crystal has an indexing (merohedral) ambiguity - a lattice symmetry higher than its Laue symmetry, e.g. P3/P4/P6/C2 - the reference also resolves it: each candidate reindexing (identity plus the twin-law cosets of the metric symmetry) is scored by its intensity correlation against the reference and the data are re-merged in the best-correlating one. This is a metric-preserving relabelling of hkl (the cell is unchanged) and a no-op for a holohedral crystal such as lysozyme.
* rugnux: `--model` validation now aligns the data to the model before scoring - the observed reflections are reindexed into the model's enantiomorph when the two differ only by hand (indistinguishable from merged intensities). A merohedral indexing ambiguity is resolved against the reference MTZ when one is given (so a whole campaign shares one indexing convention); only with a model and no reference does validation fall back to fitting each candidate reindexing and keeping the lowest R-free.
* rugnux: De-novo symmetry - recover a genuine high-symmetry group whose data are imperfectly scaled. Such a merge's within-orbit chi² lands just past the self-consistency bound (each real symmetry step adds a little systematic scatter), right where a merohedral twin also lands, so the chi² ratio alone cannot separate them. The candidate is now rescued when the extra intensity-proportional systematic error it invokes stays small relative to the confirmed subgroup - a genuine symmetry step gains multiplicity without inflating the merge error model's b, whereas a twin forces non-equivalent reflections together and b balloons. Fixes cubic insulin (I23 instead of I222) with no change to any other crystal in the test battery, including the twins that must stay in their lower symmetry.
* Docs: Document the French-Wilson amplitude estimation, R-free flagging, reference-based space-group/ambiguity resolution, and model-based validation/maps in CPU_DATA_ANALYSIS.md.
* Frontend: The status-bar pill now shows a progress bar during detector calibration (previously only during measurement), and the calibration state and its button are labelled "Calibration"/"CALIBRATE" (the internal `Pedestal` state name is unchanged for back-compatibility).Reviewed-on: #69

Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-07-13 13:54:03 +02:00

391 lines
15 KiB
C++

// Copyright 2019 Global Phasing Ltd.
//
// Metadata from coordinate files.
#ifndef GEMMI_METADATA_HPP_
#define GEMMI_METADATA_HPP_
#include <cstdint> // for uint8_t, uint16_t
#include <algorithm> // for any_of
#include <string>
#include <vector>
#include "math.hpp" // for Mat33
#include "unitcell.hpp" // for Position, Asu
#include "seqid.hpp" // for SeqId
namespace gemmi {
// corresponds to the mmCIF _software category
struct SoftwareItem {
enum Classification {
DataCollection, DataExtraction, DataProcessing, DataReduction,
DataScaling, ModelBuilding, Phasing, Refinement, Unspecified
};
std::string name;
std::string version;
std::string date;
std::string description;
std::string contact_author;
std::string contact_author_email;
Classification classification = Unspecified;
};
// Information from REMARK 200/230 is significantly expanded in PDBx/mmCIF.
// These remarks corresponds to data across 12 mmCIF categories
// including categories _exptl, _reflns, _exptl_crystal, _diffrn and others.
// _exptl and _reflns seem to be 1:1. Usually we have one experiment (_exptl),
// except for a joint refinement (e.g. X-ray + neutron data).
// Both crystal (_exptl_crystal) and reflection statistics (_reflns) can
// be associated with multiple diffraction sets (_diffrn).
// But if we use the PDB format, only one diffraction set per method
// can be described.
struct ReflectionsInfo {
double resolution_high = NAN; // _reflns.d_resolution_high
// (or _reflns_shell.d_res_high)
double resolution_low = NAN; // _reflns.d_resolution_low
double completeness = NAN; // _reflns.percent_possible_obs
double redundancy = NAN; // _reflns.pdbx_redundancy
double r_merge = NAN; // _reflns.pdbx_Rmerge_I_obs
double r_sym = NAN; // _reflns.pdbx_Rsym_value
double mean_I_over_sigma = NAN; // _reflns.pdbx_netI_over_sigmaI
};
// _exptl has no id, _exptl.method is key item and must be unique
struct ExperimentInfo {
std::string method; // _exptl.method
int number_of_crystals = -1; // _exptl.crystals_number
int unique_reflections = -1; // _reflns.number_obs
ReflectionsInfo reflections;
double b_wilson = NAN; // _reflns.B_iso_Wilson_estimate
std::vector<ReflectionsInfo> shells;
std::vector<std::string> diffraction_ids;
};
struct DiffractionInfo {
std::string id; // _diffrn.id
double temperature = NAN; // _diffrn.ambient_temp
std::string source; // _diffrn_source.source
std::string source_type; // _diffrn_source.type
std::string synchrotron; // _diffrn_source.pdbx_synchrotron_site
std::string beamline; // _diffrn_source.pdbx_synchrotron_beamline
std::string wavelengths; // _diffrn_source.pdbx_wavelength
std::string scattering_type; // _diffrn_radiation.pdbx_scattering_type
char mono_or_laue = '\0'; // _diffrn_radiation.pdbx_monochromatic_or_laue_m_l
std::string monochromator; // _diffrn_radiation.monochromator
std::string collection_date; // _diffrn_detector.pdbx_collection_date
std::string optics; // _diffrn_detector.details
std::string detector; // _diffrn_detector.detector
std::string detector_make; // _diffrn_detector.type
};
struct CrystalInfo {
std::string id; // _exptl_crystal.id
std::string description; // _exptl_crystal.description
double ph = NAN; // _exptl_crystal_grow.pH
std::string ph_range; // _exptl_crystal_grow.pdbx_pH_range
std::vector<DiffractionInfo> diffractions;
};
struct TlsGroup {
struct Selection {
std::string chain;
SeqId res_begin;
SeqId res_end;
std::string details; // _pdbx_refine_tls_group.selection_details
};
short num_id = -1; // id stored as number (optimization)
std::string id; // _pdbx_refine_tls.id
std::vector<Selection> selections;
Position origin; // _pdbx_refine_tls.origin_x/y/z
SMat33<double> T = {NAN, NAN, NAN, NAN, NAN, NAN}; // _pdbx_refine_tls.T[][]
SMat33<double> L = {NAN, NAN, NAN, NAN, NAN, NAN}; // _pdbx_refine_tls.L[][]
Mat33 S = Mat33{NAN}; // _pdbx_refine_tls.S[][]
};
// RefinementInfo corresponds to REMARK 3.
// BasicRefinementInfo is used for both total and per-bin statistics.
// For per-bin data, each values corresponds to one _refine_ls_shell.* tag.
struct BasicRefinementInfo {
double resolution_high = NAN; // _refine.ls_d_res_high, _refine_ls_shell.d_res_high
double resolution_low = NAN; // _refine.ls_d_res_low, _refine_ls_shell.d_res_low
double completeness = NAN; // _refine.ls_percent_reflns_obs, _refine_ls_shell.percent...
int reflection_count = -1; // _refine.ls_number_reflns_obs, _refine_ls_shell.number...
int work_set_count = -1; // _refine.ls_number_reflns_R_work, _refine_ls_shell.number...
int rfree_set_count = -1; // _refine.ls_number_reflns_R_free, _refine_ls_shell.number...
double r_all = NAN; // _refine.ls_R_factor_obs, _refine_ls_shell.R_factor_obs
double r_work = NAN; // _refine.ls_R_factor_R_work, _refine_ls_shell.R_factor_R_work
double r_free = NAN; // _refine.ls_R_factor_R_free, _refine_ls_shell.R_factor_R_free
double cc_fo_fc_work = NAN; // _refine.correlation_coeff_Fo_to_Fc, _refine_ls_shell.corr...
double cc_fo_fc_free = NAN; // _refine.correlation_coeff_Fo_to_Fc_free, _refine_ls_shell.c...
double fsc_work = NAN; // _refine.pdbx_average_fsc_work, _refine_ls_shell.pdbx_fsc_work
double fsc_free = NAN; // _refine.pdbx_average_fsc_free, _refine_ls_shell.pdbx_fsc_free
double cc_intensity_work = NAN; // _refine.correlation_coeff_I_to_Fcsqd_work, ...
double cc_intensity_free = NAN; // _refine.correlation_coeff_I_to_Fcsqd_free, ...
};
struct RefinementInfo : BasicRefinementInfo {
struct Restr {
std::string name;
int count = -1;
double weight = NAN;
std::string function;
double dev_ideal = NAN;
Restr() = default;
explicit Restr(const std::string& name_) : name(name_) {}
};
std::string id;
std::string cross_validation_method; // _refine.pdbx_ls_cross_valid_method
std::string rfree_selection_method; // _refine.pdbx_R_Free_selection_details
int bin_count = -1; // _refine_ls_shell.pdbx_total_number_of_bins_used
std::vector<BasicRefinementInfo> bins;
double mean_b = NAN; // _refine.B_iso_mean
SMat33<double> aniso_b{NAN, NAN, NAN, NAN, NAN, NAN}; // _refine.aniso_B[][]
double luzzati_error = NAN; // _refine_analyze.Luzzati_coordinate_error_obs
double dpi_blow_r = NAN; // _refine.pdbx_overall_SU_R_Blow_DPI
double dpi_blow_rfree = NAN; // _refine.pdbx_overall_SU_R_free_Blow_DPI
double dpi_cruickshank_r = NAN; // _refine.overall_SU_R_Cruickshank_DPI
double dpi_cruickshank_rfree = NAN; // _refine.pdbx_overall_SU_R_free_Cruickshank_DPI
std::vector<Restr> restr_stats; // _refine_ls_restr
std::vector<TlsGroup> tls_groups; // _pdbx_refine_tls
std::string remarks;
};
struct Metadata {
std::vector<std::string> authors; // _audit_author.name
std::vector<ExperimentInfo> experiments;
std::vector<CrystalInfo> crystals;
std::vector<RefinementInfo> refinement;
std::vector<SoftwareItem> software;
std::string solved_by; // _refine.pdbx_method_to_determine_struct
std::string starting_model; // _refine.pdbx_starting_model
std::string remark_300_detail; // _struct_biol.details
bool has(double RefinementInfo::*field) const {
return std::any_of(refinement.begin(), refinement.end(),
[&](const RefinementInfo& r) { return !std::isnan(r.*field); });
}
bool has(int RefinementInfo::*field) const {
return std::any_of(refinement.begin(), refinement.end(),
[&](const RefinementInfo& r) { return r.*field != -1; });
}
bool has(std::string RefinementInfo::*field) const {
return std::any_of(refinement.begin(), refinement.end(),
[&](const RefinementInfo& r) { return !(r.*field).empty(); });
}
bool has(SMat33<double> RefinementInfo::*field) const {
return std::any_of(refinement.begin(), refinement.end(),
[&](const RefinementInfo& r) { return !std::isnan((r.*field).u11); });
}
bool has_restr() const {
return std::any_of(refinement.begin(), refinement.end(),
[&](const RefinementInfo& r) { return !r.restr_stats.empty(); });
}
// TLS constraint are not specific to refinement in joint refinement,
// so they are expected to be present only in a single RefinementInfo.
// As of 2025, two PDB entries have TLS + joint refinement: 6N3U and 5NKU.
std::vector<gemmi::TlsGroup>* get_tls_groups() {
for (gemmi::RefinementInfo& ref : refinement)
if (!ref.tls_groups.empty())
return &ref.tls_groups;
return nullptr;
}
const std::vector<gemmi::TlsGroup>* get_tls_groups() const {
return const_cast<Metadata*>(this)->get_tls_groups();
}
};
// Entity description.
//
// values corresponding to mmCIF _entity.type
enum class EntityType : unsigned char {
Unknown,
Polymer,
NonPolymer,
Branched, // introduced in 2020
// _entity.type macrolide is in PDBx/mmCIF, but no PDB entry uses it
//Macrolide,
Water
};
// values corresponding to mmCIF _entity_poly.type
enum class PolymerType : unsigned char {
Unknown, // unknown or not applicable
PeptideL, // polypeptide(L) in mmCIF (168923 values in the PDB in 2017)
PeptideD, // polypeptide(D) (57 values)
Dna, // polydeoxyribonucleotide (9905)
Rna, // polyribonucleotide (4559)
DnaRnaHybrid, // polydeoxyribonucleotide/polyribonucleotide hybrid (156)
SaccharideD, // polysaccharide(D) (18)
SaccharideL, // polysaccharide(L) (0)
Pna, // peptide nucleic acid (2)
CyclicPseudoPeptide, // cyclic-pseudo-peptide (1)
Other, // other (4)
};
inline bool is_polypeptide(PolymerType pt) {
return pt == PolymerType::PeptideL || pt == PolymerType::PeptideD;
}
inline bool is_polynucleotide(PolymerType pt) {
return pt == PolymerType::Dna || pt == PolymerType::Rna ||
pt == PolymerType::DnaRnaHybrid;
}
struct Entity {
struct DbRef {
std::string db_name;
std::string accession_code;
std::string id_code;
std::string isoform; // pdbx_db_isoform
SeqId seq_begin, seq_end;
SeqId db_begin, db_end;
SeqId::OptionalNum label_seq_begin, label_seq_end;
};
std::string name;
std::vector<std::string> subchains;
EntityType entity_type = EntityType::Unknown;
PolymerType polymer_type = PolymerType::Unknown;
// In case of microheterogeneity, PDB SEQRES has only the first residue name.
bool reflects_microhetero = false;
std::vector<DbRef> dbrefs;
/// List of SIFTS Uniprot ACs referenced by SiftsUnpResidue::acc_index
std::vector<std::string> sifts_unp_acc;
/// SEQRES or entity_poly_seq with microheterogeneity as comma-separated names
std::vector<std::string> full_sequence;
Entity() = default;
explicit Entity(const std::string& name_) noexcept : name(name_) {}
static std::string first_mon(const std::string& mon_list) {
return mon_list.substr(0, mon_list.find(','));
}
};
/// Reference to UniProt residue, based on _pdbx_sifts_xref_db.
/// Used in Residue::sifts_unp. res==0 <=> unset.
struct SiftsUnpResidue {
char res = '\0'; // _pdbx_sifts_xref_db.unp_res
std::uint8_t acc_index = 0; // index of Entity::sifts_unp_acc
std::uint16_t num = 0; // _pdbx_sifts_xref_db.unp_num
};
// A connection. Corresponds to _struct_conn.
// Symmetry operators are not trusted and not stored.
// We assume that the nearest symmetry mate is connected.
struct Connection {
// in write_struct_conn() we assume that Unknown is at the end
enum Type : unsigned char { Covale=0, Disulf, Hydrog, MetalC, Unknown };
std::string name;
std::string link_id; // _struct_conn.ccp4_link_id (== _chem_link.id)
Type type = Unknown;
Asu asu = Asu::Any;
AtomAddress partner1, partner2;
double reported_distance = 0.0;
short reported_sym[4] = {}; // don't rely on it, for internal use only
};
// Corresponds to CISPEP or _struct_mon_prot_cis
struct CisPep {
AtomAddress partner_c, partner_n;
int model_num = 0;
// mmCIF has (unused by the PDB) tag _struct_mon_prot_cis.label_alt_id
// that enables defining CIS link per conformation.
char only_altloc = '\0';
double reported_angle = NAN;
};
struct ModRes {
std::string chain_name;
ResidueId res_id;
std::string parent_comp_id;
std::string mod_id; // non-standard extension used in Refmac
std::string details;
};
// Secondary structure. PDBx/mmCIF stores helices and sheets separately.
// mmCIF spec defines 32 possible values for _struct_conf.conf_type_id -
// "the type of the conformation of the backbone of the polymer (whether
// protein or nucleic acid)". But as of 2019 only HELX_P is used (not counting
// TURN_P that occurs in only 6 entries). The actual helix type is given
// by numeric value of _struct_conf.pdbx_PDB_helix_class, which corresponds
// to helixClass from the PDB HELIX record. These values are in the range 1-10.
// As of 2019 it's almost only type 1 and 5:
// 3116566 of 1 - right-handed alpha
// 16 of 2 - right-handed omega
// 84 of 3 - right-handed pi
// 79 of 4 - right-handed gamma
// 1063337 of 5 - right-handed 3-10
// 27 of 6 - left-handed alpha
// 5 of 7 - left-handed omega
// 2 of 8 - left-handed gamma
// 8 of 9 - 2-7 ribbon/helix
// 46 of 10 - polyproline
struct Helix {
enum HelixClass {
UnknownHelix, RAlpha, ROmega, RPi, RGamma, R310,
LAlpha, LOmega, LGamma, Helix27, HelixPolyProlineNone
};
AtomAddress start, end;
HelixClass pdb_helix_class = UnknownHelix;
int length = -1;
void set_helix_class_as_int(int n) {
if (n >= 1 && n <= 10)
pdb_helix_class = static_cast<HelixClass>(n);
}
};
struct Sheet {
struct Strand {
AtomAddress start, end;
AtomAddress hbond_atom2, hbond_atom1;
int sense; // 0 = first strand, 1 = parallel, -1 = anti-parallel.
std::string name; // optional, _struct_sheet_range.id if from mmCIF
};
std::string name;
std::vector<Strand> strands;
Sheet() = default;
explicit Sheet(const std::string& sheet_id) noexcept : name(sheet_id) {}
};
// bioassembly / biomolecule
struct Assembly {
struct Operator {
std::string name; // optional
std::string type; // optional (from mmCIF only)
Transform transform;
};
struct Gen {
std::vector<std::string> chains;
std::vector<std::string> subchains;
std::vector<Operator> operators;
};
enum class SpecialKind : unsigned char {
NA, CompleteIcosahedral, RepresentativeHelical, CompletePoint
};
std::string name;
bool author_determined = false;
bool software_determined = false;
SpecialKind special_kind = SpecialKind::NA;
int oligomeric_count = 0;
std::string oligomeric_details;
std::string software_name;
double absa = NAN; // TOTAL BURIED SURFACE AREA: ... ANGSTROM**2
double ssa = NAN; // SURFACE AREA OF THE COMPLEX: ... ANGSTROM**2
double more = NAN; // CHANGE IN SOLVENT FREE ENERGY: ... KCAL/MOL
std::vector<Gen> generators;
Assembly() = default;
explicit Assembly(const std::string& name_) : name(name_) {}
};
} // namespace gemmi
#endif