Files
Jungfraujoch/gemmi_gph/gemmi/seqid.hpp
T
leonarski_f dd0bffb283
Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 11m6s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 10m27s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 10m54s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 9m25s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 10m5s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m33s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 11m19s
Build Packages / build:rpm (rocky8) (push) Successful in 12m23s
Build Packages / build:rpm (rocky9) (push) Successful in 13m21s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m30s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m55s
Build Packages / DIALS test (push) Successful in 13m42s
Build Packages / XDS test (durin plugin) (push) Successful in 9m26s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 6m41s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m12s
Build Packages / Generate python client (push) Successful in 19s
Build Packages / Build documentation (push) Successful in 52s
Build Packages / Create release (push) Skipped
Build Packages / build:viewer-tgz:cpu (push) Successful in 5m29s
Build Packages / build:viewer-tgz:cuda (push) Successful in 6m12s
Build Packages / build:windows:cuda (push) Successful in 18m36s
v1.0.0-rc.159 (#69)
This is an UNSTABLE release. It includes many experimental features, as well as many AI generated fixes. We recommend using rc.152 for production use.

* rugnux: Add `--model model.pdb` - score the merged data against an atomic model and compute initial maps. It reports R-work/R-free (scaling the model to the observed amplitudes with an overall scale, an anisotropic B and a flat bulk solvent - the standard few-parameter model, so a batch of maps stays directly comparable) and writes 2Fo-Fc / Fo-Fc electron-density maps (CCP4) plus a map-coefficient MTZ. The structure itself is not refined; the model is only re-fractionalised into the data cell.
* rugnux: The merged reflection output now carries French-Wilson amplitudes (|F| and its sigma) next to the intensities - MTZ `F`/`SIGF`, mmCIF `_refln.F_meas_au`, and the text HKL - computed with the correct centric/acentric Wilson prior and epsilon multiplicity, so a downstream program (e.g. phenix.refine) can refine against amplitudes. The intensity columns are unchanged.
* rugnux: R-free test-set flags are now assigned deterministically and consistently across symmetry - a Bijvoet pair I(+)/I(-) is never split between the work and free sets, and the assignment is a reproducible per-hkl hash that depends only on the reflection index, so every dataset of one crystal form gets the same ~5% free set (what a multi-dataset campaign such as PanDDA needs). On small data the fraction is floored so the test set stays large enough for a stable R-free (~500 reflections, capped at 10%); it stays flat at 5% on ordinary data. When a reference MTZ carries a `FreeR_flag` column its test set is imported instead, letting a whole campaign inherit one shared free set.
* rugnux: A reference MTZ (`--reference-mtz`) can now fix the space group and cell for rotation data too (previously rejected), without being used to scale - the rotation merge stays self-consistent. When the crystal has an indexing (merohedral) ambiguity - a lattice symmetry higher than its Laue symmetry, e.g. P3/P4/P6/C2 - the reference also resolves it: each candidate reindexing (identity plus the twin-law cosets of the metric symmetry) is scored by its intensity correlation against the reference and the data are re-merged in the best-correlating one. This is a metric-preserving relabelling of hkl (the cell is unchanged) and a no-op for a holohedral crystal such as lysozyme.
* rugnux: `--model` validation now aligns the data to the model before scoring - the observed reflections are reindexed into the model's enantiomorph when the two differ only by hand (indistinguishable from merged intensities). A merohedral indexing ambiguity is resolved against the reference MTZ when one is given (so a whole campaign shares one indexing convention); only with a model and no reference does validation fall back to fitting each candidate reindexing and keeping the lowest R-free.
* rugnux: De-novo symmetry - recover a genuine high-symmetry group whose data are imperfectly scaled. Such a merge's within-orbit chi² lands just past the self-consistency bound (each real symmetry step adds a little systematic scatter), right where a merohedral twin also lands, so the chi² ratio alone cannot separate them. The candidate is now rescued when the extra intensity-proportional systematic error it invokes stays small relative to the confirmed subgroup - a genuine symmetry step gains multiplicity without inflating the merge error model's b, whereas a twin forces non-equivalent reflections together and b balloons. Fixes cubic insulin (I23 instead of I222) with no change to any other crystal in the test battery, including the twins that must stay in their lower symmetry.
* Docs: Document the French-Wilson amplitude estimation, R-free flagging, reference-based space-group/ambiguity resolution, and model-based validation/maps in CPU_DATA_ANALYSIS.md.
* Frontend: The status-bar pill now shows a progress bar during detector calibration (previously only during measurement), and the calibration state and its button are labelled "Calibration"/"CALIBRATE" (the internal `Pedestal` state name is unchanged for back-compatibility).Reviewed-on: #69

Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-07-13 13:54:03 +02:00

169 lines
5.4 KiB
C++

// Copyright 2017 Global Phasing Ltd.
//
// SeqId -- residue number and insertion code together.
#ifndef GEMMI_SEQID_HPP_
#define GEMMI_SEQID_HPP_
#include <climits> // for INT_MIN
#include <cstdlib> // for strtol
#include <stdexcept> // for invalid_argument
#include <string>
#include "util.hpp" // for cat
namespace gemmi {
// Optional int value. N is a special value that means not-set.
template<int N> struct OptionalInt {
enum { None=N };
int value = None;
OptionalInt() = default;
OptionalInt(int n) : value(n) {}
bool has_value() const { return value != None; }
std::string str(char null='?') const {
return has_value() ? std::to_string(value) : std::string(1, null);
}
OptionalInt& operator=(int n) { value = n; return *this; }
bool operator==(const OptionalInt& o) const { return value == o.value; }
bool operator!=(const OptionalInt& o) const { return value != o.value; }
bool operator<(const OptionalInt& o) const {
return has_value() && o.has_value() && value < o.value;
}
bool operator==(int n) const { return value == n; }
bool operator!=(int n) const { return value != n; }
OptionalInt operator+(OptionalInt o) const {
return OptionalInt(has_value() && o.has_value() ? value + o.value : N);
}
OptionalInt operator-(OptionalInt o) const {
return OptionalInt(has_value() && o.has_value() ? value - o.value : N);
}
OptionalInt& operator+=(int n) { if (has_value()) value += n; return *this; }
OptionalInt& operator-=(int n) { if (has_value()) value -= n; return *this; }
explicit operator int() const { return value; }
explicit operator bool() const { return has_value(); }
// these are defined for partial compatibility with C++17 std::optional
using value_type = int;
int& operator*() { return value; }
const int& operator*() const { return value; }
int& emplace(int n) { value = n; return value; }
void reset() noexcept { value = None; }
};
struct SeqId {
using OptionalNum = OptionalInt<INT_MIN>;
OptionalNum num; // sequence number
char icode = ' '; // insertion code
SeqId() = default;
SeqId(int num_, char icode_) { num = num_; icode = icode_; }
SeqId(OptionalNum num_, char icode_) { num = num_; icode = icode_; }
explicit SeqId(const std::string& str) {
char* endptr;
num = std::strtol(str.c_str(), &endptr, 10);
if (endptr == str.c_str() || (*endptr != '\0' && endptr[1] != '\0'))
throw std::invalid_argument("Not a seqid: " + str);
icode = (*endptr | 0x20);
}
bool operator==(const SeqId& o) const {
return num == o.num && ((icode ^ o.icode) & ~0x20) == 0;
}
bool operator!=(const SeqId& o) const { return !operator==(o); }
bool operator<(const SeqId& o) const {
return (*num * 256 + icode) < (*o.num * 256 + o.icode);
}
bool operator<=(const SeqId& o) const { return !(o < *this); }
char has_icode() const { return icode != ' '; }
std::string str(bool dot_before_icode=false) const {
std::string r = num.str();
if (icode != ' ') {
if (dot_before_icode)
r += '.';
r += icode;
}
return r;
}
};
// Sequence ID (sequence number + insertion code) + residue name + segment ID
struct ResidueId {
SeqId seqid;
std::string segment; // segid - up to 4 characters in the PDB file
std::string name;
// used for first_conformation iterators, etc.
SeqId group_key() const { return seqid; }
bool matches(const ResidueId& o) const {
return seqid == o.seqid && segment == o.segment && name == o.name;
}
bool matches_noseg(const ResidueId& o) const {
return seqid == o.seqid && name == o.name;
}
bool operator==(const ResidueId& o) const { return matches(o); }
std::string str() const { return cat(seqid.str(), '(', name, ')'); }
};
inline std::string atom_str(const std::string& chain_name,
const ResidueId& res_id,
const std::string& atom_name,
char altloc,
bool as_cid=false) {
std::string r = as_cid ? "//" + chain_name : chain_name;
r += '/';
if (!as_cid) {
r += res_id.name;
r += ' ';
}
r += res_id.seqid.str(as_cid);
if (as_cid && atom_name == "null")
return r;
r += '/';
r += atom_name;
if (altloc) {
r += as_cid ? ':' : '.';
r += altloc;
}
return r;
}
struct AtomAddress {
std::string chain_name;
ResidueId res_id;
std::string atom_name;
char altloc = '\0';
AtomAddress() = default;
AtomAddress(const std::string& ch, const ResidueId& resid,
const std::string& atom, char alt='\0')
: chain_name(ch), res_id(resid), atom_name(atom), altloc(alt) {}
AtomAddress(const std::string& ch, const SeqId& seqid, const std::string& res,
const std::string& atom, char alt='\0')
: chain_name(ch), res_id({seqid, "", res}), atom_name(atom), altloc(alt) {}
bool operator==(const AtomAddress& o) const {
return chain_name == o.chain_name && res_id.matches(o.res_id) &&
atom_name == o.atom_name && altloc == o.altloc;
}
std::string str() const {
return atom_str(chain_name, res_id, atom_name, altloc);
}
};
} // namespace gemmi
namespace std {
template <> struct hash<gemmi::ResidueId> {
size_t operator()(const gemmi::ResidueId& r) const {
size_t seqid_hash = (*r.seqid.num << 7) + (r.seqid.icode | 0x20);
return seqid_hash ^ hash<string>()(r.segment) ^ hash<string>()(r.name);
}
};
} // namespace std
#endif