Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 11m6s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 10m27s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 10m54s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 9m25s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 10m5s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m33s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 11m19s
Build Packages / build:rpm (rocky8) (push) Successful in 12m23s
Build Packages / build:rpm (rocky9) (push) Successful in 13m21s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m30s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m55s
Build Packages / DIALS test (push) Successful in 13m42s
Build Packages / XDS test (durin plugin) (push) Successful in 9m26s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 6m41s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m12s
Build Packages / Generate python client (push) Successful in 19s
Build Packages / Build documentation (push) Successful in 52s
Build Packages / Create release (push) Skipped
Build Packages / build:viewer-tgz:cpu (push) Successful in 5m29s
Build Packages / build:viewer-tgz:cuda (push) Successful in 6m12s
Build Packages / build:windows:cuda (push) Successful in 18m36s
This is an UNSTABLE release. It includes many experimental features, as well as many AI generated fixes. We recommend using rc.152 for production use. * rugnux: Add `--model model.pdb` - score the merged data against an atomic model and compute initial maps. It reports R-work/R-free (scaling the model to the observed amplitudes with an overall scale, an anisotropic B and a flat bulk solvent - the standard few-parameter model, so a batch of maps stays directly comparable) and writes 2Fo-Fc / Fo-Fc electron-density maps (CCP4) plus a map-coefficient MTZ. The structure itself is not refined; the model is only re-fractionalised into the data cell. * rugnux: The merged reflection output now carries French-Wilson amplitudes (|F| and its sigma) next to the intensities - MTZ `F`/`SIGF`, mmCIF `_refln.F_meas_au`, and the text HKL - computed with the correct centric/acentric Wilson prior and epsilon multiplicity, so a downstream program (e.g. phenix.refine) can refine against amplitudes. The intensity columns are unchanged. * rugnux: R-free test-set flags are now assigned deterministically and consistently across symmetry - a Bijvoet pair I(+)/I(-) is never split between the work and free sets, and the assignment is a reproducible per-hkl hash that depends only on the reflection index, so every dataset of one crystal form gets the same ~5% free set (what a multi-dataset campaign such as PanDDA needs). On small data the fraction is floored so the test set stays large enough for a stable R-free (~500 reflections, capped at 10%); it stays flat at 5% on ordinary data. When a reference MTZ carries a `FreeR_flag` column its test set is imported instead, letting a whole campaign inherit one shared free set. * rugnux: A reference MTZ (`--reference-mtz`) can now fix the space group and cell for rotation data too (previously rejected), without being used to scale - the rotation merge stays self-consistent. When the crystal has an indexing (merohedral) ambiguity - a lattice symmetry higher than its Laue symmetry, e.g. P3/P4/P6/C2 - the reference also resolves it: each candidate reindexing (identity plus the twin-law cosets of the metric symmetry) is scored by its intensity correlation against the reference and the data are re-merged in the best-correlating one. This is a metric-preserving relabelling of hkl (the cell is unchanged) and a no-op for a holohedral crystal such as lysozyme. * rugnux: `--model` validation now aligns the data to the model before scoring - the observed reflections are reindexed into the model's enantiomorph when the two differ only by hand (indistinguishable from merged intensities). A merohedral indexing ambiguity is resolved against the reference MTZ when one is given (so a whole campaign shares one indexing convention); only with a model and no reference does validation fall back to fitting each candidate reindexing and keeping the lowest R-free. * rugnux: De-novo symmetry - recover a genuine high-symmetry group whose data are imperfectly scaled. Such a merge's within-orbit chi² lands just past the self-consistency bound (each real symmetry step adds a little systematic scatter), right where a merohedral twin also lands, so the chi² ratio alone cannot separate them. The candidate is now rescued when the extra intensity-proportional systematic error it invokes stays small relative to the confirmed subgroup - a genuine symmetry step gains multiplicity without inflating the merge error model's b, whereas a twin forces non-equivalent reflections together and b balloons. Fixes cubic insulin (I23 instead of I222) with no change to any other crystal in the test battery, including the twins that must stay in their lower symmetry. * Docs: Document the French-Wilson amplitude estimation, R-free flagging, reference-based space-group/ambiguity resolution, and model-based validation/maps in CPU_DATA_ANALYSIS.md. * Frontend: The status-bar pill now shows a progress bar during detector calibration (previously only during measurement), and the calibration state and its button are labelled "Calibration"/"CALIBRATE" (the internal `Pedestal` state name is unchanged for back-compatibility).Reviewed-on: #69 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
123 lines
4.4 KiB
C++
123 lines
4.4 KiB
C++
// Copyright 2017 Global Phasing Ltd.
|
|
//
|
|
// Read any supported coordinate file. Usually, mmread_gz.hpp is preferred.
|
|
|
|
#ifndef GEMMI_MMREAD_HPP_
|
|
#define GEMMI_MMREAD_HPP_
|
|
|
|
#include "cif.hpp" // for cif::read
|
|
#include "fail.hpp" // for fail
|
|
#include "input.hpp" // for BasicInput
|
|
#include "json.hpp" // for read_mmjson
|
|
#include "mmcif.hpp" // for make_structure_from_block, ...
|
|
#include "model.hpp" // for Structure
|
|
#include "pdb.hpp" // for read_pdb
|
|
#include "util.hpp" // for iends_with
|
|
|
|
namespace gemmi {
|
|
|
|
inline CoorFormat coor_format_from_ext(const std::string& path) {
|
|
if (iends_with(path, ".pdb") || iends_with(path, ".ent"))
|
|
return CoorFormat::Pdb;
|
|
if (iends_with(path, ".cif") || iends_with(path, ".mmcif"))
|
|
return CoorFormat::Mmcif;
|
|
if (iends_with(path, ".json"))
|
|
return CoorFormat::Mmjson;
|
|
return CoorFormat::Unknown;
|
|
}
|
|
|
|
// If it's neither CIF nor JSON nor almost empty - we assume PDB.
|
|
inline CoorFormat coor_format_from_content(const char* buf, const char* end) {
|
|
while (buf < end - 8) {
|
|
if (std::isspace(*buf)) {
|
|
++buf;
|
|
} else if (*buf == '#') {
|
|
while (buf < end - 8 && *buf != '\n')
|
|
++buf;
|
|
} else if (*buf == '{') {
|
|
return CoorFormat::Mmjson;
|
|
} else if (ialpha4_id(buf) == ialpha4_id("data") && buf[4] == '_') {
|
|
return CoorFormat::Mmcif;
|
|
} else {
|
|
return CoorFormat::Pdb;
|
|
}
|
|
}
|
|
return CoorFormat::Unknown;
|
|
}
|
|
|
|
inline Structure make_structure_from_doc(cif::Document&& doc, bool possible_chemcomp,
|
|
cif::Document* save_doc=nullptr) {
|
|
if (possible_chemcomp) {
|
|
// check for special case - refmac dictionary or CCD file
|
|
int n = check_chemcomp_block_number(doc);
|
|
if (n != -1)
|
|
return make_structure_from_chemcomp_block(doc.blocks[n]);
|
|
}
|
|
return make_structure(std::move(doc), save_doc);
|
|
}
|
|
|
|
// when reading JSON, the input buffer is changed (as an optimization)
|
|
inline Structure read_structure_from_memory(char* data, size_t size,
|
|
const std::string& path,
|
|
CoorFormat format=CoorFormat::Unknown,
|
|
cif::Document* save_doc=nullptr) {
|
|
if (save_doc)
|
|
save_doc->clear();
|
|
if (format == CoorFormat::Unknown || format == CoorFormat::Detect)
|
|
format = coor_format_from_content(data, data + size);
|
|
if (format == CoorFormat::Pdb)
|
|
return read_pdb_from_memory(data, size, path);
|
|
if (format == CoorFormat::Mmcif)
|
|
return make_structure_from_doc(cif::read_memory(data, size, path.c_str()),
|
|
true, save_doc);
|
|
if (format == CoorFormat::Mmjson)
|
|
return make_structure(cif::read_mmjson_insitu(data, size, path), save_doc);
|
|
fail("wrong format of coordinate file " + path);
|
|
}
|
|
|
|
// deprecated
|
|
inline Structure read_structure_from_char_array(char* data, size_t size,
|
|
const std::string& path,
|
|
cif::Document* save_doc=nullptr) {
|
|
return read_structure_from_memory(data, size, path, CoorFormat::Unknown, save_doc);
|
|
}
|
|
|
|
template<typename T>
|
|
Structure read_structure(T&& input, CoorFormat format=CoorFormat::Unknown,
|
|
cif::Document* save_doc=nullptr) {
|
|
if (format == CoorFormat::Detect) {
|
|
CharArray mem = read_into_buffer(input);
|
|
return read_structure_from_memory(mem.data(), mem.size(), input.path(), format, save_doc);
|
|
}
|
|
if (save_doc)
|
|
save_doc->clear();
|
|
if (format == CoorFormat::Unknown)
|
|
format = coor_format_from_ext(input.basepath());
|
|
switch (format) {
|
|
case CoorFormat::Pdb:
|
|
return read_pdb(input);
|
|
case CoorFormat::Mmcif:
|
|
return make_structure(cif::read(input), save_doc);
|
|
case CoorFormat::Mmjson: {
|
|
Structure st = make_structure(cif::read_mmjson(input), save_doc);
|
|
st.input_format = CoorFormat::Mmjson;
|
|
return st;
|
|
}
|
|
case CoorFormat::ChemComp:
|
|
return make_structure_from_chemcomp_doc(cif::read(input), save_doc);
|
|
case CoorFormat::Unknown:
|
|
case CoorFormat::Detect:
|
|
fail("Unknown format of " +
|
|
(input.path().empty() ? "coordinate file" : input.path()) + ".");
|
|
}
|
|
unreachable();
|
|
}
|
|
|
|
inline Structure read_structure_file(const std::string& path,
|
|
CoorFormat format=CoorFormat::Unknown) {
|
|
return read_structure(BasicInput(path), format);
|
|
}
|
|
|
|
} // namespace gemmi
|
|
#endif
|