Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 11m6s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 10m27s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 10m54s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 9m25s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 10m5s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m33s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 11m19s
Build Packages / build:rpm (rocky8) (push) Successful in 12m23s
Build Packages / build:rpm (rocky9) (push) Successful in 13m21s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m30s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m55s
Build Packages / DIALS test (push) Successful in 13m42s
Build Packages / XDS test (durin plugin) (push) Successful in 9m26s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 6m41s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m12s
Build Packages / Generate python client (push) Successful in 19s
Build Packages / Build documentation (push) Successful in 52s
Build Packages / Create release (push) Skipped
Build Packages / build:viewer-tgz:cpu (push) Successful in 5m29s
Build Packages / build:viewer-tgz:cuda (push) Successful in 6m12s
Build Packages / build:windows:cuda (push) Successful in 18m36s
This is an UNSTABLE release. It includes many experimental features, as well as many AI generated fixes. We recommend using rc.152 for production use. * rugnux: Add `--model model.pdb` - score the merged data against an atomic model and compute initial maps. It reports R-work/R-free (scaling the model to the observed amplitudes with an overall scale, an anisotropic B and a flat bulk solvent - the standard few-parameter model, so a batch of maps stays directly comparable) and writes 2Fo-Fc / Fo-Fc electron-density maps (CCP4) plus a map-coefficient MTZ. The structure itself is not refined; the model is only re-fractionalised into the data cell. * rugnux: The merged reflection output now carries French-Wilson amplitudes (|F| and its sigma) next to the intensities - MTZ `F`/`SIGF`, mmCIF `_refln.F_meas_au`, and the text HKL - computed with the correct centric/acentric Wilson prior and epsilon multiplicity, so a downstream program (e.g. phenix.refine) can refine against amplitudes. The intensity columns are unchanged. * rugnux: R-free test-set flags are now assigned deterministically and consistently across symmetry - a Bijvoet pair I(+)/I(-) is never split between the work and free sets, and the assignment is a reproducible per-hkl hash that depends only on the reflection index, so every dataset of one crystal form gets the same ~5% free set (what a multi-dataset campaign such as PanDDA needs). On small data the fraction is floored so the test set stays large enough for a stable R-free (~500 reflections, capped at 10%); it stays flat at 5% on ordinary data. When a reference MTZ carries a `FreeR_flag` column its test set is imported instead, letting a whole campaign inherit one shared free set. * rugnux: A reference MTZ (`--reference-mtz`) can now fix the space group and cell for rotation data too (previously rejected), without being used to scale - the rotation merge stays self-consistent. When the crystal has an indexing (merohedral) ambiguity - a lattice symmetry higher than its Laue symmetry, e.g. P3/P4/P6/C2 - the reference also resolves it: each candidate reindexing (identity plus the twin-law cosets of the metric symmetry) is scored by its intensity correlation against the reference and the data are re-merged in the best-correlating one. This is a metric-preserving relabelling of hkl (the cell is unchanged) and a no-op for a holohedral crystal such as lysozyme. * rugnux: `--model` validation now aligns the data to the model before scoring - the observed reflections are reindexed into the model's enantiomorph when the two differ only by hand (indistinguishable from merged intensities). A merohedral indexing ambiguity is resolved against the reference MTZ when one is given (so a whole campaign shares one indexing convention); only with a model and no reference does validation fall back to fitting each candidate reindexing and keeping the lowest R-free. * rugnux: De-novo symmetry - recover a genuine high-symmetry group whose data are imperfectly scaled. Such a merge's within-orbit chi² lands just past the self-consistency bound (each real symmetry step adds a little systematic scatter), right where a merohedral twin also lands, so the chi² ratio alone cannot separate them. The candidate is now rescued when the extra intensity-proportional systematic error it invokes stays small relative to the confirmed subgroup - a genuine symmetry step gains multiplicity without inflating the merge error model's b, whereas a twin forces non-equivalent reflections together and b balloons. Fixes cubic insulin (I23 instead of I222) with no change to any other crystal in the test battery, including the twins that must stay in their lower symmetry. * Docs: Document the French-Wilson amplitude estimation, R-free flagging, reference-based space-group/ambiguity resolution, and model-based validation/maps in CPU_DATA_ANALYSIS.md. * Frontend: The status-bar pill now shows a progress bar during detector calibration (previously only during measurement), and the calibration state and its button are labelled "Calibration"/"CALIBRATE" (the internal `Pedestal` state name is unchanged for back-compatibility).Reviewed-on: #69 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
233 lines
6.5 KiB
C++
233 lines
6.5 KiB
C++
// Copyright 2018 Global Phasing Ltd.
|
|
//
|
|
// Classes for iterating over files in a directory tree, top-down,
|
|
// in alphabetical order. Wraps the tinydir library (as we cannot yet
|
|
// depend on C++17 <filesystem>).
|
|
|
|
// DirWalk<> iterates through all files and directories.
|
|
// CifWalk yields only cif files (either files that end with .cif or .cif.gz,
|
|
// or files that look like SF mmCIF files from wwPDB, e.g. r3aaasf.ent.gz).
|
|
// It's good for traversing a local copy of the wwPDB archive.
|
|
// PdbWalk: .pdb or .ent (optionally with .gz) except r????sf.ent
|
|
// CoorFileWalk: .cif, .pdb or .ent (optionally with .gz)
|
|
// except r????sf.ent and *-sf.cif
|
|
//
|
|
// Usage:
|
|
// for (const std::string& file : gemmi::DirWalk<>(top_dir))
|
|
// do_something(file);
|
|
// or
|
|
// for (const std::string& file : gemmi::CifWalk(top_dir))
|
|
// do_something(file);
|
|
// You should also catch std::runtime_error.
|
|
|
|
#ifndef GEMMI_DIRWALK_HPP_
|
|
#define GEMMI_DIRWALK_HPP_
|
|
|
|
#include <string>
|
|
#include <vector>
|
|
#include <cassert>
|
|
#if defined(_MSC_VER) && !defined(NOMINMAX)
|
|
# define NOMINMAX
|
|
#endif
|
|
#include "third_party/tinydir.h"
|
|
|
|
#include "util.hpp" // for giends_with
|
|
#include "fail.hpp" // for sys_fail
|
|
#include "pdb_id.hpp" // for is_pdb_code, expand_pdb_code_to_path
|
|
#include "glob.hpp" // for glob_match
|
|
#if defined(_WIN32) && defined(_UNICODE)
|
|
#include "utf.hpp"
|
|
#endif
|
|
|
|
namespace gemmi {
|
|
|
|
inline std::string as_utf8(const _tinydir_char_t* path) {
|
|
#if defined(_WIN32) && defined(_UNICODE)
|
|
return wchar_to_UTF8(path);
|
|
#else
|
|
return path;
|
|
#endif
|
|
}
|
|
|
|
|
|
namespace impl {
|
|
// the SF mmCIF files from PDB have names such as
|
|
// divided/structure_factors/aa/r3aaasf.ent.gz
|
|
inline bool is_rxsf_ent_filename(const std::string& filename) {
|
|
return filename[0] == 'r' && giends_with(filename, "sf.ent")
|
|
&& filename.find('.') >= 4;
|
|
}
|
|
|
|
struct IsMmCifFile { // actually we don't know what kind of cif file it is
|
|
static bool check(const std::string& filename) {
|
|
return giends_with(filename, ".cif") || giends_with(filename, ".mmcif");
|
|
}
|
|
};
|
|
|
|
struct IsCifFile {
|
|
static bool check(const std::string& filename) {
|
|
return giends_with(filename, ".cif") || is_rxsf_ent_filename(filename);
|
|
}
|
|
};
|
|
|
|
struct IsPdbFile {
|
|
static bool check(const std::string& filename) {
|
|
return giends_with(filename, ".pdb") ||
|
|
(giends_with(filename, ".ent") && !is_rxsf_ent_filename(filename));
|
|
}
|
|
};
|
|
|
|
struct IsCoordinateFile {
|
|
static bool check(const std::string& filename) {
|
|
// the SF mmCIF files from RCSB website have names such as 3AAA-sf.cif
|
|
return IsPdbFile::check(filename) ||
|
|
(IsMmCifFile::check(filename) && !giends_with(filename, "-sf.cif"));
|
|
}
|
|
};
|
|
|
|
struct IsAnyFile {
|
|
static bool check(const std::string&) { return true; }
|
|
};
|
|
|
|
struct IsMatchingFile {
|
|
bool check(const std::string& filename) const {
|
|
return glob_match(pattern, filename);
|
|
}
|
|
std::string pattern;
|
|
};
|
|
|
|
inline int utf8_tinydir_file_open(tinydir_file* file, const char* path) {
|
|
#if defined(_WIN32) && defined(_UNICODE)
|
|
return tinydir_file_open(file, UTF8_to_wchar(path).c_str());
|
|
#else
|
|
return tinydir_file_open(file, path);
|
|
#endif
|
|
}
|
|
|
|
} // namespace impl
|
|
|
|
|
|
template<bool FileOnly=true, typename Filter=impl::IsAnyFile>
|
|
class DirWalk {
|
|
public:
|
|
explicit DirWalk(const char* path, char try_pdbid='\0') {
|
|
if (impl::utf8_tinydir_file_open(&top_, path) != -1)
|
|
return;
|
|
if (try_pdbid != '\0' && is_pdb_code(path)) {
|
|
std::string epath = expand_pdb_code_to_path(path, try_pdbid, true);
|
|
if (impl::utf8_tinydir_file_open(&top_, epath.c_str()) != -1)
|
|
return;
|
|
sys_fail("Cannot open " + epath);
|
|
}
|
|
sys_fail("Cannot open " + std::string(path));
|
|
}
|
|
explicit DirWalk(const std::string& path, char try_pdbid='\0')
|
|
: DirWalk(path.c_str(), try_pdbid) {}
|
|
~DirWalk() {
|
|
for (auto& d : dirs_)
|
|
tinydir_close(&d.second);
|
|
}
|
|
void push_dir(size_t cur_pos, const _tinydir_char_t* path) {
|
|
dirs_.emplace_back();
|
|
dirs_.back().first = cur_pos;
|
|
if (tinydir_open_sorted(&dirs_.back().second, path) == -1)
|
|
sys_fail("Cannot open directory " + as_utf8(path));
|
|
}
|
|
size_t pop_dir() {
|
|
assert(!dirs_.empty());
|
|
size_t old_pos = dirs_.back().first;
|
|
tinydir_close(&dirs_.back().second);
|
|
dirs_.pop_back();
|
|
return old_pos;
|
|
}
|
|
|
|
struct Iter {
|
|
DirWalk& walk;
|
|
size_t cur;
|
|
|
|
const tinydir_dir& get_dir() const { return walk.dirs_.back().second; }
|
|
|
|
const tinydir_file& get() const {
|
|
if (walk.dirs_.empty())
|
|
return walk.top_;
|
|
assert(cur < get_dir().n_files);
|
|
return get_dir()._files[cur];
|
|
}
|
|
|
|
std::string operator*() const { return as_utf8(get().path); }
|
|
|
|
// checks for "." and ".."
|
|
bool is_special(const _tinydir_char_t* name) const {
|
|
return name[0] == '.' && (name[1] == '\0' ||
|
|
(name[1] == '.' && name[2] == '\0'));
|
|
}
|
|
|
|
size_t depth() const { return walk.dirs_.size(); }
|
|
|
|
void next() { // depth first
|
|
const tinydir_file& tf = get();
|
|
if (tf.is_dir) {
|
|
walk.push_dir(cur, tf.path);
|
|
cur = 0;
|
|
} else {
|
|
cur++;
|
|
}
|
|
while (!walk.dirs_.empty()) {
|
|
if (cur == get_dir().n_files)
|
|
cur = walk.pop_dir() + 1;
|
|
else if (is_special(get_dir()._files[cur].name))
|
|
cur++;
|
|
else
|
|
break;
|
|
}
|
|
}
|
|
|
|
void operator++() {
|
|
for (;;) {
|
|
next();
|
|
const tinydir_file& f = get();
|
|
if ((!FileOnly && f.is_dir)
|
|
|| (!f.is_dir && walk.filter.check(as_utf8(f.name)))
|
|
|| walk.is_single_file()
|
|
|| (depth() == 0 && cur == 1))
|
|
break;
|
|
}
|
|
}
|
|
|
|
// == and != is used only to compare with end()
|
|
bool operator==(const Iter& o) const { return depth()==0 && cur == o.cur; }
|
|
bool operator!=(const Iter& o) const { return !operator==(o); }
|
|
};
|
|
|
|
Iter begin() {
|
|
Iter it{*this, 0};
|
|
if (FileOnly && !is_single_file()) // i.e. the top item is a directory
|
|
++it;
|
|
return it;
|
|
}
|
|
|
|
Iter end() { return Iter{*this, 1}; }
|
|
bool is_single_file() { return !top_.is_dir; }
|
|
|
|
private:
|
|
friend struct Iter;
|
|
tinydir_file top_;
|
|
std::vector<std::pair<size_t, tinydir_dir>> dirs_;
|
|
protected:
|
|
Filter filter;
|
|
};
|
|
|
|
using CifWalk = DirWalk<true, impl::IsCifFile>;
|
|
using MmCifWalk = DirWalk<true, impl::IsMmCifFile>;
|
|
using PdbWalk = DirWalk<true, impl::IsPdbFile>;
|
|
using CoorFileWalk = DirWalk<true, impl::IsCoordinateFile>;
|
|
|
|
struct GlobWalk : public DirWalk<true, impl::IsMatchingFile> {
|
|
GlobWalk(const std::string& path, const std::string& glob) : DirWalk(path) {
|
|
filter.pattern = glob;
|
|
}
|
|
};
|
|
|
|
} // namespace gemmi
|
|
#endif
|