Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 11m6s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 10m27s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 10m54s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 9m25s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 10m5s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m33s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 11m19s
Build Packages / build:rpm (rocky8) (push) Successful in 12m23s
Build Packages / build:rpm (rocky9) (push) Successful in 13m21s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m30s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m55s
Build Packages / DIALS test (push) Successful in 13m42s
Build Packages / XDS test (durin plugin) (push) Successful in 9m26s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 6m41s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m12s
Build Packages / Generate python client (push) Successful in 19s
Build Packages / Build documentation (push) Successful in 52s
Build Packages / Create release (push) Skipped
Build Packages / build:viewer-tgz:cpu (push) Successful in 5m29s
Build Packages / build:viewer-tgz:cuda (push) Successful in 6m12s
Build Packages / build:windows:cuda (push) Successful in 18m36s
This is an UNSTABLE release. It includes many experimental features, as well as many AI generated fixes. We recommend using rc.152 for production use. * rugnux: Add `--model model.pdb` - score the merged data against an atomic model and compute initial maps. It reports R-work/R-free (scaling the model to the observed amplitudes with an overall scale, an anisotropic B and a flat bulk solvent - the standard few-parameter model, so a batch of maps stays directly comparable) and writes 2Fo-Fc / Fo-Fc electron-density maps (CCP4) plus a map-coefficient MTZ. The structure itself is not refined; the model is only re-fractionalised into the data cell. * rugnux: The merged reflection output now carries French-Wilson amplitudes (|F| and its sigma) next to the intensities - MTZ `F`/`SIGF`, mmCIF `_refln.F_meas_au`, and the text HKL - computed with the correct centric/acentric Wilson prior and epsilon multiplicity, so a downstream program (e.g. phenix.refine) can refine against amplitudes. The intensity columns are unchanged. * rugnux: R-free test-set flags are now assigned deterministically and consistently across symmetry - a Bijvoet pair I(+)/I(-) is never split between the work and free sets, and the assignment is a reproducible per-hkl hash that depends only on the reflection index, so every dataset of one crystal form gets the same ~5% free set (what a multi-dataset campaign such as PanDDA needs). On small data the fraction is floored so the test set stays large enough for a stable R-free (~500 reflections, capped at 10%); it stays flat at 5% on ordinary data. When a reference MTZ carries a `FreeR_flag` column its test set is imported instead, letting a whole campaign inherit one shared free set. * rugnux: A reference MTZ (`--reference-mtz`) can now fix the space group and cell for rotation data too (previously rejected), without being used to scale - the rotation merge stays self-consistent. When the crystal has an indexing (merohedral) ambiguity - a lattice symmetry higher than its Laue symmetry, e.g. P3/P4/P6/C2 - the reference also resolves it: each candidate reindexing (identity plus the twin-law cosets of the metric symmetry) is scored by its intensity correlation against the reference and the data are re-merged in the best-correlating one. This is a metric-preserving relabelling of hkl (the cell is unchanged) and a no-op for a holohedral crystal such as lysozyme. * rugnux: `--model` validation now aligns the data to the model before scoring - the observed reflections are reindexed into the model's enantiomorph when the two differ only by hand (indistinguishable from merged intensities). A merohedral indexing ambiguity is resolved against the reference MTZ when one is given (so a whole campaign shares one indexing convention); only with a model and no reference does validation fall back to fitting each candidate reindexing and keeping the lowest R-free. * rugnux: De-novo symmetry - recover a genuine high-symmetry group whose data are imperfectly scaled. Such a merge's within-orbit chi² lands just past the self-consistency bound (each real symmetry step adds a little systematic scatter), right where a merohedral twin also lands, so the chi² ratio alone cannot separate them. The candidate is now rescued when the extra intensity-proportional systematic error it invokes stays small relative to the confirmed subgroup - a genuine symmetry step gains multiplicity without inflating the merge error model's b, whereas a twin forces non-equivalent reflections together and b balloons. Fixes cubic insulin (I23 instead of I222) with no change to any other crystal in the test battery, including the twins that must stay in their lower symmetry. * Docs: Document the French-Wilson amplitude estimation, R-free flagging, reference-based space-group/ambiguity resolution, and model-based validation/maps in CPU_DATA_ANALYSIS.md. * Frontend: The status-bar pill now shows a progress bar during detector calibration (previously only during measurement), and the calibration state and its button are labelled "Calibration"/"CALIBRATE" (the internal `Pedestal` state name is unchanged for back-compatibility).Reviewed-on: #69 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
141 lines
4.8 KiB
C++
141 lines
4.8 KiB
C++
// Copyright 2021 Global Phasing Ltd.
|
|
//
|
|
// Finding maxima or "blobs" in a Grid (map).
|
|
// Similar to CCP4 PEAKMAX and COOT's "Unmodelled blobs".
|
|
//
|
|
// Implementation of the flood fill algorithm in find_blobs_by_flood_fill()
|
|
// differs from from FloodFill in floodfill.hpp.
|
|
// FloodFill uses more efficient scanline fill, but doesn't use symmetry.
|
|
|
|
#ifndef GEMMI_BLOB_HPP_
|
|
#define GEMMI_BLOB_HPP_
|
|
|
|
#include "grid.hpp" // for Grid
|
|
#include "asumask.hpp" // for get_asu_mask
|
|
|
|
namespace gemmi {
|
|
|
|
struct Blob {
|
|
double volume = 0.0;
|
|
double score = 0.0;
|
|
double peak_value = 0.0;
|
|
gemmi::Position centroid;
|
|
gemmi::Position peak_pos;
|
|
explicit operator bool() const { return volume != 0. ; }
|
|
};
|
|
|
|
struct BlobCriteria {
|
|
double cutoff;
|
|
double min_volume = 10.0;
|
|
double min_score = 15.0;
|
|
double min_peak = 0.0;
|
|
};
|
|
|
|
namespace impl {
|
|
|
|
struct GridConstPoint {
|
|
int u, v, w;
|
|
float value;
|
|
};
|
|
|
|
inline Blob make_blob_of_points(const std::vector<GridConstPoint>& points,
|
|
const GridMeta& grid,
|
|
const BlobCriteria& criteria) {
|
|
Blob blob;
|
|
if (points.size() < 3)
|
|
return blob;
|
|
double volume_per_point = grid.unit_cell.volume / grid.point_count();
|
|
double volume = points.size() * volume_per_point;
|
|
if (volume < criteria.min_volume)
|
|
return blob;
|
|
double sum[3] = {0., 0., 0.};
|
|
const GridConstPoint* peak_point = &points[0];
|
|
blob.peak_value = points[0].value;
|
|
double score = 0.;
|
|
for (const GridConstPoint& point : points) {
|
|
score += point.value;
|
|
if (point.value > blob.peak_value) {
|
|
blob.peak_value = point.value;
|
|
peak_point = &point;
|
|
}
|
|
sum[0] += double(point.u) * point.value;
|
|
sum[1] += double(point.v) * point.value;
|
|
sum[2] += double(point.w) * point.value;
|
|
}
|
|
if (blob.peak_value < criteria.min_peak)
|
|
return blob;
|
|
blob.score = score * volume_per_point;
|
|
if (blob.score < criteria.min_score)
|
|
return blob;
|
|
gemmi::Fractional fract(sum[0] / (score * grid.nu),
|
|
sum[1] / (score * grid.nv),
|
|
sum[2] / (score * grid.nw));
|
|
blob.centroid = grid.unit_cell.orthogonalize(fract);
|
|
blob.peak_pos = grid.get_position(peak_point->u, peak_point->v, peak_point->w);
|
|
blob.volume = volume;
|
|
return blob;
|
|
}
|
|
|
|
} // namespace impl
|
|
|
|
// with negate=true grid negatives of grid values are used
|
|
inline std::vector<Blob> find_blobs_by_flood_fill(const gemmi::Grid<float>& grid,
|
|
const BlobCriteria& criteria,
|
|
bool negate=false) {
|
|
std::vector<Blob> blobs;
|
|
std::array<std::array<int, 3>, 6> moves = {{{{-1, 0, 0}}, {{1, 0, 0}},
|
|
{{0 ,-1, 0}}, {{0, 1, 0}},
|
|
{{0, 0, -1}}, {{0, 0, 1}}}};
|
|
// the mask will be used as follows:
|
|
// -1=in blob, 0=in asu, not in blob (so far), 1=in neither
|
|
std::vector<std::int8_t> mask = gemmi::get_asu_mask(grid);
|
|
std::vector<gemmi::GridOp> ops = grid.get_scaled_ops_except_id();
|
|
size_t idx = 0;
|
|
for (int w = 0; w != grid.nw; ++w)
|
|
for (int v = 0; v != grid.nv; ++v)
|
|
for (int u = 0; u != grid.nu; ++u, ++idx) {
|
|
assert(idx == grid.index_q(u, v, w));
|
|
if (mask[idx] != 0)
|
|
continue;
|
|
float value = grid.data[idx];
|
|
if (negate)
|
|
value = -value;
|
|
if (value < criteria.cutoff)
|
|
continue;
|
|
std::vector<impl::GridConstPoint> points;
|
|
points.push_back({u, v, w, value});
|
|
mask[idx] = -1;
|
|
for (size_t j = 0; j < points.size()/*increasing!*/; ++j)
|
|
for (const std::array<int, 3>& mv : moves) {
|
|
int nabe_u = points[j].u + mv[0];
|
|
int nabe_v = points[j].v + mv[1];
|
|
int nabe_w = points[j].w + mv[2];
|
|
size_t nabe_idx = grid.index_s(nabe_u, nabe_v, nabe_w);
|
|
if (mask[nabe_idx] == -1)
|
|
continue;
|
|
float nabe_value = grid.data[nabe_idx];
|
|
if (negate)
|
|
nabe_value = -nabe_value;
|
|
if (nabe_value > criteria.cutoff) {
|
|
if (mask[nabe_idx] != 0)
|
|
for (const gemmi::GridOp& op : ops) {
|
|
auto t = op.apply(nabe_u, nabe_v, nabe_w);
|
|
size_t mate_idx = grid.index_s(t[0], t[1], t[2]);
|
|
if (mask[mate_idx] == 0)
|
|
mask[mate_idx] = 1;
|
|
}
|
|
mask[nabe_idx] = -1;
|
|
points.push_back({nabe_u, nabe_v, nabe_w, nabe_value});
|
|
}
|
|
}
|
|
if (Blob blob = impl::make_blob_of_points(points, grid, criteria))
|
|
blobs.push_back(blob);
|
|
}
|
|
std::sort(blobs.begin(), blobs.end(),
|
|
[](const Blob& a, const Blob& b) { return a.score > b.score; });
|
|
return blobs;
|
|
}
|
|
|
|
} // namespace gemmi
|
|
#endif
|