Files
Jungfraujoch/rugnux/RigidBodyRefine.cpp
T
leonarski_f a39fd29f77
Build Packages / XDS test (JFJoch plugin) (push) Successful in 11m4s
Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 17m46s
Build Packages / build:windows:cuda (push) Successful in 20m20s
Build Packages / build:viewer-tgz:cpu (push) Successful in 15m56s
Build Packages / build:viewer-tgz:cuda (push) Successful in 17m57s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 14m10s
Build Packages / build:rugnux:windows (push) Successful in 11m12s
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 7m14s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 22m13s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 19m17s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 21m21s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 17m26s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 23m56s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 20m48s
Build Packages / build:rpm (rocky8) (push) Successful in 23m43s
Build Packages / build:rpm (rocky9) (push) Successful in 20m38s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 24m57s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m58s
Build Packages / XDS test (durin plugin) (push) Successful in 10m43s
Build Packages / Generate python client (push) Successful in 47s
Build Packages / Build documentation (push) Successful in 1m5s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (neggia plugin) (push) Successful in 8m57s
Build Packages / DIALS test (push) Successful in 18m40s
v1.0.0-rc.167 (#77)
* `rugnux --model` reports CC(model, data) - the correlation of the merged intensities with the placed, scaled model - by resolution shell, on the same shells as CC1/2, with the reflection count and a significance for each.
* `rugnux --model` fits the model's scale, anisotropic B and bulk-solvent parameters on the working reflections only, so the R-free it reports is measured against a model no free reflection helped scale.
* The bulk-solvent parameters of `rugnux --model` are searched over their physically meaningful range instead of being fitted without bounds, so a model is never scaled with a solvent term that has silently switched itself off.
* The rigid-body placement of `rugnux --model` uses the same bounded bulk solvent as the reported fit, so a model is no longer placed against a target carrying a solvent term with no physical meaning.
* `rugnux --model` puts the model into the data's own description of the lattice before placing it, so a model whose cell is written on other axes - I-centred where the run indexed C-centred, a different unique axis, a permuted orthorhombic cell - is placed rather than scored where it was read; `MODEL_CHANGE_OF_BASIS=` and `MODEL_SETTING_AS_READ=` report it when it happens.
* The rugnux results report opens with a summary - `VERDICT=` (`OK`, `WARNINGS`, `UNUSABLE`, `FAILED`), `VERDICT_TEXT=`, `PATHOLOGY_FLAGS=` with one closed-vocabulary code per condition that warned, and the `WARNING:` lines, which used to close the file - and the sections after it are renumbered 1-5 with no gaps.
* `rugnux --developer` writes the full results report - the pipeline-internal keys and the long explanations the default report now leaves out - and `--finalist-ledger` adds the evidence for every space group the search considered, not only the one it adopted.
* The results report warns when the merged data carry no usable signal and when too little of reciprocal space was measured inside the fitted resolution, and omits `FITTED_RESOLUTION` where the CC1/2 curve it is fitted on never falls off.
* rugnux detects translational pseudo-symmetry and reports it under the `PSEUDO_TRANSLATION` flag as `TNCS_DETECTED=` and the `TNCS_*` keys - a translation the merged data are exactly invariant under is reported as `UNDECLARED_LATTICE_TRANSLATION=` under `LATTICE_TRANSLATION` instead - and a detected pseudo-translation can no longer buy a false screw axis in the space-group search or hide a twin from the L-test (`L_TEST_VS_TNCS=`).
* The space-group search determines glide planes from zonal systematic absences, so a non-Sohncke space group such as P 2_1/c or Pbca is named where the run previously stopped at its Sohncke subgroup; `SOHNCKE_SPACE_GROUP=` carries the best Sohncke group beside it on every run that searched, and a centre of symmetry is never claimed.
* Where the cell metric carries more rotational symmetry than the Bravais class the indexer named, the extra rotations are put to the intensities and the space-group search is asked again on the metric's own cell - adopted only where the intensities confirm the higher symmetry - so a lattice that is nearly but not exactly hexagonal, or whose reduction landed in a sub-cell, still reaches its true point group.
* Systematic-absence calls rest on the evidence rather than on counts: a screw axis whose absent class the data show extinct is no longer refused because a handful of reflections in it read as present, and `SPACE_GROUP_ALTERNATIVES=` no longer drops a candidate that differs only on a zone the sweep never measured.
* A reference correlation measured on too few reflections is refused instead of scored zero, so a run given a reference MTZ is no longer reindexed on an operator that mapped almost everything outside the reference's coverage.
* A frame counts as indexed from 6 spots on its lattice rather than 9, so a weakly diffracting crystal whose frames cannot carry 9 is no longer refused the lattice it fits; `--min-indexed-spots` overrides it.
* `-C` accepts a known cell in any equivalent description - conventional or primitive, centred or not - instead of only the reduced primitive form, so a centred cell given the way it is published no longer makes the run report that it found no lattice.
* Each reflection is corrected for the sensor's quantum efficiency at the angle it meets the detector (attenuation lengths from the NIST tables, which also fixes the spot-width parallax term on CdTe) and for the attenuation of the flight path between the sample and its pixel; `--flight-path air|helium|vacuum` declares the medium - default air, since no file states it - and the report says what was assumed and what it was worth. The unmerged MTZ records the factors in new `QE` and `FLIGHT` columns beside `LP`, so raw counts are `I / LP * QE * FLIGHT`, and `_process.h5` in new optional `qe` and `flight` datasets.
* Rotation geometry post-refinement fits the crystal and the detector at once, against the observed spot positions and the observed rocking angles together, so the refined distance depends far less on how wrong the file's distance was.
* A coarsely sliced sweep integrates correctly: partials are joined into one rocking event by angle rather than by frame count, so two crossings of the Ewald sphere are no longer summed into one full, and at 0.5 degrees per image or coarser the per-frame geometry refinement accepts a spot whose miss the exposure's own rotation accounts for.
* `rugnux --mode scale` reports the detector tilt and direct beam of the geometry it re-scaled at, instead of zeros that read as a flat detector, and no longer warns that no image was indexed on a run whose lattice came from its input file.
* Every rotation run that determined a space group and merged reports what the mounting cost: `SPINDLE_LOST_UNIQUE_FRACTION=` is the fraction (0-1) of unique reflections the mounting made unmeasurable under the measured point group, also written to the master as `/entry/MX/spindleLostUniqueFraction` and what the mounting warning fires on; `SPINDLE_SYMMETRY_AXIS_ANGLE_DEG=` / `SPINDLE_SYMMETRY_AXIS_ORDER=` describe the mounting in the `--developer` report.
* Stills and grid scans carry a per-image `spindle_blind_fraction` - how much of a rotation sweep's blind cone this orientation would make unrecoverable, 0.5 and above calling for a second orientation - through the CBOR stream, HDF5 (`/entry/MX/spindleBlindFraction`), the plot and scan-result APIs, and the viewer and frontend plots; an absent value means the frame could not be assessed and is not a 0.
* The results report's `REPORT_VERSION` is 7.

Reviewed-on: #77
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-09-09 07:25:13 +02:00

365 lines
18 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#include "RigidBodyRefine.h"
#include <algorithm>
#include <chrono>
#include <cmath>
#include <complex>
#include <vector>
#include <ceres/ceres.h>
#include <ceres/rotation.h>
#include "gemmi/dencalc.hpp" // DensityCalculator
#include "gemmi/fourier.hpp" // transform_map_to_f_phi
#include "gemmi/it92.hpp" // IT92 x-ray form factors
#include "gemmi/scaling.hpp" // Scaling (bulk solvent + anisotropic B)
#include "gemmi/solmask.hpp" // SolventMasker
#include "ModelScaling.h" // FitModelScale
#include "../common/JFJochMath.h" // PI
#include "../common/Logger.h"
namespace {
using Table = gemmi::IT92<float>;
// The ladder the placement is walked down. It starts coarse because the model arrives already placed
// but out by a cell's worth of non-isomorphism: at 6 A a few hundred reflections see the body as a
// blob and the target has one broad minimum, and each finer zone starts from the previous one's
// answer. It stops at 3.5 A, which is where rigid-body refinement is conventionally run (it is
// REFMAC's own default through dimple) - the movement being recovered is a few tenths of an
// angstrom, a tenth of that resolution, so it is well determined there, while a finer zone costs
// (1/d)^3 in grid points and reflections for a placement it cannot meaningfully sharpen.
constexpr double LADDER[] = {6.0, 4.5, 3.5};
// The step of the forward-difference Jacobian, as a fraction of the zone's resolution - so it is
// 0.06 A of atom displacement at 6 A and 0.035 A at 3.5 A. A step fixed in angstroms instead is far
// too small for the coarse zones, where a structure factor barely notices it and the derivative is
// swallowed by the jitter of the scale re-fit: measured, a fixed 0.02 A left the 6 A zone at 0.35
// degrees where this rule takes it to 2.79, which is most of the way to the answer.
constexpr double JACOBIAN_STEP_FRACTION = 0.01;
// Parameters are carried as six lengths in angstroms - the first three are the angle-axis rotation
// vector multiplied by the model's rms radius, so a unit of each of the six moves a typical atom by
// the same amount. That makes the Jacobian step isotropic in something physical, rather than mixing
// radians with angstroms.
struct Placement {
gemmi::Position centre; // the model centroid: rotating about it decorrelates R from t
double rms_radius = 1.0; // rms distance of the atoms from the centroid
void Apply(const double q[6], const std::vector<gemmi::Position> &base, gemmi::Model &model) const {
const double aa[3] = {q[0] / rms_radius, q[1] / rms_radius, q[2] / rms_radius};
size_t i = 0;
for (gemmi::Chain &ch : model.chains)
for (gemmi::Residue &r : ch.residues)
for (gemmi::Atom &a : r.atoms) {
const double p[3] = {base[i].x - centre.x, base[i].y - centre.y, base[i].z - centre.z};
double rp[3];
ceres::AngleAxisRotatePoint(aa, p, rp);
a.pos = gemmi::Position(rp[0] + centre.x + q[3],
rp[1] + centre.y + q[4],
rp[2] + centre.z + q[5]);
++i;
}
}
};
// One target evaluation: place the model, recompute Fcalc and the bulk-solvent mask to the zone's
// resolution, re-fit the scale, and hand back the amplitude residuals.
class Evaluator {
public:
Evaluator(gemmi::Model &model, const gemmi::UnitCell &cell, const gemmi::SpaceGroup &sg,
const std::vector<gemmi::Position> &base, const Placement &placement)
: model_(model), cell_(cell), sg_(sg), base_(base), placement_(placement) {}
// The zone's observations, and the scale the residuals are expressed in.
void SetZone(const gemmi::AsuData<gemmi::ValueSigma<float>> &fobs, double d_min) {
fobs_ = fobs;
d_min_ = d_min;
double sum = 0;
for (const auto &hv : fobs_.v)
sum += hv.value.value;
f_mean_ = fobs_.v.empty() ? 1.0 : sum / static_cast<double>(fobs_.v.size());
solvent_fitted_ = false;
}
size_t NumObservations() const { return fobs_.v.size(); }
double JacobianStep() const { return JACOBIAN_STEP_FRACTION * d_min_; }
int evaluations = 0;
int unmatched = 0; // zone observations with no calculated amplitude to compare against
double k_sol = 0, b_sol = 0; // the bulk solvent the zone's target was evaluated with
bool Residuals(const double q[6], double *residuals) {
++evaluations;
placement_.Apply(q, base_, model_);
gemmi::DensityCalculator<Table, float> dc;
dc.d_min = d_min_;
dc.rate = 1.5;
dc.grid.unit_cell = cell_;
dc.grid.spacegroup = &sg_;
dc.set_refmac_compatible_blur(model_);
dc.put_model_density_on_grid(model_);
gemmi::AsuData<std::complex<float>> fcalc =
gemmi::transform_map_to_f_phi(dc.grid, true).prepare_asu_data(dc.d_min, dc.blur, false, false, false);
gemmi::SolventMasker masker(gemmi::AtomicRadiiSet::Refmac);
gemmi::Grid<float> mask_grid;
mask_grid.unit_cell = cell_;
mask_grid.spacegroup = &sg_;
mask_grid.set_size_from_spacing(dc.requested_grid_spacing(), gemmi::GridSizeRounding::Up);
masker.put_mask_on_grid(mask_grid, model_);
gemmi::AsuData<std::complex<float>> fmask =
gemmi::transform_map_to_f_phi(mask_grid, true).prepare_asu_data(dc.d_min, 0);
if (fmask.size() != fcalc.size())
return false;
// Re-fitted at every evaluation: with the scale held at the starting placement's value the
// target would measure the scale as much as the placement, and the body would translate to
// repair a scale error instead of moving where the density is.
//
// The bulk solvent is not part of that scale. k_sol and b_sol describe the disordered solvent
// of the crystal rather than the fit of one placement, so they are fitted once per zone - by
// FitModelScale, inside the same physical box the reported fit is searched in - and then held
// while the overall scale and the anisotropic B follow the body. Leaving them free at every
// evaluation, which is what gemmi's unbounded Levenberg-Marquardt did here, puts a solvent
// term with no physical meaning inside the target that decides where the model goes: measured
// over a corpus of deposited models, 40% of the evaluations came out with b_sol outside
// 10-80 A^2, some of them negative, which is a solvent that GROWS with resolution.
gemmi::Scaling<float> scaling(cell_, &sg_);
scaling.use_solvent = true;
scaling.prepare_points(fcalc, fobs_, &fmask);
if (scaling.points.empty())
return false;
if (!solvent_fitted_) {
FitModelScale(scaling);
k_sol = scaling.k_sol;
b_sol = scaling.b_sol;
solvent_fitted_ = true;
}
scaling.k_sol = k_sol;
scaling.b_sol = b_sol;
scaling.fix_k_sol = true;
scaling.fix_b_sol = true;
scaling.fit_isotropic_b_approximately();
scaling.fit_parameters();
scaling.scale_data(fcalc, &fmask);
// Both are sorted and in the same ASU, so one merge pass matches them. An observation with no
// calculated amplitude gets residual 0, which drops it from the target rather than scoring it
// as a perfect fit: its Jacobian row below comes out zero as well, and the set that matches is
// fixed by the cell, the group and the zone, so it does not move as the body does. It is
// counted and reported because a large count is a statement about the model rather than about
// this refinement - a group whose reflection conditions the data do not obey leaves half of
// them with nothing to compare against.
auto c = fcalc.v.begin();
unmatched = 0;
for (size_t i = 0; i < fobs_.v.size(); ++i) {
const gemmi::Miller &h = fobs_.v[i].hkl;
while (c != fcalc.v.end() && c->hkl < h)
++c;
const bool matched = c != fcalc.v.end() && c->hkl == h;
if (!matched)
++unmatched;
residuals[i] = matched ? (fobs_.v[i].value.value - std::abs(c->value)) / f_mean_ : 0.0;
}
return true;
}
private:
gemmi::Model &model_;
const gemmi::UnitCell &cell_;
const gemmi::SpaceGroup &sg_;
const std::vector<gemmi::Position> &base_;
Placement placement_;
gemmi::AsuData<gemmi::ValueSigma<float>> fobs_;
double d_min_ = 0;
double f_mean_ = 1;
bool solvent_fitted_ = false;
};
// Ceres' own numeric differentiation steps by |x| * relative_step_size, which is zero at the start of
// every zone (the placement begins at no shift), so the Jacobian is supplied here instead, by
// forward differences at a step chosen in the parameters' units. Analytic dF/dp would need
// derivatives GEMMI's structure-factor path does not have, and at six parameters it is not worth it:
// a Jacobian costs seven evaluations, and the evaluations at 6-3.5 A are cheap.
class RigidBodyCost : public ceres::CostFunction {
public:
explicit RigidBodyCost(Evaluator &ev) : ev_(ev) {
set_num_residuals(static_cast<int>(ev.NumObservations()));
mutable_parameter_block_sizes()->push_back(6);
}
bool Evaluate(double const *const *parameters, double *residuals, double **jacobians) const override {
const int n = num_residuals();
if (!ev_.Residuals(parameters[0], residuals))
return false;
if (jacobians != nullptr && jacobians[0] != nullptr) {
std::vector<double> shifted(n);
for (int j = 0; j < 6; j++) {
double q[6];
std::copy(parameters[0], parameters[0] + 6, q);
const double step = ev_.JacobianStep();
q[j] += step;
if (!ev_.Residuals(q, shifted.data()))
return false;
for (int i = 0; i < n; i++)
jacobians[0][i * 6 + j] = (shifted[i] - residuals[i]) / step;
}
}
return true;
}
private:
Evaluator &ev_;
};
// The directions in which this space group's origin is free. Translating the whole cell content along
// one of them multiplies every F by a phase and leaves every |F| EXACTLY unchanged, so the target
// cannot determine that component: all three directions in P1, the unique axis in a polar group. The
// R-free gate cannot stand in for this - it is a function of |F| too, so along such a direction it
// sees only grid noise and commits or not by coin flip, while the other five parameters carry the
// noise in with them. The free directions are the common fixed subspace of the group's rotation
// parts, and the projector onto it is simply their average.
gemmi::Mat33 GaugeProjector(const gemmi::SpaceGroup &sg, const gemmi::UnitCell &cell) {
const gemmi::GroupOps gops = sg.operations();
double m[3][3] = {};
const double n = static_cast<double>(gops.sym_ops.size()) * gemmi::Op::DEN;
for (const gemmi::Op &op : gops.sym_ops)
for (int i = 0; i < 3; i++)
for (int j = 0; j < 3; j++)
m[i][j] += static_cast<double>(op.rot[i][j]) / n;
const gemmi::Mat33 mean(m[0][0], m[0][1], m[0][2],
m[1][0], m[1][1], m[1][2],
m[2][0], m[2][1], m[2][2]);
// Fractional projector taken into orthogonal space, where the parameters live.
return cell.orth.mat.multiply(mean).multiply(cell.frac.mat);
}
} // namespace
std::vector<gemmi::Position> ModelPositions(const gemmi::Model &model) {
std::vector<gemmi::Position> pos;
for (const gemmi::Chain &ch : model.chains)
for (const gemmi::Residue &r : ch.residues)
for (const gemmi::Atom &a : r.atoms)
pos.push_back(a.pos);
return pos;
}
void SetModelPositions(gemmi::Model &model, const std::vector<gemmi::Position> &pos) {
size_t i = 0;
for (gemmi::Chain &ch : model.chains)
for (gemmi::Residue &r : ch.residues)
for (gemmi::Atom &a : r.atoms)
a.pos = pos[i++];
}
// One rigid body, not groups: a fragment-screening model arrives already solved and isomorphous, and
// the movement to recover is the crystal's, not the molecule's. Splitting it into domains or giving a
// bound ligand its own six parameters would refine against evidence this data does not separately
// carry, and the ligand is what the difference map is meant to show rather than model away.
//
// Some of the translation would be a gauge rather than a quantity - the origin is free in all three
// directions in P1 and along the unique axis in a polar group, and |F| does not change when the whole
// content moves along it - so that component is projected out after every zone. Neither of the two
// things that might look like they cover it actually does: the R-free gate is a function of |F| and
// therefore blind to exactly this, and the LM damping follows the gauge column of the Jacobian, which
// is not zero but noise divided by the difference step.
RigidBodyRefineResult RefineRigidBody(gemmi::Model &model,
const gemmi::UnitCell &cell,
const gemmi::SpaceGroup &sg,
const gemmi::AsuData<gemmi::ValueSigma<float>> &fobs,
double d_min,
Logger &logger) {
const auto t0 = std::chrono::steady_clock::now();
RigidBodyRefineResult result;
const std::vector<gemmi::Position> base = ModelPositions(model);
if (base.empty() || fobs.v.empty())
return result;
Placement placement;
for (const gemmi::Position &p : base)
placement.centre += p;
placement.centre *= 1.0 / static_cast<double>(base.size());
double r2 = 0;
for (const gemmi::Position &p : base)
r2 += placement.centre.dist_sq(p);
placement.rms_radius = std::sqrt(r2 / static_cast<double>(base.size()));
if (!(placement.rms_radius > 0))
return result;
std::vector<double> ladder;
for (double zone : LADDER)
if (zone >= d_min)
ladder.push_back(zone);
if (ladder.empty())
ladder.push_back(d_min);
const gemmi::Mat33 gauge = GaugeProjector(sg, cell);
Evaluator ev(model, cell, sg, base, placement);
double q[6] = {0, 0, 0, 0, 0, 0};
bool any_zone_solved = false;
for (double zone : ladder) {
gemmi::AsuData<gemmi::ValueSigma<float>> zone_obs;
zone_obs.unit_cell_ = fobs.unit_cell_;
zone_obs.spacegroup_ = fobs.spacegroup_;
for (const auto &hv : fobs.v)
if (cell.calculate_d(hv.hkl) >= zone)
zone_obs.v.push_back(hv);
if (zone_obs.v.size() < 50)
continue;
result.zones.push_back(zone); // the ladder WALKED, which a thin zone drops out of
ev.SetZone(zone_obs, zone);
ceres::Problem problem;
problem.AddResidualBlock(new RigidBodyCost(ev), nullptr, q);
ceres::Solver::Options options;
options.linear_solver_type = ceres::DENSE_QR;
options.max_num_iterations = 15;
options.function_tolerance = 1e-4;
options.parameter_tolerance = 1e-4;
options.logging_type = ceres::LoggingType::SILENT;
ceres::Solver::Summary summary;
const int evaluations_before = ev.evaluations;
const auto zone_t0 = std::chrono::steady_clock::now();
ceres::Solve(options, &problem, &summary);
any_zone_solved = any_zone_solved || summary.IsSolutionUsable();
const gemmi::Vec3 along = gauge.multiply(gemmi::Vec3(q[3], q[4], q[5]));
q[3] -= along.x; q[4] -= along.y; q[5] -= along.z;
logger.Debug("Rigid body zone {:.1f} A: {} reflections ({} without a calculated amplitude), "
"solvent k_sol {:.2f} b_sol {:.0f} A^2, {} iterations, {} evaluations, {:.2f} s, "
"rotation {:.3f} deg, translation {:.3f} A", zone, zone_obs.v.size(), ev.unmatched,
ev.k_sol, ev.b_sol,
summary.iterations.empty() ? 0 : summary.iterations.size() - 1,
ev.evaluations - evaluations_before,
std::chrono::duration<double>(std::chrono::steady_clock::now() - zone_t0).count(),
std::sqrt(q[0]*q[0] + q[1]*q[1] + q[2]*q[2]) / placement.rms_radius * 180.0 / PI,
std::sqrt(q[3]*q[3] + q[4]*q[4] + q[5]*q[5]));
}
placement.Apply(q, base, model); // Ceres left the model at a Jacobian probe; put it at the answer
result.evaluations = ev.evaluations;
result.converged = any_zone_solved;
result.k_sol = ev.k_sol;
result.b_sol = ev.b_sol;
const double aa = std::sqrt(q[0] * q[0] + q[1] * q[1] + q[2] * q[2]) / placement.rms_radius;
result.angle_deg = aa * 180.0 / PI;
result.shift_A = std::sqrt(q[3] * q[3] + q[4] * q[4] + q[5] * q[5]);
result.seconds = std::chrono::duration<double>(std::chrono::steady_clock::now() - t0).count();
if (!result.zones.empty())
logger.Info("Model validation: rigid body over {} resolution zone(s) down to {:.1f} A, "
"{} evaluations in {:.2f} s: rotation {:.3f} deg, translation {:.3f} A",
result.zones.size(), result.zones.back(), result.evaluations, result.seconds,
result.angle_deg, result.shift_A);
else
logger.Info("Model validation: rigid body had no resolution zone with enough reflections to "
"run in; the model is left where it arrived");
return result;
}