Files
Jungfraujoch/rugnux/ModelScaling.cpp
T
leonarski_fandClaude Opus 5 1790abf478 rugnux --model: fit the bulk-solvent grid points in parallel
FitModelScale fits ~113 (k_sol, b_sol) grid points, each a full
Levenberg-Marquardt over every working reflection, and it is the largest single
cost of --model (every fit_model call: the first fit, the twin-law probe, the
fit after the rigid body, each null replicate). It ran on one thread.

Each point starts from fit_isotropic_b_approximately(), which sets k_overall
and b_star from the data and the point's solvent pair alone, so no point
depends on the one before it. The points of each pass (coarse, then the
refinement around the coarse winner) now run in contiguous chunks, each on its
own copy of the Scaling, and the winner is read off afterwards in grid order
with the serial rule (lowest finite R, first on a tie). Each point's arithmetic
is unchanged, so the result is the serial loop's bit for bit. Where
fit_isotropic_b_approximately() has five or fewer reflections to fit on it
returns without setting anything and points would chain, so there the grid is
still walked serially on the caller's Scaling.

Inside a null replicate (a pool worker) the chunks run inline, as before.

To check: MODEL_* keys and md5 of the maps/.mtz/_model.cif identical with and
without this commit on the audit set; model-phase time on a fit-dominated set.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:03 +02:00

124 lines
5.6 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#include "ModelScaling.h"
#include <algorithm>
#include <cmath>
#include <vector>
#include "../common/ParallelFor.h"
namespace {
// R-factor of the current parameters over the fitted reflections. This is what the grid is
// selected on, and it is the quantity the scale exists to make small.
double RFactor(const gemmi::Scaling<float> &scaling) {
double num = 0, den = 0;
for (const auto &p : scaling.points) {
num += std::fabs(p.fobs - scaling.compute_value(p));
den += p.fobs;
}
return den > 0 ? num / den : 1.0;
}
} // namespace
// Following the phenix bulk-solvent and scaling procedure: k_sol and b_sol by a grid search, with
// the overall scale and the anisotropic B refitted at every grid point - Afonine, Grosse-Kunstleve
// & Adams, Acta Cryst. D61, 850-855, 2005, which searches b_sol over 10-80 A^2 in steps of 5.
// The fit is unweighted, as in both phenix and Refmac (Murshudov, Skubak, Lebedev, Pannu, Steiner,
// Nicholls, Winn, Long & Vagin, Acta Cryst. D67, 355-367, 2011, eq. 11). The physical range and the
// starting values are those of Fokine & Urzhumtsev, Acta Cryst. D58, 1387-1392, 2002.
//
// The point of the grid is that k_sol and b_sol cannot leave the physical box: gemmi's own
// fit_parameters() is an unbounded Levenberg-Marquardt, and on this corpus it reached b_sol of
// 1707 A^2 - a solvent term switched off in all but the lowest-resolution shell. Here the solvent
// pair is held fixed at each grid point and only the overall scale and the symmetry-constrained
// anisotropic B are refined, which is the well-conditioned half of the problem and is left to
// gemmi's solver rather than reimplemented.
//
// Every grid point is its own fit: fit_isotropic_b_approximately() sets k_overall and b_star from the
// data and the point's solvent pair alone, so a point does not depend on the one fitted before it, and
// the points run in parallel, each chunk on its own copy of `scaling`. The winner is then read off in
// grid order with the serial rule (lowest finite R, the first on a tie), so the answer is the serial
// loop's bit for bit. The one exception is fit_isotropic_b_approximately() finding five or fewer
// reflections to fit on - it then returns without setting anything and a point WOULD start from where
// the previous one ended - so there the grid is walked in order, on `scaling` itself, as it always was.
ModelScaleReport FitModelScale(gemmi::Scaling<float> &scaling, ModelScaleBox box, size_t nthreads) {
ModelScaleReport report;
report.n_points = static_cast<int>(scaling.points.size());
if (scaling.points.size() < 20)
return report;
const bool had_solvent = scaling.use_solvent;
scaling.fix_k_sol = true; // the grid owns the solvent pair; the solver never sees it
scaling.fix_b_sol = true;
// The reflections fit_isotropic_b_approximately() fits on (its own filter).
int n_isotropic = 0;
for (const auto &p : scaling.points)
if (!(p.fobs < 1 || p.fobs < p.sigma))
++n_isotropic;
const bool independent = n_isotropic > 5;
double best_r = -1, best_k_sol = 0.35, best_b_sol = 46.0, best_k_overall = 1.0;
gemmi::SMat33<double> best_b_star{0, 0, 0, 0, 0, 0};
struct PointFit { double k_sol, b_sol, r = NAN, k_overall = 1.0; gemmi::SMat33<double> b_star{0, 0, 0, 0, 0, 0}; };
auto fit_point = [](gemmi::Scaling<float> &s, PointFit &pf) {
s.k_sol = pf.k_sol;
s.b_sol = pf.b_sol;
s.fit_isotropic_b_approximately(); // a fresh starting point for this solvent pair
s.fit_parameters(); // k_overall + anisotropic B only
pf.r = RFactor(s);
pf.k_overall = s.k_overall;
pf.b_star = s.b_star;
};
auto try_points = [&](std::vector<PointFit> &pts) {
if (independent)
ParallelChunks(static_cast<int>(pts.size()), nthreads, [&](int lo, int hi) {
gemmi::Scaling<float> local = scaling;
for (int i = lo; i < hi; ++i)
fit_point(local, pts[i]);
});
else
for (auto &pf : pts)
fit_point(scaling, pf);
for (const auto &pf : pts) {
++report.n_grid;
// A diverged fit gives r = NaN; latched as best_r it wins every later r < best_r.
if (std::isfinite(pf.r) && (best_r < 0 || pf.r < best_r)) {
best_r = pf.r;
best_k_sol = pf.k_sol;
best_b_sol = pf.b_sol;
best_k_overall = pf.k_overall;
best_b_star = pf.b_star;
}
}
};
// Coarse pass over the whole box, then one refinement pass around the winner.
std::vector<PointFit> coarse;
for (double ks = box.k_lo; ks <= box.k_hi + 1e-9; ks += 0.05)
for (double bs = box.b_lo; bs <= box.b_hi + 1e-9; bs += 10.0)
coarse.push_back(PointFit{ks, bs});
try_points(coarse);
const double k0 = best_k_sol, b0 = best_b_sol;
const double k_hi2 = std::min(box.k_hi, k0 + 0.05);
const double b_hi2 = std::min(box.b_hi, b0 + 10.0);
std::vector<PointFit> fine;
for (double ks = std::max(box.k_lo, k0 - 0.05); ks <= k_hi2 + 1e-9; ks += 0.025)
for (double bs = std::max(box.b_lo, b0 - 10.0); bs <= b_hi2 + 1e-9; bs += 5.0)
fine.push_back(PointFit{ks, bs});
try_points(fine);
scaling.k_sol = best_k_sol;
scaling.b_sol = best_b_sol;
scaling.k_overall = best_k_overall;
scaling.b_star = best_b_star;
scaling.use_solvent = had_solvent;
report.r_work_fit = best_r;
return report;
}