Files
Jungfraujoch/image_analysis/structure_refinement/RigidBodyGPUEngine.h
T
leonarski_fandClaude Opus 5.5 d5fcf2f05b rugnux: model validation's structure factors and maps on the GPU
ModelStructureFactorsGPU computes what compute_model_factors() and
map_from_coefficients() compute on the CPU - F_calc from the model's
density (IT92, Refmac-compatible blur, unblurred as prepare_asu_data()
does) and F_mask from the Refmac bulk-solvent mask, both on the
reflections prepare_asu_data(d_min) lists, in its order; and a map from
ASU coefficients on the grid get_size_for_hkl(coef, 0, 3.0) sizes - on a
device. Made once per cell, group, resolution and model, then evaluated
as often as the coordinates change, so refinement or MR can call it in a
loop. The device is an explicit parameter; every call leaves the calling
thread's current device as it found it.

Pieces:
- ModelDensityGPU: the rigid body's deterministic brick gather, moved
  out of RigidBodyGPU.cu into a component of its own (ModelMaskGPU's
  pattern); the rigid body uses it unchanged. MAX_BRICKS_PER_AXIS 8 ->
  16, so fine grids with high-B atoms (lysozyme at 1.2 A, a 0.9 A P1
  cell) are no longer refused; existing zones are gridded identically.
- One copy of the content is gridded and the symmetry composed in
  reciprocal space (SymmetryComposition), operators applied on the fly;
  the mask is ModelMaskGPU (every image of every atom, islands, shrink).
- Maps: gemmi's get_f_phi_on_grid() in ZYX order on the host (the
  coefficients written are the same), in-place cuFFT c2r, transposed back
  to XYZ on the device. One map at a time, in the engine's buffers.

Decided once, up front, per card, from its TOTAL memory: the engine's
bytes (16 N + cuFFT work + reflections, N the larger of the structure-
factor and map grids) must be at most half the card - the rigid body's
engines take at most a quarter beside it. Otherwise, or where the gather
cannot grid the cell, the CPU path runs, logged with needed vs total.
Anything to a resolution other than d_min (the null's 3.5 A fits) stays
on the CPU, so all replicates and the real model's side of the null are
computed the same way. A CUDA failure takes the existing path: the
validation restarts on the CPU.

Measured, model validation total per run (CPU path -> GPU), 16 GB card:
  F432 215 A cubic, 1.30 A, 500^3 grid: 47.7 -> 15.6 s (two validations;
     14.3 -> 2.4 and 33.4 -> 13.2, the rest of the second is writing the
     three 0.5 GB maps); F_calc + F_mask 5.7 s -> 0.05 s
  C2 1.11 A: 23.9 -> 13.6 s; P3_2 1.55 A: 18.2 -> 10.2 s;
  P2_1 1.25 A: 13.4 -> 6.5 s; F4_132 328 A: 13.0 -> 5.5 s;
  P6_5: 8.8 -> 4.2 s; P4_3 0.97 A: 4.6 -> 2.5 s; P1 0.92 A: 4.2 -> 2.2 s;
  small P1: 3.2 -> 1.3 s; P6_1: 8.8 -> 6.2 s; lysozyme: 1.8 -> 1.4 s.
p.mtz md5-identical on all 13 sets. Against the CPU path: FC within
1e-4 of mean |F|, phases of the strong half within 0.003 deg, maps within
1e-4 (2mFo-DFc) and 7e-4 (mFo-DFc) of their rms; every logged R, CC,
FOM, k_sol and anomalous site list identical at the printed precision,
except where a rigid-body commit sat on an exact R-free tie (0.2155 ->
0.2155) and fell the other way (R-work 0.2127 vs 0.2129). The GPU result
is bit-identical run to run and with -N 8 (maps, map MTZ, placed model).
Peak device memory of the engine: 2.5 GB at 500^3 (process total peaked
at 14.4 GB with what the merge still holds).

Tests: ModelStructureFactorsGPU_MatchesCPU (five groups, 3.5 and 1.5 A:
same reflections, F_calc <= 1e-5 of mean |F|, F_mask 2e-7 rms, repeat
bit-identical), ModelStructureFactorsGPU_MapMatchesCPU (<= 5e-6 of rms).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 12:38:37 +02:00

133 lines
6.7 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
// The device half of RigidBodyTargetGPU (RigidBodyGPU.h): one engine is one CUDA stream and the buffers
// for one resolution zone of one fit. The host half works out everything that depends only on the
// model, the cell and the zone - with gemmi, which stays out of nvcc - and hands it over in the plain
// structures below; the engine then does, per evaluation, what depends on the placement. CUDA builds
// only.
#include <array>
#include <cstddef>
#include <memory>
#include <stdexcept>
#include <vector>
#include "ModelDensityGPU.h" // ModelDensityAtom
// One term of SymmetryComposition: F1 at k = hR, times the phase of the operator's translation.
struct RigidBodyGPUTerm {
int index; // of F1(k) in the half-u transform, or of F1(-k) where k is outside the stored half
int conj; // read as the Friedel mate, conj F1(-k)
double phase[2]; // exp(+2 pi i h.t), real and imaginary
double s[3]; // k as a Cartesian reciprocal vector
};
// Everything a zone needs that does not depend on the placement.
struct RigidBodyGPUZone {
int nu = 0, nv = 0, nw = 0; // the zone's grid, u fastest
double orth[9] = {}, frac[9] = {}; // row-major
double volume = 0;
double blur = 0; // DensityCalculator's, which the rows' unblur undoes
std::vector<ModelDensityAtom> atoms; // model order
// The bulk-solvent mask's atoms: the model's index of each, and its radius (probe included).
std::vector<int> mask_atom;
std::vector<float> mask_radius;
// Every image: the group's operators, each with each centring vector, fractional.
std::vector<std::array<double, 12>> images; // rot[9] row-major, tran[3]
// The composition: rows (composed indices) and Ops() terms per row, row-major.
std::vector<std::array<int, 3>> row_hkl;
std::vector<double> row_scale; // n_cen * unblur
std::vector<double> row_stol2;
size_t ops = 1;
std::vector<RigidBodyGPUTerm> terms;
// The observations the residuals are over: each one's row (-1 without one) and amplitude.
std::vector<int> obs_row;
std::vector<float> obs_fobs, obs_sigma;
double f_mean = 1;
// The scale's points: the observations gemmi's prepare_points() would take, in order.
std::vector<int> point_obs;
std::vector<std::array<int, 3>> point_hkl;
std::vector<double> point_stol2;
std::vector<float> point_fobs, point_sigma;
std::vector<std::array<double, 6>> constraints; // adp_symmetry_constraints()
};
// Upper bounds an engine is sized for.
struct RigidBodyGPUCapacity {
size_t atoms = 0;
size_t grid_points = 0; // nu * nv * nw
size_t complex_points = 0; // (nu / 2 + 1) * nv * nw
size_t bricks = 0;
size_t pairs = 0; // (brick, atom) pairs of the gather
size_t rows = 0, terms = 0;
size_t observations = 0;
size_t fft_work_bytes = 0;
};
// FitModelScale()'s solvent grid on five or fewer strong reflections, where its fits chain from one grid
// point to the next and the device does not reproduce it (ModelScaleGPU::FitSolvent).
class ModelScaleGPUTooFewReflections : public std::runtime_error {
public:
ModelScaleGPUTooFewReflections() : std::runtime_error("ModelScaleGPU::FitSolvent: too few reflections for independent grid points") {}
};
struct RigidBodyGPUEngineImpl;
class RigidBodyGPUEngine {
public:
// The bytes an engine of this capacity reserves on the device.
static size_t DeviceBytes(const RigidBodyGPUCapacity &capacity);
// The largest cuFFT work area a (nu, nv, nw) grid needs, batch 1 or 3.
static size_t FFTWorkBytes(int nu, int nv, int nw);
// (brick, atom) pairs of the gather over at most, for a zone's atoms on its grid.
static size_t PairBound(const RigidBodyGPUZone &zone);
static size_t Bricks(int nu, int nv, int nw);
// The most gather bricks the 2 d + 1 points of an atom's box can fall in along an axis of n points.
static size_t AxisBrickBound(int d, int n);
// Whether the gather reproduces gemmi's box walk on this zone: every atom's box narrower than the cell,
// so that no point is reached by two images of one atom.
static bool Supports(const RigidBodyGPUZone &zone);
static void MemoryInfo(size_t &free, size_t &total);
static int CurrentDevice();
RigidBodyGPUEngine(const RigidBodyGPUCapacity &capacity, int device);
~RigidBodyGPUEngine();
// Per fit: each atom's position relative to the model centroid, which the placements rotate.
void SetBody(const std::vector<std::array<double, 3>> &relative);
// Per zone. Throws if the zone does not fit the capacity.
void SetZone(const RigidBodyGPUZone &zone);
// Per evaluation, at the placement x -> R x_rel + t: Fcalc and dF/dt at the rows, then the bulk-
// solvent mask of the same placement (ModelMaskGPU) and Fmask.
void Fcalc(const double rotation[9], const double translation[3]);
void Fmask();
// The scale at this evaluation's Fcalc and Fmask (ModelScaleGPU): FitModelScale()'s solvent grid, and
// the overall scale and anisotropic B at a fixed solvent. FitSolvent() throws ModelScaleGPUTooFewReflections where
// the grid's fits would chain (too few reflections for the isotropic start), which the host then does.
void FitSolvent(double &k_sol, double &b_sol);
void FitScale(double k_sol, double b_sol, double &k_overall, double b_star[6]);
// The scale's points: fcmol (Fcalc as complex<float>) and fmask, downloaded for the host's fit.
void DownloadPoints(std::vector<std::array<float, 2>> &fcmol, std::vector<std::array<float, 2>> &fmask);
// The residuals at the scale given, NumObservations() of them, into `residuals` (host).
void Residuals(double k_overall, const double b_star[6], double k_sol, double b_sol, double *residuals);
// The Jacobian at the last evaluation: `rotation[j]` and `translation[j]` place the body one step
// along rotation axis j. Writes J (fixed scale) and J_k (the scale parameters' columns) on the
// device and returns J_k^T J_k (p x p) and J_k^T J (p x 6), p = 1 + constraints, row-major.
void Jacobian(const double rotation[3][9], const double translation[3][3], double step, double k_overall,
const double b_star[6], double k_sol, double b_sol, std::vector<double> &jtj,
std::vector<double> &jtq);
// J - J_k x, x p x 6 row-major, into `jacobian` (host, NumObservations() x 6).
void ProjectJacobian(const std::vector<double> &x, double *jacobian);
private:
std::unique_ptr<RigidBodyGPUEngineImpl> impl_;
};