Files
Jungfraujoch/image_analysis/structure_refinement/ModelStructureFactorsGPU.h
T
leonarski_fandClaude Opus 5.5 d5fcf2f05b rugnux: model validation's structure factors and maps on the GPU
ModelStructureFactorsGPU computes what compute_model_factors() and
map_from_coefficients() compute on the CPU - F_calc from the model's
density (IT92, Refmac-compatible blur, unblurred as prepare_asu_data()
does) and F_mask from the Refmac bulk-solvent mask, both on the
reflections prepare_asu_data(d_min) lists, in its order; and a map from
ASU coefficients on the grid get_size_for_hkl(coef, 0, 3.0) sizes - on a
device. Made once per cell, group, resolution and model, then evaluated
as often as the coordinates change, so refinement or MR can call it in a
loop. The device is an explicit parameter; every call leaves the calling
thread's current device as it found it.

Pieces:
- ModelDensityGPU: the rigid body's deterministic brick gather, moved
  out of RigidBodyGPU.cu into a component of its own (ModelMaskGPU's
  pattern); the rigid body uses it unchanged. MAX_BRICKS_PER_AXIS 8 ->
  16, so fine grids with high-B atoms (lysozyme at 1.2 A, a 0.9 A P1
  cell) are no longer refused; existing zones are gridded identically.
- One copy of the content is gridded and the symmetry composed in
  reciprocal space (SymmetryComposition), operators applied on the fly;
  the mask is ModelMaskGPU (every image of every atom, islands, shrink).
- Maps: gemmi's get_f_phi_on_grid() in ZYX order on the host (the
  coefficients written are the same), in-place cuFFT c2r, transposed back
  to XYZ on the device. One map at a time, in the engine's buffers.

Decided once, up front, per card, from its TOTAL memory: the engine's
bytes (16 N + cuFFT work + reflections, N the larger of the structure-
factor and map grids) must be at most half the card - the rigid body's
engines take at most a quarter beside it. Otherwise, or where the gather
cannot grid the cell, the CPU path runs, logged with needed vs total.
Anything to a resolution other than d_min (the null's 3.5 A fits) stays
on the CPU, so all replicates and the real model's side of the null are
computed the same way. A CUDA failure takes the existing path: the
validation restarts on the CPU.

Measured, model validation total per run (CPU path -> GPU), 16 GB card:
  F432 215 A cubic, 1.30 A, 500^3 grid: 47.7 -> 15.6 s (two validations;
     14.3 -> 2.4 and 33.4 -> 13.2, the rest of the second is writing the
     three 0.5 GB maps); F_calc + F_mask 5.7 s -> 0.05 s
  C2 1.11 A: 23.9 -> 13.6 s; P3_2 1.55 A: 18.2 -> 10.2 s;
  P2_1 1.25 A: 13.4 -> 6.5 s; F4_132 328 A: 13.0 -> 5.5 s;
  P6_5: 8.8 -> 4.2 s; P4_3 0.97 A: 4.6 -> 2.5 s; P1 0.92 A: 4.2 -> 2.2 s;
  small P1: 3.2 -> 1.3 s; P6_1: 8.8 -> 6.2 s; lysozyme: 1.8 -> 1.4 s.
p.mtz md5-identical on all 13 sets. Against the CPU path: FC within
1e-4 of mean |F|, phases of the strong half within 0.003 deg, maps within
1e-4 (2mFo-DFc) and 7e-4 (mFo-DFc) of their rms; every logged R, CC,
FOM, k_sol and anomalous site list identical at the printed precision,
except where a rigid-body commit sat on an exact R-free tie (0.2155 ->
0.2155) and fell the other way (R-work 0.2127 vs 0.2129). The GPU result
is bit-identical run to run and with -N 8 (maps, map MTZ, placed model).
Peak device memory of the engine: 2.5 GB at 500^3 (process total peaked
at 14.4 GB with what the merge still holds).

Tests: ModelStructureFactorsGPU_MatchesCPU (five groups, 3.5 and 1.5 A:
same reflections, F_calc <= 1e-5 of mean |F|, F_mask 2e-7 rms, repeat
bit-identical), ModelStructureFactorsGPU_MapMatchesCPU (<= 5e-6 of rms).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 12:38:37 +02:00

96 lines
5.1 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
// A model's structure factors, and the maps made from coefficients on its reflections, on the GPU (CUDA
// builds only). It is what model validation computes on the CPU - F_calc from the model's density on a
// grid (gemmi's DensityCalculator, IT92, the Refmac-compatible blur and its unblur) and F_mask from the
// bulk-solvent mask (gemmi's SolventMasker with the Refmac radii), each transformed and read off as
// prepare_asu_data() lists them; and the inverse transform of a set of map coefficients, as
// get_f_phi_on_grid() and MapFromFPhi() make it - moved to the device. The two agree to rounding, not
// bit for bit: float distances and cuFFT for FFTW, and the crystal's symmetry is composed in reciprocal
// space (SymmetryComposition, RigidBodyRefine.h) instead of by symmetrizing the grid. The GPU is
// deterministic on its own.
//
// Made once for a cell, a group, a resolution and a model's atoms, then evaluated as often as the
// coordinates change: everything that depends only on the first four is worked out here, once, and
// the device buffers are reserved once. The atoms' B factors and occupancies are read at every
// evaluation, but the blur is the one the model had when the engine was made.
#include <complex>
#include <cstddef>
#include <memory>
#include <mutex>
#include <vector>
#include "gemmi/asudata.hpp"
#include "gemmi/dencalc.hpp" // DensityCalculator
#include "gemmi/grid.hpp"
#include "gemmi/it92.hpp"
#include "gemmi/model.hpp"
#include "gemmi/symmetry.hpp"
#include "gemmi/unitcell.hpp"
#include "ModelStructureFactorsGPUEngine.h"
// Each atom's density as PutModelDensityOnGrid() (ModelGrid.cpp) sets it up for `dc` - its d_min, rate
// and blur - in model order; and the atoms of the bulk-solvent mask as PutMaskOnGrid() takes them: each
// one's index in the model and its radius, probe included.
void ModelDensityAtoms(const gemmi::Model &model, const gemmi::DensityCalculator<gemmi::IT92<float>, float> &dc,
std::vector<ModelDensityAtom> &atoms, std::vector<int> &mask_atom,
std::vector<float> &mask_radius);
class ModelStructureFactorsGPU {
public:
// For `device`. Host work only: the grid, the reflections and what the device will need for them.
// Nothing is reserved until Reserve(), so DeviceBytes() can decide first whether to. Every call works
// on `device` and leaves the calling thread's current device as it found it; which card it is does not
// change a number.
ModelStructureFactorsGPU(int device, const gemmi::Model &model, const gemmi::UnitCell &cell,
const gemmi::SpaceGroup &sg, double d_min);
~ModelStructureFactorsGPU();
// Whether the GPU reproduces this case: the gather needs every atom's box narrower than the cell.
bool Supported() const { return supported_; }
// The device memory Reserve() takes: the grid and its transform, the solvent mask's labels, the cuFFT
// work area and the reflections. The largest map Map() can be given shares the same buffers.
size_t DeviceBytes() const { return device_bytes_; }
double DMin() const { return d_min_; }
int Device() const { return device_; }
// The device's total memory, which DeviceBytes() is to be judged against.
size_t DeviceTotalMemory() const { return ModelStructureFactorsGPUEngine::TotalMemory(device_); }
// The grid the structure factors are computed on.
std::array<int, 3> GridSize() const { return {setup_.grid.nu, setup_.grid.nv, setup_.grid.nw}; }
std::array<int, 3> MapSizeBound() const { return map_size_; }
// Reserves the device memory. Throws JFJochException on a CUDA failure.
void Reserve();
// F_calc and F_mask of `model` - the model this was made for, its atoms anywhere - as model validation's
// compute_model_factors() makes them: prepare_asu_data(d_min, blur) of the density and
// prepare_asu_data(d_min) of the mask. Throws JFJochException on a CUDA failure. Safe to call from
// several threads; they take turns.
void Compute(const gemmi::Model &model, gemmi::AsuData<std::complex<float>> &fcalc,
gemmi::AsuData<std::complex<float>> &fmask);
// The real-space map of ASU coefficients on this engine's reflections, as model validation's
// map_from_coefficients() makes it: on the grid get_size_for_hkl(coef, {0, 0, 0}, 3.0) sizes. Sorts
// `coef`. Throws JFJochException on a CUDA failure. Safe to call from several threads; they take turns.
gemmi::Grid<float> Map(gemmi::AsuData<std::complex<float>> &coef);
private:
int device_;
gemmi::UnitCell cell_;
const gemmi::SpaceGroup *sg_;
double d_min_;
gemmi::DensityCalculator<gemmi::IT92<float>, float> dc_;
ModelStructureFactorsGPUSetup setup_;
std::vector<gemmi::Miller> rows_;
std::array<int, 3> map_size_{};
bool supported_ = false;
size_t fft_work_bytes_ = 0, device_bytes_ = 0;
std::unique_ptr<ModelStructureFactorsGPUEngine> engine_;
std::mutex m_;
};