Files
Jungfraujoch/image_analysis/structure_refinement/ModelStructureFactorsGPUEngine.h
T
leonarski_fandClaude Opus 5.5 d5fcf2f05b rugnux: model validation's structure factors and maps on the GPU
ModelStructureFactorsGPU computes what compute_model_factors() and
map_from_coefficients() compute on the CPU - F_calc from the model's
density (IT92, Refmac-compatible blur, unblurred as prepare_asu_data()
does) and F_mask from the Refmac bulk-solvent mask, both on the
reflections prepare_asu_data(d_min) lists, in its order; and a map from
ASU coefficients on the grid get_size_for_hkl(coef, 0, 3.0) sizes - on a
device. Made once per cell, group, resolution and model, then evaluated
as often as the coordinates change, so refinement or MR can call it in a
loop. The device is an explicit parameter; every call leaves the calling
thread's current device as it found it.

Pieces:
- ModelDensityGPU: the rigid body's deterministic brick gather, moved
  out of RigidBodyGPU.cu into a component of its own (ModelMaskGPU's
  pattern); the rigid body uses it unchanged. MAX_BRICKS_PER_AXIS 8 ->
  16, so fine grids with high-B atoms (lysozyme at 1.2 A, a 0.9 A P1
  cell) are no longer refused; existing zones are gridded identically.
- One copy of the content is gridded and the symmetry composed in
  reciprocal space (SymmetryComposition), operators applied on the fly;
  the mask is ModelMaskGPU (every image of every atom, islands, shrink).
- Maps: gemmi's get_f_phi_on_grid() in ZYX order on the host (the
  coefficients written are the same), in-place cuFFT c2r, transposed back
  to XYZ on the device. One map at a time, in the engine's buffers.

Decided once, up front, per card, from its TOTAL memory: the engine's
bytes (16 N + cuFFT work + reflections, N the larger of the structure-
factor and map grids) must be at most half the card - the rigid body's
engines take at most a quarter beside it. Otherwise, or where the gather
cannot grid the cell, the CPU path runs, logged with needed vs total.
Anything to a resolution other than d_min (the null's 3.5 A fits) stays
on the CPU, so all replicates and the real model's side of the null are
computed the same way. A CUDA failure takes the existing path: the
validation restarts on the CPU.

Measured, model validation total per run (CPU path -> GPU), 16 GB card:
  F432 215 A cubic, 1.30 A, 500^3 grid: 47.7 -> 15.6 s (two validations;
     14.3 -> 2.4 and 33.4 -> 13.2, the rest of the second is writing the
     three 0.5 GB maps); F_calc + F_mask 5.7 s -> 0.05 s
  C2 1.11 A: 23.9 -> 13.6 s; P3_2 1.55 A: 18.2 -> 10.2 s;
  P2_1 1.25 A: 13.4 -> 6.5 s; F4_132 328 A: 13.0 -> 5.5 s;
  P6_5: 8.8 -> 4.2 s; P4_3 0.97 A: 4.6 -> 2.5 s; P1 0.92 A: 4.2 -> 2.2 s;
  small P1: 3.2 -> 1.3 s; P6_1: 8.8 -> 6.2 s; lysozyme: 1.8 -> 1.4 s.
p.mtz md5-identical on all 13 sets. Against the CPU path: FC within
1e-4 of mean |F|, phases of the strong half within 0.003 deg, maps within
1e-4 (2mFo-DFc) and 7e-4 (mFo-DFc) of their rms; every logged R, CC,
FOM, k_sol and anomalous site list identical at the printed precision,
except where a rigid-body commit sat on an exact R-free tie (0.2155 ->
0.2155) and fell the other way (R-work 0.2127 vs 0.2129). The GPU result
is bit-identical run to run and with -N 8 (maps, map MTZ, placed model).
Peak device memory of the engine: 2.5 GB at 500^3 (process total peaked
at 14.4 GB with what the merge still holds).

Tests: ModelStructureFactorsGPU_MatchesCPU (five groups, 3.5 and 1.5 A:
same reflections, F_calc <= 1e-5 of mean |F|, F_mask 2e-7 rms, repeat
bit-identical), ModelStructureFactorsGPU_MapMatchesCPU (<= 5e-6 of rms).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 12:38:37 +02:00

65 lines
3.3 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
// The device half of ModelStructureFactorsGPU (ModelStructureFactorsGPU.h). The host half works out with
// gemmi - which stays out of nvcc - everything that depends on the cell, the group and the resolution, and
// per evaluation the atoms' densities and positions; this half grids, transforms and composes. CUDA builds
// only.
#include <array>
#include <cstddef>
#include <memory>
#include <vector>
#include "ModelDensityGPU.h"
#include "ModelMaskGPU.h"
// Everything an engine needs that does not depend on the model's coordinates.
struct ModelStructureFactorsGPUSetup {
ModelDensityGrid grid; // the structure-factor grid, u fastest
double volume = 0;
// The group's operators as gemmi holds them: rot (row-major) and tran, in units of 1 / den.
std::vector<std::array<int, 12>> sym_ops;
int den = 24;
std::vector<ModelMaskOp> mask_ops; // every operator with every centring vector
std::vector<std::array<int, 3>> rows; // the reflections computed, in the order they come back
std::vector<double> row_scale; // n_cen * prepare_asu_data()'s unblur, per row
size_t max_atoms = 0, max_pairs = 0;
// The largest map MapFromFPhi() will be given: points of the real grid, and of its half-l transform.
size_t map_points = 0, map_complex_points = 0;
};
struct ModelStructureFactorsGPUEngineImpl;
class ModelStructureFactorsGPUEngine {
public:
// Every call names its device and leaves the calling thread's current device as it found it.
// The device's total memory.
static size_t TotalMemory(int device);
// The cuFFT work area the structure factors' r2c, and a map of (nu, nv, nw)'s c2r, take on `device`.
static size_t FFTWorkBytes(int device, const ModelDensityGrid &grid, int map_nu, int map_nv, int map_nw);
// The bytes an engine for `setup` reserves on the device, given its cuFFT work area.
static size_t DeviceBytes(const ModelStructureFactorsGPUSetup &setup, size_t fft_work_bytes);
ModelStructureFactorsGPUEngine(int device, const ModelStructureFactorsGPUSetup &setup, size_t fft_work_bytes);
~ModelStructureFactorsGPUEngine();
// F_calc and F_mask at the rows for one set of atoms: each atom's density, its grid position
// (fractional times the grid size, wrapped into [0, n); the 4th number is not read) and the atoms of
// the bulk-solvent mask. Both come back as (re, im) per row.
void Compute(const std::vector<ModelDensityAtom> &atoms, const std::vector<std::array<float, 4>> &pos,
const std::vector<ModelMaskAtom> &mask_atoms, std::vector<std::array<float, 2>> &fcalc,
std::vector<std::array<float, 2>> &fmask);
// The real-space map rho(x) = 1/V sum F exp(-2 pi i h.x) of a half-l coefficient grid in gemmi's ZYX
// order - h slowest, l fastest and halved: nu * nv * (nw / 2 + 1) complex numbers, (re, im)
// interleaved, NaN read as zero - into `map`, nu * nv * nw floats with x fastest.
void Map(int nu, int nv, int nw, double volume, const float *coefficients, float *map);
private:
std::unique_ptr<ModelStructureFactorsGPUEngineImpl> impl_;
};