Files
Jungfraujoch/image_analysis/structure_refinement/ModelStructureFactorsGPU.h
T
leonarski_fandClaude Opus 5.5 278c488c37 Merge branch 'modelpar' into gpusf
Both sides kept: modelpar's parallel basis scoring, the second validation
started on a forecast beside the first (write gate, schedule parameter),
the multi-GPU placement and delete-before-rewrite; gpusf's GPU structure
factors, maps and null engine, failure-instead-of-restart, and the merge
engine released before the validation (now just before modelpar's
ValidateAgainstModel call, after the forecast lambda is set up).

Placement: each validation's structure-factor engines (d_min and the
null's) are made on the card of the thread that runs it - the main
thread's for the first validation, card 1 % count for the speculative
second, which pins itself there - so two validations on two cards use
both, as the rigid-body pools do.

One memory rule for both, per card, from total memory, up front: a
validation plans at most half of its card - its structure-factor engines
a quarter together (was half for the d_min engine alone), its rigid-body
engines a quarter (RigidBodyGPUPool) - and the second validation runs
beside the first only where twice the first's plan (rigid-body planned
bytes + the d_min engine, twice it where a null is coming, the null's
engine being no larger) fits half of all cards' memory together (was:
twice the rigid-body plan within a quarter). The card's total is read
once when the engine is made; a CUDA error there fails the validation
like any other.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 13:51:41 +02:00

98 lines
5.1 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
// A model's structure factors, and the maps made from coefficients on its reflections, on the GPU (CUDA
// builds only). It is what model validation computes on the CPU - F_calc from the model's density on a
// grid (gemmi's DensityCalculator, IT92, the Refmac-compatible blur and its unblur) and F_mask from the
// bulk-solvent mask (gemmi's SolventMasker with the Refmac radii), each transformed and read off as
// prepare_asu_data() lists them; and the inverse transform of a set of map coefficients, as
// get_f_phi_on_grid() and MapFromFPhi() make it - moved to the device. The two agree to rounding, not
// bit for bit: float distances and cuFFT for FFTW, and the crystal's symmetry is composed in reciprocal
// space (SymmetryComposition, RigidBodyRefine.h) instead of by symmetrizing the grid. The GPU is
// deterministic on its own.
//
// Made once for a cell, a group, a resolution and a model's atoms, then evaluated as often as the
// coordinates change: everything that depends only on the first four is worked out here, once, and
// the device buffers are reserved once. The atoms' B factors and occupancies are read at every
// evaluation, but the blur is the one the model had when the engine was made.
#include <complex>
#include <cstddef>
#include <memory>
#include <mutex>
#include <vector>
#include "gemmi/asudata.hpp"
#include "gemmi/dencalc.hpp" // DensityCalculator
#include "gemmi/grid.hpp"
#include "gemmi/it92.hpp"
#include "gemmi/model.hpp"
#include "gemmi/symmetry.hpp"
#include "gemmi/unitcell.hpp"
#include "ModelStructureFactorsGPUEngine.h"
// Each atom's density as PutModelDensityOnGrid() (ModelGrid.cpp) sets it up for `dc` - its d_min, rate
// and blur - in model order; and the atoms of the bulk-solvent mask as PutMaskOnGrid() takes them: each
// one's index in the model and its radius, probe included.
void ModelDensityAtoms(const gemmi::Model &model, const gemmi::DensityCalculator<gemmi::IT92<float>, float> &dc,
std::vector<ModelDensityAtom> &atoms, std::vector<int> &mask_atom,
std::vector<float> &mask_radius);
class ModelStructureFactorsGPU {
public:
// For `device`. Host work only: the grid, the reflections and what the device will need for them.
// Nothing is reserved until Reserve(), so DeviceBytes() can decide first whether to. Throws
// JFJochException where the device cannot be asked for its memory or its transform sizes. Every call works
// on `device` and leaves the calling thread's current device as it found it; which card it is does not
// change a number.
ModelStructureFactorsGPU(int device, const gemmi::Model &model, const gemmi::UnitCell &cell,
const gemmi::SpaceGroup &sg, double d_min);
~ModelStructureFactorsGPU();
// Whether the GPU reproduces this case: the gather needs every atom's box narrower than the cell.
bool Supported() const { return supported_; }
// The device memory Reserve() takes: the grid and its transform, the solvent mask's labels, the cuFFT
// work area and the reflections. The largest map Map() can be given shares the same buffers.
size_t DeviceBytes() const { return device_bytes_; }
double DMin() const { return d_min_; }
int Device() const { return device_; }
// The device's total memory, which DeviceBytes() is to be judged against.
size_t DeviceTotalMemory() const { return device_total_; }
// The grid the structure factors are computed on.
std::array<int, 3> GridSize() const { return {setup_.grid.nu, setup_.grid.nv, setup_.grid.nw}; }
std::array<int, 3> MapSizeBound() const { return map_size_; }
// Reserves the device memory. Throws JFJochException on a CUDA failure.
void Reserve();
// F_calc and F_mask of `model` - the model this was made for, its atoms anywhere - as model validation's
// compute_model_factors() makes them: prepare_asu_data(d_min, blur) of the density and
// prepare_asu_data(d_min) of the mask. Throws JFJochException on a CUDA failure. Safe to call from
// several threads; they take turns.
void Compute(const gemmi::Model &model, gemmi::AsuData<std::complex<float>> &fcalc,
gemmi::AsuData<std::complex<float>> &fmask);
// The real-space map of ASU coefficients on this engine's reflections, as model validation's
// map_from_coefficients() makes it: on the grid get_size_for_hkl(coef, {0, 0, 0}, 3.0) sizes. Sorts
// `coef`. Throws JFJochException on a CUDA failure. Safe to call from several threads; they take turns.
gemmi::Grid<float> Map(gemmi::AsuData<std::complex<float>> &coef);
private:
int device_;
size_t device_total_;
gemmi::UnitCell cell_;
const gemmi::SpaceGroup *sg_;
double d_min_;
gemmi::DensityCalculator<gemmi::IT92<float>, float> dc_;
ModelStructureFactorsGPUSetup setup_;
std::vector<gemmi::Miller> rows_;
std::array<int, 3> map_size_{};
bool supported_ = false;
size_t fft_work_bytes_ = 0, device_bytes_ = 0;
std::unique_ptr<ModelStructureFactorsGPUEngine> engine_;
std::mutex m_;
};