Both sides kept: modelpar's parallel basis scoring, the second validation started on a forecast beside the first (write gate, schedule parameter), the multi-GPU placement and delete-before-rewrite; gpusf's GPU structure factors, maps and null engine, failure-instead-of-restart, and the merge engine released before the validation (now just before modelpar's ValidateAgainstModel call, after the forecast lambda is set up). Placement: each validation's structure-factor engines (d_min and the null's) are made on the card of the thread that runs it - the main thread's for the first validation, card 1 % count for the speculative second, which pins itself there - so two validations on two cards use both, as the rigid-body pools do. One memory rule for both, per card, from total memory, up front: a validation plans at most half of its card - its structure-factor engines a quarter together (was half for the d_min engine alone), its rigid-body engines a quarter (RigidBodyGPUPool) - and the second validation runs beside the first only where twice the first's plan (rigid-body planned bytes + the d_min engine, twice it where a null is coming, the null's engine being no larger) fits half of all cards' memory together (was: twice the rigid-body plan within a quarter). The card's total is read once when the engine is made; a CUDA error there fails the validation like any other. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
98 lines
5.1 KiB
C++
98 lines
5.1 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
// A model's structure factors, and the maps made from coefficients on its reflections, on the GPU (CUDA
|
|
// builds only). It is what model validation computes on the CPU - F_calc from the model's density on a
|
|
// grid (gemmi's DensityCalculator, IT92, the Refmac-compatible blur and its unblur) and F_mask from the
|
|
// bulk-solvent mask (gemmi's SolventMasker with the Refmac radii), each transformed and read off as
|
|
// prepare_asu_data() lists them; and the inverse transform of a set of map coefficients, as
|
|
// get_f_phi_on_grid() and MapFromFPhi() make it - moved to the device. The two agree to rounding, not
|
|
// bit for bit: float distances and cuFFT for FFTW, and the crystal's symmetry is composed in reciprocal
|
|
// space (SymmetryComposition, RigidBodyRefine.h) instead of by symmetrizing the grid. The GPU is
|
|
// deterministic on its own.
|
|
//
|
|
// Made once for a cell, a group, a resolution and a model's atoms, then evaluated as often as the
|
|
// coordinates change: everything that depends only on the first four is worked out here, once, and
|
|
// the device buffers are reserved once. The atoms' B factors and occupancies are read at every
|
|
// evaluation, but the blur is the one the model had when the engine was made.
|
|
|
|
#include <complex>
|
|
#include <cstddef>
|
|
#include <memory>
|
|
#include <mutex>
|
|
#include <vector>
|
|
|
|
#include "gemmi/asudata.hpp"
|
|
#include "gemmi/dencalc.hpp" // DensityCalculator
|
|
#include "gemmi/grid.hpp"
|
|
#include "gemmi/it92.hpp"
|
|
#include "gemmi/model.hpp"
|
|
#include "gemmi/symmetry.hpp"
|
|
#include "gemmi/unitcell.hpp"
|
|
|
|
#include "ModelStructureFactorsGPUEngine.h"
|
|
|
|
// Each atom's density as PutModelDensityOnGrid() (ModelGrid.cpp) sets it up for `dc` - its d_min, rate
|
|
// and blur - in model order; and the atoms of the bulk-solvent mask as PutMaskOnGrid() takes them: each
|
|
// one's index in the model and its radius, probe included.
|
|
void ModelDensityAtoms(const gemmi::Model &model, const gemmi::DensityCalculator<gemmi::IT92<float>, float> &dc,
|
|
std::vector<ModelDensityAtom> &atoms, std::vector<int> &mask_atom,
|
|
std::vector<float> &mask_radius);
|
|
|
|
class ModelStructureFactorsGPU {
|
|
public:
|
|
// For `device`. Host work only: the grid, the reflections and what the device will need for them.
|
|
// Nothing is reserved until Reserve(), so DeviceBytes() can decide first whether to. Throws
|
|
// JFJochException where the device cannot be asked for its memory or its transform sizes. Every call works
|
|
// on `device` and leaves the calling thread's current device as it found it; which card it is does not
|
|
// change a number.
|
|
ModelStructureFactorsGPU(int device, const gemmi::Model &model, const gemmi::UnitCell &cell,
|
|
const gemmi::SpaceGroup &sg, double d_min);
|
|
~ModelStructureFactorsGPU();
|
|
|
|
// Whether the GPU reproduces this case: the gather needs every atom's box narrower than the cell.
|
|
bool Supported() const { return supported_; }
|
|
// The device memory Reserve() takes: the grid and its transform, the solvent mask's labels, the cuFFT
|
|
// work area and the reflections. The largest map Map() can be given shares the same buffers.
|
|
size_t DeviceBytes() const { return device_bytes_; }
|
|
double DMin() const { return d_min_; }
|
|
int Device() const { return device_; }
|
|
// The device's total memory, which DeviceBytes() is to be judged against.
|
|
size_t DeviceTotalMemory() const { return device_total_; }
|
|
// The grid the structure factors are computed on.
|
|
std::array<int, 3> GridSize() const { return {setup_.grid.nu, setup_.grid.nv, setup_.grid.nw}; }
|
|
std::array<int, 3> MapSizeBound() const { return map_size_; }
|
|
|
|
// Reserves the device memory. Throws JFJochException on a CUDA failure.
|
|
void Reserve();
|
|
|
|
// F_calc and F_mask of `model` - the model this was made for, its atoms anywhere - as model validation's
|
|
// compute_model_factors() makes them: prepare_asu_data(d_min, blur) of the density and
|
|
// prepare_asu_data(d_min) of the mask. Throws JFJochException on a CUDA failure. Safe to call from
|
|
// several threads; they take turns.
|
|
void Compute(const gemmi::Model &model, gemmi::AsuData<std::complex<float>> &fcalc,
|
|
gemmi::AsuData<std::complex<float>> &fmask);
|
|
|
|
// The real-space map of ASU coefficients on this engine's reflections, as model validation's
|
|
// map_from_coefficients() makes it: on the grid get_size_for_hkl(coef, {0, 0, 0}, 3.0) sizes. Sorts
|
|
// `coef`. Throws JFJochException on a CUDA failure. Safe to call from several threads; they take turns.
|
|
gemmi::Grid<float> Map(gemmi::AsuData<std::complex<float>> &coef);
|
|
|
|
private:
|
|
int device_;
|
|
size_t device_total_;
|
|
gemmi::UnitCell cell_;
|
|
const gemmi::SpaceGroup *sg_;
|
|
double d_min_;
|
|
gemmi::DensityCalculator<gemmi::IT92<float>, float> dc_;
|
|
ModelStructureFactorsGPUSetup setup_;
|
|
std::vector<gemmi::Miller> rows_;
|
|
std::array<int, 3> map_size_{};
|
|
bool supported_ = false;
|
|
size_t fft_work_bytes_ = 0, device_bytes_ = 0;
|
|
std::unique_ptr<ModelStructureFactorsGPUEngine> engine_;
|
|
std::mutex m_;
|
|
};
|