Build Packages / Create release (push) Successful in 24s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m29s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m43s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 8m27s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 9m53s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 9m58s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 11m22s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 13m39s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 18m37s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 16m32s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 24m11s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m30s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 19m3s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 20m23s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 19m41s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 21m0s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m38s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 14m33s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 17m55s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 20m50s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 18m38s
Build Packages / Unit tests (push) Successful in 1h46m14s
* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports. * jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls. * Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results. * Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable. * Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate. * Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do. * Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence. * Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags. * Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check. * Rugnux: Clear error messages when a data set needs more GPU or host memory than is available. Reviewed-on: #83 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
141 lines
7.0 KiB
C++
141 lines
7.0 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
// The device half of RigidBodyTargetGPU (RigidBodyGPU.h): one engine is one CUDA stream and the buffers
|
|
// for one resolution zone of one fit. The host half works out everything that depends only on the
|
|
// model, the cell and the zone - with gemmi, which stays out of nvcc - and hands it over in the plain
|
|
// structures below; the engine then does, per evaluation, what depends on the placement. CUDA builds
|
|
// only.
|
|
|
|
#include <array>
|
|
#include <cstddef>
|
|
#include <memory>
|
|
#include <stdexcept>
|
|
#include <vector>
|
|
|
|
// One atom's density on one zone's grid, as PutModelDensityOnGrid() (ModelGrid.cpp) sets it up: gemmi's
|
|
// precalculated five-Gaussian sum and the radius it cuts the sum at.
|
|
struct RigidBodyGPUAtom {
|
|
float a[5];
|
|
float b[5][6]; // isotropic: b[k][0] multiplies r^2; anisotropic: the matrix, u11 u22 u33 u12 u13 u23
|
|
float occ;
|
|
float radius;
|
|
int aniso;
|
|
};
|
|
|
|
// One term of SymmetryComposition: F1 at k = hR, times the phase of the operator's translation.
|
|
struct RigidBodyGPUTerm {
|
|
int index; // of F1(k) in the half-u transform, or of F1(-k) where k is outside the stored half
|
|
int conj; // read as the Friedel mate, conj F1(-k)
|
|
double phase[2]; // exp(+2 pi i h.t), real and imaginary
|
|
double s[3]; // k as a Cartesian reciprocal vector
|
|
};
|
|
|
|
// Everything a zone needs that does not depend on the placement.
|
|
struct RigidBodyGPUZone {
|
|
int nu = 0, nv = 0, nw = 0; // the zone's grid, u fastest
|
|
double orth[9] = {}, frac[9] = {}; // row-major
|
|
double volume = 0;
|
|
double blur = 0; // DensityCalculator's, which the rows' unblur undoes
|
|
std::vector<RigidBodyGPUAtom> atoms; // model order
|
|
|
|
// The bulk-solvent mask's atoms: the model's index of each, and its radius (probe included).
|
|
std::vector<int> mask_atom;
|
|
std::vector<float> mask_radius;
|
|
// Every image: the group's operators, each with each centring vector, fractional.
|
|
std::vector<std::array<double, 12>> images; // rot[9] row-major, tran[3]
|
|
|
|
// The composition: rows (composed indices) and Ops() terms per row, row-major.
|
|
std::vector<std::array<int, 3>> row_hkl;
|
|
std::vector<double> row_scale; // n_cen * unblur
|
|
std::vector<double> row_stol2;
|
|
size_t ops = 1;
|
|
std::vector<RigidBodyGPUTerm> terms;
|
|
|
|
// The observations the residuals are over: each one's row (-1 without one) and amplitude.
|
|
std::vector<int> obs_row;
|
|
std::vector<float> obs_fobs, obs_sigma;
|
|
double f_mean = 1;
|
|
// The scale's points: the observations gemmi's prepare_points() would take, in order.
|
|
std::vector<int> point_obs;
|
|
std::vector<std::array<int, 3>> point_hkl;
|
|
std::vector<double> point_stol2;
|
|
std::vector<float> point_fobs, point_sigma;
|
|
std::vector<std::array<double, 6>> constraints; // adp_symmetry_constraints()
|
|
};
|
|
|
|
// Upper bounds an engine is sized for.
|
|
struct RigidBodyGPUCapacity {
|
|
size_t atoms = 0;
|
|
size_t grid_points = 0; // nu * nv * nw
|
|
size_t complex_points = 0; // (nu / 2 + 1) * nv * nw
|
|
size_t bricks = 0;
|
|
size_t pairs = 0; // (brick, atom) pairs of the gather
|
|
size_t rows = 0, terms = 0;
|
|
size_t observations = 0;
|
|
size_t fft_work_bytes = 0;
|
|
};
|
|
|
|
// FitModelScale()'s solvent grid on five or fewer strong reflections, where its fits chain from one grid
|
|
// point to the next and the device does not reproduce it (ModelScaleGPU::FitSolvent).
|
|
class ModelScaleGPUTooFewReflections : public std::runtime_error {
|
|
public:
|
|
ModelScaleGPUTooFewReflections() : std::runtime_error("ModelScaleGPU::FitSolvent: too few reflections for independent grid points") {}
|
|
};
|
|
|
|
struct RigidBodyGPUEngineImpl;
|
|
|
|
class RigidBodyGPUEngine {
|
|
public:
|
|
// The bytes an engine of this capacity reserves on the device.
|
|
static size_t DeviceBytes(const RigidBodyGPUCapacity &capacity);
|
|
// The largest cuFFT work area a (nu, nv, nw) grid needs, batch 1 or 3.
|
|
static size_t FFTWorkBytes(int nu, int nv, int nw);
|
|
// (brick, atom) pairs of the gather over at most, for a zone's atoms on its grid.
|
|
static size_t PairBound(const RigidBodyGPUZone &zone);
|
|
static size_t Bricks(int nu, int nv, int nw);
|
|
// The most gather bricks the 2 d + 1 points of an atom's box can fall in along an axis of n points.
|
|
static size_t AxisBrickBound(int d, int n);
|
|
// Whether the gather reproduces gemmi's box walk on this zone: every atom's box narrower than the cell,
|
|
// so that no point is reached by two images of one atom.
|
|
static bool Supports(const RigidBodyGPUZone &zone);
|
|
static void MemoryInfo(size_t &free, size_t &total);
|
|
static int CurrentDevice();
|
|
|
|
RigidBodyGPUEngine(const RigidBodyGPUCapacity &capacity, int device);
|
|
~RigidBodyGPUEngine();
|
|
|
|
// Per fit: each atom's position relative to the model centroid, which the placements rotate.
|
|
void SetBody(const std::vector<std::array<double, 3>> &relative);
|
|
// Per zone. Throws if the zone does not fit the capacity.
|
|
void SetZone(const RigidBodyGPUZone &zone);
|
|
|
|
// Per evaluation, at the placement x -> R x_rel + t: Fcalc and dF/dt at the rows, then the bulk-
|
|
// solvent mask of the same placement (ModelMaskGPU) and Fmask.
|
|
void Fcalc(const double rotation[9], const double translation[3]);
|
|
void Fmask();
|
|
// The scale at this evaluation's Fcalc and Fmask (ModelScaleGPU): FitModelScale()'s solvent grid, and
|
|
// the overall scale and anisotropic B at a fixed solvent. FitSolvent() throws ModelScaleGPUTooFewReflections where
|
|
// the grid's fits would chain (too few reflections for the isotropic start), which the host then does.
|
|
void FitSolvent(double &k_sol, double &b_sol);
|
|
void FitScale(double k_sol, double b_sol, double &k_overall, double b_star[6]);
|
|
// The scale's points: fcmol (Fcalc as complex<float>) and fmask, downloaded for the host's fit.
|
|
void DownloadPoints(std::vector<std::array<float, 2>> &fcmol, std::vector<std::array<float, 2>> &fmask);
|
|
// The residuals at the scale given, NumObservations() of them, into `residuals` (host).
|
|
void Residuals(double k_overall, const double b_star[6], double k_sol, double b_sol, double *residuals);
|
|
|
|
// The Jacobian at the last evaluation: `rotation[j]` and `translation[j]` place the body one step
|
|
// along rotation axis j. Writes J (fixed scale) and J_k (the scale parameters' columns) on the
|
|
// device and returns J_k^T J_k (p x p) and J_k^T J (p x 6), p = 1 + constraints, row-major.
|
|
void Jacobian(const double rotation[3][9], const double translation[3][3], double step, double k_overall,
|
|
const double b_star[6], double k_sol, double b_sol, std::vector<double> &jtj,
|
|
std::vector<double> &jtq);
|
|
// J - J_k x, x p x 6 row-major, into `jacobian` (host, NumObservations() x 6).
|
|
void ProjectJacobian(const std::vector<double> &x, double *jacobian);
|
|
|
|
private:
|
|
std::unique_ptr<RigidBodyGPUEngineImpl> impl_;
|
|
};
|