--model validation (battery-only for users) was 18% of the battery's time. Every number it produces is unchanged to the bit (p.mtz, maps, placed model and every model-validation line of the report md5/diff-identical on 11 open sets); only when and where the work runs changes: - The bulk-solvent grid fit (FitModelScale, most of the CPU time) fits each solvent pair on a copy of gemmi::Scaling's target that takes |Fcalc + k_sol exp(-b_sol s^2) Fmask| once per pair instead of at every solver evaluation; same expressions, same types (new test checks a grid point against gemmi's own Scaling fit with ==). - Fcalc density and the solvent mask are made on two threads; the model's structure factors beside the GPU engine reservation. - The indexing probe fits the relabellings concurrently. - The null's replicates run beside the real model's placement (they start from a snapshot of the model as read); one GPU engine per replicate plus one for the real fit instead of a cap of 4 (engines are interchangeable and deterministic). - The 2mFo-DFc, mFo-DFc and anomalous maps are made and written concurrently; the placed model is written beside the reflection files. - A rigid-body zone whose solvent-mask grid needs gemmi's shrink is sent to the CPU when the engines are reserved (ModelMaskGPU::ShrinkIsNoOp), instead of failing on the GPU and validating everything again on the CPU - the same CPU result, without the wasted first attempt. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
86 lines
3.7 KiB
C++
86 lines
3.7 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <cstddef>
|
|
#include <vector>
|
|
|
|
#include <cuda_runtime.h>
|
|
|
|
#include "../indexing/CUDAMemHelpers.h"
|
|
|
|
// The bulk-solvent mask of PutMaskOnGrid() (ModelGrid.h), i.e. gemmi's
|
|
// SolventMasker(AtomicRadiiSet::Refmac).put_mask_on_grid(), on the GPU. gemmi masks the atoms and then
|
|
// symmetrizes the grid with the minimum; the operators are isometries, so that is the same as masking
|
|
// every symmetry image of every atom, which is what is done here - no orbits, no symmetrize. The island
|
|
// removal is gemmi's (26-connected, periodic, the same size limit) as a union-find. The shrink is not
|
|
// implemented: it is a no-op on every rigid-body grid, and SetGrid() refuses a grid where it would not be.
|
|
//
|
|
// Every write of the masking is the same idempotent store and the union-find partition is unique, so
|
|
// the mask is bit-identical from run to run. It can differ from gemmi's at points lying exactly at an
|
|
// atom's radius, where float and double distances round differently.
|
|
|
|
struct ModelMaskGrid {
|
|
int nu = 0, nv = 0, nw = 0; // index = u + nu * (v + nv * w)
|
|
double orth[9]; // gemmi UnitCell::orth.mat, row-major (Cartesian = orth * fractional)
|
|
double volume = 0; // UnitCell::volume, A^3
|
|
};
|
|
|
|
// One atom to mask: fractional x, y, z wrapped into [0,1) and the mask radius in A. Double, because
|
|
// float fractional coordinates are off by up to 1e-5 A in a 200 A cell, enough to move points that lie at
|
|
// the radius; the distances themselves are float. (CUDA's double4 is deprecated from CUDA 13 and its
|
|
// replacement does not exist before, hence a struct of our own.)
|
|
struct ModelMaskAtom {
|
|
double x, y, z, radius;
|
|
};
|
|
|
|
// A fractional operator x' = rot * x + tran: every symmetry operator combined with every centring
|
|
// vector, identity included.
|
|
struct ModelMaskOp {
|
|
double rot[9]; // row-major
|
|
double tran[3];
|
|
};
|
|
|
|
class ModelMaskGPU {
|
|
public:
|
|
// Enough for Fm-3m, 48 operators times 4 centring vectors.
|
|
static constexpr int MAX_OPS = 192;
|
|
|
|
static size_t DeviceBytes(size_t max_points);
|
|
ModelMaskGPU(cudaStream_t stream, size_t max_points);
|
|
|
|
// Whether gemmi's shrink (r_shrink = 0.8 A) leaves this grid as it is - the only grids this mask is
|
|
// right on. Host code; asked before any engine is reserved, so a grid it fails goes to the CPU first.
|
|
static bool ShrinkIsNoOp(const ModelMaskGrid &grid);
|
|
|
|
// Per zone. Throws if the grid has more than max_points points, if there are more than MAX_OPS
|
|
// operators, or if ShrinkIsNoOp() fails on it.
|
|
void SetGrid(const ModelMaskGrid &grid, const std::vector<ModelMaskOp> &ops);
|
|
|
|
// d_atoms: hydrogens and unoccupied atoms already left out. d_mask: the whole grid, 1 = solvent, 0 = macromolecule. Queued on the stream.
|
|
void Compute(const ModelMaskAtom *d_atoms, int n_atoms, float *d_mask);
|
|
|
|
// The island removal alone, on a mask of 0 and 1 already on the grid. Compute() ends with it.
|
|
void RemoveIslands(float *d_mask);
|
|
|
|
// What the masking kernel needs of the grid.
|
|
struct Geometry {
|
|
int nu, nv, nw;
|
|
float orth_n[9]; // orth * diag(1/nu, 1/nv, 1/nw): Cartesian of a grid-step offset
|
|
double spacing[3]; // gemmi Grid::spacing, the interplanar distance of the grid planes
|
|
};
|
|
|
|
private:
|
|
cudaStream_t stream;
|
|
size_t max_points;
|
|
Geometry geom{};
|
|
size_t npoints = 0;
|
|
int n_ops = 0;
|
|
int island_limit = 0;
|
|
|
|
CudaDevicePtr<ModelMaskOp> ops_d;
|
|
CudaDevicePtr<int> label; // union-find parent, then the component sizes
|
|
CudaDevicePtr<int> root; // each solvent point's component, -1 elsewhere
|
|
};
|