Files
Jungfraujoch/image_analysis/structure_refinement/ModelMaskGPU.h
T
leonarski_fandClaude Opus 5.5 6025d51a48 ModelMaskGPU: gemmi's shrink step, so fine rigid-body zones stay on the GPU
The GPU bulk-solvent mask had no shrink step (SolventMasker::shrink(),
set_margin_around() with Refmac's r_shrink = 0.8 A), so every rigid-body
zone whose grid has an offset within 0.8 A was sent to the CPU. That is
not only fine grids: in an oblique setting (P2_1, beta ~141 deg) the
3.5 A zone's lattice-plane spacing is 0.69 A, and on such a set the real
fit and all nine null replicates ran their rigid bodies on the CPU.

The shrink is now two passes over the grid: mark the solvent points with
a macromolecule point among gemmi's near offsets, then turn into solvent
every macromolecule point with such an edge point at any stencil offset
(near or far, gemmi's split at the coarsest grid step). That is gemmi's
rule in both of its branches, and the stencil is built on the host in
double exactly as gemmi builds it. ModelMaskGPU::ShrinkIsNoOp() and
RigidBodyGPUEngine::MaskSupports() are gone; the rigid body no longer
refuses such zones.

Verified (ModelMaskGPU_ShrinkMatchesGemmi): the shrink alone on gemmi's
post-island mask is bit-identical, and the whole GPU mask equals gemmi's
put_mask_on_grid() bit for bit (0 differing points) on the five test
groups at 1.5 A and an oblique P2_1 cell at 3.5 A; repeats are identical.

Measured on the oblique-setting set of the open arm (P2_1, 14k atoms,
1.66 A): model validation 17.2 s -> 5.2 s with the GPU structure factors
of the next commit in place (the first validation's ten rigid bodies,
~12 s on the CPU, now take well under a second). p.mtz md5-identical;
the null moves from +73.1 to +86.1 sigma (the replicates are refined on
the GPU, which agrees with the CPU to rounding), verdict unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 12:38:37 +02:00

92 lines
4.0 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <cstddef>
#include <vector>
#include <cuda_runtime.h>
#include "../indexing/CUDAMemHelpers.h"
// The bulk-solvent mask of PutMaskOnGrid() (ModelGrid.h), i.e. gemmi's
// SolventMasker(AtomicRadiiSet::Refmac).put_mask_on_grid(), on the GPU. gemmi masks the atoms and then
// symmetrizes the grid with the minimum; the operators are isometries, so that is the same as masking
// every symmetry image of every atom, which is what is done here - no orbits, no symmetrize. The island
// removal is gemmi's (26-connected, periodic, the same size limit) as a union-find, and the shrink is
// gemmi's set_margin_around() as two passes over the grid.
//
// Every write of the masking is the same idempotent store and the union-find partition is unique, so
// the mask is bit-identical from run to run. It can differ from gemmi's at points lying exactly at an
// atom's radius, where float and double distances round differently.
struct ModelMaskGrid {
int nu = 0, nv = 0, nw = 0; // index = u + nu * (v + nv * w)
double orth[9]; // gemmi UnitCell::orth.mat, row-major (Cartesian = orth * fractional)
double volume = 0; // UnitCell::volume, A^3
};
// One atom to mask: fractional x, y, z wrapped into [0,1) and the mask radius in A. Double, because
// float fractional coordinates are off by up to 1e-5 A in a 200 A cell, enough to move points that lie at
// the radius; the distances themselves are float. (CUDA's double4 is deprecated from CUDA 13 and its
// replacement does not exist before, hence a struct of our own.)
struct ModelMaskAtom {
double x, y, z, radius;
};
// A fractional operator x' = rot * x + tran: every symmetry operator combined with every centring
// vector, identity included.
struct ModelMaskOp {
double rot[9]; // row-major
double tran[3];
};
class ModelMaskGPU {
public:
// Enough for Fm-3m, 48 operators times 4 centring vectors.
static constexpr int MAX_OPS = 192;
// Grid offsets within gemmi's r_shrink = 0.8 A: 26 at a 0.4 A spacing, 728 at 0.2 A.
static constexpr int MAX_STENCIL = 1024;
static size_t DeviceBytes(size_t max_points);
ModelMaskGPU(cudaStream_t stream, size_t max_points);
// Per zone. Throws if the grid has more than max_points points, if there are more than MAX_OPS
// operators, or if the shrink's stencil is over MAX_STENCIL offsets.
void SetGrid(const ModelMaskGrid &grid, const std::vector<ModelMaskOp> &ops);
// d_atoms: hydrogens and unoccupied atoms already left out. d_mask: the whole grid, 1 = solvent, 0 = macromolecule. Queued on the stream.
// The masking, then RemoveIslands(), then Shrink().
void Compute(const ModelMaskAtom *d_atoms, int n_atoms, float *d_mask);
// The island removal alone, on a mask of 0 and 1 already on the grid. Compute() runs it before Shrink().
void RemoveIslands(float *d_mask);
// The shrink alone, on a mask whose islands are already removed. A no-op on a grid too coarse for
// any grid offset to lie within r_shrink.
void Shrink(float *d_mask);
// What the masking kernel needs of the grid.
struct Geometry {
int nu, nv, nw;
float orth_n[9]; // orth * diag(1/nu, 1/nv, 1/nw): Cartesian of a grid-step offset
double spacing[3]; // gemmi Grid::spacing, the interplanar distance of the grid planes
};
private:
void SetShrinkStencil(const ModelMaskGrid &grid);
cudaStream_t stream;
size_t max_points;
Geometry geom{};
size_t npoints = 0;
int n_ops = 0;
int island_limit = 0;
int n_near = 0, n_stencil = 0; // the shrink's offsets: the first n_near are the near ones
CudaDevicePtr<ModelMaskOp> ops_d;
CudaDevicePtr<int> label; // union-find parent, then the component sizes
CudaDevicePtr<int> root; // each solvent point's component, -1 elsewhere
CudaDevicePtr<int3> stencil_d;
};