Files
Jungfraujoch/rugnux/ModelMaskGPU.h
T
leonarski_f 84228bf8be
Build Packages / Create release (push) Successful in 24s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m29s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m43s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 8m27s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 9m53s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 9m58s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 11m22s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 13m39s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 18m37s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 16m32s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 24m11s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m30s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 19m3s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 20m23s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 19m41s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 21m0s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m38s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 14m33s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 17m55s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 20m50s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 18m38s
Build Packages / Unit tests (push) Successful in 1h46m14s
v1.0.0-rc.173 (#83)
* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports.
* jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls.
* Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results.
* Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable.
* Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate.
* Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do.
* Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence.
* Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags.
* Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check.
* Rugnux: Clear error messages when a data set needs more GPU or host memory than is available.

Reviewed-on: #83
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-09-29 15:57:32 +02:00

82 lines
3.5 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <cstddef>
#include <vector>
#include <cuda_runtime.h>
#include "../image_analysis/indexing/CUDAMemHelpers.h"
// The bulk-solvent mask of PutMaskOnGrid() (ModelGrid.h), i.e. gemmi's
// SolventMasker(AtomicRadiiSet::Refmac).put_mask_on_grid(), on the GPU. gemmi masks the atoms and then
// symmetrizes the grid with the minimum; the operators are isometries, so that is the same as masking
// every symmetry image of every atom, which is what is done here - no orbits, no symmetrize. The island
// removal is gemmi's (26-connected, periodic, the same size limit) as a union-find. The shrink is not
// implemented: it is a no-op on every rigid-body grid, and SetGrid() refuses a grid where it would not be.
//
// Every write of the masking is the same idempotent store and the union-find partition is unique, so
// the mask is bit-identical from run to run. It can differ from gemmi's at points lying exactly at an
// atom's radius, where float and double distances round differently.
struct ModelMaskGrid {
int nu = 0, nv = 0, nw = 0; // index = u + nu * (v + nv * w)
double orth[9]; // gemmi UnitCell::orth.mat, row-major (Cartesian = orth * fractional)
double volume = 0; // UnitCell::volume, A^3
};
// One atom to mask: fractional x, y, z wrapped into [0,1) and the mask radius in A. Double, because
// float fractional coordinates are off by up to 1e-5 A in a 200 A cell, enough to move points that lie at
// the radius; the distances themselves are float. (CUDA's double4 is deprecated from CUDA 13 and its
// replacement does not exist before, hence a struct of our own.)
struct ModelMaskAtom {
double x, y, z, radius;
};
// A fractional operator x' = rot * x + tran: every symmetry operator combined with every centring
// vector, identity included.
struct ModelMaskOp {
double rot[9]; // row-major
double tran[3];
};
class ModelMaskGPU {
public:
// Enough for Fm-3m, 48 operators times 4 centring vectors.
static constexpr int MAX_OPS = 192;
static size_t DeviceBytes(size_t max_points);
ModelMaskGPU(cudaStream_t stream, size_t max_points);
// Per zone. Throws if the grid has more than max_points points, if there are more than MAX_OPS
// operators, or if gemmi's shrink (r_shrink = 0.8 A) would change anything on this grid.
void SetGrid(const ModelMaskGrid &grid, const std::vector<ModelMaskOp> &ops);
// d_atoms: hydrogens and unoccupied atoms already left out. d_mask: the whole grid, 1 = solvent, 0 = macromolecule. Queued on the stream.
void Compute(const ModelMaskAtom *d_atoms, int n_atoms, float *d_mask);
// The island removal alone, on a mask of 0 and 1 already on the grid. Compute() ends with it.
void RemoveIslands(float *d_mask);
// What the masking kernel needs of the grid.
struct Geometry {
int nu, nv, nw;
float orth_n[9]; // orth * diag(1/nu, 1/nv, 1/nw): Cartesian of a grid-step offset
double spacing[3]; // gemmi Grid::spacing, the interplanar distance of the grid planes
};
private:
cudaStream_t stream;
size_t max_points;
Geometry geom{};
size_t npoints = 0;
int n_ops = 0;
int island_limit = 0;
CudaDevicePtr<ModelMaskOp> ops_d;
CudaDevicePtr<int> label; // union-find parent, then the component sizes
CudaDevicePtr<int> root; // each solvent point's component, -1 elsewhere
};