Files
Jungfraujoch/image_analysis/structure_refinement/RigidBodyGPU.h
T
leonarski_fandClaude Opus 5.5 ed2d557bb7 ModelValidation: a CUDA failure fails the validation instead of restarting it on the CPU
A CUDA error on the validation's GPU path (rigid body, structure factors,
maps) used to start the whole validation again on the CPU - a migration
part way through, which makes how long a run takes, and in principle what
its rigid bodies converge to, depend on a device fault. It now ends the
validation with ok = false and a logged reason ("the GPU failed during the
validation (...)"), which the run reports as MODEL_NOT_VALIDATED. A
validation that did not finish decides nothing, so the reflection files
are those of a run without a model. Where there is no GPU at all, or the
up-front rule sends the work to the CPU, nothing changes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 12:47:04 +02:00

134 lines
6.0 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
// The rigid body's target evaluated on a GPU (CUDA builds only; the header itself needs no CUDA).
// It is the function RigidBodyTarget computes on the CPU - the same density, the same composition of
// Fcalc from one copy, the same bulk-solvent mask, scale fit, residuals and Jacobian - moved to the
// device so that a fit takes a fraction of a second instead of several seconds. The two agree to
// rounding, not bit for bit: float distances, cuFFT for FFTW. The GPU is deterministic on its own.
#include <condition_variable>
#include <map>
#include <memory>
#include <mutex>
#include <optional>
#include <stdexcept>
#include <string>
#include <vector>
#include "RigidBodyRefine.h"
class Logger;
class RigidBodyGPUEngine;
struct RigidBodyGPUZone;
// A CUDA failure inside the rigid body, or in model validation's structure factors and maps
// (ModelStructureFactorsGPU). Model validation catches it and reports the validation as failed.
class RigidBodyGPUFailure : public std::runtime_error {
public:
explicit RigidBodyGPUFailure(const std::string &what) : std::runtime_error(what) {}
};
// A few engines, reserved once for a whole validation: each is a stream and the buffers for one fit at
// a time, sized for the finest zone of the ladder to d_min. The real fit and the null's replicates each
// take one for the length of their fit, and wait for one when all are taken. Engines are
// interchangeable and every kernel deterministic, so which replicate gets which engine does not change
// a number. A pool belongs to one model: what a zone needs of its atoms - their densities, radii and mask
// radii, which the placement does not change - is worked out once per zone and kept.
class RigidBodyGPUPool {
public:
// Null where there is no GPU, where not even one engine fits the budget - a quarter of the card,
// and never the last gigabyte of what is free, since the merge may be running beside it - or where
// the cell is too small for the gather (an atom's box wider than the cell). Logged either way.
// `max_observations`: the most working reflections any fit will be given in its finest zone.
// `max_engines`: at most this many, however much memory there is.
static std::unique_ptr<RigidBodyGPUPool> Create(const gemmi::Model &model, const gemmi::UnitCell &cell,
const gemmi::SpaceGroup &sg, double d_min,
size_t max_observations, size_t max_engines, Logger &logger);
~RigidBodyGPUPool();
size_t Engines() const { return engines_.size(); }
RigidBodyGPUEngine &Acquire();
void Release(RigidBodyGPUEngine &engine);
// The zone to d_min without its observations, for this pool's model.
const RigidBodyGPUZone &Zone(const gemmi::Model &model, const gemmi::UnitCell &cell, const gemmi::SpaceGroup &sg,
double d_min);
// The zone with these observations and the composition of Fcalc at them. The null's replicates are all
// fitted to the same reflections, so they share it.
std::shared_ptr<const RigidBodyGPUZone> ObservedZone(const gemmi::Model &model, const gemmi::UnitCell &cell,
const gemmi::SpaceGroup &sg,
const gemmi::AsuData<gemmi::ValueSigma<float>> &fobs,
double d_min);
private:
RigidBodyGPUPool() = default;
std::vector<std::unique_ptr<RigidBodyGPUEngine>> engines_;
std::vector<RigidBodyGPUEngine *> idle_;
std::mutex m_;
std::condition_variable cv_;
std::map<double, std::unique_ptr<RigidBodyGPUZone>> zones_;
std::mutex zones_m_;
struct ObservedEntry {
double d_min;
gemmi::AsuData<gemmi::ValueSigma<float>> fobs;
std::shared_ptr<const RigidBodyGPUZone> zone;
};
std::vector<ObservedEntry> observed_;
std::mutex observed_m_;
};
// An engine taken from a pool for as long as this lives.
class RigidBodyGPULease {
public:
explicit RigidBodyGPULease(RigidBodyGPUPool &pool) : pool_(pool), engine_(pool.Acquire()) {}
~RigidBodyGPULease() { pool_.Release(engine_); }
RigidBodyGPULease(const RigidBodyGPULease &) = delete;
RigidBodyGPULease &operator=(const RigidBodyGPULease &) = delete;
RigidBodyGPUEngine &Engine() { return engine_; }
private:
RigidBodyGPUPool &pool_;
RigidBodyGPUEngine &engine_;
};
class RigidBodyTargetGPU : public RigidBodyTargetBase {
public:
// Takes an engine from `pool` for its lifetime. q = 0 is the placement `model` has now; unlike the
// CPU target the model is never moved by an evaluation, only by Place().
RigidBodyTargetGPU(RigidBodyGPUPool &pool, gemmi::Model &model, const gemmi::UnitCell &cell,
const gemmi::SpaceGroup &sg, size_t nthreads);
~RigidBodyTargetGPU() override;
void SetZone(const gemmi::AsuData<gemmi::ValueSigma<float>> &fobs, double d_min) override;
size_t NumObservations() const override;
bool Residuals(const double q[6], double *residuals) override;
bool Jacobian(const double q[6], double *jacobian) override;
private:
// The placement at q as x -> R (x - centroid) + centroid + t.
void Placement(const double q[6], double rotation[9], double translation[3]) const;
void HostScale();
RigidBodyGPUPool &pool_;
RigidBodyGPULease lease_;
RigidBodyGPUEngine &engine_;
gemmi::Model &model_;
const gemmi::UnitCell &cell_;
const gemmi::SpaceGroup &sg_;
size_t nthreads_;
double d_min_ = 0;
std::shared_ptr<const RigidBodyGPUZone> zone_;
bool solvent_fitted_ = false;
bool host_scale_ = false; // the zone's scale is fitted on the host (HostScale)
bool have_point_ = false;
std::array<double, 6> q_{};
double k_overall_ = 1;
gemmi::SMat33<double> b_star_{0, 0, 0, 0, 0, 0};
};