ModelValidation: a CUDA failure fails the validation instead of restarting it on the CPU

A CUDA error on the validation's GPU path (rigid body, structure factors,
maps) used to start the whole validation again on the CPU - a migration
part way through, which makes how long a run takes, and in principle what
its rigid bodies converge to, depend on a device fault. It now ends the
validation with ok = false and a logged reason ("the GPU failed during the
validation (...)"), which the run reports as MODEL_NOT_VALIDATED. A
validation that did not finish decides nothing, so the reflection files
are those of a run without a model. Where there is no GPU at all, or the
up-front rule sends the work to the CPU, nothing changes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
This commit is contained in:
2026-10-09 12:47:04 +02:00
co-authored by Claude Opus 5.5
parent d5fcf2f05b
commit ed2d557bb7
2 changed files with 11 additions and 7 deletions
@@ -1570,20 +1570,24 @@ ModelValidationResult ValidateAgainstModel(const std::vector<MergedReflection> &
double wavelength_A,
const std::vector<float> &report_shell_d_min) {
#ifdef JFJOCH_USE_CUDA
// A CUDA failure on the GPU path does not end the run: the validation is started again from the
// model as read, on the CPU throughout. It is re-runnable, and a dead model check should not take a
// finished merge with it.
// A CUDA failure ends the validation, not the run, and nothing is moved to the CPU part way through:
// the validation reports why it did not finish, and a validation that did not finish decides nothing,
// so the reflection files are those of a run without a model.
try {
return Validate(merged, cell, model_path, output_prefix, logger, data_space_group,
probe_indexing_ambiguity, nthreads, wavelength_A, report_shell_d_min, true);
} catch (const RigidBodyGPUFailure &e) {
cuda_clear_error();
logger.Warning("Model validation: the rigid body failed on the GPU ({}); validating again on the CPU",
e.what());
ModelValidationResult failed;
failed.model_path = model_path;
failed.failure_reason = fmt::format("the GPU failed during the validation ({})", e.what());
logger.Error("Model validation: {}", failed.failure_reason);
return failed;
}
#endif
#else
return Validate(merged, cell, model_path, output_prefix, logger, data_space_group, probe_indexing_ambiguity,
nthreads, wavelength_A, report_shell_d_min, false);
#endif
}
namespace {
@@ -25,7 +25,7 @@ class RigidBodyGPUEngine;
struct RigidBodyGPUZone;
// A CUDA failure inside the rigid body, or in model validation's structure factors and maps
// (ModelStructureFactorsGPU). Model validation catches it and starts again on the CPU.
// (ModelStructureFactorsGPU). Model validation catches it and reports the validation as failed.
class RigidBodyGPUFailure : public std::runtime_error {
public:
explicit RigidBodyGPUFailure(const std::string &what) : std::runtime_error(what) {}