The speculative geometry probe (StartSpeculativeGeometryProbe) runs an indexing-only pass on a copy of the run beside the pass's scaling merge. It holds ~4.5 GB of engines (25 image-sized spot-finding buffers, the per-engine tables, FFT indexers) plus what its streams' pool retains, for a few seconds. When a merge allocation landed in that window and did not fit, RotationScaleMergeGPU::Impl::Alloc reported "needs more GPU memory than this card has ... too large for GPU scaling", although the per-observation arrays were 1.3 GB on a 16.6 GB card and the set runs fine alone. Timing-dependent: one full-battery failure, not reproduced in ~15 plain reruns; reproduced deterministically by letting the probe hold extra device memory. - GPUWorkBeside (common/CUDAWrapper): a process-wide count of GPU work running beside the main line, with a condition variable signalled when the last one ends. The speculative probe holds one for its whole pass. - Alloc: on failure, wait for that work to end (bounded, 10 min, then a "GPU busy" error), then ask for the same buffer once more. Nothing else in flight means no wait, so a genuine shortage still fails at once. The computation is the same whenever the allocation succeeds, so results do not depend on the wait. - "Too large for this card" is now said only by the early check of the per-observation arrays against total device memory; a later shortage with nothing running beside gets its own message (buffer size, free/total, and that another program or the set's size is the cause). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
79 lines
1.9 KiB
C++
79 lines
1.9 KiB
C++
// SPDX-FileCopyrightText: 2024 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#include "CUDAWrapper.h"
|
|
|
|
#include <condition_variable>
|
|
#include <mutex>
|
|
|
|
// Build-independent: the CUDA build gets get_gpu_names() from CUDAWrapper.cu, the CPU-only build from
|
|
// the stub below, and this collapses whichever list came back. Four identical cards read better as
|
|
// "4x <name>" than as the same name four times, and a mixed machine keeps one group per model.
|
|
std::string get_gpu_description() {
|
|
const auto names = get_gpu_names();
|
|
|
|
std::string out;
|
|
for (size_t i = 0; i < names.size();) {
|
|
size_t n = 1;
|
|
while (i + n < names.size() && names[i + n] == names[i])
|
|
n++;
|
|
if (!out.empty())
|
|
out += ", ";
|
|
if (n > 1)
|
|
out += std::to_string(n) + "x ";
|
|
out += names[i];
|
|
i += n;
|
|
}
|
|
return out;
|
|
}
|
|
|
|
namespace {
|
|
std::mutex gpu_work_beside_mutex;
|
|
std::condition_variable gpu_work_beside_done;
|
|
int gpu_work_beside = 0;
|
|
}
|
|
|
|
GPUWorkBeside::GPUWorkBeside() {
|
|
std::lock_guard lock(gpu_work_beside_mutex);
|
|
++gpu_work_beside;
|
|
}
|
|
|
|
GPUWorkBeside::~GPUWorkBeside() {
|
|
{
|
|
std::lock_guard lock(gpu_work_beside_mutex);
|
|
--gpu_work_beside;
|
|
}
|
|
gpu_work_beside_done.notify_all();
|
|
}
|
|
|
|
bool wait_for_gpu_work_beside(std::chrono::seconds timeout) {
|
|
std::unique_lock lock(gpu_work_beside_mutex);
|
|
return gpu_work_beside_done.wait_for(lock, timeout, [] { return gpu_work_beside == 0; });
|
|
}
|
|
|
|
#ifndef JFJOCH_USE_CUDA
|
|
|
|
int32_t get_gpu_count() {
|
|
return 0;
|
|
}
|
|
|
|
std::vector<std::string> get_gpu_names() {
|
|
return {};
|
|
}
|
|
|
|
void set_gpu(int32_t dev_id) {}
|
|
|
|
void pin_gpu() {}
|
|
|
|
void pin_gpu(int32_t dev_id) {}
|
|
|
|
void enable_gpu_numa_binding() {}
|
|
|
|
void set_gpu_blocking_sync() {}
|
|
|
|
void cuda_clear_error() {}
|
|
|
|
void cuda_throw_if_context_lost() {}
|
|
|
|
#endif
|