Files
Jungfraujoch/common/CUDAWrapper.h
T
leonarski_fandClaude Opus 5.5 ea9d0adcde Stop on a lost CUDA context instead of falling back to the host decoder
A sticky error (an illegal address, say) is reported by whichever worker synchronises next - usually
the bslz4 device decode - and MXAnalysisWithoutFPGA::Analyze and Rugnux::MaskDefectivePixels then
logged "falling back to host" and carried on, although the context is gone and the run fails anyway,
later and less clearly. cuda_throw_if_context_lost() now throws a GPUCUDAError ("CUDA device
unusable after an unrecoverable error: ...") there first, which the image loops already treat as
fatal (IsFatalResourceError), like a CUDA out-of-memory.

The last error cannot tell: once cudaGetLastError() has returned a sticky error, it and
cudaPeekAtLastError() both report success. cudaFree(nullptr) frees nothing and does not
synchronise, but returns the sticky error (checked: after an illegal-address kernel it returns
cudaErrorIllegalAddress; with only a pending out-of-memory it returns success), so a handled,
non-sticky decode failure still falls back - MXAnalysis_HandledDeviceDecodeFailureLeavesNoError
passes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-27 17:11:38 +02:00

53 lines
2.7 KiB
C++

// SPDX-FileCopyrightText: 2024 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <cstdint>
#include <string>
#include <vector>
int32_t get_gpu_count();
// Names of the visible GPUs, in device order and one entry per device, so repeated cards repeat.
// Empty without CUDA and on a machine with no device, which is also what get_gpu_count() == 0 says.
std::vector<std::string> get_gpu_names();
// The same list collapsed for a person: "4x NVIDIA A100-SXM4-80GB", or several such groups separated
// by ", " on a mixed machine. Empty when no GPU is visible.
std::string get_gpu_description();
void set_gpu(int32_t dev_id);
// Pin the calling thread to the next GPU in round-robin order, using a process-wide counter
// (counter++ % get_gpu_count()). Call once per thread; no thread id needed. No-op when no GPU
// is visible. Honours CUDA_VISIBLE_DEVICES via get_gpu_count().
void pin_gpu();
// From here on, pin_gpu() also keeps the calling thread on the CPUs of the NUMA node its GPU is
// attached to (ThreadAffinity.h) - on a machine with one node, or without CUDA, nothing changes.
// Off unless a program asks for it.
void enable_gpu_numa_binding();
// pin_gpu() onto a given device: set_gpu(dev_id), and the NUMA binding above where it is enabled.
// For a worker that takes its card by index rather than round-robin.
void pin_gpu(int32_t dev_id);
// Have every GPU's host threads BLOCK in a synchronisation (cudaDeviceScheduleBlockingSync) instead
// of spinning on a core until the device finishes. Must be called before anything creates a CUDA
// context; a device that already has one keeps the flags it was created with. No-op without CUDA.
void set_gpu_blocking_sync();
// Drop the error CUDA has recorded for the calling thread. Call it where a CUDA failure has been
// HANDLED - a device route that fell back to the host, an indexing attempt whose failure was turned
// into a result - because the error otherwise stays as the thread's last error and the next
// cuda_err(cudaGetLastError()) after some later kernel launch reports it, over work that went fine.
// A sticky error (an illegal access, say) is not cleared by this, and nothing here pretends it is:
// the context is gone in that case and every later call fails on its own. No-op without CUDA.
void cuda_clear_error();
// Throw if the CUDA context is lost - a sticky error (an illegal access, say) that every later call
// on this device will fail with. For a device route about to fall back to the host: there is no
// fallback from that, and falling back only moves the failure somewhere less clear. A handled,
// non-sticky failure passes. No-op without CUDA.
void cuda_throw_if_context_lost();