Files
leonarski_f 84228bf8be
Build Packages / Create release (push) Successful in 24s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m29s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m43s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 8m27s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 9m53s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 9m58s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 11m22s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 13m39s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 18m37s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 16m32s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 24m11s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m30s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 19m3s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 20m23s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 19m41s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 21m0s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m38s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 14m33s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 17m55s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 20m50s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 18m38s
Build Packages / Unit tests (push) Successful in 1h46m14s
v1.0.0-rc.173 (#83)
* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports.
* jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls.
* Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results.
* Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable.
* Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate.
* Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do.
* Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence.
* Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags.
* Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check.
* Rugnux: Clear error messages when a data set needs more GPU or host memory than is available.

Reviewed-on: #83
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-09-29 15:57:32 +02:00

68 lines
3.3 KiB
C++

// SPDX-FileCopyrightText: 2024 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <chrono>
#include <cstdint>
#include <string>
#include <vector>
int32_t get_gpu_count();
// Names of the visible GPUs, in device order and one entry per device, so repeated cards repeat.
// Empty without CUDA and on a machine with no device, which is also what get_gpu_count() == 0 says.
std::vector<std::string> get_gpu_names();
// The same list collapsed for a person: "4x NVIDIA A100-SXM4-80GB", or several such groups separated
// by ", " on a mixed machine. Empty when no GPU is visible.
std::string get_gpu_description();
void set_gpu(int32_t dev_id);
// Pin the calling thread to the next GPU in round-robin order, using a process-wide counter
// (counter++ % get_gpu_count()). Call once per thread; no thread id needed. No-op when no GPU
// is visible. Honours CUDA_VISIBLE_DEVICES via get_gpu_count().
void pin_gpu();
// From here on, pin_gpu() also keeps the calling thread on the CPUs of the NUMA node its GPU is
// attached to (ThreadAffinity.h) - on a machine with one node, or without CUDA, nothing changes.
// Off unless a program asks for it.
void enable_gpu_numa_binding();
// pin_gpu() onto a given device: set_gpu(dev_id), and the NUMA binding above where it is enabled.
// For a worker that takes its card by index rather than round-robin.
void pin_gpu(int32_t dev_id);
// Have every GPU's host threads BLOCK in a synchronisation (cudaDeviceScheduleBlockingSync) instead
// of spinning on a core until the device finishes. Must be called before anything creates a CUDA
// context; a device that already has one keeps the flags it was created with. No-op without CUDA.
void set_gpu_blocking_sync();
// Drop the error CUDA has recorded for the calling thread. Call it where a CUDA failure has been
// HANDLED - a device route that fell back to the host, an indexing attempt whose failure was turned
// into a result - because the error otherwise stays as the thread's last error and the next
// cuda_err(cudaGetLastError()) after some later kernel launch reports it, over work that went fine.
// A sticky error (an illegal access, say) is not cleared by this, and nothing here pretends it is:
// the context is gone in that case and every later call fails on its own. No-op without CUDA.
void cuda_clear_error();
// Throw if the CUDA context is lost - a sticky error (an illegal access, say) that every later call
// on this device will fail with. For a device route about to fall back to the host: there is no
// fallback from that, and falling back only moves the failure somewhere less clear. A handled,
// non-sticky failure passes. No-op without CUDA.
void cuda_throw_if_context_lost();
// GPU work that runs beside the main line of a process and gives all of its device memory back when it
// ends: a probe pass run on a copy of the run while the run itself scales and merges. Hold one for as
// long as that work runs. Build-independent - it only counts.
class GPUWorkBeside {
public:
GPUWorkBeside();
~GPUWorkBeside();
GPUWorkBeside(const GPUWorkBeside &) = delete;
GPUWorkBeside &operator=(const GPUWorkBeside &) = delete;
};
// Wait until no GPUWorkBeside is held, for at most `timeout`. True once none is - at once if none was.
bool wait_for_gpu_work_beside(std::chrono::seconds timeout);