Build Packages / build:viewer-tgz:cpu (push) Successful in 8m32s
Build Packages / build:viewer-tgz:cuda (push) Successful in 10m13s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 12m56s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 13m51s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 14m2s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m25s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 15m1s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 13m2s
Build Packages / build:rpm (rocky8) (push) Successful in 12m55s
Build Packages / XDS test (durin plugin) (push) Successful in 9m41s
Build Packages / Generate python client (push) Successful in 28s
Build Packages / Build documentation (push) Successful in 47s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m13s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m35s
Build Packages / build:rpm (rocky9) (push) Successful in 13m38s
Build Packages / DIALS test (push) Successful in 13m57s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m21s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m49s
Build Packages / Unit tests (push) Successful in 1h1m3s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
One analysis engine is built per worker thread, and each uploaded its own copy of tables that are pure functions of the detector geometry: the pixel -> azimuthal bin map and the per-pixel corrections (both in AzIntEngineGPU AND again in AdaptiveSpotFinderGPU, from the same mapping), plus the pixel mask. On an 18 Mpx detector that is ~224 MB per worker; with 32 workers ~7 GB of device memory held 32 identical copies. Upload each table once per GPU instead and hand every engine on that device a shared pointer to it. The cache is keyed by (device, source-vector address) because workers are pinned round-robin across GPUs, so on a multi-GPU node each device keeps its own copy - a kernel may only read memory resident on the device it runs on - and the table is freed on the device that allocated it. Entries are held weakly, so a table goes away with the last engine using it. Measured on an 18 Mpx detector, 32 worker threads, 16 GB card: the stills path went from exhausting the card (OOM in de-novo indexing) to 8.6 GB peak, and a normal rotation run from 14.6 GB to 7.4 GB - it had been running within 1.6 GB of the limit, so any larger detector or second GPU consumer would have tipped it over. Per-worker footprint drops 403 -> 173 MB. Merge statistics are unchanged on a six-crystal regression subset, including two-pass runs where the second pass rebuilds the mapping on refined geometry, and wall time is unchanged (13.5-13.8 s vs 13.8-14.1 s). Also take the launch configuration from the current device rather than device 0 in AzIntEngineGPU and ImagePreprocessorGPU: with round-robin pinning, device 0's SM count and shared-memory size can belong to a different card than the one the kernels use. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
78 lines
3.3 KiB
C++
78 lines
3.3 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <map>
|
|
#include <memory>
|
|
#include <mutex>
|
|
#include <utility>
|
|
|
|
#include "CUDAMemHelpers.h"
|
|
|
|
// Read-only lookup tables that depend only on the detector geometry (pixel -> azimuthal bin, the
|
|
// per-pixel correction factors, the pixel mask). One analysis engine is built per worker thread, so
|
|
// each of those used to upload its own copy: on a 18 Mpx detector that is ~220 MB per thread, and
|
|
// 32 threads spent ~7 GB of device memory on identical data.
|
|
//
|
|
// Upload once per GPU instead and hand every engine on that GPU a shared pointer to the same table.
|
|
// The cache is keyed by (device, key) because a worker thread is pinned round-robin to a device
|
|
// (pin_gpu()), so on a multi-GPU node each device keeps its own copy - a kernel may only read memory
|
|
// resident on the device it runs on. `key` identifies the table's source data; use the address of the
|
|
// host vector that produced it, which lives in the experiment / integration mapping and therefore
|
|
// outlives every engine.
|
|
//
|
|
// Entries are held weakly, so the tables are released once the last engine using them is gone.
|
|
|
|
namespace jfjoch_cuda_shared_tables {
|
|
struct Registry {
|
|
std::mutex m;
|
|
std::map<std::pair<int, const void *>, std::weak_ptr<void>> tables;
|
|
};
|
|
|
|
inline Registry ®istry() {
|
|
static Registry r;
|
|
return r;
|
|
}
|
|
|
|
// Not called cuda_err: the .cu files that include this header define their own such helper in an
|
|
// anonymous namespace, and a second one at global scope would make every call ambiguous.
|
|
inline void check(cudaError_t val) {
|
|
if (val != cudaSuccess)
|
|
throw JFJochException(JFJochExceptionCategory::GPUCUDAError, cudaGetErrorString(val));
|
|
}
|
|
}
|
|
|
|
// Return the device-resident copy of `host` (`count` elements) for the calling thread's GPU,
|
|
// uploading it on `stream` the first time it is asked for.
|
|
template <typename T>
|
|
std::shared_ptr<CudaDevicePtr<T>> SharedDeviceTable(const void *key, size_t count, const T *host,
|
|
cudaStream_t stream) {
|
|
int device = 0;
|
|
jfjoch_cuda_shared_tables::check(cudaGetDevice(&device));
|
|
|
|
auto ® = jfjoch_cuda_shared_tables::registry();
|
|
// The upload happens while the lock is held: another worker must not obtain the pointer before
|
|
// its content is on the device.
|
|
std::lock_guard lock(reg.m);
|
|
auto &slot = reg.tables[{device, key}];
|
|
if (auto cached = slot.lock())
|
|
return std::static_pointer_cast<CudaDevicePtr<T>>(cached);
|
|
|
|
// Free on the device that allocated it - the last engine to drop the table may well be a worker
|
|
// pinned to a different GPU.
|
|
std::shared_ptr<CudaDevicePtr<T>> table(new CudaDevicePtr<T>(count), [device](CudaDevicePtr<T> *p) {
|
|
int current = 0;
|
|
cudaGetDevice(¤t);
|
|
cudaSetDevice(device);
|
|
delete p;
|
|
cudaSetDevice(current);
|
|
});
|
|
jfjoch_cuda_shared_tables::check(
|
|
cudaMemcpyAsync(table->get(), host, count * sizeof(T), cudaMemcpyHostToDevice, stream));
|
|
jfjoch_cuda_shared_tables::check(cudaStreamSynchronize(stream));
|
|
|
|
slot = std::shared_ptr<void>(table);
|
|
return table;
|
|
}
|