The device decoder was byte-exact on every valid input - 994 production-compressed images, 927 hand-built LZ4 blocks covering engineered (offset, matchlen) pairs across the overlap branch boundary, 18000 repeat decodes, sanitizer-clean - and an audit against LZ4_decompress_generic could not construct a valid block it mis-decodes. What it did not do was notice when the input was NOT valid, and that mattered more than it looks: the decode buffers are reused frame to frame, so a block that stopped early left the PREVIOUS image in place, and in the bitshuffled layout the untouched tail is the most significant byte-plane. A corrupt chunk therefore did not look like a missing corner. It looked like thousands of real pixels several powers of two too bright, fed to spot finding with no diagnostic, where the host decoder had raised an error. So the kernel now flags a block that fails to reach its declared length while consuming exactly its payload, and the host turns that into an exception once the caller has synchronised. Reads are clamped against the end of the payload as well as the output, both length chains are bounded exactly as read_variable_length bounds them, the two offset bytes are bounded, and LZ4's parsing restrictions are enforced. On the host side a block size that is not a multiple of 8 elements is rejected (it made the un-transpose read uninitialised shared memory), the block count is bounded by what the chunk could hold before it becomes an allocation (twelve header bytes could demand hundreds of MB of pinned memory, permanently, per worker), trailing bytes are rejected, and the stream is synchronised before any throw that happens after work is queued. An image of fewer than 8 elements is all verbatim tail and now decodes rather than throwing. When the device route fails for any reason the host decoder gets its turn, so it costs speed rather than the acquisition. The lanes cooperate on the copies and a later match can read bytes another lane wrote, which since Volta needs an explicit __syncwarp(); it worked only because ptxas happened to reconverge at the post-dominator. The prototype's offset == 1 and power-of-two fast paths are also restored - the shipped kernel ran a runtime modulo, an emulated 32-bit division per output byte, on the path its own comment calls the common case. The un-transpose is now fused with preprocessing. One thread owns one group of 8 elements across every byte-plane, so once it has transposed its 8 bytes out of each plane it holds 8 complete elements and emits 8 finished int32 pixels with the mask, the error marker, the saturation cap and the statistics applied. The decompressed image is never materialised: 0.623 -> 0.411 ms/frame at 18 Mpx, 0.523 -> 0.340 with 8 concurrent workers. Staging nothing in shared memory also drops the 48 kB ceiling, which had made any file whose bitshuffle blocks exceed it a hard failure; 64 kB blocks now decode. gpu_compressed is sized from the chunk with grow-on-demand instead of from the uncompressed size - it was reserving ~73 MB per worker to hold ~4 MB. Measured on a 1630x1553 uint32 rotation set at -N 32, peak GPU memory falls 3756 -> 3084 MiB; the same model gives ~144 MB per worker on an 18 Mpx frame. Decoding on the device also stopped reporting a decompression time, which blanked the broker's compression plot trace and filled /entry/profiling/compressionTime with NaN. The decoder brackets the decode with CUDA events and reports it again. Tests: a differential fuzz suite against the CPU decoder - incompressible and highly compressible data, engineered offsets, a size sweep hitting every rem%8 value twice, all six element sizes, an 18 Mpx frame, decoder reuse, concurrency, hand-built LZ4 blocks across the overlap boundary, 26 foreign bitshuffle block sizes from 128 B to 64 kB, corrupt payloads and malformed containers, with a coverage report that proves which LZ4 paths were reached rather than assuming it. Plus the fused path held byte for byte against ImagePreprocessorCPU, statistics included, and against the host-upload path on the same frame. Battery: 37 crystals, every merged number identical to the host-decode run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
270 lines
9.0 KiB
C++
270 lines
9.0 KiB
C++
// SPDX-FileCopyrightText: 2025 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <cuda_runtime.h>
|
|
#include <cufft.h>
|
|
#include <stdexcept>
|
|
#include <vector>
|
|
#include "../common/JFJochException.h"
|
|
|
|
class CudaStream {
|
|
cudaStream_t stream_ = nullptr;
|
|
public:
|
|
// Non-blocking by default: a stream created with cudaStreamDefault synchronises against the legacy
|
|
// NULL stream, so any NULL-stream operation anywhere in the process serialises every worker's GPU
|
|
// work against every other's. With one engine per worker thread that costs most of the parallelism.
|
|
CudaStream(unsigned int flags = cudaStreamNonBlocking) {
|
|
if (cudaStreamCreateWithFlags(&stream_, flags) != cudaSuccess)
|
|
throw JFJochException(JFJochExceptionCategory::GPUCUDAError,
|
|
"Failed to create CUDA stream");
|
|
}
|
|
~CudaStream() {
|
|
if (stream_) cudaStreamDestroy(stream_);
|
|
}
|
|
// Move-only type
|
|
CudaStream(CudaStream&& other) noexcept : stream_(other.stream_) { other.stream_ = nullptr; }
|
|
CudaStream& operator=(CudaStream&& other) noexcept {
|
|
if (this != &other) {
|
|
if (stream_) cudaStreamDestroy(stream_);
|
|
stream_ = other.stream_;
|
|
other.stream_ = nullptr;
|
|
}
|
|
return *this;
|
|
}
|
|
CudaStream(const CudaStream&) = delete;
|
|
CudaStream& operator=(const CudaStream&) = delete;
|
|
|
|
operator cudaStream_t() const { return stream_; }
|
|
cudaStream_t get() const { return stream_; }
|
|
};
|
|
|
|
// A timing event, so a phase that is queued on a stream can still report how long the device spent
|
|
// on it. cudaEventDisableTiming is deliberately NOT used - timing is the whole point here.
|
|
class CudaEvent {
|
|
cudaEvent_t event_ = nullptr;
|
|
public:
|
|
CudaEvent() {
|
|
if (cudaEventCreate(&event_) != cudaSuccess)
|
|
throw JFJochException(JFJochExceptionCategory::GPUCUDAError,
|
|
"Failed to create CUDA event");
|
|
}
|
|
~CudaEvent() {
|
|
if (event_) cudaEventDestroy(event_);
|
|
}
|
|
CudaEvent(CudaEvent&& other) noexcept : event_(other.event_) { other.event_ = nullptr; }
|
|
CudaEvent& operator=(CudaEvent&& other) noexcept {
|
|
if (this != &other) {
|
|
if (event_) cudaEventDestroy(event_);
|
|
event_ = other.event_;
|
|
other.event_ = nullptr;
|
|
}
|
|
return *this;
|
|
}
|
|
CudaEvent(const CudaEvent&) = delete;
|
|
CudaEvent& operator=(const CudaEvent&) = delete;
|
|
|
|
operator cudaEvent_t() const { return event_; }
|
|
cudaEvent_t get() const { return event_; }
|
|
};
|
|
|
|
class CudaFFTPlan {
|
|
cufftHandle plan_ = 0;
|
|
public:
|
|
CudaFFTPlan() = default;
|
|
|
|
CudaFFTPlan(
|
|
int rank,
|
|
const int* n,
|
|
const int* inembed, int istride, int idist,
|
|
const int* onembed, int ostride, int odist,
|
|
cufftType type, int batch)
|
|
{
|
|
if (cufftPlanMany(&plan_, rank, const_cast<int*>(n),
|
|
const_cast<int*>(inembed), istride, idist,
|
|
const_cast<int*>(onembed), ostride, odist,
|
|
type, batch) != CUFFT_SUCCESS)
|
|
throw JFJochException(JFJochExceptionCategory::GPUCUDAError,
|
|
"Failed to create CUFFT plan with cufftPlanMany");
|
|
}
|
|
|
|
// Convenience overload for vector input
|
|
CudaFFTPlan(
|
|
int rank,
|
|
const std::vector<int>& n,
|
|
const std::vector<int>& inembed, int istride, int idist,
|
|
const std::vector<int>& onembed, int ostride, int odist,
|
|
cufftType type, int batch)
|
|
: CudaFFTPlan(rank, n.data(), inembed.data(), istride, idist, onembed.data(), ostride, odist, type, batch)
|
|
{}
|
|
|
|
~CudaFFTPlan() {
|
|
if (plan_) cufftDestroy(plan_);
|
|
}
|
|
// Move-only type
|
|
CudaFFTPlan(CudaFFTPlan&& other) noexcept : plan_(other.plan_) { other.plan_ = 0; }
|
|
CudaFFTPlan& operator=(CudaFFTPlan&& other) noexcept {
|
|
if (this != &other) {
|
|
if (plan_) cufftDestroy(plan_);
|
|
plan_ = other.plan_;
|
|
other.plan_ = 0;
|
|
}
|
|
return *this;
|
|
}
|
|
CudaFFTPlan(const CudaFFTPlan&) = delete;
|
|
CudaFFTPlan& operator=(const CudaFFTPlan&) = delete;
|
|
|
|
operator cufftHandle() const { return plan_; }
|
|
cufftHandle get() const { return plan_; }
|
|
};
|
|
|
|
template <typename T>
|
|
class CudaDevicePtr {
|
|
T* ptr_ = nullptr;
|
|
public:
|
|
CudaDevicePtr() = default;
|
|
|
|
explicit CudaDevicePtr(size_t count) {
|
|
if (cudaMalloc(&ptr_, count * sizeof(T)) != cudaSuccess)
|
|
throw JFJochException(JFJochExceptionCategory::GPUCUDAError,
|
|
"Failed to allocate device memory");
|
|
}
|
|
~CudaDevicePtr() {
|
|
if (ptr_) cudaFree(ptr_);
|
|
}
|
|
// Move-only type
|
|
CudaDevicePtr(CudaDevicePtr&& other) noexcept : ptr_(other.ptr_) { other.ptr_ = nullptr; }
|
|
CudaDevicePtr& operator=(CudaDevicePtr&& other) noexcept {
|
|
if (this != &other) {
|
|
if (ptr_) cudaFree(ptr_);
|
|
ptr_ = other.ptr_;
|
|
other.ptr_ = nullptr;
|
|
}
|
|
return *this;
|
|
}
|
|
CudaDevicePtr(const CudaDevicePtr&) = delete;
|
|
CudaDevicePtr& operator=(const CudaDevicePtr&) = delete;
|
|
|
|
T* get() const { return ptr_; }
|
|
operator T*() const { return ptr_; }
|
|
};
|
|
|
|
template <typename T>
|
|
class CudaHostPtr {
|
|
T* ptr_ = nullptr;
|
|
public:
|
|
CudaHostPtr() = default;
|
|
explicit CudaHostPtr(size_t count) {
|
|
if (cudaMallocHost(&ptr_, count * sizeof(T)) != cudaSuccess)
|
|
throw JFJochException(JFJochExceptionCategory::GPUCUDAError,
|
|
"Failed to allocate pinned host memory");
|
|
}
|
|
~CudaHostPtr() {
|
|
if (ptr_) cudaFreeHost(ptr_);
|
|
}
|
|
// Move-only type
|
|
CudaHostPtr(CudaHostPtr&& other) noexcept : ptr_(other.ptr_) { other.ptr_ = nullptr; }
|
|
CudaHostPtr& operator=(CudaHostPtr&& other) noexcept {
|
|
if (this != &other) {
|
|
if (ptr_) cudaFreeHost(ptr_);
|
|
ptr_ = other.ptr_;
|
|
other.ptr_ = nullptr;
|
|
}
|
|
return *this;
|
|
}
|
|
CudaHostPtr(const CudaHostPtr&) = delete;
|
|
CudaHostPtr& operator=(const CudaHostPtr&) = delete;
|
|
|
|
T* get() const { return ptr_; }
|
|
operator T*() const { return ptr_; }
|
|
};
|
|
|
|
template <typename T>
|
|
class CudaRegisteredVector {
|
|
std::vector<T>* vec_ = nullptr;
|
|
bool registered_ = false;
|
|
|
|
static void registerPtr(void* ptr, size_t bytes, unsigned int flags) {
|
|
cudaError_t err = cudaHostRegister(ptr, bytes, flags);
|
|
if (err != cudaSuccess)
|
|
throw JFJochException(JFJochExceptionCategory::GPUCUDAError, "cudaHostRegister failed");
|
|
}
|
|
static void unregisterPtr(void* ptr) {
|
|
cudaError_t err = cudaHostUnregister(ptr);
|
|
if (err != cudaSuccess)
|
|
throw JFJochException(JFJochExceptionCategory::GPUCUDAError, "cudaHostUnregister failed");
|
|
}
|
|
|
|
public:
|
|
// Non-owning wrapper. Does NOT provide accessors to the vector.
|
|
CudaRegisteredVector() = default;
|
|
|
|
CudaRegisteredVector(std::vector<T>& vec, unsigned int flags = cudaHostRegisterDefault)
|
|
: vec_(&vec)
|
|
{
|
|
if (!vec.empty()) {
|
|
registerPtr(vec.data(), vec.size() * sizeof(T), flags);
|
|
registered_ = true;
|
|
}
|
|
}
|
|
|
|
~CudaRegisteredVector() {
|
|
if (registered_ && vec_ && !vec_->empty()) {
|
|
unregisterPtr(vec_->data());
|
|
}
|
|
}
|
|
|
|
// Move-only
|
|
CudaRegisteredVector(CudaRegisteredVector&& other) noexcept
|
|
: vec_(other.vec_), registered_(other.registered_) {
|
|
other.vec_ = nullptr;
|
|
other.registered_ = false;
|
|
}
|
|
|
|
CudaRegisteredVector& operator=(CudaRegisteredVector&& other) noexcept {
|
|
if (this != &other) {
|
|
// Clean current registration
|
|
if (registered_ && vec_ && !vec_->empty()) {
|
|
unregisterPtr(vec_->data());
|
|
}
|
|
vec_ = other.vec_;
|
|
registered_ = other.registered_;
|
|
other.vec_ = nullptr;
|
|
other.registered_ = false;
|
|
}
|
|
return *this;
|
|
}
|
|
|
|
CudaRegisteredVector(const CudaRegisteredVector&) = delete;
|
|
CudaRegisteredVector& operator=(const CudaRegisteredVector&) = delete;
|
|
|
|
// Re-register after vector capacity/size change. Caller must ensure
|
|
// the vector is not registered at the moment of mutation.
|
|
void rebind(std::vector<T>& vec, unsigned int flags = cudaHostRegisterDefault) {
|
|
// Unregister previous if needed
|
|
if (registered_ && vec_ && !vec_->empty()) {
|
|
unregisterPtr(vec_->data());
|
|
}
|
|
vec_ = &vec;
|
|
if (!vec.empty()) {
|
|
registerPtr(vec.data(), vec.size() * sizeof(T), flags);
|
|
registered_ = true;
|
|
} else {
|
|
registered_ = false;
|
|
}
|
|
}
|
|
|
|
// Explicit unregister (optional).
|
|
void unregister() {
|
|
if (registered_ && vec_ && !vec_->empty()) {
|
|
unregisterPtr(vec_->data());
|
|
registered_ = false;
|
|
}
|
|
}
|
|
|
|
bool isRegistered() const { return registered_; }
|
|
};
|
|
|
|
|