Build Packages / build:viewer-tgz:cpu (push) Successful in 20m32s
Build Packages / build:viewer-tgz:cuda (push) Successful in 20m40s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m24s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m8s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 27m31s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 27m38s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 29m7s
Build Packages / XDS test (durin plugin) (push) Successful in 11m12s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 22m49s
Build Packages / build:rpm (rocky9) (push) Successful in 22m51s
Build Packages / Generate python client (push) Successful in 40s
Build Packages / Build documentation (push) Successful in 1m22s
Build Packages / Create release (push) Skipped
Build Packages / DIALS test (push) Successful in 20m21s
Build Packages / build:rpm (rocky8) (push) Successful in 27m26s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m59s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m52s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m41s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m41s
Build Packages / Unit tests (push) Successful in 1h17m41s
Build Packages / build:windows:nocuda (push) Successful in 13m24s
Build Packages / build:windows:cuda (push) Successful in 17m0s
The pipeline decompressed each image on the host and uploaded the result. On an 18 Mpx rotation dataset that made the host-to-device copy the bottleneck of the whole per-image loop: nsys puts the copies at 78% of the loop against 39% for every kernel combined - 3600 transfers of 72.4 MB - and they ran at only 12.5 GB/s of an available 27-28 because the host-side decompression was itself saturating host memory bandwidth. The GPU was mostly waiting. So the compressed chunk goes across instead, about 4 MB rather than 72 MB, and is decoded on the device. That removes the transfer and the host decompression that was throttling it, in one change. Measured on an idle machine, a run goes from 45.11 s to 24.97 s - 1.81x - with the merged output unchanged. THE APPROACH IS JON WRIGHT'S (ESRF): "Experiences with GPU decompression for bitshuffle + LZ4 data", HDF5 User Group 2021, and github.com/jonwright/ bslz4decoders. The kernels here are ours, but the idea and the demonstration that it is worth doing are his. Cited in docs/ACKNOWLEDGEMENT.md and in the new section 0 of docs/CPU_DATA_ANALYSIS.md. Two kernels mirror the CPU decoder. LZ4 runs one WARP per bitshuffle block: every lane parses the same sequence stream (a broadcast read, no divergence) and the literal and match copies are split across the 32 lanes so the stores coalesce; an overlapping match is treated as a pattern of period offset sourced from bytes that already precede the write position, which keeps it parallel rather than a serial byte loop. One thread per block instead measured 13x slower. The bitshuffle inverse then un-transposes each byte-plane through shared memory and interleaves the planes back into elements. Only BSHUF_LZ4 is decoded on the device. The zstd variants have no device decoder, and neither has an uncompressed or float image; Supports() returns false for those and the caller decompresses on the host exactly as before. The fallback is explicit, so a format we cannot decode on the device is a slower path and never a wrong answer. Tests hold the device decoder against the CPU one byte for byte, on data from the production compressor, for every element size the detectors emit - including the 8-bit DECTRIS modes, which take bitshuf_decode_block's separate elem_size == 1 branch - plus a many-block frame, the formats it must decline, and malformed containers, which must throw rather than run off a buffer. Battery: 37 crystals, no failures, identical to the host-decode run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
56 lines
2.6 KiB
C++
56 lines
2.6 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <memory>
|
|
|
|
#include "../../common/CompressedImage.h"
|
|
#include "../indexing/CUDAMemHelpers.h"
|
|
|
|
// One bitshuffle block, located by the host scan and consumed by both kernels.
|
|
struct BSLZ4BlockDesc {
|
|
uint32_t in_off; // byte offset of the LZ4 payload within the chunk
|
|
uint32_t in_len; // compressed length
|
|
uint32_t out_off; // byte offset of this block's output in the image
|
|
uint32_t nelem; // elements in this block (the last one is usually shorter)
|
|
};
|
|
|
|
// Decompress a bitshuffle+LZ4 image ON THE DEVICE, so the compressed bytes are what crosses PCIe.
|
|
//
|
|
// The idea - upload the compressed chunk and decode it on the GPU rather than decompressing on the
|
|
// host - is Jon Wright's (ESRF); see https://github.com/jonwright/bslz4decoders and his 2021 HDF5
|
|
// User Group talk "Experiences with GPU decompression for bitshuffle + LZ4 data". The kernels here
|
|
// are our own, but the approach, and the observation that it is worth doing at all, are his.
|
|
//
|
|
// Why it pays: a full 18 Mpx uint32 frame is 72 MB decompressed and about 4 MB compressed, and
|
|
// profiling showed the host-to-device copy owning ~78% of the per-image loop against ~39% for
|
|
// kernels. Decoding on the device removes both that transfer and the host-side decompression,
|
|
// whose memory traffic was itself holding the copy engine well below the link rate.
|
|
//
|
|
// Only BSHUF_LZ4 is handled. The zstd variants have no device decoder, so Supports() returns false
|
|
// and the caller decompresses on the host exactly as before.
|
|
class BSLZ4DecoderGPU {
|
|
std::shared_ptr<CudaStream> stream;
|
|
|
|
CudaDevicePtr<uint8_t> gpu_compressed;
|
|
CudaDevicePtr<uint8_t> gpu_shuffled; // LZ4 output, still bitshuffled
|
|
CudaDevicePtr<BSLZ4BlockDesc> gpu_desc;
|
|
CudaHostPtr<BSLZ4BlockDesc> host_desc; // pinned, so the descriptor upload is truly async
|
|
|
|
size_t max_compressed_bytes = 0;
|
|
size_t max_uncompressed_bytes = 0;
|
|
size_t max_blocks = 0;
|
|
|
|
public:
|
|
BSLZ4DecoderGPU(size_t max_uncompressed_bytes, std::shared_ptr<CudaStream> stream);
|
|
|
|
// True when this image can be decoded on the device. Everything else must go the host route.
|
|
static bool Supports(const CompressedImage &image);
|
|
|
|
// Decode into gpu_out, which must hold image.GetUncompressedSize() bytes. Work is queued on the
|
|
// decoder's stream and the caller synchronises. Throws if the container is malformed - it comes
|
|
// off the network or off disk, so it is not trusted.
|
|
void Decode(const CompressedImage &image, uint8_t *gpu_out);
|
|
};
|