image_preprocessing: fuse the bitshuffle inverse with preprocessing, and verify the decode

The device decoder was byte-exact on every valid input - 994 production-compressed
images, 927 hand-built LZ4 blocks covering engineered (offset, matchlen) pairs across
the overlap branch boundary, 18000 repeat decodes, sanitizer-clean - and an audit
against LZ4_decompress_generic could not construct a valid block it mis-decodes. What
it did not do was notice when the input was NOT valid, and that mattered more than it
looks: the decode buffers are reused frame to frame, so a block that stopped early left
the PREVIOUS image in place, and in the bitshuffled layout the untouched tail is the
most significant byte-plane. A corrupt chunk therefore did not look like a missing
corner. It looked like thousands of real pixels several powers of two too bright, fed
to spot finding with no diagnostic, where the host decoder had raised an error.

So the kernel now flags a block that fails to reach its declared length while consuming
exactly its payload, and the host turns that into an exception once the caller has
synchronised. Reads are clamped against the end of the payload as well as the output,
both length chains are bounded exactly as read_variable_length bounds them, the two
offset bytes are bounded, and LZ4's parsing restrictions are enforced. On the host side
a block size that is not a multiple of 8 elements is rejected (it made the un-transpose
read uninitialised shared memory), the block count is bounded by what the chunk could
hold before it becomes an allocation (twelve header bytes could demand hundreds of MB
of pinned memory, permanently, per worker), trailing bytes are rejected, and the stream
is synchronised before any throw that happens after work is queued. An image of fewer
than 8 elements is all verbatim tail and now decodes rather than throwing. When the
device route fails for any reason the host decoder gets its turn, so it costs speed
rather than the acquisition.

The lanes cooperate on the copies and a later match can read bytes another lane wrote,
which since Volta needs an explicit __syncwarp(); it worked only because ptxas happened
to reconverge at the post-dominator. The prototype's offset == 1 and power-of-two fast
paths are also restored - the shipped kernel ran a runtime modulo, an emulated 32-bit
division per output byte, on the path its own comment calls the common case.

The un-transpose is now fused with preprocessing. One thread owns one group of 8
elements across every byte-plane, so once it has transposed its 8 bytes out of each
plane it holds 8 complete elements and emits 8 finished int32 pixels with the mask, the
error marker, the saturation cap and the statistics applied. The decompressed image is
never materialised: 0.623 -> 0.411 ms/frame at 18 Mpx, 0.523 -> 0.340 with 8 concurrent
workers. Staging nothing in shared memory also drops the 48 kB ceiling, which had made
any file whose bitshuffle blocks exceed it a hard failure; 64 kB blocks now decode.
gpu_compressed is sized from the chunk with grow-on-demand instead of from the
uncompressed size - it was reserving ~73 MB per worker to hold ~4 MB. Measured on a
1630x1553 uint32 rotation set at -N 32, peak GPU memory falls 3756 -> 3084 MiB; the
same model gives ~144 MB per worker on an 18 Mpx frame.

Decoding on the device also stopped reporting a decompression time, which blanked the
broker's compression plot trace and filled /entry/profiling/compressionTime with NaN.
The decoder brackets the decode with CUDA events and reports it again.

Tests: a differential fuzz suite against the CPU decoder - incompressible and highly
compressible data, engineered offsets, a size sweep hitting every rem%8 value twice,
all six element sizes, an 18 Mpx frame, decoder reuse, concurrency, hand-built LZ4
blocks across the overlap boundary, 26 foreign bitshuffle block sizes from 128 B to
64 kB, corrupt payloads and malformed containers, with a coverage report that proves
which LZ4 paths were reached rather than assuming it. Plus the fused path held byte for
byte against ImagePreprocessorCPU, statistics included, and against the host-upload
path on the same frame.

Battery: 37 crystals, every merged number identical to the host-decode run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-03 14:13:55 +02:00
co-authored by Claude Opus 5
parent 47277674fa
commit bec7e2e922
12 changed files with 2023 additions and 92 deletions
+167
View File
@@ -0,0 +1,167 @@
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute
// SPDX-License-Identifier: GPL-3.0-only
#include <catch2/catch_all.hpp>
#include "../common/CUDAWrapper.h"
#ifdef JFJOCH_USE_CUDA
#include <random>
#include <vector>
#include "../common/PixelMask.h"
#include "../compression/JFJochCompressor.h"
#include "../image_analysis/image_preprocessing/ImagePreprocessorCPU.h"
#include "../image_analysis/image_preprocessing/ImagePreprocessorGPU.h"
#include "../image_analysis/image_preprocessing/ImagePreprocessorBufferGPU.h"
// The device-decode path does NOT decompress into a buffer and then preprocess it: one kernel
// un-transposes the bitshuffle blocks and applies the mask, the error marker, the saturation cap and
// the statistics as it goes, so the decompressed image never exists. That is a different code path
// from the host-upload one, not a reordering of it, and the thing it has to reproduce is the whole
// observable output - every preprocessed pixel AND every counter - against the CPU preprocessor fed
// the host-decompressed image.
//
// Masked, error and saturated pixels are the interesting part: their priority (masked > error >
// saturated) and their sentinel outputs (INT32_MIN / INT32_MIN / INT32_MAX) are decided in the fused
// kernel now, so the image below deliberately contains all three, and the mask deliberately covers
// some of them.
namespace {
DiffractionExperiment MakeExperiment(size_t saturation) {
DiffractionExperiment x(DetJF4M());
x.DetectorDistance_mm(80).BeamX_pxl(1030).BeamY_pxl(1080);
return x;
}
template <class T>
std::vector<T> MakeImage(size_t npixels, T err_value, uint32_t seed) {
std::mt19937 rng(seed);
std::vector<T> img(npixels, 0);
// Sparse background with long runs, so LZ4 produces overlapping matches.
for (size_t i = npixels / 4; i < npixels / 2; i++)
img[i] = static_cast<T>(rng() % 11);
// Bright spots, some above any plausible saturation cap.
for (size_t s = 0; s < 500; s++) {
const size_t c = rng() % npixels;
for (size_t d = 0; d < 5 && c + d < npixels; d++)
img[c + d] = static_cast<T>(30000 + (rng() % 5000));
}
// Explicit error markers, scattered.
for (size_t s = 0; s < 300; s++)
img[rng() % npixels] = err_value;
return img;
}
bool SameStats(const ImageStatistics &a, const ImageStatistics &b) {
return a.max_value == b.max_value && a.min_value == b.min_value
&& a.masked_pixel_count == b.masked_pixel_count
&& a.error_pixel_count == b.error_pixel_count
&& a.saturated_pixel_count == b.saturated_pixel_count;
}
// One element size end to end: compress, decode+preprocess on the device, and compare against the
// host decompression fed through the CPU preprocessor.
template <class T>
void CheckFusedMatchesCPU(CompressedImageMode mode, T err_value, uint32_t seed) {
DiffractionExperiment x = MakeExperiment(32000);
const size_t npixels = x.GetPixelsNum();
PixelMask mask(x);
// Mask a deterministic scatter of pixels, so masked-vs-error-vs-saturated priority is exercised
// rather than assumed.
auto &m = const_cast<std::vector<uint32_t> &>(mask.GetMask());
for (size_t i = 0; i < npixels; i += 997) m[i] = 1;
for (size_t i = 13; i < npixels; i += 4001) m[i] = 1;
const auto img = MakeImage<T>(npixels, err_value, seed);
JFJochBitShuffleCompressor compressor(CompressionAlgorithm::BSHUF_LZ4);
const std::vector<uint8_t> compressed = compressor.Compress(img);
const CompressedImage image(compressed.data(), compressed.size(),
x.GetXPixelsNum(), x.GetYPixelsNum(), mode,
CompressionAlgorithm::BSHUF_LZ4);
REQUIRE(BSLZ4DecoderGPU::Supports(image));
// Reference: host decompression + CPU preprocessing.
ImagePreprocessorCPU cpu_pre(x, mask);
ImagePreprocessorBuffer cpu_buf(npixels);
std::vector<uint8_t> decompression_buffer;
const uint8_t *raw = image.GetUncompressedPtr(decompression_buffer);
const ImageStatistics cpu_stats = cpu_pre.Analyze(cpu_buf, raw, mode);
// Under test: compressed chunk straight to the device, decoded and preprocessed in one pass.
auto stream = std::make_shared<CudaStream>();
ImagePreprocessorGPU gpu_pre(x, mask, stream, /*copy_image_to_host=*/true);
ImagePreprocessorBufferGPU gpu_buf(npixels);
ImageStatistics gpu_stats{};
REQUIRE(gpu_pre.AnalyzeCompressed(gpu_buf, image, gpu_stats));
INFO("mode " << static_cast<int>(mode));
CHECK(SameStats(cpu_stats, gpu_stats));
size_t ndiff = 0, first = 0;
for (size_t i = 0; i < npixels; i++) {
if (gpu_buf[i] != cpu_buf[i]) {
if (ndiff == 0) first = i;
ndiff++;
}
}
INFO("first differing pixel " << first << " cpu " << cpu_buf[first] << " gpu " << gpu_buf[first]
<< " of " << ndiff << " differing");
CHECK(ndiff == 0);
}
} // namespace
TEST_CASE("ImagePreprocessorGPU_FusedDecodeMatchesCPU", "[ImagePreprocessorGPU]") {
if (get_gpu_count() == 0)
SKIP("No CUDA GPU present");
CheckFusedMatchesCPU<uint32_t>(CompressedImageMode::Uint32, UINT32_MAX, 1);
CheckFusedMatchesCPU<uint16_t>(CompressedImageMode::Uint16, UINT16_MAX, 2);
CheckFusedMatchesCPU<int32_t>(CompressedImageMode::Int32, INT32_MIN, 3);
CheckFusedMatchesCPU<int16_t>(CompressedImageMode::Int16, INT16_MIN, 4);
}
// The host-upload path must keep producing exactly what it did - it is still what every non-LZ4
// image takes - so the two entry points are held against each other on the same frame.
TEST_CASE("ImagePreprocessorGPU_FusedMatchesHostUpload", "[ImagePreprocessorGPU]") {
if (get_gpu_count() == 0)
SKIP("No CUDA GPU present");
DiffractionExperiment x = MakeExperiment(32000);
const size_t npixels = x.GetPixelsNum();
PixelMask mask(x);
auto &m = const_cast<std::vector<uint32_t> &>(mask.GetMask());
for (size_t i = 0; i < npixels; i += 1301) m[i] = 1;
const auto img = MakeImage<uint32_t>(npixels, UINT32_MAX, 77);
JFJochBitShuffleCompressor compressor(CompressionAlgorithm::BSHUF_LZ4);
const std::vector<uint8_t> compressed = compressor.Compress(img);
const CompressedImage image(compressed.data(), compressed.size(),
x.GetXPixelsNum(), x.GetYPixelsNum(), CompressedImageMode::Uint32,
CompressionAlgorithm::BSHUF_LZ4);
auto stream = std::make_shared<CudaStream>();
ImagePreprocessorGPU pre(x, mask, stream, /*copy_image_to_host=*/true);
ImagePreprocessorBufferGPU fused_buf(npixels);
ImageStatistics fused_stats{};
REQUIRE(pre.AnalyzeCompressed(fused_buf, image, fused_stats));
// Same engine, same frame, but decompressed on the host and uploaded.
std::vector<uint8_t> decompression_buffer;
const uint8_t *raw = image.GetUncompressedPtr(decompression_buffer);
ImagePreprocessorBufferGPU upload_buf(npixels);
const ImageStatistics upload_stats = pre.Analyze(upload_buf, raw, CompressedImageMode::Uint32);
CHECK(SameStats(fused_stats, upload_stats));
size_t ndiff = 0;
for (size_t i = 0; i < npixels; i++) if (fused_buf[i] != upload_buf[i]) ndiff++;
CHECK(ndiff == 0);
// Decoding on the device replaced a host decompression, so the cost is still reported as one.
CHECK(pre.GetLastDecompressionTime_s() > 0.0f);
}
#endif