Read a chunk without zeroing it first, and hold no frame the fused decoder never writes
Build Packages / build:windows:nocuda (push) Successful in 16m8s
Build Packages / build:windows:cuda (push) Successful in 18m58s
Build Packages / build:viewer-tgz:cpu (push) Successful in 20m35s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m31s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 25m9s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 25m6s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m57s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m58s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m58s
Build Packages / XDS test (durin plugin) (push) Successful in 12m3s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 22m24s
Build Packages / build:rpm (rocky9) (push) Successful in 21m45s
Build Packages / Generate python client (push) Successful in 53s
Build Packages / build:rpm (rocky8) (push) Successful in 26m9s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m37s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m34s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 22m0s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m53s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m29s
Build Packages / DIALS test (push) Successful in 23m40s
Build Packages / Unit tests (push) Successful in 1h20m1s
Build Packages / build:windows:nocuda (push) Successful in 16m8s
Build Packages / build:windows:cuda (push) Successful in 18m58s
Build Packages / build:viewer-tgz:cpu (push) Successful in 20m35s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m31s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 25m9s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 25m6s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m57s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m58s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m58s
Build Packages / XDS test (durin plugin) (push) Successful in 12m3s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 22m24s
Build Packages / build:rpm (rocky9) (push) Successful in 21m45s
Build Packages / Generate python client (push) Successful in 53s
Build Packages / build:rpm (rocky8) (push) Successful in 26m9s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m37s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m34s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 22m0s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m53s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m29s
Build Packages / DIALS test (push) Successful in 23m40s
Build Packages / Unit tests (push) Successful in 1h20m1s
Three costs before and around the image loop. Every image allocated a fresh buffer for its compressed chunk and resized it, which value-initialises, and the read then overwrote every byte. At a few megabytes a chunk the allocation is large enough to be mapped rather than reused, so the zeroing was page-fault bound and cost more than the read it preceded - twenty gigabytes of it over a long sweep. The buffer now uses an allocator that does not construct, and the two HDF5 read paths are templated on the allocator so every existing caller compiles unchanged. The rebind is deliberate: without it the vector base rebinds to the default allocator and the zeroing quietly returns. The bitshuffle decoder allocated a whole uncompressed frame in its constructor - seventy megabytes a worker, five hundred and fifty across the loop - for the route that decodes the shuffled image separately. That route is taken only when a bitshuffle block is too large for the fused kernel, which neither writer this pipeline reads produces, so on a real frame the buffer is allocated, never touched, and freed. It is now allocated where it is used. The comment two lines below already warned against sizing a buffer from the uncompressed size; the line above it had not been given the same treatment. The first call into cuFFT pays the library's one-time initialisation, and it landed in the middle of the first pass with nothing to overlap it. It is now forced on a background thread at startup, alongside the file open and the mapping build, in the manner the shadow finder already uses. Finally, the detector mask was copied into the start message whether or not a file would carry it, which a merging run does not. It is filled where a writer is constructed - both places one is constructed, the second being the fallback that writes a process file when nothing indexed. Faster on eleven of thirty-eight crystals and slower on none; the whole rotation test set falls from four minutes thirty to four minutes seventeen, with each binary repeating itself to within half a per cent. Space groups thirty-five of thirty-eight and no failures throughout, and every column of the comparison table is identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EGpGdgmJ8MyY9pCGWjktyi
This commit is contained in:
@@ -118,7 +118,6 @@ bool BSLZ4DecoderGPU::Supports(const CompressedImage &image) {
|
||||
BSLZ4DecoderGPU::BSLZ4DecoderGPU(size_t in_max_uncompressed_bytes, std::shared_ptr<CudaStream> in_stream)
|
||||
: stream(std::move(in_stream)),
|
||||
max_uncompressed_bytes(in_max_uncompressed_bytes) {
|
||||
gpu_shuffled = CudaDevicePtr<uint8_t>(max_uncompressed_bytes);
|
||||
gpu_status = CudaDevicePtr<uint32_t>(1);
|
||||
host_status = CudaHostPtr<uint32_t>(1);
|
||||
// The compressed buffer and the descriptors are grown to fit the first image instead of being
|
||||
@@ -126,13 +125,15 @@ BSLZ4DecoderGPU::BSLZ4DecoderGPU(size_t in_max_uncompressed_bytes, std::shared_p
|
||||
// tens; sizing this from the UNCOMPRESSED size cost ~73 MB per worker to hold ~4 MB.
|
||||
}
|
||||
|
||||
// The stored pixel depth is fixed within a dataset, so in practice this runs once - but it is not
|
||||
// promised anywhere, and sizing for the widest type instead would hold twice the memory a 16-bit
|
||||
// detector needs. Grown with slack because cudaMalloc and cudaFree synchronise the whole device.
|
||||
// gpu_shuffled holds a whole uncompressed frame - 72 MB at 18 Mpx, per worker - and only the
|
||||
// DecodeShuffled() route ever writes it. That route is taken when a bitshuffle block is too large for
|
||||
// the fused kernel, which neither writer this pipeline reads produces, so on a real frame the buffer
|
||||
// is never touched. So allocate it the first time it is actually asked for, at the size the caller
|
||||
// declared, rather than in the constructor.
|
||||
void BSLZ4DecoderGPU::EnsureUncompressedCapacity(size_t bytes) {
|
||||
if (bytes <= max_uncompressed_bytes)
|
||||
if (gpu_shuffled.get() && bytes <= max_uncompressed_bytes)
|
||||
return;
|
||||
const size_t want = std::max(bytes, max_uncompressed_bytes + max_uncompressed_bytes / 2);
|
||||
const size_t want = std::max(bytes, max_uncompressed_bytes);
|
||||
cuda_err(cudaStreamSynchronize(*stream));
|
||||
gpu_shuffled = CudaDevicePtr<uint8_t>(want);
|
||||
max_uncompressed_bytes = want;
|
||||
@@ -176,7 +177,6 @@ BSLZ4ShuffledImage BSLZ4DecoderGPU::PrepareChunk(const CompressedImage &image) {
|
||||
|
||||
if (clen < 12)
|
||||
throw JFJochException(JFJochExceptionCategory::Compression, "bslz4 chunk shorter than its header");
|
||||
EnsureUncompressedCapacity(total_bytes);
|
||||
if (be64(src) != total_bytes)
|
||||
throw JFJochException(JFJochExceptionCategory::Compression, "bslz4 header size does not match the image");
|
||||
|
||||
@@ -259,6 +259,7 @@ BSLZ4ShuffledImage BSLZ4DecoderGPU::PrepareChunk(const CompressedImage &image) {
|
||||
}
|
||||
|
||||
BSLZ4ShuffledImage BSLZ4DecoderGPU::DecodeShuffled(const CompressedImage &image) {
|
||||
EnsureUncompressedCapacity(image.GetUncompressedSize());
|
||||
BSLZ4ShuffledImage ret = PrepareChunk(image);
|
||||
if (ret.nblocks > 0) {
|
||||
lz4_decode_blocks<<<(ret.nblocks * 32 + 255) / 256, 256, 0, *stream>>>(
|
||||
|
||||
Reference in New Issue
Block a user