Read a chunk without zeroing it first, and hold no frame the fused decoder never writes
Build Packages / build:windows:nocuda (push) Successful in 16m8s
Build Packages / build:windows:cuda (push) Successful in 18m58s
Build Packages / build:viewer-tgz:cpu (push) Successful in 20m35s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m31s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 25m9s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 25m6s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m57s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m58s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m58s
Build Packages / XDS test (durin plugin) (push) Successful in 12m3s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 22m24s
Build Packages / build:rpm (rocky9) (push) Successful in 21m45s
Build Packages / Generate python client (push) Successful in 53s
Build Packages / build:rpm (rocky8) (push) Successful in 26m9s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m37s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m34s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 22m0s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m53s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m29s
Build Packages / DIALS test (push) Successful in 23m40s
Build Packages / Unit tests (push) Successful in 1h20m1s

Three costs before and around the image loop.

Every image allocated a fresh buffer for its compressed chunk and resized it, which
value-initialises, and the read then overwrote every byte. At a few megabytes a chunk
the allocation is large enough to be mapped rather than reused, so the zeroing was
page-fault bound and cost more than the read it preceded - twenty gigabytes of it
over a long sweep. The buffer now uses an allocator that does not construct, and the
two HDF5 read paths are templated on the allocator so every existing caller compiles
unchanged. The rebind is deliberate: without it the vector base rebinds to the
default allocator and the zeroing quietly returns.

The bitshuffle decoder allocated a whole uncompressed frame in its constructor -
seventy megabytes a worker, five hundred and fifty across the loop - for the route
that decodes the shuffled image separately. That route is taken only when a
bitshuffle block is too large for the fused kernel, which neither writer this
pipeline reads produces, so on a real frame the buffer is allocated, never touched,
and freed. It is now allocated where it is used. The comment two lines below already
warned against sizing a buffer from the uncompressed size; the line above it had not
been given the same treatment.

The first call into cuFFT pays the library's one-time initialisation, and it landed
in the middle of the first pass with nothing to overlap it. It is now forced on a
background thread at startup, alongside the file open and the mapping build, in the
manner the shadow finder already uses.

Finally, the detector mask was copied into the start message whether or not a file
would carry it, which a merging run does not. It is filled where a writer is
constructed - both places one is constructed, the second being the fallback that
writes a process file when nothing indexed.

Faster on eleven of thirty-eight crystals and slower on none; the whole rotation test
set falls from four minutes thirty to four minutes seventeen, with each binary
repeating itself to within half a per cent. Space groups thirty-five of thirty-eight
and no failures throughout, and every column of the comparison table is identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EGpGdgmJ8MyY9pCGWjktyi
This commit is contained in:
2026-08-25 01:08:38 +02:00
co-authored by Claude Opus 5
parent 641f890a40
commit 1d16dcd2b9
11 changed files with 100 additions and 48 deletions
+2 -15
View File
@@ -152,21 +152,8 @@ HDF5ImageSource::PrepareDirectRead(const HDF5ImageLocator::Location &loc) const
ds.width, ds.height, ds.mode, ds.algorithm};
}
CompressedImage HDF5ImageSource::ReadDirect(std::vector<uint8_t> &buffer, const DirectChunk &chunk) {
CompressedImage HDF5ImageSource::ReadDirect(RawByteBuffer &buffer, const DirectChunk &chunk) {
buffer.resize(chunk.size);
chunk.file->ReadAt(buffer.data(), chunk.size, chunk.address);
return {buffer, chunk.width, chunk.height, chunk.mode, chunk.algorithm};
}
CompressedImage HDF5ImageSource::ReadImageAt(std::vector<uint8_t> &buffer,
const HDF5ImageLocator::Location &loc) const {
const auto &ds = GetDataset(loc);
const std::vector<hsize_t> start = {static_cast<hsize_t>(loc.local_index), 0, 0};
if (ds.direct_chunk)
ds.dataset->ReadDirectChunk(buffer, start);
else
ds.dataset->ReadVectorToU8(buffer, start, {1, ds.height, ds.width});
return {buffer, ds.width, ds.height, ds.mode, ds.algorithm};
return {buffer.data(), buffer.size(), chunk.width, chunk.height, chunk.mode, chunk.algorithm};
}