SlurpGz cleared the reused read buffer and grew it back in 1 MB resizes, so every frame
zero-filled its whole decompressed size before gzread overwrote it. It now grows the buffer only
past what it already holds and trims it to the bytes read; the bytes are the same.
8qaw (Eiger 16M, .cbf.gz): memset 5.7% -> 0.5% of the run's user cycles (perf, GPU build). The GPU
build's image loops do not move (pre-pass / output loop 15.28 / 28.63 -> 15.16 / 28.62 s): they are
not bound by the decode. The CPU build's loops are; with the prediction change after this one, its
8qaw loops go 116.6 / 251.3 -> 94.4 / 206.8 s.
Tried and not kept: reading the next frame on a thread of its own while the worker analyses the
current one (two raw images per worker, two page-locked regions in ImagePreprocessorGPU). Exact,
but the 8qaw loops got slower (15.3 / 28.6 -> 16.1 / 29.8 s) and myob's too (2.10 -> 2.40 s): with
the decode overlapped the loop is still waiting on the uploads.
p.mtz byte-identical on 8qaw, GPU and CPU builds; the HDF5 sets do not reach this code (and are
identical with the commit after this one).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi