Every kernel, cuFFT plan and copy in the beam-centre FFT engine is queued on the legacy NULL
stream, but its buffers came from the stream-ordered pool, whose cudaFreeAsync is ordered on the
thread's non-blocking allocation stream - not after the NULL stream. The free therefore completed
at once while PointShortlist's last Suppress (and, depending on what the plan teardown syncs, the
last inverse transform) was still writing, and the pool was free to hand those bytes to another
worker or return them to the driver.
The capture runs in the background into the first pass of the main run, exactly when the spot
engines are allocating their buffers, and on a loaded shared card its NULL-stream work lags. That
is the "illegal memory access" seen 0.3-0.5 s after "Processing ... images" in four runs across
three binaries and two datasets: a sticky error, reported by whichever worker synchronised next -
the bslz4 device decode, whose host fallback cannot recover from it.
compute-sanitizer --track-stream-ordered-races all over the [BeamCenterFFTGPU] tests: 16502
use-after-free reports before, 0 after; on a rugnux run it flagged the 294 MB transform buffers
freed under a cuFFT kernel. myob output (hkl, mtz, P1 mtz, unmerged mtz) byte-identical.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C