Both queue their work (kernels, cudaMemcpy) on the legacy NULL stream - HotPixelFinderGPU for its
download, its sums otherwise on the workers' streams - while their buffers came from the
stream-ordered pool, whose cudaFreeAsync is ordered on the thread's non-blocking allocation stream
and so after none of that work. Unlike BeamCenterFFTGPU no free has overtaken a read: every entry
point waits for its work on the host (cudaDeviceSynchronize, or a blocking device-to-host copy)
before it returns. But that holds only by convention, not for a free while unwinding from a failed
call, and compute-sanitizer --track-stream-ordered-races cannot see host synchronisation, so it
reported every reassigned merge buffer as a use-after-free - noise that buries a real race like the
BeamCenterFFTGPU one.
compute-sanitizer --tool memcheck --track-stream-ordered-races all:
- rugnux --mode scale on a myob _process.h5: 22 use-after-free reports before, 0 after
(MergeAccum/MergeAccumRange/MergeRmeas buffers freed by MergeAccum's reassignment or ~Impl).
- rugnux -e 150 -N 8 on myob: 92 before, 0 after (together with the next commit).
myob, cytc, lyso: p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz byte-identical to r4-integration. Full-run
wall time unchanged within noise (17.4/25.9/18.6 s before, 17.6/25.8/18.7 s after); the scale/merge
phase alone (--mode scale, myob) 2.23 -> 2.31 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C