The engines' kernels run on their own non-blocking stream, which is not ordered after the legacy
NULL stream. The counter reset (cudaMemset, asynchronous to the host) and the radial-background
kernel upload in the Bragg engine constructor, and the direction-grid upload in FFTIndexerGPU (a
pageable cudaMemcpy returns once staged, before the DMA completes) were issued on the NULL stream,
so the first kernel reading them was not ordered after them. Now cudaMemsetAsync/cudaMemcpyAsync on
the engine's stream. In the Bragg constructor that is this->stream: the parameter of the same name
has been moved from.
compute-sanitizer --track-stream-ordered-races all flagged the Bragg ones (16 + 16 reports on
rugnux -e 150 myob; 1 in the [CUDAMemHelpers] tests); 0 after. Outputs byte-identical on myob,
cytc, lyso.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C