The wait-and-retry in RotationScaleMergeGPU::Impl::Alloc (c644c4e38) caught
the failed allocation but left its CUDA error as the thread's last error -
CudaDevicePtr leaves it there by contract (CudaDevicePtr_FailedAllocation-
LeavesErrorUntilCleared). The retry then succeeded and the next launch
check, cudaGetLastError() after ReduceGroupMeansKernel or MergeAccum,
reported that stale error over work that went fine, ending the run. Before
the retry existed the failure always ended the run, so the leftover never
mattered.
Reproduced deterministically on a set of 8.2M partials by forcing one merge
allocation to fail (first, early or later buffer of the layout) and by
making the concurrent probe pass hold extra device memory: every such run
failed at the next launch check before, and passes with identical output
now. The battery's "resource already mapped" is the same leftover with
another code from the failed cudaMalloc.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C