764f2659b7e91e8d0ef37294a226c8d3bc13db9c
The wait-and-retry in RotationScaleMergeGPU::Impl::Alloc (c644c4e38) caught
the failed allocation but left its CUDA error as the thread's last error -
CudaDevicePtr leaves it there by contract (CudaDevicePtr_FailedAllocation-
LeavesErrorUntilCleared). The retry then succeeded and the next launch
check, cudaGetLastError() after ReduceGroupMeansKernel or MergeAccum,
reported that stale error over work that went fine, ending the run. Before
the retry existed the failure always ended the run, so the leftover never
mattered.
Reproduced deterministically on a set of 8.2M partials by forcing one merge
allocation to fail (first, early or later buffer of the layout) and by
making the concurrent probe pass hold extra device memory: every such run
failed at the next launch check before, and passes with identical output
now. The battery's "resource already mapped" is the same leftover with
another code from the failed cudaMalloc.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Jungfraujoch
Application to receive data from the PSI JUNGFRAU and EIGER detectors.
All documentation is now placed in docs/ subdirectory and for the current version hosted on Jungfraujoch Read The Docs page.
Languages
C++
78.1%
HTML
6%
C
4.8%
TypeScript
3.5%
Cuda
2.7%
Other
4.8%