ca1baffef moved RotationScaleMergeGPU to synchronous cudaMalloc. cudaMalloc
cannot take memory the default pool holds, and after the image loop the pool
still reserves what the per-worker engines freed with cudaFreeAsync (the 1 GB
release threshold only applies at a synchronisation of the pool's streams).
Measured at the merge's construction on a 123.6M-partial rotation set: pool
reserved 5.64 GB, used 0.83 GB - 4.8 GB of the 16.3 GB card idle but
unavailable. The merge's per-observation arrays (about 9.5 GB there) then ran
out at a 494 MB array with 316 MB free, and the run failed with "Failed to
allocate device memory" right after the post-refine gather. The 77710e1ce
binary allocated those arrays from the pool and passed.
A failed synchronous allocation now waits for the device, trims the pool to
zero and tries once more. Memory placement only: no arithmetic changes.
p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz byte-identical on myob, cytc, lyso
and sparse; the failing set now passes (R 3:H, 1.30 A, ISa 10.9).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C