Every merge buffer now comes through Impl::Alloc. When one cannot be had
even after the pool is trimmed, the run ends with what it needed and what to
do: "Scaling N partial observations needs more GPU memory than this card
has: X GB for their per-observation arrays alone (67 bytes each), before the
merge's other buffers, on a Y GB card. This data set is too large for GPU
scaling on this card - run it on the CPU (CUDA_VISIBLE_DEVICES= rugnux ...,
or a CPU-only build)". SetPartialsLayout checks the 67 bytes per observation
of its own arrays against the card's total memory before allocating any of
them, so a set that could not fit an empty card fails before the gigabytes
are allocated.
There is deliberately no switch to CPU scaling mid-run: the two paths differ
in the last bits, so the result would depend on the card.
A 218.7M-partial rotation set (a neighbour-starved dense pattern) now fails
with this message 35 s into the run; with CUDA_VISIBLE_DEVICES= it runs to
completion on the CPU path. No arithmetic changes.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C