The speculative geometry probe (StartSpeculativeGeometryProbe) runs an
indexing-only pass on a copy of the run beside the pass's scaling merge.
It holds ~4.5 GB of engines (25 image-sized spot-finding buffers, the
per-engine tables, FFT indexers) plus what its streams' pool retains, for a
few seconds. When a merge allocation landed in that window and did not fit,
RotationScaleMergeGPU::Impl::Alloc reported "needs more GPU memory than this
card has ... too large for GPU scaling", although the per-observation arrays
were 1.3 GB on a 16.6 GB card and the set runs fine alone. Timing-dependent:
one full-battery failure, not reproduced in ~15 plain reruns; reproduced
deterministically by letting the probe hold extra device memory.
- GPUWorkBeside (common/CUDAWrapper): a process-wide count of GPU work
running beside the main line, with a condition variable signalled when the
last one ends. The speculative probe holds one for its whole pass.
- Alloc: on failure, wait for that work to end (bounded, 10 min, then a
"GPU busy" error), then ask for the same buffer once more. Nothing else in
flight means no wait, so a genuine shortage still fails at once. The
computation is the same whenever the allocation succeeds, so results do
not depend on the wait.
- "Too large for this card" is now said only by the early check of the
per-observation arrays against total device memory; a later shortage with
nothing running beside gets its own message (buffer size, free/total, and
that another program or the set's size is the cause).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C