Commit Graph
2 Commits
Author SHA1 Message Date
leonarski_fandClaude Opus 5.5 c644c4e38a rugnux: GPU merge waits for the probe pass running beside it instead of failing
The speculative geometry probe (StartSpeculativeGeometryProbe) runs an
indexing-only pass on a copy of the run beside the pass's scaling merge.
It holds ~4.5 GB of engines (25 image-sized spot-finding buffers, the
per-engine tables, FFT indexers) plus what its streams' pool retains, for a
few seconds. When a merge allocation landed in that window and did not fit,
RotationScaleMergeGPU::Impl::Alloc reported "needs more GPU memory than this
card has ... too large for GPU scaling", although the per-observation arrays
were 1.3 GB on a 16.6 GB card and the set runs fine alone. Timing-dependent:
one full-battery failure, not reproduced in ~15 plain reruns; reproduced
deterministically by letting the probe hold extra device memory.

- GPUWorkBeside (common/CUDAWrapper): a process-wide count of GPU work
  running beside the main line, with a condition variable signalled when the
  last one ends. The speculative probe holds one for its whole pass.
- Alloc: on failure, wait for that work to end (bounded, 10 min, then a
  "GPU busy" error), then ask for the same buffer once more. Nothing else in
  flight means no wait, so a genuine shortage still fails at once. The
  computation is the same whenever the allocation succeeds, so results do
  not depend on the wait.
- "Too large for this card" is now said only by the early check of the
  per-observation arrays against total device memory; a later shortage with
  nothing running beside gets its own message (buffer size, free/total, and
  that another program or the set's size is the cause).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-29 07:44:01 +02:00
leonarski_f 6dfe065365 v1.0.0-rc.172 (#82)
Build Packages / Create release (push) Successful in 16s
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 8m27s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 9m15s
Build Packages / build:viewer-tgz:cpu (push) Successful in 10m11s
Build Packages / build:viewer-tgz:cuda (push) Successful in 12m6s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 15m44s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 16m1s
Build Packages / build:windows:nocuda (push) Successful in 17m29s
Build Packages / build:windows:cuda (push) Successful in 19m58s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 24m7s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 19m8s
Build Packages / build:rugnux:windows (push) Successful in 10m58s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 20m46s
Build Packages / Generate python client (push) Successful in 53s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 20m13s
Build Packages / Build documentation (push) Successful in 1m36s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 19m57s
Build Packages / build:rpm (rocky8) (push) Successful in 18m7s
Build Packages / build:rpm (rocky9) (push) Successful in 18m54s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 19m32s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 17m30s
Build Packages / Unit tests (push) Successful in 1h39m2s
* Fixed `jfjoch_broker` cancelling every data collection with a CUDA "out of memory" error after long operation: GPU memory no longer leaks with each collection.
* Rugnux scales a rotation sweep until the per-frame scales settle instead of for a fixed three rounds, and says so when they did not - merged intensities, and the space group, resolution cut and frame rejection read off them, change accordingly; `--scaling-iterations` is now the cap on that loop (default 100).
* Rugnux places every frame of a marCCD, SMV or miniCBF series at the spindle angle its own header states, so a series with missing frames, or with angles written modulo 360, is no longer read at the wrong geometry or refused.
* Every rotation run writes two diagnostic files beside its reflections: `<prefix>_detector.jpg`, the detector projection with the pixel mask and the detected beam-stop shadow drawn on it, and `<prefix>_plot.txt`, one row per image.

Reviewed-on: #82
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-09-22 06:48:37 +02:00