Files
Jungfraujoch/docs/RUGNUX_INSTALL.md
T
leonarski_f 84228bf8be
Build Packages / Create release (push) Successful in 24s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m29s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m43s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 8m27s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 9m53s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 9m58s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 11m22s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 13m39s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 18m37s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 16m32s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 24m11s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m30s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 19m3s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 20m23s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 19m41s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 21m0s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m38s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 14m33s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 17m55s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 20m50s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 18m38s
Build Packages / Unit tests (push) Successful in 1h46m14s
v1.0.0-rc.173 (#83)
* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports.
* jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls.
* Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results.
* Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable.
* Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate.
* Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do.
* Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence.
* Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags.
* Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check.
* Rugnux: Clear error messages when a data set needs more GPU or host memory than is available.

Reviewed-on: #83
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-09-29 15:57:32 +02:00

12 KiB

Installing Rugnux

:local:
:depth: 2

rugnux is a single self-contained executable. It needs no CUDA toolkit, no Qt, and no Jungfraujoch service running anywhere; on a machine with an NVIDIA GPU it needs the NVIDIA driver, and without one it still runs on the CPU.

From the package repositories (RHEL / Rocky / Ubuntu)

On a distribution covered by the package repositories, rugnux is a package of its own:

sudo dnf install rugnux          # RHEL / Rocky 8 and 9
sudo apt install rugnux          # Ubuntu 22.04 / 24.04

It installs /usr/bin/rugnux and depends on nothing from the acquisition side — no broker, no detector libraries, no Qt — so it can go on a machine that only processes data.

Upgrading from rc.163 or earlier. /usr/bin/rugnux used to belong to the jfjoch-viewer package. The rugnux package declares that the file has moved, so installing it upgrades an old jfjoch-viewer in the same transaction instead of failing on the duplicate path. If your jfjoch-viewer is pinned to an old version, unpin it or remove it first.

From the release archive

For a machine no package manager covers — or for Windows, macOS and Arm, which have no repository — take the archive for your architecture from the Gitea release page:

Archive For
rugnux-<version>-linux-x86_64-cuda12.tgz 64-bit Intel/AMD Linux. Built on RHEL 8, so it runs on any newer Linux
rugnux-<version>-linux-aarch64-cuda13.tgz 64-bit Arm Linux — NVIDIA GH200 and DGX Spark. Built on Ubuntu 24.04, so it needs glibc 2.39 or newer. Cross-compiled and not yet exercised on Arm hardware
rugnux-<version>-win64-cuda13.zip 64-bit Windows
rugnux-<version>-macos-arm64-cpu.tgz macOS 13 (Ventura) or newer on Apple Silicon (M1 and newer). CPU only — there is no CUDA on macOS. Intel Macs are not supported

The archive has no top-level directory — it unpacks straight into bin/ and share/. Always give tar a destination of its own, or it will scatter those into whatever directory you are in:

mkdir -p /opt/rugnux-1.0.0
tar xzf rugnux-1.0.0-linux-x86_64-cuda12.tgz -C /opt/rugnux-1.0.0
/opt/rugnux-1.0.0/bin/rugnux            # prints the usage

What you get is:

bin/rugnux                                  the program
share/doc/jfjoch_rugnux/LICENSE             GPLv3
share/doc/jfjoch_rugnux/THIRD_PARTY_NOTICES.md
share/doc/jfjoch_rugnux/licenses/           verbatim licence texts of the bundled dependencies

Nothing is written outside that directory, nothing needs root, and several versions can sit side by side. To remove it, delete the directory. Put bin/ on your PATH if you want to type rugnux rather than the full path.

macOS: the first run is blocked. The release is not yet notarized by Apple, and a .tgz downloaded with a browser hands its download flag on to everything tar extracts from it, so macOS refuses to run rugnux ("cannot be opened because the developer cannot be verified"). Clear the flag once for the whole directory:

xattr -dr com.apple.quarantine ~/rugnux-1.0.0

An archive fetched with curl carries no such flag. Unpacking into a directory under your home directory (~/rugnux-<version>) rather than /opt avoids needing sudo on a Mac.

Mixing the two. If a rugnux package is also installed, /usr/bin/rugnux will normally win on PATH. Put the archive's bin/ first, or call it by its full path, to be sure which one you are running — rugnux prints its version on every run.

GPU support

The released Linux and Windows archives are CUDA builds (the macOS one is CPU-only). They need only an NVIDIA driver on the host — 525.60.13 or newer for the CUDA 12 archive, 580.65.06 or newer for the CUDA 13 ones — and no CUDA toolkit, because everything CUDA is linked statically. With no GPU or no driver, rugnux reports zero CUDA devices and falls back to the CPU path, which works but is far slower and offers only the fftw indexer. A V100 needs the CUDA 12 archive; which generations each build covers is in Release contents ▸ GPU generations and the NVIDIA driver.

Building from source

rugnux alone, without the server stack or Qt:

cmake -S . -B build -DJFJOCH_RUGNUX_ONLY=ON -DCMAKE_BUILD_TYPE=Release \
      -DCMAKE_CXX_FLAGS="-march=x86-64-v3" -DCMAKE_C_FLAGS="-march=x86-64-v3"
cmake --build build -j$(nproc) --target rugnux

The binary lands in build/rugnux/rugnux. No library has to come from the system: every dependency is downloaded and built during the first configure, which therefore needs network access. cmake --build build --target package produces the same .tgz the release ships. The -march flag is not set by the build system on purpose, so a plain build is slower than the released one on the CPU-bound stages — see the note in CMakeLists.txt.

On a Mac (Apple Silicon), leave the two -march flags out — they name an x86 level; the Apple compiler's default target is already the M1 — and use sysctl -n hw.ncpu for the job count. The build has no CUDA there and needs nothing installed beyond Xcode and CMake.

Hardware

As with the rest of Jungfraujoch, serious performance requires an NVIDIA GPU. The CUDA build provides the GPU fast-feedback indexer (ffbidx) and the GPU FFT indexer (fft); without CUDA only the CPU fftw indexer is available. With a GPU present most of the per-image pipeline runs on the device — bitshuffle+LZ4 decompression, image preprocessing, azimuthal integration, spot finding, prediction and Bragg integration — as do rotation scaling and merging and the --model rigid-body placement, with CPU implementations as the fallback where there is no GPU. The choice is automatic; a CUDA build runs on the CPU when the GPU is hidden from it (CUDA_VISIBLE_DEVICES= rugnux ...). The thread count (-N) governs the CPU side of all of it.

The released CUDA builds need only an NVIDIA driver on the host, no CUDA toolkit: 525.60.13 or newer for the CUDA 12 artefacts (RHEL 8 packages, the x86_64 rugnux archive) and 580.65.06 or newer for the CUDA 13 ones (RHEL 9, Ubuntu, the aarch64 and Windows rugnux archives). Which GPU generations each artefact supports — a V100 in particular works only with the CUDA 12 build — is in Release contents ▸ GPU generations and the NVIDIA driver.

Memory

A run's memory is set by how many reflection observations it integrates, not by how many pixels the detector has. Measured on rc.169 at -N 6, over rotation sweeps of 1200-1800 frames:

Detector Observations integrated Host (peak RSS) Device
1679 x 1475 (2.5 MP) 25 M 7.8 GB 3.8 GB
2527 x 2463 (6.2 MP) 13 M 4.9 GB 3.1 GB
3262 x 3108 (10.1 MP) 7 M 2.5 GB 3.0 GB
4371 x 4150 (18.1 MP) 13 M 4.3 GB 3.2 GB
4371 x 4150 (18.1 MP) 20 M 6.0 GB 4.4 GB
3262 x 3108 (10.1 MP) 50 M 14.2 GB 7.0 GB

Two sweeps on the same detector differ by a factor of five, so size the machine on the observation count rather than on the detector. As a rule, 0.6 GB + 0.3 GB of host memory per million observations, and 2.4 GB + 0.1 GB of device memory per million. The run reports what it actually integrated — RotationScaleMerge: ingested ... partial observations. 32 GB of host RAM and an 8 GB card cover every dataset measured here at -N 6; an ordinary sweep needs 16 GB and 6 GB.

Both peaks fall in scaling and merging, not in the per-image loop: the loop alone stays under 3.2 GB on the device at -N 6, whatever the detector. The stage that needs the most is the 3D combine.

-N multiplies the per-image loop only — about 8 bytes of device memory per detector pixel per worker (150 MB per worker at 18 MP, 87 MB at 6 MP) on top of a fixed ~2.6 GB, and it is linear with no ceiling: the same 18 MP sweep takes 2.7 GB of device memory at -N 1, 3.2 GB at -N 6 and 8.0 GB at -N 32. Host memory does not move with -N, and neither does the merge's own device memory. Where several runs share one card, -N is the knob that decides how many fit.

To use less:

  • -N — the only flag that lowers the per-image loop's device memory. Lowering it is what rescues a run on a card that is otherwise full.
  • --no-merge — on the heaviest sweep measured it took the host peak from 14.2 to 8.3 GB and the device peak from 7.0 to 2.9 GB. It also gives up the merged output.
  • --no-export-unmerged does not save memory. <prefix>_unmerged.mtz is the largest file a run writes, but its observations are resident either way; the flag saves disk, not RAM.

Nothing measures the free memory on the card before allocating. A run that does not fit degrades where it can — device decompression falls back to the host, the beam-stop projection falls back to the host, an indexer thread that cannot allocate reports the frame as not indexed, the --model rigid-body placement moves to the CPU — and fails the run otherwise, with Processing failed: CUDA (GPU) error (Failed to allocate device memory) and a non-zero exit status. That is deliberate: a CUDA out-of-memory says nothing about the frame in flight and everything about the machine, so skipping frames would leave a run that looks complete with quietly different merged numbers. A run is never silently short; it either completes or fails. On a GPU shared with another job, lower -N or wait for the card.

Very large unit cells

The rule above holds for the cells in that table, and a very large cell sits far beyond it. A long axis, a dense pattern and many frames multiply together: each frame predicts more reflections, each reflection is split over more frames, and tens of millions of partial observations are normal for a big cell. One sweep of a crystal with an axis of about 640 Å predicted 236 million reflections and peaked at 17.5 GB of host memory on a GPU build and 72.5 GB on a CPU-only build. A CPU-only build keeps in host RAM what a GPU build keeps on the card, so it is the one that needs the large machine. When such a set is too large for GPU scaling on the card, the run stops and says to run it on the CPU (CUDA_VISIBLE_DEVICES= rugnux ..., or a CPU-only build) — which then wants that much host memory.

When host memory runs out

Nothing estimates the host memory a run will need before it starts, so a set that does not fit runs until the memory is gone, and then ends in one of two ways.

Where the allocation is refused — under an address-space limit (ulimit -v, which some batch systems set), with strict overcommit, or for one request larger than the machine could ever grant — the run stops with

Processing failed: out of host memory (std::bad_alloc) - this data set needs more RAM than is available

and a non-zero exit status. The CUDA driver reserves a large address space of its own, so a GPU build under a tight ulimit -v can stop before reading anything, with CUDA (GPU) error (out of memory).

More often on Linux the allocation succeeds and the memory runs out as it is used. The kernel's OOM killer then ends the process with SIGKILL, which no program can catch: the run stops with no message of its own, the shell prints Killed (exit status 137), and dmesg or the batch system's log has an Out of memory: Killed process line. A job capped by a cgroup memory limit ends the same way. A run that disappears like this needs more RAM than the machine or the job was given, exactly as if it had said so. --no-merge lowers the host peak, as above; otherwise the run needs a larger machine.