Files
Jungfraujoch/docs/RUGNUX_INSTALL.md
T
leonarski_f 84228bf8be
Build Packages / Create release (push) Successful in 24s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m29s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m43s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 8m27s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 9m53s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 9m58s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 11m22s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 13m39s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 18m37s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 16m32s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 24m11s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m30s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 19m3s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 20m23s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 19m41s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 21m0s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m38s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 14m33s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 17m55s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 20m50s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 18m38s
Build Packages / Unit tests (push) Successful in 1h46m14s
v1.0.0-rc.173 (#83)
* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports.
* jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls.
* Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results.
* Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable.
* Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate.
* Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do.
* Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence.
* Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags.
* Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check.
* Rugnux: Clear error messages when a data set needs more GPU or host memory than is available.

Reviewed-on: #83
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-09-29 15:57:32 +02:00

209 lines
12 KiB
Markdown

# Installing Rugnux
```{contents} On this page
:local:
:depth: 2
```
`rugnux` is a **single self-contained executable**. It needs no CUDA toolkit, no Qt, and no
Jungfraujoch service running anywhere; on a machine with an NVIDIA GPU it needs the NVIDIA
**driver**, and without one it still runs on the CPU.
## From the package repositories (RHEL / Rocky / Ubuntu)
On a distribution covered by the [package repositories](REPOSITORIES.md), `rugnux` is a package of
its own:
```
sudo dnf install rugnux # RHEL / Rocky 8 and 9
sudo apt install rugnux # Ubuntu 22.04 / 24.04
```
It installs `/usr/bin/rugnux` and depends on nothing from the acquisition side — no broker, no
detector libraries, no Qt — so it can go on a machine that only processes data.
> **Upgrading from rc.163 or earlier.** `/usr/bin/rugnux` used to belong to the `jfjoch-viewer`
> package. The `rugnux` package declares that the file has moved, so installing it upgrades an old
> `jfjoch-viewer` in the same transaction instead of failing on the duplicate path. If your
> `jfjoch-viewer` is pinned to an old version, unpin it or remove it first.
## From the release archive
For a machine no package manager covers — or for Windows, macOS and Arm, which have no repository —
take the archive for your architecture from the Gitea release page:
| Archive | For |
| --- | --- |
| `rugnux-<version>-linux-x86_64-cuda12.tgz` | 64-bit Intel/AMD Linux. Built on RHEL 8, so it runs on any newer Linux |
| `rugnux-<version>-linux-aarch64-cuda13.tgz` | 64-bit Arm Linux — NVIDIA GH200 and DGX Spark. Built on Ubuntu 24.04, so it needs glibc 2.39 or newer. Cross-compiled and **not yet exercised on Arm hardware** |
| `rugnux-<version>-win64-cuda13.zip` | 64-bit Windows |
| `rugnux-<version>-macos-arm64-cpu.tgz` | macOS 13 (Ventura) or newer on Apple Silicon (M1 and newer). CPU only — there is no CUDA on macOS. Intel Macs are not supported |
**The archive has no top-level directory** — it unpacks straight into `bin/` and `share/`. Always
give `tar` a destination of its own, or it will scatter those into whatever directory you are in:
```
mkdir -p /opt/rugnux-1.0.0
tar xzf rugnux-1.0.0-linux-x86_64-cuda12.tgz -C /opt/rugnux-1.0.0
/opt/rugnux-1.0.0/bin/rugnux # prints the usage
```
What you get is:
```
bin/rugnux the program
share/doc/jfjoch_rugnux/LICENSE GPLv3
share/doc/jfjoch_rugnux/THIRD_PARTY_NOTICES.md
share/doc/jfjoch_rugnux/licenses/ verbatim licence texts of the bundled dependencies
```
Nothing is written outside that directory, nothing needs root, and several versions can sit side by
side. To remove it, delete the directory. Put `bin/` on your `PATH` if you want to type `rugnux`
rather than the full path.
> **macOS: the first run is blocked.** The release is not yet notarized by Apple, and a `.tgz`
> downloaded with a browser hands its download flag on to everything `tar` extracts from it, so
> macOS refuses to run `rugnux` ("cannot be opened because the developer cannot be verified"). Clear
> the flag once for the whole directory:
>
> ```
> xattr -dr com.apple.quarantine ~/rugnux-1.0.0
> ```
>
> An archive fetched with `curl` carries no such flag. Unpacking into a directory under your home
> directory (`~/rugnux-<version>`) rather than `/opt` avoids needing `sudo` on a Mac.
> **Mixing the two.** If a `rugnux` package is also installed, `/usr/bin/rugnux` will normally win
> on `PATH`. Put the archive's `bin/` first, or call it by its full path, to be sure which one you
> are running — `rugnux` prints its version on every run.
## GPU support
The released Linux and Windows archives are CUDA builds (the macOS one is CPU-only). They need only
an NVIDIA **driver** on the host — 525.60.13
or newer for the CUDA 12 archive, 580.65.06 or newer for the CUDA 13 ones — and no CUDA toolkit,
because everything CUDA is linked statically. With no GPU or no driver, `rugnux` reports zero CUDA
devices and falls back to the CPU path, which works but is far slower and offers only the `fftw`
indexer. A **V100 needs the CUDA 12 archive**; which generations each build covers is in
[Release contents ▸ GPU generations and the NVIDIA driver](RELEASE_CONTENTS.md#gpu-generations-and-the-nvidia-driver).
## Building from source
`rugnux` alone, without the server stack or Qt:
```
cmake -S . -B build -DJFJOCH_RUGNUX_ONLY=ON -DCMAKE_BUILD_TYPE=Release \
-DCMAKE_CXX_FLAGS="-march=x86-64-v3" -DCMAKE_C_FLAGS="-march=x86-64-v3"
cmake --build build -j$(nproc) --target rugnux
```
The binary lands in `build/rugnux/rugnux`. No library has to come from the system: every dependency
is downloaded and built during the first configure, which therefore needs network access. `cmake --build build
--target package` produces the same `.tgz` the release ships. The `-march` flag is not set by the
build system on purpose, so a plain build is slower than the released one on the CPU-bound stages —
see the note in `CMakeLists.txt`.
On a Mac (Apple Silicon), leave the two `-march` flags out — they name an x86 level; the Apple
compiler's default target is already the M1 — and use `sysctl -n hw.ncpu` for the job count. The
build has no CUDA there and needs nothing installed beyond Xcode and CMake.
## Hardware
As with the rest of Jungfraujoch, **serious performance requires an NVIDIA GPU**. The CUDA build
provides the GPU fast-feedback indexer (`ffbidx`) and the GPU FFT indexer (`fft`); without CUDA
only the CPU `fftw` indexer is available. With a GPU present most of the per-image pipeline runs on
the device — bitshuffle+LZ4 decompression, image preprocessing, azimuthal integration, spot finding,
prediction and Bragg integration — as do rotation scaling and merging and the `--model` rigid-body
placement, with CPU implementations as the fallback where there is no GPU. The choice is automatic;
a CUDA build runs on the CPU when the GPU is hidden from it (`CUDA_VISIBLE_DEVICES= rugnux ...`).
The thread count (`-N`) governs the CPU side of all of it.
The released CUDA builds need only an NVIDIA **driver** on the host, no CUDA toolkit: 525.60.13 or
newer for the CUDA 12 artefacts (RHEL 8 packages, the x86_64 `rugnux` archive) and 580.65.06 or
newer for the CUDA 13 ones (RHEL 9, Ubuntu, the aarch64 and Windows `rugnux` archives). Which GPU
generations each artefact supports — a V100 in particular works only with the CUDA 12 build — is in
[Release contents ▸ GPU generations and the NVIDIA driver](RELEASE_CONTENTS.md#gpu-generations-and-the-nvidia-driver).
### Memory
A run's memory is set by **how many reflection observations it integrates**, not by how many pixels
the detector has. Measured on rc.169 at `-N 6`, over rotation sweeps of 1200-1800 frames:
| Detector | Observations integrated | Host (peak RSS) | Device |
| --- | --- | --- | --- |
| 1679 x 1475 (2.5 MP) | 25 M | 7.8 GB | 3.8 GB |
| 2527 x 2463 (6.2 MP) | 13 M | 4.9 GB | 3.1 GB |
| 3262 x 3108 (10.1 MP) | 7 M | 2.5 GB | 3.0 GB |
| 4371 x 4150 (18.1 MP) | 13 M | 4.3 GB | 3.2 GB |
| 4371 x 4150 (18.1 MP) | 20 M | 6.0 GB | 4.4 GB |
| 3262 x 3108 (10.1 MP) | 50 M | 14.2 GB | 7.0 GB |
Two sweeps on the same detector differ by a factor of five, so size the machine on the observation
count rather than on the detector. As a rule, **0.6 GB + 0.3 GB of host memory per million
observations, and 2.4 GB + 0.1 GB of device memory per million**. The run reports what it actually
integrated — `RotationScaleMerge: ingested ... partial observations`. **32 GB of host RAM and an
8 GB card** cover every dataset measured here at `-N 6`; an ordinary sweep needs 16 GB and 6 GB.
Both peaks fall in **scaling and merging**, not in the per-image loop: the loop alone stays under
3.2 GB on the device at `-N 6`, whatever the detector. The stage that needs the most is the 3D combine.
`-N` multiplies the per-image loop only — about **8 bytes of device memory per detector pixel per
worker** (150 MB per worker at 18 MP, 87 MB at 6 MP) on top of a fixed ~2.6 GB, and it is linear
with no ceiling: the same 18 MP sweep takes 2.7 GB of device memory at `-N 1`, 3.2 GB at `-N 6` and
8.0 GB at `-N 32`. Host memory does not move with `-N`, and neither does the merge's own device
memory. Where several runs share one card, `-N` is the knob that decides how many fit.
To use less:
- **`-N`** — the only flag that lowers the per-image loop's device memory. Lowering it is what
rescues a run on a card that is otherwise full.
- **`--no-merge`** — on the heaviest sweep measured it took the host peak from 14.2 to 8.3 GB and
the device peak from 7.0 to 2.9 GB. It also gives up the merged output.
- **`--no-export-unmerged`** does **not** save memory. `<prefix>_unmerged.mtz` is the largest file a
run writes, but its observations are resident either way; the flag saves disk, not RAM.
Nothing measures the free memory on the card before allocating. A run that does not fit degrades
where it can — device decompression falls back to the host, the beam-stop projection falls back to
the host, an indexer thread that cannot allocate reports the frame as not indexed, the `--model`
rigid-body placement moves to the CPU — and **fails the run otherwise**, with `Processing failed: CUDA (GPU) error (Failed to allocate device memory)` and a
non-zero exit status. That is deliberate: a CUDA out-of-memory says nothing about the frame in
flight and everything about the machine, so skipping frames would leave a run that looks complete
with quietly different merged numbers. A run is never silently short; it either completes or fails.
On a GPU shared with another job, lower `-N` or wait for the card.
### Very large unit cells
The rule above holds for the cells in that table, and a very large cell sits far beyond it. A long
axis, a dense pattern and many frames multiply together: each frame predicts more reflections, each
reflection is split over more frames, and **tens of millions of partial observations are normal** for
a big cell. One sweep of a crystal with an axis of about 640 Å predicted 236 million reflections and
peaked at **17.5 GB of host memory on a GPU build and 72.5 GB on a CPU-only build**. A CPU-only build
keeps in host RAM what a GPU build keeps on the card, so it is the one that needs the large machine.
When such a set is too large for GPU scaling on the card, the run stops and says to run it on the
CPU (`CUDA_VISIBLE_DEVICES= rugnux ...`, or a CPU-only build) — which then wants that much host memory.
### When host memory runs out
Nothing estimates the host memory a run will need before it starts, so a set that does not fit runs
until the memory is gone, and then ends in one of two ways.
Where the allocation is refused — under an address-space limit (`ulimit -v`, which some batch
systems set), with strict overcommit, or for one request larger than the machine could ever grant —
the run stops with
```
Processing failed: out of host memory (std::bad_alloc) - this data set needs more RAM than is available
```
and a non-zero exit status. The CUDA driver reserves a large address space of its own, so a GPU build
under a tight `ulimit -v` can stop before reading anything, with `CUDA (GPU) error (out of memory)`.
More often on Linux the allocation succeeds and the memory runs out as it is used. The kernel's
**OOM killer** then ends the process with SIGKILL, which no program can catch: the run stops with no
message of its own, the shell prints `Killed` (exit status 137), and `dmesg` or the batch system's
log has an `Out of memory: Killed process` line. A job capped by a cgroup memory limit ends the same
way. A run that disappears like this needs more RAM than the machine or the job was given, exactly as
if it had said so. `--no-merge` lowers the host peak, as above; otherwise the run needs a larger
machine.