Build Packages / Create release (push) Successful in 24s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m29s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m43s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 8m27s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 9m53s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 9m58s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 11m22s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 13m39s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 18m37s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 16m32s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 24m11s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m30s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 19m3s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 20m23s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 19m41s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 21m0s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m38s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 14m33s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 17m55s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 20m50s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 18m38s
Build Packages / Unit tests (push) Successful in 1h46m14s
* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports. * jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls. * Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results. * Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable. * Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate. * Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do. * Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence. * Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags. * Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check. * Rugnux: Clear error messages when a data set needs more GPU or host memory than is available. Reviewed-on: #83 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
209 lines
12 KiB
Markdown
209 lines
12 KiB
Markdown
# Installing Rugnux
|
|
|
|
```{contents} On this page
|
|
:local:
|
|
:depth: 2
|
|
```
|
|
|
|
|
|
`rugnux` is a **single self-contained executable**. It needs no CUDA toolkit, no Qt, and no
|
|
Jungfraujoch service running anywhere; on a machine with an NVIDIA GPU it needs the NVIDIA
|
|
**driver**, and without one it still runs on the CPU.
|
|
|
|
## From the package repositories (RHEL / Rocky / Ubuntu)
|
|
|
|
On a distribution covered by the [package repositories](REPOSITORIES.md), `rugnux` is a package of
|
|
its own:
|
|
|
|
```
|
|
sudo dnf install rugnux # RHEL / Rocky 8 and 9
|
|
sudo apt install rugnux # Ubuntu 22.04 / 24.04
|
|
```
|
|
|
|
It installs `/usr/bin/rugnux` and depends on nothing from the acquisition side — no broker, no
|
|
detector libraries, no Qt — so it can go on a machine that only processes data.
|
|
|
|
> **Upgrading from rc.163 or earlier.** `/usr/bin/rugnux` used to belong to the `jfjoch-viewer`
|
|
> package. The `rugnux` package declares that the file has moved, so installing it upgrades an old
|
|
> `jfjoch-viewer` in the same transaction instead of failing on the duplicate path. If your
|
|
> `jfjoch-viewer` is pinned to an old version, unpin it or remove it first.
|
|
|
|
## From the release archive
|
|
|
|
For a machine no package manager covers — or for Windows, macOS and Arm, which have no repository —
|
|
take the archive for your architecture from the Gitea release page:
|
|
|
|
| Archive | For |
|
|
| --- | --- |
|
|
| `rugnux-<version>-linux-x86_64-cuda12.tgz` | 64-bit Intel/AMD Linux. Built on RHEL 8, so it runs on any newer Linux |
|
|
| `rugnux-<version>-linux-aarch64-cuda13.tgz` | 64-bit Arm Linux — NVIDIA GH200 and DGX Spark. Built on Ubuntu 24.04, so it needs glibc 2.39 or newer. Cross-compiled and **not yet exercised on Arm hardware** |
|
|
| `rugnux-<version>-win64-cuda13.zip` | 64-bit Windows |
|
|
| `rugnux-<version>-macos-arm64-cpu.tgz` | macOS 13 (Ventura) or newer on Apple Silicon (M1 and newer). CPU only — there is no CUDA on macOS. Intel Macs are not supported |
|
|
|
|
**The archive has no top-level directory** — it unpacks straight into `bin/` and `share/`. Always
|
|
give `tar` a destination of its own, or it will scatter those into whatever directory you are in:
|
|
|
|
```
|
|
mkdir -p /opt/rugnux-1.0.0
|
|
tar xzf rugnux-1.0.0-linux-x86_64-cuda12.tgz -C /opt/rugnux-1.0.0
|
|
/opt/rugnux-1.0.0/bin/rugnux # prints the usage
|
|
```
|
|
|
|
What you get is:
|
|
|
|
```
|
|
bin/rugnux the program
|
|
share/doc/jfjoch_rugnux/LICENSE GPLv3
|
|
share/doc/jfjoch_rugnux/THIRD_PARTY_NOTICES.md
|
|
share/doc/jfjoch_rugnux/licenses/ verbatim licence texts of the bundled dependencies
|
|
```
|
|
|
|
Nothing is written outside that directory, nothing needs root, and several versions can sit side by
|
|
side. To remove it, delete the directory. Put `bin/` on your `PATH` if you want to type `rugnux`
|
|
rather than the full path.
|
|
|
|
> **macOS: the first run is blocked.** The release is not yet notarized by Apple, and a `.tgz`
|
|
> downloaded with a browser hands its download flag on to everything `tar` extracts from it, so
|
|
> macOS refuses to run `rugnux` ("cannot be opened because the developer cannot be verified"). Clear
|
|
> the flag once for the whole directory:
|
|
>
|
|
> ```
|
|
> xattr -dr com.apple.quarantine ~/rugnux-1.0.0
|
|
> ```
|
|
>
|
|
> An archive fetched with `curl` carries no such flag. Unpacking into a directory under your home
|
|
> directory (`~/rugnux-<version>`) rather than `/opt` avoids needing `sudo` on a Mac.
|
|
|
|
> **Mixing the two.** If a `rugnux` package is also installed, `/usr/bin/rugnux` will normally win
|
|
> on `PATH`. Put the archive's `bin/` first, or call it by its full path, to be sure which one you
|
|
> are running — `rugnux` prints its version on every run.
|
|
|
|
## GPU support
|
|
|
|
The released Linux and Windows archives are CUDA builds (the macOS one is CPU-only). They need only
|
|
an NVIDIA **driver** on the host — 525.60.13
|
|
or newer for the CUDA 12 archive, 580.65.06 or newer for the CUDA 13 ones — and no CUDA toolkit,
|
|
because everything CUDA is linked statically. With no GPU or no driver, `rugnux` reports zero CUDA
|
|
devices and falls back to the CPU path, which works but is far slower and offers only the `fftw`
|
|
indexer. A **V100 needs the CUDA 12 archive**; which generations each build covers is in
|
|
[Release contents ▸ GPU generations and the NVIDIA driver](RELEASE_CONTENTS.md#gpu-generations-and-the-nvidia-driver).
|
|
|
|
## Building from source
|
|
|
|
`rugnux` alone, without the server stack or Qt:
|
|
|
|
```
|
|
cmake -S . -B build -DJFJOCH_RUGNUX_ONLY=ON -DCMAKE_BUILD_TYPE=Release \
|
|
-DCMAKE_CXX_FLAGS="-march=x86-64-v3" -DCMAKE_C_FLAGS="-march=x86-64-v3"
|
|
cmake --build build -j$(nproc) --target rugnux
|
|
```
|
|
|
|
The binary lands in `build/rugnux/rugnux`. No library has to come from the system: every dependency
|
|
is downloaded and built during the first configure, which therefore needs network access. `cmake --build build
|
|
--target package` produces the same `.tgz` the release ships. The `-march` flag is not set by the
|
|
build system on purpose, so a plain build is slower than the released one on the CPU-bound stages —
|
|
see the note in `CMakeLists.txt`.
|
|
|
|
On a Mac (Apple Silicon), leave the two `-march` flags out — they name an x86 level; the Apple
|
|
compiler's default target is already the M1 — and use `sysctl -n hw.ncpu` for the job count. The
|
|
build has no CUDA there and needs nothing installed beyond Xcode and CMake.
|
|
|
|
## Hardware
|
|
|
|
As with the rest of Jungfraujoch, **serious performance requires an NVIDIA GPU**. The CUDA build
|
|
provides the GPU fast-feedback indexer (`ffbidx`) and the GPU FFT indexer (`fft`); without CUDA
|
|
only the CPU `fftw` indexer is available. With a GPU present most of the per-image pipeline runs on
|
|
the device — bitshuffle+LZ4 decompression, image preprocessing, azimuthal integration, spot finding,
|
|
prediction and Bragg integration — as do rotation scaling and merging and the `--model` rigid-body
|
|
placement, with CPU implementations as the fallback where there is no GPU. The choice is automatic;
|
|
a CUDA build runs on the CPU when the GPU is hidden from it (`CUDA_VISIBLE_DEVICES= rugnux ...`).
|
|
The thread count (`-N`) governs the CPU side of all of it.
|
|
|
|
The released CUDA builds need only an NVIDIA **driver** on the host, no CUDA toolkit: 525.60.13 or
|
|
newer for the CUDA 12 artefacts (RHEL 8 packages, the x86_64 `rugnux` archive) and 580.65.06 or
|
|
newer for the CUDA 13 ones (RHEL 9, Ubuntu, the aarch64 and Windows `rugnux` archives). Which GPU
|
|
generations each artefact supports — a V100 in particular works only with the CUDA 12 build — is in
|
|
[Release contents ▸ GPU generations and the NVIDIA driver](RELEASE_CONTENTS.md#gpu-generations-and-the-nvidia-driver).
|
|
|
|
### Memory
|
|
|
|
A run's memory is set by **how many reflection observations it integrates**, not by how many pixels
|
|
the detector has. Measured on rc.169 at `-N 6`, over rotation sweeps of 1200-1800 frames:
|
|
|
|
| Detector | Observations integrated | Host (peak RSS) | Device |
|
|
| --- | --- | --- | --- |
|
|
| 1679 x 1475 (2.5 MP) | 25 M | 7.8 GB | 3.8 GB |
|
|
| 2527 x 2463 (6.2 MP) | 13 M | 4.9 GB | 3.1 GB |
|
|
| 3262 x 3108 (10.1 MP) | 7 M | 2.5 GB | 3.0 GB |
|
|
| 4371 x 4150 (18.1 MP) | 13 M | 4.3 GB | 3.2 GB |
|
|
| 4371 x 4150 (18.1 MP) | 20 M | 6.0 GB | 4.4 GB |
|
|
| 3262 x 3108 (10.1 MP) | 50 M | 14.2 GB | 7.0 GB |
|
|
|
|
Two sweeps on the same detector differ by a factor of five, so size the machine on the observation
|
|
count rather than on the detector. As a rule, **0.6 GB + 0.3 GB of host memory per million
|
|
observations, and 2.4 GB + 0.1 GB of device memory per million**. The run reports what it actually
|
|
integrated — `RotationScaleMerge: ingested ... partial observations`. **32 GB of host RAM and an
|
|
8 GB card** cover every dataset measured here at `-N 6`; an ordinary sweep needs 16 GB and 6 GB.
|
|
|
|
Both peaks fall in **scaling and merging**, not in the per-image loop: the loop alone stays under
|
|
3.2 GB on the device at `-N 6`, whatever the detector. The stage that needs the most is the 3D combine.
|
|
|
|
`-N` multiplies the per-image loop only — about **8 bytes of device memory per detector pixel per
|
|
worker** (150 MB per worker at 18 MP, 87 MB at 6 MP) on top of a fixed ~2.6 GB, and it is linear
|
|
with no ceiling: the same 18 MP sweep takes 2.7 GB of device memory at `-N 1`, 3.2 GB at `-N 6` and
|
|
8.0 GB at `-N 32`. Host memory does not move with `-N`, and neither does the merge's own device
|
|
memory. Where several runs share one card, `-N` is the knob that decides how many fit.
|
|
|
|
To use less:
|
|
|
|
- **`-N`** — the only flag that lowers the per-image loop's device memory. Lowering it is what
|
|
rescues a run on a card that is otherwise full.
|
|
- **`--no-merge`** — on the heaviest sweep measured it took the host peak from 14.2 to 8.3 GB and
|
|
the device peak from 7.0 to 2.9 GB. It also gives up the merged output.
|
|
- **`--no-export-unmerged`** does **not** save memory. `<prefix>_unmerged.mtz` is the largest file a
|
|
run writes, but its observations are resident either way; the flag saves disk, not RAM.
|
|
|
|
Nothing measures the free memory on the card before allocating. A run that does not fit degrades
|
|
where it can — device decompression falls back to the host, the beam-stop projection falls back to
|
|
the host, an indexer thread that cannot allocate reports the frame as not indexed, the `--model`
|
|
rigid-body placement moves to the CPU — and **fails the run otherwise**, with `Processing failed: CUDA (GPU) error (Failed to allocate device memory)` and a
|
|
non-zero exit status. That is deliberate: a CUDA out-of-memory says nothing about the frame in
|
|
flight and everything about the machine, so skipping frames would leave a run that looks complete
|
|
with quietly different merged numbers. A run is never silently short; it either completes or fails.
|
|
On a GPU shared with another job, lower `-N` or wait for the card.
|
|
|
|
### Very large unit cells
|
|
|
|
The rule above holds for the cells in that table, and a very large cell sits far beyond it. A long
|
|
axis, a dense pattern and many frames multiply together: each frame predicts more reflections, each
|
|
reflection is split over more frames, and **tens of millions of partial observations are normal** for
|
|
a big cell. One sweep of a crystal with an axis of about 640 Å predicted 236 million reflections and
|
|
peaked at **17.5 GB of host memory on a GPU build and 72.5 GB on a CPU-only build**. A CPU-only build
|
|
keeps in host RAM what a GPU build keeps on the card, so it is the one that needs the large machine.
|
|
When such a set is too large for GPU scaling on the card, the run stops and says to run it on the
|
|
CPU (`CUDA_VISIBLE_DEVICES= rugnux ...`, or a CPU-only build) — which then wants that much host memory.
|
|
|
|
### When host memory runs out
|
|
|
|
Nothing estimates the host memory a run will need before it starts, so a set that does not fit runs
|
|
until the memory is gone, and then ends in one of two ways.
|
|
|
|
Where the allocation is refused — under an address-space limit (`ulimit -v`, which some batch
|
|
systems set), with strict overcommit, or for one request larger than the machine could ever grant —
|
|
the run stops with
|
|
|
|
```
|
|
Processing failed: out of host memory (std::bad_alloc) - this data set needs more RAM than is available
|
|
```
|
|
|
|
and a non-zero exit status. The CUDA driver reserves a large address space of its own, so a GPU build
|
|
under a tight `ulimit -v` can stop before reading anything, with `CUDA (GPU) error (out of memory)`.
|
|
|
|
More often on Linux the allocation succeeds and the memory runs out as it is used. The kernel's
|
|
**OOM killer** then ends the process with SIGKILL, which no program can catch: the run stops with no
|
|
message of its own, the shell prints `Killed` (exit status 137), and `dmesg` or the batch system's
|
|
log has an `Out of memory: Killed process` line. A job capped by a cgroup memory limit ends the same
|
|
way. A run that disappears like this needs more RAM than the machine or the job was given, exactly as
|
|
if it had said so. `--no-merge` lowers the host peak, as above; otherwise the run needs a larger
|
|
machine.
|