* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports. * jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls. * Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results. * Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable. * Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate. * Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do. * Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence. * Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags. * Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check. * Rugnux: Clear error messages when a data set needs more GPU or host memory than is available. Reviewed-on: #83 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
12 KiB
Installing Rugnux
:local:
:depth: 2
rugnux is a single self-contained executable. It needs no CUDA toolkit, no Qt, and no
Jungfraujoch service running anywhere; on a machine with an NVIDIA GPU it needs the NVIDIA
driver, and without one it still runs on the CPU.
From the package repositories (RHEL / Rocky / Ubuntu)
On a distribution covered by the package repositories, rugnux is a package of
its own:
sudo dnf install rugnux # RHEL / Rocky 8 and 9
sudo apt install rugnux # Ubuntu 22.04 / 24.04
It installs /usr/bin/rugnux and depends on nothing from the acquisition side — no broker, no
detector libraries, no Qt — so it can go on a machine that only processes data.
Upgrading from rc.163 or earlier.
/usr/bin/rugnuxused to belong to thejfjoch-viewerpackage. Therugnuxpackage declares that the file has moved, so installing it upgrades an oldjfjoch-viewerin the same transaction instead of failing on the duplicate path. If yourjfjoch-vieweris pinned to an old version, unpin it or remove it first.
From the release archive
For a machine no package manager covers — or for Windows, macOS and Arm, which have no repository — take the archive for your architecture from the Gitea release page:
| Archive | For |
|---|---|
rugnux-<version>-linux-x86_64-cuda12.tgz |
64-bit Intel/AMD Linux. Built on RHEL 8, so it runs on any newer Linux |
rugnux-<version>-linux-aarch64-cuda13.tgz |
64-bit Arm Linux — NVIDIA GH200 and DGX Spark. Built on Ubuntu 24.04, so it needs glibc 2.39 or newer. Cross-compiled and not yet exercised on Arm hardware |
rugnux-<version>-win64-cuda13.zip |
64-bit Windows |
rugnux-<version>-macos-arm64-cpu.tgz |
macOS 13 (Ventura) or newer on Apple Silicon (M1 and newer). CPU only — there is no CUDA on macOS. Intel Macs are not supported |
The archive has no top-level directory — it unpacks straight into bin/ and share/. Always
give tar a destination of its own, or it will scatter those into whatever directory you are in:
mkdir -p /opt/rugnux-1.0.0
tar xzf rugnux-1.0.0-linux-x86_64-cuda12.tgz -C /opt/rugnux-1.0.0
/opt/rugnux-1.0.0/bin/rugnux # prints the usage
What you get is:
bin/rugnux the program
share/doc/jfjoch_rugnux/LICENSE GPLv3
share/doc/jfjoch_rugnux/THIRD_PARTY_NOTICES.md
share/doc/jfjoch_rugnux/licenses/ verbatim licence texts of the bundled dependencies
Nothing is written outside that directory, nothing needs root, and several versions can sit side by
side. To remove it, delete the directory. Put bin/ on your PATH if you want to type rugnux
rather than the full path.
macOS: the first run is blocked. The release is not yet notarized by Apple, and a
.tgzdownloaded with a browser hands its download flag on to everythingtarextracts from it, so macOS refuses to runrugnux("cannot be opened because the developer cannot be verified"). Clear the flag once for the whole directory:xattr -dr com.apple.quarantine ~/rugnux-1.0.0An archive fetched with
curlcarries no such flag. Unpacking into a directory under your home directory (~/rugnux-<version>) rather than/optavoids needingsudoon a Mac.
Mixing the two. If a
rugnuxpackage is also installed,/usr/bin/rugnuxwill normally win onPATH. Put the archive'sbin/first, or call it by its full path, to be sure which one you are running —rugnuxprints its version on every run.
GPU support
The released Linux and Windows archives are CUDA builds (the macOS one is CPU-only). They need only
an NVIDIA driver on the host — 525.60.13
or newer for the CUDA 12 archive, 580.65.06 or newer for the CUDA 13 ones — and no CUDA toolkit,
because everything CUDA is linked statically. With no GPU or no driver, rugnux reports zero CUDA
devices and falls back to the CPU path, which works but is far slower and offers only the fftw
indexer. A V100 needs the CUDA 12 archive; which generations each build covers is in
Release contents ▸ GPU generations and the NVIDIA driver.
Building from source
rugnux alone, without the server stack or Qt:
cmake -S . -B build -DJFJOCH_RUGNUX_ONLY=ON -DCMAKE_BUILD_TYPE=Release \
-DCMAKE_CXX_FLAGS="-march=x86-64-v3" -DCMAKE_C_FLAGS="-march=x86-64-v3"
cmake --build build -j$(nproc) --target rugnux
The binary lands in build/rugnux/rugnux. No library has to come from the system: every dependency
is downloaded and built during the first configure, which therefore needs network access. cmake --build build --target package produces the same .tgz the release ships. The -march flag is not set by the
build system on purpose, so a plain build is slower than the released one on the CPU-bound stages —
see the note in CMakeLists.txt.
On a Mac (Apple Silicon), leave the two -march flags out — they name an x86 level; the Apple
compiler's default target is already the M1 — and use sysctl -n hw.ncpu for the job count. The
build has no CUDA there and needs nothing installed beyond Xcode and CMake.
Hardware
As with the rest of Jungfraujoch, serious performance requires an NVIDIA GPU. The CUDA build
provides the GPU fast-feedback indexer (ffbidx) and the GPU FFT indexer (fft); without CUDA
only the CPU fftw indexer is available. With a GPU present most of the per-image pipeline runs on
the device — bitshuffle+LZ4 decompression, image preprocessing, azimuthal integration, spot finding,
prediction and Bragg integration — as do rotation scaling and merging and the --model rigid-body
placement, with CPU implementations as the fallback where there is no GPU. The choice is automatic;
a CUDA build runs on the CPU when the GPU is hidden from it (CUDA_VISIBLE_DEVICES= rugnux ...).
The thread count (-N) governs the CPU side of all of it.
The released CUDA builds need only an NVIDIA driver on the host, no CUDA toolkit: 525.60.13 or
newer for the CUDA 12 artefacts (RHEL 8 packages, the x86_64 rugnux archive) and 580.65.06 or
newer for the CUDA 13 ones (RHEL 9, Ubuntu, the aarch64 and Windows rugnux archives). Which GPU
generations each artefact supports — a V100 in particular works only with the CUDA 12 build — is in
Release contents ▸ GPU generations and the NVIDIA driver.
Memory
A run's memory is set by how many reflection observations it integrates, not by how many pixels
the detector has. Measured on rc.169 at -N 6, over rotation sweeps of 1200-1800 frames:
| Detector | Observations integrated | Host (peak RSS) | Device |
|---|---|---|---|
| 1679 x 1475 (2.5 MP) | 25 M | 7.8 GB | 3.8 GB |
| 2527 x 2463 (6.2 MP) | 13 M | 4.9 GB | 3.1 GB |
| 3262 x 3108 (10.1 MP) | 7 M | 2.5 GB | 3.0 GB |
| 4371 x 4150 (18.1 MP) | 13 M | 4.3 GB | 3.2 GB |
| 4371 x 4150 (18.1 MP) | 20 M | 6.0 GB | 4.4 GB |
| 3262 x 3108 (10.1 MP) | 50 M | 14.2 GB | 7.0 GB |
Two sweeps on the same detector differ by a factor of five, so size the machine on the observation
count rather than on the detector. As a rule, 0.6 GB + 0.3 GB of host memory per million
observations, and 2.4 GB + 0.1 GB of device memory per million. The run reports what it actually
integrated — RotationScaleMerge: ingested ... partial observations. 32 GB of host RAM and an
8 GB card cover every dataset measured here at -N 6; an ordinary sweep needs 16 GB and 6 GB.
Both peaks fall in scaling and merging, not in the per-image loop: the loop alone stays under
3.2 GB on the device at -N 6, whatever the detector. The stage that needs the most is the 3D combine.
-N multiplies the per-image loop only — about 8 bytes of device memory per detector pixel per
worker (150 MB per worker at 18 MP, 87 MB at 6 MP) on top of a fixed ~2.6 GB, and it is linear
with no ceiling: the same 18 MP sweep takes 2.7 GB of device memory at -N 1, 3.2 GB at -N 6 and
8.0 GB at -N 32. Host memory does not move with -N, and neither does the merge's own device
memory. Where several runs share one card, -N is the knob that decides how many fit.
To use less:
-N— the only flag that lowers the per-image loop's device memory. Lowering it is what rescues a run on a card that is otherwise full.--no-merge— on the heaviest sweep measured it took the host peak from 14.2 to 8.3 GB and the device peak from 7.0 to 2.9 GB. It also gives up the merged output.--no-export-unmergeddoes not save memory.<prefix>_unmerged.mtzis the largest file a run writes, but its observations are resident either way; the flag saves disk, not RAM.
Nothing measures the free memory on the card before allocating. A run that does not fit degrades
where it can — device decompression falls back to the host, the beam-stop projection falls back to
the host, an indexer thread that cannot allocate reports the frame as not indexed, the --model
rigid-body placement moves to the CPU — and fails the run otherwise, with Processing failed: CUDA (GPU) error (Failed to allocate device memory) and a
non-zero exit status. That is deliberate: a CUDA out-of-memory says nothing about the frame in
flight and everything about the machine, so skipping frames would leave a run that looks complete
with quietly different merged numbers. A run is never silently short; it either completes or fails.
On a GPU shared with another job, lower -N or wait for the card.
Very large unit cells
The rule above holds for the cells in that table, and a very large cell sits far beyond it. A long
axis, a dense pattern and many frames multiply together: each frame predicts more reflections, each
reflection is split over more frames, and tens of millions of partial observations are normal for
a big cell. One sweep of a crystal with an axis of about 640 Å predicted 236 million reflections and
peaked at 17.5 GB of host memory on a GPU build and 72.5 GB on a CPU-only build. A CPU-only build
keeps in host RAM what a GPU build keeps on the card, so it is the one that needs the large machine.
When such a set is too large for GPU scaling on the card, the run stops and says to run it on the
CPU (CUDA_VISIBLE_DEVICES= rugnux ..., or a CPU-only build) — which then wants that much host memory.
When host memory runs out
Nothing estimates the host memory a run will need before it starts, so a set that does not fit runs until the memory is gone, and then ends in one of two ways.
Where the allocation is refused — under an address-space limit (ulimit -v, which some batch
systems set), with strict overcommit, or for one request larger than the machine could ever grant —
the run stops with
Processing failed: out of host memory (std::bad_alloc) - this data set needs more RAM than is available
and a non-zero exit status. The CUDA driver reserves a large address space of its own, so a GPU build
under a tight ulimit -v can stop before reading anything, with CUDA (GPU) error (out of memory).
More often on Linux the allocation succeeds and the memory runs out as it is used. The kernel's
OOM killer then ends the process with SIGKILL, which no program can catch: the run stops with no
message of its own, the shell prints Killed (exit status 137), and dmesg or the batch system's
log has an Out of memory: Killed process line. A job capped by a cgroup memory limit ends the same
way. A run that disappears like this needs more RAM than the machine or the job was given, exactly as
if it had said so. --no-merge lowers the host peak, as above; otherwise the run needs a larger
machine.