Files
Jungfraujoch/docs/RUGNUX.md
T
leonarski_fandClaude Opus 5.5 2c1123c267 rugnux: calibration re-bin sums every image; clean out-of-host-memory failures
Calibration (--mode calibration): RebinAndRefit gave each worker one
AzimuthalIntegrationProfile that AzIntEngineCPU::Run clears on every
image, so the re-binned profile held only the last image each worker
read, added in completion order - on a 1800-frame LaB6 run the re-binned
fit changed with -N (197 / 181 / 205 ring points at rms ~7 px). Each
fixed block of images is now summed in image order into its own slot
(ParallelBlocks) and the slots added in block order, so the profile is
the sum over every image and the same at any thread count (rms 0.93 px,
identical at -N 1, 7 and 32). On the LaB6 runs with a sane header the
first pass still wins and the .poni is unchanged; with a header 100 px
off the re-binned fit is now adopted (218 vs 140 ring points). New test
Rugnux_CalibrationRebinsEveryImage spreads the rings over six frames
and fails on the old code (241 vs 141 points at -N 1 vs -N 4).

Host memory: out-of-memory now ends with "Processing failed: out of
host memory (...) - this data set needs more RAM than is available" and
exit 1. Tested with ulimit -v and an LD_PRELOAD allocator that refuses
allocations from a chosen phase (CPU and GPU builds, image loop through
merging and writing, calibration mode). Paths that crashed instead:
- PostIndexingRefinement ran candidate blocks on bare std::threads, so
  a bad_alloc there called std::terminate (seen: "terminate called
  recursively", SIGABRT); now std::async futures.
- FFTW aborts on its own failed allocation (CK(p) in kernel/alloc.c,
  also inside buffered transforms: BeamCenterFFTCPU, FFTIndexerCPU);
  rugnux now defines fftwf_assertion_failed to report it and exit 1.
- libjpeg's default error_exit calls exit() from the diagnostic-JPEG
  thread ("Insufficient memory (case 12)"); WriteJPEGToMem now
  longjmps back and throws (bad_alloc for JERR_OUT_OF_MEMORY).
- WorkerPool construction that fails to start a thread destroyed
  joinable threads (terminate); it now joins them and rethrows.
Also: length_error counts as a fatal resource error, pinned host
allocation failure is MemAllocFailed, and the GPU scaling fail-fast
message says the CPU path needs a lot of host memory.

Docs: very large cells and what running out of host memory looks like
(including the OOM killer) in RUGNUX_INSTALL.md, pointer in RUGNUX.md.

Validation: myob, cytc, lyso, sparse md5-identical to 82e6583e0 on
both GPU and CPU builds.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-28 23:25:54 +02:00

9.1 KiB

Rugnux

rugnux is the offline crystallographic data-analysis tool of Jungfraujoch — the data-processing half of the system (see Naming for where the name comes from). It takes an existing HDF5 dataset, runs the full analysis pipeline — spot finding, indexing, geometry refinement, Bragg integration and (optionally) scaling and merging — and writes the results to a _process.h5 file, plus reflection files (.mtz/.cif/.hkl) when merging is requested.

It runs the same analysis code as the online and interactive tools, just driven from the command line over a file rather than a live detector stream.

rugnux {<options>} <input.h5>

Run it with no arguments to print the usage.

Note. rugnux is under very active development. This page describes the tool and its options at a high level; the authoritative, always-current list of options is the program's own usage message — run rugnux with no arguments.

:local:
:depth: 2

Quick start

Four commands cover most of what people ask of rugnux. Each takes the master file of the dataset — one written by Jungfraujoch, a DECTRIS EIGER master, or an NXmx master written by another facility's toolchain; a PILATUS miniCBF, marCCD or SMV sweep works too (see What Rugnux reads) — and names its output files from -o:

# 1. everything from the data - index, integrate, scale and merge with the defaults
rugnux -o myrun dataset_master.h5

# 2. with a reference dataset of the same crystal form: it fixes the space group and the cell,
#    resolves the indexing ambiguity, and hands over its R-free set
rugnux -o myrun -z reference.mtz dataset_master.h5

# 3. with a known structure: R-work / R-free and sigma_A-weighted 2mFo-DFc / mFo-DFc maps
rugnux -o myrun --model model.pdb dataset_master.h5

# 4. with the space group and the cell pinned (-S takes either spelling: P43212 or 96)
rugnux -o myrun -S P43212 -C 79,79,38,90,90,90 dataset_master.h5

Parallelism needs no asking for: a run already uses the machine's threads. -N is there to limit that, or to lift the per-image loop's default ceiling of 16 workers per GPU.

Nothing more is needed to pick the workflow: a dataset carrying a goniometer axis is processed as a rotation sweep, one without as independent stills, and scaling and merging run by default in both. A rotation run that merges — the default — leaves seven files next to each other:

myrun.mtz            merged intensities + French-Wilson amplitudes, for CCP4 / phenix
myrun.cif            the same, as mmCIF - the self-describing format, and what to deposit
myrun.hkl            the same, as SHELX HKLF 4 - feed this to SHELXC / SHELXD / ANODE
myrun_unmerged.mtz   every observation before scaling, for pointless / aimless / careless -
                     the largest file of the run (--no-export-unmerged skips it)
myrun_P1.mtz         the same observations merged in P1, so a wrong space-group call can be
                     re-merged or re-refined without reprocessing (--no-p1-crosscheck skips it)
myrun_report.txt     what the run determined: cell, space group, statistics, warnings
myrun_plot.txt       one row per image, for plotting how the crystal behaved over the sweep

The two MTZ extras are most of the bytes a run writes — worth knowing when sizing a scratch directory for a campaign, and both have off switches.

Read myrun_report.txt first: it says which space group was chosen and on what evidence, how far the data go, and anything that needs attention.

A few things worth knowing before reaching for more flags:

  • The written reflections stop where CC1/2 falls through 0.30. Every reflection file is resolution-trimmed automatically (--resolution-cutoff cc-logistic, one shell past the crossing); --resolution-cutoff off keeps the full measured range, --scaling-high-resolution fixes the limit by hand. A Rugnux file reaching less far than another program's on the same data is usually this default at work, not lost data.
  • Rotation data are best left de novo. Pinning the cell and space group (recipe 4) is the normal thing to do for serial stills, where the ffbidx indexer needs a cell; on a rotation sweep it tends to degrade low-symmetry cases, so prefer recipe 1 and let the run determine both (see Rotation data). -S takes a Hermann-Mauguin symbol (P43212) or a space-group number (96), whichever is to hand.
  • Anomalous data are there without -A. A rotation merge always keeps the Bijvoet split: a default run's .mtz carries I(+)/I(-) beside IMEAN, and its .hkl the ±hkl mates — FRIEDELS_LAW= TRUE in the report says how the statistics were counted, not that the signal was averaged away. What -A changes is the counting basis and the error model: each hand becomes a merged observation of its own, so multiplicity, completeness and ⟨I/σ⟩ are counted anomalously and the sigmas are refitted on the Bijvoet-separated merge. Reach for it for anomalous statistics; the signal itself is in the file either way. (Stills merges carry no split by default — there -A is what creates one.)
  • A model names the enantiomorph. Where the data accept the model — it is tested against a null of the same model in random orientations, and MODEL_FIT= in the report says the verdict — --model settles which of P41212 and P43212 the merged reflections are labelled with — a choice no merged intensity can make. It is a label and nothing more: the two groups have the same rotation operations, so no reflection moves, and in particular I(+) and I(-) are left exactly as measured. Whether the model agrees with the data about the hand is then a real question, and the anomalous difference map answers it — a run says so when the density at the model's atoms comes out inverted.
  • -z and --model overlap but are not the same. A reference MTZ steers the processing from the start; a model scores the merge and settles the frame it is written in — where the data accept it; a model they reject changes nothing. Either resolves an indexing ambiguity, which on serial data decides whether the merged intensities are usable at all.
  • --scaling-high-resolution <d>, where the resolution is already known, sharpens both the space-group search and the error model.
  • A run wants memory in proportion to what it integrates, not to the detector: 2.5-14 GB of host RAM and 3-7 GB on the card over the datasets measured, both peaking in scaling and merging. Installing Rugnux ▸ Memory has the table and the two flags that lower it. A very large cell needs far more, and a CPU-only build most of all — tens of GB of host RAM (Very large unit cells).
  • Everything else is in Running Rugnux and the full Command-line options.

The rest of the manual

One page per job, so the answer needed is near the top of a short page:

  • What Rugnux does — the pipeline from images to merged reflections, in order. Read this one first.
  • Installing Rugnux — packages, the release archive, GPU drivers, building from source, hardware.
  • What Rugnux reads — will it open your data: NXmx / EIGER masters, PILATUS miniCBF, marCCD and SMV sweeps, one sweep per input.
  • Running Rugnux — a first run in detail, rotation and serial data, and every file a run writes.
  • Rugnux with other programs — the reflection-file conventions, the unmerged export, and worked command lines for phenix, REFMAC, POINTLESS / AIMLESS, careless and Phaser.
  • The results report — the KEY= value interface, sweep quality and the anisotropy section.
  • Advanced usage — reference data and the indexing ambiguity, model validation, re-merging, and the full command-line option tables.
  • Detector calibration — the geometry from a calibrant's powder rings (--mode calibration).
  • CPU/GPU data analysis — the algorithms behind all of it.

Where it fits among the three analysis tools

Tool Mode Driven by Output
jfjoch_broker Online, real-time streaming analysis on FPGA + GPU HTTP/REST + ZeroMQ Live results and statistics, images streamed to jfjoch_writer
jfjoch_viewer Interactive, on-screen exploration Qt desktop application On screen; a processing job can write the same files as rugnux
rugnux Offline batch processing of a stored dataset Command-line interface _process.h5, and .mtz/.cif/.hkl when merging

Use rugnux to re-analyse data after acquisition, to experiment with processing parameters, or to produce merged intensities for downstream structure solution.