# Rugnux `rugnux` is the **offline** crystallographic data-analysis tool of Jungfraujoch — the data-processing half of the system (see [Naming](NAMING.md) for where the name comes from). It takes an existing HDF5 dataset, runs the full analysis pipeline — spot finding, indexing, geometry refinement, Bragg integration and (optionally) scaling and merging — and writes the results to a `_process.h5` file, plus reflection files (`.mtz`/`.cif`/`.hkl`) when merging is requested. It runs the *same* analysis code as the online and interactive tools, just driven from the command line over a file rather than a live detector stream. ``` rugnux {} ``` Run it with no arguments to print the usage. > **Note.** `rugnux` is under very active development. This page describes the tool and > its options at a high level; the authoritative, always-current list of options is the program's > own usage message — run `rugnux` with no arguments. ```{contents} On this page :local: :depth: 2 ``` ## Quick start Four commands cover most of what people ask of `rugnux`. Each takes the **master** file of the dataset — one written by Jungfraujoch, a DECTRIS EIGER master, or an NXmx master written by another facility's toolchain; a PILATUS miniCBF, marCCD or SMV sweep works too (see [What Rugnux reads](RUGNUX_FORMATS.md)) — and names its output files from `-o`: ``` # 1. everything from the data - index, integrate, scale and merge with the defaults rugnux -o myrun dataset_master.h5 # 2. with a reference dataset of the same crystal form: it fixes the space group and the cell, # resolves the indexing ambiguity, and hands over its R-free set rugnux -o myrun -z reference.mtz dataset_master.h5 # 3. with a known structure: R-work / R-free and sigma_A-weighted 2mFo-DFc / mFo-DFc maps rugnux -o myrun --model model.pdb dataset_master.h5 # 4. with the space group and the cell pinned (-S takes either spelling: P43212 or 96) rugnux -o myrun -S P43212 -C 79,79,38,90,90,90 dataset_master.h5 ``` Parallelism needs no asking for: a run already uses the machine's threads. `-N` is there to *limit* that, or to lift the per-image loop's default ceiling of 16 workers per GPU. Nothing more is needed to pick the workflow: a dataset carrying a **goniometer axis** is processed as a rotation sweep, one without as **independent stills**, and scaling and merging run by default in both. A rotation run that merges — the default — leaves nine files next to each other: ``` myrun.mtz merged intensities + French-Wilson amplitudes, for CCP4 / phenix myrun.cif the same, as mmCIF - the self-describing format, and what to deposit myrun.hkl the scaled observations unmerged, as SHELX HKLF 4 - for SHELXL, SHELXC / ANODE myrun_unmerged.mtz every observation before scaling, for pointless / aimless / careless - the largest file of the run (--no-export-unmerged skips it) myrun_P1.mtz the same observations merged in P1, so a wrong space-group call can be re-merged or re-refined without reprocessing (--no-p1-crosscheck skips it) myrun_report.txt what the run determined: cell, space group, statistics, warnings myrun_plot.txt one row per image, for plotting how the crystal behaved over the sweep myrun_detector.jpg the detector with the pixel mask and the beam-stop shadow drawn on it myrun_log.txt the full timestamped log of the run (debug level) ``` The two MTZ extras are most of the bytes a run writes — worth knowing when sizing a scratch directory for a campaign, and both have off switches. Read `myrun_report.txt` first: it says which space group was chosen and on what evidence, how far the data go, and anything that needs attention. ### What the terminal shows The terminal shows the run as numbered steps, each with one-line results and the time it took, the warnings among them, and a summary at the end; everything else goes to `myrun_log.txt`. A note's time is how long the step ran since the line above it, and the rule under a step gives its total, so the steps add up to the wall time. Abridged: ``` Step 3 Indexing ... beam centre kept: measured 7.5 px off, same lattice 1.36 s ... lattice triclinic P: 28.56 35.34 63.56 106.0 90.3 90.3 0.00 s ... 52/60 validation frames, 24% of their spots on it 0.00 s ------------------------------------------------------- done in 1.36 s Step 4 Integration (geometry pre-pass), 1800 images ... 1800 images, 836 per second 2.15 s ------------------------------------------------------- done in 2.16 s ... Step 9 Scaling and merging ... merge in P1: 14464 unique 0.64 s ... space group P 1 21 1 (No. 4) 0.05 s ... merge in P21: 16847 unique 0.26 s ... resolution 1.57 A (CC1/2 fall-off at 1.64 A) 0.02 s ------------------------------------------------------- done in 1.89 s ... ============================================================================== Result WARNINGS: Merged to 1.57 A in P 1 21 1. 3 conditions need attention: SCALING_NOT_CONVERGED, MULTIPLE_LATTICES, ICE_RINGS. Space group P 1 21 1 Unit cell 35.337 28.465 62.937 90.000 105.824 90.000 Resolution 1.57 A written; the CC1/2 fall-off is fitted at 1.64 A Signal I/sigma 6.0 R_meas 17.2 % CC1/2 0.995 ISa 10.2 Wall time 11.78 s ============================================================================== With any question about this run, send myrun_report.txt and myrun_log.txt. ``` On a terminal an integration pass draws a progress bar in place; written to a file or a pipe (a batch job's log) every line starts with the time of day instead, there is no bar, and a long pass says how far it has got once a quarter. Warnings are marked `!!` and errors `** ERROR:`; a warning that comes back from a later merge is shown once more on one line and then only counted. `-v` puts the whole log on the terminal in place of the steps. The log file carries what is needed to answer a question about a run without the data: version, build, GPU, threads, the command line, the header geometry, every decision with its numbers, and a timestamp on every line. A few things worth knowing before reaching for more flags: - **The written reflections stop where CC1/2 falls through 0.30.** Every reflection file is resolution-trimmed automatically (`--resolution-cutoff cc-logistic`, one shell past the crossing); `--resolution-cutoff off` keeps the full measured range, `--scaling-high-resolution` fixes the limit by hand. A Rugnux file reaching less far than another program's on the same data is usually this default at work, not lost data. - **Rotation data are best left de novo.** Pinning the cell and space group (recipe 4) is the normal thing to do for **serial stills**, where the `ffbidx` indexer needs a cell; on a rotation sweep it tends to *degrade* low-symmetry cases, so prefer recipe 1 and let the run determine both (see [Rotation data](RUGNUX_TUTORIAL.md#rotation-data)). `-S` takes a Hermann-Mauguin symbol (`P43212`) or a space-group number (`96`), whichever is to hand. - **Anomalous data are there without `-A`.** A rotation merge always keeps the Bijvoet split: a default run's `.mtz` carries `I(+)`/`I(-)` beside `IMEAN`, and its `.hkl` every observation at the index it was measured at — `FRIEDELS_LAW= TRUE` in the report says how the *statistics* were counted, not that the signal was averaged away. What `-A` changes is the counting basis and the error model: each hand becomes a merged observation of its own, so multiplicity, completeness and ⟨I/σ⟩ are counted anomalously and the sigmas are refitted on the Bijvoet-separated merge. Reach for it for anomalous statistics; the signal itself is in the file either way. (Stills merges carry no split by default — there `-A` is what creates one.) - **A model names the enantiomorph.** Where the data accept the model — it is tested against a null of the same model in random orientations, and `MODEL_FIT=` in the report says the verdict — `--model` settles which of P41212 and P43212 the merged reflections are *labelled* with — a choice no merged intensity can make. It is a label and nothing more: the two groups have the same rotation operations, so no reflection moves, and in particular I(+) and I(-) are left exactly as measured. Whether the model agrees with the data about the hand is then a real question, and the anomalous difference map answers it — a run says so when the density at the model's atoms comes out inverted. - **`-z` and `--model` overlap but are not the same.** A reference MTZ steers the processing from the start; a model scores the merge and settles the frame it is written in — where the data accept it; a model they reject changes nothing. Either resolves an [indexing ambiguity](RUGNUX_ADVANCED.md#the-indexing-ambiguity), which on serial data decides whether the merged intensities are usable at all. - **Small molecules need no flag either.** Short axes, spots wider than the integration disk and glide-plane space groups are handled by the default run, and `myrun.hkl` goes straight to SHELXT / SHELXL ([Small-molecule data](RUGNUX_TUTORIAL.md#small-molecule-data)). - **`--scaling-high-resolution `**, where the resolution is already known, sharpens both the space-group search and the error model. - **A run wants memory in proportion to what it integrates**, not to the detector: 2.5-14 GB of host RAM and 3-7 GB on the card over the datasets measured, both peaking in scaling and merging. [Installing Rugnux ▸ Memory](RUGNUX_INSTALL.md#memory) has the table and the two flags that lower it. A very large cell needs far more, and a CPU-only build most of all — tens of GB of host RAM ([Very large unit cells](RUGNUX_INSTALL.md#very-large-unit-cells)). - Everything else is in [Running Rugnux](RUGNUX_TUTORIAL.md#running-rugnux) and the full [Command-line options](RUGNUX_ADVANCED.md#command-line-options). ## The rest of the manual One page per job, so the answer needed is near the top of a short page: - [What Rugnux does](RUGNUX_OVERVIEW.md) — the pipeline from images to merged reflections, in order. Read this one first. - [Installing Rugnux](RUGNUX_INSTALL.md) — packages, the release archive, GPU drivers, building from source, hardware. - [What Rugnux reads](RUGNUX_FORMATS.md) — will it open your data: NXmx / EIGER masters, PILATUS miniCBF, marCCD and SMV (ADSC, Rigaku d\*TREK) sweeps, one sweep per input. - [Running Rugnux](RUGNUX_TUTORIAL.md) — a first run in detail, rotation, serial and small-molecule data, and every file a run writes. - [Rugnux with other programs](RUGNUX_INTEGRATION.md) — the reflection-file conventions, the unmerged export, and worked command lines for phenix, REFMAC, POINTLESS / AIMLESS, careless, Phaser, SHELXC/D/E and, for small molecules, SHELXT / SHELXL. - [The results report](RUGNUX_REPORT.md) — the `KEY= value` interface, sweep quality and the anisotropy section. - [Advanced usage](RUGNUX_ADVANCED.md) — reference data and the indexing ambiguity, model validation, re-merging, and the full command-line option tables. - [Detector calibration](RUGNUX_CALIBRATION.md) — the geometry from a calibrant's powder rings (`--mode calibration`). - [CPU/GPU data analysis](CPU_DATA_ANALYSIS.md) — the algorithms behind all of it. ## Where it fits among the three analysis tools | Tool | Mode | Driven by | Output | | --- | --- | --- | --- | | [`jfjoch_broker`](JFJOCH_BROKER.md) | Online, real-time streaming analysis on FPGA + GPU | HTTP/REST + ZeroMQ | Live results and statistics, images streamed to [`jfjoch_writer`](JFJOCH_WRITER.md) | | [`jfjoch_viewer`](JFJOCH_VIEWER.md) | Interactive, on-screen exploration | Qt desktop application | On screen; a processing job can write the same files as `rugnux` | | **`rugnux`** | **Offline batch processing of a stored dataset** | **Command-line interface** | **`_process.h5`, and `.mtz`/`.cif`/`.hkl` when merging** | Use `rugnux` to re-analyse data after acquisition, to experiment with processing parameters, or to produce merged intensities for downstream structure solution.