Files
Jungfraujoch/docs/RUGNUX.md
T
leonarski_fandClaude Opus 5.5 37b7fdd38d rugnux: the terminal shows numbered steps with timing; the full log goes to <prefix>_log.txt
Standard mode prints only steps, one-line results with the time since the line above, a rule
closing each step with its total, warnings (!!) and errors (** ERROR:), and a summary at the
end. Every log line of every module goes to <prefix>_log.txt instead, at debug level and
timestamped, opened with version, build, GPU, threads, command line and the header geometry.
-v puts the whole log back on the terminal.

- RugnuxObserver gains OnStep / OnNote. RunPipeline posts them only on a pass that makes the
  result (the observer is held in steps_, nullptr on probe passes and on copies run beside the
  main pass); RunAllPasses opens steps for the checks between passes (rotation scale, metric
  symmetry, geometry walk) and notes its re-runs. OnPhase is written to the log as a marker.
- RugnuxConsole (CLI only): on a terminal a progress bar and a spinner status line drawn in
  place with \r, durations only; not a terminal, plain ASCII with the time of day on every line
  and quarter milestones for long passes. isatty/_isatty, no environment variables. A warning
  repeated by later merges is shown once more on one line and then counted per step.
- The CLI's former stdout blocks (space-group search, merging statistics, per-image cost,
  processing time, tNCS/twinning/anisotropy in --mode scale) go to the log; the calibration
  result block stays on the terminal as that mode's summary. The MX summary takes its rows
  from the report's own SUMMARY (ReportDocument::summary_rows); WriteResultReport returns the
  document it rendered. HDF5's automatic error-stack printing is switched off.
- The banner keeps the "Version X (git abcdef" form battery.py reads off a bare invocation;
  battery.py's failure note now reads p_log.txt, whose lines are not wrapped.

Output only: p.mtz is bit-identical on the three in-house reference sets.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-07 12:11:21 +02:00

12 KiB

Rugnux

rugnux is the offline crystallographic data-analysis tool of Jungfraujoch — the data-processing half of the system (see Naming for where the name comes from). It takes an existing HDF5 dataset, runs the full analysis pipeline — spot finding, indexing, geometry refinement, Bragg integration and (optionally) scaling and merging — and writes the results to a _process.h5 file, plus reflection files (.mtz/.cif/.hkl) when merging is requested.

It runs the same analysis code as the online and interactive tools, just driven from the command line over a file rather than a live detector stream.

rugnux {<options>} <input.h5>

Run it with no arguments to print the usage.

Note. rugnux is under very active development. This page describes the tool and its options at a high level; the authoritative, always-current list of options is the program's own usage message — run rugnux with no arguments.

:local:
:depth: 2

Quick start

Four commands cover most of what people ask of rugnux. Each takes the master file of the dataset — one written by Jungfraujoch, a DECTRIS EIGER master, or an NXmx master written by another facility's toolchain; a PILATUS miniCBF, marCCD or SMV sweep works too (see What Rugnux reads) — and names its output files from -o:

# 1. everything from the data - index, integrate, scale and merge with the defaults
rugnux -o myrun dataset_master.h5

# 2. with a reference dataset of the same crystal form: it fixes the space group and the cell,
#    resolves the indexing ambiguity, and hands over its R-free set
rugnux -o myrun -z reference.mtz dataset_master.h5

# 3. with a known structure: R-work / R-free and sigma_A-weighted 2mFo-DFc / mFo-DFc maps
rugnux -o myrun --model model.pdb dataset_master.h5

# 4. with the space group and the cell pinned (-S takes either spelling: P43212 or 96)
rugnux -o myrun -S P43212 -C 79,79,38,90,90,90 dataset_master.h5

Parallelism needs no asking for: a run already uses the machine's threads. -N is there to limit that, or to lift the per-image loop's default ceiling of 16 workers per GPU.

Nothing more is needed to pick the workflow: a dataset carrying a goniometer axis is processed as a rotation sweep, one without as independent stills, and scaling and merging run by default in both. A rotation run that merges — the default — leaves nine files next to each other:

myrun.mtz            merged intensities + French-Wilson amplitudes, for CCP4 / phenix
myrun.cif            the same, as mmCIF - the self-describing format, and what to deposit
myrun.hkl            the scaled observations unmerged, as SHELX HKLF 4 - for SHELXL, SHELXC / ANODE
myrun_unmerged.mtz   every observation before scaling, for pointless / aimless / careless -
                     the largest file of the run (--no-export-unmerged skips it)
myrun_P1.mtz         the same observations merged in P1, so a wrong space-group call can be
                     re-merged or re-refined without reprocessing (--no-p1-crosscheck skips it)
myrun_report.txt     what the run determined: cell, space group, statistics, warnings
myrun_plot.txt       one row per image, for plotting how the crystal behaved over the sweep
myrun_detector.jpg   the detector with the pixel mask and the beam-stop shadow drawn on it
myrun_log.txt        the full timestamped log of the run (debug level)

The two MTZ extras are most of the bytes a run writes — worth knowing when sizing a scratch directory for a campaign, and both have off switches.

Read myrun_report.txt first: it says which space group was chosen and on what evidence, how far the data go, and anything that needs attention.

What the terminal shows

The terminal shows the run as numbered steps, each with one-line results and the time it took, the warnings among them, and a summary at the end; everything else goes to myrun_log.txt. A note's time is how long the step ran since the line above it, and the rule under a step gives its total, so the steps add up to the wall time. Abridged:

Step 3  Indexing
        ... beam centre kept: measured 7.5 px off, same lattice         1.36 s
        ... lattice triclinic P: 28.56 35.34 63.56 106.0 90.3 90.3      0.00 s
        ... 52/60 validation frames, 24% of their spots on it           0.00 s
        ------------------------------------------------------- done in 1.36 s
Step 4  Integration (geometry pre-pass), 1800 images
        ... 1800 images, 836 per second                                 2.15 s
        ------------------------------------------------------- done in 2.16 s
   ...
Step 9  Scaling and merging
        ... merge in P1: 14464 unique                                   0.64 s
        ... space group P 1 21 1 (No. 4)                                0.05 s
        ... merge in P21: 16847 unique                                  0.26 s
        ... resolution 1.57 A (CC1/2 fall-off at 1.64 A)                0.02 s
        ------------------------------------------------------- done in 1.89 s
   ...
==============================================================================
 Result        WARNINGS: Merged to 1.57 A in P 1 21 1. 3 conditions need
               attention: SCALING_NOT_CONVERGED, MULTIPLE_LATTICES, ICE_RINGS.
 Space group   P 1 21 1
 Unit cell     35.337 28.465 62.937 90.000 105.824 90.000
 Resolution    1.57 A written; the CC1/2 fall-off is fitted at 1.64 A
 Signal        I/sigma 6.0   R_meas 17.2 %   CC1/2 0.995   ISa 10.2
 Wall time     11.78 s
==============================================================================
 With any question about this run, send myrun_report.txt and myrun_log.txt.

On a terminal an integration pass draws a progress bar in place; written to a file or a pipe (a batch job's log) every line starts with the time of day instead, there is no bar, and a long pass says how far it has got once a quarter. Warnings are marked !! and errors ** ERROR:; a warning that comes back from a later merge is shown once more on one line and then only counted. -v puts the whole log on the terminal in place of the steps. The log file carries what is needed to answer a question about a run without the data: version, build, GPU, threads, the command line, the header geometry, every decision with its numbers, and a timestamp on every line.

A few things worth knowing before reaching for more flags:

  • The written reflections stop where CC1/2 falls through 0.30. Every reflection file is resolution-trimmed automatically (--resolution-cutoff cc-logistic, one shell past the crossing); --resolution-cutoff off keeps the full measured range, --scaling-high-resolution fixes the limit by hand. A Rugnux file reaching less far than another program's on the same data is usually this default at work, not lost data.
  • Rotation data are best left de novo. Pinning the cell and space group (recipe 4) is the normal thing to do for serial stills, where the ffbidx indexer needs a cell; on a rotation sweep it tends to degrade low-symmetry cases, so prefer recipe 1 and let the run determine both (see Rotation data). -S takes a Hermann-Mauguin symbol (P43212) or a space-group number (96), whichever is to hand.
  • Anomalous data are there without -A. A rotation merge always keeps the Bijvoet split: a default run's .mtz carries I(+)/I(-) beside IMEAN, and its .hkl every observation at the index it was measured at — FRIEDELS_LAW= TRUE in the report says how the statistics were counted, not that the signal was averaged away. What -A changes is the counting basis and the error model: each hand becomes a merged observation of its own, so multiplicity, completeness and ⟨I/σ⟩ are counted anomalously and the sigmas are refitted on the Bijvoet-separated merge. Reach for it for anomalous statistics; the signal itself is in the file either way. (Stills merges carry no split by default — there -A is what creates one.)
  • A model names the enantiomorph. Where the data accept the model — it is tested against a null of the same model in random orientations, and MODEL_FIT= in the report says the verdict — --model settles which of P41212 and P43212 the merged reflections are labelled with — a choice no merged intensity can make. It is a label and nothing more: the two groups have the same rotation operations, so no reflection moves, and in particular I(+) and I(-) are left exactly as measured. Whether the model agrees with the data about the hand is then a real question, and the anomalous difference map answers it — a run says so when the density at the model's atoms comes out inverted.
  • -z and --model overlap but are not the same. A reference MTZ steers the processing from the start; a model scores the merge and settles the frame it is written in — where the data accept it; a model they reject changes nothing. Either resolves an indexing ambiguity, which on serial data decides whether the merged intensities are usable at all.
  • Small molecules need no flag either. Short axes, spots wider than the integration disk and glide-plane space groups are handled by the default run, and myrun.hkl goes straight to SHELXT / SHELXL (Small-molecule data).
  • --scaling-high-resolution <d>, where the resolution is already known, sharpens both the space-group search and the error model.
  • A run wants memory in proportion to what it integrates, not to the detector: 2.5-14 GB of host RAM and 3-7 GB on the card over the datasets measured, both peaking in scaling and merging. Installing Rugnux ▸ Memory has the table and the two flags that lower it. A very large cell needs far more, and a CPU-only build most of all — tens of GB of host RAM (Very large unit cells).
  • Everything else is in Running Rugnux and the full Command-line options.

The rest of the manual

One page per job, so the answer needed is near the top of a short page:

  • What Rugnux does — the pipeline from images to merged reflections, in order. Read this one first.
  • Installing Rugnux — packages, the release archive, GPU drivers, building from source, hardware.
  • What Rugnux reads — will it open your data: NXmx / EIGER masters, PILATUS miniCBF, marCCD and SMV (ADSC, Rigaku d*TREK) sweeps, one sweep per input.
  • Running Rugnux — a first run in detail, rotation, serial and small-molecule data, and every file a run writes.
  • Rugnux with other programs — the reflection-file conventions, the unmerged export, and worked command lines for phenix, REFMAC, POINTLESS / AIMLESS, careless, Phaser, SHELXC/D/E and, for small molecules, SHELXT / SHELXL.
  • The results report — the KEY= value interface, sweep quality and the anisotropy section.
  • Advanced usage — reference data and the indexing ambiguity, model validation, re-merging, and the full command-line option tables.
  • Detector calibration — the geometry from a calibrant's powder rings (--mode calibration).
  • CPU/GPU data analysis — the algorithms behind all of it.

Where it fits among the three analysis tools

Tool Mode Driven by Output
jfjoch_broker Online, real-time streaming analysis on FPGA + GPU HTTP/REST + ZeroMQ Live results and statistics, images streamed to jfjoch_writer
jfjoch_viewer Interactive, on-screen exploration Qt desktop application On screen; a processing job can write the same files as rugnux
rugnux Offline batch processing of a stored dataset Command-line interface _process.h5, and .mtz/.cif/.hkl when merging

Use rugnux to re-analyse data after acquisition, to experiment with processing parameters, or to produce merged intensities for downstream structure solution.