Files
leonarski_fandClaude Opus 5 b8fa8e67d5 docs: the pages catch up with the last day on rc166
An audit of docs/ against the code at HEAD, concentrating on what landed after
the previous documentation audit: the balanced half-set split and CCanom, the
declared-range completeness denominator, the one rule for a quantity nobody
measured, and the PONI the calibration now refuses to write.

Five statements the code had made false:

- The rotation merge's CCref column shows a dash, not "nan".
- Two pages still promised unweighted 2Fo-Fc / Fo-Fc maps; they have been
  sigma_A-weighted since the rigid-body work, and one of the two sat six lines
  from a paragraph that said so.
- Inserting the CCanom prose into the SigAno paragraph left the sentence about
  the PDBx items and the SigAno column with CCanom as its subject, so it read as
  a claim about a quantity that is in neither. CCanom is also rotation-only and
  is not in the mmCIF, which nothing said.
- Two cross-references did not resolve - a heading that was renumbered, and a
  slug spelled without the hyphen MyST puts in "TCP/IP".

Added where the behaviour is new and a user meets it: the report's contract for
a quantity a run did not measure - no key, and a dash in the table, which is not
the same claim as a measured zero; the half-set rule behind CC1/2, which is why
a CUDA and a non-CUDA build now agree on it and on CCanom; and the second reason
a calibration writes no .poni, a detector whose stored image is mirrored or
quarter-turned, which the PONI format cannot state.

The rc.166 changelog stays a release note. One entry is widened from the shell
table to the rule it is a case of, and two are added for output that was wrong
rather than merely undocumented: FITTED_RESOLUTION was the P1 cross-check's, and
the unmerged MTZ carried the reference setting where the merged file carried the
adopted one.

Sphinx builds clean with -W on the pinned docs/requirements.txt, and every
intra-doc anchor resolves against the generated HTML.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EFEJG6WBQv8th4UJFNe53N
2026-09-02 17:44:20 +02:00

8.6 KiB

rugnux

rugnux is the offline crystallographic data-analysis tool of Jungfraujoch — the data-processing half of the system (see Naming for where the name comes from). It takes an existing HDF5 dataset, runs the full analysis pipeline — spot finding, indexing, geometry refinement, Bragg integration and (optionally) scaling and merging — and writes the results to a _process.h5 file, plus reflection files (.mtz/.cif/.hkl) when merging is requested.

It runs the same analysis code as the online and interactive tools, just driven from the command line over a file rather than a live detector stream.

rugnux {<options>} <input.h5>

Run it with no arguments to print the usage.

Note. rugnux is under very active development. This page describes the tool and its options at a high level; the authoritative, always-current list of options is the program's own usage message — run rugnux with no arguments.

:local:
:depth: 2

Quick start

Four commands cover most of what people ask of rugnux. Each takes the master file of the dataset — one written by Jungfraujoch, a DECTRIS EIGER master, or an NXmx master written by another facility's toolchain; a PILATUS miniCBF sweep works too (see What rugnux reads) — and names its output files from -o:

# 1. everything from the data - index, integrate, scale and merge with the defaults
rugnux -o myrun dataset_master.h5

# 2. with a reference dataset of the same crystal form: it fixes the space group and the cell,
#    resolves the indexing ambiguity, and hands over its R-free set
rugnux -o myrun -z reference.mtz dataset_master.h5

# 3. with a known structure: R-work / R-free and sigma_A-weighted 2mFo-DFc / mFo-DFc maps
rugnux -o myrun --model model.pdb dataset_master.h5

# 4. with the space group and the cell pinned (-S takes either spelling: P43212 or 96)
rugnux -o myrun -S P43212 -C 79,79,38,90,90,90 dataset_master.h5

Parallelism needs no asking for: a run already uses the machine's threads. -N is there to limit that, or to lift the per-image loop's default ceiling of 16 workers per GPU.

Nothing more is needed to pick the workflow: a dataset carrying a goniometer axis is processed as a rotation sweep, one without as independent stills, and scaling and merging run by default in both. A rotation run that merges — the default — leaves seven files next to each other:

myrun.mtz            merged intensities + French-Wilson amplitudes, for CCP4 / phenix
myrun.cif            the same, as mmCIF - the self-describing format, and what to deposit
myrun.hkl            the same, as SHELX HKLF 4 - feed this to SHELXC / SHELXD / ANODE
myrun_unmerged.mtz   every observation before scaling, for pointless / aimless / careless -
                     the largest file of the run (--no-export-unmerged skips it)
myrun_P1.mtz         the same observations merged in P1, so a wrong space-group call can be
                     re-merged or re-refined without reprocessing (--no-p1-crosscheck skips it)
myrun_report.txt     what the run determined: cell, space group, statistics, warnings
myrun_image.dat      one row per image, for plotting how the crystal behaved over the sweep

The two MTZ extras are most of the bytes a run writes — worth knowing when sizing a scratch directory for a campaign, and both have off switches.

Read myrun_report.txt first: it says which space group was chosen and on what evidence, how far the data go, and anything that needs attention.

A few things worth knowing before reaching for more flags:

  • The written reflections stop where CC½ falls through 0.30. Every reflection file is resolution-trimmed automatically (--resolution-cutoff cc-logistic, one shell past the crossing); --resolution-cutoff off keeps the full measured range, --scaling-high-resolution fixes the limit by hand. A rugnux file reaching less far than another program's on the same data is usually this default at work, not lost data.
  • Rotation data are best left de novo. Pinning the cell and space group (recipe 4) is the normal thing to do for serial stills, where the ffbidx indexer needs a cell; on a rotation sweep it tends to degrade low-symmetry cases, so prefer recipe 1 and let the run determine both (see Rotation data). -S takes a Hermann-Mauguin symbol (P43212) or a space-group number (96), whichever is to hand.
  • Anomalous data are there without -A. A rotation merge always keeps the Bijvoet split: a default run's .mtz carries I(+)/I(-) beside IMEAN, and its .hkl the ±hkl mates — FRIEDELS_LAW= TRUE in the report says how the statistics were counted, not that the signal was averaged away. What -A changes is the counting basis and the error model: each hand becomes a merged observation of its own, so multiplicity, completeness and ⟨I/σ⟩ are counted anomalously and the sigmas are refitted on the Bijvoet-separated merge. Reach for it for anomalous statistics; the signal itself is in the file either way. (Stills merges carry no split by default — there -A is what creates one.)
  • A model names the enantiomorph. Where the data accept the model — it is tested against a null of the same model in random orientations, and MODEL_FIT= in the report says the verdict — --model settles which of P41212 and P43212 the merged reflections are labelled with — a choice no merged intensity can make. It is a label and nothing more: the two groups have the same rotation operations, so no reflection moves, and in particular I(+) and I(-) are left exactly as measured. Whether the model agrees with the data about the hand is then a real question, and the anomalous difference map answers it — a run says so when the density at the model's atoms comes out inverted.
  • -z and --model overlap but are not the same. A reference MTZ steers the processing from the start; a model scores the merge and settles the frame it is written in — where the data accept it; a model they reject changes nothing. Either resolves an indexing ambiguity, which on serial data decides whether the merged intensities are usable at all.
  • --scaling-high-resolution <d>, where the resolution is already known, sharpens both the space-group search and the error model.
  • Everything else is in Running rugnux and the full Command-line options.

The rest of the manual

One page per job, so the answer needed is near the top of a short page:

  • What a rugnux run does — the pipeline from images to merged reflections, in order. Read this one first.
  • Installing rugnux — packages, the release archive, GPU drivers, building from source, hardware.
  • What rugnux reads — will it open your data: NXmx / EIGER masters, PILATUS miniCBF, one sweep per input.
  • Running rugnux — a first run in detail, rotation and serial data, and every file a run writes.
  • rugnux with other programs — the reflection-file conventions, the unmerged export, and worked command lines for phenix, REFMAC, POINTLESS / AIMLESS, careless and Phaser.
  • The results report — the KEY= value interface, sweep quality and the anisotropy section.
  • Advanced usage — reference data and the indexing ambiguity, model validation, re-merging, and the full command-line option tables.
  • Detector calibration — the geometry from a calibrant's powder rings (--mode calibration).
  • CPU/GPU data analysis — the algorithms behind all of it.

Where it fits among the three analysis tools

Tool Mode Driven by Output
jfjoch_broker Online, real-time streaming analysis on FPGA + GPU HTTP/REST + ZeroMQ Live results and statistics, images streamed to jfjoch_writer
jfjoch_viewer Interactive, on-screen exploration Qt desktop application On screen; a processing job can write the same files as rugnux
rugnux Offline batch processing of a stored dataset Command-line interface _process.h5, and .mtz/.cif/.hkl when merging

Use rugnux to re-analyse data after acquisition, to experiment with processing parameters, or to produce merged intensities for downstream structure solution.