* rugnux now tells you whether a crystal diffracts anisotropically and how far it reaches in each direction, without a second program: a new `9. DIFFRACTION ANISOTROPY` section in `<prefix>_report.txt` and matching `_reflns.pdbx_aniso_B_tensor_*` / `_reflns.jfjoch_aniso_*` items in the merged mmCIF report the anisotropic deltaB, the diffraction limit along each principal direction, and a `NOT DETECTED` / `DETECTED` / `CANNOT DETERMINE` verdict measured against the data set's own systematic error. It is a description only - no intensity is corrected, no reflection is removed, and the merged data do not depend on direction.
* rugnux can hand its integrated observations to another scaling program: `--export-unmerged` writes `<prefix>_unmerged.mtz`, an unmerged MTZ readable by aimless, pointless, careless and `iotbx.merging_statistics`, in `--mode mx` and `--mode scale` alike. Each rotation reflection's partials are summed into one full; `--export-unmerged-partials` writes one row per image instead. Intensities carry the Lorentz-polarization factor and nothing else, since those programs scale the data themselves. Lattice-centring absences are not written; screw and glide absences are.
* rugnux integrates crystals with broad spots better - where it changes anything, per-shell mean I/sigma improves by up to 31% and R_meas by up to 24% - because on rotation data the integration signal radius is now taken from the crystal's own measured spot width instead of a fixed 4 px. `--adaptive-integration-radius=off` restores the fixed radius and an explicit `--integration-radius` still overrides both. The widened radius applies to the final integration pass only, and a pattern too dense for it is re-integrated at 4 px with a note in the log.
* rugnux discards fewer stills reflections for want of a background ring, improving per-shell R_meas over most of the signal-bearing range: the stills background ring now runs to 14 px instead of 12. The gain reverses in shells below a mean I/sigma of about 4.
* rugnux determines the space group with thresholds that mean the same thing on a weak crystal as on a strong one: symmetry operators are scored on resolution-normalised intensities (E squared) instead of raw merged intensities, and a reflection counts as genuinely present on its counting significance instead of on the merged I/sigma, which saturates at the merge's own ISa. The search resolution cut is no longer able to move the answer, and the twin-law H bound moves from 1.70 to 1.85, which stops one class of correct high-symmetry assignment being refused as twinning.
* rugnux says what the space-group search tested and what it could not: the twin-law disagreement H is printed for every operator together with the adopted point group's H ratio and its bound; alternatives that are not on the reported lattice are named with how their cell differs; and a lattice centring the data could not test - the crystal having been integrated on the primitive sub-cell, so the reflections it extinguishes were never measured - is marked `UNTESTED` and warned about where it is adopted, as coming from the lattice metric rather than from the intensities.
* rugnux `--mode scale` re-merges a `_process.h5` in the right symmetry without being told it: the file now records the space group on every run - a two-pass rotation run wrote none before, so re-merging defaulted to P1 - together with the change of basis under `/entry/MX/reindexMatrix` where the lattice was re-seated, and `--mode scale` also reports the Wilson B-factor estimate instead of `WILSON_B= nan`. A file written before this stops with a message naming the two cells and the override to use, instead of failing inside the merge. A third-party reader of a `_process.h5` must apply `reindexMatrix` where it is present.
* rugnux installs on its own, as a package called `rugnux` - `dnf install rugnux` or `apt install rugnux` - instead of arriving inside `jfjoch-viewer`. It pulls in none of the acquisition stack, so a machine that only processes data no longer has to carry the broker, the detector libraries or Qt to get it. Installing it over a `jfjoch-viewer` from rc.163 or earlier, which still owns `/usr/bin/rugnux`, upgrades cleanly rather than failing on the duplicate file.
* rugnux is also a standalone download, built for arm64 as well as x86_64: `rugnux-<version>-linux-{x86_64|aarch64}-cuda<major>.tgz` and `rugnux-<version>-win64-cuda<major>.zip` on the release page, for machines that are not managed by a package manager. The aarch64 build targets GH200 and DGX Spark, and is untested on hardware.
* Every portable Linux binary is now a single self-contained file: cuFFT is linked statically instead of being shipped beside the executable and found through an rpath, so `rugnux` and `jfjoch_viewer` need nothing but an NVIDIA driver, and only to use the GPU. The `.rpm`/`.deb` continue to take cuFFT from the distribution. The developer utilities `jfjoch_extract_hkl` and `jfjoch_recompress` are no longer packaged anywhere.
* Jungfraujoch needs six fewer shared libraries on the machine - libopenblas and libmetis, and libgfortran, libquadmath, libgomp and libz behind them - because the Ceres LAPACK, METIS and SuiteSparse back-ends are no longer built. Nothing in the code ever selected them, and results are unchanged.
* The PCIe driver DKMS package builds for the kernel it is being installed for instead of the running one, so a module built while a kernel update is being applied loads after the reboot.
* The PCIe driver builds on RHEL 9.5 and later, and on their CentOS Stream, Rocky and AlmaLinux equivalents, where the `vm_flags` kernel interface was backported into the 5.14 kernel.
* A data collection started with `async_start` that fails to start - a writer refusing to overwrite an existing file, for instance - is reported as an error by `/wait_until_running` and `/wait_till_done` instead of as a timeout and a successful collection respectively. The error message is the one the writer gave.
* A calibration that is cancelled or that fails to collect its pedestals is no longer reported as a successful one. The broker goes to `Inactive` with an error message and has to be initialized again, instead of sitting in `Idle` looking ready to measure while holding partial pedestals - data collected in that state was silently mis-converted.
* A failed `/initialize` is reported to `/wait_until_running` and `/wait_till_done` as soon as it happens, instead of when their timeout expires.
* `space_group_number` accepts space groups up to 230 in the API schema, so cubic space groups can be recorded. The broker always accepted them; the generated clients rejected them before the request was sent.
* The results report's `REPORT_VERSION` is 3, two sections having been added. Existing key names and table columns are unchanged.
* The merged statistics table has **9** resolution shells instead of 10, which is what XDS reports. The bins were already XDS's - equal steps in 1/d^2 between the lowest- and the highest-resolution reflection the merge kept - so at the same resolution limits the two tables now have the same shell boundaries and can be read row for row. `--resolution-shells` sets a different count.
* `rugnux --model` now settles the frame the merged reflections are written in, not only the frame the R-factors and the maps are computed in: the `.mtz`/`.cif`/`.hkl` come out in the model's indexing, and where the data were merged in the model's enantiomorph they take the model's hand and space group - which on anomalous data puts I(+) and I(-) the right way round. The indexing choice is logged with the winning R-free and the runner-up, so a decision made within noise is visible.
* `rugnux --model` can resolve the indexing ambiguity of a **serial stills** run, which a model could not do before: structure factors computed from the model become the per-image reference, the same role a reference MTZ plays. It needs the cell and space group up front (`-C` / `-S`). Without one or the other, a merohedral serial run still merges both hands together and says so.
* The rugnux documentation opens with a quick start - the default run, and runs with a reference MTZ, with a model, or with the space group and cell pinned - and explains the indexing ambiguity: what it costs on rotation and on serial data, and which of `-z` / `--model` resolves it in each case. The long reference pages now carry a table of contents.
Reviewed-on: #74
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
76 KiB
rugnux
rugnux is the offline crystallographic data-analysis tool of Jungfraujoch — the
data-processing half of the system (see Naming for where the name comes from).
It takes an existing HDF5 dataset, runs the full analysis pipeline — spot finding, indexing,
geometry refinement, Bragg integration and (optionally) scaling and merging — and writes the
results to a _process.h5 file, plus reflection files (.mtz/.cif/.hkl) when merging is
requested.
It runs the same analysis code as the online and interactive tools, just driven from the command line over a file rather than a live detector stream.
rugnux {<options>} <input.h5>
Run it with no arguments to print the usage.
Note.
rugnuxis under very active development. This page describes the tool and its options at a high level; the authoritative, always-current list of options is the program's own usage message — runrugnuxwith no arguments.
:local:
:depth: 2
Quick start
Four commands cover most of what people ask of rugnux. Each takes the master file of a
Jungfraujoch dataset and names its output files from -o:
# 1. everything from the data - index, integrate, scale and merge with the defaults
rugnux -o myrun dataset_master.h5
# 2. with a reference dataset of the same crystal form: it fixes the space group and the cell,
# resolves the indexing ambiguity, and hands over its R-free set
rugnux -o myrun -z reference.mtz dataset_master.h5
# 3. with a known structure: R-work / R-free and 2Fo-Fc / Fo-Fc maps on top of the merge
rugnux -o myrun --model model.pdb dataset_master.h5
# 4. with the space group and the cell pinned (-S takes either spelling: P43212 or 96)
rugnux -o myrun -S P43212 -C 79,79,38,90,90,90 dataset_master.h5
Parallelism needs no asking for: a run already uses the machine's threads. -N is there to limit
that, or to lift the per-image loop's default ceiling of 16 workers per GPU.
Nothing more is needed to pick the workflow: a dataset carrying a goniometer axis is processed as a rotation sweep, one without as independent stills, and scaling and merging run by default in both. A run that merges — the default — leaves five files next to each other:
myrun.mtz merged intensities + French-Wilson amplitudes, for CCP4 / phenix
myrun.cif the same, as mmCIF - the self-describing format, and what to deposit
myrun.hkl the same, as SHELX HKLF 4 - feed this to SHELXC / SHELXD / ANODE
myrun_report.txt what the run determined: cell, space group, statistics, warnings
myrun_image.dat one row per image, for plotting how the crystal behaved over the sweep
Read myrun_report.txt first: it says which space group was chosen and on what evidence, how far
the data go, and anything that needs attention.
A few things worth knowing before reaching for more flags:
- Rotation data are best left de novo. Pinning the cell and space group (recipe 4) is the
normal thing to do for serial stills, where the
ffbidxindexer needs a cell; on a rotation sweep it tends to degrade low-symmetry cases, so prefer recipe 1 and let the run determine both (see Rotation data).-Stakes a Hermann-Mauguin symbol (P43212) or a space-group number (96), whichever is to hand. - A model pins the hand. Where the model is isomorphous with the data,
--modelsettles the enantiomorph — P41212 against P43212, which no measurement can decide — and the merged reflections are written in the model's hand and its space group. On anomalous data that is what puts I(+) and I(-) the right way round. -zand--modeloverlap but are not the same. A reference MTZ steers the processing from the start; a model scores the merge and settles the frame it is written in. Either resolves an indexing ambiguity, which on serial data decides whether the merged intensities are usable at all.--scaling-high-resolution <d>, where the resolution is already known, sharpens both the space-group search and the error model.- Everything else is in Running rugnux and the full Command-line options.
Installation
rugnux is a single self-contained executable. It needs no CUDA toolkit, no Qt, and no
Jungfraujoch service running anywhere; on a machine with an NVIDIA GPU it needs the NVIDIA
driver, and without one it still runs on the CPU.
From the package repositories (RHEL / Rocky / Ubuntu)
On a distribution covered by the package repositories, rugnux is a package of
its own:
sudo dnf install rugnux # RHEL / Rocky 8 and 9
sudo apt install rugnux # Ubuntu 22.04 / 24.04
It installs /usr/bin/rugnux and depends on nothing from the acquisition side — no broker, no
detector libraries, no Qt — so it can go on a machine that only processes data.
Upgrading from rc.163 or earlier.
/usr/bin/rugnuxused to belong to thejfjoch-viewerpackage. Therugnuxpackage declares that the file has moved, so installing it upgrades an oldjfjoch-viewerin the same transaction instead of failing on the duplicate path. If yourjfjoch-vieweris pinned to an old version, unpin it or remove it first.
From the release archive
For a machine no package manager covers — or for Windows and Arm, which have no repository — take the archive for your architecture from the Gitea release page:
| Archive | For |
|---|---|
rugnux-<version>-linux-x86_64-cuda12.tgz |
64-bit Intel/AMD Linux. Built on RHEL 8, so it runs on any newer Linux |
rugnux-<version>-linux-aarch64-cuda13.tgz |
64-bit Arm Linux — NVIDIA GH200 and DGX Spark. Built on Ubuntu 24.04, so it needs glibc 2.39 or newer. Cross-compiled and not yet exercised on Arm hardware |
rugnux-<version>-win64-cuda13.zip |
64-bit Windows |
The archive has no top-level directory — it unpacks straight into bin/ and share/. Always
give tar a destination of its own, or it will scatter those into whatever directory you are in:
mkdir -p /opt/rugnux-1.0.0
tar xzf rugnux-1.0.0-linux-x86_64-cuda12.tgz -C /opt/rugnux-1.0.0
/opt/rugnux-1.0.0/bin/rugnux # prints the usage
What you get is:
bin/rugnux the program
share/doc/jfjoch_rugnux/LICENSE GPLv3
share/doc/jfjoch_rugnux/THIRD_PARTY_NOTICES.md
share/doc/jfjoch_rugnux/licenses/ verbatim licence texts of the bundled dependencies
Nothing is written outside that directory, nothing needs root, and several versions can sit side by
side. To remove it, delete the directory. Put bin/ on your PATH if you want to type rugnux
rather than the full path.
Mixing the two. If a
rugnuxpackage is also installed,/usr/bin/rugnuxwill normally win onPATH. Put the archive'sbin/first, or call it by its full path, to be sure which one you are running —rugnuxprints its version on every run.
GPU support
The released archives are CUDA builds. They need only an NVIDIA driver on the host — 525.60.13
or newer for the CUDA 12 archive, 580.65.06 or newer for the CUDA 13 ones — and no CUDA toolkit,
because everything CUDA is linked statically. With no GPU or no driver, rugnux reports zero CUDA
devices and falls back to the CPU path, which works but is far slower and offers only the fftw
indexer. A V100 needs the CUDA 12 archive; which generations each build covers is in
Release contents ▸ GPU generations and the NVIDIA driver.
Building from source
rugnux alone, without the server stack or Qt:
cmake -S . -B build -DJFJOCH_RUGNUX_ONLY=ON -DCMAKE_BUILD_TYPE=Release \
-DCMAKE_CXX_FLAGS="-march=x86-64-v3" -DCMAKE_C_FLAGS="-march=x86-64-v3"
cmake --build build -j$(nproc) --target rugnux
The binary lands in build/rugnux/rugnux. Two dependencies must come from the system — zlib and
Eigen ≥ 3.4 (zlib-devel and eigen3-devel, or their Debian equivalents); everything else is
downloaded during the first configure, which therefore needs network access. cmake --build build --target package produces the same .tgz the release ships. The -march flag is not set by the
build system on purpose, so a plain build is slower than the released one on the CPU-bound stages —
see the note in CMakeLists.txt.
Where it fits among the three analysis tools
| Tool | Mode | Driven by | Output |
|---|---|---|---|
jfjoch_broker |
Online, real-time streaming analysis on FPGA + GPU | HTTP/REST + ZeroMQ | Live results and statistics, images streamed to jfjoch_writer |
jfjoch_viewer |
Interactive, on-screen exploration | Qt desktop application | On screen; a processing job can write the same files as rugnux |
rugnux |
Offline batch processing of a stored dataset | Command-line interface | _process.h5, and .mtz/.cif/.hkl when merging |
Use rugnux to re-analyse data after acquisition, to experiment with processing
parameters, or to produce merged intensities for downstream structure solution.
Hardware
As with the rest of Jungfraujoch, serious performance requires an NVIDIA GPU. The CUDA build
provides the GPU fast-feedback indexer (ffbidx) and the GPU FFT indexer (fft); without CUDA
only the CPU fftw indexer is available. With a GPU present most of the per-image pipeline runs on
the device — bitshuffle+LZ4 decompression, image preprocessing, azimuthal integration, spot finding,
prediction and Bragg integration — as does rotation scaling and merging, with CPU implementations as
the fallback where there is no GPU. The thread count (-N) governs the CPU side of all of it.
The released CUDA builds need only an NVIDIA driver on the host, no CUDA toolkit: 525.60.13 or
newer for the CUDA 12 artefacts (RHEL 8 packages, the x86_64 rugnux archive) and 580.65.06 or
newer for the CUDA 13 ones (RHEL 9, Ubuntu, the aarch64 and Windows rugnux archives). Which GPU
generations each artefact supports — a V100 in particular works only with the CUDA 12 build — is in
Release contents ▸ GPU generations and the NVIDIA driver.
Running rugnux
A first run in detail
The Quick start command is the whole of it:
rugnux -o myrun /path/to/dataset_master.h5
-o myrun is the prefix every output file is named from and the last argument is the master file
of a Jungfraujoch dataset; -N would set the worker-thread count, which otherwise follows the
machine. Nothing is assumed about the crystal
— the goniometer axis in the file tells rugnux this is a rotation sweep, the unit cell comes from
indexing the data, the space group from its systematic absences, and the resolution limit from where
CC1/2 falls off. Progress, statistics and timing go to the terminal, and the five output files land
next to each other.
myrun_report.txt is written for a person, top to bottom: it says which space group was chosen and
on what evidence, how far the data go, and anything that needs attention. To pull one number out of
it in a script, every value is a KEY= value line:
grep '^SPACE_GROUP_NUMBER= ' myrun_report.txt
grep '^UNIT_CELL_CONSTANTS= ' myrun_report.txt
grep '^INCLUDE_RESOLUTION_RANGE= ' myrun_report.txt
grep '^ISA= ' myrun_report.txt
grep '^WARNING:' myrun_report.txt
Useful variations, each independent of the others:
# tell it where the data really stop, if you already know - this sharpens the
# space-group search and the error model
rugnux -o myrun --scaling-high-resolution 1.4 dataset_master.h5
# keep Friedel pairs apart, for anomalous work
rugnux -o myrun -A dataset_master.h5
# a quick look at the first 200 images only
rugnux -o quicklook -e 200 dataset_master.h5
# merge as usual, but also keep the per-image file so the data can be re-merged later
rugnux -o myrun --write-process-h5 dataset_master.h5
# also write an unmerged MTZ, to scale the data with aimless instead
rugnux -o myrun --export-unmerged dataset_master.h5
# check the merged data against a known structure: R-work / R-free and maps
rugnux -o myrun --model model.pdb dataset_master.h5
Re-merging is cheap and does not re-read the images. Ask the full run to keep its per-image file
with --write-process-h5, and --mode scale will then re-scale and re-merge the reflections
already integrated in it — seconds rather than minutes:
rugnux -o myrun --write-process-h5 dataset_master.h5 # integrate and merge once
rugnux --mode scale -o remerged -A myrun_process.h5 # re-merge, here anomalously
Use it to try a different resolution limit, anomalous setting or outlier rejection without
paying for integration again. --mode scale merges in the space group and cell the file
records, so the second command needs no -S. Note that --no-merge also writes a _process.h5,
but a run that never merged never determined a space group either, so re-merging that file lands in
P1 unless you pass -S yourself — --write-process-h5 is the one to use.
Rotation data
Index, integrate, scale and merge a rotation sweep, fully de novo:
rugnux rotation_master.h5 \
-o rotation_run \
--scaling-high-resolution 1.4
Because the dataset carries a rotation goniometer axis, it is processed as rotation data by
default: two-pass rotation indexing (index the sweep once, then process every frame against that
lattice) with the rot3d partiality model (rotation partials combined into 3D fulls). Scaling
and merging run by default (for both rotation and stills; --no-merge turns them off); the unit
cell is taken from the rotation indexer and the space group is determined from systematic absences,
and both are written
into the merged .cif.
Run fully de novo (no -C/-S) for the best result — supplying a cell or space group up front
tends to degrade low-symmetry cases. A -S group whose Bravais lattice the crystal turns out not to
have stops the run and names the cell that was indexed, rather than merging in a frame the reflections
are not in; where the lattice does have that group's setting, the reflections are reindexed into it. --scaling-high-resolution (set it to your expected
resolution) sharpens both the space-group search and the error model. To tune the first pass use
--two-pass-rotation=100 (or -R100 — the first-pass image count); to force the sweep to be
treated as independent stills use --force-still.
By default a rotation run also post-refines the geometry in a second pass: the first pass
integrates and merges at the header geometry, then the detector distance + beam centre and the crystal
cell / rotation-axis are refined against the merged fulls (cross-validated, and committed only for a
small < 1 % move, with the gauge-weak beam centre restrained toward the header), and the second pass
re-indexes de novo and re-integrates at the refined geometry. The refined pass is the canonical
<prefix>_* output; the header-geometry pass merges only to choose the space group and to judge the
refined pass against, and writes no merged files of its own — no <prefix>_01.mtz, .cif, .hkl or
_01_image.dat. (Where a process file is asked for at all, with --no-merge or
--write-process-h5, each pass still writes its own, so <prefix>_01_process.h5 appears beside
<prefix>_process.h5.) Disable it with --rotation-no-postrefine.
After the per-frame scale-fulls step, rotation scaling applies three correction surfaces, on by
default (--no-scaling-corrections disables all):
- Decay — a global Debye–Waller relative-B over the run, for the radiation damage that weakens
later frames more at high resolution (a resolution×time systematic the resolution-flat per-frame
scale cannot remove). It only engages when the total relative-B exceeds a physical floor (2 Ų). An
optional
--relative-b[=deg]extends this single global rate to a smooth per-batch relative-B curve (default 10°-of-rotation batches when bare, off otherwise), cross-validated like the surfaces here, for crystals whose decay is non-linear in dose. - Absorption — a smooth multiplicative factor over the diffracted-beam direction in the goniometer frame (path length through the crystal). Negligible at hard X-rays / thin crystals; it matters at low photon energy. Its benefit shows up most on model-based metrics: a smooth absorption error largely cancels among symmetry mates (little effect on the error model / ISa) but still biases the intensities, so it measurably lowers Rfree.
- Modulation — a smooth multiplicative factor over the position where a reflection lands on the detector (a flat-field: detector-response and geometric systematics that vary across the detector plane). Symmetry-equivalents of one reflection land at different detector positions as the crystal rotates, which over-determines the surface. Because it lives in the detector frame (not the rotation) the same correction concept applies to stills. This is the largest of the three on JUNGFRAU data — it lowers Rmeas by several to tens of percent on datasets that carry a detector systematic, while holding or improving CC1/2 and the anomalous signal.
All three are cross-validated — fitted on even-numbered frames and kept only if they improve the held-out odd-frame symmetry-equivalent agreement by a clear margin (and vice versa). The agreement is scored as a σ-independent, Rmeas-like fractional deviation, so a surface can never pass cross-validation by merely reshaping the sigmas; where the systematic is absent the surface is a no-op rather than a source of added noise, which is why they are safe to leave on.
Independently of any correction, a rotation run prints a radiation-damage report — the per-image
scale correlation-to-merge and mosaicity versus dose, and the relative B-factor change over the run
(first→last) together with a per-batch relative-B curve, also written to the merged mmCIF. It is a
data-quality-vs-dose diagnostic and never alters the merged intensities. A batch whose data cannot
support a measurement prints - instead of a value, and the first→last number is printed only where a
straight line describes the curve — damage is progressive, so a curve that dips and recovers is a
disturbance of the sweep, not dose, and the report says so and points at the sweep-quality section
(RADIATION_DAMAGE_RELATIVE_B= NOT_A_TREND).
Still / serial data
A dataset with no goniometer axis (e.g. a serial grid scan) is processed as independent stills automatically — no flag needed. Known-cell indexing with the GPU fast-feedback indexer, then merge against a reference structure:
rugnux serial_master.h5 \
-o serial_run \
-X ffbidx -C 79,79,38,90,90,90 -S 96 \
-z reference.mtz \
--scaling-high-resolution 1.8
A crystal form with an indexing ambiguity — P3, P4, P6 and their
relatives — needs either the -z above or --model model.pdb, and needs it on the run that
integrates: every crystal is indexed in its own hand, and the two are averaged together in the merge
unless each image is put into the same hand as it is integrated.
ffbidx requires a known cell (-C) and is the indexer of choice for sparse serial stills. The
self-calibrating spot finder is on by default for both workflows (--no-adaptive-spots turns it off), and for
serial stills leave --min-pix-per-spot unset so it is chosen per image — across the still-target battery this
combination raises the indexing rate and typically extends resolution over a fixed threshold and
fixed min-pix, at equal or better CC½. (You can still pin a fixed threshold with --spot-sigma /
--spot-threshold and a fixed min-pix with --min-pix-per-spot.) If a dataset does carry a
goniometer axis but you want per-frame stills processing anyway, add --force-still.
Input and output
Input is a single Jungfraujoch HDF5 master file (NXmx-based). Spots are always found by rugnux
itself, including for the two-pass rotation first pass — the spot lists a dataset may already carry
were found online, at the acquisition's threshold and with its ice-band spots already discarded, so
reusing them would hide the spot-finding settings from the lattice search.
Output (controlled by -o, --output-prefix, default output):
-
<prefix>_process.h5— NXmx-compliant HDF5 with derived metadata (spots, indexing, integration, azimuthal integration, per-image statistics). See HDF5 / NeXus data format for the layout. Written by default only when not merging (i.e. under--no-merge); add--write-process-h5to also write it when merging. It does not copy the images:/entry/data/datais a virtual dataset over the input files, so the input has to stay where it was for the pictures to be readable, and the pixel metadata (bit_depth_readout,underload_value, the dataset type) describes those files rather than the signed 32-bit container rugnux processes in. -
Merging is on by default (
--no-mergedisables it). The merged reflections are written in three formats — each has its uses downstream:<prefix>.mtz— CCP4 MTZ (IMEAN/I(+)/I(-), French–WilsonF,FreeR_flag) for the CCP4 / phenix reflection tools.<prefix>.cif— mmCIF, for deposition and as the self-describing native format (also carries the merging statistics, ISa, twinning and radiation-damage indicators).<prefix>.hkl— SHELX HKLF 4 text (h k l I σ(I), fixed3I4,2F8.2), the direct input for SHELXC / ANODE / SHELXD. Bijvoet mates are written separately (I(+)at+hkl,I(-)at-hkl) so the anomalous signal is preserved; intensities are put on a common scale so the largest value fits the fixed-width field (the absolute scale is irrelevant to SHELXC/ANODE), and the file ends with the0 0 0terminator record.
All three carry the refined unit cell (from rotation indexing) and the space group determined from systematic absences (constrained to the indexed lattice symmetry).
-
<prefix>_unmerged.mtz— opt-in (--export-unmerged): the integrated observations before merging, as an unmerged MTZ in POINTLESS's column layout, so the data can be scaled and merged by aimless, pointless, careless oriotbx.merging_statisticsinstead of by rugnux. See The unmerged export below.--export-unmerged-partialswrites<prefix>_unmerged_partials.mtz, one row per image, instead of summing. -
<prefix>_report.txt— the results report: what the run determined, in a form both a person and a beamline script can read. Always written, next to the files above. See The results report below.
Merged statistics (⟨I/σ⟩, CC1/2, completeness, …), the error model and timing are printed to the
console. By default the written resolution is trimmed automatically where CC1/2 falls off
(--resolution-cutoff cc-logistic, CC1/2 target 0.30); set --scaling-high-resolution to fix the
limit by hand, or --resolution-cutoff off to keep the full range.
Reflection-file conventions
mmCIF. Standard items carry their standard meanings — _refln.intensity_meas / _intensity_sigma,
the pdbx_I_plus/pdbx_I_minus and pdbx_F_plus/pdbx_F_minus anomalous pairs, _reflns.* and
_reflns_shell.* for the merging statistics, _reflns.B_iso_Wilson_estimate for the Wilson B, and
_cell.* / _diffrn_radiation_wavelength.wavelength for the geometry.
Anything rugnux reports that has no standard item is written under a jfjoch_ prefix, inside the
standard category it belongs to. That is a deliberate choice: a reader that does not know these items
ignores them, and one that does can find them without guessing.
| item | meaning |
|---|---|
_reflns.jfjoch_diffrn_ISa |
Asymptotic I/σ in XDS's sense: the whole-range 1/√(a·b) of the error model, so it can be read directly against a CORRECT.LP |
_reflns.jfjoch_diffrn_ISa_asymptotic |
The strong-reflection tier — the counting-subtracted scatter of well-measured groups. XDS has no equivalent, and it can only ever be the more optimistic of the two. Rotation path only |
_reflns.jfjoch_error_model_a, _b |
The error model in XDS's convention, σ² = a(σ₀² + b·I²), so the ISa above is re-derivable from the file rather than taken on trust |
_reflns.jfjoch_second_moment_I |
Twinning second moment ⟨I²⟩/⟨I⟩² — 2.00 untwinned, 1.50 for a perfect twin |
_reflns.jfjoch_L_test_mean_abs_L, _L_test_mean_L_squared |
Padilla–Yeates L-test. ⟨|L|⟩ is 0.500 untwinned / 0.375 for a perfect twin; ⟨L²⟩ is 0.333 / 0.200. Written only when the test found pairs |
_reflns.jfjoch_radiation_damage_relative_B |
Relative B from the first to the last rotation batch (Ų); positive is the usual direction, high-resolution intensity fading with dose |
_jfjoch_radiation_damage_batch.* |
Per-batch loop: id, rotation_start_deg, relative_B |
_diffrn_detector.jfjoch_distance_mm, _jfjoch_beam_center_x_pxl, _jfjoch_beam_center_y_pxl |
The refined detector geometry actually used, which is not otherwise recoverable from the reflection file |
_reflns.pdbx_aniso_B_tensor_eigenvalue_1..3, _pdbx_aniso_B_tensor_eigenvector_* |
The anisotropy tensor, eigen-decomposed. Eigenvalues are relative to the weakest direction (so the third is 0 and the first is the anisotropic ΔB), because only the deviatoric part is determined; eigenvectors are in the PDB orthogonalisation convention. Not written for a cubic Laue class, where symmetry forces ΔB to be zero |
_reflns.jfjoch_aniso_delta_B, _jfjoch_aniso_delta_B_linear |
The anisotropic ΔB, and the part of it that actually follows exp(−½ sᵀBs). The second is what the verdict is gated on |
_reflns.jfjoch_aniso_d_min_1..3 |
Diffraction limit (Å) along each principal direction. A comment marks a value that is the edge of the measured data rather than the crystal's own limit |
_reflns.jfjoch_aniso_shape, _jfjoch_aniso_floor, _jfjoch_aniso_significance, _jfjoch_aniso_verdict |
The resolution signature of the deficit, the data set's own systematic-error floor, ΔBlinear over that floor, and the resulting verdict. Each carries its vocabulary as a comment |
Compatibility note. Before rc.161,
_reflns.jfjoch_diffrn_ISacarried the asymptote, not the whole-range value. There is no version marker inside the file, so a number taken from an older.cifis not comparable with one taken from a newer one.
SHELX HKLF 4 (<prefix>.hkl). Fixed-format 3I4,2F8.2 — h k l I σ(I), one record per
reflection, terminated by a 0 0 0 record — which is what SHELXC, SHELXD and ANODE
expect. Two properties worth knowing before using it:
- Bijvoet mates are written separately,
I(+)at+hklandI(-)at-hkl, so the anomalous differences survive into SHELXC; a reflection with no anomalous split is written once, as its mean. - Intensities are rescaled by a single global factor so the largest value fits the
F8.2field.Iandσ(I)share that factor, so every ratio — and therefore the anomalous signal — is untouched, but the absolute scale is not meaningful. This matters only if you intend to compare magnitudes with another file; SHELXC and ANODE use ratios alone.
The unmerged export
--export-unmerged writes <prefix>_unmerged.mtz: every integrated observation, before scaling and
merging, in the column layout POINTLESS writes and aimless, pointless, careless and
iotbx.merging_statistics read. It is written in --mode mx and --mode scale alike, off by
default, and does not replace anything — rugnux still writes its own merged files in the same run.
Use it to scale the data with a different program, to have pointless give an independent opinion on the space group, or to compare rugnux's merge against another one on identical input.
Columns. H K L M/ISYM BATCH I SIGI FRACTIONCALC XDET YDET ROT LP FLAG — POINTLESS's own set —
plus four rugnux extras, DELPHI (offset from the centre of the rocking curve), ZETA (the Lorentz
geometry of that curve), BGMEAN and BGVAR (the background that was subtracted, and its variance).
BATCH is the image ordinal plus one, and a batch header is written for every batch that carries an
observation. M/ISYM records both the symmetry operation and the Friedel hand, so the index as
measured is recoverable from the index as stored.
What has been applied to the intensities, and what has not. I and SIGI carry the
Lorentz-polarization factor and nothing else; the factor itself is in the LP column, so raw
counts are I/LP. LP is applied because it is per-observation geometry that varies by more than two
orders of magnitude across a sweep and no reader can reconstruct it. Deliberately not applied:
the partiality is not divided out (it is reported in FRACTIONCALC), and the per-image scale is
not applied at all — those programs fit their own scale model, and handing them pre-scaled data
would have them fit a correction to a correction. No resolution cut, outlier rejection or ice-ring
filtering is applied either.
Partials. On a rotation run the partials of each reflection are summed into one full, using the
same rule rugnux's own 3D combine uses — consecutive frames no more than two apart — and the full is
written at the batch its rocking curve is centred on, with the summed rocking-curve fraction in
FRACTIONCALC. An event that caught less of its rocking curve than --min-partiality is not
written, exactly as in the merge. Summing is the default because a downstream program's own partial
handling is far more conservative than rugnux's: given raw partials, aimless accepted a small
fraction of the file and merged at a fraction of the multiplicity; given summed fulls it uses
essentially all of it. --export-unmerged-partials writes the unsummed form to
<prefix>_unmerged_partials.mtz for a program that would rather sum them itself. Stills have no
rocking events and are the same either way.
Systematic absences. Lattice-centring absences are not written; screw and glide absences are. Prediction runs in a primitive setting so that the space-group search can test the centring, but the interstitial reflections that leaves make a reading program take the lattice for primitive and demote the group. Screw and glide absences are kept because they are the evidence the space group was chosen on — deleting them would turn a reading program's test into an assumption. XDS and DIALS draw the line in the same place.
Scan axis. The batch headers carry the goniometer axis negated relative to the one in the
input file. This is not a correction to the file: rugnux brings an observation made at angle φ back
to zero by rotating it by +φ, so the crystal itself turns by −φ, and an MTZ batch header records the
axis a batch's own increasing PHI turns the crystal about. With the sign as exported, pointless's
independently determined orientation matrix agrees with rugnux's to well under a degree.
The results report
<prefix>_report.txt records what the run determined, next to the reflection files. It is
written on every --mode mx and --mode scale run that has an output prefix — there is no option
to enable or disable it. Two cases follow from that:
- An empty output prefix (
-o "", the "compute the statistics, persist nothing" mode) writes nothing, the report included. --no-mergestill writes a report. It determined an indexing and a geometry result, and those are recorded; the merging section then saysMERGE= NOT_PERFORMEDrather than being omitted, so the absence is a statement and not something a reader has to infer.
The report is never allowed to fail a run: if it cannot be written (unwritable path, full disk) the failure is logged as a warning and the run finishes normally.
Format
The model is XDS's CORRECT.LP: prose and tables a crystallographer reads top to bottom, with a
structure a script can consume without parsing prose.
KEY= valueassignment lines. Every number worth extracting is one, so a consumer gets it with a singlegrep '^ISA= 'and never has to read a sentence. Key names are stable.- Fixed-width tables with a stable header row for anything that is genuinely tabular — the resolution shells, the space-group candidates, the sweep-quality ranges.
WARNING:lines, one per finding, in plain English:WARNING: Frames 500-600 out of beam (10.1 deg, scale 0.12 and CC 0.30 of the run, 2% scaled).grep '^WARNING:'finds every one.- Section banners (
***…***around a numbered title) delimiting the blocks.
REPORT_VERSION= is the format's own version. Key names, table columns and the reason vocabulary
below are an interface other software may depend on: they do not change without that number moving.
Sections, in order: 1. DATA SET, 2. INDEXING, 3. GEOMETRY POST-REFINEMENT (rotation only),
4. SPACE GROUP DETERMINATION, 5. SCALING AND MERGING, 6. TWINNING, 7. RADIATION DAMAGE,
8. SWEEP QUALITY, 9. DIFFRACTION ANISOTROPY, 10. WARNINGS.
SPOT_RESOLUTION_ESTIMATE= in section 1 is how far the merged data are expected to reach, read
off the found spots alone — no lattice, no integration, no merge — so it is there on a run that never
merges, and on a run that does it can be read against INCLUDE_RESOLUTION_RANGE in section 5. It is a
prediction, good to about 0.2 Å on rotation data; nothing is cut on it.
Which pass. A rotation run integrates twice — once at the geometry in the input file, then again
at the post-refined geometry — and can integrate a third time if a guard rejects the second pass.
There is one report, for the pass that became the canonical output, and PASS= /
PASS_DECISION= in section 1 say which pass that is and on what evidence, so no number in the file
is ambiguous about which geometry produced it.
Not in the report: timing, frame rates, thread counts, per-image progress and library banners. Those are process, not result, and stay on stdout.
Sweep quality and the reason vocabulary
Section 8 lists the stretches of the sweep over which the crystal delivered much less than the rest of the run — the feedback a beamline control system needs to tell an operator that a crystal should be recentred or recollected. Nothing is excluded on the strength of it; the frames still carry signal, and this is a message for the beamline, not a filter.
SWEEP_QUALITY_STATUS= COMPUTED
SWEEP_QUALITY_COUNT= 1
SWEEP_QUALITY_REASONS= no_diffraction crystal_out_of_beam weak_diffraction loss_of_centring radiation_damage
SWEEP_ROTATION= 360.0
FLUX_PEAK_TO_TROUGH= 1.03
SCALE_MODULATION_PEAK_TO_TROUGH= 1.00
FIRST_IMAGE LAST_IMAGE N_IMAGES ROTATION REASON SEVERITY SCALE CC INDEXED
----------- ----------- --------- -------- -------------------- -------- ------ ------ --------
500 600 101 10.1 crystal_out_of_beam 0.83 0.12 0.30 0.02
----------- ----------- --------- -------- -------------------- -------- ------ ------ --------
SWEEP_QUALITY_STATUS distinguishes COMPUTED (the diagnostic ran; a count of 0 means the sweep
was clean throughout) from NOT_COMPUTED (it did not run — no scaling and merging, or stills
data). A consumer must not read a missing table or a zero count as "clean" without checking it.
SWEEP_QUALITY_REASONS lists the whole vocabulary this version can emit, so an unknown code is
distinguishable from a missing one.
| Reason code | Meaning |
|---|---|
no_diffraction |
The range recorded essentially no diffraction from the indexed lattice. |
crystal_out_of_beam |
Frames were lost: over the range a per-image scale could be fitted far less often than over the run. |
weak_diffraction |
The frames all still index, but with much less intensity — the cause was not determined. |
loss_of_centring |
One cycle of modulation per revolution: the crystal is off the rotation axis. |
radiation_damage |
The range runs to the end of a sweep whose quality was already decaying. |
The vocabulary is closed and stable: a code is never renamed, and never reused for a different
meaning. New codes are only ever added, and adding one moves REPORT_VERSION.
The columns are: FIRST_IMAGE/LAST_IMAGE — inclusive, in processed-image ordinals (the numbering
of <prefix>_image.dat and of every other per-image array rugnux writes; with -s/--stride the
source image is start + ordinal * stride); ROTATION — the width of the range in degrees;
SEVERITY — the fraction of the run's typical diffracting power missing over the range, 0 (as good
as the run) to 1 (nothing at all); SCALE and CC — the range's mean per-image scale and
CC-to-merge relative to the run median; INDEXED — the fraction of the range's frames that were
scaled at all. Every range also appears as a WARNING: sentence in section 9.
The same finding is written per image into the _process.h5 as /entry/MX/sweepQuality, when
one is written — see HDF5.
Diffraction anisotropy
Section 9 reports how much the fall-off with resolution depends on direction, and whether that is established above the data set's own systematic error. It runs automatically on every merging run — there is no flag — and it is a description only: no intensity is corrected, no reflection is removed on a directional criterion, and the merged data and the written reflection files do not depend on direction at all. The algorithm is in CPU/GPU data analysis ▸ Diffraction anisotropy.
Two different quantities are reported and they are not interchangeable. ANISOTROPY_DELTA_B is a
rate — the range of the principal components of the anisotropy tensor, on the ordinary
crystallographic B scale, so it is directly comparable with phenix.xtriage's B_cart, ctruncate's
anisotropic B and AIMLESS's anisotropic ΔB. ANISOTROPY_D_MIN_PRINCIPAL is where the signal
actually runs out along each principal direction. A crystal can have a large ΔB and almost no
spread in directional limit, or the reverse.
| key | meaning |
|---|---|
ANISOTROPY_VERDICT |
DETECTED | NOT_DETECTED | CANNOT_DETERMINE |
ANISOTROPY_FREE_DIRECTIONS |
Deviatoric directions the Laue class allows — 5 triclinic, 3 monoclinic, 2 orthorhombic, 1 tetragonal/trigonal/hexagonal, 0 cubic |
ANISOTROPY_DELTA_B |
The anisotropic ΔB (Ų), fitted on intensities with nothing dropped |
ANISOTROPY_DELTA_B_LINEAR |
The part of it that follows exp(−½ sᵀBs). This is the number the verdict is gated on, and the report says which of the two it is quoting |
ANISOTROPY_PRINCIPAL_B |
The three principal components, relative to the weakest |
ANISOTROPY_D_MIN_PRINCIPAL |
Diffraction limit (Å) along each principal direction — where ⟨I/σ(I)⟩ in a 20° cone about it falls through 2 |
ANISOTROPY_D_MIN_CENSORED |
One flag per direction. 1 means ⟨I/σ(I)⟩ never fell through 2, so the limit is the edge of the measured data, a bound and not a measurement. The prose marks it with a < |
ANISOTROPY_D_MIN_SPREAD |
Range of the three limits — itself a lower bound if any is censored |
ANISOTROPY_SHAPE |
LINEAR (a real Debye–Waller B) | FLAT (the deficit does not follow a B at all, so ΔB may be an under-estimate) | CONVEX (grows faster than a B can) | UNDETERMINED (the verdict moved on rebinning) |
ANISOTROPY_FLOOR, ANISOTROPY_SIGNIFICANCE |
The data set's own systematic-error floor (Ų) and ΔBlinear over it. Banded: below 2 not established, 2–3.5 marginal, above 3.5 established, above 5 strong |
ANISOTROPY_DETECTION_LIMIT |
The smallest ΔB that could have been established on these data. It is set by systematic error, not by counting, so it does not improve with more reflections or a longer exposure |
ANISOTROPY_N_OBSERVATIONS, ANISOTROPY_FORBIDDEN_Z, ANISOTROPY_SIGMA_SYSTEMATIC |
The unmerged observations the floor was measured on, that measurement against its own counting noise, and the floor before the counting part is added back |
CANNOT_DETERMINE is a real answer, not an evasion. The verdict is not measured against counting
statistics — real data carry systematic error far larger than that, and gating on counting error
reports anisotropy on data sets that have none. Instead the data set measures its own systematic
error in the tensor directions its Laue class forbids, where the true value is exactly zero
whatever the crystal is. Where that measurement cannot be made, the run says so and gives the
reason: a triclinic Laue class (no forbidden direction exists), an observed rotation under about
90°, merged data at the noise floor, a scale model carrying no dose term
(--no-scaling-corrections), or no unmerged observations. A cubic Laue class is different again
— symmetry forces ΔB to be exactly zero, and the run says that rather than reporting a measurement.
Where anisotropy is detected and the directional limits differ by more than 0.5 Å, a WARNING: line
says so, since refinement and map interpretation should allow for it.
Reference data and the indexing ambiguity
What a reference MTZ does (-z)
-z reference.mtz supplies known intensities of the same crystal form — a previously merged
dataset, or F-model amplitudes computed from a structure. It is read once, before processing
starts, and used for four things:
- It fixes the space group and the unit cell the run works in, unless
-S/-Coverride them. The cell is a soft reference: indexing may still drift within tolerance, so a small mismatch between reference and data is absorbed rather than rejected. - It resolves the indexing ambiguity (below) — the one thing the data cannot settle for themselves.
- It hands over its R-free test set, where the file carries one, so every dataset of a campaign is scored on the same free reflections.
- It reports CCref, the correlation of the merged intensities against the reference,
in the statistics table. Stills only — the rotation merge never scores itself against the
reference, and its table shows
nanin that column.
A reference is not a scale anchor. Both workflows scale against their own data — scaling images
against a foreign dataset injects that dataset's systematics — so -z never puts the reference's
errors into the intensities. --reference-column picks the column to read where the automatic
choice (F-model, else IMEAN/I, else FP/FOBS/F) is not the right one.
For the second of those four jobs — and only that one — an atomic model does as well: --model
computes the intensities it needs from the structure. Where a reference dataset exists, prefer it;
where only a model does, it resolves the ambiguity just the same.
The indexing ambiguity
Some crystals can be indexed in more than one way, each equally valid geometrically, and each
giving different merged intensities. This happens whenever the lattice is more symmetric than the
crystal: in P3, P4, P6, P31, C2 and their relatives (merohedral), and also where the
cell is metrically more symmetric than the Laue class by accident (pseudo-merohedral, up to 2° of
obliquity). The alternatives are related by the crystal's twin laws — reindexing operators such
as k,h,-l.
Nothing in the data breaks the tie: the merge is equally self-consistent either way, so which solution comes out is arbitrary. What that costs depends on the workflow:
- Rotation. The whole sweep is one lattice, so the whole dataset lands in one indexing, picked at random. The merge itself is sound; it may simply be the other solution from an earlier dataset of the same crystal form, and the two cannot be combined, compared or phased against the same model.
- Serial stills. Every crystal is indexed independently, so a run mixes both indexings into one merge. That is not a labelling matter — reflections that are not symmetry mates get averaged together, and CC1/2, Rmeas and the anomalous signal all degrade.
Every run that merges tests for the ambiguity, and where it exists and nothing resolves it, says so in the log and in the report's warnings:
Indexing ambiguity: this cell / space group admits alternative indexing (reindex operator(s):
-h,-k,l). Serial-stills crystals are indexed in one hand at random, and rugnux can only break this
against an external reference. WITHOUT one the merge mixes the hands and CC1/2 is degraded - supply
a reference MTZ (-z) or a model (--model, which needs -C and -S here) to resolve it.
Where something does resolve it, the run says so instead — Indexing ambiguity present (reindex operator(s): -h,-k,l); resolved against the supplied model.
How to resolve it:
| Situation | What to do |
|---|---|
| Rotation, a reference dataset exists | -z reference.mtz. Once the space group is settled, each candidate reindexing of the merged intensities is correlated with the reference and the best-correlating one is re-merged. Only the hkl labels change; the cell does not |
| Serial stills, a reference dataset exists | -z reference.mtz. Resolved per image, at integration time, by correlating each crystal's intensities with the reference — so the merge never mixes hands in the first place |
| Rotation, only a model | --model model.pdb. The merged data are fitted to the model in each candidate indexing and the lowest R-free wins; the written reflections are then reindexed into it, so the file, the R-factors and the maps agree. The log gives the winning R-free and the runner-up — a narrow margin means the data did not really decide |
| Serial stills, only a model | --model model.pdb, with the cell and space group given (-C / -S). Structure factors are computed from the model up front and used as the reference for the per-image test, exactly as a reference MTZ would be — the ambiguity has to be broken at integration time, and a model can supply the intensities to break it with |
| Neither | The run warns and merges in whichever indexing it found: for rotation data a usable dataset in an arbitrary frame, for stills a degraded one |
Two things to know about where the choice lands:
--mode scalecannot repair a stills run after the fact. The per-image test happens at integration time, so a_process.h5whose images were integrated without a reference — an MTZ or a model — has already lost the distinction, and no re-merge brings it back. On rotation data, where the ambiguity is one choice for the whole dataset,--mode scale --model model.pdbdoes resolve it.- The choice reaches the written reflections, not only the R-factors and the maps: the merged
.mtz/.cif/.hkl(and_unmerged.mtz, where it is asked for) are written in the indexing the model or the reference settled on, so the file can be refined against that model as it stands.
Two things the indexing ambiguity is not:
- Not the enantiomorph. P41212 versus P43212 (or
P31 versus P32) leaves the merged intensities unchanged, so no amount of
data can choose between them and rugnux never tries — the run reports the pair it cannot separate.
A model settles it, being the only evidence there is: with
--modelthe written reflections take the model's hand and its space group, which on anomalous data exchanges I(+) and I(-). The_process.h5, whose per-image reflections were written as they were integrated, keeps the group the run determined and is left alone. - Not twinning. The twin laws are the same operators, but twinning is a property of the crystal — two orientations diffracting at once — and is reported separately in the report's twinning section. A crystal can have an indexing ambiguity without being twinned, and usually is.
The algorithms behind both are in CPU/GPU data analysis ▸ Reference data.
Validating against a model (rugnux --model)
Given a PDB atomic model of the same structure, --model model.pdb scales the model structure
factors to the merged amplitudes — fitting a flat bulk-solvent contribution and an overall
anisotropic B — and reports R-work / R-free and the mean 2Fo-Fc density at the atom centres.
It also writes <prefix>_2fofc.ccp4, <prefix>_fofc.ccp4 and <prefix>_maps.mtz next to the
merged reflections. Nothing about the model is refined; it is only re-fractionalized into the data
cell, so a deposited model with a slightly different cell still lines up.
It is a data-quality lens, independent of the internal statistics: R-free measures the merged
intensities against external truth, where CC1/2 and Rmeas only measure them against
themselves. It also settles the two things merged intensities alone cannot: the enantiomorph (data
merged in P41212 against a P43212 model are reindexed
into the model's hand), and — when no reference MTZ has already fixed it — a merohedral
indexing ambiguity, by keeping the candidate reindexing with the lowest
R-free. Both of those reindexings are applied to the written reflections as well as to the
R-factors and the maps — validation runs before the reflection files, so the .mtz / .cif /
.hkl come out in the model's frame, and where the hand was adopted they carry the model's space
group. The log names the operator in each case, and for the indexing choice gives the winning
R-free together with the runner-up.
Re-scaling and re-merging (rugnux --mode scale)
The scale mode re-scales and merges the already-integrated reflections stored in a
_process.h5 file, without re-running spot finding or integration. Use it to re-merge quickly with a
different space group, resolution limit, anomalous setting or outlier rejection. It reuses the same
-o/-N/-s/-e/-S/-A/-B/-z/--scaling-* options as the full run, and (unlike the full pipeline) does
not run a space-group search: it merges in the space group and unit cell the file records, and -S
/ -C override them. A _process.h5 written before the group was stored carries none, and merges
in P1 unless -S says otherwise.
A reference MTZ (-z) is accepted here on stills data, where it fixes the space group and cell,
reports CCref and hands over its R-free flags; on rotation data the rotation scaler
declines it and the run stops with a message saying so. What --mode scale never does is reindex
the reflections it writes: an indexing ambiguity has to be resolved by
the run that integrates (--model still picks an indexing for its own R-free and maps here, as in
a full run).
Where the full run re-seated the lattice — the space group it settled on is in a different setting
from the one each image was indexed in — the file records the change of basis as
/entry/MX/reindexMatrix, and rugnux applies it on read, so the reflections and the stored cell
describe the same frame. An older file that was affected by this cannot be repaired (the matrix is
not recoverable after the fact); such a file now stops with a message naming both cells and the
exact -S/-C override to merge it in its own setting, instead of failing inside the merge.
Detector calibration from powder rings (rugnux --mode calibration)
The calibration mode determines the detector geometry — PONI x/y, the two tilts
rot1/rot2 and the distance — from the powder rings of a calibrant, and writes it as a
pyFAI <prefix>.poni file alongside a printed report of how far each parameter moved from
the header. Bragg data pin the beam centre worst (it is gauge-coupled to the crystal orientation);
a powder ring has no orientation to be coupled to, so this is the measurement that fixes it.
rugnux --mode calibration --calibrant lab6 -o det LaB6_master.h5
--calibrant takes lab6, agbh (silver behenate), ceo2, si or ice, case-insensitively.
ice calibrates a real experiment against its own ice rings — no calibrant exposure needed —
and is the reason a calibrant is a list of ring positions rather than a unit cell: hexagonal ice
is P63/mmc, so rings enumerated from its cell would include systematically absent ones.
--calibration picks how the rings are measured, and both use the whole dataset — -s/-e/-t
select which images:
rings(default) sums the (q × azimuth) azimuthal profile over every processed image into one map and fits the ring arcs in it. A powder ring is an arc, not a set of spots, and the summed profile measures it at every azimuth with all the run's counts behind it. It needs the profile to be binned in azimuth, so this mode defaults--azim-phi-binsto 32.spotspools the found spots of every processed image and fits those. It determines the centre from scratch (a Hough circle vote, which quantises it to a whole pixel) and then refines.
Both routes read the ring position out of a binned profile or a spot centroid, so the radial
sampling matters: at a long detector distance the default 0.01 Å⁻¹ q bin is several pixels wide
and quantises the rings route accordingly — pass a finer --azim-q-spacing there (the total
q × azimuth bin count must stay under 65534).
The report prints the fitted geometry, the scatter of the ring points about the fitted rings and the standard error that implies on the centre. That error is formal: it measures the scatter of the points, not whether the rings themselves are trustworthy, so it stays small when a fit goes wrong for a structural reason — one visible ring, or ice that is textured rather than smooth.
Both the PONI (the point of normal incidence, which is what a .poni file stores) and the direct
beam (where the beam lands, which is what most other programs call the beam centre) are printed.
They differ by distance × tan(rot) once the detector is tilted, which on a 0.3° tilt at 300 mm is
several pixels — enough to look like a disagreement with another program when there is none.
Comparing the geometry with XDS
Every run logs the detector geometry a second time in XDS's convention, so it can be read
straight across against the IDXREF.LP / CORRECT.LP of an XDS run on the same data:
XDS convention: ORGX= 1091.00 ORGY= 1137.00 DETECTOR_DISTANCE= 75.0000
XDS convention: DIRECTION_OF_DETECTOR_X-AXIS= 1.000000 0.000000 0.000000
XDS convention: DIRECTION_OF_DETECTOR_Y-AXIS= 0.000000 1.000000 0.000000
XDS convention: INCIDENT_BEAM_DIRECTION= 0 0 1 X-RAY_WAVELENGTH= 1.000000 QX= QY= 0.075000
XDS convention: ROTATION_AXIS= -1.000000 0.000000 0.000000
XDS is never given this geometry — the XDS plugin supplies image data
only, and XDS refines its own from XDS.INP — which is what makes the comparison worth having. The
two laboratory frames coincide (x along increasing detector column, y along increasing row, z along
the beam), so the numbers are directly comparable, and a tilt appears as the two detector axis
vectors rather than as angles, which is how XDS reports it after refinement. Two things to keep in
mind: ORGX/ORGY are 1-based, because XDS counts pixels from 1 and Jungfraujoch from 0; and
they are the PONI, the same quantity Jungfraujoch's beam centre is — so no correction is needed —
but not the direct beam once the detector is tilted (see above).
Command-line options
General:
| Option | Description |
|---|---|
-o, --output-prefix <txt> |
Output file prefix (default: output) |
-N, --threads <num> |
Number of worker threads (default, and for any value ≤ 0: all hardware threads). Some stages take fewer, because past a point more workers make them slower: the per-image loop of --mode mx uses at most 16 per GPU unless -N was given a positive value, and first-pass spot finding and the beam-stop pre-scan have ceilings of their own that -N does not lift. Scaling, merging and the space-group search use the full count |
-s, --start-image <num> |
First image to process (default: 0) |
-e, --end-image <num> |
Last image to process (default: all) |
-t, --stride <num> |
Process every n-th image (default: 1) |
-v, --verbose |
Verbose output |
Mode — --mode <name> (default mx):
| Value | Description |
|---|---|
mx |
Full analysis — spot finding, indexing, integration and merging |
azint |
Only azimuthal integration (no spot finding/indexing); writes <prefix>_process.h5 |
scale |
Only re-scale/merge the already-integrated reflections in the input _process.h5 (no re-integration) |
calibration |
Determine the detector geometry from powder rings; writes <prefix>.poni |
Calibration (--mode calibration):
| Option | Description |
|---|---|
--calibrant <name> |
Powder standard: lab6 | agbh | ceo2 | si | ice (default lab6, case-insensitive) |
--calibration <txt> |
How the rings are measured: rings | spots (default rings; see above). rings defaults --azim-phi-bins to 32 |
Detector mask:
| Option | Description |
|---|---|
--detect-beam-stop[=N|off] |
Find the beam stop and its holder in a projection of N images and add them to the pixel mask as bit 9, so nothing shadowed by them is integrated. On by default (60 images); =off disables. Reflections behind the stop are attenuated but not flagged, so they integrate low with a plausible sigma and no existing rejection catches them |
Geometry:
| Option | Description |
|---|---|
--estimate-beam-center |
Measure the direct beam before indexing, from the symmetry of the spots where the sweep reaches at least half a turn and from the radial background profile where it does not; the value in the file is kept where neither can measure it. Off by default |
--no-fit-spindle |
With the above, keep the rotation axis given in the file instead of fitting its skew about the beam |
Spot finding:
| Option | Description |
|---|---|
--spot-sigma <num> |
Noise sigma level for spot finding (default: 4.0) |
--spot-threshold <num> |
Photon-count threshold for spot finding (default: 10) |
--adaptive-spots |
Self-calibrating detection (default, stills and rotation alike): the strong-pixel threshold comes from each image's own per-resolution-ring noise instead of the fixed --spot-threshold, so one setting adapts across datasets (no per-dataset --spot-threshold/--spot-sigma tuning) |
--no-adaptive-spots |
Turn adaptive detection off and use the fixed --spot-threshold / --spot-sigma finder |
--spot-false-pixels <num> |
Adaptive-detection operating point: expected noise pixels tolerated per frame (default: 100; implies --adaptive-spots) |
--spot-high-resolution <num> |
High-resolution limit for spot finding, Å. Omitted (or 0): no resolution clipping — spot finding extends as far as the detector reaches, for rotation data as well as stills |
--spot-low-resolution <num> |
Low-resolution limit for spot finding, Å (default: 50; lower it, e.g. 24, to exclude the direct-beam halo on weak serial data; 0 removes the limit) |
--min-pix-per-spot <num> |
Minimum connected strong pixels per spot. If omitted, min-pix is chosen per image (stills indexing): the frame is indexed at min-pix 3/2/1 and the one maximising indexed-spot count × indexed fraction is kept. Give an explicit value to force a fixed min-pix instead. |
--max-spots <num> |
Maximum spots kept per image (the strongest ones) and handed to indexing (default: 1000) |
--detect-ice-rings[=on|off] |
Flag ice-ring spots (de-prioritised in indexing) and exclude ice-ring reflections from scaling. Default: the master file's detect_ice_rings, or — where the file carries no such key — on for rotation and off for stills |
Azimuthal integration (the radial profile behind the per-image ice-ring score):
| Option | Description |
|---|---|
-q, --azim-q-spacing <num> |
Q bin spacing, 1/Å (default: 0.01; finer resolves the narrow ice rings) |
--azim-min-q <num> |
Minimum Q, 1/Å |
--azim-max-q <num> |
Maximum Q, 1/Å. Omitted: integration extends to the highest Q the detector reaches. The adaptive spot finder shares these Q bins, so this also sets how far self-calibrating detection can see |
--azim-phi-bins <num> |
Number of azimuthal (phi) bins (default: 1) |
--polarization-correction <on|off> |
Enable/disable the azimuthal polarization correction |
--solid-angle-correction <on|off> |
Enable/disable the azimuthal solid-angle correction |
Indexing:
A dataset with a rotation goniometer axis is processed as rotation data (two-pass rotation
indexing) by default; a dataset without one is processed as independent stills. --force-still
overrides the former; the -R / --single-pass-rotation / --force-rotation-lattice flags request
rotation explicitly and pick the pass or lattice.
| Option | Description |
|---|---|
--force-still |
Treat a rotation (goniometer) dataset as independent stills instead of rotation |
-X, --indexing-algorithm <txt> |
FFBIDX | FFT | FFTW | Auto | None |
-C, --unit-cell <cell> |
Reference unit cell "a,b,c,alpha,beta,gamma" (required by ffbidx) |
-S, --space-group <num|symbol> |
Space group number (92) or Hermann-Mauguin symbol (P43212) — for indexing and scaling |
-r, --refine <txt> |
Geometry refinement: none | orientation | beam_and_lattice (default) | flex (try all three per image, keep whichever indexes the most spots; alias multi) |
-R, --two-pass-rotation[=num] |
Two-pass offline rotation indexing (default for goniometer data; optional first-pass image count, default 100) |
--single-pass-rotation[=num] |
Online-like single-pass rotation indexing (optional min angular range, deg) |
--force-rotation-lattice <vec> |
Force rotation lattice (9 floats, Å), skipping the first pass |
--rotation-no-postrefine |
Rotation: disable the default-on two-pass geometry post-refine (see the rotation section) |
--refine-geometry[=N|off] |
Stills: extra first pass that bundle-adjusts the shared beam/distance/cell from N strongly-indexed frames (default 200) then re-indexes; default ON for stills with a reference cell (-C / -z), =off disables |
--index-ice-rings[=on|off] |
Index on the spots flagged as sitting on an ice ring too, instead of setting them aside (default: off; no effect without --detect-ice-rings, which does the flagging) |
Indexer choice in brief: ffbidx (GPU) refines toward a known cell and is best for sparse
serial stills; fft (GPU) / fftw (CPU) index de novo and suit strong rotation data. See the
CPU/GPU data-analysis reference for the algorithms.
Scaling and merging:
| Option | Description |
|---|---|
--no-merge |
Skip scaling and merging (on by default); write only the per-image _process.h5 |
-A, --anomalous |
Anomalous mode (keep Friedel pairs separate) |
--scale-fulls / --no-scale-fulls |
rot3d: refit a per-frame scale on the combined fulls (XDS order, Unity model); on by default for rotation data, off for stills |
--smooth-g[=deg] |
rot3d: smooth the per-frame scale G over a degree range before the 3D combine (XDS DELPHI-like; default 5° for rotation, 0 = off) |
--no-scaling-corrections |
rot3d: disable the default-on decay + absorption + modulation correction surfaces fitted on the fulls after scale-fulls (see below) |
--relative-b[=deg] |
rot3d: fit a per-batch relative-B beyond the single decay slope over deg-degree batches, cross-validated (default 10° when bare; off otherwise) |
--simple-stills |
Stills: treat every reflection as a full (p = 1, single-pass scale/merge) — disables the default-on physical partiality post-refinement |
--no-expected-variance-merge |
Stills: disable the default expected-variance merge weighting (which rebuilds each weak observation's signal variance at the reflection mean to de-bias the inverse-variance merge); restores observed-sigma weighting |
--capture-uncertainty <num> |
rot3d: systematic sigma on under-captured fulls, ~num·(1−captured_fraction)·I (default: 1.0 for rotation, 0 otherwise) |
--min-captured-fraction <num> |
rot3d: drop a combined full whose rocking curve was captured below this fraction — edge-of-sweep truncated fulls (default: 0.7 for rotation, 0 otherwise; 0 = off) |
--scaling-high-resolution <num> |
High-resolution limit for scaling, Å — manual override (default: no limit; disables the automatic cutoff below) |
--scaling-low-resolution <num> |
Low-resolution limit for scaling and merging, Å (default: 50, the value XDS configurations use; 0 removes the limit). Reflections coarser than this sit behind or beside the beam stop and are measured on a background it has eaten into |
--resolution-cutoff <txt> |
Automatic high-resolution cutoff for the written reflections and reported shells: cc-logistic | off (default: cc-logistic; ignored when --scaling-high-resolution is set) |
--resolution-cc-target <num> |
CC1/2 target defining the cc-logistic fall-off (default: 0.30) |
--resolution-shells <num> |
Number of resolution shells in the reported statistics table (default: 9). The bins are equal steps in 1/d² between the lowest- and highest-resolution reflection merged, which is XDS's rule, and 9 is XDS's count — so at the same resolution limits the two tables have the same shells and can be read row for row |
--min-partiality <num> |
Minimum partiality to accept a reflection (default: 0.02) |
--ice-min-score <num> |
Ice-presence gate: the measured per-run ice score (1 = no ice) a dataset must reach before any ice handling is applied — the flagging and the exclusion from scaling (default: 1.5; 0 = no gate). The eleven fixed hexagonal bands cover 16–26 % of the unique reflections whether or not the crystal has ice, so handling ice on a clean crystal only costs completeness |
--ice-min-spot-ratio <num> |
The second ice-presence channel: found spots on the hexagonal rings over the same q width of ice-free flanks beside them (1 = spots spread evenly). Ice in large crystallites diffracts as discrete spots and leaves the radial profile flat, so --ice-min-score alone is blind to it (default: 2.0; 0 disables this channel) |
--reject-outliers <num> |
Per-observation outlier rejection, N σ from the per-reflection median (default: 6 for rot3d, off otherwise) |
--min-image-cc <num> |
Per-image CC limit, percent (default: no limit) |
--search-min-zeta <num> |
De-novo space-group search only: also search a merge of just the observations whose Lorentz geometry |ζ| reaches this, and report both answers (default: 0.85 for rotation, 0 = single search). Reflections crossing the Ewald sphere near-tangentially are measured worst and can make a real symmetry operator look like a twin law. Where the two searches disagree, the merge of all the observations decides — as it always has for the systematic absences |
--mosaicity <num> |
Diagnostic: fix the scaling mosaicity (°) instead of using the per-image seed |
--scaling-iterations <num> |
Scaling iterations with no reference data (default: 3) |
-z, --reference-mtz <file> |
Reference MTZ of the same crystal form: fixes the space group and cell, resolves the indexing ambiguity, hands over the R-free set and reports CCref. Not a scale anchor |
--reference-column <label> |
Reference MTZ column to use (default: auto — F-model, else IMEAN/I/…) |
--model <file.pdb> |
Validate the merged intensities against this atomic model — R-work / R-free and maps (see Validating against a model). It also settles the frame the reflections are written in: the enantiomorph, and the indexing ambiguity where no -z did. For serial stills given -C / -S, the model's structure factors become the per-image reference |
--write-process-h5 |
Also write the (large) _process.h5 when merging (default: only .mtz/.cif) |
--export-unmerged |
Also write the integrated observations as <prefix>_unmerged.mtz, an unmerged MTZ (POINTLESS column layout) for aimless / pointless / careless. Rotation partials are summed into one full per reflection. Intensities carry the Lorentz-polarization factor and nothing else — the partiality is not divided out and the per-image scale is not applied. Works in --mode mx and --mode scale. See The unmerged export |
--export-unmerged-partials |
As above, but writes <prefix>_unmerged_partials.mtz with each partial as its own row (one batch per image) for the reading program to sum |
Integration:
| Option | Description |
|---|---|
--integrator <txt> |
Spot integrator: gaussian (profile-fit, default) | empirical | boxsum (classical fallback) |
--integration-radius <r> |
Signal-box radius r1, or r1,r2,r3 (px). One value ⇒ r2=r1+2, r3=r1+4 |
--adaptive-integration-radius[=on|off] |
Set the signal radius r1 from how wide this crystal's spots actually are (default: on for rotation, off for stills). r1 is not the integration domain — that is the profile-fit grid — but it is the aperture the profile width is learned over, and a second moment over a disk of radius a saturates at a²/4, so at the shipped r1 = 4 the learned σ can never exceed 2 px and a broader spot is fitted with a profile the model cannot represent. The pre-scan reads r80, the radius holding 80 % of a spot's flux, off isolated strong spots over a fixed 14 px aperture that owes nothing to r1, fits it against 1/d and evaluates it at 5 Å; then r1 = clamp(round(2·r80), 4, 6), r2 = r1 + 2, and r3 is taken so the r2..r3 background ring keeps the area it has at the default 4,6,13. The ceiling of 6 is pattern density: r2 also drives the neighbour-ownership radius and the ring's inner edge, and past it a dense pattern starts losing reflections whose ring falls below six clean pixels. Ignored when --integration-radius is given. The widened radius applies to the final integration pass only — the two-pass geometry pre-pass keeps the radius the run started with, because the post-refinement fits the detector distance and beam to the observed reflection positions and those move with the signal disk (33 µm and 0.02 px between r1 = 4 and r1 = 6 on one crystal, enough for the second pass's de-novo lattice search to settle on a different lattice and index a fifth fewer frames). And where the widened radius leaves more than 1.1 % of the predicted reflections without a background ring — a pattern too dense for it — the final pass is re-integrated at the fixed 4 px radius, and the log says so. Over the rotation regression battery it moves 12 of 38 crystals and leaves the merged intensities of the other 26 unchanged; where it moves them, per-shell ⟨I/σ⟩ improves by up to 31 % and R_meas by up to 24 % |
--integration-stencil <k> |
Push the r2..r3 background ring out by k times the beam's radial streak bandwidth·R_px, per reflection (default 0 = the fixed circular ring). A fixed ring otherwise ends up on a streaked reflection's own tails at high resolution and measures them as background. Only the ring moves, and only radially — the r1 signal box stays a circle — and the growth is capped at 2·r3. The neighbour exclusion grows with it, so on a crowded pattern a few reflections can be left with too little background and dropped. Needs --bandwidth: on a monochromatic beam the streak is zero and this does nothing |
--background-clip <n> |
Monochromatic (rotation + still): high-side clip of the background ring at mean + n·√mean (default 4; 0 = off). The default background estimator — it rejects neighbour cores and zingers without the symmetric trim's Poisson skew bias. Broadband data always clip, at 3σ; ignored by --integrator boxsum |
--background-trim <f> |
Use the old symmetric trimmed mean for the background ring instead of the clip, 0≤f<0.5 (0.10 was the former default). Switches --background-clip off. A symmetric trim is biased low on Poisson data and adds ~5 counts to every partial, so this is for back compatibility only; 0 = plain ring mean. Rings holding more than 512 pixels fall back to the plain mean (the GPU sorts the ring in shared memory and the CPU now matches it), which the default radii never reach but wide ones do |
--background-radial[=on|off|auto] |
Correct the background ring for the curvature of the radial background (default off). Disk and ring are concentric, so a background linear in position cancels between them and only curvature survives — which on a smooth ice ring reaches +26 counts on a single reflection. auto applies it per image where that image's ice score shows a smooth powder ring, since the model is a function of radius alone: on ice made of discrete crystallite spots there is no smooth ring and the correction makes the bias worse. Ignored by --integrator boxsum (no clip pass to take the curve from) |
--integration-high-resolution <num> |
High-resolution limit for prediction and integration. Omitted (or 0) means integration extends as far as the detector reaches — which is what the predictor can place on the detector anyway, since it rejects reflections that miss it. Set a value to integrate less than the detector offers |
--max-hkl <n> |
Predict reflections with |h|,|k|,|l| ≤ n (max 511). By default this is derived per crystal from the refined cell as ceil(max(a,b,c)/d_min) + 1, which is the exact bound: the predictor keeps only |q| ≤ 1/d_min and h = a·q, so no reflection can lie outside it and no candidate inside it is wasted on a shorter axis. Set it only to override that |
--bandwidth <num> |
Relative X-ray bandwidth FWHM (e.g. 0.01 for a 1% DMM); default from file or 0 (monochromatic) |
--overlap <txt> |
What to do where two predicted reflections share signal pixels: off | reject | exclude (default exclude). A shared pixel belongs to the nearer centre; without this a crowded reflection reads high on a dense pattern. exclude drops the shared pixels from the profile fit, which renormalises itself, and keeps the reflection; reject instead drops the whole reflection when too little of its profile is cleanly its own. --integrator boxsum has no profile to renormalise, so only reject acts there |
--overlap-minpk <f> |
Least fraction of a reflection's expected profile that must be usable for it to be kept (default 0.75, XDS MINPK). Governs both the fraction that must be readable — not masked, untrusted, in a gap or overloaded — in every profile mode, and, under --overlap reject, the fraction that must be cleanly its own. Under --integrator boxsum any unreadable pixel discards the disk outright and the reject fraction goes by disk area, which cuts harder |
--prediction-mosaicity <num> |
Diagnostic: fix the rocking width (deg) the prediction window opens to, leaving partiality on the per-image σ_M. The two are one number by default, so a σ_M that moves takes the integrated reflection population with it |
Geometry overrides (defaults are taken from the input file; override them to reprocess with a corrected geometry):
| Option | Description |
|---|---|
--beam-x <num> |
Beam centre X (pixel) |
--beam-y <num> |
Beam centre Y (pixel) |
--detector-distance <num> |
Detector distance (mm) |
--wavelength <num> |
Wavelength (Å) |
--rot1 <num> |
PONI detector rotation 1 (rad) |
--rot2 <num> |
PONI detector rotation 2 (rad) |
--polarization <num> |
Polarization factor |
--rotation-scale <k> |
Goniometer rotation scale: the stage turned k times the angle stored in the file (the commanded one). Applied to both passes, and overrides the scale rugnux fits for itself |