For stills indexing the minimum-pixels-per-spot filter is now chosen per image instead of being fixed: the frame is indexed at min-pix 3/2/1 and the setting that maximises indexed-spot count weighted by indexed fraction (n_indexed^2 / n_total) is kept, then integrated once at that min-pix. The fraction factor keeps a smaller min-pix's extra spots only when the lattice actually explains them, so strong frames retain their real weak spots (extending resolution) while noise-flooded frames stay strict. The mode is selected by the presence of --min-pix-per-spot, now optional (SpotFindingSettings::min_pix_per_spot is std::optional<int64_t>): omit it for the adaptive per-image path, give a value to force a fixed min-pix. It applies only to the stills indexing path -- rotation indexing builds one global lattice and keeps a fixed min-pix, and the online receiver and the FPGA host path always carry a concrete value, so neither changes. IndexAndRefine::ProcessImage now returns whether the frame indexed, to drive the per-image selection. Exposed in the jfjoch_viewer spot-finding settings (adaptive-threshold and adaptive-min-pix checkboxes, each greying out the control it overrides); the broker uses neither. Validated on the full rotation regression battery (no regression) and the whole serial-stills target battery at full image count. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
286 lines
19 KiB
Markdown
286 lines
19 KiB
Markdown
# rugnux
|
||
|
||
`rugnux` is the **offline** crystallographic data-analysis tool of Jungfraujoch — the
|
||
data-processing half of the system (see [Naming](NAMING.md) for where the name comes from).
|
||
It takes an existing HDF5 dataset, runs the full analysis pipeline — spot finding, indexing,
|
||
geometry refinement, Bragg integration and (optionally) scaling and merging — and writes the
|
||
results to a `_process.h5` file, plus reflection files (`.mtz`/`.cif`/`.hkl`) when merging is
|
||
requested.
|
||
|
||
It runs the *same* analysis code as the online and interactive tools, just driven from the
|
||
command line over a file rather than a live detector stream.
|
||
|
||
> **Note.** `rugnux` is under very active development. This page describes the tool and
|
||
> its options at a high level; the authoritative, always-current list of options is the program's
|
||
> own usage message — run `rugnux` with no arguments.
|
||
|
||
## Where it fits among the three analysis tools
|
||
|
||
| Tool | Mode | Driven by | Output |
|
||
| --- | --- | --- | --- |
|
||
| [`jfjoch_broker`](JFJOCH_BROKER.md) | Online, real-time streaming analysis on FPGA + GPU | HTTP/REST + ZeroMQ | Live results and statistics, images streamed to [`jfjoch_writer`](JFJOCH_WRITER.md) |
|
||
| [`jfjoch_viewer`](JFJOCH_VIEWER.md) | Interactive, on-screen exploration | Qt desktop application | Displayed on screen (results not saved to disk) |
|
||
| **`rugnux`** | **Offline batch processing of a stored dataset** | **Command-line interface** | **`_process.h5`, and `.mtz`/`.cif`/`.hkl` when merging** |
|
||
|
||
Use `rugnux` to re-analyse data after acquisition, to experiment with processing
|
||
parameters, or to produce merged intensities for downstream structure solution.
|
||
|
||
## Hardware
|
||
|
||
As with the rest of Jungfraujoch, **serious performance requires an NVIDIA GPU**. The CUDA build
|
||
provides the GPU fast-feedback indexer (`ffbidx`) and the GPU FFT indexer (`fft`); without CUDA
|
||
only the CPU `fftw` indexer is available. Spot finding, integration and scaling run on the CPU and
|
||
scale with the thread count (`-N`).
|
||
|
||
## Input and output
|
||
|
||
**Input** is a single Jungfraujoch HDF5 master file (NXmx-based). If the dataset already contains
|
||
stored spot lists, two-pass rotation indexing can reuse them instead of re-running spot finding on
|
||
the first pass.
|
||
|
||
**Output** (controlled by `-o, --output-prefix`, default `output`):
|
||
|
||
- `<prefix>_process.h5` — NXmx-compliant HDF5 with derived metadata (spots, indexing,
|
||
integration, azimuthal integration, per-image statistics). See
|
||
[HDF5 / NeXus data format](HDF5.md) for the layout. Written by default only when **not** merging
|
||
(i.e. under `--no-merge`); add `--write-process-h5` to also write it when merging.
|
||
- Merging is **on by default** (`--no-merge` disables it). The merged reflections are written in
|
||
**three** formats — each has its uses downstream:
|
||
- `<prefix>.mtz` — CCP4 MTZ (`IMEAN`/`I(+)`/`I(-)`, French–Wilson `F`, `FreeR_flag`) for the CCP4 /
|
||
phenix reflection tools.
|
||
- `<prefix>.cif` — mmCIF, for deposition and as the self-describing native format (also carries the
|
||
merging statistics, ISa, twinning and radiation-damage indicators).
|
||
- `<prefix>.hkl` — SHELX **HKLF 4** text (`h k l I σ(I)`, fixed `3I4,2F8.2`), the direct input for
|
||
**SHELXC / ANODE / SHELXD**. Bijvoet mates are written separately (`I(+)` at `+hkl`, `I(-)` at
|
||
`-hkl`) so the anomalous signal is preserved; intensities are put on a common scale so the largest
|
||
value fits the fixed-width field (the absolute scale is irrelevant to SHELXC/ANODE), and the file
|
||
ends with the `0 0 0` terminator record.
|
||
|
||
All three carry the **refined unit cell** (from rotation indexing) and the **space group determined
|
||
from systematic absences** (constrained to the indexed lattice symmetry). No-reference scaling
|
||
additionally emits per-iteration `<prefix>_iterN_scale.dat`.
|
||
|
||
Merged statistics (⟨I/σ⟩, CC1/2, completeness, …), the error model and timing are printed to the
|
||
console. By default the written resolution is trimmed automatically where CC1/2 falls off
|
||
(`--resolution-cutoff cc-logistic`, CC1/2 target 0.30); set `--scaling-high-resolution` to fix the
|
||
limit by hand, or `--resolution-cutoff off` to keep the full range.
|
||
|
||
## Re-scaling and re-merging (`rugnux --scale`)
|
||
|
||
The `--scale` mode re-scales and merges the *already-integrated* reflections stored in a
|
||
`_process.h5` file, without re-running spot finding or integration. Use it to re-merge quickly with a
|
||
different space group, resolution limit, anomalous setting or reference MTZ. It reuses the same
|
||
`-o/-N/-s/-e/-S/-A/-B/-z/--scaling-*` options as the full run, and (unlike the full pipeline) does
|
||
not run a space-group search, so pass `-S` for the correct symmetry.
|
||
|
||
## Quick start
|
||
|
||
### Rotation data
|
||
|
||
Index, integrate, scale and merge a rotation sweep, fully de novo:
|
||
|
||
```
|
||
rugnux rotation_master.h5 \
|
||
-o rotation_run -N 32 \
|
||
--scaling-high-resolution 1.4
|
||
```
|
||
|
||
Because the dataset carries a rotation goniometer axis, it is processed as **rotation data by
|
||
default**: two-pass rotation indexing (index the sweep once, then process every frame against that
|
||
lattice) with the **`rot3d`** partiality model (rotation partials combined into 3D fulls). Scaling
|
||
and merging run **by default** (for both rotation and stills; `--no-merge` turns them off); the unit
|
||
cell is taken from the rotation indexer and the space group is determined from systematic absences,
|
||
and both are written
|
||
into the merged `.cif`.
|
||
|
||
Run **fully de novo** (no `-C`/`-S`) for the best result — supplying a cell or space group up front
|
||
tends to *degrade* low-symmetry cases. `--scaling-high-resolution` (set it to your expected
|
||
resolution) sharpens both the space-group search and the error model. To tune the first pass use
|
||
`--two-pass-rotation=100` (or `-R100` — the first-pass image count); to force the sweep to be
|
||
treated as independent stills use `--force-still`.
|
||
|
||
By default a rotation run also **post-refines the geometry** in a second pass: the first pass
|
||
integrates and merges at the header geometry, then the detector distance + beam centre and the crystal
|
||
cell / rotation-axis are refined against the merged fulls (cross-validated, and committed only for a
|
||
small < 1 % move, with the gauge-weak beam centre restrained toward the header), and the second pass
|
||
re-indexes de novo and re-integrates at the refined geometry. The refined pass is the canonical
|
||
`<prefix>_*` output; the header-geometry pass is kept alongside as `<prefix>_01_*` for comparison.
|
||
Disable it with `--rotation-no-postrefine`.
|
||
|
||
After the per-frame scale-fulls step, rotation scaling applies three **correction surfaces**, **on by
|
||
default** (`--no-scaling-corrections` disables all):
|
||
|
||
- **Decay** — a global Debye–Waller relative-*B* over the run, for the radiation damage that weakens
|
||
later frames more at high resolution (a resolution×time systematic the resolution-flat per-frame
|
||
scale cannot remove). It only engages when the total relative-*B* exceeds a physical floor (2 Ų). An
|
||
optional `--relative-b[=deg]` extends this single global rate to a smooth per-batch relative-*B* curve
|
||
(default 10°-of-rotation batches when bare, off otherwise), cross-validated like the surfaces here, for
|
||
crystals whose decay is non-linear in dose.
|
||
- **Absorption** — a smooth multiplicative factor over the diffracted-beam direction in the goniometer
|
||
frame (path length through the crystal). Negligible at hard X-rays / thin crystals; it matters at
|
||
low photon energy. Its benefit shows up most on model-based metrics: a smooth absorption error
|
||
largely cancels among symmetry mates (little effect on the error model / ISa) but still biases the
|
||
intensities, so it measurably lowers *R*<sub>free</sub>.
|
||
- **Modulation** — a smooth multiplicative factor over the position where a reflection lands on the
|
||
detector (a flat-field: detector-response and geometric systematics that vary across the detector
|
||
plane). Symmetry-equivalents of one reflection land at different detector positions as the crystal
|
||
rotates, which over-determines the surface. Because it lives in the detector frame (not the
|
||
rotation) the same correction concept applies to stills. This is the largest of the three on
|
||
JUNGFRAU data — it lowers *R*<sub>meas</sub> by several to tens of percent on datasets that carry a
|
||
detector systematic, while holding or improving CC<sub>1/2</sub> and the anomalous signal.
|
||
|
||
All three are **cross-validated** — fitted on even-numbered frames and kept only if they improve the
|
||
held-out odd-frame symmetry-equivalent agreement by a clear margin (and vice versa). The agreement is
|
||
scored as a σ-independent, *R*<sub>meas</sub>-like fractional deviation, so a surface can never pass
|
||
cross-validation by merely reshaping the sigmas; where the systematic is absent the surface is a no-op
|
||
rather than a source of added noise, which is why they are safe to leave on.
|
||
|
||
Independently of any correction, a rotation run prints a **radiation-damage report** — the per-image
|
||
scale correlation-to-merge and mosaicity versus dose, and the relative *B*-factor change over the run
|
||
(first→last) together with a per-batch relative-*B* curve, also written to the merged mmCIF. It is a
|
||
data-quality-vs-dose diagnostic and never alters the merged intensities.
|
||
|
||
### Still / serial data
|
||
|
||
A dataset with **no goniometer axis** (e.g. a serial grid scan) is processed as **independent
|
||
stills automatically** — no flag needed. Known-cell indexing with the GPU fast-feedback indexer,
|
||
then merge against a reference structure:
|
||
|
||
```
|
||
rugnux serial_master.h5 \
|
||
-o serial_run -N 32 \
|
||
-X ffbidx -C 79,79,38,90,90,90 -S 96 \
|
||
--adaptive-spots \
|
||
-z reference.mtz \
|
||
--scaling-high-resolution 1.8
|
||
```
|
||
|
||
`ffbidx` requires a known cell (`-C`) and is the indexer of choice for sparse serial stills. For
|
||
serial stills, prefer the self-calibrating spot finder (`--adaptive-spots`) and leave
|
||
`--min-pix-per-spot` **unset** so it is chosen per image — across the still-target battery this
|
||
combination raises the indexing rate and typically extends resolution over a fixed threshold and
|
||
fixed min-pix, at equal or better CC½. (You can still pin a fixed threshold with `--spot-sigma` /
|
||
`--spot-threshold` and a fixed min-pix with `--min-pix-per-spot`.) If a dataset *does* carry a
|
||
goniometer axis but you want per-frame stills processing anyway, add `--force-still`.
|
||
|
||
## Command-line options
|
||
|
||
General:
|
||
|
||
| Option | Description |
|
||
| --- | --- |
|
||
| `-o, --output-prefix <txt>` | Output file prefix (default: `output`) |
|
||
| `-N, --threads <num>` | Number of worker threads (default: all hardware threads) |
|
||
| `-s, --start-image <num>` | First image to process (default: 0) |
|
||
| `-e, --end-image <num>` | Last image to process (default: all) |
|
||
| `-t, --stride <num>` | Process every *n*-th image (default: 1) |
|
||
| `-v, --verbose` | Verbose output |
|
||
|
||
Modes (default: full analysis — spot finding, indexing, integration and merging):
|
||
|
||
| Option | Description |
|
||
| --- | --- |
|
||
| `--azint-only` | Only run azimuthal integration (no spot finding/indexing); writes `<prefix>_process.h5` |
|
||
| `--scale` | Only re-scale/merge the already-integrated reflections in the input `_process.h5` (no re-integration) |
|
||
|
||
Spot finding:
|
||
|
||
| Option | Description |
|
||
| --- | --- |
|
||
| `--spot-sigma <num>` | Noise sigma level for spot finding (default: 3.0) |
|
||
| `--spot-threshold <num>` | Photon-count threshold for spot finding (default: 10) |
|
||
| `--adaptive-spots` | Self-calibrating detection: replace the fixed `--spot-threshold` with a per-resolution-ring threshold derived from each image's own noise, so one setting adapts across datasets (no per-dataset `--spot-threshold`/`--spot-sigma` tuning) |
|
||
| `--spot-false-pixels <num>` | Adaptive-detection operating point: expected noise pixels tolerated per frame (default: 100; implies `--adaptive-spots`) |
|
||
| `--spot-high-resolution <num>` | High-resolution limit for spot finding, Å (default: 1.5) |
|
||
| `--spot-low-resolution <num>` | Low-resolution limit for spot finding, Å (default: 50; lower it, e.g. 24, to exclude the direct-beam halo on weak serial data) |
|
||
| `--min-pix-per-spot <num>` | Minimum connected strong pixels per spot. **If omitted, min-pix is chosen per image** (stills indexing): the frame is indexed at min-pix 3/2/1 and the one maximising indexed-spot count × indexed fraction is kept. Give an explicit value to force a fixed min-pix instead. |
|
||
| `--max-spots <num>` | Maximum spot count (default: 250) |
|
||
| `--detect-ice-rings[=on\|off]` | Flag ice-ring spots (de-prioritised in indexing) and exclude ice-ring reflections from scaling/merging; overrides the dataset/master-file setting (default: use the dataset value) |
|
||
|
||
Azimuthal integration (the radial profile behind the per-image ice-ring score):
|
||
|
||
| Option | Description |
|
||
| --- | --- |
|
||
| `-q, --azim-q-spacing <num>` | Q bin spacing, 1/Å (default: 0.01; finer resolves the narrow ice rings) |
|
||
| `--azim-min-q <num>` | Minimum Q, 1/Å |
|
||
| `--azim-max-q <num>` | Maximum Q, 1/Å |
|
||
| `--azim-phi-bins <num>` | Number of azimuthal (phi) bins (default: 1) |
|
||
| `--polarization-correction <on\|off>` | Enable/disable the azimuthal polarization correction |
|
||
| `--solid-angle-correction <on\|off>` | Enable/disable the azimuthal solid-angle correction |
|
||
|
||
Indexing:
|
||
|
||
A dataset with a **rotation goniometer axis** is processed as rotation data (two-pass rotation
|
||
indexing) by default; a dataset without one is processed as independent stills. `--force-still`
|
||
overrides the former; the `-R` / `--single-pass-rotation` / `--force-rotation-lattice` flags request
|
||
rotation explicitly and pick the pass or lattice.
|
||
|
||
| Option | Description |
|
||
| --- | --- |
|
||
| `--force-still` | Treat a rotation (goniometer) dataset as independent stills instead of rotation |
|
||
| `-X, --indexing-algorithm <txt>` | `FFBIDX` \| `FFT` \| `FFTW` \| `Auto` \| `None` |
|
||
| `-C, --unit-cell <cell>` | Reference unit cell `"a,b,c,alpha,beta,gamma"` (required by `ffbidx`) |
|
||
| `-S, --space-group <num\|symbol>` | Space group number (`92`) or Hermann-Mauguin symbol (`P43212`) — for indexing and scaling |
|
||
| `-r, --refine <txt>` | Geometry refinement: `none` \| `orientation` \| `beam_and_lattice` (default) \| `flex` (try all three per image, keep whichever indexes the most spots; alias `multi`) |
|
||
| `-R, --two-pass-rotation[=num]` | Two-pass offline rotation indexing (default for goniometer data; optional first-pass image count, default 100) |
|
||
| `--single-pass-rotation[=num]` | Online-like single-pass rotation indexing (optional min angular range, deg) |
|
||
| `--redo-rotation-spots` | Redo spot finding for the two-pass rotation first pass |
|
||
| `--force-rotation-lattice <vec>` | Force rotation lattice (9 floats, Å), skipping the first pass |
|
||
| `--rotation-no-postrefine` | Rotation: disable the default-on two-pass geometry post-refine (see the rotation section) |
|
||
| `--refine-geometry[=N\|off]` | Stills: extra first pass that bundle-adjusts the shared beam/distance/cell from N strongly-indexed frames (default 200) then re-indexes; default ON for stills with a reference cell (`-C` / `-z`), `=off` disables |
|
||
|
||
Indexer choice in brief: `ffbidx` (GPU) refines toward a **known cell** and is best for sparse
|
||
serial stills; `fft` (GPU) / `fftw` (CPU) index **de novo** and suit strong rotation data. See the
|
||
[CPU/GPU data-analysis reference](CPU_DATA_ANALYSIS.md) for the algorithms.
|
||
|
||
Scaling and merging:
|
||
|
||
| Option | Description |
|
||
| --- | --- |
|
||
| `--no-merge` | Skip scaling and merging (on by default); write only the per-image `_process.h5` |
|
||
| `-A, --anomalous` | Anomalous mode (keep Friedel pairs separate) |
|
||
| `-B, --refine-bfactor` | Refine a per-image B-factor (stills only) |
|
||
| `--scale-fulls` / `--no-scale-fulls` | rot3d: refit a per-frame scale on the combined fulls (XDS order, Unity model); on by default for rotation data, off for stills |
|
||
| `--smooth-g[=deg]` | rot3d: smooth the per-frame scale *G* over a degree range before the 3D combine (XDS DELPHI-like; default 5° for rotation, 0 = off) |
|
||
| `--no-scaling-corrections` | rot3d: disable the default-on decay + absorption + modulation correction surfaces fitted on the fulls after scale-fulls (see below) |
|
||
| `--relative-b[=deg]` | rot3d: fit a per-batch relative-*B* beyond the single decay slope over deg-degree batches, cross-validated (default 10° when bare; off otherwise) |
|
||
| `--still-partiality` | Experimental (stills): weight reflections by a Gaussian excitation-error partiality instead of treating each as a full |
|
||
| `--partiality-uncertainty <num>` | Stills: extra merge sigma ~num·(1−partiality)·⟨I⟩ on partials (use with `--still-partiality`; default 0, ~2.5 recommended) |
|
||
| `--stills-modulation` | Experimental (stills): fit a detector-plane modulation (flat-field) surface, cross-validated (default off) |
|
||
| `--capture-uncertainty <num>` | rot3d: systematic sigma on under-captured fulls, ~num·(1−captured_fraction)·I (default: 1.0 for rotation, 0 otherwise) |
|
||
| `--min-captured-fraction <num>` | rot3d: drop a combined full whose rocking curve was captured below this fraction — edge-of-sweep truncated fulls (default: 0.7 for rotation, 0 otherwise; 0 = off) |
|
||
| `--scaling-high-resolution <num>` | High-resolution limit for scaling, Å — manual override (default: no limit; disables the automatic cutoff below) |
|
||
| `--resolution-cutoff <txt>` | Automatic high-resolution cutoff for the written reflections and reported shells: `cc-logistic` \| `off` (default: `cc-logistic`; ignored when `--scaling-high-resolution` is set) |
|
||
| `--resolution-cc-target <num>` | CC1/2 target defining the `cc-logistic` fall-off (default: 0.30) |
|
||
| `--resolution-shells <num>` | Number of resolution shells in the reported statistics table (default: 10) |
|
||
| `--min-partiality <num>` | Minimum partiality to accept a reflection (default: 0.02) |
|
||
| `--reject-outliers <num>` | Per-observation outlier rejection, N σ from the per-reflection median (default: 6 for `rot3d`, off otherwise) |
|
||
| `--reject-delta-cchalf <num>` | Drop images with ΔCC1/2 below mean − N·stddev (default: off) |
|
||
| `--min-image-cc <num>` | Per-image CC limit, percent (default: no limit) |
|
||
| `--mosaicity <num>` | Diagnostic: fix the scaling mosaicity (°) instead of using the per-image seed |
|
||
| `--scaling-iterations <num>` | Scaling iterations with no reference data (default: 3) |
|
||
| `-z, --reference-mtz <file>` | Reference MTZ (enables reference-driven scaling) |
|
||
| `--reference-column <label>` | Reference MTZ column to use (default: auto — F-model, else IMEAN/I/…) |
|
||
| `--write-process-h5` | Also write the (large) `_process.h5` when merging (default: only `.mtz`/`.cif`) |
|
||
|
||
Integration:
|
||
|
||
| Option | Description |
|
||
| --- | --- |
|
||
| `--integrator <txt>` | Spot integrator: `gaussian` (profile-fit, default) \| `empirical` \| `boxsum` (classical fallback) |
|
||
| `--integration-radius <r>` | Signal-box radius `r1`, or `r1,r2,r3` (px). One value ⇒ `r2=r1+2`, `r3=r1+4` |
|
||
| `--background-trim <f>` | Monochromatic (rotation + still): symmetric trimmed-mean fraction for the background ring, 0≤f<0.5 (default 0.10; 0 = plain mean) — removes the high-side bias that over-subtracts weak high-angle spots |
|
||
| `--bandwidth <num>` | Relative X-ray bandwidth FWHM (e.g. `0.01` for a 1% DMM); default from file or 0 (monochromatic) |
|
||
|
||
Geometry overrides (defaults are taken from the input file; override them to reprocess with a corrected geometry):
|
||
|
||
| Option | Description |
|
||
| --- | --- |
|
||
| `--beam-x <num>` | Beam centre X (pixel) |
|
||
| `--beam-y <num>` | Beam centre Y (pixel) |
|
||
| `--detector-distance <num>` | Detector distance (mm) |
|
||
| `--wavelength <num>` | Wavelength (Å) |
|
||
| `--rot1 <num>` | PONI detector rotation 1 (rad) |
|
||
| `--rot2 <num>` | PONI detector rotation 2 (rad) |
|
||
| `--polarization <num>` | Polarization factor |
|