Docs: a quick start at the top of the rugnux page, and the indexing ambiguity written down
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 8m33s
Build Packages / build:windows:nocuda (push) Successful in 17m19s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 18m43s
Build Packages / build:windows:cuda (push) Successful in 18m46s
Build Packages / build:viewer-tgz:cpu (push) Successful in 21m4s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m35s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 22m36s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 27m6s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 27m38s
Build Packages / build:rugnux:windows (push) Successful in 10m45s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 19m49s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 23m20s
Build Packages / build:rpm (rocky9) (push) Successful in 23m9s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 27m18s
Build Packages / build:rpm (rocky8) (push) Successful in 28m14s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 23m18s
Build Packages / Generate python client (push) Successful in 48s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 2m12s
Build Packages / XDS test (durin plugin) (push) Successful in 10m32s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 27m54s
Build Packages / DIALS test (push) Successful in 26m59s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m59s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m50s
Build Packages / Unit tests (push) Successful in 2h41m58s
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 8m33s
Build Packages / build:windows:nocuda (push) Successful in 17m19s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 18m43s
Build Packages / build:windows:cuda (push) Successful in 18m46s
Build Packages / build:viewer-tgz:cpu (push) Successful in 21m4s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m35s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 22m36s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 27m6s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 27m38s
Build Packages / build:rugnux:windows (push) Successful in 10m45s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 19m49s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 23m20s
Build Packages / build:rpm (rocky9) (push) Successful in 23m9s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 27m18s
Build Packages / build:rpm (rocky8) (push) Successful in 28m14s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 23m18s
Build Packages / Generate python client (push) Successful in 48s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 2m12s
Build Packages / XDS test (durin plugin) (push) Successful in 10m32s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 27m54s
Build Packages / DIALS test (push) Successful in 26m59s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m59s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m50s
Build Packages / Unit tests (push) Successful in 2h41m58s
The rugnux page opened with installation and reached its quick start around line 500 of 840, which is not where someone who wants to process a dataset looks. It now opens with four commands - the defaults, a reference MTZ, a model, and a pinned cell and space group, which is what people actually ask for - followed by the five files a run leaves and a few pointers. The old quick start keeps the detailed walkthrough under "Running rugnux", moved up to sit after the orientation sections. The indexing ambiguity had no prose anywhere, and neither did -z, which had one row of a table for four different jobs. A new section says what a reference MTZ does, what the ambiguity is, that it costs a rotation run an arbitrary frame and a serial run its CC1/2 outright, how to resolve it in each of the four cases, and that it is neither the enantiomorph nor twinning - the two things it is most often taken for. The three long reference pages carry a table of contents, from the MyST contents directive rather than a hand-written list, so it cannot go stale. The theme's sidebar already lists each page's H2 headings, so this adds the level below that and an in-page map; splitting the pages was the alternative and would have broken every anchor other documents link to. Two things the code says that the docs did not. --mode scale does not accept a reference MTZ on rotation data - the rotation scaler refuses it and the run stops - where the page offered it as one of the things to re-merge with. And CPU_DATA_ANALYSIS 10.9 still said stills use the reference as their per-image scale target, which stopped being true when stills were changed to scale against their own merge; the reference is their per-image ambiguity test and nothing else. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vi1gV6Z45aZL5wLwe85Ksn
This commit is contained in:
@@ -21,6 +21,9 @@ This is an UNSTABLE release. It includes many experimental features, as well as
|
||||
* A failed `/initialize` is reported to `/wait_until_running` and `/wait_till_done` as soon as it happens, instead of when their timeout expires.
|
||||
* `space_group_number` accepts space groups up to 230 in the API schema, so cubic space groups can be recorded. The broker always accepted them; the generated clients rejected them before the request was sent.
|
||||
* The results report's `REPORT_VERSION` is 3, two sections having been added. Existing key names and table columns are unchanged.
|
||||
* `rugnux --model` now settles the frame the merged reflections are written in, not only the frame the R-factors and the maps are computed in: the `.mtz`/`.cif`/`.hkl` come out in the model's indexing, and where the data were merged in the model's enantiomorph they take the model's hand and space group - which on anomalous data puts I(+) and I(-) the right way round. The indexing choice is logged with the winning R-free and the runner-up, so a decision made within noise is visible.
|
||||
* `rugnux --model` can resolve the indexing ambiguity of a **serial stills** run, which a model could not do before: structure factors computed from the model become the per-image reference, the same role a reference MTZ plays. It needs the cell and space group up front (`-C` / `-S`). Without one or the other, a merohedral serial run still merges both hands together and says so.
|
||||
* The rugnux documentation opens with a quick start - the default run, and runs with a reference MTZ, with a model, or with the space group and cell pinned - and explains the indexing ambiguity: what it costs on rotation and on serial data, and which of `-z` / `--model` resolves it in each case. The long reference pages now carry a table of contents.
|
||||
|
||||
### 1.0.0-rc.163
|
||||
This is an UNSTABLE release. It includes many experimental features, as well as many AI generated fixes. We recommend using rc.152 for production use.
|
||||
|
||||
@@ -19,6 +19,11 @@ This document describes the crystallographic algorithms implemented in Jungfrauj
|
||||
13. amplitude estimation (French–Wilson) and R-free test-set flagging,
|
||||
14. optional model-based validation: R-free against a supplied model and 2Fo−Fc / Fo−Fc electron-density maps.
|
||||
|
||||
```{contents} On this page
|
||||
:local:
|
||||
:depth: 2
|
||||
```
|
||||
|
||||
## References
|
||||
|
||||
The methods draw on, and in places reimplement, solutions from:
|
||||
@@ -907,7 +912,7 @@ A reference dataset (`--reference-mtz`) supplies known intensities for the same
|
||||
|
||||
**Fix the space group and cell.** Unless overridden on the command line (`-S` for the space group, `-C` for the cell), the reference's space group is adopted and its cell is used as the soft reference cell — indexing may still drift the cell within tolerance, so a small mismatch between reference and data is absorbed rather than rejected. This applies to both stills and rotation data.
|
||||
|
||||
**Resolve the indexing (merohedral) ambiguity.** When the lattice symmetry is higher than the crystal's Laue symmetry (e.g. $P3$, $P4$, $P6$, $C2$), more than one indexing of the same lattice is geometrically valid, and the two solutions produce *different* merged intensities that a self-consistent scale cannot tell apart — only an external reference can. The candidate reindexings are the identity together with the twin-law cosets of the metric symmetry (from the unit-cell metric and the Laue group); each is scored by the intensity correlation $\mathrm{CC}_\mathrm{ref}$ of the reindexed merge against the reference, and the data are re-merged in the best-correlating indexing. The reindex is **metric-preserving** — only the $hkl$ labels change, the cell is unchanged — and it is a no-op for a holohedral crystal, which has no twin laws (the lattice and Laue symmetry coincide). For rotation data this is done once, after the space group is determined; the reference is *not* used to scale the rotation merge, which stays self-consistent (its $\mathrm{ISa}$ comes from the data alone). For stills the reference is the per-image scale target of the on-the-fly scaling (§10.2).
|
||||
**Resolve the indexing (merohedral) ambiguity.** When the lattice symmetry is higher than the crystal's Laue symmetry (e.g. $P3$, $P4$, $P6$, $C2$), more than one indexing of the same lattice is geometrically valid, and the two solutions produce *different* merged intensities that a self-consistent scale cannot tell apart — only an external reference can. The candidate reindexings are the identity together with the twin-law cosets of the metric symmetry (from the unit-cell metric and the Laue group); each is scored by the intensity correlation $\mathrm{CC}_\mathrm{ref}$ of the reindexed merge against the reference, and the data are re-merged in the best-correlating indexing. The reindex is **metric-preserving** — only the $hkl$ labels change, the cell is unchanged — and it is a no-op for a holohedral crystal, which has no twin laws (the lattice and Laue symmetry coincide). For rotation data this is done once, after the space group is determined, and the whole merge is then repeated in the chosen indexing. For stills it has to be done **per image**, at integration time: each crystal is indexed independently, so a run resolves the ambiguity image by image (the image's partiality/Lorentz-corrected intensities are correlated with the reference under each candidate operator, which is scale-invariant, and the best-correlating one is adopted once and for good) — otherwise the merge would average reflections that are not symmetry mates. In neither workflow is the reference a **scale** target: both scale against their own data (§10.2), so $\mathrm{ISa}$ and the merging statistics come from the data alone and no cross-dataset systematic is imported. Because the stills choice is made at integration time, a later re-merge of stored reflections cannot repair a dataset integrated without a reference. Where there is no reference dataset but there is a **model** (`--model`), the reference intensities are computed from it instead - $|F_\mathrm{model}|^2$ from the atomic structure factors with a flat bulk-solvent contribution at the standard constants ($k_\mathrm{sol}=0.35$, $B_\mathrm{sol}=46$ Å$^2$), which are not fitted because there are no observations yet. Nothing is scaled against them; they serve only to rank the candidate indexings, and the correlation that does the ranking is scale-invariant. This needs the cell and space group up front (`-C` / `-S`, as serial indexing wants anyway); on rotation data the same job is done after the merge, in §14.5, where a merge exists to fit the model to properly.
|
||||
|
||||
### 10.10 Ice rings at the scale and merge stages
|
||||
|
||||
@@ -1069,4 +1074,4 @@ Two maps are formed with the model phases $\varphi_\mathrm{model}$: a $2F_o-F_c$
|
||||
The model fixes a definite hand and indexing, but the merged data need not share them, so before comparison the observed reflections are brought into the model's frame.
|
||||
|
||||
- **Enantiomorph / screw.** When the data space group is the enantiomorph of the model's (e.g. data $P4_12_12$, model $P4_32_12$; or $P3_1/P3_2$), the two are **indistinguishable from merged intensities** — $|F_\mathrm{calc}|$ is invariant under the change of hand, so R-free cannot choose between them and probing would be meaningless. The hand is therefore taken from the model: the observed reflections are reindexed by the change-of-hand operator into the model's enantiomorph. Only the map phases (the density's hand) depend on this choice.
|
||||
- **Indexing (merohedral) ambiguity.** When the crystal has a merohedral ambiguity (§10.9), the observed intensities *do* differ between indexings, and the right one is chosen against the best available reference. **If a reference MTZ was supplied, the data were already reindexed to agree with it** (§10.9 — by the reference-intensity correlation, at the merge stage for rotation data or per image in stills scaling), and model validation keeps that authoritative choice. **Only with a model and no reference** does validation resolve the ambiguity itself, as a fallback: the scaled model is fit to each reindexing of the data (identity plus the twin-law cosets) and the one giving the **lowest R-free** is kept. This matters for a multi-dataset campaign — a single shared reference fixes one indexing convention for every dataset, whereas an independent per-dataset lowest-R-free choice could send borderline datasets to different conventions. A no-op either way for a holohedral crystal (no twin laws).
|
||||
- **Indexing (merohedral) ambiguity.** When the crystal has a merohedral ambiguity (§10.9), the observed intensities *do* differ between indexings, and the right one is chosen against the best available reference. **If a reference MTZ was supplied, the data were already reindexed to agree with it** (§10.9 — by the reference-intensity correlation, at the merge stage for rotation data or per image in stills scaling), and model validation keeps that authoritative choice. **Only with a model and no reference** does validation resolve the ambiguity itself, as a fallback: the scaled model is fit to each reindexing of the data (identity plus the twin-law cosets) and the one giving the **lowest R-free** is kept. This matters for a multi-dataset campaign — a single shared reference fixes one indexing convention for every dataset, whereas an independent per-dataset lowest-R-free choice could send borderline datasets to different conventions. A no-op either way for a holohedral crystal (no twin laws) Both reindexings — the change of hand and the ambiguity choice — are then applied to the merged reflections themselves, which are written after this step, so the reflection file, the R-factors and the maps describe one indexing. The change of hand also changes the space group the file is written in (it is the model's enantiomorph), and since it is the inversion it exchanges the Bijvoet mates: on anomalous data adopting a model's hand is what puts the anomalous differences the right way round. The ambiguity choice is reported with the R-free of the winner and of the runner-up, since the margin between them is what says whether the data decided or the two came out within noise of each other.
|
||||
|
||||
@@ -13,6 +13,11 @@ writer (running, republishing, file finalisation) is described in
|
||||
[CBOR messages](CBOR.md); fields below frequently correspond one-to-one to CBOR message fields, and
|
||||
that document is a useful companion for their meaning.
|
||||
|
||||
```{contents} On this page
|
||||
:local:
|
||||
:depth: 2
|
||||
```
|
||||
|
||||
## 1. Motivation: derived metadata and FAIR data
|
||||
|
||||
The goal of Jungfraujoch is not only to store high-throughput datasets efficiently, but to keep
|
||||
@@ -342,6 +347,14 @@ setting; a third-party reader that ignores it will index the reflections in the
|
||||
the dataset is present. Written by the offline `rugnux` path only — the broker never re-seats a
|
||||
lattice — and not carried on the CBOR stream, in the same way as the other offline-only fields.
|
||||
|
||||
A `--model` run can leave the merged reflection files (`.mtz`/`.cif`/`.hkl`) in a *different* frame
|
||||
from the `_process.h5` beside them: the model settles the enantiomorph, and the alternative indexing
|
||||
where nothing else did, and those choices are applied to the merged reflections as they are written.
|
||||
The process file is not rewritten — its per-image reflections went to disk as they were integrated —
|
||||
so it keeps the space group the run itself determined and stays self-consistent with its own data.
|
||||
The two frames describe the same measurements; an enantiomorphic pair merges identically, having the
|
||||
same Laue class and the same absences.
|
||||
|
||||
**Sweep quality.** `sweepQuality` `[n_images]` (`uint8`) says why the stretch of the sweep this
|
||||
image belongs to was flagged as delivering much less than the rest of the run: **0** means it was
|
||||
not, and any other value is a **1-based index into `sweepQualityReasons`**, a string vector written
|
||||
|
||||
+337
-178
@@ -20,6 +20,61 @@ Run it with no arguments to print the usage.
|
||||
> its options at a high level; the authoritative, always-current list of options is the program's
|
||||
> own usage message — run `rugnux` with no arguments.
|
||||
|
||||
```{contents} On this page
|
||||
:local:
|
||||
:depth: 2
|
||||
```
|
||||
|
||||
## Quick start
|
||||
|
||||
Four commands cover most of what people ask of `rugnux`. Each takes the **master** file of a
|
||||
Jungfraujoch dataset, names its output files from `-o`, and runs on `-N` worker threads:
|
||||
|
||||
```
|
||||
# 1. everything from the data - index, integrate, scale and merge with the defaults
|
||||
rugnux -o myrun -N 32 dataset_master.h5
|
||||
|
||||
# 2. with a reference dataset of the same crystal form: it fixes the space group and the cell,
|
||||
# resolves the indexing ambiguity, and hands over its R-free set
|
||||
rugnux -o myrun -N 32 -z reference.mtz dataset_master.h5
|
||||
|
||||
# 3. with a known structure: R-work / R-free and 2Fo-Fc / Fo-Fc maps on top of the merge
|
||||
rugnux -o myrun -N 32 --model model.pdb dataset_master.h5
|
||||
|
||||
# 4. with the space group and the cell pinned
|
||||
rugnux -o myrun -N 32 -S P43212 -C 79,79,38,90,90,90 dataset_master.h5
|
||||
```
|
||||
|
||||
Nothing more is needed to pick the workflow: a dataset carrying a **goniometer axis** is processed
|
||||
as a rotation sweep, one without as **independent stills**, and scaling and merging run by default
|
||||
in both. A run that merges — the default — leaves five files next to each other:
|
||||
|
||||
```
|
||||
myrun.mtz merged intensities + French-Wilson amplitudes, for CCP4 / phenix
|
||||
myrun.cif the same, as mmCIF - the self-describing format, and what to deposit
|
||||
myrun.hkl the same, as SHELX HKLF 4 - feed this to SHELXC / SHELXD / ANODE
|
||||
myrun_report.txt what the run determined: cell, space group, statistics, warnings
|
||||
myrun_image.dat one row per image, for plotting how the crystal behaved over the sweep
|
||||
```
|
||||
|
||||
Read `myrun_report.txt` first: it says which space group was chosen and on what evidence, how far
|
||||
the data go, and anything that needs attention.
|
||||
|
||||
Four things worth knowing before reaching for more flags:
|
||||
|
||||
- **Rotation data are best left de novo.** Pinning the cell and space group (recipe 4) is the
|
||||
normal thing to do for **serial stills**, where the `ffbidx` indexer needs a cell; on a rotation
|
||||
sweep it tends to *degrade* low-symmetry cases, so prefer recipe 1 and let the run determine both
|
||||
(see [Rotation data](#rotation-data)).
|
||||
- **`-z` and `--model` overlap but are not the same.** A reference MTZ steers the processing from
|
||||
the start; a model scores the merge and settles the frame it is written in. Either resolves an
|
||||
[indexing ambiguity](#the-indexing-ambiguity), which on serial data decides whether the merged
|
||||
intensities are usable at all.
|
||||
- **`--scaling-high-resolution <d>`**, where the resolution is already known, sharpens both the
|
||||
space-group search and the error model.
|
||||
- Everything else is in [Running rugnux](#running-rugnux) and the full
|
||||
[Command-line options](#command-line-options).
|
||||
|
||||
## Installation
|
||||
|
||||
`rugnux` is a **single self-contained executable**. It needs no CUDA toolkit, no Qt, and no
|
||||
@@ -133,6 +188,174 @@ newer for the CUDA 13 ones (RHEL 9, Ubuntu, the aarch64 and Windows `rugnux` arc
|
||||
generations each artefact supports — a V100 in particular works only with the CUDA 12 build — is in
|
||||
[Release contents ▸ GPU generations and the NVIDIA driver](RELEASE_CONTENTS.md#gpu-generations-and-the-nvidia-driver).
|
||||
|
||||
## Running rugnux
|
||||
|
||||
### A first run in detail
|
||||
|
||||
The [Quick start](#quick-start) command is the whole of it:
|
||||
|
||||
```
|
||||
rugnux -o myrun -N 32 /path/to/dataset_master.h5
|
||||
```
|
||||
|
||||
`-o myrun` is the prefix every output file is named from, `-N 32` is the worker-thread count, and the
|
||||
last argument is the **master** file of a Jungfraujoch dataset. Nothing is assumed about the crystal
|
||||
— the goniometer axis in the file tells rugnux this is a rotation sweep, the unit cell comes from
|
||||
indexing the data, the space group from its systematic absences, and the resolution limit from where
|
||||
CC1/2 falls off. Progress, statistics and timing go to the terminal, and the five output files land
|
||||
next to each other.
|
||||
|
||||
`myrun_report.txt` is written for a person, top to bottom: it says which space group was chosen and
|
||||
on what evidence, how far the data go, and anything that needs attention. To pull one number out of
|
||||
it in a script, every value is a `KEY= value` line:
|
||||
|
||||
```
|
||||
grep '^SPACE_GROUP_NUMBER= ' myrun_report.txt
|
||||
grep '^UNIT_CELL_CONSTANTS= ' myrun_report.txt
|
||||
grep '^INCLUDE_RESOLUTION_RANGE= ' myrun_report.txt
|
||||
grep '^ISA= ' myrun_report.txt
|
||||
grep '^WARNING:' myrun_report.txt
|
||||
```
|
||||
|
||||
**Useful variations**, each independent of the others:
|
||||
|
||||
```
|
||||
# tell it where the data really stop, if you already know - this sharpens the
|
||||
# space-group search and the error model
|
||||
rugnux -o myrun -N 32 --scaling-high-resolution 1.4 dataset_master.h5
|
||||
|
||||
# keep Friedel pairs apart, for anomalous work
|
||||
rugnux -o myrun -N 32 -A dataset_master.h5
|
||||
|
||||
# a quick look at the first 200 images only
|
||||
rugnux -o quicklook -N 32 -e 200 dataset_master.h5
|
||||
|
||||
# merge as usual, but also keep the per-image file so the data can be re-merged later
|
||||
rugnux -o myrun -N 32 --write-process-h5 dataset_master.h5
|
||||
|
||||
# also write an unmerged MTZ, to scale the data with aimless instead
|
||||
rugnux -o myrun -N 32 --export-unmerged dataset_master.h5
|
||||
|
||||
# check the merged data against a known structure: R-work / R-free and maps
|
||||
rugnux -o myrun -N 32 --model model.pdb dataset_master.h5
|
||||
```
|
||||
|
||||
Re-merging is cheap and does not re-read the images. Ask the full run to keep its per-image file
|
||||
with `--write-process-h5`, and `--mode scale` will then re-scale and re-merge the reflections
|
||||
already integrated in it — seconds rather than minutes:
|
||||
|
||||
```
|
||||
rugnux -o myrun -N 32 --write-process-h5 dataset_master.h5 # integrate and merge once
|
||||
rugnux --mode scale -o remerged -A myrun_process.h5 # re-merge, here anomalously
|
||||
```
|
||||
|
||||
Use it to try a different resolution limit, anomalous setting or outlier rejection without
|
||||
paying for integration again. `--mode scale` merges in the space group and cell the file
|
||||
records, so the second command needs no `-S`. Note that `--no-merge` also writes a `_process.h5`,
|
||||
but a run that never merged never determined a space group either, so re-merging that file lands in
|
||||
P1 unless you pass `-S` yourself — `--write-process-h5` is the one to use.
|
||||
|
||||
### Rotation data
|
||||
|
||||
Index, integrate, scale and merge a rotation sweep, fully de novo:
|
||||
|
||||
```
|
||||
rugnux rotation_master.h5 \
|
||||
-o rotation_run -N 32 \
|
||||
--scaling-high-resolution 1.4
|
||||
```
|
||||
|
||||
Because the dataset carries a rotation goniometer axis, it is processed as **rotation data by
|
||||
default**: two-pass rotation indexing (index the sweep once, then process every frame against that
|
||||
lattice) with the **`rot3d`** partiality model (rotation partials combined into 3D fulls). Scaling
|
||||
and merging run **by default** (for both rotation and stills; `--no-merge` turns them off); the unit
|
||||
cell is taken from the rotation indexer and the space group is determined from systematic absences,
|
||||
and both are written
|
||||
into the merged `.cif`.
|
||||
|
||||
Run **fully de novo** (no `-C`/`-S`) for the best result — supplying a cell or space group up front
|
||||
tends to *degrade* low-symmetry cases. A `-S` group whose Bravais lattice the crystal turns out not to
|
||||
have stops the run and names the cell that was indexed, rather than merging in a frame the reflections
|
||||
are not in; where the lattice does have that group's setting, the reflections are reindexed into it. `--scaling-high-resolution` (set it to your expected
|
||||
resolution) sharpens both the space-group search and the error model. To tune the first pass use
|
||||
`--two-pass-rotation=100` (or `-R100` — the first-pass image count); to force the sweep to be
|
||||
treated as independent stills use `--force-still`.
|
||||
|
||||
By default a rotation run also **post-refines the geometry** in a second pass: the first pass
|
||||
integrates and merges at the header geometry, then the detector distance + beam centre and the crystal
|
||||
cell / rotation-axis are refined against the merged fulls (cross-validated, and committed only for a
|
||||
small < 1 % move, with the gauge-weak beam centre restrained toward the header), and the second pass
|
||||
re-indexes de novo and re-integrates at the refined geometry. The refined pass is the canonical
|
||||
`<prefix>_*` output; the header-geometry pass merges only to choose the space group and to judge the
|
||||
refined pass against, and writes no merged files of its own — no `<prefix>_01.mtz`, `.cif`, `.hkl` or
|
||||
`_01_image.dat`. (Where a process file is asked for at all, with `--no-merge` or
|
||||
`--write-process-h5`, each pass still writes its own, so `<prefix>_01_process.h5` appears beside
|
||||
`<prefix>_process.h5`.) Disable it with `--rotation-no-postrefine`.
|
||||
|
||||
After the per-frame scale-fulls step, rotation scaling applies three **correction surfaces**, **on by
|
||||
default** (`--no-scaling-corrections` disables all):
|
||||
|
||||
- **Decay** — a global Debye–Waller relative-*B* over the run, for the radiation damage that weakens
|
||||
later frames more at high resolution (a resolution×time systematic the resolution-flat per-frame
|
||||
scale cannot remove). It only engages when the total relative-*B* exceeds a physical floor (2 Ų). An
|
||||
optional `--relative-b[=deg]` extends this single global rate to a smooth per-batch relative-*B* curve
|
||||
(default 10°-of-rotation batches when bare, off otherwise), cross-validated like the surfaces here, for
|
||||
crystals whose decay is non-linear in dose.
|
||||
- **Absorption** — a smooth multiplicative factor over the diffracted-beam direction in the goniometer
|
||||
frame (path length through the crystal). Negligible at hard X-rays / thin crystals; it matters at
|
||||
low photon energy. Its benefit shows up most on model-based metrics: a smooth absorption error
|
||||
largely cancels among symmetry mates (little effect on the error model / ISa) but still biases the
|
||||
intensities, so it measurably lowers *R*<sub>free</sub>.
|
||||
- **Modulation** — a smooth multiplicative factor over the position where a reflection lands on the
|
||||
detector (a flat-field: detector-response and geometric systematics that vary across the detector
|
||||
plane). Symmetry-equivalents of one reflection land at different detector positions as the crystal
|
||||
rotates, which over-determines the surface. Because it lives in the detector frame (not the
|
||||
rotation) the same correction concept applies to stills. This is the largest of the three on
|
||||
JUNGFRAU data — it lowers *R*<sub>meas</sub> by several to tens of percent on datasets that carry a
|
||||
detector systematic, while holding or improving CC<sub>1/2</sub> and the anomalous signal.
|
||||
|
||||
All three are **cross-validated** — fitted on even-numbered frames and kept only if they improve the
|
||||
held-out odd-frame symmetry-equivalent agreement by a clear margin (and vice versa). The agreement is
|
||||
scored as a σ-independent, *R*<sub>meas</sub>-like fractional deviation, so a surface can never pass
|
||||
cross-validation by merely reshaping the sigmas; where the systematic is absent the surface is a no-op
|
||||
rather than a source of added noise, which is why they are safe to leave on.
|
||||
|
||||
Independently of any correction, a rotation run prints a **radiation-damage report** — the per-image
|
||||
scale correlation-to-merge and mosaicity versus dose, and the relative *B*-factor change over the run
|
||||
(first→last) together with a per-batch relative-*B* curve, also written to the merged mmCIF. It is a
|
||||
data-quality-vs-dose diagnostic and never alters the merged intensities. A batch whose data cannot
|
||||
support a measurement prints `-` instead of a value, and the first→last number is printed only where a
|
||||
straight line describes the curve — damage is progressive, so a curve that dips and recovers is a
|
||||
disturbance of the sweep, not dose, and the report says so and points at the sweep-quality section
|
||||
(`RADIATION_DAMAGE_RELATIVE_B= NOT_A_TREND`).
|
||||
|
||||
### Still / serial data
|
||||
|
||||
A dataset with **no goniometer axis** (e.g. a serial grid scan) is processed as **independent
|
||||
stills automatically** — no flag needed. Known-cell indexing with the GPU fast-feedback indexer,
|
||||
then merge against a reference structure:
|
||||
|
||||
```
|
||||
rugnux serial_master.h5 \
|
||||
-o serial_run -N 32 \
|
||||
-X ffbidx -C 79,79,38,90,90,90 -S 96 \
|
||||
-z reference.mtz \
|
||||
--scaling-high-resolution 1.8
|
||||
```
|
||||
|
||||
A crystal form with an [indexing ambiguity](#the-indexing-ambiguity) — P3, P4, P6 and their
|
||||
relatives — **needs** either the `-z` above or `--model model.pdb`, and needs it on the run that
|
||||
integrates: every crystal is indexed in its own hand, and the two are averaged together in the merge
|
||||
unless each image is put into the same hand as it is integrated.
|
||||
|
||||
`ffbidx` requires a known cell (`-C`) and is the indexer of choice for sparse serial stills. The
|
||||
self-calibrating spot finder is on by default for both workflows (`--no-adaptive-spots` turns it off), and for
|
||||
serial stills leave `--min-pix-per-spot` **unset** so it is chosen per image — across the still-target battery this
|
||||
combination raises the indexing rate and typically extends resolution over a fixed threshold and
|
||||
fixed min-pix, at equal or better CC½. (You can still pin a fixed threshold with `--spot-sigma` /
|
||||
`--spot-threshold` and a fixed min-pix with `--min-pix-per-spot`.) If a dataset *does* carry a
|
||||
goniometer axis but you want per-frame stills processing anyway, add `--force-still`.
|
||||
|
||||
## Input and output
|
||||
|
||||
**Input** is a single Jungfraujoch HDF5 master file (NXmx-based). Spots are always found by `rugnux`
|
||||
@@ -411,6 +634,103 @@ reason: a **triclinic** Laue class (no forbidden direction exists), an observed
|
||||
Where anisotropy is detected and the directional limits differ by more than 0.5 Å, a `WARNING:` line
|
||||
says so, since refinement and map interpretation should allow for it.
|
||||
|
||||
## Reference data and the indexing ambiguity
|
||||
|
||||
### What a reference MTZ does (`-z`)
|
||||
|
||||
`-z reference.mtz` supplies **known intensities of the same crystal form** — a previously merged
|
||||
dataset, or `F-model` amplitudes computed from a structure. It is read once, before processing
|
||||
starts, and used for four things:
|
||||
|
||||
- **It fixes the space group and the unit cell** the run works in, unless `-S` / `-C` override them.
|
||||
The cell is a *soft* reference: indexing may still drift within tolerance, so a small mismatch
|
||||
between reference and data is absorbed rather than rejected.
|
||||
- **It resolves the indexing ambiguity** (below) — the one thing the data cannot settle for
|
||||
themselves.
|
||||
- **It hands over its R-free test set**, where the file carries one, so every dataset of a campaign
|
||||
is scored on the same free reflections.
|
||||
- **It reports CC<sub>ref</sub>**, the correlation of the merged intensities against the reference,
|
||||
in the statistics table. Stills only — the rotation merge never scores itself against the
|
||||
reference, and its table shows `nan` in that column.
|
||||
|
||||
A reference is **not** a scale anchor. Both workflows scale against their own data — scaling images
|
||||
against a foreign dataset injects that dataset's systematics — so `-z` never puts the reference's
|
||||
errors into the intensities. `--reference-column` picks the column to read where the automatic
|
||||
choice (`F-model`, else `IMEAN`/`I`, else `FP`/`FOBS`/`F`) is not the right one.
|
||||
|
||||
For the second of those four jobs — and only that one — an atomic model does as well: `--model`
|
||||
computes the intensities it needs from the structure. Where a reference dataset exists, prefer it;
|
||||
where only a model does, it resolves the ambiguity just the same.
|
||||
|
||||
### The indexing ambiguity
|
||||
|
||||
Some crystals can be indexed in **more than one way, each equally valid geometrically, and each
|
||||
giving different merged intensities**. This happens whenever the lattice is more symmetric than the
|
||||
crystal: in P3, P4, P6, P3<sub>1</sub>, C2 and their relatives (*merohedral*), and also where the
|
||||
cell is metrically more symmetric than the Laue class by accident (*pseudo-merohedral*, up to 2° of
|
||||
obliquity). The alternatives are related by the crystal's **twin laws** — reindexing operators such
|
||||
as `k,h,-l`.
|
||||
|
||||
Nothing in the data breaks the tie: the merge is equally self-consistent either way, so which
|
||||
solution comes out is arbitrary. What that costs depends on the workflow:
|
||||
|
||||
- **Rotation.** The whole sweep is one lattice, so the whole dataset lands in one indexing, picked
|
||||
at random. The merge itself is sound; it may simply be the *other* solution from an earlier
|
||||
dataset of the same crystal form, and the two cannot be combined, compared or phased against the
|
||||
same model.
|
||||
- **Serial stills.** Every crystal is indexed independently, so a run mixes both indexings into one
|
||||
merge. That is not a labelling matter — reflections that are not symmetry mates get averaged
|
||||
together, and CC<sub>1/2</sub>, R<sub>meas</sub> and the anomalous signal all degrade.
|
||||
|
||||
Every run that merges tests for the ambiguity, and where it exists and nothing resolves it, says so
|
||||
in the log and in the report's warnings:
|
||||
|
||||
```
|
||||
Indexing ambiguity: this cell / space group admits alternative indexing (reindex operator(s):
|
||||
-h,-k,l). Serial-stills crystals are indexed in one hand at random, and rugnux can only break this
|
||||
against an external reference. WITHOUT one the merge mixes the hands and CC1/2 is degraded - supply
|
||||
a reference MTZ (-z) or a model (--model, which needs -C and -S here) to resolve it.
|
||||
```
|
||||
|
||||
Where something does resolve it, the run says so instead — `Indexing ambiguity present (reindex
|
||||
operator(s): -h,-k,l); resolved against the supplied model`.
|
||||
|
||||
How to resolve it:
|
||||
|
||||
| Situation | What to do |
|
||||
| --- | --- |
|
||||
| **Rotation**, a reference dataset exists | `-z reference.mtz`. Once the space group is settled, each candidate reindexing of the merged intensities is correlated with the reference and the best-correlating one is re-merged. Only the *hkl* labels change; the cell does not |
|
||||
| **Serial stills**, a reference dataset exists | `-z reference.mtz`. Resolved **per image**, at integration time, by correlating each crystal's intensities with the reference — so the merge never mixes hands in the first place |
|
||||
| **Rotation**, only a model | `--model model.pdb`. The merged data are fitted to the model in each candidate indexing and the lowest R-free wins; the written reflections are then reindexed into it, so the file, the R-factors and the maps agree. The log gives the winning R-free and the runner-up — a narrow margin means the data did not really decide |
|
||||
| **Serial stills**, only a model | `--model model.pdb`, with the cell and space group given (`-C` / `-S`). Structure factors are computed from the model up front and used as the reference for the per-image test, exactly as a reference MTZ would be — the ambiguity has to be broken at integration time, and a model can supply the intensities to break it with |
|
||||
| Neither | The run warns and merges in whichever indexing it found: for rotation data a usable dataset in an arbitrary frame, for stills a degraded one |
|
||||
|
||||
Two things to know about where the choice lands:
|
||||
|
||||
- **`--mode scale` cannot repair a stills run after the fact.** The per-image test happens at
|
||||
integration time, so a `_process.h5` whose images were integrated without a reference — an MTZ or a
|
||||
model — has already lost the distinction, and no re-merge brings it back. On rotation data, where
|
||||
the ambiguity is one choice for the whole dataset, `--mode scale --model model.pdb` does resolve it.
|
||||
- **The choice reaches the written reflections**, not only the R-factors and the maps: the merged
|
||||
`.mtz` / `.cif` / `.hkl` (and `_unmerged.mtz`, where it is asked for) are written in the indexing
|
||||
the model or the reference settled on, so the file can be refined against that model as it stands.
|
||||
|
||||
Two things the indexing ambiguity is **not**:
|
||||
|
||||
- **Not the enantiomorph.** P4<sub>1</sub>2<sub>1</sub>2 versus P4<sub>3</sub>2<sub>1</sub>2 (or
|
||||
P3<sub>1</sub> versus P3<sub>2</sub>) leaves the merged intensities *unchanged*, so no amount of
|
||||
data can choose between them and rugnux never tries — the run reports the pair it cannot separate.
|
||||
A model settles it, being the only evidence there is: with `--model` the written reflections take
|
||||
the model's hand and its space group, which on anomalous data exchanges I(+) and I(-). The
|
||||
`_process.h5`, whose per-image reflections were written as they were integrated, keeps the group
|
||||
the run determined and is left alone.
|
||||
- **Not twinning.** The twin laws are the same operators, but twinning is a property of the crystal
|
||||
— two orientations diffracting at once — and is reported separately in the report's twinning
|
||||
section. A crystal can have an indexing ambiguity without being twinned, and usually is.
|
||||
|
||||
The algorithms behind both are in
|
||||
[CPU/GPU data analysis ▸ Reference data](CPU_DATA_ANALYSIS.md#reference-data-fixing-the-space-group-and-resolving-the-indexing-ambiguity).
|
||||
|
||||
## Validating against a model (`rugnux --model`)
|
||||
|
||||
Given a PDB atomic model of the same structure, `--model model.pdb` scales the model structure
|
||||
@@ -424,19 +744,31 @@ It is a *data-quality lens*, independent of the internal statistics: R-free meas
|
||||
intensities against external truth, where CC1/2 and R<sub>meas</sub> only measure them against
|
||||
themselves. It also settles the two things merged intensities alone cannot: the enantiomorph (data
|
||||
merged in P4<sub>1</sub>2<sub>1</sub>2 against a P4<sub>3</sub>2<sub>1</sub>2 model are reindexed
|
||||
into the model's hand), and — when no reference MTZ has already fixed it — a merohedral indexing
|
||||
ambiguity, by keeping the candidate reindexing with the lowest R-free.
|
||||
into the model's hand), and — when no reference MTZ has already fixed it — a merohedral
|
||||
[indexing ambiguity](#the-indexing-ambiguity), by keeping the candidate reindexing with the lowest
|
||||
R-free. Both of those reindexings are applied to the **written reflections** as well as to the
|
||||
R-factors and the maps — validation runs before the reflection files, so the `.mtz` / `.cif` /
|
||||
`.hkl` come out in the model's frame, and where the hand was adopted they carry the model's space
|
||||
group. The log names the operator in each case, and for the indexing choice gives the winning
|
||||
R-free together with the runner-up.
|
||||
|
||||
## Re-scaling and re-merging (`rugnux --mode scale`)
|
||||
|
||||
The `scale` mode re-scales and merges the *already-integrated* reflections stored in a
|
||||
`_process.h5` file, without re-running spot finding or integration. Use it to re-merge quickly with a
|
||||
different space group, resolution limit, anomalous setting or reference MTZ. It reuses the same
|
||||
different space group, resolution limit, anomalous setting or outlier rejection. It reuses the same
|
||||
`-o/-N/-s/-e/-S/-A/-B/-z/--scaling-*` options as the full run, and (unlike the full pipeline) does
|
||||
not run a space-group search: it merges in the space group and unit cell the file records, and `-S`
|
||||
/ `-C` override them. A `_process.h5` written before the group was stored carries none, and merges
|
||||
in P1 unless `-S` says otherwise.
|
||||
|
||||
A reference MTZ (`-z`) is accepted here on **stills** data, where it fixes the space group and cell,
|
||||
reports CC<sub>ref</sub> and hands over its R-free flags; on rotation data the rotation scaler
|
||||
declines it and the run stops with a message saying so. What `--mode scale` never does is reindex
|
||||
the reflections it writes: an [indexing ambiguity](#the-indexing-ambiguity) has to be resolved by
|
||||
the run that integrates (`--model` still picks an indexing for its own R-free and maps here, as in
|
||||
a full run).
|
||||
|
||||
Where the full run re-seated the lattice — the space group it settled on is in a different setting
|
||||
from the one each image was indexed in — the file records the change of basis as
|
||||
`/entry/MX/reindexMatrix`, and `rugnux` applies it on read, so the reflections and the stored cell
|
||||
@@ -508,179 +840,6 @@ mind: **`ORGX`/`ORGY` are 1-based**, because XDS counts pixels from 1 and Jungfr
|
||||
they are the **PONI**, the same quantity Jungfraujoch's beam centre is — so no correction is needed —
|
||||
but not the direct beam once the detector is tilted (see above).
|
||||
|
||||
## Quick start
|
||||
|
||||
### A first run, start to finish
|
||||
|
||||
Process a rotation sweep and merge it, with nothing assumed about the crystal:
|
||||
|
||||
```
|
||||
rugnux -o myrun -N 32 /path/to/dataset_master.h5
|
||||
```
|
||||
|
||||
`-o myrun` is the prefix every output file is named from, `-N 32` is the worker-thread count, and the
|
||||
last argument is the **master** file of a Jungfraujoch dataset. That is the whole command — the
|
||||
goniometer axis in the file tells rugnux this is a rotation sweep, the unit cell comes from indexing
|
||||
the data, the space group from its systematic absences, and the resolution limit from where CC1/2
|
||||
falls off.
|
||||
|
||||
Progress, statistics and timing go to the terminal. When it finishes you have five files next to
|
||||
each other:
|
||||
|
||||
```
|
||||
myrun.mtz merged intensities + French-Wilson amplitudes, for CCP4 / phenix
|
||||
myrun.cif the same, as mmCIF - the self-describing format, and what to deposit
|
||||
myrun.hkl the same, as SHELX HKLF 4 - feed this to SHELXC / SHELXD / ANODE
|
||||
myrun_report.txt what the run determined: cell, space group, statistics, warnings
|
||||
myrun_image.dat one row per image, for plotting how the crystal behaved over the sweep
|
||||
```
|
||||
|
||||
Read `myrun_report.txt` first — it is written for a person, top to bottom, and it says which space
|
||||
group was chosen and on what evidence, how far the data go, and anything that needs attention. To
|
||||
pull one number out of it in a script, every value is a `KEY= value` line:
|
||||
|
||||
```
|
||||
grep '^SPACE_GROUP_NUMBER= ' myrun_report.txt
|
||||
grep '^UNIT_CELL_CONSTANTS= ' myrun_report.txt
|
||||
grep '^INCLUDE_RESOLUTION_RANGE= ' myrun_report.txt
|
||||
grep '^ISA= ' myrun_report.txt
|
||||
grep '^WARNING:' myrun_report.txt
|
||||
```
|
||||
|
||||
**Useful variations**, each independent of the others:
|
||||
|
||||
```
|
||||
# tell it where the data really stop, if you already know - this sharpens the
|
||||
# space-group search and the error model
|
||||
rugnux -o myrun -N 32 --scaling-high-resolution 1.4 dataset_master.h5
|
||||
|
||||
# keep Friedel pairs apart, for anomalous work
|
||||
rugnux -o myrun -N 32 -A dataset_master.h5
|
||||
|
||||
# a quick look at the first 200 images only
|
||||
rugnux -o quicklook -N 32 -e 200 dataset_master.h5
|
||||
|
||||
# merge as usual, but also keep the per-image file so the data can be re-merged later
|
||||
rugnux -o myrun -N 32 --write-process-h5 dataset_master.h5
|
||||
|
||||
# also write an unmerged MTZ, to scale the data with aimless instead
|
||||
rugnux -o myrun -N 32 --export-unmerged dataset_master.h5
|
||||
|
||||
# check the merged data against a known structure: R-work / R-free and maps
|
||||
rugnux -o myrun -N 32 --model model.pdb dataset_master.h5
|
||||
```
|
||||
|
||||
Re-merging is cheap and does not re-read the images. Ask the full run to keep its per-image file
|
||||
with `--write-process-h5`, and `--mode scale` will then re-scale and re-merge the reflections
|
||||
already integrated in it — seconds rather than minutes:
|
||||
|
||||
```
|
||||
rugnux -o myrun -N 32 --write-process-h5 dataset_master.h5 # integrate and merge once
|
||||
rugnux --mode scale -o remerged -A myrun_process.h5 # re-merge, here anomalously
|
||||
```
|
||||
|
||||
Use it to try a different resolution limit, anomalous setting, outlier rejection or reference MTZ
|
||||
without paying for integration again. `--mode scale` merges in the space group and cell the file
|
||||
records, so the second command needs no `-S`. Note that `--no-merge` also writes a `_process.h5`,
|
||||
but a run that never merged never determined a space group either, so re-merging that file lands in
|
||||
P1 unless you pass `-S` yourself — `--write-process-h5` is the one to use.
|
||||
|
||||
### Rotation data
|
||||
|
||||
Index, integrate, scale and merge a rotation sweep, fully de novo:
|
||||
|
||||
```
|
||||
rugnux rotation_master.h5 \
|
||||
-o rotation_run -N 32 \
|
||||
--scaling-high-resolution 1.4
|
||||
```
|
||||
|
||||
Because the dataset carries a rotation goniometer axis, it is processed as **rotation data by
|
||||
default**: two-pass rotation indexing (index the sweep once, then process every frame against that
|
||||
lattice) with the **`rot3d`** partiality model (rotation partials combined into 3D fulls). Scaling
|
||||
and merging run **by default** (for both rotation and stills; `--no-merge` turns them off); the unit
|
||||
cell is taken from the rotation indexer and the space group is determined from systematic absences,
|
||||
and both are written
|
||||
into the merged `.cif`.
|
||||
|
||||
Run **fully de novo** (no `-C`/`-S`) for the best result — supplying a cell or space group up front
|
||||
tends to *degrade* low-symmetry cases. A `-S` group whose Bravais lattice the crystal turns out not to
|
||||
have stops the run and names the cell that was indexed, rather than merging in a frame the reflections
|
||||
are not in; where the lattice does have that group's setting, the reflections are reindexed into it. `--scaling-high-resolution` (set it to your expected
|
||||
resolution) sharpens both the space-group search and the error model. To tune the first pass use
|
||||
`--two-pass-rotation=100` (or `-R100` — the first-pass image count); to force the sweep to be
|
||||
treated as independent stills use `--force-still`.
|
||||
|
||||
By default a rotation run also **post-refines the geometry** in a second pass: the first pass
|
||||
integrates and merges at the header geometry, then the detector distance + beam centre and the crystal
|
||||
cell / rotation-axis are refined against the merged fulls (cross-validated, and committed only for a
|
||||
small < 1 % move, with the gauge-weak beam centre restrained toward the header), and the second pass
|
||||
re-indexes de novo and re-integrates at the refined geometry. The refined pass is the canonical
|
||||
`<prefix>_*` output; the header-geometry pass merges only to choose the space group and to judge the
|
||||
refined pass against, and writes no merged files of its own — no `<prefix>_01.mtz`, `.cif`, `.hkl` or
|
||||
`_01_image.dat`. (Where a process file is asked for at all, with `--no-merge` or
|
||||
`--write-process-h5`, each pass still writes its own, so `<prefix>_01_process.h5` appears beside
|
||||
`<prefix>_process.h5`.) Disable it with `--rotation-no-postrefine`.
|
||||
|
||||
After the per-frame scale-fulls step, rotation scaling applies three **correction surfaces**, **on by
|
||||
default** (`--no-scaling-corrections` disables all):
|
||||
|
||||
- **Decay** — a global Debye–Waller relative-*B* over the run, for the radiation damage that weakens
|
||||
later frames more at high resolution (a resolution×time systematic the resolution-flat per-frame
|
||||
scale cannot remove). It only engages when the total relative-*B* exceeds a physical floor (2 Ų). An
|
||||
optional `--relative-b[=deg]` extends this single global rate to a smooth per-batch relative-*B* curve
|
||||
(default 10°-of-rotation batches when bare, off otherwise), cross-validated like the surfaces here, for
|
||||
crystals whose decay is non-linear in dose.
|
||||
- **Absorption** — a smooth multiplicative factor over the diffracted-beam direction in the goniometer
|
||||
frame (path length through the crystal). Negligible at hard X-rays / thin crystals; it matters at
|
||||
low photon energy. Its benefit shows up most on model-based metrics: a smooth absorption error
|
||||
largely cancels among symmetry mates (little effect on the error model / ISa) but still biases the
|
||||
intensities, so it measurably lowers *R*<sub>free</sub>.
|
||||
- **Modulation** — a smooth multiplicative factor over the position where a reflection lands on the
|
||||
detector (a flat-field: detector-response and geometric systematics that vary across the detector
|
||||
plane). Symmetry-equivalents of one reflection land at different detector positions as the crystal
|
||||
rotates, which over-determines the surface. Because it lives in the detector frame (not the
|
||||
rotation) the same correction concept applies to stills. This is the largest of the three on
|
||||
JUNGFRAU data — it lowers *R*<sub>meas</sub> by several to tens of percent on datasets that carry a
|
||||
detector systematic, while holding or improving CC<sub>1/2</sub> and the anomalous signal.
|
||||
|
||||
All three are **cross-validated** — fitted on even-numbered frames and kept only if they improve the
|
||||
held-out odd-frame symmetry-equivalent agreement by a clear margin (and vice versa). The agreement is
|
||||
scored as a σ-independent, *R*<sub>meas</sub>-like fractional deviation, so a surface can never pass
|
||||
cross-validation by merely reshaping the sigmas; where the systematic is absent the surface is a no-op
|
||||
rather than a source of added noise, which is why they are safe to leave on.
|
||||
|
||||
Independently of any correction, a rotation run prints a **radiation-damage report** — the per-image
|
||||
scale correlation-to-merge and mosaicity versus dose, and the relative *B*-factor change over the run
|
||||
(first→last) together with a per-batch relative-*B* curve, also written to the merged mmCIF. It is a
|
||||
data-quality-vs-dose diagnostic and never alters the merged intensities. A batch whose data cannot
|
||||
support a measurement prints `-` instead of a value, and the first→last number is printed only where a
|
||||
straight line describes the curve — damage is progressive, so a curve that dips and recovers is a
|
||||
disturbance of the sweep, not dose, and the report says so and points at the sweep-quality section
|
||||
(`RADIATION_DAMAGE_RELATIVE_B= NOT_A_TREND`).
|
||||
|
||||
### Still / serial data
|
||||
|
||||
A dataset with **no goniometer axis** (e.g. a serial grid scan) is processed as **independent
|
||||
stills automatically** — no flag needed. Known-cell indexing with the GPU fast-feedback indexer,
|
||||
then merge against a reference structure:
|
||||
|
||||
```
|
||||
rugnux serial_master.h5 \
|
||||
-o serial_run -N 32 \
|
||||
-X ffbidx -C 79,79,38,90,90,90 -S 96 \
|
||||
-z reference.mtz \
|
||||
--scaling-high-resolution 1.8
|
||||
```
|
||||
|
||||
`ffbidx` requires a known cell (`-C`) and is the indexer of choice for sparse serial stills. The
|
||||
self-calibrating spot finder is on by default for both workflows (`--no-adaptive-spots` turns it off), and for
|
||||
serial stills leave `--min-pix-per-spot` **unset** so it is chosen per image — across the still-target battery this
|
||||
combination raises the indexing rate and typically extends resolution over a fixed threshold and
|
||||
fixed min-pix, at equal or better CC½. (You can still pin a fixed threshold with `--spot-sigma` /
|
||||
`--spot-threshold` and a fixed min-pix with `--min-pix-per-spot`.) If a dataset *does* carry a
|
||||
goniometer axis but you want per-frame stills processing anyway, add `--force-still`.
|
||||
|
||||
## Command-line options
|
||||
|
||||
General:
|
||||
@@ -801,9 +960,9 @@ Scaling and merging:
|
||||
| `--search-min-zeta <num>` | De-novo space-group search only: also search a merge of just the observations whose Lorentz geometry \|ζ\| reaches this, and report both answers (default: 0.85 for rotation, 0 = single search). Reflections crossing the Ewald sphere near-tangentially are measured worst and can make a real symmetry operator look like a twin law. Where the two searches disagree, the merge of all the observations decides — as it always has for the systematic absences |
|
||||
| `--mosaicity <num>` | Diagnostic: fix the scaling mosaicity (°) instead of using the per-image seed |
|
||||
| `--scaling-iterations <num>` | Scaling iterations with no reference data (default: 3) |
|
||||
| `-z, --reference-mtz <file>` | Reference MTZ (enables reference-driven scaling) |
|
||||
| `-z, --reference-mtz <file>` | Reference MTZ of the same crystal form: fixes the space group and cell, resolves the [indexing ambiguity](#the-indexing-ambiguity), hands over the R-free set and reports CC<sub>ref</sub>. Not a scale anchor |
|
||||
| `--reference-column <label>` | Reference MTZ column to use (default: auto — F-model, else IMEAN/I/…) |
|
||||
| `--model <file.pdb>` | After merging, validate the merged intensities against this atomic model (see below) |
|
||||
| `--model <file.pdb>` | Validate the merged intensities against this atomic model — R-work / R-free and maps (see [Validating against a model](#validating-against-a-model-rugnux-model)). It also settles the frame the reflections are written in: the enantiomorph, and the [indexing ambiguity](#the-indexing-ambiguity) where no `-z` did. For serial stills given `-C` / `-S`, the model's structure factors become the per-image reference |
|
||||
| `--write-process-h5` | Also write the (large) `_process.h5` when merging (default: only `.mtz`/`.cif`) |
|
||||
| `--export-unmerged` | Also write the integrated observations as `<prefix>_unmerged.mtz`, an unmerged MTZ (POINTLESS column layout) for aimless / pointless / careless. Rotation partials are summed into one full per reflection. Intensities carry the Lorentz-polarization factor and nothing else — the partiality is not divided out and the per-image scale is not applied. Works in `--mode mx` and `--mode scale`. See [The unmerged export](#the-unmerged-export) |
|
||||
| `--export-unmerged-partials` | As above, but writes `<prefix>_unmerged_partials.mtz` with each partial as its own row (one batch per image) for the reading program to sum |
|
||||
|
||||
Reference in New Issue
Block a user