Files
Jungfraujoch/docs/TESTS.md
T
leonarski_fandClaude Opus 5 bec2a10a3a docs: add the eight new public datasets to the non-SLS test list
Three IUCrData Raw Data Letters (Zenodo, CC-BY-4.0) and five SBGrid Data
Bank depositions (CC0) have been added to the data the pipeline is
exercised on. Same rules as the rest of the page: the source is the
repository and its own citable DOI, every one of which was resolved before
it was written down; beamline, resolution, space group and cell are the
values deposited with the PDB entry; the detector is read out of the image
files. None of the new detectors disagree with their PDB entry.

Two of them do not fit the page's one-row-per-sweep shape, so the shape is
described rather than flattened. The 6R72 Zenodo record holds two complete
360-degree collections on one crystal - a helical one that produced the
deposited structure and a low-dose one that has no PDB entry - and both are
listed, sharing a DOI, with the second in the no-PDB-entry table. The three
CHESS depositions are 4-11 wedges of 50 degrees per crystal plus a rotation
taken with the crystal translated out of the beam, tabulated in a new
section.

One SBGrid deposition is named but not in the table: its images are 1995 CCD
TIFFs, a format the reader does not support, so it is not processed here and
saying so is more useful than leaving it out.

ACKNOWLEDGEMENT gains the new per-repository counts and a pointer to the
Raw Data Letter citations; TESTS points at the page.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T3yNBXk4wKdMZy1ak2NY7f
2026-08-29 23:40:10 +02:00

122 lines
7.3 KiB
Markdown

# Tests
The unit and integration tests are written with [Catch2](https://github.com/catchorg/Catch2) and
collected into a single binary, `tests/jfjoch_test`. Build and run it with:
```
make -j$(nproc) jfjoch_test
cd tests
./jfjoch_test # everything
./jfjoch_test "<test name>" # one test case
./jfjoch_test "[tag]" # by tag
```
There are also benchmark and hardware routines, each printing its own usage:
* `jfjoch_hdf5_test` to measure HDF5 dataset writing speed (single threaded). It doubles as the
generator of the HDF5 files used by the external-software tests below.
* `jfjoch_lite_perf_test` to measure the CPU/GPU ("lite") analysis path - indexing, integration and
optional file writing.
* `jfjoch_fpga_test` to test quality/performance of FPGA card(s) and software routines. With `-H` it
runs the high-level-synthesis C model on the CPU, so no FPGA device is needed.
Out-of-space handling is covered separately by `jfjoch_hdf5_enospc_test`, run under the `enospc_shim`
`LD_PRELOAD` module that makes writes fail with `ENOSPC`.
In addition, tests are executed to verify that datasets written by Jungfraujoch are readable by
other MX software (see [Integration with MX data processing software](SOFTWARE_INTEGRATION.md)) -
XDS through the Jungfraujoch, Durin and Neggia plugins, and DIALS `xia2.ssx` - for each of the
NXmx layouts. Input files for these programs are placed in the `tests/xds`, `tests/xds_durin`,
`tests/xds_neggia` and `tests/crystfel` folders. See `.gitea/workflows/build_and_test.yml` for the
exact commands; the CrystFEL fixtures are run by hand rather than in the pipeline.
## Judging a change to the analysis itself
Two harnesses in the repository root run `rugnux` over a directory of stored datasets and score
the result. Neither is part of CI - run them when a change plausibly moves merged results, not as
a reflex. Both take their dataset list from **outside** the repository, because dataset and sample
identities are not committed. The public datasets the pipeline is exercised on, and the DOI to
cite for each, are listed in [Non-SLS test data](NON_SLS_TEST_DATA.md).
* `rugnux_vs_xds.py` - the rotation battery. Runs rugnux de novo over every crystal under a data
root and tabulates reflections, observations, space group, R_meas, CC1/2, ISa and wall-clock
time against the XDS `CORRECT.LP` beside each dataset.
* `rugnux_anomalous.py` - the anomalous-peak-height arbiter, below.
### The anomalous-peak-height arbiter
A change that touches **partiality** - a mosaicity estimator, a rocking-curve model, a background
change, anything that alters how partial reflections are weighted - cannot be judged by the
statistics we normally reach for:
| statistic | why it fails for this class of change |
|---|---|
| ISa, R_meas, error-model `b` | one measurement, not three; dominated by the low-resolution shells; not invariant to the uniform intensity rescale a partiality change produces |
| last-shell R_meas | moves with its denominator, i.e. the wrong way by construction |
| `rugnux --model` R-free | tracks its own zero-information floor, which moves ~22x more than R-free itself over the same sweep |
| per-shell agreement with `XDS_ASCII.HKL` | XDS never divides by partiality, so "divide less" moves us toward it mechanically; measured to put the optimum ~1.4x too low |
**Anomalous difference density at known scatterer sites** has none of these problems. It is read in
units of the map's own sigma, so a uniform intensity rescale cancels exactly, and it is referenced
to the structure rather than to another program's partiality model.
`rugnux_anomalous.py` measures it: `shelxc` + `anode -a` (CCP4) on each arm's merged reflections,
against a model that is placed **once** and then held fixed. It reports, per dataset, the mean site
height and the off-site noise floor, and, between arms, the **paired per-site** change.
```
# compare two arms (each a directory of <id>/<id>.hkl + .mtz)
./rugnux_anomalous.py --config <table>.json base=<dir-A> test=<dir-B>
# a parameter scan: numeric labels turn the arms into a curve with a per-dataset optimum
./rugnux_anomalous.py --config <table>.json \
0.85='<scan>/{name}/s0p85.hkl' 1.00='<scan>/{name}/s1.hkl' 1.20='<scan>/{name}/s1p2.hkl'
```
An arm is a rugnux output directory or a path template containing `{name}`. `--place` does the
one-off model placement, `--write-config-template` prints the config skeleton, and ANODE results
are cached under the config's `workdir` (a full 9-dataset x 11-arm scan takes under a minute).
**The gate.** A dataset counts only if its **reference arm** shows top peak > 1.5x the highest
off-site peak **and** at least 3 sites over 5 sigma. A dataset that fails is reported as
`EXCLUDED`, never as a zero - the difference between two noise measurements is not a measurement.
**Standing dataset set** (2026-08): 8 datasets from 7 crystals, 114 sulfur sites, all judged on
native sulfur signal.
| crystals | space group | photon energy | sites each |
|---|---|---|---|
| 2 | P4<sub>1</sub>2<sub>1</sub>2 | 12.4, 16.0 keV | 18 |
| 2 (lysozyme) | P4<sub>3</sub>2<sub>1</sub>2 | 13.0, 5.0 keV | 27 |
| 3 (4 datasets - one crystal contributes two energies) | cubic, I-centred | 13.0, 6.0, 5.0, 5.0 keV | 6 |
Report `n` as crystals, not datasets: two energies of one crystal are not two independent votes,
and the tool prints both counts for that reason.
**Traps this tool exists to encapsulate.** Every one of them has already cost a working day:
1. **The phasing space group comes from the config, never from the merged file.** I23 and
I2<sub>1</sub>3 have identical systematic absences (I-centring already forces the screw
condition), so no data can separate them, and phaser's automatic space-group test only tries
the *enantiomorph* - which for I23 is itself. Phasing an I-centred cubic case in the I23 that
both rugnux and XDS report gives TFZ 7-11 where the other member gives 30-50, and drops the mean
site height by a factor 3-10 - enough to make four good datasets look signal-free. Thirteen
classes of chiral space group are indistinguishable this way; `--place` tries every member of
the class and reports each one's LLG/TFZ.
2. **Place the model once, from a reference arm, and reuse it unchanged.** Re-phasing per arm lets
the model move and contaminates the comparison. Refining the placed model against the dataset's
own amplitudes is allowed (it lifts the peaks another 4-10%) as long as the *same* refined model
is then used for every arm.
3. **The gate and the measurement must use the same model.** Gating on one model and scoring the
curve with another silently changes which datasets are in the set.
4. **The off-site floor skips special positions.** A peak on the cell origin is a ripple of the
calculated phases, not a sample of the background; leaving it in inflates the floor by several
sigma and can turn a passing dataset into a failing one. Such peaks are reported in their own
`spec` column rather than dropped silently.
**Reading the result.** Judge the paired per-site change, with its standard error, pooled over
*crystals*. A per-dataset optimum whose arm does not beat the reference on the paired test is
flagged `not significant vs ref` and must not be quoted as a preference; so must one sitting on the
edge of the scanned grid (`grid edge`) - extend the grid instead.