Files
Jungfraujoch/docs/TESTS.md
T
leonarski_fandClaude Opus 5 9924dd9fc3 docs: the open test battery, updated - 82 datasets, and not only non-SLS
The public data the pipeline is exercised on has grown from 59 datasets to 82,
77 of them with a released PDB entry and released structure factors, so the
page that credits the depositors and carries the DOI to cite for each is
brought up to date with what is actually run.

Renamed from NON_SLS_TEST_DATA to EXTERNAL_TEST_DATA, because the old title
stopped being true: a few of the sets were collected at SLS beamlines, where
the data are still written by someone else's detector and someone else's
acquisition system. What the battery tests is foreign files, not a foreign
facility.

Also rewritten from the current archives rather than the earlier sample:

- Multi-collection archives: eleven are not a single continuous rotation, not
  four. Seven IRRMC archives hold more than one collection; one sweep is kept
  in six of them, and both are kept in the one whose two sweeps are at
  different wavelengths. A repository project page is not a reliable guide
  here - one describes a 900-frame sweep its own tarball does not contain.
- Detector labels: 76 rows can be compared against the PDB entry. Seven
  genuinely conflict, and one of those the file settles outright - pixel
  count, pixel size, sensor thickness and firmware string agree with an
  EIGER2 9M against the entry's PILATUS4 4M. A further 29 differ only in how
  much they state, which is not a conflict.
- One archive ships 30 placeholder files named like images that are 64-byte
  text; named on the page so a reader that globs the directory is not
  surprised by them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MxrrPcxodNiXzhNiECCVp5
2026-08-30 20:50:10 +02:00

122 lines
7.3 KiB
Markdown

# Tests
The unit and integration tests are written with [Catch2](https://github.com/catchorg/Catch2) and
collected into a single binary, `tests/jfjoch_test`. Build and run it with:
```
make -j$(nproc) jfjoch_test
cd tests
./jfjoch_test # everything
./jfjoch_test "<test name>" # one test case
./jfjoch_test "[tag]" # by tag
```
There are also benchmark and hardware routines, each printing its own usage:
* `jfjoch_hdf5_test` to measure HDF5 dataset writing speed (single threaded). It doubles as the
generator of the HDF5 files used by the external-software tests below.
* `jfjoch_lite_perf_test` to measure the CPU/GPU ("lite") analysis path - indexing, integration and
optional file writing.
* `jfjoch_fpga_test` to test quality/performance of FPGA card(s) and software routines. With `-H` it
runs the high-level-synthesis C model on the CPU, so no FPGA device is needed.
Out-of-space handling is covered separately by `jfjoch_hdf5_enospc_test`, run under the `enospc_shim`
`LD_PRELOAD` module that makes writes fail with `ENOSPC`.
In addition, tests are executed to verify that datasets written by Jungfraujoch are readable by
other MX software (see [Integration with MX data processing software](SOFTWARE_INTEGRATION.md)) -
XDS through the Jungfraujoch, Durin and Neggia plugins, and DIALS `xia2.ssx` - for each of the
NXmx layouts. Input files for these programs are placed in the `tests/xds`, `tests/xds_durin`,
`tests/xds_neggia` and `tests/crystfel` folders. See `.gitea/workflows/build_and_test.yml` for the
exact commands; the CrystFEL fixtures are run by hand rather than in the pipeline.
## Judging a change to the analysis itself
Two harnesses in the repository root run `rugnux` over a directory of stored datasets and score
the result. Neither is part of CI - run them when a change plausibly moves merged results, not as
a reflex. Both take their dataset list from **outside** the repository, because dataset and sample
identities are not committed. The public datasets the pipeline is exercised on, and the DOI to
cite for each, are listed in [External test data](EXTERNAL_TEST_DATA.md).
* `rugnux_vs_xds.py` - the rotation battery. Runs rugnux de novo over every crystal under a data
root and tabulates reflections, observations, space group, R_meas, CC1/2, ISa and wall-clock
time against the XDS `CORRECT.LP` beside each dataset.
* `rugnux_anomalous.py` - the anomalous-peak-height arbiter, below.
### The anomalous-peak-height arbiter
A change that touches **partiality** - a mosaicity estimator, a rocking-curve model, a background
change, anything that alters how partial reflections are weighted - cannot be judged by the
statistics we normally reach for:
| statistic | why it fails for this class of change |
|---|---|
| ISa, R_meas, error-model `b` | one measurement, not three; dominated by the low-resolution shells; not invariant to the uniform intensity rescale a partiality change produces |
| last-shell R_meas | moves with its denominator, i.e. the wrong way by construction |
| `rugnux --model` R-free | tracks its own zero-information floor, which moves ~22x more than R-free itself over the same sweep |
| per-shell agreement with `XDS_ASCII.HKL` | XDS never divides by partiality, so "divide less" moves us toward it mechanically; measured to put the optimum ~1.4x too low |
**Anomalous difference density at known scatterer sites** has none of these problems. It is read in
units of the map's own sigma, so a uniform intensity rescale cancels exactly, and it is referenced
to the structure rather than to another program's partiality model.
`rugnux_anomalous.py` measures it: `shelxc` + `anode -a` (CCP4) on each arm's merged reflections,
against a model that is placed **once** and then held fixed. It reports, per dataset, the mean site
height and the off-site noise floor, and, between arms, the **paired per-site** change.
```
# compare two arms (each a directory of <id>/<id>.hkl + .mtz)
./rugnux_anomalous.py --config <table>.json base=<dir-A> test=<dir-B>
# a parameter scan: numeric labels turn the arms into a curve with a per-dataset optimum
./rugnux_anomalous.py --config <table>.json \
0.85='<scan>/{name}/s0p85.hkl' 1.00='<scan>/{name}/s1.hkl' 1.20='<scan>/{name}/s1p2.hkl'
```
An arm is a rugnux output directory or a path template containing `{name}`. `--place` does the
one-off model placement, `--write-config-template` prints the config skeleton, and ANODE results
are cached under the config's `workdir` (a full 9-dataset x 11-arm scan takes under a minute).
**The gate.** A dataset counts only if its **reference arm** shows top peak > 1.5x the highest
off-site peak **and** at least 3 sites over 5 sigma. A dataset that fails is reported as
`EXCLUDED`, never as a zero - the difference between two noise measurements is not a measurement.
**Standing dataset set** (2026-08): 8 datasets from 7 crystals, 114 sulfur sites, all judged on
native sulfur signal.
| crystals | space group | photon energy | sites each |
|---|---|---|---|
| 2 | P4<sub>1</sub>2<sub>1</sub>2 | 12.4, 16.0 keV | 18 |
| 2 (lysozyme) | P4<sub>3</sub>2<sub>1</sub>2 | 13.0, 5.0 keV | 27 |
| 3 (4 datasets - one crystal contributes two energies) | cubic, I-centred | 13.0, 6.0, 5.0, 5.0 keV | 6 |
Report `n` as crystals, not datasets: two energies of one crystal are not two independent votes,
and the tool prints both counts for that reason.
**Traps this tool exists to encapsulate.** Every one of them has already cost a working day:
1. **The phasing space group comes from the config, never from the merged file.** I23 and
I2<sub>1</sub>3 have identical systematic absences (I-centring already forces the screw
condition), so no data can separate them, and phaser's automatic space-group test only tries
the *enantiomorph* - which for I23 is itself. Phasing an I-centred cubic case in the I23 that
both rugnux and XDS report gives TFZ 7-11 where the other member gives 30-50, and drops the mean
site height by a factor 3-10 - enough to make four good datasets look signal-free. Thirteen
classes of chiral space group are indistinguishable this way; `--place` tries every member of
the class and reports each one's LLG/TFZ.
2. **Place the model once, from a reference arm, and reuse it unchanged.** Re-phasing per arm lets
the model move and contaminates the comparison. Refining the placed model against the dataset's
own amplitudes is allowed (it lifts the peaks another 4-10%) as long as the *same* refined model
is then used for every arm.
3. **The gate and the measurement must use the same model.** Gating on one model and scoring the
curve with another silently changes which datasets are in the set.
4. **The off-site floor skips special positions.** A peak on the cell origin is a ripple of the
calculated phases, not a sample of the background; leaving it in inflates the floor by several
sigma and can turn a passing dataset into a failing one. Such peaks are reported in their own
`spec` column rather than dropped silently.
**Reading the result.** Judge the paired per-site change, with its standard error, pooled over
*crystals*. A per-dataset optimum whose arm does not beat the reference on the paired test is
flagged `not significant vs ref` and must not be quoted as a preference; so must one sitting on the
edge of the scanned grid (`grid edge`) - extend the grid instead.