Files
Jungfraujoch/tests/data/README.md
T
leonarski_fandClaude Opus 4.8 abbee2d4dc Remove sample identities from the repository; document the rule
Datasets may be confidential; sample names and measured unit cells committed to
the repo can leak outside the group working on them. Scrub existing occurrences
and add a "No sample identities in the repository" section to CLAUDE.md
(forbidden: sample/dataset names, internal codes, measured cells tied to a
sample; fine: space group / lattice / twinning descriptors).

- Comments: replace internal dataset codes and protein names with the
  crystallographic situation they illustrate (centred vs pseudo-symmetric,
  holohedral, cubic, F-cubic/hexagonal, ...).
- Docs: same, in the analysis/writer/stream references and example configs.
- Tests: rename sample-named identifiers, TEST_CASE names, file prefixes and
  asserted labels to neutral crystallographic names (e.g. tetragonal_uc);
  behaviour unchanged. Reduce the CrystFEL reference PDB to a bare CRYST1 cell
  file (cell.pdb) and rename the reference data file.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 17:19:43 +02:00

31 lines
1.5 KiB
Markdown

# Large reference test datasets (git-LFS)
The `[large]` Catch tests in `tests/` run the full analysis/processing pipeline over real
JUNGFRAU datasets that are too big to keep as ordinary git blobs. They are tracked with
**git-LFS** (see the `tests/data/*.h5` rule in the top-level `.gitattributes`).
These files are **not required** to build or to run the normal test suite: every test that
needs them resolves the path through `jfjoch_test::LargeDataFile()` (`tests/TestData.h`) and
`SKIP()`s when the file is absent or is still an unfetched LFS pointer. `jfjoch_test` also
prints, at start-up, whether this directory is populated.
## Fetching
```
git lfs install
git lfs pull # or: git lfs pull --include "tests/data/*.h5"
```
## Datasets
| File | Dataset | Shipped |
|----------------------------|-----------------------------------------|-----------|
| `rotation_master.h5` | protein rotation series (~1800 images) | yes (LFS) |
`rotation_master.h5` (plus its `_data_NNNNNN.h5` files) is fetched by `git lfs pull` and
drives `Rugnux_Rotation`. A separate serial dataset is intentionally **not** shipped
to keep the repository small — the rotation series can be run in serial mode (full analysis
without rotation indexing) to exercise that path. To add your own dataset, drop the master + its
data files here as real files (not symlinks, if you intend to commit them via LFS); the master
references its data files by relative name, so keep them side by side.