Datasets may be confidential; sample names and measured unit cells committed to the repo can leak outside the group working on them. Scrub existing occurrences and add a "No sample identities in the repository" section to CLAUDE.md (forbidden: sample/dataset names, internal codes, measured cells tied to a sample; fine: space group / lattice / twinning descriptors). - Comments: replace internal dataset codes and protein names with the crystallographic situation they illustrate (centred vs pseudo-symmetric, holohedral, cubic, F-cubic/hexagonal, ...). - Docs: same, in the analysis/writer/stream references and example configs. - Tests: rename sample-named identifiers, TEST_CASE names, file prefixes and asserted labels to neutral crystallographic names (e.g. tetragonal_uc); behaviour unchanged. Reduce the CrystFEL reference PDB to a bare CRYST1 cell file (cell.pdb) and rename the reference data file. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
1.5 KiB
Large reference test datasets (git-LFS)
The [large] Catch tests in tests/ run the full analysis/processing pipeline over real
JUNGFRAU datasets that are too big to keep as ordinary git blobs. They are tracked with
git-LFS (see the tests/data/*.h5 rule in the top-level .gitattributes).
These files are not required to build or to run the normal test suite: every test that
needs them resolves the path through jfjoch_test::LargeDataFile() (tests/TestData.h) and
SKIP()s when the file is absent or is still an unfetched LFS pointer. jfjoch_test also
prints, at start-up, whether this directory is populated.
Fetching
git lfs install
git lfs pull # or: git lfs pull --include "tests/data/*.h5"
Datasets
| File | Dataset | Shipped |
|---|---|---|
rotation_master.h5 |
protein rotation series (~1800 images) | yes (LFS) |
rotation_master.h5 (plus its _data_NNNNNN.h5 files) is fetched by git lfs pull and
drives Rugnux_Rotation. A separate serial dataset is intentionally not shipped
to keep the repository small — the rotation series can be run in serial mode (full analysis
without rotation indexing) to exercise that path. To add your own dataset, drop the master + its
data files here as real files (not symlinks, if you intend to commit them via LFS); the master
references its data files by relative name, so keep them side by side.