Files
Jungfraujoch/CLAUDE.md
T
leonarski_fandClaude Opus 5 b47bce7c3b
Build Packages / build:viewer-tgz:cpu (push) Successful in 6m47s
Build Packages / build:viewer-tgz:cuda (push) Successful in 7m11s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 10m9s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 9m45s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 10m39s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 11m48s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m26s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 11m48s
Build Packages / build:rpm (rocky8) (push) Successful in 10m35s
Build Packages / build:rpm (rocky9) (push) Successful in 11m22s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m41s
Build Packages / Generate python client (push) Successful in 17s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m21s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 49s
Build Packages / XDS test (durin plugin) (push) Successful in 8m19s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 7m42s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m32s
Build Packages / DIALS test (push) Successful in 12m35s
Build Packages / Unit tests (push) Successful in 1h3m57s
Build Packages / build:windows:nocuda (push) Failing after 2s
Build Packages / build:windows:cuda (push) Failing after 3s
ci: give the MSVC viewer /arch:AVX, and write down why -march lives in CI
The Linux jobs already pass -march=x86-64-v3; the MSVC viewer job passed nothing,
so it built at the x64 baseline. MSVC has no spelling for the x86-64-v2 level, but
/arch:AVX is the nearest and implies SSE4.1/4.2 - which is the part that matters,
because below SSE4.1 Eigen has no vectorised round and falls back to one libm call
per element. AVX is Sandy Bridge and up, a safe floor for a desktop viewer.

The architecture flags stay OUT of CMakeLists on purpose, so a site can build
x86-64-v4 on an AVX-512 cluster, or -march=native, or the plain baseline. That is
easy to mistake for an oversight and "fix", so say it in CLAUDE.md - together with
the consequence that catches anyone profiling: a default local Release build is not
what CI or production runs, and the gap is not uniform. GPU-bound work is
unaffected, but the CPU and Eigen bound phases - first-pass indexing and
scaling/merging - measure about 26% slower without the flags. That is enough to
make rounding look like a tenth of all cycles when a real build has it nearly free,
and to send a reader at the wrong code.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 21:58:36 +02:00

20 KiB
Raw Blame History

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

What this is

Jungfraujoch is the data-acquisition and analysis system for the PSI JUNGFRAU and EIGER X-ray detectors. It receives detector data, runs it through an FPGA-accelerated pipeline (spot finding, azimuthal/ROI integration, compression), streams images out over ZeroMQ for writing to HDF5, and runs crystallographic analysis (indexing, integration, scaling/merging). Most authoritative documentation lives in docs/ and on Read The Docs (https://jungfraujoch.readthedocs.io). When changing CLI behaviour, the program's own usage message is the source of truth, not the docs.

Build

Out-of-source CMake build, C++20, heavy use of FetchContent (spdlog, zstd, HDF5, slsDetectorPackage, Catch2, cpp-httplib, libzmq, libtiff, FFTW, Ceres, fast-feedback-indexer are downloaded and statically linked — the first configure needs network access and is slow). Two dependencies are deliberately not vendored and must come from the system: ZLIB and Eigen3 ≥ 3.4, both resolved with find_package (see the Eigen note in CMakeLists.txt for why). libjpeg-turbo is built via ExternalProject in preview/; libcurl is fetched only for viewer builds.

mkdir build && cd build
cmake -DCMAKE_BUILD_TYPE=Release ..
make -j$(nproc) jfjoch_broker      # the main service; build other targets by name

Key CMake options:

  • JFJOCH_USE_CUDA (default ON) — GPU path. Needs CUDA ≥ 12.8 (older is warned about and ignored). Provides the ffbidx and fft GPU indexers; without it only the CPU fftw indexer is available. FFTW is fetched and JFJOCH_USE_FFTW defined unconditionally, so fftw is always there. CUDA absence is not a build error — nvcc is looked for on PATH, CUDA_PATH and /usr/local/cuda.
  • JFJOCH_WRITER_ONLY (default OFF) — builds only the HDF5 writer; skips broker, FPGA, receiver, analysis, tests, frontend.
  • JFJOCH_VIEWER_BUILD (default OFF) — builds the Qt6 jfjoch_viewer desktop app in addition to the server stack.
  • JFJOCH_VIEWER_ONLY (default OFF on Linux, forced ON on Windows/macOS) — builds only jfjoch_viewer, rugnux and the libraries they link; skips receiver, FPGA, detector control, tests and the frontend.
  • SLS9 (default OFF) — build against slsDetectorPackage 9.2.0 instead of 8.0.2.
  • JFJOCH_INSTALL_DRIVER_SOURCE (default OFF) — install the PCIe driver source for DKMS/RPM.

-march is deliberately NOT set in CMakeLists

The build system sets no architecture flags, so a site can pick its own — x86-64-v4 on an AVX-512 cluster, -march=native, or the plain baseline. CI passes them explicitly (.gitea/workflows/build_and_test.yml): every Linux configure gets -march=x86-64-v3, and the MSVC viewer job gets /arch:AVX (MSVC has no x86-64-v2 spelling; /arch:AVX is the nearest and implies SSE4.1/4.2).

This matters when profiling. A plain cmake -DCMAKE_BUILD_TYPE=Release .. produces a baseline binary that is not what CI or production runs, and the difference is not uniform: GPU-bound work is unaffected, but the CPU/Eigen-bound phases — first-pass indexing and scaling/merging — are ~26% slower without the flags (measured). Eigen in particular has no vectorised round below SSE4.1 and falls back to a libm call per element, which can make rounding look like ~10% of all cycles when it is nearly free in a real build. Pass the CI flags when measuring anything CPU-side, or the profile will point at the wrong code:

cmake -DCMAKE_BUILD_TYPE=Release -DCMAKE_CXX_FLAGS="-march=x86-64-v3" \
      -DCMAKE_C_FLAGS="-march=x86-64-v3" ..

The frontend is a separate custom target: make frontend (in frontend/: npm ci, npm run build, plus the third-party-licenses, Redoc and Sphinx-docs bundling steps). It is never built automatically — make install only copies whatever already sits in frontend/dist/.

Test

Tests use Catch2 and are collected into a single binary tests/jfjoch_test.

make -j$(nproc) jfjoch_test
cd tests
./jfjoch_test                       # all tests
./jfjoch_test "<test name>"         # one test case (exact name in TEST_CASE)
./jfjoch_test "[tag]"               # by tag
./jfjoch_test -r junit -o report.xml

make jfjoch_hdf5_test builds the HDF5 write-speed benchmark (it lives in tools/, so the binary is build/tools/jfjoch_hdf5_test); CI also uses it to produce files that are validated against XDS (Durin/Neggia), DIALS and CrystFEL. jfjoch_hdf5_enospc_test + the enospc_shim module test out-of-space handling.

There is no lint config in the repo (no .clang-tidy/.clang-format, no CI lint step) — follow the conventions the code already uses: namespaces lower_case, classes CamelCase, global constants UPPER_CASE.

Code style

The overriding principle is simple, readable code — favour the smallest, most direct implementation that a reader can verify at a glance. Extra abstraction, speculative guards, and clever-but-dense constructs are treated as actively harmful, not as polish. When torn between a tidy abstraction and a flat, obvious version, pick the obvious one.

This matters most in the experimental analysis code under image_analysis/, where readability is how the physics gets verified — keep those parts especially plain. Do not add defensive/unrequested code (extra validation, rejection heuristics, "just in case" branches) without asking first; if a guard isn't clearly needed, leave it out.

Match the surrounding code's idiom, naming, and comment density rather than importing a different style.

No sample identities in the repository

Never put sample names or sample-specific measured values into source code, comments, documentation, commit messages, or test fixtures. Datasets belong to users and may be confidential or embargoed; anything committed can leak outside the group working on them. This is a hard rule, not a preference.

  • Forbidden: sample/dataset names or internal codes (a protein name, a beamline dataset ID, a run label), and measured unit-cell parameters tied to a real sample (this is the most sensitive — never hardcode "protein X has cell a,b,c").
  • Fine: general crystallographic descriptors — space group / Laue class ("a P2₁ crystal", "a holohedral 422 case"), lattice centering, twinning, "a crystal whose true axis is reported as its 3× harmonic", pseudo-symmetry, etc. Describe the crystallographic situation, not the specimen.
  • Exception — plain lysozyme. Lysozyme (HEWL) is the field's universal standard test specimen, not a user's confidential dataset, so naming it and using its well-known reference cell (tetragonal ~79/79/38, P4₃2₁2) in tests, docs and comments is allowed. This carve-out is only for generic lysozyme as a benchmark; a named user dataset that happens to be lysozyme is still off limits, and every other real sample remains forbidden.
  • Tests must otherwise use neutral names (e.g. tetragonal_uc, not a specimen-named variable) and, where a cell is needed, a synthetic cell chosen for the test — not a real dataset's parameters.

When a bug was found on a specific dataset, commit the behaviour ("de-novo indexing adopted a spurious axis-multiple supercell"), never the dataset.

Local end-to-end run (no detector / no FPGA)

The FPGA HLS logic can be simulated on the CPU (HLSSimulatedDevice), so the full software stack runs without hardware (slowly — fixed-point math on CPU). See docs/JFJOCH_BROKER.md for the canonical walkthrough.

cd build/broker
./jfjoch_broker ../../etc/broker_local.json 5232      # config JSON + HTTP port
# then, separately:
cd tests/test_data && python jfjoch_broker_test.py     # feeds a test image, starts collection
# observe at http://localhost:5232 ; HDF5 is written under build/broker

etc/broker_local.json, broker_eiger.json, broker_crmx.json are example broker configs (schema = jfjoch_settings in broker/jfjoch_api.yaml).

Architecture

Data flow (online): detector → FPGA acquisition (fpga/, acquisition_device/) → receiver/ builds full images from per-module FPGA output → image_pusher/ streams CBOR-encoded images over ZeroMQ (or TCP) → the consuming side (image_puller/) feeds jfjoch_writer (writer/), which writes NXmx HDF5. The broker also emits a low-rate preview stream and a metadata stream (preview/).

Writer file split: one acquisition produces one _master.h5 plus many _data_NNNNNN.h5 files. Dataset-wide metadata (geometry, detector config, ROI/azimuthal definitions — anything fixed for the whole run) is written to the master file in writer/HDF5NXmx.cpp (the NXmx class). Per-image arrays (one entry per frame) are written to the data files by the HDF5DataFilePlugin subclasses in writer/. Put shared metadata in NXmx, not in a data-file plugin.

The HDF5 master/data layout is one of three FileWriterFormats (common/JFJochMessages.h), all NXmx: NXmxLegacy (master + _data_NNNNNN.h5 joined by external links), NXmxVDS (master + data joined by HDF5 virtual datasets — the default), and NXmxIntegrated (a single self-contained file, no separate data files). Per-image plugins must work for all three; with NXmxIntegrated "master" and "data" are the same file. (The enum also has non-NXmx DataOnly and NoFile; values 4/5 are retired CBF/TIFF, kept only in the OpenAPI enum for back compatibility.)

Two acquisition workflows: the FPGA-accelerated path (JUNGFRAU at PSI; FPGA does masking, summation, spot finding, ROI/azimuthal integration, compression) and the DECTRIS SIMPLON path (EIGER), which has no FPGA — masking/ROI/azimuthal analysis then runs on CPU through the shared image_analysis/ library. Treat ROI and azimuthal features as available in both workflows, not FPGA-only.

jfjoch_broker (broker/) is the central online service: HTTP/REST + OpenAPI control plane, FPGA configuration, image building, ZeroMQ output. JFJochStateMachine drives acquisition state; JFJochServices wires the pieces; OpenAPIConvert/JFJochBrokerParser translate between the generated API model and internal types.

Three analysis frontends share one analysis library (image_analysis/, built as JFJochImageAnalysis):

  • jfjoch_broker — online, real-time (FPGA + GPU).
  • jfjoch_viewer — interactive Qt desktop (viewer/), results not persisted.
  • rugnux (rugnux/rugnux_cli.cpp, built on the Rugnux library in the same directory) — offline batch over a stored HDF5, invoked as rugnux {<options>} <input.h5> (it has no --help; run it with no arguments to print the usage, which is the authority on its flags). Rotation vs stills is auto-detected from the goniometer axis. Merging is on by default (--no-merge to disable); merging writes .mtz/.cif/.hkl and skips the bulky _process.h5 unless --write-process-h5, while --no-merge writes only _process.h5. --azint-only runs only azimuthal integration, and --scale re-scales/merges the already-integrated reflections in a _process.h5. (rugnux = the data-processing half of the system; see docs/NAMING.md.)

image_analysis/ pipeline (subdirs): spot_finding, indexing (ffbidx/fft GPU, fftw CPU), lattice_search, geom_refinement, bragg_prediction, bragg_integration, image_preprocessing, azint, roi, scale_merge, plus rotation_indexer/ and dark_mask_analysis/ compiled straight into JFJochImageAnalysis. (beam_stop/ is an unbuilt prototype — it is in no CMakeLists.txt.) Least-squares refinement uses Ceres (fetched in image_analysis/CMakeLists.txt, built with miniglog, no MKL, no Ceres-CUDA, CXX_THREADS). The indexer is chosen with -X (FFBIDX|FFT|FFTW|Auto|None, default Auto, which resolves to a GPU indexer when one is present and fftw otherwise): ffbidx wants a known cell (-C) and suits sparse serial stills; fft/fftw index de novo and suit strong rotation data.

FPGA (fpga/): hls/ is the Vitis HLS source (image-analysis kernels), hls_simulation/ runs that same HLS on CPU for hardware-free testing, hdl/ is the Verilog RTL, host_library/ is the host-side driver, pcie_driver/ is the kernel module. The HLS algorithms are documented in docs/FPGA_DATA_ANALYSIS.md.

Detector control (detector_control/): wrappers for SLS (JUNGFRAU) and DECTRIS SIMPLON (EIGER). jungfrau/: JUNGFRAU ADU→energy gain/pedestal calibration.

Other libs: common/ (geometry, diffraction experiment, image buffer, CUDA wrappers — the shared core, linked nearly everywhere), compression/ (vendored bitshuffle + LZ4 + zstd; the algorithms are BSHUF_LZ4, BSHUF_ZSTD, BSHUF_ZSTD_RLE, BSHUF_ZSTD_RLE_HUFF), frame_serialize/ (CBOR stream codec), reader/ (HDF5 dataset read-back), gemmi_gph/ (vendored GEMMI for MTZ/XDS_ASCII I/O), xds-plugin/ (XDS HDF5 read plugin).

Portability (jfjoch_viewer)

Cross-platform support is a goal only for jfjoch_viewer and its dependency tree — keep that code, and any shared library it transitively links (common/, image_analysis/, reader/, gemmi_gph/, etc.), MSVC-compatible so the viewer can build on Windows. The rest of the project (broker, receiver, FPGA host, detector control, …) is Linux-only and does not need to be portable; don't constrain it for portability's sake.

  • Windows/MSVC is the primary portability target. The end goal is a Windows viewer built with MSVC and CUDA (JFJOCH_USE_CUDA=ON): GPU processing is a wanted feature, not optional, so the intended Windows config is the full GPU path (ffbidx, GPU fft), not a CPU-only fallback. MSVC is required regardless, because CUDA on Windows requires it. Avoid GCC/Clang-only extensions, POSIX-only APIs, and other non-MSVC constructs in viewer-reachable code, and keep CUDA-reachable viewer code (ffbidx, GPU indexers) MSVC-buildable too.
  • macOS is a nice-to-have for the viewer. It rules out CUDA, so anything the viewer depends on must also have a working CPU-only / non-CUDA path (the JFJOCH_USE_CUDA=OFF, fftw-indexer configuration). This non-CUDA path must keep working, but it is the macOS fallback — not the intended Windows configuration.

JFJOCH_VIEWER_ONLY is forced ON on Windows and macOS, so a plain configure there already builds just the portable subset. libtiff, libjpeg-turbo (via ExternalProject in preview/) and libcurl (viewer builds only) are now brought in by the build itself, so they need nothing from the host. Two dependencies are still supplied externally on every platform, deliberately and by decision, not as an open TODO: ZLIB and Eigen (header-only; needed by Ceres, by the analysis libs directly, and by ffbidx under CUDA). Both are resolved with find_package — the system package on Linux (zlib-devel, eigen3-devel), a build prefix pointed at by CMAKE_PREFIX_PATH/Eigen3_DIR on Windows. Vendoring Eigen through FetchContent with OVERRIDE_FIND_PACKAGE is specifically ruled out: it segfaults the CMake bundled with Visual Studio (see the long note in CMakeLists.txt). Qt is supplied externally as before.

OpenAPI is the single source of truth

broker/jfjoch_api.yaml defines the entire REST API and the shared data schemas. From it, update_version.sh regenerates three clients — do not hand-edit generated code:

  • C++ server model → broker/gen/ (cpp-pistache-server generator; compiled as JFJochAPI).
  • Python client → python-client/ (and gen_python_client.sh, published as PyPI jfjoch-client).
  • TypeScript frontend client → frontend/src/client/ (hey-api openapi-ts, npm run openapi).

When you change jfjoch_api.yaml, regenerate the relevant client(s); for a version bump write the new version into VERSION and run update_version.sh (which reads it and rewrites the version: in the YAML itself, frontend/src/version.ts, frontend/package.json, docs/conf.py, the python client, the Redoc html — and also the FPGA HDL and PCIe-driver version strings). It downloads openapi-generator-cli.jar and runs npm install, so it needs network access.

Frontend

React 19 + TypeScript + MUI 6 + Vite 7 (frontend/); charts are Plotly (react-plotly.js). Data layer is generated from the OpenAPI spec (@hey-api/openapi-ts → fetch client + TanStack Query hooks + zod schemas; config in frontend/openapi-ts.config.ts). Scripts: npm start (dev server), npm run build (tsc + vite), npm run openapi (regen client), npm run redocly4broker (regen broker/redoc-static.html).

Adding a per-image scalar quantity

A per-image scalar (e.g. ice_ring_score, bkg_estimate, mosaicity) flows analysis → message → CBOR → HDF5 → scan-result/plot → API → viewer/frontend. To add one, mirror an existing float scalar (bkg_estimate is a clean template) at every layer:

  1. Compute where the azint profile is finalized: image_analysis/MXAnalysisWithoutFPGA.cpp (CPU), receiver/JFJochReceiverFPGA.cpp (FPGA), and the offline azint worker in rugnux/Rugnux.cpp.
  2. Message (common/JFJochMessages.h): std::optional<float> <name> in DataMessage; in EndMessage two members — std::vector<float> v_<name> (the per-image array) and an std::optional<float> <name> run-mean scalar. Mind the v_ prefix.
  3. CBOR: encode in frame_serialize/CBORStream2Serializer.cpp — one key in the DataMessage block (SerializeImageInternal) and two in the END block (SerializeSequenceEnd: <name> and v_<name>); decode the same three in CBORStream2Deserializer.cpp. Optional fields are back-compatible — no version bump.
  4. HDF5 write: writer/HDF5DataFilePluginMX.{h,cpp} — an AutoIncrVector<float> with reserve / per-image write / SaveVector("/entry/MX/<name>") (per-image arrays live in the data-file plugin); plus the NXmx master writes in writer/HDF5NXmx.cpp (SaveVectorIfMissing(..., end.v_<name>) and a SaveScalar(... "Mean", end.<name>)). HDF5 dataset names are camelCase even though the message fields are snake_casebkg_estimate is stored as /entry/MX/bkgEstimate (+ bkgEstimateMean). HDF5 read-back (so a stored file re-opens, e.g. in the viewer) is in reader/HDF5MetadataSource.cpp, NOT JFJochHDF5Reader.cpp: mirror the three bkgEstimate sites — master ReadOptVector, data-file ReadVector into the dataset, and the per-image message population.
  5. Scan result: common/ScanResult.h (ScanResultElem) + common/ScanResultGenerator.cpp (copy in Add, resize+fill in FillEndMessage). The name is not preserved here — the bkg_estimate member is called just bkg, and so is the API property in step 7.
  6. Receiver plot: common/Plot.h (PlotType) + common/JFJochReceiverPlots.{h,cpp} — a StatusVector member, cleared in Setup, fed in Add, plus the GetPlots / GetPlotRaw cases and (if the run-mean scalar of step 2 is wanted) a GetBkgEstimate-style accessor.
  7. API: add to the plot_type enum and the scan_result images schema in broker/jfjoch_api.yaml, regenerate the C++ model (java -jar openapi-generator-cli.jar generate -i broker/jfjoch_api.yaml -o broker/gen -g cpp-pistache-server) and the frontend client (cd frontend && npm run openapi), then wire broker/OpenAPIConvert.cpp (ConvertPlotType string→enum and the Convert(ScanResult) setter).
  8. Reader/viewer: reader/JFJochReaderDataset.h + viewer/JFJochHttpReader.cpp (GetPlot_i) and viewer/JFJochViewerDatasetInfo.cpp (combo item + ExtractMetric).
  9. Frontend: frontend/src/components/DataProcessingPlots.tsx (MenuItem) + DataProcessingPlot.tsx in the same directory (y-axis label in AxisTypeY).
  10. Docs: docs/CBOR.md and docs/HDF5.md name the fields literally; docs/CPU_DATA_ANALYSIS.md describes the quantity in prose.

Gotcha: an existing build/ dir needs a cmake . reconfigure to pick up a newly-added broker/gen/model source file — broker/CMakeLists.txt collects it with AUX_SOURCE_DIRECTORY, a configure-time directory scan. (gen/api is include-path-only and needs no reconfigure.)