Files
Jungfraujoch/docs/RELEASE_CONTENTS.md
T
leonarski_fandjungfrau 4dc2534dbf
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 18m57s
Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 16m55s
Build Packages / build:windows:cuda (push) Successful in 18m48s
Build Packages / build:viewer-tgz:cpu (push) Successful in 13m10s
Build Packages / build:viewer-tgz:cuda (push) Successful in 14m45s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 22m23s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 20m12s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 23m7s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 20m43s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 23m9s
Build Packages / XDS test (durin plugin) (push) Successful in 12m26s
Build Packages / build:rpm (rocky9) (push) Successful in 24m58s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 23m20s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (JFJoch plugin) (push) Successful in 12m37s
Build Packages / build:rpm (rocky8) (push) Successful in 27m58s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m38s
Build Packages / Build documentation (push) Successful in 59s
Build Packages / DIALS test (push) Successful in 23m16s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m38s
v1.0.0.rc-162 (#72)
**Files written by Jungfraujoch now import correctly in DIALS, XDS and pyFAI.** A tilted detector, a grid scan, a still recorded at a goniometer position, and saturated or unreadable pixels were each described in a way that a third-party program acted on wrongly. If you process Jungfraujoch data outside Jungfraujoch, prefer this release to any earlier one.

* HDF5: the detector tilt (`rot1`/`rot2`/`rot3`) is exported correctly in the NXmx transformation chain; untilted geometries are unaffected.
* HDF5: a still recorded at a goniometer position is no longer read back as a single image, and a grid scan records a stationary spindle so a program that requires a rotation axis can open it.
* HDF5: the sample transformation chain is written in mounting order, with a Smargon head position told apart from the spindle, one entry per image, `module_offset` as a float unit vector, and `offset_units` on every offset.
* HDF5: saturated, underloaded and unreadable pixels are described so a downstream program masks them - `saturation_value`, `underload_value`, `error_value` and `bit_depth_readout` are written correctly, and a data file missing next to a VDS master reads as the error marker rather than as zero counts.
* HDF5: the rotation axis is read back under whatever name it carries, and `mirror_y` records whether the assembled image is mirrored in Y relative to the detector's raw readout.
* A grid scan and a goniometer axis can both be set; they are no longer alternatives.
* `images_per_file` is chosen from the acquisition when it is not given: a rotation sweep of at most 20000 images goes into a single data file, a grid scan splits on whole fast-axis rows, and stills and serial keep 1000.
* The writer refuses a stream whose start message declares a different pixel format than its images carry, and a DECTRIS detector sending signed images is no longer declared unsigned.
* The image stream can carry the sample transformation chain (`transformations`, in the END message); a producer that does not send it gets the same chain built by the writer.
* rugnux: fixing the space group with `-S` no longer prevents the lattice from being found - a lattice indexed in a different setting is reindexed into that group's own setting, and a run whose crystal does not have that group's lattice stops and names the cell it indexed as, rather than reporting statistics that cannot describe it.
* rugnux: the per-image resolution estimate now predicts the resolution the merged data reach rather than the highest-resolution spot found, and is reported as `SPOT_RESOLUTION_ESTIMATE`.
* rugnux: two runs of the same command on the same images produce the same merged intensities; the azimuthal profile written alongside them is not yet reproducible in the same way.
* rugnux: the offline lattice refinement is bounded by iterations rather than by a wall clock, so a loaded machine can no longer refine to a different lattice; a live acquisition keeps its real-time bound.
* rugnux: the detector-frame modulation correction is fitted on a grid spanning the detector, so whether it is applied no longer depends on how far integration reached.
* rugnux: the geometry pre-pass no longer writes `<prefix>_01.mtz`, `_01.cif`, `_01.hkl` and `_01_image.dat`; the refined second pass writes those files under `<prefix>`, and that is the result to use.
* rugnux: `_process.h5` describes the pixel format of the images it links to, and is written on a thread of its own.
* rugnux: the detector geometry is also logged in XDS's convention (`ORGX`/`ORGY`, detector axis vectors, rotation axis), so it can be compared with an XDS refinement.
* rugnux: an image integrated in pyFAI through the `.poni` file written by `--mode calibration` comes out with the correct azimuth, and the file declares pyFAI's `orientation`, which needs pyFAI 2024.01 or newer. Radial integration is unchanged.
* rugnux: a rotation run is substantially faster throughout - beam-stop detection, first-pass indexing, geometry refinement, integration, scaling and merging - and observations outside the scaling resolution range are dropped as they are ingested. The refined geometry, the space group chosen and the merged statistics are unchanged.
* Faster spot finding and indexing, on the broker as well as in rugnux; the spots found and the lattices indexed are unchanged.
* A run reserves substantially less GPU memory: nothing is allocated for buffers that are never read, and a worker builds only the engines it uses.
* rugnux: with `-N` left at its default the per-image loop of `--mode mx` uses at most 16 workers per GPU, rather than one per hardware thread; an explicit `-N` is obeyed as given.
* CUDA 12 builds now contain device code for Volta, so the RHEL 8 packages and the portable Linux `.tgz` run on a V100; the CUDA 13 artefacts (RHEL 9, Ubuntu, Windows) remain Turing and newer.
* The build resolves a single Eigen for the whole project, and refuses to configure if Ceres picks up a different one; a build that mixed two Eigen versions was undefined behaviour and crashed at -O2.
* Documentation: a security page, and the supported GPU generations and minimum NVIDIA driver version of every released artefact.

**Breaking change to OpenAPI** - regenerate the client (`jfjoch-client` 1.0.0-rc.162, `frontend/src/client`):
* `dataset_settings.images_per_file` is no longer `default: 1000` and no longer accepts `0`; it is optional, and its minimum is 1. A client sending `0` (previously "one file for the whole run") is now rejected - omit the field instead, which for a rotation sweep gives the same single file.
* `file_writer_format` now defaults to `NXmxVDS`, matching the server's own default and the layout recommended for DIALS, XDS and CrystFEL. A generated client that fills in schema defaults and does not set the format explicitly will write VDS masters where it previously wrote legacy ones; set `NXmxLegacy` explicitly to keep them.

---------

Co-authored-by: jungfrau <jungfrau@mx-aare-test.psi.ch>
Reviewed-on: #72
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-08-25 08:21:39 +02:00

9.1 KiB

Release contents

This page describes what a Jungfraujoch release ships and what each artefact needs on the target machine — which CPU instruction set the binaries were compiled for, which CUDA toolkit they were built against, and which runtime libraries are bundled rather than expected from the host.

The artefacts in the table below are built and published by the continuous-integration pipeline (.gitea/workflows/build_and_test.yml) when a tag is pushed. For how to install and configure the result see Deployment; for the package-repository URLs see Linux package repositories.

Artefacts

Artefact Distributed via Contains
.rpm / .deb packages package repositories The full server stack: jfjoch (broker, frontend, FPGA and detector tools), jfjoch-writer, jfjoch-viewer (incl. the XDS plugin), jfjoch-driver-dkms
jfjoch_viewer-<version>-linux-cuda<major>.tgz, ...-linux-cpu.tgz Gitea release page Portable Linux viewer package: jfjoch_viewer, rugnux, jfjoch_extract_hkl, jfjoch_recompress and the license notices
jfjoch-viewer-<version>-win64-cuda<major>.exe, ...-win64-cpu.exe Gitea release page Windows installer with the same four programs, plus the Qt runtime
jfjoch-writer .rpm / .deb Gitea release page The writer alone, for a file-writing machine without the rest of the stack
libjfjoch_xds_plugin.so.<version> Gitea release page XDS HDF5 read plugin (built on RHEL 8); see Integration with MX software
jfjoch-client PyPI and the Gitea PyPI index Generated Python OpenAPI client
Documentation Read the Docs and the gitea-pages branch This documentation set

The FPGA firmware (.mcs) images are attached to the release as well. The firmware is stable and is carried from version to version, and is rebuilt with Vivado (see FPGA smartNIC) when it needs to change — so a card keeps its image across a software upgrade unless the release notes say otherwise.

CPU instruction set

The architecture flags live in the CI configuration rather than in CMakeLists.txt, so a site building from source picks its own (x86-64-v4 on an AVX-512 cluster, -march=native, or the plain baseline the compiler defaults to). The released binaries are compiled to a fixed floor:

Release Flags Minimum CPU
Linux (all packages, and the portable .tgz) -march=x86-64-v3 -flto=auto AVX2 + FMA + BMI2 — Intel Haswell (2013) / AMD Zen (2017) and newer
Windows installer /arch:AVX AVX — Intel Sandy Bridge (2011) / AMD Bulldozer and newer

The Windows floor is lower because MSVC has no spelling for the x86-64-v2 level; /arch:AVX is the nearest one and implies SSE4.1/4.2, which is what actually matters — without it Eigen has no vectorised round and falls back to a libm call per element. Link-time optimisation is applied on Linux only.

A binary will fault with an illegal instruction on a CPU below its floor. If you must run on older hardware, build from source without the flags.

Operating-system floor

The .rpm / .deb packages are built per distribution (RHEL/Rocky 8 and 9, Ubuntu 22.04 and 24.04) and are tied to it. The portable viewer .tgz is built on RHEL 8, the oldest supported distribution, so its glibc floor is low enough to run on any newer Linux — that is what it is for, and why it replaces the per-distro packaging of the viewer on the release page. The Windows installer is built and verified on Windows 11.

CUDA and non-CUDA builds

Every binary artefact is released in two variants, cuda<major> and cpu. The CUDA variant adds the GPU fast-feedback indexer (ffbidx), the GPU FFT indexer and GPU image processing; the CPU-only variant runs the same pipeline on the CPU with the FFTW indexer, at much lower throughput.

The CUDA toolkit used is the one on the corresponding build machine: CUDA 12 for the RHEL 8 packages, CUDA 13 for RHEL 9, Ubuntu and Windows. The major version is part of the artefact and repository name, so a download is self-identifying. Building from source needs CUDA 12.8 or newer.

A CUDA build does not require a CUDA machine. Of the CUDA components only cuFFT is linked dynamically — the CUDA runtime and the fast-feedback indexer are linked statically — and cuFFT itself has no link-time dependency on the NVIDIA driver library. Jungfraujoch asks how many CUDA devices are present at start-up and treats "none" (including "no driver installed") as zero GPUs, falling back to the CPU path. So a CUDA build starts and runs correctly on a machine with no NVIDIA GPU at all, provided the cuFFT runtime can be loaded:

  • Portable .tgz and Windows installer — cuFFT is part of the distribution, shipped next to the executable (on Linux found through an $ORIGIN rpath). Nothing else is needed: no CUDA toolkit, and on a GPU machine only the NVIDIA driver.
  • .rpm / .deb — cuFFT comes from the distribution's own CUDA packages, so that one dependency is managed centrally with the rest of CUDA. Install the cuFFT package alongside, or use the nocuda repositories on a machine where CUDA is not wanted.

The cuFFT runtime is large (the Windows DLL is ~256 MB), so the CUDA artefacts are correspondingly bigger than the CPU ones — the other reason for shipping both.

On a machine with an NVIDIA GPU, take the CUDA variant: only that one uses the GPU.

GPU generations and the NVIDIA driver

A CUDA variant carries compiled device code for a fixed set of GPU generations, and which generations those are follows from the CUDA toolkit it was built with. The CUDA runtime is linked statically, so the only NVIDIA component the target machine has to supply is the driver — there is no CUDA-toolkit version requirement on the host.

Artefact CUDA toolkit GPU generations Minimum driver
RHEL 8 packages, portable Linux .tgz 12.9 Volta (V100) through Blackwell: sm_70, 75, 80, 86, 89, 90, 100, 120, 121 525.60.13
RHEL 9 and Ubuntu packages, Windows installer 13.x Turing (T4) through Blackwell: the same list without sm_70 580.65.06 (Linux), R580 (Windows)
any cpu / nocuda variant none

A V100 needs the CUDA 12 build. CUDA 13 dropped offline compilation for Volta, and the PTX a fatbin also carries only ever JIT-compiles forwards, so a CUDA 13 artefact contains nothing a V100 can execute: every kernel launch fails with no kernel image is available for execution on the device. On a V100 host take the RHEL 8 packages or the portable Linux .tgz. Nothing older than Volta is supported.

Newer GPUs never need a newer build — the highest generation in the list ships PTX as well as SASS, which the driver JIT-compiles for a GPU that came out after the release.

The minimum driver above is the floor for the whole CUDA major version, which is what applies here because the CUDA runtime is statically linked (CUDA minor version compatibility). Newer drivers are always fine; they are backward compatible. A driver from the same release as the build toolkit (575.57.08 for the CUDA 12.9 build, 610.43.02 for a CUDA 13.3 one) additionally rules out the single caveat of minor version compatibility — a call into a driver API newer than the installed driver, which fails with cudaErrorCallRequiresNewerDriver.

Windows installer

The Windows artefact covers jfjoch_viewer and the portable analysis CLIs only; the rest of Jungfraujoch (broker, receiver, FPGA host, detector control) is Linux-only.

The toolchain bounds of the released installer are:

  • Visual Studio 2026 with the C++ (MSVC) toolset. MSVC is not optional — CUDA on Windows builds through it — and it is what the release is compiled with.
  • CUDA Toolkit 13.3 for the cuda13 variant.
  • Qt 6.11 for MSVC (msvc2022_64), including Qt Charts.
  • Ninja as the generator; zlib and Eigen 3.4 supplied from a build prefix.

The installer is generated with NSIS and bundles the Qt runtime (via windeployqt) and, on the CUDA variant, the cuFFT DLL — so the end user installs neither Qt nor a CUDA toolkit. The two variants share an install directory and Start Menu group and replace each other (CUDA is a strict superset); they are told apart by the installer filename and the Add/Remove Programs entry:

Build Installer file Add/Remove Programs
CUDA (default) jfjoch-viewer-<version>-win64-cuda<major>.exe Jungfraujoch (CUDA)
CPU-only jfjoch-viewer-<version>-win64-cpu.exe Jungfraujoch (CPU)

To build the viewer yourself on Windows, see jfjoch_viewer ▸ Building from source on Windows.

Licenses

Every package variant carries the project license, the third-party manifest and the verbatim license texts of the bundled dependencies under share/doc/jfjoch. See Third-party software notices.