Build Packages / build:rpm (rocky9_sls9) (push) Successful in 18m57s
Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 16m55s
Build Packages / build:windows:cuda (push) Successful in 18m48s
Build Packages / build:viewer-tgz:cpu (push) Successful in 13m10s
Build Packages / build:viewer-tgz:cuda (push) Successful in 14m45s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 22m23s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 20m12s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 23m7s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 20m43s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 23m9s
Build Packages / XDS test (durin plugin) (push) Successful in 12m26s
Build Packages / build:rpm (rocky9) (push) Successful in 24m58s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 23m20s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (JFJoch plugin) (push) Successful in 12m37s
Build Packages / build:rpm (rocky8) (push) Successful in 27m58s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m38s
Build Packages / Build documentation (push) Successful in 59s
Build Packages / DIALS test (push) Successful in 23m16s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m38s
**Files written by Jungfraujoch now import correctly in DIALS, XDS and pyFAI.** A tilted detector, a grid scan, a still recorded at a goniometer position, and saturated or unreadable pixels were each described in a way that a third-party program acted on wrongly. If you process Jungfraujoch data outside Jungfraujoch, prefer this release to any earlier one. * HDF5: the detector tilt (`rot1`/`rot2`/`rot3`) is exported correctly in the NXmx transformation chain; untilted geometries are unaffected. * HDF5: a still recorded at a goniometer position is no longer read back as a single image, and a grid scan records a stationary spindle so a program that requires a rotation axis can open it. * HDF5: the sample transformation chain is written in mounting order, with a Smargon head position told apart from the spindle, one entry per image, `module_offset` as a float unit vector, and `offset_units` on every offset. * HDF5: saturated, underloaded and unreadable pixels are described so a downstream program masks them - `saturation_value`, `underload_value`, `error_value` and `bit_depth_readout` are written correctly, and a data file missing next to a VDS master reads as the error marker rather than as zero counts. * HDF5: the rotation axis is read back under whatever name it carries, and `mirror_y` records whether the assembled image is mirrored in Y relative to the detector's raw readout. * A grid scan and a goniometer axis can both be set; they are no longer alternatives. * `images_per_file` is chosen from the acquisition when it is not given: a rotation sweep of at most 20000 images goes into a single data file, a grid scan splits on whole fast-axis rows, and stills and serial keep 1000. * The writer refuses a stream whose start message declares a different pixel format than its images carry, and a DECTRIS detector sending signed images is no longer declared unsigned. * The image stream can carry the sample transformation chain (`transformations`, in the END message); a producer that does not send it gets the same chain built by the writer. * rugnux: fixing the space group with `-S` no longer prevents the lattice from being found - a lattice indexed in a different setting is reindexed into that group's own setting, and a run whose crystal does not have that group's lattice stops and names the cell it indexed as, rather than reporting statistics that cannot describe it. * rugnux: the per-image resolution estimate now predicts the resolution the merged data reach rather than the highest-resolution spot found, and is reported as `SPOT_RESOLUTION_ESTIMATE`. * rugnux: two runs of the same command on the same images produce the same merged intensities; the azimuthal profile written alongside them is not yet reproducible in the same way. * rugnux: the offline lattice refinement is bounded by iterations rather than by a wall clock, so a loaded machine can no longer refine to a different lattice; a live acquisition keeps its real-time bound. * rugnux: the detector-frame modulation correction is fitted on a grid spanning the detector, so whether it is applied no longer depends on how far integration reached. * rugnux: the geometry pre-pass no longer writes `<prefix>_01.mtz`, `_01.cif`, `_01.hkl` and `_01_image.dat`; the refined second pass writes those files under `<prefix>`, and that is the result to use. * rugnux: `_process.h5` describes the pixel format of the images it links to, and is written on a thread of its own. * rugnux: the detector geometry is also logged in XDS's convention (`ORGX`/`ORGY`, detector axis vectors, rotation axis), so it can be compared with an XDS refinement. * rugnux: an image integrated in pyFAI through the `.poni` file written by `--mode calibration` comes out with the correct azimuth, and the file declares pyFAI's `orientation`, which needs pyFAI 2024.01 or newer. Radial integration is unchanged. * rugnux: a rotation run is substantially faster throughout - beam-stop detection, first-pass indexing, geometry refinement, integration, scaling and merging - and observations outside the scaling resolution range are dropped as they are ingested. The refined geometry, the space group chosen and the merged statistics are unchanged. * Faster spot finding and indexing, on the broker as well as in rugnux; the spots found and the lattices indexed are unchanged. * A run reserves substantially less GPU memory: nothing is allocated for buffers that are never read, and a worker builds only the engines it uses. * rugnux: with `-N` left at its default the per-image loop of `--mode mx` uses at most 16 workers per GPU, rather than one per hardware thread; an explicit `-N` is obeyed as given. * CUDA 12 builds now contain device code for Volta, so the RHEL 8 packages and the portable Linux `.tgz` run on a V100; the CUDA 13 artefacts (RHEL 9, Ubuntu, Windows) remain Turing and newer. * The build resolves a single Eigen for the whole project, and refuses to configure if Ceres picks up a different one; a build that mixed two Eigen versions was undefined behaviour and crashed at -O2. * Documentation: a security page, and the supported GPU generations and minimum NVIDIA driver version of every released artefact. **Breaking change to OpenAPI** - regenerate the client (`jfjoch-client` 1.0.0-rc.162, `frontend/src/client`): * `dataset_settings.images_per_file` is no longer `default: 1000` and no longer accepts `0`; it is optional, and its minimum is 1. A client sending `0` (previously "one file for the whole run") is now rejected - omit the field instead, which for a rotation sweep gives the same single file. * `file_writer_format` now defaults to `NXmxVDS`, matching the server's own default and the layout recommended for DIALS, XDS and CrystFEL. A generated client that fills in schema defaults and does not set the format explicitly will write VDS masters where it previously wrote legacy ones; set `NXmxLegacy` explicitly to keep them. --------- Co-authored-by: jungfrau <jungfrau@mx-aare-test.psi.ch> Reviewed-on: #72 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
146 lines
9.1 KiB
Markdown
146 lines
9.1 KiB
Markdown
# Release contents
|
|
|
|
This page describes **what a Jungfraujoch release ships and what each artefact needs on the target
|
|
machine** — which CPU instruction set the binaries were compiled for, which CUDA toolkit they were
|
|
built against, and which runtime libraries are bundled rather than expected from the host.
|
|
|
|
The artefacts in the table below are built and published by the continuous-integration pipeline
|
|
(`.gitea/workflows/build_and_test.yml`) when a tag is pushed. For *how* to install and configure the
|
|
result see [Deployment](DEPLOYMENT.md); for the package-repository URLs see
|
|
[Linux package repositories](REPOSITORIES.md).
|
|
|
|
## Artefacts
|
|
|
|
| Artefact | Distributed via | Contains |
|
|
| --- | --- | --- |
|
|
| `.rpm` / `.deb` packages | [package repositories](REPOSITORIES.md) | The full server stack: `jfjoch` (broker, frontend, FPGA and detector tools), `jfjoch-writer`, `jfjoch-viewer` (incl. the XDS plugin), `jfjoch-driver-dkms` |
|
|
| `jfjoch_viewer-<version>-linux-cuda<major>.tgz`, `...-linux-cpu.tgz` | Gitea release page | Portable Linux viewer package: `jfjoch_viewer`, `rugnux`, `jfjoch_extract_hkl`, `jfjoch_recompress` and the license notices |
|
|
| `jfjoch-viewer-<version>-win64-cuda<major>.exe`, `...-win64-cpu.exe` | Gitea release page | Windows installer with the same four programs, plus the Qt runtime |
|
|
| `jfjoch-writer` `.rpm` / `.deb` | Gitea release page | The writer alone, for a file-writing machine without the rest of the stack |
|
|
| `libjfjoch_xds_plugin.so.<version>` | Gitea release page | XDS HDF5 read plugin (built on RHEL 8); see [Integration with MX software](SOFTWARE_INTEGRATION.md) |
|
|
| `jfjoch-client` | [PyPI](https://pypi.org/project/jfjoch-client/) and the Gitea PyPI index | Generated Python OpenAPI client |
|
|
| Documentation | [Read the Docs](https://jungfraujoch.readthedocs.io) and the `gitea-pages` branch | This documentation set |
|
|
|
|
The FPGA firmware (`.mcs`) images are attached to the release as well. The firmware is stable and is
|
|
carried from version to version, and is rebuilt with Vivado (see [FPGA smartNIC](FPGA.md)) when it
|
|
needs to change — so a card keeps its image across a software upgrade unless the release notes say
|
|
otherwise.
|
|
|
|
## CPU instruction set
|
|
|
|
The architecture flags live in the CI configuration rather than in `CMakeLists.txt`, so a site
|
|
building from source picks its own (`x86-64-v4` on an AVX-512 cluster, `-march=native`, or the plain
|
|
baseline the compiler defaults to). The released binaries are compiled to a fixed floor:
|
|
|
|
| Release | Flags | Minimum CPU |
|
|
| --- | --- | --- |
|
|
| Linux (all packages, and the portable `.tgz`) | `-march=x86-64-v3 -flto=auto` | AVX2 + FMA + BMI2 — Intel Haswell (2013) / AMD Zen (2017) and newer |
|
|
| Windows installer | `/arch:AVX` | AVX — Intel Sandy Bridge (2011) / AMD Bulldozer and newer |
|
|
|
|
The Windows floor is lower because MSVC has no spelling for the `x86-64-v2` level; `/arch:AVX` is
|
|
the nearest one and implies SSE4.1/4.2, which is what actually matters — without it Eigen has no
|
|
vectorised `round` and falls back to a libm call per element. Link-time optimisation is applied on
|
|
Linux only.
|
|
|
|
A binary will fault with an illegal instruction on a CPU below its floor. If you must run on older
|
|
hardware, build from source without the flags.
|
|
|
|
## Operating-system floor
|
|
|
|
The `.rpm` / `.deb` packages are built per distribution (RHEL/Rocky 8 and 9, Ubuntu 22.04 and 24.04)
|
|
and are tied to it. The portable viewer `.tgz` is built on RHEL 8, the oldest supported
|
|
distribution, so its glibc floor is low enough to run on any newer Linux — that is what it is for,
|
|
and why it replaces the per-distro packaging of the viewer on the release page. The Windows
|
|
installer is built and verified on Windows 11.
|
|
|
|
## CUDA and non-CUDA builds
|
|
|
|
Every binary artefact is released in **two variants**, `cuda<major>` and `cpu`. The CUDA variant adds
|
|
the GPU fast-feedback indexer (`ffbidx`), the GPU FFT indexer and GPU image processing; the CPU-only
|
|
variant runs the same pipeline on the CPU with the FFTW indexer, at much lower throughput.
|
|
|
|
The CUDA toolkit used is the one on the corresponding build machine: **CUDA 12** for the RHEL 8
|
|
packages, **CUDA 13** for RHEL 9, Ubuntu and Windows. The major version is part of the artefact and
|
|
repository name, so a download is self-identifying. Building from source needs CUDA 12.8 or newer.
|
|
|
|
**A CUDA build does not require a CUDA machine.** Of the CUDA components only **cuFFT** is linked
|
|
dynamically — the CUDA runtime and the fast-feedback indexer are linked statically — and cuFFT
|
|
itself has no link-time dependency on the NVIDIA driver library. Jungfraujoch asks how many CUDA
|
|
devices are present at start-up and treats "none" (including "no driver installed") as zero GPUs,
|
|
falling back to the CPU path. So a CUDA build starts and runs correctly on a machine with no NVIDIA
|
|
GPU at all, provided the cuFFT runtime can be loaded:
|
|
|
|
- **Portable `.tgz` and Windows installer** — cuFFT is **part of the distribution**, shipped next to
|
|
the executable (on Linux found through an `$ORIGIN` rpath). Nothing else is needed: no CUDA
|
|
toolkit, and on a GPU machine only the NVIDIA driver.
|
|
- **`.rpm` / `.deb`** — cuFFT comes from the distribution's own CUDA packages, so that one
|
|
dependency is managed centrally with the rest of CUDA. Install the cuFFT package alongside, or use
|
|
the `nocuda` repositories on a machine where CUDA is not wanted.
|
|
|
|
The cuFFT runtime is large (the Windows DLL is ~256 MB), so the CUDA artefacts are correspondingly
|
|
bigger than the CPU ones — the other reason for shipping both.
|
|
|
|
On a machine with an NVIDIA GPU, take the CUDA variant: only that one uses the GPU.
|
|
|
|
## GPU generations and the NVIDIA driver
|
|
|
|
A CUDA variant carries compiled device code for a fixed set of GPU generations, and which
|
|
generations those are follows from the CUDA toolkit it was built with. The CUDA runtime is linked
|
|
statically, so the only NVIDIA component the target machine has to supply is the **driver** — there
|
|
is no CUDA-toolkit version requirement on the host.
|
|
|
|
| Artefact | CUDA toolkit | GPU generations | Minimum driver |
|
|
| --- | --- | --- | --- |
|
|
| RHEL 8 packages, portable Linux `.tgz` | 12.9 | Volta (V100) through Blackwell: `sm_70`, `75`, `80`, `86`, `89`, `90`, `100`, `120`, `121` | 525.60.13 |
|
|
| RHEL 9 and Ubuntu packages, Windows installer | 13.x | Turing (T4) through Blackwell: the same list **without** `sm_70` | 580.65.06 (Linux), R580 (Windows) |
|
|
| any `cpu` / `nocuda` variant | — | — | none |
|
|
|
|
**A V100 needs the CUDA 12 build.** CUDA 13 dropped offline compilation for Volta, and the PTX a
|
|
fatbin also carries only ever JIT-compiles *forwards*, so a CUDA 13 artefact contains nothing a V100
|
|
can execute: every kernel launch fails with *no kernel image is available for execution on the
|
|
device*. On a V100 host take the RHEL 8 packages or the portable Linux `.tgz`. Nothing older than
|
|
Volta is supported.
|
|
|
|
Newer GPUs never need a newer build — the highest generation in the list ships PTX as well as SASS,
|
|
which the driver JIT-compiles for a GPU that came out after the release.
|
|
|
|
The minimum driver above is the floor for the whole CUDA *major* version, which is what applies here
|
|
because the CUDA runtime is statically linked
|
|
([CUDA minor version compatibility](https://docs.nvidia.com/deploy/cuda-compatibility/minor-version-compatibility.html)).
|
|
Newer drivers are always fine; they are backward compatible. A driver from the same release as the
|
|
build toolkit (575.57.08 for the CUDA 12.9 build, 610.43.02 for a CUDA 13.3 one) additionally rules
|
|
out the single caveat of minor version compatibility — a call into a driver API newer than the
|
|
installed driver, which fails with `cudaErrorCallRequiresNewerDriver`.
|
|
|
|
## Windows installer
|
|
|
|
The Windows artefact covers `jfjoch_viewer` and the portable analysis CLIs only; the rest of
|
|
Jungfraujoch (broker, receiver, FPGA host, detector control) is Linux-only.
|
|
|
|
The toolchain bounds of the released installer are:
|
|
|
|
- **Visual Studio 2026** with the C++ (MSVC) toolset. MSVC is not optional — CUDA on Windows builds
|
|
through it — and it is what the release is compiled with.
|
|
- **CUDA Toolkit 13.3** for the `cuda13` variant.
|
|
- **Qt 6.11** for MSVC (`msvc2022_64`), including Qt Charts.
|
|
- Ninja as the generator; zlib and Eigen 3.4 supplied from a build prefix.
|
|
|
|
The installer is generated with NSIS and **bundles the Qt runtime** (via `windeployqt`) and, on the
|
|
CUDA variant, the cuFFT DLL — so the end user installs neither Qt nor a CUDA toolkit. The two
|
|
variants share an install directory and Start Menu group and replace each other (CUDA is a strict
|
|
superset); they are told apart by the installer filename and the Add/Remove Programs entry:
|
|
|
|
| Build | Installer file | Add/Remove Programs |
|
|
| --- | --- | --- |
|
|
| CUDA (default) | `jfjoch-viewer-<version>-win64-cuda<major>.exe` | `Jungfraujoch (CUDA)` |
|
|
| CPU-only | `jfjoch-viewer-<version>-win64-cpu.exe` | `Jungfraujoch (CPU)` |
|
|
|
|
To build the viewer yourself on Windows, see
|
|
[jfjoch_viewer ▸ Building from source on Windows](JFJOCH_VIEWER.md#building-from-source-on-windows).
|
|
|
|
## Licenses
|
|
|
|
Every package variant carries the project license, the third-party manifest and the verbatim
|
|
license texts of the bundled dependencies under `share/doc/jfjoch`. See
|
|
[Third-party software notices](THIRD_PARTY_NOTICES.md).
|