Files
Jungfraujoch/image_analysis/bragg_prediction/BraggPrediction.h
T
leonarski_fandjungfrau 4dc2534dbf
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 18m57s
Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 16m55s
Build Packages / build:windows:cuda (push) Successful in 18m48s
Build Packages / build:viewer-tgz:cpu (push) Successful in 13m10s
Build Packages / build:viewer-tgz:cuda (push) Successful in 14m45s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 22m23s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 20m12s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 23m7s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 20m43s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 23m9s
Build Packages / XDS test (durin plugin) (push) Successful in 12m26s
Build Packages / build:rpm (rocky9) (push) Successful in 24m58s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 23m20s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (JFJoch plugin) (push) Successful in 12m37s
Build Packages / build:rpm (rocky8) (push) Successful in 27m58s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m38s
Build Packages / Build documentation (push) Successful in 59s
Build Packages / DIALS test (push) Successful in 23m16s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m38s
v1.0.0.rc-162 (#72)
**Files written by Jungfraujoch now import correctly in DIALS, XDS and pyFAI.** A tilted detector, a grid scan, a still recorded at a goniometer position, and saturated or unreadable pixels were each described in a way that a third-party program acted on wrongly. If you process Jungfraujoch data outside Jungfraujoch, prefer this release to any earlier one.

* HDF5: the detector tilt (`rot1`/`rot2`/`rot3`) is exported correctly in the NXmx transformation chain; untilted geometries are unaffected.
* HDF5: a still recorded at a goniometer position is no longer read back as a single image, and a grid scan records a stationary spindle so a program that requires a rotation axis can open it.
* HDF5: the sample transformation chain is written in mounting order, with a Smargon head position told apart from the spindle, one entry per image, `module_offset` as a float unit vector, and `offset_units` on every offset.
* HDF5: saturated, underloaded and unreadable pixels are described so a downstream program masks them - `saturation_value`, `underload_value`, `error_value` and `bit_depth_readout` are written correctly, and a data file missing next to a VDS master reads as the error marker rather than as zero counts.
* HDF5: the rotation axis is read back under whatever name it carries, and `mirror_y` records whether the assembled image is mirrored in Y relative to the detector's raw readout.
* A grid scan and a goniometer axis can both be set; they are no longer alternatives.
* `images_per_file` is chosen from the acquisition when it is not given: a rotation sweep of at most 20000 images goes into a single data file, a grid scan splits on whole fast-axis rows, and stills and serial keep 1000.
* The writer refuses a stream whose start message declares a different pixel format than its images carry, and a DECTRIS detector sending signed images is no longer declared unsigned.
* The image stream can carry the sample transformation chain (`transformations`, in the END message); a producer that does not send it gets the same chain built by the writer.
* rugnux: fixing the space group with `-S` no longer prevents the lattice from being found - a lattice indexed in a different setting is reindexed into that group's own setting, and a run whose crystal does not have that group's lattice stops and names the cell it indexed as, rather than reporting statistics that cannot describe it.
* rugnux: the per-image resolution estimate now predicts the resolution the merged data reach rather than the highest-resolution spot found, and is reported as `SPOT_RESOLUTION_ESTIMATE`.
* rugnux: two runs of the same command on the same images produce the same merged intensities; the azimuthal profile written alongside them is not yet reproducible in the same way.
* rugnux: the offline lattice refinement is bounded by iterations rather than by a wall clock, so a loaded machine can no longer refine to a different lattice; a live acquisition keeps its real-time bound.
* rugnux: the detector-frame modulation correction is fitted on a grid spanning the detector, so whether it is applied no longer depends on how far integration reached.
* rugnux: the geometry pre-pass no longer writes `<prefix>_01.mtz`, `_01.cif`, `_01.hkl` and `_01_image.dat`; the refined second pass writes those files under `<prefix>`, and that is the result to use.
* rugnux: `_process.h5` describes the pixel format of the images it links to, and is written on a thread of its own.
* rugnux: the detector geometry is also logged in XDS's convention (`ORGX`/`ORGY`, detector axis vectors, rotation axis), so it can be compared with an XDS refinement.
* rugnux: an image integrated in pyFAI through the `.poni` file written by `--mode calibration` comes out with the correct azimuth, and the file declares pyFAI's `orientation`, which needs pyFAI 2024.01 or newer. Radial integration is unchanged.
* rugnux: a rotation run is substantially faster throughout - beam-stop detection, first-pass indexing, geometry refinement, integration, scaling and merging - and observations outside the scaling resolution range are dropped as they are ingested. The refined geometry, the space group chosen and the merged statistics are unchanged.
* Faster spot finding and indexing, on the broker as well as in rugnux; the spots found and the lattices indexed are unchanged.
* A run reserves substantially less GPU memory: nothing is allocated for buffers that are never read, and a worker builds only the engines it uses.
* rugnux: with `-N` left at its default the per-image loop of `--mode mx` uses at most 16 workers per GPU, rather than one per hardware thread; an explicit `-N` is obeyed as given.
* CUDA 12 builds now contain device code for Volta, so the RHEL 8 packages and the portable Linux `.tgz` run on a V100; the CUDA 13 artefacts (RHEL 9, Ubuntu, Windows) remain Turing and newer.
* The build resolves a single Eigen for the whole project, and refuses to configure if Ceres picks up a different one; a build that mixed two Eigen versions was undefined behaviour and crashed at -O2.
* Documentation: a security page, and the supported GPU generations and minimum NVIDIA driver version of every released artefact.

**Breaking change to OpenAPI** - regenerate the client (`jfjoch-client` 1.0.0-rc.162, `frontend/src/client`):
* `dataset_settings.images_per_file` is no longer `default: 1000` and no longer accepts `0`; it is optional, and its minimum is 1. A client sending `0` (previously "one file for the whole run") is now rejected - omit the field instead, which for a rotation sweep gives the same single file.
* `file_writer_format` now defaults to `NXmxVDS`, matching the server's own default and the layout recommended for DIALS, XDS and CrystFEL. A generated client that fills in schema defaults and does not set the format explicitly will write VDS masters where it previously wrote legacy ones; set `NXmxLegacy` explicitly to keep them.

---------

Co-authored-by: jungfrau <jungfrau@mx-aare-test.psi.ch>
Reviewed-on: #72
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-08-25 08:21:39 +02:00

119 lines
7.3 KiB
C++
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
// SPDX-FileCopyrightText: 2025 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <vector>
#include "../../common/CrystalLattice.h"
#include "../../common/DiffractionExperiment.h"
#include "../../common/Reflection.h"
struct BraggPredictionSettings {
float high_res_A = 1.5;
float ewald_dist_cutoff = 0.0005;
// Per-index half-widths of the box the predictor walks: h runs -max_h..+max_h, and so on. One limit
// per axis rather than one cube, because each index is bounded by its OWN axis (|h| <= a/d_min), so
// a cube sized for the longest axis walks the short ones far past anything the resolution cut can
// keep - on a 149/83/226 A cell that is ~16x the candidates a per-axis box generates.
int max_h = 100;
int max_k = 100;
int max_l = 100;
char centering = 'P';
float wedge_deg = 0.1f;
float mosaicity_deg = 0.2f;
float min_zeta = 0.05;
float mosaicity_multiplier = 4.0;
// Relative X-ray bandwidth Δλ/λ expressed as a Gaussian sigma (0 = monochromatic).
// Stills: the Ewald-shell acceptance is thickened radially per reflection by
// σ_bw = |recip_z|·bandwidth_sigma (= bλ/2d²), so the 1/d² pink-beam smear no
// longer clips high-resolution reflections.
// Rotation: differentiating Bragg's law at fixed d gives an extra rocking width
// Δθ = bandwidth_sigma·tan(θ_B), a spread in the same glancing angle the mosaic spread
// smears, so it adds to σ_M in quadrature. It is NOT divided by ζ: rotating the crystal
// by Δφ changes θ by ζ·Δφ, so the 1/ζ that turns an angular width into a rotation width is
// already the one the partiality applies to σ_M. CalcMosaicityXDS deconvolves the same term
// out of the fitted σ_M, so it is not counted twice.
float bandwidth_sigma = 0.0f;
};
class BraggPrediction {
protected:
// Not const: on the GPU path the buffer grows to fit a frame that predicts more than it holds.
int max_reflections;
std::vector<Reflection> reflections;
// Make room for `count` reflections. Overridden where device buffers have to follow. Called only
// when a frame predicted more than the current capacity, so a run pays for it a handful of times.
//
// NOTE: only the GPU Calc overrides call this. BraggPrediction::Calc and BraggPredictionRot::Calc
// stop filling at max_reflections instead, silently - and because that cap is applied inside the
// h/k/l walk, before the resolution test, what survives is the low-|h| block rather than the
// reflections nearest the Ewald sphere. A cell large enough to overflow 20000 therefore yields
// different merged reflections on a CPU-only build than on a GPU one.
virtual void GrowCapacity(int count);
// Deterministically cap Calc's output at output_limit: if more were predicted, keep the best-recorded
// ones - largest partiality on the rotation path, and, since partiality is 1 for every still, smallest
// excitation error on the still path - with hkl breaking what is left. Returns the kept count. Below
// the cap it is a no-op. Call at the end of every Calc override.
int TruncateToOutput(int count);
// Put the first `count` predicted reflections in an order that depends only on the reflections
// themselves. The GPU predictors append at an atomic counter, so a reflection's POSITION in the
// array is decided by the order the blocks happened to finish - and that position is not private
// to the predictor: BraggOwnerKey packs it into the owner map as the tie-break between two centres
// equidistant from a shared pixel, and every downstream sort that is not a total order (the
// ingest and post-refine bucket sorts) resolves its ties by the order it receives. Two runs of the
// same binary on the same image therefore integrated a different set of reflections. hkl is a
// property of the reflection; delta_phi separates the two rocking solutions one hkl can have.
// Call before TruncateToOutput, whose own pick is then reproducible as well.
void OrderOutput(int count);
// Scratch for OrderOutput. A Reflection is ~88 bytes and a large cell predicts tens of thousands
// of them per frame, so sorting the structs themselves moves several megabytes an image; sorting
// a 20-byte key and gathering once is the same order for a third of the traffic. Members rather
// than locals so the two allocations happen once per engine, not once per image.
struct OrderKey { int32_t h, k, l; float delta_phi_deg; int32_t index; };
std::vector<OrderKey> order_keys;
std::vector<Reflection> order_scratch;
public:
// The prediction buffer holds up to kPredictionCapacity reflections so a strong lattice does not
// overflow it. Calc returns at most output_limit, the number that flows downstream and is serialized - kept low so the
// per-image reflection list stays within the frame transport headroom.
// Starting size. On the GPU path the buffer grows to whatever a frame actually predicts
// (GrowCapacity), so a large cell is not truncated there; the CPU path still caps at this value,
// see the note on GrowCapacity. It used to be a hard cap on both, and overflowing it was lossy
// and NON-DETERMINISTIC - the GPU kernels claim slots with an atomicAdd, so which reflections
// survived depended on block scheduling and changed between runs of the same command.
static constexpr int kPredictionCapacity = 20000;
// How many reflections may flow downstream per image, offline. Sized for a large unit cell: a
// ~2.8e6 A^3 cell predicts up to ~44000 per frame at 2.4 A. Truncating below what the frame really
// has costs more than it saves - the selection keeps the best-recorded reflections, and the rotation
// combine rebuilds a full FROM the partials it drops (measured on such a crystal: CC1/2 98 -> 60,
// ISa 8.6 -> 1.8). DiffractionExperiment's image-buffer headroom is derived from this, so the
// transport can carry what the analysis produces.
static constexpr int kPredictionOutput = 65536;
// What the ONLINE path may carry per image. The acquisition transports every reflection list through
// a fixed-size image-buffer slot, and the slot size divides a fixed total - so sizing the slot for
// kPredictionOutput would cut the number of slots, and with it the receiver's ability to absorb a
// burst, by about three. Online keeps the transport-sized limit it has always had; offline, which
// has no such budget, keeps the full one. DiffractionExperiment's buffer headroom derives from THIS.
static constexpr int kOnlineMaxReflections = 10000;
// How many reflections Calc may return. A caller with a tighter cap than the offline one - online,
// whose transport slot holds kOnlineMaxReflections - sets its own here, so the surplus is dropped
// before it is integrated rather than integrated and then thrown away.
int output_limit = kPredictionOutput;
explicit BraggPrediction(int max_reflections = kPredictionCapacity);
virtual ~BraggPrediction() = default;
virtual int Calc(const DiffractionExperiment &experiment, const CrystalLattice &lattice,
const BraggPredictionSettings &settings);
const std::vector<Reflection> &GetReflections() const;
};