Build Packages / Create release (push) Successful in 21s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 9m40s
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 9m49s
Build Packages / build:viewer-tgz:cpu (push) Successful in 11m37s
Build Packages / build:viewer-tgz:cuda (push) Successful in 12m40s
Build Packages / build:windows:nocuda (push) Successful in 17m44s
Build Packages / build:windows:cuda (push) Successful in 20m13s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 14m41s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m59s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 15m5s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 14m35s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 15m53s
Build Packages / build:rugnux:windows (push) Successful in 11m29s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 18m51s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 18m43s
Build Packages / Generate python client (push) Successful in 51s
Build Packages / build:rpm (rocky8) (push) Successful in 18m51s
Build Packages / Build documentation (push) Successful in 1m21s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 18m38s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 18m24s
Build Packages / build:rpm (rocky9) (push) Successful in 19m19s
Build Packages / Unit tests (push) Successful in 1h37m15s
* Building Jungfraujoch no longer needs zlib or Eigen installed on the machine, and the dependencies the build fetches are pinned and updated to current releases. * rugnux: improvements in indexing, lattice selection and geometry post-refinement, which index crystals that previously returned no lattice and keep the better of the two geometries a run measures. * rugnux: improvements in beam-centre measurement, beam-stop detection and space-group determination. * rugnux: the unit cell reported with a determined space group now obeys that group - a cell whose symmetry was confirmed from the intensities is re-refined under it, and a cell the group cannot describe is reported with a warning rather than as it stands. * rugnux drops the stretches of a rotation sweep whose removal measurably improves the merged intensities and reports what became of every frame, and decides the resolution cut on the crystal's own diffraction rather than on its ice rings. * The rugnux results report is machine-readable - every line that is not `KEY= value` data starts with `#` - and states the build it was written by, its authorship and its terms of use (`REPORT_VERSION= 8`). * `jfjoch_viewer`: improvements in the file manager (CBF frames beside HDF5 datasets, a remembered root), the dataset plots, the inspector and the image statistics, plus a settable font size, a view of the rugnux results report, usable performance over a remote display (`ssh -X`) and a reset of all settings to defaults; the reciprocal-space window is removed. * Broker fixes around DECTRIS collections and dark-mask calibration: re-initialising after a run that never started no longer freezes the broker, a cancelled calibration is abandoned instead of reported as done, and a collection whose start message never arrives ends by itself. Reviewed-on: #79 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
50 lines
3.1 KiB
C++
50 lines
3.1 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
// Included only where CUDA is known to be present, i.e. under JFJOCH_USE_CUDA - the same rule as
|
|
// FFTIndexerGPU.h. Deliberately free of CUDA headers so a host translation unit (the parity test)
|
|
// can construct the engine.
|
|
|
|
#include "BeamCenterFFTEngine.h"
|
|
|
|
// The beam-centre transforms on cuFFT. The capture is one call per run on an image of tens of
|
|
// megapixels, so plans and device buffers are created inside the call that needs them and freed
|
|
// when it returns: the card is shared with the analysis workers, and a plan held across the run
|
|
// would be several hundred megabytes doing nothing.
|
|
class BeamCenterFFTGPU : public BeamCenterFFTEngine {
|
|
public:
|
|
BeamCenterConvSurfaces2D PointSurfaces(const std::vector<float> &a, const std::vector<float> &m,
|
|
int64_t h, int64_t w) override;
|
|
BeamCenterConvSurfaces1D LineSurfaces(const std::vector<float> &a, const std::vector<float> &m,
|
|
int64_t h, int64_t w, BeamCenterMirror mirror) override;
|
|
|
|
// The four convolutions are combined and searched where they already are, so the only thing
|
|
// that crosses PCIe is the shortlist. On a 16 Mpixel detector that is 1.16 GB of surfaces, a
|
|
// 1.4 s greedy search and a 0.4 s combination not done on the host - against 26 ms of
|
|
// transform, which is what the host work had become.
|
|
std::vector<BeamCenterFFTCandidate> PointShortlist(const std::vector<float> &a,
|
|
const std::vector<float> &m, int64_t h,
|
|
int64_t w,
|
|
const BeamCenterFFTSettings &settings,
|
|
double global_variance) override;
|
|
|
|
// The scored surface itself, brought back. Nothing in the pipeline asks for it - the shortlist
|
|
// is made on the device - but the parity test holds the device combination to
|
|
// BeamCenterPointScore element by element, which is the sharpest comparison available once the
|
|
// two paths no longer share that step.
|
|
[[nodiscard]] std::vector<float> PointScoreSurface(const std::vector<float> &a,
|
|
const std::vector<float> &m, int64_t h,
|
|
int64_t w, float min_pair_fraction,
|
|
double global_variance);
|
|
|
|
// Peak device memory the two phases need for an image of this size: buffers plus the cuFFT
|
|
// plans' own workspace.
|
|
[[nodiscard]] static size_t DeviceMemoryNeeded(int64_t width, int64_t height);
|
|
// Whether the visible device has that much free right now. The card is shared - with the
|
|
// analysis workers of this run and, on a development box, with whatever else is on it - so a
|
|
// capture that would not fit runs on the CPU instead of taking the card down with it.
|
|
[[nodiscard]] static bool FitsInDeviceMemory(int64_t width, int64_t height);
|
|
};
|