Build Packages / Create release (push) Successful in 24s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m29s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m43s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 8m27s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 9m53s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 9m58s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 11m22s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 13m39s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 18m37s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 16m32s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 24m11s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m30s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 19m3s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 20m23s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 19m41s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 21m0s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m38s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 14m33s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 17m55s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 20m50s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 18m38s
Build Packages / Unit tests (push) Successful in 1h46m14s
* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports. * jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls. * Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results. * Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable. * Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate. * Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do. * Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence. * Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags. * Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check. * Rugnux: Clear error messages when a data set needs more GPU or host memory than is available. Reviewed-on: #83 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
97 lines
4.0 KiB
C++
97 lines
4.0 KiB
C++
// SPDX-FileCopyrightText: 2025 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
#pragma once
|
|
|
|
|
|
#include <thread>
|
|
#include <mutex>
|
|
#include <condition_variable>
|
|
#include <queue>
|
|
#include <functional>
|
|
#include <future>
|
|
#include <vector>
|
|
#include <optional>
|
|
#include <memory>
|
|
#include <latch>
|
|
|
|
#include "../common/JFJochMessages.h"
|
|
#include "../common/DiffractionSpot.h"
|
|
#include "../common/DiffractionExperiment.h"
|
|
#include "Indexer.h"
|
|
|
|
// When a worker builds its indexers.
|
|
//
|
|
// Preconstruct (the default, and what the online service needs): every indexer the requested
|
|
// algorithm could resolve to is built while the pool is being constructed, so it is resident and
|
|
// planned before the first frame arrives. Building one means a cuFFT plan plus ~0.6 GB of device
|
|
// allocation, and the first cuFFT call in a process also pays the library's one-time init (0.3 s
|
|
// measured); jfjoch_broker cannot have any of that land on the live data path, so holding an indexer
|
|
// the resolved algorithm may never dispatch to is an accepted cost there.
|
|
//
|
|
// OnFirstUse (offline batch - rugnux, jfjoch_viewer): nothing is latency-critical, so the indexer is
|
|
// built on the first frame that needs it. The algorithm is only resolved from the frame's
|
|
// DiffractionExperiment (rotation -> FFT, stills with a known cell -> FFBIDX), so preconstructing
|
|
// leaves each worker holding a fully allocated FFTIndexerGPU - 0.6 GB of cuFFT plan and histograms -
|
|
// that can never be dispatched to.
|
|
enum class IndexerConstruction { Preconstruct, OnFirstUse };
|
|
|
|
// Force cuFFT's one-time library initialisation - 0.3 s measured, paid by whichever call happens to
|
|
// be first - so it can be got out of the way on a background thread at startup rather than landing in
|
|
// the middle of the first pass with nothing overlapping it. No-op without a GPU.
|
|
void WarmUpCuFFT();
|
|
|
|
class IndexerThread {
|
|
struct TaskInput {
|
|
const DiffractionExperiment &experiment;
|
|
const std::vector<Coord> &recip;
|
|
const bool severity_only;
|
|
};
|
|
|
|
// Held by value: with IndexerConstruction::OnFirstUse the worker builds its indexer long after
|
|
// the constructor returned, and pools are routinely built from a temporary - for instance
|
|
// IndexerThreadPool(experiment.GetIndexingSettings()), which returns by value.
|
|
const IndexingSettings settings_;
|
|
const IndexerConstruction construction_;
|
|
|
|
bool stop = false;
|
|
enum class TaskState {STARTING, IDLE, READY, RUNNING, COMPLETED, ERROR} state = TaskState::STARTING;
|
|
std::mutex m;
|
|
std::condition_variable c_running;
|
|
std::condition_variable c_start;
|
|
std::condition_variable c_done;
|
|
std::unique_ptr<IndexerResult> result = nullptr;
|
|
std::unique_ptr<TaskInput> task_input = nullptr;
|
|
std::thread worker_thread;
|
|
|
|
void Worker(int threadid);
|
|
public:
|
|
IndexerThread(const IndexingSettings& settings, int threadid, IndexerConstruction construction);
|
|
~IndexerThread();
|
|
std::unique_ptr<IndexerResult> Run(const DiffractionExperiment &experiment, const std::vector<Coord> &recip,
|
|
bool severity_only = false);
|
|
void Finalize();
|
|
};
|
|
|
|
class IndexerThreadPool {
|
|
std::mutex m;
|
|
std::condition_variable c;
|
|
std::vector<uint8_t> worker_busy;
|
|
size_t worker_free_count;
|
|
std::vector<std::unique_ptr<IndexerThread>> tasks;
|
|
const int64_t viable_cell_min_spots;
|
|
const bool blocking;
|
|
const IndexingSettings settings_;
|
|
int GetFreeWorker();
|
|
public:
|
|
IndexerThreadPool(const IndexingSettings& settings,
|
|
IndexerConstruction construction = IndexerConstruction::Preconstruct);
|
|
// The settings the pool's indexers were built with.
|
|
[[nodiscard]] const IndexingSettings &Settings() const { return settings_; }
|
|
// severity_only skips indexing and produces just the spindle severity - see Indexer::Run.
|
|
IndexerResult Run(const DiffractionExperiment& experiment, const std::vector<Coord>& recip,
|
|
bool severity_only = false);
|
|
};
|
|
|
|
|
|
|