Files
leonarski_f 6dfe065365
Build Packages / Create release (push) Successful in 16s
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 8m27s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 9m15s
Build Packages / build:viewer-tgz:cpu (push) Successful in 10m11s
Build Packages / build:viewer-tgz:cuda (push) Successful in 12m6s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 15m44s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 16m1s
Build Packages / build:windows:nocuda (push) Successful in 17m29s
Build Packages / build:windows:cuda (push) Successful in 19m58s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 24m7s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 19m8s
Build Packages / build:rugnux:windows (push) Successful in 10m58s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 20m46s
Build Packages / Generate python client (push) Successful in 53s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 20m13s
Build Packages / Build documentation (push) Successful in 1m36s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 19m57s
Build Packages / build:rpm (rocky8) (push) Successful in 18m7s
Build Packages / build:rpm (rocky9) (push) Successful in 18m54s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 19m32s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 17m30s
Build Packages / Unit tests (push) Successful in 1h39m2s
v1.0.0-rc.172 (#82)
* Fixed `jfjoch_broker` cancelling every data collection with a CUDA "out of memory" error after long operation: GPU memory no longer leaks with each collection.
* Rugnux scales a rotation sweep until the per-frame scales settle instead of for a fixed three rounds, and says so when they did not - merged intensities, and the space group, resolution cut and frame rejection read off them, change accordingly; `--scaling-iterations` is now the cap on that loop (default 100).
* Rugnux places every frame of a marCCD, SMV or miniCBF series at the spindle angle its own header states, so a series with missing frames, or with angles written modulo 360, is no longer read at the wrong geometry or refused.
* Every rotation run writes two diagnostic files beside its reflections: `<prefix>_detector.jpg`, the detector projection with the pixel mask and the detected beam-stop shadow drawn on it, and `<prefix>_plot.txt`, one row per image.

Reviewed-on: #82
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-09-22 06:48:37 +02:00

114 lines
4.6 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <algorithm>
#include <cstddef>
#include <mutex>
#include <new>
#include <type_traits>
#include <vector>
// Large blocks that a whole pass's per-image reflection vectors are carved from.
//
// A rotation pass keeps every frame's integrated reflections until its scaling is done - thousands of
// vectors of a few megabytes each, gigabytes together, allocated by the image workers in the
// allocator's per-thread arenas. When the pass hands them back, that memory mostly stays in those
// arenas, as holes between whatever else the workers allocated, and the next pass's workers, being
// new threads, do not reuse it: on a fine-sliced long axis gigabytes of freed reflections were carried
// to the end of the run. Here a block is allocated in one piece (large enough that it is its own
// mapping) and returned in one piece once the last vector in it is gone, so a pass's reflections
// leave nothing behind.
//
// A vector is carved from the current block by bumping a pointer; the space of a vector freed early
// is only reclaimed with its whole block. That suits the use here: the vectors are written once,
// when the frame is integrated, and handed back together.
class ReflectionArena {
public:
ReflectionArena() = default;
ReflectionArena(const ReflectionArena &) = delete;
ReflectionArena &operator=(const ReflectionArena &) = delete;
~ReflectionArena() {
for (auto &b : blocks)
::operator delete(b.data);
}
void *Allocate(size_t bytes) {
bytes = (bytes + kAlign - 1) / kAlign * kAlign;
std::unique_lock ul(m);
if (blocks.empty() || blocks.back().size - blocks.back().used < bytes) {
const size_t size = std::max(kBlockBytes, bytes);
blocks.push_back(Block{static_cast<char *>(::operator new(size)), size, 0, 0});
}
Block &b = blocks.back();
void *p = b.data + b.used;
b.used += bytes;
b.live++;
return p;
}
void Deallocate(void *p) {
std::unique_lock ul(m);
for (size_t i = 0; i < blocks.size(); ++i) {
Block &b = blocks[i];
if (static_cast<char *>(p) < b.data || static_cast<char *>(p) >= b.data + b.size)
continue;
if (--b.live == 0) {
if (i + 1 == blocks.size()) {
b.used = 0; // the block being filled: keep it for the next vector
} else {
::operator delete(b.data);
blocks.erase(blocks.begin() + static_cast<std::ptrdiff_t>(i));
}
}
return;
}
}
private:
// 64 MiB: above the largest size glibc ever serves from its heaps (32 MiB), so every block is a
// mapping of its own and goes back to the system when freed. A frame's reflections are a few
// megabytes, so the tail a block cannot fit is a few percent of it.
static constexpr size_t kBlockBytes = size_t(64) << 20;
static constexpr size_t kAlign = alignof(std::max_align_t);
struct Block {
char *data;
size_t size, used, live;
};
std::mutex m;
std::vector<Block> blocks;
};
// A std::allocator that carves from a ReflectionArena, or is plain new/delete without one. A copy of
// a container is made on the heap (select_on_container_copy_construction), so only the containers the
// arena's owner fills itself live in the arena; a move carries the arena along with the storage.
template <class T>
struct ArenaAllocator {
using value_type = T;
using propagate_on_container_move_assignment = std::true_type;
using propagate_on_container_swap = std::true_type;
ReflectionArena *arena = nullptr;
ArenaAllocator() = default;
explicit ArenaAllocator(ReflectionArena *a) : arena(a) {}
template <class U> ArenaAllocator(const ArenaAllocator<U> &o) : arena(o.arena) {}
T *allocate(size_t n) {
if (arena)
return static_cast<T *>(arena->Allocate(n * sizeof(T)));
return static_cast<T *>(::operator new(n * sizeof(T)));
}
void deallocate(T *p, size_t) {
if (arena)
arena->Deallocate(p);
else
::operator delete(p);
}
ArenaAllocator select_on_container_copy_construction() const { return ArenaAllocator(); }
template <class U> bool operator==(const ArenaAllocator<U> &o) const { return arena == o.arena; }
template <class U> bool operator!=(const ArenaAllocator<U> &o) const { return arena != o.arena; }
};