A cold run on a spinning disk waited on the disk twice over. The CBF header
scan read 256 kB from every frame on eight threads that each strode through
their own share of the sweep, so they drifted apart and the scan became a
seek storm (34 s for 2400 frames here); and after it, the pre-scan and the
first-pass indexing touch a few hundred frames and leave the disk idle until
the first image loop reads everything at seek-bound rates.
- ReadAhead (reader/): once the dataset is open, rugnux starts eight threads
that read the data files - HDF5 data files (legacy, VDS or the integrated
master) or the per-frame CBF/marCCD/SMV files - in 4 MB pieces taken
strictly in order, into a throwaway buffer. One stream reads this disk at
125 MB/s, eight in-order streams at 190 MB/s, 32 at 157 MB/s. It never gets
more than a quarter of MemAvailable (GlobalMemoryStatusEx on Windows, 4 GiB
where there is no figure) ahead of what ReadRawImage has handed out, so a
dataset bigger than the cache does not evict its own start, and it stops
with the reader. Plain ifstream reads: portable, no POSIX calls.
- Header scans (CBF, marCCD, SMV) hand the files out in order from an atomic
counter (sweep::ForEachInOrder) instead of striding: 18 s -> 12 s for 2400
cold CBF headers. The CBF header is first read with a 16 kB probe and again
with the old 256 kB one only when the separator is not in it, so the parsed
header is exactly what it was: 12 s -> 6 s.
Output unchanged: p.hkl, p.mtz and p_unmerged.mtz md5-identical to the
rc173 baseline on 6toc (CBF, 2400 frames, 6.0 GB) and 9q41 (HDF5 VDS, 900
frames, 5.1 GB), and on 6z9g (HDF5, 12.8 GB) to the unmodified branch; myob
(p.hkl p.mtz p_P1.mtz p_unmerged.mtz) md5-identical to the reference.
Measured cold (files evicted with POSIX_FADV_DONTNEED before every run),
same code without this commit vs with it, on a shared box (load 20-70, other
agents reading the same disk, so single runs scatter by +-20 s):
6toc wall 61.7/62.7 -> 49.3/49.8 s (clean pairs); all data resident
after 62/51/50 -> 45/41/42 s
9q41 wall 67.3 -> 57.8 s (clean pair); resident after 58/46/43 -> 48/37/35 s
6z9g resident after 81 -> 69 s
The first image loop can look slower with this in CBF runs: the old 256 kB
header probes pulled ~70% of the data in as kernel readahead, so the old
loop started warm - after a 34 s header scan instead of 10 s.
Warm (myob, NVMe, cached): 19.35/20.00 s without, 19.76-20.16 s with; the
read-ahead then only copies 9.3 GB out of the page cache, 0.44 s wall and
3.4 CPU-s measured standalone.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
69 lines
3.5 KiB
C++
69 lines
3.5 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <cstddef>
|
|
#include <functional>
|
|
#include <string>
|
|
#include <vector>
|
|
|
|
// Where the frames of a one-file-per-image sweep sit on the spindle.
|
|
//
|
|
// A deposited series of CBF / marCCD / SMV files is not always complete: frames go missing between
|
|
// the beamline and the archive, and the numbers that are left are scattered through the range. Every
|
|
// one of these formats writes each image's own start angle in its own header, so the sweep is fully
|
|
// recoverable - but only if the frames are placed at the angles they state rather than laid out
|
|
// end to end. Reading a gapped series as contiguous compresses the sweep: one deposited 179.8 degree
|
|
// series of 1108 files out of 1800 came out as 111 degrees, and every frame past the first gap was
|
|
// analysed at the wrong spindle angle.
|
|
//
|
|
// So the sweep is a grid of SLOTS, one per rotation step over the range the headers span, and a frame
|
|
// occupies the slot its own angle puts it in. A slot with no file is a missing image, not a compressed
|
|
// out one: the reader hands it back as "nothing to read", which every image loop already handles, and
|
|
// the goniometer's start + increment * image_number is then the true angle of every image.
|
|
namespace sweep {
|
|
|
|
// What one frame's own header says about where it sits and how the instrument stood while it was
|
|
// taken. The last four are not used to place the frame - they are what the whole series has to agree
|
|
// on for it to be one sweep at all.
|
|
struct Frame {
|
|
std::string path;
|
|
double angle_deg = 0;
|
|
double increment_deg = 0;
|
|
double distance_m = 0;
|
|
double beam_x_px = 0;
|
|
double beam_y_px = 0;
|
|
double wavelength_A = 0;
|
|
};
|
|
|
|
struct Layout {
|
|
std::vector<std::string> files; // one entry per slot; empty where the series has no frame
|
|
double start_deg = 0; // angle of slot 0
|
|
double increment_deg = 0; // one slot, signed with the direction the sweep turns
|
|
size_t present = 0; // files that are really there (files.size() - the gaps)
|
|
};
|
|
|
|
// Place the frames on the sweep they describe. frames must be in collection order (the order the
|
|
// file names sort in) and must not be empty.
|
|
//
|
|
// Throws JFJochException, naming the frames, when the series is not one sweep: when the headers
|
|
// disagree about the instrument, when two frames claim the same angle, or when the angles do not sit
|
|
// on a single rotation step - a folder of screening images taken at scattered angles is not a sweep,
|
|
// and silently averaging it into one is how such a set comes out as an indexing failure instead of a
|
|
// clear refusal.
|
|
//
|
|
// A series that does not turn at all - one angle repeated, a grid scan or a set of stills - is left
|
|
// exactly as it came, one slot per file, with the header's nominal increment.
|
|
//
|
|
// logger_name is the reader's own logger, so the gap report names the format the user passed.
|
|
Layout Place(const std::vector<Frame> &frames, const std::string &logger_name);
|
|
|
|
// Call fn(i) for every i in [0, n) on up to 8 threads, handing the indices out in order. For reading
|
|
// every frame's header: the threads then read neighbouring files at any moment, so a spinning disk
|
|
// sees one sweep through the directory rather than eight positions drifting apart (2400 CBF headers
|
|
// from a cold spinning disk: 18 s with each thread striding through its own share, 12 s in order).
|
|
void ForEachInOrder(size_t n, const std::function<void(size_t)> &fn);
|
|
|
|
} // namespace sweep
|