A cold run on a spinning disk waited on the disk twice over. The CBF header
scan read 256 kB from every frame on eight threads that each strode through
their own share of the sweep, so they drifted apart and the scan became a
seek storm (34 s for 2400 frames here); and after it, the pre-scan and the
first-pass indexing touch a few hundred frames and leave the disk idle until
the first image loop reads everything at seek-bound rates.
- ReadAhead (reader/): once the dataset is open, rugnux starts eight threads
that read the data files - HDF5 data files (legacy, VDS or the integrated
master) or the per-frame CBF/marCCD/SMV files - in 4 MB pieces taken
strictly in order, into a throwaway buffer. One stream reads this disk at
125 MB/s, eight in-order streams at 190 MB/s, 32 at 157 MB/s. It never gets
more than a quarter of MemAvailable (GlobalMemoryStatusEx on Windows, 4 GiB
where there is no figure) ahead of what ReadRawImage has handed out, so a
dataset bigger than the cache does not evict its own start, and it stops
with the reader. Plain ifstream reads: portable, no POSIX calls.
- Header scans (CBF, marCCD, SMV) hand the files out in order from an atomic
counter (sweep::ForEachInOrder) instead of striding: 18 s -> 12 s for 2400
cold CBF headers. The CBF header is first read with a 16 kB probe and again
with the old 256 kB one only when the separator is not in it, so the parsed
header is exactly what it was: 12 s -> 6 s.
Output unchanged: p.hkl, p.mtz and p_unmerged.mtz md5-identical to the
rc173 baseline on 6toc (CBF, 2400 frames, 6.0 GB) and 9q41 (HDF5 VDS, 900
frames, 5.1 GB), and on 6z9g (HDF5, 12.8 GB) to the unmodified branch; myob
(p.hkl p.mtz p_P1.mtz p_unmerged.mtz) md5-identical to the reference.
Measured cold (files evicted with POSIX_FADV_DONTNEED before every run),
same code without this commit vs with it, on a shared box (load 20-70, other
agents reading the same disk, so single runs scatter by +-20 s):
6toc wall 61.7/62.7 -> 49.3/49.8 s (clean pairs); all data resident
after 62/51/50 -> 45/41/42 s
9q41 wall 67.3 -> 57.8 s (clean pair); resident after 58/46/43 -> 48/37/35 s
6z9g resident after 81 -> 69 s
The first image loop can look slower with this in CBF runs: the old 256 kB
header probes pulled ~70% of the data in as kernel readahead, so the old
loop started warm - after a 34 s header scan instead of 10 s.
Warm (myob, NVMe, cached): 19.35/20.00 s without, 19.76-20.16 s with; the
read-ahead then only copies 9.3 GB out of the page cache, 0.44 s wall and
3.4 CPU-s measured standalone.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
64 lines
2.8 KiB
C++
64 lines
2.8 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <string>
|
|
#include <vector>
|
|
|
|
#include "JFJochReader.h"
|
|
#include "MiniCBF.h"
|
|
|
|
// Reads a rotation sweep straight from a directory of PILATUS miniCBF files - the form most
|
|
// facilities still archive - with no conversion step and no libcbf.
|
|
//
|
|
// The sweep's geometry comes from the first file's header; the rotation angle of each image comes
|
|
// from its own header, which costs only the bytes up to the binary separator. Images are decoded on
|
|
// demand, one per call, so any number of workers can read at once - there is no global lock as there
|
|
// is on the HDF5 path, HDF5 not being thread-safe.
|
|
//
|
|
// A raw CBF carries no analysis results, so this reader has no spots, no reflections and no snapshots;
|
|
// the dataset it builds is the geometry, the mask and nothing else, exactly as a plain DECTRIS file
|
|
// with no /entry/MX would give.
|
|
class JFJochCBFReader : public JFJochReader {
|
|
std::vector<std::string> files_;
|
|
std::shared_ptr<JFJochReaderDataset> dataset_;
|
|
minicbf::Header header0_;
|
|
|
|
bool LoadImage_i(std::shared_ptr<JFJochReaderDataset> &dataset,
|
|
DataMessage &message,
|
|
std::vector<uint8_t> &buffer,
|
|
int64_t image_number,
|
|
bool update_dataset) override;
|
|
|
|
// Decodes one image into the caller's byte buffer (RawByteBuffer or std::vector<uint8_t>) and
|
|
// returns the image that points at it. The compressed file is read through scratch, which the
|
|
// caller keeps between frames.
|
|
template <class Buffer>
|
|
CompressedImage DecodeInto(int64_t image_number, Buffer &buffer, std::vector<uint8_t> &scratch) const;
|
|
|
|
std::vector<std::string> DataFiles() const override { return files_; }
|
|
|
|
public:
|
|
~JFJochCBFReader() override = default;
|
|
|
|
// True if the path names something this reader can open: a miniCBF file, or a directory holding
|
|
// at least one. Cheap - it reads a few kB at most.
|
|
static bool CanRead(const std::string &path);
|
|
|
|
// path is a directory of *.cbf, or one *.cbf inside the sweep to take the whole directory from.
|
|
void ReadFiles(const std::string &path);
|
|
|
|
[[nodiscard]] uint64_t GetNumberOfImages() const override;
|
|
void Close() override;
|
|
|
|
|
|
// Whether the sweep really has an image at this point. A deposited series can be missing frames,
|
|
// and they are kept as gaps so that every image keeps the spindle angle its own header states;
|
|
// ReadRawImage hands such a slot back as "nothing to read".
|
|
[[nodiscard]] bool HasImage(int64_t image_number) const;
|
|
|
|
bool ReadRawImage(int64_t image_number, JFJochReaderRawImage &image) override;
|
|
[[nodiscard]] std::vector<SpotToSave> ReadSpots(int64_t image) const override;
|
|
};
|