Build Packages / build:windows:nocuda (push) Successful in 16m8s
Build Packages / build:windows:cuda (push) Successful in 18m58s
Build Packages / build:viewer-tgz:cpu (push) Successful in 20m35s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m31s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 25m9s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 25m6s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m57s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m58s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m58s
Build Packages / XDS test (durin plugin) (push) Successful in 12m3s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 22m24s
Build Packages / build:rpm (rocky9) (push) Successful in 21m45s
Build Packages / Generate python client (push) Successful in 53s
Build Packages / build:rpm (rocky8) (push) Successful in 26m9s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m37s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m34s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 22m0s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m53s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m29s
Build Packages / DIALS test (push) Successful in 23m40s
Build Packages / Unit tests (push) Successful in 1h20m1s
Three costs before and around the image loop. Every image allocated a fresh buffer for its compressed chunk and resized it, which value-initialises, and the read then overwrote every byte. At a few megabytes a chunk the allocation is large enough to be mapped rather than reused, so the zeroing was page-fault bound and cost more than the read it preceded - twenty gigabytes of it over a long sweep. The buffer now uses an allocator that does not construct, and the two HDF5 read paths are templated on the allocator so every existing caller compiles unchanged. The rebind is deliberate: without it the vector base rebinds to the default allocator and the zeroing quietly returns. The bitshuffle decoder allocated a whole uncompressed frame in its constructor - seventy megabytes a worker, five hundred and fifty across the loop - for the route that decodes the shuffled image separately. That route is taken only when a bitshuffle block is too large for the fused kernel, which neither writer this pipeline reads produces, so on a real frame the buffer is allocated, never touched, and freed. It is now allocated where it is used. The comment two lines below already warned against sizing a buffer from the uncompressed size; the line above it had not been given the same treatment. The first call into cuFFT pays the library's one-time initialisation, and it landed in the middle of the first pass with nothing to overlap it. It is now forced on a background thread at startup, alongside the file open and the mapping build, in the manner the shadow finder already uses. Finally, the detector mask was copied into the start message whether or not a file would carry it, which a merging run does not. It is filled where a writer is constructed - both places one is constructed, the second being the fallback that writes a process file when nothing indexed. Faster on eleven of thirty-eight crystals and slower on none; the whole rotation test set falls from four minutes thirty to four minutes seventeen, with each binary repeating itself to within half a per cent. Space groups thirty-five of thirty-eight and no failures throughout, and every column of the comparison table is identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EGpGdgmJ8MyY9pCGWjktyi
132 lines
6.2 KiB
C++
132 lines
6.2 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <cstdint>
|
|
#include <map>
|
|
#include <memory>
|
|
#include <optional>
|
|
#include <vector>
|
|
|
|
#include "HDF5ImageLocator.h"
|
|
#include "../common/CompressedImage.h"
|
|
|
|
// Raw-pixel side of the reader. Turns a global image number into a CompressedImage, using
|
|
// HDF5ImageLocator to find the file (with its open-file cache). This is the part whose "links
|
|
// to files stay" constant: switching which master the per-image metadata is read from never
|
|
// touches it. Caller must hold the global hdf5_mutex (HDF5 is not thread-safe).
|
|
// Bit depth and signedness of /entry/data/data as it is STORED. Deliberately not the same thing as
|
|
// the experiment's image format: the reader hands every image out in a signed 32-bit container
|
|
// whatever the file holds (see HDF5MetadataSource), so an output file that links to the original
|
|
// images rather than writing its own must describe them with this, not with the experiment.
|
|
struct StoredPixelFormat {
|
|
int64_t bit_depth = 0;
|
|
bool is_signed = false;
|
|
};
|
|
|
|
class HDF5ImageSource {
|
|
public:
|
|
// Plain positional-read handle on a data file, opened alongside the HDF5 one. Owns the handle.
|
|
class RawFile {
|
|
public:
|
|
explicit RawFile(const std::string &path);
|
|
~RawFile();
|
|
RawFile(const RawFile &) = delete;
|
|
RawFile &operator=(const RawFile &) = delete;
|
|
|
|
bool IsOpen() const { return handle_ != -1; }
|
|
// Read `size` bytes from byte `address`. Positional and stateless, so any number of threads
|
|
// may call it on the same handle at once. Throws on a short read.
|
|
void ReadAt(void *dst, size_t size, uint64_t address) const;
|
|
|
|
private:
|
|
intptr_t handle_ = -1; // a file descriptor on POSIX, a HANDLE on Windows
|
|
};
|
|
|
|
// Where the bytes of one image are, and what they decode to. Everything needed to read an image
|
|
// without calling HDF5 again.
|
|
struct DirectChunk {
|
|
// Shared, not borrowed: GetRawImage drops the HDF5 lock before reading through this, so a
|
|
// concurrent Clear() - which ReadFile() and Close() both do - would otherwise free the file
|
|
// and close its descriptor under the reader.
|
|
std::shared_ptr<const RawFile> file;
|
|
uint64_t address = 0;
|
|
uint32_t size = 0;
|
|
hsize_t width = 0;
|
|
hsize_t height = 0;
|
|
CompressedImageMode mode{};
|
|
CompressionAlgorithm algorithm = CompressionAlgorithm::NO_COMPRESSION;
|
|
};
|
|
|
|
void Configure(HDF5ImageLocator::Layout layout);
|
|
void Clear();
|
|
|
|
[[nodiscard]] StoredPixelFormat GetStoredPixelFormat() const;
|
|
|
|
// Where image `global` physically lives. Also used by the metadata source to find the data
|
|
// file that holds a legacy/VDS image's per-image metadata.
|
|
HDF5ImageLocator::Location Resolve(int64_t global) const;
|
|
|
|
// Read the pixels at a resolved location into a CompressedImage backed by `buffer`. Templated on
|
|
// the allocator so a caller can hand over a buffer that does not zero what it is about to
|
|
// overwrite (RawByteBuffer).
|
|
template<class Alloc>
|
|
CompressedImage ReadImageAt(std::vector<uint8_t, Alloc> &buffer, const HDF5ImageLocator::Location &loc) const {
|
|
const auto &ds = GetDataset(loc);
|
|
const std::vector<hsize_t> start = {static_cast<hsize_t>(loc.local_index), 0, 0};
|
|
|
|
if (ds.direct_chunk)
|
|
ds.dataset->ReadDirectChunk(buffer, start);
|
|
else
|
|
ds.dataset->ReadVectorToU8(buffer, start, {1, ds.height, ds.width});
|
|
|
|
return {buffer.data(), buffer.size(), ds.width, ds.height, ds.mode, ds.algorithm};
|
|
}
|
|
|
|
// Ask HDF5 where image `loc` is in the file rather than asking it for the image. This is a
|
|
// lookup in the chunk index and nothing else - no read - so the mutex is held for a fraction of
|
|
// what an actual read costs, and the read itself then happens on any number of threads at once
|
|
// through ReadDirect(). Caller must hold hdf5_mutex.
|
|
//
|
|
// Empty when this file cannot be served that way: one chunk per image is what makes an image a
|
|
// single contiguous run of bytes, and a chunk that has never been written has no address at all.
|
|
// The caller falls back to ReadImageAt() then.
|
|
std::optional<DirectChunk> PrepareDirectRead(const HDF5ImageLocator::Location &loc) const;
|
|
|
|
// Read what PrepareDirectRead() found. Touches no HDF5 and no shared state, so it needs no
|
|
// mutex; this is the whole point of the two-step split.
|
|
static CompressedImage ReadDirect(RawByteBuffer &buffer, const DirectChunk &chunk);
|
|
|
|
std::vector<HDF5DataSourceMessage> GetSourceMapping(uint64_t first_image,
|
|
std::optional<uint64_t> image_count,
|
|
uint64_t total_images,
|
|
uint64_t stride = 1) const;
|
|
|
|
private:
|
|
HDF5ImageLocator locator_;
|
|
|
|
// /entry/data/data and everything asked of it here - its rank and dimensions, its element type,
|
|
// its chunking, its compression - are properties of the file, identical for every image in it.
|
|
// They used to be looked up again for each image: four HDF5 object opens per frame, inside the
|
|
// global hdf5_mutex that every worker thread queues on. Resolve them once per file instead.
|
|
//
|
|
// The entry keeps the file alive, so the pointer it is keyed by cannot be recycled underneath it
|
|
// and the dataset handle cannot outlive the file it belongs to.
|
|
struct OpenDataset {
|
|
std::shared_ptr<HDF5ReadOnlyFile> file;
|
|
std::unique_ptr<HDF5DataSet> dataset;
|
|
std::shared_ptr<RawFile> raw;
|
|
// HDF5 addresses count from the end of the user block, so they are file offsets only once
|
|
// its size is added. Zero for everything this project writes, but not for every file.
|
|
uint64_t user_block = 0;
|
|
hsize_t width = 0;
|
|
hsize_t height = 0;
|
|
CompressedImageMode mode{};
|
|
CompressionAlgorithm algorithm = CompressionAlgorithm::NO_COMPRESSION;
|
|
bool direct_chunk = false;
|
|
};
|
|
mutable std::map<const HDF5ReadOnlyFile *, OpenDataset> dataset_cache_;
|
|
const OpenDataset &GetDataset(const HDF5ImageLocator::Location &loc) const;
|
|
};
|