Build Packages / build:windows:nocuda (push) Successful in 16m8s
Build Packages / build:windows:cuda (push) Successful in 18m58s
Build Packages / build:viewer-tgz:cpu (push) Successful in 20m35s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m31s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 25m9s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 25m6s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m57s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m58s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m58s
Build Packages / XDS test (durin plugin) (push) Successful in 12m3s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 22m24s
Build Packages / build:rpm (rocky9) (push) Successful in 21m45s
Build Packages / Generate python client (push) Successful in 53s
Build Packages / build:rpm (rocky8) (push) Successful in 26m9s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m37s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m34s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 22m0s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m53s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m29s
Build Packages / DIALS test (push) Successful in 23m40s
Build Packages / Unit tests (push) Successful in 1h20m1s
Three costs before and around the image loop. Every image allocated a fresh buffer for its compressed chunk and resized it, which value-initialises, and the read then overwrote every byte. At a few megabytes a chunk the allocation is large enough to be mapped rather than reused, so the zeroing was page-fault bound and cost more than the read it preceded - twenty gigabytes of it over a long sweep. The buffer now uses an allocator that does not construct, and the two HDF5 read paths are templated on the allocator so every existing caller compiles unchanged. The rebind is deliberate: without it the vector base rebinds to the default allocator and the zeroing quietly returns. The bitshuffle decoder allocated a whole uncompressed frame in its constructor - seventy megabytes a worker, five hundred and fifty across the loop - for the route that decodes the shuffled image separately. That route is taken only when a bitshuffle block is too large for the fused kernel, which neither writer this pipeline reads produces, so on a real frame the buffer is allocated, never touched, and freed. It is now allocated where it is used. The comment two lines below already warned against sizing a buffer from the uncompressed size; the line above it had not been given the same treatment. The first call into cuFFT pays the library's one-time initialisation, and it landed in the middle of the first pass with nothing to overlap it. It is now forced on a background thread at startup, alongside the file open and the mapping build, in the manner the shadow finder already uses. Finally, the detector mask was copied into the start message whether or not a file would carry it, which a merging run does not. It is filled where a writer is constructed - both places one is constructed, the second being the fallback that writes a process file when nothing indexed. Faster on eleven of thirty-eight crystals and slower on none; the whole rotation test set falls from four minutes thirty to four minutes seventeen, with each binary repeating itself to within half a per cent. Space groups thirty-five of thirty-eight and no failures throughout, and every column of the comparison table is identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EGpGdgmJ8MyY9pCGWjktyi
80 lines
3.6 KiB
C++
80 lines
3.6 KiB
C++
// SPDX-FileCopyrightText: 2025 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <cstddef>
|
|
#include <cstdint>
|
|
#include <memory>
|
|
#include <new>
|
|
#include <string>
|
|
#include <utility>
|
|
#include <vector>
|
|
|
|
#include "../compression/CompressionAlgorithmEnum.h"
|
|
#include "ColorScale.h"
|
|
|
|
// std::vector value-initialises whatever it resizes into. A buffer that the very next read fills in
|
|
// full has no use for that, and on a compressed image chunk it is megabytes of memset per image. This
|
|
// allocator default-initialises instead, which for bytes is no initialisation at all. The rebind is
|
|
// not optional: without it the container rebinds to std::allocator and the zeroing comes back.
|
|
template <class T>
|
|
struct NoInitAllocator : std::allocator<T> {
|
|
NoInitAllocator() = default;
|
|
template <class U> NoInitAllocator(const NoInitAllocator<U> &) {}
|
|
template <class U> struct rebind { using other = NoInitAllocator<U>; };
|
|
template <class U> void construct(U *p) { ::new (static_cast<void *>(p)) U; }
|
|
template <class U, class... Args> void construct(U *p, Args &&...args) {
|
|
::new (static_cast<void *>(p)) U(std::forward<Args>(args)...);
|
|
}
|
|
};
|
|
|
|
// Bytes of an image as they are stored: sized by the reader and then overwritten by the read.
|
|
using RawByteBuffer = std::vector<uint8_t, NoInitAllocator<uint8_t>>;
|
|
|
|
enum class CompressedImageMode {Int8, Int16, Int32, Uint8, Uint16, Uint32, RGB, Float16, Float32, Float64};
|
|
|
|
CompressedImageMode CalcImageMode(size_t byte_depth, bool is_float, bool is_signed);
|
|
|
|
class CompressedImage {
|
|
size_t xpixel;
|
|
size_t ypixel;
|
|
CompressedImageMode mode;
|
|
CompressionAlgorithm algorithm;
|
|
const uint8_t *data;
|
|
size_t size; // After compression
|
|
std::string channel;
|
|
public:
|
|
CompressedImage();
|
|
CompressedImage(const void* data, size_t size,
|
|
size_t xpixel, size_t ypixel,
|
|
CompressedImageMode mode,
|
|
CompressionAlgorithm algorithm = CompressionAlgorithm::NO_COMPRESSION,
|
|
std::string channel = "default");
|
|
|
|
CompressedImage(const std::vector<rgb>& input, size_t xpixel, size_t ypixel);
|
|
CompressedImage(const std::vector<float>& input, size_t xpixel, size_t ypixel);
|
|
CompressedImage(const std::vector<uint8_t>& input, size_t xpixel, size_t ypixel,
|
|
CompressedImageMode mode = CompressedImageMode::Uint8,
|
|
CompressionAlgorithm algorithm = CompressionAlgorithm::NO_COMPRESSION);
|
|
CompressedImage(const std::vector<uint16_t>& input, size_t xpixel, size_t ypixel);
|
|
CompressedImage(const std::vector<uint32_t>& input, size_t xpixel, size_t ypixel);
|
|
CompressedImage(const std::vector<int16_t>& input, size_t xpixel, size_t ypixel);
|
|
CompressedImage(const std::vector<int32_t>& input, size_t xpixel, size_t ypixel);
|
|
CompressedImage &Channel(std::string channel);
|
|
|
|
[[nodiscard]] size_t GetUncompressedSize() const;
|
|
[[nodiscard]] bool IsSigned() const;
|
|
void GetUncompressed(std::vector<uint8_t> &buffer) const;
|
|
[[nodiscard]] const uint8_t* GetUncompressedPtr(std::vector<uint8_t> &buffer) const;
|
|
[[nodiscard]] size_t GetByteDepth() const;
|
|
[[nodiscard]] size_t GetNumChannels() const;
|
|
[[nodiscard]] std::string GetChannel() const;
|
|
[[nodiscard]] size_t GetWidth() const;
|
|
[[nodiscard]] size_t GetHeight() const;
|
|
[[nodiscard]] CompressedImageMode GetMode() const;
|
|
[[nodiscard]] CompressionAlgorithm GetCompressionAlgorithm() const;
|
|
[[nodiscard]] const uint8_t *GetCompressed() const;
|
|
[[nodiscard]] size_t GetCompressedSize() const;
|
|
};
|