From 2ab8c55dfab0d6af65eeb04127366ce58130df16 Mon Sep 17 00:00:00 2001 From: Filip Leonarski Date: Wed, 16 Sep 2026 22:35:05 +0200 Subject: [PATCH] rugnux and the viewer read SMV and gzipped miniCBF Two more of the formats deposited data actually arrives in, found by processing a corpus of it: every PETRA III EMBL set is .cbf.gz, and NSRRC and the whole ADSC Quantum era are SMV. Both were previously "no native input". SMV is an ASCII "KEY=value;" block between braces, then the pixels - no container, no compression, nothing to decode by offset - so reader/SMV.{h,cpp} and JFJochSMVReader are a smaller job than the marCCD pair they sit beside, and need no new dependency at all. Two things the format does not give us, both said out loud rather than papered over: * It states no saturation value, so overloads are judged on the 16-bit container alone. That can only fail to call a pixel saturated, never condemn a good one, but a CCD at the top of its range does saturate, so the reader warns once. * Its beam centre is in MILLIMETRES and which of X/Y is the fast direction is a convention rather than a rule. Measured on one ALS ADSC sweep the file's value is TRANSPOSED: as stated it indexes 2/60 frames, and the run's own beam-centre measurement (which adopts the right one automatically) indexes 60/60. Swapping it here would fit that writer and might break another, so the header is read as the format defines it and the measurement stays the arbiter. Revisit with a second vendor's SMV in hand. .cbf.gz needed only Slurp() in MiniCBF.cpp, through which every read already passes: it sniffs the two-byte gzip magic - not the file name - and takes a zlib path when it is there, leaving the plain path free of zlib's buffer copy. zlib-ng is already in the build, so this is a link line, not a dependency. The sweep template grew a suffix, because ".cbf" and ".cbf.gz" are separate sweeps and std::filesystem cannot split the double extension on its own. The viewer's single cbf_reader becomes three, dispatched by CanRead() in the same order as rugnux. Dispatch is by CONTENT in both: ".img" is used by miniCBF, marCCD AND SMV depending on the writer, and a PDB detector label has now been wrong about the format four times, so an extension decides nothing. Measured, de novo, no flags: 9fcg (1800 gzipped frames) gives P4 and a cell 0.06% from the deposited one at 1.37 A against a deposited 1.54; 6oel (ADSC SMV) gives F4132 - 96 operations, the most a protein space group can have - and a cell 0.05% out, 100% indexed. Tests cover both formats and the transposed-beam-centre case with fixtures written byte for byte, so they need no external data. Co-Authored-By: Claude Opus 5 (1M context) --- docs/CHANGELOG.md | 3 + docs/RUGNUX_FORMATS.md | 16 +- reader/CMakeLists.txt | 6 +- reader/JFJochCBFReader.cpp | 39 +++- reader/JFJochSMVReader.cpp | 191 +++++++++++++++++++ reader/JFJochSMVReader.h | 57 ++++++ reader/MiniCBF.cpp | 49 +++++ reader/SMV.cpp | 283 ++++++++++++++++++++++++++++ reader/SMV.h | 70 +++++++ rugnux/rugnux_cli.cpp | 10 +- tests/JFJochReaderTest.cpp | 116 ++++++++++++ viewer/JFJochImageReadingWorker.cpp | 22 ++- viewer/JFJochImageReadingWorker.h | 13 +- viewer/JFJochProcessController.cpp | 17 +- viewer/JFJochViewerMenu.cpp | 2 +- 15 files changed, 867 insertions(+), 27 deletions(-) create mode 100644 reader/JFJochSMVReader.cpp create mode 100644 reader/JFJochSMVReader.h create mode 100644 reader/SMV.cpp create mode 100644 reader/SMV.h diff --git a/docs/CHANGELOG.md b/docs/CHANGELOG.md index b1e31d853..8d14c61ab 100644 --- a/docs/CHANGELOG.md +++ b/docs/CHANGELOG.md @@ -4,6 +4,9 @@ ### 1.0.0-rc.171 * rugnux reads marCCD images natively: naming any frame of a Rayonix or mar Mosaic sweep processes the whole sweep, with no conversion step. +* rugnux reads SMV images natively - the format ADSC Quantum detectors wrote and Rayonix still writes. +* rugnux reads gzipped miniCBF (`.cbf.gz`) directly, so a sweep archived that way needs no decompression pass. +* `jfjoch_viewer` opens marCCD, SMV and gzipped miniCBF frames alongside HDF5 and plain miniCBF. ### 1.0.0-rc.170 diff --git a/docs/RUGNUX_FORMATS.md b/docs/RUGNUX_FORMATS.md index 6102d3e8d..cd74c274d 100644 --- a/docs/RUGNUX_FORMATS.md +++ b/docs/RUGNUX_FORMATS.md @@ -1,10 +1,10 @@ # What rugnux reads Which data `rugnux` opens, before anything is typed. The short answer: an HDF5 master (NXmx or -DECTRIS, from any facility), a PILATUS miniCBF sweep or a marCCD sweep — nothing else is read, so -any other format has to be converted to one of these three first. +DECTRIS, from any facility), a PILATUS miniCBF sweep, a marCCD sweep or an SMV sweep — nothing else +is read, so any other format has to be converted to one of these first. -**Input** is an HDF5 master file, or a directory of PILATUS miniCBF or marCCD frames. One input is +**Input** is an HDF5 master file, or a directory of PILATUS miniCBF, marCCD or SMV frames. One input is **one sweep of one crystal** — rugnux does not combine sweeps or crystals in a run; process each sweep to its own `_unmerged.mtz` and merge them downstream (see [Taking the data onward](RUGNUX_INTEGRATION.md#taking-the-data-onward)). @@ -26,6 +26,9 @@ sweep to its own `_unmerged.mtz` and merge them downstream header, including the imgCIF axis table where the header carries one (see [Detector geometry](DETECTOR_GEOMETRY.md)). A raw CBF carries no analysis results, so `--mode scale` — which re-scales the reflections stored in a `_process.h5` — does not accept one. + **Gzipped frames (`.cbf.gz`) are read directly**, with no decompression step: EMBL Hamburg's + beamlines write that form by default, and expanding a sweep first would double the disk it needs. + The two forms are separate sweeps, so a directory holding both is not spliced into one crystal. * **marCCD sweep** — what Rayonix MX-series and mar Mosaic detectors write, and what a decade of deposited CCD data is archived as: an uncompressed TIFF with the instrument header in the gap before the pixels. One frame per file, read natively. Naming a frame or its directory selects the @@ -36,6 +39,13 @@ sweep to its own `_unmerged.mtz` and merge them downstream from the data as it does for a miniCBF that carries no axis table. A CCD frame marks no untrusted pixels, so the sweep starts with nothing masked. Like a raw CBF, it carries no analysis results and `--mode scale` does not accept one. +* **SMV sweep** — what ADSC Quantum detectors wrote and what Rayonix and others still write, so most + archived CCD data from the 2000s is in this form: an ASCII `KEY=value;` block between braces, then + the pixels. One frame per file, read natively, the sweep selected exactly as above. The beam centre + it states is in **millimetres** and is converted here; its convention varies between writers, and a + file that states it transposed indexes nothing until the run's own beam-centre measurement replaces + it, which it does automatically. SMV states no saturation value at all, so overloads are judged on + the 16-bit container alone and the reader says so once per sweep. Spots are always found by `rugnux` itself, including for the two-pass rotation first pass — the spot lists a dataset may already carry were found online, at the acquisition's threshold and with diff --git a/reader/CMakeLists.txt b/reader/CMakeLists.txt index 3f336aae5..23029e7a7 100644 --- a/reader/CMakeLists.txt +++ b/reader/CMakeLists.txt @@ -10,6 +10,10 @@ ADD_LIBRARY(JFJochReader STATIC MarCCD.h JFJochMarCCDReader.cpp JFJochMarCCDReader.h + SMV.cpp + SMV.h + JFJochSMVReader.cpp + JFJochSMVReader.h HDF5ImageLocator.cpp HDF5ImageLocator.h HDF5ImageSource.cpp @@ -29,5 +33,5 @@ ADD_LIBRARY(JFJochReader STATIC # tiff: the marCCD reader's files are ordinary uncompressed TIFFs, so libtiff - which the build # already fetches for JFJochPreview in every mode - reads the pixels and there is no new dependency. TARGET_LINK_LIBRARIES(JFJochReader JFJochImageAnalysis JFJochAPI JFJochCommon JFJochZMQ JFJochLogger - JFJochHDF5Wrappers CBORStream2FrameSerialize tiff + JFJochHDF5Wrappers CBORStream2FrameSerialize tiff ZLIB::ZLIB ${CMAKE_DL_LIBS}) diff --git a/reader/JFJochCBFReader.cpp b/reader/JFJochCBFReader.cpp index 2556df484..f8dbddf88 100644 --- a/reader/JFJochCBFReader.cpp +++ b/reader/JFJochCBFReader.cpp @@ -12,6 +12,7 @@ #include #include #include +#include #include "../common/JFJochException.h" #include "../common/Logger.h" @@ -19,10 +20,23 @@ namespace { +// ".cbf", or ".cbf.gz" - EMBL Hamburg's beamlines write the gzipped form by default, and the +// reader decompresses it in place, so a sweep of those is named here rather than converted first. +// Returns the suffix that matched, because the sweep template needs its length. +std::optional CBFSuffix(const std::filesystem::path &p) { + std::string name = p.filename().string(); + std::transform(name.begin(), name.end(), name.begin(), + [](unsigned char c) { return std::tolower(c); }); + for (const char *suffix : {".cbf.gz", ".cbf"}) { + const std::string s(suffix); + if (name.size() > s.size() && name.compare(name.size() - s.size(), s.size(), s) == 0) + return s; + } + return {}; +} + bool HasCBFExtension(const std::filesystem::path &p) { - std::string ext = p.extension().string(); - std::transform(ext.begin(), ext.end(), ext.begin(), [](unsigned char c) { return std::tolower(c); }); - return ext == ".cbf"; + return CBFSuffix(p).has_value(); } // The sweep a file belongs to, as a template: everything before the trailing run of digits, the @@ -32,12 +46,15 @@ bool HasCBFExtension(const std::filesystem::path &p) { struct Template { std::string prefix; size_t digits = 0; + std::string suffix = ".cbf"; // ".cbf" or ".cbf.gz"; the two are not one sweep bool Matches(const std::string &name) const { - if (name.size() != prefix.size() + digits + 4) // + ".cbf" + if (name.size() != prefix.size() + digits + suffix.size()) return false; if (name.compare(0, prefix.size(), prefix) != 0) return false; + if (name.compare(prefix.size() + digits, suffix.size(), suffix) != 0) + return false; for (size_t i = 0; i < digits; i++) if (!std::isdigit(static_cast(name[prefix.size() + i]))) return false; @@ -47,15 +64,19 @@ struct Template { std::optional