Most facilities still archive rotation data as a directory of miniCBF frames, which until now had
to be converted to HDF5 before rugnux could see it. Nothing in that format needs a CIF parser or a
library: it is an ASCII header, four separator bytes, then one byte-offset compressed image, and
every value the reader wants sits on a "# " comment line or a MIME line.
MiniCBF holds the format itself - header parse and the byte-offset decoder, which is a running
value with deltas stored smallest-container-first. Verified byte-exact against dxtbx on PILATUS 6M,
6M-F, 300K, silicon and CdTe sensors, and three sensor thicknesses.
JFJochCBFReader is a sibling of JFJochHDF5Reader under the JFJochReader base. NAMING ANY FRAME
READS ITS WHOLE SWEEP: the sweep is identified by the template (prefix + digit count) the named
frame belongs to, not by "every .cbf in the directory", so a directory holding two sweeps does not
splice two crystals together. Naming a directory takes the sweep with the most frames in it.
Images decode on demand, one per call, so any number of workers can read at once - there is no
global lock as there is on the HDF5 path, HDF5 not being thread-safe. A raw CBF carries no analysis
results, so the dataset it builds is the geometry, the mask and nothing else, exactly as a plain
DECTRIS file with no /entry/MX gives.
Two header quirks are handled because real files have them: the sensor material is written
"Silicon" where the rest of the code compares against "CdTe", and the thickness unit is sometimes
omitted. Headers are not a fixed size either - one set carries 6335 bytes - so the parse runs to
the binary separator rather than over a fixed prefix.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>