5 Commits
Author SHA1 Message Date
leonarski_fandClaude Opus 5 7c3ab72fed reader: refuse a CBF that is not a detector image instead of misreading it
Three ways a miniCBF opened silently wrong.

A byte-offset CBF with no PILATUS header at all was accepted: Count_cutoff then
defaulted to 0, SaturationLimitFromValue(0) is 1, and every pixel at or above one
count was flagged saturated - the integration accept gate drops the whole
reflection, so the run comes out empty for a reason nothing reports. The pixel
size defaulted to 0 with no validation anywhere downstream, which collapses every
resolution, every scattering vector and the beam centre in millimetres. XDS
writes its correction files in exactly this shape, so this is not hypothetical.
A pixel size is now required to claim the file at all, and a missing Count_cutoff
leaves the saturation limit unset - falling back to the container's own overflow,
which can only fail to call a pixel saturated - with a warning saying so.

And the header captures are character classes, not number grammars: "[\d.eE+-]+"
matches a bare "." and "(\d+)" matches a digit string too long for int64.
std::stod and std::stoll answer both with a raw std:: exception, which escaped
the format probe - CanRead catches JFJochException only - so merely LOOKING at a
corrupt file threw out of the viewer's open path. A header field that does not
parse is now reported as a malformed header.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EFEJG6WBQv8th4UJFNe53N
2026-08-31 07:17:40 +02:00
leonarski_fandClaude Opus 5 a27c4cf26f reader: take a miniCBF's mounting from the imgCIF axis table its header states
A miniCBF header states three things about how the instrument is put together
that the reader was assuming instead: which laboratory direction the image's
columns run along, which its rows run along, and which the spindle turns about.
Some beamlines append a CBF template block holding the full imgCIF axis table,
which says all three outright.

Two instruments in the corpus are not what was assumed, in two different ways.
One mounts its detector a quarter turn round, so the image's columns run
vertically. Another turns its spindle about the VERTICAL, with the image mounted
the usual way; its table says so, and its "# Oscillation_axis" line says so a
second way, by naming the image direction the spindle runs along rather than a
vector. Either error leaves the spindle 90 degrees from the image. That is not a
sign, so the run's axis-sign rescue cannot reach it, and no refinement recovers
it: all three affected sweeps indexed nothing usable.

So the table is read. The element axes give the image orientation, matched against
the eight discrete mountings exactly as the NXmx module directions already are -
the match itself moves to DetectorOrientation, so both readers share one
definition rather than two copies. The goniometer axis with no parent gives the
spindle DIRECTION; its sign stays the rescue's business, which is the part a
convention can legitimately differ on. The detector axis with no parent gives the
2theta arm, replacing the assumption that the arm shares the spindle's axis - the
one header stating both states them with the same vector, so this changes no
answer, only what it rests on. imgCIF's frame differs from the internal one by a
half turn about x, a rotation and not a mirror, as writer/HDF5NXmx.cpp already
records from the other side.

Where a header carries no table, a "+SLOW" on the Oscillation_axis line still
says the spindle runs along the image's slow direction. That is the only thing one
of the three affected sets says about it. The axis NAME on that line stays
unusable - the header that carries both says "X.CW" where its own table says Y -
but the direction token is not: where both are present they agree, which is what
makes reading it evidence rather than a guess.

Also: naming a frame with no directory at all now finds its sweep. parent_path()
of a bare filename is empty and iterating an empty path finds nothing, so running
from inside the data directory reported that no images were found.

Measured, with nothing on the command line. The vertical-spindle protein set goes
from no usable lattice to 100% indexed, P 6(3) 2 2 with a cell 0.43% from
deposited, 87846 reflections at 86.3% completeness and CC(1/2) 0.995. Its
companion from the same detector, which has no table and only the +SLOW token,
goes from a spurious monoclinic cell at 2.3% completeness and I/sigma 0.21 to the
right orthorhombic lattice, 97.7% indexed, 59.7% complete, CC(1/2) 0.996. The
quarter-turned set's three sweeps, at three arm positions, now all index without
the hand-passed quarter turn they needed and agree on one cell to 0.03 A. Six
miniCBF sets that state no table and no +SLOW - including one whose
Oscillation_axis line names an axis in a third dialect - are byte-identical in
.hkl, .mtz, .cif and the image statistics, as are two NXmx sets, which is the
shared orientation matcher moving nothing on that path either.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T3yNBXk4wKdMZy1ak2NY7f
2026-08-30 14:03:32 +02:00
leonarski_fandClaude Opus 5 5c44e544dc reader: place a detector swung out on a 2theta arm where the file says it stands
Chemical crystallography reaches high angle by swinging the detector out on a
2theta arm. Both readers had the number and neither used it: the miniCBF header's
Detector_2theta was parsed into a struct member nothing ever read, and on the NXmx
side the rotation was in the depends_on chain, which was not followed at all. A
sweep taken at 30 degrees was therefore processed with its detector plane 30
degrees from where it stood, and nothing indexed.

The geometry could already express it, and needed no change: the arm turns the
detector about the sample, so the distance is still measured along the detector
normal and the beam centre is still the point of normal incidence - which is
exactly the PONI convention, and a swung detector is one PONI rotation. What moves
is the direct beam, by distance*tan(2theta), off the beam centre and often off the
detector.

NXmx is the harder half, because the swing has no field of its own: it is one
rotation in the chain the detector's position depends on, and "two_theta" is only
one beamline's name for that dataset. So the chain is followed and its rotations
composed, rather than a field of one name being looked for - each transformation
states its vector in the frame of the one it depends on, which is why the product
is the whole placement. Translations are skipped; they are the distance and the
beam centre, which the file states separately in the square-on frame. Vectors come
from McStas through the same 180-degree turn about z the module directions already
use, a proper rotation, so an axis carried through it turns the same way.

The three rotations a file this system writes ARE that chain, and are also read as
the PONI angles - so those three paths are skipped, or every tilted file we have
ever written would come back tilted twice. That is the one way this change could
have broken existing data, and the test for it writes a tilted file and reads it
back.

For miniCBF the arm turns about the base spindle axis: on the four-circle geometry
those headers describe the two are one axis, and the imgCIF axis table such a
header carries states them with the same vector. Both now come from one constant,
so a later correction to the frame moves them together.

Measured. On a swung NXmx sweep the chain gives rot2 = -0.34907 rad for the 20
degrees it states, and the sweep goes from "nothing was integrated" to 25000
reflections at 82.2% completeness and CC(1/2) 0.9993, in the same space group and
the same cell to 0.03 A as the square-on sweep of that crystal; the opposite sign
indexes nothing. A miniCBF sweep at 30 degrees goes the same way, to 0.585 A, and
a second sweep of that crystal at 55 degrees reaches 0.476 A and reproduces the
cell again - with a low-resolution limit of 2.36 A rather than 13 A, which is what
a detector swung that far records. On all of them post-refinement recovers the
header's own beam centre and distance, and the beam stop shadow sits within four
pixels of where the swung geometry puts the direct beam, 417 and 537 pixels from
where the unswung one does. Seven sets whose detector is square to the beam, three
of them carrying a chain whose 2theta is zero, are byte-identical in .hkl, .mtz,
.cif and the image statistics.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T3yNBXk4wKdMZy1ak2NY7f
2026-08-30 12:36:46 +02:00
leonarski_fandClaude Opus 5 367eba55b1 reader: take a miniCBF rotation axis from the goniometer the header states
Every miniCBF sweep was handed the same hardcoded axis regardless of what its header said, and
Chi/Kappa/Phi/Omega were not parsed at all - there were no members for them. A sweep collected on a
tilted chi cradle therefore ran with an axis that is 54.7 degrees wrong.

The header angles and their increments are now read, the scanned axis is identified from the
non-zero increment (the name is consulted only when no increment is stated, which is what absorbs the
five different spellings the corpus contains, including one file that states no axis name at all),
and a phi scan composes the head chain. An omega scan returns the base axis untouched, because a
fixed chi cannot tilt the axis it hangs from.

-9999 is a sentinel meaning "not set", not an angle. It is treated as absent, so it can never reach
the geometry.

The direction and sense are not invented: these files append an imgCIF _axis loop stating their own
vectors, and SOURCE with GRAVITY fix the imgCIF-to-internal transform, which independently reproduces
the transform this repository already documents for NXmx. Under it the file's own stated phi axis is
exactly the composed one, to four decimals.

Driving the real reader over all 39 corpus sweeps, 37 return the previous axis bit-identically -
including every sweep carrying a large fixed chi, every sentinel header and every axis-name spelling.
Only the two genuine phi scans move, and an unrelated rotation dataset is unchanged end to end.

This is necessary but not sufficient for the one dataset that motivates it: with the axis corrected
it still does not index, because that detector is also mounted rotated 90 degrees in its own plane,
which the reader does not yet read. Compensating both takes its phi sweep from no indexed validation
frames to 90.89% indexed and a complete merge, which is what shows this half is load-bearing. The
detector mount and the two-theta swing belong to the detector-frame work.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Lc5JG6kJqZoCWaoZ43JGTW
2026-08-29 19:58:50 +02:00
leonarski_fandClaude Opus 5 dc16a00271 reader: read a PILATUS miniCBF sweep natively, without libcbf
Most facilities still archive rotation data as a directory of miniCBF frames, which until now had
to be converted to HDF5 before rugnux could see it. Nothing in that format needs a CIF parser or a
library: it is an ASCII header, four separator bytes, then one byte-offset compressed image, and
every value the reader wants sits on a "# " comment line or a MIME line.

MiniCBF holds the format itself - header parse and the byte-offset decoder, which is a running
value with deltas stored smallest-container-first. Verified byte-exact against dxtbx on PILATUS 6M,
6M-F, 300K, silicon and CdTe sensors, and three sensor thicknesses.

JFJochCBFReader is a sibling of JFJochHDF5Reader under the JFJochReader base. NAMING ANY FRAME
READS ITS WHOLE SWEEP: the sweep is identified by the template (prefix + digit count) the named
frame belongs to, not by "every .cbf in the directory", so a directory holding two sweeps does not
splice two crystals together. Naming a directory takes the sweep with the most frames in it.

Images decode on demand, one per call, so any number of workers can read at once - there is no
global lock as there is on the HDF5 path, HDF5 not being thread-safe. A raw CBF carries no analysis
results, so the dataset it builds is the geometry, the mask and nothing else, exactly as a plain
DECTRIS file with no /entry/MX gives.

Two header quirks are handled because real files have them: the sensor material is written
"Silicon" where the rest of the code compares against "CdTe", and the thickness unit is sometimes
omitted. Headers are not a fixed size either - one set carries 6335 bytes - so the parse runs to
the binary separator rather than over a fixed prefix.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 20:12:36 +02:00