v1.0.0.rc-161 #71

Merged
leonarski_f merged 297 commits from adaptive-spot-finding into main 2026-08-13 17:03:10 +02:00
Owner

This is an UNSTABLE release. It includes many experimental features, as well as many AI generated fixes. We recommend using rc.152 for production use.

  • rugnux: significantly better quality of results, and faster. A large rework of integration, scaling, merging, geometry refinement and space-group determination, together with measurements the program previously made no attempt at - the direct beam before indexing, the beam stop, the goniometer rotation scale, and the stretches of a sweep the crystal did not deliver. A rotation dataset typically gains observations at better <I/sigma> and R_meas, and every mx and scale run writes a <prefix>_report.txt results report modelled on XDS's CORRECT.LP. Many defaults moved with it: spot detection is self-calibrating, beam-stop detection and rotation geometry post-refinement are on, resolution limits default to as far as the detector reaches, and ice-ring handling engages only where the crystal is measured to have ice.
  • jfjoch_viewer: the beam-stop shadow, the detector calibration and the beam-centre measurement are reachable from "Analyze dataset"; the settings panel reports how the sample moved and how polarized the beam was; image rendering and interaction are faster.
  • Performance: bitshuffle+LZ4 images are decoded on the GPU rather than on the host, with the bitshuffle inverse fused into preprocessing so the decompressed frame is never held in device memory.
  • Broker, writer, packaging and build: image-slot lifetime and locking fixes, per-image datasets sized by the images actually written, the Debian/Ubuntu broker package renamed to jfjoch, and image_analysis compiling under MSVC again.

Breaking change to the rugnux command line:

  • --azint-only and --scale are removed, replaced by --mode azint and --mode scale; the full pipeline is --mode mx and remains the default. A script passing the old flags now fails with the list of valid modes rather than silently running the wrong one.
  • -t/--stride is refused on rotation data: skipping frames cuts every reflection's rocking curve, so the combined fulls and their partiality would be measured over frames the sweep never recorded. Select a contiguous range with -s/-e instead. --mode azint and --force-still still take a stride.

Breaking changes to OpenAPI - regenerate the client (jfjoch-client 1.0.0-rc.161, frontend/src/client) or read the affected fields as optional:

  • image_scale_b is removed from the plot_type enum, so a client requesting that plot now gets an error rather than a curve.
  • azim_int_settings.high_q_recipA, spot_finding_settings.high_resolution_limit and spot_finding_settings.low_resolution_limit are no longer required. All three mean "no limit at that end" when unset and are omitted from the response instead of carrying a placeholder value, which raises in a client generated from an rc.160-or-earlier spec. A value of 0 is still accepted and means the same thing.

Breaking changes to the stored formats - a consumer reading these fields must treat them as optional:

  • The per-image image-scale B factor is no longer computed, so /entry/MX/imageScaleBFactor is absent from newly written HDF5 files and the corresponding key is absent from the CBOR DataMessage and END blocks. Files written by rc.160 and earlier still contain it and still open; nothing in the pipeline reads it any more.
  • _reflns.jfjoch_diffrn_ISa now carries the whole-range 1/sqrt(a*b) that XDS's ISa denotes, and the error-model a and b are reported in XDS's convention; the strong-reflection asymptote moves to _reflns.jfjoch_diffrn_ISa_asymptotic. A file written by an earlier version carries the asymptote under the plain ISa name.
This is an UNSTABLE release. It includes many experimental features, as well as many AI generated fixes. We recommend using rc.152 for production use. * **rugnux: significantly better quality of results, and faster.** A large rework of integration, scaling, merging, geometry refinement and space-group determination, together with measurements the program previously made no attempt at - the direct beam before indexing, the beam stop, the goniometer rotation scale, and the stretches of a sweep the crystal did not deliver. A rotation dataset typically gains observations at better <I/sigma> and R_meas, and every `mx` and `scale` run writes a `<prefix>_report.txt` results report modelled on XDS's `CORRECT.LP`. Many defaults moved with it: spot detection is self-calibrating, beam-stop detection and rotation geometry post-refinement are on, resolution limits default to as far as the detector reaches, and ice-ring handling engages only where the crystal is measured to have ice. * **jfjoch_viewer:** the beam-stop shadow, the detector calibration and the beam-centre measurement are reachable from "Analyze dataset"; the settings panel reports how the sample moved and how polarized the beam was; image rendering and interaction are faster. * **Performance:** bitshuffle+LZ4 images are decoded on the GPU rather than on the host, with the bitshuffle inverse fused into preprocessing so the decompressed frame is never held in device memory. * **Broker, writer, packaging and build:** image-slot lifetime and locking fixes, per-image datasets sized by the images actually written, the Debian/Ubuntu broker package renamed to `jfjoch`, and `image_analysis` compiling under MSVC again. **Breaking change to the rugnux command line:** * `--azint-only` and `--scale` are **removed**, replaced by `--mode azint` and `--mode scale`; the full pipeline is `--mode mx` and remains the default. A script passing the old flags now fails with the list of valid modes rather than silently running the wrong one. * `-t`/`--stride` is **refused on rotation data**: skipping frames cuts every reflection's rocking curve, so the combined fulls and their partiality would be measured over frames the sweep never recorded. Select a contiguous range with `-s`/`-e` instead. `--mode azint` and `--force-still` still take a stride. **Breaking changes to OpenAPI** - regenerate the client (`jfjoch-client` 1.0.0-rc.161, `frontend/src/client`) or read the affected fields as optional: * `image_scale_b` is removed from the `plot_type` enum, so a client requesting that plot now gets an error rather than a curve. * `azim_int_settings.high_q_recipA`, `spot_finding_settings.high_resolution_limit` and `spot_finding_settings.low_resolution_limit` are no longer `required`. All three mean "no limit at that end" when unset and are omitted from the response instead of carrying a placeholder value, which raises in a client generated from an rc.160-or-earlier spec. A value of 0 is still accepted and means the same thing. **Breaking changes to the stored formats** - a consumer reading these fields must treat them as optional: * The per-image image-scale B factor is no longer computed, so `/entry/MX/imageScaleBFactor` is absent from newly written HDF5 files and the corresponding key is absent from the CBOR DataMessage and END blocks. Files written by rc.160 and earlier still contain it and still open; nothing in the pipeline reads it any more. * `_reflns.jfjoch_diffrn_ISa` now carries the whole-range `1/sqrt(a*b)` that XDS's ISa denotes, and the error-model `a` and `b` are reported in XDS's convention; the strong-reflection asymptote moves to `_reflns.jfjoch_diffrn_ISa_asymptotic`. **A file written by an earlier version carries the asymptote under the plain `ISa` name.**
leonarski_f added 297 commits 2026-08-13 17:02:45 +02:00
GenerateSpotPlot iterated msg.spots, but SpotAnalyze called it before
assigning output.spots. In the online path the DataMessage is fresh per
frame, so the plot was built from an empty list and spot_plot_intensity
/ spot_plot_count came out all zeros. Pass the finished spots vector
explicitly instead of relying on the field being set: the live path
passes the full pre-truncation list, the HDF5 read-back path passes
message.spots.

GetResolution scaled the 5th-percentile index by spots.size() (which
includes ice-ring spots) while indexing the ice-filtered resolutions
vector, biasing the estimate and reading out of bounds on ice-heavy
frames. Index by resolutions.size() instead.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
On very-low-multiplicity data (e.g. EP_cs_01-24, mult ~1.4) the merge has too
few symmetry equivalents to measure the asymptotic I/sigma: both the (a, b)
error-model fit and the per-group strong-reflection scatter collapse toward
zero, so 1/error_model_b_asymptotic either explodes to an impossibly high ISa
(tiny positive b) or is left as 0. Real macromolecular data does not exceed
ISa ~50, so clamp the reported asymptote at a generous cap (ISa 100) and treat
anything past it as unmeasured (result.isa undetermined) rather than emitting a
spurious extreme. No-op for all well-measured data (b_asy well above the cap).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The serial-stills merge (MergeOnTheFly::CorrectedSigma) weighted each
observation by 1/sigma^2 using the observation's OWN sigma. Below ~1
photon the Poisson signal part of that sigma correlates with the
observation's up/down fluctuation, so the inverse-variance mean is
biased low: an up-fluctuated observation acquires a larger sigma and is
over-downweighted. The rotation combine (RotationScaleMerge::
process_rawrun) already avoids this by rebuilding the signal variance at
the pooled estimate; the stills path did not.

Decompose each observation's variance into a background/read part (kept
per-observation) and a Poisson signal part, and rebuild the signal part
at the reflection's expected <I>. Bit-identical when an observation sits
at its reflection mean; only weak-shell weights move. Now default on, so
the stills path matches the rotation path;
--no-expected-variance-merge restores the old observed-sigma weighting.

Validated by paired refinement (phenix, 5 free-set seeds, byte-identical
free flags across arms): R-free-neutral on strong lysozyme and lower
R-free on weak serial-stills data checked against an independent
deposited model (6/6 seeds). The CC1/2 dip on strong data reflects
precision, not accuracy. Applies to both offline rugnux and the online
broker stills merge.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Replace the frozen scalar-sigma stills partiality with a physical, refined model.
Per crystal, refine an orientation tilt (dpsi_x, dpsi_y) against the running merge
and recompute each reflection's partiality analytically from the refined geometry
(angular Ewald-proximity model, sigma(d*) = gamma_e*d*), with the per-crystal scale
G profiled out by the existing robust IRLS - no re-integration. A soft Gaussian
prior on dpsi tames weak-data overfit while staying inert on strong data. The
merge <-> refine loop iterates a few times.

This is now the stills default via ScalingSettings::stills_partiality_refine (on).
A single opt-out flag `--simple-stills` reverts to treating every reflection as a
full (p=1, single pass). Retires the experimental `--still-partiality` flag. The
viewer gains a "Partiality post-refinement (stills)" checkbox in Scaling settings.

Validated (integrate-once / --scale): CC1/2 and R_meas both improve on three
monochromatic serial-stills datasets (+2.8 / -10, +5.6 / -3.4, +2.1 / -4);
neutral on a pink-beam DMM set (already-full reflections); R-free/R-work down vs
a fixed model; competitive with CrystFEL partialator on matched frames.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The offline CPU spot finder marks a pixel strong when it clears a fixed photon
count AND a local-window SNR. The fixed photon floor forces per-dataset tuning:
its sweet spot tracks the background level (weak sets want a low threshold,
strong or high-background sets a high one) and the usable window is narrow, so
users hand-tune --spot-threshold/--spot-sigma per dataset.

Add an opt-in --adaptive-spots mode (AdaptiveSpotFinderCPU) that replaces the
fixed floor with a per-resolution-ring threshold derived from each image's own
noise. Per ring it computes a peak-excluded background mean and sigma (one plain
pass + two sigma-clip passes over the assembled photon image, binned by the
azimuthal-integration ring index) and sets

    thr = max( PoissonTail(mean, p), mean + z * sqrt(sigma^2 + read^2) )

with p = false_pixels_per_frame / n_pixels the single portable knob (default
100) and z = Phi^-1(1 - p). The Poisson arm is the correct significance where
the background is countable (it carries the sqrt(mean) shot noise, so a bright
low-resolution ring gets a high threshold); the read-noise-floored Gaussian arm
keeps the threshold physical where the background vanishes (empty high-resolution
rings), without which those rings flood. read is a detector-level constant, not
a per-dataset knob. Both arms are needed: Poisson alone floods near-zero
background, Gaussian alone drops the shot-noise term and under-thresholds bright
rings.

One --adaptive-spots setting then adapts across a wide range of serial datasets
with no per-dataset threshold, matching or beating hand-tuned thresholds and the
peakfinder8/xgandalf reference on both weak large-cell and strong serial data,
with equal merged R-free.

The finder runs on the CPU (offline/viewer path) and reads the host image, which
the GPU pipeline already keeps in sync, so it works in either build. The default
(non-adaptive) path and the online/FPGA path are unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add --persistence-spots, a second parameter-free detector alongside --adaptive-spots.
Instead of a hard per-ring threshold it builds the noise-normalised image
z = (I - ring_mean) / sqrt(ring_sigma^2 + read^2) (same per-ring background as the
hard variant) and scores every intensity maximum by its 0-D topological persistence:
sweeping the height from high to low, each maximum is born and, when its basin meets
a taller one at a saddle, dies with persistence = birth - saddle, in sigma. A lone
noise spike merges into the background almost immediately (persistence ~1 sigma); a
real peak stands many sigma proud. Emitting maxima whose persistence clears the same
z(E) significance bar needs no photon threshold and no min-pix, and it deblends
touching peaks (each keeps its own maximum). Implemented with the same union-find
idiom as the connected-component labeller.

On serial stills this auto-adapts with no per-dataset tuning like --adaptive-spots,
finding fewer but cleaner (deblended) spots; the hard-threshold variant remains more
sensitive on the very weakest data. Both share the per-ring background and read-noise
floor. comp_of is allocated lazily so the default and hard-adaptive paths pay nothing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add --soft-weight (implies --adaptive-spots): give every detected spot a
continuous quality weight in (0,1] and keep the highest-weight spots rather than
the brightest, so a deliberately loose detector self-cleans -- bright ice / salt
/ jet blobs and single-pixel noise no longer evict faint clean Bragg spots from
the max-spots cut.

The weight is a product of dimensionless gates (AdaptiveSpotFinderCPU::ApplyWeights,
computed against the per-ring background the adaptive finder already builds): a
logistic ramp in the spot's SNR and a soft size band (rises from one pixel,
plateaus, falls for oversized ice/salt/streak blobs). It carries on
DiffractionSpot -> SpotToSave and is consumed by FilterSpotsByCount, which ranks
by {non-ice, weight, intensity} when requested and by intensity otherwise, so the
classic and FPGA paths are unchanged.

Honest result: on the serial-stills battery this is index-rate-NEUTRAL. The
weighted ranking only changes the outcome when the spot count exceeds the
max-spots cap and the weight disagrees with intensity in a way that affects
indexing; the adaptive detectors already produce clean spot lists and the weak
sets sit under the cap, so re-ranking is a wash there (and a wash, not a
regression, on the one set that floods). Its intended benefit -- robustness to
ice/jet-contaminated frames and to a loosened detector -- is not exercised by
this battery; kept opt-in as the substrate for that.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
On flooded or noisy still frames (weakly-diffracting detectors, XFEL background,
ice) the full spot list derails the known-cell indexer: its many spurious peaks
compete with the true reflections for the search, so genuinely diffracting frames
fail to index.

Seed the indexer with a few spot-count subsets (30 / 80 / all) and keep the lattice
that explains the largest FRACTION of its own seed -- a lean, clean seed that a good
lattice indexes almost fully beats a flooded seed it fits only in small part. This
auto-selects a lean seed on noisy frames and the full seed where the extra spots are
real signal, with no per-dataset setting. Geometry refinement and integration still
use the full spot list (the orientation refiner filters spots by lattice match, so
the flood is ignored while high-resolution spots are kept), so resolution is
preserved. Costs at most ~3 indexer calls per frame, only on frames that do not
index on the first, lean seed.

Lifts the indexed-crystal yield on mildly-flooded synchrotron serial data with no
regression elsewhere. Stills only; the rotation indexing path is unaffected.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The per-image CC-to-reference filter (--min-image-cc) was only honoured on
the rotation merge; the stills merge added every crystal unconditionally.
Extend it to stills so the flag is meaningful there too: on flooded frames
that produce many spurious lattices (large-cell serial data), the crystals
whose per-image CC to the reference falls below the limit are dropped,
keeping only the coherent ones in the merge.

Opt-in and default-off (limit 0 -> the loop passes cc_filter=false and the
merge is bit-identical to before), so no existing behaviour changes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two opt-in tools for weak serial-stills tuning; both default-off, so the
default pipeline is bit-identical (verified: a serial-stills reference run
reproduces HEAD's 7.85% indexing rate exactly).

--local-snr <sigma> (AdaptiveSpotFinderCPU::FilterByLocalSNR): after the loose
per-ring adaptive threshold builds connected-component spots, drop any spot that
does not stand this many sigmas above its OWN LOCAL background (robust median/MAD
of a square annulus), not just the azimuthal ring mean. On structured-background
(XFEL) frames the ring mean underestimates the local diffuse level in some
sectors, so the ring threshold floods; a real Bragg peak still stands many local
sigmas proud. Validated on XFEL stills to separate real peaks from flood at the
pixel level (real median local-SNR ~70 vs flood ~2.6; SNR>=5 keeps ~99.8% of
real peaks, ~14% of flood). GPU-portable (a per-spot local reduction). NOTE: on
the current serial-stills battery it is index-rate/CC1/2 neutral -- the flood that
survives as CC clusters overlaps weak-real spots, and only lattice-fit separates
those -- but it is the correct tool for genuinely floody data (ice/jet/loosened
detector) and the right substrate for the online FPGA path.

--min-indexed-fraction <f>: exposes the previously hardcoded 0.20 minimum
indexed-spot fraction (AnalyzeIndexing) as a per-run setting. Lowering it admits
weaker/sparser crystals; on flooded XFEL data the extra lattices are spurious
(pair with --min-image-cc to gate them), on clean synchrotron data there are no
marginal frames so it is a no-op -- useful as a gating-experiment primitive.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Drops --persistence-spots and AdaptiveSpotFinderCPU::RunPersistence (the 0-D
topological-persistence detector added in 5a33b0743). It was a research variant
that never beat the hard-threshold adaptive detector on a CC1/2 basis and is a
GPU dead-end (global candidate sort + union-find), so it is not a production
path. The hard-threshold --adaptive-spots detector is unaffected.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Drops the --reject-delta-cchalf flag and MergeOnTheFly::DeltaCChalfReject.
The CLI value was parsed but never consumed (the method had no call site), so
the flag was already a no-op. Wiring it up and testing against an external
reference structure showed it is confirmation bias: on a spurious-crystal flood it
raised internal CC1/2 while CCref (correlation to the true structure) fell, and
it never improved R_meas. The merge weights are already correct; per-crystal
merge-side rejection has no genuine lever here.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Both were opt-in adaptive-spot refinements that did not help. Soft per-spot
weighting was index-rate neutral across the battery (re-ranking only bites when
spots exceed the max-spot cap, which weak serial data does not reach). The
local-SNR gate was neutral on index rate and degraded merged CC1/2 on flooded
XFEL data. Drops the flags, ApplyWeights/FilterByLocalSNR, the per-spot weight
field, and the by-weight FilterSpotsByCount branch (now strongest-first only).
--adaptive-spots itself is unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Trims three opt-in stills parameters that did not improve data quality on the
external-reference (PDB R-free) battery and only added code:

- --partiality-uncertainty: the (1-p)/p merge-sigma term was null on all four
  serial-stills datasets of the battery vs their reference structures (and
  neutral-to-harmful at higher coefficients); removed the flag, setting and
  CorrectedSigma term.
- --stills-modulation: the detector-plane flat-field surface was net-negative
  on flooded data; removed the flag, setting and MergeOnTheFly::RefineModulation
  (the rotation modulation in RotationScaleMerge is unaffected).
- --min-indexed-fraction: every value other than the 0.20 default collapsed
  CC1/2; removed the override flag/setter, keeping the fixed 0.20 acceptance
  floor.

Default behaviour is unchanged (all three were off / at their default).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
AdaptiveSpotFinderGPU does the per-resolution-ring reduction once on the GPU and
drives both products from it: the azimuthal-integration profile (corrected space)
and the self-calibrating adaptive spot-detection threshold (raw counts). This
replaces the separate GPU azint pass and the host-side adaptive spot finder that
runs on the GPU path today. On a ~4.5 MP detector it does both jobs in ~1 ms/frame
versus ~40 ms for the CPU adaptive finder (~42x), with an identical spot list and
azimuthal profile.

The per-ring threshold math (Poisson tail + read-floored Gaussian, operating point
from the false-pixels-per-frame knob) is factored into AdaptiveThreshold.h so the
CPU and GPU finders share one source of truth and cannot drift.

Wired opt-in via a MXAnalysisWithoutFPGA constructor flag, default on for the rugnux
offline path and the interactive viewer, off for the online receiver (so the broker
path is unchanged). When on, Analyze() skips the separate azint pass and lifts the
profile from the fused engine. The viewer gains an "Adaptive threshold" checkbox that
greys out the signal/noise and photon-count sliders (the adaptive finder uses neither).

Dedicated tests exercise both products (spot-finding parity vs the CPU finder,
azimuthal profile vs a standalone GPU azint) plus a speed benchmark. Validated
end-to-end on lysozyme serial stills: fused == CPU-adaptive index rate and merge stats.

Docs: new section 3.2 in docs/CPU_DATA_ANALYSIS.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
For stills indexing the minimum-pixels-per-spot filter is now chosen per image
instead of being fixed: the frame is indexed at min-pix 3/2/1 and the setting that
maximises indexed-spot count weighted by indexed fraction (n_indexed^2 / n_total)
is kept, then integrated once at that min-pix. The fraction factor keeps a smaller
min-pix's extra spots only when the lattice actually explains them, so strong frames
retain their real weak spots (extending resolution) while noise-flooded frames stay
strict.

The mode is selected by the presence of --min-pix-per-spot, now optional
(SpotFindingSettings::min_pix_per_spot is std::optional<int64_t>): omit it for the
adaptive per-image path, give a value to force a fixed min-pix. It applies only to
the stills indexing path -- rotation indexing builds one global lattice and keeps a
fixed min-pix, and the online receiver and the FPGA host path always carry a concrete
value, so neither changes. IndexAndRefine::ProcessImage now returns whether the frame
indexed, to drive the per-image selection.

Exposed in the jfjoch_viewer spot-finding settings (adaptive-threshold and
adaptive-min-pix checkboxes, each greying out the control it overrides); the broker
uses neither.

Validated on the full rotation regression battery (no regression) and the whole
serial-stills target battery at full image count.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The qimg_buffer_ member was added to avoid reallocating the full-size image
every recolour, but GeneratePixmap still built a local QImage and the member
was never referenced. Wire it up: the buffer is reallocated only when the
image dimensions change.

The data pointer is taken once, before the parallel loop. scanLine() is
non-const and would otherwise have every worker detach the buffer at the same
time, which is a data race as soon as the buffer is shared with the pixmap.

18.1 Mpx recolour: 28.0 -> 22 ms (measured on the colouring path alone).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
GeneratePixmap wrote every pixel twice: once into the QImage and once into
image_rgb. The only reader was writePixelLabels, which needs a colour for at
most 5000 pixels and only above 30x zoom, so the mirror cost a W*H*3 buffer
and a second store per pixel to serve a fraction of a percent of them.

Read the colour back from the rendered image instead. 18.1 Mpx colouring loop:
9.5 -> 6.2 ms.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every recolour ended with QPixmap::fromImage(), which allocates a second
full-size buffer and converts the whole image into the screen format. That
conversion was the largest single cost left in the colouring path.

Replace QGraphicsPixmapItem with a small item that paints qimg_buffer_ with
QPainter::drawImage. The buffer is already what the raster engine wants, so
nothing is converted or copied. The item declares its opaque area, as the
pixmap item did, so the view still skips the background fill underneath it,
and it turns SmoothPixmapTransform off before drawing to keep the
nearest-neighbour sampling QGraphicsPixmapItem gave us by default -- zoomed-in
detector pixels stay sharp squares.

GeneratePixmap is renamed RenderImage: it no longer makes a pixmap.

18.1 Mpx recolour: 22 -> 5.6 ms (28.0 ms before this series).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ctrl+wheel, Shift+wheel and the foreground slider each recoloured the whole
image synchronously, once per input event. On a large detector the recolour is
slower than the events arrive, so they queued up and the view lagged behind the
cursor for as long as the user kept scrolling.

Defer the recolour to a zero-delay single shot and drop the intermediate
values: at most one recolour is in flight, and it always uses the newest
foreground/background.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
"auto img = image->Image()" deduced std::vector<int32_t> by value, so every
frame copied the whole detector image before converting it -- 72 MB on a 16 Mpx
detector. Bind a const reference instead.

The sentinel-to-float conversion also ran single-threaded on the GUI thread;
spread it over rows the same way RenderImage does. 18.1 Mpx: 11.8 -> ~1 ms,
plus the copy that is now gone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Panning called updateOverlay() three times per mouse move: once for each
scrollbar's valueChanged -> onScroll(), then once explicitly. Zooming was the
same. Every one of those tore down and rebuilt every overlay item.

Suppress onScroll() for the duration of the gesture instead, and let the
gesture do its single rebuild at the end. Note this cannot be done by blocking
the scrollbars' signals: QAbstractScrollArea drives the actual scrolling off
valueChanged, so blocking it would stop the view moving at all.

updateOverlay() also refreshed the image item unconditionally, which marks the
whole item dirty and forces a full-viewport repaint even though pan and zoom
never change the pixels. Track whether RenderImage has run since the last
refresh and skip it otherwise.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
DrawResolutionRings traced every ring point by point on each overlay rebuild:
361 ResPhiToPxl calls per ring, so about 4000 geometry evaluations per rebuild
with the 11 ice rings shown -- and a rebuild happens on every pan step.

The contours depend only on the ring list and the geometry, neither of which
changes while the view moves, so keep them. The cache is keyed on the ring list
(which RingMode::Auto recomputes from the visible area, so it still re-traces
when it should) and cleared in loadImage for a possibly-new geometry. Labels
are still placed per rebuild: they depend on the visible rect.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
image_fp is the base class's one pixel representation, and it earns that for
three of the four image widgets: the azimuthal image is already float, the grid
scan holds computed 1/sigma^2 floats, and the calibration viewer accepts eight
source types from uint8 to float64. The diffraction image is the odd one out --
its source is a large int32 buffer -- and it is the one paying: a full
int32 -> float pass plus a second resident copy of the image, on every frame.

Split the mapping from the source. PixelColorMap holds the precomputed LUT
constants and does value -> colour; a virtual ColorRow() picks the pixels out of
whatever buffer the subclass has. Both paths now go through the same Apply(), so
only the gap/bad/saturated dispatch differs, and it lines up exactly with the
encoding LoadImageInternal used:

    GAP_PXL_VALUE       -> NAN  -> gap
    ERROR_PXL_VALUE     -> -INF -> bad
    SATURATED_PXL_VALUE -> +INF -> saturated

The base class still needs real pixel values for ROI statistics and per-pixel
labels, so image_fp is filled on demand instead of per frame -- and only when
something reads it: a non-empty scratch ROI, or labels above 30x zoom. Neither
happens while simply looking at frames, and nothing else routinely sets roiBox
(the named ROIs are computed in the reading worker, not here).

18.1 Mpx: 10.5 -> 5.9 ms per frame and 72 MB less resident. 4.5 Mpx: 1.8 -> 1.0 ms
and 18 MB.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The status bar, the resolution readout and the magnifier were all regenerated
on every single mouse motion event. Each regeneration repaints, and on a remote
X session a repaint uploads the whole window regardless of how little changed,
so the pointer merely crossing the image saturates the link: measured with a
counting relay in front of the X server, 50 motions over the image cost 273 MB,
and a build with the hover work removed cost 18 KB.

Rate-limit it. Two details matter:

- The limit is applied inline, not from a timer. Running the update inside the
  mouse event keeps its damage in the same repaint as anything else that event
  triggers (a pan). A first attempt deferred the work to a timer instead, which
  split one repaint into two and made panning measurably worse.
- The catch-up that reports the final position is debounced, not queued per
  skipped motion, so it fires once after the pointer stops rather than
  repeatedly mid-gesture.

mouseHover() now takes the scene position and modifiers instead of the event,
which also removes the identical mapToScene() from all four implementations.

Hover traffic over 3 repeats: 173 MB mean -> 140 MB, and the run-to-run spread
drops from +-14% to +-2%. The harness tops out near 30 motions/s, barely above
the 15 Hz limit; a real mouse reports far faster, where the cap does more.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The hovered "d = ... A" readout was a QGraphicsTextItem flagged
ItemIgnoresTransformations, repositioned on every mouse motion. Qt cannot
compute a tight dirty rect for an item that ignores the view transform, so it
marks the entire viewport dirty whenever such an item moves or changes text --
and this one moved constantly.

Paint it in drawForeground() in viewport pixels instead, and repaint only the
union of its old and new rectangles. That also removes the item lifetime
special-casing: it was deliberately kept out of overlay_items_, had to be
nulled by hand after scene()->clear(), and carried comments in three places
warning about the dangling pointer.

This does not reduce raw X11 traffic -- there every repaint uploads the whole
window whatever the damage -- but it cuts the work per hover, and it does
matter under a compressing remote protocol (VNC/NX/xpra), which encodes only
the region that actually changed.

Verified against the previous build: same text, colour and position.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
FullViewportUpdate redraws the whole viewport on any change. The attached
comment ("keep overlays in pixel units independent of zoom") does not describe
what the setting does, and nothing here needs it: SmartViewportUpdate repaints
the changed rectangles and falls back to a full repaint by itself once there
are too many to be worth tracking.

This is the view used by the calibration window and the magnifier, and the
magnifier is driven from every hover, so on a remote session it repainted
its whole viewport per pointer motion.

Note: not exercised visually -- both windows are opened from menus, which the
headless harness does not drive. The change is a repaint-mode switch with no
effect on what is drawn.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
centerAt() checked isVisible(), but imageLoaded() did not, so every frame built
a SimpleImage over the whole detector image and ran it through the full
JFJochSimpleImage path -- convert to float, colour every pixel, redraw -- to
feed a 320x320 window that is closed by default and stays closed most of the
time.

Remember the frame instead and do the work in showEvent(). Holding the
shared_ptr also keeps alive the buffer that the SimpleImage's CompressedImage
points into, which it did not own.

Stepping 30 frames with the magnifier closed: 5545 -> 4770 ms CPU (-14%), on a
2.8 Mpx detector; the saving is per-pixel, so it grows with detector size. With
the magnifier open the cost is unchanged (5500 ms), which is what was being
paid unconditionally before.

Verified in the GUI: opening the magnifier still populates it, and it still
refreshes when the frame changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The window is a placeholder for future functionality and is closed almost all
of the time, but it extracted the frame's spots and rebuilt and uploaded its
vertex arrays on every image, whether or not anything was on screen.

Guard it in rebuildGL() rather than at each of the eight call sites, so any
future caller inherits the behaviour: while hidden it only records that a
rebuild is owed, and showEvent() pays it. imageLoaded() additionally skips
extracting the frame's spots, which is the other half of the per-frame work.

The OpenGL code path is untouched and still built and exercised the moment the
window is opened.

Note: I could not show a CPU saving for this on the headless test machine --
there, ~74% of the process CPU is Mesa llvmpipe software rasterisation that I
was unable to attribute to any per-frame code path, and it swamps the effect.
The work being skipped is nonetheless unambiguously unnecessary.

Verified in the GUI: after stepping frames with the window closed, opening it
shows the current frame's spots, and it keeps updating while open.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The magnifier and the main view are two views of the same image at different
position and zoom, but the magnifier ran the whole pipeline again on its own
copy: it wrapped the same int32 buffer in a SimpleImage, converted it to float,
coloured every pixel and kept its own full-size QImage. That is a second
conversion and two extra full-detector buffers (20 MB at 2.8 Mpx, 138 MB at
18 Mpx) to feed a 320x320 window.

Separate producing a frame from displaying one:

- JFJochImage keeps the rendered frame in a shared_ptr<QImage> (the pointer is
  stable for the widget's lifetime; only the contents change, so the existing
  buffer reuse is unaffected), publishes it via Frame() and announces new
  pixels with frameRendered().
- JFJochImageItem holds that shared_ptr instead of a reference to a member of
  its owner, which also removes a lifetime coupling.
- JFJochFollowerImage is a small read-only view of such a frame with its own
  zoom and centre. It shows only the image: overlays, ROI tools and per-pixel
  labels belong to the view that owns the data.
- The magnifier becomes one of those, fed from frameRendered().

Consequences beyond the saving: the magnifier now agrees with the main view on
colour map, contrast and HDR mode, which it never did -- it was wired to
neither, so it always drew with its own defaults. And the visibility guard
added in 6d1af4921 is gone: there is no longer any per-frame work to skip, so
nothing needs guarding. That guard was a workaround for this design.

Stepping 30 frames with the magnifier open: 5550 -> 4810 ms CPU, which is what
it costs with the magnifier closed (4770 ms) -- it is now free either way.

Verified: main image panel and a drag-pan stay pixel-identical to the
pre-refactor binary (AE=0); the magnifier follows the cursor, updates on a new
frame, and now tracks a colour-map change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Users expect a magnifier to tell them the counts, which the follower view could
not do: it has the rendered pixels but not the numbers behind them.

Take them from the detector's int32 buffer directly, the same source the main
view colours from, so no float copy of the image is needed - the magnifier
still holds nothing full-size of its own, only a shared_ptr to the frame and
one to the reader image.

The labels are painted in drawForeground() rather than as scene items. The main
view creates up to 5000 QGraphicsSimpleTextItems per overlay rebuild for this;
here they are just drawn, so there is no item churn and no scene invalidation.
Text is laid out in viewport pixels so it stays a constant readable size, and
black/white is chosen from the luminance of the rendered pixel underneath, as
the main view does.

Threshold is the same 30x as the main view, so the default 12x magnification
shows no labels until the user wheels in; a cap keeps pathological window sizes
from drawing thousands of them.

Verified in the GUI at 32x: counts drawn per pixel with white text over the
dark centre of a Bragg peak and black elsewhere, and "Gap" across a module gap.
Main image panel still pixel-identical to the pre-series baseline.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two changes to the per-pixel value labels, which appear above 30x zoom.

They were up to 5000 QGraphicsSimpleTextItems created and destroyed on every
overlay rebuild - so on every pan step while zoomed in. Paint them in
drawForeground() instead: no item churn, no scene invalidation, and the text is
laid out in viewport pixels so it is a constant readable size rather than a
scene-space font scaled by 0.2. Same approach as the magnifier's labels.

The value text becomes a virtual, PixelLabel(). The base still formats from
image_fp, which is what the genuinely float-valued views hold (azimuthal
profile, grid-scan 1/sigma^2, the calibration viewer's eight source types).
JFJochDiffractionImage overrides it to read the int32 image directly: counts are
exact integers, so routing them through float32 is a detour that also cannot
represent summed values above 2^24 exactly.

Verified at 38 wheel clicks over a module edge: identical values and
gap/contrast handling to the previous float path, now centred in each pixel.
Fit-view panel still pixel-identical to the pre-series baseline.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Drawing an ROI only means something where there are detector counts to
accumulate. Gate the gesture on a virtual AllowROI(), true only for
JFJochDiffractionImage: shift-drag, the resize handles, the hover cursor and the
"Clear ROI" context entry now do nothing in the azimuthal, grid-scan and
calibration views, which cannot report anything about a box anyway.

The statistics move out of the base class into the diffraction view and read the
int32 image directly, so no float copy of the detector image is built for them
either. With the labels already converted, image_fp is now untouched by the
diffraction view, and the lazy EnsurePixelValues machinery it needed is gone.
image_fp stays as the base's representation for the views whose data really is
float: the azimuthal profile, the grid-scan 1/sigma^2 map, and the calibration
viewer's eight source types.

Removed with it: the ROI readouts in the calibration and 2D azimuthal windows,
which were the only two consumers of roiCalculated -- the diffraction view
emitted it and nothing listened. Nothing surfaces ROI statistics now; the
pixel-mask case wants rectangles counting excluded pixels and deserves its own
design. JFJochViewerROIResult is still used by the side-panel ROI list, so the
widget stays.

Verified in the GUI: shift-drag in the diffraction view still draws the box,
turns it into a named ROI and runs the statistics; fit-view panel remains
pixel-identical to the pre-series baseline.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The statistics shown for a drawn ROI do not come from the view at all. Drawing
one promotes it to a named ROI, roiGeometryEdited goes to the reading worker, and
the worker's per-image results arrive in ImageData().roi, which is what the
Inspector's ROI section displays. So accumulateROI/CalcROI/roiCalculated were a
second implementation of the same thing whose output nothing read -- and the
worker's version is the better one: it handles the mask and it persists per
image.

Remove them. The view now owns only the ROI's geometry and gestures, which is
all the worker needs from it.

This corrects the previous commit's claim that nothing surfaces ROI statistics:
the Inspector does, via the worker. Verified by drawing a box over the beam
centre: Sum 65453, Max 1634, Mean 0.468, centre of mass (786.3, 843.4) against a
beam centre of (764, 850). Note the numbers appear from the next analysed frame
onward, since the worker attaches them at analysis time and the displayed frame
was analysed before the ROI existed -- that behaviour is unchanged here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Drawing an ROI reported Sum 0 and Max 0 until the next frame was loaded, with
only the pixel count looking right because that comes from the ROI map rather
than from the data.

SetROIDefinition_i called RunROIOnly, which integrates whatever the preprocessor
buffer already holds. LoadImage_i only preprocesses an image when a ROI is
already defined, so the very first ROI is drawn on an image that was never
preprocessed: the buffer is empty and every sum integrates to zero. From the next
frame on a ROI exists, LoadImage_i preprocesses, and the numbers look correct --
which is what made this look like a refresh problem rather than a wrong call.

Use AnalyzeROIOnly, which preprocesses the image before integrating. It costs a
pass over the image per ROI edit, of the same order as one recolour, and ROI
edits already keep at most one recompute in flight (live_pending_ in
JFJochDiffractionImage), so the editing rate is bounded. I did not measure the
drag rate specifically.

Verified with a single frame loaded and no frame step: the first ROI drawn over
the beam centre reports Sum 66065, Max 2761, Mean 0.473, centre of mass
(787.3, 844.5). The Max equals the image's own reported maximum of 2761, as it
must for a box containing the brightest pixel.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds Help > Data Analysis Algorithms, showing docs/CPU_DATA_ANALYSIS.md in a
window. The document is baked into the binary through the Qt resource system
(aliased to :/cpu_data_analysis.md), so it needs no docs/ directory at runtime
and cannot drift from the build it shipped with. Same shape as the existing
third-party licences window, created once and raised thereafter.

Limitation worth knowing: QTextBrowser::setMarkdown renders the headings, lists,
emphasis and inline code well, but it has no math support, so the inline LaTeX in
the more quantitative sections appears as raw "$...$" source. The descriptive
material - which is most of the 744 lines - reads fine. Fixing that properly
means either pre-rendering the document to HTML with a math filter at build time,
or sending the user to the Read The Docs copy instead; neither seemed worth doing
without knowing which you would prefer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Revert "Viewer: read the data-analysis algorithm documentation from Help"
Build Packages / build:viewer-tgz:cpu (push) Successful in 7m14s
Build Packages / build:viewer-tgz:cuda (push) Successful in 8m3s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 11m29s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 11m30s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 11m42s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 9m20s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 10m16s
Build Packages / build:rpm (rocky8) (push) Successful in 10m49s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 11m25s
Build Packages / build:rpm (rocky9) (push) Successful in 11m46s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 11m21s
Build Packages / Generate python client (push) Successful in 25s
Build Packages / Build documentation (push) Successful in 1m2s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m30s
Build Packages / XDS test (durin plugin) (push) Successful in 7m52s
Build Packages / DIALS test (push) Successful in 12m25s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m38s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m16s
Build Packages / Unit tests (push) Successful in 1h2m14s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
38f1c1a387
This reverts commit 681e80e46.

Dropped for now: QTextBrowser::setMarkdown leaves the inline LaTeX in
CPU_DATA_ANALYSIS.md as raw "$...$" source, and the document's own HTML/math
rendering wants sorting out first. Reverted rather than rewritten so the
implementation stays available in history to bring back afterwards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
One changeset, developed together in response to a review of this branch, so the
files carry several of the changes at once. Full test suite passes (733 cases).

Spot finding
- Split ImageSpotFinder into Detect() (flag strong pixels - the expensive
  per-pixel pass) and ExtractSpots() (CCL + min/max-pix + resolution mask), with
  Run() = both. The per-image min-pix escalation now detects ONCE and repeats
  only the cheap extraction, instead of re-running the whole finder four times
  per frame as it did on the default path. It also keeps the winning attempt's
  spot list rather than re-extracting it, so the frame that is integrated is
  exactly the frame that was scored - which a GPU re-extract could not guarantee
  (float atomic ordering).
- spot_finding_time_s no longer swallows indexing time, and indexing_time_s now
  sums every escalation call instead of reporting only the last.

Detection limits follow the detector
- The azimuthal-integration upper q and the spot-finding high-resolution limit
  are now std::optional, in the C++ structs AND in the OpenAPI schema, and
  resolve to the detector's own maximum (DiffractionExperiment::GetDetectorMaxQ_
  recipA). Adaptive detection reads a pixel's ring from the azimuthal bins, so a
  pixel outside that q range could never be strong - the integration range
  silently bounded what detection could see, regardless of the requested
  resolution limit. Regenerated the C++ and TypeScript clients; the viewer and
  the web frontend each gained a "to detector edge" switch.

Detection defaults are now per workflow (measured, not assumed)
- Stills: adaptive detection, min-pix chosen per image, no resolution clipping.
- Rotation: fixed-threshold finder, min-pix 2, 1.5 A limit.
  On a 33-crystal rotation battery, adaptive detection helped four hard crystals
  but deterministically broke three (a lost space group, a halved indexing rate,
  a collapsed merge), and the detector-edge limit cost indexing on a strong
  rotation set (100.0 -> 96.8%). Each is still overridable by its flag, and
  --no-adaptive-spots is new.

Indexer seed escalation
- Stop escalating once a seed's lattice explains >= 90% of the seed spots.
  Previously any frame with >= 80 spots always paid three indexer calls, online
  broker included.

Merge-consistency filter
- --min-image-cc gated on a per-image CC computed BEFORE the stills partiality
  post-refinement and never refreshed; the refiner now recomputes it, so the
  reported CC describes the data that are actually merged.
- Replaced the per-call cc_mask argument with one MergeOnTheFly flag, so the
  merge, the error model and MergeStats can no longer disagree about which
  images are in (the --scale path merged unfiltered while its statistics were
  filtered).

Per-image B-factor refinement (-B) removed
- Measured on four serial-stills datasets: it is a no-op where the per-image fit
  is well conditioned and actively harmful where it is not (CC1/2 -8.1, R_meas
  +23.2 on the weakest large-cell set, whose fits hit their [-50, 200] bounds on
  14-25% of images). It had also been silently DISCARDED since the partiality
  post-refinement landed - reported but not applied. Rather than fix and keep a
  knob with no demonstrated benefit, the flag and the whole image_scale_b_factor
  chain are gone: setting, scaling fit, message field, CBOR, HDF5 write and
  read-back, per-image plot, OpenAPI enum, viewer column and checkbox, docs.
  ScaleOnTheFly no longer needs Ceres at all - the fit is a linear IRLS.
  (The Wilson per-image b_factor is a different quantity and stays.)

Stills partiality width now fits both of its components
- sigma^2 = gamma0^2 + (gamma_e*d*)^2 instead of a purely angular gamma_e*d*
  with gamma0 pinned to 0. Fitted per crystal by least squares of dist_ewald^2
  on d*^2. The angular-only width is fitted over a d*^2-dense population, so it
  was pinned by the high-resolution edge and collapsed at low d*: median
  partiality 0.008 beyond 13 A for reflections that were plainly recorded, 55%
  of them under the merge's partiality floor, and the survivors divided by those
  values - which inflated the merged low-resolution intensity scale 3.6x
  (~ +9 A^2 of apparent B). Measured on 5000 stills: the ramp flattens to 0.89x,
  no observation is dropped any more (701750 -> 716811), shell-mean CC1/2 and
  R-free improve slightly. Note CC1/2, R_meas, completeness and a B-refining
  R-free are all blind to that ramp, which is why it survived earlier validation;
  the cost is high-resolution R_meas (98.5 -> 101.9 shell-averaged).

Removed dead code from add-then-remove churn
- Prediction-time "still partiality" (unreachable: no setter), the phantom
  IndexingSettings::min_indexed_spot_fraction knob (getter, no setter - now the
  constant it always was), StillsPartialityRefine's caller-less Settings
  constructor and its reference to a long-gone env var, ProcessImage's unread
  bool return, an unused include, and a dead viewer overlay hook.

Also
- Viewer: the magnifier compared a QImage with itself, so its scene rect was set
  once ever and it could not pan into a larger dataset; the hover tail timer
  could fire after leaveEvent and resurrect the resolution readout outside the
  image.
- update_version.sh regenerated the frontend lock file BEFORE bumping the
  version (every release shipped an off-by-one lock), and did git rm/git add on
  a path that has not existed since the client moved to src/client - with no
  set -e, both failed silently.
- fpga/pcie_driver/postinstall.sh tested "[ ! occurrences > 0 ]", which is a
  redirect, not a test, so dkms add never ran.
- Unit tests for the adaptive-threshold host functions, which had none.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
--min-image-cc is consumed only by the stills merge (MergeOnTheFly); RotationScaleMerge
never reads it. On rotation data it was accepted and then silently did nothing, so a run
that looked filtered was not. It now says so.

FitProfileRadius_MAD had zero callers - a robust twin sitting uncalled next to the
non-robust estimator that is actually used is a trap, so it goes.

Neither changes any result: verified on a rotation dataset (indexing rate, cell, space
group and merge statistics identical, warning emitted).

Context for anyone tempted to wire that estimator in: I tested exactly that today and it
is NOT justified. The population it would clip is truncated by construction - a spot is
only marked `indexed` when its fractional-Miller norm is inside the indexing tolerance -
and is measurably shorter-tailed than Gaussian (kurtosis 2.85). Across four serial-stills
datasets a MAD-clipped variant only narrowed the prediction window (-17% integrated
reflections everywhere), which was neutral on strong data and destroyed real signal on
weak data (one set lost completeness 96.0 -> 93.9%), with R-free 0.3753 -> 0.3767.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The point-group decision moved with the AMOUNT of data at fixed physics: a
partially twinned trigonal crystal was promoted into the twin's holohedry
whenever the search happened to see a larger first-pass merge, and kept its
true subgroup when it saw a smaller one. Simulation over 6 noise draws with
only the merge multiplicity varying: the twin is promoted 0/6 at multiplicity
2 and 6/6 at 18, while the genuine control is promoted 6/6 throughout. The
cause is that every existing gate is a ratio to the merge error model -
b_parent grows toward the true systematic scatter as sigma shrinks with
1/sqrt(N), while b_cand is already saturated by the twin's disagreement, so
the ratio slides down through a fixed veto. The parent statistic moves with
data amount and the candidate statistic does not.

Gate promotions on the operator disagreement H = <|I1-I2|/(I1+I2)> instead,
as the ratio of the operators a promotion ADDS to the parent group's own
operators on the same reflections. There is no sigma in it, so it cannot
drift with the error model, and the parent normalisation cancels data
quality. Measured over 27 runs, 5 promotion types and 450-1800 images:
genuine symmetry 0.862-1.219, merohedral twins 1.270-2.084. On the synthetic
grid it is flat across a 9x change in multiplicity - genuine pinned at 1.00,
twins 3-12x the bound - which is precisely the property the old gates lacked.
chi^2 and the systematic-b stay as secondary vetoes; they protect against
non-crystallographic pseudo-symmetry, which is where correlation-based
scoring is weak.

Pick the parent carefully: 422 has two maximal subgroups of order 4, and on a
tetragonal crystal twinned by 2[100] the rival (222) is CC-confirmed too and
CONTAINS the twin laws, so normalising against it hides the twin among the
promotion's own real operators (ratio 8.19 against the true parent, 0.78
against the rival). Where several parents tie, judge on the most damning.

Also:

- Report a refused promotion instead of silently processing lower. Merging a
  twin in the twin's holohedry averages non-equivalent reflections into each
  other and cannot be undone from the output; keeping the subgroup costs only
  redundancy. The refusal names the group and the number that caused it.

- Stop the twinning report from arguing in a circle. It ran after adoption and
  conditioned on the adopted group, so a promotion into a holohedral Laue
  class made it print "no merohedral twin law exists" - the test was
  conditioned on the decision it should audit. Twinning is now also measured
  on the subgroup merge before adoption, and the post-adoption text says when
  its own conclusion is not authoritative.

- Compare PRIMITIVE cell volumes in the first-pass scheme tie-break. A centred
  setting's cell is an exact integer multiple of its primitive one (a
  rhombohedral lattice in hexagonal axes is exactly 3x), so the
  integer-supercell test fired on a pure setting difference and demoted a good
  scheme to a threefold-smaller merge - which is what let the twin see the
  small merge to begin with.

Rotation battery, 33 crystals: point-group agreement 30/33 -> 29/33, one
crystal moved. That crystal (P422 -> P222) is the one with the known
unresolved integration defect where reflections near the rotation-axis plane
are wildly mis-integrated; its symmetry mates genuinely disagree, and its
lower-symmetry merge is measurably better (ISa 2.72 -> 3.63, high-shell CC
75.4 -> 86.0). The threshold was not moved to accommodate it: 1.25 sits inside
the measured gap and widening it would admit real twins. Separately the
tie-break improved one crystal's CC1/2 from 77.7 to 84.0.

Tests: a synthetic twin-fraction x multiplicity grid, which is what the search
had never had - the existing tests are noise-free and exercise only Stage B
absences.

A NOTE ON WHAT WAS TRIED AND REJECTED, so it is not rebuilt: the obvious
"physics-anchored" statistic is the disattenuated cross-validated correlation
rho = corr(I_half0(h), I_half1(Rh)) / corr(I_half0, I_half1), which is 1 for
real symmetry at any data quality and 2a(1-a)/((1-a)^2+a^2) for a twin. It
passes the synthetic grid perfectly and FAILS ON REAL DATA IN BOTH
DIRECTIONS - five false refusals of genuine symmetry on the battery, and it
waves through a twin (rho 0.998) that H refuses. The reason is that cc_half
correlates the two halves of the SAME reflection and so measures only random
error, while cc_cross compares DIFFERENT reflections carrying different
systematic error; dividing by cc_half removes the noise and leaves a
systematic floor that varies by crystal AND by operator. Genuine rho measures
0.9987 on strong data and 0.73 on weak. A synthetic generator validates a
statistic's arithmetic, never its premise, and this premise - that the only
departure from exact symmetry is noise - is false for every real crystal.
Any per-operator agreement statistic needs a same-crystal reference; an
absolute threshold on one cannot be made to work by tuning.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The refusal message fell through to the chi^2 branch whenever the
systematic-b balloon veto was the binding test, so it reported a chi^2 ratio
that did not justify the refusal at all - on one battery crystal it printed
"merge chi^2 is 1.25x the subgroup's (bound 1.85)", i.e. a number comfortably
inside its own bound, as the reason for processing in the lower symmetry. A
diagnostic that names the wrong cause is worse than none: it sends the reader
after the wrong statistic.

Report the b test when it is what fired, with both b values and the bound.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The per-frame scale enters every intensity as 1/G, and SolveScaleIRLS floors G
at zero and nothing else. A frame whose fit is not determined by its data can
return G ~ 0.002 against a run median of 0.865, and every observation it
carries is then multiplied by ~500 - sigma by the identical factor, which is
why no sigma-based outlier test can see it and why this looked for a long
time like a partiality problem. (The 1/partiality path is in fact guarded:
min_captured_fraction floors it at 0.7 by default on rotation.)

The window smoothing that should have absorbed such a frame instead made it
permanent. It averages log G over a window, so a scale collapsing toward zero
does not merely corrupt its own frame - its logarithm drags the whole window
down. Worse, where a run has a stretch of frames too sparse to fit at all, the
only FITTED frames in a window can be the collapsed ones, and the geometric
mean then averages the fault with itself. Measured on a multi-lattice dataset:
frames 816 and 818 fitted G = 0.0023 and 0.0014 with every neighbour from 800
to 839 unfitted, so smoothing set G = 0.0018 across the whole neighbourhood -
a 546x amplification. About 500 observations of 152000 (0.66%) then carried
99% of sum(I^2), and the merged CC1/2 read 17.2% where the same data with the
classic finder read 93.7%.

Treat a fitted scale far below the run's median as what it is - an
undetermined scale, exactly like the too-few-reflections case the code already
handles - rather than as a successful fit. Such frames no longer contribute to
the smoothing mean, and a frame whose own scale is not credible takes the
neighbourhood's, or the run's typical scale when the neighbourhood holds
nothing credible either.

The bound is a RATIO to the run's own median because the rotation per-frame G
is not gauge-fixed: G and the group means have an exact global multiplicative
degeneracy, and the fitted median drifts over 0.745-1.358 across the battery.
An absolute floor would reject everything in a run that drifted low.
MIN_CREDIBLE_SCALE_RATIO = 0.02 was chosen from measurement over 12 crystals
in the default configuration, where the smallest legitimate min(G)/median(G)
is 0.070; the failing case sat at 0.0017. It is 3.5x below anything real and
12x above the failure.

Effect on the intensity tail of the failing case: max I 10224 -> 438, and the
top 1000 observations' share of sum(I^2) 0.990 -> 0.421 (the classic-finder
reference is 0.632, so the tail is now cleaner than the run this was compared
against). Rotation battery, 33 crystals in the default configuration: ZERO
crystals differ - no space group, CC1/2, high-shell CC or ISa change anywhere.
The guard fires only on the pathology.

It does NOT rescue that dataset: with the amplification gone its CC1/2 is
26.2% and R_meas 49.2% against the classic finder's 93.7% and 27.8%. Adaptive
detection degrades those intensities for a second, independent reason that is
still open. This commit removes a latent hazard for any run with a sparse
stretch of frames; it is not the fix for that dataset.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
704098712 guarded the per-frame scale on the partials, but it put the check
inside ComputeSmoothGWindow - which only the partials path calls. Step 4
refits the scale from scratch on the COMBINED FULLS (Unity model, so corr is
exactly 1/G), with no smoothing and no floor, and that refit was still free to
collapse toward zero.

It is the same failure and it is worse here, because there is no window
average to dilute it: the collapsed frame's own fulls are multiplied directly.
Measured on a dataset where the previous commit had already fixed the partials
stage, the fulls refit put 1/G = 559x and 175x on two frames carrying 517
observations, and the merged CC1/2 read 26.2% where the intensities ENTERING
that stage were fine - better, in fact, than the comparison run's in all ten
resolution shells (R_meas 34.5% vs 39.5%, CC1/2 89.1% vs 83.0%).

Reject a collapsed scale here as well, on the same measured criterion, and let
those fulls merge unscaled - the state the combine left them in, and the same
fallback the fit already uses for a frame with too few reflections. The check
reads the host fulls after both the CPU loop and the GPU ScaleFulls, so one
implementation covers both paths; the corrected corr is pushed back to the
device exactly as the correction surfaces already do.

THIS IS NOT A DETECTION-MODE PROBLEM. Over 8 configurations (both spot
finders x 4 frame ranges) the separation is exact: every run with a collapsed
scale had CC1/2 <= 58.5%, every run without had CC1/2 >= 74.1%, and nothing
else predicted it. On one frame range it is the DEFAULT finder that collapses
(CC1/2 58.5%) while the other is clean at 93.8%. The instability was never
specific to the finder; it was latent in the scaling stage and either finder
could trip it.

Same 8 configurations, with this commit:

  finder A  full    93.7 -> 93.7   (untouched)
  finder A  -s 1    58.5 -> 91.3   (recovered)
  finder A  -e 899  94.0 -> 94.0   (untouched)
  finder A  -e 898  87.2 -> 87.2   (untouched)
  finder B  full    26.2 -> 91.1   (recovered)
  finder B  -s 1    93.8 -> 93.8   (untouched)
  finder B  -e 899   8.1 -> 91.7   (recovered)
  finder B  -e 898  74.1 -> 74.1   (untouched)

Every collapse recovers; every healthy run is unchanged. The CC1/2 spread over
the four frame ranges falls from 35.5 to 6.8 points for one finder and from
85.7 to 19.7 for the other - this dataset was not sampling a deep instability
when its CC1/2 swung between 17 and 94 across frame ranges, it was sampling
whether this bug happened to fire.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`RefineGeometryIfNeeded` hands XtalOptimizer the WHOLE spot list, not the
indexed subset, and the first pass admits anything within 0.3 fractional-Miller
units of an integer - which is 11.3% of RANDOMLY placed spots, since the
admitted volume is (4/3)*pi*t^3. Every one of them then enters an unweighted L2
fit with an arbitrary rounded index. On images with many detections the
refined orientation ends up 2.3-2.8 degrees from the goniometer-consistent one
and explains 14 of its own 250 spots where the undragged orientation explains
68; mosaicity and profile radius inherit the error and integration follows.

Weight every spot by its intensity divided by the median intensity of its own
equal-count resolution shell, applied as w^2 on the squared residual with
w^2 = r/(1+r). The shell normalisation is the point: refinement needs the
high-resolution spots because they carry the cell and distance, and those are
LEGITIMATELY weaker, so a raw intensity weight would suppress exactly the
spots the fit depends on. Measured, the weight is resolution-neutral - median
exactly 0.707 in every shell, and corr(w, 1/d^2) = -0.20 / -0.11 against
-0.32 / -0.34 for the same function of un-normalised intensity.

This is a PRIOR: it is computed from the spot alone and never looks at the
current residual, so unlike a robust loss it cannot mistake a genuine spot for
an outlier while the starting geometry is still far off and leave the fit
unable to move. That failure is not hypothetical - a CauchyLoss on this same
residual, at the scale the multi-frame GeometryRefiner uses, collapsed one
crystal's indexing rate from 99.89% to 19.83% and was rejected.

It does not work by telling good spots from bad, and it does not need to. No
per-spot property separates spots that index from spots that do not: measured
AUC is 0.53 for peak pixel, 0.53 for total intensity, 0.51 for pixel count,
0.45 for peakedness, and a logistic regression on all twelve available
features with pairwise interactions reaches only 0.64. What the weight does is
halve the EFFECTIVE COUNT of every spot (mean w^2 = 0.517), and the damage
scales with the absolute count of unexplained spots in the objective - 80.6
per frame here against 36.8 for the finder that was never damaged. That is
also why an empirical `--max-spots 66` cap works while leaving the list no
purer than before: it reaches the same operating point by discarding spots.
This reaches it without discarding any, and without a tuned constant.

Rotation battery, 33 crystals, both spot finders:

  finder A   29/33 -> 30/33 point groups   (one crystal P222 -> P4212 = XDS,
                                            its high-shell CC1/2 86.0 -> 98.4)
  finder B   28/33 -> 29/33 point groups   (one crystal I222 -> I23,
                                            its high-shell CC1/2 14.8 -> 38.0)

No crystal lost its point group in either mode and no run failed. On the
meta-stable multi-lattice dataset the CC1/2 spread over four frame ranges
falls 19.7 -> 13.1 for finder B, and the indexing rate rises in 8 of 8
configurations. The crystal that the rejected robust loss destroyed keeps its
99.89% indexing rate exactly.

The cost, stated plainly: ISa falls by 0.2-1.7 on about five crystals (and
rises on two). Point-group correctness is worth more than that - merging in
the wrong symmetry cannot be undone from the output, whereas ISa is a quality
metric of data that remain correct - but it is a real trade and not a free win.

Off by default. The indexers pass a spot list they have already selected, so
their calls are unchanged; only the per-image refinement, which gets the raw
list, turns it on.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A merohedral twin mixes EVERY reflection with its twin mate, so it shifts the
whole distribution of |I1-I2|/(I1+I2). A minority of badly measured
reflections shifts only the tail. The mean cannot tell those apart; the median
is blind to the second and just as sensitive to the first.

Measured on real crystals, moving the statistic from the mean to the median
leaves genuine promotions where they are and pushes every twin up:

  genuine tetragonal    1.016 -> 1.013
  genuine lysozyme      1.051 -> 1.067
  genuine tetragonal    1.238 -> 1.231
  twin (-e 1050)        1.272 -> 1.447
  twin (-e 450)         1.280 -> 1.622
  twin (full)           1.441 -> 1.522
  twin (-e 600)         1.427 -> 2.010

The margin around the 1.25 bound widens from 2.7% (genuine 1.238 against twin
1.272 - uncomfortably tight for a decision that cannot be undone downstream)
to 17.5% (1.231 against 1.447). The bound itself does not move.

Rotation battery, 33 crystals in both detection modes: no point group changed
in either (30/33 and 29/33, as before), and only one crystal's numbers move at
all - the one already documented as nondeterministic between repeat runs of
the same binary. The synthetic twin-fraction x multiplicity grid passes
unchanged. So this buys margin, not outcomes.

Found while testing a different hypothesis, which the same measurement refuted:
a tetragonal crystal whose 422 promotion is wrongly refused reads 1.484 by the
mean and 1.472 by the median, i.e. its disagreement is distribution-wide and is
NOT a badly-integrated minority. That crystal's cause is elsewhere and is not
addressed here - see the note below.

  Its indexing-ambiguity operator (-k,-h,-l) lies INSIDE 422 but OUTSIDE 222,
  so the subgroup merge the search is given mixes lattices indexed in the two
  alternative hands. That corrupts exactly the 4-fold relationships and leaves
  the 2-fold ones intact - measured, the 222 step reads 0.917 and the 422 step
  1.484 - and the corruption is indistinguishable from a twin law. Forcing the
  tetragonal group merges the two hands as equivalent and the same data give
  CC1/2 99.2% at multiplicity 10.7, matching XDS. The failure is worse the
  BETTER the frames index (99.9% vs 63.3% for the run that gets it right),
  because indexing more frames picks up more of both hands.

  So no statistic computed on a subgroup merge can arbitrate a promotion whose
  added operators include an indexing-ambiguity operator. Fixing that means
  resolving the ambiguity before the search, or detecting the coincidence and
  deciding another way; the operators needed to detect it are already computed
  (the run warns about them).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The flag was accepted on rotation data and did nothing - it is read only by the
stills merge (Merge.cpp), and the CLI warned about that rather than fixing it.
Meanwhile RotationScaleMerge already COMPUTES a per-frame correlation against
the merged reference and writes it to the per-image table; nothing acted on it.

Wire the two together. A rejected frame has its partials' corr set to 0, which
is how a frame already leaves the pipeline - every consumer requires corr > 0,
so the combine, the merge and the error model all drop it together. The GPU
path reuses the SmoothCorr kernel with a ratio of 0, so one implementation
covers both. Off by default (0), and verified bit-identical to the previous
binary when off.

What it catches, on the two rotation datasets that have a population to catch:

  a two-lattice crystal - two lattices in two physical AREAS of the sample, so
  the sweep passes from one to the other and whole blocks of frames measure a
  different crystal from the one being merged (frames 500-700 index perfectly
  well at a per-frame CC of 0.22 against 0.47-0.56 either side, in 11 contiguous
  runs). R_meas 28.6 -> 24.6%, CC1/2 93.6 -> 95.1, high-shell CC 23.4 -> 38.3.

  a second dataset with 9.5% of frames below CC 0.30: R_meas 24.3 -> 23.4%,
  CC1/2 92.6 -> 93.4.

The criterion is "this frame disagrees with the merged reference", NOT "this
frame is off-crystal". It happens to catch both, because a frame that measures
nothing and a frame that measures a DIFFERENT crystal fail the same test, and it
does not need to know which. For the two-area case that is a workaround, not a
treatment: it recovers one crystal by discarding the other, where processing the
two as separate sweeps would keep both. The frame-block structure is clean
enough that such a split could be detected automatically.

WHY THERE IS NO DEFAULT. The per-frame CC is not comparable between datasets -
it is as much a measure of data quality as of frame validity. Measured medians
across the battery run from 0.30 to 0.81, so one absolute bound removes 13
frames from one dataset and 584 of 1800 from another:

  battery at --min-image-cc 30, 33 crystals: no point group changed (30/33),
  four crystals clearly better (one +5.4 CC1/2 points, the two-lattice case
  above, and ISa gains of 1.3-4.6 on three others) - and one healthy crystal
  lost a third of its frames and with them its high-resolution shell
  (CC1/2_hi 26.2 -> 2.0).

This is the same trap as an absolute bound on any per-operator or per-frame
agreement statistic, and the same one the per-frame scale guard avoids by
measuring against the run's own median. A principled version would cut on the
SHAPE of the per-frame CC distribution - a dataset with a bad subpopulation is
bimodal, a uniformly weak one is not - rather than on an absolute value. Until
that exists this stays opt-in, and the per-image CC it keys on is already in the
_image.dat table for anyone choosing a value.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rugnux: adaptive spot detection is the default for rotation data too
Build Packages / build:viewer-tgz:cpu (push) Successful in 8m0s
Build Packages / build:viewer-tgz:cuda (push) Successful in 9m13s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 13m40s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 13m48s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m42s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 14m27s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m43s
Build Packages / build:rpm (rocky8) (push) Successful in 11m32s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 13m9s
Build Packages / XDS test (durin plugin) (push) Successful in 7m22s
Build Packages / Generate python client (push) Successful in 27s
Build Packages / Build documentation (push) Successful in 1m7s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m57s
Build Packages / build:rpm (rocky9) (push) Successful in 13m25s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 13m50s
Build Packages / DIALS test (push) Successful in 13m52s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m0s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m55s
Build Packages / Unit tests (push) Successful in 1h2m59s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
6f4917dcee
It was held back because a 33-crystal rotation battery showed it breaking three
crystals deterministically - a lost space group, a halved indexing rate and a
collapsed merge. None of those causes turned out to be in detection.

The extra spots adaptive finds are real. Measured per spot against a
finder-neutral local background: 64% recur at the same position on the adjacent
frame (chance rate 0.5%) with 2-frame rocking curves, and 0.00% would fail a
conventional local SNR >= 4 test, median local SNR 34. What they include is
genuine peaks belonging to no lattice the indexer found, and the damage they did
scaled with their absolute COUNT (80.6 per frame against 36.8 for the fixed
finder), not with their quality - which is why nothing aimed at judging
individual spots ever worked.

The three failures fell to fixes elsewhere:

  merge collapsed  - a per-frame scale free to collapse toward zero amplified
                     two junk frames by 546x (704098712, ec7a82613). Not a
                     detection problem at all: the fixed-threshold finder trips
                     the same bug on a different frame range.
  space group lost - the per-image geometry refinement was dragged 2.3-2.8 deg
                     off by the weak-spot tail in an unweighted fit; weighting
                     each spot by how strong it is FOR ITS RESOLUTION fixed it
                     (ae126c3d5), and gained a point group for the fixed finder
                     too.
  indexing halved  - gone with the same two; that crystal is now better under
                     adaptive (CC1/2 92.6 -> 95.9, high-shell 44.9 -> 56.4).

Battery, 33 crystals, adaptive vs the fixed finder:

  exact space group matching XDS        26/33  vs  25/33
  point group matching XDS              29/33  vs  30/33
  ISa better on                         6 crystals
  recovers a screw axis the other misses (P321 -> P3121)

The one point group it loses is a tetragonal crystal where adaptive collects
2.4x the observations at better R_meas (21.4% vs 28.3%) and better ISa (4.58 vs
2.91), and forced to the right group gives CC1/2 99.2% at multiplicity 10.7 -
matching XDS. Only the automatic symmetry call fails there, and six candidate
causes have been measured and refuted (mixed indexing hands, off-crystal frames,
a badly integrated minority, radiation damage, pseudo-tetragonality, uncorrected
anisotropy). It is left as the subgroup, which is the recoverable direction: -S
gives XDS-quality data from the same run, whereas the failures this unblocks
were not recoverable.

--no-adaptive-spots reverts to the fixed-threshold finder.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
DetectorSetup hardcodes bit_depth_image = 16 for every DECTRIS detector
(DetectorSetup.cpp:78), and GetByteDepthImage() consults that BEFORE the depth
the reader takes from the file - so a file storing 32-bit images had its
overflow computed as a 16-bit one. With the reader also declaring the images
signed, GetOverflow() returned INT16_MAX and GetSaturationLimit() became
min(file value, 32767).

Every count above 32767 was therefore marked saturated, and because the
integration accept gate requires ALL inner pixels valid, the whole reflection
was discarded. That silently removes the strongest reflections of a strong
crystal - the low-resolution ones that anchor scaling - while the file itself
declares saturation at 105000-133000.

Measured on a lysozyme rotation set (200 frames), before -> after:

  saturated pixels per frame   0.815 -> 0.000
  brightest accepted pixel     32738 -> 87633
  mean per-frame maximum       26218 -> 37390

i.e. the ceiling was exactly INT16_MAX and nothing genuine reached it.

Which datasets this touches depends on how bright they are: measured pixels
above the old ceiling range from 0.0 per frame on some rotation sets to 6.1 on
others, so the fix is a no-op on weak data and only ever adds reflections.

Rotation battery, 33 crystals: no point group changed (30/33 before and after)
and no run failed. Four crystals move on quality, in both directions - ISa
1.90 -> 2.40 and 3.29 -> 4.80 on two, 2.97 -> 1.85 and 20.83 -> 18.47 on two
others; three of the four are the battery's known weak or run-to-run-unstable
crystals. The one strong crystal that moves gains 136 observations out of
1.9 million and loses 2.4 ISa: the reflections restored are by construction the
brightest ones, and they carry the systematic error that the strongest
reflections always carry. That is a real cost, but it is the cost of MEASURING
them rather than discarding them unseen, and a lower asymptotic I/sigma on data
that are now complete is preferable to a flattering one on data that quietly
are not.

Only the offline file reader is affected; the online path builds its detector
setup from configuration.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rugnux: report how close a symmetry axis lies to the spindle
Build Packages / build:viewer-tgz:cpu (push) Successful in 7m11s
Build Packages / build:viewer-tgz:cuda (push) Successful in 8m48s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m57s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 14m20s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 14m22s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m31s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 14m36s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 12m59s
Build Packages / build:rpm (rocky8) (push) Successful in 11m47s
Build Packages / XDS test (durin plugin) (push) Successful in 7m42s
Build Packages / Generate python client (push) Successful in 27s
Build Packages / Build documentation (push) Successful in 1m8s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2404) (push) Successful in 13m6s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 13m18s
Build Packages / build:rpm (rocky9) (push) Successful in 13m58s
Build Packages / DIALS test (push) Successful in 14m9s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m11s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m6s
Build Packages / Unit tests (push) Successful in 1h2m16s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
eb70684fa9
A rotation sweep never records the reflections whose reciprocal vector lies
within the Bragg angle of the spindle - the blind cusp. Symmetry normally
supplies them from an equivalent elsewhere in reciprocal space, so the hole
closes. It cannot when a symmetry axis IS the spindle: the cusp is then mapped
onto itself, every reflection in it is equivalent only to other reflections in
it, and it stays empty however long the sweep runs. The user can fix this at
the microscope - re-mount, or add a sweep on another axis - but only if they
are told, and nothing in the output mentioned it.

Report the smallest angle between any proper rotation axis of the adopted
space group and the goniometer axis, always on rotation data, and warn when it
falls under 15 deg. The axis is found by projecting onto each operator's
invariant direction (the sum of its powers annihilates everything else) and
mapping that fractional direction through the refined lattice into the lab
frame; the angle is invariant under the sweep, so the reference orientation is
enough. Cross-check: this reports 30.6 deg for a crystal whose 4-fold an
independent analysis of the XDS orientation matrix put at 30.5 deg.

Measured on three rotation sets: 13.6 deg (2-fold, warns), 16.2 deg (2-fold,
99.7% complete) and 30.6 deg (4-fold). The 15 deg bound is practical rather
than derived - the blind cone's half-angle is the maximum Bragg angle, ~15 deg
for 2 A data at 1 A wavelength - and the wording says what the diagnostic can
honestly support: the angle is a risk indicator, the loss is confined to the
cone rather than spread over the data, and overall completeness may still look
reasonable while the region near the spindle is empty. It does not promise a
completeness number, because across those three sets the overall figure does
not track the angle (99.7% at 16.2 deg, 92.6% at 30.6 deg).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rugnux: --search-min-zeta drops badly-measured observations from the symmetry search
Build Packages / build:viewer-tgz:cpu (push) Successful in 8m30s
Build Packages / build:viewer-tgz:cuda (push) Successful in 9m3s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m52s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 14m4s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 14m23s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 14m27s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m59s
Build Packages / build:rpm (rocky8) (push) Successful in 11m45s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 12m59s
Build Packages / XDS test (durin plugin) (push) Successful in 7m14s
Build Packages / Generate python client (push) Successful in 29s
Build Packages / Build documentation (push) Successful in 1m7s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (rocky9) (push) Successful in 13m21s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 13m3s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 13m38s
Build Packages / DIALS test (push) Successful in 14m4s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m4s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m49s
Build Packages / Unit tests (push) Successful in 1h1m28s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
f2b92e3f4d
zeta is the sine of the angle between a reflection's rocking path and the
spindle. Near 0 the reflection crosses the Ewald sphere almost tangentially,
spends many frames in diffracting position and is measured worst. The de-novo
space-group search asks how EQUAL an operator's paired intensities are, so its
answer is dominated by whichever reflections are measured worst - and when the
spindle lies in a lattice plane, an operator that permutes the two in-plane
axes samples a different mixture of measurement qualities than one that only
flips signs. That is not a fair comparison, and it can make a real symmetry
operator look like a twin law.

Measured on a thaumatin set mounted that way (its 4-fold is 88.9 deg from the
spindle), the added operators' disagreement is 1.74x the parent's over pairs
where both reflections have zeta < 0.85 and 1.003x - i.e. the symmetry is
exact - over pairs where both are above it. The search consequently refuses the
422 promotion and merges the crystal in P222, while the same data forced to the
right group give CC1/2 99.2% at multiplicity 10.7, matching XDS.

With the option the de-novo pass ignores those observations (the final merge
keeps everything - there completeness is the point):

  zeta cut   observations ignored   H ratio   adopted
  0 (off)                       -      1.47   P222
  0.5                     1620648      1.44   P222
  0.7                     3006013      1.34   P21212
  0.85                    4536724   promoted  P4212   (correct point group)

OFF BY DEFAULT, and it must stay off, because the same cut costs four other
crystals their space group (P41212 -> P212121, I23 -> P2, I23 -> I222 twice):
at 0.85 it discards 40-80% of all observations, which on a crystal whose
geometry is not the problem simply starves the search. Two independent
implementations - filtering the pairs that enter the statistic, and filtering
the observations that enter the merge - trade exactly the same crystals, so
this is a property of the cut and not of where it is applied. Verified
bit-identical to the previous binary when off.

The companion diagnostic is already there: the run now reports how close a
symmetry axis lies to the spindle, which is the geometry that makes this
option worth reaching for.

Implementation note for anyone tempted by the cheaper route: excluding these
observations from the ASU grouping alone does NOT work. The 3D combine selects
partials on corr, not on their group, so their intensity still reaches the
fulls and the merged intensities are unchanged - measured, the statistic did
not move by 0.03 while 67% of observations were nominally excluded. Zeroing
corr is what removes an observation from the combine, the merge and the error
model alike.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Space-group search: ask twice - all observations, and only the well-measured ones
Build Packages / build:viewer-tgz:cpu (push) Successful in 8m20s
Build Packages / build:viewer-tgz:cuda (push) Successful in 9m7s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m32s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 13m58s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 14m0s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 14m9s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m22s
Build Packages / build:rpm (rocky8) (push) Successful in 11m51s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 12m47s
Build Packages / XDS test (durin plugin) (push) Successful in 8m57s
Build Packages / Generate python client (push) Successful in 39s
Build Packages / Build documentation (push) Successful in 1m7s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (rocky9) (push) Successful in 12m46s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m32s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m31s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m52s
Build Packages / DIALS test (push) Successful in 14m42s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m32s
Build Packages / Unit tests (push) Successful in 1h15m42s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
25458265d3
--search-min-zeta rescues a point group that the full merge cannot confirm, but
used on its own it is a trade: on the crystal it was built for it recovers the
correct 422, and on four others it costs the space group outright, because
discarding 40-80% of the observations starves operator correlations that were
perfectly healthy. Both ways of applying it - filtering the pairs that enter
the statistic, and filtering the observations that enter the merge - trade the
SAME crystals, so the cut itself is the problem, not where it is applied.

Filip's observation makes it one-way: every disagreement between the two is a
LOST operator, never an invented one. Discarding observations can starve a
correlation; it cannot manufacture symmetry that is not there. So run the
search on both merges and keep whichever found MORE symmetry, and the failure
mode disappears - each arm rescues the other exactly where it fails.

  crystal            all observations   Lorentz-filtered   adopted
  thaumatin (weak)         222                422            422
  tetragonal lysozyme      422                222            422
  cubic insulin x3          23              2 / 222           23

The filtered merge is used ONLY to rescue the point group. The screw and
centering determination always comes from the merge with all the observations,
because systematic absences are decided by the WEAK reflections and the filter
throws most of them away. Preferring the filtered arm on a tie is not a
conservative choice, it is a wrong one: it cost four crystals their screw axes
(P2(1) read as P2, P4(1)2(1)2 as P42(1)2) with the point group and every
intensity statistic identical - a regression invisible to CC1/2, R_meas and ISa.

Where the two find the same ORDER but different symmetry, nothing can prefer
one, so the run says so: it names both space groups, states that the data do
not decide, reports which one processing continued in, and gives the flag to
force the other. Two candidates of the same order imply different molecular
replacement searches, and trying both is cheap next to reprocessing - much
cheaper than a confident wrong answer.

Rotation battery, 33 crystals, both spot finders:

  fixed-threshold finder   30/33 - ZERO crystals differ from the single search
  adaptive finder          30/33 - the same three mismatches, gap CLOSED

The adaptive finder now matches the fixed-threshold one exactly, which it has
not done before: its last remaining loss was the thaumatin set whose 4-fold
sits 88.9 deg from the spindle, and it now reads P42(1)2 (all-observation merge
-> 222, Lorentz-filtered -> 422, higher taken). A merohedral twin stays refused
in BOTH arms at all three frame ranges where it over-promotes, and at one of
them the second opinion is strictly better than shipping behaviour - the full
merge collapses to P1 where the filtered one finds the correct H3.

Cost is the extra scale-combine-merge on already-ingested partials, with no
re-integration: 47.2 s against 47.8 s on the same crystal back to back.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
--search-min-zeta now defaults to 0.85 for rotation, so the de-novo search runs
on a merge of all the observations AND on a merge of only the well-measured
ones, and keeps whichever found more symmetry. Previously it shipped off and
the second opinion had to be asked for.

Rotation battery, 33 crystals, NO flags beyond the resolution limit:

  fixed-threshold finder   30/33 - zero crystals differ from the single search
  adaptive finder          30/33 - the same three mismatches

Both arms now agree crystal for crystal, which they have not done before. The
last disagreement was a thaumatin set whose 4-fold sits 88.9 deg from the
spindle: at defaults it now reads P42(1)2 (all-observation merge -> 222,
Lorentz-filtered -> 422, higher taken) where it read P222. The classic arm is a
strict no-op - zero differences against both the explicitly-flagged run and the
run predating the dual search - so the default costs nothing where the geometry
is not the problem, and 47.2 s against 47.8 s on the same crystal back to back.

The default is safe to set because the two searches can only disagree by a LOST
operator: discarding observations starves an operator correlation, it cannot
invent one. That also makes the 0.85 itself uncritical - too aggressive a cut
only means the second opinion contributes nothing and the full merge wins.
--search-min-zeta 0 restores the single search.

Docs: CHANGELOG gains a 1.0.0-rc.161 section covering the branch, and
CPU_DATA_ANALYSIS records the four analysis changes of this work - the
confidence-weighted per-image refinement, the collapsed per-frame scale guard,
the opt-in per-image rejection, and the operator-disagreement criterion with
the two-search rule.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
VERSION: 1.0.0-rc.161
Build Packages / build:viewer-tgz:cpu (push) Successful in 7m50s
Build Packages / build:viewer-tgz:cuda (push) Successful in 8m41s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 12m17s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 12m23s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 12m50s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 12m52s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m21s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 11m53s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m34s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m3s
Build Packages / build:rpm (rocky8) (push) Successful in 12m47s
Build Packages / build:rpm (rocky9) (push) Successful in 13m1s
Build Packages / Generate python client (push) Successful in 18s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 52s
Build Packages / XDS test (durin plugin) (push) Successful in 7m39s
Build Packages / DIALS test (push) Successful in 12m1s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m21s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 6m52s
Build Packages / Unit tests (push) Successful in 1h2m5s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
43e9de9573
A reflection the group predicts absent counted as a violation when
I/sigma > 3 AND E^2 = I/<I>(shell) > 0.3. Neither half survives contact
with real data:

  * merged sigma is floored at b|I|, so merged I/sigma saturates at ISa
    for nearly every reflection - the I/sigma half is an on/off switch
    keyed on ISa vs 3, not a per-reflection test. On one crystal the
    absent class read <I/s> 4.10 against 3.73 for the present class while
    being genuinely extinct;

  * <I>(shell) decays with resolution while a systematically-absent
    reflection keeps a small NON-decaying residual (background / profile
    leakage), so absent reflections drift over an absolute E^2 cut at high
    resolution. That cost a tetragonal 42_12 crystal its 4_1: 18 of its 47
    absent 00l crossed the cut, all beyond 3.7 A, at absolute intensities
    identical to the low-resolution ones correctly judged absent, while
    their l=4n row-mates sat 20-60x higher at the same resolution.

A screw extinguishes only the reflections that lie ON its axis, so the
fair yardstick is the rest of that same row. The threshold is now
0.3 * max(1, median E^2 of the reflection's own row), the row being the
gcd-reduced reciprocal-space direction and the control class the same-row
reflections the group predicts present. Floored at 1, so it only ever
relaxes: a screw can be recovered by it, never lost.

Per row, not pooled. A 4_1 along c and a 2_1 along a are separate
conditions with separate controls; pooling let the weak a/b rows (median
E^2 ~0.5) set the threshold for a strong c row (8.4) and the rescue never
fired.

The candidate table now reports the screw evidence (median E^2 of the
absent class and of its rows) - the <I/s> columns are the centering
evidence and say nothing about screws, for the sigma-floor reason above.

Rotation battery, 33 crystals: 31 decisions bit-identical, the 42_12
crystal recovers its 4_1 (0 violations, row E^2 8.4 vs absent 0.12), and
one crystal with a long axis and heavy 00l overlap moves to a 4_1 group at
exactly 10.0% violations - marginal, and its sister crystal of the same
form sits at 13.3% and does not move. Real screws now span 0-9.3%
violations, so max_absent_violation_fraction cannot be tightened below
0.10 without risking a genuine one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rugnux: show the space-group search on rotation data
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
Build Packages / build:viewer-tgz:cpu (push) Successful in 7m12s
Build Packages / build:viewer-tgz:cuda (push) Successful in 8m47s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m27s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 13m54s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 13m58s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m16s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 14m20s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 12m53s
Build Packages / build:rpm (rocky8) (push) Successful in 11m27s
Build Packages / XDS test (durin plugin) (push) Successful in 7m49s
Build Packages / Generate python client (push) Successful in 34s
Build Packages / Build documentation (push) Successful in 1m6s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m49s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 13m23s
Build Packages / build:rpm (rocky9) (push) Successful in 13m51s
Build Packages / DIALS test (push) Successful in 14m5s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m12s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m7s
Build Packages / Unit tests (push) Successful in 1h0m25s
8cb1cacadf
The search runs in pass 1 of the rotation two-pass; pass 2 only reuses the
group it decided, so the result the CLI renders carried no search at all
and the whole report - operator correlations, the space-group candidate
table, the refused higher symmetry - was silently dropped on every
rotation dataset. Most costly of all, the "or <group> (indistinguishable
from these data)" line never appeared, so an enantiomorphic pair the
intensities genuinely cannot separate was reported as a single answer.
Carry pass 1's search into the returned result.

Also de-duplicate the alternatives when the centred-lattice test swaps the
metric-matching candidate into the answer: the group it displaced was left
out and the chosen one listed twice ("C2 or P21 or C2").

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Clear() reallocated sum and count but left sum2 as the constructor had
sized it. Every caller reuses one profile across the images of a dataset
(rugnux keeps one per worker, the viewer one per view), so the squares
kept accumulating while the means restarted at zero: the per-image
standard deviation written to /entry/azint and shown in the plots was
meaningless from the second image on, and grew without bound over a run.

The size mismatch was the sharper edge. Clearing to a mapping with a
different bin count left sum2 shorter than sum, and GetStd() and
operator+= then read past its end - reachable in the viewer by opening a
dataset with a wider q range than the one before it.

Both are covered by tests: the same frame twice with a Clear() in
between has to give the same standard deviation, and a profile cleared
to a wider mapping has to report the right value in a bin that only the
wider mapping has.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The mapping took its width and height from the CONVERTED geometry
unconditionally, while pixel_to_bin is sized per mode: converted when the
geometry is transformed, raw module layout when it is not. In raw mode
the two disagreed - 2068x2162 reported against a 1024x4096 map on a JF4M.

Only the adaptive spot finders read those dimensions, and they read them
for exactly the thing that breaks: the CPU finder derives npix = w*h and
then indexes the image, pixel_to_bin and the resolution mask with it, so
it walked ~277k pixels past the end of all three; the GPU finder stays in
bounds but decodes the strong-pixel bit index with the wrong row stride
and reports spots at wrong coordinates. Nothing combines raw geometry
with adaptive detection today, so this was latent rather than live.

Take them from GetXPixelsNum()/GetYPixelsNum(), which already follow the
geometry mode. The converted path is unchanged - it is the same number
there - and every internal use is inside SetupConvGeom, which only runs
when the geometry is transformed.

Covered by two tests: the mapping's dimensions must match pixel_to_bin in
both modes, and the CPU adaptive finder must return a spot planted on a
raw-geometry image at that raw pixel.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The two-arm search compared its arms by the order of the space group each
had picked, but Stage B leaves best_space_group unset whenever no
candidate is eligible - no candidate had enough observed absences to
trust. That is not rare on the Lorentz-filtered arm, and for a systematic
reason: the filter removes the badly-measured observations, which is
where the weak systematically-absent reflections are.

An arm that confirmed 422 but stopped short of naming a space group
therefore scored order 0 and lost to an arm supporting P2, and the
demotion was logged as "taking the higher symmetry" - the comparison and
the message both wrong, in the one direction the design says cannot
happen.

Carry the point-group order in the result, set from the order Stage A
actually adopted, and compare on that.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The two-arm check for "same order, different symmetry" ran after the
all-observation arm had already been adopted by value, so it compared
that arm against a copy of itself: the point-group names were always
equal and the branch was dead. In the other direction the orders were
always unequal, so it was dead there too. The AMBIGUOUS warning - written
for the case where the two merges support different symmetries of the
same order, which implies two different molecular-replacement searches -
could never be emitted.

Check before the adoption, while both arms still hold their own result.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Passing 0 reset the limit to "unset" and logged "No high resolution limit
for spot finding: as far as the detector reaches" - and then, 700 lines
later, the rotation default put 1.5 A back, because unset carried two
different requests: the user said nothing, or the user asked for none.
The log said one thing and detection did another, and there was no way to
lift the limit on rotation data at all.

Remember whether the option was given, and apply the rotation default
only when it was not. Documented in the usage message.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The expected-variance weights decompose an observation's sigma^2 into a
background part and a Poisson signal part, then rebuild the signal part
at the reflection's merged mean. The decomposition subtracted corr*I with
I taken as-is, so a negative I ADDED to the background part: an
observation at I = -1.5 with sigma^2 = 1 came out with a base variance of
2.7 rather than 1.

That inflates the variance of precisely the down-fluctuated observations
the correction exists for. Below about one photon they are then
under-weighted and the merged mean is biased high - the same direction of
error, in the same regime, that weighting by the observation's own sigma
produces. Subtract max(0, I) instead: a negative intensity has no Poisson
signal to remove.

Both users of the decomposition are fixed - the stills merge, where
expected-variance weighting is now the default, and the rotation combine
it was mirrored from, which had it first.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The first-pass spot cache constructed an MXAnalysisWithoutFPGA and an
AzimuthalIntegrationProfile inside the per-image lambda, so every cache
miss allocated a CUDA stream, the preprocessing buffer, the spot finder,
the azimuthal integrator and the Bragg engine, used them for one frame,
and freed them again - hundreds of times, serially, on the
--redo-rotation-spots path. Both worker loops already hoist the same
object out of their loop; only this path did not.

Build them once for the whole first pass, and only when spots actually
have to be found (with --reuse-rotation-spots there is nothing to
allocate). Reuse is safe because every azimuthal-integration path -
CPU, GPU and the fused adaptive engine - clears the caller's profile
before adding to it, so each frame's output is unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Making azim_int_settings.high_q_recipA and
spot_finding_settings.high_resolution_limit optional changes the wire
format: when unset they are omitted rather than sent with a placeholder,
and a client generated from an older spec does j.at() on them. The
azimuthal-integration limit now defaults to unset, so a stock broker
omits it out of the box - the break needs no operator action to hit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ten places where the document and the implementation had drifted apart.
Each was checked against the source before rewriting:

  * 7.5 rotation post-refinement: it is TWO separate cross-validated
    steps (cell+axis from the angles, then distance+beam from the
    positions with the cell fixed), not one joint fit against the merged
    fulls; the held-out split is an hkl hash, not a frame split; the
    bounds are +-5% on distance and +-15 px on the beam, not "under
    ~1%"; and only the distance and beam centre reach pass 2, which
    re-indexes de novo.
  * 9.2 the trimmed-mean background: it is computed in the shared
    background pass, so it DOES apply to --integrator boxsum. Only the
    broadband sigma-clip is excluded. The section previously said both,
    contradicting itself two paragraphs apart.
  * 9.3 per-reflection profile rebuild, sub-pixel centring and radial
    elongation are gaussian-only; the empirical profile keeps the fixed
    per-shell grid and is accumulated on rounded predicted positions,
    not centroids.
  * 10.5 the asymptotic ISa and the b_ISa sigma floor are rotation-only;
    stills report 1/b and floor with the whole-range b.
  * 8.4 centering absences are applied only when the user fixes the
    space group - de novo, prediction runs in P so the search can
    confirm the centering from the intensities.
  * 10.6 per-batch relative-B cross-validates on ASU-group parity, not
    the frame parity the other surfaces use.
  * 13 the resolution cutoff sits one reported-shell width PAST the
    CC1/2 = 0.30 crossing, so data below 0.30 are kept.
  * 7.1 the orientation-only prior penalises all three components of the
    angle-axis vector.
  * 10.1 rlp is the RECIPROCAL Lorentz factor, L = 1/rlp.
  * 14.3 the model scaling is fitted over work and free reflections
    alike, so R-free is free of refinement, not of the scaling fit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rugnux: drop the 1.5 A spot-finding limit on rotation data
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
Build Packages / build:viewer-tgz:cpu (push) Successful in 8m16s
Build Packages / build:viewer-tgz:cuda (push) Successful in 9m21s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m55s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 14m9s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 14m19s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m26s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 14m31s
Build Packages / build:rpm (rocky8) (push) Successful in 11m51s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 13m17s
Build Packages / XDS test (durin plugin) (push) Successful in 7m37s
Build Packages / Generate python client (push) Successful in 29s
Build Packages / Build documentation (push) Successful in 1m9s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m33s
Build Packages / build:rpm (rocky9) (push) Successful in 13m37s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 13m40s
Build Packages / DIALS test (push) Successful in 14m11s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m59s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m39s
Build Packages / Unit tests (push) Successful in 1h1m40s
0f1851cf82
Rotation kept a 1.5 A high-resolution limit for spot finding on the
strength of one indexing-rate measurement (100.0 -> 96.8% on a strong
set). Measured properly, over the whole 33-crystal rotation battery, it
does not earn its place:

  * no space-group decision changes - the same 30/33 agree with XDS, and
    the three that differ are the known pre-existing cases, unchanged;
  * 29 of 33 crystals are identical to the digit - same indexing rate,
    R_meas, CC1/2, ISa. The limit was doing nothing on the large
    majority;
  * where it does bite, the limit is the worse setting. The one crystal
    that loses appreciable indexing rate without it (99.50 -> 94.22%)
    comes back with lower R_meas (29.4 -> 28.1), higher high-resolution
    CC1/2 (27.9 -> 29.1) and higher ISa (5.77 -> 6.17). Another loses
    0.4% of frames and gains 2.8 points of CC1/2_hi. Fewer frames
    indexed, better data from them;
  * runtime is unchanged (16m46s vs 17m32s over the battery).

So the indexing-rate cost is real but does not carry through to the
merged data, which is what the limit was protecting. Unset now means "as
far as the detector reaches" for rotation as well as stills;
--spot-high-resolution still sets a limit for weak, high-background data
where the extra high-resolution spots are genuinely noise.

This also removes the flag that distinguished "the user asked for no
limit" from "the user said nothing" - with no rotation default left,
both mean the same thing. While rewriting the comment block, corrects
its neighbouring claim that rotation keeps the fixed-threshold finder;
adaptive detection has been the default for both workflows since
6f4917dce.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
docs: CPU_DATA_ANALYSIS describes the algorithms, not their history
Build Packages / Build documentation (push) Successful in 1m3s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m40s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 13m28s
Build Packages / build:rpm (rocky9) (push) Successful in 13m31s
Build Packages / DIALS test (push) Successful in 14m18s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m57s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m47s
Build Packages / build:viewer-tgz:cpu (push) Successful in 7m53s
Build Packages / build:viewer-tgz:cuda (push) Successful in 9m23s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 13m54s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m49s
Build Packages / Unit tests (push) Successful in 1h1m51s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 14m9s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m16s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 14m17s
Build Packages / build:rpm (rocky8) (push) Successful in 11m23s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 13m17s
Build Packages / XDS test (durin plugin) (push) Successful in 7m36s
Build Packages / Generate python client (push) Successful in 26s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
a3feb1271c
The document had accumulated development narrative: what was tried and
rejected, which datasets a change rescued or cost, measured percentages
from test batteries. That belongs in commit messages and reports, not in
a reference description of the pipeline - it dates quickly, and a reader
looking up what an algorithm does has to sort it out from how it came to
be.

Removed throughout, keeping the algorithmic content and the design
reasoning that explains a choice on principle:

  * 3.2 the whole paragraph justifying the rotation spot-finding limit
    from battery measurements, and the CPU-vs-GPU per-frame timings;
  * 3.3 "a significance/z-score was considered but is uninformative";
  * 7.4 / 7.5 the comparisons to a robust loss and to joint refinement
    as approaches that had failed;
  * 9.2 the R_meas / CC1/2 outcomes attributed to the trimmed-mean
    background;
  * 9.3 "per-detector-region and crystal-anisotropy profiles were
    evaluated and add nothing";
  * 10.2 the stills tilt "succeeds where a freely-fitted width
    collapses";
  * 10.5 the CC_anom argument, trimmed to why the statistic behaves as
    it does;
  * 10.6 the survey of per-frame correlation medians across datasets;
  * 13 the space-group bullet, restructured into the three gates it
    actually applies, dropping the dataset anecdotes;
  * 14.2 the free-form per-shell rescale, stated as a design choice
    rather than an experiment.

Section 13's space-group text was one 20-line paragraph; it is now a
numbered list of the three tests, which is what the code does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
6be94f2be stopped subtracting a negative intensity's Poisson term from the
background variance, but only in the host Combine(). The CUDA combine is the
path that actually runs: Run() selects it whenever a device is present and no
observation dump was asked for, so the correction never took effect on a normal
run, and a --dump-observations run merged differently from a normal one - the
two are meant to be identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The cap that refuses to report an impossible ISa was zeroing the asymptotic b
itself, and that same value is the floor passed to SigmaWithSystematicFloor -
where zero means "no floor". So on the degenerate low-multiplicity fit the guard
is written for, instead of capping merged I/sigma at 100 it removed the cap
entirely. Report the asymptote as unmeasured, keep the fitted value for the
floor.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A candidate can have several confirmed subgroups of the same order - 422 has
both 4 and 222 - and on a twinned crystal the rival is not a harmless
alternative: a P4 crystal twinned by 2[100] has 222 confirmed too, and 222
CONTAINS the twin laws, so its own merge b is already ballooned. The H test
already answers to every tied parent; the systematic-b veto and rescue took
whichever one the enumeration happened to list first (222 before 4, by space-
group number), which disabled the veto on exactly the case it exists for. Take
the smallest parent b, which is the conservative direction for both tests.

The refusal message also quoted the raw parent b rather than the floored value
the veto actually compared against.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The row-relative cut scales the "too strong to be absent" threshold by the axial
row's own median E^2, floored at the plain value - so it can only raise the bar,
and a row whose control class holds a single strong reflection sets it from that
one reflection. That direction invents screws: a genuine 4_2 whose 00l happen to
be observed only at l=4n reads its l=4n+2 reflections as absent and ranks
4_1/4_3 above the truth. Require three controls before the row may set the
scale; below that the row keeps the plain cut.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The two-arm search is meant to use the Lorentz-filtered merge for the point
group only - systematic absences live in the weak reflections a |zeta| cut
removes, and reading them off the filtered arm is what cost four crystals their
screw axes. That is what the code comment and CPU_DATA_ANALYSIS both say, but
the filtered-arm-wins branch kept its whole result, screws and centering
included.

Let a search be pinned to a point group decided elsewhere (fixed_point_group)
and re-run Stage B on the all-observation merge when the filtered arm rescues
the point group. The point group is passed as its symmorphic representative, not
by name: gemmi calls both P321 and P312 "32". Reporting that representative also
lets the ambiguity check see two arms that disagree about which 2-folds are real
- by name they looked identical - and the advice it prints now names a space
group -S can actually be given.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The two-pass rotation run reinstates pass-1's space group for pass 2's merge, so
pass 2 skips the search block entirely - and with it the flag that records that
the Laue class was CHOSEN by the search rather than given. The canonical output
therefore printed the plain "no twinning: the Laue class is holohedral, so no
merohedral twin law exists", which is exactly the circular conclusion the flag
was added to replace; only the throwaway _01 output carried the caveat.

The text is written per pass, inside RunPipeline, so the flag has to travel with
prepass_merge_sg_ rather than being patched onto the returned result.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every other reader of spot_finding.high_resolution_limit spells "unset" as
value_or(0) and compares, so 0 and nullopt are interchangeable - except in
SpotAnalyze, which passed the 0 straight to ResolutionShells and threw
"Resolution must be above zero" on every image. Reachable over the REST API,
where 0 is the natural way to say "no limit" and the settings check lets it
through; the rugnux CLI already maps 0 to unset before this point.

While here, check that a limit that IS set is finite regardless of its sign -
NaN fails the > 0 test and was skipping validation entirely.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Making the high limit optional removed the implicit upper bound on the low one
(it used to follow from high <= maxQ and high > low), so --azim-min-q 50 with no
maximum is accepted and ResolveHighQ then calls std::clamp with its lower bound
above its upper bound, which is undefined.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
calc_std uses the cancellation-prone (sum2 - sum^2/n) form on float accumulators
summed over millions of pixels, so a flat ring - true variance near zero - comes
out negative as often as positive and GetStd() returns NaN. Both adaptive
spot-finder ring accumulators already floor this at zero; this one did not, and
1a0774eed made sum2 correct, so the path is now actually exercised.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The host sized the grid from 2*max_hkl while the kernel guards against
2*max_hkl+1, so whenever the rounded-up grid landed exactly on 2*max_hkl threads
(max_hkl a multiple of 4, with the 8x8x8 block) the h = +max_hkl plane was never
launched. The CPU loop runs -max_hkl..+max_hkl inclusive, so the GPU predicted a
strict subset.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
b81c6f00b took the container depth from bit_depth_image in the file, which fixed
32-bit EIGER2 files but got the general case wrong: the reader converts every
image to SIGNED int32 (PixelSigned(true) a few lines up), while bit_depth_image
describes an unsigned container, and GetOverflow() combines the two. So a 16-bit
file still capped at INT16_MAX rather than 65535, and an 8-bit file newly capped
at 127 - flagging counts 127..254 as saturated, which drops the whole reflection
at the integration accept gate.

Declare 32 bits, matching what the reader actually hands out. The saturation cap
then comes from the file's own saturation_value, which is what it is for, and
the error value reported for a read dataset becomes INT32_MIN - the sentinel the
reader really uses. A file whose bit_depth_image is not 8/16/32 also stops
throwing on open.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The ring sigma is the cancelling difference sum2/n - m^2, and both sums were
float accumulated by atomics whose order is arbitrary. Two costs: the
cancellation left only ~4 digits in the variance, and the ordering moved the
resulting threshold by ~0.05 counts between runs - enough to flip a pixel
sitting on the hard "value >= threshold" test, and with it a connected
component's size. So the GPU engine did not reproduce the CPU one and did not
reproduce itself.

Only the accumulators that span blocks are widened. The per-block staging stays
float, because a block contributes a few dozen similar-magnitude pixels to a
ring and there is nothing to lose there - that also keeps the shared-memory
footprint of the hot loop, and hence its occupancy, exactly as it was: measured
on a 4.5 MP frame, 0.960 vs 0.966 ms/frame (40.9x over the CPU path, unchanged).
finalize_rings now does the cancellation in double and rounds to float last,
which is what AdaptiveSpotFinderCPU::AccumulateRings does.

The device properties are also read from the current device rather than device
0; callers round-robin engines across GPUs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The existing cases plant blobs at 200 on a background of 8..12, so any threshold
between 12 and 200 passes them - replacing RingThreshold with a constant leaves
them all green. Two cases that do not:

- the CPU threshold has to track the background: a frame and the same frame
  scaled ten times must give the same spots, with a pixel a few sigma above the
  background staying unfound in both. A constant threshold, or one that drops
  the sigma term, fails one scale or the other.
- the GPU engine has to agree with itself across runs, which is what the ring
  sums being order-independent buys.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two faults in the same block, both of which let a search pass corrupt the
production merge that follows it.

The device's corr was only copied back to the host for the diagnostic dump, but
the |zeta| filter runs on the host and then uploaded the whole host array - so on
a CUDA build it wrote the values ingested BEFORE scaling over the scaled and
smoothed corr the device had just computed. With the rotation default
--search-min-zeta 0.85 that means the space-group search was deciding the
symmetry from an unscaled merge. Copy corr back first, and upload once after
both filters instead.

Zeroing corr also has no owner: it is how an observation leaves the merge, but
the only thing that ever rewrites it is the scaling loop, which skips frames it
cannot fit. A frame left with too few well-measured reflections therefore kept
its dropped observations at zero for the rest of the object's life - and the
final production merge re-uses the same object without re-ingesting. Snapshot
corr before the filters and restore it at the start of the next pass, so each
pass decides for itself and the final merge keeps everything, as documented.

The frame rejection (--min-image-cc) is now applied on the host for both paths;
its separate device path did nothing whenever the CPU combine was in use.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
SolveScaleIRLS returns whatever it converged to and both writers accept any
G > 0, so a fit that collapsed to ~1e-3 multiplies that image's intensities by
a thousand. Nothing downstream notices, because the sigmas are multiplied by the
same factor and the merge's n-sigma outlier test is therefore blind to it - only
a total collapse self-heals, by overflowing corr to inf.

The rotation path refuses a per-frame scale this far below its neighbours; the
stills path had no guard. Judge each image against the median of the images that
did scale, and put a collapsed one back to G = 1 - the same state as an image
with too few reflections to fit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The Ceres summary was discarded, so a solve that diverged or aborted left its
last iterate in psi and that tilt was written onto the partiality of every
reflection of the crystal. Restore the tilt the crystal came in with and stop
refining it; the scale fit alone is still a usable model, which is what the
other three early returns in this function fall back to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
main had no enclosing try/catch, but plenty of ordinary input reaches a setter
that throws: --polarization 2, --detector-distance 0, -q 0, --azim-max-q 20,
--scale combined with a reference MTZ, and every failure inside the pipeline
itself. All of them ended as "terminate called after throwing an instance of
'JFJochException'" and exit 134, with the message nowhere to be seen. Move the
body into RunRugnux and let main report what was thrown, exit 1.

Two options also still bypassed the numeric parser that exists to prevent this:
--scaling-high-resolution used atof, which turns a typo into 0 and then throws
from the setter, and --integration-radius used std::stof, which throws on
non-numeric input.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
With merging on, the _process.h5 is skipped because the merged reflections are
the wanted output and that file is large (113 MB for 200 images here). But if
nothing indexes there are no merged reflections either, so the run finished
successfully having written no file at all - the one case where the user most
needs something to look at.

Write it in that case. The per-image messages have already gone past unwritten,
so this carries the dataset metadata, the mask, the azimuthal profile and the
summary scalars rather than the full per-image tables - and it is small for the
same reason it is needed (124 kB on a zero-index run). A run that does index is
untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The "d = ... A" readout moved from a scene item flagged
ItemIgnoresTransformations to a fixed viewport position painted in
drawForeground, which is what made a hover update dirty a small rect instead of
the whole viewport. But QGraphicsView pans by blitting the viewport: the painted
text is shifted along with the image and left there, and the pending update for
its old position is translated away too, so dragging the image smears ghost
copies of the readout across the corner. Dirty the old and the new rect when the
view scrolls - still a couple of hundred pixels, not the viewport.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
drawPixelLabels took its range from the whole viewport rather than from the
exposed rect it was given, so a 200x40 px hover repaint still walked up to 5000
cells doing mapFromScene + QImage::pixel + drawText for each, only to have the
result clipped away. With hover feedback now rate-limited to 15 Hz that ran ~75
times a second at high zoom, against once per overlay rebuild before the
rendering rework. Intersect with the exposed rect; same in the magnifier.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
centerAt returns early while the window is hidden, and nothing replays the last
position when it comes back, so re-opening the magnifier showed whatever region
the cursor was over when it was closed - with current pixels, which makes it
look like a live view of the wrong place. Remember the position while hidden and
apply it on show.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adaptive spot detection became the default for rotation data as well in
6f4917dce; RUGNUX.md still said rotation kept the fixed-threshold finder, which
is also the opposite of what the usage message and CPU_DATA_ANALYSIS say.
--search-min-zeta had no entry in the option tables at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The script now removes the generated C++ model, the frontend client and the
published python docs before regenerating, but not python-client/ itself - and
docs/python_client/docs is filled by copying that directory. So a schema dropped
from the API kept its generated model in the PyPI package and its .md page in
the published docs, linked from no index. Eight such pages are in the tree
today, JfjochSettingsSsl among them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
None of this has a reader:

- ScalingSettings::scaling_regularize and its setter/getter
- ScaleOnTheFlyResult::succesful (never set) and ::time_s (set, never read),
  with the timing that only fed the latter
- JFJochImage::last_fit_viewport_ (written twice, read nowhere) and the
  comment claiming the retry uses it - the retry keys off initial_fit_done_
- JFJochDiffractionImage::ice_ring_width_Q_recipA, and a QtConcurrent include
  in a file that uses none
- an unused gemmi::Op accumulator in the spindle-angle helper
- <random> in Merge.{h,cpp}, from before the half-set split became a hash
- an orphaned comment describing the Ceres B-factor residual deleted in
  014e43a4c, and two trailing comments that had collided on one line

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Eight pages under docs/python_client/docs describe schemas that appear nowhere
in jfjoch_api.yaml and are linked from no index - left behind because the
regeneration step never cleared python-client/, which these are copied from.
That is fixed in update_version.sh; this removes what accumulated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Also corrects the reader bit-depth entry: taking the depth from the file was
the wrong fix for the 32-bit EIGER2 case and broke 8-bit files.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The protection against a per-frame scale collapsing toward zero lived inside
ComputeSmoothGWindow, so it only existed when smooth-G did: --smooth-g=0, a
dataset whose oscillation width is unknown, and any caller that never sets a
smoothing range - the viewer among them - merged with no guard at all. A
collapsed G multiplies that frame's intensities by 1/G and its sigmas by the
same factor, so nothing downstream can see it; the merge's n-sigma cut scales
with the number that is wrong.

Pull it out into ReplaceCollapsedScales, called unconditionally right after the
partial scaling loop, and let the smooth-G window assume what it now guarantees
instead of computing its own median and floor.

The fulls guard built its median from every frame including those never fitted -
those sit at the combine's corr = 1, so a run with many unfitted frames dragged
the median toward 1 and the floor with it. It also reported the absolute
amplification where the message says "below the run median".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The error model says sigma -> b*I for strong reflections, so merged I/sigma
flattens off at 1/b - the number reported as ISa. Plotting I/sigma against I
with that asymptote drawn on it is what shows whether the reported ISa
describes the data or comes from a degenerate fit, which nothing in the window
could show before. A third page next to the per-shell plot and table.

The merge carries a few thousand strided (I, sigma) pairs to the viewer for it -
a shape, not a reflection list; the reflections themselves are in the .mtz/.cif.

The hero row gains the Wilson B and the radiation-damage Delta-B. Both were
already computed and already in MergeStatistics, so they only needed showing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
"Analyze dataset" and `rugnux` with no options are two front ends onto the same
library and are meant to agree, but they decided their defaults separately and
the two lists had drifted. The viewer was missing:

- the de-novo starting point. The CLI discards the cell and space group stored
  in the input file before it does anything; the viewer left them on the
  experiment. Rugnux only searches for a space group when none is set, so the
  search was skipped entirely and the stored group was reported straight back.
  That is self-reinforcing: a finished job's own _process.h5 becomes the active
  snapshot, so a run that ended in P1 pinned every later run to P1 - which is
  what "lysozyme keeps coming out P1 in the viewer" was.
- the polarization factor, so the Lp correction was omitted altogether. The
  missing factor is azimuthal and intensity-proportional, and symmetry mates sit
  at the same 2-theta but different azimuth - the exact "unequal intensities
  forced together" signature the space-group search vetoes as pseudo-symmetry,
  which can land a genuinely de-novo run in P1 on its own.
- five rotation scaling defaults: the smooth-G range, the minimum captured
  fraction, the capture-aware sigma, outlier rejection and --search-min-zeta.
  The CLI's own comments tie the captured-fraction default to a crystal
  recovering its true space group instead of P1.

Put the policy in one place (RugnuxDefaults) and have both front ends start from
it. The CLI now takes its defaults from there and applies user options on top;
its output is unchanged, verified bit-for-bit on four battery crystals.

The stored cell/group is still available: the job dialog offers "Use the stored
unit cell / space group", off by default, shown only when the file has one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Found while chasing a 12% run-to-run spread in the merged reflection count of one
crystal. The GPU kernels claim output slots with an atomicAdd and, on overflow,
undid the increment with an atomicSub - so the counter saturated at the capacity
and the host could not tell a full buffer from an overflowing one. Which
reflections survived was then decided by CUDA block scheduling and changed every
run. Measured on that dataset: every frame predicts 23000-44000 against a 20000
buffer, and the spread reached the merged output (161591 / 165193 / 166110 /
166479 unique across four runs of the same command). Single-threaded runs diverge
too - this is entirely GPU-side.

Stop clamping the counter, so the true number predicted reaches the host, and
warn once per predictor when it exceeds the buffer. Which reflections are kept is
unchanged: making that reproducible means deciding what to keep when a frame
predicts more than the pipeline carries, and the obvious answers are worse - the
capacity is not the real limit, kPredictionOutput (10000, selected by smallest
excitation error) is, and on this crystal both a bigger buffer and a strided
selection collapse the merge, because the rotation combine rebuilds fulls from
exactly the partials that a smallest-excitation-error cut throws away.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Digging into the selection logic showed the caps were not deciding the science -
the two-pass geometry post-refinement was, and the caps only fed it randomness.

Caps. The prediction buffer now grows to whatever a frame predicts instead of
keeping an arbitrary subset of it, and the per-image reflection limit is raised
to 65536, with the image-buffer transport headroom derived from the same
constant so the two cannot drift. Measured: bit-identical output on five battery
crystals, because a normal cell never approached the old limits - only a large
cell (~2.8e6 A^3, ~30000-44000 predictions per frame) ever did.

Pass-2 guard. The refined pass is normally the better answer, which is why it is
the canonical output, but it was adopted whatever it produced. On that same
crystal it merged more unique reflections than its own cell can hold -
completeness "117%", which is arithmetically impossible - while the header-
geometry pass sat at 92.6% and CC1/2 0.98. Compare the two and, when the refined
pass is not credible, go back to the header geometry and re-run so the canonical
files are the ones that are kept. Both bounds are set where only a failure
reaches them.

Together on that crystal: 111639 unique against XDS's 118730 (was 88000-99000
and different every run), CC1/2 98.0% (was 96.9-97.7%), ISa 8.54, and two runs
now agree bit for bit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The second pass re-indexes de novo and can land in a different setting from the
first - most often on the PRIMITIVE sub-cell of a centred lattice. The reindex
that exists to undo that declines when the metric does not match, and the code
then went on to stamp pass 1's group onto the cell regardless.

That is not a small error. A C-centred group on an already-primitive cell means
the centring absence rule removes half the reflections that genuinely exist, so
the merge holds more unique reflections than its own cell can - measured here as
"117% complete" with CC1/2 0.62, against the first pass's 92.6% and 0.98, on a
cell of exactly half the C-centred volume.

Detect the conflict where it happens and feed it to the pass-2 credibility guard
rather than acting on it locally: letting the pass re-search its own group
instead produced a P1 answer on a crystal XDS and the first pass both call C2,
which is a worse outcome than simply not trusting the pass. The completeness and
CC1/2 tests stay as the symptom-side net.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The per-image refinement stopped on a wall-clock budget (40 ms, and 20 ms for
the rotation-only extra pass). Online that is exactly right - the budget is real
and an image that overruns it costs the acquisition. Offline it means the same
file refines to a different lattice depending on what else the machine was doing
at the time, which is not a property reprocessing should have.

Bound it by iteration count instead when the caller is offline. IndexAndRefine
takes the workflow as a constructor argument: the receiver asks for the
wall-clock bound, rugnux and the viewer get the reproducible one. 50 iterations
is Ceres' own default; the per-image problem converges well inside it, so it
bounds the pathological case rather than the normal one - measured on five
battery crystals, every number is unchanged from the timed version.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Stills partiality: adopt the refined tilt only when it fits better
Build Packages / build:viewer-tgz:cpu (push) Successful in 8m17s
Build Packages / build:viewer-tgz:cuda (push) Successful in 9m26s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m43s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 13m59s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 14m4s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 14m18s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m24s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 12m49s
Build Packages / build:rpm (rocky8) (push) Successful in 12m8s
Build Packages / XDS test (durin plugin) (push) Successful in 8m54s
Build Packages / Generate python client (push) Successful in 37s
Build Packages / Build documentation (push) Successful in 1m4s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m29s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m51s
Build Packages / build:rpm (rocky9) (push) Successful in 13m41s
Build Packages / DIALS test (push) Successful in 14m33s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m43s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m34s
Build Packages / Unit tests (push) Successful in 1h1m35s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
3e56d96921
RefineOne re-measured the image's correlation to the reference after writing the
refined partialities - because --min-image-cc drops images by it - and then
ignored what it measured. A crystal the tilt model suits worse than the fixed
partiality it replaces kept the refined model anyway, and the refinement is on by
default. Compare against the CC the crystal arrived with and put it back
untouched when the refinement does not improve it, which is the same state a
crystal with too few reflections to fit ends in.

Also four things noted in review and left until now: AdaptiveThresholdTest.cpp
was listed twice in the test target, AdaptiveThreshold.h was the one header in
image_analysis/spot_finding not in its library's source list, CLAUDE.md said
update_version.sh rewrites VERSION when it only reads it, and the CHANGELOG did
not mention that image_scale_b is gone from the plot_type enum - which breaks a
client that asks for that plot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Raising the per-image reflection limit to 65536 for offline reprocessing also
raised the image-buffer headroom derived from it, and that headroom divides a
FIXED total buffer - so every slot grew from compressed+4 MB to compressed+16.7
MB and the receiver's slot count, i.e. how much of a burst it can absorb, fell by
about three. Online never needed the raised limit: measured on three serial
stills datasets the worst frame predicts 1380 reflections, 14% of even the old
cap.

So split them, the same way the geometry refinement's stopping rule is split:
online keeps the transport-sized 10000, offline gets the full 65536, and the
buffer headroom derives from the online one. Both still come from BraggPrediction
so the cap, the prediction and the headroom cannot drift apart.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ScaleOnTheFly's collapsed-scale guard ran, and then StillsPartialityRefine
re-fitted every crystal's scale with no floor and adopted it unconditionally
whenever the image had no prior CC - which is exactly the state the guard leaves
behind. So the guard was protecting almost nothing. Measured on a lysozyme jet
dataset: of 367 images it left unscaled, only 8 were still unscaled in the
output, and 53 reached the merge at or below a fiftieth of the run median, the
worst at a 4525th; on a second run of the same sample, 297 images, worst at a
75000th. Those intensities are what the merge saw - up to 94x too high in the
written file.

Run the guard again on the refined scales. It now reports 367 then 51 on that
dataset, and the merge improves: R-meas 117.4 -> 110.8%, CC1/2 96.3 -> 96.5%.

Measurements confirm the rest of the guard is right as it stands: 0.02 is ~5x
below the lowest scale ever seen on an image that correlates with the merge
(no image at CC >= 0.4 falls below a tenth of the median), and leaving the image
at G = 1 beats both dropping it and replacing its scale with the median - on a
run where 30% of images are affected, dropping costs 1.8 CC1/2 and 29%
multiplicity.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CUDA: let worker streams run concurrently
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 12m51s
Build Packages / build:rpm (rocky8) (push) Successful in 12m23s
Build Packages / XDS test (durin plugin) (push) Successful in 8m38s
Build Packages / Generate python client (push) Successful in 31s
Build Packages / Build documentation (push) Successful in 1m7s
Build Packages / Unit tests (push) Successful in 1h19m8s
Build Packages / Create release (push) Skipped
Build Packages / build:viewer-tgz:cpu (push) Successful in 6m53s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m25s
Build Packages / build:viewer-tgz:cuda (push) Successful in 8m31s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 13m9s
Build Packages / build:rpm (rocky9) (push) Successful in 13m46s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m47s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m3s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 14m23s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m19s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 14m25s
Build Packages / DIALS test (push) Successful in 14m46s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m56s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 14m56s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
2be8680422
Every per-thread stream was created with cudaStreamDefault, and the 20 MB raw
image upload went to the legacy NULL stream. A NULL-stream operation implicitly
synchronises with every blocking stream in the process, so with one engine per
worker thread no two workers' GPU work could ever overlap - the whole GPU
pipeline ran serially however many threads were asked for.

Create the streams non-blocking and put the upload on the engine's own stream.
Measured on 2000 serial stills, interleaved, medians of three: 24.6 -> 19.2 s at
-N 32 (-22%), 32.7 -> 21.0 s at -N 16 (-36%), CPU utilisation 436-570% -> 723-859%.
Output bit-identical - same observations, uniques, completeness, R-meas, CC1/2,
error model and cell. The stream is synchronised at the end of the same function,
so the ordering the code relies on is unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Stills geometry refinement: stop sampling once there are enough strong frames
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m18s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m12s
Build Packages / Generate python client (push) Successful in 15s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 10m55s
Build Packages / build:rpm (rocky9) (push) Successful in 12m56s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 50s
Build Packages / XDS test (durin plugin) (push) Successful in 8m2s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 11m0s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 7m12s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 10m23s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 11m35s
Build Packages / DIALS test (push) Successful in 12m50s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m16s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m11s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 11m52s
Build Packages / build:rpm (rocky8) (push) Successful in 11m24s
Build Packages / build:viewer-tgz:cpu (push) Successful in 6m43s
Build Packages / build:viewer-tgz:cuda (push) Successful in 7m44s
Build Packages / Unit tests (push) Successful in 1h2m10s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
7786fc1af3
The first pass sampled min(n, max(refine_frames * 50, 8000)) images to keep the
200 strongest, so any serial run of 10000 frames or fewer indexed every frame
TWICE - and 99.3% of the pass was that sampling, the bundle adjust itself taking
0.24 s. The budget is sized for a low-hit-rate dataset; on data that indexes well
almost all of it was wasted.

Stop once four times the bundle size has been found, which still leaves the
"strongest N" selection a real pool and still spans the run, because the sample
is equally spaced. On a lysozyme jet dataset that is 1500 frames examined instead
of 4000, the same 200 bundled, and the same refined geometry - beam and distance
to the pixel, cell to 0.01 A. Warm cache: the pass drops 33 s -> 9.1 s and the
whole run 65 s -> 35 s, with CC1/2 and R-meas unchanged inside replicate noise.

Stills only - rotation has its own two-pass and returns from this function early.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The kernel guards against 2*max_hkl+1 and maps thread i to h = i - max_hkl, but
the host launched a grid sized 2*max_hkl. The h = k = l = +max_hkl planes were
therefore never launched while -max_hkl was, so the GPU predicted an asymmetric
subset of what the CPU loop (inclusive on both ends) does. The same bug was fixed
on the stills twin when the whole hkl range moved to the GPU; the rotation
predictor kept the old expression.

It only bites where the cell actually reaches |h| = 100 inside d_min - a ~150 A
axis at 1.5 A - so most data never noticed. Over the 33-crystal rotation battery
29 crystals are bit-identical and 4 gain observations, all of them large-cell or
high-resolution: +8519, +4693, +901 and +758 observations, with the high-shell
CC1/2 up 15.0->15.1%, 52.0->52.2%, 76.3->76.6% and 51.6->52.1%. Nothing is lost
anywhere, and R-meas and ISa move by at most 0.01.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
--polarization was applied with the other geometry overrides, but
configure_offline_output runs afterwards and calls ApplyRugnuxExperimentDefaults,
which sets the polarization factor unconditionally. Every full-analysis run used
0.99 whatever was asked for, so the Lp correction was wrong at a beamline with
different polarization. Apply it after the defaults instead, and stop claiming in
RugnuxDefaults.h that nothing here is user-selectable.

--scale built a bare ScalingSettings and re-derived the rotation/stills split by
hand rather than calling RugnuxDefaultScalingSettings, which is what the split was
factored out for. It got scale-fulls, smooth-G, min-captured-fraction and outlier
rejection right and dropped CaptureUncertaintyCoeff on the floor: 1.0 in the
pipeline, 0.0 here. So re-scaling a rotation _process.h5 gave different sigmas and
ISa than the run that wrote it - the exact failure the block's own comment says it
exists to prevent. Start from the shared defaults and apply the overrides on top,
which also picks up --mosaicity and --search-min-zeta, and let -C bind here too.

REJECT_OUTLIERS_DEFAULT_NSIGMA had no reader left afterwards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
"Analyze dataset" cleared the stored cell and space group unless "Use the stored
unit cell / space group" was ticked, and that checkbox defaulted off. But the
settings panel writes the user's own cell and space group onto the experiment, so
a cell typed into the panel was discarded too - while the checkbox label said
"stored", implying it came from the file.

It also contradicted the dialog next to it: "Refine geometry (stills)" is offered
and default-ticked precisely because a cell is present, and the run then removed
that cell. The default dialog state on a stills dataset with a known cell ran the
bundle adjustment with nothing to anchor on and dropped indexing off ffbidx, which
needs a cell, onto de-novo FFT.

Drop the checkbox and take the crystal from the panel, which already has exactly
the right semantics: "Unit cell known" ticked writes the cell and group, unticked
clears both, and a space group of 0 means none. So ticked = -C/-S, unticked =
bare rugnux, and what a run will use is always what is on screen. That also keeps
the copied command line honest, since RugnuxCommandLine emits -C/-S from the same
experiment. The panel is refilled from the file when one is opened, so a finished
job's _process.h5 becoming the active snapshot now shows its group and can be
cleared, instead of silently pinning every later run to it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The "keep what the crystal came in with" gate required std::isfinite(cc) before it
would reject, so a refined model whose CC could not be measured at all was adopted.
ImageReferenceCC returns NaN when fewer than 20 reflections clear the partiality
cut - which is exactly what a refinement that collapsed the partialities produces,
since the cut is on the partialities it just rewrote. The gate therefore failed
open on precisely the crystals it exists to catch, and wrote the NaN into
image_scale_cc, on which --min-image-cc then drops the image from the merge, the
error model and the statistics.

Treat a CC that cannot be measured as worse than one that can, so the crystal is
put back exactly as it arrived.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Leaving it at G = 1 looked like the conservative choice and is the more damaging
of the two errors. The per-image scale enters as rlp/(partiality*G) and multiplies
intensity and sigma alike, so substituting 1 for a scale that was really 1/200 of
the run median puts the intensities in 200x too low with sigmas 200x too low too -
1/G^2 times the weight they deserve. The merge cannot defend itself against that,
because the number that is wrong is the number the weight is built from. And if
the collapsed value was instead a failed fit, G = 1 merges the image mis-scaled by
an unknown factor. Per-crystal scales on serial stills genuinely span orders of
magnitude, unlike frames of one rotation sweep, so both readings are live.

An image whose scale is not believable has no usable scale. Write NaN into its
image_scale_corr, which every merge path already skips on, so it drops out of the
merged intensities, the error model and the statistics consistently.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Making the worker streams non-blocking removed the implicit ordering that the
constructors were still relying on. Each engine uploads its static inputs - the
pixel mask, the pixel-to-bin map, the corrections, the ROI map - with a blocking
NULL-stream cudaMemcpy, and then reads them from kernels on its own stream. A
pageable host-to-device cudaMemcpy returns once the source has been staged, with
the DMA still in flight, and a non-blocking stream no longer waits for the NULL
stream. The failure mode is a silently unapplied mask or a stale mapping, not a
crash, so it would not have announced itself.

Put them on the stream the engine already owns, and synchronise once at the end of
the constructor - that is required for the preprocessor, whose source is a local
vector, and leaves the others settled rather than in flight for the cost of one
one-time sync. The GPU spot-finder test uploaded its image the same way.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Dropping the fixed spot-finding limit left the reader still generating the
spot-vs-resolution plot over shells that stop at 1.5 A, so a stored file reopened
in the viewer showed a plot truncated at exactly the limit that was removed -
GenerateSpotPlot drops every spot outside its shells. Pass the detector's own
maximum resolution, as SpotAnalyze already does.

That value is 0 when the geometry gives no scattering angle at all (no distance or
no wavelength), and ResolutionShells throws on a non-positive d_min, once per
image. There is no resolution axis to plot against in that case, so skip the plot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
max_hkl was hardcoded to 100 at the one place production builds the prediction
settings, so the only way to change it was to edit and rebuild - and it is not a
constant of the method, it is a property of the cell. An axis is truncated once
a/d_min exceeds it: 100 covers a 150 A axis at 1.5 A, but the same axis at 1.0 A,
or a 250 A axis anywhere, loses its outermost reflections with nothing said.

Move it into BraggIntegrationSettings next to the other prediction/integration
parameters and add rugnux --max-hkl (1..511, default 100 - no behaviour change).
Like the integration radii and the background trim it stays out of the OpenAPI, so
the broker keeps the default it has today and live analysis cannot be handed a
range that would not finish; the offline front end, which knows its cell, can ask
for more. RugnuxCommandLine emits it when it is not the default.

Measured on five rotation crystals at --max-hkl 200: two are bit-identical at no
cost, and three were being truncated - one gains 419k observations (+17%) and
takes its high-shell CC1/2 from 15.1% to 25.8% for +14% wall clock, the other two
gain 12k and 5.8k observations with CC1/2 76.6->82.4% and 52.1->55.3% for +9% and
+1%. ISa is unchanged throughout, and no frame overflowed the prediction buffer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Bragg prediction: derive the lattice walk from the cell, and expose it in the API
Build Packages / build:viewer-tgz:cpu (push) Successful in 7m15s
Build Packages / build:viewer-tgz:cuda (push) Successful in 8m42s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m47s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 14m1s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 14m11s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m19s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 14m31s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 12m39s
Build Packages / build:rpm (rocky8) (push) Successful in 11m41s
Build Packages / XDS test (durin plugin) (push) Successful in 7m52s
Build Packages / Generate python client (push) Successful in 34s
Build Packages / Build documentation (push) Successful in 1m3s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m10s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 13m0s
Build Packages / build:rpm (rocky9) (push) Successful in 13m39s
Build Packages / DIALS test (push) Successful in 14m29s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m33s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m25s
Build Packages / Unit tests (push) Successful in 1h36m21s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
b0e315e73c
Follow-up to making max_hkl a setting: it is now an optional, and unset means "take
it from this crystal". The predictor keeps only |q| <= 1/d_min and h = a.q for the
real-space axis a, so |h| <= a/d_min exactly - and likewise |k| <= b/d_min and
|l| <= c/d_min. max(a,b,c)/d_min therefore bounds all three at once: nothing that
could be predicted lies outside it, and nothing inside it is reached by a shorter
axis. It applies to rotation and stills alike, both going through the one place the
prediction settings are built.

Offline (rugnux, viewer) the default is unset, so every crystal gets its own range;
--max-hkl overrides it. Online the broker holds a concrete number, because the cost
is the cube of it per image and a live acquisition should not have its frame rate
decided by whichever sample is mounted: max_hkl joins bragg_integration_settings in
the OpenAPI with a default of 100, so an omitted field arrives as that default (the
generated model carries it) rather than as "derive it", and the frontend exposes it
next to the integration model.

Measured against a fixed 100 on six rotation crystals: three are bit-identical, two
were being truncated and recover 419k and 5.8k observations with the high-shell
CC1/2 going 15.1 -> 25.8% and 52.1 -> 55.3%, and the space group is unchanged 6/6.
It reproduces a fixed 200 exactly, which is the bound being tight rather than merely
safe.

The sixth is worth recording: a 149/83/226 A cell derives 227, and because a single
scalar has to cover the longest axis the cube is ~16x what a per-axis box would be -
22% wall clock, for a net 22 observations out of 364k (the per-frame 65536-reflection
cap re-selects at the margin when more candidates are offered) and identical CC1/2,
ISa and space group. Per-axis limits would remove that; the predictors already map a
thread index to h, k and l separately.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The working tree accumulates merged reflection files, models and per-image dumps
while testing, and they sit in the repository root next to the source. They are
user data: a merged .mtz/.cif carries a sample's measured unit cell and its
filename usually carries the sample's name, neither of which may enter this
repository. Only build*/ and python-client/ were ignored, so a `git add -A` would
have picked all of it up - which is exactly what happened while preparing this
branch, caught before the commit was made.

Ignore the file types rather than rely on everyone typing the right paths, and
un-ignore tests/ so checked-in fixtures still work (git add -f for anything else
that genuinely belongs). Also widen the rugnux_vs_xds.py output dir to rugnux_cmp*/,
which is where the ad-hoc comparison runs land.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Bragg prediction: one limit per index, not one cube
Build Packages / build:viewer-tgz:cpu (push) Successful in 6m47s
Build Packages / build:viewer-tgz:cuda (push) Successful in 6m42s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 10m46s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 10m35s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 11m1s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 11m37s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m1s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 11m59s
Build Packages / build:rpm (rocky8) (push) Successful in 12m3s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m9s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m52s
Build Packages / Generate python client (push) Successful in 15s
Build Packages / build:rpm (rocky9) (push) Successful in 13m31s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m12s
Build Packages / XDS test (durin plugin) (push) Successful in 9m3s
Build Packages / DIALS test (push) Successful in 12m48s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 7m46s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m39s
Build Packages / Unit tests (push) Successful in 1h16m34s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
406c406988
Each Miller index is bounded by its OWN axis - |h| <= a/d_min, |k| <= b/d_min,
|l| <= c/d_min - so a single half-width has to be sized for the longest axis and
then walks the short ones far past anything the resolution cut can keep. Give the
predictor max_h, max_k and max_l instead, in all four implementations (CPU and GPU,
stills and rotation), and derive each from its own axis.

On a 149/83/226 A cell that is 23.1M candidates per frame instead of 94.2M, 4.1x
fewer. Results are bit-identical, as they must be - the candidates removed are only
ones the |q| <= 1/d_min cut rejected anyway: over six rotation crystals every merged
observation count, high-shell CC1/2 and space group matches the cube exactly, 6/6
space groups correct.

It buys almost no time, and the earlier claim that the cube cost 22% of that
crystal's wall clock was wrong. Removing 4.1x of the candidates moves it 1m58s ->
1m57s, so the whole prediction sweep is ~1% of the run. The 22% that crystal costs
relative to a fixed max_hkl of 100 is genuine extra work at max_l = 227: real
reflections inside the resolution sphere along the long axis, predicted and
integrated either way. Per-axis limits do not reduce that and cannot.

The user-facing setting stays a single number: it exists to bound the work, not to
describe the crystal, and applies to all three indices when set.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Bragg integration: integrate as far as the detector reaches, not to a fixed 1.0 A
Build Packages / build:viewer-tgz:cpu (push) Successful in 7m28s
Build Packages / build:viewer-tgz:cuda (push) Successful in 7m50s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 10m44s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 10m30s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 10m9s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 9m5s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 10m1s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 11m39s
Build Packages / build:rpm (rocky8) (push) Successful in 10m52s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 10m50s
Build Packages / build:rpm (rocky9) (push) Successful in 11m45s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m11s
Build Packages / Generate python client (push) Successful in 16s
Build Packages / Build documentation (push) Successful in 53s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (durin plugin) (push) Successful in 7m16s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 7m22s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m56s
Build Packages / DIALS test (push) Successful in 10m58s
Build Packages / Unit tests (push) Successful in 1h2m58s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
0ca159449f
BraggIntegrationSettings::DMinLimit_A had a setter that nothing anywhere called, so
it was always its 1.0 A default - in rugnux, the viewer and the broker alike, with
no option or API field to change it. It feeds the predictor as high_res_A, which
discards any reflection with |q| > 1/d_min, so integration simply stopped at 1.0 A
however far the detector reached.

Five of the 33 rotation test datasets have detectors reaching past it, down to
0.981 A. On one of them, run with no resolution limit, the shell table ended dead
at 1.00 A with that shell still at CC1/2 55.6% and <I/sig> 3.4 - cut mid-shell
rather than fading out. This branch had already made the sibling limits
detector-driven (spot finding, scaling), so the pipeline was finding spots the
detector could see and then refusing to integrate them.

Make it a std::optional: unset means as far as the detector reaches, a value limits.
The limit is only a bound on how far the lattice walk goes, never a second opinion
on what is measurable - both predictors independently drop reflections that miss the
detector (BraggPrediction.cpp, BraggPredictionRot.cpp) - which is what makes the
detector's own reach the right default. rugnux gains --integration-high-resolution
(0 = no limit, as for --spot-high-resolution); the derived per-axis prediction range
resolves against the same number, so the two cannot drift.

Full battery: 30/33 space groups, unchanged from before, 0 failures and the same
three known mismatches; 22 of 32 crystals bit-identical and nothing worse than 5
observations in ~500k. The datasets that gain do so because their detector reached
past 1.0 A - the effect is understated here because the harness caps each merge at
the XDS resolution anyway.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The per-image ice ring score was encoded for every DataMessage but never
for the END message, although docs/CBOR.md has always listed it there and
both NXmx::EndResultVectors and HDF5MetadataSource expect it. Any dataset
written over the stream therefore had no /entry/MX/iceRingScore, and under
NXmxIntegrated - where the whole-run vector is the only copy - the score
was lost entirely.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
SendZeroCopy closes the message when zmq_msg_send fails, and closing a
message built with zmq_msg_init_data runs its free function - here
zmq_socket_free, which already calls release(). The writer thread then
released the same slot a second time, under a comment claiming the
callback would not run.

The second release put a slot back on the free list while the receiver
had already taken it for the next image, so two threads wrote the same
buffer and the sending/preparation counters drifted permanently.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
release() published the handle through ReleaseSlot and only then wrote
status = InPreparation. ReleaseSlot makes the handle available to
GetImageSlot immediately, and GetImageSlot hands back this very object
without resetting it, so the receiver could observe the stale Sending
status and throw "Trying to send image that is not in preparation",
aborting the collection - or take the opposite interleaving and leak the
slot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
International Tables A 3.1.3.1 gives 100 / -110 / -1-13 for character 9;
the last element was -3. With a negative determinant the transform is
left-handed and the "conventional" rhombohedral cell is not hexagonal -
beta came out around 110-134 degrees instead of 90 and c was far too
long. Any R lattice tall enough to reduce to character 9 was affected,
and the downstream Trigonal->Hexagonal promotion then forced 90/90/120
onto that wrong cell.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Character 40 carried a verbatim copy of character 35's matrix
(0-10 / -100 / 00-1), whose determinant is 1. A C-centred conventional
cell needs determinant 2, so a genuine oC lattice was returned as its
primitive monoclinic cell while still being labelled Orthorhombic 'C':
the refiner then clamped a ~117 degree beta to 90 and prediction dropped
half the reflections of a cell that has no centring.

International Tables A 3.1.3.1 gives 0-10 / 012 / -100 for character 40.
Character 35 is correct as it stands and is left alone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rowsPerWave is rounded up, so with 32 waves the last waves can start at
or past the last row: rmin was never clamped and only the drain loop
checked front against height. On any detector below about 1500 rows -
including the module-converted 500K and 1M geometries and the kernel's
own unit tests - the priming and steady-state loops read whole rows past
the end of the image buffer, and those garbage rows entered the sliding
background window of the bottom rows.

Blocks with no rows to write now return before the first __syncthreads
(rmin depends only on blockIdx.y, so the block leaves together and the
collective ops stay well formed), and both remaining reads are bounded by
height. Rows past the end keep the INT32_MIN sentinel, which the window
already treats as "not counted".

The raw read in the steady-state loop is left as it is: making it apply
the prev_out substitution that the other two read sites use would change
which pixels are found, which is a separate question from this fix.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The processing VDS mapping is built from the number of images a run set
out to process, while total_images comes from the end message and is the
number it actually finished. Those differ whenever a run is cancelled or
skips an unreadable frame, and the mismatch was a hard throw - which
NXmx::Finalize catches by deleting the temporary master, so rugnux lost
the entire _process.h5 and every completed image with it. Ctrl-C after
5000 of 100000 images produced no output file at all.

Map what was written and drop the remainder instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Deactivate() holds m for the whole power-off sequence, which is right -
nothing else should touch the detector while it is being turned off - but
it had no state check, unlike every other entry point. Called during a
measurement, calibration or initialisation it waited on measurement.get()
while holding m, and those threads re-acquire m to finish: a deadlock
that wedged every endpoint, /cancel included.

IsRunning() is exactly the set of states with a live background thread,
so deactivating stays possible from Error - otherwise a failed initialise
would leave no way to power the detector down.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
JFJochDecompressHperfPtr took source_size and never looked at it: every
per-block length was read out of the stream and passed straight to
LZ4_decompress_safe/ZSTD_decompress as the source length, with src_ptr
advanced by it. The only check happened after the whole buffer had
already been walked. A truncated frame, or a block header claiming
0x7fffffff, read far past the end of a heap buffer - reachable from the
ZeroMQ CBOR path and from any HDF5 chunk the XDS plugin is handed.

block_size == 0 satisfied the "% BSHUF_BLOCKED_MULT" test and then
divided at nelements / block_size, and source_size < 12 underflowed
source_size - 12 to about 2^64. Both are now rejected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This was the only panel firing the generated call directly, with the
rejection routed to console.log. Because the poll then kept returning the
unchanged server value - which still equalled lastDownloadedS - the
resync effect never fired, so the dashboard showed a threshold the broker
was not using, indefinitely and silently. It now goes through useUpload
like its siblings, so a failure raises the snackbar, and onError puts the
server's value back in the panel.

The eight sliders also applied from onChange, which MUI fires for every
intermediate position while dragging: one drag across the ice-ring width
sent ~200 PUTs plus ~200 forced /statistics refetches, each reconfiguring
spot finding on the running acquisition. Dragging now only moves the
panel; the value is sent from onChangeCommitted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The GPU merge kernel rejects outliers on the device and keeps a per-full
flag there, but only returned the per-group counts. The host array the
CPU path fills stayed all zero, and the anomalous I(+)/I(-) accumulator
is host-side and unconditional - so with --reject-outliers and a GPU
present, the observations the merged IMEAN dropped were still averaged
into I(+) and I(-). The same command on a CPU-only host excluded them:
the exported anomalous differences depended on whether a GPU was there.

R_meas was unaffected, having its own device-side path that reads the
flags in place. MergeAccum now hands the per-full flags back so every
host-side reduction sees the same rejections. The comment claiming
reject_outliers was excluded from the GPU path was never true.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
DrawSaturation walked the whole saturated set adding two QGraphicsLineItems
each, with none of the viewport culling DrawSpots and DrawPredictions do
directly above it, and no upper bound - and it is rebuilt on every pan,
zoom and frame change. Unlike spots, that set is not bounded by a setting:
an over-exposed frame or a missing beamstop saturates a large fraction of
the detector, which meant hundreds of thousands of scene items and a
multi-second freeze on each mouse drag.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The queue-level fix for the live-follow OOM bounded how many datasets are
in flight, but not what each tick costs. Three handlers did full-dataset or
full-detector work per tick regardless of whether their window was open:

- the calibration window copied the whole pixel mask (GetMask returns a
  reference; it was taken by value), memcpy'd it and ran a full-resolution
  recolour on the GUI thread;
- the image-list window rebuilt one row of eight QStandardItems per image,
  and then repainted every cell of the model on every frame to move a
  one-row highlight;
- the dataset-info plot was rebuilt twice per tick, because setCurrentIndex
  fires currentIndexChanged -> comboBoxSelected -> UpdatePlot and the
  caller then called UpdatePlot again.

The first two now defer to showEvent while hidden, following the pattern
JFJochViewerReciprocalSpaceWindow::rebuildGL already uses; the highlight
repaints only the two rows that change; and the combo is blocked around
setCurrentIndex so the plot is built once.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Verified section by section. The corrections that matter:

- image_analysis/pixel_refinement/ is gone (38ea0ec23). The style rule it
  anchored - no defensive or unrequested code - stays, now attached to the
  experimental analysis code generally.
- rugnux_cli.cpp lives in rugnux/, not tools/ (f737424bd).
- There is no .clang-tidy in the tree and never has been; the naming
  conventions are kept, described as what the code already does.
- compression/ has no sqrt codec - the algorithms are BSHUF_LZ4 and the
  three BSHUF_ZSTD variants. The square-root transform is an FPGA pipeline
  stage.
- jfjoch_hdf5_test is defined in tools/, not tests/.
- JFJOCH_VIEWER_ONLY was undocumented, and is forced ON on Windows/macOS.
- The per-image-scalar recipe pointed at reader/JFJochHttpReader.cpp, which
  does not exist (it is viewer/), and missed that EndMessage carries both a
  per-image vector and a run-mean scalar, so the CBOR END block needs two
  keys. Added the camelCase-dataset vs snake_case-field trap.
- The portability notes described work already done (libjpeg-turbo) and
  recommended fetching Eigen, which CMakeLists explicitly rules out.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four of the seven ceres::Solve calls in image_analysis obtained a
Solver::Summary and never looked at it, so a solve that failed numerically
had its parameters written back and was reported as success.
StillsPartialityRefine and both PostRefine solves already gated on
IsSolutionUsable(); this brings the rest to the same contract.

IsSolutionUsable() is the right test rather than checking for CONVERGENCE:
it accepts a solve that ran out of iterations or wall-clock time but still
descended, which is exactly what the real-time callers depend on when they
set max_solver_time instead of max_num_iterations. Only FAILURE and
USER_FAILURE are rejected.

XtalOptimizer checks before the write-back, so a failed refinement now
leaves the caller's geom and latt untouched instead of half-updated.
GeometryRefiner folds it into result.ok, which previously reported success
from spot and frame counts alone. RingOptimizer returns a geometry by
value that both callers assign straight back over their input, so it hands
back the unchanged reference rather than a diverged beam centre.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The azimuthal bin limit is FPGA_INTEGRATION_BIN_COUNT = 2048, not 1024.
There are 16 ROIs, not 64, and the map is a 16-bit per-pixel mask, so a
pixel belongs to any subset of them rather than to exactly one.

The lossy transform is round(sqrt(N*N*X)) = round(N*sqrt(X)): the HLS
squares the sqrtmult register before multiplying. The doc said sqrt(N*X),
which is off by sqrt(N), and the register comment claimed the value was
"minus one" and "should be square of the coeff" - both wrong, the host
writes N and the FPGA squares it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
BraggPrediction.h claimed the buffer "GROWS to whatever a frame actually
predicts, so a large cell is never truncated here". Only the two GPU Calc
overrides call GrowCapacity; both CPU predictors stop at max_reflections.
The cap is applied inside the h/k/l walk and before the resolution test,
so what survives is the low-|h| block, not the reflections nearest the
Ewald sphere - a cell large enough to overflow 20000 gives different
merged reflections with and without a GPU. Documented rather than
silently claimed otherwise.

Also removed a paragraph describing a once-per-predictor overflow warning
that no longer exists, and fixed the rugnux_cli.cpp path in HDF5.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The declaration still said "stills on, rotation off". Adaptive detection
has since been turned on for both workflows - adaptive_spots.value_or(true)
- which the comment at the assignment already explains.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Deactivate() called measurement.get() inside the try that guards the
power-off, so an exception stored by a previously failed run was rethrown
before services.Off() ever ran: the detector stayed powered while the
state reported Error, and the operator had no way to turn it off. The
future is still reaped - it has to be - but its failure is logged and
dropped. It was already reported when it happened, and leaving the
detector on is the worse outcome.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The header has said since it was written that blocking queue operations
must never run under connections_mutex; three code paths did exactly that.
KeepaliveThread held it while sending a keepalive to every connection,
which blocks until the peer-liveness or backpressure timeout - so one
half-dead writer socket could stall SendImage and every /statistics poll
for up to a minute, from an idle-time heartbeat. AcceptorThread and
StartDataCollection held it across RemoveDeadConnections, which joins a
writer thread that may itself be inside such a send.

RemoveDeadConnections is split in two: DetachDeadConnections unlinks them
from the pool under the mutex, which is quick, and CloseDeadConnections
tears them down afterwards with the mutex released - safe because they are
no longer reachable by anyone else. The keepalive loop copies the pool out
and sends outside the lock, the pattern EndDataCollection already used.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
spot_finding: run the same two passes on the CPU as on the GPU
Build Packages / build:viewer-tgz:cpu (push) Successful in 7m50s
Build Packages / build:viewer-tgz:cuda (push) Successful in 8m38s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m32s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 14m17s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 14m21s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 14m27s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m39s
Build Packages / build:rpm (rocky8) (push) Successful in 11m59s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 13m8s
Build Packages / XDS test (durin plugin) (push) Successful in 7m15s
Build Packages / Generate python client (push) Successful in 24s
Build Packages / Build documentation (push) Successful in 1m5s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (rocky9) (push) Successful in 12m28s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m50s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 13m18s
Build Packages / DIALS test (push) Successful in 14m17s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m9s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m52s
Build Packages / Unit tests (push) Successful in 59m1s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
1a1e05ad14
ImageSpotFinderGPU::Detect launches its kernel twice, feeding the first
pass's strong-pixel bitmap back in so the second recomputes each local
background with those pixels excluded and keeps them strong. The CPU
finder ran a single pass, so the two returned different spot lists for the
same frame and a dataset processed without a GPU did not match one
processed with it.

It matters for any spot wide enough to reach into its own 31x31 background
box: the spot inflates the mean and variance it is then tested against, so
its outer pixels fail the SNR test. On the test image added here - a 5x5
core at 300 counts with a one-pixel ring at 25 - a single pass returns the
25-pixel core and 7500 counts where two passes return the full 49 pixels
and 8100.

pxl_val also becomes int64_t, matching the GPU's pixel_result signature.
It was int32_t, so pxl_val * pxl_val overflowed above 46341 counts even
though the surrounding sums were already 64-bit.

The new parity test compares PixelCount and Count, not just the centroid,
which does not move for a symmetric spot whether or not the ring was
picked up; it was confirmed to fail against the old single-pass CPU.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
image_analysis: share the read-only GPU lookup tables per device
Build Packages / build:viewer-tgz:cpu (push) Successful in 8m32s
Build Packages / build:viewer-tgz:cuda (push) Successful in 10m13s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 12m56s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 13m51s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 14m2s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m25s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 15m1s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 13m2s
Build Packages / build:rpm (rocky8) (push) Successful in 12m55s
Build Packages / XDS test (durin plugin) (push) Successful in 9m41s
Build Packages / Generate python client (push) Successful in 28s
Build Packages / Build documentation (push) Successful in 47s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m13s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m35s
Build Packages / build:rpm (rocky9) (push) Successful in 13m38s
Build Packages / DIALS test (push) Successful in 13m57s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m21s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m49s
Build Packages / Unit tests (push) Successful in 1h1m3s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
0b1fb6c870
One analysis engine is built per worker thread, and each uploaded its own copy of
tables that are pure functions of the detector geometry: the pixel -> azimuthal bin
map and the per-pixel corrections (both in AzIntEngineGPU AND again in
AdaptiveSpotFinderGPU, from the same mapping), plus the pixel mask. On an 18 Mpx
detector that is ~224 MB per worker; with 32 workers ~7 GB of device memory held 32
identical copies.

Upload each table once per GPU instead and hand every engine on that device a shared
pointer to it. The cache is keyed by (device, source-vector address) because workers
are pinned round-robin across GPUs, so on a multi-GPU node each device keeps its own
copy - a kernel may only read memory resident on the device it runs on - and the
table is freed on the device that allocated it. Entries are held weakly, so a table
goes away with the last engine using it.

Measured on an 18 Mpx detector, 32 worker threads, 16 GB card: the stills path went
from exhausting the card (OOM in de-novo indexing) to 8.6 GB peak, and a normal
rotation run from 14.6 GB to 7.4 GB - it had been running within 1.6 GB of the limit,
so any larger detector or second GPU consumer would have tipped it over. Per-worker
footprint drops 403 -> 173 MB. Merge statistics are unchanged on a six-crystal
regression subset, including two-pass runs where the second pass rebuilds the mapping
on refined geometry, and wall time is unchanged (13.5-13.8 s vs 13.8-14.1 s).

Also take the launch configuration from the current device rather than device 0 in
AzIntEngineGPU and ImagePreprocessorGPU: with round-robin pinning, device 0's SM count
and shared-memory size can belong to a different card than the one the kernels use.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rugnux: make spot settings reach rotation indexing, and stop over-claiming
Build Packages / build:viewer-tgz:cpu (push) Successful in 8m55s
Build Packages / build:viewer-tgz:cuda (push) Successful in 10m20s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 12m57s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 13m57s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 13m58s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m0s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 14m56s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 12m35s
Build Packages / build:rpm (rocky8) (push) Successful in 12m34s
Build Packages / XDS test (durin plugin) (push) Successful in 9m18s
Build Packages / Generate python client (push) Successful in 34s
Build Packages / Build documentation (push) Successful in 1m3s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m44s
Build Packages / build:rpm (rocky9) (push) Successful in 13m54s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 13m29s
Build Packages / DIALS test (push) Successful in 14m46s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m10s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m11s
Build Packages / Unit tests (push) Successful in 1h1m18s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
2238290d6d
Three small honesty and cost fixes on the two-pass rotation path.

Spot-finding settings did not reach the step that determines the unit cell. The
two-pass first pass reuses the spots stored in the file whenever it has them, and
reuse is the default, so both sampling schemes and the validation loop ran on
acquisition-time spots while only the per-image pass saw the command line. Every
--spot-* option was therefore a no-op for the lattice search on any file written by
this software, silently, and the lattice was cross-validated against one spot set and
applied to another. Giving spot settings now implies re-finding them for the first
pass as well, and plain reuse says so in the log.

The summary printed a space group and unit cell even when nothing indexed. With a
zero indexing rate the cell is whatever the lattice search happened to return, no
reflection was ever measured on it, and no output file is written - so stating it as
the run's answer claims a result the data do not support. Say that no lattice was
determined instead.

The second pass re-indexes de novo so the cell comes out self-consistent with the
post-refined geometry, and its result was already checked against the first pass -
once by the supercell test and once by the centring test - but only after every image
had been integrated with it, so a disagreement cost a whole extra pass on a dataset
that ended up on the first pass's lattice regardless. Compare them at the point the
lattice is adopted instead, using the same two tests and the same fallback. A
triclinic de-novo cell is left alone, being the demotion the merge reindexes.

Measured over the 37-crystal regression set: merge statistics are unchanged on every
crystal (the two that move are the known rotation-indexing non-determinism - one
observation in 2.9 million, and a zero-score lattice landing on no partials instead of
a few). Crystals whose data were already cached in the reference run are unchanged in
wall time. The one dataset that was burning a discarded pass went from three passes to
two, 509 s to 198 s, against 0.81x for the same-detector dataset that was already
running two passes - so about 214 s of the saving is the removed pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
indexing: complete a rank-deficient direction set, and keep the higher-symmetry setting
Build Packages / build:viewer-tgz:cpu (push) Successful in 6m54s
Build Packages / build:viewer-tgz:cuda (push) Successful in 7m54s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m50s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 14m8s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m14s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 14m21s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 14m33s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 12m45s
Build Packages / build:rpm (rocky8) (push) Successful in 11m56s
Build Packages / XDS test (durin plugin) (push) Successful in 6m43s
Build Packages / Generate python client (push) Successful in 27s
Build Packages / Build documentation (push) Successful in 1m4s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (rocky9) (push) Successful in 13m4s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m51s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 13m6s
Build Packages / DIALS test (push) Successful in 13m36s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m27s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m18s
Build Packages / Unit tests (push) Successful in 1h4m14s
Build Packages / build:windows:nocuda (push) Failing after 4s
Build Packages / build:windows:cuda (push) Failing after 3s
0ae1a307bc
The FFT shortlist could be rank-deficient, and then no cell could be formed at
all. FilterFFTResults takes the strongest max_vectors RAW directions and only
then prunes ones within 5 degrees of each other, but a single lattice row is
sampled by many neighbouring directions of the 16k half-sphere, so thirty raw
peaks routinely prune down to four or five distinct directions - the strongest,
hence shortest, rows. When a crystal's densest rows share a plane, every
surviving direction is coplanar, every triple the reduction forms is degenerate,
and the indexer returns nothing. On such a crystal the weak third axis was the
eighth distinct direction, at raw rank 78. Keep walking the same magnitude order
for up to four more directions that are 5 degrees clear of everything kept,
appended after the length sort so the earlier entries hold their positions and
the reduction still forms every triple it formed before - the shortlist only
gains candidates at its end.

That exposed two ways a change of SETTING was mistaken for a different lattice.
A centred conventional cell is an exact integer multiple of its primitive one,
so the same lattice described two ways differs by that factor: comparing
conventional volumes reads a setting change as a sub-cell or a supercell. Both
the candidate selection in the rotation indexer and the pass-2 comparison in the
driver did exactly that, and between them they discarded a correctly-classified
cubic F cell in favour of the body-centred tetragonal description of the very
same lattice. Compare primitive volumes in both, as the scheme comparison
already did.

Fixing the volumes alone was not enough, because the indexed fraction is also
biased across crystal systems: a subgroup setting holds fewer cell parameters
fixed than its supergroup, so it can never index fewer spots and will always
look better by that measure. Where a candidate has a lower lattice point-group
order at the same primitive volume - the signature of the same lattice in less
symmetry - require it to index markedly better, not merely better, before it
displaces the incumbent.

A general metric-symmetry promotion was implemented and rejected on evidence: it
raised a correct body-centred orthorhombic cell to triclinic and a monoclinic
one to C-centred orthorhombic, and no threshold separates the cases, because a
false pseudo-orthorhombic degeneracy measured tighter than a true cubic one on
obliquity and on alternative-basis axis excess alike. Metric alone cannot decide
this; only the intensities can, which is what the space-group search is for.

Measured over the 37-crystal regression set: one crystal goes from failing
outright to 91% indexed with 91% completeness and a better R_meas than the
reference, one keeps the cubic setting it had before, and every other crystal is
byte-identical. Full unit suite passes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rugnux: report per-image cost honestly instead of per-worker blocked time
Build Packages / build:viewer-tgz:cpu (push) Successful in 7m1s
Build Packages / build:viewer-tgz:cuda (push) Successful in 7m54s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m44s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 13m59s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m12s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 14m18s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 14m26s
Build Packages / build:rpm (rocky8) (push) Successful in 12m10s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 13m21s
Build Packages / XDS test (durin plugin) (push) Successful in 7m6s
Build Packages / Generate python client (push) Successful in 28s
Build Packages / Build documentation (push) Successful in 1m3s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m38s
Build Packages / build:rpm (rocky9) (push) Successful in 13m41s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 13m38s
Build Packages / DIALS test (push) Successful in 13m45s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m33s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m14s
Build Packages / Unit tests (push) Successful in 1h2m17s
Build Packages / build:windows:nocuda (push) Failing after 2s
Build Packages / build:windows:cuda (push) Failing after 3s
bfb8cb813c
Each stage timer measures wall time inside one worker, so it counts whatever that
worker spent blocked on a contended resource - above all the single GPU - as well
as its own work. Those waits overlap across workers, so the mean was printed as if
it were the per-image cost when it is roughly the per-image cost times the worker
count. At the default thread count on a large detector the reported total came out
more than twenty times the truth, and single stages were printed as several times
the entire per-image budget of the run. That is the one output anyone tuning
performance reads, and it sent this investigation at the wrong stage for a while.

Divide by the worker count. It is a lower bound - a worker idle rather than blocked
is not counted - so rather than hide the remainder, report the image loop's own wall
time next to it, and with it the time spent OUTSIDE the loop. Nothing measured the
latter before, yet on a rotation run the first-pass indexing and the scaling and
merging can be more of the run than the per-image work is: on a large-detector run
here it is 5.1 s against 3.0 s. Both figures are for the last pass, and a two-pass
rotation run does all of it twice.

Also stop printing nan. The per-image indexing and scaling timers are never fed on
the two-pass rotation path, because the lattice is forced rather than searched per
image and the merge happens outside the loop, so every default rotation run reported
"indexing nan scaling nan". A stage that did not run is now simply absent.

Measured against the loop's own wall clock on a 18 Mpx dataset: 5% at one worker,
11% at eight, 29% at thirty-two, versus 23x too high before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
image_analysis: stop paying for work that is thrown away
Build Packages / build:viewer-tgz:cpu (push) Successful in 8m17s
Build Packages / build:viewer-tgz:cuda (push) Successful in 9m11s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m38s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 13m57s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 13m57s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 14m13s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m15s
Build Packages / build:rpm (rocky8) (push) Successful in 11m22s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 12m51s
Build Packages / XDS test (durin plugin) (push) Successful in 7m56s
Build Packages / Generate python client (push) Successful in 32s
Build Packages / Build documentation (push) Successful in 1m4s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (rocky9) (push) Successful in 13m23s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 13m15s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 13m53s
Build Packages / DIALS test (push) Successful in 14m21s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m36s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m16s
Build Packages / Unit tests (push) Successful in 1h15m16s
Build Packages / build:windows:nocuda (push) Failing after 2s
Build Packages / build:windows:cuda (push) Failing after 2s
6e805f53c0
Three independent costs, each measured, none changing a result. Across the
37-crystal regression set the run time halves (median per crystal 2.0x, total
2.3x) and every crystal's merge statistics are unchanged.

The image copy back from the device moved the whole preprocessed frame - 72 MB
on a large detector, every frame, per worker - to serve a single host consumer
that reads only the strong pixels, at most a few hundred kilobytes of it. Give
the buffer a Gather() so that consumer asks for the values it actually wants (a
host loop on the CPU, a small kernel on the GPU), and copy the frame back only
when a CPU spot finder will genuinely read it. The copy the other way was worse:
it came from an unregistered vector, so the driver staged it through its own
pinned pool with a host-side memcpy on the calling thread, which does not overlap
and collapses under concurrency - 11.6 GB/s at one worker, 1.6 GB/s at eight.
That, not any hardware limit, is why throughput stopped improving past four to
eight workers. Pinning the decompression buffer once per worker fixes it: on a
18 Mpx dataset the image loop goes from 13.6 to 7.9 ms per image at 32 workers,
and 32 workers now beat 8 instead of losing to them.

Ceres was computing seventeen partial derivatives where five are free. The
per-image rotation refinement frees the beam and the orientation and holds
distance, detector angles, rotation axis and cell constant, but the cost
function declared all seven blocks, so every residual evaluated in Jet<17>
arithmetic. A residual exposing only the two free blocks - the same arithmetic,
the constants baked in - halves refinement, and it is exact rather than merely
close: dual coordinates evolve independently, so the residuals and the free
Jacobian columns are unchanged bit for bit.

The merge sorted an index array with a comparator that dereferenced a 1.6 GB
array of 72-byte records, i.e. a random walk over memory, single-threaded, twice
per two-pass run. Sorting a packed key instead is 2.4x. French-Wilson allocated
its integration scratch per reflection and ran serially; it now takes caller-owned
scratch and runs over chunks, 4.2x. The correction surfaces re-tested every
observation for usability and parity on each of ~22 passes and re-allocated their
accumulators each time; bucket the indices once and hoist the buffers.

Also convert std::round to std::rint where the rounded value only ever enters a
squared residual. The tie rules differ - away from zero against to even - so this
is safe exactly where a tie flips the sign but not the magnitude, and unsafe
wherever the value becomes a Miller index; those sites keep std::round. Verified
over all 2^32 float bit patterns: 8388608 exact ties exist, and the squared
residual is bitwise equal for every one of them. Worth little on its own here,
because the rounding that dominates is in candidate refinement, where the value
is an index and the substitution is not available.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ci: give the MSVC viewer /arch:AVX, and write down why -march lives in CI
Build Packages / build:viewer-tgz:cpu (push) Successful in 6m47s
Build Packages / build:viewer-tgz:cuda (push) Successful in 7m11s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 10m9s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 9m45s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 10m39s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 11m48s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m26s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 11m48s
Build Packages / build:rpm (rocky8) (push) Successful in 10m35s
Build Packages / build:rpm (rocky9) (push) Successful in 11m22s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m41s
Build Packages / Generate python client (push) Successful in 17s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m21s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 49s
Build Packages / XDS test (durin plugin) (push) Successful in 8m19s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 7m42s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m32s
Build Packages / DIALS test (push) Successful in 12m35s
Build Packages / Unit tests (push) Successful in 1h3m57s
Build Packages / build:windows:nocuda (push) Failing after 2s
Build Packages / build:windows:cuda (push) Failing after 3s
b47bce7c3b
The Linux jobs already pass -march=x86-64-v3; the MSVC viewer job passed nothing,
so it built at the x64 baseline. MSVC has no spelling for the x86-64-v2 level, but
/arch:AVX is the nearest and implies SSE4.1/4.2 - which is the part that matters,
because below SSE4.1 Eigen has no vectorised round and falls back to one libm call
per element. AVX is Sandy Bridge and up, a safe floor for a desktop viewer.

The architecture flags stay OUT of CMakeLists on purpose, so a site can build
x86-64-v4 on an AVX-512 cluster, or -march=native, or the plain baseline. That is
easy to mistake for an oversight and "fix", so say it in CLAUDE.md - together with
the consequence that catches anyone profiling: a default local Release build is not
what CI or production runs, and the gap is not uniform. GPU-bound work is
unaffected, but the CPU and Eigen bound phases - first-pass indexing and
scaling/merging - measure about 26% slower without the flags. That is enough to
make rounding look like a tenth of all cycles when a real build has it nearly free,
and to send a reader at the wrong code.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
indexing: select predicted reflections by partiality, build indexers where it pays
Build Packages / build:viewer-tgz:cpu (push) Successful in 7m34s
Build Packages / build:viewer-tgz:cuda (push) Successful in 8m42s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 13m24s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m31s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 13m44s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 14m5s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 14m16s
Build Packages / build:rpm (rocky8) (push) Successful in 11m28s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 12m45s
Build Packages / XDS test (durin plugin) (push) Successful in 7m39s
Build Packages / Generate python client (push) Successful in 36s
Build Packages / Build documentation (push) Successful in 1m4s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (rocky9) (push) Successful in 12m20s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m35s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 13m9s
Build Packages / DIALS test (push) Successful in 13m57s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m57s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m39s
Build Packages / Unit tests (push) Successful in 1h1m5s
Build Packages / build:windows:nocuda (push) Failing after 2s
Build Packages / build:windows:cuda (push) Failing after 3s
639fbb3fbc
When more reflections are predicted for a frame than the output can hold, the
surplus was dropped by keeping those closest to the Ewald sphere. On the rotation
path that quantity is identically zero by construction - the rocking coordinate is
chosen so the scattering vector lands exactly on the sphere - so the comparison
fell through to h, k and l and the survivors were whichever came first in
lexicographic order. Measured on a large cell: every value within one float ulp of
zero, and the kept set had a MEAN PARTIALITY BELOW that of the full set, i.e. worse
than choosing at random. Rank by partiality instead, which the predictor already
computes and which is what the header always claimed was being kept. On the one
regression crystal large enough to cross the cap this lifts completeness from 84.8%
to 90.2% on the same observations; multiplicity and R_meas move the way they must
when the same measurements cover more of reciprocal space.

The online path asked for a cap of ten thousand but the truncation was hardcoded to
the offline limit, so the broker predicted and integrated up to six times what it
could transport and discarded the rest after paying for it. Honour the caller's
limit, which also makes the post-integration re-truncation dead code.

Indexer pool construction becomes a policy. The online service needs every indexer
resident before data arrives, because a cuFFT plan built on the first frame is
planning time inside the measurement; spending memory to be ready is the intended
trade there and stays the default. Offline there is no such deadline, and a stills
run with a known cell was holding a fully allocated FFT indexer per worker that the
algorithm resolution can never dispatch - 2.8 GB where 0.4 GB is needed. rugnux and
the viewer opt into building on first use; the broker, the receiver and the tests
are untouched. This also removes a dangling reference that was latent: the worker
held the settings by reference although the pool is routinely constructed from a
temporary, which only survived because eager construction finished inside the
constructor call.

Finally, refuse a first-pass lattice that indexes fewer than a sixth of the
validation frames. It fires on nothing in the regression set - the weakest real
crystal sits at 22 of 60, more than twice the floor - so it is a backstop, but the
failure it prevents is one the set does contain: a dataset with no crystal at all
adopts a lattice from its powder rings, integrates every image against it, and dies
much later inside the merge complaining about resolution. It now stops in the first
pass and says what to try.

Regression set: 36 of 37 crystals byte-identical, the exception being the
completeness gain above; 34 of 37 space groups, no failures. Full unit suite passes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
image_preprocessing: inline the buffer accessors
Build Packages / build:viewer-tgz:cpu (push) Successful in 7m6s
Build Packages / build:viewer-tgz:cuda (push) Successful in 8m15s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m41s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 13m53s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m15s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 14m27s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 14m44s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 13m6s
Build Packages / build:rpm (rocky8) (push) Successful in 12m1s
Build Packages / XDS test (durin plugin) (push) Successful in 6m58s
Build Packages / Generate python client (push) Successful in 35s
Build Packages / Build documentation (push) Successful in 1m3s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m41s
Build Packages / build:rpm (rocky9) (push) Successful in 14m0s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 13m59s
Build Packages / DIALS test (push) Successful in 13m49s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m38s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m20s
Build Packages / Unit tests (push) Successful in 1h1m46s
Build Packages / build:windows:nocuda (push) Failing after 2s
Build Packages / build:windows:cuda (push) Failing after 2s
e4d70f0e55
operator[], size(), data() and getBuffer() are one-line accessors that were defined
in the .cpp. The build sets no link-time optimisation, so out of line each of them is
a real call - once per pixel, from the CPU preprocessor, the CPU azimuthal integrator
and the CPU spot finder - and they stop those loops vectorising at all. They show up
in a profile directly: about six per cent of a whole azimuthal-integration-only run
is spent in the call overhead of two accessors that do nothing but index a vector.

Moving them into the header retires 30% fewer instructions on that run and takes the
per-image CPU cost on a GPU-less pass from 34.6 to 24.2 ms, with the output bit for
bit unchanged - same observation count, same cell, same merge statistics. It is worth
nothing on the GPU path, where the image stays on the device, and everything on the
paths that have no GPU to fall back on.

This also explains a measurement that had been blamed on the pixel mask being a
vector<bool>: a microbenchmark of that loop indexed a raw pointer and came out far
faster than the same loop in the binary, and the difference was this call, not the
mask. Measured properly the mask costs about 14% single-threaded rather than the 41%
claimed, and at the thread counts this actually runs at the bit mask is FASTER than
the byte mask it was proposed to become, because it moves eight times less traffic
and the loop is bandwidth bound. That change should not be made.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
spot_finding: find connected components on the GPU
Build Packages / build:viewer-tgz:cpu (push) Successful in 7m46s
Build Packages / build:viewer-tgz:cuda (push) Successful in 9m14s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m51s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 14m17s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 14m14s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m43s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 14m45s
Build Packages / build:rpm (rocky8) (push) Successful in 11m44s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 13m24s
Build Packages / XDS test (durin plugin) (push) Successful in 8m33s
Build Packages / Generate python client (push) Successful in 28s
Build Packages / Build documentation (push) Successful in 1m4s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (rocky9) (push) Successful in 12m45s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m25s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 13m1s
Build Packages / DIALS test (push) Successful in 14m29s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m17s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m5s
Build Packages / Unit tests (push) Successful in 1h16m19s
Build Packages / build:windows:nocuda (push) Failing after 2s
Build Packages / build:windows:cuda (push) Failing after 3s
4bdb229fb8
The spot finder flagged strong pixels on the device and then labelled them on the
host, so every frame sent the packed bitmask back - 2.26 MB on a large detector -
and the host walked all of it to recover a few hundred pixels. Do the labelling on
the device instead: compact the bitmask into a flat-index-sorted list, find each
pixel's backward neighbours by binary search, union them lock-free with path
halving, then label, accumulate and filter in one kernel. Only the spot list comes
back, and only one stream synchronisation per frame.

The gain in the ordinary case is modest - about a quarter off per-image spot
finding - because the host algorithm is genuinely fast on a normal frame. What
justifies it is the frame that is not ordinary. The host labels a sorted sparse
list through a window spanning two detector lines, so its cost is quadratic in how
many strong pixels share a line. A lit band of detector rows - a hot module, a
panel edge - costs 33 ms at two rows and 377 ms at fifteen, all of it under the
pixel cap that was supposed to bound this, and none of it maskable when the cause
is a diffraction ring rather than a defect: a ring runs tangent to a row at its
top and bottom, which is exactly the shape that hurts. The device version is flat
at 0.05 to 0.64 ms across every geometry tried, so an online run no longer stalls
a quarter of a second on an ice ring. Rejecting an over-cap frame is now free too,
since the count is known before any pixel is written.

Also label once and filter three times. The per-image minimum-pixel search runs the
extraction at three settings, but that setting only decides which components are
kept - it does not change the components - so the search itself need not be
repeated. This helps the host path as much as the device one.

The resolution mask moves to the device as a bit mask, uploaded when the limits
change rather than per frame, since the compaction needs it there.

Parity is asserted permanently rather than argued: five cases covering realistic
frames, occupancy from a hundred pixels to past the cap, the pathological
geometries including rings, the resolution mask, and a hundred-repeat determinism
check - requiring the same partition, the same spot order, and identical counts.
The centroid is a float sum and therefore order-dependent, so the device walks each
component from its root in ascending order and fuses its multiply-add the way the
host's does; note that whether the host fuses at all depends on the architecture
flags, so exact centroid equality is asserted where the compiler fuses and a
two-ulp bound otherwise. Making those accumulators integer would remove that
dependence entirely and is worth doing separately.

Regression set: all 37 crystals identical to the last printed digit. Unit suite
passes with the new cases.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
reader: read the frame number back with the reflections
Build Packages / build:viewer-tgz:cpu (push) Successful in 7m3s
Build Packages / build:viewer-tgz:cuda (push) Successful in 7m40s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 10m59s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 10m21s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 10m24s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 11m40s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m8s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 12m26s
Build Packages / build:rpm (rocky8) (push) Successful in 11m32s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m0s
Build Packages / build:rpm (rocky9) (push) Successful in 12m44s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m39s
Build Packages / Generate python client (push) Successful in 24s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 54s
Build Packages / XDS test (durin plugin) (push) Successful in 8m19s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 7m27s
Build Packages / DIALS test (push) Successful in 12m54s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m20s
Build Packages / Unit tests (push) Successful in 1h1m7s
Build Packages / build:windows:nocuda (push) Failing after 3s
Build Packages / build:windows:cuda (push) Failing after 2s
ccbc366e2f
The per-reflection frame number is written to the process file and always has been,
but the reader's designated initialiser simply omitted it, so every reflection came
back at frame zero. Nothing complained, because zero is a valid frame.

It matters because the 3D combine splits a reflection's partials into rocking events
by frame contiguity. With every observation claiming frame zero there are no gaps to
split on, so a reflection's entire rotation range collapses into ONE event: measured
on a rotation dataset, 216066 fulls against 216705 distinct reflections, where the
pipeline finds 367416. Forty-two per cent of the observations disappear, the
goniometer-frame absorption surface evaluates every observation at a single angle,
and the radiation-damage estimate is computed over a run that appears to last no
time at all.

The reason this survived is that the damage flatters: fewer, better-agreeing
observations per reflection give R_meas sixteen per cent lower, ISa thirty per cent
higher and a slightly better CC1/2 than the real merge. Anyone re-scaling a stored
file was reading numbers that looked better than the pipeline's while standing on
less than two thirds of the data, and one radiation-damage figure that was pure
artefact.

Read it as mandatory rather than optional-with-default, like h/k/l and the
intensities: a silent zero is precisely the failure being fixed, and every file this
function can read carries the dataset.

After the fix the combine reproduces the pipeline exactly. The normal path does not
go through this reader and is byte-identical before and after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
receiver: stop copying every frame back from the device on the Lite path
Build Packages / Unit tests (push) Successful in 1h1m55s
Build Packages / build:viewer-tgz:cpu (push) Successful in 8m10s
Build Packages / build:viewer-tgz:cuda (push) Successful in 9m20s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 14m6s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m9s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 14m13s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m43s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 14m22s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 13m21s
Build Packages / build:rpm (rocky8) (push) Successful in 12m0s
Build Packages / build:rpm (rocky9) (push) Successful in 13m23s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 13m18s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 13m6s
Build Packages / DIALS test (push) Successful in 13m59s
Build Packages / XDS test (durin plugin) (push) Successful in 8m4s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m40s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m1s
Build Packages / Generate python client (push) Successful in 32s
Build Packages / Build documentation (push) Successful in 1m9s
Build Packages / Create release (push) Skipped
Build Packages / build:windows:nocuda (push) Failing after 13m23s
Build Packages / build:windows:cuda (push) Failing after 12m24s
e7be5447d3
The Lite workflow built its analysis with the fused GPU engine disabled, which is
also what decides whether the preprocessed image is copied device-to-host after
every frame. So on a machine with a GPU the online path was moving the whole image
back - 72 MB on a large detector, every frame, per worker - for a host reader that
does not exist on that path.

It was left off deliberately when the fused engine was added, to keep the online
path unchanged in that commit, and never revisited. Nothing depends on it: the FPGA
workflow uses a different analysis class, and strong-pixel values are read through a
device gather rather than from the host image.

Turning it on changes no result, and cannot: adaptive detection is unreachable
online, because the REST schema exposes no way to enable it, so the classic GPU
finder runs either way. Measured anyway, both engines on the same frames across five
datasets including very weak ones: 2400 frames, 638260 spots, not one difference -
identical lists, identical indexing rate, identical merge statistics to every
printed digit.

On a large detector with eight workers the median per-image cost falls from 94 to
59 ms and preprocessing from 21 to 6 ms; throughput rises from about 48 to 55 Hz. No
percentile regresses, which is what matters for a service - the ninetieth improves
from 128 to 74 ms and the tail with it. Spot finding gets faster too, because the
large copy no longer contends with the device gather.

Correct two statements while here. The flag's comment and the data-analysis
document both said the online receiver uses the CPU adaptive finder; online never
runs an adaptive finder at all, and the copy the flag really controls was not
mentioned. That copy would be better expressed as what it is - whether a host engine
will read the image, which the constructor already knows - rather than inferred from
which spot finder is wanted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
MSVC does not define M_PI without _USE_MATH_DEFINES, and rugnux is part of
the portable subset that JFJOCH_VIEWER_ONLY builds. JFJochMath.h already
carries PI for exactly this reason.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The photon-weighted position sums were floats, so the centroid's last bit
depended on the build rather than on the data: gcc contracts the multiply-add
in AddPixel into an FMA under -march=x86-64-v3 and cannot at the baseline,
and MSVC does not contract at all under /fp:precise. The GPU extractor had to
match with __fmaf_rn, and the parity test still needed a two-ulp slack for
hosts that do not fuse.

Column, line and the per-pixel count are all integral, so the sums are exact
in int64 and both implementations reach the same bits with nothing to match.
The parity test now demands exact equality unconditionally and gets it,
including on a baseline build.

ConvertToImageCoordinates keeps the sums integral too: the raw -> image map
is a signed axis swap plus an integer translation, so it is applied to the
sums instead of to the centroid.

Drops the SpotToSave constructor, which had no callers and could not have
been converted without quantising the stored centroid.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rotation_indexer: demand a decisive margin before adopting an axis multiple
Build Packages / build:viewer-tgz:cpu (push) Successful in 7m55s
Build Packages / build:viewer-tgz:cuda (push) Successful in 9m0s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m54s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 13m59s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 14m18s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 14m27s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m38s
Build Packages / build:windows:nocuda (push) Successful in 17m15s
Build Packages / build:rpm (rocky8) (push) Successful in 11m49s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 13m9s
Build Packages / XDS test (durin plugin) (push) Successful in 7m54s
Build Packages / Generate python client (push) Successful in 31s
Build Packages / Build documentation (push) Successful in 1m7s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2404) (push) Successful in 13m9s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 13m40s
Build Packages / build:rpm (rocky9) (push) Successful in 14m46s
Build Packages / DIALS test (push) Successful in 14m16s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m33s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m47s
Build Packages / build:windows:cuda (push) Successful in 15m50s
Build Packages / Unit tests (push) Successful in 1h41m18s
b6d3dcc6fe
Candidate selection promoted a later cell whenever it indexed 0.05 more of
the accumulated spots. That margin is not meaningful when the candidate is a
near-integer volume multiple of the incumbent: multiplying an axis halves the
reciprocal spacing, so the multiple has a lattice point wherever its sub-cell
has one and another in between, and it collects spots the sub-cell leaves
unindexed for reasons that have nothing to do with the crystal. The indexed
fraction is biased in its favour, and a small lead is not evidence.

On one rotation dataset the true cell and a spurious 5x supercell were
separated by 0.003 of indexed fraction against a bar of 0.05 - close enough
that the -march flags the binary happened to be built with decided it. The
baseline build kept the true cell and merged to an R-free of 0.24 against an
external model; an -march=x86-64-v3 build (what CI uses) took the supercell,
carried it into a doubled cell and a different space group, and merged to an
R-free of 0.58, which is noise. Both were reproducible, five runs each, and
independent of thread count.

An integer multiple now has to index 1.5x the incumbent, the same shape the
lower-symmetry-setting guard next to it already uses. A real superstructure's
satellite rows are a large share of its spots and clear that comfortably.
Both builds now agree on the true cell with a wide margin.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
tests: hold the integer coordinate conversion to the module map
Build Packages / build:viewer-tgz:cpu (push) Successful in 6m40s
Build Packages / build:viewer-tgz:cuda (push) Successful in 7m28s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 10m51s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 12m20s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 11m38s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 12m35s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m39s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 12m10s
Build Packages / build:rpm (rocky8) (push) Successful in 11m55s
Build Packages / build:windows:nocuda (push) Successful in 19m48s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m4s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m42s
Build Packages / Generate python client (push) Successful in 20s
Build Packages / build:rpm (rocky9) (push) Successful in 13m33s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m8s
Build Packages / XDS test (durin plugin) (push) Successful in 8m51s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 7m17s
Build Packages / DIALS test (push) Successful in 13m39s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m5s
Build Packages / build:windows:cuda (push) Successful in 21m16s
Build Packages / Unit tests (push) Successful in 1h47m1s
5f30f0d1ac
ConvertToImageCoordinates now transforms the photon-weighted sums instead of
the centroid, which is only equivalent because a module's raw -> image map is
a signed axis swap plus an integer translation. Check that against the map
itself on every module of a detector whose modules do not share an
orientation, and either side of the 256-column multipixel gaps where the
translation changes. A wrong sign or a dropped gap term on any single module
would otherwise only show up as mispositioned spots on that module.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rugnux: handle ice rings in --scale as the full pipeline does
Build Packages / build:viewer-tgz:cpu (push) Successful in 7m11s
Build Packages / build:viewer-tgz:cuda (push) Successful in 7m42s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 9m36s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 10m35s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 11m20s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 9m14s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 10m40s
Build Packages / build:rpm (rocky8) (push) Successful in 11m41s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 10m38s
Build Packages / build:rpm (rocky9) (push) Successful in 11m41s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 12m42s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m18s
Build Packages / Generate python client (push) Successful in 26s
Build Packages / Build documentation (push) Successful in 1m0s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (neggia plugin) (push) Successful in 7m7s
Build Packages / XDS test (durin plugin) (push) Successful in 7m31s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 7m53s
Build Packages / build:windows:nocuda (push) Successful in 16m47s
Build Packages / DIALS test (push) Successful in 10m22s
Build Packages / build:windows:cuda (push) Successful in 17m37s
Build Packages / Unit tests (push) Successful in 1h42m32s
ea667cb306
--scale did none of the ice handling the run that wrote the _process.h5 had
done, so re-scaling a stored dataset silently produced a different - and
flatteringly more complete - answer than the pipeline it was meant to
reproduce. Three separate gaps:

  * --detect-ice-rings was accepted and ignored. The --scale block returns
    before the line that applies it.
  * Reflections were never flagged as sitting on an ice ring, so the per-image
    scale fit included them. The flag is not stored per reflection, so it has
    to be recomputed from the resolution.
  * RotationScaleMerge was constructed with the ice half-width hardcoded to
    zero. That is what turns a resolution into a ring index, so every ice test
    inside the merge was a no-op whatever was passed to it.

The CC1/2 ring test that decides which rings to drop moves into
FindDecorrelatedIceRings, shared with the full pipeline so both reach the same
verdict on the same data, and --scale now re-merges with the mask the way the
pipeline does. The stills branch re-runs only the merge: the scaling has
already been applied to the reflections and repeating it would compound it.

Measured on a rotation dataset with three decorrelated rings, --scale went
from 8765 unique / 36.3% completeness / R-meas 18.5% / <I/sig> 1.1 to
7638 / 31.6% / 18.0% / 1.3, against the full pipeline's 7692 / 31.8% / 17.9% /
1.3 - the reported completeness had been inflated by reflections the pipeline
drops. The full pipeline is bit-identical across the refactor.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The round-trip test checked h, k, l, I and the two predicted coordinates. The
other ten fields were written and read back unasserted, which is how the
offline --scale path came to lose image_number without a test noticing - it is
the field 3D-integrated reflections carry a fractional value in, and rocking
events cannot be grouped without it.

Reflections are now built by a helper that puts a distinct value in every
field that is meant to survive, keyed on the image and the reflection index so
a value read back from the wrong place cannot match, and checked by one that
asserts all of them. Verified by reintroducing the image_number loss, which
fails six assertions and passes none of them silently.

dist_ewald, observed and on_ice_ring are deliberately excluded and the test
says why: the first two are prediction/integration scratch that is never
written, and the third is recomputed from the resolution by whoever scales.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The self-calibrating detection threshold was reachable from rugnux and the
viewer but not online: spot_finding_settings carried no adaptive_threshold, so
the receiver always ran the fixed-threshold finder and the fused GPU engine
sat unused behind it.

adaptive_threshold and false_pixels_per_frame are now part of the schema, both
optional so an existing client that sends neither is unaffected, wired through
OpenAPIConvert in both directions and surfaced in the frontend panel, where
turning the mode on greys out the count threshold it replaces and reveals the
operating point it uses instead. The C++ server model and the TypeScript
client are regenerated from the spec; the Python client is generated but not
tracked.

Enabling it is REFUSED where spots are found on the FPGA - the JUNGFRAU and
EIGER workflows - rather than accepted and ignored, because a detection
setting that silently had no effect cannot be told apart from one that did.
The DECTRIS/SIMPLON workflow, which analyses images in software, accepts it.
Verified against a running broker: adaptive_threshold true is rejected with
that message and leaves the stored settings untouched, while false and omitted
both succeed.

It stays off by default online, unlike rugnux and the viewer. The broker
serves both workflows and the default has to be the one that works on either.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
update_version.sh had not been run for the adaptive spot-finding schema
change. Running it leaves the C++ server model and the TypeScript client
byte-identical to what the generators produced directly, but it also
regenerates two artefacts the direct calls do not touch and which are tracked:
the Python client's published documentation and the Redoc bundle. Both now
carry adaptive_threshold and false_pixels_per_frame.

LTO joins -march in the CI flags, which is why MARCH_CMAKE_FLAGS is now
LINUX_CMAKE_FLAGS - it no longer describes only the architecture. Measured on
rugnux against an otherwise identical build: 7-10% fewer retired instructions
and a 9% smaller binary, but only ~1.5% off the wall clock, because the
pipeline is GPU- and I/O-bound. It costs about 3x on an incremental rebuild
(9.8 s -> 30.1 s for one file plus link), so it stays out of CMakeLists and
out of a developer's edit cycle: CI builds from scratch and ships the result,
paying the link once. It links against CUDA with no special handling, and both
CI images already put gcc-toolset-13 on PATH, which -flto=auto requires.

MSVC is left alone: its LTO is a different flag (/GL + /LTCG) and nothing here
measured it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ci: record why LTO is a flag and not CMAKE_INTERPROCEDURAL_OPTIMIZATION
Build Packages / build:viewer-tgz:cpu (push) Successful in 15m33s
Build Packages / build:viewer-tgz:cuda (push) Successful in 16m1s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 17m16s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 18m56s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 18m30s
Build Packages / build:windows:nocuda (push) Successful in 14m54s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 16m10s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 17m47s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 21m31s
Build Packages / build:rpm (rocky8) (push) Successful in 20m26s
Build Packages / build:rpm (rocky9) (push) Successful in 15m58s
Build Packages / XDS test (durin plugin) (push) Successful in 9m8s
Build Packages / Generate python client (push) Successful in 35s
Build Packages / build:windows:cuda (push) Successful in 20m42s
Build Packages / Build documentation (push) Successful in 1m23s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2404) (push) Successful in 16m32s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 20m3s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m47s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m55s
Build Packages / DIALS test (push) Successful in 15m9s
Build Packages / Unit tests (push) Successful in 1h17m19s
87f31fa06c
The CMake variable is the tidier spelling and would cover the MSVC job too, so
it is the obvious thing to reach for and worth saying why it was not.

It builds and links, CUDA included - and it does not reach .cu targets either
way, so there is no -dlto risk on either route. But it optimises less: 396.9 G
retired instructions against 384.9 G for -flto=auto, three runs each, with a
0.45% run-to-run spread, so a 3% gap is not measurement luck. Of 107 static
libraries the two routes agree within 5% on 105; the flag additionally covers
FFTW and libzmq. And CMAKE_AR stayed plain ar under the variable, so the
archive-handling argument for it did not hold here either.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
indexing: stop computing angles the candidate filter only compares
Build Packages / build:windows:nocuda (push) Successful in 16m15s
Build Packages / build:viewer-tgz:cpu (push) Successful in 19m47s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m4s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m49s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m7s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m2s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m7s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m19s
Build Packages / build:windows:cuda (push) Successful in 15m41s
Build Packages / XDS test (durin plugin) (push) Successful in 10m48s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 19m55s
Build Packages / build:rpm (rocky9) (push) Successful in 21m11s
Build Packages / Generate python client (push) Successful in 40s
Build Packages / Build documentation (push) Successful in 1m44s
Build Packages / Create release (push) Skipped
Build Packages / DIALS test (push) Successful in 20m53s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 26m15s
Build Packages / build:rpm (rocky8) (push) Successful in 27m37s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m36s
Build Packages / XDS test (neggia plugin) (push) Successful in 10m11s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 11m5s
Build Packages / Unit tests (push) Successful in 1h21m20s
83e95b0c5a
Candidate cell filtering called acos three times per candidate to turn dot
products into degrees, then compared those against the min/max angle bounds.
acos is strictly decreasing on [-1, 1], so "angle outside [min, max]" is
exactly "cosine outside [cos(max), cos(min)]" with the ends swapped - the
bounds convert once, and the three acos calls per candidate disappear.

The same loop also re-derived every already-accepted candidate's unit cell on
each new triple, inside the duplicate scan: three more acos each, for every
candidate accepted so far. Those cells are now kept alongside the candidates.

Measured on de-novo serial stills, where the indexer runs once per image:
34.43 s -> 14.17 s on one dataset and 21.92 s -> 6.59 s on another, with the
indexing rate and the merged reflection count unchanged (one gained 0.25
points of indexing rate). acos had been 40% of the whole process there.

Scope is narrower than that number suggests, and worth stating: the win is on
the de-novo path, which Auto selects for stills only when NO cell is known.
With a known cell Auto picks ffbidx, which reaches the same filter but feeds
it few candidates - measured neutral there (+0.5% instructions, -1.6% wall,
identical output), and that path already runs 14x faster in absolute terms.
Rotation runs the indexer twice per dataset rather than per image, so it is
unaffected: the full 37-crystal battery is identical, crystal for crystal.

Comparing cosines instead of angles can only move a candidate that sits on the
bound, so the filter's behaviour is unchanged except at that measure-zero
boundary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two bugs in analyze_pixel, both confined to the middle stage of the wave.

The kernel walks each wave's rows in three stages. The priming and drain
loops read prev_out and substitute INT32_MAX for a pixel the previous pass
found strong, exactly as the CPU finder's value_at() does on every read. The
main loop did not - it read the image raw. So in the second pass the pixels
the first pass found strong stayed in the background statistics, inflating the
local mean and variance, and the halo of every broad spot failed the
signal-to-noise test. The two engines therefore did not agree, despite
1a1e05ad1 having set out to make them.

Separately, shared_sum2 is an int64 accumulator but val*val and old*old were
computed in int32 at three sites. That wraps above 46340 while the detector
overloads around 1e6, so any window containing a bright pixel got a corrupted
variance. pixel_result already did the same arithmetic in 64 bits.

The reason this survived is worth recording: numberOfWaves is fixed at 32, so
a wave owns ceil(height/32) rows, and the main loop only runs while
front < rmax with front starting NBX+1 = 16 rows ahead. At the existing test's
100 rows a wave owns 4 rows and THE MAIN LOOP NEVER EXECUTES - every row goes
through priming or drain, and the parity test passed on the broken kernel. On
a 4362-row detector frame a wave owns 137 rows and the main loop carries about
121 of them, so the bug covered roughly 88% of a real image.

The new tall-image test is sized to the partitioning rather than to
convenience: 1024 rows, spots placed inside the main-loop region, one core at
100000 to exercise the overflow. Against the unfixed kernel it reports GPU 25
pixels where the CPU finds 49 and fails six assertions.

The default rugnux path is unaffected because it uses the adaptive finder, and
the 37-crystal battery is identical crystal for crystal. The classic finder is
what SpotFindingSettings defaults to, so this is the online receiver's path;
measured there with --no-adaptive-spots, R-meas improves 11.2% -> 10.9% and
CC1/2 98.8% -> 98.9% on one crystal, with multiplicity up on both tried.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
reduce_rings_shared is the largest kernel in the per-image loop - 73% of GPU
kernel time on an 18 Mpx rotation run, launched three times per image - and it
is bound by shared-memory atomic replay rather than by bandwidth: it reaches
156 GB/s against a measured 913 GB/s ceiling, and removing the atomics while
keeping the same loads makes it five times faster.

That is the case that wants resident warps to hide the serialisation, and four
blocks per SM left only 512 of the 1536 threads an SM can hold. The per-block
histogram is nbins * 20 B, about 9.6 kB at the default 0.01 1/A spacing, so
eight blocks fit in shared memory with room to spare. Both kernels are
grid-stride loops, so any grid is correct and a device that cannot co-schedule
eight simply queues the rest.

Measured: 9.21 s -> 5.33 s of kernel time over a run (852 -> 493 us per
launch), cutting total kernel time from 12.57 s to about 8.85 s.

flag_strong keeps four. It is bandwidth-shaped rather than atomic-bound and
eight measured no better (181 vs 175 us).

Wall clock is unchanged, and that is expected rather than disappointing:
kernels are 39% of the image loop while the host-to-device copy is 78%, so
faster kernels idle the GPU more without shortening the loop. This is
groundwork for the transfer work, not a speedup on its own.

The shared accumulators are float and summed with atomics, so the block count
changes the summation order and with it the last bits. The 37-crystal battery
is identical crystal for crystal except one observation in 925850 on a single
dataset.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
image_preprocessing: decode bitshuffle+LZ4 on the GPU
Build Packages / build:viewer-tgz:cpu (push) Successful in 20m32s
Build Packages / build:viewer-tgz:cuda (push) Successful in 20m40s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m24s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m8s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 27m31s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 27m38s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 29m7s
Build Packages / XDS test (durin plugin) (push) Successful in 11m12s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 22m49s
Build Packages / build:rpm (rocky9) (push) Successful in 22m51s
Build Packages / Generate python client (push) Successful in 40s
Build Packages / Build documentation (push) Successful in 1m22s
Build Packages / Create release (push) Skipped
Build Packages / DIALS test (push) Successful in 20m21s
Build Packages / build:rpm (rocky8) (push) Successful in 27m26s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m59s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m52s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m41s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m41s
Build Packages / Unit tests (push) Successful in 1h17m41s
Build Packages / build:windows:nocuda (push) Successful in 13m24s
Build Packages / build:windows:cuda (push) Successful in 17m0s
6e4c0ce202
The pipeline decompressed each image on the host and uploaded the result. On
an 18 Mpx rotation dataset that made the host-to-device copy the bottleneck of
the whole per-image loop: nsys puts the copies at 78% of the loop against 39%
for every kernel combined - 3600 transfers of 72.4 MB - and they ran at only
12.5 GB/s of an available 27-28 because the host-side decompression was itself
saturating host memory bandwidth. The GPU was mostly waiting.

So the compressed chunk goes across instead, about 4 MB rather than 72 MB, and
is decoded on the device. That removes the transfer and the host decompression
that was throttling it, in one change. Measured on an idle machine, a run goes
from 45.11 s to 24.97 s - 1.81x - with the merged output unchanged.

THE APPROACH IS JON WRIGHT'S (ESRF): "Experiences with GPU decompression for
bitshuffle + LZ4 data", HDF5 User Group 2021, and github.com/jonwright/
bslz4decoders. The kernels here are ours, but the idea and the demonstration
that it is worth doing are his. Cited in docs/ACKNOWLEDGEMENT.md and in the new
section 0 of docs/CPU_DATA_ANALYSIS.md.

Two kernels mirror the CPU decoder. LZ4 runs one WARP per bitshuffle block:
every lane parses the same sequence stream (a broadcast read, no divergence)
and the literal and match copies are split across the 32 lanes so the stores
coalesce; an overlapping match is treated as a pattern of period offset sourced
from bytes that already precede the write position, which keeps it parallel
rather than a serial byte loop. One thread per block instead measured 13x
slower. The bitshuffle inverse then un-transposes each byte-plane through
shared memory and interleaves the planes back into elements.

Only BSHUF_LZ4 is decoded on the device. The zstd variants have no device
decoder, and neither has an uncompressed or float image; Supports() returns
false for those and the caller decompresses on the host exactly as before. The
fallback is explicit, so a format we cannot decode on the device is a slower
path and never a wrong answer.

Tests hold the device decoder against the CPU one byte for byte, on data from
the production compressor, for every element size the detectors emit -
including the 8-bit DECTRIS modes, which take bitshuf_decode_block's separate
elem_size == 1 branch - plus a many-block frame, the formats it must decline,
and malformed containers, which must throw rather than run off a buffer.

Battery: 37 crystals, no failures, identical to the host-decode run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
EnsureCapacity resized its 13 device arrays to exactly the current image's
predicted-reflection count, so every image that set a new record freed and
reallocated all of them. cudaMalloc and cudaFree take a device-wide lock in the
CUDA driver, so those images stalled every other worker: sampling the worker
threads during the per-image loop found 21-24 of 32 parked in cuMemAlloc_v2 or
cuMemFree_v2, all called from this one function, and the running maximum makes
32 workers do far more allocator work than one does.

Grow by half again instead. All transfers and kernel launches are sized by the
per-image reflection count rather than by the capacity, and the member is
already documented as holding at least that many, so over-allocating changes no
result. On an 18 Mpx rotation set the integration stage drops from 1.37 to
1.25 ms per image at 32 workers; merged statistics, error model and adopted
space group are unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rugnux: parallelise candidate-cell refinement, and stop repeating work in the tail
Build Packages / build:viewer-tgz:cpu (push) Successful in 20m26s
Build Packages / build:viewer-tgz:cuda (push) Successful in 21m30s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m36s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 24m4s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m10s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m12s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m23s
Build Packages / XDS test (durin plugin) (push) Successful in 11m21s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 20m56s
Build Packages / build:rpm (rocky9) (push) Successful in 21m10s
Build Packages / Generate python client (push) Successful in 40s
Build Packages / Build documentation (push) Successful in 1m34s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (rocky8) (push) Successful in 25m28s
Build Packages / DIALS test (push) Successful in 21m15s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m26s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m51s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m53s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m41s
Build Packages / Unit tests (push) Successful in 2h21m29s
Build Packages / build:windows:nocuda (push) Successful in 1m15s
Build Packages / build:windows:cuda (push) Successful in 28m0s
7e47afe47f
Three independent changes to the CPU-bound parts of an offline rotation run, none
of which alters a result.

Candidate-cell refinement now splits across threads. RefineCandidateCells already
took a (block, nblocks) partition, but the only call site passed nblocks=1, so the
whole first pass of a two-pass rotation run sat on one thread per scheme - two
threads, unchanged at every -N, for a third of the run. A block touches only its
own scores(j) and cells rows and holds its own scratch, so the split is exact.
The budget is a new IndexingSettings::RefineThreads, left at 1 by default and set
only where few indexer threads exist: raising it unconditionally would
oversubscribe the paths that already run one indexer per image across all workers.

The mmCIF writer built a std::ostringstream per formatted number, twelve per
reflection. snprintf gives the same digits for 0.535 -> 0.220 s per file.

The space-group search built the same orbit mapping twice per candidate point
group - once for the merge chi^2 and once for the systematic-error b, an
apply_to_hkl and Canonicalize per observation per operator each time. Build it
once and hand it to both.

18 Mpx rotation set 24.6 -> 18.7 s, 2.5 Mpx 13.0 -> 10.7 s, and the 37-crystal
battery 13m55s -> 10m47s with no failures, the same 34/37 space groups, and
statistics unchanged on 30 of 37 (the rest drift within the run-to-run spread the
binary already had, which a control build with the split disabled reproduces).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
docs: tighten the changelog and the decoding section for release
Build Packages / build:viewer-tgz:cpu (push) Successful in 11m6s
Build Packages / build:viewer-tgz:cuda (push) Successful in 15m41s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 19m45s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 20m58s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 24m16s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 24m47s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 22m49s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 17m45s
Build Packages / build:rpm (rocky9) (push) Successful in 18m59s
Build Packages / XDS test (durin plugin) (push) Successful in 11m18s
Build Packages / build:rpm (rocky8) (push) Successful in 25m1s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 19m31s
Build Packages / Generate python client (push) Successful in 49s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m30s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 23m30s
Build Packages / DIALS test (push) Successful in 17m23s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m53s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m21s
Build Packages / Unit tests (push) Successful in 1h56m58s
Build Packages / build:windows:nocuda (push) Successful in 18m13s
Build Packages / build:windows:cuda (push) Successful in 22m29s
737cbde3ff
The rc.161 changelog had grown entries several hundred words long and listed the
same area three or four times over. Collapse them by subject - spot finding,
resolution limits, space-group search, scaling, performance, correctness - and
hold each to one line, keeping the actionable detail in the breaking API entry.
Add the performance work that had not been written up: device-side image
decoding and the parallel first-pass candidate-cell refinement.

Section 0 of the CPU analysis document was the longest thing in it after two
core algorithm sections, and most of that was a profiling narrative and the
measurements that motivated the change rather than a description of what runs.
Cut it to the two kernels, the host-side block scan and the fallback rule. The
attribution stays; it is also in the reference list.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
docs: split the breaking API changes out of the rc.161 change list
Build Packages / Unit tests (push) Successful in 1h20m27s
Build Packages / build:windows:nocuda (push) Successful in 15m19s
Build Packages / build:viewer-tgz:cpu (push) Successful in 11m44s
Build Packages / build:viewer-tgz:cuda (push) Successful in 13m1s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 19m11s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 16m35s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 20m51s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 15m28s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 19m40s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 17m44s
Build Packages / build:rpm (rocky8) (push) Successful in 20m45s
Build Packages / build:rpm (rocky9) (push) Successful in 18m12s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 19m32s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 15m51s
Build Packages / DIALS test (push) Successful in 14m57s
Build Packages / XDS test (durin plugin) (push) Successful in 8m53s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m30s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m20s
Build Packages / Generate python client (push) Successful in 38s
Build Packages / Build documentation (push) Successful in 1m4s
Build Packages / Create release (push) Skipped
Build Packages / build:windows:cuda (push) Successful in 19m11s
47277674fa
Earlier releases put breaking changes in their own paragraph after the bullets
(rc.139, rc.29) rather than as one item among them. Follow that: the OpenAPI
changes now sit under their own heading below the list, with the client-side
action in the lead line, and are listed one per change instead of run together.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The device decoder was byte-exact on every valid input - 994 production-compressed
images, 927 hand-built LZ4 blocks covering engineered (offset, matchlen) pairs across
the overlap branch boundary, 18000 repeat decodes, sanitizer-clean - and an audit
against LZ4_decompress_generic could not construct a valid block it mis-decodes. What
it did not do was notice when the input was NOT valid, and that mattered more than it
looks: the decode buffers are reused frame to frame, so a block that stopped early left
the PREVIOUS image in place, and in the bitshuffled layout the untouched tail is the
most significant byte-plane. A corrupt chunk therefore did not look like a missing
corner. It looked like thousands of real pixels several powers of two too bright, fed
to spot finding with no diagnostic, where the host decoder had raised an error.

So the kernel now flags a block that fails to reach its declared length while consuming
exactly its payload, and the host turns that into an exception once the caller has
synchronised. Reads are clamped against the end of the payload as well as the output,
both length chains are bounded exactly as read_variable_length bounds them, the two
offset bytes are bounded, and LZ4's parsing restrictions are enforced. On the host side
a block size that is not a multiple of 8 elements is rejected (it made the un-transpose
read uninitialised shared memory), the block count is bounded by what the chunk could
hold before it becomes an allocation (twelve header bytes could demand hundreds of MB
of pinned memory, permanently, per worker), trailing bytes are rejected, and the stream
is synchronised before any throw that happens after work is queued. An image of fewer
than 8 elements is all verbatim tail and now decodes rather than throwing. When the
device route fails for any reason the host decoder gets its turn, so it costs speed
rather than the acquisition.

The lanes cooperate on the copies and a later match can read bytes another lane wrote,
which since Volta needs an explicit __syncwarp(); it worked only because ptxas happened
to reconverge at the post-dominator. The prototype's offset == 1 and power-of-two fast
paths are also restored - the shipped kernel ran a runtime modulo, an emulated 32-bit
division per output byte, on the path its own comment calls the common case.

The un-transpose is now fused with preprocessing. One thread owns one group of 8
elements across every byte-plane, so once it has transposed its 8 bytes out of each
plane it holds 8 complete elements and emits 8 finished int32 pixels with the mask, the
error marker, the saturation cap and the statistics applied. The decompressed image is
never materialised: 0.623 -> 0.411 ms/frame at 18 Mpx, 0.523 -> 0.340 with 8 concurrent
workers. Staging nothing in shared memory also drops the 48 kB ceiling, which had made
any file whose bitshuffle blocks exceed it a hard failure; 64 kB blocks now decode.
gpu_compressed is sized from the chunk with grow-on-demand instead of from the
uncompressed size - it was reserving ~73 MB per worker to hold ~4 MB. Measured on a
1630x1553 uint32 rotation set at -N 32, peak GPU memory falls 3756 -> 3084 MiB; the
same model gives ~144 MB per worker on an 18 Mpx frame.

Decoding on the device also stopped reporting a decompression time, which blanked the
broker's compression plot trace and filled /entry/profiling/compressionTime with NaN.
The decoder brackets the decode with CUDA events and reports it again.

Tests: a differential fuzz suite against the CPU decoder - incompressible and highly
compressible data, engineered offsets, a size sweep hitting every rem%8 value twice,
all six element sizes, an 18 Mpx frame, decoder reuse, concurrency, hand-built LZ4
blocks across the overlap boundary, 26 foreign bitshuffle block sizes from 128 B to
64 kB, corrupt payloads and malformed containers, with a coverage report that proves
which LZ4 paths were reached rather than assuming it. Plus the fused path held byte for
byte against ImagePreprocessorCPU, statistics included, and against the host-upload
path on the same frame.

Battery: 37 crystals, every merged number identical to the host-decode run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
RefineThreads rejects anything above 64, and the two-pass rotation path - the default -
passed nthreads/2 straight from the machine. On a host with 130 or more hardware
threads rugnux therefore exited 1 with the bare message "Candidate-cell refinement
thread count" before reading a single image, and the only workaround was to discover -N
and pass a smaller value. Dual 64-core parts are squarely in range.

The limit is now a named constant the setter and the callers share, so the two cannot
drift apart again.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
How far the predictor walks the lattice became a setting defaulting to "derive it from
the refined cell", which is what is wanted offline. Online it is not: per-image
prediction cost then depends on whichever crystal is mounted, and the cube (2n+1)^3
grows quickly. The broker was only ever handed a concrete value when a client sent a
bragg_integration block, and none of the shipped configs has one, so the derived path
was the normal deployment.

Four places in the tree - the conversion, the predictor, the OpenAPI description and the
changelog - already state that the broker keeps a fixed bootstrap. Make that true.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two leftovers from earlier fixes of the same shape. BraggIntegrationEngineGPU still read
device 0's shared-memory size to decide whether its profile grid fits; workers are pinned
round-robin across GPUs, so on a heterogeneous node that check can pass on a different
card than the one the kernel launches on. SpotExtractorGPU still uploaded its default
resolution mask with a pageable copy on the NULL stream, which is not ordered against the
engine stream now that streams are created non-blocking.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The per-image image-scale B factor was dropped from the CBOR stream and from the written
HDF5, which is a change for anything reading those files, but the changelog listed it only
under the OpenAPI breaking changes.

The GPU adaptive finder test claimed both finders sum the rings in double. The CPU one
does; the GPU one stages a block's contribution in float before reducing across blocks in
double, deliberately, to keep the hot loop's shared footprint down. Say so, and say what
follows from it - detection compares integer pixel values, so a threshold that crosses an
integer flips every pixel of that value in the ring at once.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The raw-bytes path assembled each element a byte at a time, which on a full frame cost
about 4x against writing the 8 contiguous elements a thread owns through an
element-typed pointer. They are 8*ES-byte aligned, so the compiler merges them.
72.4 MB frame: 1.524 -> 0.406 ms for upload plus both kernels.

The test now also times the LZ4 pass on its own, so the bounds and validity checks in
the hot loop can be costed rather than guessed at. They are free: 0.231 ms against
0.2297 ms measured for the kernel before any of them existed - the restored offset == 1
and power-of-two fast paths pay for them. compute-sanitizer memcheck reports no error
over 400 single-bit-corrupted payloads and nine malformed containers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two futures per connection are joined from more than one path: the acceptor reaping a
dead connection, and the control plane starting or ending a run. Calling get() on one
future from two threads at once is undefined and invalidates it, and the guard against
it was a non-atomic test-then-set of an atomic flag - `if (!c.active) return; c.active
= false;` - so both callers could pass it.

Holding connections_mutex across the joins is what the previous commit removed on
purpose, and rightly: the joins block on a writer that may be inside a send. So the
lock is per connection and covers only the teardown. Nothing it joins takes it, so it
cannot deadlock.

The comment claiming a detached connection is unreachable by anyone else was wrong: a
control-plane call that copied the pool before the erase still holds a shared_ptr to it.
That is precisely how the two teardowns meet.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The stills first pass drew frames from a shared cursor and stopped when a shared counter
reached its target, which got two things wrong at once. The cursor walked the equally
spaced sample in ascending order, so stopping early read only its leading PREFIX - the
beam centre, distance and cell were fitted to the beginning of the run, not across it,
and the comment claiming otherwise was wrong. And where the stop landed depended on how
the workers happened to interleave, so the set of frames varied run to run: on the same
data at -N 32 and -N 8 the pass examined 483 and 457 frames and refined the detector
distance to 168.0481 and 168.0530 mm.

The sample is now cut into a fixed number of interleaved stripes, each stopping once it
has contributed its share. Every stripe spans the whole run, so an early stop no longer
biases the fit, and a stripe is processed identically whichever worker claims it - so
what gets examined depends only on the data, not on timing and not on -N. The same three
runs now give 451 frames examined and 168.0452 mm, identically.

The bundle selection was order-dependent too: frames are collected in worker-completion
order and sorted by spot count with a non-stable sort, so equally strong frames swapped
places between runs. They carry their image ordinal now and it breaks the tie.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The per-ring sums were floats reduced by atomics, so the ring sigma - and with it the
detection threshold - depended on the order the blocks happened to arrive in. Detection
compares an INTEGER pixel value against that threshold, so a threshold that drifts
across an integer flips every pixel of that value in the ring at once, which is how a
last-bit difference turned into a different spot list.

A preprocessed pixel is an exact int32 and the masked and saturated sentinels are
skipped, so v and v*v are exact in 64 bits, and integer addition is associative: the
sums no longer care about arrival order. Both engines now accumulate the same way, so
they agree exactly rather than approximately, and the GPU spot list is bit-identical
across runs. The corrected sums that feed the reported azimuthal profile stay float -
a pixel value times a float correction has no exact integer form - but they do not
enter the detection decision.

Cost: the ring reduction needs 28 bytes per bin instead of 20 in the plain pass, which
drops it from eight co-resident blocks per SM to seven and costs about 11% of that
kernel (0.582 -> 0.650 ms/frame on a 4.5 Mpx frame). End to end it does not show:
alternating runs on three rotation crystals came out the same or slightly faster, and
the battery is unchanged in every number. The CPU engine got 30% faster (32.2 -> 22.6
ms/frame), integers being cheaper than doubles.

Tests: exact CPU/GPU agreement on the spot list, and 50 repeats of bit-identical output
where there were four.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`frac > RATIO * best_frac` cannot be satisfied once best_frac passes 1/RATIO - above
0.667 for a ratio of 1.5, which is ordinary for good rotation data. Above that the two
guards do not raise the bar, they close the branch: no axis multiple and no
lower-symmetry setting can displace the incumbent however much better it fits, so a
genuine superstructure is kept as its sub-cell and its satellite rows go unindexed,
silently.

The obvious repair - restate the bar on the fraction left UNINDEXED, which is well
defined over the whole range - was implemented and measured. It regressed the
37-crystal battery from 34/37 to 32/37 correct space groups: a C2 lattice fell to P1,
and a P2 case went to C222 keeping 2923 of 22440 reflections with CC1/2 in the last
shell at -35%. The indexed fraction is too noisy to carry a looser test.

So the unreachable-but-safe form stays, and the limitation is recorded at the comparison
rather than left to be rediscovered. Fixing it properly needs the selection to be
decided on something better than the indexed fraction.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Each image is written at its own ordinal, so a frame that fails to load or analyse leaves
a HOLE - the frames after it keep their positions rather than shifting up. The end
message nevertheless reported the number of frames that SUCCEEDED as the image count,
and the writer sizes /entry/data/data and every per-image array from that.

So one failed frame in the middle of a run made the declared extent one short, and the
image it dropped was the LAST one written, not the one that failed. Two failures dropped
two, and so on: the file quietly ends before the data does, with the per-image metadata
still carrying rows for images the VDS no longer maps. It also under-counted the data
files when the run was split.

Track the highest ordinal actually written and use that. A hole then reads as the fill
value, which is what a frame that was never written should look like, and the counts of
collected and written images stay counts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The cache returned a device copy for a (device, host address) pair and cast it to
whatever the caller asked for, with nothing checking that the bytes behind that address
were still the same bytes. A host buffer can be mutated in place - PixelMask::LoadMask
does exactly that - or freed and reallocated at the same address, and either hands the
caller a device copy of something else. Nothing would report it: the tables are read-only
geometry, so the engine would simply mask the wrong pixels for the rest of the run while
the azimuthal mapping, the written pixel_mask dataset and the viewer overlay used the new
one. Today that is unreachable, but only because of two guards in unrelated files that
neither state nor assert the requirement.

The byte length and an FNV-1a checksum of the bytes being uploaded are now part of the
key. Both are computed once per engine construction, over a buffer that is about to be
copied to the device anyway, so the cost does not show. Expired entries are pruned on
insert, since distinct content now means distinct entries.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
docs: bring the rc.161 change list up to date
Build Packages / build:viewer-tgz:cpu (push) Successful in 20m10s
Build Packages / build:viewer-tgz:cuda (push) Successful in 21m31s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 23m16s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m36s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m7s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m8s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m10s
Build Packages / XDS test (durin plugin) (push) Successful in 11m32s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 20m27s
Build Packages / build:rpm (rocky9) (push) Successful in 22m10s
Build Packages / Generate python client (push) Successful in 38s
Build Packages / build:rpm (rocky8) (push) Successful in 25m37s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m33s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m25s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 26m6s
Build Packages / DIALS test (push) Successful in 21m34s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m38s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m35s
Build Packages / Unit tests (push) Successful in 1h19m26s
Build Packages / build:windows:nocuda (push) Successful in 21m29s
Build Packages / build:windows:cuda (push) Successful in 27m54s
ceb92fc4cc
Covers the GPU decode work (fused un-transpose, the memory it frees, corrupt-chunk
detection, large bitshuffle blocks, host fallback), the two reproducibility fixes
(integer ring statistics, striped geometry-refinement sampling), the connection-teardown
and written-extent fixes, and the thread-count, max_hkl and compression-time repairs.
Folded into the existing entries where they belong rather than added as new ones.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CPACK_DEBIAN_MAIN_COMPONENT does not exist. Only CPackRPM.cmake has a MAIN_COMPONENT
variable, which is why the RPM came out as "jfjoch" while the .deb of the same component
came out as "jfjoch-jfjoch" - the DEB generator names every component package
"<CPACK_PACKAGE_NAME>-<component>" unless CPACK_DEBIAN_<COMPONENT>_PACKAGE_NAME overrides
it, and the line we set was read by nobody. It also named a component ("broker") that
does not exist, so it could not have matched anything either way. Set the name the
generator actually reads, and declare Replaces/Conflicts on the old one: the new package
owns the same files, so dpkg would otherwise refuse to unpack it over an installation
that already has jfjoch-jfjoch.

The DKMS component's postinst and prerm are the driver's postinstall.sh and
preuninstall.sh, configure_file'd into place. They were mode 644, and
CPACK_DEBIAN_PACKAGE_CONTROL_STRICT_PERMISSION is off, so they were packaged as they are
- and dpkg cannot execute a maintainer script it cannot execute. The RPM path is
unaffected: it reads the same files as text into the spec.

The viewer's Freedesktop menu entry and its icon are installed on Linux only now. On
Windows the Start Menu shortcut comes from CPACK_PACKAGE_EXECUTABLES and on macOS from
the .app bundle, so on those two the installer was carrying a share/applications entry
and a share/pixmaps icon that nothing reads.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
PROJECT() carried a hardcoded 1.0.0 next to a JFJOCH_VERSION read from the VERSION file,
so PROJECT_VERSION was free to drift from the version everything else uses. It cannot
simply be handed the same string - project(VERSION) accepts numeric major.minor.patch
only, and rejects a pre-release suffix such as -rc.161 - so cut the numeric part out of
the same file instead of writing it down a second time.

common/ then read ../VERSION again into PACKAGE_VERSION, purely to interpolate it into
GitInfo.cpp. That is the same file read twice with two variable names, one of them a
common enough name to be set by something else in the parent scope. Use JFJOCH_VERSION,
which is already in scope there.

The CUDA architecture note claimed the list "embeds no PTX". A bare entry in
CMAKE_CUDA_ARCHITECTURES emits SASS and PTX both, so the newest entry has been the
forward-compatibility path all along: on a GPU newer than anything listed, the driver
JIT-compiles that PTX at first launch. Adding sm_121 still buys native code on Spark
instead of a JIT, which is what the comment should have said.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The openapi-generator invocation still passed --git-host=git.psi.ch and a user id of
jungfraujoch, from before the move to gitea.psi.ch/mx/jungfraujoch. Those properties are
not cosmetic: they become the source URL in the generated README and pyproject, so the
published client documentation - docs/python_client/README.md, which is copied out of the
generated tree - told readers to pip install from a host that no longer answers.
Regenerating with the corrected flags changes those two lines and nothing else, verified
against the committed tree.

update_version.sh, make_doc.sh and gen_python_client.sh were all mode 644, so the
"run update_version.sh" the documentation asks for fails on the shebang. CMake and the CI
both work around it by invoking them through bash.

make_doc.sh builds a throw-away venv in the working tree and deletes it on the last line,
which set -e skips whenever pip or sphinx fails - so a failed docs build left tmp_venv/
behind. Delete it from a trap instead, and ignore it along with the default output
directory and the sdist directory gen_python_client.sh creates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
--verbose was declared required_argument while -v carries no colon in the getopt string,
so the long form consumed whatever followed it. "jfjoch_writer --verbose tcp://host:5400"
swallowed the address as the flag's argument and then failed for want of a data source -
the short form was fine, which is presumably why it went unnoticed.

The usage line for the root directory asked for <int>. It is a path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
--no-scaling-corrections said it disabled the decay and absorption surfaces. It disables
the modulation surface too - GetCorrectionSurfaces() gates all three - and the flag has
worked that way since the detector-plane modulation was added; only its description did
not follow.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rugnux_scale and azint have not been separate binaries since they became rugnux --scale
and rugnux --azint-only; the viewer component carries rugnux, jfjoch_extract_hkl and
jfjoch_recompress.

Two job conditions tested github.ref_type against 'workflow_dispatch'. ref_type is only
ever 'branch' or 'tag', so that test never matched and never did anything: build-rpm's
whole condition was that test, and half of the unit-test one was. Both jobs ran on a
dispatch, as the release flow needs them to. Drop the dead tests rather than repair them
- the behaviour they read as intending is not the behaviour that is wanted - and write
down why unit-tests is skipped on a tag, which is the part that is deliberate: the
dispatch run is the one that tests, and the tag it creates only rebuilds and uploads.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Nothing said what is in a release or what it needs of the machine it lands on: that the
Linux binaries are built -march=x86-64-v3 and the Windows ones /arch:AVX, so each has a
CPU floor; that the portable .tgz is built on RHEL 8 for its glibc; that the Windows
installer is MSVC (Visual Studio 2026), CUDA 13.3, Qt 6.11 and carries the Qt runtime;
and above all what the CUDA variants need. Only cuFFT is linked dynamically, and it has
no link-time dependency on the driver library, so a CUDA build starts on a machine with
no NVIDIA GPU at all and falls back to the CPU path - as long as cuFFT can be loaded,
which the .tgz and the installer arrange by shipping it and the distribution packages
arrange through the distribution's own CUDA packages. Collected into a new page rather
than scattered over the install instructions.

The repository page had the RHEL 9 rows pointing at el8 paths under the wrong slsdet
number, no rows at all for the two slsdet9 repositories the pipeline uploads, a driver
package named jfjoch-driver where it is jfjoch-driver-dkms, and a note that RPMs are
unsigned from before the pipeline started uploading them with sign=true.

The FPGA page had a paragraph that stopped mid-sentence, in the middle of a link, and a
section describing a firmware build triggered by commit message. The firmware is stable
and carried from version to version now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rugnux gained --model - R-free and 2Fo-Fc/Fo-Fc maps against an atomic model, and with it
the resolution of the enantiomorph and of a merohedral indexing ambiguity - without the
page ever mentioning it. It was the only option missing; the two lists now agree in both
directions, checked against the usage the binary prints.

The viewer page still said results are never saved and that no Windows package exists.
Both have been false for a while: the Processing panel runs full rugnux jobs on the open
dataset, writes _process.h5 and the merged reflections, registers each run as a
selectable view so runs can be compared, and can hand out the equivalent command line for
a cluster; and the installer is published with every release. The mask menu also loads
TIFFs now, and the View menu has layout presets.

The writer page documented -R for the root directory, which is the back-compatibility
alias for -d, and an HTTP status interface that no longer exists - status reaches the
broker over the writer notification socket, and a writer is stopped with a signal.

The test page pointed at .gitlab-ci.yml and at jfjoch_offline_process, which is not a
binary any more; the CrystFEL fixture pointed at HDF5DatasetWriteTest, which is not
either. The broker page linked ../broker/redoc-static.html, which MyST resolved by
copying the 700 kB file into _downloads/ rather than using the copy already in _static.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Covers the Debian package rename and the DKMS scripts, the writer's --verbose, the
version plumbing, and the documentation pass - the release-contents page, the corrected
repository URLs and package names, and the rugnux, viewer and writer pages.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
bitshuffle_hperf is an x86-only implementation - its entire vector body sits behind
__i386__/__x86_64__, so on aarch64 every entry point compiles down to the scalar
fallback. Measured against its own SIMD path that costs 8.3x on encode and 3.8x on
decode, and it is the transform behind every compressed image the writer produces and
every one the reader, preview and XDS plugin take apart again.

The classic bitshuffle vendored beside it does have an aarch64 NEON path, and is already
compiled into the same target, so this costs nothing new. BitShuffleBlock.h picks
bshuf_trans_bit_elem/bshuf_untrans_bit_elem there and keeps bitshuf_encode_block /
bitshuf_decode_block everywhere else, where hperf is about twice classic SSE2 and remains
the better choice. The expected aarch64 gain is ~2.5x encode / ~1.7x decode: classic NEON
is 128-bit and carries an extra pass, so it recovers part of the gap rather than all of
it. The condition mirrors USEARMNEON in bitshuffle_core.c exactly, because with NEON off
the classic scalar path is slower than hperf's and must not be selected.

Swapping implementations is only safe while the two agree bit for bit - otherwise an ARM
build would write files an x86 build could not read. They do: verified byte-identical
output and mutual cross-decoding for elem_size 1/2/4/8 over block sizes from 8 to 65536
elements. Both are always compiled in, so the new test holds them to it on every
architecture, not just the one that would notice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The two front ends disagreed on the fixed-threshold spot finder: rugnux started
from 3.0, jfjoch_viewer from the SpotFindingSettings default of 4.0, so the same
file processed either way could give different spots.

Inert on the default path - the adaptive finder derives its threshold from each
image's own per-resolution-ring noise and never reads signal_to_noise_threshold
(only ImageSpotFinderCPU/GPU and DetModuleSpotFinder do). It changes behaviour
only under --no-adaptive-spots, and there it now matches the viewer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Offline reprocessing is not bound by the online spot budget, and the cap is
applied at the end of SpotAnalyze, so it is exactly the spot list the indexer
and the per-image refinement see. jfjoch_viewer already sends 1000, so the two
front ends now agree on the same file.

Measured as a paired A/B over the 37-crystal rotation battery, de novo, with the
resolution and Friedel setting matched to the XDS reference, both arms from the
same binary bar this constant:

  R_meas low shell   16 better    0 worse   19 unchanged
  R_meas             14 better    4 worse   17 unchanged
  ISa                14 better    6 worse   15 unchanged
  CC1/2               6 better    3 worse   26 unchanged

Low-resolution R_meas is a clean sweep. Around half the battery is bit-identical:
those frames never reach 250 spots, so the cap never bound. Wall clock is
unchanged (10m00s vs 10m44s, uncontrolled for page cache).

Known cost, and the reason this is its own commit: one crystal in the battery
reproducibly loses symmetry, tetragonal 422 -> orthorhombic 222, doubling its
asymmetric unit. Its R_meas and ISa "improve" there, but that is what merging in
too low a symmetry always does, and the lower symmetry then admits a merohedral
indexing ambiguity. An intermediate cap of 500 demotes it too, so it buys none of
the safety. This is the known point-group-decision-moves-with-data-amount
fragility of the space-group search rather than an argument for starving the
indexer of spots - the search is the thing to fix.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rugnux: fit the mosaicity from the strongest spots only
Build Packages / build:viewer-tgz:cpu (push) Successful in 21m50s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m21s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m34s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 24m7s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m34s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m30s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m37s
Build Packages / XDS test (durin plugin) (push) Successful in 11m30s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 22m6s
Build Packages / build:rpm (rocky9) (push) Successful in 21m35s
Build Packages / Generate python client (push) Successful in 43s
Build Packages / Build documentation (push) Successful in 1m17s
Build Packages / Create release (push) Skipped
Build Packages / DIALS test (push) Successful in 20m20s
Build Packages / build:rpm (rocky8) (push) Successful in 27m13s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m40s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m19s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m37s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m5s
Build Packages / Unit tests (push) Successful in 1h19m31s
Build Packages / build:windows:nocuda (push) Successful in 19m12s
Build Packages / build:windows:cuda (push) Successful in 22m29s
2c94f3013e
The per-image mosaicity MLE ran over the whole indexed spot list, so it rode
on --max-spots, which is an indexing budget. A spot is detected when
I_full * R(tau) clears the finder threshold, so selecting by intensity censors
on R(tau): a deeper list holds proportionally more large-|tau| partially
recorded spots and the fit widens with it. Raising the budget 250 -> 1000
widened sigma_M 0.059 -> 0.075 deg on a rotation dataset whose measured rocking
width says 0.054.

That is not cosmetic. An over-wide mosaicity mis-states every partiality in
scaling: forcing the mosaicity across that range moved the merge error model
from b 0.039 / ISa 26 to b 0.167 / ISa 6, and the space-group search lost a
genuine 422 with it, merging the crystal in 222 instead.

Cap the fit at the strongest 250 spots. FilterSpotsByCount leaves the list
strongest-first, so this selects exactly the spots a smaller --max-spots would,
and the mosaicity becomes invariant: 0.0538 deg at 250, 500, 1000 and 2000
spots, with the correct space group at each. Trimming or down-weighting the
tau tail does not work - the censoring is multiplicative in R(tau), so it
widens the whole distribution rather than adding a tail.

Battery over 37 crystals: exactly one change, the demoted crystal repaired
(33 space groups matching XDS -> 34). 23 of 37 are bit-identical, never
reaching 250 spots. Unaffected elsewhere: the default spot count is 250, and
stills have no goniometer so they return before the fit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rugnux: fit the profile radius from the strongest spots too
Build Packages / build:viewer-tgz:cpu (push) Successful in 18m20s
Build Packages / build:viewer-tgz:cuda (push) Successful in 20m23s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m35s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m47s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m24s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m35s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 29m9s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 19m23s
Build Packages / XDS test (durin plugin) (push) Successful in 10m17s
Build Packages / build:rpm (rocky9) (push) Successful in 20m45s
Build Packages / Generate python client (push) Successful in 33s
Build Packages / Build documentation (push) Successful in 1m7s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (rocky8) (push) Successful in 26m5s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 11m15s
Build Packages / DIALS test (push) Successful in 20m23s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m50s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m36s
Build Packages / XDS test (neggia plugin) (push) Successful in 10m3s
Build Packages / Unit tests (push) Successful in 1h17m43s
Build Packages / build:windows:nocuda (push) Successful in 16m24s
Build Packages / build:windows:cuda (push) Successful in 17m50s
457b1bfd1d
Same defect as the mosaicity in 2c94f3013, in the same file's sibling fit. The
profile radius is an RMS of the excitation error over whatever spots were kept,
and weaker spots sit further off the Ewald sphere, so it grows with the depth of
the list: measured over a 150 -> unlimited spot budget it rises ~60%, and on a
clean dataset as much as on a hard one, so this is general rather than something
one awkward crystal provoked. That made it a function of --max-spots, which is
an indexing budget, rather than of the crystal.

Its one consumer treats it as a membership gate (ewald_dist_cutoff is twice it)
where reflections at the cutoff carry near-zero partiality and are removed
downstream anyway, so the integrated data barely notices: with the mosaicity
already pinned, the partial count moves 0.2% across a 34% change in the radius.
It is also reported per image as a diagnostic, though, and a number that slides
with an unrelated setting is misleading to anyone comparing two runs - and it
would stop being benign the moment anything used it as a width rather than a
gate.

Battery over 37 crystals: no space group changes, 36 of 37 bit-identical, no
failures, one crystal marginally better.

Also in the comparison script: report XDS's mosaicity next to rugnux's. XDS has
two and they are not interchangeable - CORRECT.LP's REFLECTING_RANGE_E.S.D. is
post-refined, while INTEGRATE.LP's per-batch SIGMAR is its integration-stage
estimate, and the two differ by up to 2.3x. The XDS cell now prints both as
postrefined|MLE so a per-image estimate is compared against the one measured the
same way. Fixes a scoping bug in the same addition where every crystal read the
last directory's INTEGRATE.LP.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Geometry is re-refined independently on every frame, against that frame's spots
alone - as few as a dozen on a sparse crystal, where XDS fits its equivalent to
about sixty times more data. Measured over ten datasets the per-frame orientation
carries two components: a slow drift that is real, with rugnux and XDS agreeing to
R^2 0.83-0.88 on the two crystals that genuinely slip by 1.5 and 0.54 degrees, and
a fast jitter that is fit noise, scaling with spots-per-frame at exponent -0.79
where counting noise alone would give -0.5. The jitter is worth 1-8% on merged
intensities, 24% on the sparsest crystal.

It cannot be fixed by refining less. Turning per-image refinement off entirely
loses six space groups and a whole crystal, and even a 624-spot-per-frame crystal
collapses; dropping the beam-centre terms holds the space groups but is worse on
31 of 37 crystals. The freedom is earning its keep, so keep it and suppress only
the band that cannot be physical - a crystal does not re-orient and snap back from
one frame to the next.

So smooth the orientation in frame order after integration and recompute each
partial's delta_phi, and hence its partiality, from the smoothed lattice. Batching
at integration time was not an option: frames are processed independently and the
online path depends on that. This runs before the GPU upload, so the device path
picks it up with no separate kernel.

The window is chosen per dataset by leave-one-out cross-validation, because the
two components vary far too much for one number - drift spans 0.018 to 1.288
degrees and jitter 0.005 to 0.221, so any fixed window over-smooths one crystal
while under-smoothing another. Chosen windows range from +-1 to +-20 frames. It is
capped: cross-validation scores how well neighbours predict a frame's orientation,
which on a barely-drifting crystal keeps improving with width, but the per-frame
fit is also absorbing a real per-frame systematic and smoothing too wide destroys
it - uncapped, one crystal chose +-60 and lost 16% of its ISa.

Battery over 37 crystals: space groups unchanged at 34 matching XDS, R_meas better
on 31 and worse on 6, low-resolution R_meas 30/7, ISa 26/10, high-resolution CC1/2
23/12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
docs: describe the per-frame geometry smoothing
Build Packages / build:viewer-tgz:cpu (push) Successful in 19m9s
Build Packages / build:viewer-tgz:cuda (push) Successful in 21m0s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 23m12s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m58s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m17s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m53s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m59s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 18m55s
Build Packages / XDS test (durin plugin) (push) Successful in 11m20s
Build Packages / build:rpm (rocky9) (push) Successful in 21m37s
Build Packages / Generate python client (push) Successful in 38s
Build Packages / build:rpm (rocky8) (push) Successful in 24m58s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m29s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 11m23s
Build Packages / DIALS test (push) Successful in 20m30s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m7s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m26s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 26m1s
Build Packages / Unit tests (push) Successful in 1h27m56s
Build Packages / build:windows:nocuda (push) Successful in 16m37s
Build Packages / build:windows:cuda (push) Successful in 17m39s
5eb386e333
Goes in §10.3 next to the per-frame scale and mosaicity smoothing, since it is
the same mechanism applied for the same reason, and trims the changelog line to
one sentence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Revert "rugnux: fit the profile radius from the strongest spots too"
Build Packages / build:viewer-tgz:cpu (push) Successful in 16m51s
Build Packages / build:viewer-tgz:cuda (push) Successful in 18m44s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 20m42s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 21m37s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 24m37s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 24m48s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 20m13s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 25m19s
Build Packages / build:rpm (rocky9) (push) Successful in 23m23s
Build Packages / DIALS test (push) Successful in 21m35s
Build Packages / Generate python client (push) Successful in 40s
Build Packages / build:rpm (rocky8) (push) Successful in 29m16s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2404) (push) Successful in 23m17s
Build Packages / XDS test (durin plugin) (push) Successful in 11m5s
Build Packages / Build documentation (push) Successful in 1m14s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 27m29s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m21s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m24s
Build Packages / build:windows:nocuda (push) Successful in 13m58s
Build Packages / build:windows:cuda (push) Successful in 16m6s
Build Packages / Unit tests (push) Successful in 1h18m59s
fb55645b81
Reverts the profile-radius part of 457b1bfd1; the comparison-script and
mosaicity-column changes from that commit are kept.

The cap was validated on the rotation battery, which cannot test it: the profile
radius feeds `ewald_dist_cutoff` in IndexAndRefine, and that is read only by the
STILLS predictors (BraggPrediction/BraggPredictionGPU). The rotation predictors
gate on the mosaicity window instead and never look at it. So "no space-group
changes, 36 of 37 crystals bit-identical" showed the quantity is inert for
rotation, not that capping it is safe - and the one regime where it does act was
never exercised.

Validating it needs the serial-stills battery, which is a much larger exercise.
Until then the arbitrary constant is not worth carrying in a code path nobody
measured.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The correlation stage kept only reflections with I/sigma >= present_i_over_sigma
(3.0). That statistic is taken on the P1-MERGED intensities, whose sigma is
floored at b|I| (Merge.h, SigmaWithSystematicFloor) so that ISa = 1/b is the
asymptotic I/sigma ceiling - no reflection in a merge can read above it.
Verified over the rotation battery: max I/sigma equals 1/b on every merge.

So a fixed cut is not a per-reflection test at all. Every reflection sitting at
the floor reads 1/b exactly, however strong, and on a merge whose ISa falls
below the cut NOTHING passes: every operator is left with no pairs, its CC is
NaN, and the point group collapses to 1. The predicate "search-merge ISa < 3"
identifies the affected crystals exactly.

It is latent today - no crystal in the battery starves on the shipped
integration background - but it fires on four as soon as an additive intensity
bias is removed, and it is not a data-quality verdict: the crystals it silences
have final merges at ISa 19-22 while their low-multiplicity search merge sits
at 3.5-4.0, just above the cut.

Cap the cut at the merge's own I/sigma quantile so the correlation stage always
keeps at least its strongest quarter. A no-op wherever the fixed cut already
keeps that many - the cut stays exactly 3.000 on healthy merges. Battery
unchanged at 34 space groups matching XDS / 3 differing, with merged
observations identical to 0.000% on all 37 crystals.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Running bare `./jfjoch_test` costs far more time than it is worth on a
developer machine, and CI runs the full suite on every push regardless; the
useful local run is the cases or tags covering the code that changed. Same
reasoning for the 37-crystal rotation battery - it belongs on changes that
plausibly move merged results, not on every edit.

Also record that pushing is the maintainer's decision, not part of "make the
change".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The r2..r3 background ring was averaged with a 10% SYMMETRIC trimmed mean. A
symmetric trim is not a consistent estimator of the mean of a right-skewed
(Poisson) sample: on a clean Poisson ring it sits ~0.1 ct/px BELOW the true
mean at every level, and with ~50 signal pixels in the r1 disk that
under-subtraction adds ~5 counts to every partial on every frame. Measured two
independent ways on four rotation datasets - stored background_mean against a
plain ring mean over the same pixels on reflection-free frames, and directly on
apertures that provably hold no reflection. Empty-aperture pedestal, counts:
plain mean -0.03..-0.20, 10% symmetric trim +5.05..+6.34, 4 sigma clip
+0.02..+0.54.

Replace it with a high-side-only sigma clip at mean + n*sqrt(mean), n = 4 for
monochromatic data. It rejects the same one-sided contamination the trim was
there for - better, in fact: a 40 px neighbour core at +100 ct shifts the trim
by +10.1 ct/px, because a symmetric trim collapses once contamination exceeds
~10% of the ring, versus +0.009 ct/px at 4 sigma. False rejection on a clean
ring is 0.04-0.39%. Broadband data keep their tuned 3 sigma clip unchanged. The
trim stays reachable with --background-trim for back compatibility; setting
either estimator clears the other, so they can never stack. --integrator boxsum
does not take the clip (matching what the shipped clip already did), so it now
uses the plain ring mean unless --background-trim is given.

The intensities get measurably more accurate: per-shell agreement with an
independent processing of the same images improves on 14 of 16 crystals
(weighted -0.0347, outermost shell 12/4), the outermost-shell R_meas NUMERATOR
- absolute scatter, not a denominator effect - falls 13.5% median on 16/5, and
CC1/2 in the outer shell improves on 14/7.

EXPECT <I/sigma> TO FALL AND EDGE R_meas TO RISE. Both are inflated by
information-free counts, so both get worse when the bias is removed; neither is
evidence against this change. That fingerprint is exactly how the trimmed mean
was accepted in the first place.

Known cost: over the 37-crystal rotation battery the de-novo space-group count
goes 34 OK / 3 DIFF to 33 / 4. The single regression is a two-lattice crystal
whose merge fails the absolute-sanity gate under either background (R_meas
63.5%, CC1/2 72.2%) and which carries an unresolved indexing ambiguity on the
very operator being scored, so its operator CC is diluted by construction. No
other crystal changes space group, and twin protection is not weakened - the
H-ratio veto that refuses genuinely twinned crystals gets MORE decisive
(1.63 -> 1.84, 2.83 -> 3.99).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
--detect-ice-rings did two unrelated jobs at once: flagging ice spots so
indexing de-prioritises them and keeping ice reflections out of the scale fit,
AND gating the merge-time mask that drops a decorrelated ice ring and re-merges.
Turning it off to de-confound a merge-stage experiment therefore also changed
how the data were indexed - measured, that breaks indexing outright on two of
the 37 rotation battery crystals - while leaving it on lets the mask land
differently between two arms of an experiment and contaminate the comparison
(measured on up to 19 of 37 crystals in response to a small intensity change).

Add --ice-ring-mask[=on|off], default on, gating only the merge-time mask.
Verified with =off: ice-spot flagging and the scaling exclusion still log and
still apply, no mask line, no second merge, and the first error model is
bit-identical to the =on arm. The full pipeline and the offline --scale path
reach the same verdict on the same data, as they must.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The mask drops a hexagonal-ice ring when its merged half-set CC1/2 falls a fixed
0.05 below its resolution shoulders. That margin is not a significance level: at
the populations these rings actually have, 0.05 spans 1.1 to 7.3 sigma across
firings, and a nominal Fisher-z error understates the real scatter of these
heavy-tailed intensities by ~2.7x, so the null has to be measured rather than
derived.

Measured it with decoy bands - the identical ring/shoulder statistic evaluated
at q positions carrying no ice ring - over the 37-crystal rotation battery: the
gap's empirical null is p95 +0.032, p99 +0.095. So 0.05 sits near the 96th
percentile, about 4% of ice-free bands clear it, and roughly half the 22
observed firings are indistinguishable from bands with no ice in them. The
firing gaps are continuous, not bimodal, with 12 of 22 in [0.05, 0.10).

Raise it to 0.10, the 99th percentile of that null. Firings 22 -> 10, crystals
12 -> 5, decoy false-positive rate 3.4% -> 0.8%. An independent check against
XDS - which integrates through ice rings and so measures exactly what we delete
- agrees: of the firings with a usable comparison, 9 true / 9 false becomes
7 true / 2 false.

Battery: space groups 34 OK / 3 DIFF, the same three crystals as baseline, and
no other discrete decision changes on 37/37. The heavily iced crystal keeps all
five of its rings and its CC1/2 of 96.6; eight others recover 3.9-11.9% more
unique reflections and up to 10.4 completeness points. Cost is CC1/2 -0.84 on
one crystal, -0.35 on another, and agreement with XDS on the common reflections
worse by a median 0.0004.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The search already knows when several groups share an absence pattern - it
scores them identically, marks them all in the candidate table and prints
"Best space group: I23 or I213 (indistinguishable from these data)". The
one-line summary then dropped that and reported only the representative, so the
run's headline answer claimed a decision the diffraction had not made.

Carry the alternatives through to the summary. It already has them:
ProcessResult holds the whole SearchSpaceGroupResult.

  Space group:     I23 (No. 197) or I213 (No. 199) - indistinguishable from these data

Some of these pairs are enantiomorphs (P4_1 vs P4_3), where the choice needs
phasing or anomalous signal. Others are not, and are worth naming because they
surprise: I23 vs I2_13 and I222 vs I2_12_12_1 differ only by a screw whose
condition h00: h=2n is ALREADY implied by the I-centering condition
h+k+l=2n, so the screw has no observable signature at all. Checked over 35936
reflections with gemmi, the two absence patterns are identical - not nearly, but
exactly. Of the 65 chiral space groups, 13 classes are indistinguishable this
way, the largest being the four-way P3_112 / P3_121 / P3_212 / P3_221.

The representative stays the lowest space-group number, which is why a cubic
insulin comes out I23 where the deposited convention is I2_13. That choice is a
convention and the summary now says so instead of implying it was measured.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A change that touches partiality - a mosaicity estimator, a rocking-curve
model, a background change - cannot be judged by the statistics we normally
reach for, and this was learned the expensive way. ISa is anti-correlated with
external accuracy and is the largest mover of any statistic; last-shell R_meas
moves with its denominator, i.e. the wrong way by construction; `--model`
R-free tracks its own zero-information floor, which shifts ~22x more than
R-free itself over the same sweep; and per-shell agreement with XDS is biased,
because XDS never divides by partiality, so "divide less" moves us toward it
mechanically - measured to put the optimum ~1.4x too low.

Anomalous difference density at known scatterer sites has none of those
problems. It is read in units of the map's own sigma, so the uniform intensity
rescale a partiality change produces cancels exactly, and it is referenced to
the structure rather than to another program's partiality model.

The script runs SHELXC + ANODE per arm against a model placed ONCE and then held
fixed, and reports the mean site height, the off-site noise floor, and the
paired per-site change between arms. Numeric arm labels turn a set of arms into
a curve with a per-dataset optimum. The dataset table lives outside the
repository, as rugnux_vs_xds.py already does for the battery, because dataset
and sample identities are not committed.

It reproduces the measurements it was built from: all nine points of three
pooled curves, every per-crystal optimum, the site heights, the paired t
statistics, and the adversarial control in which a model refined against the
worst arm reproduces the curves to <=0.005 and the same optimum on 4/4.

Four things the ad-hoc scripts it replaces got wrong, all now handled:

* Keying sites on the ANODE atom label silently drops an alternate conformation
  sharing that label - one dataset class has 18 sulfur sites, not 17, and the
  uncorrected mean read 15.03 against a true 14.49.
* The off-site floor skipped any peak within 1.0 A of ANY atom, so a ripple
  sitting on a light atom was not counted as background; requiring 1.5 A from an
  anomalous scatterer raises one floor from 7.12 to 9.48 sigma.
* Special-position peaks are Fourier ripples, not background. Excluding them is
  load-bearing on 3 of 9 datasets and they are now reported in their own column
  rather than dropped silently.
* Enantiomorph care turned out to be unnecessary - passing the merged file's
  screw label to ANODE while the model sits in the other hand gives byte
  identical peaks. What does matter is the pair whose absences are identical,
  I23 vs I2_13, which phaser's automatic hand test does not cover.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ANODE names its ranked peaks after the SFAC element it was given, so the labels
read S1, S2, ... on a sulfur case but FE1, MN1, ZN1, SE1 as soon as the model
carries a heavier scatterer. The parser matched "S" followed by digits, so on
any such dataset it dropped the ENTIRE peak list and the dataset failed the gate
as "no peaks" - however strong its signal actually was. One heme case reported
no peaks when its true top peak is 13.79 sigma at 3.02x the off-site floor.

The bug is silent and it mis-gates exactly the datasets most likely to widen the
arbiter set, since a heavy scatterer is what makes a weakly diffracting crystal
usable as an arbiter in the first place.

Byte-identical on the standing all-sulfur set, verified by diffing the gate table
before and after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
SmoothGeometry de-rotates each frame's lattice to a common reference, averages
in frame order and rotates back. A frame that never indexed keeps a
default-constructed CrystalLattice whose vectors are all ZERO - and zero is
finite, so the isfinite guard let it through. Those zero vectors were averaged
into their neighbours' smoothed orientation, pulling it toward the origin, and
they were scored in the leave-one-out cross-validation that picks the smoothing
window.

On a crystal where 374 of 900 frames fail to index, the effect on the window
choice is not subtle. Measured:

  before   n_scored 900 (only 526 indexed)   CV score ~504-542 A^2   window +-12
  after    n_scored 516-526                  CV score  0.160-0.175   window +-2

The score was inflated ~3000x and the choice among windows was noise. It settled
on +-12 frames - 9.6 degrees of goniometer rotation - on a crystal whose
orientation genuinely drifts by ~8 degrees over the sweep, so every partial's
delta_phi was recomputed from an orientation averaged across that drift.

Require a real cell. Exactly inert when every frame indexes, and no threshold is
touched.

The crystal that exposed it goes P1 -> P2_1, observations 60107 -> 77021,
completeness 64.1% -> 93.0%, multiplicity 1.10 -> 2.0, CC1/2 70.0% -> 84.9%,
R_meas low shell 37.3% -> 22.1%, and its 2-fold operator CC 0.330 -> 0.669,
comfortably clear of the 0.5 gate. Battery over 37 crystals: space groups
33/37 -> 34/37, and that crystal is the ONLY flip - no losses. Another crystal
is rescued from near-total collapse (4402 -> 139213 observations) because the
two-pass "going back to the header geometry" fallback stops firing. Anomalous
peak height +0.043 +- 0.022 sigma over 7 crystals, so the background clip's gain
is intact. Merged quality is otherwise neutral (CC1/2 6 better/6 worse,
R_meas_lo 9/6) with observations up on 18 crystals.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two independent pieces in the same code path.

The background-estimate variance was never propagated. A reflection's background comes
from a finite ring of n_b pixels, so subtracting it adds var(B)/n_b per signal pixel -
sqrt(1 + n_d/n_b) = 1.109 with the shipped stencil. Both engines omitted it, which is
exactly the 1.11-1.19 gap measured between the off-ring scatter and the reported sigma.
Three lines each; it affects every dataset, not only iced ones.

The radial correction is new and OFF by default (--background-radial). The signal disk
and the background ring are concentric, so for any background LINEAR in position
<B>_ann == <B>_disk identically and a plane fit buys nothing; the leading error is the
CURVATURE of the radial background, which on a sharp ice ring reaches +26 counts on a
single reflection. Since every reflection uses the same stencil, that error is a fixed
kernel over radial offset - one short dot product per reflection and no extra pixel
reads. Validated on empty apertures before any C++: mean |bias| over 9 bands / 3
crystals 4.33 -> 0.79 counts with the scatter unchanged.

Three things it cost a battery each to learn, all now in the code:
 - the radial curve must be accumulated from CLIPPED annulus pixels, inside the clip
   pass, or it carries neighbour tails and zingers (so it is inert under --integrator
   boxsum, which has no clip pass);
 - the GPU version was a 1.8x slowdown from atomicAdd contention on a small radial
   array - staged in shared memory per block it now costs nothing measurable;
 - it is battery-NEUTRAL as a default, because the reflections whose bias it fixes are
   the ones the ice handling already excludes. Hence off by default.

CPU/GPU parity extended with two radial sections: 9002 assertions.

Also fixes a latent French-Wilson quadrature collapse: j_max = I + 8 sigma on a fixed
400-point grid degenerates to a single cell once sigma >> 50 <I>, giving F = 0.1 sqrt(sigma)
with sigmaF -> 0. Harmless today, but any sigma-inflation scheme detonates it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The per-image ice score was read off the PLAIN azimuthal profile. That profile is a
per-ring mean, so a few strong Bragg reflections landing in a ring's q bin lift it
exactly as ice would. Measured over 37 rotation crystals, that did not merely add
noise - it INVERTED the metric: the two highest-scoring crystals had no ice at all
(4.23 and 4.06), while a clean control read 1.57. A decoy null - the identical
statistic evaluated at q positions where hexagonal ice cannot be - reaches 1.51 at its
99th percentile and 2.70 at its maximum, so that metric cannot support any absolute
threshold whatsoever.

The adaptive spot finder already computes the right input for its own threshold: a
sigma-clipped per-resolution-ring background, in the same bins. A powder ring is
azimuthally smooth and survives the clip; Bragg peaks do not. On the clipped profile
the clean population tightens to 1.00-1.22 and the crystals with confirmed ice sit at
2.08-2.37, against a decoy null that never exceeds 1.29.

That channel is blind to one thing: ice in large crystallites diffracts as DISCRETE
spots and leaves the radial profile flat. So a second channel counts found spots on the
rings against the same q width of ice-free flanks beside them. The two barely overlap -
the smooth-ice crystals read 2.1-2.4 / ~1.0 and the textured ones ~1.1 / 3.8-17.6,
while a clean crystal reads 1.04 on both.

Both are then used as a GATE (--ice-min-score 1.5, --ice-min-spot-ratio 2.0, both
calibrated on the battery, 0 disables): the eleven fixed hexagonal bands cover 16-26 %
of the unique reflections at typical resolutions whether or not the crystal has ice, so
flagging, the exclusion from the scale fit and the merge-time CC1/2 ring mask are now
all skipped when neither channel sees any. The gate is applied in the full pipeline and
in --scale, which reads the stored per-image values back out of the _process.h5.

Also fixes the merge-time mask's control: the shoulder now excludes reflections that
are themselves on an ice ring. The rings are not evenly spaced - 1.947/1.916/1.882 A
sit 0.05-0.06 apart in q - so for those three the [w,3w) shoulder landed squarely on
the neighbours and the test compared ice against ice. Measured, that is the only thing
this changes: it removes firings on those three rings and leaves every other firing's
CC pair identical to three decimals.

And the online ice half-width, which was 0.02 in the API against 0.03 offline, so the
same data got a narrower band online than the measured ~0.06 ring FWHM justifies.

Battery (37 rotation crystals, against the previous behaviour): space groups 34/37 in
both and NO crystal's space group changes; 6 crystals gain unique reflections, 1 loses.
Best of them gains 7082 unique reflections with R_meas 16.0 -> 14.3, CC1/2 95.9 -> 97.3
and ISa 13.7 -> 19.0; another goes R_meas 54.9 -> 42.9, CC1/2 84.0 -> 90.4, ISa
3.9 -> 5.5; a third reaches CC1/2 99.4 from 95.7 at an unchanged reflection count. The
one crystal that loses reflections improves on both R_meas and CC1/2.

Not done here: the ScanResult/API/plot-type/frontend/viewer layers for the new
spot_count_ice_control (they need the OpenAPI regeneration). Message, CBOR, HDF5
write/read and the receiver plots are.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Post-refine: report the goniometer rotation scale it already fits
Build Packages / build:viewer-tgz:cpu (push) Successful in 20m2s
Build Packages / build:viewer-tgz:cuda (push) Successful in 20m10s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m20s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m33s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m3s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m7s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m50s
Build Packages / XDS test (durin plugin) (push) Successful in 11m12s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 21m35s
Build Packages / build:rpm (rocky9) (push) Successful in 21m32s
Build Packages / Generate python client (push) Successful in 39s
Build Packages / Build documentation (push) Successful in 1m29s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (rocky8) (push) Successful in 26m3s
Build Packages / DIALS test (push) Successful in 20m19s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m18s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m11s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m17s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m2s
Build Packages / Unit tests (push) Successful in 1h18m44s
Build Packages / build:windows:nocuda (push) Failing after 12m15s
Build Packages / build:windows:cuda (push) Failing after 11m57s
17eb6ef091
A stage that turns further than commanded is invisible in the file, because the stored
omega values ARE the commanded ones - both XDS and rugnux then read the discrepancy as
the crystal drifting. Measured on one dataset in 37, a ~1.2 % over-rotation costs it
unique 9.9k -> 29k and CC1/2 68 -> 98 % when corrected by hand.

No new degree of freedom is added, because the one needed is already there and being
thrown away: step A's residual rotates by -angle_rad * axis[] with axis an UNNORMALISED
3-vector, so the length it fits IS the factor by which the stage actually turned.
GoniometerAxis::Axis then normalises it away (with the `increment *= len` line sitting
commented out). This only reports it.

Guarded by the same cross-validation that gates the cell move - a fold that merely
soaked up noise cannot raise the flag - and by a 0.5 % tolerance, which is where a
direct scan of this factor puts 36 of 37 datasets (all at exactly 1.0000). The known
fault reads 1.00604 and warns; clean controls read 0.99958 and 0.99954.

It UNDER-reads the true magnitude: the fit only sees reflections already indexed at the
nominal angle, per-frame orientation refinement has absorbed part of the error, and the
axis components are bounded. Treat it as a detector, not a calibration - nothing here
corrects the data.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Whenever the merge-time ice-ring mask dropped a band, the per-shell observation
count and hence the reported multiplicity were wrong. On one crystal the lowest
resolution shell read 40780 observations over 1932 unique reflections - 21.1x -
where the truth is 27007 and 13.98x, and the overall redundancy read 12.52
against 12.29. Only counts were affected: intensities, sigmas, R_meas, CC1/2,
completeness and ISa were right throughout, because a masked group carries
merged_I = NaN and never enters those sums.

It looked like double counting and was not - it is a MOVE. Two independent
faults, both in three lines:

total_obs rides on the R_meas re-walk, whose filter deliberately ignores the
ring mask (and, on a search pass, the ice flag) so that R_meas is computed on
the same reflections either way. RmeasUsable therefore differs from MergeUsable
by exactly those two tests, and the observations they admit were being counted
against a `unique` that excludes them.

On the GPU path that count is binned by the GROUP's resolution, and a group
every one of whose observations is masked never has one written - acc[g].d stays
NaN. ResolutionShells::GetShell(NaN) then returned shell 0 rather than nothing:
NaN fails both bound comparisons, falls through to the arithmetic, and
static_cast<int32_t>(NaN) is INT_MIN, which the clamp maps to 0. So the masked
ring's observations were re-labelled into the lowest-resolution shell, four
shells from the ring they came from.

The two paths disagreeing on the same run is what settled it: with the mask on,
the GPU statistics gave shell 0 = 752 and the CPU statistics 423, while the
merged intensities were identical.

Count the merged population instead - acc[g].nh, which the merge already
accumulates per group - and guard the CPU increment with usable_merge. The
rnusable skip stays: any group present in the merged output has at least one
observation passing MergeUsable, and MergeUsable is a subset of RmeasUsable, so
it cannot drop a group that contributes to `unique`.

With the mask off and for_search false the two predicates are identical, so this
is provably inert on every shipped configuration - demonstrated on four
configurations, including one where ice handling is active but the mask does not
fire: the statistics blocks are unchanged. (The reflection lists differ in the
last ulp on 3-12% of lines, but so do two runs of the same binary; that is the
known rotation nondeterminism, and the statistics block is what is stable.)

The NaN guard also removes a silent contamination nobody was looking for. Four
call sites validate a resolution with `d <= 0`, which NaN passes: the Wilson-B
fit and per-shell <I/sigma> (CalcISigma), the per-image resolution plot
(SpotUtils) and the shell Wilson prior (FrenchWilson) were all binning
non-finite d into their lowest-resolution shell. French-Wilson now falls back to
the global mean rather than to that shell's, which is the worst prior available.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three defaults, each settled by measurement rather than by argument. The
arbiter throughout is structure-referenced - anomalous peak height where a
crystal can carry it, and otherwise the agreement of the ice bands with a fixed
external model against resolution-matched DECOY bands carrying no ice. The
band-versus-decoy contrast is used because R-free here tracks completeness, and
every one of these switches moves completeness.

The damage is real and it localizes: over the rotation battery the ice bands'
excess amplitude reaches +9.6% on a smooth-ice crystal and +35% on the worst,
while a clean control sits at +0.6% (z +0.45). On the worst crystal, nine of the
ten largest excess peaks in a q scan land on hexagonal ring positions. Turning
ice handling off leaves the contrast unchanged and forcing it on a clean crystal
does not create one, so it is the ice and not the machinery.

MERGE-TIME RING MASK -> OFF. It deletes reflections, which no other program does
by default - AIMLESS, DIALS, xia2, XDS and CrystFEL all keep ice-band
reflections in the merge and exclude them only from the model fit; autoPROC is
the sole exception. On the one battery crystal where the mask fires and an
anomalous arbiter can score it, dropping the band moved the mean peak height at
the known sites by -0.001 +- 0.018 sigma, 2% of the site height, while removing
1149 unique reflections whose mean I/sigma was 3.62 against the dataset's own
3.05 - better than average data - and costing 17 completeness points in that
shell. It fires on 5 of 37 crystals, changes no space group, and those 5
disagree in sign: it clearly helps the two most heavily iced, is a wash on two
and costs a third. So it stays as a switch, worth setting by hand on a badly
iced crystal where it shows in the high shell, but it is not a default.

RADIAL BACKGROUND -> AUTO, gated per image. The correction models the background
as a function of radius alone, and that is exactly when it works. On a crystal
with pure smooth powder ice it removes 43% of the bands' excess amplitude, with
the improvement 7x larger inside the bands than outside; on a crystal whose ice
is discrete crystallite spots - no smooth ring to model - the excess amplitude
GREW by half; on clean data it is inert to four decimals. The two ice channels
already separate those morphologies, so --background-radial takes on|off|auto
and auto applies it to an image when that image's peak-excluded score reaches
--ice-min-score. Auto never engages without such a score, because the plain
profile carries the Bragg peaks and cannot support an absolute threshold.

Per image rather than per run, and that was tested rather than assumed: the
gate fires on 100% and 94% of frames on the two crystals that want it, and on
1.5% of frames - 32 blocks, 23 of them single frames - on the textured-ice
crystal. A seam statistic against off + f*(on - off) is null on both mixed runs,
every merge statistic is bracketed by the pure arms, and the textured crystal's
auto arm lands on `off` rather than on `on`'s harm. A run-level gate would need
the score before the pass that integrates, i.e. rotation-only plumbing, and buys
nothing measurable.

The kernel was already built unconditionally, so flipping the flag per image is
free - except on the GPU, where the launches were gated on a construction-time
n_rad. That is why the buffers are now allocated whenever the correction could
run, and Run() decides per image.

DETECTION -> the geometry's default when the file is silent: on for rotation,
off for stills, with the command line and then the file taking precedence. A
rotation sweep sits on the same rings for the whole run, so ice there is a
coherent systematic and the presence gate keeps it inert on a clean crystal; a
serial stills run has too few spots per image to spend any on flagging. The
master file's key is kept as written rather than collapsed to a bool, so "the
file said nothing" is distinguishable from "the file said no" - it used to fall
silently to off, taking the exclusion from the scale fit with it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
First-pass rotation indexing reuses the spots the acquisition wrote, and does
NOT re-mark them - their ice flags come from the file. So --detect-ice-rings can
only invalidate those spots when it asks for something the file did not do.

It was flagged as a spot-finding option, which forced reuse_rotation_spots off
whenever it appeared, so naming it swapped the acquisition's spots for this
program's own and moved the first-pass lattice by itself. Measured over the
37-crystal rotation battery, passing the SEMANTICALLY NULL --detect-ice-rings=on
to files that already carry detect_ice_rings=1 changed the merged data on every
crystal, sent one crystal's ISa from 1.66 to 0.38, and lost MyoB_13 to indexing
failure outright. Both arms of any A/B on the flag therefore moved for a reason
that had nothing to do with ice, which made the flag impossible to test.

Re-find only when the requested value differs from the file's, and say so when
it happens. A file with no key at all counts as "did not mark", which is what
its stored spots show - such a dataset carries no per-spot ice flags to reuse.

Verified over the battery, with the merge mask and the radial background pinned
off so this is the only variable. --detect-ice-rings=on on the 36 keyed
crystals: all 36 log "using the spots stored in the file", the re-finding line
appears nowhere in the arm, unique counts are identical on every crystal, the
largest mean |dI|/sigma is 3.4e-5 against a repeat-run floor of 9.3e-5, and not
one of R_meas, CC1/2, CC1/2_hi, ISa, completeness, SigAno, d_min or space group
differs anywhere. MyoB_13 indexes again. The one file carrying no key reuses
under =off and re-finds under =on, as it should.

--detect-ice-rings=off still re-finds, since it does differ from those files,
and two crystals still fail to index there. That is not this change: a control
that re-finds with ice marking ON indexes both. With the marking off, ice spots
are no longer ordered last, so they consume the --max-spots budget and the
first pass collapses.

With the confound removed the flag can finally be measured, and on a comparison
whose spot source is identical on both arms it is clearly worth having - though
the win is at INDEXING rather than at the scale fit. Three crystals are saved
outright (one would otherwise collapse to P1 at CC1/2 19%, one loses half its
completeness and its screw axis, one loses its F-centred cubic lattice), two
more only index with it on, and the remaining eight gate-fired crystals differ
by well under 1% in R_meas and CC1/2 in both directions.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ice handling was gated on a measurement the run only made AFTER the images had
been processed, so the per-image pass could not use it. The flagging therefore
ran unconditionally: ice-band spots were ordered last in the --max-spots budget
and held out of the indexer seed and the geometry refinement on every crystal,
iced or not. The eleven bands are fixed geometry holding 16-26 % of the unique
reflections whether or not there is ice, so on a clean crystal that discards a
fifth of the spots - the strongest first - for nothing. Measured on a crystal
whose gate never fires, that moved the merged data by a mean of 0.85 sigma
against a run-to-run floor of 9.3e-5.

Measure it in the first pass instead. That pass already looks at ~100 images
spread over the sweep, and it already stops at the spot finder, so it sees the
azimuthal profile for the smooth channel and the unfiltered connected components
for the spot channel. Both counts SpotAnalyze takes are pre-filter, so pooling
them there is the run's own verdict, reached before anything has been discarded
and in time for the pass that acts on it. Where the sample sees no ice, the run
indexes on the ice-band spots too.

It has to be the whole sample: the spot channel is a ratio pooled over images,
because one frame carries a handful of control spots. A per-image gate is not an
alternative - two of the crystals whose indexing this rescues fire on that
channel alone, at profile scores of 1.12 and 1.22, so gating per image on the
profile score would drop exactly the cases that matter.

This also removes the first-pass spot reuse, and with it --redo-rotation-spots
and the reuse path. Finding the ~100 first-pass spots costs little, and reusing
was actively wrong here: the stored spots were found online at the acquisition's
threshold and have already had their ice-band entries ordered last and dropped
by its spot budget, so counting ice from them under-reads it by construction,
and the lattice search never saw the spot-finding settings at all. It also
removes the need for the machinery that re-found spots whenever a spot-finding
option was named, which made those options impossible to A/B.

IndexAndRefine cached index_ice_rings at construction, which happens before the
first pass; it holds a reference to the experiment, so it now reads the setting
where it uses it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
M_PI is not standard C++. MSVC defines it only when _USE_MATH_DEFINES is set
before <cmath>, so the radial background kernel's azimuth loop does not compile
there:

  error C2065: 'M_PI': undeclared identifier
  error C2737: 'phi': const object must be initialized   (cascade from the first)

GCC and Clang define it anyway, which is why the Linux build stayed green.
image_analysis is viewer-reachable, so it has to build under MSVC.

common/JFJochMath.h already carries a constexpr PI for exactly this reason - its
comment names this case - so use that. Same value to the last digit, so the
integration results are unchanged; the CPU/GPU parity test passes unaltered
(9002 assertions).

This was the only M_PI left in the viewer-reachable tree. The remaining uses are
in tests/, which Windows does not build (JFJOCH_VIEWER_ONLY is forced there).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
"Two-pass rotation indexing found only a lattice that indexes 2/60 validation
frames" tells a user that something went wrong and nothing about what. The
commonest cause is a metric symmetry promoted one class too far: the constrained
cell then misses every reflection by the small angle the constraint snapped
away, and the run dies with no way to see that a centred supercell setting was
chosen over the primitive one it should have kept.

Print the centring, crystal system and cell that was rejected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The ring calibration already here (AssignSpotsToRings + RingOptimizer, driven from
the viewer's powder panel) is given a SPOT LIST from a single image. A powder ring
is not a set of spots - it is a smooth arc - so a spot finder samples it wherever
its threshold happens to bite, and one image carries only the counts that image
collected. An azimuthally-binned profile summed over a run measures the same ring
directly, at every azimuth, with the whole run behind it.

RingsFromAzimuthalProfile turns such a profile into the (x, y, q_expected) triples
RingOptimizer already consumes, so nothing downstream changes: for each calibrant
ring and each azimuthal sector it fits the radial peak against a locally
interpolated background, and maps the measured (q, phi) back through the current
geometry to the pixel it came from.

What this is for is the BEAM CENTRE. A powder ring is a conic centred on the beam,
so a wrong centre makes its apparent radius oscillate once per turn and a detector
tilt twice - and neither depends on the calibrant's d-spacings or on the detector
distance. That matters, because the beam centre is otherwise the weakest parameter
we have: fitted from Bragg spots it is gauge-coupled to the crystal orientation,
which is why PostRefine has to restrain it toward the header and commit only a
sub-1 % move, and why XtalOptimizer carries a soft prior noting the beam is "only
LaB6-monitored to ~a few px". A ring does not know about the crystal.

Two things the peak fit is careful about, both of which would otherwise show up as
a spurious cos(phi) - i.e. as a beam-centre shift:

 - the sector's CENTRE is used, not its lower edge. GetBin() floors phi into the
   sector, so a bin stands for [j, j+1), and taking its edge rotates every ring
   point by half a sector.
 - a peak has to stand clear of the scatter of the background either side of it,
   or a sector with no ring in it contributes its largest noise excursion as
   though it were a measurement.

Refuses a single-azimuthal-bin profile outright: that is a plain radial profile,
the ring has been averaged over every direction, and there is nothing left to say
where its centre is.

Tested by round trip against the forward model, as the existing calibration tests
are: synthesise the profile the azimuthal integration would build with the rings
where a shifted geometry puts them but every pixel binned with the unshifted one,
then extract and fit. A 6.0 / -4.0 px beam offset is recovered as 6.13 / -4.03
from 192 ring points. Only the beam centre is exercised here; the tilt path is
covered by the existing DetGeomCalibTest round trips.

This is the extraction only - nothing calls it yet, and the run-scoped accumulator
it is meant to read (JFJochReceiverPlots::az_int_profile, already summed over a run
and written to /entry/azint/dataset) is still integrated with one azimuthal bin by
default.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A detector tilt does NOT appear as a cos(2 phi) modulation of the ring radius, as
the previous comment claimed. To first order a misalignment beta gives

    r(phi) = R + (R^2 / F) (beta_x cos phi + beta_y sin phi)

which is a cos(phi) term - the same harmonic a wrong beam centre produces. What
separates them is the radius dependence: the centre's amplitude is the same on
every ring, the tilt's grows as R^2. So they are told apart across rings, not
within one, and on a single ring they are exactly degenerate. Measured on a powder
standard the true cos(2 phi) term is of order R^3 beta^2 / F^2 - hundredths of a
pixel, at the noise floor - so it carries nothing usable.

Also add the tilted round trip, which was missing. It doubles as a check that
RingOptimizer's open-coded rotation agrees with DiffractionGeometry's: the fitter
applies Rx(-rot2) Ry(+rot1) by hand rather than going through the geometry's
Rz(-rot3) Rx(-rot2) Ry(+rot1), and those had never been held against each other.
They agree - 0.020 / -0.015 rad recovered as 0.0197 / -0.0148. Dropping rot3 is
right rather than an omission, since rings cannot constrain in-plane roll.

The tilted case yields fewer ring points than the centred one, which is expected
and worth knowing: the extractor searches a window centred on where each ring is
EXPECTED, so a large enough geometry error carries part of a ring out of it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The Bravais class is decided from the UNREFINED FFT candidate against a fixed
3 degree angular tolerance (LatticeSearch). A lattice that is pseudo-symmetric to
a few tenths of a degree is therefore promoted a class too far, and the constraint
then snaps a real angle to the ideal one - which throws nearly every reflection of
every frame out of tolerance. Measured on a monoclinic crystal that is
pseudo-C-orthorhombic to 0.42 degrees: the promoted cell indexes 2 of 60
validation frames and the run dies, where its own primitive cell indexes 39. It is
the same lattice in a different setting, b_oC = -(a + 2c), volume exactly 2.00x.

The perverse part is that BETTER SPOTS MAKE IT WORSE. LatticeSearch applied to the
true cell returns the promoted class deterministically; runs that succeed escape
only because the raw FFT candidate is inaccurate enough to miss the promotion
window. So it is bistable and non-monotone in every knob - 190 spots per image
gives 44/60, 195 gives 12/60, 200 gives 2/60 - and it will bite harder as spot
finding improves.

The indexer already refines a free triclinic cell alongside each constrained
candidate, but decides between them on the fraction of the accumulated first-pass
cloud that indexes, where the two differ by less than a factor 2 (measured 0.243
vs 0.135, missing both of that guard's bars). The caller has a far sharper
statistic: it already counts how many of 60 validation frames a candidate indexes,
and there the same pair differs by more than 20x. So keep the triclinic cell
instead of dropping it, and let the first pass settle it.

The bar is a clear majority, not a margin, and that is the part that took a
battery to get right: the unconstrained refinement holds NO cell parameter fixed,
so it can only index at least as many frames as the constrained one, and on
genuine symmetry it does index a few more. A 10 % margin - the bar a later scheme
needs to displace an earlier one - demoted a real I-centred orthorhombic crystal
to P1 (47 -> 54 frames) and perturbed an F-cubic one (49 -> 58). Only a
constrained cell that fails outright while its unconstrained cell works is
evidence of a false promotion, so demand exactly that. It is the same "fails to
index half the frames" test the long-axis rescue below already uses.

Battery over 37 rotation crystals: 33/37 space groups matching XDS with one hard
failure becomes 34/37 with none, and the other 36 crystals are identical in every
printed statistic (checked against a repeat run of the previous binary, which
itself differs on one crystal by one observation). The extra validation pass runs
only where the constrained cell already failed - 71 ms in a 15 s run - and not at
all on the 34 crystals whose constrained cell indexes a majority.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Changelog entries for the three changes above, and a new CPU_DATA_ANALYSIS
section on determining detector geometry from powder rings: why a ring is an
independent constraint on the beam centre (it has no crystal orientation to be
gauge-coupled to, unlike everything else that fits geometry here), what a ring
can and cannot determine, and how the ring points are obtained.

The section states the harmonics correctly, which is worth writing down because
the intuitive version is wrong: a detector tilt shows up as cos(phi), the same
harmonic as a beam-centre error, and the two are separated by the amplitude
growing as the ring radius SQUARED - so it takes at least two rings, and on one
ring they are exactly degenerate. The genuine cos(2 phi) term is hundredths of a
pixel.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The profile is the MEAN of each bin, so a few strong reflections landing in a bin
lift it exactly as a smooth powder ring does. That is the wrong quantity whenever
the profile is wanted as a background rather than as a measurement of what is in
the bin - the ice score being the case in point, where reading a plain profile
INVERTED the metric: over 37 rotation crystals the two highest-scoring crystals
had no ice at all.

The adaptive spot finder already computes the right thing, a sigma-clipped
per-resolution-ring background, as a byproduct of its own threshold. Where it
runs, the ice score uses that. Where it does not - --no-adaptive-spots,
--azint-only, and anything reading the profile the broker wrote - there was no way
to get it. This adds one: azim_int_settings.sigma_clip (rugnux --azim-sigma-clip),
0 = off, minimum 2 because a tighter clip rejects a large part of a clean Gaussian
bin and biases the estimate low rather than removing outliers.

Two clip passes follow the plain one, matching the finder's recipe - the first
pass's standard deviation is itself inflated by the peaks being removed, so one
pass leaves a threshold that is still too generous. A bin with fewer than eight
pixels is left alone: at the detector edge and behind the beam stop there is no
spread to clip on.

Both engines do it. On the GPU the accept range is computed by a small kernel and
stays resident, so a clip pass is one more read of the same pixels and no round
trip; the two accumulation kernels take the range as a pointer that is null on the
plain pass. Measured on a JUNGFRAU rotation dataset, non-adaptive path: azimuthal
integration 0.02 -> 0.06 ms per image, exactly the 3x the extra passes predict,
against a 0.34 ms per-image total.

Note what the result IS: the smooth background under the peaks, not the bin mean.
It should not be switched on where a ring's integrated intensity is wanted - the
powder-ring geometry fit reads ring peaks, and those are what a clip is designed
to remove. Off by default, so nothing changes unless it is asked for.

Not exposed over the REST API - that needs the generated model regenerated, which
is a separate step.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rugnux: say what the radial background correction decided, even in auto
Build Packages / build:windows:nocuda (push) Successful in 11m50s
Build Packages / build:viewer-tgz:cpu (push) Successful in 20m3s
Build Packages / build:viewer-tgz:cuda (push) Successful in 21m22s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m56s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 24m3s
Build Packages / build:windows:cuda (push) Successful in 14m55s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 27m44s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m4s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 29m18s
Build Packages / XDS test (durin plugin) (push) Successful in 12m4s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 22m1s
Build Packages / build:rpm (rocky9) (push) Successful in 22m49s
Build Packages / Generate python client (push) Successful in 38s
Build Packages / Build documentation (push) Successful in 1m11s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m10s
Build Packages / build:rpm (rocky8) (push) Successful in 27m56s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m37s
Build Packages / DIALS test (push) Successful in 21m31s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m1s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m53s
Build Packages / Unit tests (push) Successful in 2h30m11s
e9e3dac1b8
The "radial curvature correction on/off" line was printed only when
--background-radial was given explicitly. In auto - the default - it said nothing,
so a run's log carried no record of whether the correction had been applied.

That is not cosmetic. Auto decides per image from that image's ice score, so two
runs of the same data with different flags can differ substantially with nothing
in either log to explain it: a crystal whose high-shell CC1/2 read 8.1 % with the
correction pinned off and 4.5 % under the default looked like a regression for
some time before the flag turned out to be the whole difference.

Log the effective mode unconditionally.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Auto rode in with the ice work rather than on its own evidence, and measured over
the 37-crystal rotation battery it does not carry itself yet. It TARGETS
correctly - it fires on ten crystals and every one is ice-positive, no failures,
no space-group changes - but it costs 1.35x the wall clock (median +3 s per
crystal, worst +29 s) and on the merge statistics it is the familiar sign-mixed
trade: high-shell CC1/2 worse on three of the four crystals that move materially,
mean -0.76.

The case for it is real but rests on agreement with a fixed external model - 43 %
of the ice bands' excess amplitude removed on smooth ice, the effect 7x stronger
inside the bands than outside - which is the better arbiter and also the narrower
one. That deserves settling on its own, not riding along with a set of ice
defaults. `--background-radial=auto` keeps it a flag away.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Regenerate the API documentation from the spec
Build Packages / build:windows:nocuda (push) Successful in 14m15s
Build Packages / build:windows:cuda (push) Successful in 20m17s
Build Packages / build:viewer-tgz:cpu (push) Successful in 16m9s
Build Packages / build:viewer-tgz:cuda (push) Successful in 17m41s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 18m44s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 20m49s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 24m20s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 22m56s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 19m52s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 22m35s
Build Packages / build:rpm (rocky9) (push) Successful in 19m44s
Build Packages / build:rpm (rocky8) (push) Successful in 25m16s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m48s
Build Packages / Generate python client (push) Successful in 43s
Build Packages / Build documentation (push) Successful in 1m5s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (durin plugin) (push) Successful in 9m52s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 24m7s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m36s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m3s
Build Packages / DIALS test (push) Successful in 15m45s
Build Packages / Unit tests (push) Successful in 1h54m32s
1b5e2e85fd
update_version.sh at 1.0.0-rc.161. The only substantive change is the one that had
drifted: the spot-finding ice-ring half-width was still documented as 0.02 in the
generated Python client and its docs while broker/jfjoch_api.yaml has said 0.03
since the band was widened to the measured ring FWHM. Anyone reading the client
docs - or relying on the client's default when omitting the field - got a band
two-thirds the width the pipeline actually uses.

The TypeScript frontend client regenerates identically (the spec itself did not
move), and python-client/ is not tracked here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
docs: cover the azimuthal sigma clip, the indexing flag, and the metric-symmetry check
Build Packages / build:windows:nocuda (push) Successful in 17m17s
Build Packages / build:windows:cuda (push) Successful in 21m20s
Build Packages / build:viewer-tgz:cpu (push) Successful in 11m54s
Build Packages / build:viewer-tgz:cuda (push) Successful in 12m55s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 16m59s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 20m19s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 19m55s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 16m47s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 19m57s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 16m31s
Build Packages / build:rpm (rocky9) (push) Successful in 17m38s
Build Packages / build:rpm (rocky8) (push) Successful in 19m29s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 19m16s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 17m38s
Build Packages / Generate python client (push) Successful in 11s
Build Packages / DIALS test (push) Successful in 15m9s
Build Packages / Build documentation (push) Successful in 42s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (durin plugin) (push) Successful in 9m32s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m24s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m20s
Build Packages / Unit tests (push) Successful in 1h19m27s
cfd3697ddb
Three things had reached the code without reaching the documentation.

The azimuthal sigma clip had a RUGNUX.md row and a CPU_DATA_ANALYSIS section but
no changelog entry - and the only "sigma clip" the changelog mentioned was the
background ring's, which is a different thing at a different stage.

--index-ice-rings was in the options table but nowhere in the changelog, so the
entry describing the ice gate still implied that whether indexing uses the
ice-band spots is decided per run, which it no longer is.

CPU_DATA_ANALYSIS section 6 still described the Bravais class as simply "the
highest-symmetry class that matches within tolerances", which is the behaviour
that lost a crystal outright. It now records that the class is chosen from the
UNREFINED candidate against a fixed angular tolerance, that a pseudo-symmetric
lattice therefore gets promoted a class too far, and that the first pass settles
it on validation-frame counts with a clear-majority bar - including why the bar is
a majority rather than a margin, since that distinction is the whole reason the
check is safe. It also records that the first pass finds its own spots rather than
reading the acquisition's, which was not written down anywhere.

Also a build note: M_PI is not standard C++ and MSVC does not define it, so the
Bragg integrator's use of it broke the Windows viewer build.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CPU_DATA_ANALYSIS describes how the pipeline works; the account of which bar was
tried first belongs in the commit that changed it. Same facts, no narrative.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rugnux: --mode, and detector calibration from powder rings
Build Packages / Unit tests (push) Failing after 6m24s
Build Packages / build:rpm (rocky9_nocuda) (push) Failing after 14m5s
Build Packages / build:viewer-tgz:cpu (push) Failing after 14m34s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Failing after 14m54s
Build Packages / build:viewer-tgz:cuda (push) Failing after 16m14s
Build Packages / build:rpm (rocky8_nocuda) (push) Failing after 16m24s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Failing after 18m53s
Build Packages / build:rpm (rocky9_sls9) (push) Failing after 13m2s
Build Packages / build:rpm (rocky8_sls9) (push) Failing after 19m34s
Build Packages / build:rpm (rocky9) (push) Failing after 14m54s
Build Packages / Generate python client (push) Successful in 42s
Build Packages / build:rpm (ubuntu2404) (push) Failing after 14m10s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (durin plugin) (push) Successful in 12m14s
Build Packages / XDS test (neggia plugin) (push) Successful in 11m52s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 12m15s
Build Packages / Build documentation (push) Successful in 2m10s
Build Packages / build:rpm (rocky8) (push) Failing after 18m29s
Build Packages / build:rpm (ubuntu2204) (push) Failing after 17m55s
Build Packages / DIALS test (push) Successful in 17m4s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
6468dd13be
--azint-only and --scale are replaced by --mode mx|azint|scale|calibration, with
mx the default. The old flags are removed rather than aliased.

Calibration mode fits the detector geometry - PONI x/y, the two tilts and the
distance - to a calibrant's powder rings and writes a pyFAI .poni alongside a
report of how far each parameter moved from the header. Bragg data constrain the
beam centre worst, because it is gauge-coupled to the crystal orientation; a
powder ring has no orientation to couple to.

--calibrant takes lab6, agbh, ceo2, si or ice. A calibrant is a list of ring
positions rather than a unit cell, because hexagonal ice is P6_3/mmc: rings
enumerated from its cell would include systematically absent ones. So the
crystalline standards generate their rings from a cell and ice carries the
measured list, and RingsFromAzimuthalProfile, GuessGeometry and OptimizeGeometry
all take ring q. The calibrant table is shared with the viewer's powder panel,
which previously carried its own copy.

--calibration picks how the rings are measured: rings (default) sums the
(q x azimuth) profile over every processed image and fits the arcs in it; spots
pools the found spots and fits those. Both use the whole run, with -s/-e/-t
selecting images. rings defaults --azim-phi-bins to 32, since a profile with one
azimuthal bin has averaged the ring over every direction and cannot locate it.

Two fixes this exposed:

The extraction window is capped at half the gap to the neighbouring ring. The
background under a peak is taken from the ends of its window, so a window wider
than half that gap measures the next ring's flank as this ring's background -
and hexagonal ice has three rings within 0.06 1/A. Ice calibration was 3.5 px
out before this and 0.29 px after; LaB6 is unaffected.

RingOptimizer holds rot1/rot2 fixed when only one ring is present. A tilt and a
centre offset both move a ring as cos(phi) and are separated only by the tilt's
amplitude growing as the ring radius squared, so on a single ring they are
exactly degenerate.

Measured. LaB6 at five distances: the fitted direct beam is within 0.36 px of an
independent implementation out to 300 mm, and D = -0.046 + 1.000788 dtz with an
rms of 0.011 mm. At 500 mm one ring is fully on the detector and a second only
clips the corners, which is not enough to constrain a tilt - restricting the q
range to the resolved ring recovers 0.06 px. Ice: 5.53 -> 0.29 px on one crystal
and 4.71 -> 0.80 px on another, against XDS's refined direct beam. On an ice-free
crystal the fit is worse than the header, which is the correct outcome.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
common: expose the run-summed azimuthal profile object
Build Packages / build:viewer-tgz:cpu (push) Successful in 18m43s
Build Packages / build:viewer-tgz:cuda (push) Successful in 19m50s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 21m55s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m26s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 27m42s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 29m0s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m33s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 21m8s
Build Packages / XDS test (durin plugin) (push) Successful in 10m55s
Build Packages / build:rpm (rocky8) (push) Successful in 25m7s
Build Packages / build:rpm (rocky9) (push) Successful in 21m53s
Build Packages / Generate python client (push) Successful in 37s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 2m2s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m45s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m31s
Build Packages / DIALS test (push) Successful in 20m46s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m23s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m32s
Build Packages / Unit tests (push) Successful in 2h32m3s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
942e978ffc
Belongs with the previous commit - powder calibration reads the ring positions
off the accumulated profile, and without this accessor rugnux/Rugnux.cpp does not
compile. It was left out of that commit by a staging mistake, not by intent.

GetAzIntProfile() flattens the profile to an array; the calibration wants the
object's own GetResult(), which leaves a bin no pixel fell in as NaN. Flattened to
zero, such a bin reads as a deep hole in the ring rather than as no measurement.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
viewer: calibrate the whole dataset from "Analyze dataset"
Build Packages / build:viewer-tgz:cpu (push) Successful in 11m45s
Build Packages / build:viewer-tgz:cuda (push) Successful in 17m11s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 18m59s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 21m9s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 24m51s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 25m1s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 21m32s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 17m49s
Build Packages / build:rpm (rocky8) (push) Successful in 23m8s
Build Packages / build:rpm (rocky9) (push) Successful in 20m8s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m4s
Build Packages / XDS test (durin plugin) (push) Successful in 10m53s
Build Packages / Generate python client (push) Successful in 28s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m30s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 23m22s
Build Packages / DIALS test (push) Successful in 18m9s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m55s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m56s
Build Packages / Unit tests (push) Successful in 1h54m52s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
6194fe6fbf
The powder panel could only calibrate the image on screen. Calibration is now a
third page beside MX and AzInt, so the dataset button runs it over every image the
same way it runs the other two - which is the point, since a powder ring is
measured far better by summing a run than by one frame.

The page carries the calibrant and the method (rings or spots); the interactive
Guess/Refine buttons stay where they were and now share the one calibrant
selection, so there is no second combo to drift. analyzeDataset() carries the
ProcessMode rather than a bool: a third state was coming, and two bools would have
had one combination that cannot be valid.

The calibrant list gains ICE, which it could not offer before: the widget worked
in unit cells, and hexagonal ice has none that generates its rings correctly
(P6_3/mmc would include systematically absent ones). FindCenter now takes the ring
list its first line used to derive, so the interactive path gets ice as well.

The result window leads with the residual rms rather than the fitted sigma. The
sigma is a formal scatter estimate and understates a bad fit badly - measured on
ice, 0.215 px reported against a 1.70 px residual - while the rms separates a
usable fit from one that has locked onto the wrong thing.

A rings run needs the profile binned in azimuth; below four sectors it returns
nothing at all. The viewer raises the count to 32 exactly as the CLI does, and
says so in the panel and in the job dialog rather than doing it silently.

Also fixes a CLI inconsistency this comparison exposed: rugnux's calibration
branch never applied the standard offline analysis defaults, so it measured the
rings in a profile built with the file's polarization factor while every other
mode - and the viewer - uses 0.99. Found because the two disagreed by 0.005 px in
PONI x, and confirmed by reproducing the viewer exactly with --polarization 0.99.
With it applied the CLI and the viewer write byte-identical .poni files on LaB6
by rings, LaB6 by spots, and an iced dataset over 1800 images.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Powder calibration: write rot2/rot3 in pyFAI's frame, not ours
Build Packages / Unit tests (push) Successful in 1h23m35s
Build Packages / build:viewer-tgz:cpu (push) Successful in 11m23s
Build Packages / build:viewer-tgz:cuda (push) Successful in 14m56s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 20m33s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 18m7s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 19m59s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 15m33s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 19m50s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 19m7s
Build Packages / build:rpm (rocky8) (push) Successful in 20m39s
Build Packages / build:rpm (rocky9) (push) Successful in 18m55s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 20m18s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 15m27s
Build Packages / DIALS test (push) Successful in 13m49s
Build Packages / XDS test (durin plugin) (push) Successful in 8m43s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m52s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m59s
Build Packages / Generate python client (push) Successful in 12s
Build Packages / Build documentation (push) Successful in 42s
Build Packages / Create release (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 20m47s
Build Packages / build:windows:cuda (push) Successful in 23m0s
a72ba82f48
The .poni file carried rot2 and rot3 with our sign, which is not pyFAI's. pyFAI
has the slow axis increasing bottom to top; the MX convention runs top to bottom,
so the two frames differ by a reflection in y. Conjugating a rotation by a
reflection gives R(n, theta) -> R(Mn, -theta), so for rot2 (about x) and rot3
(about the beam) the sense reverses, while for rot1 the axis IS y and the axis and
the sense reverse together and cancel. Negate the first two, leave rot1 alone.

Poni1/Poni2 are unaffected: they are distances from pixel (0, 0) along each axis,
which the direction the axis runs in does not change.

Caught by integrating a LaB6 image in pyFAI with the file we had just written.
Unflipped, the rings come out BROADER than they do with no tilt at all - peak
height 42 against 30, mean ring-position error 0.0045 1/A against 0.0027 - which
is the signature of a tilt applied the wrong way. Flipped, they sharpen to 132 and
0.0005, and flipping rot1 as well makes it far worse (peak 5), so the asymmetry is
real and not a fitting artefact.

The unit test pinned the old signs, so it passed throughout. It now pins the
verified ones and says why, since the stored values and the written ones
disagreeing looks like a bug unless the reason is written down.

Only the exported file was wrong. Nothing internal changes: the fitted geometry
and everything downstream of it in Jungfraujoch were always self-consistent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The same frame mismatch as the rot2/rot3 fix, in the other two fields. Our pixel
coordinates are pixel-centred - 948.0 is the CENTRE of pixel 948 - while pyFAI
measures from the edge of the sensor and puts the centre of pixel i at (i + 0.5) *
pixel size. Poni1/Poni2 went out as beam * pixel size, so anything reading the
file placed the pattern half a pixel (37.5 um at 75 um pixels) off ours. The
previous commit's "Poni1/Poni2 need no such change" was right about the axis
directions and wrong about the origin.

The proof was already in the tree. The pyFAI reference values in
DiffractionGeometryTest were computed for a .poni with Poni2: 0.150 and a 75 um
pixel, which the tests translate to beam_x = 2000 - but pyFAI's numbers are
reproduced only at 1999.5. At 2000 every one of them is out by 2.6e-3 nm^-1, which
the 1e-2 tolerance hid. The tests now use the beam centre those headers actually
mean, and agree with pyFAI to 1e-6 - float precision - across untilted q, azimuth,
rot1, rot1+rot2, rot3, rot1+rot2+rot3 and the solid-angle correction. Tolerances
drop to 1e-4 (1e-5 for solid angle): ~100x the observed float noise, and 26x
tighter than the half pixel they were blind to.

The viewer's calibration window printed "PONI x = ... mm" from the un-offset value
beside the path of the file it disagreed with; it now matches the file.

Also moves the viewer's beam-centre cross half a pixel down and right, where the
spot, prediction, top-pixel and saturation markers already are. Our coordinates
are pixel-centred and the Qt scene's are pixel-cornered, so the map between them
is +0.5, and DrawBeamCenter was the one overlay missing it.

The convention itself is now written down in docs/DETECTOR_GEOMETRY.md, with the
conversions to XDS ORGX/ORGY and to the edge-of-sensor programs, this being the
second bug to come out of it.

Only exported and displayed values change; the fitted geometry, spot positions and
integration were always self-consistent. A .poni written by an earlier build is
half a pixel off.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The changelog is for users. It is not a place for commit messages or developer
notes, so: one line per entry, say what changed rather than why or how it was
arrived at, and no sample identities. Rationale and measurements belong in the
commit message.

Pushing gets its own section rather than a line inside Test, since it is a rule
about the repository and not about running the suite.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Removes azim_int_settings.sigma_clip / rugnux --azim-sigma-clip and the clipping
machinery in AzIntEngine. This is a partial revert of a6be35ccd - the ice-ring-mask
removal that commit also carried stays. Sigma clipping remains where it started and
where it is needed: inside the adaptive spot finder, at a fixed 3 sigma on raw
counts, feeding the detection threshold and the ice score.

The option made the workflow harder to reason about than the quantity was worth. It
gave azimuthal integration two meanings behind one setting - the bin mean and the
background under the peaks - which the azimuthal-integration workflows do not need.
It also did not compose with the fused GPU engine, which supplies the profile from
its PLAIN pass: on the default rugnux, viewer and receiver path the setting was
silently doing nothing (measured, the profile came out identical to the unclipped
run to 1e-6 with identical per-bin pixel counts). Making it correct is not a matter
of gating that one shortcut - it means separating the workflows (azimuthal
integration, MX rotation, MX stills, geometry calibration) and deciding per workflow
what the profile is for, which is a larger change than the option earns.

The default path is unaffected: over 20 images of a rotation dataset the radial
profile, the per-bin pixel counts and the spot counts are unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The GPU engine copies six small per-ring arrays back to the host every frame - the
clipped raw sum/sum2/count that the threshold is computed from, and the plain
corrected sum/sum2/count that become the azimuthal profile. They were plain
std::vectors, so the copies landed in pageable memory, and a device-to-host copy
into pageable memory blocks the calling thread until it has completed whatever
stream it was issued on. The profile snapshot sits between the plain pass and the
two sigma-clip passes, so Detect() stopped there and the device then sat idle while
the host caught up and enqueued the rest.

Register them, as AzIntEngineGPU already does with its own, and the copies are
genuinely asynchronous. Measured on a 4.5 Mpixel frame: 0.647 -> 0.621 ms per
frame. Nothing else changes - the spot list and the profile are unaffected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Fix the defects found reviewing the branch before merge
Build Packages / build:viewer-tgz:cpu (push) Successful in 19m32s
Build Packages / build:windows:nocuda (push) Successful in 19m57s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m45s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m38s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m24s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m8s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m9s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m18s
Build Packages / XDS test (durin plugin) (push) Successful in 11m10s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 20m21s
Build Packages / build:windows:cuda (push) Successful in 22m5s
Build Packages / build:rpm (rocky9) (push) Successful in 20m57s
Build Packages / Generate python client (push) Successful in 34s
Build Packages / Build documentation (push) Successful in 1m29s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m41s
Build Packages / DIALS test (push) Successful in 21m19s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m34s
Build Packages / build:rpm (rocky8) (push) Successful in 27m4s
Build Packages / XDS test (neggia plugin) (push) Successful in 10m19s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m58s
Build Packages / Unit tests (push) Successful in 1h17m36s
df9a9c2a2c
Image buffer: the per-image CBOR metadata headroom had been re-derived from the
online reflection cap alone, which cut it from 4 MiB to 2.55 MB while the measured
worst case - reflections plus the capped spot list plus the three azimuthal arrays -
is 2.9 MB, so the receiver dropped the frames with the most to say. Restore it and
give it a name that both the code and its guard test read: written down twice, the
two had drifted and the test kept passing against the value the code had left.

Spot finding: an unset low_resolution_limit means no limit at that end, as an unset
high_resolution_limit already did. An optional rather than a zero sentinel, because
zero is not a natural "no limit" here - every pixel lies above it, so the plain
comparison masked the whole image instead of none of it, and nothing validated the
zero. The API field is no longer required; a zero is folded into the unset case at
the boundary, where older clients still send it, so one spelling reaches the
analysis code. The FPGA takes its fixed-point ceiling instead, since ap_ufixed<16,9>
wraps above 512 A and would have masked everything.

image_preprocessing: check the CUDA calls on the fused decode path - the one new GPU
file with none, and the path fed by bytes we did not produce. An unchecked
synchronise returned the host-written sentinel as if it were a measurement, so the
decode looked successful and the fallback to the host decoder never fired.

rugnux: --stride no longer writes one past the end of the per-image arrays, whose
count floored where the worker loop ceils, and the written process file links the
images actually processed rather than the first N - each frame's picture now sits
next to its own analysis.

Powder calibration: the face-centred calibrants no longer list their systematically
absent rings, so the distance fit starts from a reflection that exists rather than
an extinct one; the triclinic calibrant covers both signs of h and k instead of a
single octant, which is only valid for a diagonal metric. The test asserted the old
behaviour - one ring formula for every cubic standard - and is rewritten.

CBOR: skip an unknown tagged value in the end block, as the other four blocks
already do. One advance lands on the tagged item rather than past it, so an older
reader fed a newer end message threw and never finalized its file.

Viewer: a settings value the setter rejects no longer escapes as an uncaught throw
from a worker slot, and the field offers only what the setter accepts.

Space-group search: judge stage B on the same "present" cut stage A already computes.
Merged sigma is floored so no reflection reads above ISa, so on a low-ISa merge the
fixed cut left both stage B tests unsatisfiable - every screw axis passed unchallenged
and the centering rescue switched itself off on exactly the weak data it exists for.
Where the fixed cut is the smaller of the two they are equal and this is inert: over
the 37-crystal rotation battery every crystal reports the identical space group and
identical merge statistics, so it is a no-op there and the low-ISa case it targets
remains unmeasured.

rugnux: --polarization reaches --mode azint, which parsed the flag and then dropped
it; that mode also applies the same polarization default as every other mode.

Acknowledge the ACTS/traccc project, whose sparse connected-component labelling both
spot extractors take their algorithm from, with its citation and its license.

The rc.161 change list is brought back to one line per entry, and the user-visible
changes that were missing from it added.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The systematic-b veto compares a candidate merge's fitted b against its
parent's, and both move with data quality. Removing genuinely bad observations
improved both merges but the subgroup more than the supergroup (parent
0.1644 -> 0.1480, candidate 0.3187 -> 0.3056), so the ratio crossed its 2.00
bound at 2.065 and a correct cubic promotion was refused - while the H
statistic, which has no sigma in it, did not move at all (0.898 either way).
Better data demoting a crystal is the wrong behaviour.

The veto now fires only where the H test has not confirmed the promotion. H is
the statistic that was measured to separate a real symmetry operator from a
twin law; b's genuine and twin ranges are interleaved. A twin fails both.

No bound moved and no option was added. Battery: 34/37 point-group agreement
with XDS before and after with no crystal changing; with the beam-stop mask
33/37 -> 34/37, the single change being a cubic crystal recovering its true
I23. Both real merohedral twins stay refused on H in every arm.

Gating the guards on the L-test / second moment was tried and rejected: those
indicators do not flag a real twin on the P1 pre-promotion merge, only after
merging in its true symmetry, so the gate promoted a twin into its holohedry.

Left alone deliberately: merge_systematic_b divides its reduced chi^2 by the
observation count rather than by the degrees of freedom, which inflates the
ratio more for small-orbit parents. Fixing it requires re-deriving all three b
bounds, which were calibrated on the biased statistic.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rugnux finds the beam stop and its holder in a projection of 60 images and
marks them in the pixel mask as bit 9 (--detect-beam-stop[=N|off], on by
default). Reflections behind the stop are attenuated but not flagged, so they
integrate low with a plausible sigma and nothing downstream catches them: the
signal-box gate requires 100% valid pixels and shadow pixels are valid, the
background clip is high-side only, and the |zeta| cut applies only to the
space-group search merge.

The detection compares each pixel's background against the typical background
at the same radius on two channels. An azimuthal one (the ring median) finds
the holder arm, which is a minority of its ring; a radial one (the background
just outside) finds the disk, which the ring median cannot see because inside a
fully blocked ring the median is the shadow itself. Pixels are pooled over a
5x5 box and tested only where the background has actually been counted, so
low-background data no longer masks the whole detector. Recorded reflections
are carved back out - a beam stop cannot block a reflection that was measured.

Bit 9 belongs to the run that found it, not to the dataset: it is cleared when
a run starts, so a mask read back from a file that carries one starts clear.
The user mask (bit 8) is left alone.

Scaling and merging gain a low-resolution limit, default 50 A
(--scaling-low-resolution <num>, 0 removes it), applied per observation before
scaling so it also protects the per-frame scale fit and the space-group search.
50 A is the value XDS configurations use; rugnux_vs_xds.py now matches both of
XDS's resolution limits instead of only the high one, so the lowest shell is
the same shell in the two programs.

The viewer draws the detected shadow in coral with a "Show beam stop" switch in
the side panel, exposes the low-resolution limit in the settings dock, and
offers detection in its processing jobs. Adding an image marker meant giving
the reader a MIN_REAL_PXL_VALUE, because several places classify a pixel by
range rather than by equality and would otherwise read the new marker as a very
negative intensity.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The background belongs to the beam and the shadow to the stop, and the two are
not concentric - fitting the stop edge per azimuth gives offsets of 13.4 px on
an 85.8 px disk, 22.2 px on 67.3 px and 6.9 px on 23.7 px, 8 to 33 per cent of
the stop radius on every crystal measured. The finder bridged that gap with a
radial envelope, the largest ring background over an outward window, used as the
reference for an individual pixel. That quantity exceeds the local background
wherever the background rises outward, so sound pixels near the stop scored below
the penumbra threshold and were masked. Measured against the fitted edge on a
long-distance disk stop, the mask was displaced rather than mis-sized: short by
up to 20 px on one side, over-reaching by up to 45 px on the other, with eight of
twenty-four azimuth sectors falling short.

The ring median is already the right reference wherever a ring still has
unshadowed pixels to measure, which is every ring except those lying wholly
inside the disk - and it needs no assumption about where the stop sits. So the
envelope is gone from the per-pixel test, and the rings it existed to cover are
handled directly: walking outward, a ring whose background is a fraction of the
background further out is shadow in its entirety. That comparison is only ever
asked whether a whole ring is inside the stop, never to judge a pixel, which is
where its failure mode lives. Blockage is deliberately not a counting test - on a
bright dataset the shadow interior is still well counted.

Detection is now one channel instead of two, and 113 lines shorter.

Measured: no azimuth sector falls short by more than 3.4 px, over-reach drops on
all three fitted crystals, and mask area moves by at most 0.04 per cent of the
detector on six crystals, so this corrects the shape rather than resizing.
Battery: space-group agreement with XDS unchanged at 34/37, median change in
R_meas and in the lowest shell 0.000 pp. The crystal that suffered worst when
masking was introduced recovers to its unmasked quality - R_meas 25.1 -> 17.2 per
cent, ISa 4.45 -> 10.04 - which is what removing the over-masking should do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The "median rocking width -> estimated mosaicity" line took an intensity-weighted
second moment of the frame-centre angles with max(0, I) weights, per event, then a
median over events. For a two-frame event that moment is exactly zero whenever only
one frame has I > 0 - probability 2/3 for a reflection carrying no signal - so on a
noise-dominated dataset the median lands in the degenerate spike and prints 0.0000.
Simulated against a known width it is wrong by 0.23x to 13x, in both directions, and
on a pure-noise null it returns a plausible-looking 0.06 deg.

est_mosaicity_deg was read nowhere, so nothing downstream was affected; the number
only misled whoever read the log. It was built to measure a signal for a mosaicity
refinement that was then abandoned, and the estimator that replaced it is the
per-image one that already drives prediction.

Report instead the frames per rocking event, which is what the block could honestly
say: near 2.0 the reflections barely rock, so the observed angle this refinement is
fitted to is under-determined. It is a geometry count, so noise cannot inflate it.

Also drop phi_rms_deg, which is never assigned anywhere, and Partial::zeta, which is
only written.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two changes to the same variance chain; they are in one commit because the second
exists to remove an assumption the first was breaking, and separating them leaves a
tree that is correct only by luck.

The reported sigma was floored at 2% of the intensity, a per-partial I/sigma cap of
50. It applied only to the box-sum seed, never to the profile fit, so the shipped
default was unaffected - but the combine back-derives each partial's non-signal
variance as sigma^2 - I, and a floored sigma makes that quantity mean nothing. It
then read corr^2 * (0.0004 I^2 - I), which is not a background variance. Measured on
--integrator boxsum: the reported sigma understated the true scatter by up to 16x at
I ~ 21000 counts per partial, and pooled_I amplified a 1 ct/px background drift into
an 11.5% intensity error on the strongest reflections.

What the floor stood in for - that at high intensity the error is systematic rather
than counting - is already carried downstream, twice: the fitted b in
v = a*sigma^2 + (b*I)^2, measured from the data rather than assumed, and
SigmaWithSystematicFloor on the merged sigma. The floor was that idea applied one
level too early with a hardcoded b of 0.02. It arrived without a test or a setter and
was unreachable from the CLI, the API and the config.

The merge now takes the non-signal variance the integrator actually measured instead
of inverting sigma^2 = I + N. That identity is exact for a box sum once the floor is
gone and was never exact for a profile fit, whose sigma^2 = 1/den + (wsum/den)^2 *
bkg_var is formed against a fitted intensity. The value is carried through
BraggFitResult, Reflection and Obs, both engines, both merges, and the process-file
round trip; files written before this change are read with the term absent, which is
what they had.

Battery, 37 crystals, paired: space groups unchanged, reflection sets unchanged,
median delta zero on R_meas and CC1/2. --integrator boxsum on the reference crystal
goes ISa 8.9 -> 20.2 with a 0.947 -> 1.032.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The beam is not constant. On one beamline it oscillates +-9.8% with a ~5.9-frame
period, confirmed four ways: the raw images, our own azimuthal-integration total, the
per-frame mean background, and XDS's per-image SCALE, which correlates +0.999 with the
first three. XDS removes it inside INTEGRATE, per image.

The fitted per-frame G could not: --smooth-g defaults to 5 degrees, which is 25 frames
at 0.2 deg/frame, so a 5.9-frame signal is smoothed away. Measured, the applied scale
carried 0.70% rms against a 9.8% modulation and correlated 0.66 with XDS's SCALE. The
only thing removing the oscillation was the refit on fulls, which acts after several
partials spanning most of a period have already been summed, so it removes the mean and
leaves the dispersion inside each event.

Take the flux from the per-frame mean background, gauge it to the run median, and divide
it out of rlp as the partials are ingested, so the fitted G sees only the residual and
smooth-G smooths only the residual. The background mean tracks our own azimuthal
background at r = +0.971 and XDS's SCALE at |r| = 0.93, with 95% of its detrended power
in the 3-8 frame band. Slower background movers - ice, a drifting shadow, absorption
against the goniometer angle, radiation damage - are still absorbed by G, which keeps its
low frequencies through the smoothing.

The applied scale now carries 9.41% rms at |r| = 0.93 against XDS. On the affected
dataset R_meas 9.9 -> 9.5%, low-resolution R_meas 6.5 -> 6.0%, ISa 12.8 -> 13.8, and the
anomalous peak height rises 0.423 +- 0.069 sigma over 18 sites (p < 0.001) - the only
significant move in the arbiter. Over the 38-crystal battery the space groups and the
merged reflection sets are unchanged and every metric has median delta zero.

A monochromatic dataset carries the same modulation at 2.0% rms, confirmed by the same
three proxies; the correction engages there too but no merged statistic moves at that
amplitude.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The rotation predictor and RotationPartiality used the mosaicity alone. Energy
bandwidth broadens a reflection's rocking curve as (dlambda/lambda)*tan(theta_B),
resolution-dependent and negligible at low angle, so on a large-bandwidth beam the
modelled reflecting range was too narrow exactly where the crystal still diffracts:
0.064 deg of broadening against a fitted 0.083, i.e. 26% at the detector edge. The
stills predictor has carried the term since it was written; only rotation was missing
it.

Add it in the three places that have to agree. The predictor widens both its
acceptance window and the partiality it hands to integration; the merge widens the
partiality it recomputes from the smoothed mosaicity; and the per-image mosaicity fit
subtracts the same term before fitting, so what it returns is the intrinsic mosaicity
rather than the mosaicity plus the beam. Without that last part the bandwidth would be
counted twice.

The term goes in without the 1/zeta of the usual expression: the erf already divides
by zeta, so adding a per-reflection width that itself carries 1/zeta would divide by it
twice - up to 20x at the minimum zeta. dphi = delta*tan(theta_B), and the zeta stays
where it was. The rotation identity dtheta/dphi = zeta was checked against a numerical
solve of the diffraction condition at four resolutions and three orientations.

Monochromatic data is untouched by construction - the term is guarded on a non-zero
bandwidth and is an assignment, not arithmetic, when there is none. Verified: 246456
reflections byte-identical through the predictor, 4.7 million rocking-fraction
evaluations with no bitwise difference, and identical merge tables end to end. The
bandwidth is read from the file (incident_wavelength_spread) or from --bandwidth, and
is absent from every dataset in the rotation battery.

On the bandwidth dataset the fitted mosaicity becomes resolution-independent
(0.0745 -> 0.0719 deg), the prediction window widens, frames per rocking event go
4.6 -> 5.3, per-image correlation to the merge rises 0.710 -> 0.725, and R_meas
improves 0.1-0.7 pp in every shell while CC1/2 falls 0.6-0.9 pp in the outer two.
Merged quality is net neutral: the combine normalises by sum(partiality), so a uniform
widening largely cancels.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The merged sigma was floored at b*|I|, so I/sigma could never exceed the reported ISa.
On one dataset every merged reflection came out at I/sigma <= 12.96 with a 99th
percentile of 12.77 in every resolution shell alike, while the scatter of the
observations implied about 44 and XDS reported 58.

The floor is wrong in principle. `b` is fitted from the scatter BETWEEN a reflection's
symmetry equivalents, i.e. from the part that is not common to them, so it averages
down with multiplicity exactly like the counting term. 1/sqrt(sum_w) with the
b-inflated per-observation sigma already gives b*I/sqrt(n); flooring at b*|I| puts the
sqrt(n) back. That is the whole effect: 12.96 * sqrt(21.6) = 60, against XDS's 58.

It was introduced on a comparison of our MERGED I/sigma against XDS's UNMERGED
I/sigma. XDS's own merged low-resolution I/sigma exceeds its reported ISa on 30 of the
39 reference datasets here, median ratio 1.78 and up to 4.23.

Merged low-shell I/sigma now lands where XDS's does: 22.4 -> 46.2 against 46.2 on one
crystal, 26.7 -> 115.7 against 96.6 on another, 12.5 -> 45.0 against 58.0 on a third.
Over the 38-crystal battery the space groups, the merged reflection sets, R_meas and
CC1/2 are all unchanged - every one of them is sigma-independent, which is what makes
them the right control - and <I/sigma> rises on 35 crystals with none worse.

The asymptotic estimator that fed the floor stays, for the reported ISa only, and is
repaired in the process: it subtracts a*sigma^2 rather than the raw sigma^2 (at a < 1
the difference is the same size as the b^2 being measured, which is what made it
flip between 10.9 and 62.7 on consecutive passes of the same data), it rescales each
group's variance median-unbiased before subtracting an unbiased counting term, its
I/sigma gate uses the same convention, and it is bounded by the whole-range b - an
asymptote exists to refine 1/b upward, not to report 0.3 because "strong" was selected
on a sigma scale the fit itself rejects.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The (a, b) fit ran over the whole merged range and the automatic resolution cutoff was
applied afterwards, so the sigma correction applied to the reflections that survive was
calibrated largely on reflections that do not. Measured on one dataset: a = 0.286
fitted over 843k reflections, 22k written. A manual --scaling-high-resolution already
restricts the population at ingest, so only the automatic path was affected.

Fit over the full range, merge, read the cutoff from that merge, refit (a, b) on the
samples the cutoff keeps, merge again. The circularity resolves by direction: the
cutoff comes from CC1/2, a correlation of the two half-set means, which the sigma scale
barely moves, so the cutoff can be read first and the sigmas calibrated on the
population it chose. One refinement, not an iteration; one extra merge pass.

Note this is invisible to the rotation battery, which passes an explicit high
resolution limit matched to XDS and so never exercises the automatic cutoff. With a
manual limit the fitted (a, b) are byte-identical to before.

The equivalent defect in the stills / offline --scale path is untouched; it is a
different engine and needs its own validation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The profile fit weights each pixel by 1/v with v = max(bkg, floor) + max(0, I)*P, where
I is the fit's own current estimate. Rectifying it means that at true zero the plug-in
is E[max(0,I)] = 0.4*sigma rather than 0, and with sum(P^3)/sum(P^2)^2 = 4/3 for a
Gaussian the reported sigma comes out about 0.2 counts too large - always, additively.
That is nothing at sigma ~ 7 counts and 11% at sigma ~ 2, so it only shows on data
measured against roughly one background count.

Clamp the whole weight instead of the intensity: v = max(bkg + I*P, bkg/2). Simulation
of the real integrator gives claimed/true sigma 0.92-1.01 at zero intensity across
backgrounds 0.02-2.0 ct/px and 1.000-1.007 above I = 30, where the clamp never binds.
Dropping the signal term entirely instead (v = max(bkg, floor)) is exact at zero and
wrong everywhere else - 1.91 at I = 5, 4.29 at I = 30, 13.3 at I = 300 - and a test
built on systematically absent reflections cannot see that, because it only measures
zero. Removing the clamp altogether overshoots and biases the intensity, since a
downward fluctuation shrinks v at the peak and over-weights it.

The pixel variance floor was 1/12, documented as the rounding of a continuous energy.
That does not describe a photon counter: measured on raw frames at 0.065-0.082 ct/px,
var/mean is 1.042-1.045, i.e. Poisson with no digitisation term, and a digitisation
term would be additive rather than a floor. What the floor really protects is the
background estimate, which a small ring can read as exactly zero, so it belongs at the
resolution of that estimate, ~1/n_bkg. At 1/12 it multiplied the reported variance by
floor/bkg below 0.083 ct/px - a factor of two at 0.04. Set to 0.01.

Measured on systematically absent reflections, whose true intensity is zero, as
std(I)/rms(sigma) binned by background - not std(I/sigma), which is deflated by the
correlation between the plug-in sigma and the reflection's own fluctuation. On 2.78 M
absent observations at 0.16-3 ct/px the ratio goes 1.04-1.07 to 0.99-1.00. On 2.58 M at
0.005-0.6 ct/px, decomposed: the clamp carries it above 0.08 ct/px, the floor below it.
Intensities move 0.4%; this changes sigma, not I.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Setting a bandwidth flipped three unrelated switches at once: it changed the profile's
radial capture term, it moved the width measurement from the signal disk to the whole
fit grid, and it silently overrode the background clip and trim, so --background-clip
under --bandwidth was ignored - the two runs were bit-identical.

The width measurement was the damaging one. The fit grid is an azimuthally averaged
stack, so its second moment is sigma_r^2 + sigma_t^2 and the radial smear of a
bandwidth leaked into the tangential model - a tangential width of 3.04 px against a
1.06 px truth, inflating the effective background pixel count where the weak signal is.
The result was a step rather than a slope: on genuinely monochromatic data, declaring a
0.2% bandwidth cost ISa 28.4 -> 22.2.

Measure the two widths separately, accumulated in each spot's own radial/tangential
frame over the signal disk, from the signed profile cells - away from the peak a
learned cell is background noise centred on zero, so the signed sum is unbiased, while
clamping it at zero turns that noise into a pedestal the r^2 weight reads as width. The
radial term is then the measured excess or the analytic floor, whichever is larger.

With the two widths separated there is nothing left for the broadband switch to select,
so it is gone - which is the proof the three were independent. The background clip and
trim now come from the settings in every case; the tuned 3-sigma broadband default
moves to the rugnux front end, which is the only place that knows whether the user gave
a value.

Monochromatic data: declaring a 0.2% bandwidth now costs ISa 28.4 -> 27.9 rather than
22.2, and forcing the old 3-sigma clip in the new build reproduces the good result, so
none of the step came from the clip. On large-bandwidth data CC1/2 improves in 8 of 10
shells. Across 12 monochromatic crystals the space groups are unchanged and CC1/2 moves
by at most 0.2 points.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
f4e281b2f described this change in full but committed only one of its six files.
What went in was RotationScaleMerge.cpp - the merge widening the partiality it
recomputes from the smoothed mosaicity. That is precisely the part which is unsafe on
its own, by the original message's own argument: without the mosaicity fit subtracting
the term before fitting, the bandwidth is counted twice, and without the predictor
widening its acceptance window, the partiality the merge recomputes no longer matches
the one integration measured.

Add the five files that were left behind: the rotation predictor and its GPU twin
widen the acceptance window and the partiality handed to integration, the settings
struct carries the term, and CalcMosaicityXDS deconvolves it before fitting so what it
returns is the intrinsic mosaicity rather than the mosaicity plus the beam.

Monochromatic data is untouched by construction - every hunk is guarded on a non-zero
bandwidth, which is read from incident_wavelength_spread or --bandwidth and is absent
from every dataset in the rotation battery. Verified on the one dataset that has a
bandwidth: at --bandwidth 0, the merge table is identical to the branch tip; with the
bandwidth set, the fitted mosaicity drops 0.0718 -> 0.0694 deg as the deconvolution
takes effect and CC1/2 in the outermost shell recovers 30.3 -> 31.4%.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
max_operator_h_ratio was 1.25. Instrumented over the rotation battery, the statistic it bounds reads
0.85-1.57 on GENUINE promotions - and 2.48 on a genuine orthorhombic step in an arm left short of
pairs - while the two real merohedral twins read 1.82 and 4.01. There is a wide empty gap between the
two populations and 1.25 was not in it: it sat inside the genuine range.

Four genuine promotions already exceeded it and survived only because the two-arm rule happened to
offer cover from the other arm; a cubic case with no such cover was refused outright, by a margin of
0.4%, and merged in the orthorhombic subgroup with twice the unique reflections. That refusal is
invisible to the standard battery, which passes an explicit resolution limit: the limit also
constrains the merge the search's internal cutoff is derived from, and lands it just under the
crossing. It appears only when the automatic cutoff runs.

Set the bound to 1.70, in the gap. On the automatic-cutoff arm the cubic case returns to its true
group (unique reflections 49277 -> 23330, CC1/2 in the outermost shell 25.7 -> 50.4) and no other
crystal changes symmetry. The 38-crystal battery at matched limits is unchanged, space group included.
All ten SearchSpaceGroup test cases pass, including the twin decision table and the H-margin case -
worth checking explicitly, because a looser bound also confirms the operator agreement more often and
so suppresses the systematic-absence veto more often.

The header note already predicted this failure mode: the test's known limit is angular coverage, not
data quality, and a refusal on a lopsided merge says more about the coverage than the symmetry.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The per-detector-plane modulation surface was learned on a 16x16 grid. Fitting the same surface to
rugnux's own symmetry mates and cross-validating on held-out frames shows the grid was the binding
constraint, not the data: held-out R_meas improves monotonically to 24x24 and then stops -
none 12.64%, 8x8 11.13%, 16x16 10.71%, 24x24 10.58%, 32x32 10.60%, 48x48 10.58%, 64x64 10.60%.
At d > 4.4 A the same ladder reads 5.29 / 4.92 / 4.86 / 4.66 / 4.72 / 4.64 / 4.74%. Frame-parity,
random 50/50 and 4-fold splits agree.

The structure being fitted is ours, not a reference program's: the same surface fitted to the other
program's observations of the SAME events moves it 7.10 -> 7.07%, against 13.46 -> 12.81% for ours,
and its amplitude is 6.5% robust sd against 1.2%.

Measured across seven crystals spanning multiplicity 3.7-9.4, two detector types and 75-100%
completeness, 24 never clearly hurts and mildly helps six of them; 32 adds nothing beyond it. An
earlier in-sample ladder suggested 48x48 was worth twice as much - that was in-sample, and it
overstated the gain about threefold.

Nothing else needs adjusting: the Tikhonov shrinkage already adapts to thinly-populated cells, and the
cross-validation gate already refuses the surface outright where the finer grid is too fine for the
data - on the weakest crystal tested its held-out gain falls 4.3% -> 3.1% -> 1.7% and the surface is
skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The P1 merge that feeds the space-group search is cut at the first 1/40 shell whose mean I/sigma falls
below 1. If even the lowest-resolution shell fails, the cut was abandoned altogether and the search
was handed the whole merge out to the detector corner.

That inverts under its own feedback. The bound is absolute, while the merged I/sigma it tests is
capped by the merge's own asymptote: once noise-dominated high-resolution reflections have inflated
the error model's b, no shell can reach 1 - so the case where the cut is abandoned is exactly the case
where the merge is worst. Measured on one crystal: b 0.239 -> 0.782, ISa 4.2 -> 1.3, no shell above
the bound, 46853 reflections into the search become 3989103, the added operator's agreement falls
0.978 -> 0.429, and a C-centred monoclinic crystal merges as triclinic.

Fall back to the lowest shell's own high-resolution edge instead, when that shell is populated enough
to define one. It is what the healthy case does anyway - shell 0 is the lowest-resolution fortieth of
reciprocal volume, 46852 reflections on the crystal above.

Only reachable when rugnux chooses its own resolution limit: an explicit --scaling-high-resolution
also constrains the merge this cut is derived from and lands it above the crossing, which is why the
battery at matched limits cannot see any of this. Both battery arms are therefore unchanged - space
group identical on all 38 crystals, every quality column identical, whether rugnux picks the limit or
takes it from the reference. On a deliberately degraded integration the crystal above returns to its
true symmetry (294655 unique reflections back to 102514, R_meas 41.4 -> 29.3%) and its error model
recovers.

Four crystals in the battery sit within a factor of two of the bound, and one sits 0.19% from it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The centering test compares the absent class's mean intensity against half the present class's. With
a present mean at or below zero - which happens on a merge dominated by noise - the bound is
non-positive, and the comparison stops measuring whether the absences are weak and starts turning on
the sign of the absent mean. Seen in an uncut merge: absent -0.16 against present -0.03, where a more
negative absent class passes and one nearer zero fails, both by accident.

Require a positive present mean before the mean-ratio branch can confirm a centering. The rate branch
below it counts violations rather than averaging intensities, so it cannot change sign, and it already
exists for exactly the weak-data case this leaves to it.

No crystal in the battery changes, at matched limits or with the automatic cutoff.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
var_bkg was added to Reflection and to the HDF5 writer but never to the CBOR reflection map, so it
survived only where rugnux drives the writer in-process. Everything that reaches the writer over the
wire - i.e. every acquisition the broker records - wrote /entry/reflections/*/background_variance as
an array of zeros, presented as a measured quantity, and re-processing such a file fed the merge a
non-signal variance of zero.

Encode and decode it. The key is optional on both sides, so a stream from an older version still
reads and one from this version still reads on an older client.

The round-trip test only checked h, k, l, the predicted position and d - which is why a missing float
was invisible. It now gives every field a distinct value and checks all of them, so the next field
added to Reflection and forgotten here fails immediately.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A _process.h5 written before background_variance existed was read with var_bkg = 0, on the reasoning
that zero leaves the combine with the signal term alone, "which is what it had before". It does not.
Before, the combine back-derived the non-signal variance from sigma itself, and on a weak reflection
that is essentially the whole of sigma^2; zero deletes the dominant term and weights the reflection by
roughly 1/I instead of 1/sigma^2.

Recover it from the integrator's own identity, sigma^2 = I + var_bkg, when the dataset is absent.

Measured by re-scaling a stills _process.h5 with the dataset deleted, against the same file with it
intact: the automatic resolution cutoff was reading 1.66 A where the intact file reads 1.81, with
14731 unique reflections against 11508 - i.e. the zeroed file looked good enough to merge 0.15 A past
its own limit. Reconstructed, it reads 1.78 A and 11926, within 0.03 A of the intact file. The
residue is the profile-fit path, where var_bkg is not exactly sigma^2 - I and only the box-sum
identity is exact; that is recoverable to a closer approximation only by storing it, which is what
files written from now on do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The signal disk and the r2..r3 background ring were fixed pixel circles, identical for every
reflection at every resolution. A reflection is not round: a finite bandwidth streaks it radially by
bw_sigma*Rpx, so at high resolution the ring sits within 1.3-2.2 sigma of the reflection's own
profile and measures its tails as background.

--integration-stencil <k> makes the RING an ellipse, elongated along the beam->reflection direction
by k times that streak, capped at 2*r3. The tangential half-widths stay r2 and r3, and the r1 signal
disk stays a circle: r1 drives the all-or-nothing n_inner_valid == n_inner gate, so growing it
rejects any reflection carrying one bad pixel along a long streak, and the flux a circular r1 loses
is a function of resolution alone, which the per-shell scale absorbs.

The geometry lives in one shared header compiled by both the host compiler and nvcc, so the seven
pixel-classification sites - the CPU mask/main/clip loops and the GPU mark_mask/main/trim/clip
kernels - cannot drift apart. Rather than evaluate an ellipse, each pixel's squared distance has its
radial part scaled down, d2 - q*rad^2 against r2^2/r3^2 with q = 1 - (r/(r+grow))^2, so grow = 0
gives q = 0 and both tests collapse onto d2 exactly in floating point.

The width is the bandwidth streak alone, not the profile's full radial variance, which also carries
the sensor parallax and weak-spot capture terms. Deriving the growth from those was implemented
first and measured on the rotation battery: at k=1 it took Thau_9's high-shell CC1/2 from 75.8 to
27.9 and Benas_3's from 14.1 to 6.0, against cytC_10 +1.2 and lyso_ref flat. On a monochromatic beam
they are the only terms there are, and C_CAPTURE is 64% of them. Keeping only the streak also makes
the option exactly inert without a bandwidth, rather than merely small.

Default 0. Measured on broadband rotation data with the bandwidth set to its spectroscopic value,
matched resolution limits: high-shell CC1/2 30.6 -> 46.4 at k=4, and better in EVERY shell in both
CC1/2 and R_meas (top shell R_meas 194.7% -> 138.7%), with completeness, multiplicity and space
group unchanged and 28 of 98833 unique reflections lost. Anomalous peak height over 18 sites
+0.107 +- 0.039 sigma (p = 0.013). The full 38-crystal rotation battery is unchanged to every
reported digit, base against k=3.

Two consequences of an elongated ring are handled rather than inherited. The neighbour exclusion
marks the inner ELLIPSE in each neighbour's own frame, or an elongated neighbour leaks its tails
into this reflection's ring. And the radial-background curvature kernel becomes a small table
indexed by the growth, because its azimuthal average makes one kernel serve every reflection only
while their stencils are identical; the GPU's radial window, previously a fixed 32 bins, is now
sized on the host from the widest ring on the detector.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The kernel table is sized and built only where the correction can ever run - explicitly on, or auto,
which is the same condition the GPU allocates its radial buffers under. BackgroundRadial(true) on any
other engine therefore asked the CPU to correct with a single CIRCULAR kernel for rings that may be
elongated, while the GPU, having no buffers, did not correct at all: a wrong kernel on one engine and
silence on the other, from the same call. Only the auto path calls it today, so it was unreachable,
but the setter is public and the invariant it depends on is not local to it.

Remember whether the table was built and refuse to raise the flag otherwise.

Also treat a zero stencil cap as "uncapped" rather than "no growth". The engine always sets
max_grow, so this changes nothing that runs; it makes a caller that forgets it fail loudly instead
of silently disabling the feature.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rugnux fits sigma^2 = a*sigma0^2 + (b*<I>)^2, so its `b` is a fraction of the intensity. XDS fits
sigma^2 = a*(sigma0^2 + b*I^2) and prints ISa = 1/sqrt(a*b). The two `a` are the same number, but the
two `b` are not - b_xds = b^2/a - so the pair rugnux printed could not be read against a CORRECT.LP,
which is the only reason anyone looks at it.

Convert at the report. The fit, the merge weights and both engines' variance expressions are
untouched, so this is a re-expression and not a change: on a rotation dataset the merged intensities
move strictly less between before and after than they do between two runs of the SAME binary (99.9%
identical, max |dI/I| 9.1e-4 against the run-to-run control's 7.5e-3), with the same reflection set.

The rotation path also printed the wrong ISa for the comparison it invites. What it calls ISa is the
strong-reflection asymptote, a tier XDS has no equivalent of and which can only ever be the more
optimistic of the two; XDS's ISa is the whole-range 1/sqrt(a*b), which in rugnux units is exactly
1/b. Print both, labelled. On a broadband rotation dataset that is 13.2 (whole range) and 15.6
(asymptote) against XDS's 21.18 - so the number previously compared was flattering rugnux by 2.4.

A third, unrelated `b` lives in the space-group search: fitted with the sigma^2 coefficient held at 1,
with gate constants calibrated in that convention, and a ratio bound does not survive the mapping
(1.90 would have to become 3.61) while the absolute floor has no correct value at all, there being no
`a`. It is now commented as such, since making the three consistent is the obvious wrong move.

Also corrects three comments and two doc passages that still described a merged-sigma systematic
floor deleted in 72efb75a8.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The mmCIF's _reflns.jfjoch_diffrn_ISa carried the strong-reflection asymptote, a tier XDS has no
equivalent of, while the name invites comparison with XDS's ISa - which is the whole-range
1/sqrt(a*b). rugnux_vs_xds.py reads that item for the battery's ISa column, so the comparison that
column exists to make was between two different quantities, flattering rugnux by the difference
between the tiers.

Write the whole-range value there, move the asymptote to _reflns.jfjoch_diffrn_ISa_asymptotic, and
add _reflns.jfjoch_error_model_a and _b in XDS's convention so the number can be re-derived from the
file rather than taken on trust.

On a broadband rotation dataset the battery column now reads 13.25 against XDS's 21.18 where it read
15.6 before, and the two error models can be compared term by term for the first time: a 1.538 vs
1.249 and b 3.71e-03 vs 1.78e-03, so the gap is in BOTH the counting and the systematic term
(1.23x and 2.08x, and sqrt(1.23*2.08) = 1.60 = 21.18/13.25).

This is a deliberate redefinition of an exported item, not an addition: a file written by an earlier
version carries the asymptote under the old name and there is no version marker to tell them apart.
Noted in the changelog and in docs/CPU_DATA_ANALYSIS.md. Nothing reads the item back into the
pipeline - it is written and never parsed by rugnux itself - so no stored file is reinterpreted in a
way that changes a result.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The mmCIF carries eleven items rugnux invents, and the rule they follow - a jfjoch_ prefix inside
whichever standard category the quantity belongs to - was nowhere written down, so the only way to
learn what was in a merged .cif was to read WriteReflections.cpp. Tabulate them, with the values a
reader needs in order to interpret each one (the untwinned and perfect-twin values for the L test and
the second moment, the sign convention for the radiation-damage B).

Two of them need more than a name. The compatibility note records that jfjoch_diffrn_ISa changed
meaning and that a file carries no marker saying which. And the HKLF-4 .hkl has two properties that
are invisible in the file and change what a comparison means: Bijvoet mates are separate records, and
the intensities carry a single global rescale so the largest fits F8.2 - harmless to SHELXC and ANODE,
which use ratios, but not something to compare magnitudes across.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A reflection was accepted onto an image when |delta_phi| * zeta was within the
mosaicity window, where delta_phi is the offset from the frame's mid-exposure
angle to the exact diffracting condition. That asks whether the frame's CENTRE
lies inside the rocking curve, which is a stricter question than the one that
matters: whether any of the curve lies inside the frame's exposure. The two
differ by half a wedge, and the partiality computed a few lines further down
already integrates over that half wedge on both sides - so the acceptance test
and the quantity it gates disagreed about where the frame is.

The consequence is not a clipped intensity but a lost reflection. Consecutive
frame centres are one wedge apart, so the nearest centre can be half a wedge
away; once the window is narrower than that, the reflection fails the test on
its best frame and on every other, and is never predicted at all. That happens
when sigma_eff < zeta * wedge / (2 * mosaicity_multiplier) - coarse slicing on a
sharp crystal at high zeta, which is where a reflection is fully recorded on one
image and measured best.

Subtracting the half wedge from the tested offset restores the intended
question. On a crystal that reaches the regime (0.4 deg per image, fitted
sigma_M 0.051 deg) low-resolution R_meas goes 6.8% -> 5.4% and ISa 13.3 -> 14.1.
Elsewhere the window merely widens by half a wedge, which admits partials whose
partiality is a few parts in a thousand; those are correctly measured and
correctly down-weighted, and four of the six crystals tested do not move, while
one loses 1.2 ISa. Both engines carry the same test and both are changed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The background is estimated from the r2..r3 ring and then subtracted from every
pixel of the r1 disk, so the ring mean's own error enters the intensity n_inner
times over: var(I) carries n_inner^2 * bkg / n_B. That term is first-order in
sigma, and it is set by how many pixels the ring holds - not by anything about
the reflection. At r3 = 10 the ring holds about 200 px against the disk's 50.
Widening it to 13 roughly doubles that. The signal disk is untouched, and the
pixels gained lie further from the reflection rather than nearer, so nothing is
traded for them.

The effect is not subtle once looked for. Matched observation by observation on
one crystal, halving the ring's pixel count leaves the intensity alone and
inflates sigma by 4.7%, and the inflation rank-orders with the ring collapse
across the battery.

Over the whole rotation battery, against the same binary at r3 = 10: ISa better
on 14 crystals and worse on 4, the summed shortfall against the reference
164.7 -> 155.9, the summed low-resolution R_meas excess 69.7 -> 59.1 percentage
points, and one more crystal reaching the reference space group (33/37 -> 34/37,
a trigonal case that was over-promoting). Largest gains where the ring was
starved worst; the four losses are 0.25 to 2.16 in ISa and none of them changes
a space group.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The comparison read each shell as float(d), float(R_meas), float(CC1/2) inside
a single try, so a shell with any column unset was dropped whole. R_meas is
unset exactly when the mean intensity in that shell has gone non-positive -
which happens on the WORST arm - and dropping the row took its CC1/2 with it.
The high-resolution CC1/2 then silently came from the next shell in, and the
arm whose outer data had collapsed was reported as the better one.

Measured on one crystal where two integration settings were being compared: at
the same 1.74 A shell the two arms are 61.5% and 17.8%, and the table printed
61.5% against 36.9% - the second number being the other arm's 1.83 A shell.
The error is not a rounding matter and it points the wrong way.

Each column is now read on its own, and each statistic is taken from the
outermost (or innermost) shell that actually carries it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Nothing kept a neighbour's flux out of a reflection's own signal disk. The union mask
keeps neighbour cores out of the BACKGROUND ring, but the r1 disk was read whole, so on
a dense pattern a crowded reflection measures part of its neighbour as its own.

Ownership is decided once per image into a per-pixel (quantised distance, reflection)
key written with an atomic minimum, so the nearest predicted centre wins whatever order
the writes arrive in and the lowest index breaks a tie. `--overlap exclude`, now the
default, drops the pixels a nearer neighbour owns from the profile fit. A profile fit is
the amplitude of a normalised profile, so leaving pixels out renormalises the estimator
by construction and the reflection stays unbiased rather than being discarded; the
summation-fallback guard is scaled back to the disk the box-sum seed actually read, so
it still compares like with like. `--overlap reject` is the XDS MINPK alternative - drop
the reflection when less than `--overlap-minpk` of its expected profile is cleanly its
own. A box sum has no profile to renormalise with, so `exclude` is a no-op there and
only `reject` acts on it.

Widening the split - keeping a pixel only where no other centre is within its distance
PLUS a margin - was built and measured, and it is worse monotonically: the residual bias
of the pixels that were kept grows from +0.072 to +0.209 in ln intensity at 0 to 3 px of
margin. What the margin removes is the reflection's own profile, not the neighbour's
tail, so the plain nearest-centre split is the rule.

Measured on the full 38-crystal rotation battery against the same binary with the
treatment off: ISa better 15 / worse 8, summed shortfall against XDS 39.7 -> 28.1. Three
of the losses are the two-pass loop taking its other branch - their median mosaicity
moves between the two known attractors - rather than the change under test; excluding
those it is better 15 / worse 5 and the shortfall goes 31.3 -> 14.4. The two crowded
crystals gain 38% and 52% of their ISa, one of them passing XDS. High-shell CC1/2 over
the 35 crystals that neither flipped branch nor carry a collapsed error model is better
7 / worse 7. Space groups unchanged at 35/38. The owner map is built only when a
treatment is asked for and costs 1.1% of the battery's wall clock - 23% on a genuinely
crowded crystal, nothing where no two predictions touch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every partial's delta_phi was solved from its OWN frame's lattice - by the predictor,
and again by SmoothGeometry. Per-frame geometry is re-refined against that frame's spots
alone, so what is left of its jitter entered each frame of an event independently and the
frames of one rocking event stopped sitting exactly one oscillation apart on the curve.
Their partialities then no longer tile it, and because a broad rocking curve spans more
frames, the error grows as 1/zeta - which is how it has been showing up: a zeta-graded
systematic that nothing in the integrator could reach.

For the frames of one event the geometry is exact. Each frame has already turned one
oscillation further, so delta_phi is linear in frame number with slope minus the
increment; the sign is checked against the data rather than derived, the measured mean
frame-to-frame slope being -0.19996 deg/frame at an increment of 0.20000. Fit the one
free number, the offset, over the event and lay its partials back on that line. The rms
departure removed is 0.25 deg - larger than the oscillation itself, because a small
orientation wobble is amplified by 1/zeta. The raw-hkl runs the merge already builds give
the grouping, so this costs one pass over the partials and no extra sort.

Full 38-crystal rotation battery against the same binary without it, on unchanged data
(observations +0.20%, unique reflections +0.03%, so none of this is selection):

  R_meas_lo   better 22 / worse 5, summed -41.0 pp; excess against XDS -46.9 -> -87.9
  R_meas      better 22 / worse 5, summed -24.7
  CC1/2       better 17 / worse 2,  summed +36.8
  ISa         better 15 / worse 22, summed +6.87; shortfall against XDS 28.1 -> 21.2
  space groups unchanged at 35/38

The low-resolution R_meas gains land on the crystals that have carried this gap: 19.3 ->
10.8, 20.7 -> 13.8 (now past XDS), 25.6 -> 19.7, 17.5 -> 12.4 per cent. Exactly one
crystal shows any change in the two-pass branch fingerprint, so unlike most changes on
this path the result is not confounded by that bistability.

ISa falls on more crystals than it rises, and that is the estimator becoming honest
rather than the data getting worse: every crystal whose ISa dropped materially was
over-optimistic against its own R_meas_lo and moved toward consistency, and the median
ratio of reported ISa to the value its own R_meas_lo implies goes 1.09 -> 1.01, against
1.19 for XDS. The one real loss is a crystal going 1.11 -> 0.95 on that ratio.

High-shell CC1/2 is worse on 22 crystals, by about 1.2 points each. It is the one metric
that dissents, and it is also the one that has failed as an arbiter repeatedly on this
data, while overall CC1/2, R_meas, R_meas_lo and reflection count all improve on an
unchanged number of observations.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The error model's systematic term b is identified only by the spread of I^2/sigma^2
across the intensity bins the fit uses, and those bins hold equal COUNTS. So when fewer
reflections are strong than one bin holds - a sixteenth of the pool - the top bin's
median sits at an intensity where b cannot be measured at all, and the fit hands it the
bins' own noise-selection slope instead: sorting noise by its group mean squared makes
dev2 rise with I2 even when the true b is zero, and with no strong bin to out-vote it
that slope becomes b.

The result is not a small error. On the battery's weakest crystal, 2.2% of whose fulls
reach I/sigma 2, the fit returns b = 5.6 - sigma -> 2*I at the strong end - and since
corrected_sigma applies b at the GROUP MEAN, sigma^2 = a*sigma^2 + (b*mean)^2 is a
per-group constant that caps merged |I/sigma| at sqrt(n)/b. The cap lands at 2.3, so 98.8%
of merged reflections come out below 3 and the reported ISa is 0.50, on data whose CC1/2
is 99.3% at multiplicity 18.7. XDS fits 6.13 from the same images. Feeding XDS's own
scaled observations through this estimator returns 0.84, so it is the estimator and not
the data; synthetic data built with b = 0 and 1.8% strong reproduces a = 0.51 and ISa 0.50
to two digits, and recovers the truth as soon as the strong fraction passes one bin.

So refuse to report what was not measured: when the strongest bin's own (I/sigma)^2 is
below 4, fit a alone, hold b at zero and warn that ISa is unmeasured. The threshold is not
delicate - the two crystals it fires on sit at 0.22 and 0.84 while the next crystal in the
battery is at 31.7 and a healthy one at 342, so anything from 4 to 25 selects the same two.

Full 38-crystal rotation battery: it fires on those two crystals and no others, and space
groups are unchanged at 35/38. Dropping the spurious term also fixes the merge weights it
had been distorting - on the worse of the two, R_meas 19.2 -> 13.6%, low-resolution R_meas
13.8 -> 6.9% against XDS's 14.1%, CC1/2 98.7 -> 100.0%, with chi2 1.11 on the
one-parameter model. Two further crystals move slightly; the guard never fires on either,
and they are marginal crystals of the kind whose two-pass branch any recompilation can
shift.

This reports the parameter as unmeasured rather than clamping it to something plausible,
because the honest statement is that the data do not reach far enough for a systematic
error to be seen - not that there is none.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A screw's predicted-absent class was required to hold min_absent_observed = 8 reflections before the
screw could be claimed. That count is the wrong measure of evidence, and it is wrong in both
directions.

A screw extinguishes one row of reciprocal space, and that row is often the one a rotation sweep
records least: it lies near the spindle, where the blind cusp maps onto itself and symmetry cannot
fill it in. Counting it measures the geometry of the sweep. A monoclinic crystal whose 2-fold sits
7.6 deg from the spindle contributed six 0k0-odd reflections, every one of them measured between
-0.013 and 4e-5 of the shell mean with zero violations, against a 0k0 row averaging 1.44x the shell
mean - and was refused its 2_1 for being six rather than eight. XDS's own integration of the same
images finds seventeen of those reflections and every one of them is likewise dead.

The count is equally wrong the other way: a uniformly weak axial row produces no violations at all,
so with enough reflections on it a screw is claimed from no evidence whatsoever. The second new test
section demonstrates exactly that on the old gate.

Judge the class by how unlikely it would be if the screw did not exist. Under "no screw" the absent
class and the rest of its row are both Wilson-distributed with the same mean, so with each absent
intensity taken in units of its row's control mean, sum_u/(sum_u + n_control) follows Beta(n_absent,
n_control) exactly; the reported evidence is -log of that lower tail. The row's own strength cancels,
which is the property the count lacks, and the scale is set by the number of reflections, so
few-but-decisive and many-but-marginal are told apart. It is sigma-free by design: the merged sigma
carries the error model's intensity-proportional term and so shrinks with I, reading much the same on
an absent reflection as on a present one.

This follows POINTLESS (Evans, Acta Cryst D67, 282-292 (2011), Appendix A3), which likewise scores an
absence against the rest of its own axial row rather than against a global mean or a fixed cut, and
likewise lets confidence fall away with the number of axial reflections instead of refusing outright
below a count. POINTLESS calibrates its null width from control transforms of non-axial reflections;
the Beta tail here is an analytic null in its place. XDS is not a reference for this: it "deliberately
avoids any test for the presence of screw axes as these tests would depend strongly on the
completeness of the data" (Kabsch, Acta Cryst D66, 133-144 (2010), section 6), so a screw axis in a
CORRECT.LP was supplied to it, not determined by it.

Measured over five probe crystals, genuine screw conditions read 34-800 nats and false ones - the
4_1/4_3 conditions of a cubic crystal that has no screw, whose predicted-absent class is STRONGER
than its control row - read -7 to -8.5. The bound is set at 20, in the gap, at p <= 2e-9: three
well-measured dead axial reflections clear it and two do not.

min_absent_observed keeps its job for CENTERING, where a count is a fair measure - that class is a
third to a half of every reflection in the data set and the bound is never binding on a centering
that exists.

The candidate table now prints the screw-absent count and this evidence in place of the two E^2
medians that were its raw ingredients, so a refusal can be read off the log.

Measured on the five probes: the monoclinic crystal above returns to P2_1 with every merge statistic
unchanged (R_meas 58.8 -> 58.7%, CC1/2 49.1 -> 49.3%, ISa 6.61 -> 6.59 - P2 and P2_1 share a point
group, so only the symbol and the absent reflections differ). The other four are untouched, space
group included, and the two-pass branch fingerprint (indexed frames, distance, mosaicity) is
identical on all five. The full battery has not been run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Several methods adopted recently came from other crystallographic packages - the screw-absence test
from POINTLESS, MINPK and the profile-fit reweighting from XDS/Otwinowski, the CC1/2 cutoff and merge
outlier rejection from DIALS, the per-frame indexing gate from CrystFEL - and nothing in the
repository said where such a debt is recorded. The licence side was already worked out (licences beside
the vendored code, verbatim texts in licenses/ collected by COLLECT.sh, a row in THIRD_PARTY_NOTICES.md,
all installed under share/doc/jfjoch); the credit side was ad hoc.

Write the rule into CLAUDE.md. It states the distinction that matters: vendoring or linking someone's
CODE creates a LICENCE obligation, discharged in licenses/ and THIRD_PARTY_NOTICES.md; reimplementing
an algorithm from a PAPER creates none of that but creates an obligation of academic CREDIT, discharged
in docs/ACKNOWLEDGEMENT.md and in a comment at the algorithm. Neither substitutes for the other, and
taking both source and paper incurs both. It also fixes the citation form (authors, title, year,
journal, volume, pages, verified DOI), and says in-source credit goes at the algorithm, not the file
header, in the one-line style the code already uses.

Then bring the repository into compliance for the works concerned: docs/ACKNOWLEDGEMENT.md gains a
section acknowledging XDS, DIALS, POINTLESS/CCP4, MOSFLM, CrystFEL, GEMMI, the Kabsch/Otwinowski
profile fit, the Diederichs & Karplus statistics and the IUCr nomenclature reports, each with a DOI
checked against Crossref; docs/CPU_DATA_ANALYSIS.md's reference list gains the ones it was missing;
and four algorithms gain a line naming their source where no adjacent comment carried one.

No licence change. licenses/ and THIRD_PARTY_NOTICES.md are untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The SparseCCL connected-component labelling adapted from traccc is MPL-2.0 source in this tree, and
MPL-2.0 requires the notice to be conveyed with it. The licence text has always shipped
(licenses/traccc.txt, installed with the rest of licenses/ into share/doc/jfjoch) and
docs/ACKNOWLEDGEMENT.md has always credited it - but the root THIRD_PARTY_NOTICES.md, which is the
manifest a recipient actually reads, never named it.

The generated docs/ copy DID name it, and more completely than a fresh row would have: both files
it touches, and a note explaining that only one of them adapts the source while the other follows
the design. So the fix is to carry that entry back into the canonical file rather than write a new
one - the generated copy had been edited directly at some point, against its own banner, and the
edit never reached the file it is generated from. Regenerating now reproduces docs/ byte for byte,
which is the check that the two are finally in step.

update_version.sh gains the link rewrite that entry needs: the acknowledgement is docs/ACKNOWLEDGEMENT.md
from the root and ACKNOWLEDGEMENT.md from inside docs/, so without the rule the next regeneration
would have written a path that resolves nowhere.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A weighted mean is only unbiased while the weights are independent of the values
being averaged. The IUCr's own nomenclature report (Schwarzenbach et al., Acta
Cryst A45 (1989) 63-75) puts it directly: weights in averaging "should not be
based on the counting statistics of the individual observations whose estimated
variances are biased and result in larger weights for accidentally low
intensities". Two places in the rotation pipeline were doing exactly that, and
between them they drove whole resolution shells of merged intensity negative.

1. The profile fit computed its non-signal variance as

       var_bkg = max(0, 1/den - max(0, I) + bkg-estimate term)

   The point of a separate var_bkg is that it does NOT move with the
   reflection's own fluctuation, and 1/den - I is the quantity that does not:
   1/den is the fit variance taken at the fitted intensity and grows with it
   roughly one for one. Clamping the subtrahend at zero left a down-fluctuated
   reflection's own deflated variance standing as its background variance.
   Measured over 6.9 M partials of one weak rotation dataset, var_bkg/bkg came
   out at 3.7-5.4 for observations with I < 0 against 11.4-13.7 for I > 0 - the
   down-fluctuated half of every reflection carried a variance ~2.7x too small
   and was weighted up by the same factor, first in the 3D combine and then
   again in the merge. Removing the clamp makes var_bkg flat in I (~13 x bkg
   across the whole range).

2. The merge then weighted each combined full by 1/sigma_full^2, and sigma_full
   is by construction a function of the full's own answer: the combine's
   variance carries a corr*max(0, F) signal term, so every full with F <= 0 got
   the smallest variance the model allows while the strongest quartile got
   2.26x more. The merge now rebuilds that variance at the reflection's mean
   instead, from a linear model var(I) = var_bkg + var_per_I * I that the
   combine measures and stores on the full. This mirrors
   MergeOnTheFly::CorrectedSigma, whose comment already claimed to mirror the
   rotation combine.

Verified against an estimator that cannot see the fluctuation - summing the
partials and dividing by the summed partiality, the classical construction every
other program uses (Greenhough & Suddath, J. Appl. Cryst. 19 (1986) 400-409, via
Leslie, Acta Cryst D55 (1999) 1696-1702: profile fitting biases the individual
partials but not their sum). Reproducing the merge on dumped observations, the
shipped weighting sat ~1.9 sigma below that reference in the noise shells; the
two changes recover most of it, and every intensity-independent weighting
scheme agrees with the reference once (1) is in.

Four-crystal probe, XDS resolution limits, branch fingerprint identical on all
four (so none of these is a two-pass branch flip):

  weak cubic case   last shell <I/sig> -1.6 -> +0.2 (XDS +0.10), last shell
                    R_meas 478% -> 250% (XDS 246%), overall <I/sig> 6.1 -> 7.5
                    (XDS 7.18), R_meas 18.3% -> 18.1%, CC1/2_hi 38.2% -> 43.7%
  tetragonal case   outer shells <I/sig> -0.4/-0.8/-0.9/-1.0 -> +1.8/+1.2/
                    +0.9/+0.4, R_meas 184%/595%/7614%/nan -> 95%/119%/135%/232%
                    (the nan was the shell mean crossing zero), R_meas 33.3% ->
                    32.9%, CC1/2_hi 38.3% -> 56.5%
  trigonal case     R_meas 13.0% -> 12.5%, CC1/2_hi 14.4% -> 16.5%
  strong control    unchanged to every printed digit but ISa

Cost: ISa falls (17.2 -> 14.0 and 16.7 -> 14.9 on the two mid-strength cases,
28.3 -> 27.8 on the control). Strong reflections are untouched by (1) - their
partials are all positive, so var_bkg is bit-identical - but the joint a/b fit
redistributes: honest weak sigmas lower a, and b rises to keep the strong bins
fitted. The median reduced chi^2 improves (1.25 -> 1.14, 1.35 -> 1.28) so the
new split describes the scatter better, but ISa is the one headline metric that
moves the wrong way and it should be watched over the full battery.

The integrator change is shared, so the stills merge sees it too; there it feeds
GetExpectedVarianceMerge, which had been handed the same contaminated var_bkg.
That path is untested here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A rotation dataset has ONE lattice. Once the first pass has found it and the
goniometer gives each frame its orientation, every frame of the sweep is a
frame of that crystal - yet integration was gated on each frame re-indexing on
its own, a test that carries an absolute floor of 9 indexed spots. A weakly
diffracting crystal shows a handful of spots per image while the geometry still
puts ~1500 reflections on the detector, so the floor threw away whole frames
that had nothing wrong with them.

Measured on a 360-degree battery crystal: 1484 of its 1800 frames failed that
gate, all of them on the spot-count floor alone and none on the consistency
test - the median failing frame had 4 spots and the lattice indexed all 4.
Integration therefore ran on 17.7% of the sweep and the merge came out 35.7%
complete at multiplicity 1.1, against XDS's 97.7% at 2.81 from the same images.
XDS's own INTEGRATE.LP shows why the floor is the wrong test there: 964 of its
frames have fewer than 9 strong spots and it predicts ~1483 reflections near
the Ewald sphere on every one of them, because INTEGRATE works from the global
orientation and has no per-frame indexing gate at all. Neither does
dials.integrate.

Split the one verdict into the two questions it was answering. "Does this frame
index?" - what the indexing rate reports and what the first pass scores
candidate lattices on - keeps the floor, because a handful of spots sit on
almost any lattice by chance. "Is this frame worth integrating?" keeps only the
consistency part, and only where the lattice does not come from this frame. A
frame whose spots largely MISS the lattice is still refused: on another battery
crystal that is 35% of the sweep, and integrating those collapsed the space
group to P1 - the floor had been shielding the merge from frames the model does
not describe, which is a different defect and not one to paper over here.

Two consequences had to be handled. A frame that is too sparse to index is also
too sparse to fit its own rocking width, and the placeholder it used to predict
with was being reported onward as if measured, into the frame-order average
that recomputes every partiality; report nothing instead, and fill the gaps in
that average with the run's median rather than a fixed default.

Probe (XDS in brackets): the crystal above goes 9 700 -> 81 956 observations,
8 618 -> 23 960 unique [23 576], 35.7% -> 99.4% complete [97.7%], R_meas
21.2% -> 68.6% [76.7%], CC1/2 96.0% -> 86.4% [81.1%], low-shell R_meas
7.2% -> 14.3% [20.6%], ISa unmeasurable -> 13.8 [10.4] - better than XDS on
every statistic, where before it was merging a third of the data. A second
crystal gains 41% more observations with R_meas 12.6% -> 8.5% and ISa
3.3 -> 3.7. The high-multiplicity control is unchanged to 2 observations in
924 782, and four further crystals move within recompilation noise.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two guards catch a per-frame scale far below the run median. Both then invented a value for it -
one substituted the run median, the other set corr = 1 and merged the frame "unscaled". For a frame
whose scale really is 1/17402 of its neighbours', asserting 1 is worse than asserting nothing, and
it is the assertion that does the damage: those observations enter the merge at full weight carrying
an intensity scale that is wrong by four orders of magnitude.

It surfaced when rotation started integrating every frame the sweep's lattice explains, but it is
not caused by that change - six crystals in the battery already tripped these guards before it. What
the extra frames did was find a crystal where the collapsed population is large enough to dominate:
R_meas 19.1 -> 90.1%, ISa 25.60 -> 4.58, from 17% more observations.

The frames are not sparse and the fit is not running away. A per-frame dump shows 3394 observations
on the median collapsed frame against 3469 on live ones - the scale is over-determined 3400:1 for
one parameter - and 113 frames fit exactly zero. They form one contiguous arc of about 68 degrees
once the sweep's wrap is accounted for, over which the per-frame correlation to the merge is 0.035
against 0.85 elsewhere, while the flux measured from the background varies by only 1.55x. So the
fitted zero is a well-determined measurement that the frame holds no diffraction from this lattice,
not a failure to measure. The frames are empty, not under-determined.

That is also why the smooth or shrunk alternatives do not apply, and both were built and measured
rather than argued away: giving a collapsed frame the geometric mean of its credible neighbours is
worse than the baseline (R_meas 115.3%), because it merges noise at the weight of a good frame, and
a dead region 112 and 232 frames wide has no local neighbourhood to borrow from in any case.

Dropping them: R_meas 90.1 -> 38.3%, low-resolution R_meas 26.9 -> 10.2% (past XDS's 14.3), ISa
4.58 -> 22.00, CC1/2 99.0 -> 99.9, with 5.4% more observations retained than before frames were
integrated at all. Over the full battery, against the same binary without either change, ISa moves
from -22.5 to -3.4 summed, CC1/2 +24.6, and 216247 more observations. The crystal that motivated the
integration change is untouched by this one, bit for bit.

The detection and the MIN_CREDIBLE_SCALE_RATIO threshold are unchanged. Note that threshold is now
marginal: its own comment records 0.070 as the smallest legitimate ratio seen, and one crystal here
has a legitimate live frame at 0.026, so it cannot be raised to catch the partly-dead transition
frames at the edges of an arc without risking real data.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Everything a run determines went to stdout and nowhere else. The space group and the evidence behind
it, the error model, the post-refine commit-or-reject decisions and their held-out residuals, the
two-pass adopt-or-roll-back, the resolution cut, the merging statistics - all of it scrolled past
interleaved with progress lines and was gone. A user who was not watching had no record, and nothing
could read it. `rugnux` had no log file at all; the `rugnux.log` in the regression harness is that
harness capturing stdout.

Write `<prefix>_report.txt` alongside the .cif/.mtz/.hkl, always, with no option to ask for it. It
holds what the run DETERMINED; timing, rates, per-image progress and engine chatter stay on stdout,
where they belong. Every line rugnux logs was classified result-or-process against the regression
corpus to decide what crosses over.

The format follows XDS's CORRECT.LP, which has been read by people and parsed by other programs for
twenty years: `KEY= value` assignment lines a script greps one at a time, fixed-width tables with
stable headers and a total row, `WARNING:` sentences in plain English, section banners. REPORT_VERSION
says when that interface last changed. It is assembled from results the pipeline already computed, so
an unconditional file costs nothing, and a failure to write it is logged and swallowed - a run that
produced good reflections must not be lost to a side file.

One thing CORRECT.LP does not have to solve: a rotation run integrates twice and writes both passes,
so every report says which pass it describes and why that pass was adopted.

`--no-merge` gets a report too, saying MERGE= NOT_PERFORMED rather than leaving a reader to infer it
from absent sections. An empty output prefix still writes nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Scaling: report the stretches of a sweep the crystal did not deliver
Build Packages / build:viewer-tgz:cpu (push) Successful in 19m5s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m34s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 23m21s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 24m15s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m59s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 29m12s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 29m4s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 19m35s
Build Packages / XDS test (durin plugin) (push) Successful in 11m6s
Build Packages / build:rpm (rocky9) (push) Successful in 20m57s
Build Packages / Generate python client (push) Successful in 46s
Build Packages / Build documentation (push) Successful in 1m29s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m36s
Build Packages / build:rpm (rocky8) (push) Successful in 27m38s
Build Packages / DIALS test (push) Successful in 21m0s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 11m35s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m16s
Build Packages / XDS test (neggia plugin) (push) Successful in 10m27s
Build Packages / Unit tests (push) Successful in 1h18m5s
Build Packages / build:windows:nocuda (push) Successful in 19m18s
Build Packages / build:windows:cuda (push) Successful in 23m19s
ce3199748d
Rotation processing no longer refuses to integrate a frame that fails to index on its own, which is
right - no other program does that - but it means a genuinely bad stretch of a sweep is now
integrated instead of quietly dropped. Some sweeps have a real problem behind that stretch: the
crystal partly or wholly out of the illuminated volume, off the rotation axis, or dying of dose.
That is actionable at the beamline ("recollect", "re-centre"), and until now nothing said it.

MeasureSweepQuality reports it as contiguous RANGES, never per-frame flags, and reports only - no
observation is excluded on the strength of it. A single weak frame is noise; forty consecutive ones
are a fact about the experiment, and the frames still carry signal worth merging.

The discriminator is that the incident flux is already out of the per-image scale before that scale
is fitted (DivideOutIncidentFlux runs from Ingest), so a drop in G that the beam does not explain is
on the sample side by elimination. Measured on one crystal with a dead arc: the flux proxy spans
1.4x across the run where the fitted scale spans 246x.

A range needs BOTH per-frame channels down: the scale, and the CC to merge. The CC channel is what
keeps a merely attenuated stretch out - absorption and flux scale a frame's intensities without
changing how well they correlate with the merged reference. Without it the clean high-multiplicity
control, whose per-image scale swings 4x on a 180 degree period, would be reported as a bad crystal.
It is not: it produces no ranges at all, and neither does the other control.

Five codes, each the field's own words and each a phrase a report can print:

  no diffraction    - essentially nothing was recorded from the indexed lattice over the range
  out of beam       - frames were lost: the range gets a scale far less often than the run does
  weak diffraction  - the frames all still index, with much less intensity; cause not determined
  loss of centring  - one cycle of modulation per revolution (autoPROC's words for the phenomenon)
  radiation damage  - the range runs to the end of a sweep whose quality was already decaying

Only the last two claim a cause, and each rests on its own evidence. Damage is progressive, so it
must have been setting in before the range and must not recover. Loss of centring rests on the one
signature that breaks a documented degeneracy: Evans (Acta Cryst. D62, 72-82) notes that illuminated
volume and absorption are indistinguishable, but a crystal's own shape absorbs on a 180 degree
period, so a dominant 360 degree fundamental over a full turn cannot be the crystal's shape. That
test runs on the total scale, flux included, unlike everything else here - the flux proxy is a
background, a crystal leaving the beam takes its own scattering with it, and the beam cannot be
periodic in an angle it does not know. Where the evidence does not reach, weak diffraction says so
rather than guessing.

Frame numbers are processed-image ordinals, inclusive at both ends, the numbering of _image.dat.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The angles a rotation dataset stores are the COMMANDED ones, so a stage whose travel is
miscalibrated leaves no trace in the header - every angle is self-consistently wrong. No
existing parameter can absorb it either: the cell scale, the axis direction, the detector
distance and the beam centre are all orthogonal to an error in rotation MAGNITUDE.

So fit it as what it is - one scalar k, the ratio of the travel to the commanded angle -
on the rocking events the geometry post-refinement already builds, after step A so the
cell scale and the axis direction are fixed and k is the only free quantity. Two details
decide whether the number means anything. The angle enters measured from the CENTRE of
the sweep: the reference orientation was fitted against the commanded angles and has
already absorbed their mean error, so measured from the goniometer's zero instead a
constant missetting about the spindle leaks into k with a gain of <phi>/<phi^2>, which
depends only on where the sweep happens to sit - on a short sweep starting near zero a
0.14 deg missetting fakes 1.4 % of k. Referred to the sweep centre that leak is
identically zero at any width. And the robust loss is scaled to the scatter the events
actually have, which varies by more than a decade between datasets, so any fixed constant
is either inert or throws away real data.

A stage fault is rare and a 1 % angle correction applied to a healthy dataset would damage
it silently, so the correction is committed only when every test passes: at least 30 deg
of sweep and 5000 events, |k-1| over 0.5 %, a misorientation of at least 0.5 deg at each
end of the sweep, and the same k from every fifth of the sweep left out. The last test is
not optional. A second lattice that dominates ONE END of a sweep - exactly what happens
where the primary stops indexing - fakes a k that passes the other two, and the hkl-hash
split used elsewhere in this file cannot see it, because both of its folds sit at the same
angles and anything structured in phi survives in both.

When it commits, the second pass re-integrates against the corrected angles. The pre-pass
mosaicity is dropped with it: that is a width in degrees fitted against angles the second
pass has just stopped using, and since the override can only ever raise the second pass's
own estimate, carrying it over would hold the second pass at the rocking width the
uncorrected angles produced - the correction half-applied.

--rotation-scale asserts a known stage calibration by hand and overrides the fit.

On the 38-crystal rotation battery the gate fires on exactly one dataset, at k = 1.01318
with 0.74 of that k surviving every fifth left out. The largest of the other 37 is
1.00211, which fails the end-error test; 34 of them sit below 1.0006. On the one that
fires:

  R_meas          39.2 -> 23.9 %   (XDS 37.1)
  CC1/2           86.5 -> 96.0 %   (XDS 94.3)
  CC1/2 outer      1.4 -> 53.4 %   (XDS 42.5)
  unique refl    40990 -> 41540    (XDS 41322)
  observations   74975 -> 103858   (XDS 129322)
  mosaicity      0.181 -> 0.159 deg

which takes it from losing to XDS on R_meas, CC1/2 and outer-shell CC1/2 to beating it on
all three, and the mosaicity drop is the inflation the uncorrected angles were producing.
Its low-resolution R_meas is the one number that moves the wrong way, 12.0 -> 13.9 %,
still well inside XDS's 18.3. No space group moves anywhere, and every other crystal's
merge is unchanged beyond the two-pass loop's own jitter - measured here as the spread of
the post-refined distance across arms that do not touch post-refinement at all, which is
larger than anything this commit produces.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Prediction and partiality read one number, the per-image sigma_M, so a sigma_M
that moves takes the integrated reflection population with it and there is no
way to ask which of the two uses carries a downstream difference.
--prediction-mosaicity fixes the width the prediction window opens to while the
partiality keeps using the measured sigma_M, which separates them.

Measured with it on a rotation crystal whose lattice search returns two
different cells a few tens of microns of detector distance apart: widening the
prediction window from the narrower branch's 0.203 deg to the wider branch's
0.272 deg adds 28% more partials and moves the merged statistics by less than
half a percent (I/sigma 6.1 -> 6.2, R_meas 12.7 -> 12.6%, ISa 10.7 -> 10.9);
narrowing the wide branch the other way removes 26% of its partials and
recovers nothing. On a tetragonal reference crystal a 4.7x over-wide window
costs 11%. The prediction window is not where a mosaicity difference turns into
a merged-data difference - the captured-fraction gate and the partiality
weighting downstream absorb a generous window.

Diagnostic only; off by default, so nothing changes unless it is asked for.
The second pass of the two-pass rotation loop widens its prediction window to
the frame-order-smoothed mosaicity the first pass fitted, taking the max with
this frame's own estimate so no reflection is dropped. That widened value was
then reported onward as the frame's mosaicity, so it also became the divisor
RotationScaleMerge recomputes every partiality from - a number chosen for
prediction safety, applied to the intensities.

The two uses are not symmetric. Prediction only decides membership: a
reflection just inside a generous window arrives with a partiality near zero
and is weighted as such, and widening the window by 28% was measured to move
the merged statistics by under half a percent. The divisor multiplies every
partial, and forcing it 34% wide on a rotation crystal cost ISa 10.7 -> 8.0 and
R_meas 12.7 -> 14.5%. Pass 2 has also just re-measured the rocking width
against the post-refined cell, which is the better of the two numbers - the
literature's own remedy for a mosaicity fitted against a stale cell is exactly
to re-estimate it after post-refinement (XDS documents this as a manual second
INTEGRATE/CORRECT round).

So keep the widened value where it was wanted, on the prediction window, and
let the partiality divide by what the frame measured. The predicted reflection
population is unchanged.

Measured on three rotation crystals at their XDS resolution limits, identical
partial counts in every arm: a weak monoclinic ISa 6.3 -> 6.7, R_meas
19.1 -> 18.7%, <I/sigma> 3.3 -> 3.5; tetragonal lysozyme ISa 27.1 -> 27.4,
R_meas unchanged at 4.5%; a second monoclinic ISa 13.9 -> 13.8, R_meas
unchanged at 12.3%.

The full 38-crystal rotation battery then says the defect was almost never
active: 36 of 38 crystals are untouched, and the reported per-image mosaicity
moves on exactly one of them (0.3380 -> 0.3385 deg). The three-crystal probe
above does not reproduce against the current baseline - lysozyme reads
27.81 -> 27.80 - because the max() rarely bites: pass 1's per-frame width is
already smooth in frame order, its median frame-to-frame step being 0.0005 deg
against a run spread of 0.028, so it seldom exceeds what pass 2 measures for
itself. The two crystals that do move are marginal ones whose two-pass lattice
search takes a different branch (validation frames 25/60 -> 26/60 and
48/60 -> 49/60); on that evidence their merge numbers measure the branch, not
this change. No space group moves.

So this lands as a correctness fix with no measurable effect on today's data,
not as an improvement: the widened window must not become the divisor, whether
or not the two happen to coincide on the crystals we have.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A predicted reflection was discarded outright if ANY pixel of its signal disk was
unreadable - masked, untrusted, in a detector gap, or overloaded. On a battery
crystal that is 11.1% of all predictions, thrown away for a defect in one pixel
of fifty, and the pixels concerned sit at fixed places on the detector, so the
loss is systematic in reciprocal space rather than random.

Neither XDS nor dials does that. Both estimate the missing part from the profile
instead and keep the reflection while enough of it was seen: XDS's MINPK (default
75%, "the missing intensity is estimated from the learned profiles"), dials'
integration.profile.valid_foreground_threshold (default 0.75). MOSFLM is the one
program that rejects by default, and even it relaxes to 50% with PROFILE EDGE.

We already had the argument and the machinery: a profile fit is the amplitude of
a NORMALISED profile, so leaving pixels out renormalises the estimator by
construction - it costs information, which sum P^2/v duly loses and sigma duly
gains, and biases nothing. That is exactly why --overlap exclude drops a
neighbour's pixels from the fit rather than the reflection. Unreadable pixels are
the same case with a different reason, so they take the same treatment, cut on
the same threshold, in the same place: the readable fraction of the expected
profile, measured against the profile mass that lands on the detector at all so a
reflection is judged on the pixels that exist. A box sum has no profile to
renormalise with and keeps the all-or-nothing rule.

Two consequences handled. The summation seed and its variance now count the
pixels actually read, and the runaway guard scales the fit back to that same disk
before comparing - both exactly as before wherever nothing is missing. (Its
fallback then hands back that partial sum unrescaled, which would read low; the
guard fires on 8 of 96 688 recovered reflections, and on none at all on a weak
crystal, so it is not worth a branch.) And the profile, its resolution shells and
their widths are learned from COMPLETE reflections only, as is the box-sum
centroid post-refinement reads as an observed position: a disk with a hole gives
a centroid pulled away from the hole, and the hole does not move between frames.

That sigma gains what the missing pixels carried is the claim the whole change
rests on, and it is measurable. Force the conventional CENTRED cell of a
body-centred crystal in P1: the predictor then enumerates every lattice point,
and the reflections the centring makes systematically absent have a true
intensity of exactly zero, so their scatter about zero must equal their reported
sigma. Over 7.1 M such observations, matched by resolution shell, the trimmed
std(I)/rms(sigma) of the recovered reflections is 0.99 / 1.20 / 2.33 / 1.04 /
1.69 against 0.98 / 1.22 / 2.29 / 1.03 / 1.56 for the reflections that were
complete - the same calibration to a few percent. The lever there is small,
because the typical recovered reflection is missing only 5% of its disk. Lowering
the threshold to 0.50 admits a band missing 25-50%, which is a real lever: there
sigma comes out 8-43% larger than a complete reflection's in the same shell, and
the scatter about zero tracks it, 0.97 / 1.09 / 1.92 / 0.99 / 1.37, at or below
the complete population. Sigma grows, and by the amount it should.

The threshold stays at XDS's and dials' 0.75, on that evidence and on quality.
Below it the estimator starts to run out: on those same zero-intensity
reflections the recovered ones read +0.8 counts high at 0.75 and +1.9 counts high
in the 0.50-0.75 band, against a sigma of 12-17, and at 0.25 the fit degenerates
outright, single reflections carrying sigma in the thousands. Above it there is
nothing to buy: 0.90 leaves a fifth of the recoverable observations behind and
measures no better for them. On the high-multiplicity control, R_rim over
as-shipped / 0.90 / 0.75 / 0.50 runs 4.49% / 4.51% / 4.56% / 4.78% while
<I/sigma> runs 33.47 / 34.02 / 33.89 / 33.43 - 0.50 is where the recovered
observations stop paying for themselves.

Probe against the previous commit, six crystals. The high-multiplicity control
gains 4.2% more observations, 924 803 -> 963 946, which lands it on XDS's 961 379
from the same images, for <I/sigma> 33.47 -> 33.89, R_rim 4.49% -> 4.56% at 4.3%
more multiplicity, CC1/2 unchanged at 0.9998 and ISa 27.80 -> 27.12. Five weaker
crystals gain 3.3-4.8% of their observations and up to 1.0 point of completeness,
for <I/sigma> +0.4 to +3.6%, R_rim between -8.1% and +5.8% relative, CC1/2
+6.6 / +0.3 / +0.2 / -0.0 / -1.2 points, and ISa between +0.3% and -3.4%. Some of
that ISa is the point rather than the price: a reflection integrated over fewer
pixels carries less information, and the absence test above says the sigma that
reports so is honest. The GPU and CPU engines agree as before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Full 38-crystal rotation battery, against the same binary without it:

  observations     better 38 / worse  0,  +937 100
  unique refl      better 30 / worse  0,    +9 229
  overall <I/sig>  better 33 / worse  1,     +7.00
  CC1/2            better  5 / worse  1,      +6.2
  space groups     unchanged at 35/38

Every crystal gains observations and not one loses a unique reflection. The two
costs are small and both are understood. Low-resolution R_meas is worse on eight
crystals, by +0.8 pp at most and +3.2 pp summed - a reflection whose own peak
pixel is unreadable loses the part of the profile that carries most of the
amplitude, and that population sits at low resolution; the following commit
handles it. And ISa falls on 32 crystals, by 10.9 summed, which is what admitting
937 000 further observations does to the strong-reflection asymptote: R_meas
excluding the one crystal whose thread-count noise is 1.5 pp is flat.
CLAUDE.md said `beam_stop/` was in no CMakeLists. It is in
image_analysis/CMakeLists.txt:29-30, included at rugnux/Rugnux.cpp:27 and
constructed at :126, where the shadow detection runs by default over 60 frames.
An agent sent to look for a pre-indexing beam-centre estimator read the line,
correctly took the machinery to be absent, and had to measure the code to find
out otherwise.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The document opened with "It is not yet wired into any workflow" and carried a
"Wiring plan (deferred - implement later)" whose first step was to add the source
to image_analysis/CMakeLists.txt. That step, and the rugnux consumer, are both
done: the finder runs by default over 60 equally-spaced frames and writes
PixelMask bit 9. This is the same stale claim just corrected in CLAUDE.md, and
this file is where it came from. The broker consumer sketched further down is
genuinely still open, so that part is left as a plan.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
MINPK asks how MUCH of the expected profile is readable. It does not ask WHERE, and
the two are not the same question. The renormalisation argument the rescue rests on -
a fit over a subset of a normalised profile is unbiased - needs the pixels to go
missing for reasons unrelated to the reflection. A gap, a mask or the edge of the
sensor is such a reason: the loss is set by the detector, and the fit renormalises
over what is left. A pixel invalidated BY THE FLUX IT SAW is not: it goes missing
because the reflection was bright, and it is the peak.

Measured on the combined fulls, against the mean of the complete observations of the
same reflection, in the innermost resolution shell of the high-multiplicity control
and of a weaker crystal: a rescued reflection whose unreadable pixel sits within a
pixel of the predicted centre reads |I - <I>|/I of 0.50 and 0.53, against 0.073 and
0.212 for a complete observation - 6.8x and 2.5x - and carries several times the mean
intensity of its shell. On the control that is 0.21% of the shell's observations
supplying 1.77% of the R_meas numerator; on the weaker crystal 0.52% supplying 6.82%.
Rescues that lost only rim pixels are unremarkable by the same measure, 1.19x and
0.88x. Dropping the peak-losers alone takes the shell's R_meas from 7.440% back to
7.315% (unrescued: 7.307%) and from 22.03% to 21.06% (unrescued: 21.28%) - which is
the whole of the low-resolution R_meas the rescue cost, and on the second crystal
rather more.

Raw frames say what they are. The pattern is a dead-centre invalid pixel with 5878,
9875 and 27583 counts around it: the detector's per-frame invalid marker on the
brightest reflections. MINPK cannot catch them because it cuts on profile MASS, and
the peak of a broad spot is a few percent of the mass.

So a second condition, in the loop that already measures the readable fraction: no
unreadable pixel may carry more than 0.9 of the profile's own peak value. A fraction
of the peak rather than a radius in pixels because the peak is as wide as the spot -
for a Gaussian the cut is at sqrt(-2 ln f) sigma, 0.46 sigma here, which is the peak
pixel alone where sigma is 0.8 px and the crest of the ridge where the profile is a
bandwidth streak. Swept against the alternatives on two crystals: a fixed radius
needs 1.0-1.5 px to do the same work and costs 3-9x more observations for it, and
0.5 px does not reach the peak of a sub-pixel-offset prediction at all; tightening
the fraction to 0.5 or 0.2 buys nothing beyond 0.9 and costs 7x more.

Six crystals, three detectors, against the rescue as it stands: the rule keeps
99.86-99.96% of the recovered observations and returns R_meas to its unrescued value
or below (4.6 -> 4.5%, 6.7 -> 6.6%, 25.1 -> 25.0%), R_meas in the innermost shell
likewise (2.7 -> 2.6%, 5.9 -> 5.3% against 5.4% unrescued, 16.5 -> 16.4%), <I/sigma>
up or level everywhere, and every unique reflection the rescue won is kept. Raising
--overlap-minpk to 0.90 instead reaches the same place on two of them and short of it
on the third, while discarding 0.8% of the recovered observations rather than 0.05%.
An elongated pink-beam profile on a 9M detector and an EIGER2 16M dataset are both
untouched at 99.9%, so the crest protection does not over-reject a streak.

One crystal is not improved: a dataset whose error model rugnux declines to fit for
want of strong reflections, whose <I/sigma> is <= 0 in eight of its ten shells and
whose R_meas is undefined in as many. There the rule costs about 3% of <I/sigma> in
the one shell that has signal, reproducibly, on top of the 9% the rescue itself costs
there - while its overall R_meas moves 1.5 points on nothing but the thread count.

The parity test gains four sections. Unreadable pixels were only ever punched into
empty sky, so neither the rescue nor this rule had any CPU/GPU coverage at all;
they now go into the signal disks - the peak of every fifth reflection, ~1.1 sigma
out of every seventh, the disk edge of every eleventh - for both profile modes, a
box sum and an elongated stencil, with a check that the clipping actually costs
reflections so the coverage cannot go quietly vacuous.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Full 38-crystal rotation battery against the rescue without this rule, both on the
same base:

  ISa         better 19 / worse  4,   +0.73
  CC1/2       better  3 / worse  1,   +1.3
  R_meas_lo   better  4 / worse  3,   -0.3
  space groups unchanged

for 17 770 observations, 0.09 % of the run total and under 2 % of what the rescue
had won. The two crystals whose peak-loss population was measured beforehand
land on their predicted values: a tetragonal reference goes R_meas_lo 2.7 ->
2.6 % and ISa 27.11 -> 27.42, a cubic insulin 5.9 -> 5.3 % and 20.34 -> 20.65.

One crystal pays: a cubic case with 2381 unique reflections goes R_meas 8.8 ->
9.6 % and ISa 4.08 -> 3.49. It is the crystal in the battery with the fewest
uniques, so its rescued population is small and its shell statistics are coarse,
but the loss is real and not noise in the R_meas. The R_meas sum over the battery
reads +1.3, of which +3.2 is one crystal whose R_meas moves 1.5 points on thread
count alone; without it the sum is negative.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two experiment properties the panel could not state, both of which the analysis it
drives has an opinion about anyway.

The polarization factor is one number that corrects both communities' output - the
azimuthal profile through AzimuthalIntegrationMapping and the integrated intensities
through BraggIntegrationEngine - so it goes in the Geometry section, which the MX and
AzInt pages already share. Files carry no polarization factor at all, so before this
the interactive analysis integrated with none while "Analyze dataset" applied 0.99
from the rugnux defaults: the panel showed nothing and the two front ends disagreed.
The viewer's starting experiment now takes the same rugnux defaults, so the panel
shows what the analysis actually uses, and a processing job takes the panel's value
over the default - as it already did for the scaling fields.

The Goniometer section states which of the three things a dataset is - a still, a
rotation, or a grid scan - which is exactly the choice a file makes at
/entry/sample/transformations/omega vs /entry/sample/grid_scan vs neither. That
choice IS the rotation/stills switch, so "Process as stills" is gone from the
Indexing section rather than sitting beside it as a second control. It is not tied
to what the file says: a still file can be given an axis or a grid, and a rotation
file can be processed as stills. The inactive modes grey out but keep their values,
so switching away and back does not lose an axis; a file that names none offers
omega / -1 0 0 / 0.1 deg and a 10-point, 20 um raster.

Only the fast axis of a grid gets a count field, because that is all there is:
GridScanSettings derives the slow one from the image count, and the file stores
n_fast alone. The grid the settings make is spelled out under them instead. The
steps are signed - the sign is the direction the scan runs in - so only zero is
rejected, and a half-typed entry falls back to the default rather than throwing out
of a widget signal, which would abort the viewer.

JFJochReader::UpdateGeomMetadata carries a fixed whitelist of the fields a panel edit
may change, and it grew by three. Without them the new controls reset themselves on
the next dataset refresh, since a non-whitelisted field comes back as the file's.
The function has two call sites, both in the viewer's reading worker, and it is not
virtual: objdump on the rugnux binary shows the code linked in and never called, so
offline processing is untouched.

RugnuxCommandLine emits --force-still from the axis the experiment carries, which the
panel now clears when the mode is not Rotation - so the worker remembers the axis the
file was opened with, and the copied command line is built with that put back. rugnux
has no flag for the axis itself, so an axis edited or invented in the panel still
cannot be expressed on a command line; the in-process run gets the experiment object
and is unaffected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ApplyCellSurface fits one multiplicative factor per cell by least squares with the
OBSERVATION on the regressor side - A = sum w Is Iref / sum w Is^2, the slope that
carries Is onto Iref. A least-squares slope is attenuated by the noise in its own
regressor, here by 1/(1 + (sigma/I)^2), and an observation carries all of a
reflection's noise where the reference carries about 1/n of it. So every cell is
pulled towards zero by an amount set by its own signal-to-noise - and on a detector
that is a function of radius, which is to say of resolution. The gauge fix then
spreads the ramp over the whole surface and the alternating rounds compound it.

Nothing downstream catches it. The cross-validation splits by frame parity, and a
bias that depends only on a cell's signal-to-noise is identical in both halves. And
the held-out score is sum|Is - Iref| / sum Iref over the whole resolution range,
which any resolution-dependent scale lowers without tightening a single reflection:
applying a scale that is purely a function of d to a merged 360 deg sweep leaves
every resolution shell's R_meas unchanged to 0.05 pp and takes the run's overall
R_meas from 42.4% to 27.5%. The surface finds that manipulation because its own bias
points exactly along it, and reports it as a 40% held-out gain.

The cost is large wherever a sweep was taken past its signal. On one such run the
fitted detector-plane "flat field" ran from 0.25 to 4.0 with 37% of its cells pinned
at the low clamp - an 11x centre-to-edge ramp - and on the same combined fulls it
moved the merge: low-resolution R_meas 8.8% -> 15.5%, <I/sigma> 33.2 -> 13.4, CC1/2
99.65 -> 98.87, error model b 5.0e-03 -> 4.3e-02, ISa 11.0 -> 3.5. The program's own
--no-scaling-corrections run agrees on the same fulls and the same space group
(b 5.4e-03 ISa 10.6 against b 4.1e-02 ISa 3.6). Regressing the other way round
restores all of it - 8.6%, 33.6, 99.65, 4.9e-03, ISa 11.1 - and keeps the surface's
real gain in the middle shells, where CC1/2 rises by 1-2 points.

Where the data are well measured everywhere the two estimators are indistinguishable:
on a control crystal the two surfaces agree to 0.02 pp in every shell and 0.02 in
ISa, which is what a flat field should look like. Recovering a synthetic +-20% ripple
imposed on the same fulls: 0.18 rms in log against 0.91 for the shipped form on the
weak-outer-shell crystal, 0.059 against 0.082 on the well-measured one, with the
spurious correlation between the fitted factor and detector radius falling from -0.30
to -0.01.

The same estimator serves the absorption surface and any resolution-indexed surface
fitted through this function, where the bias lands directly on the resolution axis:
on a crystal the program itself reports as having no radiation damage (total dB
0.00 A^2), a batch x resolution-shell surface fitted the old way manufactures a
monotone 12% falloff from low to high resolution out of nothing, and the new way
gives 1%.

The per-frame scale (FitPerFrameG) already regresses this way round.

Full 38-crystal rotation battery against the same binary without it. Read it with
the next commit, which supplies the correction this one stops faking; alone it
is a partial state:

  observations  better 34 / worse  4,  +22 940
  R_meas_lo     better  4 / worse  8,     -5.9
  ISa           better 20 / worse 14,   +32.66
  CC1/2         better  2 / worse 10,    -11.4
  R_meas        better  1 / worse 30,   +114.3
  space groups  unchanged at 35/38

Overall R_meas RISES, and that is the artefact leaving rather than arriving: it
is a ratio of sums across every shell, so a resolution-dependent scale lowers it
without one reflection getting tighter - measured, a scale that is purely f(d)
leaves every shell's R_meas unchanged to 0.05 pp while moving the run's overall
value from 42.4 to 27.5 per cent. That is precisely the shape of the bias, which
is why the surface's own held-out score read it as a 40 per cent gain.

CC1/2 falling on ten crystals is not covered by that argument and is the reason
this commit is not defensible on its own: the ramp was partly standing in for a
real time-dependent absorption that nothing else modelled. With that correction
supplied by the following commit the same battery gives CC1/2 -1.4 and
R_meas_lo -26.0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
RefineAbsorption indexes its surface by the diffracted direction
de-rotated into the crystal frame, deliberately without a time axis, and
RefineModulation indexes its by detector position, also without one.
Nothing is indexed by (rotation, detector position), so the part of the
absorption that changes as the crystal turns has no parameter at all.

For a rigid absorber illuminating a fixed volume that is the right
model: the incident path is a function of the spindle angle alone and
the per-image scale takes it, and the exit path is then fixed in the
crystal frame.  What breaks the factorisation is the diffracting volume
moving - a crystal larger than the beam, a mis-centred loop, ice
building up.  The exit path then depends on the spindle angle as well as
the direction, and no time-independent surface reaches it.

Measured on 34 rotation datasets, on XDS's own uncorrected intensities,
as what is left after the crystal-frame absorption and detector
modulation surfaces have taken what they can.  The cross-validation gate
lets the surface engage on 22 of them.  Scored per resolution shell -
the gate's whole-range ratio is lowered by any resolution-dependent
scale without a reflection getting tighter, so the honest readout is
each shell's own ratio, which a per-shell scale leaves unchanged - the
median engaged crystal gains 4.1 %, the set gains 117 % summed against
19 % of damage, and 3 of the 22 are hurt.

The surface has to be smooth in rotation angle to be absorption at all,
and it is: the lag-1 autocorrelation of the fitted factor along the time
axis runs +0.32 to +0.71 on the crystals it engages, against -0.08 for
the same surface with its time bins shuffled.  Where it is not smooth it
is fitting something else, and says so - on a sweep whose beam was
obstructed for a 70 deg wedge the autocorrelation is +0.16 and the
profile is a cliff at the wedge, not a turn.

Two null controls.  Assign every observation a random cell and the gate
refuses it (-1.2 % to -3.8 %).  Keep the detector bin and shuffle only
the time bin - a surface that cannot contain any time-dependent
information - and the gate refuses that too, at +0.04 %, -0.60 % and
+0.33 % on three crystals.  Against the real surface's +3.7 % to
+14.8 % on the same three.

12 time bins x a 10 x 10 detector grid = 1200 factors.  On the per-shell
score the median gain moves only between 3.1 % and 4.1 % across grids
from 216 to 2400 cells, so the grid is second order; 12 x 10 has the
largest net and the fewest crystals hurt.  Equal-occupancy detector
bins, not equal width: an equal-width grid starves the edges and the
corners, and a starved cell is where a free surface over-fits.  Fitted
last, so the two time-independent surfaces get first claim on what they
can explain.

QUALIFICATION, measured after this was written: the "33 better / 0 worse" above is
overall R_meas, which is a ratio of sums across every shell and is therefore
lowered by any resolution-dependent scale without a reflection getting tighter -
the same property that let the estimator bias pass its own gate. Scored per
resolution shell instead, this surface HURTS 6 of 18 crystals under the
acceptance gate as it currently stands, because that gate shares the defect and
admits the surface where it should not. With a per-shell gate the surface is
refused on exactly those crystals and its net over the chain goes from +86.3 to
+159.5 per cent with none worse. The correction is right; the gate that decides
where to apply it is the next commit's problem, not this one's.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Full 38-crystal rotation battery against its own matched baseline - the same
binary with the corrected estimator and without this surface:

  R_meas        better 33 / worse  0,   -78.7
  R_meas_lo     better 24 / worse  2,   -20.7
  CC1/2         better  8 / worse  0,    +9.6
  ISa           better 30 / worse  2,  +99.15
  space groups  unchanged at 35/38

The low-energy datasets gain most, which is what absorption should do: at 5-6 keV
one crystal goes ISa 24.68 -> 37.42 and another 14.21 -> 22.08, while the same
protein measured at 13 keV moves 13.41 -> 14.83.

Taken with the estimator fix it precedes, against a clean baseline: R_meas_lo
better 25 / worse 3 summed -26.0, ISa better 31 / worse 2 summed +113.1,
outer-shell CC1/2 +83.0, observations +25 880 on 36 crystals of 38, CC1/2 flat at
-1.4 and no space group moved. That last number is the point of the pair: the
estimator fix alone costs CC1/2 -11.4, because the ramp it removes was partly
standing in for this correction.

One cost, predicted in advance and still unexplained: outer-shell CC1/2 falls on
three of the four low-energy crystals, by 15.8 points on the worst, while every
other statistic on those same crystals improves. The fourth goes up. On 5000-9000
That outer-shell fall has since been attributed, and it is not this surface: with
the merge's 6-sigma outlier rejection turned off, the sign flips on every crystal
that lost, +20.5 and +20.7 where it read -15.8 and -8.0. The surface removes most
of the deviants in sample - it is fitted on all the data and applied to it, with
no robustness of its own - so the merge's cut stops firing and the survivors land
in a shell whose multiplicity is about three. Last-shell CC1/2 is largely made by
that cut: one crystal's baseline goes 10.4 to 92.7 purely by dropping 19 per cent
of the shell.

A second qualification, measured after the numbers above were taken: overall
R_meas is a ratio of sums across every shell, so any resolution-dependent scale
lowers it without a reflection getting tighter - the same property that let the
estimator bias pass its own gate. Scored per resolution shell instead, this
surface hurts 6 of 18 crystals under the acceptance gate as it stands, because
that gate shares the defect and admits the surface where it should not. Under a
per-shell gate it is refused on exactly those crystals and its net over the
correction chain goes from +86.3 to +159.5 per cent with none worse. The
correction is right; where to apply it is the gate's problem.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The monitor fitted each batch's relative-B on SINGLE observations -
ln(I_ref/I_obs) regressed on s^2, weighted by (I_obs/sigma)^2, with the
logarithm requiring I_obs > 0.  The observation therefore sits in the
response and in its own weight, and the positivity requirement keeps only
the upward half of the noise, so the estimate is biased downwards wherever
I/sigma approaches 1 and is unbounded in the limit.  Simulated: on a batch
with no relative-B at all and <I/sigma> = 0.3 it reads -28 A^2; on a batch
whose true relative-B is +30 A^2 it reads -29.  The bias grows with dose,
so it inverts the answer on exactly the data the number exists for.

That is not a corner case.  Re-measured on stored integrated intensities,
a 360 deg sweep obstructed over a 60 deg wedge - whose honest curve is flat
for 130 deg, dips over the wedge and comes back - printed a per-batch curve
saturated at -31.4 A^2 for twelve consecutive batches (the +-50 A^2 clamp
less the low-dose anchor, "no data here" reported as a measurement) under a
headline of -19 A^2 of radiation damage.  A deliberately dosed dataset
printed -39 A^2 where the honest measurement is about +34: the one crystal
with real damage got the sign wrong.  Eight of thirty-eight datasets
reported |dB| > 5 A^2 and their curves oscillate by tens of A^2.

So pool the observations into ten equal-occupancy resolution shells per
batch before taking the logarithm, and fit slope AND intercept over the
shell means, weighting each shell by its own pooled (I/sigma)^2.  A shell
mean is well determined where a single observation is not, it admits
negative intensities, and it carries the I/sigma that says whether the
batch can be measured at all.  The same simulations then reproduce the
truth to under 1 A^2 at every signal level.  The intercept keeps a batch
that is merely dimmer than the run - an attenuated beam, a mis-fitted frame
scale - out of the damage number: a batch mis-scaled by 2x read +12 A^2 of
"damage" without it and +0.05 with it.

A batch whose shells are too weak to fit is now absent from the curve,
printed as "-", instead of pinned to the clamp.  And the clamp itself is
now an argument of the solve rather than one shared constant: it guards
against divergence, and the correction keeps the bound it was tuned with,
but with the estimator fixed a heavily dosed crystal's honest relative-B
runs past it - the monitor pinned thirteen consecutive batches at +49 A^2,
which is the same defect in the other direction.  The monitor is given room
a real relative-B cannot reach and drops any batch that lands on it anyway.

The shells are laid inside the range the run actually diffracted to,
not across the whole merged range: a resolution limit taken from another
program or left generous spends most of an equal-occupancy grid on noise
and leaves a batch with too few shells to fit at all - on the battery that
silenced three crystals outright and cost two of them nineteen batches of
thirty-six.  Where the merged range already sits inside the signal the grid
is unchanged and so is every number.

The first->last headline
is reported only where a straight line explains at least half of the
curve's variance, or where the curve is flat to within a couple of A^2 and
the answer is simply "no damage"; otherwise there is no headline and the
report says the loss was not dose and points at the sweep-quality section.
The three shapes separate cleanly - progressive damage R^2 0.97, the
obstructed sweep 0.24, the clean control flat at +0.35 A^2.  And the label
now follows the sign: damage fades the high-resolution intensity, so only a
positive change is dose, where before any |dB| > 5 was called damage.

Report-only throughout - the monitor never touches corr, and the per-batch
curve's only consumer beyond the report is a sweep-quality field no reader
reads; classification runs on the per-frame scale and CC, and is unmoved.
The decay correction's global slope and the opt-in per-batch relative-B
share this estimator and are left alone here: they fold into the scale, so
Full 38-crystal rotation battery, twice (the second confirming the shell
placement), against a clean baseline at the same base:

  space groups   unchanged at 35/38
  merge metrics  move on three crystals only - the same three whose two-pass
                 lattice search takes a different branch on nearly every arm run
                 this session, one of which moves its own R_meas by 1.5 points on
                 thread count alone

A report-only change ought to be bit-identical and this is not quite, which is
worth saying plainly: the three crystals that move are the known unstable ones
and no space group moves, but "identical except where nothing is ever identical"
is a weaker statement than "identical", and the residue has not been chased to
ground.

The three validation cases behave as they must:

  60 deg beam-obstructed wedge, no decay   -23.02, labelled damage, twelve
                                           batches printing the clamp
                                        -> NOT_A_TREND, curve within 3 A^2, the
                                           two unmeasurable batches absent, and a
                                           pointer to the sweep-quality section
  genuine progressive damage               -39.46, sign inverted
                                        -> +94.50, monotone, corroborated by a
                                           per-image CC that falls 0.608 -> 0.159
                                           and never recovers
  clean control                            +0.18 -> +0.76, flat within 1 A^2

Across the battery the report now names four crystals as radiation-damaged
instead of ten; the other three are the two lowest-energy datasets and the
pink-beam one, each showing a monotone rise of about ten square Angstroms.

Sweep-quality classification is untouched, and the coupling that was assumed to
exist does not: rad_damage_b_batch reaches it through one field that is written
and never read. Ranges and reasons are identical on 35 of 38, the three that
differ by one to seven frames are the same unstable crystals, and the census of
stretches called radiation damage is one before and one after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A rotation run makes the space-group decision twice: once on a merge of only
the well-measured observations (--search-min-zeta) and once on all of them, and
the rule was to keep whichever search found more symmetry. Its justification
was that the filtered arm can only ever LOSE an operator - discarding 40-80% of
the observations starves the operator correlations - and never invent one.

That premise was checked at one set of geometries, and there it holds exactly:
over 1232 stored battery runs (38 geometries x 32 code variants) the two arms
disagree 59 times and the all-observation arm is the higher one every single
time. Away from those geometries it fails. Over 68 runs whose first pass was
given a displaced beam centre, the filtered arm confirms an operator the full
merge refuses 14 times, and the two arms never once both confirm the promotion.
On one rhombohedral crystal the all-observation merge refuses a 3 -> 32
promotion at the correct beam centre, on the same twin-law statistic the search
uses everywhere (1.78 against a bound of 1.70), and the filtered merge - which
had thrown two thirds of the observations away - overrides it into the wrong
space group. The rule was validated on the only data that cannot test it.

So the filtered arm no longer promotes. Where the two disagree the answer is
the one the merge with all the observations supports, as the systematic
absences already were, and both are still reported. The filter keeps the job it
was added for - stopping a near-tangential measurement from making a real
operator look like a twin law - it simply cannot outvote the merge that has
every observation in it.

Costless on the 38-crystal battery, as predicted: the changed branch is never
taken there (41 searches: 39 agree, 2 with the all-observation arm higher, none
the other way), the space groups stay at 35/38 with the same three misses, and
the crystals the second search does rescue - where the all-observation merge is
the one finding the higher symmetry - are untouched. Injecting the post-refined
beam centre into the crystal above now yields the right space group, 27900
unique reflections at R_meas 12.6% against 14078 at 21.3% before.
The second pass is meant to re-integrate at the refined geometry, not to re-decide the
symmetry: it reuses pass 1's space group so a borderline determination cannot flip between
the two passes. A group that came from the intensity-based centred-lattice test was excluded
from that, on the grounds that its conventional setting is not the frame the de-novo indexer
returns - but the primitive->conventional reindex that handles exactly that case, and the
centring check that refuses a pass whose lattice the group cannot describe, were both already
in place. The exclusion only switched them off, and with them the reuse itself, leaving pass 2
free to search again.

On a crystal whose C-centring is pseudo-symmetric to 2.4% of axis-length equality - against
the 3% tolerance the Bravais classification allows - that re-search is a coin toss. Pass 1
confirmed the centring from the intensities and pass 2, a quarter of a millimetre of refined
detector distance away, classified the same lattice as triclinic and merged the run in P1,
doubling the asymmetric unit. With the group carried over, pass 2's own lattice is reindexed
into it where the metric allows and the pass is dropped where it does not, which is what
happens for every non-reindexed group already.

Only a run whose pass-1 group came from that test can change; every other run takes the
identical branch.
The 44 lattice characters come in two types - all-acute reduced cells and all-obtuse ones - and
the search skips the type-1 characters whenever the reduced beta is within the angle tolerance of
90 degrees, because at that point the two types are no longer distinguishable and only the type-2
statement of a character is safe to test. But a cell whose reduced beta really is 90 within
tolerance can itself reduce EITHER way, and an all-acute one then matches no character at all: the
type-1 ones were skipped and the type-2 ones are written for the obtuse setting. It falls through
to triclinic, and a centred lattice loses its centring.

Measured on a C-centred monoclinic crystal whose reduced beta sits 0.07 degrees from 90. Its
free-refined cell came back acute or obtuse depending on the last digits of the refinement - a
sub-pixel change in the beam centre was enough - and with it the lattice was read as C-centred
monoclinic or as triclinic, and the run merged in C2 or in P1 with twice the asymmetric unit. In
the obtuse setting the character matches with residuals of 0.01 to 0.14 degrees against a 3 degree
tolerance, at every geometry tried; there was never any doubt about the lattice, only about which
side of the boundary the reduction landed on.

So when nothing matched and the cell is acute with beta at the boundary, present it in the obtuse
setting and match once more. Negating a and c leaves the lattice and beta alone and turns alpha
and gamma into their supplements. The second attempt runs only where the first found nothing, so
no lattice that is classified today can be re-classified by this.
The retry added in the previous commit was unreachable - character 44 fits any cell, so the match
always succeeds and "nothing fits" arrives as a triclinic answer, not as no answer. Retry on a
triclinic answer instead, and only replace it when the second attempt finds a real class.

That makes the gate matter, and the angle tolerance is the wrong one to use for it. It says how far
a metric may sit from an ideal one and still be called it; the question here is whether the two
Niggli types are interchangeable at all, which they are only when the reduced beta is 90. Two
degrees off the boundary is a real type-1 cell, and presenting it in the obtuse setting promotes a
general triclinic lattice to C-centred monoclinic on residuals of about two degrees - which is what
happened to the triclinic case in the unit tests. Half a degree separates that from the crystal this
was found on, which sits 0.07 degrees from the boundary and matches on 0.006 to 0.135.

With the retry reachable and gated, that crystal's lattice comes out C-centred monoclinic from the
first pass at every geometry tried, including the one where its cell reduces to the acute setting
and the run used to merge in P1.
docs: changelog for the lattice-search and two-pass symmetry fixes
Build Packages / build:viewer-tgz:cpu (push) Successful in 19m38s
Build Packages / build:viewer-tgz:cuda (push) Successful in 19m52s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 22m58s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 24m36s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m39s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m57s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 29m5s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 20m40s
Build Packages / XDS test (durin plugin) (push) Successful in 11m15s
Build Packages / build:rpm (rocky9) (push) Successful in 21m10s
Build Packages / Generate python client (push) Successful in 35s
Build Packages / build:rpm (rocky8) (push) Successful in 25m19s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m24s
Build Packages / DIALS test (push) Successful in 19m58s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m45s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 24m51s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m34s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m30s
Build Packages / Unit tests (push) Successful in 1h21m48s
Build Packages / build:windows:nocuda (push) Successful in 18m37s
Build Packages / build:windows:cuda (push) Successful in 39m25s
b34aabe39f
Both are user-visible: a centred lattice whose reduced beta sits on the Niggli type boundary
keeps its centring, and the second rotation pass reuses a group that came from the
centred-lattice test instead of re-deciding the symmetry at the refined geometry.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The centre in the file is often a placeholder, and nothing measures it until post-refinement
has already indexed the sweep - by which time a wrong centre has chosen the lattice. Two exact
facts about a rotation sweep give it from spot positions alone, with no cell, no orientation
matrix and nothing indexed.

Rotating 180 degrees about the spindle and taking -h negates a reflection's component along the
spindle and leaves the rest, so with the spindle perpendicular to the beam the Laue condition is
preserved and the spots recorded half a turn apart are mirror images along the spindle. Those are
Friedel mates, not the same reflection. The same reflection appears twice for a different reason:
it meets the Ewald sphere on two crossings, generally not half a turn apart, differing only in the
sign of the component perpendicular to both the spindle and the beam. The first observable gives
the coordinate along the spindle, the second the coordinate across it. Each candidate pairing
votes and the true value accumulates while wrong pairings scatter.

Both observables need guarding, because a vote is a comb and the tallest tooth is not always the
right one. Along the spindle a false pairing cannot fake the equality of Friedel amplitudes.
Across it, the two crossings of one reflection are separated by a sweep angle its own position
fixes, which no accidental pair reproduces.

The mirror is exact in the laboratory frame, so it is only as good as the rotation axis. Every
file here states an ideal axis and none of them has one; a skew about the beam spreads the vote
instead of shifting it, and past a milliradian it moves an otherwise correct answer by pixels
while every internal statistic still looks healthy. It is therefore fitted, not assumed. A tilt of
the axis towards the beam is measured and reported but not applied, being confounded with the
detector rotation until that is fitted too.

Nothing inside the fit can see a wrong tooth - when the vote flips, every frame pair flips with
it - so the answer is checked from outside, by asking whether it depends on where the search
began. That, and a floor on the angular span the pairs cover, are what refuse the cases this
cannot measure: a sweep barely past half a turn is the dangerous one, not the short one, because
at exactly half a turn there is nothing to fit and just past it there is almost nothing.

Where the sweep is too short for any of this the radial background profile gives a coarser centre
from a handful of images, and where neither can measure it the file's value is kept.

The beam-stop projection now takes its own frames rather than sharing the sample, so turning this
on cannot change the mask; and both samples keep away from the ends of the sweep, where shutter
synchronisation spoils an image. Reading twice as many frames as before costs a few seconds once,
and is what makes the answer independent of which frames were drawn.

Off by default. Over the 38-crystal rotation battery it serves every dataset, agrees with XDS's
refined direct beam to 0.116 px in the median against 0.135 for the value in the file, and changes
no space group.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The Laue condition fixes the component along the beam; a half turn about the spindle negates the
two perpendicular to it and taking -h negates all three, so the composition leaves the beam
component alone and the mate is on the sphere exactly, not nearly. Says that they are separated
by half a turn rather than recorded together precisely because the Ewald sphere is curved - a
near-flat one would excite both at once - and that the positions need only the reciprocal lattice
to be centrosymmetric, Friedel's law entering as the test of a pairing rather than its basis.

Also states the yield: a sweep of S degrees gives S-180 degrees' worth of pairs, so half a turn
gives none and a little over half a turn gives few.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The radial walk decides how far out the beam stop blocks a ring entirely by comparing each ring
median against the largest background further out, and it refuses to judge a ring holding fewer
than MIN_RING_PIXELS valid pixels - but it was building that outward maximum over every ring,
including the ones it had just refused. At the corner of the detector a ring holds nine pixels and
its median is one pixel's mean, so a single recorded reflection out there becomes the background
that every ring inside it is compared against.

Measured on the 38 rotation regression crystals, sampling 59 frames instead of 60: on one crystal
the outward maximum moved from a 2544 pixel ring at r = 437 (median 12.53) to a 9 pixel ring at
r = 1155 (median 32.37), the threshold went up 2.6x with it, and the walk marched from r = 32 to
r = 391 - 33 223 masked pixels became 182 661, everything out to about 13 A declared to be inside
the beam stop. The ring medians the walk actually reads changed by less than 1% between the two
samples; only the reference did.

Taking the maximum over the countable rings alone leaves the mask bit-identical on 38 of 38 at the
default 60 frames and on 37 of 38 at 59 frames, the exception being the repair. Over the 114
samples of 38 crystals at 59, 60 and 61 frames, the outward maximum was set by an unjudgeable ring
8 times; 7 of those were within six pixels of the beam centre, where the suffix maximum does not
reach the rings that matter, and the eighth is the failure above.
Viewer: offer the beam-centre measurement in the new-job dialog
Build Packages / build:viewer-tgz:cpu (push) Successful in 19m13s
Build Packages / build:viewer-tgz:cuda (push) Successful in 21m7s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 23m6s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m54s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m14s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m29s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 29m6s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 20m32s
Build Packages / XDS test (durin plugin) (push) Successful in 11m26s
Build Packages / build:rpm (rocky9) (push) Successful in 21m25s
Build Packages / Generate python client (push) Successful in 38s
Build Packages / build:rpm (rocky8) (push) Successful in 25m24s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m21s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m59s
Build Packages / DIALS test (push) Successful in 21m21s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 26m24s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m44s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m48s
Build Packages / build:windows:nocuda (push) Successful in 28m31s
Build Packages / build:windows:cuda (push) Successful in 23m44s
Build Packages / Unit tests (push) Successful in 2h32m35s
914be292ca
The measurement that landed in rugnux as --estimate-beam-center is a
whole-dataset pre-scan, so it belongs with the other run options in the
"Analyze dataset" dialog rather than in the always-visible settings
panel, whose contents are all properties of the experiment. It sits
under "Detect beam stop", which shares the same pre-scan and is the
checkbox it is modelled on, and it is off by default as in the CLI.

The CLI's --estimate-beam-center also fits the spindle's skew about the
beam (--no-fit-spindle is the deviation from it), while the bare
ProcessConfig default leaves that off; the one checkbox therefore sets
both, so the viewer runs what the flag it is named after runs. The
sub-option itself is not exposed.

RugnuxCommandLine emits the flag too, so the dialog's "Copy command"
reproduces what "Run locally" would do.
A stride cuts every rocking curve. The angles still come out right - the goniometer is
shifted so one ordinal is one stride - but a reflection's partials are then sampled every
k-th frame, so the combine sees a fraction of each event, the captured fraction collapses,
and the partiality divides by a width the sweep never delivered. There is no reading of a
strided rotation sweep worth having, so it is refused rather than processed into a
plausible wrong answer. --mode azint, --mode calibration and --force-still are unaffected;
each returns before the rotation decision or does not combine partials.

The beam-centre pre-scan is corrected with it. PreScan was the only place in the file that
asked the goniometer for an angle by ORIGINAL image number rather than by ordinal, while
the goniometer has already been shifted so that an ordinal maps to the original image it
came from - so its angles carried a constant offset whenever -s was given, and, before the
refusal above, counted the stride a second time on top of the increment that already has
it. Harmless in both cases for the pairing, which consumes differences only, but it is the
one site that disagrees with the rest and the next reader would copy it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The comment justified the rule by informative missingness alone - a pixel over the
detector's range goes missing BECAUSE the reflection was bright - and explicitly excused a
gap, a mask and the sensor edge as unbiased. The code cuts on all of them, so the comment
described a narrower rule than the one implemented.

The code is right and the reason was incomplete. Renormalising over the surviving pixels is
unbiased only while the profile MODEL is exact; lose the peak and the amplitude is set by
the wings alone, so the result stops being a measurement of the reflection and becomes a
measurement of how well the fitted shape describes it. That holds whatever made the pixel
unreadable. The overload remains what motivated the rule and what biases hardest, since the
loss then concentrates on the strong low-resolution reflections that are the largest terms
of R_meas.

Comment only; no logic changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
docs: summarise the rc.161 changelog, and give CPU_DATA_ANALYSIS its proportions back
Build Packages / build:viewer-tgz:cpu (push) Successful in 10m2s
Build Packages / build:viewer-tgz:cuda (push) Successful in 17m41s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 19m53s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 19m55s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 24m47s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 25m36s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 22m44s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 16m52s
Build Packages / build:rpm (rocky9) (push) Successful in 19m45s
Build Packages / XDS test (durin plugin) (push) Successful in 10m37s
Build Packages / build:rpm (rocky8) (push) Successful in 25m12s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m33s
Build Packages / Generate python client (push) Successful in 35s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m30s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 23m6s
Build Packages / DIALS test (push) Successful in 18m59s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m10s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m28s
Build Packages / build:windows:nocuda (push) Successful in 17m55s
Build Packages / build:windows:cuda (push) Successful in 22m35s
Build Packages / Unit tests (push) Successful in 1h55m18s
9f86384210
The rc.161 changelog had grown to 88 entries and 3080 words, two to eight times any other
release in the file and unreadable as a release note. It is now four: rugnux, the viewer,
GPU image decoding, and the broker/writer/packaging fixes. The three breaking-change blocks
are kept in full - a consumer upgrading needs them - and gain the stride refusal.

CPU_DATA_ANALYSIS.md had drifted the same way in miniature. Box summation, whose own title
says it is a seed and a fallback, ran to 1204 words, more than the whole of indexing;
measuring the direct beam, which is off by default, outweighed both post-refinement and the
lattice search. The trims take out the evidence for past decisions - battery ranges,
empty-aperture pedestals, percentages of recovered observations - and keep every statement
they supported, on the rule that a measurement belongs in the commit that made it. Two
textbook asides and a note about bit-compatibility with earlier builds go entirely.

The reverse imbalance is fixed too: space-group determination, twinning, outlier rejection
and the resolution cutoff were bullets inside "Practical notes and limitations", which is
not what any of them is. They are now 13.1 to 13.4, with the practical notes as 13.5, so
section 14 does not move and no cross-reference breaks. One that was already broken is
repaired - the error-model section pointed at the space-group search as 11, which is
mosaicity - along with a typo, a tautological formula and a duplicated scope-list number.
Net 18200 to 16700 words.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rugnux: the pre-scan says what it is doing
Build Packages / build:viewer-tgz:cpu (push) Successful in 12m15s
Build Packages / build:viewer-tgz:cuda (push) Successful in 15m45s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 16m36s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 19m47s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 20m16s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 14m47s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 17m13s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 20m55s
Build Packages / build:rpm (rocky8) (push) Successful in 22m2s
Build Packages / build:rpm (rocky9) (push) Successful in 18m21s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 20m26s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 15m17s
Build Packages / Generate python client (push) Successful in 13s
Build Packages / XDS test (durin plugin) (push) Successful in 9m1s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 40s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m42s
Build Packages / DIALS test (push) Successful in 15m17s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m35s
Build Packages / build:windows:nocuda (push) Successful in 18m14s
Build Packages / Unit tests (push) Successful in 1h23m22s
Build Packages / build:windows:cuda (push) Successful in 20m59s
cac669495d
The beam-stop projection and the beam-centre measurement both run before the first image of
the run proper is processed, and neither reported a phase - so a viewer job sat at "queued"
for the whole pre-scan, which is tens of seconds once the beam centre asks for its own
frames on top of the projection, and longer again on the sparse re-read that reads several
times as many. The observer reached RunPipeline but was never passed down.

Two phases, because the two long reads are separated by a fit that can decide it needs more
data: the sample the projection and the symmetry share, named for whichever of the two asked
for it, and the sparse re-read. The fits themselves are fast and are not called out.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Viewer: every panel section starts folded, and calibration carries the azimuthal settings
Build Packages / build:viewer-tgz:cpu (push) Successful in 19m17s
Build Packages / build:viewer-tgz:cuda (push) Successful in 20m58s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 21m57s
Build Packages / build:windows:nocuda (push) Successful in 23m25s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 24m14s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m25s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m26s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 29m26s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 20m52s
Build Packages / XDS test (durin plugin) (push) Successful in 12m1s
Build Packages / build:rpm (rocky9) (push) Successful in 21m4s
Build Packages / Generate python client (push) Successful in 40s
Build Packages / Build documentation (push) Successful in 1m12s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (rocky8) (push) Successful in 25m9s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m53s
Build Packages / DIALS test (push) Successful in 21m14s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m44s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m32s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m13s
Build Packages / build:windows:cuda (push) Successful in 32m13s
Build Packages / build:windows:nocuda (pull_request) Waiting to run
Build Packages / build:windows:cuda (pull_request) Waiting to run
Build Packages / Unit tests (pull_request) In progress
Build Packages / build:viewer-tgz:cpu (pull_request) Successful in 16m9s
Build Packages / Unit tests (push) Successful in 1h30m48s
Build Packages / build:viewer-tgz:cuda (pull_request) Successful in 18m19s
Build Packages / build:rpm (ubuntu2404_nocuda) (pull_request) Successful in 20m42s
Build Packages / build:rpm (rocky9_nocuda) (pull_request) Successful in 21m5s
Build Packages / build:rpm (rocky8_nocuda) (pull_request) Successful in 24m42s
Build Packages / build:rpm (ubuntu2204_nocuda) (pull_request) Successful in 24m50s
Build Packages / build:rpm (rocky9_sls9) (pull_request) Successful in 22m6s
Build Packages / build:rpm (rocky8_sls9) (pull_request) Successful in 27m44s
Build Packages / build:rpm (rocky9) (pull_request) Successful in 24m45s
Build Packages / DIALS test (pull_request) Successful in 21m17s
Build Packages / build:rpm (rocky8) (pull_request) Successful in 28m21s
Build Packages / Generate python client (pull_request) Successful in 39s
Build Packages / Create release (pull_request) Skipped
Build Packages / Build documentation (pull_request) Successful in 1m38s
Build Packages / build:rpm (ubuntu2404) (pull_request) Successful in 23m38s
Build Packages / build:rpm (ubuntu2204) (pull_request) Successful in 27m43s
Build Packages / XDS test (durin plugin) (pull_request) Successful in 10m5s
Build Packages / XDS test (JFJoch plugin) (pull_request) Successful in 8m48s
Build Packages / XDS test (neggia plugin) (pull_request) Successful in 7m39s
93093c5f21
A CollapsibleSection was born expanded and each caller folded it again, so the three that
never got round to it - geometry, unit cell, goniometer - greeted every start with three
open accordions and the rest of the settings pushed below the fold. The default is now
folded, both panels open as a list of headers, and the ten setExpanded(false) calls that
existed only to undo the old default are gone. The one remaining call is the ROI section
opening itself when ROIs appear, which is a real behaviour and still works.

Azimuthal integration was a page of the MX/AzInt/Calib stack, which left a calibration run
unable to reach the settings it depends on: calibrating from powder rings integrates the run
in azimuthal sectors and over the same Q range and spacing as any other integration, and the
Calib page could only report what the AzInt page had been set to. The section now lives
beside geometry, outside the stack, shown for both pages and hidden on MX - one set of
widgets over one AzimuthalIntegrationSettings, so there is nothing to keep in sync. The
too-few-sectors note moves into the powder section and appears only when the count is below
the four the rings fit needs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
leonarski_f merged commit 538f3504d3 into main 2026-08-13 17:03:10 +02:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: mx/Jungfraujoch#71