Commit Graph
272 Commits
Author SHA1 Message Date
leonarski_fandClaude Opus 5 bc1c4c6800 Rotation: land the rest of the bandwidth term
f4e281b2f described this change in full but committed only one of its six files.
What went in was RotationScaleMerge.cpp - the merge widening the partiality it
recomputes from the smoothed mosaicity. That is precisely the part which is unsafe on
its own, by the original message's own argument: without the mosaicity fit subtracting
the term before fitting, the bandwidth is counted twice, and without the predictor
widening its acceptance window, the partiality the merge recomputes no longer matches
the one integration measured.

Add the five files that were left behind: the rotation predictor and its GPU twin
widen the acceptance window and the partiality handed to integration, the settings
struct carries the term, and CalcMosaicityXDS deconvolves it before fitting so what it
returns is the intrinsic mosaicity rather than the mosaicity plus the beam.

Monochromatic data is untouched by construction - every hunk is guarded on a non-zero
bandwidth, which is read from incident_wavelength_spread or --bandwidth and is absent
from every dataset in the rotation battery. Verified on the one dataset that has a
bandwidth: at --bandwidth 0, the merge table is identical to the branch tip; with the
bandwidth set, the fitted mosaicity drops 0.0718 -> 0.0694 deg as the deconvolution
takes effect and CC1/2 in the outermost shell recovers 30.3 -> 31.4%.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 22:26:26 +02:00
leonarski_f 0b5fb4fb92 Merge branch 'fix56-work' into integration-variance-fixes 2026-08-09 21:08:39 +02:00
leonarski_fandClaude Opus 5 1239c49731 Bragg integration: separate the three things a bandwidth used to switch
Setting a bandwidth flipped three unrelated switches at once: it changed the profile's
radial capture term, it moved the width measurement from the signal disk to the whole
fit grid, and it silently overrode the background clip and trim, so --background-clip
under --bandwidth was ignored - the two runs were bit-identical.

The width measurement was the damaging one. The fit grid is an azimuthally averaged
stack, so its second moment is sigma_r^2 + sigma_t^2 and the radial smear of a
bandwidth leaked into the tangential model - a tangential width of 3.04 px against a
1.06 px truth, inflating the effective background pixel count where the weak signal is.
The result was a step rather than a slope: on genuinely monochromatic data, declaring a
0.2% bandwidth cost ISa 28.4 -> 22.2.

Measure the two widths separately, accumulated in each spot's own radial/tangential
frame over the signal disk, from the signed profile cells - away from the peak a
learned cell is background noise centred on zero, so the signed sum is unbiased, while
clamping it at zero turns that noise into a pedestal the r^2 weight reads as width. The
radial term is then the measured excess or the analytic floor, whichever is larger.

With the two widths separated there is nothing left for the broadband switch to select,
so it is gone - which is the proof the three were independent. The background clip and
trim now come from the settings in every case; the tuned 3-sigma broadband default
moves to the rugnux front end, which is the only place that knows whether the user gave
a value.

Monochromatic data: declaring a 0.2% bandwidth now costs ISa 28.4 -> 27.9 rather than
22.2, and forcing the old 3-sigma clip in the new build reproduces the good result, so
none of the step came from the clip. On large-bandwidth data CC1/2 improves in 8 of 10
shells. Across 12 monochromatic crystals the space groups are unchanged and CC1/2 moves
by at most 0.2 points.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 21:08:29 +02:00
leonarski_fandClaude Opus 5 3d3fb0e58b Bragg integration: stop rectifying the fitted intensity into its own variance
The profile fit weights each pixel by 1/v with v = max(bkg, floor) + max(0, I)*P, where
I is the fit's own current estimate. Rectifying it means that at true zero the plug-in
is E[max(0,I)] = 0.4*sigma rather than 0, and with sum(P^3)/sum(P^2)^2 = 4/3 for a
Gaussian the reported sigma comes out about 0.2 counts too large - always, additively.
That is nothing at sigma ~ 7 counts and 11% at sigma ~ 2, so it only shows on data
measured against roughly one background count.

Clamp the whole weight instead of the intensity: v = max(bkg + I*P, bkg/2). Simulation
of the real integrator gives claimed/true sigma 0.92-1.01 at zero intensity across
backgrounds 0.02-2.0 ct/px and 1.000-1.007 above I = 30, where the clamp never binds.
Dropping the signal term entirely instead (v = max(bkg, floor)) is exact at zero and
wrong everywhere else - 1.91 at I = 5, 4.29 at I = 30, 13.3 at I = 300 - and a test
built on systematically absent reflections cannot see that, because it only measures
zero. Removing the clamp altogether overshoots and biases the intensity, since a
downward fluctuation shrinks v at the peak and over-weights it.

The pixel variance floor was 1/12, documented as the rounding of a continuous energy.
That does not describe a photon counter: measured on raw frames at 0.065-0.082 ct/px,
var/mean is 1.042-1.045, i.e. Poisson with no digitisation term, and a digitisation
term would be additive rather than a floor. What the floor really protects is the
background estimate, which a small ring can read as exactly zero, so it belongs at the
resolution of that estimate, ~1/n_bkg. At 1/12 it multiplied the reported variance by
floor/bkg below 0.083 ct/px - a factor of two at 0.04. Set to 0.01.

Measured on systematically absent reflections, whose true intensity is zero, as
std(I)/rms(sigma) binned by background - not std(I/sigma), which is deflated by the
correlation between the plug-in sigma and the reflection's own fluctuation. On 2.78 M
absent observations at 0.16-3 ct/px the ratio goes 1.04-1.07 to 0.99-1.00. On 2.58 M at
0.005-0.6 ct/px, decomposed: the clamp carries it above 0.08 ct/px, the floor below it.
Intensities move 0.4%; this changes sigma, not I.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 21:08:29 +02:00
leonarski_f 1283e04ada Merge branch 'fix23-work' into integration-variance-fixes 2026-08-09 20:58:41 +02:00
leonarski_fandClaude Opus 5 97dbbc50b4 Merging: fit the error model on the reflections the cutoff keeps
The (a, b) fit ran over the whole merged range and the automatic resolution cutoff was
applied afterwards, so the sigma correction applied to the reflections that survive was
calibrated largely on reflections that do not. Measured on one dataset: a = 0.286
fitted over 843k reflections, 22k written. A manual --scaling-high-resolution already
restricts the population at ingest, so only the automatic path was affected.

Fit over the full range, merge, read the cutoff from that merge, refit (a, b) on the
samples the cutoff keeps, merge again. The circularity resolves by direction: the
cutoff comes from CC1/2, a correlation of the two half-set means, which the sigma scale
barely moves, so the cutoff can be read first and the sigmas calibrated on the
population it chose. One refinement, not an iteration; one extra merge pass.

Note this is invisible to the rotation battery, which passes an explicit high
resolution limit matched to XDS and so never exercises the automatic cutoff. With a
manual limit the fitted (a, b) are byte-identical to before.

The equivalent defect in the stills / offline --scale path is untouched; it is a
different engine and needs its own validation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 20:58:35 +02:00
leonarski_fandClaude Opus 5 72efb75a8c Merging: do not floor the merged sigma at the systematic term
The merged sigma was floored at b*|I|, so I/sigma could never exceed the reported ISa.
On one dataset every merged reflection came out at I/sigma <= 12.96 with a 99th
percentile of 12.77 in every resolution shell alike, while the scatter of the
observations implied about 44 and XDS reported 58.

The floor is wrong in principle. `b` is fitted from the scatter BETWEEN a reflection's
symmetry equivalents, i.e. from the part that is not common to them, so it averages
down with multiplicity exactly like the counting term. 1/sqrt(sum_w) with the
b-inflated per-observation sigma already gives b*I/sqrt(n); flooring at b*|I| puts the
sqrt(n) back. That is the whole effect: 12.96 * sqrt(21.6) = 60, against XDS's 58.

It was introduced on a comparison of our MERGED I/sigma against XDS's UNMERGED
I/sigma. XDS's own merged low-resolution I/sigma exceeds its reported ISa on 30 of the
39 reference datasets here, median ratio 1.78 and up to 4.23.

Merged low-shell I/sigma now lands where XDS's does: 22.4 -> 46.2 against 46.2 on one
crystal, 26.7 -> 115.7 against 96.6 on another, 12.5 -> 45.0 against 58.0 on a third.
Over the 38-crystal battery the space groups, the merged reflection sets, R_meas and
CC1/2 are all unchanged - every one of them is sigma-independent, which is what makes
them the right control - and <I/sigma> rises on 35 crystals with none worse.

The asymptotic estimator that fed the floor stays, for the reported ISa only, and is
repaired in the process: it subtracts a*sigma^2 rather than the raw sigma^2 (at a < 1
the difference is the same size as the b^2 being measured, which is what made it
flip between 10.9 and 62.7 on consecutive passes of the same data), it rescales each
group's variance median-unbiased before subtracting an unbiased counting term, its
I/sigma gate uses the same convention, and it is bounded by the whole-range b - an
asymptote exists to refine 1/b upward, not to report 0.3 because "strong" was selected
on a sigma scale the fit itself rejects.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 20:58:35 +02:00
leonarski_fandClaude Opus 5 f4e281b2f5 Rotation: give prediction and partiality the energy bandwidth
The rotation predictor and RotationPartiality used the mosaicity alone. Energy
bandwidth broadens a reflection's rocking curve as (dlambda/lambda)*tan(theta_B),
resolution-dependent and negligible at low angle, so on a large-bandwidth beam the
modelled reflecting range was too narrow exactly where the crystal still diffracts:
0.064 deg of broadening against a fitted 0.083, i.e. 26% at the detector edge. The
stills predictor has carried the term since it was written; only rotation was missing
it.

Add it in the three places that have to agree. The predictor widens both its
acceptance window and the partiality it hands to integration; the merge widens the
partiality it recomputes from the smoothed mosaicity; and the per-image mosaicity fit
subtracts the same term before fitting, so what it returns is the intrinsic mosaicity
rather than the mosaicity plus the beam. Without that last part the bandwidth would be
counted twice.

The term goes in without the 1/zeta of the usual expression: the erf already divides
by zeta, so adding a per-reflection width that itself carries 1/zeta would divide by it
twice - up to 20x at the minimum zeta. dphi = delta*tan(theta_B), and the zeta stays
where it was. The rotation identity dtheta/dphi = zeta was checked against a numerical
solve of the diffraction condition at four resolutions and three orientations.

Monochromatic data is untouched by construction - the term is guarded on a non-zero
bandwidth and is an assignment, not arithmetic, when there is none. Verified: 246456
reflections byte-identical through the predictor, 4.7 million rocking-fraction
evaluations with no bitwise difference, and identical merge tables end to end. The
bandwidth is read from the file (incident_wavelength_spread) or from --bandwidth, and
is absent from every dataset in the rotation battery.

On the bandwidth dataset the fitted mosaicity becomes resolution-independent
(0.0745 -> 0.0719 deg), the prediction window widens, frames per rocking event go
4.6 -> 5.3, per-image correlation to the merge rises 0.710 -> 0.725, and R_meas
improves 0.1-0.7 pp in every shell while CC1/2 falls 0.6-0.9 pp in the outer two.
Merged quality is net neutral: the combine normalises by sum(partiality), so a uniform
widening largely cancels.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 20:16:39 +02:00
leonarski_fandClaude Opus 5 e352227a2d Scaling: divide out the incident flux before fitting the per-frame scale
The beam is not constant. On one beamline it oscillates +-9.8% with a ~5.9-frame
period, confirmed four ways: the raw images, our own azimuthal-integration total, the
per-frame mean background, and XDS's per-image SCALE, which correlates +0.999 with the
first three. XDS removes it inside INTEGRATE, per image.

The fitted per-frame G could not: --smooth-g defaults to 5 degrees, which is 25 frames
at 0.2 deg/frame, so a 5.9-frame signal is smoothed away. Measured, the applied scale
carried 0.70% rms against a 9.8% modulation and correlated 0.66 with XDS's SCALE. The
only thing removing the oscillation was the refit on fulls, which acts after several
partials spanning most of a period have already been summed, so it removes the mean and
leaves the dispersion inside each event.

Take the flux from the per-frame mean background, gauge it to the run median, and divide
it out of rlp as the partials are ingested, so the fitted G sees only the residual and
smooth-G smooths only the residual. The background mean tracks our own azimuthal
background at r = +0.971 and XDS's SCALE at |r| = 0.93, with 95% of its detrended power
in the 3-8 frame band. Slower background movers - ice, a drifting shadow, absorption
against the goniometer angle, radiation damage - are still absorbed by G, which keeps its
low frequencies through the smoothing.

The applied scale now carries 9.41% rms at |r| = 0.93 against XDS. On the affected
dataset R_meas 9.9 -> 9.5%, low-resolution R_meas 6.5 -> 6.0%, ISa 12.8 -> 13.8, and the
anomalous peak height rises 0.423 +- 0.069 sigma over 18 sites (p < 0.001) - the only
significant move in the arbiter. Over the 38-crystal battery the space groups and the
merged reflection sets are unchanged and every metric has median delta zero.

A monochromatic dataset carries the same modulation at 2.0% rms, confirmed by the same
three proxies; the correction engages there too but no merged statistic moves at that
amplitude.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 20:01:03 +02:00
leonarski_fandClaude Opus 5 f60768d49c Bragg integration: drop the 2% sigma floor and carry the background variance
Two changes to the same variance chain; they are in one commit because the second
exists to remove an assumption the first was breaking, and separating them leaves a
tree that is correct only by luck.

The reported sigma was floored at 2% of the intensity, a per-partial I/sigma cap of
50. It applied only to the box-sum seed, never to the profile fit, so the shipped
default was unaffected - but the combine back-derives each partial's non-signal
variance as sigma^2 - I, and a floored sigma makes that quantity mean nothing. It
then read corr^2 * (0.0004 I^2 - I), which is not a background variance. Measured on
--integrator boxsum: the reported sigma understated the true scatter by up to 16x at
I ~ 21000 counts per partial, and pooled_I amplified a 1 ct/px background drift into
an 11.5% intensity error on the strongest reflections.

What the floor stood in for - that at high intensity the error is systematic rather
than counting - is already carried downstream, twice: the fitted b in
v = a*sigma^2 + (b*I)^2, measured from the data rather than assumed, and
SigmaWithSystematicFloor on the merged sigma. The floor was that idea applied one
level too early with a hardcoded b of 0.02. It arrived without a test or a setter and
was unreachable from the CLI, the API and the config.

The merge now takes the non-signal variance the integrator actually measured instead
of inverting sigma^2 = I + N. That identity is exact for a box sum once the floor is
gone and was never exact for a profile fit, whose sigma^2 = 1/den + (wsum/den)^2 *
bkg_var is formed against a fitted intensity. The value is carried through
BraggFitResult, Reflection and Obs, both engines, both merges, and the process-file
round trip; files written before this change are read with the term absent, which is
what they had.

Battery, 37 crystals, paired: space groups unchanged, reflection sets unchanged,
median delta zero on R_meas and CC1/2. --integrator boxsum on the reference crystal
goes ISa 8.9 -> 20.2 with a 0.947 -> 1.032.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 19:10:57 +02:00
leonarski_fandClaude Opus 5 d565b66916 Post-refine: drop the rocking-width diagnostic, report frames per event
The "median rocking width -> estimated mosaicity" line took an intensity-weighted
second moment of the frame-centre angles with max(0, I) weights, per event, then a
median over events. For a two-frame event that moment is exactly zero whenever only
one frame has I > 0 - probability 2/3 for a reflection carrying no signal - so on a
noise-dominated dataset the median lands in the degenerate spike and prints 0.0000.
Simulated against a known width it is wrong by 0.23x to 13x, in both directions, and
on a pure-noise null it returns a plausible-looking 0.06 deg.

est_mosaicity_deg was read nowhere, so nothing downstream was affected; the number
only misled whoever read the log. It was built to measure a signal for a mosaicity
refinement that was then abandoned, and the estimator that replaced it is the
per-image one that already drives prediction.

Report instead the frames per rocking event, which is what the block could honestly
say: near 2.0 the reflections barely rock, so the observed angle this refinement is
fitted to is under-determined. It is a geometry count, so noise cannot inflate it.

Also drop phi_rms_deg, which is never assigned anywhere, and Partial::zeta, which is
only written.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 19:10:18 +02:00
leonarski_fandClaude Opus 5 4d3434e2a5 Beam stop: compare each pixel only against its own ring
The background belongs to the beam and the shadow to the stop, and the two are
not concentric - fitting the stop edge per azimuth gives offsets of 13.4 px on
an 85.8 px disk, 22.2 px on 67.3 px and 6.9 px on 23.7 px, 8 to 33 per cent of
the stop radius on every crystal measured. The finder bridged that gap with a
radial envelope, the largest ring background over an outward window, used as the
reference for an individual pixel. That quantity exceeds the local background
wherever the background rises outward, so sound pixels near the stop scored below
the penumbra threshold and were masked. Measured against the fitted edge on a
long-distance disk stop, the mask was displaced rather than mis-sized: short by
up to 20 px on one side, over-reaching by up to 45 px on the other, with eight of
twenty-four azimuth sectors falling short.

The ring median is already the right reference wherever a ring still has
unshadowed pixels to measure, which is every ring except those lying wholly
inside the disk - and it needs no assumption about where the stop sits. So the
envelope is gone from the per-pixel test, and the rings it existed to cover are
handled directly: walking outward, a ring whose background is a fraction of the
background further out is shadow in its entirety. That comparison is only ever
asked whether a whole ring is inside the stop, never to judge a pixel, which is
where its failure mode lives. Blockage is deliberately not a counting test - on a
bright dataset the shadow interior is still well counted.

Detection is now one channel instead of two, and 113 lines shorter.

Measured: no azimuth sector falls short by more than 3.4 px, over-reach drops on
all three fitted crystals, and mask area moves by at most 0.04 per cent of the
detector on six crystals, so this corrects the shape rather than resizing.
Battery: space-group agreement with XDS unchanged at 34/37, median change in
R_meas and in the lowest shell 0.000 pp. The crystal that suffered worst when
masking was introduced recovers to its unmasked quality - R_meas 25.1 -> 17.2 per
cent, ISa 4.45 -> 10.04 - which is what removing the over-masking should do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 05:31:57 +02:00
leonarski_fandClaude Opus 5 a29c36600f Beam-stop shadow detection, and a low-resolution limit for scaling
rugnux finds the beam stop and its holder in a projection of 60 images and
marks them in the pixel mask as bit 9 (--detect-beam-stop[=N|off], on by
default). Reflections behind the stop are attenuated but not flagged, so they
integrate low with a plausible sigma and nothing downstream catches them: the
signal-box gate requires 100% valid pixels and shadow pixels are valid, the
background clip is high-side only, and the |zeta| cut applies only to the
space-group search merge.

The detection compares each pixel's background against the typical background
at the same radius on two channels. An azimuthal one (the ring median) finds
the holder arm, which is a minority of its ring; a radial one (the background
just outside) finds the disk, which the ring median cannot see because inside a
fully blocked ring the median is the shadow itself. Pixels are pooled over a
5x5 box and tested only where the background has actually been counted, so
low-background data no longer masks the whole detector. Recorded reflections
are carved back out - a beam stop cannot block a reflection that was measured.

Bit 9 belongs to the run that found it, not to the dataset: it is cleared when
a run starts, so a mask read back from a file that carries one starts clear.
The user mask (bit 8) is left alone.

Scaling and merging gain a low-resolution limit, default 50 A
(--scaling-low-resolution <num>, 0 removes it), applied per observation before
scaling so it also protects the per-frame scale fit and the space-group search.
50 A is the value XDS configurations use; rugnux_vs_xds.py now matches both of
XDS's resolution limits instead of only the high one, so the lowest shell is
the same shell in the two programs.

The viewer draws the detected shadow in coral with a "Show beam stop" switch in
the side panel, exposes the low-resolution limit in the settings dock, and
offers detection in its processing jobs. Adding an image marker meant giving
the reader a MIN_REAL_PXL_VALUE, because several places classify a pixel by
range rather than by equality and would otherwise read the new marker as a very
negative intensity.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 01:05:31 +02:00
leonarski_fandClaude Opus 5 c673521b76 Space-group search: do not veto on systematic-b where the H test confirms
The systematic-b veto compares a candidate merge's fitted b against its
parent's, and both move with data quality. Removing genuinely bad observations
improved both merges but the subgroup more than the supergroup (parent
0.1644 -> 0.1480, candidate 0.3187 -> 0.3056), so the ratio crossed its 2.00
bound at 2.065 and a correct cubic promotion was refused - while the H
statistic, which has no sigma in it, did not move at all (0.898 either way).
Better data demoting a crystal is the wrong behaviour.

The veto now fires only where the H test has not confirmed the promotion. H is
the statistic that was measured to separate a real symmetry operator from a
twin law; b's genuine and twin ranges are interleaved. A twin fails both.

No bound moved and no option was added. Battery: 34/37 point-group agreement
with XDS before and after with no crystal changing; with the beam-stop mask
33/37 -> 34/37, the single change being a cubic crystal recovering its true
I23. Both real merohedral twins stay refused on H in every arm.

Gating the guards on the L-test / second moment was tried and rejected: those
indicators do not flag a real twin on the P1 pre-promotion merge, only after
merging in its true symmetry, so the gate promoted a twin into its holohedry.

Left alone deliberately: merge_systematic_b divides its reduced chi^2 by the
observation count rather than by the degrees of freedom, which inflates the
ratio more for small-orbit parents. Fixing it requires re-deriving all three b
bounds, which were calibrated on the biased statistic.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 01:05:14 +02:00
leonarski_fandClaude Opus 5 df9a9c2a2c Fix the defects found reviewing the branch before merge
Build Packages / build:viewer-tgz:cpu (push) Successful in 19m32s
Build Packages / build:windows:nocuda (push) Successful in 19m57s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m45s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m38s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m24s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m8s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m9s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m18s
Build Packages / XDS test (durin plugin) (push) Successful in 11m10s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 20m21s
Build Packages / build:windows:cuda (push) Successful in 22m5s
Build Packages / build:rpm (rocky9) (push) Successful in 20m57s
Build Packages / Generate python client (push) Successful in 34s
Build Packages / Build documentation (push) Successful in 1m29s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m41s
Build Packages / DIALS test (push) Successful in 21m19s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m34s
Build Packages / build:rpm (rocky8) (push) Successful in 27m4s
Build Packages / XDS test (neggia plugin) (push) Successful in 10m19s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m58s
Build Packages / Unit tests (push) Successful in 1h17m36s
Image buffer: the per-image CBOR metadata headroom had been re-derived from the
online reflection cap alone, which cut it from 4 MiB to 2.55 MB while the measured
worst case - reflections plus the capped spot list plus the three azimuthal arrays -
is 2.9 MB, so the receiver dropped the frames with the most to say. Restore it and
give it a name that both the code and its guard test read: written down twice, the
two had drifted and the test kept passing against the value the code had left.

Spot finding: an unset low_resolution_limit means no limit at that end, as an unset
high_resolution_limit already did. An optional rather than a zero sentinel, because
zero is not a natural "no limit" here - every pixel lies above it, so the plain
comparison masked the whole image instead of none of it, and nothing validated the
zero. The API field is no longer required; a zero is folded into the unset case at
the boundary, where older clients still send it, so one spelling reaches the
analysis code. The FPGA takes its fixed-point ceiling instead, since ap_ufixed<16,9>
wraps above 512 A and would have masked everything.

image_preprocessing: check the CUDA calls on the fused decode path - the one new GPU
file with none, and the path fed by bytes we did not produce. An unchecked
synchronise returned the host-written sentinel as if it were a measurement, so the
decode looked successful and the fallback to the host decoder never fired.

rugnux: --stride no longer writes one past the end of the per-image arrays, whose
count floored where the worker loop ceils, and the written process file links the
images actually processed rather than the first N - each frame's picture now sits
next to its own analysis.

Powder calibration: the face-centred calibrants no longer list their systematically
absent rings, so the distance fit starts from a reflection that exists rather than
an extinct one; the triclinic calibrant covers both signs of h and k instead of a
single octant, which is only valid for a diagonal metric. The test asserted the old
behaviour - one ring formula for every cubic standard - and is rewritten.

CBOR: skip an unknown tagged value in the end block, as the other four blocks
already do. One advance lands on the tagged item rather than past it, so an older
reader fed a newer end message threw and never finalized its file.

Viewer: a settings value the setter rejects no longer escapes as an uncaught throw
from a worker slot, and the field offers only what the setter accepts.

Space-group search: judge stage B on the same "present" cut stage A already computes.
Merged sigma is floored so no reflection reads above ISa, so on a low-ISa merge the
fixed cut left both stage B tests unsatisfiable - every screw axis passed unchallenged
and the centering rescue switched itself off on exactly the weak data it exists for.
Where the fixed cut is the smaller of the two they are equal and this is inert: over
the 37-crystal rotation battery every crystal reports the identical space group and
identical merge statistics, so it is a no-op there and the low-ISa case it targets
remains unmeasured.

rugnux: --polarization reaches --mode azint, which parsed the flag and then dropped
it; that mode also applies the same polarization default as every other mode.

Acknowledge the ACTS/traccc project, whose sparse connected-component labelling both
spot extractors take their algorithm from, with its citation and its license.

The rc.161 change list is brought back to one line per entry, and the user-visible
changes that were missing from it added.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 18:18:46 +02:00
leonarski_fandClaude Opus 5 3ccb97e31b Adaptive spot finder: pin the per-ring host buffers
The GPU engine copies six small per-ring arrays back to the host every frame - the
clipped raw sum/sum2/count that the threshold is computed from, and the plain
corrected sum/sum2/count that become the azimuthal profile. They were plain
std::vectors, so the copies landed in pageable memory, and a device-to-host copy
into pageable memory blocks the calling thread until it has completed whatever
stream it was issued on. The profile snapshot sits between the plain pass and the
two sigma-clip passes, so Detect() stopped there and the device then sat idle while
the host caught up and enqueued the rest.

Register them, as AzIntEngineGPU already does with its own, and the copies are
genuinely asynchronous. Measured on a 4.5 Mpixel frame: 0.647 -> 0.621 ms per
frame. Nothing else changes - the spot list and the profile are unaffected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 15:33:24 +02:00
leonarski_fandClaude Opus 5 5830f78d57 Revert the azimuthal-integration sigma clip
Removes azim_int_settings.sigma_clip / rugnux --azim-sigma-clip and the clipping
machinery in AzIntEngine. This is a partial revert of a6be35ccd - the ice-ring-mask
removal that commit also carried stays. Sigma clipping remains where it started and
where it is needed: inside the adaptive spot finder, at a fixed 3 sigma on raw
counts, feeding the detection threshold and the ice score.

The option made the workflow harder to reason about than the quantity was worth. It
gave azimuthal integration two meanings behind one setting - the bin mean and the
background under the peaks - which the azimuthal-integration workflows do not need.
It also did not compose with the fused GPU engine, which supplies the profile from
its PLAIN pass: on the default rugnux, viewer and receiver path the setting was
silently doing nothing (measured, the profile came out identical to the unclipped
run to 1e-6 with identical per-bin pixel counts). Making it correct is not a matter
of gating that one shortcut - it means separating the workflows (azimuthal
integration, MX rotation, MX stills, geometry calibration) and deciding per workflow
what the profile is for, which is a larger change than the option earns.

The default path is unaffected: over 20 images of a rotation dataset the radial
profile, the per-bin pixel counts and the spot counts are unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 15:31:19 +02:00
leonarski_fandClaude Opus 5 6468dd13be rugnux: --mode, and detector calibration from powder rings
Build Packages / Unit tests (push) Failing after 6m24s
Build Packages / build:rpm (rocky9_nocuda) (push) Failing after 14m5s
Build Packages / build:viewer-tgz:cpu (push) Failing after 14m34s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Failing after 14m54s
Build Packages / build:viewer-tgz:cuda (push) Failing after 16m14s
Build Packages / build:rpm (rocky8_nocuda) (push) Failing after 16m24s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Failing after 18m53s
Build Packages / build:rpm (rocky9_sls9) (push) Failing after 13m2s
Build Packages / build:rpm (rocky8_sls9) (push) Failing after 19m34s
Build Packages / build:rpm (rocky9) (push) Failing after 14m54s
Build Packages / Generate python client (push) Successful in 42s
Build Packages / build:rpm (ubuntu2404) (push) Failing after 14m10s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (durin plugin) (push) Successful in 12m14s
Build Packages / XDS test (neggia plugin) (push) Successful in 11m52s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 12m15s
Build Packages / Build documentation (push) Successful in 2m10s
Build Packages / build:rpm (rocky8) (push) Failing after 18m29s
Build Packages / build:rpm (ubuntu2204) (push) Failing after 17m55s
Build Packages / DIALS test (push) Successful in 17m4s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
--azint-only and --scale are replaced by --mode mx|azint|scale|calibration, with
mx the default. The old flags are removed rather than aliased.

Calibration mode fits the detector geometry - PONI x/y, the two tilts and the
distance - to a calibrant's powder rings and writes a pyFAI .poni alongside a
report of how far each parameter moved from the header. Bragg data constrain the
beam centre worst, because it is gauge-coupled to the crystal orientation; a
powder ring has no orientation to couple to.

--calibrant takes lab6, agbh, ceo2, si or ice. A calibrant is a list of ring
positions rather than a unit cell, because hexagonal ice is P6_3/mmc: rings
enumerated from its cell would include systematically absent ones. So the
crystalline standards generate their rings from a cell and ice carries the
measured list, and RingsFromAzimuthalProfile, GuessGeometry and OptimizeGeometry
all take ring q. The calibrant table is shared with the viewer's powder panel,
which previously carried its own copy.

--calibration picks how the rings are measured: rings (default) sums the
(q x azimuth) profile over every processed image and fits the arcs in it; spots
pools the found spots and fits those. Both use the whole run, with -s/-e/-t
selecting images. rings defaults --azim-phi-bins to 32, since a profile with one
azimuthal bin has averaged the ring over every direction and cannot locate it.

Two fixes this exposed:

The extraction window is capped at half the gap to the neighbouring ring. The
background under a peak is taken from the ends of its window, so a window wider
than half that gap measures the next ring's flank as this ring's background -
and hexagonal ice has three rings within 0.06 1/A. Ice calibration was 3.5 px
out before this and 0.29 px after; LaB6 is unaffected.

RingOptimizer holds rot1/rot2 fixed when only one ring is present. A tilt and a
centre offset both move a ring as cos(phi) and are separated only by the tilt's
amplitude growing as the ring radius squared, so on a single ring they are
exactly degenerate.

Measured. LaB6 at five distances: the fitted direct beam is within 0.36 px of an
independent implementation out to 300 mm, and D = -0.046 + 1.000788 dtz with an
rms of 0.011 mm. At 500 mm one ring is fully on the detector and a second only
clips the corners, which is not enough to constrain a tilt - restricting the q
range to the resolved ring recovers 0.06 px. Ice: 5.53 -> 0.29 px on one crystal
and 4.71 -> 0.80 px on another, against XDS's refined direct beam. On an ice-free
crystal the fit is worse than the header, which is the correct outcome.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 10:00:03 +02:00
leonarski_fandClaude Opus 5 a6be35ccdb Azimuthal integration: optional sigma clipping of the reported profile
The profile is the MEAN of each bin, so a few strong reflections landing in a bin
lift it exactly as a smooth powder ring does. That is the wrong quantity whenever
the profile is wanted as a background rather than as a measurement of what is in
the bin - the ice score being the case in point, where reading a plain profile
INVERTED the metric: over 37 rotation crystals the two highest-scoring crystals
had no ice at all.

The adaptive spot finder already computes the right thing, a sigma-clipped
per-resolution-ring background, as a byproduct of its own threshold. Where it
runs, the ice score uses that. Where it does not - --no-adaptive-spots,
--azint-only, and anything reading the profile the broker wrote - there was no way
to get it. This adds one: azim_int_settings.sigma_clip (rugnux --azim-sigma-clip),
0 = off, minimum 2 because a tighter clip rejects a large part of a clean Gaussian
bin and biases the estimate low rather than removing outliers.

Two clip passes follow the plain one, matching the finder's recipe - the first
pass's standard deviation is itself inflated by the peaks being removed, so one
pass leaves a threshold that is still too generous. A bin with fewer than eight
pixels is left alone: at the detector edge and behind the beam stop there is no
spread to clip on.

Both engines do it. On the GPU the accept range is computed by a small kernel and
stays resident, so a clip pass is one more read of the same pixels and no round
trip; the two accumulation kernels take the range as a pointer that is null on the
plain pass. Measured on a JUNGFRAU rotation dataset, non-adaptive path: azimuthal
integration 0.02 -> 0.06 ms per image, exactly the 3x the extra passes predict,
against a 0.34 ms per-image total.

Note what the result IS: the smooth background under the peaks, not the bin mean.
It should not be switched on where a ring's integrated intensity is wanted - the
powder-ring geometry fit reads ring peaks, and those are what a clip is designed
to remove. Off by default, so nothing changes unless it is asked for.

Not exposed over the REST API - that needs the generated model regenerated, which
is a separate step.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 00:18:52 +02:00
leonarski_fandClaude Opus 5 2c51e00aae Rotation indexing: do not keep a metric symmetry that indexes almost nothing
The Bravais class is decided from the UNREFINED FFT candidate against a fixed
3 degree angular tolerance (LatticeSearch). A lattice that is pseudo-symmetric to
a few tenths of a degree is therefore promoted a class too far, and the constraint
then snaps a real angle to the ideal one - which throws nearly every reflection of
every frame out of tolerance. Measured on a monoclinic crystal that is
pseudo-C-orthorhombic to 0.42 degrees: the promoted cell indexes 2 of 60
validation frames and the run dies, where its own primitive cell indexes 39. It is
the same lattice in a different setting, b_oC = -(a + 2c), volume exactly 2.00x.

The perverse part is that BETTER SPOTS MAKE IT WORSE. LatticeSearch applied to the
true cell returns the promoted class deterministically; runs that succeed escape
only because the raw FFT candidate is inaccurate enough to miss the promotion
window. So it is bistable and non-monotone in every knob - 190 spots per image
gives 44/60, 195 gives 12/60, 200 gives 2/60 - and it will bite harder as spot
finding improves.

The indexer already refines a free triclinic cell alongside each constrained
candidate, but decides between them on the fraction of the accumulated first-pass
cloud that indexes, where the two differ by less than a factor 2 (measured 0.243
vs 0.135, missing both of that guard's bars). The caller has a far sharper
statistic: it already counts how many of 60 validation frames a candidate indexes,
and there the same pair differs by more than 20x. So keep the triclinic cell
instead of dropping it, and let the first pass settle it.

The bar is a clear majority, not a margin, and that is the part that took a
battery to get right: the unconstrained refinement holds NO cell parameter fixed,
so it can only index at least as many frames as the constrained one, and on
genuine symmetry it does index a few more. A 10 % margin - the bar a later scheme
needs to displace an earlier one - demoted a real I-centred orthorhombic crystal
to P1 (47 -> 54 frames) and perturbed an F-cubic one (49 -> 58). Only a
constrained cell that fails outright while its unconstrained cell works is
evidence of a false promotion, so demand exactly that. It is the same "fails to
index half the frames" test the long-axis rescue below already uses.

Battery over 37 rotation crystals: 33/37 space groups matching XDS with one hard
failure becomes 34/37 with none, and the other 36 crystals are identical in every
printed statistic (checked against a repeat run of the previous binary, which
itself differs on one crystal by one observation). The extra validation pass runs
only where the constrained cell already failed - 71 ms in a 15 s run - and not at
all on the 34 crystals whose constrained cell indexes a majority.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 22:11:11 +02:00
leonarski_fandClaude Opus 5 e2de790867 Powder calibration: cover the tilt round trip, and correct how a tilt shows itself
A detector tilt does NOT appear as a cos(2 phi) modulation of the ring radius, as
the previous comment claimed. To first order a misalignment beta gives

    r(phi) = R + (R^2 / F) (beta_x cos phi + beta_y sin phi)

which is a cos(phi) term - the same harmonic a wrong beam centre produces. What
separates them is the radius dependence: the centre's amplitude is the same on
every ring, the tilt's grows as R^2. So they are told apart across rings, not
within one, and on a single ring they are exactly degenerate. Measured on a powder
standard the true cos(2 phi) term is of order R^3 beta^2 / F^2 - hundredths of a
pixel, at the noise floor - so it carries nothing usable.

Also add the tilted round trip, which was missing. It doubles as a check that
RingOptimizer's open-coded rotation agrees with DiffractionGeometry's: the fitter
applies Rx(-rot2) Ry(+rot1) by hand rather than going through the geometry's
Rz(-rot3) Rx(-rot2) Ry(+rot1), and those had never been held against each other.
They agree - 0.020 / -0.015 rad recovered as 0.0197 / -0.0148. Dropping rot3 is
right rather than an omission, since rings cannot constrain in-plane roll.

The tilted case yields fewer ring points than the centred one, which is expected
and worth knowing: the extractor searches a window centred on where each ring is
EXPECTED, so a large enough geometry error carries part of a ring out of it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 21:32:30 +02:00
leonarski_fandClaude Opus 5 5b5bed4f66 Powder calibration: read the rings off an azimuthal profile, not off a spot list
The ring calibration already here (AssignSpotsToRings + RingOptimizer, driven from
the viewer's powder panel) is given a SPOT LIST from a single image. A powder ring
is not a set of spots - it is a smooth arc - so a spot finder samples it wherever
its threshold happens to bite, and one image carries only the counts that image
collected. An azimuthally-binned profile summed over a run measures the same ring
directly, at every azimuth, with the whole run behind it.

RingsFromAzimuthalProfile turns such a profile into the (x, y, q_expected) triples
RingOptimizer already consumes, so nothing downstream changes: for each calibrant
ring and each azimuthal sector it fits the radial peak against a locally
interpolated background, and maps the measured (q, phi) back through the current
geometry to the pixel it came from.

What this is for is the BEAM CENTRE. A powder ring is a conic centred on the beam,
so a wrong centre makes its apparent radius oscillate once per turn and a detector
tilt twice - and neither depends on the calibrant's d-spacings or on the detector
distance. That matters, because the beam centre is otherwise the weakest parameter
we have: fitted from Bragg spots it is gauge-coupled to the crystal orientation,
which is why PostRefine has to restrain it toward the header and commit only a
sub-1 % move, and why XtalOptimizer carries a soft prior noting the beam is "only
LaB6-monitored to ~a few px". A ring does not know about the crystal.

Two things the peak fit is careful about, both of which would otherwise show up as
a spurious cos(phi) - i.e. as a beam-centre shift:

 - the sector's CENTRE is used, not its lower edge. GetBin() floors phi into the
   sector, so a bin stands for [j, j+1), and taking its edge rotates every ring
   point by half a sector.
 - a peak has to stand clear of the scatter of the background either side of it,
   or a sector with no ring in it contributes its largest noise excursion as
   though it were a measurement.

Refuses a single-azimuthal-bin profile outright: that is a plain radial profile,
the ring has been averaged over every direction, and there is nothing left to say
where its centre is.

Tested by round trip against the forward model, as the existing calibration tests
are: synthesise the profile the azimuthal integration would build with the rings
where a shifted geometry puts them but every pixel binned with the unshifted one,
then extract and fit. A 6.0 / -4.0 px beam offset is recovered as 6.13 / -4.03
from 192 ring points. Only the beam centre is exercised here; the tilt path is
covered by the existing DetGeomCalibTest round trips.

This is the extraction only - nothing calls it yet, and the run-scoped accumulator
it is meant to read (JFJochReceiverPlots::az_int_profile, already summed over a run
and written to /entry/azint/dataset) is still integrated with one azimuthal bin by
default.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 21:30:24 +02:00
leonarski_fandClaude Opus 5 b71e8c6a56 Bragg integration: use the project's PI, not M_PI, in the radial kernel
M_PI is not standard C++. MSVC defines it only when _USE_MATH_DEFINES is set
before <cmath>, so the radial background kernel's azimuth loop does not compile
there:

  error C2065: 'M_PI': undeclared identifier
  error C2737: 'phi': const object must be initialized   (cascade from the first)

GCC and Clang define it anyway, which is why the Linux build stayed green.
image_analysis is viewer-reachable, so it has to build under MSVC.

common/JFJochMath.h already carries a constexpr PI for exactly this reason - its
comment names this case - so use that. Same value to the last digit, so the
integration results are unchanged; the CPU/GPU parity test passes unaltered
(9002 assertions).

This was the only M_PI left in the viewer-reachable tree. The remaining uses are
in tests/, which Windows does not build (JFJOCH_VIEWER_ONLY is forced there).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 20:52:33 +02:00
leonarski_fandClaude Opus 5 b5f5879a1d rugnux: measure the ice in the first pass, and always find its own spots
Ice handling was gated on a measurement the run only made AFTER the images had
been processed, so the per-image pass could not use it. The flagging therefore
ran unconditionally: ice-band spots were ordered last in the --max-spots budget
and held out of the indexer seed and the geometry refinement on every crystal,
iced or not. The eleven bands are fixed geometry holding 16-26 % of the unique
reflections whether or not there is ice, so on a clean crystal that discards a
fifth of the spots - the strongest first - for nothing. Measured on a crystal
whose gate never fires, that moved the merged data by a mean of 0.85 sigma
against a run-to-run floor of 9.3e-5.

Measure it in the first pass instead. That pass already looks at ~100 images
spread over the sweep, and it already stops at the spot finder, so it sees the
azimuthal profile for the smooth channel and the unfiltered connected components
for the spot channel. Both counts SpotAnalyze takes are pre-filter, so pooling
them there is the run's own verdict, reached before anything has been discarded
and in time for the pass that acts on it. Where the sample sees no ice, the run
indexes on the ice-band spots too.

It has to be the whole sample: the spot channel is a ratio pooled over images,
because one frame carries a handful of control spots. A per-image gate is not an
alternative - two of the crystals whose indexing this rescues fire on that
channel alone, at profile scores of 1.12 and 1.22, so gating per image on the
profile score would drop exactly the cases that matter.

This also removes the first-pass spot reuse, and with it --redo-rotation-spots
and the reuse path. Finding the ~100 first-pass spots costs little, and reusing
was actively wrong here: the stored spots were found online at the acquisition's
threshold and have already had their ice-band entries ordered last and dropped
by its spot budget, so counting ice from them under-reads it by construction,
and the lattice search never saw the spot-finding settings at all. It also
removes the need for the machinery that re-found spots whenever a spot-finding
option was named, which made those options impossible to A/B.

IndexAndRefine cached index_ice_rings at construction, which happens before the
first pass; it holds a reference to the experiment, so it now reads the setting
where it uses it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 20:48:22 +02:00
leonarski_fandClaude Opus 5 f0cdb027e1 Ice: default the merge mask off, gate the radial background on smooth ice, and pick detection by geometry
Three defaults, each settled by measurement rather than by argument. The
arbiter throughout is structure-referenced - anomalous peak height where a
crystal can carry it, and otherwise the agreement of the ice bands with a fixed
external model against resolution-matched DECOY bands carrying no ice. The
band-versus-decoy contrast is used because R-free here tracks completeness, and
every one of these switches moves completeness.

The damage is real and it localizes: over the rotation battery the ice bands'
excess amplitude reaches +9.6% on a smooth-ice crystal and +35% on the worst,
while a clean control sits at +0.6% (z +0.45). On the worst crystal, nine of the
ten largest excess peaks in a q scan land on hexagonal ring positions. Turning
ice handling off leaves the contrast unchanged and forcing it on a clean crystal
does not create one, so it is the ice and not the machinery.

MERGE-TIME RING MASK -> OFF. It deletes reflections, which no other program does
by default - AIMLESS, DIALS, xia2, XDS and CrystFEL all keep ice-band
reflections in the merge and exclude them only from the model fit; autoPROC is
the sole exception. On the one battery crystal where the mask fires and an
anomalous arbiter can score it, dropping the band moved the mean peak height at
the known sites by -0.001 +- 0.018 sigma, 2% of the site height, while removing
1149 unique reflections whose mean I/sigma was 3.62 against the dataset's own
3.05 - better than average data - and costing 17 completeness points in that
shell. It fires on 5 of 37 crystals, changes no space group, and those 5
disagree in sign: it clearly helps the two most heavily iced, is a wash on two
and costs a third. So it stays as a switch, worth setting by hand on a badly
iced crystal where it shows in the high shell, but it is not a default.

RADIAL BACKGROUND -> AUTO, gated per image. The correction models the background
as a function of radius alone, and that is exactly when it works. On a crystal
with pure smooth powder ice it removes 43% of the bands' excess amplitude, with
the improvement 7x larger inside the bands than outside; on a crystal whose ice
is discrete crystallite spots - no smooth ring to model - the excess amplitude
GREW by half; on clean data it is inert to four decimals. The two ice channels
already separate those morphologies, so --background-radial takes on|off|auto
and auto applies it to an image when that image's peak-excluded score reaches
--ice-min-score. Auto never engages without such a score, because the plain
profile carries the Bragg peaks and cannot support an absolute threshold.

Per image rather than per run, and that was tested rather than assumed: the
gate fires on 100% and 94% of frames on the two crystals that want it, and on
1.5% of frames - 32 blocks, 23 of them single frames - on the textured-ice
crystal. A seam statistic against off + f*(on - off) is null on both mixed runs,
every merge statistic is bracketed by the pure arms, and the textured crystal's
auto arm lands on `off` rather than on `on`'s harm. A run-level gate would need
the score before the pass that integrates, i.e. rotation-only plumbing, and buys
nothing measurable.

The kernel was already built unconditionally, so flipping the flag per image is
free - except on the GPU, where the launches were gated on a construction-time
n_rad. That is why the buffers are now allocated whenever the correction could
run, and Run() decides per image.

DETECTION -> the geometry's default when the file is silent: on for rotation,
off for stills, with the command line and then the file taking precedence. A
rotation sweep sits on the same rings for the whole run, so ice there is a
coherent systematic and the presence gate keeps it inert on a clean crystal; a
serial stills run has too few spots per image to spend any on flagging. The
master file's key is kept as written rather than collapsed to a bool, so "the
file said nothing" is distinguishable from "the file said no" - it used to fall
silently to off, taking the exclusion from the scale fit with it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 19:06:30 +02:00
leonarski_fandClaude Opus 5 06b8c8ed66 Merge statistics: count the observations the merge kept, not the ones it walked
Whenever the merge-time ice-ring mask dropped a band, the per-shell observation
count and hence the reported multiplicity were wrong. On one crystal the lowest
resolution shell read 40780 observations over 1932 unique reflections - 21.1x -
where the truth is 27007 and 13.98x, and the overall redundancy read 12.52
against 12.29. Only counts were affected: intensities, sigmas, R_meas, CC1/2,
completeness and ISa were right throughout, because a masked group carries
merged_I = NaN and never enters those sums.

It looked like double counting and was not - it is a MOVE. Two independent
faults, both in three lines:

total_obs rides on the R_meas re-walk, whose filter deliberately ignores the
ring mask (and, on a search pass, the ice flag) so that R_meas is computed on
the same reflections either way. RmeasUsable therefore differs from MergeUsable
by exactly those two tests, and the observations they admit were being counted
against a `unique` that excludes them.

On the GPU path that count is binned by the GROUP's resolution, and a group
every one of whose observations is masked never has one written - acc[g].d stays
NaN. ResolutionShells::GetShell(NaN) then returned shell 0 rather than nothing:
NaN fails both bound comparisons, falls through to the arithmetic, and
static_cast<int32_t>(NaN) is INT_MIN, which the clamp maps to 0. So the masked
ring's observations were re-labelled into the lowest-resolution shell, four
shells from the ring they came from.

The two paths disagreeing on the same run is what settled it: with the mask on,
the GPU statistics gave shell 0 = 752 and the CPU statistics 423, while the
merged intensities were identical.

Count the merged population instead - acc[g].nh, which the merge already
accumulates per group - and guard the CPU increment with usable_merge. The
rnusable skip stays: any group present in the merged output has at least one
observation passing MergeUsable, and MergeUsable is a subset of RmeasUsable, so
it cannot drop a group that contributes to `unique`.

With the mask off and for_search false the two predicates are identical, so this
is provably inert on every shipped configuration - demonstrated on four
configurations, including one where ice handling is active but the mask does not
fire: the statistics blocks are unchanged. (The reflection lists differ in the
last ulp on 3-12% of lines, but so do two runs of the same binary; that is the
known rotation nondeterminism, and the statistics block is what is stable.)

The NaN guard also removes a silent contamination nobody was looking for. Four
call sites validate a resolution with `d <= 0`, which NaN passes: the Wilson-B
fit and per-shell <I/sigma> (CalcISigma), the per-image resolution plot
(SpotUtils) and the shell Wilson prior (FrenchWilson) were all binning
non-finite d into their lowest-resolution shell. French-Wilson now falls back to
the global mean rather than to that shell's, which is the worst prior available.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 19:06:00 +02:00
leonarski_fandClaude Opus 5 17eb6ef091 Post-refine: report the goniometer rotation scale it already fits
Build Packages / build:viewer-tgz:cpu (push) Successful in 20m2s
Build Packages / build:viewer-tgz:cuda (push) Successful in 20m10s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m20s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m33s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m3s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m7s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m50s
Build Packages / XDS test (durin plugin) (push) Successful in 11m12s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 21m35s
Build Packages / build:rpm (rocky9) (push) Successful in 21m32s
Build Packages / Generate python client (push) Successful in 39s
Build Packages / Build documentation (push) Successful in 1m29s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (rocky8) (push) Successful in 26m3s
Build Packages / DIALS test (push) Successful in 20m19s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m18s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m11s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m17s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m2s
Build Packages / Unit tests (push) Successful in 1h18m44s
Build Packages / build:windows:nocuda (push) Failing after 12m15s
Build Packages / build:windows:cuda (push) Failing after 11m57s
A stage that turns further than commanded is invisible in the file, because the stored
omega values ARE the commanded ones - both XDS and rugnux then read the discrepancy as
the crystal drifting. Measured on one dataset in 37, a ~1.2 % over-rotation costs it
unique 9.9k -> 29k and CC1/2 68 -> 98 % when corrected by hand.

No new degree of freedom is added, because the one needed is already there and being
thrown away: step A's residual rotates by -angle_rad * axis[] with axis an UNNORMALISED
3-vector, so the length it fits IS the factor by which the stage actually turned.
GoniometerAxis::Axis then normalises it away (with the `increment *= len` line sitting
commented out). This only reports it.

Guarded by the same cross-validation that gates the cell move - a fold that merely
soaked up noise cannot raise the flag - and by a 0.5 % tolerance, which is where a
direct scan of this factor puts 36 of 37 datasets (all at exactly 1.0000). The known
fault reads 1.00604 and warns; clean controls read 0.99958 and 0.99954.

It UNDER-reads the true magnitude: the fit only sees reflections already indexed at the
nominal angle, per-frame orientation refinement has absorbed part of the error, and the
axis components are bounded. Treat it as a detector, not a calibration - nothing here
corrects the data.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 16:17:23 +02:00
leonarski_fandClaude Opus 5 61a7c91b90 Ice: detect it on two channels, and only handle it when it is there
The per-image ice score was read off the PLAIN azimuthal profile. That profile is a
per-ring mean, so a few strong Bragg reflections landing in a ring's q bin lift it
exactly as ice would. Measured over 37 rotation crystals, that did not merely add
noise - it INVERTED the metric: the two highest-scoring crystals had no ice at all
(4.23 and 4.06), while a clean control read 1.57. A decoy null - the identical
statistic evaluated at q positions where hexagonal ice cannot be - reaches 1.51 at its
99th percentile and 2.70 at its maximum, so that metric cannot support any absolute
threshold whatsoever.

The adaptive spot finder already computes the right input for its own threshold: a
sigma-clipped per-resolution-ring background, in the same bins. A powder ring is
azimuthally smooth and survives the clip; Bragg peaks do not. On the clipped profile
the clean population tightens to 1.00-1.22 and the crystals with confirmed ice sit at
2.08-2.37, against a decoy null that never exceeds 1.29.

That channel is blind to one thing: ice in large crystallites diffracts as DISCRETE
spots and leaves the radial profile flat. So a second channel counts found spots on the
rings against the same q width of ice-free flanks beside them. The two barely overlap -
the smooth-ice crystals read 2.1-2.4 / ~1.0 and the textured ones ~1.1 / 3.8-17.6,
while a clean crystal reads 1.04 on both.

Both are then used as a GATE (--ice-min-score 1.5, --ice-min-spot-ratio 2.0, both
calibrated on the battery, 0 disables): the eleven fixed hexagonal bands cover 16-26 %
of the unique reflections at typical resolutions whether or not the crystal has ice, so
flagging, the exclusion from the scale fit and the merge-time CC1/2 ring mask are now
all skipped when neither channel sees any. The gate is applied in the full pipeline and
in --scale, which reads the stored per-image values back out of the _process.h5.

Also fixes the merge-time mask's control: the shoulder now excludes reflections that
are themselves on an ice ring. The rings are not evenly spaced - 1.947/1.916/1.882 A
sit 0.05-0.06 apart in q - so for those three the [w,3w) shoulder landed squarely on
the neighbours and the test compared ice against ice. Measured, that is the only thing
this changes: it removes firings on those three rings and leaves every other firing's
CC pair identical to three decimals.

And the online ice half-width, which was 0.02 in the API against 0.03 offline, so the
same data got a narrower band online than the measured ~0.06 ring FWHM justifies.

Battery (37 rotation crystals, against the previous behaviour): space groups 34/37 in
both and NO crystal's space group changes; 6 crystals gain unique reflections, 1 loses.
Best of them gains 7082 unique reflections with R_meas 16.0 -> 14.3, CC1/2 95.9 -> 97.3
and ISa 13.7 -> 19.0; another goes R_meas 54.9 -> 42.9, CC1/2 84.0 -> 90.4, ISa
3.9 -> 5.5; a third reaches CC1/2 99.4 from 95.7 at an unchanged reflection count. The
one crystal that loses reflections improves on both R_meas and CC1/2.

Not done here: the ScanResult/API/plot-type/frontend/viewer layers for the new
spot_count_ice_control (they need the OpenAPI regeneration). Message, CBOR, HDF5
write/read and the receiver plots are.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 16:17:23 +02:00
leonarski_fandClaude Opus 5 0e23fd3ab9 Bragg integration: propagate the background-estimate uncertainty, add an opt-in radial background correction
Two independent pieces in the same code path.

The background-estimate variance was never propagated. A reflection's background comes
from a finite ring of n_b pixels, so subtracting it adds var(B)/n_b per signal pixel -
sqrt(1 + n_d/n_b) = 1.109 with the shipped stencil. Both engines omitted it, which is
exactly the 1.11-1.19 gap measured between the off-ring scatter and the reported sigma.
Three lines each; it affects every dataset, not only iced ones.

The radial correction is new and OFF by default (--background-radial). The signal disk
and the background ring are concentric, so for any background LINEAR in position
<B>_ann == <B>_disk identically and a plane fit buys nothing; the leading error is the
CURVATURE of the radial background, which on a sharp ice ring reaches +26 counts on a
single reflection. Since every reflection uses the same stencil, that error is a fixed
kernel over radial offset - one short dot product per reflection and no extra pixel
reads. Validated on empty apertures before any C++: mean |bias| over 9 bands / 3
crystals 4.33 -> 0.79 counts with the scatter unchanged.

Three things it cost a battery each to learn, all now in the code:
 - the radial curve must be accumulated from CLIPPED annulus pixels, inside the clip
   pass, or it carries neighbour tails and zingers (so it is inert under --integrator
   boxsum, which has no clip pass);
 - the GPU version was a 1.8x slowdown from atomicAdd contention on a small radial
   array - staged in shared memory per block it now costs nothing measurable;
 - it is battery-NEUTRAL as a default, because the reflections whose bias it fixes are
   the ones the ice handling already excludes. Hence off by default.

CPU/GPU parity extended with two radial sections: 9002 assertions.

Also fixes a latent French-Wilson quadrature collapse: j_max = I + 8 sigma on a fixed
400-point grid degenerates to a single cell once sigma >> 50 <I>, giving F = 0.1 sqrt(sigma)
with sigmaF -> 0. Harmless today, but any sigma-inflation scheme detonates it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 15:44:13 +02:00
leonarski_fandClaude Opus 5 52f0e58cae rugnux: do not smooth un-indexed frames into the per-frame geometry
SmoothGeometry de-rotates each frame's lattice to a common reference, averages
in frame order and rotates back. A frame that never indexed keeps a
default-constructed CrystalLattice whose vectors are all ZERO - and zero is
finite, so the isfinite guard let it through. Those zero vectors were averaged
into their neighbours' smoothed orientation, pulling it toward the origin, and
they were scored in the leave-one-out cross-validation that picks the smoothing
window.

On a crystal where 374 of 900 frames fail to index, the effect on the window
choice is not subtle. Measured:

  before   n_scored 900 (only 526 indexed)   CV score ~504-542 A^2   window +-12
  after    n_scored 516-526                  CV score  0.160-0.175   window +-2

The score was inflated ~3000x and the choice among windows was noise. It settled
on +-12 frames - 9.6 degrees of goniometer rotation - on a crystal whose
orientation genuinely drifts by ~8 degrees over the sweep, so every partial's
delta_phi was recomputed from an orientation averaged across that drift.

Require a real cell. Exactly inert when every frame indexes, and no threshold is
touched.

The crystal that exposed it goes P1 -> P2_1, observations 60107 -> 77021,
completeness 64.1% -> 93.0%, multiplicity 1.10 -> 2.0, CC1/2 70.0% -> 84.9%,
R_meas low shell 37.3% -> 22.1%, and its 2-fold operator CC 0.330 -> 0.669,
comfortably clear of the 0.5 gate. Battery over 37 crystals: space groups
33/37 -> 34/37, and that crystal is the ONLY flip - no losses. Another crystal
is rescued from near-total collapse (4402 -> 139213 observations) because the
two-pass "going back to the header geometry" fallback stops firing. Anomalous
peak height +0.043 +- 0.022 sigma over 7 crystals, so the background clip's gain
is intact. Merged quality is otherwise neutral (CC1/2 6 better/6 worse,
R_meas_lo 9/6) with observations up on 18 crystals.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 01:21:21 +02:00
leonarski_fandClaude Opus 5 b22e1b6822 rugnux: raise the ice-ring mask margin to the measured null
The mask drops a hexagonal-ice ring when its merged half-set CC1/2 falls a fixed
0.05 below its resolution shoulders. That margin is not a significance level: at
the populations these rings actually have, 0.05 spans 1.1 to 7.3 sigma across
firings, and a nominal Fisher-z error understates the real scatter of these
heavy-tailed intensities by ~2.7x, so the null has to be measured rather than
derived.

Measured it with decoy bands - the identical ring/shoulder statistic evaluated
at q positions carrying no ice ring - over the 37-crystal rotation battery: the
gap's empirical null is p95 +0.032, p99 +0.095. So 0.05 sits near the 96th
percentile, about 4% of ice-free bands clear it, and roughly half the 22
observed firings are indistinguishable from bands with no ice in them. The
firing gaps are continuous, not bimodal, with 12 of 22 in [0.05, 0.10).

Raise it to 0.10, the 99th percentile of that null. Firings 22 -> 10, crystals
12 -> 5, decoy false-positive rate 3.4% -> 0.8%. An independent check against
XDS - which integrates through ice rings and so measures exactly what we delete
- agrees: of the firings with a usable comparison, 9 true / 9 false becomes
7 true / 2 false.

Battery: space groups 34 OK / 3 DIFF, the same three crystals as baseline, and
no other discrete decision changes on 37/37. The heavily iced crystal keeps all
five of its rings and its CC1/2 of 96.6; eight others recover 3.9-11.9% more
unique reflections and up to 10.4 completeness points. Cost is CC1/2 -0.84 on
one crystal, -0.35 on another, and agreement with XDS on the common reflections
worse by a median 0.0004.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 19:00:10 +02:00
leonarski_fandClaude Opus 5 09fb8e0306 Bragg integration: clip the background ring high side instead of trimming it
The r2..r3 background ring was averaged with a 10% SYMMETRIC trimmed mean. A
symmetric trim is not a consistent estimator of the mean of a right-skewed
(Poisson) sample: on a clean Poisson ring it sits ~0.1 ct/px BELOW the true
mean at every level, and with ~50 signal pixels in the r1 disk that
under-subtraction adds ~5 counts to every partial on every frame. Measured two
independent ways on four rotation datasets - stored background_mean against a
plain ring mean over the same pixels on reflection-free frames, and directly on
apertures that provably hold no reflection. Empty-aperture pedestal, counts:
plain mean -0.03..-0.20, 10% symmetric trim +5.05..+6.34, 4 sigma clip
+0.02..+0.54.

Replace it with a high-side-only sigma clip at mean + n*sqrt(mean), n = 4 for
monochromatic data. It rejects the same one-sided contamination the trim was
there for - better, in fact: a 40 px neighbour core at +100 ct shifts the trim
by +10.1 ct/px, because a symmetric trim collapses once contamination exceeds
~10% of the ring, versus +0.009 ct/px at 4 sigma. False rejection on a clean
ring is 0.04-0.39%. Broadband data keep their tuned 3 sigma clip unchanged. The
trim stays reachable with --background-trim for back compatibility; setting
either estimator clears the other, so they can never stack. --integrator boxsum
does not take the clip (matching what the shipped clip already did), so it now
uses the plain ring mean unless --background-trim is given.

The intensities get measurably more accurate: per-shell agreement with an
independent processing of the same images improves on 14 of 16 crystals
(weighted -0.0347, outermost shell 12/4), the outermost-shell R_meas NUMERATOR
- absolute scatter, not a denominator effect - falls 13.5% median on 16/5, and
CC1/2 in the outer shell improves on 14/7.

EXPECT <I/sigma> TO FALL AND EDGE R_meas TO RISE. Both are inflated by
information-free counts, so both get worse when the bias is removed; neither is
evidence against this change. That fingerprint is exactly how the trimmed mean
was accepted in the first place.

Known cost: over the 37-crystal rotation battery the de-novo space-group count
goes 34 OK / 3 DIFF to 33 / 4. The single regression is a two-lattice crystal
whose merge fails the absolute-sanity gate under either background (R_meas
63.5%, CC1/2 72.2%) and which carries an unresolved indexing ambiguity on the
very operator being scored, so its operator CC is diluted by construction. No
other crystal changes space group, and twin protection is not weakened - the
H-ratio veto that refuses genuinely twinned crystals gets MORE decisive
(1.63 -> 1.84, 2.83 -> 3.99).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 18:03:11 +02:00
leonarski_fandClaude Opus 5 8ecd126e93 rugnux: stop the space-group search starving on a low-ISa merge
The correlation stage kept only reflections with I/sigma >= present_i_over_sigma
(3.0). That statistic is taken on the P1-MERGED intensities, whose sigma is
floored at b|I| (Merge.h, SigmaWithSystematicFloor) so that ISa = 1/b is the
asymptotic I/sigma ceiling - no reflection in a merge can read above it.
Verified over the rotation battery: max I/sigma equals 1/b on every merge.

So a fixed cut is not a per-reflection test at all. Every reflection sitting at
the floor reads 1/b exactly, however strong, and on a merge whose ISa falls
below the cut NOTHING passes: every operator is left with no pairs, its CC is
NaN, and the point group collapses to 1. The predicate "search-merge ISa < 3"
identifies the affected crystals exactly.

It is latent today - no crystal in the battery starves on the shipped
integration background - but it fires on four as soon as an additive intensity
bias is removed, and it is not a data-quality verdict: the crystals it silences
have final merges at ISa 19-22 while their low-multiplicity search merge sits
at 3.5-4.0, just above the cut.

Cap the cut at the merge's own I/sigma quantile so the correlation stage always
keeps at least its strongest quarter. A no-op wherever the fixed cut already
keeps that many - the cut stays exactly 3.000 on healthy merges. Battery
unchanged at 34 space groups matching XDS / 3 differing, with merged
observations identical to 0.000% on all 37 crystals.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 16:48:40 +02:00
leonarski_fandClaude Opus 5 fb55645b81 Revert "rugnux: fit the profile radius from the strongest spots too"
Build Packages / build:viewer-tgz:cpu (push) Successful in 16m51s
Build Packages / build:viewer-tgz:cuda (push) Successful in 18m44s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 20m42s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 21m37s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 24m37s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 24m48s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 20m13s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 25m19s
Build Packages / build:rpm (rocky9) (push) Successful in 23m23s
Build Packages / DIALS test (push) Successful in 21m35s
Build Packages / Generate python client (push) Successful in 40s
Build Packages / build:rpm (rocky8) (push) Successful in 29m16s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2404) (push) Successful in 23m17s
Build Packages / XDS test (durin plugin) (push) Successful in 11m5s
Build Packages / Build documentation (push) Successful in 1m14s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 27m29s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m21s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m24s
Build Packages / build:windows:nocuda (push) Successful in 13m58s
Build Packages / build:windows:cuda (push) Successful in 16m6s
Build Packages / Unit tests (push) Successful in 1h18m59s
Reverts the profile-radius part of 457b1bfd1; the comparison-script and
mosaicity-column changes from that commit are kept.

The cap was validated on the rotation battery, which cannot test it: the profile
radius feeds `ewald_dist_cutoff` in IndexAndRefine, and that is read only by the
STILLS predictors (BraggPrediction/BraggPredictionGPU). The rotation predictors
gate on the mosaicity window instead and never look at it. So "no space-group
changes, 36 of 37 crystals bit-identical" showed the quantity is inert for
rotation, not that capping it is safe - and the one regime where it does act was
never exercised.

Validating it needs the serial-stills battery, which is a much larger exercise.
Until then the arbitrary constant is not worth carrying in a code path nobody
measured.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 21:35:17 +02:00
leonarski_fandClaude Opus 5 017f64690c rugnux: smooth the per-frame geometry before scaling
Geometry is re-refined independently on every frame, against that frame's spots
alone - as few as a dozen on a sparse crystal, where XDS fits its equivalent to
about sixty times more data. Measured over ten datasets the per-frame orientation
carries two components: a slow drift that is real, with rugnux and XDS agreeing to
R^2 0.83-0.88 on the two crystals that genuinely slip by 1.5 and 0.54 degrees, and
a fast jitter that is fit noise, scaling with spots-per-frame at exponent -0.79
where counting noise alone would give -0.5. The jitter is worth 1-8% on merged
intensities, 24% on the sparsest crystal.

It cannot be fixed by refining less. Turning per-image refinement off entirely
loses six space groups and a whole crystal, and even a 624-spot-per-frame crystal
collapses; dropping the beam-centre terms holds the space groups but is worse on
31 of 37 crystals. The freedom is earning its keep, so keep it and suppress only
the band that cannot be physical - a crystal does not re-orient and snap back from
one frame to the next.

So smooth the orientation in frame order after integration and recompute each
partial's delta_phi, and hence its partiality, from the smoothed lattice. Batching
at integration time was not an option: frames are processed independently and the
online path depends on that. This runs before the GPU upload, so the device path
picks it up with no separate kernel.

The window is chosen per dataset by leave-one-out cross-validation, because the
two components vary far too much for one number - drift spans 0.018 to 1.288
degrees and jitter 0.005 to 0.221, so any fixed window over-smooths one crystal
while under-smoothing another. Chosen windows range from +-1 to +-20 frames. It is
capped: cross-validation scores how well neighbours predict a frame's orientation,
which on a barely-drifting crystal keeps improving with width, but the per-frame
fit is also absorbing a real per-frame systematic and smoothing too wide destroys
it - uncapped, one crystal chose +-60 and lost 16% of its ISa.

Battery over 37 crystals: space groups unchanged at 34 matching XDS, R_meas better
on 31 and worse on 6, low-resolution R_meas 30/7, ISa 26/10, high-resolution CC1/2
23/12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 20:06:41 +02:00
leonarski_fandClaude Opus 5 457b1bfd1d rugnux: fit the profile radius from the strongest spots too
Build Packages / build:viewer-tgz:cpu (push) Successful in 18m20s
Build Packages / build:viewer-tgz:cuda (push) Successful in 20m23s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m35s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m47s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m24s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m35s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 29m9s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 19m23s
Build Packages / XDS test (durin plugin) (push) Successful in 10m17s
Build Packages / build:rpm (rocky9) (push) Successful in 20m45s
Build Packages / Generate python client (push) Successful in 33s
Build Packages / Build documentation (push) Successful in 1m7s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (rocky8) (push) Successful in 26m5s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 11m15s
Build Packages / DIALS test (push) Successful in 20m23s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m50s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m36s
Build Packages / XDS test (neggia plugin) (push) Successful in 10m3s
Build Packages / Unit tests (push) Successful in 1h17m43s
Build Packages / build:windows:nocuda (push) Successful in 16m24s
Build Packages / build:windows:cuda (push) Successful in 17m50s
Same defect as the mosaicity in 2c94f3013, in the same file's sibling fit. The
profile radius is an RMS of the excitation error over whatever spots were kept,
and weaker spots sit further off the Ewald sphere, so it grows with the depth of
the list: measured over a 150 -> unlimited spot budget it rises ~60%, and on a
clean dataset as much as on a hard one, so this is general rather than something
one awkward crystal provoked. That made it a function of --max-spots, which is
an indexing budget, rather than of the crystal.

Its one consumer treats it as a membership gate (ewald_dist_cutoff is twice it)
where reflections at the cutoff carry near-zero partiality and are removed
downstream anyway, so the integrated data barely notices: with the mosaicity
already pinned, the partial count moves 0.2% across a 34% change in the radius.
It is also reported per image as a diagnostic, though, and a number that slides
with an unrelated setting is misleading to anyone comparing two runs - and it
would stop being benign the moment anything used it as a width rather than a
gate.

Battery over 37 crystals: no space group changes, 36 of 37 bit-identical, no
failures, one crystal marginally better.

Also in the comparison script: report XDS's mosaicity next to rugnux's. XDS has
two and they are not interchangeable - CORRECT.LP's REFLECTING_RANGE_E.S.D. is
post-refined, while INTEGRATE.LP's per-batch SIGMAR is its integration-stage
estimate, and the two differ by up to 2.3x. The XDS cell now prints both as
postrefined|MLE so a per-image estimate is compared against the one measured the
same way. Fixes a scoping bug in the same addition where every crystal read the
last directory's INTEGRATE.LP.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 10:52:30 +02:00
leonarski_fandClaude Opus 5 2c94f3013e rugnux: fit the mosaicity from the strongest spots only
Build Packages / build:viewer-tgz:cpu (push) Successful in 21m50s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m21s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m34s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 24m7s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m34s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m30s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m37s
Build Packages / XDS test (durin plugin) (push) Successful in 11m30s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 22m6s
Build Packages / build:rpm (rocky9) (push) Successful in 21m35s
Build Packages / Generate python client (push) Successful in 43s
Build Packages / Build documentation (push) Successful in 1m17s
Build Packages / Create release (push) Skipped
Build Packages / DIALS test (push) Successful in 20m20s
Build Packages / build:rpm (rocky8) (push) Successful in 27m13s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m40s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m19s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m37s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m5s
Build Packages / Unit tests (push) Successful in 1h19m31s
Build Packages / build:windows:nocuda (push) Successful in 19m12s
Build Packages / build:windows:cuda (push) Successful in 22m29s
The per-image mosaicity MLE ran over the whole indexed spot list, so it rode
on --max-spots, which is an indexing budget. A spot is detected when
I_full * R(tau) clears the finder threshold, so selecting by intensity censors
on R(tau): a deeper list holds proportionally more large-|tau| partially
recorded spots and the fit widens with it. Raising the budget 250 -> 1000
widened sigma_M 0.059 -> 0.075 deg on a rotation dataset whose measured rocking
width says 0.054.

That is not cosmetic. An over-wide mosaicity mis-states every partiality in
scaling: forcing the mosaicity across that range moved the merge error model
from b 0.039 / ISa 26 to b 0.167 / ISa 6, and the space-group search lost a
genuine 422 with it, merging the crystal in 222 instead.

Cap the fit at the strongest 250 spots. FilterSpotsByCount leaves the list
strongest-first, so this selects exactly the spots a smaller --max-spots would,
and the mosaicity becomes invariant: 0.0538 deg at 250, 500, 1000 and 2000
spots, with the correct space group at each. Trimming or down-weighting the
tau tail does not work - the censoring is multiplicative in R(tau), so it
widens the whole distribution rather than adding a tail.

Battery over 37 crystals: exactly one change, the demoted crystal repaired
(33 space groups matching XDS -> 34). 23 of 37 are bit-identical, never
reaching 250 spots. Unaffected elsewhere: the default spot count is 250, and
stills have no goniometer so they return before the fit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 07:34:40 +02:00
leonarski_fandClaude Opus 5 d79b20e268 indexing: key the shared device tables on their content, not only on an address
The cache returned a device copy for a (device, host address) pair and cast it to
whatever the caller asked for, with nothing checking that the bytes behind that address
were still the same bytes. A host buffer can be mutated in place - PixelMask::LoadMask
does exactly that - or freed and reallocated at the same address, and either hands the
caller a device copy of something else. Nothing would report it: the tables are read-only
geometry, so the engine would simply mask the wrong pixels for the rest of the run while
the azimuthal mapping, the written pixel_mask dataset and the viewer overlay used the new
one. Today that is unreachable, but only because of two guards in unrelated files that
neither state nor assert the requirement.

The byte length and an FNV-1a checksum of the bytes being uploaded are now part of the
key. Both are computed once per engine construction, over a buffer that is about to be
copied to the device anyway, so the cost does not show. Expired entries are pruned on
insert, since distinct content now means distinct entries.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 16:33:24 +02:00
leonarski_fandClaude Opus 5 5727cb68a4 rotation_indexer: write down why the supercell bar is unreachable, and what failed to fix it
`frac > RATIO * best_frac` cannot be satisfied once best_frac passes 1/RATIO - above
0.667 for a ratio of 1.5, which is ordinary for good rotation data. Above that the two
guards do not raise the bar, they close the branch: no axis multiple and no
lower-symmetry setting can displace the incumbent however much better it fits, so a
genuine superstructure is kept as its sub-cell and its satellite rows go unindexed,
silently.

The obvious repair - restate the bar on the fraction left UNINDEXED, which is well
defined over the whole range - was implemented and measured. It regressed the
37-crystal battery from 34/37 to 32/37 correct space groups: a C2 lattice fell to P1,
and a P2 case went to C222 keeping 2923 of 22440 reflections with CC1/2 in the last
shell at -35%. The indexed fraction is too noisy to carry a looser test.

So the unreachable-but-safe form stays, and the limitation is recorded at the comparison
rather than left to be rediscovered. Fixing it properly needs the selection to be
decided on something better than the indexed fraction.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 15:32:55 +02:00
leonarski_fandClaude Opus 5 abb94ca450 spot_finding: accumulate the adaptive ring statistics in integers
The per-ring sums were floats reduced by atomics, so the ring sigma - and with it the
detection threshold - depended on the order the blocks happened to arrive in. Detection
compares an INTEGER pixel value against that threshold, so a threshold that drifts
across an integer flips every pixel of that value in the ring at once, which is how a
last-bit difference turned into a different spot list.

A preprocessed pixel is an exact int32 and the masked and saturated sentinels are
skipped, so v and v*v are exact in 64 bits, and integer addition is associative: the
sums no longer care about arrival order. Both engines now accumulate the same way, so
they agree exactly rather than approximately, and the GPU spot list is bit-identical
across runs. The corrected sums that feed the reported azimuthal profile stay float -
a pixel value times a float correction has no exact integer form - but they do not
enter the detection decision.

Cost: the ring reduction needs 28 bytes per bin instead of 20 in the plain pass, which
drops it from eight co-resident blocks per SM to seven and costs about 11% of that
kernel (0.582 -> 0.650 ms/frame on a 4.5 Mpx frame). End to end it does not show:
alternating runs on three rotation crystals came out the same or slightly faster, and
the battery is unchanged in every number. The CPU engine got 30% faster (32.2 -> 22.6
ms/frame), integers being cheaper than doubles.

Tests: exact CPU/GPU agreement on the spot list, and 50 repeats of bit-identical output
where there were four.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 15:19:11 +02:00
leonarski_fandClaude Opus 5 212fbf9bab rugnux: make the geometry-refinement sample deterministic, and spread it over the run
The stills first pass drew frames from a shared cursor and stopped when a shared counter
reached its target, which got two things wrong at once. The cursor walked the equally
spaced sample in ascending order, so stopping early read only its leading PREFIX - the
beam centre, distance and cell were fitted to the beginning of the run, not across it,
and the comment claiming otherwise was wrong. And where the stop landed depended on how
the workers happened to interleave, so the set of frames varied run to run: on the same
data at -N 32 and -N 8 the pass examined 483 and 457 frames and refined the detector
distance to 168.0481 and 168.0530 mm.

The sample is now cut into a fixed number of interleaved stripes, each stopping once it
has contributed its share. Every stripe spans the whole run, so an early stop no longer
biases the fit, and a stripe is processed identically whichever worker claims it - so
what gets examined depends only on the data, not on timing and not on -N. The same three
runs now give 451 frames examined and 168.0452 mm, identically.

The bundle selection was order-dependent too: frames are collected in worker-completion
order and sorted by spot count with a non-stable sort, so equally strong frames swapped
places between runs. They carry their image ordinal now and it breaks the tie.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 15:19:11 +02:00
leonarski_fandClaude Opus 5 2d3c39c9dd image_preprocessing: write whole elements out of the un-transpose
The raw-bytes path assembled each element a byte at a time, which on a full frame cost
about 4x against writing the 8 contiguous elements a thread owns through an
element-typed pointer. They are 8*ES-byte aligned, so the compiler merges them.
72.4 MB frame: 1.524 -> 0.406 ms for upload plus both kernels.

The test now also times the LZ4 pass on its own, so the bounds and validity checks in
the hot loop can be costed rather than guessed at. They are free: 0.231 ms against
0.2297 ms measured for the kernel before any of them existed - the restored offset == 1
and power-of-two fast paths pay for them. compute-sanitizer memcheck reports no error
over 400 single-bit-corrupted payloads and nine malformed containers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 14:16:58 +02:00
leonarski_fandClaude Opus 5 4b1c611bdf image_analysis: query the current device, and upload the resolution mask on the engine stream
Two leftovers from earlier fixes of the same shape. BraggIntegrationEngineGPU still read
device 0's shared-memory size to decide whether its profile grid fits; workers are pinned
round-robin across GPUs, so on a heterogeneous node that check can pass on a different
card than the one the kernel launches on. SpotExtractorGPU still uploaded its default
resolution mask with a pageable copy on the NULL stream, which is not ordered against the
engine stream now that streams are created non-blocking.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 14:14:24 +02:00
leonarski_fandClaude Opus 5 bec7e2e922 image_preprocessing: fuse the bitshuffle inverse with preprocessing, and verify the decode
The device decoder was byte-exact on every valid input - 994 production-compressed
images, 927 hand-built LZ4 blocks covering engineered (offset, matchlen) pairs across
the overlap branch boundary, 18000 repeat decodes, sanitizer-clean - and an audit
against LZ4_decompress_generic could not construct a valid block it mis-decodes. What
it did not do was notice when the input was NOT valid, and that mattered more than it
looks: the decode buffers are reused frame to frame, so a block that stopped early left
the PREVIOUS image in place, and in the bitshuffled layout the untouched tail is the
most significant byte-plane. A corrupt chunk therefore did not look like a missing
corner. It looked like thousands of real pixels several powers of two too bright, fed
to spot finding with no diagnostic, where the host decoder had raised an error.

So the kernel now flags a block that fails to reach its declared length while consuming
exactly its payload, and the host turns that into an exception once the caller has
synchronised. Reads are clamped against the end of the payload as well as the output,
both length chains are bounded exactly as read_variable_length bounds them, the two
offset bytes are bounded, and LZ4's parsing restrictions are enforced. On the host side
a block size that is not a multiple of 8 elements is rejected (it made the un-transpose
read uninitialised shared memory), the block count is bounded by what the chunk could
hold before it becomes an allocation (twelve header bytes could demand hundreds of MB
of pinned memory, permanently, per worker), trailing bytes are rejected, and the stream
is synchronised before any throw that happens after work is queued. An image of fewer
than 8 elements is all verbatim tail and now decodes rather than throwing. When the
device route fails for any reason the host decoder gets its turn, so it costs speed
rather than the acquisition.

The lanes cooperate on the copies and a later match can read bytes another lane wrote,
which since Volta needs an explicit __syncwarp(); it worked only because ptxas happened
to reconverge at the post-dominator. The prototype's offset == 1 and power-of-two fast
paths are also restored - the shipped kernel ran a runtime modulo, an emulated 32-bit
division per output byte, on the path its own comment calls the common case.

The un-transpose is now fused with preprocessing. One thread owns one group of 8
elements across every byte-plane, so once it has transposed its 8 bytes out of each
plane it holds 8 complete elements and emits 8 finished int32 pixels with the mask, the
error marker, the saturation cap and the statistics applied. The decompressed image is
never materialised: 0.623 -> 0.411 ms/frame at 18 Mpx, 0.523 -> 0.340 with 8 concurrent
workers. Staging nothing in shared memory also drops the 48 kB ceiling, which had made
any file whose bitshuffle blocks exceed it a hard failure; 64 kB blocks now decode.
gpu_compressed is sized from the chunk with grow-on-demand instead of from the
uncompressed size - it was reserving ~73 MB per worker to hold ~4 MB. Measured on a
1630x1553 uint32 rotation set at -N 32, peak GPU memory falls 3756 -> 3084 MiB; the
same model gives ~144 MB per worker on an 18 Mpx frame.

Decoding on the device also stopped reporting a decompression time, which blanked the
broker's compression plot trace and filled /entry/profiling/compressionTime with NaN.
The decoder brackets the decode with CUDA events and reports it again.

Tests: a differential fuzz suite against the CPU decoder - incompressible and highly
compressible data, engineered offsets, a size sweep hitting every rem%8 value twice,
all six element sizes, an 18 Mpx frame, decoder reuse, concurrency, hand-built LZ4
blocks across the overlap boundary, 26 foreign bitshuffle block sizes from 128 B to
64 kB, corrupt payloads and malformed containers, with a coverage report that proves
which LZ4 paths were reached rather than assuming it. Plus the fused path held byte for
byte against ImagePreprocessorCPU, statistics included, and against the host-upload
path on the same frame.

Battery: 37 crystals, every merged number identical to the host-decode run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 14:13:55 +02:00
leonarski_fandClaude Opus 5 7e47afe47f rugnux: parallelise candidate-cell refinement, and stop repeating work in the tail
Build Packages / build:viewer-tgz:cpu (push) Successful in 20m26s
Build Packages / build:viewer-tgz:cuda (push) Successful in 21m30s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m36s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 24m4s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m10s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m12s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m23s
Build Packages / XDS test (durin plugin) (push) Successful in 11m21s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 20m56s
Build Packages / build:rpm (rocky9) (push) Successful in 21m10s
Build Packages / Generate python client (push) Successful in 40s
Build Packages / Build documentation (push) Successful in 1m34s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (rocky8) (push) Successful in 25m28s
Build Packages / DIALS test (push) Successful in 21m15s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m26s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m51s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m53s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m41s
Build Packages / Unit tests (push) Successful in 2h21m29s
Build Packages / build:windows:nocuda (push) Successful in 1m15s
Build Packages / build:windows:cuda (push) Successful in 28m0s
Three independent changes to the CPU-bound parts of an offline rotation run, none
of which alters a result.

Candidate-cell refinement now splits across threads. RefineCandidateCells already
took a (block, nblocks) partition, but the only call site passed nblocks=1, so the
whole first pass of a two-pass rotation run sat on one thread per scheme - two
threads, unchanged at every -N, for a third of the run. A block touches only its
own scores(j) and cells rows and holds its own scratch, so the split is exact.
The budget is a new IndexingSettings::RefineThreads, left at 1 by default and set
only where few indexer threads exist: raising it unconditionally would
oversubscribe the paths that already run one indexer per image across all workers.

The mmCIF writer built a std::ostringstream per formatted number, twelve per
reflection. snprintf gives the same digits for 0.535 -> 0.220 s per file.

The space-group search built the same orbit mapping twice per candidate point
group - once for the merge chi^2 and once for the systematic-error b, an
apply_to_hkl and Canonicalize per observation per operator each time. Build it
once and hand it to both.

18 Mpx rotation set 24.6 -> 18.7 s, 2.5 Mpx 13.0 -> 10.7 s, and the 37-crystal
battery 13m55s -> 10m47s with no failures, the same 34/37 space groups, and
statistics unchanged on 30 of 37 (the rest drift within the run-to-run spread the
binary already had, which a control build with the split disabled reproduces).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 07:30:33 +02:00
leonarski_fandClaude Opus 5 13aa20a528 bragg_integration: grow the GPU reflection arrays with slack
EnsureCapacity resized its 13 device arrays to exactly the current image's
predicted-reflection count, so every image that set a new record freed and
reallocated all of them. cudaMalloc and cudaFree take a device-wide lock in the
CUDA driver, so those images stalled every other worker: sampling the worker
threads during the per-image loop found 21-24 of 32 parked in cuMemAlloc_v2 or
cuMemFree_v2, all called from this one function, and the running maximum makes
32 workers do far more allocator work than one does.

Grow by half again instead. All transfers and kernel launches are sized by the
per-image reflection count rather than by the capacity, and the member is
already documented as holding at least that many, so over-allocating changes no
result. On an 18 Mpx rotation set the integration stage drops from 1.37 to
1.25 ms per image at 32 workers; merged statistics, error model and adopted
space group are unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 01:02:07 +02:00
leonarski_fandClaude Opus 5 6e4c0ce202 image_preprocessing: decode bitshuffle+LZ4 on the GPU
Build Packages / build:viewer-tgz:cpu (push) Successful in 20m32s
Build Packages / build:viewer-tgz:cuda (push) Successful in 20m40s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m24s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m8s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 27m31s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 27m38s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 29m7s
Build Packages / XDS test (durin plugin) (push) Successful in 11m12s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 22m49s
Build Packages / build:rpm (rocky9) (push) Successful in 22m51s
Build Packages / Generate python client (push) Successful in 40s
Build Packages / Build documentation (push) Successful in 1m22s
Build Packages / Create release (push) Skipped
Build Packages / DIALS test (push) Successful in 20m21s
Build Packages / build:rpm (rocky8) (push) Successful in 27m26s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m59s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m52s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m41s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m41s
Build Packages / Unit tests (push) Successful in 1h17m41s
Build Packages / build:windows:nocuda (push) Successful in 13m24s
Build Packages / build:windows:cuda (push) Successful in 17m0s
The pipeline decompressed each image on the host and uploaded the result. On
an 18 Mpx rotation dataset that made the host-to-device copy the bottleneck of
the whole per-image loop: nsys puts the copies at 78% of the loop against 39%
for every kernel combined - 3600 transfers of 72.4 MB - and they ran at only
12.5 GB/s of an available 27-28 because the host-side decompression was itself
saturating host memory bandwidth. The GPU was mostly waiting.

So the compressed chunk goes across instead, about 4 MB rather than 72 MB, and
is decoded on the device. That removes the transfer and the host decompression
that was throttling it, in one change. Measured on an idle machine, a run goes
from 45.11 s to 24.97 s - 1.81x - with the merged output unchanged.

THE APPROACH IS JON WRIGHT'S (ESRF): "Experiences with GPU decompression for
bitshuffle + LZ4 data", HDF5 User Group 2021, and github.com/jonwright/
bslz4decoders. The kernels here are ours, but the idea and the demonstration
that it is worth doing are his. Cited in docs/ACKNOWLEDGEMENT.md and in the new
section 0 of docs/CPU_DATA_ANALYSIS.md.

Two kernels mirror the CPU decoder. LZ4 runs one WARP per bitshuffle block:
every lane parses the same sequence stream (a broadcast read, no divergence)
and the literal and match copies are split across the 32 lanes so the stores
coalesce; an overlapping match is treated as a pattern of period offset sourced
from bytes that already precede the write position, which keeps it parallel
rather than a serial byte loop. One thread per block instead measured 13x
slower. The bitshuffle inverse then un-transposes each byte-plane through
shared memory and interleaves the planes back into elements.

Only BSHUF_LZ4 is decoded on the device. The zstd variants have no device
decoder, and neither has an uncompressed or float image; Supports() returns
false for those and the caller decompresses on the host exactly as before. The
fallback is explicit, so a format we cannot decode on the device is a slower
path and never a wrong answer.

Tests hold the device decoder against the CPU one byte for byte, on data from
the production compressor, for every element size the detectors emit -
including the 8-bit DECTRIS modes, which take bitshuf_decode_block's separate
elem_size == 1 branch - plus a many-block frame, the formats it must decline,
and malformed containers, which must throw rather than run off a buffer.

Battery: 37 crystals, no failures, identical to the host-decode run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 00:16:36 +02:00
leonarski_fandClaude Opus 5 59702b0123 spot_finding: give the ring reduction eight blocks per SM instead of four
reduce_rings_shared is the largest kernel in the per-image loop - 73% of GPU
kernel time on an 18 Mpx rotation run, launched three times per image - and it
is bound by shared-memory atomic replay rather than by bandwidth: it reaches
156 GB/s against a measured 913 GB/s ceiling, and removing the atomics while
keeping the same loads makes it five times faster.

That is the case that wants resident warps to hide the serialisation, and four
blocks per SM left only 512 of the 1536 threads an SM can hold. The per-block
histogram is nbins * 20 B, about 9.6 kB at the default 0.01 1/A spacing, so
eight blocks fit in shared memory with room to spare. Both kernels are
grid-stride loops, so any grid is correct and a device that cannot co-schedule
eight simply queues the rest.

Measured: 9.21 s -> 5.33 s of kernel time over a run (852 -> 493 us per
launch), cutting total kernel time from 12.57 s to about 8.85 s.

flag_strong keeps four. It is bandwidth-shaped rather than atomic-bound and
eight measured no better (181 vs 175 us).

Wall clock is unchanged, and that is expected rather than disappointing:
kernels are 39% of the image loop while the host-to-device copy is 78%, so
faster kernels idle the GPU more without shortening the loop. This is
groundwork for the transfer work, not a speedup on its own.

The shared accumulators are float and summed with atomics, so the block count
changes the summation order and with it the last bits. The 37-crystal battery
is identical crystal for crystal except one observation in 925850 on a single
dataset.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 00:02:34 +02:00
leonarski_fandClaude Opus 5 a047275760 spot_finding: fix the GPU finder's main loop, which the tests could not reach
Two bugs in analyze_pixel, both confined to the middle stage of the wave.

The kernel walks each wave's rows in three stages. The priming and drain
loops read prev_out and substitute INT32_MAX for a pixel the previous pass
found strong, exactly as the CPU finder's value_at() does on every read. The
main loop did not - it read the image raw. So in the second pass the pixels
the first pass found strong stayed in the background statistics, inflating the
local mean and variance, and the halo of every broad spot failed the
signal-to-noise test. The two engines therefore did not agree, despite
1a1e05ad1 having set out to make them.

Separately, shared_sum2 is an int64 accumulator but val*val and old*old were
computed in int32 at three sites. That wraps above 46340 while the detector
overloads around 1e6, so any window containing a bright pixel got a corrupted
variance. pixel_result already did the same arithmetic in 64 bits.

The reason this survived is worth recording: numberOfWaves is fixed at 32, so
a wave owns ceil(height/32) rows, and the main loop only runs while
front < rmax with front starting NBX+1 = 16 rows ahead. At the existing test's
100 rows a wave owns 4 rows and THE MAIN LOOP NEVER EXECUTES - every row goes
through priming or drain, and the parity test passed on the broken kernel. On
a 4362-row detector frame a wave owns 137 rows and the main loop carries about
121 of them, so the bug covered roughly 88% of a real image.

The new tall-image test is sized to the partitioning rather than to
convenience: 1024 rows, spots placed inside the main-loop region, one core at
100000 to exercise the overflow. Against the unfixed kernel it reports GPU 25
pixels where the CPU finds 49 and fails six assertions.

The default rugnux path is unaffected because it uses the adaptive finder, and
the 37-crystal battery is identical crystal for crystal. The classic finder is
what SpotFindingSettings defaults to, so this is the online receiver's path;
measured there with --no-adaptive-spots, R-meas improves 11.2% -> 10.9% and
CC1/2 98.8% -> 98.9% on one crystal, with multiplicity up on both tried.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 23:31:44 +02:00
leonarski_fandClaude Opus 5 83e95b0c5a indexing: stop computing angles the candidate filter only compares
Build Packages / build:windows:nocuda (push) Successful in 16m15s
Build Packages / build:viewer-tgz:cpu (push) Successful in 19m47s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m4s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m49s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m7s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m2s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m7s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m19s
Build Packages / build:windows:cuda (push) Successful in 15m41s
Build Packages / XDS test (durin plugin) (push) Successful in 10m48s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 19m55s
Build Packages / build:rpm (rocky9) (push) Successful in 21m11s
Build Packages / Generate python client (push) Successful in 40s
Build Packages / Build documentation (push) Successful in 1m44s
Build Packages / Create release (push) Skipped
Build Packages / DIALS test (push) Successful in 20m53s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 26m15s
Build Packages / build:rpm (rocky8) (push) Successful in 27m37s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m36s
Build Packages / XDS test (neggia plugin) (push) Successful in 10m11s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 11m5s
Build Packages / Unit tests (push) Successful in 1h21m20s
Candidate cell filtering called acos three times per candidate to turn dot
products into degrees, then compared those against the min/max angle bounds.
acos is strictly decreasing on [-1, 1], so "angle outside [min, max]" is
exactly "cosine outside [cos(max), cos(min)]" with the ends swapped - the
bounds convert once, and the three acos calls per candidate disappear.

The same loop also re-derived every already-accepted candidate's unit cell on
each new triple, inside the duplicate scan: three more acos each, for every
candidate accepted so far. Those cells are now kept alongside the candidates.

Measured on de-novo serial stills, where the indexer runs once per image:
34.43 s -> 14.17 s on one dataset and 21.92 s -> 6.59 s on another, with the
indexing rate and the merged reflection count unchanged (one gained 0.25
points of indexing rate). acos had been 40% of the whole process there.

Scope is narrower than that number suggests, and worth stating: the win is on
the de-novo path, which Auto selects for stills only when NO cell is known.
With a known cell Auto picks ffbidx, which reaches the same filter but feeds
it few candidates - measured neutral there (+0.5% instructions, -1.6% wall,
identical output), and that path already runs 14x faster in absolute terms.
Rotation runs the indexer twice per dataset rather than per image, so it is
unaffected: the full 37-crystal battery is identical, crystal for crystal.

Comparing cosines instead of angles can only move a candidate that sits on the
bound, so the filter's behaviour is unchanged except at that measure-zero
boundary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 21:24:57 +02:00