jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports.
jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls.
Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results.
Rugnux: --model rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable.
Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate.
Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do.
Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence.
Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags.
Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check.
Rugnux: Clear error messages when a data set needs more GPU or host memory than is available.
* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports.
* jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls.
* Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results.
* Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable.
* Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate.
* Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do.
* Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence.
* Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags.
* Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check.
* Rugnux: Clear error messages when a data set needs more GPU or host memory than is available.
Ice-ring reflections are flagged and "excluded from scaling, kept for merging",
but ApplyCellSurface never read the flag. While the fit dropped Is <= 0 this was
mostly hidden: the over-subtracted half of the on-ring observations went with the
filter. Once the negatives were admitted (1d1a40d90), a set with strong, azimuthally
uneven rings fitted its surfaces to the rings: with the surfaces off the error
model's b is 0.018, and the surfaces no longer took it down.
On the two in-house lysozyme sets with heavy ice (25% of reflections flagged):
ISa 7.57 -> 10.08 and 7.08 -> 10.16, b 0.0155 -> 0.0090; on a third, 3.57 -> 4.37
(4.41 before the filter was removed). A run where no ice is detected sets no flag
and is unchanged. The remaining gap to the filtered surfaces (ISa ~15) and to
XDS (~20) is not ice: extending the flag to the measured powder rings moved
nothing, and a symmetric residual cut and a variance rebuilt at the reference did
not either.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
lyso_micromax_{mono,pink} and lysoI_micromax_{mono,pink}: 1800 x 0.1 deg on a
JUNGFRAU 9M, P4(3)2(1)2. XDS references (FRIEDEL'S_LAW=FALSE) cut where CC1/2
falls through ~30%: 1.50 / 1.45 A native, 1.65 / 1.65 A iodine; ISa 39.9 / 37.4 /
31.4 / 29.2. The iodine sets carry the anomalous signal (low-resolution CCanom
~0.67). XDS.INP and the full-range CORRECT.LP.edge are kept in each set's xds/.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
5ky6 5mln 5t39 6cdl 6f3p 6g1f 6jgh 6nen 6qaj 6s1u 6w75 6zqr 6zqy 6zr0 7ou1 7raa
9fcf 9h0q, references from the deposition. Five are primitive screw-free groups
(6zqr 6zqy 6zr0 9fcf in P4, 6nen in P312) as negative controls for screw
detection; 6qaj, 9h0q and 6g1f have a 330-375 A axis.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
Adds 5ky6 5mln 5t39 6cdl 6f3p 6g1f 6jgh 6nen 6qaj 6s1u 6w75 6zqr 6zqy 6zr0
7ou1 7raa 9fcf 9h0q in the existing row form: deposited values from RCSB, the
detector from the image files, DOIs resolved (6NEN's is a Crossref DOI whose
landing page refuses scripted access). New notes: the five screw-free negative
controls, the 5KY6 RAR set, the damaged 6ZR0 zip, the hand-downloaded 6NEN
archive, two detector conflicts (5MLN, 9H0Q), third-round counts.
The table rows, which were grouped by download round, are now sorted by PDB id;
the round moves into a Round column, and the prose that referred to "the first
95 rows" or "the last 51 rows" refers to the round instead. Counts updated to
171 datasets / 164 PDB entries (all 18 new entries have released structure
factors).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
Drops the Round column and its explanation. Statements that depended on a
round or on table position now describe the whole table: the detector
comparison covers all 163 comparable rows in one conflict table (18
conflicts, sorted by PDB id) with one paragraph on the marCCD/SMV headers;
the per-round "in numbers" sections become one whole-table section
(repository, facility, crystal system, long axes), dropping the per-round
size and file-format tallies; the archive section no longer counts archives
per round.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The row's alternative was labelled P 31 1 2 (151) while its own reason
argued for P 32 1 2; the deposited P 32 has P 32 1 2 (153) as its 312
supergroup, and P 31 1 2 is accepted from it as the hand, so scoring is
unchanged (rugnux's P 31 1 2 still passes as an accepted alternative).
Evidence re-checked on the rc173 run: the added two-folds correlate at
0.98-0.99 against 0.98 for the three-folds and 0.32-0.40 for the 321/622
operators, POINTLESS on the P1 merge picks P -3 1 m at likelihood 1.000,
and the zone centric in 312 but not in 3 reads centric (+435 nats). The
deposition's 55502 unique reflections to 1.747 A are what Laue class -3
holds, so the depositor merged in point group 3; the model breaks the
two-fold at a level it can resolve (R 0.22 vs 0.25 for the two
indexings). Both answers stay accepted.
The row is pinned to the first sweep (0-180 deg at 200 mm); the second
(180-360 deg at 300 mm, half the exposure, files numbered by phi) is the
same crystal's continuation, and the two were tied on frame count, so
the choice rested on the alphabetical tie-break. The earlier P 6 2 2
over-promotion was rugnux's and is refused by the current code.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The image file is authoritative for the detector and the PDB entries'
detector fields are known to be unreliable, so the page no longer keeps
a per-entry conflict table (eighteen rows, among them 5MLN and 9H0Q) or
the paragraphs arguing individual cases. The intro bullet now says in
one clause that the file wins where the entry disagrees; the paragraph
on how marCCD and SMV rows fill the column stays. The table's values
are unchanged - 5MLN (miniCBF "PILATUS3 2M, S/N 24-0118") and 9H0Q (NXmx
description "Dectris EIGER1 Si 9M", E-18-0102) were re-read from the
headers and already matched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The field states the marker a saturated pixel carries, not the last valid count
an NXmx saturation_value states, so it is already the exclusive limit and must
not be raised by one. At the usual 65535 in a 16-bit image the limit became
65536, which no pixel can reach, and no marCCD pixel was ever called saturated.
On one open-arm Rayonix set that costs the space group: ~60k pixels per frame sit
at the marker in one detector block, at 1.37-1.65 A, and integrate as if they
were signal - I/sigma ~11 with CC1/2 ~0. The symmetry operators' correlations
drop below the 0.30 gate on that band alone, so the run merges in P1 and cuts at
1.65 A. With the limit right it is P 2 21 21 to 1.19 A with ISa 10.3.
The miniCBF path is left alone: there Count_cutoff is of order 1e6, so the same
one-count difference cannot reach a pixel.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The count-rate correction clips an over-range pixel to Count_cutoff rather than
marking it - "if the observed rate exceeds the corresponding observed cutoff
rate, the corrected rate is set to the cutoff rate", Trueb et al. (2015)
J. Synchrotron Rad. 22, 701-707 - so the cutoff is a value a pixel can hold and
the limit is exclusive. It is not an NXmx saturation_value, which states the last
VALID count and keeps its +1.
Three of the miniCBF sets in the corpus carry a clipped pixel in every frame:
isolated, in flat background, at exactly the cutoff, with the next-highest count
in the image a tenth of it. Those pixels were integrated as valid measurements.
Only pixels exactly at the cutoff change: 0-2 per frame in 3 of ~30 sets, none in
any EIGER-written CBF, and no set has a pixel within 100 counts below it.
XDS and DIALS read the field as the last valid count and pass such a pixel
through; imgCIF's own overload item states the valid values are those below it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
CCD_IMAGE_SATURATION is in the header of every SMV file the corpus holds, and it
is the value a saturated pixel carries, as the marCCD field is. The reader never
read it, and said in a comment that saturation was "judged on the 16-bit
container alone" - which the line below it made untrue: the pixels are handed out
as 32-bit, so the fallback limit was INT32_MAX and no SMV pixel could ever be
called saturated, whatever the detector did.
The limit is now the header's value, or the container's maximum where a writer
states none - said explicitly rather than left to the fallback, for the reason
above. One of the new open-arm sets carries pixels at exactly 65535.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The length sort in FFTIndexer::FilterFFTResults recomputed Coord::Length() inside
its comparator. Under LTO with -march=x86-64-v3 (CI and production builds) GCC
inlines Length() at several sites in std::__introsort_loop and contracts
x*x+y*y+z*z into FMAs differently at each (fma(x,x,y*y) for the element keys,
fma(y,y,x*x) for the pivot recomputed after a swap). The same element's key then
differs by one ulp between comparisons. Shortlist lengths come from quantised FFT
bins, so exact ties are common on noise frames; on such a tie the unguarded
partition scan passes its sentinel and runs off the index array.
Evidence:
- rc.172 jfjoch_broker disassembly: pivot key after swap at 0x925a83 is rounded
differently from the scan keys.
- The deployed rc.172 introsort, called directly on finite golden-spiral
directions x binned lengths, faults at binary +0x5259ca (the journal's crash
address) in up to ~0.5% of sorts. No NaN is involved; the earlier NaN guards
could not help.
- New test FFTIndexer_ManyNoiseFrames: unfixed rc.172 built with the CI flags
(-march=x86-64-v3 -flto=auto) segfaults in the same introsort from
FilterFFTResults on the GPU FFT path (noise frame 1632), 3/3 runs; passes with
this fix. A non-LTO build keeps Length() out of line and cannot crash, which
is why the suite never caught it.
Keys are now precomputed and sorted with stable_sort. The same pattern was
fixed in SpindleBlindFraction (broker-reachable, 62 lattice rows, |a|=|b| ties)
and LePageLattice::PlaneBasis (rugnux, v/-v exact ties); tie order in the
latter may change.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sort keys computed once, not in the comparator: an FMA-contracted Length() can
round differently at different inlined call sites, which breaks the strict weak
ordering std::sort needs and lets its partition scan run off the array.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
6z9g, P 1 21 1, 120.306 93.815 126.951 / 105.201, 1.76 A, 2300 frames on an
EIGER X 4M (ESRF ID30B; Zenodo 10.5281/zenodo.3874714, CC BY, 12.8 GB).
Its deposited cell is an index-4 superstructure of the sublattice a/2, b, c/2:
four copies of each of two entities, whose NCS pairs superpose as the XOR-closed
trio of near-pure translations (1/2,0,0), (0,0,1/2) and (1/2,0,1/2), rmsd
0.18-0.22 A with 5-9 deg of libration. It is the only such entry among the 283
multi-copy PDB entries that declare a raw-data DOI, and the corpus had no index-4
case at all: every pseudo-translation it carries is index 2 or a setting artefact.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The first-pass rotation fit refines the spindle-perpendicular detector tilt
freely, and on a sweep whose seed spots reach only a few degrees of 2theta the
keystone that would determine it is a pixel or two. The fit then commits
whatever the centroids' own systematics prefer - measured 2.5 deg on one sweep
seeded to 11.7 A - and that tilt mispredicts the detector corners by 20-60 px
against a 6 px integration disc, collapsing the run from 2.8 A to 6.35 A and
biasing the metric-symmetry arbiter to a = b on the way. On synthetic data the
same runaway is reproducible: a 1.5 deg detector error on 6 A reflections walks
the free fit to 3.9 deg, on 8 A reflections to 37.
Three data-driven discriminators were tried first and refuted on the corpus:
a blanket restraint (0.13-0.16 A and up to 28% of ISa lost on the crystals
whose tilt is largest and real, and in-house runs pinned at a header tilt known
to be wrong), the distance's held-out excitation criterion (the rocking angles
move by a tenth of their noise whether the tilt is real or not), and a keystone
comparison of the fitted and header tilts with the beam free in both arms
(236-set battery: ~30 sets moved, one collapsed from P1 to C2, one lost a
screw axis; differences of 0.007-0.6 sigma held the header on right and wrong
cases alike). For a tilt of a few tenths of a degree the keystone is a fraction
of a pixel and nothing in the spots says whether it is real. What separated the
cases was the tilt's size: every set the keystone arm moved carries 0.08-0.46
deg, and a survey of 211 corpus sets leaves the file's tilt by more than 0.56
deg on exactly one - a 2theta arm swung out 12.8 deg that its file records as
square, which the free fit recovers and the held arm refuses by 90 sigma.
So the hardware prior decides whether to ask, and the spots decide. Below
1 deg of walk from the tilt the pass started at the fit is trusted as before,
by construction. Beyond it the winning candidate is refined again from its
start with the tilt held there - the beam centre taking the shift the tilt is
equivalent to, so the header tilt is never paired with a beam fitted beside a
refused tilt - and the two are judged as the rounds of one chain already are,
on the spots each indexes inside the wide gate: the walked tilt stands only
when it leads by more than the count's own noise. A real tilt of degrees has a
keystone of tens of pixels over the seed spots and wins outright; an artefact
has none and loses on a tie. Held rather than bounded, because a box the fit
lands on is the same wrong answer at a smaller size. Decided at the fit, so
every pass and arm of a run handles itself and nothing is carried between
passes; the unconstrained alternative is re-solved beside a refused tilt too.
The chain that drives a candidate to its fixed point becomes a lambda so it
can be run on the held candidate.
The run logs the walk, what it would have moved the far corner by, the shift
the beam took instead and both spot counts; where it was refused the report's
REFINED_DETECTOR_TILT is the starting tilt and REFUSED_DETECTOR_TILT what the
fit had walked to. Tests: a half-degree tilt is found, kept and not reported as
a walk; a walk past the prior is made on synthetic data, and the result's
verdict, counts and geometry agree with each other.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
depdata_check.py compares each open-arm merge with the depositor's own data, per resolution
shell, using the deposited model as the common yardstick: the rank correlation of our IMEAN with
|Fc|^2 minus that of the deposited intensities (or F^2), on the reflections both carry (d < 4 A,
eight equal-count shells), with the data carried into the model's setting by model_check's
change-of-basis search. New row keys dep_cc_delta_all, dep_cc_delta_outer (two outermost common
shells), dep_beyond_cc / dep_beyond_d (our correlation with |Fc|^2 past the deposited data's
limit), dep_kind, dep_d_min, dep_n_common, dep_status, dep_reason. gemmi only, no CCP4; runs on
every open-arm set with a model, and `report` fills it in for older runs (four processes).
Reported in its own report section and the per-set table, never scored.
Rank rather than Pearson correlation: on the phase-1 run a handful of gross outliers (I/sigma
above 1000 at 2.4 A, or ~1000x the shell median at 1.41 A) decided the outer-shell Pearson CC of
several merges on their own.
On the phase-1 run (38 of 40 open-arm sets compared; the two without are the lattice failures)
the outer-shell delta has median -0.012; 22 s for the whole report fill-in.
model_check.py: REFMAC's mmCIF reader stopped with "rdaniso_cif: Atom symbol mismatch" on 12 of 35
entries whose atoms carry ANISOU only in part. The model is now written in the PDB format where it
fits (under 100000 atoms, chain names of at most two characters), where each ANISOU follows its
own ATOM line; all 12 then score. R-factors of the entries that worked before move by < 0.001.
The runner also records model_check's baseline note (e.g. an intensity-only deposition) in
refmac_reason instead of leaving it empty.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Unmasked persistent hot pixels are integrated into whichever reflection's box they fall in; under
rotation one pixel collects a different reflection on every frame that reaches it, and the merge
carries intensities hundreds to thousands of times their shell mean on one or two observations.
HotPixelFinder (rugnux/HotPixels.{h,cpp}) reads the pre-scan sample a second time, once the
beam-stop projection has measured where the background puts the beam, and on each frame calls a
pixel lit when it exceeds its 2 px iso-2theta ring level (max of the ring and 1/16-sector medians)
by 3.3 sigma (sqrt(level) or the ring's MAD) + 2. A pixel lit on at least max(k1, kB) frames is
persistent: k1 = 1 + ceil((osc + 5 deg)/|zeta| / frame spacing) is more than one reflection can
light, kB the binomial bound (0.01 family-wise over the detector) from the ring's own lit rate.
A persistent pixel is masked, as the new PixelMask bit 10, only if it stands alone (component of
persistent pixels <= 2), reads on average >= 10x its ring and its mean excess is above the Poisson
bound; pixels holding the error value on most frames are masked with them. Counting sensors (thickness > 0)
and rotation data only; a CCD is left alone. One log line reports the counts.
Drawing the rings about the file's centre, as a first version did, masked pixels along the
background fall-off on a sweep whose file centre is 171 px from the background's and cost it 14%
ISa; about the measured centre that sweep is within 1%. Masking the detector's outermost row and
column unconditionally was tried and dropped: the persistence test already catches the hot pixels
there, and the whole lines bought nothing measurable.
Numbers below are from the looser first criterion (no isolation / 10x gate), against rc173-final
on the same base: merged reflections > 30x their shell mean gone on the
sets with proven hot pixels (7brr 21 -> 0, 9ih9 15 -> 0, 8xte 10 -> 0, 6z8o worst 1706x -> 40x);
6z8o CC1/2 0.50 -> 0.995, ISa 8.6 -> 12.8, CC to model 0.82 -> 0.90; 8xte ISa 8.0 -> 10.8; 6u7g
ISa 9.8 -> 12.0; controls (lyso_x06da_ref, marCCD) unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Per isolated strong spot of the spot-width sample, second moments along and across its radius, with
the exact per-spot Jacobians (pixels per radian of 2theta and of the angle across the scattering
plane, from the geometry) and the sensor parallax fixed from physics (conversion depth exponential
with length L cos(psi), truncated at the thickness, smeared z tan(psi) along the ray's in-plane
direction). y = (m_rad - 1/12 - par_u) - (jr/jt)^2 (m_tan - 1/12 - par_v) is linear in
(2 jr tan theta)^2 with slope sigma^2; fitted over eight equal-count 20%-trimmed bins, error from 200
seeded bootstrap re-draws inflated by the reduced chi^2, significant at z > 3. The pre-scan logs the
FWHM, its standard error, z and chi2 next to the bandwidth the run uses; nothing consumes it, so
output is unchanged.
The truncated-exponential depth variance moves into SensorAbsorption.h
(ConversionDepthVariance_um2), shared with the integrator's parallax_var_px2 (same arithmetic).
Catch2: a synthetic spot population painted with 0.45% FWHM bandwidth and 450 um Si parallax
recovers 0.449% (z 82); the same spots without bandwidth give 0.06% at z 0.7.
On the battery sets (67d49f base, estimate as logged): MicroMAX pink lyso 0.41% z 7.3, thau 0.42%
z 6.9, CHESS 7B2 9q41 0.36% z 14.6, ALS 8.2.1 8u0i 0.27% z 5.6; mono controls lyso_micromax_mono,
lyso_x06da_ref, 5mln and near misses 9z44, 7orr, lyso_x06da_5keV all below z 3.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The broker rebuilds the outgoing StartMessage from DiffractionExperiment,
and FillMessage hard-coded countrate_correction_enabled and
flatfield_enabled to false. JFJochReceiverLite parsed the true values from
the detector's stream2 start message and dropped them, so every DECTRIS
file written through the broker said neither correction was applied -
DECTRIS enables both by default. pixel_mask_applied had the same defect:
it reported the local apply_mask setting (an FPGA feature), not what the
detector did to the pixels, which ReceiverLite forwards byte for byte.
These now come from the stream: DetectorSetup carries them, ReceiverLite
copies them in Configure, FillMessage reads them. Two more fields the
DECTRIS stream sends and NXmx defines are passed through to the master
file: countrate_correction_lookup_table (uint32, possibly bslz4/bszstd
compressed in the stream) and virtual_pixel_interpolation_applied. The
flatfield is deliberately not written - it makes the master file too
large. PSI EIGER is unchanged: jfjoch never enables rate correction and
the detector server starts with it off.
The reader also opens the DECTRIS "hdf5 nexus v2024.2 nxmx" layout, where
/entry/data/data is 4D [image, channel, y, x] and the pixel mask is kept
per channel. Only the first channel is read; this is compatibility, not
full multichannel support. A _process.h5 made from such a file links its
pictures to that channel.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The mask written with the data is Jungfraujoch's (the detector's own plus
user masking) and is never uploaded to the detector. The stream's
pixel_mask_enabled only says the detector applied ITS mask, so it does not
describe the mask in the file; on this path that mask is never applied.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The correction-surface gate averaged the per-shell change of the held-out half-set CC over 20
equal-occupancy shells and adopted a surface on its sign. On the shells where a multiplicative error
shows, CC1/2 sits at 0.99-0.9996 and a large reduction of the error moves it in the fourth decimal;
the shells beyond the data's reach (up to 12 of 20 at CC ~ 0) each add +-0.01, so the sign was set
by noise (measured spread of the mean +-0.001-0.003 against effects of -0.0003..-0.0008). It refused
the 24x24 detector surface on four EIGER2 16M sweeps. The change is now averaged on atanh(CC) - the
candidate from the hq-bisect investigation.
The order of the three overlapping surfaces also decides what is adopted. With the gate fixed,
goniometer-frame first took a 5 keV insulin sweep from ISa 31.8 to 20.1 (the time surface then
refused; R_meas 7.8 -> 8.9%, CC to the model 0.8307 -> 0.8282) and cost a 0.1 deg-sliced thaumatin
sweep R_meas 18.2 -> 19.6%; time x detector first refused the detector surface on the 16 keV
thaumatin sweep (ISa 33 instead of 40). Modulation, time, goniometer frame avoids both.
Bare runs, --report-resolution at XDS's range, against rc173 (0252880e0): ISa / lowest-shell R_meas
thaumatin 16 keV 32.9/.030 -> 39.8/.028, lysozyme 90 deg sweeps 10.1/.059 -> 14.1/.048 and
10.1/.060 -> 13.8/.049, cytochrome c 19.2 -> 20.4 and 17.3 -> 19.3, thaumatin 3.8 keV 13.3 -> 14.8;
CC to a deposited model (standard-protein models on 14 in-house sets, the deposition on 13 open sets)
within +-0.0045, R_model_shell_scaled within +-0.002 except where d_min moved. 46 sets, one regression:
a half-image lysozyme test set, CC1/2 0.878 -> 0.862.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The median outlier test needs three observations, so a reflection measured
once or twice (one good observation and one artefact, the usual case after
Friedel merging) was never tested, and where there were more, a precise-looking
artefact (a hot pixel, thousands of counts) could outweigh mates from weak
frames and become the weighted median itself.
Every observation of the written merge is now also judged against Wilson
statistics beyond 4 A: E^2 = I / (epsilon <I/epsilon>_shell), centric and
acentric laws, a bound set by alpha = 0.01 expected false rejections per
dataset (ln(2N/alpha); twice that for centrics), the lower confidence limit
I - z*sigma (own sigma, z from the same budget) required to exceed it, shells
whose <I> is not established at that significance not judged, and the bound
widened by the dataset's measured tail scale (peaks over threshold on the
well-measured shells; 1 for a Wilson crystal, larger under tNCS/anisotropy).
An improbable singleton is rejected; one with company only when most of its
reflection's other observations are probable and it disagrees with them. The
reflection is the group, or the Friedel pair under -A. Search merges are not
tested (epsilon is 1 in P1, and the screw rows are the reflections a wrong
epsilon would call improbable). The flags are handed to the device merge as
pre-rejections, so CPU and GPU agree.
Report: OBSERVATIONS_REJECTED_WILSON= (part of OBSERVATIONS_REJECTED=), and the
developer report lists each rejected observation (hkl, d, E^2, image, x, y).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The help text claimed 0.99 is right for every synchrotron in the corpus.
Bending-magnet, superbend and wiggler sources are plausibly 0.8-0.95, no file
format read states the value reliably, and it cannot be fitted from a single
dataset (it correlates with spindle-symmetric absorption). Say what the number
is (degree p; XDS FRACTION_OF_POLARIZATION = (1+p)/2) and that a known
beamline value can be passed. Documentation only; the default stays 0.99.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The pair rule drops the improbable (higher) member of a discordant pair. If the
lower member lost its core to saturation or the mask, its profile estimate of
what remained may be low, and a genuinely strong reflection would be replaced
by the clipped value - the failure Aimless guards against.
Integration now marks a reflection whose signal disk was not fully readable
(Reflection::clipped, from the engines' existing full-disk test); the flag is
carried through the partials to the fulls (OR over an event's partials, CPU and
GPU combine alike), and the Wilson test never counts a clipped observation as
a probable witness against a larger one.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The pass-1 post-refinement fits the rotation scale k only on the frames the stored angles still
track, and a rate error is exactly what stops them tracking the rest: on a sweep whose stage turned
~3 % slow the fit read 0.979 over 94 deg, failed its leave-a-fifth-out test and was thrown away,
leaving half the frames unscaled.
The pass no longer decides. Between the passes, at the pass-2 detector geometry, the lattice is
indexed (index-only probe) under the stored angles and under the fitted k and scored on the
validation frames of the whole sweep: share of the spots on the lattice beyond the wrong-spindle
null. k is adopted only where it scores higher by more than the binomial noise of the two
(ValidationEvidencePrefers - the test the beam-centre arms already used, now one function); the
run then integrates and post-refines at k (post-refine-only probe), fits again on top of it and
repeats until the next k no longer scores better (WalkRotationScale). The stored angles are the
first hypothesis. Measured: 1.000 25.1 %, 0.97874 58.9 %, 0.97041 90.0 %, 0.97006 90.7 % (not
significant) -> 0.97041 adopted.
Probes restore the experiment, the pass-2 geometry, pass-1 mosaicity and the beam-centre-search
flag; a probe opens no beam-centre search. Forced pass-1 results get their axis scaled; the header
revert drops the scale. GONIOMETER_ROTATION_SCALE reports the adopted k (SUSPECT = adopted). The
leave-a-fifth-out figure stays in the log only.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
OSCILLATION_RANGE was read off the caller's experiment, so a run that adopted a goniometer rotation
scale still reported the file's oscillation; it is now scaled by the adopted k (the starting angle is
the stage's and stays). The post-refinement logs the range of k over the leave-a-fifth-out folds
instead of a ratio to k - 1, which is meaningless near k = 1.
PostRefine_RotationScale: partials of a 180 deg sweep simulated at a known stage rate are fitted
back to it (0.97 and 1.0, to 1e-3).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The pre-scan's spot-shape estimate (spot_width::EstimateBandwidth) becomes the run's bandwidth where it
is significant (z > 3) and neither the file nor --bandwidth states one; an explicit value, including
--bandwidth 0, still wins. Everything that already reads GetBandwidthFWHM() - prediction, partiality,
the profile's radial width, the stencil growth, scaling - takes it from there. Below z 3 the bandwidth
stays unset, so monochromatic runs are bit-identical (12 mono sets checked on p.hkl md5).
The width tiers stop as soon as r80 settles, often at 15 frames and a few hundred spots, which is
too few for the slope. The pre-scan now keeps measuring spot shapes for the bandwidth alone until the
pool holds 1000 spots or the sample ends. Those extra frames feed neither r80 nor the powder
measurement, so r1..r3, powder and the spot-resolution quantiles are unchanged. Over the 68-set
battery plus the known mono controls (47 sets re-probed) no monochromatic set crosses z 3 (highest
lyso_x06da_5keV 2.0, insu 5-6 keV 1.7); the one new positive is 9z44 (ALS 8.2.1, multilayer
beamline, 0.38% at z 3.1 from 382 spots instead of z 1.0 from 94).
Two hidden b > 0 mode switches go, so the bandwidth acts continuously:
- the background clip default dropped from 4 to 3 sigma when a bandwidth was set. With the measured
bandwidth, clip 3 vs 4 on the MicroMAX pink sets: REFRES R_meas 0.0507/0.0507 (lyso),
0.0842/0.0844 (thau), REFRES ISa 28.83/28.77, 14.12/14.12 - no bandwidth-specific effect.
- the "--integration-stencil has no effect without --bandwidth" message, which is decided before the
pre-scan and would now be wrong.
The one set where clip 3 helps (9z44, 7 A: R_meas 0.252 -> 0.201) gains the same at zero bandwidth, so
that is a property of the clip, not of the beam.
Effect of the measured bandwidth (stencil 0, clip 4), against hq-integ 9b6736dbd at the same
resolution range (REFRES; open sets rerun with --report-resolution at the base d_min): R_meas up on
all six consumers - MicroMAX pink lyso 0.0499 -> 0.0507, thau 0.0824 -> 0.0844, 9q41 0.250 -> 0.265,
8u0i 0.0796 -> 0.0816, 9sl0 0.0758 -> 0.0773, 9z44 0.252 -> 0.255; CC1/2 unchanged to 1e-4; CC_MODEL
overall +0.001/+0.0004/+0.0005/-0.0025 (9q41/8u0i/9sl0/9z44); no space-group change.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Where neither centre indexes the whole spot list, the check ran the lean-rung ladder at the file's
centre and at the background-measured one and kept whichever indexed more validation frames. On
the same lattice those counts are two readings of one hypothesis, and the frame count is the
statistic the check itself says must not arbitrate centres: on one rotation set the measured
centre read 56/60 against the file's 50/60, then 48/60 against 50/60 once the hot-pixel mask took
out ~3k defective pixels. The run then kept a centre 12.6 px out, the post-refinement moved it
0.15 px, and the merge lost it: 32% of frames rejected (17%), low-res R_meas 14.7% (8.1%),
CC_MODEL 0.42 (0.89). The P1 setting change seen alongside it (alpha/beta exchanged) is an a~b
Niggli-boundary flip of the same lattice and was not causal; model validation reindexes it.
Now, when both ladders index a majority on the same lattice (class + primitive volume within
2%), the file's centre stays for the pass and the measured centre is carried into the existing
two-arm arbitration in RunAllPasses. The measured arm starts from the depth (rings/seed/d_min)
the check found it at - a fresh pass at that centre walked 1 px onto a 2x supercell instead -
and the file arm's spot-finding settings are restored with the rest of its state when kept.
The arbitration moves to MeasuredCentreWins(): arms on the same lattice are judged on the
reflections at I/sigma >= 2 of their search merges (the quality guard's signal statistic),
without the guard's 10% margin, which covers header-vs-refined passes integrated with different
spot width and mosaicity - the two arms here integrate identically. Arms on different lattices
keep the CC1/2 + 0.05 rule. On the set above: 10095 vs 9116 -> measured centre, CC_MODEL 0.866,
R-free 0.333 (rc173 0.333), R_MODEL_SHELL_SCALED 0.316. 27 other sets bit-identical.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
A rotation frame records the part of a reflection's rocking curve inside its own oscillation, and
while the crystal turns through the curve the spot walks along its Debye ring. The predictor put
every partial at the exact diffracting condition, so a partial recorded on the curve's flank was
integrated pixels away from where its flux landed. The walk is largest where the reflection moves
nearly tangent to the Ewald sphere (low |zeta|), whose curves are widest.
The prediction now turns S about the beam by the rotation's component along the ring times the
flux-weighted centre of the frame's slice of the curve (a truncated-normal mean, RockingSlice.h),
on the CPU and the GPU predictor alike. 2theta is unchanged; a frame that straddles the condition
symmetrically gets no shift.
Measured (observed r1 centroid minus prediction, tangential, I/sig > 10): against the slice model
correlation 0.93-0.98 before, 0.0 after; residual rms 1.0-1.8 px -> 0.35-0.48 px at |zeta| < 0.3.
Emulated capture of low-|zeta| partials at high angle 0.50 -> 0.94 of a 12 px aperture. CPU runs
on four rotation sets (400-800 frames): R_meas -0.03..-0.09 pp, ISa +1.2..+1.8, rejected
observations roughly halved, space groups unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The rotation merge fitted var = a*s2 + b^2*<I>^2 from three separate medians (s2, I2, dev2) per
bin of I2. A median of dev2 over observations whose variances differ is not 0.455 times their mean
variance, so the ratio of medians read a too low and b too high: on the scaled fulls of 28 sets
(in-house, open and private) the core of the normalised deviations scattered at up to 1.8x its
stated variance in the weak and middle bins and at 0.1-0.7x in the strongest. Reproduced on
synthetic samples with a known model (a 1.3 read as 1.12, ISa 33 read as 31).
Now (ErrorModel.h/.cpp, host-only, so the GPU and CPU paths share it):
- bins are equal counts in counting I/sigma (I2/s2), where b is identified;
- each bin is calibrated on the median of dev2/var with var from the previous iteration, iterated
to a fixed point - heterogeneity inside a bin no longer biases it, and the median keeps it robust
to tails;
- s2 is the counting variance the merge actually applies the model to (rebuilt at the reflection's
mean), not the observation's own sigma^2.
The separate 6-sigma misfit refit is gone: the median does not need it.
A mean-based fit (misfits cut at z^2 > 2 ln N) was tried first: it calibrates the total variance
best (median rms log chi2 over the 28 sets 0.14 vs 0.19 here) but on heavy-tailed data it sizes
the sigmas on the tails, the merge's outlier test widens with them, and CC1/2 fell 0.80 -> 0.71 on
a powder-contaminated set (0.83 with this fit). Rejected for that.
Offline, 28 sets: rms log chi2 of the median normalised deviation over (counting I/sigma x
resolution) 0.208 -> 0.139 (better on 23), of the mean 0.239 -> 0.191 (better on 22).
Battery, 45 of 48 sets against the 9b6736 run (3 lost to CUDA OOM from GPU contention): space group
unchanged on all; d_min unchanged except two poor multi-lattice sets (1.69 -> 1.56, 1.96 -> 1.90);
ISa x1.13 (median), in-house ISa/XDS 0.73 -> 0.93; ISa*R_meas_lo/0.8 0.93 -> 1.05 (XDS ~1.2);
CC1/2 over the XDS range +0.002 (mean; up 0.014-0.031 on the three poorest sets, else +-0.0001);
CC_model +0.0011, R_model_shell_scaled -0.0007 (mean over 17 open sets); CC_anom +0.003 (mean).
Six private sets: space group, d_min and CC1/2 unchanged, ISa up by 14-67% towards XDS's.
Remaining misfit, not addressed: the excess variance grows slower than <I>^2 (the effective
fractional error falls 2-2.5x from counting I/sigma 5 to 200), so the strongest reflections still
scatter below their sigma on open sets. A third, linear term (as in Aimless) fits it better on most
sets but leaves b unidentified on some (b -> 0 on 4 of 28); not landed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
A pass 2 that re-indexes pass 1's axis-multiple supercell to the true cell
reports it primitive-centred; the guard read that as a dropped centring and
forced pass 1's supercell back (completeness 99.7% -> 65% on a 7x case).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
ComputeSmoothGWindow leaves a frame below MIN_CREDIBLE_SCALE_RATIO of the
median out of every window, and DropCollapsedScales drops it after the loop,
but RunScalingLoop still counted its step. Such a frame has lost its vote in
its own references and keeps moving, so the loop was stopped as "unsettled".
On 8xtf (hq-pool battery) the first-pass partials loop stopped after 16 rounds
at rms dlogG 1.8e-01 (base 9.5e-04); with the frame left out of the step it
settles in 23 rounds (8.0e-04), and every other loop of that run that had
stopped unsettled (pass 2 at 1.6e-03) now settles too. Every loop of that
sweep that stopped unsettled was one that went on to drop a frame. The final
8xtf merge is bit-identical; the false SCALING_NOT_CONVERGED warning goes.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The delta-CC1/2 disposition keeps a weak stretch in the merge (downgraded)
wherever removing it would not raise CC1/2, and the merge then carries each
of its observations at the small 1/sigma^2 its scaled-up counting error gives
it. R_meas counted those observations at full weight, so the frames that add
next to nothing to the intensities set the number. hq-pool battery: 8xtf kept
115 weak frames (scale 0.1-0.2 of the run's) with the same CC1/2, <I/sigma>
and CC_model per shell as rc173 with 84 frames rejected, and R_meas was
1.3-1.5x higher in every shell; 9w3y (33 deg rejected -> 0, every model
metric better) and lyso_x10sa_strong read the
same way. The disposition itself is not segmentation-dependent: conviction
is on the batch grid, not the ledger ranges, and those sets' dispositions
changed because the corrected data changed.
R_meas now weights each observation by its merge weight v = 1/sigma^2 under
the error model (corrected_sigma on the host, ModelSigma on the GPU, with the
error model of the last MergeAccum), normalised per reflection to Kish's
effective count, so equal sigmas give the ordinary formula
(WeightedRmeasTerms). Where the proportional term dominates - strong
reflections - frames are weighted alike, as in the merge: 8xtf's lowest shell
reads 15.0%, the rc173 run with 84 frames rejected 15.2%. The per-hand table
is weighted the same way. R_MEAS_UNWEIGHTED / REFRES_R_MEAS_UNWEIGHTED keep
the XDS/AIMLESS convention and are what to set beside XDS; the battery scorer
records them. MULTIPLICITY stays a count.
Effect (weighted / unweighted): lyso_x06da_ref 0.0454 / 0.0479, 8xtf
0.225 / 0.907, insu_I_x06da_ref REFRES 0.072 / 0.232. GPU and CPU paths agree
to the last printed digit on lyso_x06da_ref.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Where the file's and the measured beam centre index lattices whose primitive
volumes differ by an integer factor 2-4, the pooled validation evidence decided
only when the measured centre held the LARGER cell. The other direction was
left to the second pass on the assumption that it drops to the smaller cell
by itself. It does so only when the pass-2 indexer's free beam refinement
walks the whole centre error, and that rescue is chaotic: on 7ris a 0.1 %
change in the pass-1 distance switched off a 6.5 px walk, the doubled c axis
(380 A) won and the run failed.
Both directions are now decided on the same measurement: pooled validation
spots on each lattice against its own wrong-spindle null
(ValidationEvidencePrefers). No new threshold. 7ris: 44.3 % at the file's
centre vs 83.5 % at the measured one; the run adopts the measured centre and
the 190 A cell (R_free 0.81 -> 0.20, 1.52 A). 8pqd (3x, measured centre
holds the true cell): 16.6 % vs 21.0 %, adopted in pass 1 instead of being
recovered by pass 2.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The joint crystal + detector fit was committed only when the held-out residual
fell by at least a fixed 2 %. On a crystal whose geometry is already right
that bound cuts through the noise: two builds read 1.9 % and 2.05 % for the
same 0.03 px / 0.01 % move (lyso_x06da_half_image), one committed and the
other did not, and everything downstream (smoothing window, stretch
segmentation) followed the coin.
The bar is now the residual's own noise - the standard errors of the two
held-out means combined - which is the bar a round of the geometry walk
already has to clear (HeldOutResidualFell). The log line carries the joint
residual and that noise. Noise on the battery is 2-7 % of the residual, so
the gate is about as strict as before but no longer at a fixed edge.
Checked on 57 open/in-house sets and 11 private ones against the pool
battery. The verdict changed on 8 open/in-house sets besides the passes that
follow the 7ris/8pqd lattice change: half_image, 6qaj, lalanine, 8xtf, 9ig7,
8sqt (commit -> reject of a 1-step move) and myob_x06da_split, 6cdl (reject
-> commit). Merge statistics are equal to the last digit everywhere except
half_image (REFRES R_meas 1.16 -> 0.95, CC1/2 0.79 -> 0.86), 9ig7 (5 fewer
rejected observations) and 8xtf, whose stretch disposition followed the
coin: multiplicity 20.8 -> 15.6, R_meas 0.91 -> 0.60, R_free 0.193 -> 0.189,
REFMAC R_free 0.179 -> 0.176. Private arm unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Every lattice decision counts spots or frames, and the sub-lattice's strong
reflections win every count, so a crystal whose cell is doubled by a weak
superstructure class (9min, 6z9g) is adopted at the half cell. This asks the
question in intensities instead. On 60 frames spread over the sweep, after
each frame's own integration, the 2a x 2b x 2c supercell of its primitive
lattice is predicted to 3 A with the frame's own refined orientation and
geometry and integrated on the same engine; the reflections are summed per
parity class in two shells (20-5, 5-3 A), with a fit of intensity against
partiality (I = a + b p) that separates what rocks like a Bragg reflection
from what sits at the node whatever the rocking. Only the tested frames are
predicted and nothing is retained, so memory is bounded (the previous
prototype predicted the whole run through the merge and ran out of GPU
memory). The probe's integrations are kept out of the engine's own counts,
which the two-pass stencil guard reads. Results are bit-identical with and
without it.
REPORT ONLY - it decides nothing, because on the battery it does not yet
separate a weak real class from what sits at the half-integer nodes of
crystals whose cell is right. Real classes: 6z9g class 101 at 24 % of the
lattice's intensity (29 % rocking), 9min 100 at 19 % (4 % rocking - its real
class does not rock like the lattice either). On correct cells the largest
classes reach 9-12 % raw (7n2s, 9i0a, 7os3) and 3-4 % rocking (7dkp, 7os3),
and 7mzt reads 40 % / 19 % on a class the deposition does not have. The log
line is the population a decision has to be calibrated on.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
R_MEAS is what users set beside XDS/AIMLESS, so its meaning does not change; the merge-weighted value
from hq-cure-merge is reported beside it (R_MEAS_WEIGHTED, REFRES_R_MEAS_WEIGHTED). The per-hand
table keeps the ordinary statistic.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The between-pass geometry walk kept a round only when the realised held-out residual fell by more
than its noise, and started only on a fit move of a trust-region step or more. The residual is
dominated by low-resolution reflections, where a distance and the compensating cell scale move every
spot alike, and its centroids are taken inside a disc centred on the prediction, so it barely sees a
distance error that costs the high-resolution shells. On 8pqd the canonical pass ran at 96.456 mm; the
fit asked for 95.878 mm (0.6 %, less than a step, so no walk). Forced, that round read the residual
only 0.74 sigma lower but put 70.4 % of the validation spots on the lattice against 58.1 %.
Now a round is kept when either the residual falls beyond its noise (HeldOutResidualFell) or the
validation evidence prefers it (ValidationEvidencePrefers, z = 3.29). A move of less than a step is
first tried as two index-only probes on the validation frames (fit's geometry vs the one in hand) and
pays for a re-integrated round only where the fit's geometry scores higher; walks started by a large
move run as before.
8pqd: 96.456 -> 95.878 mm, d_min 1.374 -> 1.306 A, R_free .218 -> .206, REFMAC R_free .209 -> .202,
Wilson B 34.4 -> 30.7 (pool had 1.325 A / .209 / .203). 6vww 7n0i 7ris 8egn 8xtf 9qw8 9w3y 6qaj
6cdl 9i0a myob_x06da_split lyso_x06da_half_image unchanged (probes say the geometry in hand stands);
probe cost 1-10 s per run.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The report-only supercell probe now reaches the results report: SUPERCELL_CLASS (parity of the
primitive indices), its occupancy and Bragg-like (rocking) part against the lattice's own reflections,
its <I/sigma>, and SUPERCELL_DOUBLED_CELL, the Niggli-reduced cell to give with -C. No decision is
taken on it: correct cells whose half-integer class is diffuse, and cells whose depositor kept the
sub-cell, cannot be told from a real doubling by the data alone. On a crystal with a real doubled
axis the reported cell matched the deposited one, and processing on it with -C brought R_free
against the deposited model from 0.60 to 0.29.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Advisory only: the class is measured (<I/sigma> >= 0.5) and its rocking part is at least 2% of the
lattice's, three standard errors clear. The warning asks the user to process both settings - the
run's cell and the reported doubled cell with -C - and compare them in refinement. Of 105 battery
rotation sets it names 8: both crystals whose accepted cell is the doubled one, one pseudo-translation
whose doubled description is an accepted alternative, and five correct sub-cells with weak ordered
half-integer intensity.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The magnifier moves from a helper window into its own dock below the
inspector. It has three fixed zoom levels (x64 and x32 with the pixel
values written on the pixels, x10 without) instead of wheel zoom, and
follows the cursor only while Shift is held over the diffraction image.
Meanwhile the diffraction view draws a frame around the area it covers,
just outside that area so it stays visible at low zoom without hiding the
pixels; releasing Shift (seen by the application-wide key filter, so
wherever the focus is), moving without it or leaving the image hides it.
A "Pop out" button moves the magnifier into a separate window and back.
Qt's own dock floating stays disabled: a floated dock is placed
off-screen on WSLg. Whether it was popped out, and the window geometry,
persist across sessions. The inspector's image statistics become a
collapsible section so a small screen can give the space to the
magnifier. kLayoutVersion is bumped for the new dock.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The figures, their value pairs, the excluded sets with reasons, the gnuplot script and the runs'
provenance are written together, so a published figure can be audited and redrawn from the
battery run it came from. Facility and beamline counts come from the entries' _diffrn_source.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Presentation only; the data, counts and audit files are unchanged.
Follows the IUCr artwork guide (journals.iucr.org/services/help/artwork/guide.html)
and the Acta D notes for authors (journals.iucr.org/d/services/notesforauthors.html):
- each figure is a single-column 8.8 cm square (guide: 8.85 cm; notes: 8.8 cm), so a
two-panel composite also fits the 18 cm page width
- lettering 8 pt upright Helvetica, embedded by pdfcairo (guide: ~8 pt, standard fonts
Arial/Courier/Helvetica/Symbol/Times, fonts embedded; notes: 1.5-3 mm lettering)
- line weights 0.75 pt for border, ticks and the y = x line (guide: 0.35-1.5 pt;
pdfcairo lw 1 is 0.5 pt)
- the PNG is the PDF rasterised with pdftoppm at 600 d.p.i. (guide: 400 d.p.i. colour,
600 d.p.i. line art), so the two files are the same figure
- no grid (notes: grids avoided where not required), one Okabe-Ito blue for the points,
which stays distinct from the black dashed line in greyscale and to colour-blind readers
- ticks short, outward and not mirrored; d_min ticks 0.5/1/2/4/8 A all at one decimal
- italic R and d with roman subscripts; the REFMAC qualifier moves to the caption
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Owner's choice for the 8.8 cm single-column figures; still inside the IUCr notes'
1.5-3 mm lettering range (10 pt Helvetica capitals are ~2.5 mm). Points enlarged
slightly to match. Presentation only.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Protocol changes in tools/battery/model_check.py (the --model-check R factors):
- Model preparation: atoms of unknown element (UNX, element X) are left out
(REFMAC stops on them: "Atom does not seem to be an atom : X"); chain names
longer than two characters are shortened so the model is written in PDB
format, whose ANISOU records REFMAC reads (its mmCIF ANISOU reader failed with
"rdaniso_cif: Atom symbol mismatch" on two entries). A model too large for the
PDB format is written as mmCIF with isotropic B only. What was changed is
recorded in model_modification.
- Depositor's structure factors: amplitudes from F_meas, else the mean of
F(+)/F(-); where only intensities were deposited (I, else the mean of
I(+)/I(-)) they are converted by French-Wilson (ctruncate) instead of being
skipped. Recorded in depdata_kind. A numeric pdbx_r_free_flag with more than
two values is read with the CCP4 convention (0 = free); status 'f' is still
preferred when present.
- The depositor's data go through the same change-of-basis choice as Rugnux's
(reindex_op_depdata): a twinned entry can be deposited in the other branch
of a merohedral ambiguity relative to its model.
- An entry that declares twinning (_pdbx_reflns_twin, more than one domain) is
scored with REFMAC's twin refinement for both data sets; the untwinned
numbers are kept under "untwinned" (battery rows: refmac_twin,
refmac_untwinned).
- The work directory is made absolute (REFMAC runs inside it; a relative
--workdir failed).
Unchanged: rigid-body first-cycle R factors, d_min_used = max(Rugnux d_min,
deposited d_min), the field names refmac_rfree / refmac_rfree_depflags /
refmac_rfree_depdata / refmac_rfree_ratio / refmac_reason.
Re-run on the existing p.mtz of all 165 open-arm PDB sets of the last full
battery: 131 give identical numbers; 3 former REFMAC failures now score; 23
intensity-only or F(+)/F(-)-only depositions gain the depositor baseline; 7
twin-declared entries move to twinned R; one untwinned P3x entry whose model is
near-symmetric under the merohedral twofold picks the other (tied) branch for
the depositor's data (R-free 0.2928 -> 0.2904).
README: which R-free is which, the common resolution limit, and the twin and
French-Wilson handling.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Third figure: rugnux's own WALL_TIME per dataset (rugnux_wall_s, not the runner's elapsed_s,
which adds the wait for a GPU slot under --gpulock) as a 25 s linear histogram with the median
marked and the open/in-house counts. Every set that ran to a report is counted; crashed sets and
the no-crystal controls (no report, so no WALL_TIME) go to excluded.txt. Sets that shared the GPU
(gpu_others > 0) are drawn as the lighter top segment of their bar rather than dropped, and
--uncontended RUN sizes their slowdown against a run with the GPU to itself, reported in
summary.txt with the percentiles, the arm medians and the page-cache caveat. time.dat carries
set, arm, seconds, images and gpu_others beside the figure.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The time histogram now covers the open arm, like the other two figures, and can come from a run
whose sets had the GPU to themselves while the quality figures use the latest run; summary.txt
lists the provenance of both. The median label sits above the bars.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The files were written on the axes the space-group search named, which for
P2221/P21212 puts the unique axis wherever the a<b<c indexing put it (a 52.51
87.87 137.72 crystal came out P 2 21 21 where XDS writes 87.87 137.72 52.51
P 21 21 2), and a reference MTZ was matched in the data's frame: on permuted
axes its free-R flags landed on unrelated reflections while the log reported a
high matched count.
- New CrystalSetting (scale_merge): changes of basis between settings of one
lattice (cell, group, index operator, basis matrix), the {-1,0,1} det +1
candidates (CellMappingOperators, moved from ModelValidation), MetricViolation
(moved from Rugnux), ChooseOutputSetting and SeatGroupByAbsences.
- Output setting: after every decision the merge, the integrated reflections
(unmerged MTZ), the P1 cross-check, the lattice and the _process.h5 reindex
matrix are relabelled into the ITA standard setting - or, in priority order,
a reference MTZ's, a fitting model's, the -C axis order, a non-standard -S
symbol's. Free-R flags are drawn again on the written axes. Reported as
SETTING_OPERATOR / SETTING_SOURCE.
- Reference MTZ: the group is kept in its setting; after the merge every cell
mapping onto the reference cell (times the twin laws) is scored by the
reference CC, the best is re-seated and re-merged, and the free flags are
inherited only where CC >= 0.5 over >= 50% of the reference range
(REFERENCE_MISMATCH otherwise; --mode scale gates the same way).
REFERENCE_OPERATOR / _CC / _MATCHED_FRACTION / _FREE_FLAGS_INHERITED.
- -S: a fixed group is put on the axes its absences name before merging
(SeatGroupByAbsences), fixing -S 18 on a cell whose pure axis is not c.
- --model: a model in another setting is now a claim the null tests; where it
fits, the data are written in its setting and the validation is remade on
those axes (KeepModelVerdict carries the decisions over).
Tests: [setting] (synthetic #18/#17/I222/C222/c-unique P21/C2 beta/I2->C2/P1,
-C and -S order, absence seating, permuted reference with flags).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
RUGNUX_ADVANCED (new section "The setting the files are written in", -z and
--model paragraphs), RUGNUX_REPORT (SETTING_* keys, REFERENCE_MISMATCH, model
change-of-basis keys), CPU_DATA_ANALYSIS_DECISIONS / _INTEGRATION, CHANGELOG
entries under rc.173.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
-z (now also --reference; --reference-mtz still works) takes an SF-mmCIF
(e.g. a deposited -sf.cif, gzipped or not) as well as an MTZ. The format is
recognised by content: a file starting with "MTZ " is an MTZ, anything else
is parsed as CIF. The first merged reflection block with the requested
column (or, by default, one the auto choice accepts) is converted to a
gemmi::Mtz in memory with GEMMI's CifToMtz and then read by the unchanged
MTZ loader, so the in-memory reference is exactly what the MTZ path yields.
Unmerged (_diffrn_refln) and anomalous-only blocks are passed over; the log
names the block used.
The R-free set comes from _refln.status (f -> FreeR_flag 0, o -> 1: the
CCP4 convention the loader already reads) and is preferred to
_refln.pdbx_r_free_flag, whose convention varies by program; a status
column with no 'f' is ignored and pdbx_r_free_flag is used instead.
Checked on two open-arm sets in the deposited setting (one with F_meas_au
only, one with intensity_meas): reference loaded with the deposited cell and
group, the inherited free set agrees with the deposited status 'f' on every
common reflection, and the merge correlates at CC 0.99 with the deposition.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
`battery.py recheck RUN [--only a,b] [--jobs N]` reruns model_check.py
(REFMAC) and depdata_check.py on every open-arm row of a finished run from
its own p.mtz, and updates the refmac_* / dep_* fields of results.json in
place (verdicts untouched). Previous results.json, report and each set's
model_check directory are kept with a .pre-recheck-<stamp> suffix; the
manifest gains a "rechecks" entry {date, end, runner_git, runner_dirty, sets,
fields, previous_results, command} so the tools version behind each published
number is auditable. The report is re-rendered and the run made read-only
again.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Adds a second REFMAC protocol to model_check.py beside the first-cycle one
(whose numbers and field names are unchanged):
- the same deposited model is refined against Rugnux's data and against the
depositor's data by the same protocol - restrained, 10 cycles, automatic
weight, isotropic B for every atom, twin refinement where the entry declares
twinning - each data set over its OWN resolution range;
- R-free on a SHARED free set: the depositor's free reflections present in
both data sets after the change of basis (so the lower limit in every
direction). Depositor free reflections outside the intersection are removed
from both; every other reflection is a work reflection of its data set;
- R-free recomputed from each output MTZ over exactly that set (F vs
FC_ALL_LS); for a twinned refinement REFMAC's own R-free (its free set is
the shared set), as the output holds no twinned Fc. REFMAC's own R-free is
kept for reference.
Isotropic B: refining deposited (often TLS-derived) ANISOU atom by atom
diverged (-LL rising every cycle at 1.4 A on a model deposited at 1.7 A).
A ligand whose code REFMAC's monomer library describes with other atoms
("atom ... is absent in the library", a stopped refinement) is renamed to a
code the library lacks, so REFMAC restrains it from the model's coordinates.
Battery row fields: refmac_refined_rfree_shared,
refmac_refined_rfree_shared_depdata, refmac_refined_ratio,
refmac_refined_rwork, refmac_refined_rwork_depdata,
refmac_refined_rfree_refmac, refmac_refined_rfree_refmac_depdata,
refmac_shared_free_n, refmac_refined_reason. model_check.py --no-refine
skips the protocol. README documents both protocols.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The R-free figure now draws refmac_refined_rfree_shared(_depdata): the deposited model refined
against each data set at its own resolution, R-free over the depositor's free reflections both data
sets share. --rfree first-cycle keeps the previous first-cycle rigid-body pair at the common
resolution. With the rechecked hq-pool3 runs the figure covers 160 structures (was 135).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Rotation, report only: once the first pass has its lattice, the spots it leaves over (ice left
out) are indexed again with a fresh rotation indexer, up to three times. A lattice is kept when
the leftover validation spots sit on it far more often than at a displaced spindle angle (z >= 5,
the first pass's own null), and classified against the main lattice: a misoriented domain of the
same cell (split crystal), one turned by 180 deg (non-merohedral twin domain), one concentrated
in part of the sweep (>= 60% of its spots in 2 of 8 blocks), a related cell (volume n or 1/n) or
an unrelated one. Misorientation is the minimum over the bases whose metric matches, so a lattice
symmetry is not read as one. EXTRA_LATTICE_* keys, a summary row, and a MULTIPLE_LATTICES warning
where domains or an unrelated lattice hold >= 10% of the non-ice spot intensity (29 of 200
battery sets in the exploration census). Its own IndexAndRefine instances and indexer: merged
MTZ byte-identical on three sets; 0.2-2.5 s per run.
Also MOSAICITY_DEG (median, P10, P90 over frames) - the frame-order-smoothed sigma_M the merge
computes partiality from, Kabsch's sigma_M as XDS's REFLECTING_RANGE_E.S.D. - in the report.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
MULTIPLE_LATTICES and EXTRA_LATTICE_INTENSITY_PCT now count DOMAIN, TWIN_DOMAIN and SEGMENTED
lattices only. A FOREIGN lattice (an unrelated cell) is still listed but no longer warned about: it
is as often a wrong main lattice or tNCS as a second crystal (one of the four such sets in the
battery census is a tNCS crystal). On the census this names 25 of 200 rotation sets instead of 29.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The lattice walk admits an angle up to 3 deg from 90 and the constrained refinement then holds it
there. The existing two-arm test (promoted vs demoted class, judged on the held-out residual) only
covered length equalities (a = b). It now also covers monoclinic/orthorhombic classes: pass 1 reads
how far the free refinement leaves a held angle from 90 (AngleEqualityDeparture_deg), turns it into
pixels at the far corner, and where that exceeds the integration disc runs the demoted arm on the
indexer's free (triclinic) cell. A primitive monoclinic crystal with beta 1.4 deg from 90 was
indexed and integrated as primitive orthorhombic on 60/60 validation frames, so the per-frame guard
never fired; demoted, the held-out residual halves (1.54e-6 -> 8.16e-7), ISa 5.5 -> 6.1, d_min
3.70 -> 3.37 A, CC1/2 at 3.9 A 0.54 -> 0.76. Unchanged on five controls (three oP/oI, two near-90
monoclinic).
Model validation scored alternative cell frames against the data in the merged indexing, before
the indexing-ambiguity probe. On a pseudo-orthorhombic P2 crystal whose data need h,-k,-l to match
the model, every frame read random (R 0.77-0.84) and noise picked an axis swap into P1121; each
frame is now scored at its best over the data's alternative indexings. CC_model 0.23 -> 0.45,
R-free 0.66 -> 0.44 on that set.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
A merged reflection's d came from whichever frame first observed it, while the possible
reflections per shell are enumerated from the reference cell, so reflections near a shell edge
fell on the other side from their possible twin and shells read 100.1-100.9 % complete (17 of
210 battery sets). The group's d is now computed from the reference cell before the export, and
the per-hand table reads the group's d instead of an observation's.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Compared with ctruncate (CCP4 9, version 1.17.29) on rugnux's own merged
intensities of the open battery arm, three differences:
- ctruncate gives no amplitude to an intensity below -3.7 sigma (exactly
that bound on every set checked; up to 2498 reflections on one set).
Rugnux turned them into small, confident amplitudes (F/Fc ~0.3 on the
sets whose background is over-subtracted on powder/ice rings). They now
get F = NaN (missing in MTZ, '?' in mmCIF); IMEAN is kept. They are
also left out of the shell mean that sets the Wilson prior.
- The switch to sqrt(I) at I/sigma = 4 left 4-6 sigma amplitudes 2-3%
above ctruncate's on every set (1.019-1.028). The posterior now applies
up to 20 sigma (emulated: 0.995-1.000).
- A shell whose mean intensity is not positive gave a prior at the 1e-10
clamp and amplitudes of ~0 (one set's outer shell); it now takes the
nearest lower-resolution shell's mean.
The remaining gap (weak amplitudes 2-5% below ctruncate's in the outer
shells) is ctruncate's anisotropy-corrected prior; not attempted. An
offline R-free ablation put the whole ctruncate conversion at -0.0006
median (15/18 sets better) and dropping its rejected negatives at -0.0009
mean (-0.009 on the worst set).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
A new correction surface, fitted after the time x detector surface and
before the goniometer-frame 8x8 grid: log A is a sum of real spherical
harmonics (l = 1..6, 48 terms) of the diffracted-beam direction de-rotated
into the crystal frame. The incident-beam path depends on phi alone and is
in the per-frame scale already.
It runs through ApplyCellSurface unchanged in everything but the update:
the cells are 32 x 64 equal-solid-angle direction bins, and each round the
per-cell sums (ref2, cross, the same damping) become one ridge-regularised
Gauss-Newton step on the coefficients (prior width 0.1/l per degree-l
coefficient) instead of independent per-cell steps. The half-set Fisher-z
gate adopts or refuses it exactly as it does the grids; where it is
refused, the grid after it sees what it saw before.
Why: the folded 8x8 grid (hemispheres share a cell) is the weak basis for
long-wavelength absorption. Offline, held out by unique reflection, this
basis lowered held-out scatter 4-8% on 6 of 11 long-wavelength sets where
no cell grid did, raised model-phased anomalous peaks 0.02-0.2 sigma, and
was neutral on hard-X-ray controls.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The local radial background at ring radii (--background-radial) already
existed and is gated per image on the smooth-ice channel of the ice score,
but it was off unless asked for. Systematically negative merged
intensities at ring radii (I/sigma < -4 on 0.4-1.6% of reflections of the
affected open sets) come from the flat annulus mean over-subtracting a
sharp ring. rugnux now runs with =auto unless the flag says otherwise; the
library default (and so the broker) stays off.
Measured on the open arm against rc173 (commit 727bd0d6d as reference):
auto: 8agq (smooth ice, fires) CC_model .9327->.9476, R shell-scaled
.1987->.1795, I/sig<-4 1.55%->1.15%; 8rud R shell-scaled
.2507->.2456, ISa 10.0->13.3; 5reo, 7kcn, 9fhc bit-identical
(gate does not fire); wall time unchanged.
forced on (not adopted): same 8agq gain and 8rud .2414, 9fhc/8qq7
small gains, but 7kcn (no ice) CC_model .9491->.9411, ISa
11.0->7.5 - which is why the gate stays.
ISa falls on 8agq (15.6->13.3) while agreement with the model rises: the
merge-statistics vs external-model trade the settings comment describes.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
After the existing mask is complete, a second step takes the pixels it left out that are
significantly dimmer than their ring, joins them through a 6 px bridge and across module gaps of
any width, and adds a piece whole when it holds >= 2000 dim pixels of which >= 200 are deep. The
existing mask is never touched, so a sweep with no such piece keeps its mask bit for bit.
Quick subset tests (battery --only, no model check) against the rc173 a0518abe6 full battery:
152 of 233 masks identical; 10 of 10 controls (including the sets where earlier shadow changes
regressed through marginal decisions) give byte-identical merges. On the 21 sets that gain a
piece the space group never changes, merge outlier rejections fall on 19, R_meas falls and ISa
rises on 17, and the shell-scaled model R improves on 16 of 18; a transmitting arm with an
over-subtracted background strip is recovered (R_meas 13.3 -> 12.0 %, ISa 13.4 -> 15.5).
Known costs: on one set with background bumps at ring radii the low-resolution agreement with
the model falls (CC 0.89 -> 0.85) while its internal statistics improve, and two sets cut
slightly coarser (1.50 -> 1.55 A, 1.72 -> 1.80 A) with a better model R.
Squashed from branch hq-beamstop (cfa080a59..60b6c6822), whose history records the variants
tried and dropped: a straight port of the earlier arm finder moved every mask, a two-shape
construction filled almost a whole sweep through hole filling, and a deep-fraction floor
rejected a real arm that is dim along its whole length.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Report-layer fixes from the warnings audit of the rc173 full battery. Processing is
unchanged (merged MTZ/HKL bit-identical on the sets checked).
- SUPERCELL_POSSIBLE stays a key and a summary line; a warning only where the rocking part
reaches 20% (sub-cells that refine normally rocked at 2.1-16.2%, the accepted doubled cells
at 4.2% and 29.1%). The summary says where further lattice domains may put spots on the
half-integer nodes.
- LATTICE_TRANSLATION only where the Patterson at the vector reaches 90% of the origin
(UNDECLARED_LATTICE_TRANSLATION_PCT, new); the two admitted vectors on the battery read 79%
and 85% on crystals refining in the deposited cell, and are now reported as a very strong
pseudo-translation. PSEUDO_TRANSLATION warns from a 20% peak (xtriage convention; detected
peaks were 8.1% on one set, 24.8-63.2% on the rest).
- <|L|> outside 0.375-0.55 (on the adopted or the pre-search merge) is TWINNING_VERDICT=
NOT_READABLE with no warning, instead of a twin or SYMMETRY_SUSPECT reading.
- NO_LATTICE no longer fires on a rotation run that indexed the sweep with no frame indexed
on its own; INDEXING_RATE is --developer on rotation.
- A rotation run that finds no lattice writes a report (VERDICT= FAILED, NO_LATTICE) before
exiting 1, instead of an input-parameter error and no report (NoLatticeFound).
- Powder and Ice summary lines, ICE_* keys, and a POWDER_RINGS flag from a 0.5 spot fraction.
- SWEEP_GAPS warns only where the degraded ranges cover 10% of the sweep (77 -> 48 sets).
- A single rotation sweep no longer carries an INDEXING_AMBIGUITY warning (one orientation
matrix, consistent hand); INDEXING_AMBIGUITY_OPERATORS and an Indexing choice line instead.
- A DOMAIN more than 10 deg away is a second crystal; the MULTIPLE_LATTICES sum is documented.
- Mosaicity line and docs name XDS's suggested REFLECTING_RANGE_E.S.D. as the comparable
number (0.90x median over 34 in-house sets; per-image SIGMAR runs ~1.35x higher).
- HARMONIC_CONTAMINATION, SCALING_NOT_CONVERGED, POWDER_RINGS in the documented vocabulary.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Owner decision: a warning is a prompt to check and must catch the real cases (9min) at the
cost of some spurious ones. The physically motivated corrections stay (<|L|> outside its
physical range is not twinning; a single sweep's indexing choice; NO_LATTICE on rotation;
the no-crystal report); thresholds raised only to cut noise come back down:
- SUPERCELL_POSSIBLE warns wherever the class measures and rocks (as before rc173's audit
fix), worded as "check the cell", naming weak ordered intensity of a correct cell and spots
of further lattice domains as the other readings. 9min (rock 4.2%) warns again.
- LATTICE_TRANSLATION warns on every admitted vector (>=75% of the origin); below 90% the
wording names a very strong pseudo-translation as the other reading.
- PSEUDO_TRANSLATION warns on every detection; below a 20% peak it is worded as weak, check.
- SWEEP_GAPS warns where the degraded ranges cover at least 1% of the sweep (the 4 sets of 77
below that had 1-2 frames, 0.4-0.6% of the sweep).
- Powder rings are split between hexagonal-ice positions and the rest (MeasurePowderRings,
report-only fields); ICE_RINGS and POWDER_RINGS warn separately from 5% of the spots, and
ICE_RINGS also where the merge's ice gate found ice.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The local-box SNR kernel (analyze_pixel) was 55% of all GPU kernel time on a
rotation run and held the image loop GPU-bound. Three exact changes:
- analyze_pixel is rewritten as one warp per 32 output columns, each lane
holding the vertical sums of two input columns in registers and the 31-wide
horizontal window taken from warp prefix scans. No shared memory and no
block-wide synchronisation; the sums are modular 64-bit integers, so the
result is the same bits as before.
- The adaptive finder keeps only (local-test & ring threshold), and a pixel's
second-pass result depends only on its own window, so the second pass is
evaluated at the ring pixels alone (analyze_candidates) instead of densely
followed by and_bits.
- The first pass is then only read within NBX of a ring pixel, so a warp whose
tile no ring pixel can reach skips it; it also no longer reads an all-zero
previous-pass buffer.
Checked bit for bit against the old kernel on every frame of a 16M rotation
run, and p.hkl / p_unmerged.mtz md5-identical on two inhouse EIGER2 16M sets.
Total kernel time 19.8 -> 12.1 s, image loop 2.69 -> 1.37 ms/image; wall
44.0 -> 37.2 s and 56.0 -> 48.7 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
RunImpl allocated the reflection mask (1 byte/pixel) and the owner map (4
bytes/pixel) afresh for every image - 90 MB at 18 Mpixel, above the malloc
mmap threshold, so each call paid an mmap, a zero-fill page fault per 4 kB and
an munmap (with its TLB shootdown across the other workers). On a CPU-only
rotation run of a 16M set that was 79 M page faults and 794 s of system time.
Both are now members of the engine, and each call clears the rectangles the
previous one wrote before it starts. Contents at every read are unchanged, so
the output is byte-identical; CPU-only wall 297 -> 262 s, system time
794 -> 214 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The adaptive finder keeps (local test & ring threshold), so the local test's
second pass is only wanted at the ring pixels. Its window there is the first
pass's window with the first pass's strong pixels taken out, and those are
few: the first pass now records its window sums at the ring pixels, and the
second pass subtracts the strong pixels in each window instead of sweeping
the whole image again. Integer sums and the same acceptance test (factored
into StrongInWindow), so the bits are exactly the dense pass's.
CPU-only rotation run of a 16M set: md5-identical output, 262 -> 242 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The geometry walk after the canonical pass starts with an indexing probe at
the geometry in hand (probe_at_best), which is the canonical pass's own first
pass run over again - a full first-pass indexing whose answer is already
known (identical validation-spot counts in the logs).
The first pass is deterministic in its inputs, so each canonical or probe
first pass now leaves its validation evidence behind keyed by what it read
(geometry, goniometer, spot budget, cell, indexing and spot-finding settings,
and the run state it consults), and an indexing-only probe with the same key
takes it. Nothing is stored for the geometry pre-pass, for a pass that
changed its own inputs on the way (a rescue adopted), or for one that ran the
beam-centre ladder, which a probe never runs. RUGNUX_VERIFY_FIRST_PASS_MEMO
runs the probe anyway and fails the run on any difference.
md5-identical output; 37.2 -> 34.0 s on a 16M rotation set with a geometry
walk (one probe fewer).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The short-axis pass (the low-FFT-floor hypothesis for small-molecule cells)
ran after the standard first-pass indexing and its rescues, and it is itself
a full indexing of both schemes - about a second of mostly serial refinement,
paid in every first pass, probes included.
It is now started as soon as the standard schemes are fed, on a copy of the
experiment, and runs alongside them. At its own place it is taken only if what
it read is still what the run has there - the experiment and spot-finding
settings (ExperimentKey, split out of FirstPassInputKey), the spot list's
mapping and the per-image spot budget; a rescue that moved any of them makes
it run again there as before. Same inputs through the same code, so the
result is the same.
md5-identical output on two 16M rotation sets; 34.0 -> 29.5 s and
48.2 -> 47.1 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The Rugnux library and the rugnux executable differ only in case, so on a
case-insensitive filesystem (macOS, Windows) their CMakeFiles/<target>.dir
directories collide and the executable's build.make overwrites the library's
("No rule to make target rugnux/CMakeFiles/Rugnux.dir/depend"). The library is
renamed JFJochRugnux, in line with the other libraries.
Apple's linker rejects the TLS wrapper clang emits for the inline thread_local
static member WorkerPool::in_worker as a duplicate symbol once it is included
from several libraries. It is now a function-local thread_local.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The viewer's light (salmon/white) theme was built on the system palette, so
under a dark macOS theme the text roles stayed white on light backgrounds; with
Fusion's standard palette macOS still supplies grey text. Start from Fusion's
standard palette and set window, entry, button and tooltip text to black for
the active and inactive groups, as on Linux and Windows. Disabled text keeps
Fusion's grey.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
View > Theme offers Follow system / Light / Dark, stored in the settings.
Follow system uses the desktop's colour scheme where Qt reports it (6.5+)
and is light otherwise. The dark theme is a deep indigo (#1A1D3A panels,
#12142B entry fields and charts).
The palettes live in the new ViewerTheme module. A switch applies at once:
the application stylesheet is set again after the palette so every widget
is re-polished (with a stylesheet active, polished widgets otherwise keep
their old palette); widgets that bake colours into stylesheets use
SetThemedStyleSheet, which re-applies them on ThemeNotifier::changed;
toolbar icons draw their glyph at paint time through an icon engine, so the
ink follows the theme (and stays sharp at any size); charts re-theme and
rebuild. Hard-coded white backgrounds on entry fields and tables are
removed in favour of the palette's base colour.
The idle detector-status badge in the status bar is now transparent instead
of an empty default-styled progress bar.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The right end of the display toolbar now carries three "A" buttons for the
100/125/150 % font sizes and a light/dark switch (a sun with a half-filled
disc), so both settings are visible without opening the View menu. The size
buttons and View > Font size stay in sync through JFJochViewerMenu's new
SetFontZoom / fontZoomChanged; the View > Theme tick follows ThemeNotifier,
so it moves when the toolbar switches the theme.
With Qt 6.8+ an explicit Light or Dark choice is also requested from the OS
(QStyleHints::setColorScheme), so the title bar macOS and Windows draw
matches the window, as apps with their own theme switch do. Follow system
withdraws the request first, since Qt otherwise reports the requested scheme
instead of the desktop's.
The image counter reads "1 / 1800" instead of "1/1800".
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The inspector's width bounds were 75-100 average character widths, which on
macOS (8 px) came to 600 px for content that needs 397. They are now
measured: at least the panel's minimum size hint plus the scroll bar, at
most a third more. CollapsibleSection reports its folded content's width
in minimumSizeHint, so the panel keeps its width when a section opens. On a
font change the measurement waits a turn, until the children have the new
font.
When shrinking the window would leave the image narrower than the
inspector, the inspector folds away, with the magnifier if it shares the
inspector's column; they come back once the image would still be at least
as wide as the inspector (plus 40 px of hysteresis). Only a shrink folds, a
dock shown by anyone else is no longer the window's to restore, and folded
docks are saved as open. Measured on a 1920 px screen: folds at 1140 px,
returns at 1260 px; at 960 px the image gets 620 px instead of ~145.
Dark theme: disabled text is set explicitly (#858AAE, ~4.8:1 on the panels)
instead of the derived colour, which was nearly the panel colour and made
greyed-out fields look empty. The splash text is black in either theme:
its message is rich text, which takes the palette's text colour rather
than showMessage's.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Where a pseudo-translation is detected, the class vector is re-refined up a
ladder on the whole data by a greedy sweep over the 26 neighbours of the
current vector, each a cosine correlation over up to every acentric
reflection - seconds of one thread on the main path of the run (3.4 s and
1.5 s for the two calls on a P3_121 16M set).
The neighbours of the current vector are now evaluated together; the first
improvement in the sweep's own order is taken, and the neighbours after it
are evaluated again about the new vector - exactly the moves the serial
sweep makes. Each correlation is still one serial sum, so the result is
bit-identical. AnalyzeTranslationalNCS takes the thread count.
md5-identical output and report; 47.1 -> 44.4 s on that set.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Two whole-dataset sorts sat on the main thread at the end of a rotation run:
WilsonOutliers orders every full by resolution, and the unmerged MTZ is put
in H K L M/ISYM BATCH order by Mtz::sort(5) - together about 2 s of one
thread on a 1.4 M-observation set.
ParallelSort (common/ParallelFor.h) sorts one piece per worker and merges
them pairwise. It is only for comparators that are a strict total order,
where the sorted sequence is unique and the result is the serial sort's bit
for bit: WilsonOutliers already breaks ties on the index, and the unmerged
writer now sorts the rows itself on the five key columns and then the row
number - the order Mtz::sort's stable sort gives - and sets sort_order as it
did. An empty table still fails the way Mtz::sort does.
md5-identical p.hkl, p.mtz and p_unmerged.mtz; 44.4 -> 43.6 s on a 16M set.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The .dmg cpack produced did not start: jfjoch_viewer links Qt as
@rpath/Qt*.framework, install strips the build tree's rpath, and no install
rpath was set, so dyld found none ("no LC_RPATH's found"); macdeployqt's
"Cannot resolve rpath" errors were the same gap. INSTALL_RPATH is now
@executable_path/../Frameworks. Info.plist had an empty identifier and
version; it now carries ch.psi.jfjoch.viewer, the numeric version (the rc
suffix in the long version string) and the name "JFJoch Viewer".
On macOS the notices go into jfjoch_viewer.app/Contents/Resources instead of
a share/ folder beside the app in the .dmg, which a user dragging the app
out would leave behind. They are installed before the Qt deploy script,
which signs the bundle - a file added afterwards invalidates the signature.
JFJOCH_NOTICE_FILES is set before the subdirectories for that.
Names: jfjoch-viewer-<version>-macos-<arch>.dmg (volume "Jungfraujoch
Viewer <version>"), and a rugnux-only build on macOS packs as
rugnux-<version>-macos-<arch>-cpu.tar.gz instead of claiming linux.
CI: build:macos:viewer and build:rugnux:macos on a host-mode runner
labelled macos-arm64. Each checks its artifact (the bundle finds its own
Qt, holds no path into the build machine and is validly signed; rugnux
links only system libraries and starts) and uploads it on a release build.
The Developer ID signing / notarytool / stapler steps are left as a comment
until the project has an Apple Developer account.
Verified locally: the DMG mounts, the app starts from it and loads every Qt
framework and plugin from inside the bundle; the rugnux tarball runs.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CMAKE_OSX_DEPLOYMENT_TARGET 12.0 -> 13.0: the Qt 6.11 frameworks the .dmg
bundles are built for macOS 13, so the app could not start on 12 anyway, and
the linker warned about every framework ("building for macOS-12.0, but
linking with dylib ... built for newer version 13.0").
CI job names follow one scheme - build:viewer|rugnux|jfjoch, then the
platform (<os>-<arch> for the portable products, the distribution for the
packages), then cuda / nocuda (cuda-sls9 for the slsDetectorPackage 9
builds). The Linux viewer matrix's "cpu" variant is now "nocuda" like the
Windows one; the RPM/DEB matrix gains platform and variant beside the
distro key its upload step tests. Only display names change: job ids, and
so needs:, and artifact file names are as before.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The rotation two-pass runs several first-pass indexing probes one after the
other although some of them do not depend on each other:
- The rotation-scale walk always indexes at the stored angles and then at
the pass-1 fit; both are known before it starts. The probe at the fit now
runs at the same time, on a copy of the run as it stands before either.
Its answer is taken only if the probe at the stored angles left behind
none of the state the next pass reads (the spot-finding settings a
first-pass ladder rung adopts; each probe puts the experiment back
itself) - otherwise it is run again in sequence, as before.
- The geometry walk's indexing probe at the canonical pass's post-refined
geometry depends on nothing that pass does after its post-refinement. It
is now started there, on a copy, and runs beside the pass's scaling, merge
and output; its evidence comes back through the first-pass memo (now a
short list) and the walk's own probe takes it through the usual key check,
so RUGNUX_VERIFY_FIRST_PASS_MEMO covers it too. It is started only for the
passes whose post-refinement the walk probes with an indexing pass.
The copies are plain copies of Rugnux: the cancel flag becomes a shared
flag (so cancelling the run cancels them) and everything else was already
copyable. Probe passes run on a copy get no observer.
md5-identical output on two 16M sets; the walk's two probes take ~2.3 s
together instead of ~3 s, and the last probe of the geometry walk is hidden
behind the merge.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
- set_gpu_blocking_sync(): every device is put in
cudaDeviceScheduleBlockingSync before its context exists, so a host thread
waiting on the GPU sleeps instead of spinning on a core. On a 16M rotation
run a fifth of all CPU time was that spinning; wall time unchanged within
noise. Called first thing in rugnux.
- enable_gpu_numa_binding(): from then on pin_gpu() (and the new
pin_gpu(dev), used by the first-pass spot workers that take a card by
index) also keeps the thread on the CPUs of the NUMA node the card hangs
off. The node and its CPUs come from /sys (no libnuma), intersected with
the process's own mask; Linux only, and nothing happens on a machine with a
single node. rugnux turns it on; the broker does not.
- A thread inherits its creator's affinity, so the shared ParallelFor pool
would run every later pass on one socket if a pinned worker created it:
its threads now reset to the mask the process started with
(common/ThreadAffinity).
Byte-identical output. The NUMA part is a no-op on the single-node test box
and still has to be measured on a two-socket machine.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
RELEASE_CONTENTS gains the macOS .dmg and rugnux .tgz in the artefact
table, the CPU floor (any Apple Silicon Mac), the OS floor (macOS 13), the
CPU-only note in the CUDA/GPU tables, and a macOS section: Apple Silicon
only (Rosetta does not run arm64 code on Intel), drag-to-Applications,
notices inside the bundle, and how to open the not-yet-notarized release
(Open Anyway on macOS 15+, Control-click Open on 13/14, or xattr).
JFJOCH_VIEWER states the platforms and requirements, that the Mac build is
CPU-only, that D-Bus is Linux-only, and adds Building from source on macOS.
RUGNUX_INSTALL adds the macOS archive and the quarantine note - checked: a
browser-downloaded .tgz hands its quarantine flag to everything tar
extracts, Gatekeeper rejects rugnux, and xattr -dr clears it. DEPLOYMENT
points to the pre-built Windows/macOS viewers and the macOS rugnux archive.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
FFTW's CMake build knows only SSE/SSE2/AVX/AVX2; NEON exists only in its
autotools build (--enable-neon). The ENABLE_NEON this file set was silently
ignored, so on Apple Silicon and the Linux aarch64 build FFTW ran its scalar
codelets (config.h: HAVE_NEON undefined, no simd/neon objects). The NEON
sources are now added to fftw3f the way its CMakeLists adds the SSE2/AVX
ones, with HAVE_NEON as a target definition (the config.h template carries
it only as a comment). No compiler flag: NEON is part of every aarch64 CPU.
FFTW 3.3.11 would not help - its CMake build has no NEON either, and its
Apple ARM cycle counter only matters to FFTW_MEASURE planning, while every
plan here is FFTW_ESTIMATE.
FFTW's runtime NEON probe on unix/linux executes .long 0xf2000150; on
aarch64 that decodes as ands x16, x10, #0x100000001 and runs without
SIGILL (checked on Apple Silicon), so the codelets are used on Linux too.
Measured on Apple Silicon against the scalar build: batched 1D r2c 1.8x
(power-of-two length) and 1.3x (730), 2D r2c 1024^2 1.9x, 3D c2r 128^3
2.6x, spectra equal to 1.6e-8 relative. rugnux -X FFTW on 360 frames of
the rotation test set: merged .hkl byte-identical, report identical; CPU
time unchanged (300 s), so FFTs are a small share of that run.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
PointSurfaces made three forward and four inverse transforms of the padded
detector (8748 x 8400 on a 16M) one after the other on the main thread -
~3.5 s of one core in a CPU-only run. The forwards do not depend on each
other, nor do the inverse products; each group now runs at the same time,
one workspace per transform. A workspace is built as the single one was -
its own std::vector buffers and its own FFTW_ESTIMATE plan for them - so the
same plan runs on the same data and the result is bit-identical. Costs about
2 GB more transient memory on a 16M detector.
CPU-only 16M rotation run: md5-identical, 216 -> 210 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
A rotation run asks the rotation indexer the same question many times: the
canonical pass's first pass repeats the rotation-scale probe at the stored
angles exactly (identical validation evidence on every set checked), and on
a crystal that does not index, pass 2 and every probe repeat pass 1's rescue
ladder rung for rung. Each such RunIndexing is an FFT search plus a serial
Ceres fixed-point chain, ~1-3 s on the GPU build.
RunIndexing is deterministic in its inputs, so its outcome - every member it
sets - is now kept process-wide under a key of all of them: the accumulated
spots (every field) and their angles, both geometries and the axis, the
experiment's indexing settings, cell and space group, and the settings of the
pool it indexes with (IndexerThreadPool::Settings). A RotationIndexer asking
with the same key takes the outcome. RUGNUX_VERIFY_FIRST_PASS_MEMO recomputes
and throws on a difference.
md5-identical output on four sets; GPU wall 28.9 -> 26.1 s, 43.9 -> 41.0 s,
30.3 -> 26.9 s and 143.6 -> 120.5 s (the set that does not index); the
verify mode found no difference on the last.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Three serial stretches of the canonical pass's tail, each made parallel with
the same arithmetic in the same order:
- MeasureBatchDeltaCCHalf measured every batch of the curve, every open
candidate of the rejection loop and every ledger range one after another.
Each measurement is a pure function of its range and the fixed totals, so
they now run side by side (measure_with, one scratch set per worker) and the
decisions scan the results in the original order; the edge walk (locate)
stays one at a time.
- The twin-immune zone evidence (CentricOverAcentric, a few hundred
exponentials per reflection) is evaluated in parallel and summed in the
original order.
- The P1 cross-check's AnalyzeTranslationalNCS was still called with one
thread; it gets the run's thread count like the other two calls.
md5-identical p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz and report on four sets;
GPU wall 41.0 -> 38.3 s and 26.9 -> 24.8 s on the two sets with long merges.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
On a 16M rotation run the pre-scan was ~11 s on the critical path, the first
3 s of it the spot-width tiers (CPU spot finding on the sample frames),
which the shadow, the defective-pixel mask and the beam-centre capture
waited for although none of them reads what the tiers measure.
The sample is now read twice. The first pass builds the projection and
nothing else; the shadow, the diagnostic, the defective pixels and the
capture start as soon as it is in. The second finds the spots - the width
tiers exactly as before, on copies of the experiment and the pre-shadow mask
it always read - in the background beside them, and reads only the frames
it has something to measure on. What it measured (radii, bandwidth, powder
rings, spot quantiles, the spot-symmetry pool) is applied once both are done,
before anything that reads it. The diagnostic JPEG is rendered in the
background from copies too.
md5-identical p.hkl / p.mtz / p_P1.mtz / p_unmerged.mtz, report and
p_detector.jpg on four sets. GPU 26.2 -> 23.8 s and 38.3 -> 36.4 s on two
16M sets; CPU-only 210 -> 203 s and 114 -> 103 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Two whole-image passes per frame out of the CPU spot finder (the pre-scan's
finder on every build, and every image on the CPU-only build):
- AccumulateRings ran three passes over the frame - the plain ring statistics
and two sigma clips. The plain pass now also counts each ring's valid
values in a histogram (0..1023, the rest in a short list), and the clip
passes sum over the distinct values: each meets the same float test its
pixels would, and the sums are integers, so the totals are the same.
- The local test's first pass is read by DetectAt only inside a candidate's
window. It now marks the row/32-column blocks those windows reach and
keeps its sliding sums everywhere but skips the per-pixel test elsewhere;
the bits it leaves unset are never read.
md5-identical output on four sets (GPU) and three (CPU-only). CPU-only
16M: 203 -> 177 s and 303 -> 263 s; GPU 16M 23.8 -> 22.8 s (the pre-scan's
finder).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
- FFTIndexerCPU::ExecuteFFT built the per-direction histograms and scanned the
per-direction spectra for their most prominent peak in one thread; every
direction has its own histogram and spectrum, so the directions are now
split over the refinement threads, each filled and scanned in the same
order as before.
- BeamCenterShortlist2D scanned the whole padded surface (73 M points on a 16M
detector) once per candidate. It now keeps each row's maximum and the first
index holding it and rescans only the rows a suppression touched; rows in
order, first index within a row, is the same first maximum.
md5-identical output; CPU-only 177 -> 169 s and 101 -> 96 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Every first pass found the spots of its ~200 scheme and validation frames
afresh, although the canonical pass follows a probe on the same frames under
the same settings, and a pass repeating an earlier pass's rescue ladder asks
for the same frames again: on the CPU-only build that is several seconds per
first pass.
The spot lists are now also kept in a store shared by the run and its
copies, keyed by everything a frame's list depends on (SpotFindingKey: the
experiment key - geometry, goniometer, spot budget, settings - the rest of the
spot-finding settings including the measured rings, the ice-ring switch, and
the checksums of the pixel mask and of the spot mapping). A first pass takes
what an earlier one found under the same key. RUGNUX_VERIFY_FIRST_PASS_MEMO
finds the spots again and throws on any difference, field by field.
md5-identical output (four sets GPU, two CPU-only); CPU-only 169 -> 163 s and
96 -> 94 s, GPU 23.0 -> 22.4 s and 119 -> 116 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The census (report only: what the crystal's lattice leaves over on the
scheme and validation frames, up to three further lattices) ran on the
critical path of the canonical pass between the first pass and the image
loop - ~1.5 s on the GPU build for a set where it finds lattices.
Its spots are still found in place; the rest now runs in the background on
copies of what it reads (the experiment, those frames' spot lists, the
lattice, the validation settings) and is collected into the result before
the pass returns. It no longer runs on a pass that only post-refines, whose
result nothing reads.
md5-identical output and identical report on four sets; GPU 24.1 -> 22.8 s
on a set with leftover lattices.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
On the CPU path every image made a separate azimuthal-integration pass
(AzIntEngineCPU) over the 72 MB frame although the adaptive finder's first
ring pass reads the same pixels in the same order under the same rules
(skip the INT32_MIN/MAX sentinels, bins below the mapping's count). That
pass now also accumulates the corrected profile - the same statements as
AzIntEngineCPU, so the same float sums - and MXAnalysisWithoutFPGA takes the
profile from the finder instead of running the separate pass, as the fused
GPU engine already does. Only where the azimuthal engine would be the CPU
one; the finder the pre-scan uses does not accumulate it.
md5-identical p.hkl, p_unmerged.mtz, p_plot.txt and report; CPU-only
163 -> 149 s and 94 -> 91 s. GPU unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Every rotation merge measured CC1/2 before the correction surfaces - a merge
of its own - and the P1 cross-check's value was never read. scale_and_merge
takes whether it is wanted, and the cross-check says no.
md5-identical p.hkl, p_P1.mtz, p_unmerged.mtz and report on four sets
(one log line fewer); ~0.3-0.6 s on the sets with large merges.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
macOS turns Shift + a vertical mouse wheel into horizontal scrolling, so
the notch arrives in angleDelta().x() and y() is 0; the contrast step
read only y and did nothing. The wheel handler now takes whichever axis
carries the notch, for zoom, foreground and background alike.
Help > Mouse Shortcuts names Cmd for Ctrl on macOS (Qt maps Command to
Ctrl), Ctrl-click for right click and Fn + arrows for Home/End/Page
Up/Down, and lists B held + wheel, which was missing. The magnifier's
tooltip says how to drive it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every non-search merge of a rotation pass spent a whole extra merge on cc_half_before_corrections,
but only the pass's first merge (search_merge_cc_half) is read; the ProcessResult copy was never
read at all and is removed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The unmerged file reads the integration outcomes and the determined group, neither of which the P1
merge changes, except each image's mosaicity, which the batch headers carry and the merge rewrites.
UnmergedMtz builds the file without it on a second thread, and SetUnmergedMtzMosaicity fills it in
after the merge, so the file is the same bytes as before.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
A run on data that index poorly asks over a hundred indexing questions - every rung of the
spot-budget ladder, in every pass and walk probe - and a later pass repeats a probe's ladder, which
the 32-entry memo had long evicted (about 20 s on one such sweep). The key holds every spot, so it is
now kept as two independent 64-bit hashes and its length instead of whole.
RUGNUX_VERIFY_FIRST_PASS_MEMO still recomputes and compares.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
zeta >= min_zeta, so the rocking-curve test rejects every solution with
min_zeta * (|phi| - half wedge) > multiplier * mosaicity. Tested right after phi, with a 1e-5 rad
margin (float rounding of the full test is below 1e-6), it skips the cross product, normalisation
and zeta of the 85-95% of solutions no frame keeps, and cannot reject one the full test would pass.
CPU build: on the order of 10 s of a 1.2 A sweep's prediction.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
ReduceGroupMeans ran once or twice per scaling iteration as one serial scatter over every partial
(about 8 s of a CPU-only run's main thread on a 16M sweep). ComputeAsuGroups already builds the
partials' group CSR - a stable counting sort, observation order within each group - for the GPU
reduction; it is now kept on the CPU path too, and each group is summed over it in the order the
serial loop added it, so the means are the same numbers.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
RunScalingLoop rescales after every iteration before anything reads corr, so the iteration's
UpdateCorr and the rescale are one pass over the partials instead of two: the G and fitted frames
as the iteration left them, then the ratio on top - the same two roundings in the same order.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
At the smoothing window every Run has just restored corr to what Ingest built, and nothing else the
measurement reads changes after Ingest, so its answer is the same on every Run; it walked all the
partials twice per Run on the CPU path.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
nth_element compared through pointers into gigabytes of partials, a cache miss per comparison on one
thread (0.3 s on a 16M sweep). The ratio is now computed beside each pointer on all threads; the
comparisons and so the selection and its order are the same.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The ASU-key sample over every partial ran on one thread (0.2-1 s of the merge tail). Each image's
parts are now taken on the workers and joined in image order, so the sort sees the same parts in the
same order.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Each growth is 11 pinned host and 19 device allocations under the driver's device-wide lock; at the
start of an image loop 16 workers growing 1.5x at a time spent about 0.3 s in them.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Spot finding reads whether there is a goniometer, never its axis or angles, and the rotation-scale
walk's probes change only those - so each probe found every frame's spots again (0.2-0.3 s on a 16M
sweep). RUGNUX_VERIFY_FIRST_PASS_MEMO still recomputes and compares.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
A probe whose inputs a first pass has already run on returned that pass's evidence only after the
pass had built its azimuthal mapping, writer setup and indexer (0.15 s each on a 16M detector, twice
at the end of a run). The lookup now comes first; nothing in between changes what its key reads.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
60-70% of a large detector lies outside the resolution band the walk bins, and every one paid a
square root and two arctangents on every iteration. A test on tan(2theta) = rho / lz with a 0.1%
margin, far above float rounding, drops them first; every pixel the exact test keeps still reaches it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Four per-pixel passes over char masks ran on the pre-scan's critical path on one thread; each
pixel's result depends on that pixel alone, so splitting them changes nothing.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The four Kabsch reweighting iterations each walked the reflection's grid again with the same bounds,
validity and ownership tests the p_valid pass had just applied. That pass now keeps the profile value
and background-subtracted count of the pixels the fit reads, in grid order, and the iterations run over
them - the same terms summed in the same order (about 6 s of a CPU-only 16M run).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The sigma-clip pass recomputed the stencil distances over the whole box to find the same background
ring pixels pass A had just summed. Pass A now keeps their values and radial offsets in its own
order, and the clip runs over them - the same pixels, the same sums.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The capture (about 1.5 s on a 16M detector) sat at the end of the pre-scan, but nothing reads its
answer before the first pass's beam-centre check. It now runs on copies of the experiment, mask and
projection it measured, and is taken (JoinBeamCenterCapture) wherever background_center_ or
measured_beam_center_ is read. The first-pass memo key counts a capture still running as a centre
present, which it will be by the time the pass reads it. Its log lines now come from its own thread,
among the first pass's.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
With candidates, the first pass already tested only the 32-column blocks within reach of a candidate,
but still slid the horizontal window across every column of every row. It now slides it only over
the runs of needed blocks, starting each run from the vertical sums it covers - integers, so the same
window sums - and sets the bits there directly; the rest stay 0 as before.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Every valid pixel of every sampled frame read and wrote its valid count and level sum, 20 of the 44
bytes the accumulation moved per pixel, on a pass that is bound by memory bandwidth. Both depend on
the pixel only through its ring-sector and its own error frames: they are now summed once per frame
per ring-sector, and a pixel keeps only what its error frames take out of them. Integers, so the same
sums.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Each frame packed every ring's pixels together and selected in them twice - the median, then the
median absolute deviation - about 40% of the pass. A ring's background is a few counts, so both are
now read off a histogram of the values 0..1023, exactly, wherever the statistic lands inside it (and,
for the spread, no count is negative); otherwise the ring is packed and selected as before.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Wrong: every spot carries the goniometer angle of its frame (DiffractionSpot phi), so the rotation-
scale probes' spots differ from the stored ones - RUGNUX_VERIFY_FIRST_PASS_MEMO caught it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The MTZ points at the experiment's space group, and the background build used a copy that died with
the task: the file was written through a dangling pointer (a verify-mode run failed with a garbled
space-group name; the ordinary runs happened to read the freed memory intact).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Not bit-exact: the radial offset d.rad is ddx*ux + ddy*uy inlined at two sites, which the compiler may
contract to an FMA differently, so the clip's bin (lround(r0 + rad)) moved on a few pixels of one
sweep's CPU run.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
A pass's P1 merges - the search merge, the all-observations merge, the P1 cross-check - each ran the
partial scaling loop from scratch on identical inputs (about 5 s each on a 16M CPU-only run). The loop
restarts from corr_ingested, fits every frame from its own observations and the group means alone,
and writes G and corr only on the frames it fits; everything else it reads is fixed at Ingest. So
over the same ASU grouping and settings (space group, Friedel, resolution limits, partiality floor,
window, iteration cap) it ends where it ended before, and the frames, G and corr it left are kept for
the next Run to take. Frames it does not fit keep whatever G they had, as before.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
One chain runs at a time there, after the parallel candidate solves, on 1-3.5 cores of an otherwise
idle machine. Ceres sums per-thread cost and gradient pieces, so the thread count may move a result at
rounding level; measured, it did not: identical outputs on the four reference sweeps and all 22 smoke
sets, and a sweep that indexes poorly (a chain after every spot-budget rung) went from 96 to 62 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
a1816e905 adds the partly transmitting holder-arm pieces to the beam-stop mask, and the
background beam-centre fit read that mask too. A thin arm running into the beam, masked, takes
cells out of the rings it crosses; which radial bins stay above the sector-coverage gate then
changes as the trial centre moves, and a walk seeded beside the arm (the FFT capture sat 48 px off)
limit-cycled on that edge and stopped 44 px short (sigma 0.34 -> 2.27 px). On an open-arm set whose
header centre is 171 px out, the measured centre no longer indexed (60/60 -> 0/60), the run kept the
file's centre and merged P1 at CC1/2 0.004 instead of P212121 at 0.995.
The arm's dimming is multiplicative in azimuth, which the fit's per-sector amplitude already
absorbs, so the fit now uses the mask without those pieces (ShadowFinder marks them TRANSMITTING);
integration still masks them. The fitted centre is back to its a0518abe6 value on every set checked
(the failing one plus four arm sets); the failing set merges P212121, CC1/2 0.995, ISa 22.6.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
On a 16 Mpx detector the pre-scan's main thread spent most of its kernel time faulting in memory.
GetMask made some thirty per-pixel arrays (1.6 GB) as value-initialised std::vectors, so the calling
thread zeroed - and took the page faults of - every one of them before the parallel passes ran: about
0.75 s on a 16M sweep. They are now Planes (NoInitAllocator vectors) that the ParallelChunks writing
them first touches; each pass writes every element it used to leave at the initial value, and the
three that are filled piecemeal start from a parallel fill.
The projection was also reduced three times - for the ring-centre fit, for GetMask and for the
beam-centre capture - each a 360 MB Download into fresh memory. It is reduced once on the first read
and kept until the caller releases it after GetMask, and PreScan shares the one mean projection
read-only (shared_ptr<const>) with the ring-centre fit, the diagnostic picture and the capture
instead of copying it into each.
Output identical (p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz) on the three GPU reference sweeps;
ShadowFinder and BeamCenterFromBackground tests pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Building a mapping evaluates the geometry of every pixel (resolution, azimuth, solid-angle and
polarization corrections, bin) - about 2 CPU-s and 0.1-0.25 s of wall on a 16 Mpx detector - and a
rotation run built seven, several for a geometry it had already evaluated: the first pass repeats the
pre-scan's, and the canonical pass and the second of two concurrent probes repeat the probes'. None of
it depends on the mask beyond skipping masked pixels, so AzimuthalIntegrationGeometryCache keeps the
unmasked tables of the last geometry (keyed bit for bit on everything SetupPixel reads), and a
mapping for the same geometry copies them and blanks its own masked pixels to what SetupPixel leaves
there. One entry, shared by the run's copies; on myob four of the seven builds now evaluate, three
copy (~40-120 ms against ~240+ ms on a loaded box).
The pre-scan's mapping is only read by the spot measurement, so it is built at the start of that
background task rather than on the main thread ahead of the projection read. It also fills the cache
for the first pass.
Output identical on the three GPU reference sweeps; a new test checks cached against fresh mappings
under different masks and centres.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
MaskDefectivePixels re-read its 60 sample frames with 8 CPU workers - host
bitshuffle/LZ4 decode, CPU preprocessing, then per frame a scatter by ring-sector,
nth_element medians and per-pixel sums over ~22 B/px of state - and was bound by
memory bandwidth: 2.9 s wall on an EIGER2 16M sweep, the largest single step left in
the pre-scan.
With a GPU each worker now decodes and preprocesses the frame on the device
(ImagePreprocessorGPU::AnalyzeCompressed, as the image loops do), and
HotPixelFinderGPU takes the order statistics there: one block per ring-sector and
per ring, radix selection eight bits at a time, which is exact for any values, so
the histogram-plus-fallback of the host path is not needed. Only the per-key
statistics (count, sector median, ring median and MAD) come back; levels and
thresholds are computed on the host by the same function the CPU path now calls
(AddLevels), uploaded, and a per-pixel kernel does AddImage's loop line for line,
including the float comparison. The per-pixel sums stay on the device and are
downloaded once when the mask is read. The CPU path is unchanged. The constructor's
~470 MB of per-pixel arrays are no longer zeroed on the calling thread: they are
allocated unwritten and first written by the existing parallel key pass.
Exactness: a temporary check that ran both finders side by side on the same
frames (CPU decode + preprocess against GPU decode + preprocess) found every
per-pixel and per-key sum identical on myob, cytc, lyso and sparse; the new test
HotPixelFinder_DeviceMatchesHost compares the masks on synthetic frames that take
every branch (histogram fallback, negatives, saturated/error/masked pixels,
134 masked). p.hkl, p.mtz, p_P1.mtz and p_unmerged.mtz are byte-identical to the
previous GPU build on myob and cytc; the masked counts (3/2/0) are unchanged.
myob GPU, quiet box: MaskDefectivePixels 2.90 s -> 0.42 s, total 20.4 -> 19.4 s,
peak RSS 3.52 -> 3.60 GB.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The transmitting-piece step added in a1816e905 called a pixel dim when it sat below 0.75 of its
ring's median. When the polarization factor the background is divided by is not the beam's, every
outer ring is left with a cos 2phi modulation: on one open-arm marCCD sweep (0.75 A, 90 mm, default
p = 0.99) the corrected background runs 1.5-1.8x the ring median along x and 0.6-0.75x along y at
r = 1600-2200 px. Joined, the y lobes made two pieces of 1.01 M and 0.39 M dim pixels holding only
0.17 % deep pixels (noise and a real ~20 % partial shadow) - past the absolute gate of 200 - and the
mask grew from 16,584 to 1,428,328 pixels, covering the top and bottom of the detector.
The step now fits, per 64 px radial band, the 24 sector medians with m + p cos 2phi + q sin 2phi
(sectors below 0.75 of the fit left out, three times over) and judges dim against min(1, model).
Capped at 1 the model only removes dim pixels, so a sweep with no piece keeps its mask bit for bit
and a piece can only shrink or go.
Shadow pixels (b90b377b8 -> this), same binary otherwise:
that sweep 1,428,328 -> 29,821 (the stop's stick extension kept); arm sets 7atg 724,032 -> 723,708,
6oel 223,087 -> 216,126, 7rji 311,951 -> 309,816, 9fhc 173,597 -> 171,372, 5ebi 793,217 -> 785,477,
6zqr 289,324 -> 251,764, 5j23 530,541 -> 526,115, metformin 341,368 -> 231,123 (two speckled
pieces in the high-angle lobe dropped, the arm below the beam kept).
Merges: 6oel, 7rji, 6zqr, metformin unchanged in space group and resolution, R_meas/ISa within
0.1 %/0.2; 7atg fails either way on this base (the beam-centre defect fixed on atg-arm-centre).
The sweep above: same space group, resolution, CC1/2 and ISa (6.72 -> 6.73); R_meas 21.7 -> 22.7 %,
because the spurious mask also hid part of a real ~20 % partial shadow at the top right that neither
version masks (it is above penumbra_ratio).
New test ShadowFinder_AWrongPolarizationFactorIsNotAnArm (unpolarized background, 0.99 stated,
small absorbing speck in the dim lobe) masks 4547 px before and 0 after; all *Shadow* tests pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
On a sweep that indexes poorly the rotation first pass runs every rung of the spot-budget ladder (up to
12 rungs, two schemes each) and then the beam-centre walk (four centres a ring), each a RotationIndexer
run of a few parallel candidate solves and a serial Ceres chain, one after the other on a mostly idle
machine. The rungs are independent - the ladder decides over all of them together - and a ring's
centres are read in a fixed order, so pick_best is split into feed / index / score: every rung (every
centre of a ring) is fed in turn under its own settings, all of them are indexed at once, and they are
then scored and decided in the original order, each under its settings again. A centre is fed on its
own copy of the experiment, since try_beam_center moves experiment_. The indexers read only what they
were fed and their own experiment, so nothing but the log order changes; a centre after the adopted
one is indexed for nothing.
Measured on the weak sweep that nothing indexes well (GPU, load ~5): 62 -> 40 s; each ladder 9.8 ->
4.8 s, the two beam-centre rings 11.6 -> 5 s. Under load 40-65, three interleaved pairs: 442 -> 126,
289 -> 64, 114 -> 85 s. p.hkl, p.mtz, p_P1.mtz and p_unmerged.mtz byte-identical on the four reference
sweeps (GPU) and on two of them (CPU-only); RUGNUX_VERIFY_FIRST_PASS_MEMO passes on two.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Every round of the rotation geometry walk was a full canonical pass - image loop, scale/merge,
space-group search, model validation, P1 cross-check and every output file - while the walk reads
only the round's post-refinement and validation-spot evidence, both in hand before the scaling
engine is built. Rounds now stop there (postrefine_probe_only_, as the lattice arms already do), and
the round the walk keeps is run once more to write the canonical files.
Exact: a round's probe and its full pass are the same code up to the probe's return, and past that
point a pass changes nothing on the run but the space group, which the next pass resets. A pass
does change the run before that point (spot budget, axis sign, beam-centre search), so the kept
LAST round is re-run from the experiment and beam_center_searched_ it started from, not from where
its probe left them; a kept earlier round is gone back to exactly as before. The final pass runs
without the post-refinement measurement the last round carried, which only reads the integrated
reflections.
p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz and p.cif byte-identical against the parent on the three
open-arm sets whose walk takes rounds here (8pqd 1 round, 5m17 2, 7ph1 3; without --model, and 7ph1
also with it, where the report's model section matches too), and on myob and cytc (no rounds). The
walk's log lines (distances, probe percentages, stop reason) are unchanged; the report's PASS count
rises by one, the probe that now runs beside the full pass.
Cost: a walk of n rounds runs n probes and one full pass instead of n full passes. Measured on a
shared, loaded box (load 30-60): a round's full pass took 32-47 s on 7ph1 (56-88 s with --model)
and 44-77 s on 5m17, its probe 6-11 s (16-18 s with --model); 7ph1 run back to back 182.7 -> 113.5 s.
A one-round walk pays one probe more (8pqd: 24 s).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
As in the image loops: AnalyzeCompressed throwing no longer drops the frame from the defective-pixel
statistics; it is logged, the CUDA error is cleared, and the frame is decompressed on the host and
preprocessed as any format without a device decoder already is.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
jfjoch_broker: /start takes an optional `tokens` list (any number, all equivalent; writeOnly in the
API). While the current dataset has any, the endpoints that expose it - /statistics/data_collection,
/result/scan, /image_buffer/{start.cbor,image.cbor,image.jpeg,image.tiff}, /preview/plot{,.bin} -
answer 401 unless the request carries `Authorization: Bearer <token>`; /statistics keeps serving
the instrument view and only omits its `measurement` block. Enforcement is one pre-routing hook
over a named path set (the same set carries `bearerAuth` in jfjoch_api.yaml); tokens are compared
in constant time and never read back or logged. The tokens are replaced only by an accepted start,
and atomically with clearing the previous run's status, plots and image buffer, in this order:
clear, swap, import the new settings - so no moment serves the old run under the new tokens or the
new run's name under the old ones. Settings are validated on a copy first so a refused start
changes nothing.
jfjoch_viewer: a Token field (password echo) next to the http/https scheme in Open HTTP Connection,
JUNGFRAUJOCH_HTTP_TOKEN, and D-Bus LoadFile(..., token) / SetHttpToken; the dialog overrides the
others, nothing is persisted. A 401 clears the display, stops following and puts a line on the
status bar - no dialog, since a dataset changing hands is the normal cause.
Web frontend: key button in the top bar (token in sessionStorage, applied to the bearerAuth
operations by the generated client, and to the raw preview fetch), a tokens field in the start
form, and a "token required" hint on the plots. Python client: Configuration(access_token=...)
after regeneration.
Viewer: reference dataset accepts a structure-factor mmCIF (already read by content; the dialog
now offers it) and a model file can be chosen next to it (ProcessConfig::model_path, as
rugnux --model); the job also carries the reference's free-R flags, cell and setting, and the
copied command line states -z / --reference-column / --model. A grid scan keeps its own preferred
dataset-info plot and "Spots + background" means the spot count there. Dark theme: the navy hero
buttons, checked segments, warning texts and chart guide lines follow the theme instead of their
light-theme colours.
Docs: SECURITY.md section 3 is now the implemented scheme; a guided tour with four screenshots
in JFJOCH_VIEWER.md; broker, OpenAPI, Python client and frontend pages mention the token.
Tests: BearerTokens unit test and an HTTP round trip over the real broker HTTP layer (jfjoch_test
now compiles JFJochBrokerHttp.cpp and links httplib).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A cold run on a spinning disk waited on the disk twice over. The CBF header
scan read 256 kB from every frame on eight threads that each strode through
their own share of the sweep, so they drifted apart and the scan became a
seek storm (34 s for 2400 frames here); and after it, the pre-scan and the
first-pass indexing touch a few hundred frames and leave the disk idle until
the first image loop reads everything at seek-bound rates.
- ReadAhead (reader/): once the dataset is open, rugnux starts eight threads
that read the data files - HDF5 data files (legacy, VDS or the integrated
master) or the per-frame CBF/marCCD/SMV files - in 4 MB pieces taken
strictly in order, into a throwaway buffer. One stream reads this disk at
125 MB/s, eight in-order streams at 190 MB/s, 32 at 157 MB/s. It never gets
more than a quarter of MemAvailable (GlobalMemoryStatusEx on Windows, 4 GiB
where there is no figure) ahead of what ReadRawImage has handed out, so a
dataset bigger than the cache does not evict its own start, and it stops
with the reader. Plain ifstream reads: portable, no POSIX calls.
- Header scans (CBF, marCCD, SMV) hand the files out in order from an atomic
counter (sweep::ForEachInOrder) instead of striding: 18 s -> 12 s for 2400
cold CBF headers. The CBF header is first read with a 16 kB probe and again
with the old 256 kB one only when the separator is not in it, so the parsed
header is exactly what it was: 12 s -> 6 s.
Output unchanged: p.hkl, p.mtz and p_unmerged.mtz md5-identical to the
rc173 baseline on 6toc (CBF, 2400 frames, 6.0 GB) and 9q41 (HDF5 VDS, 900
frames, 5.1 GB), and on 6z9g (HDF5, 12.8 GB) to the unmodified branch; myob
(p.hkl p.mtz p_P1.mtz p_unmerged.mtz) md5-identical to the reference.
Measured cold (files evicted with POSIX_FADV_DONTNEED before every run),
same code without this commit vs with it, on a shared box (load 20-70, other
agents reading the same disk, so single runs scatter by +-20 s):
6toc wall 61.7/62.7 -> 49.3/49.8 s (clean pairs); all data resident
after 62/51/50 -> 45/41/42 s
9q41 wall 67.3 -> 57.8 s (clean pair); resident after 58/46/43 -> 48/37/35 s
6z9g resident after 81 -> 69 s
The first image loop can look slower with this in CBF runs: the old 256 kB
header probes pulled ~70% of the data in as kernel readahead, so the old
loop started warm - after a 34 s header scan instead of 10 s.
Warm (myob, NVMe, cached): 19.35/20.00 s without, 19.76-20.16 s with; the
read-ahead then only copies 9.3 GB out of the page cache, 0.44 s wall and
3.4 CPU-s measured standalone.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
They read only the search merge, which nothing writes until the adopted group's merge replaces it,
and nothing reads them before that merge returns (the refused-promotion arm after it is their first
reader). So they run on their own thread while the engine re-scales in the adopted group, and are
joined before `sm` is replaced. The cell is captured at launch, since a committed change of basis
replaces it. The merges share one engine (one set of device buffers, corr, groups and memos), so the
candidate merges themselves stay serial; this is the analysis that could overlap one of them.
The gap from the all-observation search to "Adopting" went 0.81 -> 0.20 s on a 16M cytochrome C
sweep (median of 3 paired runs, GPU; the adopted merge beside it 1.89 -> 1.71 s with the later
commits). Outputs byte-identical on myoglobin, cytochrome C and lysozyme (GPU) and cytochrome C
(CPU-only); logs and reports identical up to line order.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
What the main thread did alone in each written merge (adopted group, pinned pair, P1 cross-check),
each rewritten so that every sum is still taken over the same terms in the same order:
- WilsonOutliers: the observations are gathered once in resolution order and each shell is one task
(the shells are contiguous runs of that order); the reflection sort is on a total order, so it is
a ParallelSort.
- MergeAndStats: the per-group sums (Bijvoet excess, ISa asymptote, the CPU error-model means) and the
per-pair anomalous sums split the GROUPS over the threads, each walking every full and taking only
its own, so each group is summed in index order; the error-model samples, the Wilson records and
the anomalous cells are filled per observation after an in-order compaction.
- FitErrorModel: the per-bin rounds on up to 16 threads instead of 5 (one bin per task already).
- MeasureBatchDeltaCCHalf: the three fields it reads are taken out of the 80-byte fulls once, on all
threads.
- ScaledObservations (anisotropy): the rocking events walked in blocks cut at hkl boundaries and
joined in block order. Its sort stays serial: parts can tie on (hkl, image), and the event sums
depend on introsort's order of the tied ones.
Median of 3 paired runs (GPU, no other rugnux running), with the other two commits of this series:
cytochrome C tail 11.1 -> 8.5 s (P1 cross-check segment 4.0 -> 2.9 s, reports 0.76 -> 0.70),
lysozyme 9.2 -> 8.2 s, myoglobin 2.57 -> 2.40 s. Outputs byte-identical on myoglobin, cytochrome C
and lysozyme (GPU) and cytochrome C (CPU-only); logs and reports identical up to line order.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Ingest and every Run allocated per-observation arrays as sized vectors (the ingest sort keys, the
remap, the ingest-time delta_phi / partiality / image number / usability arrays, and per Run the
group ids and the group permutation), each zeroed on the main thread and faulted in page by page
before the parallel loop that writes every element of it: 0.6 s of main-thread memset in a loaded
16M profile of ComputeAsuGroups alone. They are plain arrays now, first touched by the threads that
fill them - the idiom the ingest staging buffers already use. RockingEventFrames' walk over every
raw-hkl run (0.3 s alone on the main thread at ingest) counts per block of runs on all threads;
integer histograms, so the sum is the serial one.
Item 2 of this round - building part of the ingest incrementally while the last image loop runs -
is not done: the per-frame pieces (background mean, h range, h histogram) are one of the five sweeps
over the reflections, and the sort and the grouping need the whole sweep.
Median of 3 paired runs (GPU, no other rugnux running), with the other two commits of this series:
ingest segment 0.94 -> 0.67 s on cytochrome C (17M partials), merge segments down with it. Peak RSS
unchanged (4.74 vs 4.77 GB cytochrome C GPU). Outputs byte-identical on myoglobin, cytochrome C and
lysozyme (GPU) and cytochrome C (CPU-only); logs and reports identical up to line order.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Every kernel, cuFFT plan and copy in the beam-centre FFT engine is queued on the legacy NULL
stream, but its buffers came from the stream-ordered pool, whose cudaFreeAsync is ordered on the
thread's non-blocking allocation stream - not after the NULL stream. The free therefore completed
at once while PointShortlist's last Suppress (and, depending on what the plan teardown syncs, the
last inverse transform) was still writing, and the pool was free to hand those bytes to another
worker or return them to the driver.
The capture runs in the background into the first pass of the main run, exactly when the spot
engines are allocating their buffers, and on a loaded shared card its NULL-stream work lags. That
is the "illegal memory access" seen 0.3-0.5 s after "Processing ... images" in four runs across
three binaries and two datasets: a sticky error, reported by whichever worker synchronised next -
the bslz4 device decode, whose host fallback cannot recover from it.
compute-sanitizer --track-stream-ordered-races all over the [BeamCenterFFTGPU] tests: 16502
use-after-free reports before, 0 after; on a rugnux run it flagged the 294 MB transform buffers
freed under a cuFFT kernel. myob output (hkl, mtz, P1 mtz, unmerged mtz) byte-identical.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Both queue their work (kernels, cudaMemcpy) on the legacy NULL stream - HotPixelFinderGPU for its
download, its sums otherwise on the workers' streams - while their buffers came from the
stream-ordered pool, whose cudaFreeAsync is ordered on the thread's non-blocking allocation stream
and so after none of that work. Unlike BeamCenterFFTGPU no free has overtaken a read: every entry
point waits for its work on the host (cudaDeviceSynchronize, or a blocking device-to-host copy)
before it returns. But that holds only by convention, not for a free while unwinding from a failed
call, and compute-sanitizer --track-stream-ordered-races cannot see host synchronisation, so it
reported every reassigned merge buffer as a use-after-free - noise that buries a real race like the
BeamCenterFFTGPU one.
compute-sanitizer --tool memcheck --track-stream-ordered-races all:
- rugnux --mode scale on a myob _process.h5: 22 use-after-free reports before, 0 after
(MergeAccum/MergeAccumRange/MergeRmeas buffers freed by MergeAccum's reassignment or ~Impl).
- rugnux -e 150 -N 8 on myob: 92 before, 0 after (together with the next commit).
myob, cytc, lyso: p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz byte-identical to r4-integration. Full-run
wall time unchanged within noise (17.4/25.9/18.6 s before, 17.6/25.8/18.7 s after); the scale/merge
phase alone (--mode scale, myob) 2.23 -> 2.31 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The engines' kernels run on their own non-blocking stream, which is not ordered after the legacy
NULL stream. The counter reset (cudaMemset, asynchronous to the host) and the radial-background
kernel upload in the Bragg engine constructor, and the direction-grid upload in FFTIndexerGPU (a
pageable cudaMemcpy returns once staged, before the DMA completes) were issued on the NULL stream,
so the first kernel reading them was not ordered after them. Now cudaMemsetAsync/cudaMemcpyAsync on
the engine's stream. In the Bragg constructor that is this->stream: the parameter of the same name
has been moved from.
compute-sanitizer --track-stream-ordered-races all flagged the Bragg ones (16 + 16 reports on
rugnux -e 150 myob; 1 in the [CUDAMemHelpers] tests); 0 after. Outputs byte-identical on myob,
cytc, lyso.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
A sticky error (an illegal address, say) is reported by whichever worker synchronises next - usually
the bslz4 device decode - and MXAnalysisWithoutFPGA::Analyze and Rugnux::MaskDefectivePixels then
logged "falling back to host" and carried on, although the context is gone and the run fails anyway,
later and less clearly. cuda_throw_if_context_lost() now throws a GPUCUDAError ("CUDA device
unusable after an unrecoverable error: ...") there first, which the image loops already treat as
fatal (IsFatalResourceError), like a CUDA out-of-memory.
The last error cannot tell: once cudaGetLastError() has returned a sticky error, it and
cudaPeekAtLastError() both report success. cudaFree(nullptr) frees nothing and does not
synchronise, but returns the sticky error (checked: after an illegal-address kernel it returns
cudaErrorIllegalAddress; with only a pending out-of-memory it returns success), so a handled,
non-sticky decode failure still falls back - MXAnalysis_HandledDeviceDecodeFailureLeavesNoError
passes.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
rugnux will mostly run on GPFS/Lustre-type storage. GPFS caches in its own
fixed-size pagepool, often smaller than a dataset, so streaming the files
ahead of the consumer can evict data before it is used; and choosing the
behaviour by the underlying filesystem is not wanted. This removes the
ReadAhead streamer of aa3fa6b9c: reader/ReadAhead.{h,cpp},
JFJochReader::StartReadAhead/NoteImageRead, the DataFiles() lists that only
it used (HDF5 reader, HDF5ImageSource, HDF5ImageLocator, CBF/marCCD/SMV),
and its start in rugnux_cli.cpp.
Kept from the same commit: the in-order CBF/marCCD/SMV header scan
(sweep::ForEachInOrder) and the 16 kB -> 256 kB CBF header probe, with its
test. They only change the order and size of reads the program makes
anyway, hold no memory and read less (cold 2400-frame CBF header scan
18 s -> 6 s on the measured HDD).
Output unchanged: myob p.hkl, p.mtz, p_P1.mtz and p_unmerged.mtz
byte-identical to the rc173 reference (17.1 s wall warm, 3.8 GB peak RSS).
Reader tests: [HDF5] 108 cases, CBF/marCCD/SMV/sweep/VDS/GetRawImage cases
all pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
MeasureBatchDeltaCCHalf re-formed the reference and re-measured every open
candidate on every round of its removal loop - including the rounds where the
picked candidate was only marked spent (not corroborated / no longer
convicted), which change nothing. On a 7200-frame sweep that is most of its
cost. All changes below keep every sum in its original order, so every
delta-CC1/2 is bit-identical:
- candidates are re-measured only after frames are actually removed; a
spent candidate leaves the others' measurements standing. The curve is
read off the first round of candidates (the width-1 ones are the batches),
and the redundant build_totals after the loop is gone.
- measure_all measures ranges that start on the same frame as one chain,
shortest first, carrying on the per-reflection sums - the same additions
in the same order, but shared frames are walked once.
- the usable observations are laid out once in frame order as
{group, I, info}, and the per-reflection sums are one struct per
reflection, so a measurement streams contiguous memory instead of three
indirect loads and four arrays per observation.
- locate measures the (up to) four edge moves of a step side by side and
takes them in the original order; the candidate's own measurement and the
conviction's are reused rather than repeated.
Verified with a temporary in-process check that re-ran every measurement
through the old code (sparse 324+825, cytc 68+68 measurements, all
bit-identical), and p.hkl/p.mtz/p_P1.mtz/p_unmerged.mtz plus the logged
delta-CC1/2 lines byte-identical to rc173 on myob/cytc/lyso (GPU), cytc and
sparse (CPU, sparse removes 288 and 72 frames) and 8a1a (vs the r4int
baseline binary).
Measured (perf, loaded box): 8a1a delta-CC1/2 20.1+7.4 s -> 2.2+1.7 s wall
(123 -> 13 CPU-s), run 129.9 -> 107.1 s; lyso (CPU build) 0.76+0.30 ->
0.16+0.11 s; cytc (CPU build) 0.15+0.25 -> 0.10+0.16 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The canonical pass at the widened integration radius (and the walk's
write pass) is thrown away when the integrator finds more than
BKG_STARVED_MAX_FRACTION of the reflections without a background ring,
and RunAllPasses re-runs it at the fixed radius. The pass knew this at
integration but still scaled, merged, searched the space group and
relabelled before marking itself superseded.
Now RunAllPasses flags the two passes the starvation check can replace
(rerun_if_starved_), and such a pass returns at the same point a probe
pass does. Readers audited: the geometry walk reads only post_refine
and validation_evidence (both in hand there); RunAllPasses reads only
bkg_starved_fraction_ (set at integration) before replacing pass2; the
space group the search leaves on experiment_ is reset at the top of
every RunPipeline; the speculative walk probe is started before the
exit point. The supercell-guard and header-geometry re-runs are not
flagged, so they merge as before. Skipped when a _process.h5 writer is
open.
Measured on 5 starved open-arm sets (8agq, 8sqq, 9h0q, 6moj, 7qij, no
--model) against rc173 and the r4-integration binary: p.hkl, p.mtz,
p_P1.mtz, p_unmerged.mtz and p.cif byte-identical on all five, and on
myob (not starved). Integration-to-re-run-decision time summed over the
five: 82.8 s -> 19.0 s (8sqq 21.2 -> 3.9 s, 6moj 19.6 -> 5.3 s);
wall totals were dominated by GPU-lock contention on a shared box.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
On aarch64 JFJochBitUnshuffleBlock went through bshuf_untrans_bit_elem_NEON, which mallocs and
frees a block-sized temporary for every block, while the caller (JFJochDecompressHperfPtr) already
owns an unused scratch buffer of exactly that size. Call the two NEON stages it wraps
(bshuf_trans_byte_bitrow_NEON, bshuf_shuffle_bit_eightelem_NEON) directly with that scratch.
Same code, same output; x86 is untouched.
Checked byte-identical under qemu-aarch64 against the classic entry point, hperf and the x86
decoders: element sizes 1/2/3/4/8, every block length 8..1040 elements plus 1536..16384, seven
data patterns, and real EIGER2/PILATUS4 bslz4 chunks.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The nine null replicates ran under ParallelFor, i.e. on pool workers, and a parallel pass reached
from a pool worker runs inline. So the scale fit (FitModelScale's grid) and the rigid-body Jacobian
of every replicate but the one on the calling thread ran serially, and the null lasted as long as
its slowest serial replicate. Sampled from /proc every 0.5 s through the null of a 1.63 A P2 set
(36793 atoms, 4 indexing candidates): CPU use sat at exactly 8.0 cores of 32 for most of its 220 s,
falling to 1-4 at the tail; on every set measured the null averaged 6-8 busy cores. The replicates
now run on std::async threads outside the pool, so their passes queue on the pool as the real
fit's do.
Exact: every pass splits its work the same way wherever it runs (ParallelChunks is split on the
thread count, not on the caller), so p.hkl, p.mtz, p.cif, p_maps.mtz, p_model.cif are byte-identical
and the model-validation section of p_report.txt unchanged on 9hnc, 8sa8, 8xtg, 9fhc, 5epe and 6oel
(battery commands, --model). (6oel's report also differs in SUPERCELL_DOUBLED_CELL, 106.695/119.998
vs the equivalent 73.305/60.002 - that line flips between runs of the same code already, and 6oel
never reaches the null.)
Null wall clock / mean busy cores (shared 16-core/32-thread box with other jobs, so indicative):
9hnc 220 s / 7.8 -> 152 s / 15.9
8sa8 167 s / 8.2 -> 60 s / 23.7
8xtg 88 s / 7.3 -> 46 s / 20.3
9fhc 77 s / 6.7 -> 57 s / 17.7
5epe 45 s / 6.1 -> 36 s / 17.0
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Ceres evaluates some points twice - the residuals alone to decide whether to take a step, then
residuals and Jacobian together at the same point - and every evaluation is a full Fcalc, a solvent
mask and a scale fit. The residuals are a function of the placement alone (the zone's bulk solvent
is fixed by its first evaluation), so the second computation is replaced by a copy of the first. On
the null of a 1.63 A P2 set the replicates' evaluations drop 46/64/72/89/136 -> 44/59/68/82/123, and
each saved evaluation is a serial one on the replicate's critical path. With the previous commit the
null there goes 152 s -> 87 s (27 busy cores); 8sa8 60 -> 53 s.
Exact: p.hkl, p.mtz, p.cif, p_maps.mtz, p_model.cif byte-identical to rc173 and the model-validation
report unchanged on 9hnc, 8sa8, 8xtg, 9fhc, 5epe, 6oel (battery commands, --model); rigid-body
rotations and translations in the log identical. [ModelValidation] tests pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
With --model, model validation and the P1 cross-check merge ran one after the other at the end of
the run. Neither reads anything the other writes, so the first ValidateAgainstModel now runs on its
own thread while the main thread makes the P1 merge (rsm->Run in P1); the merge's result is kept and
reported at the cross-check, where it used to be made, through a new `already_run` argument of
scale_and_merge.
Why it is exact: the validation reads the adopted merge, the cell, the group and the wavelength,
all captured before the thread starts, and nothing else. The scaling engine merges its own copy of
the observations under the group set on the experiment (P1 for the merge, restored after, as at the
cross-check), holds its own copy of the cell, reads the outcomes' lattices and mosaicities only at
Ingest, and writes back per-image G, CC and mosaicity - which nothing between the validation and
the cross-check reads (the unmerged MTZ build reads the mosaicity but replaces it after the merge,
as before). Every relabelling the validation decides reaches the P1 merge afterwards through
merge_to_written, exactly as when it was made later. The result fields scale_and_merge sets are
still set at the cross-check, so the report and the finalist ledger read the same values. Without
--model, or without a P1 cross-check, nothing changes.
The validation logs into a Logger::Buffered() and the lines are replayed as one block when it is
done, so the two do not interleave; the P1 merge's log lines now come before the validation's.
Checked against rc173 (rugnux_r4int_77710e1ce, output-identical to it) on 8tyy, 9hnc, 8sa8, 6oel,
7ph1 with --model and on myob: p.hkl, p.mtz, p.cif, p_P1.mtz, p_unmerged.mtz, the maps and the
placed model are byte-identical, and so is p_report.txt apart from date/command/timing lines. A
last-digit flip of one image's bkg in p_plot.txt (8sa8, one of two runs) is image-loop
nondeterminism the baseline shows too (8tyy, base vs base).
Tail wall (sweep-quality line to unmerged MTZ written), paired back-to-back runs on a loaded box:
8tyy 110.3 -> 87.3 s, 6oel 77.8 -> 63.3 s, 8sa8 171.4 -> 154.6 s; unpaired 9hnc 368.8 -> 337.0 s,
7ph1 21.1 -> 15.7 s. The saving is the P1 merge's time; the validation still sets the floor.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
bitshuffle_hperf's vector code is x86-only, so aarch64 decoded with the classic NEON path, whose
bit stage emulates x86 movemask (8 and + 8 cmeq + 2 x 9-instruction pairwise-add reductions per
16 bytes, then eight scattered 16-bit stores): ~29 instructions per 8x8 bit block plus a byte
transpose pass with scattered 8-byte stores, and a scalar path for 1-byte elements.
bitshuffle_neon.c follows hperf's structure (per byte plane bit-untranspose, then byte interleave)
but does the 8x8 bit transpose NEON-natively: the eight rows stay in eight registers, bits are
exchanged between them with shift + bit-select in three levels, and one zip level plus two st4
write the 128 contiguous output bytes. GCC 13 -O3 gives ~75 instructions per 16 blocks (~4.7 per
block) in the bit stage; the byte interleave is one st2/st4 per 16 elements. Other element sizes
keep the classic stages with the caller's scratch. x86 is unchanged (the file compiles empty).
Byte-identical under qemu-aarch64 (GCC 13, -O3 and -O0) to the classic decoder, hperf and the
x86 hperf/AVX2 path: element sizes 1/2/3/4/8, every block of 8..1040 elements plus 1536..16384,
seven data patterns, and real EIGER2/PILATUS4 bslz4 chunks (whose decoded image hash also
matches an hdf5plugin read). Same licence and component as bitshuffle_hperf, so licenses/ and
THIRD_PARTY_NOTICES.md are unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Exact speed-ups of the per-pass geometry post-refinement; outputs are byte-identical.
Joint fit: the distance-held solve starts from nominal and reads nothing the free solve
produces, so it now runs on its own thread (std::async) beside the free solve and its
covariance, and is used only where the free fit converged, as before. Each solve has its own
copy of the tilt block and one shared, non-owned CauchyLoss instead of one per residual. The
covariance's Jacobian evaluation is threaded (Covariance::Options::num_threads): each residual
block writes its own rows, so the values do not depend on the thread count.
Rotation scale: the all-data fit and the five leave-a-fifth-out folds differ only in which of
the five per-fifth sums they add, so their golden-section searches are stepped together, one
event sweep per step for all of them, and fits whose brackets still coincide share the point.
Each fit takes the same steps on the same sums as before: cost_grid's per-(k, fifth) sums are
bit-identical whatever k it is handed alongside. About 23 event sweeps instead of ~140.
Measured per pass (rc173 baseline -> this, GPU build, box shared with other jobs; post-refine
log timestamps, joint fit / rotation-scale fit):
lyso 0.36 / 0.28 s -> 0.28 / 0.19 s
cytc 0.42 / 0.35 s -> 0.36 / 0.22 s
myob 0.29 / 0.04 s -> 0.23 / 0.04 s
sparse 0.20 / 0.06 s -> 0.14 / 0.07 s
p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz byte-identical to the rc173 outputs on GPU myob, cytc,
lyso, sparse and CPU-only cytc, and the logged post-refinement lines (geometry, residuals, k)
identical.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The background-starvation check (the widened integration radius left too many
reflections without a background ring) followed only the canonical pass and the
geometry walk's write pass. The two later re-runs - the supercell guard's pass
with pass-1's lattice forced and the header-geometry redo - ran at whatever
radius was in hand, which is still the widened one whenever the canonical pass
was not starved.
A starved supercell-guard re-run marks itself superseded (quality_guard_pass1_
is still set), writes no report and no files, and nothing re-ran it: the run
exited 0 with the files of the spurious-supercell pass it was meant to replace
and a report from the re-run. A starved header-geometry redo was not superseded
but shipped the starved intensities.
The starvation check is now a lambda asked after each of the three; after the
first re-run the radius is the fixed one and it cannot fire again. The
supercell re-run keeps pass-1's lattice forced through its own starvation
re-run. Both late passes set rerun_if_starved_, so a starved one stops after
integration as the canonical pass does.
Demonstrated on myob with temporary switches forcing the supercell guard and a
starved forced pass (removed before commit): before, PASS 3 of 3 superseded and
the files came from the supercell pass; after, PASS 4 of 4 at r1=4 with the
forced lattice. Normal runs byte-identical (myob and a starved open-arm set:
p.hkl p.mtz p_P1.mtz p_unmerged.mtz). No battery run of 4 checked ever fired
the supercell guard.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The GPU finder read the whole image (int32 values + pixel_to_bin) three times per frame for its
per-ring background: a plain pass and two 3-sigma clip passes. As the CPU finder already does
(87cfc879b), the plain pass now also records every valid pixel in a per-ring value histogram
(HIST_VALUES = 1024 bins per ring, global memory, integer atomics aggregated per warp by
(ring, value)), and pixels outside [0, 1024) on an overflow list. Each clip pass is then taken from
the histogram and the list (clip_rings_from_hist, one warp per ring) - integer sums of the same
values, so the ring statistics, thresholds and spots are exactly the image passes'. If the list
overflows its capacity (npix / 32 entries), the clip passes read the image as before; the choice is
made on the device, so there is no extra host sync.
The clip bounds are now written as explicit fused multiply-adds (clip_keep). ptxas already fused
`mean -/+ clip_k * sigma` into FFMA in the image-pass kernel (checked in the sm_120 SASS); spelling it
out keeps the histogram clip and the image clip on exactly the same line.
Memory: nbins * 4 KiB histogram + 8 B * npix / 32 overflow list per engine, only on the
shared-memory path (at most ~1750 rings, so <= 7 MB of histogram); the list is 4.5 MB for a 16M EIGER2.
Exactness: a temporary side-by-side check (histogram clips vs forced image clips on every frame)
gave 0 differing ring sums/counts on myob, cytc, lyso and sparse (overflow list at most 308
entries); p.hkl, p.mtz, p_P1.mtz and p_unmerged.mtz are byte-identical to rc173 on all four.
New test AdaptiveSpotFinderGPU_HistogramClipMatchesImagePasses covers rings above/straddling the
histogram range and the overflow fallback.
Timing (RTX 5080, nsys, GPU locked; myob 16M EIGER2): reduce_rings 230 us x 3 launches/frame
(3.20 s) -> plain pass 308 us + 2 x 4 us early-exit launches + 2 x 12 us clip_rings_from_hist
(1.58 s total); summed kernel time 11.28 -> 9.18 s; image loops 5.89/5.09 -> 5.49/4.69 s. cytc:
kernels 12.86 -> 10.75 s. Small detectors gain less (lyso 0.40 -> 0.32 s, sparse 1.19 -> 1.01 s of
ring reduction), as the histogram costs relatively more in the plain pass there.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
record_fit set r_work_sigma to 0 whenever the null's sd fell under NULL_SD_FLOOR = 1e-3, and the
indexing margin had the same guard at 1e-4. The floor was there so that sd == 0 and sd == 1e-7 do not
land on opposite verdicts, but a zero sigma is itself a verdict: DOES NOT FIT. A large, well-measured
crystal has a null that narrow for real - the replicates of a random placement agree to 0.0006-0.0010
in R-work over hundreds of thousands of reflections - so the crystal's own model, at R-work 0.17-0.29
against a null of 0.50-0.55, read "+0.00 sigma, the model DOES NOT FIT", and the decisions it gates
(setting, indexing, enantiomorph label) were refused. The spread is now floored in the denominator,
(mean - R) / max(sd, floor): continuous at sd -> 0, and a null with no spread because the model is
too small or symmetric for a rotation to matter still reads near zero, since the real fit is then no
better than its own random placements.
Changed, measured on the battery commands with --model (every set in the last battery whose sigma
was 0.00 for this reason; all three were REJECTED):
8sa8 +0.00 -> +328 sigma, FITS; the model's setting is adopted, R-free 0.2273 -> 0.1747
(deposited 0.1525)
9bn8 +0.00 -> +378 sigma, FITS; reindexed to the model's indexing (+137 sigma), R-free
0.5462 -> 0.1721 (deposited 0.1575)
8xtg +0.00 -> +215 sigma, FITS; the indexing it gates was already the model's ("kept"), so the
reflection files are byte-identical and only the verdict changes
Unchanged (files byte-identical, sigmas identical): 9hnc, 9fhc, 5epe; no margin null in the battery
was under its own 1e-4 floor. [ModelValidation] tests pass.
5epe stays DOES NOT FIT for a different reason, not touched here: in F23 one of the nine random
placements lies within the rigid body's reach of the twin-related orientation, walks there (held-out
R-free 0.51 -> 0.18), and the null's sd becomes 0.13.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
QUALITY-AFFECTING, separate from the exact commits before it. The nine null replicates were fitted
at the data's full resolution: a full-resolution Fcalc and mask, 1 + n(indexing candidates) scale
fits over every reflection, the rigid body (which already stops at 3.5 A), then another Fcalc and
fit. Those full-resolution scale fits were most of the null's CPU. Now the replicates, the real
model's R-work and the real indexing margin they are compared with are all computed to
max(d_min, 3.5 A) - the same fits on the same reflections on both sides - but never on fewer than
1000 free reflections, since the margin is read in R-free: on a small cell the limit moves out to
where 1000 of them are, or to the data's own. The reported R-factors, maps, placement and files are
still made at full resolution; only the null test moves.
Why the free-reflection floor: without it the [ModelValidation] reindexing test (a 34x34x38 A cell,
1553 reflections) lost its decision - at 3.5 A its margin was read on ~30 free reflections, the
null margin sd went to 0.053 and the real margin to +2.6 sigma, under the gate of 3.
Measured against the floor-fixed build before this commit (battery commands, --model; CPU seconds
in the null - wall clock on this shared box moved with other jobs, 2-4x less where it was quiet):
set CPU s R-work sigma indexing margin sigma verdict / decision / files
9hnc 2307 -> 583 +10.4 -> +15.3 +106 -> +21 (reindexed) same / same / identical
8sa8 1382 -> 313 +328 -> +203 - same / same / identical
8xtg 968 -> 496 +215 -> +172 +77 -> +27 (kept) same / same / identical
9fhc 1042 -> 737 +201 -> +139 +76 -> +41 (reindexed) same / same / identical
9bn8 393 -> 153 +378 -> +235 +137 -> +56 (reindexed) same / same / identical
5epe 631 -> 600 +0.01 -> -0.03 +26 -> +25 (not decided) same / same / identical
Sets where the null does not run (6oel, 5reo) are untouched by construction. The sigmas drop
because the null's sd at 3.5 A is about twice as wide; the smallest margin here is still 8x the gate,
but a borderline indexing decision would be decided less often, not more.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
A Catch2 executable that builds in the portable configurations (JFJOCH_VIEWER_ONLY /
JFJOCH_RUGNUX_ONLY) and links only what those build - JFJochRugnux, JFJochReader,
JFJochImageAnalysis, JFJochWriter, JFJochCommon - so it can run on the macOS arm64 and
Windows x64 jobs, where the receiver/broker/FPGA/HLS sources are not built and there is no GPU.
EXCLUDE_FROM_ALL, so a product build does not pay for it; catch2 is now made available in the
portable configure as well (it is EXCLUDE_FROM_ALL too).
The cases tagged [portable] cover what depends on the architecture, the compiler or the
standard library: bitshuffle/LZ4/zstd, HDF5 read-back (legacy/VDS/integrated, the direct-chunk
path), miniCBF/marCCD/SMV header parsing, CBOR, CPU spot finding, azimuthal mapping, Bragg
prediction/integration, gemmi MTZ/mmCIF. New in tests/PortableTest.cpp:
- a golden FNV-1a hash of two frames of compression_benchmark.h5, decoded from the raw chunk by
the hperf and the classic bitshuffle and through the HDF5 filter (x86 hashes
affea29c511b6ec2 / e46913009c95a1f1);
- a golden hash of a bitshuffle/LZ4 encode (the writer must produce the same bytes everywhere);
- the shipped bitshuffle block selector against the classic reference over elem 1/2/4/8 and
block tails;
- the FFTW indexer, named explicitly, on a synthetic orthorhombic lattice (the existing FFT
indexer lattice tests run only under CUDA);
- a 4-frame end-to-end rugnux run on the git-LFS rotation dataset (HDF5 via external links, CPU
spot finding, indexing, integration); SKIPs when LFS was not pulled.
All 45 take ~4 s on Linux (~2 s without the LFS case), CPU-only build.
M_PI replaced by PI (common/JFJochMath.h) in the two tagged files that used it, as MSVC does not
define M_PI. The CBF gzip test shells out to gzip and is left untagged.
CI: build and run jfjoch_portable_test "[portable]" in build-windows (both variants),
build-rugnux-windows, build-macos-viewer and build-rugnux-macos, after the build and before
packaging.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The r2..r3 background ring excludes every pixel inside any predicted
reflection's r2 region, and a reflection whose ring kept five or fewer
clean pixels was dropped. The mask marks every prediction in the +-4 sigma
rocking window, most of which put a few percent of their flux on the
frame. On a finely sliced, dense pattern (0.05 deg frames, sigma_M 0.067
deg, 0.35 A, 7200 frames) they fill every ring: ~9400 predictions per frame
but only ~900 with partiality above 0.2, and 88.9% of all partials were
dropped at the fixed radius - including the peak frames. The merge was 17.9%
complete.
Such a ring now falls back on all its readable pixels; the existing
high-side clip removes the neighbour cores that are really there. The
starvation counts are unchanged, so the adaptive-radius guard reads the
same numbers. Detector-starved rings are still dropped.
On that set (fixed radius, no --model): 7.4M -> 67M partials ingested,
completeness 17.9 -> 79.7% (88% to 0.65 A), d_min 0.69 -> 0.61 A,
ISa 12.8 -> 14.3, reflections in common with the deposition 60k -> 252k,
rank CC vs |Fc|^2 minus the depositor's, outer shells -0.10 -> +0.07.
A dense 360-frame set: 14.9M -> 17.6M partials, dCC outer -0.116 -> -0.096.
myob_x10sa MTZ byte-identical.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
"RotationIndexer judges a tilt no mounting can have on the spots" (the cheapest of the four,
~3.5 s on FFTW) joins the portable set, and RotationIndexerTest.cpp is compiled into
jfjoch_portable_test. The three rotation-indexer setups now ask for Auto instead of naming FFT
under CUDA and FFTW otherwise: for rotation indexing DiffractionExperiment::GetIndexingAlgorithm
resolves FFT/FFTW on get_gpu_count() whatever was requested, so this changes nothing on either
build, and states what a run without a GPU relies on.
Checked on a CUDA build with CUDA_VISIBLE_DEVICES="": all 45 [portable] cases pass (Auto falls
back to FFTW - the rotation case takes 3.8 s against 1.4 s with the GPU visible - and the 4-frame
rugnux wedge runs on the CPU path). No change to the indexer selection was needed. The FFTIndexer*
unit tests are not added: none goes through Auto - each names FFT (and FFTW) explicitly, and the
explicit FFT request fails with "cudaHostRegister failed" when no GPU is visible, as it should.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The CPU-only image loop is DRAM-bound on 16M frames: decode, preprocess, the adaptive finder's plain
ring pass and FlagRings each streamed the whole frame through memory.
- JFJochDecompressHperfBlocks hands each decoded bitshuffle block to a callback; with no output
buffer the block is unshuffled into a reused block-sized scratch (JFJochDecompressBlocks).
- MXAnalysisWithoutFPGA::PreprocessCPU preprocesses each block into the int32 buffer
(ImagePreprocessorCPU::AnalyzeBlock) and, when the fused CPU finder runs, puts it through the plain
ring pass + fused azint (AdaptiveSpotFinderCPU::AccumulateRingsBlock) while it is in cache. Detect()
then starts from those sums. The per-worker decompression buffer is no longer allocated for
bitshuffled data.
- FlagRings becomes FlagRow, called by DetectAt's first pass for row y+NBX just before that row enters
the vertical sums; first_pass_needed is marked from each row's candidates at the same point.
Exact: blocks arrive in pixel order, so the float azint sums see the same pixels in the same order;
the per-pixel expressions are unchanged; everything else is integer. p.hkl, p.mtz, p_P1.mtz and
p_unmerged.mtz byte-identical to rc173 on myob, cytc, lyso, sparse (CPU-only build), GPU myob
identical (the GPU path does not take this route).
CPU-only, 32 workers, under the gpulock on a shared (loaded) machine, base -> fused, two rounds
(second in reversed order):
myob 155.9 -> 97.7 s, 139.9 -> 88.1 s (loop 46.1 -> 26.8 s/pass; user 3690 -> 2221 s)
cytc 220.8 -> 156.6 s, 216.6 -> 155.1 s (user 5160 -> 4128 s)
lyso 134.5 -> 123.3 s, 69.5 -> 64.8 s
peak RSS myob 15.2 -> 10.8 GB, cytc 12.6 -> 10.4 GB, lyso 7.0 -> 6.7 GB
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The rigid-body target re-grids the model at every evaluation (gemmi's
put_model_density_on_grid with the refmac blur, then the Refmac solvent
mask), and that gridding was ~80% of an evaluation on low-symmetry cells.
The central evaluation ran it on one thread and the six Jacobian columns
on one pool worker each, so most of a 32-thread machine sat idle through
the real fit.
New rugnux/ModelGrid.{h,cpp} reimplements the two gemmi routines so that
they are bit-identical to gemmi on any number of threads:
- Atoms are put on the grid per w plane: each plane is one task and walks
every atom whose box reaches it, in model order, visiting only its own
points (a copy of gemmi's do_use_points_in_box restricted to one plane).
Every grid point therefore receives the same additions in the same
order as in gemmi's serial loop. Per-atom coefficients and radius are
computed once up front, and each atom's density function is copied to a
local so the compiler keeps it in registers.
- Symmetrization runs over orbits: gemmi reduces each orbit into its
lowest index, so the leaders are the points with no lower mate. They
depend only on the grid size and group, so RefineRigidBody finds them
once per zone; each orbit is then reduced by one thread with gemmi's
operands in gemmi's order. Orbits share no points.
- The solvent mask uses the same two passes (setting points to 0 is order
independent anyway); gemmi's own island removal and shrink follow.
vendored gemmi is untouched.
The six Jacobian columns now run on std::async threads, not on the pool,
so each column's gridding can spread over the pool (a pass reached from a
pool worker runs inline). With one thread they run deferred, serially.
Evidence: the grids are memcmp-identical to gemmi's at nt 1/5/32 on 41
deposited models from the battery's PDB cache plus the 8 battery sets
(23 of them with anisotropic atoms; P1 up to F4132, R3/R32, I and C
centring) at 6/4.5/3.5 A; new test ModelValidation_ParallelGriddingMatchesGemmi
covers P1, C2, P212121, I23 and F4132 with iso and aniso atoms. A
standalone RefineRigidBody benchmark (synthetic |F| from the model,
displaced model) gives identical evaluations/angle/shift/coordinates to
the unchanged code: 5lzl 24-33 s -> 12 s real fit on a loaded machine;
null (9 concurrent replicates) unchanged within noise.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
AdaptiveSpotFinderCPU::AccumulateRings: consecutive pixels mostly share a ring, so the ring's
sum/sum2/count and the fused azint sums are held in locals while they do and stored when the ring
changes - the same additions in the same order (az_sum2 is still contracted to the same FMA), without
a store-and-reload chain through memory on every pixel.
ImageSpotFinderCPU::DetectPass: the vertical-sum update (add the entering row, take out the leaving
one) is one branch-free loop over the raw image that GCC vectorises (int64 lanes). A pixel strong in
the previous pass used to be substituted per pixel through a bit test, which kept the loop scalar;
it is now added with its row and taken out again from the few set bits of prev_strong. Integer sums,
so the same totals. (A first, fully branch-free version that kept the per-pixel bit test did not
vectorise on the prev_strong path and was measured slower; this is its replacement.)
ImagePreprocessorCPU: the per-engine std::vector<bool> built bit by bit from the 32-bit mask
(~10 core-s per cytc run, one per worker per pass) is replaced by 32-pixel mask words that PixelMask
derives once beside its binary mask; each engine copies 2 MB. A branch-free rewrite of the Analyze
loop was measured and dropped: the loop is bound by reading the decompressed image (330 vs 328
core-s on cytc), so only the mask test changed.
Measured (perf, 499 Hz, CPU-only build, cytc, first version of this change): AccumulateRings
591 -> 539 core-s. Byte-identical p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz on myob, cytc, lyso,
sparse (CPU) and myob, lyso (GPU).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Most rotation solutions are far from the frame, and the existing phi prefilter rejected them only
after computing phi = atan2(sin, cos). The same test is now asked of the (cos, sin) pair first:
|phi| > phi_limit is cos(phi) < cos(phi_limit), i.e. c < cos(phi_limit) * |(c, s)| in double.
phi_limit is the existing test's threshold at the largest rocking width any reflection can have
(bandwidth term at the resolution limit) plus a 1e-3 rad margin, far above float rounding, so every
solution rejected here is rejected by the existing test too, and the survivors go through it
unchanged. One cos per Calc; min_zeta <= 0 disables it, as it disables the existing test.
Measured (perf, 499 Hz, CPU-only build, cytc): atanf32 + atan2f32 + __atan2f_finite
196 -> 12 core-s, Calc itself 161 -> 199 core-s (the sqrt of the new test), net ~ -150 core-s.
Byte-identical p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz on myob, cytc, lyso, sparse (CPU) and myob,
lyso (GPU); Bragg* tests pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Each iteration of the partial scaling loop passes over all partials three times (group means,
per-frame G fit, corr update + rescale), each time reading a few fields of the 80-byte Obs, and the
group means gathered whole Obs records through the group permutation. The loop now works on
ScalingObs (frame, group, I, sigma, partiality, prescaling_corr, zeta, on_ice; 36 bytes, partials'
order) with corr in its own float array, and the group means read ScalingRef (I, sigma, partiality)
laid out in the CSR's group order, so only corr is gathered. corr is written back to the partials
(and kept by the memo) when the loop ends. ReferenceWeight takes the fields, FitPerFrameG is a
template over Obs/ScalingObs; the same values, so the same arithmetic in the same order. Transient
memory: 52 bytes per partial during the loop (~0.4 GB on 8M partials).
Measured (perf, 499 Hz, CPU-only build, cytc): update/rescale lambda 104 -> 35, group means
87 -> 24, FitPerFrameG 63 -> 27 core-s (-167 core-s, of which the copy costs ~10). CPU-only path;
the GPU scaling path is untouched. Byte-identical p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz on myob,
cytc, lyso, sparse (CPU) and myob, lyso (GPU); rotation scale/merge tests pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
refl_mask (uint8) and owner (uint32) were full-frame arrays per engine, ~90 MB at 18 Mpixel, of
which a call writes only the neighbourhoods of its reflections; over a sweep every page of both got
touched in every worker, and the dirty-rectangle clears were row-strided memsets. They are now a
TiledFrame: a per-tile directory plus a pool of 16x16 tiles allocated on first write and dropped by
Clear(); reading a tile nothing wrote gives the empty value (0 / BRAGG_OWNER_NONE). Writes and reads
are exactly those of the flat arrays, so the same ownership decisions.
Measured (perf, 499 Hz, CPU-only build, cytc): RunImpl 1633 -> 1557 core-s, memset 88 -> 32 core-s.
Peak RSS and page faults did not move measurably (myob 15.3 GB; the peak is set elsewhere).
Byte-identical p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz on myob, cytc, lyso, sparse (CPU) and myob,
lyso (GPU); Bragg* tests pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The null is nine fixed-seed random orientations of the model, each placed by the same rigid body as
the real fit and fitted to the data AS MERGED. Where the indexing probe reindexed the data to the
model's, the model's solution against the merged data is its twin-related orientation - and more
generally the crystal looks the same from every orientation the lattice's rotations (the space
group's, composed with the twin laws) take the model to. A draw within the rigid body's reach of one
of those is walked onto it and is not a sample of the null.
On 5epe (F 2 3, twin law y,x,-z, 24 equivalent orientations) replicate 8 of 9 was drawn 7.9 deg from
one of them; the rigid body moved it 7.9 deg, held-out R-free 0.53 -> 0.20, R-work 0.197 against
0.56-0.58 for the other eight. The null's sd became 0.125 and the crystal's own model read +0.01
sigma, DOES NOT FIT, so the indexing it had decided at +25.8 sigma was refused and the files kept the
merged indexing (R-free 0.52). This happens on every binary whose indexer picks the other hand.
The fix is at the draw, not the score: a draw whose angle to the nearest equivalent orientation is
under the rigid body's reach is drawn again from the same generator, so everywhere no draw falls
there the nine rotations - and every output - are exactly the ones before. The reach is the rotation
about the centroid that moves the model's rms-radius atom by the ladder's first zone (6 A, or d_min
where coarser): 18.7 deg on 5epe, whose other draws sat at 21.7 deg and beyond and did not converge.
Twin laws come from the lattice (ReindexAmbiguityOperators on the data cell and group) regardless of
whether the probe ran. Redraws are capped at 1000 so that a model small enough for the reach to cover
the whole of SO(3) takes its draws as they come instead of looping.
Scoring the real model against the best of both hands instead was not taken: the real fit is
already scored in the hand that wins, and the contamination is in the null, not in the real fit.
Measured, battery commands with --model, rc173 27e154e5d vs this:
5epe 1 redraw; DOES NOT FIT +0.01 -> FITS +55.86 sigma (null 0.5730 +- 0.0072); indexing
decided at +40.4 sigma and reindexed y,x,-z; R-free 0.5246 -> 0.1749 (deposited 0.1671)
9fhc 1 redraw (reach 8.8 deg; the replaced draw had not converged); files byte-identical,
MODEL_FIT_SIGMA +201.27 -> +199.39, verdict unchanged
9hnc, 8sa8, 8xtg, 9bn8: no redraw; p.hkl, p.mtz, p.cif (bar the audit date), p_maps.mtz,
p_model.cif and the model-validation report byte-identical
[ModelValidation] tests pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Conflicts in the three CPU pixel loops the fused pipeline split per block, resolved by carrying this
branch's loop bodies into the new functions with the per-pixel arithmetic unchanged:
- AdaptiveSpotFinderCPU: the register-held ring sums now live in AccumulateRingsBlock (stored at the
end of each block, so the additions stay in pixel order).
- ImagePreprocessorCPU::AnalyzeBlock reads the PixelMask-derived 32-pixel mask words at first + i.
- ImageSpotFinderCPU::DetectPass: rc173's new_row() marking/fill_row calls kept in place, the
vertical update replaced by the vectorised slide with the prev_strong fix-up.
Byte-identical p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz vs rc173 references on myob, cytc, lyso,
sparse (CPU-only build) and myob (GPU build); targeted tests (ImageSpotFinderCPU*, AdaptiveSpotFinder,
SpotFinding, PixelMask, Bragg*, RotationScale, AzimuthalIntegration, portable) pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Two fixes found in the post-merge profile, where the previous commit's loops did not do what they
were written to do:
- AccumulateRingsBlock: the ring_overflow.emplace_back call inside the loop left no callee-saved
registers for the locals, so GCC kept sum/sum2/count on the stack and the store-reload chain was
still there (29% of the function's samples on one stack add). The azint sums get a loop of their own
over the block, and the out-of-histogram values are listed by a second loop run only when the block
has any, in the same pixel order; now the sums live in registers. Same additions in the same order.
- DetectPass: inlined into DetectPass's main loop the vertical slide was not vectorised (the
vectorised copies GCC made were for the other call sites). It is now a free function,
SlideVertical, which GCC vectorises on its own.
Measured (perf, 75 s of myob, CPU-only, relative to the unchanged FlagRow/AnalyzeBlock):
AccumulateRingsBlock -4..-11%, DetectPass + SlideVertical -7..-14%. Byte-identical p.hkl, p.mtz,
p_P1.mtz, p_unmerged.mtz on myob, cytc, lyso, sparse (CPU) and myob (GPU); ImageSpotFinderCPU*,
AdaptiveSpotFinder, SpotFinding, AzimuthalIntegration and portable tests pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
RunImpl on the CPU path is memory-latency bound, not arithmetic bound:
on the 18 Mpx rotation sets ~60% of its cycles sat on the first load of
an image, reflection-mask or owner-map row (each box row is its own
cache line and 4 kB page, in a frame far larger than the caches), and
the function took 59% of all DRAM demand fills of a cytc run (5.0e9).
Pass A now asks for the next reflection's box rows (image and
reflection mask over its outer-ellipse box, owner map over its claim)
while the current one is summed; pass B asks for the fit grid's image
and owner rows as soon as the grid size Rf is known, before the
Gaussian profile is built. A prefetch is only a hint, so every result
is unchanged: p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz byte-identical on
cytc, lyso, myob and sparse; Bragg* and [integration] tests pass.
The prefetch loops are written out in place: GCC 13 treats a function
that only prefetches as having no effect and deletes calls to it (a
lambda version compiled to nothing).
Measured on cytc (CPU build, -march=x86-64-v3, 32 threads, perf at
199 Hz): RunImpl 1645/1665 -> 1375 core-s (-17%); RunImpl DRAM demand
fills 5.0e9 -> 3.3e9. Walking the reflections in detector order instead
(bands of rows) changed neither the fills nor the time, so the misses
are compulsory, not a matter of order - that variant was dropped.
Remaining stalls: the owner-map build and pass B's image rows.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
rc173 stores refl_mask and owner as 16x16 tiles (cpu-loops-tail), so
the rows of those two can no longer be prefetched as flat rows; the
prefetch now covers the image only, and the box helper uses a local
struct in place of the removed Rect.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
B+C of the rigid-body speed-up plan (not bit-exact, by design):
- Fcalc is composed from the transform of ONE copy of the model, gridded without
symmetrization: F(h) = n_cen * sum_S exp(2 pi i h.t) F1(hR), through a per-zone gather table
(half-l index, Friedel flag, phase, Cartesian hR). Only the indices prepare_asu_data() would
give a value are composed (ASU, strict d_min, Nyquist box, model-group absences, not 000),
so the unmatched-observation semantics are unchanged. The mask stays symmetrized and is read
at h directly.
- Translation columns of the Jacobian are exact (the phase factor's derivative), computed with
every evaluation; rotation columns are forward differences needing only one copy gridded and
transformed (3 per Jacobian instead of 6 full evaluations).
- The Jacobian holds the scale and the bulk-solvent mask at the evaluation's; the residuals keep
both exact. The per-evaluation scale re-fit is folded in by Kaufman's variable projection.
- Composition and projection reductions are per-reflection / serial, so the result does not
depend on the thread count.
Tests: RigidBodyP1FcalcMatchesSymmetrized (hkl set identical, <= 4.3e-6 of mean |F|),
RigidBodyTranslationDerivativeIsExact (5-point difference, <= 6.3e-4, the grid sampling),
RigidBodyRotationDerivativeMatchesDifferences (q != 0, anisotropic atoms, cos >= 0.9968),
RigidBodyJacobianMatchesNumericScaleRefit (cos >= 0.998 at 0.07 A; >= 0.983 at 0.35 A, the
residual-weighted term Kaufman drops). Existing rigid-body tests pass unchanged.
Validation on 12 open-arm sets against 20260928-0459_24ae26_r6-pooled: 9hnc's real fit, which
moved 0.000 deg / 0.000 A before, now moves 1.46 deg / 0.65 A and R-free 0.5159 -> 0.4651
(MODEL_FIT_SIGMA 15.4 -> 43.3). STOP finding on 9yzk: a null replicate drawn outside the
redraw guard's 10.6 deg reach was walked 23.9 deg onto an equivalent orientation (R-free
0.5578 -> 0.3917), the null SD went 0.016 -> 0.066 and MODEL_FIT flipped ACCEPTED (12.2 sigma)
-> REJECTED (2.6 sigma). The stronger optimiser's capture radius outgrows the null's redraw
guard; not ready to merge until that is addressed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
ca1baffef moved RotationScaleMergeGPU to synchronous cudaMalloc. cudaMalloc
cannot take memory the default pool holds, and after the image loop the pool
still reserves what the per-worker engines freed with cudaFreeAsync (the 1 GB
release threshold only applies at a synchronisation of the pool's streams).
Measured at the merge's construction on a 123.6M-partial rotation set: pool
reserved 5.64 GB, used 0.83 GB - 4.8 GB of the 16.3 GB card idle but
unavailable. The merge's per-observation arrays (about 9.5 GB there) then ran
out at a 494 MB array with 316 MB free, and the run failed with "Failed to
allocate device memory" right after the post-refine gather. The 77710e1ce
binary allocated those arrays from the pool and passed.
A failed synchronous allocation now waits for the device, trims the pool to
zero and tries once more. Memory placement only: no arithmetic changes.
p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz byte-identical on myob, cytc, lyso
and sparse; the failing set now passes (R 3:H, 1.30 A, ISa 10.9).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
- The background (black-point) slider is hidden by default. It is shown from
View > Background slider, or as soon as the background becomes non-zero
(B + wheel), so an active background is never invisible. The choice is kept
in QSettings ("backgroundSlider") and cleared by "Reset all settings".
- Auto contrast, continuous or the one-shot A key, sets the background back
to zero.
- A dataset-info plot without a run colour kept the default QPen: black, and
near-invisible on the dark theme. This is every plot of a live HTTP dataset,
whose run list is empty. It now takes #1F77B4 on the light theme and
#FF7F0E on the dark one (the spots / background colours of the split plot).
Checked headlessly: the rendered line pixels are exactly those colours on
the #FFFFFF and #12142B plot backgrounds.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The pre-draw guard keeps a null replicate from STARTING within RigidBodyReachDeg of an orientation
equivalent to the model's (space-group rotations x {identity + twin laws}), but the analytic-Jacobian
rigid body walks further than that reach: on 9yzk one replicate drawn outside the 10.6 deg reach was
walked 23.9 deg onto an equivalent orientation (R-free 0.5578 -> 0.3917), the null SD went
0.016 -> 0.066 and MODEL_FIT flipped ACCEPTED -> REJECTED.
Now where the replicate ENDS is checked too: RefineRigidBody reports the rotation it applied, the
final orientation is that times the drawn one, and a replicate that ended within reach is drawn again
(through the same pre-draw filter) and placed again, capped at NULL_MAX_REDRAWS. The redraws come from
a generator per replicate (NULL_SEED + 1 + i), so they do not depend on the order the concurrent
replicates finish in; the count is logged once per run. The real fit's procedure is unchanged.
Validation, 12 open-arm sets against 20260928-0459_24ae26_r6-pooled: post-placement redraws on 9yzk (1)
and 5epe (1), none elsewhere. 9yzk back to ACCEPTED, +14.15 sigma, null 0.5880 +- 0.0138 (base +12.21,
0.5864 +- 0.0158), p.mtz byte-identical to base. 5epe: MODEL_FIT +41.62 -> +46.39, indexing margin
+40.94 -> +31.97 sigma, same decision. Every other set identical to the previous rigid-bc commit.
-N 1 and the default thread count give identical outputs and logs on 9yzk and 6toc.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The dataset-info and per-image plots used matplotlib's tab10 (blue, orange,
...) or the Qt chart theme's defaults, unrelated to the salmon/navy and
indigo/light-blue themes. They now come from ViewerTheme:
- PlotSeriesColor(i): light/dark pairs of one hue each. First a steel navy
#3A6EA5 on light (the heading navy #1F3A5F drawn 3 px wide reads as black)
and the headings' light blue #9DBBE3 on dark; then the caution amber
#B8860B / #E9C46A; then teal, plum, green, rose, slate, brown softened to
the same weight. Contrast on the plot area: 3.2-5.7:1 light, 8-11:1 dark.
The navy/amber pair (Spots + background) is also colour-blind safe.
Coral and red are kept out: hover lines, the marker and alerts use them.
- The old pair measured 2.5:1 (orange on white) and 3.8:1 (blue on the dark
#12142B plot); navy itself would be 1.6:1 there.
- Current-image marker: coral (the sliders' fill) with a ring of plot
background, instead of black/near-white.
- StyleChartAxes: grid lines take the slider groove colour (salmon #F3D9D4 /
indigo #363A66) and the dark theme's axis lines are muted; Qt's light grey
grid outshone the data on dark.
- Run colours are picked again on a theme change (they used to keep the
previous theme's).
Verified headlessly in both themes, including a live theme switch: rendered
line pixels are exactly #3A6EA5/#B8860B and #9DBBE3/#E9C46A.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every merge buffer now comes through Impl::Alloc. When one cannot be had
even after the pool is trimmed, the run ends with what it needed and what to
do: "Scaling N partial observations needs more GPU memory than this card
has: X GB for their per-observation arrays alone (67 bytes each), before the
merge's other buffers, on a Y GB card. This data set is too large for GPU
scaling on this card - run it on the CPU (CUDA_VISIBLE_DEVICES= rugnux ...,
or a CPU-only build)". SetPartialsLayout checks the 67 bytes per observation
of its own arrays against the card's total memory before allocating any of
them, so a set that could not fit an empty card fails before the gigabytes
are allocated.
There is deliberately no switch to CPU scaling mid-run: the two paths differ
in the last bits, so the result would depend on the card.
A 218.7M-partial rotation set (a neighbour-starved dense pattern) now fails
with this message 35 s into the run; with CUDA_VISIBLE_DEVICES= it runs to
completion on the CPU path. No arithmetic changes.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
RigidBodyTarget becomes the CPU backend behind RigidBodyTargetBase, numerics
unchanged. RigidBodyTargetGPU evaluates the same target on a GPU engine: atoms
placed on the device, density by a deterministic brick gather (atoms in model
order per point), cuFFT R2C, the SymmetryComposition gather for Fcalc and
dF/dt in double, the numeric rotation columns as a batch-3 transform, the
Jacobian rows and the Kaufman sums by fixed-order reductions. The mask and the
scale fit are still computed on the host (stubs until ModelMaskGPU and
ModelScaleGPU land).
A pool of 1-4 engines is reserved once per validation, capped at min(25% of
the card, free - 1 GB); GPU or CPU is decided once for the whole validation,
and a CUDA failure restarts it on the CPU. JFJOCH_RIGID_BODY_CPU forces the CPU.
PutMaskOnGrid skips gemmi's shrink where its stencil is empty (every rigid-body
zone): exact, it cannot change a point.
GPU vs CPU on the five test groups: residuals within 7e-6 of <F>, Jacobian
columns within 4e-5 relative, fits 4e-7 A apart.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The unmerged export (_unmerged.mtz) was serial apart from its final row
sort: gathering the partials, sorting them by (h,k,l,image), summing the
rocking events and filling the rows all ran on one thread. On a 17 M
partial rotation set that was 2.8 s (sort 1.6 s, event sums 0.6 s,
gather 0.25 s, rows 0.19 s); on a synthetic 122 M partial sweep 23 s.
- The partials are gathered into preallocated slots per outcome, left
uninitialised so the workers first-touch the pages, and sorted with
ParallelSort. The sort key now ends with the part's input position
(outcome, index), so the order is total and the same on any number
of workers. It differs from the previous std::sort order only among
parts with equal (h,k,l,image_number); measured on four in-house
sweeps (464-6138 such positions each), every summed full is bitwise
the same.
- Events are summed in chunks whose bounds move forward to the next
event start, so no event is split; the chunks are returned as pieces
in order rather than concatenated.
- Rows are filled by index into a preallocated table; absences and the
batch range are still decided in one serial pass. The final row
gather after the sort is parallel too.
Measured: the 17 M set 2.82 s -> 0.86 s, the 8 M set 1.37 -> 0.49 s,
all while the P1 cross-check merge shares the pool; a synthetic 122 M
partial sweep 22.7 s -> 7.7 s at 32 workers. What remains is mostly
ParallelSort's pairwise merge levels. The file write (0.1 s for 112 MB)
is I/O and left as it is.
p.hkl, p.mtz, p_P1.mtz and p_unmerged.mtz are md5-identical to the base
on four in-house sets, on the GPU and the CPU build, and at -N 1.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
gemmi's Scaling<float> fit as RigidBodyTarget uses it - fit_isotropic_b_approximately() and the
Levenberg-Marquardt of fit_parameters() with k_sol and b_sol fixed - and FitModelScale's k_sol/b_sol
grid, with the sums over the reflections on the device (double, fixed launch shape, shuffle tree per
warp, warps and then blocks summed in order: no float atomics, bit-identical repeats). The LevMar
control is gemmi's, ported line for line to the host and unrolled into its requests, so the 88 coarse
and up to 25 fine grid fits share one launch per step.
Against gemmi on real zone hkl sets with synthetic amplitudes: k_overall and b* within 1e-9..1e-5
relative, the same grid winner every time, R within 1e-7.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The equivalent of PutMaskOnGrid() / gemmi SolventMasker(Refmac).put_mask_on_grid() for the GPU
rigid body. Every symmetry image of every atom is masked directly (one block per image, gemmi's
box and !(d2 > r2) rule), which equals gemmi's mask-then-symmetrize-with-min because the operators
are isometries. Islands are removed by a 26-connected periodic union-find (atomicCAS, larger root
under smaller, so each component is rooted at its lowest index), a read-only root pass, a size
count capped at the island limit and a removal pass - no float atomics, bit-identical repeats.
The shrink is not implemented: SetGrid() throws when gemmi's 0.8 A stencil would be non-empty,
which it is not on any rigid-body grid (spacing d/3 >= 1.17 A).
Atoms come in as double fractional coordinates (ModelMaskAtom): with float input, 9fhc at 3.5 A
differed from gemmi at one 24-fold orbit of radius points (|d - r| ~ 2e-6 A); with double input
the mask equals gemmi's on all 8 prototype sets at 3.5 and 6 A, including 6oel, 8t7r, 9hnc, 9fhc.
Time per mask (RTX 5080, Compute incl. islands): 6oel 288^3 4.2 ms, 8t7r 18.4M 2.8 ms,
9fhc 8.0M 1.5 ms, 9hnc 1.9M 0.45 ms.
Union-find credited to Playne & Hawick (2018) at the kernel and in ACKNOWLEDGEMENT.md.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The solvent grid once per zone and the overall scale and anisotropic B at every
evaluation now run through ModelScaleGPU on the engine's stream, on the points
packed on the device. A zone whose grid fits would chain (five or fewer
reflections for the isotropic start) is scaled on the host exactly as the CPU
target scales it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The r2 regions set aside from a reflection's r2..r3 background ring were
those of every prediction in the +-4 sigma rocking window. On a finely
sliced dense pattern most of them are the tails of reflections recorded on
the frames either side, which fill every ring while the frame shows nothing
there. 9ac2ca677 kept the reflections those rings starved by taking the ring
whole, neighbour pixels included, and relying on the high-side clip.
A prediction now masks the ring only where it puts at least 5% of its flux
on the frame (partiality >= 0.05). A ring still starved by those neighbours
has real flux in it, and the reflection is dropped, as before 9ac2ca677:
taking it whole let the neighbours' wings into the background. The CPU mask
holds two levels (tail, flux); the GPU writes them in two launches so the
flux mark wins, with no atomics.
The count the widened-radius guard reads keeps measuring the pattern's
density over every prediction, tails included, so the guard decides on
the quantity its 1.13% bound was read off. Read on the gated mask it let a
cubic set keep r1=6, and R_meas went 36.9 -> 43.2%. Masked but
unreadable ring pixels are no longer counted as the neighbours' doing.
Fixed radius, same code base, --model, the 0.05 deg / 7200-frame set:
pre-9ac2ca677 9ac2ca677 this
rings starved 88.9% 88.9% 21.5%
partials ingested 7.4M 67M 53M
completeness 17.9% 79.7% 85.1%
ISa 12.8 14.3 14.7
R_meas 6.8% 8.4% 7.6%
R_free 0.139 0.176 0.171
radial misfit 0.11 0.185 0.066
R vs model, common hkl to 0.69 A (61k):
0.133 0.130 0.129
(the R_free rise over pre-9ac2ca677 is composition: 63k -> 319k
reflections to 0.69 A, the added ones weaker.) Gate at 0.2 instead:
0.1% starved but radial misfit 0.35, R_free 0.180.
A crystal whose header-geometry pass indexes a 7x supercell:
9ac2ca677's whole rings raised that pass's I/sigma >= 2 count past the
refined pass's by more than 10%. RefinedPassIsWorse then sent the run
back to the supercell. Now: the true cell and space group as before
9ac2ca677, same ISa, 104M partials in the header pass against 219M.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
BraggPrediction::Calc and BraggPredictionRot::Calc stopped filling at
max_reflections (kPredictionCapacity = 20000), silently, inside the h/k/l
walk - so a frame that predicted more kept whichever reflections the walk
reached first, while the GPU predictors grow their buffer (GrowCapacity) and
keep them all. A large cell therefore merged a different reflection set on a
CPU-only build (the macOS/Windows viewer path, or CUDA_VISIBLE_DEVICES=) than
on a GPU one.
Both CPU predictors now double the buffer when the walk fills it, so they
return the same set as the GPU; output_limit and TruncateToOutput are
unchanged and remain the one cap on what leaves Calc, on both paths.
Tests: BraggPrediction_GrowsPastStartingCapacity ([portable], CPU only) - a
large cell predicting 28k (still) / 54k (rotation) reflections on one frame
gives exactly what a buffer large enough from the outset gives, and an
output_limit keeps the best-recorded half. BraggPrediction_CPU_GPU_
consistency_large_cell - the same frame on CPU and GPU: 28384/28384 and
53671/53671 matched by hkl, positions within 0.1 px. Both fail on the old
code (count == 20000).
Validation (CPU build): four in-house sets md5-identical to the base. A
dense rotation set from the private arm predicted 72.0M (20000 x 3600
frames) before and 235.9M now, against the GPU's 235.9M; it ingests
218.65M partials, as the GPU does, and now reaches the GPU's cell, space
group and merging statistics, where the capped CPU run settled on a
different lattice. Cost on the CPU: 72.5 GB peak RSS (17.5 GB before),
44.5 min wall on a shared machine (7.2 min before).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
CI configures MSVC with -DCMAKE_CXX_FLAGS="/arch:AVX" (since rc161). Setting CMAKE_CXX_FLAGS
on the command line replaces CMake's MSVC default "/DWIN32 /D_WINDOWS /GR /EHsc" instead of
appending to it, so every C++ target on Windows - viewer, rugnux, all FetchContent deps - was
compiled without /EHsc: no stack unwinding on a throw (destructors do not run, C4530) and no
_CPPUNWIND. Catch2 detects exceptions from _CPPUNWIND and so built itself with
CATCH_CONFIG_DISABLE_EXCEPTIONS; REQUIRE_THROWS became a no-op and the first thrown exception
(JFJochReader_DataI16 deliberately loads an out-of-range image) terminated jfjoch_portable_test
with "Attempted to translate active exception under CATCH_CONFIG_DISABLE_EXCEPTIONS".
ADD_COMPILE_OPTIONS($<$<COMPILE_LANGUAGE:CXX>:/EHsc>) under IF(MSVC), right after PROJECT()
and before any FetchContent, so it reaches every C++ target whatever CMAKE_CXX_FLAGS holds.
CUDA keeps -Xcompiler=/EHsc from its untouched default CMAKE_CUDA_FLAGS. No effect off MSVC
(Linux flags.make identical before/after).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
ParallelChunks cuts a range into one chunk per worker, so a floating-point
sum folded per chunk and then added up rounds differently at another -N.
ParallelBlocks cuts it by n alone (n / min_per_block blocks, at most 256);
each block folds into its own slot and the slots are added in block order,
so the sum has the same bits at any thread count.
Converted:
- RotationScaleMerge::ApplyCellSurface: the per-cell cross/ref2 sums of the
surface fit (modulation, absorption), previously per-thread partials cut by
ThreadsForWork(idx_all) threads.
- PostRefine: the scale-scan cost grid (per-chunk slots cut by -N).
- PostRefine joint solve: Ceres at a fixed 16 threads (as the rotation
indexer's chain) instead of -N. Ceres sums cost and gradient in
4 * num_threads pieces, so the split no longer follows -N. The pieces are
handed to its threads in scheduling order, which only one thread makes
exact; one thread was measured at +1.0-1.4 s of post-refinement on myob and
lyso (0.5 -> 1.9 s, 1.1 -> 2.1 s), so that channel is left.
- IndexAndRefine supercell probe: per-frame probes kept by image number and
summed in frame order, instead of added under a mutex in completion order;
the primitive is the highest probed frame's instead of the last finisher's.
Integer reductions and per-item passes are unchanged; the three ParallelSort
callers already break ties on the index.
Validation (myob, cytc, lyso, sparse; GPU and CPU builds; -N 8/16/32):
p.hkl, p.mtz, p_P1.mtz and p_unmerged.mtz are md5-identical across -N, and
identical to the base 351be7de0 at every -N. The base was already -N
invariant on these four sets (GPU -N 1..32, CPU lyso -N 2..32), so no output
moved: the converted sums differ from the old ones only in last bits that
never reached a written float.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The gather stages each brick's atoms cooperatively with the Cartesian position
of the image nearest the brick, so a point no longer wraps and transforms
every atom it visits (a narrow cell keeps the per-point image). What a zone
needs of the atoms is worked out once per pool and zone rather than once per
fit. Engine streams run at the device's highest priority, because the first
validation runs beside the P1 cross-check merge.
8t7r-sized fit (63.6k atoms, 30 evaluations, 14 Jacobians): 0.80 -> 0.45 s;
6oel-sized: 0.50 -> 0.27 s.
Tests: a non-CUDA build compiles the GPU test helpers away; the GPU-vs-CPU
residual bound is 2e-4 of <Fobs>, the resolution of gemmi's own scale fit.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
cudaMemcpy from pageable memory can return before the DMA lands, and the legacy stream it runs on
does not order a non-blocking stream, so RemoveIslands could read a partly uploaded mask. Both
uploads are now cudaMemcpyAsync on the test's stream.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The null's replicates are all fitted to the same reflections, so the zone
with its composition table (in the device's half-u layout) and the scale's
points is built once per validation and zone and shared; each fit only
uploads it. 8t7r-sized fit 0.45 -> 0.39 s, 6oel-sized 0.27 -> 0.20 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
- |Fcalc + solvent| is taken once per fit instead of at every step: the solvent pair is fixed for the
whole of a fit.
- The per-step anisotropic factor and the derivatives are computed per point in float, and summed in
double as before. Double arithmetic was most of the cost on a card with little double throughput. The
final R of the grid stays in gemmi's double arithmetic.
- Each mode reduces only the slots it fills; 48 blocks per fit instead of 128.
- The first upload of a batch no longer adds a stream synchronisation.
FitSolvent per zone at 3.5 A (RTX 5080, real zone hkl sets, synthetic amplitudes): 3.2 ms at 3.3k
points, 9.2 ms at 72k and 21.6 ms at 213k, against 9.7 / 29.6 / 62.9 ms before. Fit() is 0.25-0.66 ms.
Against gemmi, k_overall and b* agree to 1e-9 - 1.4e-5 relative, and the grid winner is the same.
New test ModelScaleGPU_MatchesGemmiOnAModelsPoints: the rigid body's own points (the ClusterPdb
fixture with anisotropic atoms, its own amplitudes, a displaced placement, the rigid body's Fcalc and
mask), five groups at 6 and 3.5 A. The R tolerance of the synthetic cases is now 1e-5. That is the
resolution of gemmi's own fit: its Levenberg-Marquardt stops at a WSSR change of 1e-5, and perturbing
gemmi's own Fcalc by 1e-5 moves its answer by up to 1.2e-4 of <Fobs>.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Validated on the 12-set --model comparison against the CPU B+C run
(20260928-1206_e0e3e9_rigidbc-redraw): p.mtz, p.hkl and p_unmerged.mtz
identical on every set, MODEL_FIT, the indexing decisions and the redraw
counts unchanged, no rigid-body commit flipped, MODEL_FIT sigma within -5%
to +3%; two GPU runs identical to the byte on all five output files; the CPU
fallback (JFJOCH_RIGID_BODY_CPU, and CUDA_VISIBLE_DEVICES=) identical to the
CPU build.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The GPU rigid body is chosen by whether a card is there, nothing else. A run
that must stay on the CPU says so at the process level (CUDA_VISIBLE_DEVICES=),
which RigidBodyGPUPool::Create already honours through get_gpu_count().
Parity tests ([ModelValidation][gpu], compiled into jfjoch_test on a CUDA
build):
- every GPU case now SKIPs without a card, instead of passing silently;
- ModelMaskGPU_MatchesGemmi also checks that the CPU path's PutMaskOnGrid is
gemmi's mask bit for bit on the same grid, so the GPU mask is held against
the CPU path itself;
- RigidBodyGPU_MatchesCPU: Jacobian columns 2e-3 -> 5e-4 relative (measured
up to 5e-5; the floor is the 1e-4 the scale's near-tie can move every
column by). Residuals stay at 2e-4 of <Fobs> (measured up to 3.9e-5);
- RigidBodyGPU_FitAgreesWithCPU: fit endpoint 2e-3 -> 1e-5 A rmsd (measured
up to 5e-7 A).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The recompute-and-compare check of the first-pass memos (a probe's
validation evidence, the per-frame spot lists, the RotationIndexer outcome)
is now ProcessConfig::verify_first_pass_memo, passed to RotationIndexer with
VerifyMemo(). No command-line option; only tests set it.
A mismatch now throws MemoMismatch, which the passes' catch-and-continue
sites treat as fatal: before, a mismatch inside a rotation-scale probe was
caught and logged as a probe that did not complete, so verification could
pass without having checked anything.
Rugnux_FirstPassMemoMatchesRecomputation ([large], the LFS rotation set,
~15 s on a GPU) runs the two-pass with post-refinement and scaling, which
takes the walk's probes and the canonical pass through every memo, with the
verification on. Checked that a forced mismatch fails it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Calibration (--mode calibration): RebinAndRefit gave each worker one
AzimuthalIntegrationProfile that AzIntEngineCPU::Run clears on every
image, so the re-binned profile held only the last image each worker
read, added in completion order - on a 1800-frame LaB6 run the re-binned
fit changed with -N (197 / 181 / 205 ring points at rms ~7 px). Each
fixed block of images is now summed in image order into its own slot
(ParallelBlocks) and the slots added in block order, so the profile is
the sum over every image and the same at any thread count (rms 0.93 px,
identical at -N 1, 7 and 32). On the LaB6 runs with a sane header the
first pass still wins and the .poni is unchanged; with a header 100 px
off the re-binned fit is now adopted (218 vs 140 ring points). New test
Rugnux_CalibrationRebinsEveryImage spreads the rings over six frames
and fails on the old code (241 vs 141 points at -N 1 vs -N 4).
Host memory: out-of-memory now ends with "Processing failed: out of
host memory (...) - this data set needs more RAM than is available" and
exit 1. Tested with ulimit -v and an LD_PRELOAD allocator that refuses
allocations from a chosen phase (CPU and GPU builds, image loop through
merging and writing, calibration mode). Paths that crashed instead:
- PostIndexingRefinement ran candidate blocks on bare std::threads, so
a bad_alloc there called std::terminate (seen: "terminate called
recursively", SIGABRT); now std::async futures.
- FFTW aborts on its own failed allocation (CK(p) in kernel/alloc.c,
also inside buffered transforms: BeamCenterFFTCPU, FFTIndexerCPU);
rugnux now defines fftwf_assertion_failed to report it and exit 1.
- libjpeg's default error_exit calls exit() from the diagnostic-JPEG
thread ("Insufficient memory (case 12)"); WriteJPEGToMem now
longjmps back and throws (bad_alloc for JERR_OUT_OF_MEMORY).
- WorkerPool construction that fails to start a thread destroyed
joinable threads (terminate); it now joins them and rethrows.
Also: length_error counts as a fatal resource error, pinned host
allocation failure is MemAllocFailed, and the GPU scaling fail-fast
message says the CPU path needs a lot of host memory.
Docs: very large cells and what running out of host memory looks like
(including the OOM killer) in RUGNUX_INSTALL.md, pointer in RUGNUX.md.
Validation: myob, cytc, lyso, sparse md5-identical to 82e6583e0 on
both GPU and CPU builds.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Two decisions in the rotation two-pass flipped whenever the integration
changed, because each compared a raw number with no account of its noise.
The quality guard (RefinedPassIsWorse) counted reflections at I/sigma >= 2 in
each pass's search merge. Across two lattices that count is not like for like:
a header pass on a 7x supercell merges seven times the reflections, six in
seven empty, and their noise reached the strong count - it out-counted the
refined pass on the crystal's own lattice and the run went back to the
supercell, while that merge held more reflections at I/sigma <= -2 than at
>= 2. The guard now reads the signal as the strong count less the count at
<= -2 (the merge's own noise tail), and its dilution clause scales the
reflection count of the smaller lattice by the integer index of one lattice in
the other, so the same lattice (index 1) compares exactly as before. Measured
on 37 sets: no guard decision changes except the supercell case, which now
keeps the refined pass and the crystal's own lattice both in the configuration
that flipped and at HEAD; the axial-row arm still fires on the crystal it was
written for.
The post-refinement commit gate required each residual family to fall by any
amount. The positional family is what sees the detector geometry a commit hands
the next pass, and its sign was a coin toss where it did not move: on 9yzk it
read +7.1 % in one build and -0.14 % (0.1 of its paired noise) in the next, and
the second committed a geometry the positions had no opinion about (R-free
0.382 -> 0.412, radial misfit 0.30 -> 0.39). The gate now asks the positional
residual to fall by more than its paired standard error (both geometries read on
the same held-out reflections) and the excitation residual not to rise by more
than its own. On the 37-set validation this moves five fit decisions: 9yzk
(R-free 0.412 -> 0.394, radial misfit 0.39 -> 0.29, ISa 9.7 -> 7.8), two
canonical-pass probes whose output is unchanged, and one pre-pass commit on 9pbb
whose excitation moved +0.1 of its noise (ISa 14.77 -> 14.54, R-free 0.2353 ->
0.2357, inside the noise). The in-house myob, cytc, lyso and sparse sets are
md5-identical.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
RUGNUX_ADVANCED: --bandwidth defaults to the pre-scan's measured value where
the file states none; --background-clip is 4 whatever the bandwidth (the
broadband 3-sigma default is gone); --integration-stencil no longer needs
--bandwidth; --polarization text follows the usage (XDS fraction, undulator
default); --beam-center-check describes the both-centres merge arbitration;
the --model placement runs on the GPU when present, CPU otherwise.
RUGNUX_INSTALL: the rigid-body placement is on the GPU, with automatic CPU
fallback; CUDA_VISIBLE_DEVICES= runs a CUDA build on the CPU, as the GPU
scaling fail-fast message says.
JFJOCH_VIEWER: the magnifier follows the cursor only while Shift is held;
the inspector folds away on a narrow window; B + wheel shows the background
slider.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
- RigidBodyGPU gather: AxisBrickBound undercounted by one where a box wraps
past a partial last brick (n = 17: points 15, 16, 0 fall in bricks 1, 2, 0);
the bound is now +2, and axis_bricks reports an overflow that Density()
turns into a failure (-> CPU fallback) instead of dropping a brick.
Exhaustive host test over n <= 64.
- RigidBodyTargetGPU::Residuals tests for an empty zone before launching
Fcalc/Fmask (a 0-block grid).
- The pool engine is held by an RAII lease, returned if the constructor throws.
- The engine's stream is a CudaStream member (priority argument added), so it
is not leaked when an allocation in the constructor throws.
- FitSolvent's "five or fewer strong reflections" case throws its own type,
and only that is caught for the host scale.
- Hot pixels: the per-pixel counters are 16-bit, so the pre-scan sample is
clamped to 65535 frames; no rings -> no 0-block sector/ring launches.
- PostRefine commit gate: fewer than two held-out positions refuse with
their own reason; an excitation family of fewer than two values no longer
refuses (its SE was NaN). No set of the r6-pooled battery reaches either
case (smallest joint fit: ~2400 events, all SEs finite).
- Tests: ModelScaleGPU uploads on the fits' stream and accepts a near-tie's
neighbouring grid point at the CPU winner's R; [gpu] rigid-body cases SKIP
when the card is busy rather than fail.
- THIRD_PARTY_NOTICES: bitshuffle h-perf copyright holder is Kal Conley.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The Windows CUDA 13.3 build failed on rugnux/RigidBodyGPU.cu, the first of our
own .cu files to include CCCL (<cub/cub.cuh>, for the gather's stable radix sort):
CCCL #errors when nvcc's host compiler cl.exe runs its traditional preprocessor.
Pass -Xcompiler=/Zc:preprocessor to every CUDA compile under MSVC, via a
COMPILE_LANGUAGE:CUDA add_compile_options next to the /EHsc one, so it reaches all
CUDA targets including fetched ones. ffbidx already sets the same flag for its own
CUB-using kernel, which is why the older CUDA targets built. C++ sources are left
alone: none includes CCCL (cuda_runtime.h/cufft.h do not pull it in).
Linux: all 243 generated flags.make files are byte-identical before and after.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The speculative geometry probe (StartSpeculativeGeometryProbe) runs an
indexing-only pass on a copy of the run beside the pass's scaling merge.
It holds ~4.5 GB of engines (25 image-sized spot-finding buffers, the
per-engine tables, FFT indexers) plus what its streams' pool retains, for a
few seconds. When a merge allocation landed in that window and did not fit,
RotationScaleMergeGPU::Impl::Alloc reported "needs more GPU memory than this
card has ... too large for GPU scaling", although the per-observation arrays
were 1.3 GB on a 16.6 GB card and the set runs fine alone. Timing-dependent:
one full-battery failure, not reproduced in ~15 plain reruns; reproduced
deterministically by letting the probe hold extra device memory.
- GPUWorkBeside (common/CUDAWrapper): a process-wide count of GPU work
running beside the main line, with a condition variable signalled when the
last one ends. The speculative probe holds one for its whole pass.
- Alloc: on failure, wait for that work to end (bounded, 10 min, then a
"GPU busy" error), then ask for the same buffer once more. Nothing else in
flight means no wait, so a genuine shortage still fails at once. The
computation is the same whenever the allocation succeeds, so results do
not depend on the wait.
- "Too large for this card" is now said only by the early check of the
per-observation arrays against total device memory; a later shortage with
nothing running beside gets its own message (buffer size, free/total, and
that another program or the set's size is the cause).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The wait-and-retry in RotationScaleMergeGPU::Impl::Alloc (c644c4e38) caught
the failed allocation but left its CUDA error as the thread's last error -
CudaDevicePtr leaves it there by contract (CudaDevicePtr_FailedAllocation-
LeavesErrorUntilCleared). The retry then succeeded and the next launch
check, cudaGetLastError() after ReduceGroupMeansKernel or MergeAccum,
reported that stale error over work that went fine, ending the run. Before
the retry existed the failure always ended the run, so the leftover never
mattered.
Reproduced deterministically on a set of 8.2M partials by forcing one merge
allocation to fail (first, early or later buffer of the layout) and by
making the concurrent probe pass hold extra device memory: every such run
failed at the next launch check before, and passes with identical output
now. The battery's "resource already mapped" is the same leftover with
another code from the failed cudaMalloc.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The logo is black only, so the dark-theme negative is the same image with
its colours inverted, redrawn on every theme change. It is now a flat tool
button opening https://www.psi.ch. Button padding is removed and the logo
is scaled to the menu bar's natural height, so it no longer makes the bar
taller than its entries.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
--modelrigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable.lyso_micromax_{mono,pink} and lysoI_micromax_{mono,pink}: 1800 x 0.1 deg on a JUNGFRAU 9M, P4(3)2(1)2. XDS references (FRIEDEL'S_LAW=FALSE) cut where CC1/2 falls through ~30%: 1.50 / 1.45 A native, 1.65 / 1.65 A iodine; ISa 39.9 / 37.4 / 31.4 / 29.2. The iodine sets carry the anomalous signal (low-resolution CCanom ~0.67). XDS.INP and the full-range CORRECT.LP.edge are kept in each set's xds/. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByATUnmasked persistent hot pixels are integrated into whichever reflection's box they fall in; under rotation one pixel collects a different reflection on every frame that reaches it, and the merge carries intensities hundreds to thousands of times their shell mean on one or two observations. HotPixelFinder (rugnux/HotPixels.{h,cpp}) reads the pre-scan sample a second time, once the beam-stop projection has measured where the background puts the beam, and on each frame calls a pixel lit when it exceeds its 2 px iso-2theta ring level (max of the ring and 1/16-sector medians) by 3.3 sigma (sqrt(level) or the ring's MAD) + 2. A pixel lit on at least max(k1, kB) frames is persistent: k1 = 1 + ceil((osc + 5 deg)/|zeta| / frame spacing) is more than one reflection can light, kB the binomial bound (0.01 family-wise over the detector) from the ring's own lit rate. A persistent pixel is masked, as the new PixelMask bit 10, only if it stands alone (component of persistent pixels <= 2), reads on average >= 10x its ring and its mean excess is above the Poisson bound; pixels holding the error value on most frames are masked with them. Counting sensors (thickness > 0) and rotation data only; a CCD is left alone. One log line reports the counts. Drawing the rings about the file's centre, as a first version did, masked pixels along the background fall-off on a sweep whose file centre is 171 px from the background's and cost it 14% ISa; about the measured centre that sweep is within 1%. Masking the detector's outermost row and column unconditionally was tried and dropped: the persistence test already catches the hot pixels there, and the whole lines bought nothing measurable. Numbers below are from the looser first criterion (no isolation / 10x gate), against rc173-final on the same base: merged reflections > 30x their shell mean gone on the sets with proven hot pixels (7brr 21 -> 0, 9ih9 15 -> 0, 8xte 10 -> 0, 6z8o worst 1706x -> 40x); 6z8o CC1/2 0.50 -> 0.995, ISa 8.6 -> 12.8, CC to model 0.82 -> 0.90; 8xte ISa 8.0 -> 10.8; 6u7g ISa 9.8 -> 12.0; controls (lyso_x06da_ref, marCCD) unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5CThe files were written on the axes the space-group search named, which for P2221/P21212 puts the unique axis wherever the a<b<c indexing put it (a 52.51 87.87 137.72 crystal came out P 2 21 21 where XDS writes 87.87 137.72 52.51 P 21 21 2), and a reference MTZ was matched in the data's frame: on permuted axes its free-R flags landed on unrelated reflections while the log reported a high matched count. - New CrystalSetting (scale_merge): changes of basis between settings of one lattice (cell, group, index operator, basis matrix), the {-1,0,1} det +1 candidates (CellMappingOperators, moved from ModelValidation), MetricViolation (moved from Rugnux), ChooseOutputSetting and SeatGroupByAbsences. - Output setting: after every decision the merge, the integrated reflections (unmerged MTZ), the P1 cross-check, the lattice and the _process.h5 reindex matrix are relabelled into the ITA standard setting - or, in priority order, a reference MTZ's, a fitting model's, the -C axis order, a non-standard -S symbol's. Free-R flags are drawn again on the written axes. Reported as SETTING_OPERATOR / SETTING_SOURCE. - Reference MTZ: the group is kept in its setting; after the merge every cell mapping onto the reference cell (times the twin laws) is scored by the reference CC, the best is re-seated and re-merged, and the free flags are inherited only where CC >= 0.5 over >= 50% of the reference range (REFERENCE_MISMATCH otherwise; --mode scale gates the same way). REFERENCE_OPERATOR / _CC / _MATCHED_FRACTION / _FREE_FLAGS_INHERITED. - -S: a fixed group is put on the axes its absences name before merging (SeatGroupByAbsences), fixing -S 18 on a cell whose pure axis is not c. - --model: a model in another setting is now a claim the null tests; where it fits, the data are written in its setting and the validation is remade on those axes (KeepModelVerdict carries the decisions over). Tests: [setting] (synthetic #18/#17/I222/C222/c-unique P21/C2 beta/I2->C2/P1, -C and -S order, absence seating, permuted reference with flags). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C`battery.py recheck RUN [--only a,b] [--jobs N]` reruns model_check.py (REFMAC) and depdata_check.py on every open-arm row of a finished run from its own p.mtz, and updates the refmac_* / dep_* fields of results.json in place (verdicts untouched). Previous results.json, report and each set's model_check directory are kept with a .pre-recheck-<stamp> suffix; the manifest gains a "rechecks" entry {date, end, runner_git, runner_dirty, sets, fields, previous_results, command} so the tools version behind each published number is auditable. The report is re-rendered and the run made read-only again. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5CAdds a second REFMAC protocol to model_check.py beside the first-cycle one (whose numbers and field names are unchanged): - the same deposited model is refined against Rugnux's data and against the depositor's data by the same protocol - restrained, 10 cycles, automatic weight, isotropic B for every atom, twin refinement where the entry declares twinning - each data set over its OWN resolution range; - R-free on a SHARED free set: the depositor's free reflections present in both data sets after the change of basis (so the lower limit in every direction). Depositor free reflections outside the intersection are removed from both; every other reflection is a work reflection of its data set; - R-free recomputed from each output MTZ over exactly that set (F vs FC_ALL_LS); for a twinned refinement REFMAC's own R-free (its free set is the shared set), as the output holds no twinned Fc. REFMAC's own R-free is kept for reference. Isotropic B: refining deposited (often TLS-derived) ANISOU atom by atom diverged (-LL rising every cycle at 1.4 A on a model deposited at 1.7 A). A ligand whose code REFMAC's monomer library describes with other atoms ("atom ... is absent in the library", a stopped refinement) is renamed to a code the library lacks, so REFMAC restrains it from the model's coordinates. Battery row fields: refmac_refined_rfree_shared, refmac_refined_rfree_shared_depdata, refmac_refined_ratio, refmac_refined_rwork, refmac_refined_rwork_depdata, refmac_refined_rfree_refmac, refmac_refined_rfree_refmac_depdata, refmac_shared_free_n, refmac_refined_reason. model_check.py --no-refine skips the protocol. README documents both protocols. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5CThe local radial background at ring radii (--background-radial) already existed and is gated per image on the smooth-ice channel of the ice score, but it was off unless asked for. Systematically negative merged intensities at ring radii (I/sigma < -4 on 0.4-1.6% of reflections of the affected open sets) come from the flat annulus mean over-subtracting a sharp ring. rugnux now runs with =auto unless the flag says otherwise; the library default (and so the broker) stays off. Measured on the open arm against rc173 (commit 727bd0d6d as reference): auto: 8agq (smooth ice, fires) CC_model .9327->.9476, R shell-scaled .1987->.1795, I/sig<-4 1.55%->1.15%; 8rud R shell-scaled .2507->.2456, ISa 10.0->13.3; 5reo, 7kcn, 9fhc bit-identical (gate does not fire); wall time unchanged. forced on (not adopted): same 8agq gain and 8rud .2414, 9fhc/8qq7 small gains, but 7kcn (no ice) CC_model .9491->.9411, ISa 11.0->7.5 - which is why the gate stays. ISa falls on 8agq (15.6->13.3) while agreement with the model rises: the merge-statistics vs external-model trade the settings comment describes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5CThe Rugnux library and the rugnux executable differ only in case, so on a case-insensitive filesystem (macOS, Windows) their CMakeFiles/<target>.dir directories collide and the executable's build.make overwrites the library's ("No rule to make target rugnux/CMakeFiles/Rugnux.dir/depend"). The library is renamed JFJochRugnux, in line with the other libraries. Apple's linker rejects the TLS wrapper clang emits for the inline thread_local static member WorkerPool::in_worker as a duplicate symbol once it is included from several libraries. It is now a function-local thread_local. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>The .dmg cpack produced did not start: jfjoch_viewer links Qt as @rpath/Qt*.framework, install strips the build tree's rpath, and no install rpath was set, so dyld found none ("no LC_RPATH's found"); macdeployqt's "Cannot resolve rpath" errors were the same gap. INSTALL_RPATH is now @executable_path/../Frameworks. Info.plist had an empty identifier and version; it now carries ch.psi.jfjoch.viewer, the numeric version (the rc suffix in the long version string) and the name "JFJoch Viewer". On macOS the notices go into jfjoch_viewer.app/Contents/Resources instead of a share/ folder beside the app in the .dmg, which a user dragging the app out would leave behind. They are installed before the Qt deploy script, which signs the bundle - a file added afterwards invalidates the signature. JFJOCH_NOTICE_FILES is set before the subdirectories for that. Names: jfjoch-viewer-<version>-macos-<arch>.dmg (volume "Jungfraujoch Viewer <version>"), and a rugnux-only build on macOS packs as rugnux-<version>-macos-<arch>-cpu.tar.gz instead of claiming linux. CI: build:macos:viewer and build:rugnux:macos on a host-mode runner labelled macos-arm64. Each checks its artifact (the bundle finds its own Qt, holds no path into the build machine and is validly signed; rugnux links only system libraries and starts) and uploads it on a release build. The Developer ID signing / notarytool / stapler steps are left as a comment until the project has an Apple Developer account. Verified locally: the DMG mounts, the app starts from it and loads every Qt framework and plugin from inside the bundle; the rugnux tarball runs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>CMAKE_OSX_DEPLOYMENT_TARGET 12.0 -> 13.0: the Qt 6.11 frameworks the .dmg bundles are built for macOS 13, so the app could not start on 12 anyway, and the linker warned about every framework ("building for macOS-12.0, but linking with dylib ... built for newer version 13.0"). CI job names follow one scheme - build:viewer|rugnux|jfjoch, then the platform (<os>-<arch> for the portable products, the distribution for the packages), then cuda / nocuda (cuda-sls9 for the slsDetectorPackage 9 builds). The Linux viewer matrix's "cpu" variant is now "nocuda" like the Windows one; the RPM/DEB matrix gains platform and variant beside the distro key its upload step tests. Only display names change: job ids, and so needs:, and artifact file names are as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>jfjoch_broker: /start takes an optional `tokens` list (any number, all equivalent; writeOnly in the API). While the current dataset has any, the endpoints that expose it - /statistics/data_collection, /result/scan, /image_buffer/{start.cbor,image.cbor,image.jpeg,image.tiff}, /preview/plot{,.bin} - answer 401 unless the request carries `Authorization: Bearer <token>`; /statistics keeps serving the instrument view and only omits its `measurement` block. Enforcement is one pre-routing hook over a named path set (the same set carries `bearerAuth` in jfjoch_api.yaml); tokens are compared in constant time and never read back or logged. The tokens are replaced only by an accepted start, and atomically with clearing the previous run's status, plots and image buffer, in this order: clear, swap, import the new settings - so no moment serves the old run under the new tokens or the new run's name under the old ones. Settings are validated on a copy first so a refused start changes nothing. jfjoch_viewer: a Token field (password echo) next to the http/https scheme in Open HTTP Connection, JUNGFRAUJOCH_HTTP_TOKEN, and D-Bus LoadFile(..., token) / SetHttpToken; the dialog overrides the others, nothing is persisted. A 401 clears the display, stops following and puts a line on the status bar - no dialog, since a dataset changing hands is the normal cause. Web frontend: key button in the top bar (token in sessionStorage, applied to the bearerAuth operations by the generated client, and to the raw preview fetch), a tokens field in the start form, and a "token required" hint on the plots. Python client: Configuration(access_token=...) after regeneration. Viewer: reference dataset accepts a structure-factor mmCIF (already read by content; the dialog now offers it) and a model file can be chosen next to it (ProcessConfig::model_path, as rugnux --model); the job also carries the reference's free-R flags, cell and setting, and the copied command line states -z / --reference-column / --model. A grid scan keeps its own preferred dataset-info plot and "Spots + background" means the spot count there. Dark theme: the navy hero buttons, checked segments, warning texts and chart guide lines follow the theme instead of their light-theme colours. Docs: SECURITY.md section 3 is now the implemented scheme; a guided tour with four screenshots in JFJOCH_VIEWER.md; broker, OpenAPI, Python client and frontend pages mention the token. Tests: BearerTokens unit test and an HTTP round trip over the real broker HTTP layer (jfjoch_test now compiles JFJochBrokerHttp.cpp and links httplib). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>A cold run on a spinning disk waited on the disk twice over. The CBF header scan read 256 kB from every frame on eight threads that each strode through their own share of the sweep, so they drifted apart and the scan became a seek storm (34 s for 2400 frames here); and after it, the pre-scan and the first-pass indexing touch a few hundred frames and leave the disk idle until the first image loop reads everything at seek-bound rates. - ReadAhead (reader/): once the dataset is open, rugnux starts eight threads that read the data files - HDF5 data files (legacy, VDS or the integrated master) or the per-frame CBF/marCCD/SMV files - in 4 MB pieces taken strictly in order, into a throwaway buffer. One stream reads this disk at 125 MB/s, eight in-order streams at 190 MB/s, 32 at 157 MB/s. It never gets more than a quarter of MemAvailable (GlobalMemoryStatusEx on Windows, 4 GiB where there is no figure) ahead of what ReadRawImage has handed out, so a dataset bigger than the cache does not evict its own start, and it stops with the reader. Plain ifstream reads: portable, no POSIX calls. - Header scans (CBF, marCCD, SMV) hand the files out in order from an atomic counter (sweep::ForEachInOrder) instead of striding: 18 s -> 12 s for 2400 cold CBF headers. The CBF header is first read with a 16 kB probe and again with the old 256 kB one only when the separator is not in it, so the parsed header is exactly what it was: 12 s -> 6 s. Output unchanged: p.hkl, p.mtz and p_unmerged.mtz md5-identical to the rc173 baseline on 6toc (CBF, 2400 frames, 6.0 GB) and 9q41 (HDF5 VDS, 900 frames, 5.1 GB), and on 6z9g (HDF5, 12.8 GB) to the unmodified branch; myob (p.hkl p.mtz p_P1.mtz p_unmerged.mtz) md5-identical to the reference. Measured cold (files evicted with POSIX_FADV_DONTNEED before every run), same code without this commit vs with it, on a shared box (load 20-70, other agents reading the same disk, so single runs scatter by +-20 s): 6toc wall 61.7/62.7 -> 49.3/49.8 s (clean pairs); all data resident after 62/51/50 -> 45/41/42 s 9q41 wall 67.3 -> 57.8 s (clean pair); resident after 58/46/43 -> 48/37/35 s 6z9g resident after 81 -> 69 s The first image loop can look slower with this in CBF runs: the old 256 kB header probes pulled ~70% of the data in as kernel readahead, so the old loop started warm - after a 34 s header scan instead of 10 s. Warm (myob, NVMe, cached): 19.35/20.00 s without, 19.76-20.16 s with; the read-ahead then only copies 9.3 GB out of the page cache, 0.44 s wall and 3.4 CPU-s measured standalone. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5CA sticky error (an illegal address, say) is reported by whichever worker synchronises next - usually the bslz4 device decode - and MXAnalysisWithoutFPGA::Analyze and Rugnux::MaskDefectivePixels then logged "falling back to host" and carried on, although the context is gone and the run fails anyway, later and less clearly. cuda_throw_if_context_lost() now throws a GPUCUDAError ("CUDA device unusable after an unrecoverable error: ...") there first, which the image loops already treat as fatal (IsFatalResourceError), like a CUDA out-of-memory. The last error cannot tell: once cudaGetLastError() has returned a sticky error, it and cudaPeekAtLastError() both report success. cudaFree(nullptr) frees nothing and does not synchronise, but returns the sticky error (checked: after an illegal-address kernel it returns cudaErrorIllegalAddress; with only a pending out-of-memory it returns success), so a handled, non-sticky decode failure still falls back - MXAnalysis_HandledDeviceDecodeFailureLeavesNoError passes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5CMeasureBatchDeltaCCHalf re-formed the reference and re-measured every open candidate on every round of its removal loop - including the rounds where the picked candidate was only marked spent (not corroborated / no longer convicted), which change nothing. On a 7200-frame sweep that is most of its cost. All changes below keep every sum in its original order, so every delta-CC1/2 is bit-identical: - candidates are re-measured only after frames are actually removed; a spent candidate leaves the others' measurements standing. The curve is read off the first round of candidates (the width-1 ones are the batches), and the redundant build_totals after the loop is gone. - measure_all measures ranges that start on the same frame as one chain, shortest first, carrying on the per-reflection sums - the same additions in the same order, but shared frames are walked once. - the usable observations are laid out once in frame order as {group, I, info}, and the per-reflection sums are one struct per reflection, so a measurement streams contiguous memory instead of three indirect loads and four arrays per observation. - locate measures the (up to) four edge moves of a step side by side and takes them in the original order; the candidate's own measurement and the conviction's are reused rather than repeated. Verified with a temporary in-process check that re-ran every measurement through the old code (sparse 324+825, cytc 68+68 measurements, all bit-identical), and p.hkl/p.mtz/p_P1.mtz/p_unmerged.mtz plus the logged delta-CC1/2 lines byte-identical to rc173 on myob/cytc/lyso (GPU), cytc and sparse (CPU, sparse removes 288 and 72 frames) and 8a1a (vs the r4int baseline binary). Measured (perf, loaded box): 8a1a delta-CC1/2 20.1+7.4 s -> 2.2+1.7 s wall (123 -> 13 CPU-s), run 129.9 -> 107.1 s; lyso (CPU build) 0.76+0.30 -> 0.16+0.11 s; cytc (CPU build) 0.15+0.25 -> 0.10+0.16 s. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5Crecord_fit set r_work_sigma to 0 whenever the null's sd fell under NULL_SD_FLOOR = 1e-3, and the indexing margin had the same guard at 1e-4. The floor was there so that sd == 0 and sd == 1e-7 do not land on opposite verdicts, but a zero sigma is itself a verdict: DOES NOT FIT. A large, well-measured crystal has a null that narrow for real - the replicates of a random placement agree to 0.0006-0.0010 in R-work over hundreds of thousands of reflections - so the crystal's own model, at R-work 0.17-0.29 against a null of 0.50-0.55, read "+0.00 sigma, the model DOES NOT FIT", and the decisions it gates (setting, indexing, enantiomorph label) were refused. The spread is now floored in the denominator, (mean - R) / max(sd, floor): continuous at sd -> 0, and a null with no spread because the model is too small or symmetric for a rotation to matter still reads near zero, since the real fit is then no better than its own random placements. Changed, measured on the battery commands with --model (every set in the last battery whose sigma was 0.00 for this reason; all three were REJECTED): 8sa8 +0.00 -> +328 sigma, FITS; the model's setting is adopted, R-free 0.2273 -> 0.1747 (deposited 0.1525) 9bn8 +0.00 -> +378 sigma, FITS; reindexed to the model's indexing (+137 sigma), R-free 0.5462 -> 0.1721 (deposited 0.1575) 8xtg +0.00 -> +215 sigma, FITS; the indexing it gates was already the model's ("kept"), so the reflection files are byte-identical and only the verdict changes Unchanged (files byte-identical, sigmas identical): 9hnc, 9fhc, 5epe; no margin null in the battery was under its own 1e-4 floor. [ModelValidation] tests pass. 5epe stays DOES NOT FIT for a different reason, not touched here: in F23 one of the nine random placements lies within the rigid body's reach of the twin-related orientation, walks there (held-out R-free 0.51 -> 0.18), and the null's sd becomes 0.13. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5CThe rigid-body target re-grids the model at every evaluation (gemmi's put_model_density_on_grid with the refmac blur, then the Refmac solvent mask), and that gridding was ~80% of an evaluation on low-symmetry cells. The central evaluation ran it on one thread and the six Jacobian columns on one pool worker each, so most of a 32-thread machine sat idle through the real fit. New rugnux/ModelGrid.{h,cpp} reimplements the two gemmi routines so that they are bit-identical to gemmi on any number of threads: - Atoms are put on the grid per w plane: each plane is one task and walks every atom whose box reaches it, in model order, visiting only its own points (a copy of gemmi's do_use_points_in_box restricted to one plane). Every grid point therefore receives the same additions in the same order as in gemmi's serial loop. Per-atom coefficients and radius are computed once up front, and each atom's density function is copied to a local so the compiler keeps it in registers. - Symmetrization runs over orbits: gemmi reduces each orbit into its lowest index, so the leaders are the points with no lower mate. They depend only on the grid size and group, so RefineRigidBody finds them once per zone; each orbit is then reduced by one thread with gemmi's operands in gemmi's order. Orbits share no points. - The solvent mask uses the same two passes (setting points to 0 is order independent anyway); gemmi's own island removal and shrink follow. vendored gemmi is untouched. The six Jacobian columns now run on std::async threads, not on the pool, so each column's gridding can spread over the pool (a pass reached from a pool worker runs inline). With one thread they run deferred, serially. Evidence: the grids are memcmp-identical to gemmi's at nt 1/5/32 on 41 deposited models from the battery's PDB cache plus the 8 battery sets (23 of them with anisotropic atoms; P1 up to F4132, R3/R32, I and C centring) at 6/4.5/3.5 A; new test ModelValidation_ParallelGriddingMatchesGemmi covers P1, C2, P212121, I23 and F4132 with iso and aniso atoms. A standalone RefineRigidBody benchmark (synthetic |F| from the model, displaced model) gives identical evaluations/angle/shift/coordinates to the unchanged code: 5lzl 24-33 s -> 12 s real fit on a loaded machine; null (9 concurrent replicates) unchanged within noise. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C- The background (black-point) slider is hidden by default. It is shown from View > Background slider, or as soon as the background becomes non-zero (B + wheel), so an active background is never invisible. The choice is kept in QSettings ("backgroundSlider") and cleared by "Reset all settings". - Auto contrast, continuous or the one-shot A key, sets the background back to zero. - A dataset-info plot without a run colour kept the default QPen: black, and near-invisible on the dark theme. This is every plot of a live HTTP dataset, whose run list is empty. It now takes #1F77B4 on the light theme and #FF7F0E on the dark one (the spots / background colours of the split plot). Checked headlessly: the rendered line pixels are exactly those colours on the #FFFFFF and #12142B plot backgrounds. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>The pre-draw guard keeps a null replicate from STARTING within RigidBodyReachDeg of an orientation equivalent to the model's (space-group rotations x {identity + twin laws}), but the analytic-Jacobian rigid body walks further than that reach: on 9yzk one replicate drawn outside the 10.6 deg reach was walked 23.9 deg onto an equivalent orientation (R-free 0.5578 -> 0.3917), the null SD went 0.016 -> 0.066 and MODEL_FIT flipped ACCEPTED -> REJECTED. Now where the replicate ENDS is checked too: RefineRigidBody reports the rotation it applied, the final orientation is that times the drawn one, and a replicate that ended within reach is drawn again (through the same pre-draw filter) and placed again, capped at NULL_MAX_REDRAWS. The redraws come from a generator per replicate (NULL_SEED + 1 + i), so they do not depend on the order the concurrent replicates finish in; the count is logged once per run. The real fit's procedure is unchanged. Validation, 12 open-arm sets against 20260928-0459_24ae26_r6-pooled: post-placement redraws on 9yzk (1) and 5epe (1), none elsewhere. 9yzk back to ACCEPTED, +14.15 sigma, null 0.5880 +- 0.0138 (base +12.21, 0.5864 +- 0.0158), p.mtz byte-identical to base. 5epe: MODEL_FIT +41.62 -> +46.39, indexing margin +40.94 -> +31.97 sigma, same decision. Every other set identical to the previous rigid-bc commit. -N 1 and the default thread count give identical outputs and logs on 9yzk and 6toc. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5CCalibration (--mode calibration): RebinAndRefit gave each worker one AzimuthalIntegrationProfile that AzIntEngineCPU::Run clears on every image, so the re-binned profile held only the last image each worker read, added in completion order - on a 1800-frame LaB6 run the re-binned fit changed with -N (197 / 181 / 205 ring points at rms ~7 px). Each fixed block of images is now summed in image order into its own slot (ParallelBlocks) and the slots added in block order, so the profile is the sum over every image and the same at any thread count (rms 0.93 px, identical at -N 1, 7 and 32). On the LaB6 runs with a sane header the first pass still wins and the .poni is unchanged; with a header 100 px off the re-binned fit is now adopted (218 vs 140 ring points). New test Rugnux_CalibrationRebinsEveryImage spreads the rings over six frames and fails on the old code (241 vs 141 points at -N 1 vs -N 4). Host memory: out-of-memory now ends with "Processing failed: out of host memory (...) - this data set needs more RAM than is available" and exit 1. Tested with ulimit -v and an LD_PRELOAD allocator that refuses allocations from a chosen phase (CPU and GPU builds, image loop through merging and writing, calibration mode). Paths that crashed instead: - PostIndexingRefinement ran candidate blocks on bare std::threads, so a bad_alloc there called std::terminate (seen: "terminate called recursively", SIGABRT); now std::async futures. - FFTW aborts on its own failed allocation (CK(p) in kernel/alloc.c, also inside buffered transforms: BeamCenterFFTCPU, FFTIndexerCPU); rugnux now defines fftwf_assertion_failed to report it and exit 1. - libjpeg's default error_exit calls exit() from the diagnostic-JPEG thread ("Insufficient memory (case 12)"); WriteJPEGToMem now longjmps back and throws (bad_alloc for JERR_OUT_OF_MEMORY). - WorkerPool construction that fails to start a thread destroyed joinable threads (terminate); it now joins them and rethrows. Also: length_error counts as a fatal resource error, pinned host allocation failure is MemAllocFailed, and the GPU scaling fail-fast message says the CPU path needs a lot of host memory. Docs: very large cells and what running out of host memory looks like (including the OOM killer) in RUGNUX_INSTALL.md, pointer in RUGNUX.md. Validation: myob, cytc, lyso, sparse md5-identical to82e6583e0on both GPU and CPU builds. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C