Commit Graph
200 Commits
Author SHA1 Message Date
leonarski_f 58a4bf7302 Merge branch 'dq-sym-pg' into rc174-all 2026-10-04 15:58:43 +02:00
leonarski_f 4c7d8a0385 Merge branch 'dq-sym-screw' into rc174-all
# Conflicts:
#	docs/CHANGELOG.md
2026-10-04 15:58:43 +02:00
leonarski_f 0c5ac8c8e5 Merge branch 'dq-sm-scale' into rc174-all 2026-10-04 15:58:36 +02:00
leonarski_fandClaude Opus 5.5 ac179e549e Rotation scaling: reconcile with the ingest partiality fix; no flux on the fulls-only path
Merged dq-sm-absorb, whose ingest fix (corr_ingested carries the recomputed partiality, not the
predictor's) repairs the same defect as this branch's unit partial scale on the path without
partial scaling. The two did not double-apply - the unit scale rebuilt corr from scratch - but
they differ by the incident flux, and measured head to head the flux costs the sparse sweeps:
SHELXL R1 on four in-house organic sweeps 0.0459/0.0396/0.0709/0.0413 with it, 0.0421/0.0393/
0.0697/0.0405 without; on five open small-molecule sets equal or better without (0.0707 ->
0.0660 on one). Only the cubic absorbing sweep prefers it (0.119 vs 0.139): there the
background really does fall with the absorbed beam. So the unit partial scale and its device
kernel are removed and the ingest fix stands alone; the flux meter's smoothing, which only
mattered on that path, is reverted, so the partial-scaling path sees the flux exactly as
before. The penalised per-frame scale of the fulls stays.

Also restores two characters of CPU_DATA_ANALYSIS_INTEGRATION.md that the previous commit's
rewrite had turned from a stray carriage return into a line break.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-04 15:43:15 +02:00
leonarski_fandClaude Opus 5.5 1c3ac29b08 SearchSpaceGroup: refuse a lattice-holohedral group whose merge narrows the L-test like a twin's
Stage A's gates all compare the operators' disagreement with a reference - a parent group, the best
operator, the noise floor - and a pseudo-symmetric structure or a twin defeats every one of them: the
added operators correlate nearly as well as real ones. On the open arm that over-called 2wnz and 2xfw
(deposited P2_1, pseudo-merohedral C-orthorhombic metric) as C222_1, and 7bgu (deposited P1 with
b = c to 2.7%) as C2.

New gate, for a candidate that holds every rotation of its lattice (gemmi find_lattice_symmetry at
3 deg obliquity, with the lattice centring now passed in SearchSpaceGroupOptions::lattice_centring):
no twin law exists for such a group, so intensities merged under it that read like a twin's are not a
twin. AnalyzeLTestUnderMerge reads the Padilla-Yeates L-test twice on the same pairs of the P1 merge -
as measured, and with every intensity replaced by its plain orbit mean under the candidate - so the
narrowing is the merge's own doing and not a property of the data. Refused when the merged <|L|>
reads twin-like (0.375 <= <|L|> < 0.44) and the averaging closed a third or more of the gap from the
unmerged <|L|> to 0.375.

Calibration on the C++ search merges (targeted battery sympg-pg1): genuine lattice-holohedral groups
close 0.01-0.25 of the gap (highest 6z9g, indexed on half its deposited cell, 0.23-0.25; 5vml 0.16-
0.17; 7kcn 0.06-0.09); the over-calls 2wnz 0.47, 2xfw 0.46-0.47, 7bgu 0.79. A merge reading below
0.375 is not read - nothing a twin or a false operator does reaches it, so something else compresses
the intensities: 8c3e (0.353) and the genuine 9zmu (0.351, 494 A axis) look alike there. Near-
perfect twins (4bwl 0.14-0.16, 2wnq 0.13-0.20) are out of reach by construction, and 6p8j (0.19-0.30,
deposited as a twin of P2_1) is left alone. Cost: 170 ms per test on a 730k-reflection merge, only for
candidates that pass every other gate and hold the lattice's full symmetry.

Targeted battery sympg-final (44 open/in-house sets: the over-calls and a control panel of genuine
lattice-holohedral, twinned and twin-like sets) and the private arm: 2wnz C2221 -> P1211 (R_free with
the deposited model 0.227), 2xfw C2221 -> P1211 (0.215), 7bgu C121 -> P1 (R_meas 0.28 -> 0.14), all
the deposited groups. Every other set keeps its space group; against the rc174-cand runs of the same
sets, open and private, nothing else moves beyond run time. Still over-called: 8c3e, 4bwl, 2wnq,
6p8j, 3r6o (reasons above; 3r6o and 8c3e read below 0.375).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-04 15:38:00 +02:00
leonarski_f 21e5ce38db Merge branch 'dq-sm-absorb' into dq-sm-scale 2026-10-04 15:25:28 +02:00
leonarski_fandClaude Opus 5.5 7d6c201e95 Rotation merge: partial corr carries the recomputed partiality, not the predictor's
Every partial's corr is formed at integration as image_scale_corr = LP*QE*flight / partiality with
the predictor's per-frame partiality. Ingest then recomputes every partiality from the
frame-order-smoothed mosaicity (and the one exact-Bragg angle per rocking event), but corr_ingested
kept the predictor's 1/p. Where partials are scaled every fitted frame's corr is rebuilt from the
recomputed partiality, so that path only saw it in the first reference. Where they are not - fewer
than 50 rocking events per frame: small molecules, and long-wavelength protein sweeps - corr is what
the 3D combine divides each partial by, and the full became a weighted mean of I/p_pred while the
event's recomputed fractions summed to 1.00: the combined fulls tracked 1/sum(p_pred) (r = -0.96)
and scattered 13% rms about the plain sum of the same partials.

corr_ingested is now rescaled by p_pred/p_recomputed once at ingest.

Targeted battery against rc174-cand: every set on the partial-scaling path is bit-identical
(5reo, 9qw8, lyso_x06da_ref, thau_x10sa_0p1deg, lyso_x06da_atten_wedge, myob_x06da_split,
cytc_x06da_1, insu_I_x06da_5keV/6keV). Small molecules, SHELXL R1 against the COD model:
aspirin 20 keV 0.0515 -> 0.0432 (ISa 10.3 -> 30.9), aspirin 25 keV 0.0451 -> 0.0396, HEPES
0.0452 -> 0.0420, citric acid 0.0747 -> 0.0726, dnba 0.0472 -> 0.0273, metformin 0.0405 -> 0.0326,
nidppe 0.0498 -> 0.0426, lalanine 0.081 -> 0.057, cytidine 0.077 -> 0.071. Long-wavelength
proteins on the fulls-alone path: lyso_x06da_5keV ISa 17.9 -> 27.5, thau_bl1a_3p8keV 21.7 -> 28.4,
thau_bl1a_4p6keV 26.4 -> 33.5. One regression: lcystine_x10sa_25keV (a polycrystalline aggregate)
now refuses 622 and reports P31 instead of P6122.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-04 15:21:20 +02:00
leonarski_fandClaude Opus 5.5 4979f8aa68 French-Wilson: anisotropic Wilson prior from the fitted anisotropy tensor
The French-Wilson prior of each reflection is now epsilon * K_shell * a(h), with
a(h) = exp(-1/2 s^T B s) from the deviatoric tensor AnalyzeAnisotropy already fits
(the form it is fitted in) and K_shell = sum(I/eps) / sum(a), so a shell's priors still
average to its measured mean. The amplitudes are made isotropically at the merge as
before and made again once the tensor exists (full pipeline and --mode scale). Only
F/SIGF and F(+)/F(-) change; IMEAN/I(+)/I(-) are bit-identical. Applied whenever a
tensor was fitted, with no detection gate: a near-isotropic tensor gives a(h) ~ 1 and
the isotropic prior back, and the prior wants the best estimate of <I> along h whatever
its cause. Follows ctruncate's anisotropic prior (Ballard & Stein, CCP4); credit in
ACKNOWLEDGEMENT.md, CPU_DATA_ANALYSIS.md and at the algorithm.

--model scaling (ModelScaling.cpp) was checked: k_overall + symmetry-constrained
anisotropic B + flat bulk solvent fitted on the working set, against the same FW F
written to the MTZ - as REFMAC/phenix.refine do. No change needed.

Evidence (REFMAC 10-cycle restrained refinement of the deposited model, R-free on
the depositor's free reflections shared by both data sets; base = rc174 processing,
same IMEAN):
  set   base    new     d          set   base    new     d
  9rcs  0.3475  0.3494  +0.0019    8qq7  0.4452  0.4503  +0.0051
  9yzk  0.3192  0.3171  -0.0021    9hs7  0.2898  0.2465  -0.0433
  6yqf  0.4642  0.4543  -0.0099    5nw5  0.3256  0.3206  -0.0050
  7n2s  0.3126  0.2998  -0.0128    6z8o  0.2927  0.2892  -0.0035
  6qaj  0.3390  0.3038  -0.0352    6moj  0.2756  0.2651  -0.0105
  6r72  0.3826  0.3797  -0.0029    7qij  0.3239  0.3140  -0.0099
  anisotropic sets: median -0.0075, mean -0.0107, 10/12 better
  isotropic controls: 5reo -0.0014, 7kcn +0.0003, 6fid +0.0002, 11if 0.0000
rugnux's own --model R-free moves the same way (median about -0.019; 5nw5 +0.006),
R_model shell-scaled too; dep_cc_delta unchanged (intensity based). The adoption rule
(median gain >= 0.005 on the anisotropic sets, no control worse than +0.002) is met.
The two sets that lose are the one with a FLAT resolution signature (8qq7) and 9rcs,
where the exp form drives the dead direction's prior to ~0 beyond 3.7 A.
For scale: ctruncate's own anisotropic prior on the same merges moved the same
REFMAC R-free by a median of only -0.0008 (9hs7 +0.026).
Inhouse lyso_x06da_ref, thau_x10sa_0p1deg: every battery metric unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-04 14:24:14 +02:00
leonarski_fandClaude Opus 5.5 495f6a97b3 Rotation scaling: penalised per-frame scale of the fulls; fulls-only partials get flux and partiality
Three defects in the per-frame scaling of sparse (small-molecule, weak) rotation sweeps:

1. The fulls' per-frame scale pooled sparse frames by their RAW full count, but in a
   high-symmetry group with many systematic absences most usable counts stayed under the
   minimum, so most frames were never fitted and kept corr = 1 beside pinned, fitted frames
   (two gauges in one reference). The release step then read the gauge offset (a constant
   134x on a cubic Ia-3d small-molecule sweep) as signal and gave each frame exp(kept_f * 4.9) - the e^-5
   errors on a quarter of the frames, R1 0.62. A fixed box window also cannot follow a
   100x absorption ramp over a few degrees, and frames under the credible floor, exempt
   from the window, ran away to 1e-7.
   Now each round fits every frame on its own fulls (no pooling, no minimum), and the scale
   is a penalised second-difference smoother of log G (Whittaker/Eilers), each frame at
   the information of its fit, lambda by cross-validation over blocks one rocking curve
   wide (interleaved single frames leak through shared rocking curves and chose to follow
   every frame). After convergence the existing ShrinkToRestrained hands back the per-frame
   deviation its neighbour shares. Pooling and the box window are gone from the fulls loop;
   the partials loop is unchanged.

2. With partial scaling off (< 50 rocking events per frame) the partials kept the
   integration-time corr: no incident-flux correction and not the partiality of the
   ingest-smoothed geometry, because only the partial scaling loop rewrote corr. They now
   get corr = prescaling_corr / partiality at G = 1 (host, and a device kernel).

3. The flux meter (per-frame mean background) jumped 30x between neighbouring frames of a
   sparse sweep - on a few reflections it measures which reflections the frame holds. It
   is read through the same smoother at the precision of each frame's mean.

SHELXL R1(>4sig) against the published structures, rc174-cand -> this, on the in-house
small-molecule sweeps: cubic Ia-3d 0.615 -> 0.119 (XDS 0.088); four organic sweeps
(monoclinic / orthorhombic) 0.0515 -> 0.0459, 0.0451 -> 0.0396, 0.0747 -> 0.0709,
0.0452 -> 0.0413; ISa up to 11.6 -> 31. Raw flux instead of smoothed costs 0.002-0.003 R1 on
the first two. A weak, decaying protein sweep on the fulls-only path: ISa 6.7 -> 27.1.
CPU and GPU paths agree.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-04 14:18:08 +02:00
leonarski_fandClaude Opus 5.5 3f89759cb5 Space group: screw-absence bound 20 -> 8 nats, recalibrated on the zones
Mapped every single-axis screw candidate of every primitive-lattice set in
battery 3 (open + in-house) and the private arm against the deposited group.
False zones with no violation read at most +4 nats (plus one pseudo-
translation row at +30.5 that no bound separates); true zones go down to
+10. The true zones under 20 are all short monoclinic rows (b ~ 25-30 A:
four to six 0k0-odd reflections inside the search's resolution range, on
weak data at a few percent of their row), refused at 20 on five crystals
that are P2_1: four myoglobin sweeps (10.0-16.3 nats) and 6cs9 (10.8).
The rc174-cand myob_x06da_split call sat at 21.1, one refit away from P2.

Re-scoring the stored candidate tables at the new bound changes only those
five plus 3r6o (I4_1 2 2 newly eligible, toward the deposited I4_1 screw);
no adopted group elsewhere. New test section pins both sides: four 0k0-odd
at 2% of their row are claimed (11.3 nats), at 20% they are not.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-04 11:56:32 +02:00
leonarski_fandClaude Opus 5.5 15fb2543ca Revert "RotationScaleMerge: refit the error model about the merge's own mean"
This reverts commit 1baf92606. The refit fixed one small-molecule set
but cost a split crystal its screw axis and moved ISa down across the
battery; the rc173 error model stays.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-04 11:26:58 +02:00
leonarski_f 86ca86f1d1 Merge branch 'hkl-unmerged-fulls' into overnight-cint 2026-10-04 03:41:26 +02:00
leonarski_f 9bd92962c0 Merge branch 'overnight-cand' into dq-integration 2026-10-04 02:02:15 +02:00
leonarski_fandClaude Opus 5.5 8394b5988d Integration follows the measured spot footprint where a spot outgrows the r1 disk
The integrator's r1 disk and r2..r3 background ring are fixed in pixels and chosen from spots near
the beam. On small-molecule data at 20-25 keV a spot's standard deviation grows from ~1 px near the
beam to ~5 px at the edge (radially from parallax/obliquity, tangentially from the crystal's
azimuthal spread), so the r1 = 4 disk holds a quarter of the flux there, the background ring a third
of it, and the in-disk second moments the Gaussian is built from saturate near r1^2/4. On top of
that, the profile/summation runaway guard sent 20-30% of these reflections - the strong, wide ones -
back to the truncated r1 box sum.

- SpotFootprint: every pre-scan spot (width frames) is measured with a window that follows it
  (3 sigma, iterated, re-centred), radially and tangentially; the medians per distance-from-beam bin
  become BraggIntegrationSettings::Footprint. Installed only where some bin outgrows r1, and on the
  adaptive side like the radius (pre-pass without; the starvation guard falls back to the settings
  without it).
- BraggStencil: where 3 sigma > r1 the background ring starts at 3 sigma along and across the radius,
  the summation region is the r1 disk plus the 3-sigma footprint ellipse (so the guard's fallback is a
  complete intensity), and the per-reflection Gaussian takes the footprint widths. Compact spots keep
  the stencil bit for bit. Both engines build it from the same header.

SHELXL against COD (R1 / fixed-XDS-model R1(F)): citric acid .101/.230 -> .077/.055, HEPES
.070/.179 -> .048/.050, aspirin 20 keV .059/.070 -> .052/.061, aspirin 25 keV unchanged, L-cystine
25 keV unchanged (.145 -> .144).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-04 02:02:15 +02:00
leonarski_fandClaude Opus 5.5 9874db5d7f rugnux: .hkl holds the unmerged scaled fulls on rotation data
The SHELX HKLF 4 file now has one record per full reflection - its partials summed,
the per-frame scale and every correction applied, sigma(I) as the merge weighted it -
at the index it was measured at, not averaged with its equivalents: the chemical
crystallographer's convention, so SHELXL computes Rint and Rsigma itself. Outliers
the merge rejected and fulls beyond its resolution cut are left out; no batch column
(it would select a BASF scale in SHELXL). The engine hands the fulls back only for a
merge that may be written (RotationScaleMerge::SetExportScaledFulls), so the search
merges, the pre-pass and the P1 cross-check carry no copy; the fulls follow the same
relabelling as the merged reflections (merge_to_written). Stills keep the merged
file. --mode scale writes the unmerged form too.

Validation: p.mtz md5 unchanged on myob/cytc/thau (GPU). SHELXL on the same runs,
merged-old vs unmerged-new (COD models, harness /data/tmp/sm_shared):
aspirin 20 keV Rint 0 -> 0.071, Rsigma 0.036 -> 0.043, R1 0.0964 -> 0.0969,
wR2 0.312 -> 0.310, GooF 1.53 -> 1.47; HEPES 20 keV Rint 0 -> 0.163, Rsigma 0.057
-> 0.068, R1 0.0903 -> 0.0897, wR2 0.318 -> 0.263, GooF 1.57 -> 1.10; the "input
data appear to be merged" warning is gone.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-04 01:16:50 +02:00
leonarski_f 236cb49a99 Merge branch 'dq-sm-errormodel' into overnight-cand 2026-10-04 00:57:01 +02:00
leonarski_fandClaude Opus 5.5 1baf926066 RotationScaleMerge: refit the error model about the merge's own mean
The error model was fitted once, on deviations from a mean that weighs every full by its COUNTING
variance, and the outlier test's median took the same weights. Where equivalents disagree beyond
counting statistics, the low-count observations dominate that centre: on a strongly absorbing crystal
(YAG, Ia-3d, equivalents spread over a factor 100 after scaling) the mean of (0 4 0) sat at 21k among
observations from 11k to 1.9M, a ran into its bound (100), b came out at 800% internally, and the
six-sigma test about the biased median removed 63% of the observations - the strong ones.

Now, after the first fit, the model is refitted about the model-weighted mean (each full weighted by
the variance the fitted model gives it at the reflection's mean) until a and b settle, and the
rejection median takes the same weights. Where counting statistics are right nothing moves. Following
Blessing (1997) J. Appl. Cryst. 30, 421-426. GPU path: the refitted means are uploaded (SetEmMean).

Measured (SHELXL R1(>4sigma) against the COD model, sm-a's harness; rc174 scaling):
  YAG 0.556 -> 0.127 (XDS 0.083 merged), rejected 8055 -> 53, normalised deviations calibrated
  (median |z| 0.62-0.69 in every intensity decile); aspirin 20 keV 0.0964 -> 0.0958; citric acid
  0.161 -> 0.159; HEPES 0.0903 -> 0.0899; L-cystine 25 keV 0.1425 -> 0.1456;
  aspirin 25 keV 0.094 -> 0.106 (fixed-model R1 0.107 -> 0.173): its strong equivalents split into two
  frame-dependent populations from the per-frame partial scaling (sm-a's dq-smallmol), which the old
  under-sized sigmas happened to cut; with that scaling fixed (f69339ce6 + pooling, --no-scale-partials)
  this change is neutral to better on every small molecule (aspirin 20 .0618 -> .0586, aspirin 25
  .0456 -> .0454, citric .1093 -> .1006, HEPES .0703 -> .0697, YAG .649 -> .222; SHELXL GooF ~1.1).
  => ship together with the scaling fix.
Proteins (GPU full runs): CC1/2 and R_meas unchanged to 0.002; ISa myob 9.06 -> 8.20, thau 52.5 -> 47.5,
cytc 25.8 -> 25.4, lyso 29.4 -> 29.4 (still above XDS's 5.2 / 44.5 / 31.8 / 28.3 except cytc).
CPU build gives the same statistics as the GPU build on aspirin 20 keV and myob.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-04 00:48:35 +02:00
leonarski_f 45aa32d1d9 Merge branch 'dq-3r6o-twin' into overnight-cand 2026-10-04 00:43:34 +02:00
leonarski_f eb25c47db6 Merge branch 'dq-smallmol' into overnight-cand 2026-10-04 00:43:34 +02:00
leonarski_f ad6ec82c39 Merge branch 'dq-5cc8-screw' into overnight-cand 2026-10-04 00:43:34 +02:00
leonarski_f 88e773ee85 Merge branch 'dq-flag-stability' into overnight-cand 2026-10-04 00:43:34 +02:00
leonarski_fandClaude Opus 5.5 ba92417210 RotationScaleMerge: per-frame scale from the fulls alone on sparse sweeps (data-decided)
Replaces the prototype --no-scale-partials switch with a rule read off the data: a merge whose
sweep holds fewer than 50 rocking events per frame takes its per-frame scale from the fulls alone
(scale-fulls, a sparse frame fitted over its neighbours - PoolHalfWidth), and logs that it did.
The rocking-event walk already run for the smoothing window now also returns its event count.

Within one rocking curve a partial's scale and an error of the partiality model are the same
thing; on a sparse fine-sliced sweep the partial fit takes the one for the other and imprints an
hkl-dependent bias common to all equivalents. Measured populations: small-molecule sweeps 2.6-23
events per frame, protein sets 84-900; on the proteins the fulls-only scale leaves model R-free
unchanged (+-0.002 over the smoke tier's open-arm sets) but lowers ISa, so they keep the partial
scale. p.mtz md5 unchanged on the three profiling sets, GPU and CPU builds.

SHELXL against the COD structures, R1(>4sig) rc174 -> this: aspirin 20 keV 0.096 -> 0.062,
aspirin 25 keV 0.094 -> 0.046, citric acid 0.161 -> 0.109, HEPES 0.090 -> 0.070 (XDS 0.030-0.038);
SHELXL's weight a comes off its 0.2 cap on all four. CPU and GPU paths agree.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-04 00:34:39 +02:00
leonarski_fandClaude Opus 5.5 20ab92fcab Build: no FMA contraction; first-pass cell refinement independent of SIMD width
The same source built with and without -march=x86-64-v3 gave different results on battery sets
(9zmu axis-harmonic supercell arbiter fired in one build only; 8rud resolution cut 1.69 vs 1.70 A;
myob_x06da_powder_1 cut 1.07 vs 1.42 A; 7mzt short-axis first pass; lcystine_x10sa_20keV indexed
vs no lattice), against the rule that no decision may depend on compiler flags.

Two mechanisms, found by building the merged tree four ways (baseline, x86-64-v2, x86-64-v3,
x86-64-v3 -mno-fma) with and without -ffp-contract=off and comparing p.mtz:

- FMA contraction. GCC contracts a*b+c whenever the target has FMA. With -ffp-contract=off the
  x86-64-v3 build gives a p.mtz byte-identical to the baseline build on 8 of 9 sets (7mzt, 8rud,
  9zmu, insu_I_x06da_5keV_2, myob_x06da_powder_1, myob_x10sa, cytc_x10sa, thau_x10sa_16keV), on
  both the GPU and the CPU build. Cost: none measurable (user core-s, CPU build, x86-64-v3 vs the
  same with -ffp-contract=off: 1877/1874, 3285/3253, 2212/2195 on myob/cytc/thau; GPU likewise
  within noise). Set project-wide for C, C++ and CUDA host code; MSVC does not contract under
  /fp:precise.

- SIMD width. lcystine still differed: baseline and x86-64-v2 (128-bit) agreed, x86-64-v3 with or
  without FMA (256-bit) disagreed - Eigen's HouseholderQR in the FFT indexer's candidate refinement
  (PostIndexingRefinement.cpp) reduces column norms over all spots in packets of the target width.
  Replaced by the 3x3 normal equations summed in spot order in double. All four builds now agree on
  all nine sets.

Changes results of the default x86-64-v3 build (contraction off); lcystine_x10sa_20keV now gives no
lattice in every build (its first pass is a knife-edge: 0/60 vs 9/60 validation frames before).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-04 00:05:36 +02:00
leonarski_fandClaude Opus 5.5 f69339ce63 RotationScaleMerge: fulls-only per-frame scaling (prototype), pooled over sparse frames
Prototype switch --no-scale-partials: skip the per-frame scaling of the partials and let the
fulls (scale-fulls) carry the per-frame scale, as XDS does. Without partial scaling a frame whose
fulls number fewer than MIN_REFLECTIONS is fitted over the nearest frames on either side that
together hold enough (PoolHalfWidth), on the host and in the GPU kernel. Default path unchanged
(p.mtz md5 identical on the three profiling sets).

Why: on fine-sliced small-molecule sweeps the per-frame partial scale and the partiality model are
degenerate within a rocking curve, and the fit swings G 0.23..1.2 with a 180 deg period (XDS's own
frame scale: 0.79..0.99). That imprints an hkl-dependent bias common to all equivalents, which
R_meas/CC1/2/ISa cannot see but a refinement against the known structure does. And scale-fulls
never fitted a frame on such data: a full is filed under one frame, about 8 per frame, below
MIN_REFLECTIONS, so every frame kept G = 1.

Measured with SHELXL refining the COD structures (R1 >4sig), default -> switch:
aspirin 20 keV 0.096 -> 0.062, aspirin 25 keV 0.094 -> 0.046, citric acid 0.161 -> 0.109,
HEPES 0.090 -> 0.070 (XDS 0.030-0.038). Proteins lose ISa with the switch (myob 9.1 -> 7.6,
cytc 25.8 -> 13.7), so it is not a default; the choice is to be made from the data.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-03 23:08:12 +02:00
leonarski_fandClaude Opus 5.5 42441ac328 Twinning report: warn when a promotion stands on intensities below a perfect twin's; flag compressed twin-immune zone controls
A promoted point group whose <|L|> reads below 0.375 (NOT_READABLE) previously
raised no TWINNING warning at all; a twinned subgroup whose law the promotion
absorbed predicts the same data, so the adopted group is not confirmed. Warn.

The twin-immune zone control reading more compressed than a perfect twin's
acentric population (0.541) is something no twin fraction produces (overlap or
neighbour correlation); genuine symmetry then reads acentric in its zones too
(measured: a genuine 622 with its control at 0.528 read its 2-folds at -950 to
-2640 nats). Such zones are now marked ambiguous in the text and the zone
decision line. Report-only: no decision and no output file but the report changes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-03 21:40:51 +02:00
leonarski_fandClaude Opus 5.5 da5ef0c503 SearchSpaceGroup: offer every setting of the chosen rotation set, not only on a near-exact cell
The non-reference settings of the adopted point group were offered only when the cell hosted their
rotations to within ~0.1 deg, while the reference settings - which hold the very same rotations -
were offered without asking. A cell refined free after integration a few tenths of a degree off 90
therefore lost every setting but the reference one. On an orthorhombic set whose measured screws
lie on a and c and whose b row was never recorded, that left P2(1)2(1)2(1) as the only candidate
covering both screws, and it was reported as determined. With P 21 2 21 offered, the two tie, the
b-axis screw is reported as undetermined, and the model check uses the setting the data describe.

The point-group stage still asks the cell whether a rotation set it adds is hosted; only the
setting enumeration within an already chosen set stops asking.

Validation: myob/cytc/thau x10sa p.mtz byte-identical (GPU); 7mzt fail -> unscored (b screw
undetermined, P 21 21 21 or P 21 2 21); 5cc8 unchanged; [SearchSpaceGroup] 23 cases pass incl. a
new section with a cell 0.2-0.3 deg off 90.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-03 21:40:42 +02:00
leonarski_fandClaude Opus 5.5 0f728ecabf rugnux: make the tail's two P1 merges on the run's own engine again
The all-observation arm of the space-group search and the P1 cross-check
were made on a second RotationScaleMerge engine beside the run's own
(871347b7a). That engine holds a second device copy of every observation,
and on the largest sets the two no longer fit a 16 GB card: 8a1a, 8qaw and
8tyy (55-125 M partial observations) stopped with an out-of-memory error
in scaling. Measured on a quiet box the engine bought 0.7 s (cytc) and
1.0 s (thau) of tail and nothing on myob, which does not justify a memory
budget, so it is removed and both merges run on rsm in sequence, as
before 871347b7a.

Kept from 871347b7a: the GPU scaling's own non-blocking stream, the
cross-check not writing per-frame G/CC/mosaicity back (the per-image table
still describes the merge that was written), the restored scaling
iteration counts, and the anisotropy analysis beside the other analyses.

p.mtz byte-identical to before on myob/cytc/thau, GPU and CPU builds;
p_plot.txt unchanged apart from the GPU bkg column.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-03 20:32:59 +02:00
leonarski_f 581e1c8ca2 Merge branch 'perf-f-gpu-surfaces' into perf-merge 2026-10-03 13:55:04 +02:00
leonarski_fandClaude Opus 5.5 eb1e6fa410 RotationScaleMerge: correction-surface fit passes on the GPU, same bits
ApplyCellSurface (detector modulation, time x detector, crystal-frame SH and
goniometer-frame absorption) spends most of its time in two passes per round:
the per-group reference sums and the per-(block, cell) fit sums. With a GPU
both now run on the device (RotationScaleMergeGPU::Surface*) over the same
terms in the same order:
- reference: one thread per ASU group, walking a group-order permutation of
  the terms in fulls order;
- fit: each subset cut into the host's reduction blocks with every block's
  terms ordered by cell (stable counting sort, host); a per-term kernel forms
  w*Is, w*Iref, Iref and one thread per (block, cell) runs the two fma chains;
  the block slots are added per cell in block order.
Every rounding is spelled out (__dmul_rn/__dadd_rn/fma) to be the one the host
build makes: GCC at -march=x86-64-v3 fuses swI's multiply-add only in the
parity-filtered copy of the reference loop, and both fit sums.
Host side, exact on both paths: the 19 serial nth_element selections of the
shell edges become one parallel sort (same order statistics), the per-term shell
lookup runs on all threads, and the gate's per-shell CC is one walk over the
groups instead of one per shell.

Exact: p.mtz md5 identical to the oracle on myob/cytc/thau x10sa, GPU build
(all CUDA architectures) and CPU build. CorrectionSurfaceGPU test checks the
device sums bit for bit against an explicitly rounded host loop.
Measured (cytc/thau, two interleaved A/B pairs, box at load 13-25 from sibling
work): ApplyCellSurface host core-seconds -83% on cytc; SG adoption -> writing
reflections 4.58->3.75 and 3.85->2.54 s (cytc), 3.15->2.28 and 2.29->2.03 s
(thau); RSM final merge -0.7..-0.9 s and P1 cross-check -1.5..-1.8 s on cytc.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-03 13:35:43 +02:00
leonarski_f d89efeed65 Merge branch 'perf-b2-gpu-prescan' into perf-merge
# Conflicts:
#	image_analysis/geom_refinement/CMakeLists.txt
2026-10-03 13:34:29 +02:00
leonarski_fandClaude Opus 5.5 f849e2d1be Pre-scan on the GPU: background beam-centre walk and beam-stop mask
The two pre-scan steps that were still CPU-bound in a GPU build now run where the
projection already is.

- FindBeamCenterFromBackground: the per-iteration binning pass and the two clip rounds
  run on the device (BeamCenterBackgroundGPU); the fit itself stays on the host. Each
  cell is summed in the host's order (pixel order within the host's row blocks, blocks
  in order), and the per-pixel cell/derivative formula is shared (BackgroundBand.h).
  The angles come from BackgroundAtan2 (IEEE ops only) instead of atan2f, and both
  translation units are compiled without FMA contraction, so host and device give the
  same bits: 0 of 6.5 M pixels in a different cell, identical walks on the three
  in-house rotation sets. With glibc/CUDA atan2f and default contraction ~30 pixels per
  16 Mpx sweep changed cell and the fitted centre moved by up to 0.05 px.
- ShadowFinder::GetMask: the whole mask (pooling, ring medians, components, morphology,
  hole fill, arm search) runs on the device from ShadowAccumulatorGPU's projection
  (ShadowMaskGPU), so the 360 MB projection no longer comes back; the mean projection is
  divided on the device too (same bits). The two small fits over rings and sectors
  (BlockedOutTo, HarmonicFit) are shared with the host path in ShadowFinderInternal.h.
  Integers, comparisons, sorts and components are exact; the polarization trig, the
  Poisson log and the arm-search azimuth are not, so a pixel at a threshold can differ.

The one-time change against the previous CPU arithmetic (BackgroundAtan2, no
contraction), measured on the myoglobin, cytochrome C and thaumatin rotation sets:
ring centre moves 0.002-0.045 px (fit sigma 0.75-1.2 px), beam-centre capture
0.01-0.04 px; beam-stop mask differs on 31 / 144 / 53 pixels of 259k / 144k / 198k
(25 of the myoglobin ones are GPU-vs-CPU arithmetic in the mask, the rest follow the
centre); hot-pixel mask identical. Spot width, integration radii, bandwidth, beam-centre
arbitration, indexing, space group, cell, resolution and the merged statistics table
are identical; only the error model moves in its 4th digit. CPU build: the same
centres and decisions.

Timing (GPU, box at load 30-38): ring walk 0.54 -> 0.23-0.27 s, mask 1.24-1.44 ->
0.18-0.22 s, beam-centre capture walk 1.1-1.3 -> 0.31-0.35 s.

Tests: ShadowFinder_DeviceMaskMatchesHost, BeamCenterFromBackground_DeviceMatchesHost
(bit-exact), plus [ShadowFinder], [BeamCenter], [HotPixelFinder].

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-03 12:19:05 +02:00
leonarski_f 4c0b9ff84d Merge branch 'perf-d-azint-kernel' into perf-merge 2026-10-03 11:49:30 +02:00
leonarski_fandClaude Opus 5.5 03dee2dda7 AzIntEngineGPU: four pixels per thread, one shared add per ring run
The standalone GPU azimuthal integration did three shared-memory atomics per
pixel on the same few ring addresses, which is what it was limited by. It now
reads four pixels per thread as vector loads and keeps a running total per ring,
flushed when the ring changes - the scheme the adaptive finder's ring pass
(reduce_rings_shared) already uses. The npix % 4 leftovers are done one at a time.

Used wherever the fused adaptive engine is not (fixed-threshold spot finding, the
broker's non-adaptive path). Measured on a 16 Mpx sweep (1800 frames,
--no-adaptive-spots, RTX 5080): 843 -> 295 us per call (min 621 -> 196 us).

Not bit-identical, and the old kernel was not either: float atomics arrive in any
order, so two runs of the OLD kernel already differ by up to 1.7e-6 relative in
the per-frame profile; new vs old differs by up to 1.9e-6, the same order. Per-ring
pixel counts are identical. Default rugnux runs do not reach this kernel (p.mtz
md5 unchanged on three sets); on the fixed-threshold path p_unmerged.mtz is
md5-identical to the old kernel's. New test: GPU vs CPU engine on a pixel count
that is not a multiple of four, with masked and saturated pixels.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-03 11:43:12 +02:00
leonarski_f 8ab9284dad Merge branch 'perf-e-tail-overlap' into perf-merge 2026-10-03 11:20:46 +02:00
leonarski_f 2a945f3fbc Merge branch 'perf-a-cpu-kernels' into perf-merge 2026-10-03 11:20:46 +02:00
leonarski_f f3dd64e8ed Merge branch 'perf-b-prescan' into perf-merge 2026-10-03 11:20:46 +02:00
leonarski_fandClaude Opus 5.5 24e36ae740 Pre-scan: parallel beam-stop mask, leaner background beam-centre fit
Exact: p.mtz and the pre-scan products (shadow mask, mean projection,
defective-pixel mask, ring and capture centres, compared as hashes and
hex floats) are bit-identical to rc174 on three in-house rotation sets,
GPU and CPU builds.

- ShadowFinder::GetMask: the serial parts run in parallel - connected
  components by row band joined with union-find (both the shadow and
  the transmitting-arm searches, and the hole fill), ring binning and
  the harmonic sector gather by blocks, gap bridging by line; ring pixel
  counts read off the ring offsets. Mean projection filled in parallel.
- ShadowFinder host accumulation: one band-locked projection instead of
  a 20 B/px shard per pre-scan worker (2.7 GB zeroed and folded on a
  16M detector); SetShardCount and the shard argument are gone.
- FindBeamCenterFromBackground: the usable-pixel test is made once, the
  in-band pixels are kept in pixel order so the clipping rounds no
  longer sweep the whole detector, the 67 MB cell map is gone and the
  per-iteration block fold runs in parallel - same sums, same order.
- HotPixelFinder::GetMask: the chance-rate counts in parallel (integers).

Measured on a loaded box (load ~25 from other jobs), pre-scan window:
GPU 5.9-6.5 s -> 3.2-3.4 s, CPU 8.4-9.0 s -> 6.1-7.4 s. The GPU-build
pre-scan now ends with its background spot measurement (CPU spot finder
on ~120 frames, ~13 core-s on 8 workers).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-03 10:46:49 +02:00
leonarski_fandClaude Opus 5.5 871347b7ad rugnux: make the P1 merges of the tail beside the critical path
The tail of the canonical pass made five merges one after another. Two of
them read nothing the space-group search or the in-symmetry merge decides:
the all-observation arm of the search and the P1 cross-check. They are now
made on a second RotationScaleMerge engine, ingested beside the run's own
before any merge writes per-frame values back, and taken where they were
made before; where the run re-ingests (a reindex, a cell change) they are
made on the run's engine as before.

- RotationScaleMerge::SetWriteBackPerFrameScale(false) keeps a merge that
  is not the run's answer from writing G/CC/mosaicity onto the outcomes.
  The P1 cross-check no longer overwrites them, so _plot.txt's scale_G,
  cc_to_merge and cc_n now describe the merge that was written (in the
  determined group) instead of the P1 cross-check; the cross-check also
  no longer leaks its scaling iteration count into the report.
- RotationScaleMergeGPU runs on its own non-blocking stream instead of
  the legacy NULL stream, so the two engines (and a probe pass beside
  them) do not serialise at every launch and synchronisation.
- The anisotropy analysis (mostly ScaledObservations) runs beside tNCS,
  twinning and the other report-only analyses.

p.mtz, p_P1.mtz, p.cif, p.hkl and p_unmerged.mtz are byte-identical on
myob/cytc/thau (GPU and CPU builds). Tail on cytc GPU 8.1 -> ~6.8-7.7 s
under a loaded box.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-03 10:06:22 +02:00
leonarski_fandClaude Opus 5.5 56f0c5dd0a CPU spot finder and rotation prediction: same results, less work per pixel
AdaptiveSpotFinderCPU::AccumulateRingsBlock reads each pixel once for both the
ring histogram and the fused azimuthal profile (was two loops), and no longer
keeps the per-ring integer sums: they are taken from the histogram, as the two
sigma-clip passes already were (ClipRings(INFINITY)). Integer sums, so the same
totals; the profile's float sums keep their pixel order.

FlagRow is branch-free and works a 32-pixel word at a time, so it vectorises;
pixels outside every ring meet a +inf threshold in an extra ring_thr entry.

BraggPredictionRot::Calc takes A*h, A*h + B*k, C*l and 4*S0*S0 out of the inner
loops; p0 is the same ((A*h) + (B*k)) + (C*l) as before.

Measured on cytc (CPU build, both binaries run concurrently on a loaded box):
AccumulateRingsBlock -35%, FlagRow -47%, Calc + Coord ops -30% cycles; whole run
-5% cycles. p.mtz md5 unchanged on myob/cytc/thau, CPU and GPU builds.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-03 10:06:05 +02:00
leonarski_fandClaude Opus 5.5 a4bb6f74f1 XtalOptimizer: own Levenberg-Marquardt solver in place of Ceres
The crystal refinement (XtalOptimizer, both the seven-block and the reduced
beam+orientation form, and XtalOptimizerRotationOnly) no longer builds a
ceres::Problem. XtalRefine holds the problem as data and solves it with
LMSolver, which follows Ceres' trust-region LM step for step - Jacobi scaling,
damping and radius updates, stopping rules, box projection, the projected
Armijo line search with cubic interpolation on bounded problems, the
SphereManifold for the spindle - but takes J^T J and J^T r directly instead of
a Jacobian. The residual is the same XtalResidual code, now Ceres-free and
evaluated on a forward-mode Dual (Dual.h); everything that depends on
parameters alone (detector-angle trig, per-frame back-rotation, reciprocal
basis, orientation rotation) is worked out once per evaluation, and the
observed and predicted halves carry 6 and 9 derivative lanes rather than 16.
The sums are cut into blocks that depend on the residual count alone, so the
answer does not depend on the thread count. Because the line-search trial point
is the candidate point, a bounded iteration costs one evaluation instead of
Ceres' three.

Validation (rc174 + this, -march=x86-64-v3):
- p.mtz md5 identical to the Ceres build on myob/cytc/thau x10sa, GPU and CPU
  builds, and on the lyso8 stills reference.
- Solve corpus (every 16-parameter solve and every 10th per-image solve of the
  three sets, 8.3k problems, inputs and Ceres results dumped from a run that
  reproduced the md5s): usable/failed agree on all, iteration counts identical
  on all, parameters agree to <2e-11 (in px / rad / 0.01 A units), costs to
  1e-13.
- Same process, same threads: 7-9x faster per solve than Ceres.
- In-run (GPU, loaded box): xtal 16-parameter solves myob 22.2 -> 6.4 core-s,
  cytc 88 -> 26 core-s; per-image solves 5.4 -> 1.2 core-s (myob); cytc first
  pass indexing windows 1.9 -> 0.85 s, myob 1.1 -> 0.45 s; solver share of the
  whole cytc run 17% -> 4% of CPU samples.

New tests compare the solver with Ceres on synthetic rotation problems
(full/weighted/reduced) and check thread-count independence.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
2026-10-03 09:46:26 +02:00
leonarski_f 84228bf8be v1.0.0-rc.173 (#83)
Build Packages / Create release (push) Successful in 24s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m29s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m43s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 8m27s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 9m53s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 9m58s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 11m22s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 13m39s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 18m37s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 16m32s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 24m11s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m30s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 19m3s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 20m23s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 19m41s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 21m0s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m38s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 14m33s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 17m55s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 20m50s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 18m38s
Build Packages / Unit tests (push) Successful in 1h46m14s
* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports.
* jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls.
* Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results.
* Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable.
* Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate.
* Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do.
* Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence.
* Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags.
* Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check.
* Rugnux: Clear error messages when a data set needs more GPU or host memory than is available.

Reviewed-on: #83
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-09-29 15:57:32 +02:00
leonarski_f 6dfe065365 v1.0.0-rc.172 (#82)
Build Packages / Create release (push) Successful in 16s
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 8m27s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 9m15s
Build Packages / build:viewer-tgz:cpu (push) Successful in 10m11s
Build Packages / build:viewer-tgz:cuda (push) Successful in 12m6s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 15m44s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 16m1s
Build Packages / build:windows:nocuda (push) Successful in 17m29s
Build Packages / build:windows:cuda (push) Successful in 19m58s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 24m7s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 19m8s
Build Packages / build:rugnux:windows (push) Successful in 10m58s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 20m46s
Build Packages / Generate python client (push) Successful in 53s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 20m13s
Build Packages / Build documentation (push) Successful in 1m36s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 19m57s
Build Packages / build:rpm (rocky8) (push) Successful in 18m7s
Build Packages / build:rpm (rocky9) (push) Successful in 18m54s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 19m32s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 17m30s
Build Packages / Unit tests (push) Successful in 1h39m2s
* Fixed `jfjoch_broker` cancelling every data collection with a CUDA "out of memory" error after long operation: GPU memory no longer leaks with each collection.
* Rugnux scales a rotation sweep until the per-frame scales settle instead of for a fixed three rounds, and says so when they did not - merged intensities, and the space group, resolution cut and frame rejection read off them, change accordingly; `--scaling-iterations` is now the cap on that loop (default 100).
* Rugnux places every frame of a marCCD, SMV or miniCBF series at the spindle angle its own header states, so a series with missing frames, or with angles written modulo 360, is no longer read at the wrong geometry or refused.
* Every rotation run writes two diagnostic files beside its reflections: `<prefix>_detector.jpg`, the detector projection with the pixel mask and the detected beam-stop shadow drawn on it, and `<prefix>_plot.txt`, one row per image.

Reviewed-on: #82
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-09-22 06:48:37 +02:00
leonarski_f 30b6800289 v1.0.0-rc.171 (#81)
Build Packages / build:windows:nocuda (push) Successful in 17m20s
Build Packages / build:windows:cuda (push) Successful in 19m52s
Build Packages / build:viewer-tgz:cpu (push) Successful in 9m38s
Build Packages / build:viewer-tgz:cuda (push) Successful in 11m18s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 9m34s
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 5m42s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 20m33s
Build Packages / Create release (push) Successful in 33s
Build Packages / build:rugnux:windows (push) Successful in 12m0s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 15m42s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 14m59s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 16m8s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 14m35s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 16m55s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 16m58s
Build Packages / Generate python client (push) Successful in 16s
Build Packages / build:rpm (rocky8) (push) Successful in 15m21s
Build Packages / Build documentation (push) Successful in 54s
Build Packages / build:rpm (rocky9) (push) Successful in 16m23s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m2s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 10m6s
Build Packages / Unit tests (push) Successful in 1h10m26s
* Rugnux: basic support for CCD images (marCCD, SMV) and for gzipped miniCBF.
* `jfjoch_viewer`: opens the CCD formats, and fixes to the dataset plots.
* Documentation updates.

Reviewed-on: #81
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-09-17 14:42:52 +02:00
leonarski_f cb5a2f032a v1.0.0-rc.170 (#80)
Build Packages / Create release (push) Successful in 23s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 10m6s
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 9m6s
Build Packages / build:viewer-tgz:cpu (push) Successful in 11m15s
Build Packages / build:viewer-tgz:cuda (push) Successful in 12m21s
Build Packages / build:windows:nocuda (push) Successful in 17m9s
Build Packages / build:windows:cuda (push) Successful in 19m49s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m42s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 16m0s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 14m54s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 17m7s
Build Packages / build:rugnux:windows (push) Successful in 10m47s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 17m4s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 17m8s
Build Packages / Generate python client (push) Successful in 45s
Build Packages / Build documentation (push) Successful in 1m45s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 19m23s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 19m20s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 18m43s
Build Packages / build:rpm (rocky8) (push) Successful in 19m31s
Build Packages / build:rpm (rocky9) (push) Successful in 20m16s
Build Packages / Unit tests (push) Successful in 1h41m19s
* Fixed a `jfjoch_broker` crash during indexing: sorting no longer misbehaves on non-finite values, and GPU FFT indexer kernel launches are now error-checked.
* rugnux needs about a third less peak memory to scale, merge and post-refine rotation data, with identical results.
* `rugnux --model`: the placed coordinate file carries the space group its own coordinates obey, and says so when that is not the group the reflection files beside it carry.
* `jfjoch_viewer`: fixes in the dataset plots, inspector and layout; spot markers lose their black outline by default (a checkbox under "Image features" restores it) and the highest-pixel markers are white boxes around the pixel.

Reviewed-on: #80
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-09-16 18:17:46 +02:00
leonarski_f 9aae0c2ba7 v1.0.0-rc.169 (#79)
Build Packages / Create release (push) Successful in 21s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 9m40s
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 9m49s
Build Packages / build:viewer-tgz:cpu (push) Successful in 11m37s
Build Packages / build:viewer-tgz:cuda (push) Successful in 12m40s
Build Packages / build:windows:nocuda (push) Successful in 17m44s
Build Packages / build:windows:cuda (push) Successful in 20m13s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 14m41s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m59s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 15m5s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 14m35s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 15m53s
Build Packages / build:rugnux:windows (push) Successful in 11m29s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 18m51s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 18m43s
Build Packages / Generate python client (push) Successful in 51s
Build Packages / build:rpm (rocky8) (push) Successful in 18m51s
Build Packages / Build documentation (push) Successful in 1m21s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 18m38s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 18m24s
Build Packages / build:rpm (rocky9) (push) Successful in 19m19s
Build Packages / Unit tests (push) Successful in 1h37m15s
* Building Jungfraujoch no longer needs zlib or Eigen installed on the machine, and the dependencies the build fetches are pinned and updated to current releases.
* rugnux: improvements in indexing, lattice selection and geometry post-refinement, which index crystals that previously returned no lattice and keep the better of the two geometries a run measures.
* rugnux: improvements in beam-centre measurement, beam-stop detection and space-group determination.
* rugnux: the unit cell reported with a determined space group now obeys that group - a cell whose symmetry was confirmed from the intensities is re-refined under it, and a cell the group cannot describe is reported with a warning rather than as it stands.
* rugnux drops the stretches of a rotation sweep whose removal measurably improves the merged intensities and reports what became of every frame, and decides the resolution cut on the crystal's own diffraction rather than on its ice rings.
* The rugnux results report is machine-readable - every line that is not `KEY= value` data starts with `#` - and states the build it was written by, its authorship and its terms of use (`REPORT_VERSION= 8`).
* `jfjoch_viewer`: improvements in the file manager (CBF frames beside HDF5 datasets, a remembered root), the dataset plots, the inspector and the image statistics, plus a settable font size, a view of the rugnux results report, usable performance over a remote display (`ssh -X`) and a reset of all settings to defaults; the reciprocal-space window is removed.
* Broker fixes around DECTRIS collections and dark-mask calibration: re-initialising after a run that never started no longer freezes the broker, a cancelled calibration is abandoned instead of reported as done, and a collection whose start message never arrives ends by itself.

Reviewed-on: #79
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-09-15 17:09:31 +02:00
leonarski_f 77bc0cfe52 v1.0.0-rc.168 (#78)
Build Packages / build:viewer-tgz:cpu (push) Successful in 11m35s
Build Packages / build:windows:nocuda (push) Successful in 16m52s
Build Packages / build:windows:cuda (push) Successful in 20m24s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 18m38s
Build Packages / build:rugnux:windows (push) Successful in 10m38s
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 9m4s
Build Packages / build:viewer-tgz:cuda (push) Successful in 14m44s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 20m46s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 21m24s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 24m13s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 26m2s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 18m34s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 23m55s
Build Packages / build:rpm (rocky9) (push) Successful in 21m9s
Build Packages / XDS test (durin plugin) (push) Successful in 12m8s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m6s
Build Packages / build:rpm (rocky8) (push) Successful in 25m46s
Build Packages / Generate python client (push) Successful in 38s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 58s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 24m36s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m44s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m17s
Build Packages / DIALS test (push) Successful in 19m45s
Build Packages / Unit tests (push) Successful in 1h26m13s
* rugnux is substantially faster - a corpus of 145 rotation datasets processes in about two thirds of the time - with identical results.
* A crystal whose lattice looks more symmetric than it is because the beam centre is off is no longer processed on the wrong cell.
* rugnux prints at startup, and writes at the foot of every results report, a short acknowledgement of the X-ray research community whose methods it implements and of the open-source projects it builds on; `ACKNOWLEDGEMENT.md` now ships in every package beside `LICENSE` and `THIRD_PARTY_NOTICES.md`.

Reviewed-on: #78
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-09-10 13:51:16 +02:00
leonarski_f a39fd29f77 v1.0.0-rc.167 (#77)
Build Packages / XDS test (JFJoch plugin) (push) Successful in 11m4s
Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 17m46s
Build Packages / build:windows:cuda (push) Successful in 20m20s
Build Packages / build:viewer-tgz:cpu (push) Successful in 15m56s
Build Packages / build:viewer-tgz:cuda (push) Successful in 17m57s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 14m10s
Build Packages / build:rugnux:windows (push) Successful in 11m12s
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 7m14s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 22m13s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 19m17s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 21m21s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 17m26s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 23m56s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 20m48s
Build Packages / build:rpm (rocky8) (push) Successful in 23m43s
Build Packages / build:rpm (rocky9) (push) Successful in 20m38s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 24m57s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m58s
Build Packages / XDS test (durin plugin) (push) Successful in 10m43s
Build Packages / Generate python client (push) Successful in 47s
Build Packages / Build documentation (push) Successful in 1m5s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (neggia plugin) (push) Successful in 8m57s
Build Packages / DIALS test (push) Successful in 18m40s
* `rugnux --model` reports CC(model, data) - the correlation of the merged intensities with the placed, scaled model - by resolution shell, on the same shells as CC1/2, with the reflection count and a significance for each.
* `rugnux --model` fits the model's scale, anisotropic B and bulk-solvent parameters on the working reflections only, so the R-free it reports is measured against a model no free reflection helped scale.
* The bulk-solvent parameters of `rugnux --model` are searched over their physically meaningful range instead of being fitted without bounds, so a model is never scaled with a solvent term that has silently switched itself off.
* The rigid-body placement of `rugnux --model` uses the same bounded bulk solvent as the reported fit, so a model is no longer placed against a target carrying a solvent term with no physical meaning.
* `rugnux --model` puts the model into the data's own description of the lattice before placing it, so a model whose cell is written on other axes - I-centred where the run indexed C-centred, a different unique axis, a permuted orthorhombic cell - is placed rather than scored where it was read; `MODEL_CHANGE_OF_BASIS=` and `MODEL_SETTING_AS_READ=` report it when it happens.
* The rugnux results report opens with a summary - `VERDICT=` (`OK`, `WARNINGS`, `UNUSABLE`, `FAILED`), `VERDICT_TEXT=`, `PATHOLOGY_FLAGS=` with one closed-vocabulary code per condition that warned, and the `WARNING:` lines, which used to close the file - and the sections after it are renumbered 1-5 with no gaps.
* `rugnux --developer` writes the full results report - the pipeline-internal keys and the long explanations the default report now leaves out - and `--finalist-ledger` adds the evidence for every space group the search considered, not only the one it adopted.
* The results report warns when the merged data carry no usable signal and when too little of reciprocal space was measured inside the fitted resolution, and omits `FITTED_RESOLUTION` where the CC1/2 curve it is fitted on never falls off.
* rugnux detects translational pseudo-symmetry and reports it under the `PSEUDO_TRANSLATION` flag as `TNCS_DETECTED=` and the `TNCS_*` keys - a translation the merged data are exactly invariant under is reported as `UNDECLARED_LATTICE_TRANSLATION=` under `LATTICE_TRANSLATION` instead - and a detected pseudo-translation can no longer buy a false screw axis in the space-group search or hide a twin from the L-test (`L_TEST_VS_TNCS=`).
* The space-group search determines glide planes from zonal systematic absences, so a non-Sohncke space group such as P 2_1/c or Pbca is named where the run previously stopped at its Sohncke subgroup; `SOHNCKE_SPACE_GROUP=` carries the best Sohncke group beside it on every run that searched, and a centre of symmetry is never claimed.
* Where the cell metric carries more rotational symmetry than the Bravais class the indexer named, the extra rotations are put to the intensities and the space-group search is asked again on the metric's own cell - adopted only where the intensities confirm the higher symmetry - so a lattice that is nearly but not exactly hexagonal, or whose reduction landed in a sub-cell, still reaches its true point group.
* Systematic-absence calls rest on the evidence rather than on counts: a screw axis whose absent class the data show extinct is no longer refused because a handful of reflections in it read as present, and `SPACE_GROUP_ALTERNATIVES=` no longer drops a candidate that differs only on a zone the sweep never measured.
* A reference correlation measured on too few reflections is refused instead of scored zero, so a run given a reference MTZ is no longer reindexed on an operator that mapped almost everything outside the reference's coverage.
* A frame counts as indexed from 6 spots on its lattice rather than 9, so a weakly diffracting crystal whose frames cannot carry 9 is no longer refused the lattice it fits; `--min-indexed-spots` overrides it.
* `-C` accepts a known cell in any equivalent description - conventional or primitive, centred or not - instead of only the reduced primitive form, so a centred cell given the way it is published no longer makes the run report that it found no lattice.
* Each reflection is corrected for the sensor's quantum efficiency at the angle it meets the detector (attenuation lengths from the NIST tables, which also fixes the spot-width parallax term on CdTe) and for the attenuation of the flight path between the sample and its pixel; `--flight-path air|helium|vacuum` declares the medium - default air, since no file states it - and the report says what was assumed and what it was worth. The unmerged MTZ records the factors in new `QE` and `FLIGHT` columns beside `LP`, so raw counts are `I / LP * QE * FLIGHT`, and `_process.h5` in new optional `qe` and `flight` datasets.
* Rotation geometry post-refinement fits the crystal and the detector at once, against the observed spot positions and the observed rocking angles together, so the refined distance depends far less on how wrong the file's distance was.
* A coarsely sliced sweep integrates correctly: partials are joined into one rocking event by angle rather than by frame count, so two crossings of the Ewald sphere are no longer summed into one full, and at 0.5 degrees per image or coarser the per-frame geometry refinement accepts a spot whose miss the exposure's own rotation accounts for.
* `rugnux --mode scale` reports the detector tilt and direct beam of the geometry it re-scaled at, instead of zeros that read as a flat detector, and no longer warns that no image was indexed on a run whose lattice came from its input file.
* Every rotation run that determined a space group and merged reports what the mounting cost: `SPINDLE_LOST_UNIQUE_FRACTION=` is the fraction (0-1) of unique reflections the mounting made unmeasurable under the measured point group, also written to the master as `/entry/MX/spindleLostUniqueFraction` and what the mounting warning fires on; `SPINDLE_SYMMETRY_AXIS_ANGLE_DEG=` / `SPINDLE_SYMMETRY_AXIS_ORDER=` describe the mounting in the `--developer` report.
* Stills and grid scans carry a per-image `spindle_blind_fraction` - how much of a rotation sweep's blind cone this orientation would make unrecoverable, 0.5 and above calling for a second orientation - through the CBOR stream, HDF5 (`/entry/MX/spindleBlindFraction`), the plot and scan-result APIs, and the viewer and frontend plots; an absent value means the frame could not be assessed and is not a 0.
* The results report's `REPORT_VERSION` is 7.

Reviewed-on: #77
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-09-09 07:25:13 +02:00
leonarski_f 680c36c20d v1.0.0-rc.166 (#76)
Build Packages / Unit tests (push) Successful in 1h22m15s
Build Packages / build:windows:nocuda (push) Successful in 18m0s
Build Packages / build:windows:cuda (push) Successful in 20m30s
Build Packages / build:viewer-tgz:cpu (push) Successful in 10m32s
Build Packages / build:viewer-tgz:cuda (push) Successful in 11m39s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 8m55s
Build Packages / build:rugnux:windows (push) Successful in 11m25s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 20m6s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 16m27s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 20m19s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 15m34s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 20m25s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 19m36s
Build Packages / build:rpm (rocky8) (push) Successful in 17m43s
Build Packages / build:rpm (rocky9) (push) Successful in 13m34s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 21m28s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 18m19s
Build Packages / DIALS test (push) Successful in 12m36s
Build Packages / XDS test (durin plugin) (push) Successful in 6m56s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 6m48s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m7s
Build Packages / Generate python client (push) Successful in 11s
Build Packages / Build documentation (push) Successful in 36s
Build Packages / Create release (push) Skipped
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 5m11s
* `rugnux --mode calibration` writes `<prefix>.json` beside the `.poni`, whose `dataset_settings` member is a `jfjoch_broker` `dataset_settings` body as it stands.
* `rugnux` and `jfjoch_viewer` read PILATUS miniCBF sweeps natively, without conversion.
* Masters written by other facilities open, including Eiger 1.x and third-party NXmx variants.
* `rugnux` measures the beam centre on every run, and indexes with it when the file's value indexes nothing.
* A detector swung out on a 2theta arm is placed where the file says it stands, and the calibration can hold the tilt fixed.
* `rugnux` writes the unmerged MTZ by default, and a P1 merge beside it, so a wrong space group can be re-merged without reprocessing.
* Significant improvements to symmetry handling in `rugnux`: the lattice, the point group, the setting and the systematic absences.
* The `rugnux` report gives the resolution the CC1/2 fit reached, beside the range the reflections were written to.
* The `rugnux` report gives the twinning statistics measured before the space group was decided, beside the ones measured after.
* The `rugnux` report gives the strong-direction diffraction limit, and warns when CC1/2 is not monotone with resolution.
* `rugnux` ranks screw axes on the evidence their absences carry, rather than on how many control reflections a candidate happens to have.
* Twinning is no longer reported when the L-test contradicts it.
* The `rugnux` report gives the detector tilt, the measured tilt and the direct beam beside the beam centre, and a post-refined beam centre is judged against the run's own measurement rather than the file's.
* `--no-refine-tilt` holds the detector tilt at the value in the file, instead of zeroing it, when the calibration starts from the spots.
* The `jfjoch_viewer` grid scan view draws the cells in the proportion of the scan steps, so the map has the shape of the scanned area.

Reviewed-on: #76
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-09-02 21:17:31 +02:00
leonarski_f 511be0c366 v1.0.0-rc.165 (#75)
Build Packages / build:rpm (rocky8) (push) Successful in 24m0s
Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 16m54s
Build Packages / build:windows:cuda (push) Successful in 19m25s
Build Packages / build:viewer-tgz:cpu (push) Successful in 14m44s
Build Packages / build:viewer-tgz:cuda (push) Successful in 16m3s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 13m15s
Build Packages / build:rugnux:windows (push) Successful in 10m45s
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 9m34s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 19m7s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 18m9s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 24m48s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 18m13s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 24m51s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 22m58s
Build Packages / build:rpm (rocky9) (push) Successful in 21m23s
Build Packages / Generate python client (push) Successful in 1m2s
Build Packages / Build documentation (push) Successful in 1m23s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (durin plugin) (push) Successful in 9m45s
Build Packages / XDS test (neggia plugin) (push) Successful in 10m19s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 11m10s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 22m15s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 17m37s
Build Packages / DIALS test (push) Successful in 17m16s
* `rugnux --model` adopts the model's space group as a label where the data were merged in its enantiomorph, instead of reindexing the reflections - which swapped I(+) with I(-).
* `rugnux --model` warns, naming the atom, when the anomalous density at the model's atoms comes out inverted, which means the data and the model are in opposite hands.
* `rugnux --model` writes an anomalous difference map (`<prefix>_anom.ccp4`) when the merge kept the Bijvoet split, and names the ten model atoms it peaks highest on as `ANOMALOUS_SITE_01`..`_10`.
* `MEAN_ATOM_DENSITY_SIGMA` is read from the map by cubic rather than linear interpolation and comes out around a tenth higher; it is no longer comparable with the figure earlier versions printed.
* `rugnux --model` reads an mmCIF coordinate file as well as a PDB one, gzipped or not, taking the format from the file's content rather than its name.
* A model `rugnux --model` cannot use is reported as a `WARNING:` line in the results report instead of only in the log.
* The rugnux results report has a `10. MODEL VALIDATION` section when `--model` was given; `REPORT_VERSION` is 4, `WARNINGS` moves to section 11 and no existing key changed.
* The rugnux results report records how the run was invoked, what it cost and what it ran on: `COMMAND_LINE=`, `WALL_TIME=` and `GPU_COUNT=` / `GPU=`.
* rugnux says which GPUs it can see before it starts processing.
* `rugnux --export-unmerged` also writes `<prefix>_unmerged.mtz` on a `--no-merge` run, and is ignored on a run with no output prefix instead of writing a file called `_unmerged.mtz`.
* `/start` asks the writer whether the run can be written before the detector is armed, so a run whose master file already exists, or whose output directory cannot be created, is refused up front with the writer's own message. This needs the TCP image stream or the built-in HDF5 writer; the ZeroMQ stream is unchanged.
* A calibration that fails goes to `Error` carrying the reason instead of `Inactive`, so `/wait_till_done` and `/wait_until_running` report it; a cancelled calibration still goes to `Inactive`.
* `/wait_till_done` answers 500 with the message when a collection ended in an error. A cancelled collection and a collection that only triggered a warning still answer 200.
* A pending start failure is discarded by `/cancel` and `/deactivate`, as it already was by `/start` and `/initialize`.
* `/scan_result` no longer reports the previous run's images after a collection that failed to start, or after `/deactivate`.
* The TCP image stream protocol version is 4. `jfjoch_writer` and `jfjoch_broker` have to be of the same release, as before.

Reviewed-on: #75
2026-08-27 22:16:54 +02:00
leonarski_f 749db470ca v1.0.0-rc.164 (#74)
Build Packages / build:rpm (rocky9) (push) Successful in 19m56s
Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 16m57s
Build Packages / build:windows:cuda (push) Successful in 19m18s
Build Packages / build:viewer-tgz:cpu (push) Successful in 14m48s
Build Packages / build:viewer-tgz:cuda (push) Successful in 16m18s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 14m19s
Build Packages / build:rugnux:windows (push) Successful in 10m34s
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 8m49s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 20m55s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 17m4s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 20m48s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 19m15s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 24m26s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 20m32s
Build Packages / build:rpm (rocky8) (push) Successful in 23m39s
Build Packages / Generate python client (push) Successful in 46s
Build Packages / Build documentation (push) Successful in 1m45s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (durin plugin) (push) Successful in 11m3s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 11m30s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m10s
Build Packages / XDS test (neggia plugin) (push) Successful in 10m17s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 23m12s
Build Packages / DIALS test (push) Successful in 20m12s
* rugnux now tells you whether a crystal diffracts anisotropically and how far it reaches in each direction, without a second program: a new `9. DIFFRACTION ANISOTROPY` section in `<prefix>_report.txt` and matching `_reflns.pdbx_aniso_B_tensor_*` / `_reflns.jfjoch_aniso_*` items in the merged mmCIF report the anisotropic deltaB, the diffraction limit along each principal direction, and a `NOT DETECTED` / `DETECTED` / `CANNOT DETERMINE` verdict measured against the data set's own systematic error. It is a description only - no intensity is corrected, no reflection is removed, and the merged data do not depend on direction.
* rugnux can hand its integrated observations to another scaling program: `--export-unmerged` writes `<prefix>_unmerged.mtz`, an unmerged MTZ readable by aimless, pointless, careless and `iotbx.merging_statistics`, in `--mode mx` and `--mode scale` alike. Each rotation reflection's partials are summed into one full; `--export-unmerged-partials` writes one row per image instead. Intensities carry the Lorentz-polarization factor and nothing else, since those programs scale the data themselves. Lattice-centring absences are not written; screw and glide absences are.
* rugnux integrates crystals with broad spots better - where it changes anything, per-shell mean I/sigma improves by up to 31% and R_meas by up to 24% - because on rotation data the integration signal radius is now taken from the crystal's own measured spot width instead of a fixed 4 px. `--adaptive-integration-radius=off` restores the fixed radius and an explicit `--integration-radius` still overrides both. The widened radius applies to the final integration pass only, and a pattern too dense for it is re-integrated at 4 px with a note in the log.
* rugnux discards fewer stills reflections for want of a background ring, improving per-shell R_meas over most of the signal-bearing range: the stills background ring now runs to 14 px instead of 12. The gain reverses in shells below a mean I/sigma of about 4.
* rugnux determines the space group with thresholds that mean the same thing on a weak crystal as on a strong one: symmetry operators are scored on resolution-normalised intensities (E squared) instead of raw merged intensities, and a reflection counts as genuinely present on its counting significance instead of on the merged I/sigma, which saturates at the merge's own ISa. The search resolution cut is no longer able to move the answer, and the twin-law H bound moves from 1.70 to 1.85, which stops one class of correct high-symmetry assignment being refused as twinning.
* rugnux says what the space-group search tested and what it could not: the twin-law disagreement H is printed for every operator together with the adopted point group's H ratio and its bound; alternatives that are not on the reported lattice are named with how their cell differs; and a lattice centring the data could not test - the crystal having been integrated on the primitive sub-cell, so the reflections it extinguishes were never measured - is marked `UNTESTED` and warned about where it is adopted, as coming from the lattice metric rather than from the intensities.
* rugnux `--mode scale` re-merges a `_process.h5` in the right symmetry without being told it: the file now records the space group on every run - a two-pass rotation run wrote none before, so re-merging defaulted to P1 - together with the change of basis under `/entry/MX/reindexMatrix` where the lattice was re-seated, and `--mode scale` also reports the Wilson B-factor estimate instead of `WILSON_B= nan`. A file written before this stops with a message naming the two cells and the override to use, instead of failing inside the merge. A third-party reader of a `_process.h5` must apply `reindexMatrix` where it is present.
* rugnux installs on its own, as a package called `rugnux` - `dnf install rugnux` or `apt install rugnux` - instead of arriving inside `jfjoch-viewer`. It pulls in none of the acquisition stack, so a machine that only processes data no longer has to carry the broker, the detector libraries or Qt to get it. Installing it over a `jfjoch-viewer` from rc.163 or earlier, which still owns `/usr/bin/rugnux`, upgrades cleanly rather than failing on the duplicate file.
* rugnux is also a standalone download, built for arm64 as well as x86_64: `rugnux-<version>-linux-{x86_64|aarch64}-cuda<major>.tgz` and `rugnux-<version>-win64-cuda<major>.zip` on the release page, for machines that are not managed by a package manager. The aarch64 build targets GH200 and DGX Spark, and is untested on hardware.
* Every portable Linux binary is now a single self-contained file: cuFFT is linked statically instead of being shipped beside the executable and found through an rpath, so `rugnux` and `jfjoch_viewer` need nothing but an NVIDIA driver, and only to use the GPU. The `.rpm`/`.deb` continue to take cuFFT from the distribution. The developer utilities `jfjoch_extract_hkl` and `jfjoch_recompress` are no longer packaged anywhere.
* Jungfraujoch needs six fewer shared libraries on the machine - libopenblas and libmetis, and libgfortran, libquadmath, libgomp and libz behind them - because the Ceres LAPACK, METIS and SuiteSparse back-ends are no longer built. Nothing in the code ever selected them, and results are unchanged.
* The PCIe driver DKMS package builds for the kernel it is being installed for instead of the running one, so a module built while a kernel update is being applied loads after the reboot.
* The PCIe driver builds on RHEL 9.5 and later, and on their CentOS Stream, Rocky and AlmaLinux equivalents, where the `vm_flags` kernel interface was backported into the 5.14 kernel.
* A data collection started with `async_start` that fails to start - a writer refusing to overwrite an existing file, for instance - is reported as an error by `/wait_until_running` and `/wait_till_done` instead of as a timeout and a successful collection respectively. The error message is the one the writer gave.
* A calibration that is cancelled or that fails to collect its pedestals is no longer reported as a successful one. The broker goes to `Inactive` with an error message and has to be initialized again, instead of sitting in `Idle` looking ready to measure while holding partial pedestals - data collected in that state was silently mis-converted.
* A failed `/initialize` is reported to `/wait_until_running` and `/wait_till_done` as soon as it happens, instead of when their timeout expires.
* `space_group_number` accepts space groups up to 230 in the API schema, so cubic space groups can be recorded. The broker always accepted them; the generated clients rejected them before the request was sent.
* The results report's `REPORT_VERSION` is 3, two sections having been added. Existing key names and table columns are unchanged.
* The merged statistics table has **9** resolution shells instead of 10, which is what XDS reports. The bins were already XDS's - equal steps in 1/d^2 between the lowest- and the highest-resolution reflection the merge kept - so at the same resolution limits the two tables now have the same shell boundaries and can be read row for row. `--resolution-shells` sets a different count.
* `rugnux --model` now settles the frame the merged reflections are written in, not only the frame the R-factors and the maps are computed in: the `.mtz`/`.cif`/`.hkl` come out in the model's indexing, and where the data were merged in the model's enantiomorph they take the model's hand and space group - which on anomalous data puts I(+) and I(-) the right way round. The indexing choice is logged with the winning R-free and the runner-up, so a decision made within noise is visible.
* `rugnux --model` can resolve the indexing ambiguity of a **serial stills** run, which a model could not do before: structure factors computed from the model become the per-image reference, the same role a reference MTZ plays. It needs the cell and space group up front (`-C` / `-S`). Without one or the other, a merohedral serial run still merges both hands together and says so.
* The rugnux documentation opens with a quick start - the default run, and runs with a reference MTZ, with a model, or with the space group and cell pinned - and explains the indexing ambiguity: what it costs on rotation and on serial data, and which of `-z` / `--model` resolves it in each case. The long reference pages now carry a table of contents.

Reviewed-on: #74
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-08-26 22:47:00 +02:00