Non-Sohncke candidates were enumerated only when their PROPER rotations equalled the measured
point group's, which only centrosymmetric groups satisfy. Groups without a centre (Pc, Pna2_1,
I-42d, I4_1md, Fdd2, P-42_1c, I-43d) were never candidates: where a centrosymmetric group shares
their absences rugnux wrote it silently (Pc -> P2/c), and where none does the dead glide zone was
dropped and a Sohncke subgroup written (KDP-type I-42d data -> I4_122).
- Enumerate non-Sohncke groups by Laue class. The Sohncke-signature dedup, CellHostsRotations and
the glide-evidence bar are unchanged. Groups with identical absences in the same Laue class are
scored once and carried as `same_absences`; a selected candidate brings them into
`alternatives`.
- Convention for which twin is written: the centrosymmetric one where it exists (missed centres are
the common error, Baur & Kassner 1992), otherwise the lowest-numbered (I4_1md over I-42d).
- The centrosymmetric default gives way only when, on the general reflections of the Laue class of
the final merge, <|L|>, <|E^2-1|> and N(0.1), each calibrated against acentric/centric intensities
simulated with the reflections' own sigmas (fixed-seed mt19937_64, own deviate transforms), read
acentric (L f <= 0.3, others <= 0.5), every f lies in [-0.3, 1.3], the always-centric control reads
centric (f >= 0.5), and the lattice excludes twinning (gemmi Le Page metric admits no rotation
beyond the Laue class, no TWIN_DOMAIN leftover lattice). Then the group is switched and re-merged.
- Report: SPACE_GROUP_CENTRE (IMPLIED_BY_ABSENCES / ABSENT_BY_ABSENCES / NOT_DETERMINED /
ABSENT_BY_STATISTICS), CENTRE_TWINNING_EXCLUDED, CENTRE_STATISTICS_* on every searched run, and
prose for the absence-equivalent alternatives.
Battery (targeted): all 14 small-molecule sets vs 58a4bf (smt-all, kdp-all): every group
unchanged except kdp_x10sa_20keV I 41 2 2 -> I 41 m d with I -4 2 d as the alternative; dnba stays
C 1 2/c 1, now NOT_DETERMINED with C 1 c 1 listed (statistics read centric, f 0.96-1.24). The
P2_1/c, Pbca and Ia-3d sets read IMPLIED_BY_ABSENCES with no alternative.
Protein panel (19 sets, twinned panel + controls, run glide-prot): every space group equals the
rc174-all base (3r6o re-run with the base binary: I 41 2 2 both); no SPACE_GROUP_CENTRE on any.
Twinned proteins read the centre statistics NOT_APPLICABLE (L f -0.3 to -1.7) or ACENTRIC.
Private subset (8 sets): 8/8 groups unchanged vs 58a4bf.
The glide-zone evidence scan over battery 3 (every non-Sohncke group of each run's Laue class,
centrosymmetric or not) peaks at 0.69 nats/reflection on proteins against the 2.0 bar.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
A crystal of slightly misaligned domains (a ferroelastic domain twin below a phase transition, a
split crystal) records each reflection as two or more compact spots around the averaged lattice's
prediction, moving apart with resolution. The pre-scan footprint is measured about each spot, so it
saw compact spots; the r1 disk held the gap between them and the background ring sat on them.
The geometry pre-pass now compares every indexed spot with the predicted position of its own
reflection on the same frame and adds the mean square offset (radial and tangential, by distance
from the beam) to the pre-scan widths; the canonical pass integrates with that table. Where spots sit
on their predictions this moves the widths by the prediction error alone (lysozyme: 0.2-1.0 px, no
reflection outgrows r1); on a 100 K KDP domain twin the offsets reach 8-16 px at the edge.
KDP (kdp_x10sa_20keV), battery SHELXL recipe on the COD model: R1 0.272 -> 0.048, wR2 0.685 ->
0.129, EXTI 26.9 -> 0.025, GooF 3.2 -> 1.29 (XDS: 0.112 / 0.333 / 0.054 / 1.25); R_meas 21.6% -> 5.3%.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The multiplicative outlier band (c8de5f8d6) took its width from the error
model's b. b is a Gaussian width in I and grows with whatever tail the fulls
have: on a low-ISa hexagonal small-molecule set with a population of fulls
lost to near zero, b = 0.54 against a core ln-spread of 0.25, and exp(6 b) = 25
let the high outliers of the weak reflections through (SHELXL R1 0.191 ->
0.333 there, 0.111 -> 0.124 on its 25 keV sweep).
The spread is now measured on the merge's own fulls: the weighted median of
|ln(I / median)|, weighted by (<I>/sigma_counting)^2 so the fulls whose ratio
counting noise does not blur carry it; b remains the fallback where nothing
can be measured.
Battery (29 sets: 13 small-molecule, 16 protein), SHELXL R1 against the
median fix alone: the strongly absorbing cubic set 0.116 -> 0.0998 (b-band
0.103); the low-ISa hexagonal set 0.191 -> 0.174 at 20 keV, 0.111 -> 0.101 at
25 keV; organics within +-0.0003 or better (cytidine 0.0617 -> 0.0613,
lalanine 0.0541 -> 0.0533). Proteins unchanged except two low-ISa sets whose
CC1/2 cutoff moves: 6yqf 3.32 -> 2.89 A (placement R-free at a fixed 3.32 A
0.4329 -> 0.4326, at 2.89 A 0.4642 -> 0.4630), and the split myoglobin set
1.77 -> 1.94 A (low-resolution R_meas 33.0% -> 27.3%). Private subset: two
sets' R_meas lower, the rest unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The 6-sigma outlier band was symmetric in I about the reflection's median,
while the error model's systematic term b*<I> is a multiplicative error
(absorption, illuminated volume, a scale that is off), as likely a factor 1/f
as f. With b of a few percent the two agree; with a large b the linear band
reaches far below the median and only a little above it. On a strongly
absorbing cubic crystal (internal b = 0.34) 96% of the 828 rejected fulls read
2-7x their median in well-illuminated frames - the least absorbed, nearest the
true intensity - and almost none read low.
The b part of the band is now exp(+-n*b) about the expected intensity, the
counting part stays linear (OutlierBand.h, one formula for the host loop and
the device kernel; the two exp constants are taken once on the host). To
first order in n*b it is the old band.
That set's SHELXL R1 0.116 -> 0.103 (--reject-outliers 12 / 20 on the old
band: 0.107 / 0.101; no rejection at all: 0.31).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The per-frame scale was fitted on the partials from 50 rocking events per
frame and taken from the fulls alone below. The counts of small-molecule
sweeps (3-43) and proteins (5-900) overlap, and a weak protein at 36
events per frame scaled from its fulls had no resolved error model at all
(ISa undetermined, d_min 5.10 A) where its partials gave ISa 32 and 4.88 A.
The first merge of a pass that is not a space-group search is now made
both ways, and the count stands as the prior unless its merge has no
resolved error model while the other merge has one. Search merges (P1 /
subgroups) take the prior.
The two ISa values are deliberately not compared beyond that. Tried as
the arbiter, "higher ISa wins" (and weighted R_meas as fallback) agreed
with the external yardsticks on 10 of 11 crystals but chose the partials
on a 6-events-per-frame small-molecule sweep: ISa 10.5 vs 8.7, weighted
R_meas 0.102 vs 0.115, SHELXL R1 0.105 vs 0.062 - the partiality error
the partial scale absorbs is shared by symmetry mates at the same
rocking geometry, so their agreement cannot see it. Statistics of the
difference between the two arms' frame scales did not separate that
sweep from the proteins that want the partials either.
Effect: only crystals whose prior arm fails change; every small-molecule
set and every protein control keeps its arm. Private weak protein: ISa
undetermined -> 32.3, d_min 5.10 -> 4.88 A, weighted R_meas .343 -> .337.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
A saturated pixel in a spot means the brightest part of the reflection was
not measured. The integration used to drop the peak frame's partial (its
peak pixel is unreadable) and keep the flanks, so the combine extrapolated
the event from its tails by the partiality model: on a strongly
diffracting small-molecule crystal the strongest low-order reflections
read 2-3x low and were the largest SHELXL misfits. XDS drops such a
reflection (OVERLOAD); so does rugnux now.
- Integration (CPU + GPU engines): a reflection is `overloaded` when a
signal-disk pixel is saturated, or unreadable on this frame but not in
the run's pixel mask - EIGER/PILATUS write their error value for a
pixel they could not count, which the preprocessor turns into a masked
pixel like a gap's. The engines now receive the PixelMask to tell the
two apart (an earlier attempt that re-classified the marker as
saturation in the preprocessor broke a dataset whose gaps are not in
the file's mask). An overloaded reflection is kept with its box sum,
unfitted, only so its event can be recognised.
- Rotation combine (CPU + GPU): an event with any overloaded partial is
dropped whole; counted in the log and the report
(OBSERVATIONS_REJECTED_OVERLOAD=). The unmerged MTZ export drops it too.
- Everything else that reads reflections leaves an overloaded one out:
AcceptReflection (stills merge, per-image scaling), the post-refinement
gather, the axial-row sums.
- Capture uncertainty: the merge rebuilds each full's variance at the
reflection's mean (counting_variance / ModelSigma) and dropped the
capture term the combine had put into sigma, so a full extrapolated
from part of its rocking curve merged at the weight of a whole one.
Fulls now carry it (Obs::capture) and the rebuilt variance adds
(capture * <I>)^2, host and device.
SHELXL R1 on rugnux's own integration (harness), median fix -> this:
citric acid .0648 -> .0420 (XDS .051; 221 events dropped, EXTI 1.02 -> 0.29),
HEPES .0396 -> .0381 (184), aspirin 20 keV .0387 -> .0385 (6),
aspirin 25 keV .0376 -> .0375 (5); metformin/nidppe/dnba/lalanine/cytidine
no overloads, unchanged. YAG .116 -> .128 (87 dropped; its scale loop does
not settle either way). Proteins and private subset: see the branch report.
Tests: BraggIntegrationEngineCPU_SaturatedPeakIsFlaggedNotDropped (new),
BraggIntegrationEngineGPU_MatchesCPU (overloaded flag compared),
AcceptReflection_ResolutionLimits, [write_reflections], [large].
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The per-reflection median the 6-sigma outlier test is centred on was weighted
by each observation's own counting variance. That variance falls with the
observation's own intensity, so the median went to whichever equivalents came
out low. On a strongly absorbing cubic crystal, where frames near the edge-on
orientation read some equivalents 10-100x down by their diffracted direction,
the median sat on the absorbed ones and the test removed ~10% of the
observations, all of them correct measurements 3-20x above it.
The median is now weighted by the counting variance at the reflection's
expected intensity - the rule the merge weights already use - which keeps the
down-weighting of frames scaled up by 1/G and drops the dependence on the
observation's own fluctuation. In-house garnet (Ia-3d) SHELXL R1 0.139 -> 0.116
(--mode scale on the same integration).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
A spot that outgrows the r1 disk is summed over its measured footprint ellipse and its background
ring starts beyond it. That reach was the same 3 sigma that decides whether the spot outgrew the
disk at all. Wide spots are not Gaussian - mosaic streaks and diffuse halos carry flux past 3 sigma -
so the ring started on the spot's own tails and read them as background.
The install test stays at 3 sigma (BRAGG_FOOTPRINT_NSIGMA); the new BRAGG_FOOTPRINT_REACH = 4 sets
how far the summation ellipse, the ring start and the profile grid go. Compact protein spots never
install the footprint, so they are unchanged bit for bit.
Evidence (rugnux's own combined fulls put through XDS's own per-observation corrections, so only
the integration differs; SHELXL R1(>4sig) on ~93% of observations matched to XDS by hkl):
citric acid: 3 sigma .0521, 4 sigma .0519, 5 sigma .0526 (XDS 3D summation .0484)
HEPES: 3 sigma .0348, 4 sigma .0338, 5 sigma .0344 (XDS .0311)
Full pipeline SHELXL R1: citric .0697 -> .0685, HEPES .0405 -> .0370.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
Merged dq-sm-absorb, whose ingest fix (corr_ingested carries the recomputed partiality, not the
predictor's) repairs the same defect as this branch's unit partial scale on the path without
partial scaling. The two did not double-apply - the unit scale rebuilt corr from scratch - but
they differ by the incident flux, and measured head to head the flux costs the sparse sweeps:
SHELXL R1 on four in-house organic sweeps 0.0459/0.0396/0.0709/0.0413 with it, 0.0421/0.0393/
0.0697/0.0405 without; on five open small-molecule sets equal or better without (0.0707 ->
0.0660 on one). Only the cubic absorbing sweep prefers it (0.119 vs 0.139): there the
background really does fall with the absorbed beam. So the unit partial scale and its device
kernel are removed and the ingest fix stands alone; the flux meter's smoothing, which only
mattered on that path, is reverted, so the partial-scaling path sees the flux exactly as
before. The penalised per-frame scale of the fulls stays.
Also restores two characters of CPU_DATA_ANALYSIS_INTEGRATION.md that the previous commit's
rewrite had turned from a stray carriage return into a line break.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
Stage A's gates all compare the operators' disagreement with a reference - a parent group, the best
operator, the noise floor - and a pseudo-symmetric structure or a twin defeats every one of them: the
added operators correlate nearly as well as real ones. On the open arm that over-called 2wnz and 2xfw
(deposited P2_1, pseudo-merohedral C-orthorhombic metric) as C222_1, and 7bgu (deposited P1 with
b = c to 2.7%) as C2.
New gate, for a candidate that holds every rotation of its lattice (gemmi find_lattice_symmetry at
3 deg obliquity, with the lattice centring now passed in SearchSpaceGroupOptions::lattice_centring):
no twin law exists for such a group, so intensities merged under it that read like a twin's are not a
twin. AnalyzeLTestUnderMerge reads the Padilla-Yeates L-test twice on the same pairs of the P1 merge -
as measured, and with every intensity replaced by its plain orbit mean under the candidate - so the
narrowing is the merge's own doing and not a property of the data. Refused when the merged <|L|>
reads twin-like (0.375 <= <|L|> < 0.44) and the averaging closed a third or more of the gap from the
unmerged <|L|> to 0.375.
Calibration on the C++ search merges (targeted battery sympg-pg1): genuine lattice-holohedral groups
close 0.01-0.25 of the gap (highest 6z9g, indexed on half its deposited cell, 0.23-0.25; 5vml 0.16-
0.17; 7kcn 0.06-0.09); the over-calls 2wnz 0.47, 2xfw 0.46-0.47, 7bgu 0.79. A merge reading below
0.375 is not read - nothing a twin or a false operator does reaches it, so something else compresses
the intensities: 8c3e (0.353) and the genuine 9zmu (0.351, 494 A axis) look alike there. Near-
perfect twins (4bwl 0.14-0.16, 2wnq 0.13-0.20) are out of reach by construction, and 6p8j (0.19-0.30,
deposited as a twin of P2_1) is left alone. Cost: 170 ms per test on a 730k-reflection merge, only for
candidates that pass every other gate and hold the lattice's full symmetry.
Targeted battery sympg-final (44 open/in-house sets: the over-calls and a control panel of genuine
lattice-holohedral, twinned and twin-like sets) and the private arm: 2wnz C2221 -> P1211 (R_free with
the deposited model 0.227), 2xfw C2221 -> P1211 (0.215), 7bgu C121 -> P1 (R_meas 0.28 -> 0.14), all
the deposited groups. Every other set keeps its space group; against the rc174-cand runs of the same
sets, open and private, nothing else moves beyond run time. Still over-called: 8c3e, 4bwl, 2wnq,
6p8j, 3r6o (reasons above; 3r6o and 8c3e read below 0.375).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
Every partial's corr is formed at integration as image_scale_corr = LP*QE*flight / partiality with
the predictor's per-frame partiality. Ingest then recomputes every partiality from the
frame-order-smoothed mosaicity (and the one exact-Bragg angle per rocking event), but corr_ingested
kept the predictor's 1/p. Where partials are scaled every fitted frame's corr is rebuilt from the
recomputed partiality, so that path only saw it in the first reference. Where they are not - fewer
than 50 rocking events per frame: small molecules, and long-wavelength protein sweeps - corr is what
the 3D combine divides each partial by, and the full became a weighted mean of I/p_pred while the
event's recomputed fractions summed to 1.00: the combined fulls tracked 1/sum(p_pred) (r = -0.96)
and scattered 13% rms about the plain sum of the same partials.
corr_ingested is now rescaled by p_pred/p_recomputed once at ingest.
Targeted battery against rc174-cand: every set on the partial-scaling path is bit-identical
(5reo, 9qw8, lyso_x06da_ref, thau_x10sa_0p1deg, lyso_x06da_atten_wedge, myob_x06da_split,
cytc_x06da_1, insu_I_x06da_5keV/6keV). Small molecules, SHELXL R1 against the COD model:
aspirin 20 keV 0.0515 -> 0.0432 (ISa 10.3 -> 30.9), aspirin 25 keV 0.0451 -> 0.0396, HEPES
0.0452 -> 0.0420, citric acid 0.0747 -> 0.0726, dnba 0.0472 -> 0.0273, metformin 0.0405 -> 0.0326,
nidppe 0.0498 -> 0.0426, lalanine 0.081 -> 0.057, cytidine 0.077 -> 0.071. Long-wavelength
proteins on the fulls-alone path: lyso_x06da_5keV ISa 17.9 -> 27.5, thau_bl1a_3p8keV 21.7 -> 28.4,
thau_bl1a_4p6keV 26.4 -> 33.5. One regression: lcystine_x10sa_25keV (a polycrystalline aggregate)
now refuses 622 and reports P31 instead of P6122.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The French-Wilson prior of each reflection is now epsilon * K_shell * a(h), with
a(h) = exp(-1/2 s^T B s) from the deviatoric tensor AnalyzeAnisotropy already fits
(the form it is fitted in) and K_shell = sum(I/eps) / sum(a), so a shell's priors still
average to its measured mean. The amplitudes are made isotropically at the merge as
before and made again once the tensor exists (full pipeline and --mode scale). Only
F/SIGF and F(+)/F(-) change; IMEAN/I(+)/I(-) are bit-identical. Applied whenever a
tensor was fitted, with no detection gate: a near-isotropic tensor gives a(h) ~ 1 and
the isotropic prior back, and the prior wants the best estimate of <I> along h whatever
its cause. Follows ctruncate's anisotropic prior (Ballard & Stein, CCP4); credit in
ACKNOWLEDGEMENT.md, CPU_DATA_ANALYSIS.md and at the algorithm.
--model scaling (ModelScaling.cpp) was checked: k_overall + symmetry-constrained
anisotropic B + flat bulk solvent fitted on the working set, against the same FW F
written to the MTZ - as REFMAC/phenix.refine do. No change needed.
Evidence (REFMAC 10-cycle restrained refinement of the deposited model, R-free on
the depositor's free reflections shared by both data sets; base = rc174 processing,
same IMEAN):
set base new d set base new d
9rcs 0.3475 0.3494 +0.0019 8qq7 0.4452 0.4503 +0.0051
9yzk 0.3192 0.3171 -0.0021 9hs7 0.2898 0.2465 -0.0433
6yqf 0.4642 0.4543 -0.0099 5nw5 0.3256 0.3206 -0.0050
7n2s 0.3126 0.2998 -0.0128 6z8o 0.2927 0.2892 -0.0035
6qaj 0.3390 0.3038 -0.0352 6moj 0.2756 0.2651 -0.0105
6r72 0.3826 0.3797 -0.0029 7qij 0.3239 0.3140 -0.0099
anisotropic sets: median -0.0075, mean -0.0107, 10/12 better
isotropic controls: 5reo -0.0014, 7kcn +0.0003, 6fid +0.0002, 11if 0.0000
rugnux's own --model R-free moves the same way (median about -0.019; 5nw5 +0.006),
R_model shell-scaled too; dep_cc_delta unchanged (intensity based). The adoption rule
(median gain >= 0.005 on the anisotropic sets, no control worse than +0.002) is met.
The two sets that lose are the one with a FLAT resolution signature (8qq7) and 9rcs,
where the exp form drives the dead direction's prior to ~0 beyond 3.7 A.
For scale: ctruncate's own anisotropic prior on the same merges moved the same
REFMAC R-free by a median of only -0.0008 (9hs7 +0.026).
Inhouse lyso_x06da_ref, thau_x10sa_0p1deg: every battery metric unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
Three defects in the per-frame scaling of sparse (small-molecule, weak) rotation sweeps:
1. The fulls' per-frame scale pooled sparse frames by their RAW full count, but in a
high-symmetry group with many systematic absences most usable counts stayed under the
minimum, so most frames were never fitted and kept corr = 1 beside pinned, fitted frames
(two gauges in one reference). The release step then read the gauge offset (a constant
134x on a cubic Ia-3d small-molecule sweep) as signal and gave each frame exp(kept_f * 4.9) - the e^-5
errors on a quarter of the frames, R1 0.62. A fixed box window also cannot follow a
100x absorption ramp over a few degrees, and frames under the credible floor, exempt
from the window, ran away to 1e-7.
Now each round fits every frame on its own fulls (no pooling, no minimum), and the scale
is a penalised second-difference smoother of log G (Whittaker/Eilers), each frame at
the information of its fit, lambda by cross-validation over blocks one rocking curve
wide (interleaved single frames leak through shared rocking curves and chose to follow
every frame). After convergence the existing ShrinkToRestrained hands back the per-frame
deviation its neighbour shares. Pooling and the box window are gone from the fulls loop;
the partials loop is unchanged.
2. With partial scaling off (< 50 rocking events per frame) the partials kept the
integration-time corr: no incident-flux correction and not the partiality of the
ingest-smoothed geometry, because only the partial scaling loop rewrote corr. They now
get corr = prescaling_corr / partiality at G = 1 (host, and a device kernel).
3. The flux meter (per-frame mean background) jumped 30x between neighbouring frames of a
sparse sweep - on a few reflections it measures which reflections the frame holds. It
is read through the same smoother at the precision of each frame's mean.
SHELXL R1(>4sig) against the published structures, rc174-cand -> this, on the in-house
small-molecule sweeps: cubic Ia-3d 0.615 -> 0.119 (XDS 0.088); four organic sweeps
(monoclinic / orthorhombic) 0.0515 -> 0.0459, 0.0451 -> 0.0396, 0.0747 -> 0.0709,
0.0452 -> 0.0413; ISa up to 11.6 -> 31. Raw flux instead of smoothed costs 0.002-0.003 R1 on
the first two. A weak, decaying protein sweep on the fulls-only path: ISa 6.7 -> 27.1.
CPU and GPU paths agree.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
Mapped every single-axis screw candidate of every primitive-lattice set in
battery 3 (open + in-house) and the private arm against the deposited group.
False zones with no violation read at most +4 nats (plus one pseudo-
translation row at +30.5 that no bound separates); true zones go down to
+10. The true zones under 20 are all short monoclinic rows (b ~ 25-30 A:
four to six 0k0-odd reflections inside the search's resolution range, on
weak data at a few percent of their row), refused at 20 on five crystals
that are P2_1: four myoglobin sweeps (10.0-16.3 nats) and 6cs9 (10.8).
The rc174-cand myob_x06da_split call sat at 21.1, one refit away from P2.
Re-scoring the stored candidate tables at the new bound changes only those
five plus 3r6o (I4_1 2 2 newly eligible, toward the deposited I4_1 screw);
no adopted group elsewhere. New test section pins both sides: four 0k0-odd
at 2% of their row are claimed (11.3 nats), at 20% they are not.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The integrator's r1 disk and r2..r3 background ring are fixed in pixels and chosen from spots near
the beam. On small-molecule data at 20-25 keV a spot's standard deviation grows from ~1 px near the
beam to ~5 px at the edge (radially from parallax/obliquity, tangentially from the crystal's
azimuthal spread), so the r1 = 4 disk holds a quarter of the flux there, the background ring a third
of it, and the in-disk second moments the Gaussian is built from saturate near r1^2/4. On top of
that, the profile/summation runaway guard sent 20-30% of these reflections - the strong, wide ones -
back to the truncated r1 box sum.
- SpotFootprint: every pre-scan spot (width frames) is measured with a window that follows it
(3 sigma, iterated, re-centred), radially and tangentially; the medians per distance-from-beam bin
become BraggIntegrationSettings::Footprint. Installed only where some bin outgrows r1, and on the
adaptive side like the radius (pre-pass without; the starvation guard falls back to the settings
without it).
- BraggStencil: where 3 sigma > r1 the background ring starts at 3 sigma along and across the radius,
the summation region is the r1 disk plus the 3-sigma footprint ellipse (so the guard's fallback is a
complete intensity), and the per-reflection Gaussian takes the footprint widths. Compact spots keep
the stencil bit for bit. Both engines build it from the same header.
SHELXL against COD (R1 / fixed-XDS-model R1(F)): citric acid .101/.230 -> .077/.055, HEPES
.070/.179 -> .048/.050, aspirin 20 keV .059/.070 -> .052/.061, aspirin 25 keV unchanged, L-cystine
25 keV unchanged (.145 -> .144).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The SHELX HKLF 4 file now has one record per full reflection - its partials summed,
the per-frame scale and every correction applied, sigma(I) as the merge weighted it -
at the index it was measured at, not averaged with its equivalents: the chemical
crystallographer's convention, so SHELXL computes Rint and Rsigma itself. Outliers
the merge rejected and fulls beyond its resolution cut are left out; no batch column
(it would select a BASF scale in SHELXL). The engine hands the fulls back only for a
merge that may be written (RotationScaleMerge::SetExportScaledFulls), so the search
merges, the pre-pass and the P1 cross-check carry no copy; the fulls follow the same
relabelling as the merged reflections (merge_to_written). Stills keep the merged
file. --mode scale writes the unmerged form too.
Validation: p.mtz md5 unchanged on myob/cytc/thau (GPU). SHELXL on the same runs,
merged-old vs unmerged-new (COD models, harness /data/tmp/sm_shared):
aspirin 20 keV Rint 0 -> 0.071, Rsigma 0.036 -> 0.043, R1 0.0964 -> 0.0969,
wR2 0.312 -> 0.310, GooF 1.53 -> 1.47; HEPES 20 keV Rint 0 -> 0.163, Rsigma 0.057
-> 0.068, R1 0.0903 -> 0.0897, wR2 0.318 -> 0.263, GooF 1.57 -> 1.10; the "input
data appear to be merged" warning is gone.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The error model was fitted once, on deviations from a mean that weighs every full by its COUNTING
variance, and the outlier test's median took the same weights. Where equivalents disagree beyond
counting statistics, the low-count observations dominate that centre: on a strongly absorbing crystal
(YAG, Ia-3d, equivalents spread over a factor 100 after scaling) the mean of (0 4 0) sat at 21k among
observations from 11k to 1.9M, a ran into its bound (100), b came out at 800% internally, and the
six-sigma test about the biased median removed 63% of the observations - the strong ones.
Now, after the first fit, the model is refitted about the model-weighted mean (each full weighted by
the variance the fitted model gives it at the reflection's mean) until a and b settle, and the
rejection median takes the same weights. Where counting statistics are right nothing moves. Following
Blessing (1997) J. Appl. Cryst. 30, 421-426. GPU path: the refitted means are uploaded (SetEmMean).
Measured (SHELXL R1(>4sigma) against the COD model, sm-a's harness; rc174 scaling):
YAG 0.556 -> 0.127 (XDS 0.083 merged), rejected 8055 -> 53, normalised deviations calibrated
(median |z| 0.62-0.69 in every intensity decile); aspirin 20 keV 0.0964 -> 0.0958; citric acid
0.161 -> 0.159; HEPES 0.0903 -> 0.0899; L-cystine 25 keV 0.1425 -> 0.1456;
aspirin 25 keV 0.094 -> 0.106 (fixed-model R1 0.107 -> 0.173): its strong equivalents split into two
frame-dependent populations from the per-frame partial scaling (sm-a's dq-smallmol), which the old
under-sized sigmas happened to cut; with that scaling fixed (f69339ce6 + pooling, --no-scale-partials)
this change is neutral to better on every small molecule (aspirin 20 .0618 -> .0586, aspirin 25
.0456 -> .0454, citric .1093 -> .1006, HEPES .0703 -> .0697, YAG .649 -> .222; SHELXL GooF ~1.1).
=> ship together with the scaling fix.
Proteins (GPU full runs): CC1/2 and R_meas unchanged to 0.002; ISa myob 9.06 -> 8.20, thau 52.5 -> 47.5,
cytc 25.8 -> 25.4, lyso 29.4 -> 29.4 (still above XDS's 5.2 / 44.5 / 31.8 / 28.3 except cytc).
CPU build gives the same statistics as the GPU build on aspirin 20 keV and myob.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
Replaces the prototype --no-scale-partials switch with a rule read off the data: a merge whose
sweep holds fewer than 50 rocking events per frame takes its per-frame scale from the fulls alone
(scale-fulls, a sparse frame fitted over its neighbours - PoolHalfWidth), and logs that it did.
The rocking-event walk already run for the smoothing window now also returns its event count.
Within one rocking curve a partial's scale and an error of the partiality model are the same
thing; on a sparse fine-sliced sweep the partial fit takes the one for the other and imprints an
hkl-dependent bias common to all equivalents. Measured populations: small-molecule sweeps 2.6-23
events per frame, protein sets 84-900; on the proteins the fulls-only scale leaves model R-free
unchanged (+-0.002 over the smoke tier's open-arm sets) but lowers ISa, so they keep the partial
scale. p.mtz md5 unchanged on the three profiling sets, GPU and CPU builds.
SHELXL against the COD structures, R1(>4sig) rc174 -> this: aspirin 20 keV 0.096 -> 0.062,
aspirin 25 keV 0.094 -> 0.046, citric acid 0.161 -> 0.109, HEPES 0.090 -> 0.070 (XDS 0.030-0.038);
SHELXL's weight a comes off its 0.2 cap on all four. CPU and GPU paths agree.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The same source built with and without -march=x86-64-v3 gave different results on battery sets
(9zmu axis-harmonic supercell arbiter fired in one build only; 8rud resolution cut 1.69 vs 1.70 A;
myob_x06da_powder_1 cut 1.07 vs 1.42 A; 7mzt short-axis first pass; lcystine_x10sa_20keV indexed
vs no lattice), against the rule that no decision may depend on compiler flags.
Two mechanisms, found by building the merged tree four ways (baseline, x86-64-v2, x86-64-v3,
x86-64-v3 -mno-fma) with and without -ffp-contract=off and comparing p.mtz:
- FMA contraction. GCC contracts a*b+c whenever the target has FMA. With -ffp-contract=off the
x86-64-v3 build gives a p.mtz byte-identical to the baseline build on 8 of 9 sets (7mzt, 8rud,
9zmu, insu_I_x06da_5keV_2, myob_x06da_powder_1, myob_x10sa, cytc_x10sa, thau_x10sa_16keV), on
both the GPU and the CPU build. Cost: none measurable (user core-s, CPU build, x86-64-v3 vs the
same with -ffp-contract=off: 1877/1874, 3285/3253, 2212/2195 on myob/cytc/thau; GPU likewise
within noise). Set project-wide for C, C++ and CUDA host code; MSVC does not contract under
/fp:precise.
- SIMD width. lcystine still differed: baseline and x86-64-v2 (128-bit) agreed, x86-64-v3 with or
without FMA (256-bit) disagreed - Eigen's HouseholderQR in the FFT indexer's candidate refinement
(PostIndexingRefinement.cpp) reduces column norms over all spots in packets of the target width.
Replaced by the 3x3 normal equations summed in spot order in double. All four builds now agree on
all nine sets.
Changes results of the default x86-64-v3 build (contraction off); lcystine_x10sa_20keV now gives no
lattice in every build (its first pass is a knife-edge: 0/60 vs 9/60 validation frames before).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
Prototype switch --no-scale-partials: skip the per-frame scaling of the partials and let the
fulls (scale-fulls) carry the per-frame scale, as XDS does. Without partial scaling a frame whose
fulls number fewer than MIN_REFLECTIONS is fitted over the nearest frames on either side that
together hold enough (PoolHalfWidth), on the host and in the GPU kernel. Default path unchanged
(p.mtz md5 identical on the three profiling sets).
Why: on fine-sliced small-molecule sweeps the per-frame partial scale and the partiality model are
degenerate within a rocking curve, and the fit swings G 0.23..1.2 with a 180 deg period (XDS's own
frame scale: 0.79..0.99). That imprints an hkl-dependent bias common to all equivalents, which
R_meas/CC1/2/ISa cannot see but a refinement against the known structure does. And scale-fulls
never fitted a frame on such data: a full is filed under one frame, about 8 per frame, below
MIN_REFLECTIONS, so every frame kept G = 1.
Measured with SHELXL refining the COD structures (R1 >4sig), default -> switch:
aspirin 20 keV 0.096 -> 0.062, aspirin 25 keV 0.094 -> 0.046, citric acid 0.161 -> 0.109,
HEPES 0.090 -> 0.070 (XDS 0.030-0.038). Proteins lose ISa with the switch (myob 9.1 -> 7.6,
cytc 25.8 -> 13.7), so it is not a default; the choice is to be made from the data.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
A promoted point group whose <|L|> reads below 0.375 (NOT_READABLE) previously
raised no TWINNING warning at all; a twinned subgroup whose law the promotion
absorbed predicts the same data, so the adopted group is not confirmed. Warn.
The twin-immune zone control reading more compressed than a perfect twin's
acentric population (0.541) is something no twin fraction produces (overlap or
neighbour correlation); genuine symmetry then reads acentric in its zones too
(measured: a genuine 622 with its control at 0.528 read its 2-folds at -950 to
-2640 nats). Such zones are now marked ambiguous in the text and the zone
decision line. Report-only: no decision and no output file but the report changes.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The non-reference settings of the adopted point group were offered only when the cell hosted their
rotations to within ~0.1 deg, while the reference settings - which hold the very same rotations -
were offered without asking. A cell refined free after integration a few tenths of a degree off 90
therefore lost every setting but the reference one. On an orthorhombic set whose measured screws
lie on a and c and whose b row was never recorded, that left P2(1)2(1)2(1) as the only candidate
covering both screws, and it was reported as determined. With P 21 2 21 offered, the two tie, the
b-axis screw is reported as undetermined, and the model check uses the setting the data describe.
The point-group stage still asks the cell whether a rotation set it adds is hosted; only the
setting enumeration within an already chosen set stops asking.
Validation: myob/cytc/thau x10sa p.mtz byte-identical (GPU); 7mzt fail -> unscored (b screw
undetermined, P 21 21 21 or P 21 2 21); 5cc8 unchanged; [SearchSpaceGroup] 23 cases pass incl. a
new section with a cell 0.2-0.3 deg off 90.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The all-observation arm of the space-group search and the P1 cross-check
were made on a second RotationScaleMerge engine beside the run's own
(871347b7a). That engine holds a second device copy of every observation,
and on the largest sets the two no longer fit a 16 GB card: 8a1a, 8qaw and
8tyy (55-125 M partial observations) stopped with an out-of-memory error
in scaling. Measured on a quiet box the engine bought 0.7 s (cytc) and
1.0 s (thau) of tail and nothing on myob, which does not justify a memory
budget, so it is removed and both merges run on rsm in sequence, as
before 871347b7a.
Kept from 871347b7a: the GPU scaling's own non-blocking stream, the
cross-check not writing per-frame G/CC/mosaicity back (the per-image table
still describes the merge that was written), the restored scaling
iteration counts, and the anisotropy analysis beside the other analyses.
p.mtz byte-identical to before on myob/cytc/thau, GPU and CPU builds;
p_plot.txt unchanged apart from the GPU bkg column.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
ApplyCellSurface (detector modulation, time x detector, crystal-frame SH and
goniometer-frame absorption) spends most of its time in two passes per round:
the per-group reference sums and the per-(block, cell) fit sums. With a GPU
both now run on the device (RotationScaleMergeGPU::Surface*) over the same
terms in the same order:
- reference: one thread per ASU group, walking a group-order permutation of
the terms in fulls order;
- fit: each subset cut into the host's reduction blocks with every block's
terms ordered by cell (stable counting sort, host); a per-term kernel forms
w*Is, w*Iref, Iref and one thread per (block, cell) runs the two fma chains;
the block slots are added per cell in block order.
Every rounding is spelled out (__dmul_rn/__dadd_rn/fma) to be the one the host
build makes: GCC at -march=x86-64-v3 fuses swI's multiply-add only in the
parity-filtered copy of the reference loop, and both fit sums.
Host side, exact on both paths: the 19 serial nth_element selections of the
shell edges become one parallel sort (same order statistics), the per-term shell
lookup runs on all threads, and the gate's per-shell CC is one walk over the
groups instead of one per shell.
Exact: p.mtz md5 identical to the oracle on myob/cytc/thau x10sa, GPU build
(all CUDA architectures) and CPU build. CorrectionSurfaceGPU test checks the
device sums bit for bit against an explicitly rounded host loop.
Measured (cytc/thau, two interleaved A/B pairs, box at load 13-25 from sibling
work): ApplyCellSurface host core-seconds -83% on cytc; SG adoption -> writing
reflections 4.58->3.75 and 3.85->2.54 s (cytc), 3.15->2.28 and 2.29->2.03 s
(thau); RSM final merge -0.7..-0.9 s and P1 cross-check -1.5..-1.8 s on cytc.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The two pre-scan steps that were still CPU-bound in a GPU build now run where the
projection already is.
- FindBeamCenterFromBackground: the per-iteration binning pass and the two clip rounds
run on the device (BeamCenterBackgroundGPU); the fit itself stays on the host. Each
cell is summed in the host's order (pixel order within the host's row blocks, blocks
in order), and the per-pixel cell/derivative formula is shared (BackgroundBand.h).
The angles come from BackgroundAtan2 (IEEE ops only) instead of atan2f, and both
translation units are compiled without FMA contraction, so host and device give the
same bits: 0 of 6.5 M pixels in a different cell, identical walks on the three
in-house rotation sets. With glibc/CUDA atan2f and default contraction ~30 pixels per
16 Mpx sweep changed cell and the fitted centre moved by up to 0.05 px.
- ShadowFinder::GetMask: the whole mask (pooling, ring medians, components, morphology,
hole fill, arm search) runs on the device from ShadowAccumulatorGPU's projection
(ShadowMaskGPU), so the 360 MB projection no longer comes back; the mean projection is
divided on the device too (same bits). The two small fits over rings and sectors
(BlockedOutTo, HarmonicFit) are shared with the host path in ShadowFinderInternal.h.
Integers, comparisons, sorts and components are exact; the polarization trig, the
Poisson log and the arm-search azimuth are not, so a pixel at a threshold can differ.
The one-time change against the previous CPU arithmetic (BackgroundAtan2, no
contraction), measured on the myoglobin, cytochrome C and thaumatin rotation sets:
ring centre moves 0.002-0.045 px (fit sigma 0.75-1.2 px), beam-centre capture
0.01-0.04 px; beam-stop mask differs on 31 / 144 / 53 pixels of 259k / 144k / 198k
(25 of the myoglobin ones are GPU-vs-CPU arithmetic in the mask, the rest follow the
centre); hot-pixel mask identical. Spot width, integration radii, bandwidth, beam-centre
arbitration, indexing, space group, cell, resolution and the merged statistics table
are identical; only the error model moves in its 4th digit. CPU build: the same
centres and decisions.
Timing (GPU, box at load 30-38): ring walk 0.54 -> 0.23-0.27 s, mask 1.24-1.44 ->
0.18-0.22 s, beam-centre capture walk 1.1-1.3 -> 0.31-0.35 s.
Tests: ShadowFinder_DeviceMaskMatchesHost, BeamCenterFromBackground_DeviceMatchesHost
(bit-exact), plus [ShadowFinder], [BeamCenter], [HotPixelFinder].
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The standalone GPU azimuthal integration did three shared-memory atomics per
pixel on the same few ring addresses, which is what it was limited by. It now
reads four pixels per thread as vector loads and keeps a running total per ring,
flushed when the ring changes - the scheme the adaptive finder's ring pass
(reduce_rings_shared) already uses. The npix % 4 leftovers are done one at a time.
Used wherever the fused adaptive engine is not (fixed-threshold spot finding, the
broker's non-adaptive path). Measured on a 16 Mpx sweep (1800 frames,
--no-adaptive-spots, RTX 5080): 843 -> 295 us per call (min 621 -> 196 us).
Not bit-identical, and the old kernel was not either: float atomics arrive in any
order, so two runs of the OLD kernel already differ by up to 1.7e-6 relative in
the per-frame profile; new vs old differs by up to 1.9e-6, the same order. Per-ring
pixel counts are identical. Default rugnux runs do not reach this kernel (p.mtz
md5 unchanged on three sets); on the fixed-threshold path p_unmerged.mtz is
md5-identical to the old kernel's. New test: GPU vs CPU engine on a pixel count
that is not a multiple of four, with masked and saturated pixels.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
Exact: p.mtz and the pre-scan products (shadow mask, mean projection,
defective-pixel mask, ring and capture centres, compared as hashes and
hex floats) are bit-identical to rc174 on three in-house rotation sets,
GPU and CPU builds.
- ShadowFinder::GetMask: the serial parts run in parallel - connected
components by row band joined with union-find (both the shadow and
the transmitting-arm searches, and the hole fill), ring binning and
the harmonic sector gather by blocks, gap bridging by line; ring pixel
counts read off the ring offsets. Mean projection filled in parallel.
- ShadowFinder host accumulation: one band-locked projection instead of
a 20 B/px shard per pre-scan worker (2.7 GB zeroed and folded on a
16M detector); SetShardCount and the shard argument are gone.
- FindBeamCenterFromBackground: the usable-pixel test is made once, the
in-band pixels are kept in pixel order so the clipping rounds no
longer sweep the whole detector, the 67 MB cell map is gone and the
per-iteration block fold runs in parallel - same sums, same order.
- HotPixelFinder::GetMask: the chance-rate counts in parallel (integers).
Measured on a loaded box (load ~25 from other jobs), pre-scan window:
GPU 5.9-6.5 s -> 3.2-3.4 s, CPU 8.4-9.0 s -> 6.1-7.4 s. The GPU-build
pre-scan now ends with its background spot measurement (CPU spot finder
on ~120 frames, ~13 core-s on 8 workers).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The tail of the canonical pass made five merges one after another. Two of
them read nothing the space-group search or the in-symmetry merge decides:
the all-observation arm of the search and the P1 cross-check. They are now
made on a second RotationScaleMerge engine, ingested beside the run's own
before any merge writes per-frame values back, and taken where they were
made before; where the run re-ingests (a reindex, a cell change) they are
made on the run's engine as before.
- RotationScaleMerge::SetWriteBackPerFrameScale(false) keeps a merge that
is not the run's answer from writing G/CC/mosaicity onto the outcomes.
The P1 cross-check no longer overwrites them, so _plot.txt's scale_G,
cc_to_merge and cc_n now describe the merge that was written (in the
determined group) instead of the P1 cross-check; the cross-check also
no longer leaks its scaling iteration count into the report.
- RotationScaleMergeGPU runs on its own non-blocking stream instead of
the legacy NULL stream, so the two engines (and a probe pass beside
them) do not serialise at every launch and synchronisation.
- The anisotropy analysis (mostly ScaledObservations) runs beside tNCS,
twinning and the other report-only analyses.
p.mtz, p_P1.mtz, p.cif, p.hkl and p_unmerged.mtz are byte-identical on
myob/cytc/thau (GPU and CPU builds). Tail on cytc GPU 8.1 -> ~6.8-7.7 s
under a loaded box.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
AdaptiveSpotFinderCPU::AccumulateRingsBlock reads each pixel once for both the
ring histogram and the fused azimuthal profile (was two loops), and no longer
keeps the per-ring integer sums: they are taken from the histogram, as the two
sigma-clip passes already were (ClipRings(INFINITY)). Integer sums, so the same
totals; the profile's float sums keep their pixel order.
FlagRow is branch-free and works a 32-pixel word at a time, so it vectorises;
pixels outside every ring meet a +inf threshold in an extra ring_thr entry.
BraggPredictionRot::Calc takes A*h, A*h + B*k, C*l and 4*S0*S0 out of the inner
loops; p0 is the same ((A*h) + (B*k)) + (C*l) as before.
Measured on cytc (CPU build, both binaries run concurrently on a loaded box):
AccumulateRingsBlock -35%, FlagRow -47%, Calc + Coord ops -30% cycles; whole run
-5% cycles. p.mtz md5 unchanged on myob/cytc/thau, CPU and GPU builds.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB