Its deposited sweep lacks 40 deg (frames 561-720 are not in the
archive; the depositors processed it as two sweeps), so it tests
gap handling rather than processing.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
SRS Daresbury PX10.1 (6RYM), EMBL Hamburg DORIS X13 (7BGU), NSLS X29A and X25
(8V4J, 8V2T), ESRF ID14-2 (5JK4), Australian Synchrotron MX2 (6CS9), CLSI 08B1-1
(7UDI), SSRF BL17UM (9LXL), MAX IV BioMAX (9S02), ALBA XALOC (6GVK) - marCCD,
ADSC (Quantum 4/210r/315), PILATUS and EIGER2 files. Each read and processed with
the current rugnux; 7BGU (C2 imposed on a P1 lattice) and 7UDI (P41 under-called
to P2) disagree with the deposition. Multi-sweep archives pinned to one sweep.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
Pseudo-merohedral and merohedral twins, pseudo-symmetric cells and their
negative controls, chosen to test the space-group search's promotions:
- 8c3e (P3121, merohedral twin ~0.23, home-source HyPix CBF)
- 4bwl, 2wnq, 2xfw, 2wnn (P21 NAL crystals with the pseudo-merohedral
twin law -h,-k,h+l at fractions 0.49/0.46/0.10/0.33; ADSC SMV) and 2wnz,
the same crystal form deposited untwinned
- 6oww, 6p8j, 9qw2 (P21 with near-90 beta or a~c), 9qvv (I222, b=c),
5ojv (P21212, a~c), 7q6j (P212121, b~c)
Data staged in /data/scout_twin, linked from the open data root.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The XDS references of aspirin (20, 25 keV), citric acid, HEPES and YAG were made with
OVERLOAD=65534, half of the EIGER2's saturation value (133202). CORRECT ignored 13-313
reflections per set whose peak pixel lay between the two - the strongest, most
extinguished reflections - so the references described a data set without them.
Re-ran only the CORRECT step on the unchanged INTEGRATE.HKL with every input parameter
reconstructed from the old CORRECT.LP and OVERLOAD=133201 (saturation - 1, as the other
in-house sets already use); a rerun at 65534 reproduced the old CORRECT.LP. The old outputs
are kept in xds_overload_half/ beside each data set. `refs --write` picked up the new values.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
BraggIntegrationEngineGPU_MatchesCPU gains three footprint sections (spaced, crowded under overlap
exclude, with the radial background correction): both engines classify the summation ellipse, the
grown ring and the footprint Gaussian alike. Integration chapter and changelog describe the measured
footprint.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The integrator's r1 disk and r2..r3 background ring are fixed in pixels and chosen from spots near
the beam. On small-molecule data at 20-25 keV a spot's standard deviation grows from ~1 px near the
beam to ~5 px at the edge (radially from parallax/obliquity, tangentially from the crystal's
azimuthal spread), so the r1 = 4 disk holds a quarter of the flux there, the background ring a third
of it, and the in-disk second moments the Gaussian is built from saturate near r1^2/4. On top of
that, the profile/summation runaway guard sent 20-30% of these reflections - the strong, wide ones -
back to the truncated r1 box sum.
- SpotFootprint: every pre-scan spot (width frames) is measured with a window that follows it
(3 sigma, iterated, re-centred), radially and tangentially; the medians per distance-from-beam bin
become BraggIntegrationSettings::Footprint. Installed only where some bin outgrows r1, and on the
adaptive side like the radius (pre-pass without; the starvation guard falls back to the settings
without it).
- BraggStencil: where 3 sigma > r1 the background ring starts at 3 sigma along and across the radius,
the summation region is the r1 disk plus the 3-sigma footprint ellipse (so the guard's fallback is a
complete intensity), and the per-reflection Gaussian takes the footprint widths. Compact spots keep
the stencil bit for bit. Both engines build it from the same header.
SHELXL against COD (R1 / fixed-XDS-model R1(F)): citric acid .101/.230 -> .077/.055, HEPES
.070/.179 -> .048/.050, aspirin 20 keV .059/.070 -> .052/.061, aspirin 25 keV unchanged, L-cystine
25 keV unchanged (.145 -> .144).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The SHELX HKLF 4 file now has one record per full reflection - its partials summed,
the per-frame scale and every correction applied, sigma(I) as the merge weighted it -
at the index it was measured at, not averaged with its equivalents: the chemical
crystallographer's convention, so SHELXL computes Rint and Rsigma itself. Outliers
the merge rejected and fulls beyond its resolution cut are left out; no batch column
(it would select a BASF scale in SHELXL). The engine hands the fulls back only for a
merge that may be written (RotationScaleMerge::SetExportScaledFulls), so the search
merges, the pre-pass and the P1 cross-check carry no copy; the fulls follow the same
relabelling as the merged reflections (merge_to_written). Stills keep the merged
file. --mode scale writes the unmerged form too.
Validation: p.mtz md5 unchanged on myob/cytc/thau (GPU). SHELXL on the same runs,
merged-old vs unmerged-new (COD models, harness /data/tmp/sm_shared):
aspirin 20 keV Rint 0 -> 0.071, Rsigma 0.036 -> 0.043, R1 0.0964 -> 0.0969,
wR2 0.312 -> 0.310, GooF 1.53 -> 1.47; HEPES 20 keV Rint 0 -> 0.163, Rsigma 0.057
-> 0.068, R1 0.0903 -> 0.0897, wR2 0.318 -> 0.263, GooF 1.57 -> 1.10; the "input
data appear to be merged" warning is gone.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
shelx_check.py refines the row's published structure (manifest key "cod", cached from the
Crystallography Open Database into the site's cod_cache) against rugnux's p.hkl with one fixed
recipe: data reindexed into the COD setting (lowest-R1 integer matrix), non-H anisotropic, H fixed,
EXTI, MERG 2, three rounds of SHELXL's suggested weights. It records R1/wR2/GooF/EXTI/WGHT/residual
density/R(int)/R(sigma)/K of the strongest bin and a fixed-model R1(F) (|Fc| of the COD model as
published, gemmi). Reported, never scored. The report gets a small-molecule table; compare lists
SHELXL R1/wR2/GooF/EXTI deltas; report/compare fill the check in for older runs.
COD entries matched by Niggli-reduced cell and space group: aspirin 7050897, citric acid 5000063,
HEPES 2224210, YAG 2003066, L-cystine 1513328 (2005, replaces the 1959 model for refinement),
cytidine 2001311, 3,5-dinitrobenzoic acid 4510615, L-alanine 2104782, metformin HCl 2108029,
NiCl2(dppe) 2012031. cuhf2 has no reference cell and no match: left without one.
SHELXL is called, not shipped (site key "shelxl" or PATH; it comes with CCP4).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The error model was fitted once, on deviations from a mean that weighs every full by its COUNTING
variance, and the outlier test's median took the same weights. Where equivalents disagree beyond
counting statistics, the low-count observations dominate that centre: on a strongly absorbing crystal
(YAG, Ia-3d, equivalents spread over a factor 100 after scaling) the mean of (0 4 0) sat at 21k among
observations from 11k to 1.9M, a ran into its bound (100), b came out at 800% internally, and the
six-sigma test about the biased median removed 63% of the observations - the strong ones.
Now, after the first fit, the model is refitted about the model-weighted mean (each full weighted by
the variance the fitted model gives it at the reflection's mean) until a and b settle, and the
rejection median takes the same weights. Where counting statistics are right nothing moves. Following
Blessing (1997) J. Appl. Cryst. 30, 421-426. GPU path: the refitted means are uploaded (SetEmMean).
Measured (SHELXL R1(>4sigma) against the COD model, sm-a's harness; rc174 scaling):
YAG 0.556 -> 0.127 (XDS 0.083 merged), rejected 8055 -> 53, normalised deviations calibrated
(median |z| 0.62-0.69 in every intensity decile); aspirin 20 keV 0.0964 -> 0.0958; citric acid
0.161 -> 0.159; HEPES 0.0903 -> 0.0899; L-cystine 25 keV 0.1425 -> 0.1456;
aspirin 25 keV 0.094 -> 0.106 (fixed-model R1 0.107 -> 0.173): its strong equivalents split into two
frame-dependent populations from the per-frame partial scaling (sm-a's dq-smallmol), which the old
under-sized sigmas happened to cut; with that scaling fixed (f69339ce6 + pooling, --no-scale-partials)
this change is neutral to better on every small molecule (aspirin 20 .0618 -> .0586, aspirin 25
.0456 -> .0454, citric .1093 -> .1006, HEPES .0703 -> .0697, YAG .649 -> .222; SHELXL GooF ~1.1).
=> ship together with the scaling fix.
Proteins (GPU full runs): CC1/2 and R_meas unchanged to 0.002; ISa myob 9.06 -> 8.20, thau 52.5 -> 47.5,
cytc 25.8 -> 25.4, lyso 29.4 -> 29.4 (still above XDS's 5.2 / 44.5 / 31.8 / 28.3 except cytc).
CPU build gives the same statistics as the GPU build on aspirin 20 keV and myob.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
Replaces the prototype --no-scale-partials switch with a rule read off the data: a merge whose
sweep holds fewer than 50 rocking events per frame takes its per-frame scale from the fulls alone
(scale-fulls, a sparse frame fitted over its neighbours - PoolHalfWidth), and logs that it did.
The rocking-event walk already run for the smoothing window now also returns its event count.
Within one rocking curve a partial's scale and an error of the partiality model are the same
thing; on a sparse fine-sliced sweep the partial fit takes the one for the other and imprints an
hkl-dependent bias common to all equivalents. Measured populations: small-molecule sweeps 2.6-23
events per frame, protein sets 84-900; on the proteins the fulls-only scale leaves model R-free
unchanged (+-0.002 over the smoke tier's open-arm sets) but lowers ISa, so they keep the partial
scale. p.mtz md5 unchanged on the three profiling sets, GPU and CPU builds.
SHELXL against the COD structures, R1(>4sig) rc174 -> this: aspirin 20 keV 0.096 -> 0.062,
aspirin 25 keV 0.094 -> 0.046, citric acid 0.161 -> 0.109, HEPES 0.090 -> 0.070 (XDS 0.030-0.038);
SHELXL's weight a comes off its 0.2 cap on all four. CPU and GPU paths agree.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The same source built with and without -march=x86-64-v3 gave different results on battery sets
(9zmu axis-harmonic supercell arbiter fired in one build only; 8rud resolution cut 1.69 vs 1.70 A;
myob_x06da_powder_1 cut 1.07 vs 1.42 A; 7mzt short-axis first pass; lcystine_x10sa_20keV indexed
vs no lattice), against the rule that no decision may depend on compiler flags.
Two mechanisms, found by building the merged tree four ways (baseline, x86-64-v2, x86-64-v3,
x86-64-v3 -mno-fma) with and without -ffp-contract=off and comparing p.mtz:
- FMA contraction. GCC contracts a*b+c whenever the target has FMA. With -ffp-contract=off the
x86-64-v3 build gives a p.mtz byte-identical to the baseline build on 8 of 9 sets (7mzt, 8rud,
9zmu, insu_I_x06da_5keV_2, myob_x06da_powder_1, myob_x10sa, cytc_x10sa, thau_x10sa_16keV), on
both the GPU and the CPU build. Cost: none measurable (user core-s, CPU build, x86-64-v3 vs the
same with -ffp-contract=off: 1877/1874, 3285/3253, 2212/2195 on myob/cytc/thau; GPU likewise
within noise). Set project-wide for C, C++ and CUDA host code; MSVC does not contract under
/fp:precise.
- SIMD width. lcystine still differed: baseline and x86-64-v2 (128-bit) agreed, x86-64-v3 with or
without FMA (256-bit) disagreed - Eigen's HouseholderQR in the FFT indexer's candidate refinement
(PostIndexingRefinement.cpp) reduces column norms over all spots in packets of the target width.
Replaced by the 3x3 normal equations summed in spot order in double. All four builds now agree on
all nine sets.
Changes results of the default x86-64-v3 build (contraction off); lcystine_x10sa_20keV now gives no
lattice in every build (its first pass is a knife-edge: 0/60 vs 9/60 validation frames before).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
Prototype switch --no-scale-partials: skip the per-frame scaling of the partials and let the
fulls (scale-fulls) carry the per-frame scale, as XDS does. Without partial scaling a frame whose
fulls number fewer than MIN_REFLECTIONS is fitted over the nearest frames on either side that
together hold enough (PoolHalfWidth), on the host and in the GPU kernel. Default path unchanged
(p.mtz md5 identical on the three profiling sets).
Why: on fine-sliced small-molecule sweeps the per-frame partial scale and the partiality model are
degenerate within a rocking curve, and the fit swings G 0.23..1.2 with a 180 deg period (XDS's own
frame scale: 0.79..0.99). That imprints an hkl-dependent bias common to all equivalents, which
R_meas/CC1/2/ISa cannot see but a refinement against the known structure does. And scale-fulls
never fitted a frame on such data: a full is filed under one frame, about 8 per frame, below
MIN_REFLECTIONS, so every frame kept G = 1.
Measured with SHELXL refining the COD structures (R1 >4sig), default -> switch:
aspirin 20 keV 0.096 -> 0.062, aspirin 25 keV 0.094 -> 0.046, citric acid 0.161 -> 0.109,
HEPES 0.090 -> 0.070 (XDS 0.030-0.038). Proteins lose ISa with the switch (myob 9.1 -> 7.6,
cytc 25.8 -> 13.7), so it is not a default; the choice is to be made from the data.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The first version judged "the same lattice" by class and centring, so an
F-centred cubic re-index (57/60) and pass 1's triclinic primitive of the very
same lattice (60/60) counted as different, and the unconstrained primitive won
the frame count - 6oel lost its cubic setting (R_meas 36.6 -> 42.1 %). The
two are now compared on their Niggli-reduced primitive edges (within 2 %, the
battery's lattice-identity test), so a symmetric setting never loses to its own
primitive, while 9qw8's C-centred re-index (reduced edges 5.7 % off pass 1's)
is still judged against pass 1's lattice on the frames.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
A promoted point group whose <|L|> reads below 0.375 (NOT_READABLE) previously
raised no TWINNING warning at all; a twinned subgroup whose law the promotion
absorbed predicts the same data, so the adopted group is not confirmed. Warn.
The twin-immune zone control reading more compressed than a perfect twin's
acentric population (0.541) is something no twin fraction produces (overlap or
neighbour correlation); genuine symmetry then reads acentric in its zones too
(measured: a genuine 622 with its control at 0.528 read its 2-folds at -950 to
-2640 nats). Such zones are now marked ambiguous in the text and the zone
decision line. Report-only: no decision and no output file but the report changes.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The non-reference settings of the adopted point group were offered only when the cell hosted their
rotations to within ~0.1 deg, while the reference settings - which hold the very same rotations -
were offered without asking. A cell refined free after integration a few tenths of a degree off 90
therefore lost every setting but the reference one. On an orthorhombic set whose measured screws
lie on a and c and whose b row was never recorded, that left P2(1)2(1)2(1) as the only candidate
covering both screws, and it was reported as determined. With P 21 2 21 offered, the two tie, the
b-axis screw is reported as undetermined, and the model check uses the setting the data describe.
The point-group stage still asks the cell whether a rotation set it adds is hosted; only the
setting enumeration within an already chosen set stops asking.
Validation: myob/cytc/thau x10sa p.mtz byte-identical (GPU); 7mzt fail -> unscored (b screw
undetermined, P 21 21 21 or P 21 2 21); 5cc8 unchanged; [SearchSpaceGroup] 23 cases pass incl. a
new section with a cell 0.2-0.3 deg off 90.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The space-group hero card (and the "indistinguishable" line) render the short
Hermann-Mauguin symbol as rich text: screw axes subscripted, rotoinversions
overlined (P2_1/c, P4_12_12, Fd-3m). It is built from gemmi's spaced symbol
with short_name()'s monoclinic shortening applied; checked against all 564
gemmi settings to reduce to short_name() once the markup is stripped.
The report's SUMMARY now also records its pathology rows as typed checks
(ReportDocument::checks): PRESENT exactly when one of the row's codes is in
PATHOLOGY_FLAGS, ABSENT when it was measured and did not fire, UNDETERMINED
when it could not be measured. A flag no row stands for gets a check named
after its code, so the list never shows less than PATHOLOGY_FLAGS. The text
report is unchanged. The viewer builds the document once, renders the text
from it and shows the checks phenix.xtriage-style: a green/red/grey light per
pathology, the summary words and fired warnings in the tooltip.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Two knife-edges let a sub-0.05 px change of the measured beam centre turn a
P1 crystal (9qw8) from a pass into a merge with R_meas 60000% / ISa 0:
- The refined-geometry pass re-indexes de novo and fell back to pass 1's
lattice only when the re-index scored under the 1/6 validation floor. The
long-axis rescue lifted a related C-centred 303 A cell to 41/60 - over the
floor - while pass 1's lattice indexes 55/60 at the same geometry. Pass 1's
lattice is now scored as a hypothesis of its own whenever the re-index
found a DIFFERENT lattice (class + primitive volume within 2 %), and is
integrated when it indexes more validation frames. The same lattice found
again is kept as the re-index refined it, so ordinary crystals are
untouched.
- The two-pass quality guard's "going back to the header geometry" re-ran
pass 1 de novo, a hypothesis nobody had judged: at the centre where pass 1
indexed 56/60 it found 5/60, flipped the axis sign and shipped garbage. It
now forces pass 1's whole indexing result, which is what the guard preferred.
9qw8: P1 1.71 A on -march=x86-64-v3 and on no-march GPU builds (both failed
before, one catastrophically). myob/cytc/thau p.mtz md5 unchanged (GPU).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
Battery: 20261003-1424_581e1c_perf-merge-full (+_private) vs rc174 built with the same flags:
no verdict change except the second scaling engine's GPU OOM on the three largest sets,
removed in 0f728ecab (8a1a/8qaw/8tyy pass again: 20261003-2032_581e1c_perf-oom-fix).
One-time result change from the GPU/CPU-parity beam-centre walk (b2).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The search merge was cut at the first thin shell whose mean <I/sigma> fell under 1. On a merge
whose <I/sigma> is flat near 1 - its ISa has collapsed, as on a strongly absorbing garnet - the first
single shell to dip is decided by noise: the same crystal cut at 1.19 A without -march and 0.85 A with
-march=x86-64-v3, and the coarser cut left each glide zone fewer than 20 absences, so Ia-3d became
unjudgeable and I4(1)32 was adopted. The cut now comes from a non-increasing (pool-adjacent-violators)
fit of the shell means, which moves only as much as its input does; that garnet's profile never falls
under 1, so its search sees the full range and adopts Ia-3d under both builds with identical
candidate tables. On monotone profiles nothing changes: myob/cytc/thau x10sa keep their cuts and
p.mtz byte for byte.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The all-observation arm of the space-group search and the P1 cross-check
were made on a second RotationScaleMerge engine beside the run's own
(871347b7a). That engine holds a second device copy of every observation,
and on the largest sets the two no longer fit a 16 GB card: 8a1a, 8qaw and
8tyy (55-125 M partial observations) stopped with an out-of-memory error
in scaling. Measured on a quiet box the engine bought 0.7 s (cytc) and
1.0 s (thau) of tail and nothing on myob, which does not justify a memory
budget, so it is removed and both merges run on rsm in sequence, as
before 871347b7a.
Kept from 871347b7a: the GPU scaling's own non-blocking stream, the
cross-check not writing per-frame G/CC/mosaicity back (the per-image table
still describes the merge that was written), the restored scaling
iteration counts, and the anisotropy analysis beside the other analyses.
p.mtz byte-identical to before on myob/cytc/thau, GPU and CPU builds;
p_plot.txt unchanged apart from the GPU bkg column.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
cpuinfo_max_freq is the highest boost state, not the base clock -
mislabeling it "nominal" understated the gap between the two on an AMD
box using acpi-cpufreq (this machine: 5.08 GHz boost vs 3.40 GHz base).
Prefer, in order: the model-name string, cpufreq's own base_frequency
(intel_pstate, kHz), ACPI CPPC's nominal_freq (MHz; what acpi-cpufreq
exposes instead), else "unknown". cpuinfo_max_freq is now reported
alongside, separately, as "max boost", never folded into "nominal".
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Collect CPU model, nominal/base frequency (with its source), physical
core/thread/socket counts, RAM and GPU name(s)/memory at run time
(Linux via /proc, /sys/.../cpufreq, nvidia-smi; a short best-effort on
macOS via sysctl), store it under manifest.json's new "hardware" key,
and render it as a line in the report header. Runs from before this
change have no such key; the report prints "not recorded" for them.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
ApplyCellSurface (detector modulation, time x detector, crystal-frame SH and
goniometer-frame absorption) spends most of its time in two passes per round:
the per-group reference sums and the per-(block, cell) fit sums. With a GPU
both now run on the device (RotationScaleMergeGPU::Surface*) over the same
terms in the same order:
- reference: one thread per ASU group, walking a group-order permutation of
the terms in fulls order;
- fit: each subset cut into the host's reduction blocks with every block's
terms ordered by cell (stable counting sort, host); a per-term kernel forms
w*Is, w*Iref, Iref and one thread per (block, cell) runs the two fma chains;
the block slots are added per cell in block order.
Every rounding is spelled out (__dmul_rn/__dadd_rn/fma) to be the one the host
build makes: GCC at -march=x86-64-v3 fuses swI's multiply-add only in the
parity-filtered copy of the reference loop, and both fit sums.
Host side, exact on both paths: the 19 serial nth_element selections of the
shell edges become one parallel sort (same order statistics), the per-term shell
lookup runs on all threads, and the gate's per-shell CC is one walk over the
groups instead of one per shell.
Exact: p.mtz md5 identical to the oracle on myob/cytc/thau x10sa, GPU build
(all CUDA architectures) and CPU build. CorrectionSurfaceGPU test checks the
device sums bit for bit against an explicitly rounded host loop.
Measured (cytc/thau, two interleaved A/B pairs, box at load 13-25 from sibling
work): ApplyCellSurface host core-seconds -83% on cytc; SG adoption -> writing
reflections 4.58->3.75 and 3.85->2.54 s (cytc), 3.15->2.28 and 2.29->2.03 s
(thau); RSM final merge -0.7..-0.9 s and P1 cross-check -1.5..-1.8 s on cytc.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The hot-pixel step of the pre-scan (MaskDefectivePixels) on a GPU build:
- The device half is built once, before the workers start (HotPixelFinder::PrepareDevice),
instead of by the first worker's frame while the others waited. The unmasked pixels
grouped by key are sorted on the device (stable radix sort: the same order the host
fill gave) instead of scattered on the host and uploaded.
- No per-frame host round trip: the ring-sector levels and lit thresholds are made on
the device from the order statistics, and the per-key frame and level sums are kept
there too; frames queue on their workers' streams and the per-pixel accumulation is
ordered by an event instead of a host synchronisation.
- The mask: the chance rate's per-ring counts are summed on the device, and only the
pixels the tests can pass (error value on most frames, or lit on at least
min(max(2, k_chance), valid frames)) come back with their sums - not the five
per-pixel arrays (470 MB pageable on a 16 Mpx detector). The host tests run on them
unchanged.
The threshold is written as fma(nsigma, noise, level) + offset on the host - what GCC
already contracted it to - and the device takes the same two roundings, so the levels
are bit-identical (and no longer depend on whether a compiler contracts).
Exact: hot-pixel mask and p.mtz byte-identical to f849e2d1b on myoglobin, cytochrome C
and thaumatin, GPU and CPU builds. Hot-pixel step (GPU, box at load 18-23, interleaved
A/B, two pairs each): 0.94-1.32 s -> 0.65-0.81 s; frames 0.37-0.46 -> 0.18-0.23 s,
mask 0.21-0.34 -> 0.08-0.18 s. Device memory of the finder: 543 MB as before, plus a
~0.36 GB transient for the sort while it is built.
Tests: [HotPixelFinder] (HotPixelFinder_DeviceMatchesHost bit-exact), [ShadowFinder],
[BeamCenter].
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The pre-scan's spot measurement (spot width and bandwidth, powder rings,
spot-symmetry beam centre) ran the CPU preprocessor and adaptive spot finder
on every sampled frame, ~13-19 core-seconds per 16M-pixel run. In a GPU build
it now uses the same engines as the image loops: the frame is decoded and
preprocessed on the device (with the host decoder as fallback), the spots
are found and extracted there, and the preprocessed image comes back only
when the spot width needs it.
This is a choice of where, not of what: the device finders reproduce the
host's spot list exactly (integer ring sums, the same connected components
in the same order). Checked frame by frame with both engines side by side on
myob/cytc/thau x10sa and all 21 smoke-tier sets (~540k spots, HDF5, CBF,
marCCD, SMV, pink beam, 3.8 keV): identical spot lists in identical order on
every frame, and an identical preprocessed image on every width frame.
p.mtz md5 identical to the oracle on myob/cytc/thau, GPU and CPU builds.
Measured on a loaded box (load 20-30): the spot measurement takes 0.8-0.95 s
instead of 2.5-3.6 s and no longer competes for the cores the beam-stop mask
and the background beam-centre fit run on; the pre-scan falls from 3.0-4.9 s
to 1.7-2.4 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The two pre-scan steps that were still CPU-bound in a GPU build now run where the
projection already is.
- FindBeamCenterFromBackground: the per-iteration binning pass and the two clip rounds
run on the device (BeamCenterBackgroundGPU); the fit itself stays on the host. Each
cell is summed in the host's order (pixel order within the host's row blocks, blocks
in order), and the per-pixel cell/derivative formula is shared (BackgroundBand.h).
The angles come from BackgroundAtan2 (IEEE ops only) instead of atan2f, and both
translation units are compiled without FMA contraction, so host and device give the
same bits: 0 of 6.5 M pixels in a different cell, identical walks on the three
in-house rotation sets. With glibc/CUDA atan2f and default contraction ~30 pixels per
16 Mpx sweep changed cell and the fitted centre moved by up to 0.05 px.
- ShadowFinder::GetMask: the whole mask (pooling, ring medians, components, morphology,
hole fill, arm search) runs on the device from ShadowAccumulatorGPU's projection
(ShadowMaskGPU), so the 360 MB projection no longer comes back; the mean projection is
divided on the device too (same bits). The two small fits over rings and sectors
(BlockedOutTo, HarmonicFit) are shared with the host path in ShadowFinderInternal.h.
Integers, comparisons, sorts and components are exact; the polarization trig, the
Poisson log and the arm-search azimuth are not, so a pixel at a threshold can differ.
The one-time change against the previous CPU arithmetic (BackgroundAtan2, no
contraction), measured on the myoglobin, cytochrome C and thaumatin rotation sets:
ring centre moves 0.002-0.045 px (fit sigma 0.75-1.2 px), beam-centre capture
0.01-0.04 px; beam-stop mask differs on 31 / 144 / 53 pixels of 259k / 144k / 198k
(25 of the myoglobin ones are GPU-vs-CPU arithmetic in the mask, the rest follow the
centre); hot-pixel mask identical. Spot width, integration radii, bandwidth, beam-centre
arbitration, indexing, space group, cell, resolution and the merged statistics table
are identical; only the error model moves in its 4th digit. CPU build: the same
centres and decisions.
Timing (GPU, box at load 30-38): ring walk 0.54 -> 0.23-0.27 s, mask 1.24-1.44 ->
0.18-0.22 s, beam-centre capture walk 1.1-1.3 -> 0.31-0.35 s.
Tests: ShadowFinder_DeviceMaskMatchesHost, BeamCenterFromBackground_DeviceMatchesHost
(bit-exact), plus [ShadowFinder], [BeamCenter], [HotPixelFinder].
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
The standalone GPU azimuthal integration did three shared-memory atomics per
pixel on the same few ring addresses, which is what it was limited by. It now
reads four pixels per thread as vector loads and keeps a running total per ring,
flushed when the ring changes - the scheme the adaptive finder's ring pass
(reduce_rings_shared) already uses. The npix % 4 leftovers are done one at a time.
Used wherever the fused adaptive engine is not (fixed-threshold spot finding, the
broker's non-adaptive path). Measured on a 16 Mpx sweep (1800 frames,
--no-adaptive-spots, RTX 5080): 843 -> 295 us per call (min 621 -> 196 us).
Not bit-identical, and the old kernel was not either: float atomics arrive in any
order, so two runs of the OLD kernel already differ by up to 1.7e-6 relative in
the per-frame profile; new vs old differs by up to 1.9e-6, the same order. Per-ring
pixel counts are identical. Default rugnux runs do not reach this kernel (p.mtz
md5 unchanged on three sets); on the fixed-threshold path p_unmerged.mtz is
md5-identical to the old kernel's. New test: GPU vs CPU engine on a pixel count
that is not a multiple of four, with masked and saturated pixels.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
Exact: p.mtz and the pre-scan products (shadow mask, mean projection,
defective-pixel mask, ring and capture centres, compared as hashes and
hex floats) are bit-identical to rc174 on three in-house rotation sets,
GPU and CPU builds.
- ShadowFinder::GetMask: the serial parts run in parallel - connected
components by row band joined with union-find (both the shadow and
the transmitting-arm searches, and the hole fill), ring binning and
the harmonic sector gather by blocks, gap bridging by line; ring pixel
counts read off the ring offsets. Mean projection filled in parallel.
- ShadowFinder host accumulation: one band-locked projection instead of
a 20 B/px shard per pre-scan worker (2.7 GB zeroed and folded on a
16M detector); SetShardCount and the shard argument are gone.
- FindBeamCenterFromBackground: the usable-pixel test is made once, the
in-band pixels are kept in pixel order so the clipping rounds no
longer sweep the whole detector, the 67 MB cell map is gone and the
per-iteration block fold runs in parallel - same sums, same order.
- HotPixelFinder::GetMask: the chance-rate counts in parallel (integers).
Measured on a loaded box (load ~25 from other jobs), pre-scan window:
GPU 5.9-6.5 s -> 3.2-3.4 s, CPU 8.4-9.0 s -> 6.1-7.4 s. The GPU-build
pre-scan now ends with its background spot measurement (CPU spot finder
on ~120 frames, ~13 core-s on 8 workers).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB