Floating-point summation order only (tier F): the seven global sums of RefineDecay (slope fits,
held-out disagreements), the per-batch sums of FitRelativeBCurve, and the per-shell and
per-(batch, shell) sums of MeasureRadiationDamageB were single-threaded walks over every full.
They are now FixedBlockSum: the split depends on the number of fulls alone (ReductionBlocks,
ParallelBlocks), the block partials are merged in block order, so the sums have the same bits at
every thread count and on every machine - but not those of the serial walk.
Checked: p.mtz, p_P1.mtz, p.cif, p.hkl and the report identical to production on myob/cytc/thau
(GPU) and myob (CPU), and identical between -N 4/8/default (GPU cytc, myob; CPU myob). On these
sets the decay correction is not adopted (negligible or not cross-validated) and the monitor's
printed values do not move, so the F-tier change is not visible in their outputs; a set where the
decay correction is adopted will move in the last bits of corr.
Clean branch: p.mtz md5 identical on myob/cytc/thau/8a1a (GPU), cytc -N 8, myob -N 4 (CPU).
Measured (prototype with spans, 8a1a, 2 interleaved pairs, load 8-10): decay sections 2.70 -> 1.08 s,
radiation-damage monitor 1.04 -> 0.86 s, RSM total 37.1 -> 35.6 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi