FitErrorModelGPU puts the samples in the host's rank order with four stable
radix passes, so the equal-count bins hold the same samples; computes every
normalised deviation with the host's rounding, so every bin median is the
host's to the bit; and hands the per-bin terms to the host's own update step,
now ErrorModelUpdate (a pure extraction from FitErrorModel). The one thing
that differs is the order each bin's two sums are added in: a fixed tree on the
device, the order nth_element happened to leave the bin in on the host. a and
b^2 can therefore differ from the host fit in their last bits - measured on two
synthetic pools: relative difference 2e-16 and 3e-13 in a, 5e-16 in b^2. The
device fit is deterministic (same pool, same bits).
p.mtz md5 nevertheless unchanged on myob, cytc, 8a1a and 8qaw (GPU build): the
difference does not survive the float rounding of the merged sigmas there. A
knife-edge decision elsewhere could still see it, which is why this is kept as
its own commit.
The host fit took 0.4-0.6 s per fit on the large merges, about half of it the
initial equal-count split, with only 16-way parallelism in the iterations.
Device peak about 72 bytes per sample (the rank sort; each stage's buffers are
freed when it is done); the free memory is checked first and a device without
it stops with a message naming the CPU build.
Kept last on the branch so that it can be dropped on its own.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi