ComputeAsuGroups sorted one key per distinct raw hkl on one thread, once per
merge (6-8 merges a run). Ties are now broken on the run index, which makes the
order total, so ParallelSort returns the same sequence; and the result never
depended on the order within a tie (every run in it reduces to the same ASU
reflection). Measured on the GPU build, both sorts side by side on the same
input in every merge of 6 in-house runs: keys identical in all calls; a cytochrome C
run 0.230 -> 0.058 s, a thaumatin run 0.186 -> 0.058 s, summed over its merges.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB