ParallelChunks cuts a range into one chunk per worker, so a floating-point
sum folded per chunk and then added up rounds differently at another -N.
ParallelBlocks cuts it by n alone (n / min_per_block blocks, at most 256);
each block folds into its own slot and the slots are added in block order,
so the sum has the same bits at any thread count.
Converted:
- RotationScaleMerge::ApplyCellSurface: the per-cell cross/ref2 sums of the
surface fit (modulation, absorption), previously per-thread partials cut by
ThreadsForWork(idx_all) threads.
- PostRefine: the scale-scan cost grid (per-chunk slots cut by -N).
- PostRefine joint solve: Ceres at a fixed 16 threads (as the rotation
indexer's chain) instead of -N. Ceres sums cost and gradient in
4 * num_threads pieces, so the split no longer follows -N. The pieces are
handed to its threads in scheduling order, which only one thread makes
exact; one thread was measured at +1.0-1.4 s of post-refinement on myob and
lyso (0.5 -> 1.9 s, 1.1 -> 2.1 s), so that channel is left.
- IndexAndRefine supercell probe: per-frame probes kept by image number and
summed in frame order, instead of added under a mutex in completion order;
the primitive is the highest probed frame's instead of the last finisher's.
Integer reductions and per-item passes are unchanged; the three ParallelSort
callers already break ties on the index.
Validation (myob, cytc, lyso, sparse; GPU and CPU builds; -N 8/16/32):
p.hkl, p.mtz, p_P1.mtz and p_unmerged.mtz are md5-identical across -N, and
identical to the base 351be7de0 at every -N. The base was already -N
invariant on these four sets (GPU -N 1..32, CPU lyso -N 2..32), so no output
moved: the converted sums differ from the old ones only in last bits that
never reached a written float.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C