The space-group search was almost entirely single-threaded. The per-reflection work now runs in
parallel over fixed chunks, while every floating-point accumulation stays serial and in the same
order as before, so the output is bit-identical.
- AnalyzeLTestUnderMerge takes nthreads (default 1): the usable test, the orbit means and the
partner search run in parallel; the 20 resolution-shell passes become one walk.
- Stage B ClassifyReflection: one byte per reflection, candidates scored in batches of nthreads.
- FindRowModulation: cone test and cosines in parallel.
- PrepareMerge: parallel flatten; the sorts use inline {d, index} keys with the same comparison.
Measured: SearchSpaceGroup 14.7 -> 4.8 s on 8tyy (instrumented build). Wall clock, median of 3
interleaved runs vs rc175, timed inside the GPU lock (-march=x86-64-v3, CUDA): 8tyy 136.9 -> 127.0 s,
8a1a 81.8 -> 80.4 s, cytochrome C 18.8 -> 18.9 s (unchanged). 8qaw max RSS 21.7 -> 21.9 GB.
p.mtz md5-identical to rc175 on myoglobin, cytochrome C, thaumatin, 8a1a, 8tyy, 9gdj and 8qaw
(GPU build), and on myoglobin and cytochrome C in a CUDA-off build.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi