The null - nine random placements of the model, and the real model's side
of the comparison - is scored to a coarser limit (3.5 A, or where the free
set allows), and those structure factors still ran on the CPU, one model
at a time on each replicate thread. Where the d_min structure factors are
on the GPU, a second ModelStructureFactorsGPU engine at the null's limit
is now made, on the same card, before any replicate starts; every
replicate and the real side go through it, so all of them are computed
the same way. It is decided like the first engine, from the card's total
memory: the two together must stay within half of it, else the null stays
on the CPU (logged). The replicates' threads may sit on other cards
(modelpar's 8d6e9d3a4); each call works on the engine's own device. Where
the null is scored to d_min it shares the d_min engine, as before.
Measured (validation time, before -> after; CPU path for reference):
P3_2 1.55 A: 11.2 -> 9.0 s (CPU 17.7)
P2_1 oblique: 5.1 -> 4.5 s (CPU 7.9)
C2 1.11 A: 12.9 -> 13.2 s (CPU 23.3) - its null time is in the
scale fits, which stay on the CPU
p.mtz md5-identical in all four variants per set (before, after, CPU,
-N 8); maps and map MTZ bit-identical after vs -N 8; the null's sigma
moves only in the second decimal (+113.39 -> +113.38, +86.13 -> +86.14).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi