The rigid-body pool put all of its engines on the calling thread's current
card, so a validation used one GPU whatever the machine had. Now, by fixed
rules decided up front and never by momentary free memory:
- RigidBodyGPUPool::Create puts engine i on card (d + i) % count, d being
the calling thread's card; the pool restores that card afterwards
(an engine's constructor sets its own) and an engine is released on
its own card.
- The threads that run the null's replicates are pinned with pin_gpu
(which also binds them to the card's NUMA node where that is enabled),
replicate thread t to card (d + 1 + t) % count, and Acquire() hands a
thread an idle engine on its own card where there is one, any other
otherwise. With one card this is exactly the previous back-of-the-list
choice.
- The validation started on the forecast runs on card 1 % count, so with
two cards or more it is on the other card from the first; the up-front
memory rule becomes: twice the planned engine bytes within a quarter of
the cards' total memory taken together.
The engines' kernels are deterministic (no floating-point atomics) and an
engine's result does not depend on which engine it is, so on cards of one
model the numbers are those of one card; what changes is only where the
work runs. On this one-card workstation the multi-card path cannot be
exercised: md5 of p.mtz and the validation outputs are identical to the
base, and the card-count arithmetic was checked by reading only.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
Where a model that fits is written on other axes than the data, the files
take its setting and the validation is made a second time on the
relabelled data. That second validation depends on the first only through
the setting and the indexing the first settles, and both are known long
before the first has finished: the change of basis right after the frame
scoring, the indexing the probe prefers before the null. So ValidateAgainstModel
now reports them (ModelFrameForecast, through ModelValidationSchedule::
on_forecast), and the second validation is started there, on the
relabelling AdoptModelFrame and relabel_output would make, applied to a
copy of the merge - beside the first one's null, real fit and maps.
It is kept only where the first decides exactly what was forecast (the
model fits, same change of basis, same indexing; a model asserting the
other enantiomorph is not forecast, as the label is decided last, on the
anomalous map). Until then its log is held (Logger::Buffered, replayed
where the serial run logged it) and its files wait on a gate
(ModelValidationSchedule::write_gate) placed before the first map is
written; otherwise it is released with false and returns unwritten, and
the serial validation runs as before.
GPU memory: two validations at once take twice the rigid-body engines.
The parallel start is decided up front from sizes, never from what is
free: allowed where twice what the first pool asked for (bytes per engine
times the engines wanted) fits a quarter of the card's TOTAL memory, the
share one validation may take. Threads: both validations submit to the
one ParallelFor pool from threads outside it (the second runs on a
std::async thread, as the null's replicates do), so no pass runs inline
on a pool worker and the pool's size bounds the workers.
Measured on the loaded 16-core workstation, TIMING model validation:
8sa8 30.6 -> 19.7 s, 8xtg 19.9 -> 13.9 s, 9ea5 22.1 -> 16.4 s (this and
the previous commit together). p.mtz, p.hkl, p.cif, the three maps,
p_maps.mtz and p_model.cif/pdb md5-identical to the base on 8sa8, 8xtg,
9ea5, 7qis and myob_x10sa; the validation's log lines identical as a set
(the null's replicate lines were already in completion order).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
--model validation (battery-only for users) was 18% of the battery's time. Every
number it produces is unchanged to the bit (p.mtz, maps, placed model and every
model-validation line of the report md5/diff-identical on 11 open sets); only
when and where the work runs changes:
- The bulk-solvent grid fit (FitModelScale, most of the CPU time) fits each
solvent pair on a copy of gemmi::Scaling's target that takes
|Fcalc + k_sol exp(-b_sol s^2) Fmask| once per pair instead of at every
solver evaluation; same expressions, same types (new test checks a grid
point against gemmi's own Scaling fit with ==).
- Fcalc density and the solvent mask are made on two threads; the model's
structure factors beside the GPU engine reservation.
- The indexing probe fits the relabellings concurrently.
- The null's replicates run beside the real model's placement (they start
from a snapshot of the model as read); one GPU engine per replicate plus
one for the real fit instead of a cap of 4 (engines are interchangeable
and deterministic).
- The 2mFo-DFc, mFo-DFc and anomalous maps are made and written
concurrently; the placed model is written beside the reflection files.
- A rigid-body zone whose solvent-mask grid needs gemmi's shrink is sent to
the CPU when the engines are reserved (ModelMaskGPU::ShrinkIsNoOp), instead
of failing on the GPU and validating everything again on the CPU - the
same CPU result, without the wasted first attempt.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
A pure move. ModelValidation, RigidBodyRefine, RigidBodyGPU, ModelFFT, ModelGrid,
ModelScaling, ModelMaskGPU, ModelScaleGPU and SigmaA - everything that works on an
atomic model - become the JFJochStructureRefinement library, linked by
JFJochImageAnalysis. WriteModel (the placed-model mmCIF/PDB writer) goes to writer/
as its own small JFJochModelWriter target, so JFJochWriter, which a writer-only build
compiles, does not gain a gemmi dependency. Only include paths and CMake lists change.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi