278c488c375dcc2f2805101056c9a8a4499fccd3
7
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
278c488c37 |
Merge branch 'modelpar' into gpusf
Both sides kept: modelpar's parallel basis scoring, the second validation started on a forecast beside the first (write gate, schedule parameter), the multi-GPU placement and delete-before-rewrite; gpusf's GPU structure factors, maps and null engine, failure-instead-of-restart, and the merge engine released before the validation (now just before modelpar's ValidateAgainstModel call, after the forecast lambda is set up). Placement: each validation's structure-factor engines (d_min and the null's) are made on the card of the thread that runs it - the main thread's for the first validation, card 1 % count for the speculative second, which pins itself there - so two validations on two cards use both, as the rigid-body pools do. One memory rule for both, per card, from total memory, up front: a validation plans at most half of its card - its structure-factor engines a quarter together (was half for the d_min engine alone), its rigid-body engines a quarter (RigidBodyGPUPool) - and the second validation runs beside the first only where twice the first's plan (rigid-body planned bytes + the d_min engine, twice it where a null is coming, the null's engine being no larger) fits half of all cards' memory together (was: twice the rigid-body plan within a quarter). The card's total is read once when the engine is made; a CUDA error there fails the validation like any other. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi |
||
|
|
d5fcf2f05b |
rugnux: model validation's structure factors and maps on the GPU
ModelStructureFactorsGPU computes what compute_model_factors() and
map_from_coefficients() compute on the CPU - F_calc from the model's
density (IT92, Refmac-compatible blur, unblurred as prepare_asu_data()
does) and F_mask from the Refmac bulk-solvent mask, both on the
reflections prepare_asu_data(d_min) lists, in its order; and a map from
ASU coefficients on the grid get_size_for_hkl(coef, 0, 3.0) sizes - on a
device. Made once per cell, group, resolution and model, then evaluated
as often as the coordinates change, so refinement or MR can call it in a
loop. The device is an explicit parameter; every call leaves the calling
thread's current device as it found it.
Pieces:
- ModelDensityGPU: the rigid body's deterministic brick gather, moved
out of RigidBodyGPU.cu into a component of its own (ModelMaskGPU's
pattern); the rigid body uses it unchanged. MAX_BRICKS_PER_AXIS 8 ->
16, so fine grids with high-B atoms (lysozyme at 1.2 A, a 0.9 A P1
cell) are no longer refused; existing zones are gridded identically.
- One copy of the content is gridded and the symmetry composed in
reciprocal space (SymmetryComposition), operators applied on the fly;
the mask is ModelMaskGPU (every image of every atom, islands, shrink).
- Maps: gemmi's get_f_phi_on_grid() in ZYX order on the host (the
coefficients written are the same), in-place cuFFT c2r, transposed back
to XYZ on the device. One map at a time, in the engine's buffers.
Decided once, up front, per card, from its TOTAL memory: the engine's
bytes (16 N + cuFFT work + reflections, N the larger of the structure-
factor and map grids) must be at most half the card - the rigid body's
engines take at most a quarter beside it. Otherwise, or where the gather
cannot grid the cell, the CPU path runs, logged with needed vs total.
Anything to a resolution other than d_min (the null's 3.5 A fits) stays
on the CPU, so all replicates and the real model's side of the null are
computed the same way. A CUDA failure takes the existing path: the
validation restarts on the CPU.
Measured, model validation total per run (CPU path -> GPU), 16 GB card:
F432 215 A cubic, 1.30 A, 500^3 grid: 47.7 -> 15.6 s (two validations;
14.3 -> 2.4 and 33.4 -> 13.2, the rest of the second is writing the
three 0.5 GB maps); F_calc + F_mask 5.7 s -> 0.05 s
C2 1.11 A: 23.9 -> 13.6 s; P3_2 1.55 A: 18.2 -> 10.2 s;
P2_1 1.25 A: 13.4 -> 6.5 s; F4_132 328 A: 13.0 -> 5.5 s;
P6_5: 8.8 -> 4.2 s; P4_3 0.97 A: 4.6 -> 2.5 s; P1 0.92 A: 4.2 -> 2.2 s;
small P1: 3.2 -> 1.3 s; P6_1: 8.8 -> 6.2 s; lysozyme: 1.8 -> 1.4 s.
p.mtz md5-identical on all 13 sets. Against the CPU path: FC within
1e-4 of mean |F|, phases of the strong half within 0.003 deg, maps within
1e-4 (2mFo-DFc) and 7e-4 (mFo-DFc) of their rms; every logged R, CC,
FOM, k_sol and anomalous site list identical at the printed precision,
except where a rigid-body commit sat on an exact R-free tie (0.2155 ->
0.2155) and fell the other way (R-work 0.2127 vs 0.2129). The GPU result
is bit-identical run to run and with -N 8 (maps, map MTZ, placed model).
Peak device memory of the engine: 2.5 GB at 500^3 (process total peaked
at 14.4 GB with what the merge still holds).
Tests: ModelStructureFactorsGPU_MatchesCPU (five groups, 3.5 and 1.5 A:
same reflections, F_calc <= 1e-5 of mean |F|, F_mask 2e-7 rms, repeat
bit-identical), ModelStructureFactorsGPU_MapMatchesCPU (<= 5e-6 of rms).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
|
||
|
|
6025d51a48 |
ModelMaskGPU: gemmi's shrink step, so fine rigid-body zones stay on the GPU
The GPU bulk-solvent mask had no shrink step (SolventMasker::shrink(), set_margin_around() with Refmac's r_shrink = 0.8 A), so every rigid-body zone whose grid has an offset within 0.8 A was sent to the CPU. That is not only fine grids: in an oblique setting (P2_1, beta ~141 deg) the 3.5 A zone's lattice-plane spacing is 0.69 A, and on such a set the real fit and all nine null replicates ran their rigid bodies on the CPU. The shrink is now two passes over the grid: mark the solvent points with a macromolecule point among gemmi's near offsets, then turn into solvent every macromolecule point with such an edge point at any stencil offset (near or far, gemmi's split at the coarsest grid step). That is gemmi's rule in both of its branches, and the stencil is built on the host in double exactly as gemmi builds it. ModelMaskGPU::ShrinkIsNoOp() and RigidBodyGPUEngine::MaskSupports() are gone; the rigid body no longer refuses such zones. Verified (ModelMaskGPU_ShrinkMatchesGemmi): the shrink alone on gemmi's post-island mask is bit-identical, and the whole GPU mask equals gemmi's put_mask_on_grid() bit for bit (0 differing points) on the five test groups at 1.5 A and an oblique P2_1 cell at 3.5 A; repeats are identical. Measured on the oblique-setting set of the open arm (P2_1, 14k atoms, 1.66 A): model validation 17.2 s -> 5.2 s with the GPU structure factors of the next commit in place (the first validation's ten rigid bodies, ~12 s on the CPU, now take well under a second). p.mtz md5-identical; the null moves from +73.1 to +86.1 sigma (the replicates are refined on the GPU, which agrees with the CPU to rounding), verdict unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi |
||
|
|
8d6e9d3a44 |
ModelValidation: use every visible GPU - the null's engines and the second validation
The rigid-body pool put all of its engines on the calling thread's current card, so a validation used one GPU whatever the machine had. Now, by fixed rules decided up front and never by momentary free memory: - RigidBodyGPUPool::Create puts engine i on card (d + i) % count, d being the calling thread's card; the pool restores that card afterwards (an engine's constructor sets its own) and an engine is released on its own card. - The threads that run the null's replicates are pinned with pin_gpu (which also binds them to the card's NUMA node where that is enabled), replicate thread t to card (d + 1 + t) % count, and Acquire() hands a thread an idle engine on its own card where there is one, any other otherwise. With one card this is exactly the previous back-of-the-list choice. - The validation started on the forecast runs on card 1 % count, so with two cards or more it is on the other card from the first; the up-front memory rule becomes: twice the planned engine bytes within a quarter of the cards' total memory taken together. The engines' kernels are deterministic (no floating-point atomics) and an engine's result does not depend on which engine it is, so on cards of one model the numbers are those of one card; what changes is only where the work runs. On this one-card workstation the multi-card path cannot be exercised: md5 of p.mtz and the validation outputs are identical to the base, and the card-count arithmetic was checked by reading only. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi |
||
|
|
b5980c9ba0 |
rugnux: start the validation in the model's setting on a forecast, beside the first
Where a model that fits is written on other axes than the data, the files take its setting and the validation is made a second time on the relabelled data. That second validation depends on the first only through the setting and the indexing the first settles, and both are known long before the first has finished: the change of basis right after the frame scoring, the indexing the probe prefers before the null. So ValidateAgainstModel now reports them (ModelFrameForecast, through ModelValidationSchedule:: on_forecast), and the second validation is started there, on the relabelling AdoptModelFrame and relabel_output would make, applied to a copy of the merge - beside the first one's null, real fit and maps. It is kept only where the first decides exactly what was forecast (the model fits, same change of basis, same indexing; a model asserting the other enantiomorph is not forecast, as the label is decided last, on the anomalous map). Until then its log is held (Logger::Buffered, replayed where the serial run logged it) and its files wait on a gate (ModelValidationSchedule::write_gate) placed before the first map is written; otherwise it is released with false and returns unwritten, and the serial validation runs as before. GPU memory: two validations at once take twice the rigid-body engines. The parallel start is decided up front from sizes, never from what is free: allowed where twice what the first pool asked for (bytes per engine times the engines wanted) fits a quarter of the card's TOTAL memory, the share one validation may take. Threads: both validations submit to the one ParallelFor pool from threads outside it (the second runs on a std::async thread, as the null's replicates do), so no pass runs inline on a pool worker and the pool's size bounds the workers. Measured on the loaded 16-core workstation, TIMING model validation: 8sa8 30.6 -> 19.7 s, 8xtg 19.9 -> 13.9 s, 9ea5 22.1 -> 16.4 s (this and the previous commit together). p.mtz, p.hkl, p.cif, the three maps, p_maps.mtz and p_model.cif/pdb md5-identical to the base on 8sa8, 8xtg, 9ea5, 7qis and myob_x10sa; the validation's log lines identical as a set (the null's replicate lines were already in completion order). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi |
||
|
|
f956eb25e8 |
rugnux: model validation does the same work in less time
--model validation (battery-only for users) was 18% of the battery's time. Every number it produces is unchanged to the bit (p.mtz, maps, placed model and every model-validation line of the report md5/diff-identical on 11 open sets); only when and where the work runs changes: - The bulk-solvent grid fit (FitModelScale, most of the CPU time) fits each solvent pair on a copy of gemmi::Scaling's target that takes |Fcalc + k_sol exp(-b_sol s^2) Fmask| once per pair instead of at every solver evaluation; same expressions, same types (new test checks a grid point against gemmi's own Scaling fit with ==). - Fcalc density and the solvent mask are made on two threads; the model's structure factors beside the GPU engine reservation. - The indexing probe fits the relabellings concurrently. - The null's replicates run beside the real model's placement (they start from a snapshot of the model as read); one GPU engine per replicate plus one for the real fit instead of a cap of 4 (engines are interchangeable and deterministic). - The 2mFo-DFc, mFo-DFc and anomalous maps are made and written concurrently; the placed model is written beside the reflection files. - A rigid-body zone whose solvent-mask grid needs gemmi's shrink is sent to the CPU when the engines are reserved (ModelMaskGPU::ShrinkIsNoOp), instead of failing on the GPU and validating everything again on the CPU - the same CPU result, without the wasted first attempt. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi |
||
|
|
9ad92b6bfe |
Move the atomic-model code to image_analysis/structure_refinement/ and WriteModel to writer/
A pure move. ModelValidation, RigidBodyRefine, RigidBodyGPU, ModelFFT, ModelGrid, ModelScaling, ModelMaskGPU, ModelScaleGPU and SigmaA - everything that works on an atomic model - become the JFJochStructureRefinement library, linked by JFJochImageAnalysis. WriteModel (the placed-model mmCIF/PDB writer) goes to writer/ as its own small JFJochModelWriter target, so JFJochWriter, which a writer-only build compiles, does not gain a gemmi dependency. Only include paths and CMake lists change. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi |