Commit Graph
8 Commits
Author SHA1 Message Date
leonarski_fandClaude Opus 5.5 672e182d6a GPU engines wait for their stream before their buffers go; a lost context fails where it is seen
A pooled CudaDevicePtr frees on the thread's allocation stream, not on the engine's stream, and the
pool may hand the memory to another engine - or, past its release threshold, unmap it - as soon as
that free is reached, which on an idle allocation stream is at once. An engine destroyed with work
still queued (FFTIndexerGPU after SearchCap's last DirectionsChanged upload, a spot finder between
DetectAt and Extract, a shadow accumulator after a pending fold, any engine on an exception path)
thus had kernels or copies writing memory that was someone else's or no longer mapped. Now:
- CudaStream synchronises before cudaStreamDestroy (destructor and move-assignment), which covers
  engines whose own stream is declared after their buffers (FFTIndexerGPU, the gather buffer);
- every engine holding pooled buffers and a stream (shared or own, declared before the buffers)
  synchronises it in its destructor; BraggIntegrationEngineGPU also before EnsureCapacity
  reallocates, where a Run that threw leaves work queued.

A GPU failure that is handled no longer hides a lost context: ShadowFinder, BeamCenterFFT, the
rigid-body pool and model validation call cuda_throw_if_context_lost() before cuda_clear_error(),
as the device-decode fallbacks already did; RotationScaleMergeGPU's Alloc does so before waiting up
to ten minutes for GPU work beside it and then reporting a lost device as out of memory; and a
failed cudaMalloc says why. BeamCenterFFT logs the failure it used to drop silently, the
speculative geometry probe logs the exception it swallowed (its GPU fault was otherwise reported by
the merge beside it, under the merge's name), and RotationScaleMergeGPU's DeviceGuard no longer
throws from its destructor.

Only synchronisation and error paths change: p.hkl md5 and the MTZ data (gemmi) are identical to
the b530c2d full battery on myob_x10sa, cytc_x10sa, 8a1a, 9gdj, 11if, kdp_x10sa_20keV and 6z9g.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-10 10:59:24 +02:00
leonarski_fandClaude Opus 5.5 278c488c37 Merge branch 'modelpar' into gpusf
Both sides kept: modelpar's parallel basis scoring, the second validation
started on a forecast beside the first (write gate, schedule parameter),
the multi-GPU placement and delete-before-rewrite; gpusf's GPU structure
factors, maps and null engine, failure-instead-of-restart, and the merge
engine released before the validation (now just before modelpar's
ValidateAgainstModel call, after the forecast lambda is set up).

Placement: each validation's structure-factor engines (d_min and the
null's) are made on the card of the thread that runs it - the main
thread's for the first validation, card 1 % count for the speculative
second, which pins itself there - so two validations on two cards use
both, as the rigid-body pools do.

One memory rule for both, per card, from total memory, up front: a
validation plans at most half of its card - its structure-factor engines
a quarter together (was half for the d_min engine alone), its rigid-body
engines a quarter (RigidBodyGPUPool) - and the second validation runs
beside the first only where twice the first's plan (rigid-body planned
bytes + the d_min engine, twice it where a null is coming, the null's
engine being no larger) fits half of all cards' memory together (was:
twice the rigid-body plan within a quarter). The card's total is read
once when the engine is made; a CUDA error there fails the validation
like any other.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 13:51:41 +02:00
leonarski_fandClaude Opus 5.5 d5fcf2f05b rugnux: model validation's structure factors and maps on the GPU
ModelStructureFactorsGPU computes what compute_model_factors() and
map_from_coefficients() compute on the CPU - F_calc from the model's
density (IT92, Refmac-compatible blur, unblurred as prepare_asu_data()
does) and F_mask from the Refmac bulk-solvent mask, both on the
reflections prepare_asu_data(d_min) lists, in its order; and a map from
ASU coefficients on the grid get_size_for_hkl(coef, 0, 3.0) sizes - on a
device. Made once per cell, group, resolution and model, then evaluated
as often as the coordinates change, so refinement or MR can call it in a
loop. The device is an explicit parameter; every call leaves the calling
thread's current device as it found it.

Pieces:
- ModelDensityGPU: the rigid body's deterministic brick gather, moved
  out of RigidBodyGPU.cu into a component of its own (ModelMaskGPU's
  pattern); the rigid body uses it unchanged. MAX_BRICKS_PER_AXIS 8 ->
  16, so fine grids with high-B atoms (lysozyme at 1.2 A, a 0.9 A P1
  cell) are no longer refused; existing zones are gridded identically.
- One copy of the content is gridded and the symmetry composed in
  reciprocal space (SymmetryComposition), operators applied on the fly;
  the mask is ModelMaskGPU (every image of every atom, islands, shrink).
- Maps: gemmi's get_f_phi_on_grid() in ZYX order on the host (the
  coefficients written are the same), in-place cuFFT c2r, transposed back
  to XYZ on the device. One map at a time, in the engine's buffers.

Decided once, up front, per card, from its TOTAL memory: the engine's
bytes (16 N + cuFFT work + reflections, N the larger of the structure-
factor and map grids) must be at most half the card - the rigid body's
engines take at most a quarter beside it. Otherwise, or where the gather
cannot grid the cell, the CPU path runs, logged with needed vs total.
Anything to a resolution other than d_min (the null's 3.5 A fits) stays
on the CPU, so all replicates and the real model's side of the null are
computed the same way. A CUDA failure takes the existing path: the
validation restarts on the CPU.

Measured, model validation total per run (CPU path -> GPU), 16 GB card:
  F432 215 A cubic, 1.30 A, 500^3 grid: 47.7 -> 15.6 s (two validations;
     14.3 -> 2.4 and 33.4 -> 13.2, the rest of the second is writing the
     three 0.5 GB maps); F_calc + F_mask 5.7 s -> 0.05 s
  C2 1.11 A: 23.9 -> 13.6 s; P3_2 1.55 A: 18.2 -> 10.2 s;
  P2_1 1.25 A: 13.4 -> 6.5 s; F4_132 328 A: 13.0 -> 5.5 s;
  P6_5: 8.8 -> 4.2 s; P4_3 0.97 A: 4.6 -> 2.5 s; P1 0.92 A: 4.2 -> 2.2 s;
  small P1: 3.2 -> 1.3 s; P6_1: 8.8 -> 6.2 s; lysozyme: 1.8 -> 1.4 s.
p.mtz md5-identical on all 13 sets. Against the CPU path: FC within
1e-4 of mean |F|, phases of the strong half within 0.003 deg, maps within
1e-4 (2mFo-DFc) and 7e-4 (mFo-DFc) of their rms; every logged R, CC,
FOM, k_sol and anomalous site list identical at the printed precision,
except where a rigid-body commit sat on an exact R-free tie (0.2155 ->
0.2155) and fell the other way (R-work 0.2127 vs 0.2129). The GPU result
is bit-identical run to run and with -N 8 (maps, map MTZ, placed model).
Peak device memory of the engine: 2.5 GB at 500^3 (process total peaked
at 14.4 GB with what the merge still holds).

Tests: ModelStructureFactorsGPU_MatchesCPU (five groups, 3.5 and 1.5 A:
same reflections, F_calc <= 1e-5 of mean |F|, F_mask 2e-7 rms, repeat
bit-identical), ModelStructureFactorsGPU_MapMatchesCPU (<= 5e-6 of rms).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 12:38:37 +02:00
leonarski_fandClaude Opus 5.5 6025d51a48 ModelMaskGPU: gemmi's shrink step, so fine rigid-body zones stay on the GPU
The GPU bulk-solvent mask had no shrink step (SolventMasker::shrink(),
set_margin_around() with Refmac's r_shrink = 0.8 A), so every rigid-body
zone whose grid has an offset within 0.8 A was sent to the CPU. That is
not only fine grids: in an oblique setting (P2_1, beta ~141 deg) the
3.5 A zone's lattice-plane spacing is 0.69 A, and on such a set the real
fit and all nine null replicates ran their rigid bodies on the CPU.

The shrink is now two passes over the grid: mark the solvent points with
a macromolecule point among gemmi's near offsets, then turn into solvent
every macromolecule point with such an edge point at any stencil offset
(near or far, gemmi's split at the coarsest grid step). That is gemmi's
rule in both of its branches, and the stencil is built on the host in
double exactly as gemmi builds it. ModelMaskGPU::ShrinkIsNoOp() and
RigidBodyGPUEngine::MaskSupports() are gone; the rigid body no longer
refuses such zones.

Verified (ModelMaskGPU_ShrinkMatchesGemmi): the shrink alone on gemmi's
post-island mask is bit-identical, and the whole GPU mask equals gemmi's
put_mask_on_grid() bit for bit (0 differing points) on the five test
groups at 1.5 A and an oblique P2_1 cell at 3.5 A; repeats are identical.

Measured on the oblique-setting set of the open arm (P2_1, 14k atoms,
1.66 A): model validation 17.2 s -> 5.2 s with the GPU structure factors
of the next commit in place (the first validation's ten rigid bodies,
~12 s on the CPU, now take well under a second). p.mtz md5-identical;
the null moves from +73.1 to +86.1 sigma (the replicates are refined on
the GPU, which agrees with the CPU to rounding), verdict unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 12:38:37 +02:00
leonarski_fandClaude Opus 5.5 8d6e9d3a44 ModelValidation: use every visible GPU - the null's engines and the second validation
The rigid-body pool put all of its engines on the calling thread's current
card, so a validation used one GPU whatever the machine had. Now, by fixed
rules decided up front and never by momentary free memory:

- RigidBodyGPUPool::Create puts engine i on card (d + i) % count, d being
  the calling thread's card; the pool restores that card afterwards
  (an engine's constructor sets its own) and an engine is released on
  its own card.
- The threads that run the null's replicates are pinned with pin_gpu
  (which also binds them to the card's NUMA node where that is enabled),
  replicate thread t to card (d + 1 + t) % count, and Acquire() hands a
  thread an idle engine on its own card where there is one, any other
  otherwise. With one card this is exactly the previous back-of-the-list
  choice.
- The validation started on the forecast runs on card 1 % count, so with
  two cards or more it is on the other card from the first; the up-front
  memory rule becomes: twice the planned engine bytes within a quarter of
  the cards' total memory taken together.

The engines' kernels are deterministic (no floating-point atomics) and an
engine's result does not depend on which engine it is, so on cards of one
model the numbers are those of one card; what changes is only where the
work runs. On this one-card workstation the multi-card path cannot be
exercised: md5 of p.mtz and the validation outputs are identical to the
base, and the card-count arithmetic was checked by reading only.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 12:11:12 +02:00
leonarski_fandClaude Opus 5.5 b5980c9ba0 rugnux: start the validation in the model's setting on a forecast, beside the first
Where a model that fits is written on other axes than the data, the files
take its setting and the validation is made a second time on the
relabelled data. That second validation depends on the first only through
the setting and the indexing the first settles, and both are known long
before the first has finished: the change of basis right after the frame
scoring, the indexing the probe prefers before the null. So ValidateAgainstModel
now reports them (ModelFrameForecast, through ModelValidationSchedule::
on_forecast), and the second validation is started there, on the
relabelling AdoptModelFrame and relabel_output would make, applied to a
copy of the merge - beside the first one's null, real fit and maps.

It is kept only where the first decides exactly what was forecast (the
model fits, same change of basis, same indexing; a model asserting the
other enantiomorph is not forecast, as the label is decided last, on the
anomalous map). Until then its log is held (Logger::Buffered, replayed
where the serial run logged it) and its files wait on a gate
(ModelValidationSchedule::write_gate) placed before the first map is
written; otherwise it is released with false and returns unwritten, and
the serial validation runs as before.

GPU memory: two validations at once take twice the rigid-body engines.
The parallel start is decided up front from sizes, never from what is
free: allowed where twice what the first pool asked for (bytes per engine
times the engines wanted) fits a quarter of the card's TOTAL memory, the
share one validation may take. Threads: both validations submit to the
one ParallelFor pool from threads outside it (the second runs on a
std::async thread, as the null's replicates do), so no pass runs inline
on a pool worker and the pool's size bounds the workers.

Measured on the loaded 16-core workstation, TIMING model validation:
8sa8 30.6 -> 19.7 s, 8xtg 19.9 -> 13.9 s, 9ea5 22.1 -> 16.4 s (this and
the previous commit together). p.mtz, p.hkl, p.cif, the three maps,
p_maps.mtz and p_model.cif/pdb md5-identical to the base on 8sa8, 8xtg,
9ea5, 7qis and myob_x10sa; the validation's log lines identical as a set
(the null's replicate lines were already in completion order).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 12:10:46 +02:00
leonarski_fandClaude Opus 5.5 f956eb25e8 rugnux: model validation does the same work in less time
--model validation (battery-only for users) was 18% of the battery's time. Every
number it produces is unchanged to the bit (p.mtz, maps, placed model and every
model-validation line of the report md5/diff-identical on 11 open sets); only
when and where the work runs changes:

- The bulk-solvent grid fit (FitModelScale, most of the CPU time) fits each
  solvent pair on a copy of gemmi::Scaling's target that takes
  |Fcalc + k_sol exp(-b_sol s^2) Fmask| once per pair instead of at every
  solver evaluation; same expressions, same types (new test checks a grid
  point against gemmi's own Scaling fit with ==).
- Fcalc density and the solvent mask are made on two threads; the model's
  structure factors beside the GPU engine reservation.
- The indexing probe fits the relabellings concurrently.
- The null's replicates run beside the real model's placement (they start
  from a snapshot of the model as read); one GPU engine per replicate plus
  one for the real fit instead of a cap of 4 (engines are interchangeable
  and deterministic).
- The 2mFo-DFc, mFo-DFc and anomalous maps are made and written
  concurrently; the placed model is written beside the reflection files.
- A rigid-body zone whose solvent-mask grid needs gemmi's shrink is sent to
  the CPU when the engines are reserved (ModelMaskGPU::ShrinkIsNoOp), instead
  of failing on the GPU and validating everything again on the CPU - the
  same CPU result, without the wasted first attempt.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 00:01:14 +02:00
leonarski_fandClaude Opus 5.5 9ad92b6bfe Move the atomic-model code to image_analysis/structure_refinement/ and WriteModel to writer/
A pure move. ModelValidation, RigidBodyRefine, RigidBodyGPU, ModelFFT, ModelGrid,
ModelScaling, ModelMaskGPU, ModelScaleGPU and SigmaA - everything that works on an
atomic model - become the JFJochStructureRefinement library, linked by
JFJochImageAnalysis. WriteModel (the placed-model mmCIF/PDB writer) goes to writer/
as its own small JFJochModelWriter target, so JFJochWriter, which a writer-only build
compiles, does not gain a gemmi dependency. Only include paths and CMake lists change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-07 14:05:37 +02:00