Commit Graph
14 Commits
Author SHA1 Message Date
leonarski_fandClaude Opus 5.5 672e182d6a GPU engines wait for their stream before their buffers go; a lost context fails where it is seen
A pooled CudaDevicePtr frees on the thread's allocation stream, not on the engine's stream, and the
pool may hand the memory to another engine - or, past its release threshold, unmap it - as soon as
that free is reached, which on an idle allocation stream is at once. An engine destroyed with work
still queued (FFTIndexerGPU after SearchCap's last DirectionsChanged upload, a spot finder between
DetectAt and Extract, a shadow accumulator after a pending fold, any engine on an exception path)
thus had kernels or copies writing memory that was someone else's or no longer mapped. Now:
- CudaStream synchronises before cudaStreamDestroy (destructor and move-assignment), which covers
  engines whose own stream is declared after their buffers (FFTIndexerGPU, the gather buffer);
- every engine holding pooled buffers and a stream (shared or own, declared before the buffers)
  synchronises it in its destructor; BraggIntegrationEngineGPU also before EnsureCapacity
  reallocates, where a Run that threw leaves work queued.

A GPU failure that is handled no longer hides a lost context: ShadowFinder, BeamCenterFFT, the
rigid-body pool and model validation call cuda_throw_if_context_lost() before cuda_clear_error(),
as the device-decode fallbacks already did; RotationScaleMergeGPU's Alloc does so before waiting up
to ten minutes for GPU work beside it and then reporting a lost device as out of memory; and a
failed cudaMalloc says why. BeamCenterFFT logs the failure it used to drop silently, the
speculative geometry probe logs the exception it swallowed (its GPU fault was otherwise reported by
the merge beside it, under the merge's name), and RotationScaleMergeGPU's DeviceGuard no longer
throws from its destructor.

Only synchronisation and error paths change: p.hkl md5 and the MTZ data (gemmi) are identical to
the b530c2d full battery on myob_x10sa, cytc_x10sa, 8a1a, 9gdj, 11if, kdp_x10sa_20keV and 6z9g.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-10 10:59:24 +02:00
leonarski_fandClaude Opus 5.5 8442d55f6f rugnux --model: score the twin hypothesis too, and take the maps from the merge the model prefers
Where a twin hypothesis (<prefix>_twin_<SG>) was written, the model is also
validated against that merge: as merged, and as twinned by its law -
|F_twin(h)|^2 = (1-a)|F_model(h)|^2 + a|F_model(Th)|^2 from the placed, scaled
model, one overall scale refitted per a, a scanned 0..0.5 on the working set
and R-free read at it (ValidateAgainstModel's new twin_law argument; GPU
structure factors as before). The maps and the placed model are then written
from whichever merge has the lower R-free; the report gives MAPS_FROM and the
subgroup's plain and twinned R-work/R-free and fraction. The subgroup file
follows its own validation's indexing, as the answer follows its own.

The model chooses the maps only: the space group, <prefix>.mtz and R_FREE stay
the answer's. Documented in RUGNUX_REPORT.md, RUGNUX_ADVANCED.md and
CPU_DATA_ANALYSIS_DECISIONS.md (14.9); Yeates (1997) credited.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 19:40:22 +02:00
leonarski_fandClaude Opus 5.5 32e78823fb ModelValidation: the non-CUDA build compiles again
The multi-GPU placement (replicate threads pinned with pin_gpu) and the
second-validation memory rule read RigidBodyGPUPool and get_gpu_count(),
which a build without CUDA does not have; both are now inside
JFJOCH_USE_CUDA, where the pool can exist at all. Nothing changes in a
CUDA build. Non-CUDA rugnux and jfjoch_test build; [ModelValidation] 22
cases pass there.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 14:38:20 +02:00
leonarski_fandClaude Opus 5.5 9a1f729993 rugnux: do not make the first validation's maps where the forecast one is kept
Where the validation in the model's setting, started beside the first on
the forecast (b5980c9ba), is going to be kept, it writes the maps, the map
coefficients and the placed model on the written axes, and the first
validation's p_2fofc/p_fofc/p_anom.ccp4 and p_maps.mtz were overwritten
unread. ModelValidationSchedule::maps_superseded is asked once the
verdict, the setting and the indexing are decided, before any map is
made; the first validation then makes none (the map coefficients and the
MTZ are not built, the three map transforms not run, nothing written).
The answer is the same rule that keeps the forecast validation
afterwards (forecast_kept, now one lambda for both). The kept validation
also brings back the reflections it relabelled, free set included, so
relabel_output does not reindex and redraw them a second time. Where the
forecast is discarded or was never started, everything runs as before,
and modelpar's delete-before-rewrite stays for those cases. If the kept
validation then fails, a warning says no maps were written.

Log lines that disappear (first validation only, kept-forecast runs):
"mean 2mFo-DFc density at atom centres = ...", "anomalous difference map
from ... Bijvoet pairs; strongest density ...", "wrote p_2fofc.ccp4, ...";
in their place: "no maps made here - the validation in the model's
setting writes them". (The "anomalous density ... inverted" warning would
also go; it cannot be raised where a forecast is kept.)

Verified on 8sa8, 8xtg, 9ea5 (forecast kept), 8tyy, 7qis, myob_x10sa (no
forecast): p.mtz, p.hkl, p.cif, the three maps, p_maps.mtz and
p_model.cif/pdb md5-identical to the merge commit, repeated and at -N 8;
the discard path (8sa8 with a rotated model) identical files and log.
Time saved is small with the maps already on the GPU: 8sa8 7.6 -> 7.4 s,
8xtg 5.1 -> 5.0 s, 9ea5 within noise.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 14:34:46 +02:00
leonarski_fandClaude Opus 5.5 278c488c37 Merge branch 'modelpar' into gpusf
Both sides kept: modelpar's parallel basis scoring, the second validation
started on a forecast beside the first (write gate, schedule parameter),
the multi-GPU placement and delete-before-rewrite; gpusf's GPU structure
factors, maps and null engine, failure-instead-of-restart, and the merge
engine released before the validation (now just before modelpar's
ValidateAgainstModel call, after the forecast lambda is set up).

Placement: each validation's structure-factor engines (d_min and the
null's) are made on the card of the thread that runs it - the main
thread's for the first validation, card 1 % count for the speculative
second, which pins itself there - so two validations on two cards use
both, as the rigid-body pools do.

One memory rule for both, per card, from total memory, up front: a
validation plans at most half of its card - its structure-factor engines
a quarter together (was half for the d_min engine alone), its rigid-body
engines a quarter (RigidBodyGPUPool) - and the second validation runs
beside the first only where twice the first's plan (rigid-body planned
bytes + the d_min engine, twice it where a null is coming, the null's
engine being no larger) fits half of all cards' memory together (was:
twice the rigid-body plan within a quarter). The card's total is read
once when the engine is made; a CUDA error there fails the validation
like any other.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 13:51:41 +02:00
leonarski_fandClaude Opus 5.5 0f707caab9 ModelValidation: the null's structure factors on the GPU too
The null - nine random placements of the model, and the real model's side
of the comparison - is scored to a coarser limit (3.5 A, or where the free
set allows), and those structure factors still ran on the CPU, one model
at a time on each replicate thread. Where the d_min structure factors are
on the GPU, a second ModelStructureFactorsGPU engine at the null's limit
is now made, on the same card, before any replicate starts; every
replicate and the real side go through it, so all of them are computed
the same way. It is decided like the first engine, from the card's total
memory: the two together must stay within half of it, else the null stays
on the CPU (logged). The replicates' threads may sit on other cards
(modelpar's 8d6e9d3a4); each call works on the engine's own device. Where
the null is scored to d_min it shares the d_min engine, as before.

Measured (validation time, before -> after; CPU path for reference):
  P3_2 1.55 A:  11.2 -> 9.0 s (CPU 17.7)
  P2_1 oblique: 5.1 -> 4.5 s (CPU 7.9)
  C2 1.11 A:   12.9 -> 13.2 s (CPU 23.3) - its null time is in the
               scale fits, which stay on the CPU
p.mtz md5-identical in all four variants per set (before, after, CPU,
-N 8); maps and map MTZ bit-identical after vs -N 8; the null's sigma
moves only in the second decimal (+113.39 -> +113.38, +86.13 -> +86.14).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 13:08:09 +02:00
leonarski_fandClaude Opus 5.5 ed2d557bb7 ModelValidation: a CUDA failure fails the validation instead of restarting it on the CPU
A CUDA error on the validation's GPU path (rigid body, structure factors,
maps) used to start the whole validation again on the CPU - a migration
part way through, which makes how long a run takes, and in principle what
its rigid bodies converge to, depend on a device fault. It now ends the
validation with ok = false and a logged reason ("the GPU failed during the
validation (...)"), which the run reports as MODEL_NOT_VALIDATED. A
validation that did not finish decides nothing, so the reflection files
are those of a run without a model. Where there is no GPU at all, or the
up-front rule sends the work to the CPU, nothing changes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 12:47:04 +02:00
leonarski_fandClaude Opus 5.5 d5fcf2f05b rugnux: model validation's structure factors and maps on the GPU
ModelStructureFactorsGPU computes what compute_model_factors() and
map_from_coefficients() compute on the CPU - F_calc from the model's
density (IT92, Refmac-compatible blur, unblurred as prepare_asu_data()
does) and F_mask from the Refmac bulk-solvent mask, both on the
reflections prepare_asu_data(d_min) lists, in its order; and a map from
ASU coefficients on the grid get_size_for_hkl(coef, 0, 3.0) sizes - on a
device. Made once per cell, group, resolution and model, then evaluated
as often as the coordinates change, so refinement or MR can call it in a
loop. The device is an explicit parameter; every call leaves the calling
thread's current device as it found it.

Pieces:
- ModelDensityGPU: the rigid body's deterministic brick gather, moved
  out of RigidBodyGPU.cu into a component of its own (ModelMaskGPU's
  pattern); the rigid body uses it unchanged. MAX_BRICKS_PER_AXIS 8 ->
  16, so fine grids with high-B atoms (lysozyme at 1.2 A, a 0.9 A P1
  cell) are no longer refused; existing zones are gridded identically.
- One copy of the content is gridded and the symmetry composed in
  reciprocal space (SymmetryComposition), operators applied on the fly;
  the mask is ModelMaskGPU (every image of every atom, islands, shrink).
- Maps: gemmi's get_f_phi_on_grid() in ZYX order on the host (the
  coefficients written are the same), in-place cuFFT c2r, transposed back
  to XYZ on the device. One map at a time, in the engine's buffers.

Decided once, up front, per card, from its TOTAL memory: the engine's
bytes (16 N + cuFFT work + reflections, N the larger of the structure-
factor and map grids) must be at most half the card - the rigid body's
engines take at most a quarter beside it. Otherwise, or where the gather
cannot grid the cell, the CPU path runs, logged with needed vs total.
Anything to a resolution other than d_min (the null's 3.5 A fits) stays
on the CPU, so all replicates and the real model's side of the null are
computed the same way. A CUDA failure takes the existing path: the
validation restarts on the CPU.

Measured, model validation total per run (CPU path -> GPU), 16 GB card:
  F432 215 A cubic, 1.30 A, 500^3 grid: 47.7 -> 15.6 s (two validations;
     14.3 -> 2.4 and 33.4 -> 13.2, the rest of the second is writing the
     three 0.5 GB maps); F_calc + F_mask 5.7 s -> 0.05 s
  C2 1.11 A: 23.9 -> 13.6 s; P3_2 1.55 A: 18.2 -> 10.2 s;
  P2_1 1.25 A: 13.4 -> 6.5 s; F4_132 328 A: 13.0 -> 5.5 s;
  P6_5: 8.8 -> 4.2 s; P4_3 0.97 A: 4.6 -> 2.5 s; P1 0.92 A: 4.2 -> 2.2 s;
  small P1: 3.2 -> 1.3 s; P6_1: 8.8 -> 6.2 s; lysozyme: 1.8 -> 1.4 s.
p.mtz md5-identical on all 13 sets. Against the CPU path: FC within
1e-4 of mean |F|, phases of the strong half within 0.003 deg, maps within
1e-4 (2mFo-DFc) and 7e-4 (mFo-DFc) of their rms; every logged R, CC,
FOM, k_sol and anomalous site list identical at the printed precision,
except where a rigid-body commit sat on an exact R-free tie (0.2155 ->
0.2155) and fell the other way (R-work 0.2127 vs 0.2129). The GPU result
is bit-identical run to run and with -N 8 (maps, map MTZ, placed model).
Peak device memory of the engine: 2.5 GB at 500^3 (process total peaked
at 14.4 GB with what the merge still holds).

Tests: ModelStructureFactorsGPU_MatchesCPU (five groups, 3.5 and 1.5 A:
same reflections, F_calc <= 1e-5 of mean |F|, F_mask 2e-7 rms, repeat
bit-identical), ModelStructureFactorsGPU_MapMatchesCPU (<= 5e-6 of rms).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 12:38:37 +02:00
leonarski_fandClaude Opus 5.5 89f52fef95 ModelValidation: remove a map file before writing it again, rather than truncating it
The validation in the model's setting writes its three maps and the map
MTZ over the ones the first validation wrote seconds before, and a
two-pass run writes over the first pass's. XFS (like ext4) flushes a file
that was truncated and rewritten when it is closed, so each overwrite
forced ~155 MB out to the disk on the spot. Measured on /data (XFS on a
hard disk): three 155 MB files rewritten 3.16 s, written fresh 0.62 s,
removed and written again 0.66 s. The contents are the same bytes either
way.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 12:32:10 +02:00
leonarski_fandClaude Opus 5.5 8d6e9d3a44 ModelValidation: use every visible GPU - the null's engines and the second validation
The rigid-body pool put all of its engines on the calling thread's current
card, so a validation used one GPU whatever the machine had. Now, by fixed
rules decided up front and never by momentary free memory:

- RigidBodyGPUPool::Create puts engine i on card (d + i) % count, d being
  the calling thread's card; the pool restores that card afterwards
  (an engine's constructor sets its own) and an engine is released on
  its own card.
- The threads that run the null's replicates are pinned with pin_gpu
  (which also binds them to the card's NUMA node where that is enabled),
  replicate thread t to card (d + 1 + t) % count, and Acquire() hands a
  thread an idle engine on its own card where there is one, any other
  otherwise. With one card this is exactly the previous back-of-the-list
  choice.
- The validation started on the forecast runs on card 1 % count, so with
  two cards or more it is on the other card from the first; the up-front
  memory rule becomes: twice the planned engine bytes within a quarter of
  the cards' total memory taken together.

The engines' kernels are deterministic (no floating-point atomics) and an
engine's result does not depend on which engine it is, so on cards of one
model the numbers are those of one card; what changes is only where the
work runs. On this one-card workstation the multi-card path cannot be
exercised: md5 of p.mtz and the validation outputs are identical to the
base, and the card-count arithmetic was checked by reading only.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 12:11:12 +02:00
leonarski_fandClaude Opus 5.5 b5980c9ba0 rugnux: start the validation in the model's setting on a forecast, beside the first
Where a model that fits is written on other axes than the data, the files
take its setting and the validation is made a second time on the
relabelled data. That second validation depends on the first only through
the setting and the indexing the first settles, and both are known long
before the first has finished: the change of basis right after the frame
scoring, the indexing the probe prefers before the null. So ValidateAgainstModel
now reports them (ModelFrameForecast, through ModelValidationSchedule::
on_forecast), and the second validation is started there, on the
relabelling AdoptModelFrame and relabel_output would make, applied to a
copy of the merge - beside the first one's null, real fit and maps.

It is kept only where the first decides exactly what was forecast (the
model fits, same change of basis, same indexing; a model asserting the
other enantiomorph is not forecast, as the label is decided last, on the
anomalous map). Until then its log is held (Logger::Buffered, replayed
where the serial run logged it) and its files wait on a gate
(ModelValidationSchedule::write_gate) placed before the first map is
written; otherwise it is released with false and returns unwritten, and
the serial validation runs as before.

GPU memory: two validations at once take twice the rigid-body engines.
The parallel start is decided up front from sizes, never from what is
free: allowed where twice what the first pool asked for (bytes per engine
times the engines wanted) fits a quarter of the card's TOTAL memory, the
share one validation may take. Threads: both validations submit to the
one ParallelFor pool from threads outside it (the second runs on a
std::async thread, as the null's replicates do), so no pass runs inline
on a pool worker and the pool's size bounds the workers.

Measured on the loaded 16-core workstation, TIMING model validation:
8sa8 30.6 -> 19.7 s, 8xtg 19.9 -> 13.9 s, 9ea5 22.1 -> 16.4 s (this and
the previous commit together). p.mtz, p.hkl, p.cif, the three maps,
p_maps.mtz and p_model.cif/pdb md5-identical to the base on 8sa8, 8xtg,
9ea5, 7qis and myob_x10sa; the validation's log lines identical as a set
(the null's replicate lines were already in completion order).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 12:10:46 +02:00
leonarski_fandClaude Opus 5.5 aabac1541a ModelValidation: score the candidate changes of basis at the same time
Each candidate setting of the model was scored one after another on the
coarse shell (a density, an FFT and an isotropic scale fit, 0.2-0.5 s
each), and a run with a model in another setting scores them in both of
its validations. They share nothing but what they read, so they now run
on ParallelFor, each on its own copy of the model, and are read back in
their own order: the log lines, the ranking and the tie-break (first of
equal R wins) are those of the serial loop.

Measured (16-core workstation, loaded): 8sa8 3 candidates 0.64 s ->
0.22 s; 9ea5 3 candidates 1.47 s -> 0.50 s; 8xtg 1.45 s -> 0.40 s.
p.mtz, the maps, the map MTZ and the placed model md5-identical.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 12:08:39 +02:00
leonarski_fandClaude Opus 5.5 f956eb25e8 rugnux: model validation does the same work in less time
--model validation (battery-only for users) was 18% of the battery's time. Every
number it produces is unchanged to the bit (p.mtz, maps, placed model and every
model-validation line of the report md5/diff-identical on 11 open sets); only
when and where the work runs changes:

- The bulk-solvent grid fit (FitModelScale, most of the CPU time) fits each
  solvent pair on a copy of gemmi::Scaling's target that takes
  |Fcalc + k_sol exp(-b_sol s^2) Fmask| once per pair instead of at every
  solver evaluation; same expressions, same types (new test checks a grid
  point against gemmi's own Scaling fit with ==).
- Fcalc density and the solvent mask are made on two threads; the model's
  structure factors beside the GPU engine reservation.
- The indexing probe fits the relabellings concurrently.
- The null's replicates run beside the real model's placement (they start
  from a snapshot of the model as read); one GPU engine per replicate plus
  one for the real fit instead of a cap of 4 (engines are interchangeable
  and deterministic).
- The 2mFo-DFc, mFo-DFc and anomalous maps are made and written
  concurrently; the placed model is written beside the reflection files.
- A rigid-body zone whose solvent-mask grid needs gemmi's shrink is sent to
  the CPU when the engines are reserved (ModelMaskGPU::ShrinkIsNoOp), instead
  of failing on the GPU and validating everything again on the CPU - the
  same CPU result, without the wasted first attempt.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 00:01:14 +02:00
leonarski_fandClaude Opus 5.5 9ad92b6bfe Move the atomic-model code to image_analysis/structure_refinement/ and WriteModel to writer/
A pure move. ModelValidation, RigidBodyRefine, RigidBodyGPU, ModelFFT, ModelGrid,
ModelScaling, ModelMaskGPU, ModelScaleGPU and SigmaA - everything that works on an
atomic model - become the JFJochStructureRefinement library, linked by
JFJochImageAnalysis. WriteModel (the placed-model mmCIF/PDB writer) goes to writer/
as its own small JFJochModelWriter target, so JFJochWriter, which a writer-only build
compiles, does not gain a gemmi dependency. Only include paths and CMake lists change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-07 14:05:37 +02:00