The GPU rigid body is chosen by whether a card is there, nothing else. A run
that must stay on the CPU says so at the process level (CUDA_VISIBLE_DEVICES=),
which RigidBodyGPUPool::Create already honours through get_gpu_count().
Parity tests ([ModelValidation][gpu], compiled into jfjoch_test on a CUDA
build):
- every GPU case now SKIPs without a card, instead of passing silently;
- ModelMaskGPU_MatchesGemmi also checks that the CPU path's PutMaskOnGrid is
gemmi's mask bit for bit on the same grid, so the GPU mask is held against
the CPU path itself;
- RigidBodyGPU_MatchesCPU: Jacobian columns 2e-3 -> 5e-4 relative (measured
up to 5e-5; the floor is the 1e-4 the scale's near-tie can move every
column by). Residuals stay at 2e-4 of <Fobs> (measured up to 3.9e-5);
- RigidBodyGPU_FitAgreesWithCPU: fit endpoint 2e-3 -> 1e-5 A rmsd (measured
up to 5e-7 A).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
- |Fcalc + solvent| is taken once per fit instead of at every step: the solvent pair is fixed for the
whole of a fit.
- The per-step anisotropic factor and the derivatives are computed per point in float, and summed in
double as before. Double arithmetic was most of the cost on a card with little double throughput. The
final R of the grid stays in gemmi's double arithmetic.
- Each mode reduces only the slots it fills; 48 blocks per fit instead of 128.
- The first upload of a batch no longer adds a stream synchronisation.
FitSolvent per zone at 3.5 A (RTX 5080, real zone hkl sets, synthetic amplitudes): 3.2 ms at 3.3k
points, 9.2 ms at 72k and 21.6 ms at 213k, against 9.7 / 29.6 / 62.9 ms before. Fit() is 0.25-0.66 ms.
Against gemmi, k_overall and b* agree to 1e-9 - 1.4e-5 relative, and the grid winner is the same.
New test ModelScaleGPU_MatchesGemmiOnAModelsPoints: the rigid body's own points (the ClusterPdb
fixture with anisotropic atoms, its own amplitudes, a displaced placement, the rigid body's Fcalc and
mask), five groups at 6 and 3.5 A. The R tolerance of the synthetic cases is now 1e-5. That is the
resolution of gemmi's own fit: its Levenberg-Marquardt stops at a WSSR change of 1e-5, and perturbing
gemmi's own Fcalc by 1e-5 moves its answer by up to 1.2e-4 of <Fobs>.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
gemmi's Scaling<float> fit as RigidBodyTarget uses it - fit_isotropic_b_approximately() and the
Levenberg-Marquardt of fit_parameters() with k_sol and b_sol fixed - and FitModelScale's k_sol/b_sol
grid, with the sums over the reflections on the device (double, fixed launch shape, shuffle tree per
warp, warps and then blocks summed in order: no float atomics, bit-identical repeats). The LevMar
control is gemmi's, ported line for line to the host and unrolled into its requests, so the 88 coarse
and up to 25 fine grid fits share one launch per step.
Against gemmi on real zone hkl sets with synthetic amplitudes: k_overall and b* within 1e-9..1e-5
relative, the same grid winner every time, R within 1e-7.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C