Commit Graph
3 Commits
Author SHA1 Message Date
leonarski_fandClaude Opus 5.5 23812b902c Rigid body: no JFJOCH_RIGID_BODY_CPU; GPU/CPU parity tests tightened
The GPU rigid body is chosen by whether a card is there, nothing else. A run
that must stay on the CPU says so at the process level (CUDA_VISIBLE_DEVICES=),
which RigidBodyGPUPool::Create already honours through get_gpu_count().

Parity tests ([ModelValidation][gpu], compiled into jfjoch_test on a CUDA
build):
- every GPU case now SKIPs without a card, instead of passing silently;
- ModelMaskGPU_MatchesGemmi also checks that the CPU path's PutMaskOnGrid is
  gemmi's mask bit for bit on the same grid, so the GPU mask is held against
  the CPU path itself;
- RigidBodyGPU_MatchesCPU: Jacobian columns 2e-3 -> 5e-4 relative (measured
  up to 5e-5; the floor is the 1e-4 the scale's near-tie can move every
  column by). Residuals stay at 2e-4 of <Fobs> (measured up to 3.9e-5);
- RigidBodyGPU_FitAgreesWithCPU: fit endpoint 2e-3 -> 1e-5 A rmsd (measured
  up to 5e-7 A).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-28 22:46:00 +02:00
leonarski_fandClaude Opus 5.5 b94d7b9bbd ModelMaskGPUTest: upload on the mask's stream
cudaMemcpy from pageable memory can return before the DMA lands, and the legacy stream it runs on
does not order a non-blocking stream, so RemoveIslands could read a partly uploaded mask. Both
uploads are now cudaMemcpyAsync on the test's stream.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-28 19:04:02 +02:00
leonarski_fandClaude Opus 5.5 63fde9246a ModelMaskGPU: the Refmac bulk-solvent mask on the GPU
The equivalent of PutMaskOnGrid() / gemmi SolventMasker(Refmac).put_mask_on_grid() for the GPU
rigid body. Every symmetry image of every atom is masked directly (one block per image, gemmi's
box and !(d2 > r2) rule), which equals gemmi's mask-then-symmetrize-with-min because the operators
are isometries. Islands are removed by a 26-connected periodic union-find (atomicCAS, larger root
under smaller, so each component is rooted at its lowest index), a read-only root pass, a size
count capped at the island limit and a removal pass - no float atomics, bit-identical repeats.
The shrink is not implemented: SetGrid() throws when gemmi's 0.8 A stencil would be non-empty,
which it is not on any rigid-body grid (spacing d/3 >= 1.17 A).

Atoms come in as double fractional coordinates (ModelMaskAtom): with float input, 9fhc at 3.5 A
differed from gemmi at one 24-fold orbit of radius points (|d - r| ~ 2e-6 A); with double input
the mask equals gemmi's on all 8 prototype sets at 3.5 and 6 A, including 6oel, 8t7r, 9hnc, 9fhc.

Time per mask (RTX 5080, Compute incl. islands): 6oel 288^3 4.2 ms, 8t7r 18.4M 2.8 ms,
9fhc 8.0M 1.5 ms, 9hnc 1.9M 0.45 ms.

Union-find credited to Playne & Hawick (2018) at the kernel and in ACKNOWLEDGEMENT.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-28 17:13:09 +02:00