The rigid-body target re-grids the model at every evaluation (gemmi's
put_model_density_on_grid with the refmac blur, then the Refmac solvent
mask), and that gridding was ~80% of an evaluation on low-symmetry cells.
The central evaluation ran it on one thread and the six Jacobian columns
on one pool worker each, so most of a 32-thread machine sat idle through
the real fit.
New rugnux/ModelGrid.{h,cpp} reimplements the two gemmi routines so that
they are bit-identical to gemmi on any number of threads:
- Atoms are put on the grid per w plane: each plane is one task and walks
every atom whose box reaches it, in model order, visiting only its own
points (a copy of gemmi's do_use_points_in_box restricted to one plane).
Every grid point therefore receives the same additions in the same
order as in gemmi's serial loop. Per-atom coefficients and radius are
computed once up front, and each atom's density function is copied to a
local so the compiler keeps it in registers.
- Symmetrization runs over orbits: gemmi reduces each orbit into its
lowest index, so the leaders are the points with no lower mate. They
depend only on the grid size and group, so RefineRigidBody finds them
once per zone; each orbit is then reduced by one thread with gemmi's
operands in gemmi's order. Orbits share no points.
- The solvent mask uses the same two passes (setting points to 0 is order
independent anyway); gemmi's own island removal and shrink follow.
vendored gemmi is untouched.
The six Jacobian columns now run on std::async threads, not on the pool,
so each column's gridding can spread over the pool (a pass reached from a
pool worker runs inline). With one thread they run deferred, serially.
Evidence: the grids are memcmp-identical to gemmi's at nt 1/5/32 on 41
deposited models from the battery's PDB cache plus the 8 battery sets
(23 of them with anisotropic atoms; P1 up to F4132, R3/R32, I and C
centring) at 6/4.5/3.5 A; new test ModelValidation_ParallelGriddingMatchesGemmi
covers P1, C2, P212121, I23 and F4132 with iso and aniso atoms. A
standalone RefineRigidBody benchmark (synthetic |F| from the model,
displaced model) gives identical evaluations/angle/shift/coordinates to
the unchanged code: 5lzl 24-33 s -> 12 s real fit on a loaded machine;
null (9 concurrent replicates) unchanged within noise.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C