All exact to the bit (bench against the previous FitModelScale on P1, P2_1,
P2_12_12_1, P4_32_12, P6_122, P2_13, I23 and the dependent n<=5 path; the
ModelScaling cases; validation outputs identical on real data):
- the start of a grid point (fit_isotropic_b_approximately) is computed from the
same |Fcalc| the fit takes, once instead of twice, and without copying the
Scaling's points per chunk;
- gemmi's Levenberg-Marquardt is followed in a copy (LevMarFit) whose
compute_lm_matrices is templated on the parameter count (2..7), so alpha/beta
stay in registers; a zero derivative row adds +0 instead of being skipped
(alpha/beta are never -0), and the full square is summed (the lower half is
gemmi's, the upper is its mirror) - branch-free and vectorised;
- grid points are scheduled one per task (ParallelFor) instead of fixed chunks;
- a fine-pass pair bit-identical to a coarse-pass pair reuses that fit.
Bench (single thread, 48k reflections, P2_1): 1.58 -> 1.33 s; ~15-25% on the
other groups.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi