Commit Graph
1197 Commits
Author SHA1 Message Date
leonarski_f 508fcfb400 Merge branch 'c-screw-undetermined' into pool-C 2026-09-20 18:45:17 +02:00
leonarski_f d3990ca2e0 Merge branch 'c-zone-fix' into pool-C 2026-09-20 18:45:17 +02:00
leonarski_fandClaude Opus 5 7a6df3c24a Space group: say which axis a screw could not be decided on
A screw axis whose row the sweep never recorded - it lies in the spindle's
blind cone, or outside the resolution range - is not a group the data refused,
it is a question nobody asked. The search already offered the whole set in
SPACE_GROUP_ALTERNATIVES, but the per-zone screw table that would say it in
words is printed for the SELECTED candidate only, and the selected candidate
in exactly this case is the one with no screw zones, so the run's account of
the open axis was a blank.

The search now names the axes on which two SELECTED candidates disagree about
whether the row carries screw absences at all, with why the row could not be
judged (never recorded / no control class). The report writes the axes as
SPACE_GROUP_SCREW_UNDETERMINED= beside SPACE_GROUP_ALTERNATIVES and explains
them in prose; the adoption logs a warning naming the axis and the set. An
enantiomorphic or origin-ambiguous pair predicts the same absences on every
row and is not named here - that ambiguity is the hand, or the origin.

Nothing about the decision moves: the group adopted, the alternatives and the
written .mtz/.cif/.hkl are exactly as before, because a reflection file cannot
hold "maybe a screw".

The battery scorer mirrors its existing "hand only" rule: a set differing from
its reference only by a screw the run reports as undeterminable, with the
reference among the groups it offered, scores unscored/screw_undetermined
instead of a sym_screw failure. All three conditions are necessary, so a screw
called wrongly where the row WAS measured stays a failure.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:17 +02:00
leonarski_f ee442eb5a2 Merge branch 'c-speed' into pool-C 2026-09-20 18:45:17 +02:00
leonarski_fandClaude Fable 5.1 40874c2c2d Space-group search: the twin-immune zone is read on a normalisation the control certifies
The zone verdict that arbitrates a refused promotion read its zone absolutely, against the
centric and acentric Wilson expectations, and an absolute reading is only as good as the
normalisation under it. On two refused 622 promotions - a 6/m crystal and a 312 one, whose added
operators disagreed at 4.7x and 5.2x the parent's H and merged to an R_meas of 0.34 under the
higher group - every class read centric: a 67 A^2 anisotropy, isotropically normalised in bins
of 100, spread one shell's expected intensity over a factor of 15 between its directions, and
the acentric control read 1.00 against its 0.74. The zone read the same as the control, +0.10
nats per reflection, and +124 nats over 947 reflections rescued a twin law.

Two things change. The anisotropy is fitted on the acentric reflections of the shells read
(ln I = c + s^T Q s, by least squares) and its deviatoric part taken out of every intensity
before anything is normalised; that alone brings the controls of the crystals measured to
0.73-0.77. What no normalisation removes, the control then certifies: an acentric population
reads -0.130 nats per reflection when the normalisation is right, a twinned one reads below
that, so whatever the control reads above it is the normalisation's - anisotropy, a
pseudo-translation, a pseudo-centring, noise all inflate every class towards centric alike -
and the zone, normalised the same way, carries the same per reflection; the calibrated evidence
has it taken off, and that is what the verdict reads. The two rescued twin laws now read -130
and -78 nats and stay refused, with the law named; the genuine promotions measured read +137 to
+580 (a pseudo-centred orthorhombic crystal whose control reads 0.99 still +334); the trigonal
and hexagonal partial twins -129 to -420 as before. The report prints the control's excess and
the calibrated evidence beside the raw one.

The metric-lattice re-ask carried a second copy of the two-arm rule without the zone: a P31
partial twin whose 32 the main rule had refused on the zone was promoted by that ask's
Lorentz-filtered arm. It now puts the same verdict to a refusal the other arm would outvote.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:17 +02:00
leonarski_fandClaude Opus 5 bac2d01c44 Space group: a screw zone's evidence no longer hangs on its largest absence
A screw's predicted-absent class is one axial row - half a dozen to a few dozen
reflections - and its evidence is a SUM over them, so it is decided by its
largest member. The file's own LIMIT comment said so; 7n2s is that limit firing
on real data. Between two scaling passes that differed only in which weak frames
were rejected, one of eight dead 0k0 moved from 14 +- 9 to 99 +- 10 while the
other seven did not move at all, and the zone fell from 30.1 nats to 17.1 and
lost the 2(1) under a bound of 20. That reflection was never measured to the
precision its sigma claimed: its two half-set merges read 198 and 2.5.

The zone's sum is now taken with its single largest member dropped and rescaled
for the trim - divided by n - H_n, the expected sum of the other n-1 under the
null, and multiplied back by n. ScrewZoneEvidence reads the result exactly as
before: same statistic, same floor, same bound, same calibration, with a robust
estimate of the zone's deadness in place of a fragile one. One member only,
whatever the zone holds: a zone with two strong absences is a zone that is not
extinct. On a uniformly dead zone the rescale under-states by 2.2 nats at eight
absences and 3.9 at sixty-four - it only ever refuses, never claims. Glide zones
keep the untrimmed sum: a plane holds hundreds to thousands of reflections and
no single one can carry the verdict.

7n2s -> P 1 21 1 (zone 27.6 nats, set by the seven reflections that did not
move), matching its deposit; a second monoclinic crystal decided six nats under
the bound (7 absent, 1 violation, 13.9 nats) reaches 21.9 and its 2(1) as well.
Unchanged on 7mzt, 7k1l, 11if, 9hs7, 9zlo and four in-house reference sets.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:17 +02:00
leonarski_fandClaude Opus 5 1149c541fc rugnux: the distance arm is decided on probes, not on written passes
Where the pre-pass cannot tell whether its distance move paid, the two hypotheses were judged by
running each one as a full canonical pass - scaling, space-group search, correction surfaces,
every merge, the model step and the output files - although the comparison reads exactly one
number from each: the held-out residual the post-refinement measures right after integration,
before the scaling engine is even built. Where the header arm won, a third full pass then
reproduced the second's numbers digit for digit to write the files.

Both arms are now measured before the canonical pass by passes that stop as soon as that
post-refinement has measured, and only the arm that wins is run as a canonical pass. So a run
that takes the arm pays two image loops plus one canonical pass instead of two or three
canonical passes.

Result-neutral by construction - the residuals compared, the comparison, the geometry adopted
and the pass that writes are unchanged. Verified byte-identical p.mtz, p.hkl, p.cif and p_P1.mtz,
and identical reports bar the command line and the clock, on eight rotation datasets covering
both arms (seven where the header wins, one where the post-refined distance does).

Measured on the datasets whose logs show the arm: it takes an open-arm corpus pass from 261 to
218 minutes, the in-house arm from 25 to 22 and the private arm from 24 to 19, with the worst
single set going 594 -> 316 s.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:17 +02:00
leonarski_f 3ca22c07c1 Merge branch 'c-mem' into pool-C 2026-09-20 18:45:17 +02:00
leonarski_f 4721a91e2f Merge branch 'c-core' into pool-C 2026-09-20 18:45:17 +02:00
leonarski_fandClaude Opus 5 eb6589f647 rugnux: sixteen-bit Miller indices in the ingest sort key and the post-refine partial
Both arrays are one record per integrated observation - tens of millions on a fine-sliced long axis,
gigabytes each - and both carry the raw hkl only to sort and group on. A Miller index needs sixteen
bits (|h| <= a / d_min, in the hundreds even on the longest axis at atomic resolution), which takes
the ingest sort key from 24 to 20 bytes and the post-refine partial from 32 to 28.

Same comparisons, same order; merged output byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:17 +02:00
leonarski_f af1b5dd9aa Merge branch 'a23-dtz' into pool-C 2026-09-20 18:45:17 +02:00
leonarski_fandClaude Fable 5.1 ddc7855900 Rotation scaling: frame guards and delta-CC1/2 read the frames by what they carry
Three follow-ups to the converging scale loop, all on sweeps with a stretch the crystal barely
diffracted on:

- The sweep ledger's scale channel (and with it delta-CC1/2's "normal frame" test) and the
  space-group search's scale floor measure a frame against the precision-weighted typical frame
  G_ref (TypicalFrameScale) instead of the run median: on a sweep that spent most of its turn out
  of the beam the median frame is itself a dead one, and every stretch then reads as typical. The
  two collapse guards stay on the median on purpose - a frame that collapsed toward zero and was
  not quite dropped has its fulls re-fitted with a scale of 1/G, a cubic mean is then theirs, and
  the floor read against it dropped three quarters of every live frame's fulls (measured).

- delta-CC1/2's sigma-tau statistic enters each reflection with the information it carries (the
  same per-frame factor as the CC1/2 weight), so a dead stretch, whose scaled-up noise flooded the
  mean error variance of every mixed reflection and read as harm up to the 25% cap, now costs
  about nothing and is left to the ledger.

- The scale loop's step test weighs each frame by its merge weight (its observations at its scale
  squared): a dead frame's scale is fitted on noise and wanders by orders of magnitude every
  iteration, carries nothing into the merge, is dropped after the loop, and must not hold the loop
  open.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:17 +02:00
leonarski_fandClaude Opus 5 9de53ad015 RotationScaleMerge: download the GPU merge accumulators in slices
The device merge's per-group sums were downloaded into a full-length host array per field and then
unpacked into the per-group accumulators, so every merge held two full copies of them - on a large
cell searched in P1 that is most of a gigabyte beside the accumulators themselves. MergeAccum now
leaves the sums on the device and MergeAccumRange downloads a million groups at a time, each slice
unpacked straight into the accumulators.

Same values, same integer reject count; merged output byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:17 +02:00
leonarski_f b5b86f0266 Merge branch 'c-beamcentre-arbiter' into pool-C 2026-09-20 18:45:17 +02:00
leonarski_fandClaude Fable 5.1 36c24b6e5e rugnux: a distance the data cannot place stays at the header
The joint post-refinement is asked twice, with the distance free and with it held at
the header. Where freeing it lowers the fit's held-out residual by more than that
residual's own standard error, the free fit is committed exactly as before. Where it
does not, the fit cannot tell - its own gain need not survive re-integration - so
both hypotheses are carried: the canonical pass runs at the free geometry as before
and once more at the header distance, and the run keeps the one whose RE-INTEGRATED
held-out residual is lower, the header's on a tie. The report says which
(POSTREFINE_DISTANCE_HELD), the pass decision names both realised residuals, and the
log prints the free fit's formal sd(distance) and its correlation with the cell
lengths, from the covariance of the solution.

Why: at a detector far enough away that no reflection reaches more than a few
degrees of 2theta, a longer distance and a larger cell move every spot the same way
to first order (the difference is of order sin^2 theta of the spot's position -
0.4 px rms per per cent over a 2M detector at 820 mm, against 2-3 px at usual
distances), so the fit finds a distance/cell pair that fits its own observations a
little better than the header, commits it, and the pass re-integrated there asks for
the next pair: a walk along the degenerate direction the realised residual never
ratifies. XDS leaves the distance out of IDXREF by default and its documentation
says to remove it from CORRECT where it drifts on low-resolution data; DIALS fixes
the wavelength for the same reason. Here the data are asked instead of a rule.

Measured: a sweep at 820 mm reaching 2theta ~ 9 deg no longer walks to 846 mm with
the cell 3.3 % too large (pool-B stopped it at 825.6 mm, still +1.1 %); it keeps
820 mm and its cell agrees with the reference at that distance to 0.2 %, R_meas
101 % -> 70 %, two passes fewer. 9yl4 keeps 600 mm (cell 0.5 % -> 0.1 % from the
deposition), 7mzt keeps 400 mm (R_meas 169 % -> 145 %, CC1/2 0.961 -> 0.986), 5lzl
keeps 689 mm on a tie (cell -0.3 % instead of +0.25 %, ISa 13.9 -> 11.6). Merged
output byte-identical on a lysozyme reference sweep at 110 mm (the free distance
wins its realised comparison), 8egn and 8pqd (decisive in-fit gains, no second arm)
and 8qq7 (the quality guard had already reverted it); one extra canonical pass
wherever the two arms are run.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:17 +02:00
leonarski_fandClaude Fable 5.1 004839a8ee Space-group search: the twin-immune zone arbitrates a refused promotion
A promotion the operator correlations confirm and a gate refuses (the H ratio, the added-operator
R, the merge chi^2) used to be settled two ways that both read the per-frame scales: the two-arm
rule let the other arm's confirmation outvote the refusal, and the remerge arbiter compared the
R_meas and ISa of the two pinned merges. Neither holds once the scales have settled: a partial
twin at a fraction one arm's coverage hides passes there (an H3 twin promoted to R32, a P31 one to
P3121), and the lower group's converged merge - the same scales fitted against fewer equivalents -
agrees with itself better than the true group on a genuine step (a 2 -> 222 step reading R_meas
0.094 -> 0.103), so the arbiter refuses real symmetry.

The refusal is now decided by the twin-immune zone of the operators the promotion adds: the
reflections centric in the higher group and acentric in the lower are their own twin mates, so
they read centric if the operators are real and acentric if they are a twin law or a
pseudo-symmetry, whatever the twin fraction and however the scales were fitted. Read on the
all-observation P1 merge in the two-arm rule and on the adopted group's own merge at the arbiter
(where a weak crystal's P1 search merge has no shell with signal to read), for a group two orders
above the other, with the pseudo-translation normalisation the L-test uses, and decided on the
sign with the screw-axis convention of 20 nats: acentric and the refusal stands on all
observations, centric and the higher group is adopted, undecided and the rules stay as they were.
Measured: a genuine 222 step reads +150 to +1630 nats, the trigonal twins -130 nats and below.

A refusal that stood on the zone names the law: the first operator the refused group adds, at the
fraction the operator's own H implies (else the pre-search L-test), reported as TWIN_LAW and
written as the mmCIF _pdbx_reflns_twin loop of the merge in the true group. Every lattice twin law
outside the adopted group is reported with its own H, implied fraction, R and correlation on the
merge expanded to P1 (TWIN_LAW_n), following Yeates (1997) Methods Enzymol. 276, 344-358.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:17 +02:00
leonarski_fandClaude Opus 5 ca9a21c45e rugnux: keep a pass's per-image reflections in large blocks
The whole-run passes retain every frame's integrated reflections until scaling is done - thousands
of vectors of a few megabytes each, allocated by the image workers in the allocator's per-thread
arenas. When a pass hands them back, most of that memory stays in those arenas as holes, and the
next pass's workers (new threads) do not reuse it, so on a fine-sliced long axis gigabytes of freed
reflections were carried to the end of the run.

IndexAndRefine now copies each retained frame's reflections into a ReflectionArena: 64 MiB blocks,
each its own mapping, carved by a bump pointer and returned to the system in one piece when the
last vector in them is gone. IntegrationOutcome::reflections becomes a std::vector with an allocator
that uses the arena when given one and plain new/delete otherwise (copies go to the heap), so the
read sites are unchanged; the few functions that took the vector by type now take a span.

No arithmetic changes; merged output byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:17 +02:00
leonarski_fandClaude Opus 5 1b414909cb rugnux: the model maps come from FFTW too, so the model path uses one FFT library
The map-coefficient to real-space transform was still gemmi's bundled pocketfft while the
map to structure-factor direction had moved to FFTW. MapFromFPhi does that inverse with
FFTW - the same half-l XYZ layout, the same 1/V scale, NaN coefficients read as zero, the
same conjugation convention - and the output maps are now built with it. Plans are
FFTW_ESTIMATE, cached per grid size and direction, and made under the shared planner lock.

No pocketfft code is instantiated in rugnux any more (75 symbols -> 0); gemmi's fourier.hpp
is still included for its non-FFT helpers (get_size_for_hkl, get_f_phi_on_grid), so the
vendored header and its notice stay.

Tested element-wise against gemmi on odd and even grids, with arbitrary phases, negative
indices and symmetry/Friedel expansion in three space groups, plus a round trip; both new
tests fail if the conjugation is dropped. On the eight --model audit sets the merged MTZ is
byte-identical, the map coefficients are identical, the CCP4 maps agree to 3e-7 relative
(map CC 1.00000000) and every reported key is unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:17 +02:00
leonarski_fandClaude Opus 5 8be83841f4 rugnux: the beam-centre arms are arbitrated only while they still disagree
The beam-centre check carries both centres forward when indexing at each
returns the same cell in a DIFFERENT metric symmetry, and RunAllPasses then
adopts whichever arm merges better on its P1 search merge. That question is
asked before the short-axis pass and before the geometry post-refinement,
either of which can replace the lattice it was asked about: on a crystal whose
shortest axis is under the FFT floor the check compares two supercells that the
short-axis pass then throws away, and both arms end on the same sub-cell in the
same class, post-refined to the same centre within a fifth of a pixel.

The arbiter was then deciding between two integrations of ONE hypothesis on a
statistic whose difference between them is noise. On a small-molecule rotation
sweep whose per-frame scales collapse on a fifth of the frames, the two arms'
search merges came out 0.543 and 0.594 - a 0.051 margin against the 0.05 the
run moves off the file's centre for - and the adopted arm took the run from
P 21 21 21 at 0.65 A to P 1 2 1 at 1.79 A on the same cell.

The arms are now compared on the lattices they ENDED on: where both hold the
same class, the same centring and primitive volumes within 2 %, the centre did
not decide the metric symmetry after all, the merges are not asked to arbitrate
it, and the file's centre stays (the post-refinement moves it wherever the spots
put it). Where they still differ the merge comparison runs exactly as before,
and both branches now log the two arms' final classes. ProcessResult carries the
rotation lattice's class for the comparison.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:17 +02:00
leonarski_fandClaude Opus 5 8bc4c6e8a6 Merge pool-B into rc172
Pooled run B: twin report-only evidence and free-R on the lattice holohedry,
information-weighted resolution cutoff with weighted outlier median and ISa
rework, sparse-lattice integration and beam-centre arbiter, GPU/CPU collapsed
scale upload, geometry walk on the realised residual, --model speed-ups with
FFTW structure factors, memory reductions for large-cell sweeps, -A decisions
Friedel-merged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:17 +02:00
leonarski_fandClaude Fable 5.1 61afff891b Rotation scaling: the per-frame scale loop runs to convergence, gauge pinned, plain fit
The alternating per-frame scaling used to run a fixed three rounds of a Cauchy-reweighted fit
against a reference that included half a sweep's rocking-curve tails. Three rounds left a short
sweep merged in P1 far from its answer (a 90 deg tetragonal sweep refused its 422 with the P1
scales anti-correlated with the converged ones), and more rounds did not help: the fit had no
fixed point. Two things made it walk. The objective is invariant under G -> cG with the reference
-> reference/c, so every round moved every scale by a constant factor; and the robust loss, iterated
against a reference refitted each round, drops the strong reflections of a frame whose scale is off
by a third (ten-sigma residuals) and lets the weak ones carry it further off - measured on a 360 deg
sweep the scales shrank 10-30% per round for thirty rounds and the H ratio of a genuine 222 read 21x.

Now the loop pins its gauge every iteration (G divided by the precision-weighted typical frame
scale G_ref = sum G^3 / sum G^2, one definition shared with the CC1/2 weight), fits the plain
weighted least-squares slope with the weights the reference uses and only on the observations the
reference is built from (the partiality floor), and stops when the rms |log(G_new/G_old)| over the
frames falls below 1e-3. With the same weights on both sides the alternating fit is exact
coordinate descent on one objective and cannot climb; measured, the 360 deg sweep settles in 19
rounds and the 90 deg one in 30-50, each pass's partials and fulls loops alike.

--scaling-iterations is now the cap (default 100). A loop that reaches it is logged, the report
prints SCALING_ITERATIONS and raises SCALING_NOT_CONVERGED, and the correction surfaces run to the
same tolerance under their own cap. On the GPU the loop runs one iteration per call so the pin and
the step test read the same numbers as on the host; the fulls' reset is split out of ScaleFulls.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:17 +02:00
leonarski_fandClaude Opus 5 bb55715d28 RotationScaleMerge: stream the resident ingest instead of materialising a record per observation
The ingest-time geometry smoothing (delta_phi from the smoothed lattice, one exact-Bragg angle per
rocking event, partiality from the smoothed mosaicity) used to run over a 32-byte host record per
observation, copied from the source reflections and held beside them. It now reads hkl, frame, d
and zeta straight from the source reflections, frame by frame, and keeps arrays only for what it
rewrites (delta_phi, partiality) and what the rocking-event walks read in raw-hkl order
(image_number, and on the resident path the usability flag): 13 bytes an observation instead of 32.
The full-Obs path runs the same code and copies the two rewritten fields back.

Same arithmetic in the same order; merged output byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:17 +02:00
leonarski_fandClaude Opus 5 c63f9973bd v1.0.0-rc.172
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:04 +02:00
leonarski_fandClaude Opus 5 1545bd2bbb tools/battery: verdicts judge lattice and symmetry only
Data quality - resolution, CC1/2, R_meas, completeness, ISa - is no longer a
pass/fail criterion anywhere; it is reported beside the reference in the tables
and plots for a human to judge. The CC1/2 noise ratio against XDS stays as a
reported guide.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:04 +02:00
leonarski_f 5c8011fe67 compression: verify the bitshuffle seam for Apple Silicon; SIMD test knows NEON
The x86/ARM split asked for on macOS already exists: BitShuffleBlock.h (fbc507839) sends
the per-block transform to bitshuffle_hperf everywhere except aarch64 with NEON, where it
uses the classic bitshuffle's NEON path, and both call sites go through it. Nothing new is
needed for Apple Silicon - clang --target=arm64-apple-macos defines __aarch64__ and
__ARM_NEON exactly as aarch64 Linux does, both vendored files compile to arm64 Mach-O
objects, the seam resolves to bshuf_(un)trans_bit_elem and bshuf_using_NEON() is 1. The
header now says why the switch is a preprocessor one: a universal build compiles it once
per architecture, which a CMake-time answer cannot follow.

What had never happened is the NEON code actually running. It has now, under
qemu-aarch64 in the project's cross image: 504 blocks (elem 1/2/4/8, 8 elements up to the
128 kB block, odd multiples of 8, random and detector-like data) plus 54 whole-buffer
bitshuffle/LZ4 streams with element counts that are no multiple of 8 or of the block. The
output is byte-identical (same md5) across aarch64 NEON, aarch64 without NEON (seam falls
back to hperf's portable code), x86 hperf AVX2 via ifunc, hperf SSE2, hperf portable
fallback, and classic SSE2 / AVX2 / scalar; every implementation decodes every other.

Throughput, one thread on a loaded Zen 3, 16/32-bit, GB/s encode/decode: hperf AVX2
~9-12/~10-11, classic AVX2 ~8/~5.5-8, classic SSE2 ~4-5/~4.5-5.5, hperf portable fallback
auto-vectorised ~2.8/~5, not vectorised ~1/~1.8. Classic 128-bit SIMD beating hperf's
fallback is the x86 stand-in for the choice the seam makes on ARM; no ARM hardware was
available, so the NEON-vs-fallback ranking on a real core is still unmeasured.

The one thing that would have failed on aarch64 is the Bshuf_SSE test, which required
SSE2 outright. It now accepts NEON as well, which also makes it the check that a Mac or
DGX Spark build did not end up on the scalar path.

Not built: no CMake configure or jfjoch_test build was run on this machine (busy); the
test expression was compiled and run standalone on x86 only.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 5957b15493dfd0cc8a18fb780a18c2573a6eccde)
2026-09-20 18:45:04 +02:00
leonarski_fandClaude Opus 5 57ab5e640e tools/battery: merge quality is CC1/2 against XDS only
The owner's scoring decision: no R_meas criterion, and merge quality
judged only by CC1/2 relative to XDS.

- The global rule (R_meas > 60% or CC1/2 < 0.5 fails) is gone.
- XDS arms: a set fails `merge` when (1/CC1/2 - 1) of the reference-range
  table is more than twice XDS's, XDS's CC1/2 taken at the bottom of its
  0.1% rounding. CC1/2 = S/(S+E), so 1/CC1/2 - 1 = E/S at any CC1/2, and
  E goes as 1/observations: 2x is XDS's merge with half its observations.
  Stored as cc_half_noise_ratio. No REFRES table or no XDS CC1/2: no
  criterion.
- Open arm: no merge criterion; CC1/2 and R_meas stay reported numbers.
- Low-resolution R_meas is reported, not scored: the lowest shell of the
  own table, the reference-range table and XDS's (lowres_*), with
  lowres_r_meas_ratio in the like-for-like table and a ratio plot.
- Every schema-3 run is re-scored when it is read (report, compare,
  baseline delta) with today's scorer and today's manifest rows, so both
  sides of a comparison are scored alike. results.json keeps the verdicts
  as scored at run time; rows without a lattice keep them.

Re-scoring the rc172 reference run: inhouse 29 -> 27 pass (three new
merge fails, one R_meas fail lifted), open 137 -> 138 (one R_meas fail
lifted).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:04 +02:00
leonarski_f 6a21d453ba macOS groundwork: what a static read says will stop an Apple Clang / libc++ build
None of this has been built on a Mac - there is none yet. It is the list a read-only audit of the
viewer/rugnux subtree produced, plus a serial -fsyntax-only pass of every reachable .cpp with
clang 16 + libc++ on Linux, which found exactly one error (the first item).

- JFJochDatasetInfoChartView: std::vector<fftwf_complex> does not compile with libc++, whose
  construct_at refuses an array element type (float[2]). Use std::vector<std::complex<float>> and
  the reinterpret_cast every other FFTW call site already uses.
- libcurl: GSSAPI off on Linux, and neither TLS nor GSSAPI on macOS. The viewer never sets
  CURLOPT_HTTPAUTH, so Negotiate was dead weight that cost a krb5-devel build dependency; on macOS
  curl's FindGSS refuses the system Heimdal outright, and with Secure Transport gone from curl
  (8.15) TLS would mean a Homebrew OpenSSL - the only host library a Mac build would need.
  Linux keeps OpenSSL. CURL_USE_GSSAPI is forced OFF rather than left unset so an existing build
  tree drops its cached ON.
- libjpeg-turbo ExternalProject: CMAKE_SYSTEM_NAME/PROCESSOR were forwarded unconditionally, which
  puts even a native sub-build into cross-compiling mode, and CMAKE_OSX_ARCHITECTURES / SYSROOT /
  DEPLOYMENT_TARGET were not forwarded at all. Now the same rule the zlib-ng sub-build follows.
- ShadowAccumulatorGPU.cu was added on the JFJOCH_USE_CUDA option (default ON) instead of
  JFJOCH_CUDA_AVAILABLE like every other .cu, so a machine without nvcc got a CUDA source in a
  target with no CUDA language.
- CMAKE_OSX_DEPLOYMENT_TARGET defaults to 12.0 (overridable). Left unset, CMake takes the build
  machine's OS version and the .dmg starts nowhere older.
- Standard headers that were only arriving transitively (<chrono>, <cmath>, <limits>, <cstring>,
  <atomic>, <thread>, <string>); newer libc++ releases keep removing such transitive includes.

Checked: the seven changed sources pass clang 16 + libc++ -fsyntax-only. The CMake changes are
not configured or built.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eAE2K7i5JGDwgwifiCfuA
(cherry picked from commit 5a7282759a)
2026-09-20 18:45:04 +02:00
leonarski_fandClaude Opus 5 377d1e1b25 tools/battery: the XDS references carry the lowest shell's R_meas
parse_correct_lp also reads the first row of CORRECT.LP's last resolution
table: its high-resolution limit (dmin_low) and R_meas (r_meas_low).
inhouse.json regenerated with `refs --arm inhouse --write`; the only
change is the two new keys per set (the private manifest was regenerated
the same way, outside the repository).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:04 +02:00
leonarski_f 08771694b3 writer: GetCompressionAlgorithm passed a byte count as the filter-parameter count
H5Pget_filter2 takes cd_nelmts as the number of elements the cd_values buffer can hold and
writes up to that many. It was given sizeof(cd_values) = 32 for an 8-element array, so a
filter carrying more than 8 parameters would have been written past the end of the buffer.
Bitshuffle carries 5, so nothing overflowed in practice.

The "Weird value" message printed cd_values[1] while the test above it is on cd_values[4].

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eAE2K7i5JGDwgwifiCfuA
(cherry picked from commit dd1af9576c)
2026-09-20 18:45:04 +02:00
leonarski_fandClaude Opus 5 9e2a7a13c2 rugnux: a rotation run decides with Friedel mates merged, -A or not
-A set MergeFriedel(false), and RotationScaleMerge builds one grouping for the
whole pass from it, so every decision on an anomalous rotation run was taken on
hand-separated groups: per-frame scaling, the correction surfaces, the outlier
median, the error model, the CC1/2 half-sets and the resolution cut, the
delta-CC1/2 frame ledger and the space-group search. That halves the
multiplicity each of those gates was calibrated at, and flipped the bare run's
symmetry, resolution and frame-rejection calls on several sets - none of them
through the anomalous signal (a 422 refused on an added-operator R, a 23
refused on a chi^2 gate, a screw row read on hand rows, thin hands exempt from
the outlier median letting one hot pair own a CC1/2 bin).

A rotation run now always merges with Friedel's law: the Rugnux constructor
turns MergeFriedel back on when the run is rotation (IsRotationIndexing, which
--force-still clears), and --mode scale does the same for its rotation merge.
Nothing is lost for anomalous use: the merge always keeps the Bijvoet split, and
I(+)/I(-), F(+)/F(-), SigAno and CCanom are written and reported from it
whether or not -A was given. -A on a rotation run is therefore a no-op on the
processing, which is the point; stills keep -A as their only anomalous route.
A bare run is untouched.

This is step 1 of the -A design (decide with pairs merged, present with them
split). The per-hand statistics table that -A should select is not built yet,
so an -A rotation run now reports the Friedel-merged table (FRIEDELS_LAW=
TRUE, Laue multiplicity and completeness). The thin-hand pair median of the
previous commit is unreachable from rugnux on rotation data but kept for any
direct MergeFriedel(false) caller of RotationScaleMerge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:04 +02:00
leonarski_fandClaude Opus 5 e016bc6c78 tools/battery: verdicts judge lattice and symmetry only
Data quality - resolution, CC1/2, R_meas, completeness, ISa - is no longer a
pass/fail criterion anywhere; it is reported beside the reference in the tables
and plots for a human to judge. The CC1/2 noise ratio against XDS stays as a
reported guide.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:04 +02:00
leonarski_fandClaude Opus 5 a827e20bec tools/battery: merge quality is CC1/2 against XDS only
The owner's scoring decision: no R_meas criterion, and merge quality
judged only by CC1/2 relative to XDS.

- The global rule (R_meas > 60% or CC1/2 < 0.5 fails) is gone.
- XDS arms: a set fails `merge` when (1/CC1/2 - 1) of the reference-range
  table is more than twice XDS's, XDS's CC1/2 taken at the bottom of its
  0.1% rounding. CC1/2 = S/(S+E), so 1/CC1/2 - 1 = E/S at any CC1/2, and
  E goes as 1/observations: 2x is XDS's merge with half its observations.
  Stored as cc_half_noise_ratio. No REFRES table or no XDS CC1/2: no
  criterion.
- Open arm: no merge criterion; CC1/2 and R_meas stay reported numbers.
- Low-resolution R_meas is reported, not scored: the lowest shell of the
  own table, the reference-range table and XDS's (lowres_*), with
  lowres_r_meas_ratio in the like-for-like table and a ratio plot.
- Every schema-3 run is re-scored when it is read (report, compare,
  baseline delta) with today's scorer and today's manifest rows, so both
  sides of a comparison are scored alike. results.json keeps the verdicts
  as scored at run time; rows without a lattice keep them.

Re-scoring the rc172 reference run: inhouse 29 -> 27 pass (three new
merge fails, one R_meas fail lifted), open 137 -> 138 (one R_meas fail
lifted).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:04 +02:00
leonarski_fandClaude Opus 5 1a85fbfbee tools/battery: the XDS references carry the lowest shell's R_meas
parse_correct_lp also reads the first row of CORRECT.LP's last resolution
table: its high-resolution limit (dmin_low) and R_meas (r_meas_low).
inhouse.json regenerated with `refs --arm inhouse --write`; the only
change is the two new keys per set (the private manifest was regenerated
the same way, outside the repository).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:04 +02:00
leonarski_fandClaude Opus 5 ffe649efae tools/battery: a set with a known-wrong reference can be left unscored
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:04 +02:00
leonarski_fandClaude Opus 5 629ea58239 RotationScaleMerge: a 32-byte resident ingest record, and smaller staging slices
On the resident (GPU) path the ingest's narrow host record sits on top of the source reflections
for the whole smoothing and partiality recompute, and on a long axis that moment is the run's
memory high-water mark. Store its Miller indices in 16 bits and move the rocking-event flag into
their padding: 40 -> 32 bytes a partial. Stage the device upload in slices of two million
observations rather than eight, which takes the thirteen staging arrays from ~400 MB to ~100 MB.
Neither changes a value.

Measured on a 3600-frame long-axis rotation set (-N 6): VmHWM 13.79 -> 12.80 GiB on top of the
previous commits (15.0-15.1 GiB before any of them); merged MTZ/HKL/CIF, P1 MTZ, per-image table and
the report are byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:03 +02:00
leonarski_fandClaude Opus 5 2f612d6aae tools/battery: the 16 keV thaumatin reference obeys Friedel's law
At 16 keV a native crystal has no usable anomalous signal, so XDS's CORRECT
was rerun with FRIEDEL'S_LAW=TRUE and the set no longer runs with -A.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 18:45:03 +02:00
leonarski_fandClaude Opus 5 66db3353c7 rugnux: measure the canonical pass's geometry probe before the scaling engine is built
The post-refinement probe on a canonical pass only measures, and it reads nothing the scaling
engine produces - the reflections' hkl, I, sigma and positions are final once the images are
integrated, and the engine writes back per-frame fields only. It ran after the first merge,
beside the engine's arrays, where its gathered observations and rocking events set the pass's
memory high-water mark. Run it before the engine is built instead; the pre-pass keeps its
original place (it consumes the smoothed mosaicity the merge writes back).

Measured on a 3600-frame long-axis rotation set: the canonical pass's high-water drops from
14.8 GB to under 12.8 GB (sampled RSS, -N 6); merged MTZ/HKL/CIF, P1 MTZ, per-image table and the
report are byte-identical, the probe's own result included.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:03 +02:00
leonarski_fandClaude Opus 5 764add208d tools/battery: the like-for-like table covers XDS's own range
--report-resolution now gets XDS's own d_min, the range its CORRECT.LP totals
cover, so the comparison is exact; shells past rugnux's cut print as not
merged. The derived reference d_min stays what rugnux's own cut is scored
against.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 18:45:03 +02:00
leonarski_fandClaude Opus 5 3e5e010984 RotationScaleMerge: ingest without the index-order copy of the sort keys
Ingest wrote every observation's 24-byte SortKey into an index-order array and then scattered that
array into its h buckets - two key arrays alive together at the top of the ingest. Count the
buckets per chunk of frames straight from the source reflections and scatter from them instead:
a frame's reflections are consecutive in the key numbering, so the chunks are consecutive index
ranges and each bucket lands in the same index order as before. The key array is byte-identical.

BuildInRangeObservations also hands back the remap temporaries (old runs, per-run maps, new_idx)
before the observation array is built rather than at the end of the function.

Measured on a 3600-frame long-axis rotation set (95.7 M partials in the geometry pre-pass): the
pre-pass ingest high-water drops 15.45 -> 14.1 GB (sampled RSS, -N 6); merged MTZ/HKL/CIF, the P1
cross-check MTZ and the per-image table are byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:03 +02:00
leonarski_fandClaude Opus 5 665338d3d1 tools/battery: one rugnux run per set
--model and --report-resolution are both processing-inert (merged reflections
bit-identical with and without them), so the bare/model and bare/xds variant
pairs collapse into one command per arm:

  open:          rugnux -o p --no-export-unmerged --model <deposited coords> <input>
                 (no --model where there is no deposited model)
  inhouse/private: rugnux -o p --no-export-unmerged [-A] --report-resolution
                 <reference d_min>,<XDS d_max> <input>

Every row is rugnux's own result (own cut vs the reference, as `bare` was), and
on the XDS arms also carries the REFRES_* table, scored against CORRECT.LP's
completeness, multiplicity, R_meas, CC1/2 and ISa, with ISa and R_meas ratio
plots and a like-for-like table; REFRES_SHELLS_PAST_LIMIT > 0 marks the row as
coverage (rugnux's cut coarser than the reference). Nothing is forced on the
processing any more. On the open arm the space group is scored on the data's
own determination (SOHNCKE_SPACE_GROUP where SPACE_GROUP_ENANTIOMORPH is
ASSUMED_FROM_MODEL), the model's label kept as sg_label; R_FREE/R_WORK/CC_MODEL
/MODEL_FIT are trend fields (placement-only, own free set); REFMAC stays opt-in.

results.json schema 3: one row per set, no variant/first_read; older runs are
read through their bare rows (open-arm model R-factors folded in), so compare
and report keep working against them. Every time is now a first read.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 18:45:03 +02:00
leonarski_fandClaude Opus 5 8eaa4b9df7 Bragg integration: size each image's reflection list to what is kept
Finalize reserved npredicted reflections and kept only the ones whose fit succeeded. On a dense
long-axis pattern more than half the predicted reflections lose their background ring, so each
retained per-image vector carried more dead capacity than data for the rest of the pass. Count the
kept ones first and reserve exactly that. No result changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:03 +02:00
leonarski_fandClaude Opus 5 0ef6b3f00b rugnux: the run's own N_obs and R_meas stop at the same cut as N_uniq
The own statistics table's grid bottoms 0.1% below the finest kept reflection,
and the observation-level counts (N_obs, R_meas, the Bijvoet split behind SigAno
and CCanom) re-walk the fulls, which still hold every ingested group. Groups in
that band just below the auto cut were erased from the merge, so they were not
in N_uniq, but their observations still counted in N_obs and R_meas of the
finest shell. The own table now floors those counts at the cut by group d, the
rule the erase applies - the same floor the reference-range table already used
(now passed as a double, so the comparison is exactly the erase's).

Report-only: on an auto-cut cubic in-house set the written MTZ, HKL, P1 MTZ, unmerged
MTZ and image table are byte-identical; TOTAL_OBSERVATIONS 378754 -> 377933,
MULTIPLICITY 40.31 -> 40.22 (now equal to the reference-range table's), finest
shell N_obs 60873 -> 60052, R_meas 1136.9% -> 1129.2%, CCanom -0.6% -> -0.3%;
overall R_MEAS, CC1/2, SigAno unchanged at printed precision; the mmCIF's
pdbx_number_measured_all / pdbx_redundancy / last shell row follow. A
detector-limited lysozyme run (no auto cut) is byte-identical throughout.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 18:45:03 +02:00
leonarski_fandClaude Opus 5 90e84718fe tools/battery: a set with a known-wrong reference can be left unscored
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:03 +02:00
leonarski_fandClaude Opus 5 6a26591786 rugnux: the indexing-ambiguity warning is withdrawn only where the model settled it
The INDEXING_AMBIGUITY warning was suppressed whenever --model was given on
rotation data, decided before the model was read. A model that then decided
nothing - not tested, rejected, or a probe winner that did not beat its own null
- still silenced it, so WARNING_COUNT, PATHOLOGY_FLAGS and possibly VERDICT
differed from a run without the model although the written reflections were
identical.

The warning is now issued exactly as without a model, and withdrawn (from the
warnings and from the statistics text) after model validation only where the
indexing probe decided the indexing: the model fits and the winner's R-free
margin beats the random-placement null (ModelValidationResult::indexing_decided,
set where the decision is taken). A reference MTZ, or the model reference on
serial stills with -C and -S, suppresses it up front as before.

Verified bare vs --model on three open-arm sets: merged MTZ data identical in
all three; warnings identical where the model decided nothing (identity probe
without null; no twin law); withdrawn where the probe decided (+33 sigma).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 18:45:03 +02:00
leonarski_fandClaude Opus 5 65e7b658e5 rugnux: with Friedel mates merged separately, a thin hand takes its outlier median from the pair
The rot3d merge tests each observation against the weighted median of its
reflection, and only from three observations up: at two the median is one of
them. With -A each hand of a Friedel pair is its own reflection, so a P1
anomalous merge sits at about two observations a hand and almost nothing is
tested - on one P1 sweep the merge rejected 0 observations with -A against
tens of thousands without it. A single observation a hundred times its
partner's intensity (300 sigma apart, a 3.9 A reflection) then survived into
the merge, carried 94% of the weighted variance of its CC1/2 bin, took that
bin from 0.80 to 0.18, and the automatic cut with it: 3.46 A and UNUSABLE,
where the same data without -A cut at 1.40 A.

A hand with fewer than three observations now takes the weighted median of
both hands together, where the pair has three. I(+) and I(-) differ by the
anomalous signal, a few percent of I and far inside the six-sigma test, so
the mate supplies the observations the hand is missing. A hand with three of
its own keeps its own median exactly as before, and a Friedel-merged run is
unchanged. The medians are formed on the host, so the device merge reads the
same ones.

Measured with -A: the P1 sweep 3.46 A / UNUSABLE -> 1.49 A (1.63 A XDS
reference; 1.40 A without -A, unchanged), 0 -> 3,699 rejections. Two lysozyme
and one insulin set: same cut, space group, CC1/2 and ISa. On one of the
lysozyme sets 16 more observations are rejected, every shell's CCanom is
unchanged, and the overall CCanom goes 0.31 -> -0.02: the whole-range figure
was being carried by a handful of wild pairs, while the shells ranged -18% to
+12% and XDS's overall anomalous correlation is 1%.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:03 +02:00
leonarski_fandClaude Fable 5.1 7547df77de rugnux: --report-resolution, a second statistics table at a reference range
Comparing a run with another program's table has meant running rugnux AT that
program's resolution range (--scaling-high-resolution), which is a different
run: the range moves the cut, the space-group decision and everything after
them, so the comparison buys itself a different answer. --report-resolution
<dmin>[,<dmax>] instead leaves the run alone and adds a second table to
section 3 of the report - the REFRES_* keys and a shell table - binned from the
same merged reflections over the range given, with the completeness
denominator enumerated over that range and the shells in equal steps of 1/d^2
so they read row for row against a CORRECT.LP at the same range. Report-only:
the merged files and every decision are byte-identical with and without it.

The table holds only what the run kept. Where the reference range is finer
than the run's own limit, the shells past it are printed as not merged (with
their possible count) rather than as zeros, REFRES_SHELLS_PAST_LIMIT counts
them so a consumer can tell "not merged" from a measured zero, REFRES_
COMPLETENESS counts their reflections as missing, and the other overall numbers
are over the shells the run reached; nothing is read from the observations the
run judged to carry no signal. REFRES_ISA is the error model refitted on the
reflections of the table alone, in XDS's convention (rotation only; the stills
model is fitted over the whole range already).

On the rotation path the statistics block of MergeAndStats becomes a lambda
over a shell grid, called once for the run's own grid and once for the
reference one; the reference call floors every observation-level count at the
cut by group d, the rule the erase applied. The stills MergeStats takes a
declared range, whose bounds are the grid's whether or not any reflection
reaches them. Both --mode mx and --mode scale report it, the viewer's command
line echoes it, and the docs describe the keys.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 18:45:03 +02:00
leonarski_fandClaude Opus 5 440e19343c rugnux --model: MapToFPhi returns the conjugate, as gemmi's transform does
gemmi::transform_map_to_f_phi conjugates the forward FFT at the end (a
structure factor is the sum over exp(+2 pi i h.x), a forward FFT the sum over
exp(-2 pi i h.x)); MapToFPhi did not, so every model-path Fcalc and Fmask came
out with its phase negated. Amplitudes, and so the R-factors and the scale,
were unaffected; the phases of the maps and everything read off them (map
coefficients, density at atom centres) were not. Caught by
ModelValidation_MapToFPhiMatchesGemmi, which now passes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:03 +02:00
leonarski_fandClaude Fable 5.1 53715d010d rugnux: overall SigAno and CCanom over the written range only
The per-shell SigAno and CCanom of the rotation merge were accumulated on the
report grid, but the overall pair of numbers - the SIGANO / CC_ANOM keys and
the mmCIF's pdbx_absDiff_over_sigma_anomalous - were summed over every Bijvoet
pair the fulls hold, including the ones past the automatic resolution cut that
the written reflections do not contain. On a run the cut trims, the overall
line therefore described more data than the shells above it add up to. Both
overall numbers now count only the pairs that land in a shell, as the rest of
the table does; a run with a manual limit, whose ingest already ends at the
limit, is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 18:45:03 +02:00
leonarski_fandClaude Opus 5 7dea9eaf9c One process-wide lock for FFTW planning
FFTW's planner (every fftwf_plan_* and fftwf_destroy_plan) shares global state
and is not thread-safe; executing a plan is. FFTIndexerCPU and BeamCenterFFTCPU
each guarded their planning with a lock of their own, TranslationalNCS and the
viewer's spectrum with none, and ModelFFT with a third - which does not stop
two of them planning at once. common/FFTWPlannerLock.h holds the one mutex they
all now take. No numerical change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
2026-09-20 18:45:03 +02:00
leonarski_fandClaude Fable 5.1 060a89cd7d CUDA: allocation streams are handed back, not leaked; a handled pool failure leaves no error
A long-running broker began cancelling every data collection with

    Device decoding failed (CUDA (GPU) error (out of memory)), falling back to host decompression
    CUDA (GPU) error (out of memory)

while nvidia-smi showed the cards less than a fifth full. Two defects, both in CUDAMemHelpers.h and
both from the pooled allocator (cudaMallocAsync) that came with rc.162.

The leak. cuda_allocation_stream() kept one stream per (thread, device) in a thread_local map of raw
cudaStream_t and never destroyed them. That was written against rugnux, where the worker threads live
as long as the process. The broker starts fresh std::async threads for every data collection - 16 or
64 of them - so every collection left that many streams behind. Measured: 0.56 MB of device memory
per leaked stream, linear to 4928 streams, never returned, with nothing on the host side growing.

The RAII wrapper (CudaStream) was there but not used at this site, and using it as-is - a stream
destroyed when its thread exits - would not have been safe: ShadowFinder builds its GPU accumulator
on a throw-away std::async thread and frees it from another thread long after, and that free is
ordered on the allocating thread's stream. So the streams are still never destroyed, but a thread
now only borrows one: CudaStream objects live in a process-wide per-device idle list, a thread takes
one on first use and hands it back when it exits. Their number is bounded by the threads that were
ever alive at once instead of by the threads ever started. The list itself is deliberately leaked, so
that nothing calls into CUDA during static destruction.

Replaying the broker's pattern against the real header, 60 collections of 64 threads:
    before   3840 streams, 260 -> 2424 MB of device memory
    after      64 streams, 260 ->  358 MB

The stale error. Every helper here throws a named message, yet the log carried the raw CUDA string,
so the failure came through a cuda_err() and not from an allocation. CudaDevicePtr falls back to
cudaMalloc when cudaMallocAsync fails, silently - but the failed call stays behind as the thread's
last error, and the cudaGetLastError() that follows the next kernel launch reports it. The buffers
were all allocated; the frame was lost anyway, once on the device-decode route (caught, hence the
warning) and once more on the host fallback (fatal). The pooled attempt failing, and a stream that
cannot be created, are both handled by falling back, so both now clear the error they leave.

What finite resource the production cards ran out of at under 4 GB used was not established - no
cap on the number of streams was found up to 4928 on the card this was measured on. The leak is the
only thing on this path that grows with uptime.

tests/CUDAMemHelpersTest.cpp: later threads end up on the same stream, concurrent threads on
different ones, a buffer is freed cleanly after its allocating thread has exited (and another has
borrowed its stream), and a pool that cannot serve a request leaves no error behind - the last by
capping a memory pool at 4 MB so that the pooled attempt fails and the fallback succeeds.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 18:45:03 +02:00