The twin-law report block dereferenced the adopted space group without the
guard every other use of it carries; a de-novo run whose search refused a
point group but left no adopted group would crash there, after the merge.
The joint post-refinement's distance-correlation array was read on every run
that holds the header distance, but only written when the solve is usable, so
an unusable first solve printed four indeterminate values as correlations.
The merge-degradation gate ran neither of its two branches when the data
confirm more than one operator and the best of them is itself close to the
random-pairing end: the contrast has no scale to read on such a merge, but the
candidate was then promoted on no merge-degradation evidence at all. It now
falls back to the noise-floor ratio, as it does when only one operator is
confirmed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The departure from a = b that decides whether the metric-symmetry arms run at
all was read by pushing the free (triclinic) refinement through the search
result's `reindex`. That matrix is stale for a whole population: where the walk
matches a metrically hexagonal lattice on its C-centred orthorhombic character,
RotationIndexer re-expresses `conventional` in hexagonal axes but leaves
`reindex` describing the setting it replaced. The C-centred setting has
b = a*sqrt(3) BY CONSTRUCTION, so a against b there is
2(sqrt(3)-1)/(sqrt(3)+1) = 53.6 % apart on every hexagonal lattice there is,
whatever the crystal.
So the gate fired on the whole hexagonal and trigonal population - eighteen
sets in one corpus run, all reporting 53.1-53.7 % - and the arms then decided
two classes that fit equally well on differences in the third or fourth
significant figure. Three of those coin flips landed on the ortho-hexagonal
supercell and took a P 61 2 2, a P 65 2 2 and a P 31 with them.
The departure is now computed by LengthEqualityDeparture, which derives the
primitive-to-conventional map from the pair of cells the search result carries
instead of trusting `reindex`, and reads a against b in that basis. The two
cells describe one lattice, so the map is the integer relabelling of its basis
vectors. The measurement no longer depends on which character the walk happened
to match.
Re-measured, the hexagonal and trigonal sets report 0.00-0.41 %, the band the
tetragonal and cubic proteins were already in, and the gate does not fire on any
of them; the genuine pseudo-symmetric small-molecule cases stand alone at 0.98,
1.85 and 2.61 %. The three lost groups come back, and the sets that merely
tolerated the noise are bit-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
A screw zone's evidence drops its largest member to defend against one badly-measured
reflection whose sigma lies about it, and such a reflection is by construction weak - a
fraction of the row it sits on. A member that has passed the violation test AND stands
at or above the mean of its row's own present class is a different animal: not a
measurement that moved, but a reflection that is there. Trimming it removed the single
datum that refutes the claim, and the violation-count deferral then read the inflated
evidence to forgive the very violation that had been trimmed out of it.
Nested screw ORDERS are decided entirely on this. 6_1 extinguishes l != 6n and 6_2/6_4
extinguish l != 3n, so the two differ only on l = 3n not 6n. On a hexagonal crystal
whose 00l row holds five present reflections, the strongest of the whole row lay in
that difference: trimmed, 6_1 read the row as perfectly dead and won on the count of
absences alone - nine at 50.6 nats with one violation against seven at 43.3 with none
- and the run reported the wrong screw order with the right one ranked below it. With
the violation left in, 6_1 reads 6.8 and is refused. An independent POINTLESS run on
the same P1 merge puts the 6_1 condition at probability 0.000 and the 3n condition at
0.998.
Trim only among members that are not both flagged present and at full row strength.
Both halves of the condition are needed and the corpus separates them: the reflection
above stands at 1.94 of its row's mean, where two monoclinic crystals whose 0k0 are
genuinely dead carry one violation each at 0.31 and 0.74 of their row - the
mis-measurement the trim exists for, and one that costs a real 2_1 if it stays in.
Zones with no violations are bit-identical, and so is every candidate whose absent
class is clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
Smoothing the per-frame scale inside the loop runs before the guard that drops frames whose scale
collapsed, and the window is a geometric mean - so a frame the crystal did not diffract on was both
dragging its whole window down by its logarithm and being lifted toward its neighbours, which is how
it escaped the guard meant to drop it. Measured: on sweeps with a genuinely dead stretch the guard
went from eight fires at up to a million times below the median to five at six hundred times, and on
a long clean sweep with a dead block from six fires to none.
A frame already below the guard's floor is now left out of every window and keeps its own fitted
scale, so what reaches the guard is the scale the fit produced. The floor is the guard's own, read
the same way - no new criterion and no new constant.
The guard's fires come back where the fit still collapses those frames (on the two dead-stretch
sets to seven and six, at up to three and seven million times below the median; on the long clean
sweep to six, the count it had before), and with them the error-model asymptote: ISa 7.9 back to 9.2
on one dead-stretch set, 36.7 to 38.0 on the clean one. Where the restraint stops the fit from
driving a frame below the floor at all there is nothing to exclude and nothing changes. No result
moves the wrong way: the fine-sliced, small-cell and powder sets keep the merges the restraint
recovered - two of them improve, one by CC1/2 0.916 to 0.943 - and every space group is unchanged,
including the short-sweep tetragonal pair and the de-novo orthorhombic set the loop was fixed for.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The alternating scale fit has a second degenerate direction beyond the gauge, and that one is not
harmless. A frame whose scale drifts below its neighbours' loses its vote in the references it is
fitted against - an observation's weight there goes as G^2 - which moves the references away from
it, which drives its scale further down. Where every frame is measured against hundreds of
reflections the references barely notice one frame and the loop settles; on a fine-sliced sweep of
a small cell, where a frame holds two or three reflections and a rocking curve spans a dozen or
more frames, the frames at the tails of the curves lose against the frames at their peaks and the
loop walks the bulk of the sweep down fifteen decades. The convergence test then said the loop had
settled, because its own weight (frame observations at G^2) vanishes for exactly the frames that
are collapsing.
So the scale is now smoothed over the rotation INSIDE the loop, at the window the merge already
smoothed G over afterwards (the post-loop smoothing is the same operation and is not repeated):
every frame in a window shares the window's geometric mean, no frame can lose its vote alone, and
what a frame's own observations cannot determine its neighbourhood's do. The window is floored at
six rocking curves - within one curve a change of scale and an error of the partiality model are
the same thing, and a window holding a curve or two fits the model's error as scale. The step test
is weighted by each frame's observation count alone, and a loop whose step has not fallen below
nine tenths of its smallest value for five rounds stops and says so rather than walking further.
Measured against the previous behaviour: a fine-sliced sweep that collapsed to P1 with CC1/2 0.000
merges in its true orthorhombic group at CC1/2 0.95; a monoclinic small-cell set goes from an
unusable 2.13 A merge to 0.81 A at R_meas 11.5%; another from R_meas 33% / ISa 2.2 to 6.1% / 14.2;
two powder-ring sets and a short-sweep tetragonal one improve. The space group is unchanged on
every protein control measured, and the short-sweep and de-novo orthorhombic gains of the
convergent loop are kept.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The Bravais class is chosen by a walk that reads two axes as EQUAL when they
agree to a fixed relative tolerance, and the class carrying that equality is
then imposed on everything below: the cell is refined with a = b, reflections
are predicted from it, and they are integrated at those predicted positions.
On a small cell the tolerance is far wider than the spot positions resolve, so
a genuinely orthorhombic crystal whose a and b differ by 2 % is integrated as
tetragonal and every decision downstream is read off the wreckage.
Measure the equality instead of assuming it. The indexer already refines each
candidate a second time with nothing held, so the first pass now reports how
far that free refinement leaves the two axes apart. A relative split of eps
displaces a reflection at radius r by eps/2 * r pixels: where that displacement
at the far corner of the detector stays inside the integration disc, imposing
the equality moves nothing out of its own box and the higher symmetry is kept
with no extra pass. Where it does not, both hypotheses are run as probe-only
passes - the promoted class, and the class the same walk carries when it is
granted no length equality at all - and whichever realises the lower held-out
positional residual is kept, the promoted class on a tie.
Two small-molecule sets whose axes differ by 2 % now index, refine and merge in
their own orthorhombic lattice instead of a tetragonal mean: one goes from a
cell 1 % wrong and P 1 at CC1/2 0.19 to the deposited cell within 0.2 % and
P 2 2 2 at CC1/2 0.95, the other from a tetragonal mean to a cell matching its
reference to 0.4 %. Protein sets whose symmetry is real (P41212, P4222, I23,
P6422, F4132) are unchanged: their free refinements leave the axes 0.02-0.10 %
apart, a few tenths of a pixel, so the question is never asked. Where it was
asked on a weak sweep whose free refinement diverged, the arms decided for the
higher symmetry and the output was identical.
No new threshold: the comparison is the integration radius the run already
integrates at, and the arms are judged by HeldOutResidualFell, as the geometry
walk's rounds are. postrefine_probe_only_ returns for the arms' sake - a pass
run only to measure stops before the scaling engine is built, so neither arm
pays for a merge or a space-group search.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The translational-NCS vector was refined only on d >= 4 A, the band its
detection gate reads, where an error of 0.02 costs the contrast almost
nothing. The L-test's partner steps and the twin-immune zone's per-class
normalisation then assign a class, cos(2 pi h.u), to reflections at full
resolution, where the phase error is |h| times the error in the vector: past
about |h| = 12 the classes were no longer the physical ones, and on one corpus
crystal they were anti-correlated with them over a whole band.
The vector is now re-refined against all the intensities, up a ladder that
doubles the number of reflections read at each step and starts each step from
the previous one's answer, so the phase is never extrapolated further than it
is known. The objective is the correlation between E^2 and cos(2 pi h.u) -
neither the gate's max/min bin ratio nor the fitted amplitude survives a
full-resolution population, both being ratios that run away where the cosine
has little variance. The gate itself is untouched: which crystals are called
is unchanged, only where the vector points.
Measured as that correlation in bands of |h|, before against after:
0.57/0.61, 0.42/0.52, 0.19/0.30, 0.14/0.18 on one crystal and 0.14/0.52,
-0.15/0.42, 0.00/0.20, -0.02/0.09 on another. The acentric control of the
twin-immune zone moves towards its analytic 0.736 where the classes change
(0.671 -> 0.748 on one), and no space-group or twinning verdict moves.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
Whether a sweep determines its detector distance at all was decided by running the canonical pass
twice - once at the post-refined distance, once at the header's - and keeping whichever realised the
lower held-out residual, the header's on a tie. The first-stage test that sent a run down that arm
compared the free and held fits on the POOLED held-out residual, and pooling is what made it
uninformative: the positional family cannot see the distance along the degenerate direction, so it
dilutes the one family that can, while the unpaired standard error of a heavy-tailed mean of squares
is 2-25 % of the mean against in-fit differences of 0.1-1 %. The test therefore said "cannot tell" on
four fifths of the fitted sweeps and the arm ran on most of a corpus, buying by prior what it could
not measure.
The two hypotheses are now compared on the one residual family that can tell them apart. The
excitation residual never involves the detector, so it is blind to the distance itself; what it sees
is the cell scale, and a held fit at a wrong header distance is forced into a wrong cell scale by the
spot positions, which the rocking angles then refuse - measured on a sweep whose header was 1.4 %
long, the held fit's held-out excitation residual is seventeen times the free fit's. Where freeing
the distance lowers that residual by more than its own standard error the free fit is committed as
before; where it does not, the held fit is committed - header distance, refined beam, cell,
orientation and axis - and the walk and the commit run held. The positional residual is deliberately
not consulted: its in-fit gain along the degenerate direction is the one re-integration erases.
The question is asked only while the run is still at the file's distance. A run that has walked off
the header has already refuted that hypothesis by re-integrating, and asking it again at every round
stalls a walk short of its fixed point, because the excitation standard error at the walk's tail is
outlier-dominated (measured: a walk stopped 0.4 % early, seven passes, cell 0.65 % off against
0.26 %, ISa 10.6 against 14.0).
So there is no arm, and with it go the two probe passes that measured it and the canonical pass
the losing arm used to cost: the decision costs two Ceres solves. On thirteen
rotation sweeps covering both verdicts, every decision the arm took by evidence or by its tie rule is
reproduced at the fit, except where the excitation family sees what the pooled test could not and the
free distance - the better cell against an external reference - is taken instead; merged intensities
are unchanged where the verdict is.
POSTREFINE_DISTANCE_HELD now means "the committed fit held the header distance", and is cleared where
a later geometry walk left it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The full-resolution merge-degradation gate divided the added operators' mean
intensity-weighted R by the smallest R anywhere in the crystal. That denominator is an
extreme order statistic, so it has a pole: one unusually clean operator condemns every
other genuine one. It is not a rare accident. A rotation about an axis near the spindle
maps a reflection onto one recorded at nearly the same detector position, so the lab-frame
systematics cancel for that operator alone; measured over the rotation corpus the spread
WITHIN a genuine group reaches 2.0-2.6x, the whole width of the old bound of 2.0. On a weak
cubic crystal whose near-spindle 3-fold read 0.074 against 0.16-0.22 for its ten other
genuine operators, every one of them was refused, the search kept only the group generated
by the reference operator itself (which reads 1.00 by construction), and the same 222
hypothesis on the same merge read 2.45 under a cubic enumeration and 1.02 under an
orthorhombic one - a statistic that moves with which OTHER operators were enumerated.
The added operators are now placed on the scale the crystal itself defines, between the R
of unrelated reflections and the R of the best-agreeing operator:
contrast = (random_pairing_r - r_added) / (random_pairing_r - global_best_operator_r)
random_pairing_r is new: the same intensity-weighted R over shell-matched pairs of
reflections no symmetry relates, measured on the merge (deterministic, no RNG). Both ends
are hypothesis-free, and the form has no pole - an unusually clean reference widens the
denominator by a few per cent instead of driving the divisor towards zero. The test
abstains where the best-agreeing operator is itself no better than half way to unrelated
reflections, i.e. where the merge holds no clean end to measure from, as it already
abstains when there is no second operator at all.
Calibrated over 149 rotation datasets with a known answer, read as the gate reads it:
genuine promotions reach down to 0.77 and the worst false one reads 0.695, so the bound is
0.72. It keeps every refusal the old ratio made on that corpus except the weak cubic one,
which now promotes to its cubic group; the twins that read above the genuine range are
refused by the H gate and the twin-immune zone, as before.
The reported value, the point-group report line, the finalist ledger and the
gate-fired-and-was-overridden note all follow the contrast.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
A screw axis whose row the sweep never recorded - it lies in the spindle's
blind cone, or outside the resolution range - is not a group the data refused,
it is a question nobody asked. The search already offered the whole set in
SPACE_GROUP_ALTERNATIVES, but the per-zone screw table that would say it in
words is printed for the SELECTED candidate only, and the selected candidate
in exactly this case is the one with no screw zones, so the run's account of
the open axis was a blank.
The search now names the axes on which two SELECTED candidates disagree about
whether the row carries screw absences at all, with why the row could not be
judged (never recorded / no control class). The report writes the axes as
SPACE_GROUP_SCREW_UNDETERMINED= beside SPACE_GROUP_ALTERNATIVES and explains
them in prose; the adoption logs a warning naming the axis and the set. An
enantiomorphic or origin-ambiguous pair predicts the same absences on every
row and is not named here - that ambiguity is the hand, or the origin.
Nothing about the decision moves: the group adopted, the alternatives and the
written .mtz/.cif/.hkl are exactly as before, because a reflection file cannot
hold "maybe a screw".
The battery scorer mirrors its existing "hand only" rule: a set differing from
its reference only by a screw the run reports as undeterminable, with the
reference among the groups it offered, scores unscored/screw_undetermined
instead of a sym_screw failure. All three conditions are necessary, so a screw
called wrongly where the row WAS measured stays a failure.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The zone verdict that arbitrates a refused promotion read its zone absolutely, against the
centric and acentric Wilson expectations, and an absolute reading is only as good as the
normalisation under it. On two refused 622 promotions - a 6/m crystal and a 312 one, whose added
operators disagreed at 4.7x and 5.2x the parent's H and merged to an R_meas of 0.34 under the
higher group - every class read centric: a 67 A^2 anisotropy, isotropically normalised in bins
of 100, spread one shell's expected intensity over a factor of 15 between its directions, and
the acentric control read 1.00 against its 0.74. The zone read the same as the control, +0.10
nats per reflection, and +124 nats over 947 reflections rescued a twin law.
Two things change. The anisotropy is fitted on the acentric reflections of the shells read
(ln I = c + s^T Q s, by least squares) and its deviatoric part taken out of every intensity
before anything is normalised; that alone brings the controls of the crystals measured to
0.73-0.77. What no normalisation removes, the control then certifies: an acentric population
reads -0.130 nats per reflection when the normalisation is right, a twinned one reads below
that, so whatever the control reads above it is the normalisation's - anisotropy, a
pseudo-translation, a pseudo-centring, noise all inflate every class towards centric alike -
and the zone, normalised the same way, carries the same per reflection; the calibrated evidence
has it taken off, and that is what the verdict reads. The two rescued twin laws now read -130
and -78 nats and stay refused, with the law named; the genuine promotions measured read +137 to
+580 (a pseudo-centred orthorhombic crystal whose control reads 0.99 still +334); the trigonal
and hexagonal partial twins -129 to -420 as before. The report prints the control's excess and
the calibrated evidence beside the raw one.
The metric-lattice re-ask carried a second copy of the two-arm rule without the zone: a P31
partial twin whose 32 the main rule had refused on the zone was promoted by that ask's
Lorentz-filtered arm. It now puts the same verdict to a refusal the other arm would outvote.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
A screw's predicted-absent class is one axial row - half a dozen to a few dozen
reflections - and its evidence is a SUM over them, so it is decided by its
largest member. The file's own LIMIT comment said so; 7n2s is that limit firing
on real data. Between two scaling passes that differed only in which weak frames
were rejected, one of eight dead 0k0 moved from 14 +- 9 to 99 +- 10 while the
other seven did not move at all, and the zone fell from 30.1 nats to 17.1 and
lost the 2(1) under a bound of 20. That reflection was never measured to the
precision its sigma claimed: its two half-set merges read 198 and 2.5.
The zone's sum is now taken with its single largest member dropped and rescaled
for the trim - divided by n - H_n, the expected sum of the other n-1 under the
null, and multiplied back by n. ScrewZoneEvidence reads the result exactly as
before: same statistic, same floor, same bound, same calibration, with a robust
estimate of the zone's deadness in place of a fragile one. One member only,
whatever the zone holds: a zone with two strong absences is a zone that is not
extinct. On a uniformly dead zone the rescale under-states by 2.2 nats at eight
absences and 3.9 at sixty-four - it only ever refuses, never claims. Glide zones
keep the untrimmed sum: a plane holds hundreds to thousands of reflections and
no single one can carry the verdict.
7n2s -> P 1 21 1 (zone 27.6 nats, set by the seven reflections that did not
move), matching its deposit; a second monoclinic crystal decided six nats under
the bound (7 absent, 1 violation, 13.9 nats) reaches 21.9 and its 2(1) as well.
Unchanged on 7mzt, 7k1l, 11if, 9hs7, 9zlo and four in-house reference sets.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
Both arrays are one record per integrated observation - tens of millions on a fine-sliced long axis,
gigabytes each - and both carry the raw hkl only to sort and group on. A Miller index needs sixteen
bits (|h| <= a / d_min, in the hundreds even on the longest axis at atomic resolution), which takes
the ingest sort key from 24 to 20 bytes and the post-refine partial from 32 to 28.
Same comparisons, same order; merged output byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
Three follow-ups to the converging scale loop, all on sweeps with a stretch the crystal barely
diffracted on:
- The sweep ledger's scale channel (and with it delta-CC1/2's "normal frame" test) and the
space-group search's scale floor measure a frame against the precision-weighted typical frame
G_ref (TypicalFrameScale) instead of the run median: on a sweep that spent most of its turn out
of the beam the median frame is itself a dead one, and every stretch then reads as typical. The
two collapse guards stay on the median on purpose - a frame that collapsed toward zero and was
not quite dropped has its fulls re-fitted with a scale of 1/G, a cubic mean is then theirs, and
the floor read against it dropped three quarters of every live frame's fulls (measured).
- delta-CC1/2's sigma-tau statistic enters each reflection with the information it carries (the
same per-frame factor as the CC1/2 weight), so a dead stretch, whose scaled-up noise flooded the
mean error variance of every mixed reflection and read as harm up to the 25% cap, now costs
about nothing and is left to the ledger.
- The scale loop's step test weighs each frame by its merge weight (its observations at its scale
squared): a dead frame's scale is fitted on noise and wanders by orders of magnitude every
iteration, carries nothing into the merge, is dropped after the loop, and must not hold the loop
open.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The device merge's per-group sums were downloaded into a full-length host array per field and then
unpacked into the per-group accumulators, so every merge held two full copies of them - on a large
cell searched in P1 that is most of a gigabyte beside the accumulators themselves. MergeAccum now
leaves the sums on the device and MergeAccumRange downloads a million groups at a time, each slice
unpacked straight into the accumulators.
Same values, same integer reject count; merged output byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The joint post-refinement is asked twice, with the distance free and with it held at
the header. Where freeing it lowers the fit's held-out residual by more than that
residual's own standard error, the free fit is committed exactly as before. Where it
does not, the fit cannot tell - its own gain need not survive re-integration - so
both hypotheses are carried: the canonical pass runs at the free geometry as before
and once more at the header distance, and the run keeps the one whose RE-INTEGRATED
held-out residual is lower, the header's on a tie. The report says which
(POSTREFINE_DISTANCE_HELD), the pass decision names both realised residuals, and the
log prints the free fit's formal sd(distance) and its correlation with the cell
lengths, from the covariance of the solution.
Why: at a detector far enough away that no reflection reaches more than a few
degrees of 2theta, a longer distance and a larger cell move every spot the same way
to first order (the difference is of order sin^2 theta of the spot's position -
0.4 px rms per per cent over a 2M detector at 820 mm, against 2-3 px at usual
distances), so the fit finds a distance/cell pair that fits its own observations a
little better than the header, commits it, and the pass re-integrated there asks for
the next pair: a walk along the degenerate direction the realised residual never
ratifies. XDS leaves the distance out of IDXREF by default and its documentation
says to remove it from CORRECT where it drifts on low-resolution data; DIALS fixes
the wavelength for the same reason. Here the data are asked instead of a rule.
Measured: a sweep at 820 mm reaching 2theta ~ 9 deg no longer walks to 846 mm with
the cell 3.3 % too large (pool-B stopped it at 825.6 mm, still +1.1 %); it keeps
820 mm and its cell agrees with the reference at that distance to 0.2 %, R_meas
101 % -> 70 %, two passes fewer. 9yl4 keeps 600 mm (cell 0.5 % -> 0.1 % from the
deposition), 7mzt keeps 400 mm (R_meas 169 % -> 145 %, CC1/2 0.961 -> 0.986), 5lzl
keeps 689 mm on a tie (cell -0.3 % instead of +0.25 %, ISa 13.9 -> 11.6). Merged
output byte-identical on a lysozyme reference sweep at 110 mm (the free distance
wins its realised comparison), 8egn and 8pqd (decisive in-fit gains, no second arm)
and 8qq7 (the quality guard had already reverted it); one extra canonical pass
wherever the two arms are run.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
A promotion the operator correlations confirm and a gate refuses (the H ratio, the added-operator
R, the merge chi^2) used to be settled two ways that both read the per-frame scales: the two-arm
rule let the other arm's confirmation outvote the refusal, and the remerge arbiter compared the
R_meas and ISa of the two pinned merges. Neither holds once the scales have settled: a partial
twin at a fraction one arm's coverage hides passes there (an H3 twin promoted to R32, a P31 one to
P3121), and the lower group's converged merge - the same scales fitted against fewer equivalents -
agrees with itself better than the true group on a genuine step (a 2 -> 222 step reading R_meas
0.094 -> 0.103), so the arbiter refuses real symmetry.
The refusal is now decided by the twin-immune zone of the operators the promotion adds: the
reflections centric in the higher group and acentric in the lower are their own twin mates, so
they read centric if the operators are real and acentric if they are a twin law or a
pseudo-symmetry, whatever the twin fraction and however the scales were fitted. Read on the
all-observation P1 merge in the two-arm rule and on the adopted group's own merge at the arbiter
(where a weak crystal's P1 search merge has no shell with signal to read), for a group two orders
above the other, with the pseudo-translation normalisation the L-test uses, and decided on the
sign with the screw-axis convention of 20 nats: acentric and the refusal stands on all
observations, centric and the higher group is adopted, undecided and the rules stay as they were.
Measured: a genuine 222 step reads +150 to +1630 nats, the trigonal twins -130 nats and below.
A refusal that stood on the zone names the law: the first operator the refused group adds, at the
fraction the operator's own H implies (else the pre-search L-test), reported as TWIN_LAW and
written as the mmCIF _pdbx_reflns_twin loop of the merge in the true group. Every lattice twin law
outside the adopted group is reported with its own H, implied fraction, R and correlation on the
merge expanded to P1 (TWIN_LAW_n), following Yeates (1997) Methods Enzymol. 276, 344-358.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The whole-run passes retain every frame's integrated reflections until scaling is done - thousands
of vectors of a few megabytes each, allocated by the image workers in the allocator's per-thread
arenas. When a pass hands them back, most of that memory stays in those arenas as holes, and the
next pass's workers (new threads) do not reuse it, so on a fine-sliced long axis gigabytes of freed
reflections were carried to the end of the run.
IndexAndRefine now copies each retained frame's reflections into a ReflectionArena: 64 MiB blocks,
each its own mapping, carved by a bump pointer and returned to the system in one piece when the
last vector in them is gone. IntegrationOutcome::reflections becomes a std::vector with an allocator
that uses the arena when given one and plain new/delete otherwise (copies go to the heap), so the
read sites are unchanged; the few functions that took the vector by type now take a span.
No arithmetic changes; merged output byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
Pooled run B: twin report-only evidence and free-R on the lattice holohedry,
information-weighted resolution cutoff with weighted outlier median and ISa
rework, sparse-lattice integration and beam-centre arbiter, GPU/CPU collapsed
scale upload, geometry walk on the realised residual, --model speed-ups with
FFTW structure factors, memory reductions for large-cell sweeps, -A decisions
Friedel-merged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The alternating per-frame scaling used to run a fixed three rounds of a Cauchy-reweighted fit
against a reference that included half a sweep's rocking-curve tails. Three rounds left a short
sweep merged in P1 far from its answer (a 90 deg tetragonal sweep refused its 422 with the P1
scales anti-correlated with the converged ones), and more rounds did not help: the fit had no
fixed point. Two things made it walk. The objective is invariant under G -> cG with the reference
-> reference/c, so every round moved every scale by a constant factor; and the robust loss, iterated
against a reference refitted each round, drops the strong reflections of a frame whose scale is off
by a third (ten-sigma residuals) and lets the weak ones carry it further off - measured on a 360 deg
sweep the scales shrank 10-30% per round for thirty rounds and the H ratio of a genuine 222 read 21x.
Now the loop pins its gauge every iteration (G divided by the precision-weighted typical frame
scale G_ref = sum G^3 / sum G^2, one definition shared with the CC1/2 weight), fits the plain
weighted least-squares slope with the weights the reference uses and only on the observations the
reference is built from (the partiality floor), and stops when the rms |log(G_new/G_old)| over the
frames falls below 1e-3. With the same weights on both sides the alternating fit is exact
coordinate descent on one objective and cannot climb; measured, the 360 deg sweep settles in 19
rounds and the 90 deg one in 30-50, each pass's partials and fulls loops alike.
--scaling-iterations is now the cap (default 100). A loop that reaches it is logged, the report
prints SCALING_ITERATIONS and raises SCALING_NOT_CONVERGED, and the correction surfaces run to the
same tolerance under their own cap. On the GPU the loop runs one iteration per call so the pin and
the step test read the same numbers as on the host; the fulls' reset is split out of ScaleFulls.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The ingest-time geometry smoothing (delta_phi from the smoothed lattice, one exact-Bragg angle per
rocking event, partiality from the smoothed mosaicity) used to run over a 32-byte host record per
observation, copied from the source reflections and held beside them. It now reads hkl, frame, d
and zeta straight from the source reflections, frame by frame, and keeps arrays only for what it
rewrites (delta_phi, partiality) and what the rocking-event walks read in raw-hkl order
(image_number, and on the resident path the usability flag): 13 bytes an observation instead of 32.
The full-Obs path runs the same code and copies the two rewritten fields back.
Same arithmetic in the same order; merged output byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
None of this has been built on a Mac - there is none yet. It is the list a read-only audit of the
viewer/rugnux subtree produced, plus a serial -fsyntax-only pass of every reachable .cpp with
clang 16 + libc++ on Linux, which found exactly one error (the first item).
- JFJochDatasetInfoChartView: std::vector<fftwf_complex> does not compile with libc++, whose
construct_at refuses an array element type (float[2]). Use std::vector<std::complex<float>> and
the reinterpret_cast every other FFTW call site already uses.
- libcurl: GSSAPI off on Linux, and neither TLS nor GSSAPI on macOS. The viewer never sets
CURLOPT_HTTPAUTH, so Negotiate was dead weight that cost a krb5-devel build dependency; on macOS
curl's FindGSS refuses the system Heimdal outright, and with Secure Transport gone from curl
(8.15) TLS would mean a Homebrew OpenSSL - the only host library a Mac build would need.
Linux keeps OpenSSL. CURL_USE_GSSAPI is forced OFF rather than left unset so an existing build
tree drops its cached ON.
- libjpeg-turbo ExternalProject: CMAKE_SYSTEM_NAME/PROCESSOR were forwarded unconditionally, which
puts even a native sub-build into cross-compiling mode, and CMAKE_OSX_ARCHITECTURES / SYSROOT /
DEPLOYMENT_TARGET were not forwarded at all. Now the same rule the zlib-ng sub-build follows.
- ShadowAccumulatorGPU.cu was added on the JFJOCH_USE_CUDA option (default ON) instead of
JFJOCH_CUDA_AVAILABLE like every other .cu, so a machine without nvcc got a CUDA source in a
target with no CUDA language.
- CMAKE_OSX_DEPLOYMENT_TARGET defaults to 12.0 (overridable). Left unset, CMake takes the build
machine's OS version and the .dmg starts nowhere older.
- Standard headers that were only arriving transitively (<chrono>, <cmath>, <limits>, <cstring>,
<atomic>, <thread>, <string>); newer libc++ releases keep removing such transitive includes.
Checked: the seven changed sources pass clang 16 + libc++ -fsyntax-only. The CMake changes are
not configured or built.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015eAE2K7i5JGDwgwifiCfuA
(cherry picked from commit 5a7282759a)
On the resident (GPU) path the ingest's narrow host record sits on top of the source reflections
for the whole smoothing and partiality recompute, and on a long axis that moment is the run's
memory high-water mark. Store its Miller indices in 16 bits and move the rocking-event flag into
their padding: 40 -> 32 bytes a partial. Stage the device upload in slices of two million
observations rather than eight, which takes the thirteen staging arrays from ~400 MB to ~100 MB.
Neither changes a value.
Measured on a 3600-frame long-axis rotation set (-N 6): VmHWM 13.79 -> 12.80 GiB on top of the
previous commits (15.0-15.1 GiB before any of them); merged MTZ/HKL/CIF, P1 MTZ, per-image table and
the report are byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
Ingest wrote every observation's 24-byte SortKey into an index-order array and then scattered that
array into its h buckets - two key arrays alive together at the top of the ingest. Count the
buckets per chunk of frames straight from the source reflections and scatter from them instead:
a frame's reflections are consecutive in the key numbering, so the chunks are consecutive index
ranges and each bucket lands in the same index order as before. The key array is byte-identical.
BuildInRangeObservations also hands back the remap temporaries (old runs, per-run maps, new_idx)
before the observation array is built rather than at the end of the function.
Measured on a 3600-frame long-axis rotation set (95.7 M partials in the geometry pre-pass): the
pre-pass ingest high-water drops 15.45 -> 14.1 GB (sampled RSS, -N 6); merged MTZ/HKL/CIF, the P1
cross-check MTZ and the per-image table are byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
Finalize reserved npredicted reflections and kept only the ones whose fit succeeded. On a dense
long-axis pattern more than half the predicted reflections lose their background ring, so each
retained per-image vector carried more dead capacity than data for the rest of the pass. Count the
kept ones first and reserve exactly that. No result changes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
The own statistics table's grid bottoms 0.1% below the finest kept reflection,
and the observation-level counts (N_obs, R_meas, the Bijvoet split behind SigAno
and CCanom) re-walk the fulls, which still hold every ingested group. Groups in
that band just below the auto cut were erased from the merge, so they were not
in N_uniq, but their observations still counted in N_obs and R_meas of the
finest shell. The own table now floors those counts at the cut by group d, the
rule the erase applies - the same floor the reference-range table already used
(now passed as a double, so the comparison is exactly the erase's).
Report-only: on an auto-cut cubic in-house set the written MTZ, HKL, P1 MTZ, unmerged
MTZ and image table are byte-identical; TOTAL_OBSERVATIONS 378754 -> 377933,
MULTIPLICITY 40.31 -> 40.22 (now equal to the reference-range table's), finest
shell N_obs 60873 -> 60052, R_meas 1136.9% -> 1129.2%, CCanom -0.6% -> -0.3%;
overall R_MEAS, CC1/2, SigAno unchanged at printed precision; the mmCIF's
pdbx_number_measured_all / pdbx_redundancy / last shell row follow. A
detector-limited lysozyme run (no auto cut) is byte-identical throughout.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The rot3d merge tests each observation against the weighted median of its
reflection, and only from three observations up: at two the median is one of
them. With -A each hand of a Friedel pair is its own reflection, so a P1
anomalous merge sits at about two observations a hand and almost nothing is
tested - on one P1 sweep the merge rejected 0 observations with -A against
tens of thousands without it. A single observation a hundred times its
partner's intensity (300 sigma apart, a 3.9 A reflection) then survived into
the merge, carried 94% of the weighted variance of its CC1/2 bin, took that
bin from 0.80 to 0.18, and the automatic cut with it: 3.46 A and UNUSABLE,
where the same data without -A cut at 1.40 A.
A hand with fewer than three observations now takes the weighted median of
both hands together, where the pair has three. I(+) and I(-) differ by the
anomalous signal, a few percent of I and far inside the six-sigma test, so
the mate supplies the observations the hand is missing. A hand with three of
its own keeps its own median exactly as before, and a Friedel-merged run is
unchanged. The medians are formed on the host, so the device merge reads the
same ones.
Measured with -A: the P1 sweep 3.46 A / UNUSABLE -> 1.49 A (1.63 A XDS
reference; 1.40 A without -A, unchanged), 0 -> 3,699 rejections. Two lysozyme
and one insulin set: same cut, space group, CC1/2 and ISa. On one of the
lysozyme sets 16 more observations are rejected, every shell's CCanom is
unchanged, and the overall CCanom goes 0.31 -> -0.02: the whole-range figure
was being carried by a handful of wild pairs, while the shells ranged -18% to
+12% and XDS's overall anomalous correlation is 1%.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
Comparing a run with another program's table has meant running rugnux AT that
program's resolution range (--scaling-high-resolution), which is a different
run: the range moves the cut, the space-group decision and everything after
them, so the comparison buys itself a different answer. --report-resolution
<dmin>[,<dmax>] instead leaves the run alone and adds a second table to
section 3 of the report - the REFRES_* keys and a shell table - binned from the
same merged reflections over the range given, with the completeness
denominator enumerated over that range and the shells in equal steps of 1/d^2
so they read row for row against a CORRECT.LP at the same range. Report-only:
the merged files and every decision are byte-identical with and without it.
The table holds only what the run kept. Where the reference range is finer
than the run's own limit, the shells past it are printed as not merged (with
their possible count) rather than as zeros, REFRES_SHELLS_PAST_LIMIT counts
them so a consumer can tell "not merged" from a measured zero, REFRES_
COMPLETENESS counts their reflections as missing, and the other overall numbers
are over the shells the run reached; nothing is read from the observations the
run judged to carry no signal. REFRES_ISA is the error model refitted on the
reflections of the table alone, in XDS's convention (rotation only; the stills
model is fitted over the whole range already).
On the rotation path the statistics block of MergeAndStats becomes a lambda
over a shell grid, called once for the run's own grid and once for the
reference one; the reference call floors every observation-level count at the
cut by group d, the rule the erase applied. The stills MergeStats takes a
declared range, whose bounds are the grid's whether or not any reflection
reaches them. Both --mode mx and --mode scale report it, the viewer's command
line echoes it, and the docs describe the keys.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The per-shell SigAno and CCanom of the rotation merge were accumulated on the
report grid, but the overall pair of numbers - the SIGANO / CC_ANOM keys and
the mmCIF's pdbx_absDiff_over_sigma_anomalous - were summed over every Bijvoet
pair the fulls hold, including the ones past the automatic resolution cut that
the written reflections do not contain. On a run the cut trims, the overall
line therefore described more data than the shells above it add up to. Both
overall numbers now count only the pairs that land in a shell, as the rest of
the table does; a run with a manual limit, whose ingest already ends at the
limit, is unchanged.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
FFTW's planner (every fftwf_plan_* and fftwf_destroy_plan) shares global state
and is not thread-safe; executing a plan is. FFTIndexerCPU and BeamCenterFFTCPU
each guarded their planning with a lock of their own, TranslationalNCS and the
viewer's spectrum with none, and ModelFFT with a third - which does not stop
two of them planning at once. common/FFTWPlannerLock.h holds the one mutex they
all now take. No numerical change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT
A long-running broker began cancelling every data collection with
Device decoding failed (CUDA (GPU) error (out of memory)), falling back to host decompression
CUDA (GPU) error (out of memory)
while nvidia-smi showed the cards less than a fifth full. Two defects, both in CUDAMemHelpers.h and
both from the pooled allocator (cudaMallocAsync) that came with rc.162.
The leak. cuda_allocation_stream() kept one stream per (thread, device) in a thread_local map of raw
cudaStream_t and never destroyed them. That was written against rugnux, where the worker threads live
as long as the process. The broker starts fresh std::async threads for every data collection - 16 or
64 of them - so every collection left that many streams behind. Measured: 0.56 MB of device memory
per leaked stream, linear to 4928 streams, never returned, with nothing on the host side growing.
The RAII wrapper (CudaStream) was there but not used at this site, and using it as-is - a stream
destroyed when its thread exits - would not have been safe: ShadowFinder builds its GPU accumulator
on a throw-away std::async thread and frees it from another thread long after, and that free is
ordered on the allocating thread's stream. So the streams are still never destroyed, but a thread
now only borrows one: CudaStream objects live in a process-wide per-device idle list, a thread takes
one on first use and hands it back when it exits. Their number is bounded by the threads that were
ever alive at once instead of by the threads ever started. The list itself is deliberately leaked, so
that nothing calls into CUDA during static destruction.
Replaying the broker's pattern against the real header, 60 collections of 64 threads:
before 3840 streams, 260 -> 2424 MB of device memory
after 64 streams, 260 -> 358 MB
The stale error. Every helper here throws a named message, yet the log carried the raw CUDA string,
so the failure came through a cuda_err() and not from an allocation. CudaDevicePtr falls back to
cudaMalloc when cudaMallocAsync fails, silently - but the failed call stays behind as the thread's
last error, and the cudaGetLastError() that follows the next kernel launch reports it. The buffers
were all allocated; the frame was lost anyway, once on the device-decode route (caught, hence the
warning) and once more on the host fallback (fatal). The pooled attempt failing, and a stream that
cannot be created, are both handled by falling back, so both now clear the error they leave.
What finite resource the production cards ran out of at under 4 GB used was not established - no
cap on the number of streams was found up to 4928 on the card this was measured on. The leak is the
only thing on this path that grows with uptime.
tests/CUDAMemHelpersTest.cpp: later threads end up on the same stream, concurrent threads on
different ones, a buffer is freed cleanly after its allocating thread has exited (and another has
borrowed its stream), and a pool that cannot serve a request leaves no error behind - the last by
capping a memory pool at 4 MB so that the pooled attempt fails and the fallback succeeds.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Each post-refinement now records, beside what it commits, what the integration
it was handed realises at the geometry it started from: the held-out residual
at nominal (the same value the log prints on the left of "held-out ... ->"),
its standard error over the held-out residual values, and the cell the pass's
own indexing refined. Nothing here changes a decision.
Two helpers read them: HeldOutResidualFell (a pass at a new geometry realises a
lower residual than another by more than the standard error of the difference)
and ReindexPushesCellBack (re-indexing at a committed geometry returns, on the
length the fit moved most, a cell on the side the fit moved away from). The
geometry walk in RunAllPasses uses both in the next commit.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
DropCollapsedFullScales zeroes the corr of fulls whose frame scale collapsed on
the host copy of the fulls; on the resident-fulls (GPU) path that zero reached
the device only with the corrected corr pushed after the surfaces. The merge
that measures CC1/2 before the correction surfaces runs in between, so on the
GPU it still merged the dropped frames' fulls while the host path excluded
them - the two builds measured different observations.
The corr is now pushed to the device right after the drop, and the later push
covers only the surfaces and the frame rejection.
This changes the input of the two-pass quality guard (cc_half_before_corrections)
on GPU runs where a full scale collapsed; the final merge is unchanged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A weak crystal inside a powder of its own microcrystals plus ice puts only
~5% of each frame's spots on its lattice. The pooled first-pass test accepted
that lattice (5.5% of validation spots against 0.6% at a wrong spindle angle),
but then:
- every frame failed the 20% per-frame floor (LATTICE_MIN_INDEXED_FRACTION),
so nothing was integrated and the merge was skipped (indexing rate 0);
- the background-measured beam centre, ~3 px from the file's and correct,
was not adopted because the arbiter counts validation frames, which are 0/60
at every centre.
Changes:
- When fewer than 1/6 of the validation frames clear the per-frame floor AND
the sweep's pooled on-lattice fraction is itself below that floor, integrate
every frame from the sweep's lattice (as XDS/DIALS do) and leave frame
selection to scaling. Sweeps sparse only in spots per frame keep the floor.
- Merge when a rotation lattice was found even if no frame "indexed" on its own.
- When neither centre indexes a validation frame, adopt the measured centre if
its pooled excess over chance beats the file's by 3.29 sigma.
On the case above: P2 at the XDS cell, merged to 1.77 A, CC1/2 0.95, ISa 3.4
(XDS with the same lattice: ISa 3.3-5.5). Three healthy/partially-indexing
rotation sets are bit-identical; a two-wavelength CBF set that currently
merges 42 frames on a wrong cell moves (beam centre adopted, 755 frames).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The (a, b) error-model fit already refuses to report ISa when its top intensity
bin has no leverage on b. It still printed 1/b when b had leverage but came out
at zero within its own error - the strongest bins scatter no more than
counting statistics say (the fitted b^2 is negative and clamps to 0), or there
are too few samples for the bin medians to mean anything. ISa then reports the
noise in b as an I/sigma: on one low-resolution sweep consecutive merges of the
same data gave ISa 0, 28, 46, 113 and 130.
The standard error of b^2 is taken from the 16 bins' own scatter about the
fitted line, and when b^2 is less than two standard errors above zero the
merge result carries isa_resolved = false. Only what is printed follows it:
the report's ISA key reads "undetermined", the summary line says so, the
asymptote is not printed and the mmCIF carries "?". The fitted ISa itself is
unchanged and is still what the space-group search's present-reflection cut
and the refused-point-group arbitration read, so no decision moves.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The resolution cutoff, the shell table and the overall CC1/2 counted every
unique reflection equally. The merge itself is inverse-variance weighted, so
an observation from a frame the crystal barely diffracted on enters it at
1/G^2 of a good one - honestly, with its sigma - but a reflection measured
only on such frames is scaled-up noise that then counts as much as a
well-measured pair in every Pearson CC1/2 read off the merge. On a sweep
where half the frames are weak the curve collapses at every resolution: the
cut lands at 3.7 A on a P1 crystal whose good frames reach 1.5 A, and the
UNUSABLE verdict fires (CC1/2 0.40 beside I/sigma 10).
Each merged reflection now carries cc_weight: the precision its half-sets
would have had with every observation at the run's typical frame scale, over
the precision they have. G_ref = sum G^3 / sum G^2 over the usable
observations is the precision-weighted typical scale, which the dead frames
cannot drag down however many there are; the factor per observation is
max(1, (G_ref/G)^2), with G the frame's total scale (partial scale, flux and
the fulls' own G) taken before the correction surfaces and before collapsed
frames are dropped, so nothing intensity- or resolution-dependent enters it.
On a sweep without a weak stretch every weight is 1 and the CC1/2 is the
plain Pearson it was. The cutoff fit, the shell table and the overall CC1/2
(and so the UNUSABLE verdict and the report's shell checks) all read the same
weighted statistic.
The two extra per-group sums are accumulated on both merge paths, the host
loop and MergeAccumKernel, from one per-frame factor array; a host recompute
of the device sums agrees to 1e-15 relative on every merge after the
corrected corr is uploaded, and a host-merge run gives the same cut, space
group, CC1/2 and ISa on five sets.
Weighting by the half-set error variance alone (1/(v0+v1)) is not this: v
grows with the intensity, so it weights the weak end of the intensity
distribution and biases homogeneous data coarser.
Measured (written resolution, together with the weighted outlier median):
a P1 sweep with a long weak stretch 3.73 -> 1.40 A against a 1.63 A XDS
reference, UNUSABLE withdrawn (overall CC1/2 0.40 -> 0.97); a second 3.61 ->
3.00 A; one with most of the sweep out of beam 7.35 -> 5.20 A. Lysozyme,
thaumatin and two insulin sets unchanged to 0.01 A (one lysozyme sweep with a
weak wedge 1.13 -> 1.16 A, from the median), same space groups throughout.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The per-reflection median the merge rejects outliers against was an
unweighted median of the scaled intensities. On a sweep where the crystal
barely diffracts over a long stretch, observations from those frames are
scaled up by 1/G together with their sigmas, and in P1 at multiplicity ~3
two of them outvote one well-measured observation: the median becomes
noise and the well-measured observation is rejected against its own small
sigma. On one P1 sweep with ~44% such frames this rejected 52,574
observations (2,131 once those frames are dropped) and took the
correlation of the merged intensities with an external reference from
0.92 to 0.68 at low resolution. The median is now weighted by
1/(sigma*corr)^2, which leaves it unchanged where the observations are
comparably precise: rejections 3,932, correlation 0.92.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Under a twin law T a reflection whose twin mate is itself (up to the true
group and Friedel) is untouched at every twin fraction. For each index-2
subgroup H of the adopted point group, those are the reflections centric in
the group and acentric in H: centric if the operators the group adds over H
are real, acentric if they are a twin law or a pseudo-symmetry - the one
intensity statistic that still separates the two at fraction 0.5, where
every operator statistic reads "real".
Read on the P1 cross-check merge: epsilon-1 reflections in shells with
<I/sigma> >= 5 (noise inflates every class towards centric), each class
normalised against its own mean in bins of ~100 reflections and within the
two phase classes of a detected pseudo-translation, Wilson outliers above
E^2 = 20 dropped. Reported per subgroup: the added operators, n,
<|E^2-1|> +- SE read absolutely against 0.968 / 0.736, and the centric-over-
acentric Wilson log-likelihood in nats with both densities convolved with
each reflection's error (flooring E^2 at its sigma instead read a genuine
1.2 A lysozyme zone as acentric), with the acentric control beside it.
Log, report prose and TWIN_ZONE_n keys. Nothing reads it back; no decision
changes.
On the reference sets: 6toc P4222, all three subgroups centric (1.04-1.07,
+253 to +534 nats); 6iu9 P3121 over P31 acentric (0.789, -294 nats); 5j23
R32 over R3 acentric (0.805, -304 nats, tNCS-class normalised); lysozyme
P41212 centric (0.90, +476 to +1297); a P21 myoglobin centric (0.930, +135).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The free-set hash was keyed on the Laue ASU of the merging group, so a twin
law - a lattice symmetry the crystal lacks - put nearly every free
reflection's twin mate in the working set (measured: 97-98% of the
free-touching twin pairs mixed, on the main P3121 output of 6iu9 and on
every law of the 6toc P1 cross-check), and each file of one crystal carried
a different free set (P1 vs merged agreed on 88-92% of reflections).
With the cell given, the key is now the reflection's orbit, Friedel mate
included, under the lattice holohedry: the metric point group of the cell
(gemmi Le Page two-folds, 3 deg obliquity, lattice of the cell's own basis
vectors so every file gets the same group; this contains the merging
group). Twin mates share a flag, and the merged MTZ/mmCIF, the P1
cross-check and any subgroup re-merge carry one free set (nested where the
small-data floor lifts the fraction of one file more than another). This is
phenix.refine's default (use_lattice_symmetry). Where the cell does not
carry the merging group, the merging group's key is used as before. The
small-data floor still counts reflections of the merging group, so the free
fraction is unchanged. Reference free sets are untouched.
On 6toc, 6iu9, 5j23: mixed twin pairs 0 for every law; P1 and merged flags
agree on 100% (6iu9, 5j23) and nested on 6toc; free fraction 0.050-0.051 as
before (6toc merged 0.086, floor unchanged). Intensities and space groups
identical to rc171.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The L-test now divides each intensity by its resolution-shell mean before
forming pairs, over the shells whose <I/sigma> reaches 1 (the floor the
second moment already used). Two index steps are not the same resolution,
and on a small cell with a steep fall-off the raw pairs read <|L|> up to
0.08 high: an untwinned crystal 0.568 and a partial twin 0.427 where
phenix.xtriage reads 0.486 and 0.360. Normalised, the four reference merges
(6toc, 6iu9, 5j23, a lysozyme) agree with xtriage within 0.02 (0.487,
0.376, 0.374, 0.480). Selection stays by shell, never by the reflection's
own I/sigma, which biases <|L|> down. Equal-count shells were tried for
the second moment and widen its gap to xtriage, so the shells are kept.
The L-test is no longer switched off in holohedral Laue classes. Merging
I(h) with I(Th) under a false operator gives (I(h)+I(Th))/2 for every twin
fraction, the perfect-twin distribution, so <|L|> < 0.42 there is reported
as an adopted operator averaging unequal intensities (promotion suspect, or
a twin law absorbed into the point group) - a warning, not a veto. Over 42
holohedral merges of the corpus the genuine ones read 0.436-0.513, the
over-promoted H3 twin 0.374, and one ~490 A-axis crystal 0.365.
The verdict is one line (TwinningVerdictLine) used by the stats text, the
report prose, the summary row, the warning and the viewer; a new
TWINNING_VERDICT key names it. The twin fraction comes from the statistic
that carries the verdict (the L-test unless the call rests on the second
moment alone), is not quoted from a second moment under a detected
pseudo-translation, and is not quoted at all on a merge under a suspect
operator.
The pre-search numbers (measured on the P1 search merge) are user-visible,
are measured with that merge's own pseudo-translation declared, and leave
out the reflections the lattice centring extinguishes - an R lattice in
its hexagonal cell otherwise pairs present with absent reflections (5j23:
0.633 with them, 0.40 without).
Report-only: no space-group decision, merge or intensity changes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Rugnux: basic support for CCD images (marCCD, SMV) and for gzipped miniCBF.
* `jfjoch_viewer`: opens the CCD formats, and fixes to the dataset plots.
* Documentation updates.
Reviewed-on: #81
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
* Fixed a `jfjoch_broker` crash during indexing: sorting no longer misbehaves on non-finite values, and GPU FFT indexer kernel launches are now error-checked.
* rugnux needs about a third less peak memory to scale, merge and post-refine rotation data, with identical results.
* `rugnux --model`: the placed coordinate file carries the space group its own coordinates obey, and says so when that is not the group the reflection files beside it carry.
* `jfjoch_viewer`: fixes in the dataset plots, inspector and layout; spot markers lose their black outline by default (a checkbox under "Image features" restores it) and the highest-pixel markers are white boxes around the pixel.
Reviewed-on: #80
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>