- set_gpu_blocking_sync(): every device is put in
cudaDeviceScheduleBlockingSync before its context exists, so a host thread
waiting on the GPU sleeps instead of spinning on a core. On a 16M rotation
run a fifth of all CPU time was that spinning; wall time unchanged within
noise. Called first thing in rugnux.
- enable_gpu_numa_binding(): from then on pin_gpu() (and the new
pin_gpu(dev), used by the first-pass spot workers that take a card by
index) also keeps the thread on the CPUs of the NUMA node the card hangs
off. The node and its CPUs come from /sys (no libnuma), intersected with
the process's own mask; Linux only, and nothing happens on a machine with a
single node. rugnux turns it on; the broker does not.
- A thread inherits its creator's affinity, so the shared ParallelFor pool
would run every later pass on one socket if a pinned worker created it:
its threads now reset to the mask the process started with
(common/ThreadAffinity).
Byte-identical output. The NUMA part is a no-op on the single-node test box
and still has to be measured on a two-socket machine.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The rotation two-pass runs several first-pass indexing probes one after the
other although some of them do not depend on each other:
- The rotation-scale walk always indexes at the stored angles and then at
the pass-1 fit; both are known before it starts. The probe at the fit now
runs at the same time, on a copy of the run as it stands before either.
Its answer is taken only if the probe at the stored angles left behind
none of the state the next pass reads (the spot-finding settings a
first-pass ladder rung adopts; each probe puts the experiment back
itself) - otherwise it is run again in sequence, as before.
- The geometry walk's indexing probe at the canonical pass's post-refined
geometry depends on nothing that pass does after its post-refinement. It
is now started there, on a copy, and runs beside the pass's scaling, merge
and output; its evidence comes back through the first-pass memo (now a
short list) and the walk's own probe takes it through the usual key check,
so RUGNUX_VERIFY_FIRST_PASS_MEMO covers it too. It is started only for the
passes whose post-refinement the walk probes with an indexing pass.
The copies are plain copies of Rugnux: the cancel flag becomes a shared
flag (so cancelling the run cancels them) and everything else was already
copyable. Probe passes run on a copy get no observer.
md5-identical output on two 16M sets; the walk's two probes take ~2.3 s
together instead of ~3 s, and the last probe of the geometry walk is hidden
behind the merge.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Two whole-dataset sorts sat on the main thread at the end of a rotation run:
WilsonOutliers orders every full by resolution, and the unmerged MTZ is put
in H K L M/ISYM BATCH order by Mtz::sort(5) - together about 2 s of one
thread on a 1.4 M-observation set.
ParallelSort (common/ParallelFor.h) sorts one piece per worker and merges
them pairwise. It is only for comparators that are a strict total order,
where the sorted sequence is unique and the result is the serial sort's bit
for bit: WilsonOutliers already breaks ties on the index, and the unmerged
writer now sorts the rows itself on the five key columns and then the row
number - the order Mtz::sort's stable sort gives - and sets sort_order as it
did. An empty table still fails the way Mtz::sort does.
md5-identical p.hkl, p.mtz and p_unmerged.mtz; 44.4 -> 43.6 s on a 16M set.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Where a pseudo-translation is detected, the class vector is re-refined up a
ladder on the whole data by a greedy sweep over the 26 neighbours of the
current vector, each a cosine correlation over up to every acentric
reflection - seconds of one thread on the main path of the run (3.4 s and
1.5 s for the two calls on a P3_121 16M set).
The neighbours of the current vector are now evaluated together; the first
improvement in the sweep's own order is taken, and the neighbours after it
are evaluated again about the new vector - exactly the moves the serial
sweep makes. Each correlation is still one serial sum, so the result is
bit-identical. AnalyzeTranslationalNCS takes the thread count.
md5-identical output and report; 47.1 -> 44.4 s on that set.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The short-axis pass (the low-FFT-floor hypothesis for small-molecule cells)
ran after the standard first-pass indexing and its rescues, and it is itself
a full indexing of both schemes - about a second of mostly serial refinement,
paid in every first pass, probes included.
It is now started as soon as the standard schemes are fed, on a copy of the
experiment, and runs alongside them. At its own place it is taken only if what
it read is still what the run has there - the experiment and spot-finding
settings (ExperimentKey, split out of FirstPassInputKey), the spot list's
mapping and the per-image spot budget; a rescue that moved any of them makes
it run again there as before. Same inputs through the same code, so the
result is the same.
md5-identical output on two 16M rotation sets; 34.0 -> 29.5 s and
48.2 -> 47.1 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The geometry walk after the canonical pass starts with an indexing probe at
the geometry in hand (probe_at_best), which is the canonical pass's own first
pass run over again - a full first-pass indexing whose answer is already
known (identical validation-spot counts in the logs).
The first pass is deterministic in its inputs, so each canonical or probe
first pass now leaves its validation evidence behind keyed by what it read
(geometry, goniometer, spot budget, cell, indexing and spot-finding settings,
and the run state it consults), and an indexing-only probe with the same key
takes it. Nothing is stored for the geometry pre-pass, for a pass that
changed its own inputs on the way (a rescue adopted), or for one that ran the
beam-centre ladder, which a probe never runs. RUGNUX_VERIFY_FIRST_PASS_MEMO
runs the probe anyway and fails the run on any difference.
md5-identical output; 37.2 -> 34.0 s on a 16M rotation set with a geometry
walk (one probe fewer).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The adaptive finder keeps (local test & ring threshold), so the local test's
second pass is only wanted at the ring pixels. Its window there is the first
pass's window with the first pass's strong pixels taken out, and those are
few: the first pass now records its window sums at the ring pixels, and the
second pass subtracts the strong pixels in each window instead of sweeping
the whole image again. Integer sums and the same acceptance test (factored
into StrongInWindow), so the bits are exactly the dense pass's.
CPU-only rotation run of a 16M set: md5-identical output, 262 -> 242 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
RunImpl allocated the reflection mask (1 byte/pixel) and the owner map (4
bytes/pixel) afresh for every image - 90 MB at 18 Mpixel, above the malloc
mmap threshold, so each call paid an mmap, a zero-fill page fault per 4 kB and
an munmap (with its TLB shootdown across the other workers). On a CPU-only
rotation run of a 16M set that was 79 M page faults and 794 s of system time.
Both are now members of the engine, and each call clears the rectangles the
previous one wrote before it starts. Contents at every read are unchanged, so
the output is byte-identical; CPU-only wall 297 -> 262 s, system time
794 -> 214 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The local-box SNR kernel (analyze_pixel) was 55% of all GPU kernel time on a
rotation run and held the image loop GPU-bound. Three exact changes:
- analyze_pixel is rewritten as one warp per 32 output columns, each lane
holding the vertical sums of two input columns in registers and the 31-wide
horizontal window taken from warp prefix scans. No shared memory and no
block-wide synchronisation; the sums are modular 64-bit integers, so the
result is the same bits as before.
- The adaptive finder keeps only (local-test & ring threshold), and a pixel's
second-pass result depends only on its own window, so the second pass is
evaluated at the ring pixels alone (analyze_candidates) instead of densely
followed by and_bits.
- The first pass is then only read within NBX of a ring pixel, so a warp whose
tile no ring pixel can reach skips it; it also no longer reads an all-zero
previous-pass buffer.
Checked bit for bit against the old kernel on every frame of a 16M rotation
run, and p.hkl / p_unmerged.mtz md5-identical on two inhouse EIGER2 16M sets.
Total kernel time 19.8 -> 12.1 s, image loop 2.69 -> 1.37 ms/image; wall
44.0 -> 37.2 s and 56.0 -> 48.7 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Owner decision: a warning is a prompt to check and must catch the real cases (9min) at the
cost of some spurious ones. The physically motivated corrections stay (<|L|> outside its
physical range is not twinning; a single sweep's indexing choice; NO_LATTICE on rotation;
the no-crystal report); thresholds raised only to cut noise come back down:
- SUPERCELL_POSSIBLE warns wherever the class measures and rocks (as before rc173's audit
fix), worded as "check the cell", naming weak ordered intensity of a correct cell and spots
of further lattice domains as the other readings. 9min (rock 4.2%) warns again.
- LATTICE_TRANSLATION warns on every admitted vector (>=75% of the origin); below 90% the
wording names a very strong pseudo-translation as the other reading.
- PSEUDO_TRANSLATION warns on every detection; below a 20% peak it is worded as weak, check.
- SWEEP_GAPS warns where the degraded ranges cover at least 1% of the sweep (the 4 sets of 77
below that had 1-2 frames, 0.4-0.6% of the sweep).
- Powder rings are split between hexagonal-ice positions and the rest (MeasurePowderRings,
report-only fields); ICE_RINGS and POWDER_RINGS warn separately from 5% of the spots, and
ICE_RINGS also where the merge's ice gate found ice.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Report-layer fixes from the warnings audit of the rc173 full battery. Processing is
unchanged (merged MTZ/HKL bit-identical on the sets checked).
- SUPERCELL_POSSIBLE stays a key and a summary line; a warning only where the rocking part
reaches 20% (sub-cells that refine normally rocked at 2.1-16.2%, the accepted doubled cells
at 4.2% and 29.1%). The summary says where further lattice domains may put spots on the
half-integer nodes.
- LATTICE_TRANSLATION only where the Patterson at the vector reaches 90% of the origin
(UNDECLARED_LATTICE_TRANSLATION_PCT, new); the two admitted vectors on the battery read 79%
and 85% on crystals refining in the deposited cell, and are now reported as a very strong
pseudo-translation. PSEUDO_TRANSLATION warns from a 20% peak (xtriage convention; detected
peaks were 8.1% on one set, 24.8-63.2% on the rest).
- <|L|> outside 0.375-0.55 (on the adopted or the pre-search merge) is TWINNING_VERDICT=
NOT_READABLE with no warning, instead of a twin or SYMMETRY_SUSPECT reading.
- NO_LATTICE no longer fires on a rotation run that indexed the sweep with no frame indexed
on its own; INDEXING_RATE is --developer on rotation.
- A rotation run that finds no lattice writes a report (VERDICT= FAILED, NO_LATTICE) before
exiting 1, instead of an input-parameter error and no report (NoLatticeFound).
- Powder and Ice summary lines, ICE_* keys, and a POWDER_RINGS flag from a 0.5 spot fraction.
- SWEEP_GAPS warns only where the degraded ranges cover 10% of the sweep (77 -> 48 sets).
- A single rotation sweep no longer carries an INDEXING_AMBIGUITY warning (one orientation
matrix, consistent hand); INDEXING_AMBIGUITY_OPERATORS and an Indexing choice line instead.
- A DOMAIN more than 10 deg away is a second crystal; the MULTIPLE_LATTICES sum is documented.
- Mosaicity line and docs name XDS's suggested REFLECTING_RANGE_E.S.D. as the comparable
number (0.90x median over 34 in-house sets; per-image SIGMAR runs ~1.35x higher).
- HARMONIC_CONTAMINATION, SCALING_NOT_CONVERGED, POWDER_RINGS in the documented vocabulary.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
After the existing mask is complete, a second step takes the pixels it left out that are
significantly dimmer than their ring, joins them through a 6 px bridge and across module gaps of
any width, and adds a piece whole when it holds >= 2000 dim pixels of which >= 200 are deep. The
existing mask is never touched, so a sweep with no such piece keeps its mask bit for bit.
Quick subset tests (battery --only, no model check) against the rc173 a0518abe6 full battery:
152 of 233 masks identical; 10 of 10 controls (including the sets where earlier shadow changes
regressed through marginal decisions) give byte-identical merges. On the 21 sets that gain a
piece the space group never changes, merge outlier rejections fall on 19, R_meas falls and ISa
rises on 17, and the shell-scaled model R improves on 16 of 18; a transmitting arm with an
over-subtracted background strip is recovered (R_meas 13.3 -> 12.0 %, ISa 13.4 -> 15.5).
Known costs: on one set with background bumps at ring radii the low-resolution agreement with
the model falls (CC 0.89 -> 0.85) while its internal statistics improve, and two sets cut
slightly coarser (1.50 -> 1.55 A, 1.72 -> 1.80 A) with a better model R.
Squashed from branch hq-beamstop (cfa080a59..60b6c6822), whose history records the variants
tried and dropped: a straight port of the earlier arm finder moved every mask, a two-shape
construction filled almost a whole sweep through hole filling, and a deep-fraction floor
rejected a real arm that is dim along its whole length.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The local radial background at ring radii (--background-radial) already
existed and is gated per image on the smooth-ice channel of the ice score,
but it was off unless asked for. Systematically negative merged
intensities at ring radii (I/sigma < -4 on 0.4-1.6% of reflections of the
affected open sets) come from the flat annulus mean over-subtracting a
sharp ring. rugnux now runs with =auto unless the flag says otherwise; the
library default (and so the broker) stays off.
Measured on the open arm against rc173 (commit 727bd0d6d as reference):
auto: 8agq (smooth ice, fires) CC_model .9327->.9476, R shell-scaled
.1987->.1795, I/sig<-4 1.55%->1.15%; 8rud R shell-scaled
.2507->.2456, ISa 10.0->13.3; 5reo, 7kcn, 9fhc bit-identical
(gate does not fire); wall time unchanged.
forced on (not adopted): same 8agq gain and 8rud .2414, 9fhc/8qq7
small gains, but 7kcn (no ice) CC_model .9491->.9411, ISa
11.0->7.5 - which is why the gate stays.
ISa falls on 8agq (15.6->13.3) while agreement with the model rises: the
merge-statistics vs external-model trade the settings comment describes.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
A new correction surface, fitted after the time x detector surface and
before the goniometer-frame 8x8 grid: log A is a sum of real spherical
harmonics (l = 1..6, 48 terms) of the diffracted-beam direction de-rotated
into the crystal frame. The incident-beam path depends on phi alone and is
in the per-frame scale already.
It runs through ApplyCellSurface unchanged in everything but the update:
the cells are 32 x 64 equal-solid-angle direction bins, and each round the
per-cell sums (ref2, cross, the same damping) become one ridge-regularised
Gauss-Newton step on the coefficients (prior width 0.1/l per degree-l
coefficient) instead of independent per-cell steps. The half-set Fisher-z
gate adopts or refuses it exactly as it does the grids; where it is
refused, the grid after it sees what it saw before.
Why: the folded 8x8 grid (hemispheres share a cell) is the weak basis for
long-wavelength absorption. Offline, held out by unique reflection, this
basis lowered held-out scatter 4-8% on 6 of 11 long-wavelength sets where
no cell grid did, raised model-phased anomalous peaks 0.02-0.2 sigma, and
was neutral on hard-X-ray controls.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Compared with ctruncate (CCP4 9, version 1.17.29) on rugnux's own merged
intensities of the open battery arm, three differences:
- ctruncate gives no amplitude to an intensity below -3.7 sigma (exactly
that bound on every set checked; up to 2498 reflections on one set).
Rugnux turned them into small, confident amplitudes (F/Fc ~0.3 on the
sets whose background is over-subtracted on powder/ice rings). They now
get F = NaN (missing in MTZ, '?' in mmCIF); IMEAN is kept. They are
also left out of the shell mean that sets the Wilson prior.
- The switch to sqrt(I) at I/sigma = 4 left 4-6 sigma amplitudes 2-3%
above ctruncate's on every set (1.019-1.028). The posterior now applies
up to 20 sigma (emulated: 0.995-1.000).
- A shell whose mean intensity is not positive gave a prior at the 1e-10
clamp and amplitudes of ~0 (one set's outer shell); it now takes the
nearest lower-resolution shell's mean.
The remaining gap (weak amplitudes 2-5% below ctruncate's in the outer
shells) is ctruncate's anisotropy-corrected prior; not attempted. An
offline R-free ablation put the whole ctruncate conversion at -0.0006
median (15/18 sets better) and dropping its rejected negatives at -0.0009
mean (-0.009 on the worst set).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
A merged reflection's d came from whichever frame first observed it, while the possible
reflections per shell are enumerated from the reference cell, so reflections near a shell edge
fell on the other side from their possible twin and shells read 100.1-100.9 % complete (17 of
210 battery sets). The group's d is now computed from the reference cell before the export, and
the per-hand table reads the group's d instead of an observation's.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The lattice walk admits an angle up to 3 deg from 90 and the constrained refinement then holds it
there. The existing two-arm test (promoted vs demoted class, judged on the held-out residual) only
covered length equalities (a = b). It now also covers monoclinic/orthorhombic classes: pass 1 reads
how far the free refinement leaves a held angle from 90 (AngleEqualityDeparture_deg), turns it into
pixels at the far corner, and where that exceeds the integration disc runs the demoted arm on the
indexer's free (triclinic) cell. A primitive monoclinic crystal with beta 1.4 deg from 90 was
indexed and integrated as primitive orthorhombic on 60/60 validation frames, so the per-frame guard
never fired; demoted, the held-out residual halves (1.54e-6 -> 8.16e-7), ISa 5.5 -> 6.1, d_min
3.70 -> 3.37 A, CC1/2 at 3.9 A 0.54 -> 0.76. Unchanged on five controls (three oP/oI, two near-90
monoclinic).
Model validation scored alternative cell frames against the data in the merged indexing, before
the indexing-ambiguity probe. On a pseudo-orthorhombic P2 crystal whose data need h,-k,-l to match
the model, every frame read random (R 0.77-0.84) and noise picked an axis swap into P1121; each
frame is now scored at its best over the data's alternative indexings. CC_model 0.23 -> 0.45,
R-free 0.66 -> 0.44 on that set.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
MULTIPLE_LATTICES and EXTRA_LATTICE_INTENSITY_PCT now count DOMAIN, TWIN_DOMAIN and SEGMENTED
lattices only. A FOREIGN lattice (an unrelated cell) is still listed but no longer warned about: it
is as often a wrong main lattice or tNCS as a second crystal (one of the four such sets in the
battery census is a tNCS crystal). On the census this names 25 of 200 rotation sets instead of 29.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Rotation, report only: once the first pass has its lattice, the spots it leaves over (ice left
out) are indexed again with a fresh rotation indexer, up to three times. A lattice is kept when
the leftover validation spots sit on it far more often than at a displaced spindle angle (z >= 5,
the first pass's own null), and classified against the main lattice: a misoriented domain of the
same cell (split crystal), one turned by 180 deg (non-merohedral twin domain), one concentrated
in part of the sweep (>= 60% of its spots in 2 of 8 blocks), a related cell (volume n or 1/n) or
an unrelated one. Misorientation is the minimum over the bases whose metric matches, so a lattice
symmetry is not read as one. EXTRA_LATTICE_* keys, a summary row, and a MULTIPLE_LATTICES warning
where domains or an unrelated lattice hold >= 10% of the non-ice spot intensity (29 of 200
battery sets in the exploration census). Its own IndexAndRefine instances and indexer: merged
MTZ byte-identical on three sets; 0.2-2.5 s per run.
Also MOSAICITY_DEG (median, P10, P90 over frames) - the frame-order-smoothed sigma_M the merge
computes partiality from, Kabsch's sigma_M as XDS's REFLECTING_RANGE_E.S.D. - in the report.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The R-free figure now draws refmac_refined_rfree_shared(_depdata): the deposited model refined
against each data set at its own resolution, R-free over the depositor's free reflections both data
sets share. --rfree first-cycle keeps the previous first-cycle rigid-body pair at the common
resolution. With the rechecked hq-pool3 runs the figure covers 160 structures (was 135).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Adds a second REFMAC protocol to model_check.py beside the first-cycle one
(whose numbers and field names are unchanged):
- the same deposited model is refined against Rugnux's data and against the
depositor's data by the same protocol - restrained, 10 cycles, automatic
weight, isotropic B for every atom, twin refinement where the entry declares
twinning - each data set over its OWN resolution range;
- R-free on a SHARED free set: the depositor's free reflections present in
both data sets after the change of basis (so the lower limit in every
direction). Depositor free reflections outside the intersection are removed
from both; every other reflection is a work reflection of its data set;
- R-free recomputed from each output MTZ over exactly that set (F vs
FC_ALL_LS); for a twinned refinement REFMAC's own R-free (its free set is
the shared set), as the output holds no twinned Fc. REFMAC's own R-free is
kept for reference.
Isotropic B: refining deposited (often TLS-derived) ANISOU atom by atom
diverged (-LL rising every cycle at 1.4 A on a model deposited at 1.7 A).
A ligand whose code REFMAC's monomer library describes with other atoms
("atom ... is absent in the library", a stopped refinement) is renamed to a
code the library lacks, so REFMAC restrains it from the model's coordinates.
Battery row fields: refmac_refined_rfree_shared,
refmac_refined_rfree_shared_depdata, refmac_refined_ratio,
refmac_refined_rwork, refmac_refined_rwork_depdata,
refmac_refined_rfree_refmac, refmac_refined_rfree_refmac_depdata,
refmac_shared_free_n, refmac_refined_reason. model_check.py --no-refine
skips the protocol. README documents both protocols.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
`battery.py recheck RUN [--only a,b] [--jobs N]` reruns model_check.py
(REFMAC) and depdata_check.py on every open-arm row of a finished run from
its own p.mtz, and updates the refmac_* / dep_* fields of results.json in
place (verdicts untouched). Previous results.json, report and each set's
model_check directory are kept with a .pre-recheck-<stamp> suffix; the
manifest gains a "rechecks" entry {date, end, runner_git, runner_dirty, sets,
fields, previous_results, command} so the tools version behind each published
number is auditable. The report is re-rendered and the run made read-only
again.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
-z (now also --reference; --reference-mtz still works) takes an SF-mmCIF
(e.g. a deposited -sf.cif, gzipped or not) as well as an MTZ. The format is
recognised by content: a file starting with "MTZ " is an MTZ, anything else
is parsed as CIF. The first merged reflection block with the requested
column (or, by default, one the auto choice accepts) is converted to a
gemmi::Mtz in memory with GEMMI's CifToMtz and then read by the unchanged
MTZ loader, so the in-memory reference is exactly what the MTZ path yields.
Unmerged (_diffrn_refln) and anomalous-only blocks are passed over; the log
names the block used.
The R-free set comes from _refln.status (f -> FreeR_flag 0, o -> 1: the
CCP4 convention the loader already reads) and is preferred to
_refln.pdbx_r_free_flag, whose convention varies by program; a status
column with no 'f' is ignored and pdbx_r_free_flag is used instead.
Checked on two open-arm sets in the deposited setting (one with F_meas_au
only, one with intensity_meas): reference loaded with the deposited cell and
group, the inherited free set agrees with the deposited status 'f' on every
common reflection, and the merge correlates at CC 0.99 with the deposition.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
RUGNUX_ADVANCED (new section "The setting the files are written in", -z and
--model paragraphs), RUGNUX_REPORT (SETTING_* keys, REFERENCE_MISMATCH, model
change-of-basis keys), CPU_DATA_ANALYSIS_DECISIONS / _INTEGRATION, CHANGELOG
entries under rc.173.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The files were written on the axes the space-group search named, which for
P2221/P21212 puts the unique axis wherever the a<b<c indexing put it (a 52.51
87.87 137.72 crystal came out P 2 21 21 where XDS writes 87.87 137.72 52.51
P 21 21 2), and a reference MTZ was matched in the data's frame: on permuted
axes its free-R flags landed on unrelated reflections while the log reported a
high matched count.
- New CrystalSetting (scale_merge): changes of basis between settings of one
lattice (cell, group, index operator, basis matrix), the {-1,0,1} det +1
candidates (CellMappingOperators, moved from ModelValidation), MetricViolation
(moved from Rugnux), ChooseOutputSetting and SeatGroupByAbsences.
- Output setting: after every decision the merge, the integrated reflections
(unmerged MTZ), the P1 cross-check, the lattice and the _process.h5 reindex
matrix are relabelled into the ITA standard setting - or, in priority order,
a reference MTZ's, a fitting model's, the -C axis order, a non-standard -S
symbol's. Free-R flags are drawn again on the written axes. Reported as
SETTING_OPERATOR / SETTING_SOURCE.
- Reference MTZ: the group is kept in its setting; after the merge every cell
mapping onto the reference cell (times the twin laws) is scored by the
reference CC, the best is re-seated and re-merged, and the free flags are
inherited only where CC >= 0.5 over >= 50% of the reference range
(REFERENCE_MISMATCH otherwise; --mode scale gates the same way).
REFERENCE_OPERATOR / _CC / _MATCHED_FRACTION / _FREE_FLAGS_INHERITED.
- -S: a fixed group is put on the axes its absences name before merging
(SeatGroupByAbsences), fixing -S 18 on a cell whose pure axis is not c.
- --model: a model in another setting is now a claim the null tests; where it
fits, the data are written in its setting and the validation is remade on
those axes (KeepModelVerdict carries the decisions over).
Tests: [setting] (synthetic #18/#17/I222/C222/c-unique P21/C2 beta/I2->C2/P1,
-C and -S order, absence seating, permuted reference with flags).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The time histogram now covers the open arm, like the other two figures, and can come from a run
whose sets had the GPU to themselves while the quality figures use the latest run; summary.txt
lists the provenance of both. The median label sits above the bars.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Third figure: rugnux's own WALL_TIME per dataset (rugnux_wall_s, not the runner's elapsed_s,
which adds the wait for a GPU slot under --gpulock) as a 25 s linear histogram with the median
marked and the open/in-house counts. Every set that ran to a report is counted; crashed sets and
the no-crystal controls (no report, so no WALL_TIME) go to excluded.txt. Sets that shared the GPU
(gpu_others > 0) are drawn as the lighter top segment of their bar rather than dropped, and
--uncontended RUN sizes their slowdown against a run with the GPU to itself, reported in
summary.txt with the percentiles, the arm medians and the page-cache caveat. time.dat carries
set, arm, seconds, images and gpu_others beside the figure.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Protocol changes in tools/battery/model_check.py (the --model-check R factors):
- Model preparation: atoms of unknown element (UNX, element X) are left out
(REFMAC stops on them: "Atom does not seem to be an atom : X"); chain names
longer than two characters are shortened so the model is written in PDB
format, whose ANISOU records REFMAC reads (its mmCIF ANISOU reader failed with
"rdaniso_cif: Atom symbol mismatch" on two entries). A model too large for the
PDB format is written as mmCIF with isotropic B only. What was changed is
recorded in model_modification.
- Depositor's structure factors: amplitudes from F_meas, else the mean of
F(+)/F(-); where only intensities were deposited (I, else the mean of
I(+)/I(-)) they are converted by French-Wilson (ctruncate) instead of being
skipped. Recorded in depdata_kind. A numeric pdbx_r_free_flag with more than
two values is read with the CCP4 convention (0 = free); status 'f' is still
preferred when present.
- The depositor's data go through the same change-of-basis choice as Rugnux's
(reindex_op_depdata): a twinned entry can be deposited in the other branch
of a merohedral ambiguity relative to its model.
- An entry that declares twinning (_pdbx_reflns_twin, more than one domain) is
scored with REFMAC's twin refinement for both data sets; the untwinned
numbers are kept under "untwinned" (battery rows: refmac_twin,
refmac_untwinned).
- The work directory is made absolute (REFMAC runs inside it; a relative
--workdir failed).
Unchanged: rigid-body first-cycle R factors, d_min_used = max(Rugnux d_min,
deposited d_min), the field names refmac_rfree / refmac_rfree_depflags /
refmac_rfree_depdata / refmac_rfree_ratio / refmac_reason.
Re-run on the existing p.mtz of all 165 open-arm PDB sets of the last full
battery: 131 give identical numbers; 3 former REFMAC failures now score; 23
intensity-only or F(+)/F(-)-only depositions gain the depositor baseline; 7
twin-declared entries move to twinned R; one untwinned P3x entry whose model is
near-symmetric under the merohedral twofold picks the other (tied) branch for
the depositor's data (R-free 0.2928 -> 0.2904).
README: which R-free is which, the common resolution limit, and the twin and
French-Wilson handling.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Owner's choice for the 8.8 cm single-column figures; still inside the IUCr notes'
1.5-3 mm lettering range (10 pt Helvetica capitals are ~2.5 mm). Points enlarged
slightly to match. Presentation only.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Presentation only; the data, counts and audit files are unchanged.
Follows the IUCr artwork guide (journals.iucr.org/services/help/artwork/guide.html)
and the Acta D notes for authors (journals.iucr.org/d/services/notesforauthors.html):
- each figure is a single-column 8.8 cm square (guide: 8.85 cm; notes: 8.8 cm), so a
two-panel composite also fits the 18 cm page width
- lettering 8 pt upright Helvetica, embedded by pdfcairo (guide: ~8 pt, standard fonts
Arial/Courier/Helvetica/Symbol/Times, fonts embedded; notes: 1.5-3 mm lettering)
- line weights 0.75 pt for border, ticks and the y = x line (guide: 0.35-1.5 pt;
pdfcairo lw 1 is 0.5 pt)
- the PNG is the PDF rasterised with pdftoppm at 600 d.p.i. (guide: 400 d.p.i. colour,
600 d.p.i. line art), so the two files are the same figure
- no grid (notes: grids avoided where not required), one Okabe-Ito blue for the points,
which stays distinct from the black dashed line in greyscale and to colour-blind readers
- ticks short, outward and not mirrored; d_min ticks 0.5/1/2/4/8 A all at one decimal
- italic R and d with roman subscripts; the REFMAC qualifier moves to the caption
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The figures, their value pairs, the excluded sets with reasons, the gnuplot script and the runs'
provenance are written together, so a published figure can be audited and redrawn from the
battery run it came from. Facility and beamline counts come from the entries' _diffrn_source.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The magnifier moves from a helper window into its own dock below the
inspector. It has three fixed zoom levels (x64 and x32 with the pixel
values written on the pixels, x10 without) instead of wheel zoom, and
follows the cursor only while Shift is held over the diffraction image.
Meanwhile the diffraction view draws a frame around the area it covers,
just outside that area so it stays visible at low zoom without hiding the
pixels; releasing Shift (seen by the application-wide key filter, so
wherever the focus is), moving without it or leaving the image hides it.
A "Pop out" button moves the magnifier into a separate window and back.
Qt's own dock floating stays disabled: a floated dock is placed
off-screen on WSLg. Whether it was popped out, and the window geometry,
persist across sessions. The inspector's image statistics become a
collapsible section so a small screen can give the space to the
magnifier. kLayoutVersion is bumped for the new dock.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Advisory only: the class is measured (<I/sigma> >= 0.5) and its rocking part is at least 2% of the
lattice's, three standard errors clear. The warning asks the user to process both settings - the
run's cell and the reported doubled cell with -C - and compare them in refinement. Of 105 battery
rotation sets it names 8: both crystals whose accepted cell is the doubled one, one pseudo-translation
whose doubled description is an accepted alternative, and five correct sub-cells with weak ordered
half-integer intensity.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The report-only supercell probe now reaches the results report: SUPERCELL_CLASS (parity of the
primitive indices), its occupancy and Bragg-like (rocking) part against the lattice's own reflections,
its <I/sigma>, and SUPERCELL_DOUBLED_CELL, the Niggli-reduced cell to give with -C. No decision is
taken on it: correct cells whose half-integer class is diffuse, and cells whose depositor kept the
sub-cell, cannot be told from a real doubling by the data alone. On a crystal with a real doubled
axis the reported cell matched the deposited one, and processing on it with -C brought R_free
against the deposited model from 0.60 to 0.29.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The between-pass geometry walk kept a round only when the realised held-out residual fell by more
than its noise, and started only on a fit move of a trust-region step or more. The residual is
dominated by low-resolution reflections, where a distance and the compensating cell scale move every
spot alike, and its centroids are taken inside a disc centred on the prediction, so it barely sees a
distance error that costs the high-resolution shells. On 8pqd the canonical pass ran at 96.456 mm; the
fit asked for 95.878 mm (0.6 %, less than a step, so no walk). Forced, that round read the residual
only 0.74 sigma lower but put 70.4 % of the validation spots on the lattice against 58.1 %.
Now a round is kept when either the residual falls beyond its noise (HeldOutResidualFell) or the
validation evidence prefers it (ValidationEvidencePrefers, z = 3.29). A move of less than a step is
first tried as two index-only probes on the validation frames (fit's geometry vs the one in hand) and
pays for a re-integrated round only where the fit's geometry scores higher; walks started by a large
move run as before.
8pqd: 96.456 -> 95.878 mm, d_min 1.374 -> 1.306 A, R_free .218 -> .206, REFMAC R_free .209 -> .202,
Wilson B 34.4 -> 30.7 (pool had 1.325 A / .209 / .203). 6vww 7n0i 7ris 8egn 8xtf 9qw8 9w3y 6qaj
6cdl 9i0a myob_x06da_split lyso_x06da_half_image unchanged (probes say the geometry in hand stands);
probe cost 1-10 s per run.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
R_MEAS is what users set beside XDS/AIMLESS, so its meaning does not change; the merge-weighted value
from hq-cure-merge is reported beside it (R_MEAS_WEIGHTED, REFRES_R_MEAS_WEIGHTED). The per-hand
table keeps the ordinary statistic.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Every lattice decision counts spots or frames, and the sub-lattice's strong
reflections win every count, so a crystal whose cell is doubled by a weak
superstructure class (9min, 6z9g) is adopted at the half cell. This asks the
question in intensities instead. On 60 frames spread over the sweep, after
each frame's own integration, the 2a x 2b x 2c supercell of its primitive
lattice is predicted to 3 A with the frame's own refined orientation and
geometry and integrated on the same engine; the reflections are summed per
parity class in two shells (20-5, 5-3 A), with a fit of intensity against
partiality (I = a + b p) that separates what rocks like a Bragg reflection
from what sits at the node whatever the rocking. Only the tested frames are
predicted and nothing is retained, so memory is bounded (the previous
prototype predicted the whole run through the merge and ran out of GPU
memory). The probe's integrations are kept out of the engine's own counts,
which the two-pass stencil guard reads. Results are bit-identical with and
without it.
REPORT ONLY - it decides nothing, because on the battery it does not yet
separate a weak real class from what sits at the half-integer nodes of
crystals whose cell is right. Real classes: 6z9g class 101 at 24 % of the
lattice's intensity (29 % rocking), 9min 100 at 19 % (4 % rocking - its real
class does not rock like the lattice either). On correct cells the largest
classes reach 9-12 % raw (7n2s, 9i0a, 7os3) and 3-4 % rocking (7dkp, 7os3),
and 7mzt reads 40 % / 19 % on a class the deposition does not have. The log
line is the population a decision has to be calibrated on.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The joint crystal + detector fit was committed only when the held-out residual
fell by at least a fixed 2 %. On a crystal whose geometry is already right
that bound cuts through the noise: two builds read 1.9 % and 2.05 % for the
same 0.03 px / 0.01 % move (lyso_x06da_half_image), one committed and the
other did not, and everything downstream (smoothing window, stretch
segmentation) followed the coin.
The bar is now the residual's own noise - the standard errors of the two
held-out means combined - which is the bar a round of the geometry walk
already has to clear (HeldOutResidualFell). The log line carries the joint
residual and that noise. Noise on the battery is 2-7 % of the residual, so
the gate is about as strict as before but no longer at a fixed edge.
Checked on 57 open/in-house sets and 11 private ones against the pool
battery. The verdict changed on 8 open/in-house sets besides the passes that
follow the 7ris/8pqd lattice change: half_image, 6qaj, lalanine, 8xtf, 9ig7,
8sqt (commit -> reject of a 1-step move) and myob_x06da_split, 6cdl (reject
-> commit). Merge statistics are equal to the last digit everywhere except
half_image (REFRES R_meas 1.16 -> 0.95, CC1/2 0.79 -> 0.86), 9ig7 (5 fewer
rejected observations) and 8xtf, whose stretch disposition followed the
coin: multiplicity 20.8 -> 15.6, R_meas 0.91 -> 0.60, R_free 0.193 -> 0.189,
REFMAC R_free 0.179 -> 0.176. Private arm unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Where the file's and the measured beam centre index lattices whose primitive
volumes differ by an integer factor 2-4, the pooled validation evidence decided
only when the measured centre held the LARGER cell. The other direction was
left to the second pass on the assumption that it drops to the smaller cell
by itself. It does so only when the pass-2 indexer's free beam refinement
walks the whole centre error, and that rescue is chaotic: on 7ris a 0.1 %
change in the pass-1 distance switched off a 6.5 px walk, the doubled c axis
(380 A) won and the run failed.
Both directions are now decided on the same measurement: pooled validation
spots on each lattice against its own wrong-spindle null
(ValidationEvidencePrefers). No new threshold. 7ris: 44.3 % at the file's
centre vs 83.5 % at the measured one; the run adopts the measured centre and
the 190 A cell (R_free 0.81 -> 0.20, 1.52 A). 8pqd (3x, measured centre
holds the true cell): 16.6 % vs 21.0 %, adopted in pass 1 instead of being
recovered by pass 2.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C