Commit Graph
214 Commits
Author SHA1 Message Date
leonarski_fandClaude Opus 5.5 d3afe0cb07 ImageSpotFinderCPU: the first pass slides its window only where its result is read
With candidates, the first pass already tested only the 32-column blocks within reach of a candidate,
but still slid the horizontal window across every column of every row. It now slides it only over
the runs of needed blocks, starting each run from the vertical sums it covers - integers, so the same
window sums - and sets the bits there directly; the rest stay 0 as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 23:01:21 +02:00
leonarski_fandClaude Opus 5.5 7e1be959a6 BraggIntegrationEngineCPU: the background clip walks the ring pixels pass A read
The sigma-clip pass recomputed the stencil distances over the whole box to find the same background
ring pixels pass A had just summed. Pass A now keeps their values and radial offsets in its own
order, and the clip runs over them - the same pixels, the same sums.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:56:29 +02:00
leonarski_fandClaude Opus 5.5 967d5691cc BraggIntegrationEngineCPU: gather the profile fit's pixels once, not on every reweighting
The four Kabsch reweighting iterations each walked the reflection's grid again with the same bounds,
validity and ownership tests the p_valid pass had just applied. That pass now keeps the profile value
and background-subtracted count of the pixels the fit reads, in grid order, and the iterations run over
them - the same terms summed in the same order (about 6 s of a CPU-only 16M run).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:55:48 +02:00
leonarski_fandClaude Opus 5.5 a6c27e9a04 ShadowFinder: the element-wise full-detector mask updates on all threads
Four per-pixel passes over char masks ran on the pre-scan's critical path on one thread; each
pixel's result depends on that pixel alone, so splitting them changes nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:55:17 +02:00
leonarski_fandClaude Opus 5.5 2746c50822 BeamCenterFromBackground: skip pixels clearly outside the band before sqrt and atan2
60-70% of a large detector lies outside the resolution band the walk bins, and every one paid a
square root and two arctangents on every iteration. A test on tan(2theta) = rho / lz with a 0.1%
margin, far above float rounding, drops them first; every pixel the exact test keeps still reaches it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:54:37 +02:00
leonarski_fandClaude Opus 5.5 b119b19707 BraggIntegrationEngineGPU: grow the per-reflection buffers to twice the record
Each growth is 11 pinned host and 19 device allocations under the driver's device-wide lock; at the
start of an image loop 16 workers growing 1.5x at a time spent about 0.3 s in them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:52:27 +02:00
leonarski_fandClaude Opus 5.5 cfe202e9a1 ScaledObservations: sample each image on its own thread
The ASU-key sample over every partial ran on one thread (0.2-1 s of the merge tail). Each image's
parts are now taken on the workers and joined in image order, so the sort sees the same parts in the
same order.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:52:09 +02:00
leonarski_fandClaude Opus 5.5 dd02e65693 PostRefine: rank the observation cap on I/sigma worked out once, not per comparison
nth_element compared through pointers into gigabytes of partials, a cache miss per comparison on one
thread (0.3 s on a 16M sweep). The ratio is now computed beside each pointer on all threads; the
comparisons and so the selection and its order are the same.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:51:31 +02:00
leonarski_fandClaude Opus 5.5 2c48ed3fcd RotationScaleMerge: measure the rocking-event span once, not twice per Run
At the smoothing window every Run has just restored corr to what Ingest built, and nothing else the
measurement reads changes after Ingest, so its answer is the same on every Run; it walked all the
partials twice per Run on the CPU path.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:50:11 +02:00
leonarski_fandClaude Opus 5.5 45a8db03d4 RotationScaleMerge: without a GPU, fold the partials' corr update into the rescale pass
RunScalingLoop rescales after every iteration before anything reads corr, so the iteration's
UpdateCorr and the rescale are one pass over the partials instead of two: the G and fitted frames
as the iteration left them, then the ratio on top - the same two roundings in the same order.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:49:32 +02:00
leonarski_fandClaude Opus 5.5 6df56e661b RotationScaleMerge: without a GPU, sum the partials' group means on all threads
ReduceGroupMeans ran once or twice per scaling iteration as one serial scatter over every partial
(about 8 s of a CPU-only run's main thread on a 16M sweep). ComputeAsuGroups already builds the
partials' group CSR - a stable counting sort, observation order within each group - for the GPU
reduction; it is now kept on the CPU path too, and each group is summed over it in the order the
serial loop added it, so the means are the same numbers.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:48:41 +02:00
leonarski_fandClaude Opus 5.5 698d06a0e2 BraggPredictionRot: reject solutions far from the frame before the rest of the arithmetic
zeta >= min_zeta, so the rocking-curve test rejects every solution with
min_zeta * (|phi| - half wedge) > multiplier * mosaicity. Tested right after phi, with a 1e-5 rad
margin (float rounding of the full test is below 1e-6), it skips the cross product, normalisation
and zeta of the 85-95% of solutions no frame keeps, and cannot reject one the full test would pass.
CPU build: on the order of 10 s of a 1.2 A sweep's prediction.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:47:12 +02:00
leonarski_fandClaude Opus 5.5 79bc90e512 RotationIndexer: keep 1024 indexing outcomes, keyed by a digest of the input
A run on data that index poorly asks over a hundred indexing questions - every rung of the
spot-budget ladder, in every pass and walk probe - and a later pass repeats a probe's ladder, which
the 32-entry memo had long evicted (about 20 s on one such sweep). The key holds every spot, so it is
now kept as two independent 64-bit hashes and its length instead of whole.
RUGNUX_VERIFY_FIRST_PASS_MEMO still recomputes and compares.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:46:20 +02:00
leonarski_fandClaude Opus 5.5 92877e5eb7 Rugnux: build the unmerged MTZ beside the P1 cross-check merge
The unmerged file reads the integration outcomes and the determined group, neither of which the P1
merge changes, except each image's mosaicity, which the batch headers carry and the merge rewrites.
UnmergedMtz builds the file without it on a second thread, and SetUnmergedMtzMosaicity fills it in
after the merge, so the file is the same bytes as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:45:48 +02:00
leonarski_fandClaude Opus 5.5 cd98787728 CPU analysis: take the azimuthal profile in the adaptive finder's ring pass
On the CPU path every image made a separate azimuthal-integration pass
(AzIntEngineCPU) over the 72 MB frame although the adaptive finder's first
ring pass reads the same pixels in the same order under the same rules
(skip the INT32_MIN/MAX sentinels, bins below the mapping's count). That
pass now also accumulates the corrected profile - the same statements as
AzIntEngineCPU, so the same float sums - and MXAnalysisWithoutFPGA takes the
profile from the finder instead of running the separate pass, as the fused
GPU engine already does. Only where the azimuthal engine would be the CPU
one; the finder the pre-scan uses does not accumulate it.

md5-identical p.hkl, p_unmerged.mtz, p_plot.txt and report; CPU-only
163 -> 149 s and 94 -> 91 s. GPU unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 21:18:04 +02:00
leonarski_fandClaude Opus 5.5 ef4090c830 CPU FFT indexer and beam-centre shortlist: less serial work
- FFTIndexerCPU::ExecuteFFT built the per-direction histograms and scanned the
  per-direction spectra for their most prominent peak in one thread; every
  direction has its own histogram and spectrum, so the directions are now
  split over the refinement threads, each filled and scanned in the same
  order as before.
- BeamCenterShortlist2D scanned the whole padded surface (73 M points on a 16M
  detector) once per candidate. It now keeps each row's maximum and the first
  index holding it and rescans only the rows a suppression touched; rows in
  order, first index within a row, is the same first maximum.

md5-identical output; CPU-only 177 -> 169 s and 101 -> 96 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 20:48:08 +02:00
leonarski_fandClaude Opus 5.5 87cfc879b6 CPU adaptive spot finder: sigma clips from a histogram, first pass only where read
Two whole-image passes per frame out of the CPU spot finder (the pre-scan's
finder on every build, and every image on the CPU-only build):

- AccumulateRings ran three passes over the frame - the plain ring statistics
  and two sigma clips. The plain pass now also counts each ring's valid
  values in a histogram (0..1023, the rest in a short list), and the clip
  passes sum over the distinct values: each meets the same float test its
  pixels would, and the sums are integers, so the totals are the same.
- The local test's first pass is read by DetectAt only inside a candidate's
  window. It now marks the row/32-column blocks those windows reach and
  keeps its sliding sums everywhere but skips the per-pixel test elsewhere;
  the bits it leaves unset are never read.

md5-identical output on four sets (GPU) and three (CPU-only). CPU-only
16M: 203 -> 177 s and 303 -> 263 s; GPU 16M 23.8 -> 22.8 s (the pre-scan's
finder).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 20:40:46 +02:00
leonarski_fandClaude Opus 5.5 94505d10c4 Merge tail: parallel delta-CC1/2 ranges, twin-immune zone evidence and P1 tNCS
Three serial stretches of the canonical pass's tail, each made parallel with
the same arithmetic in the same order:

- MeasureBatchDeltaCCHalf measured every batch of the curve, every open
  candidate of the rejection loop and every ledger range one after another.
  Each measurement is a pure function of its range and the fixed totals, so
  they now run side by side (measure_with, one scratch set per worker) and the
  decisions scan the results in the original order; the edge walk (locate)
  stays one at a time.
- The twin-immune zone evidence (CentricOverAcentric, a few hundred
  exponentials per reflection) is evaluated in parallel and summed in the
  original order.
- The P1 cross-check's AnalyzeTranslationalNCS was still called with one
  thread; it gets the run's thread count like the other two calls.

md5-identical p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz and report on four sets;
GPU wall 41.0 -> 38.3 s and 26.9 -> 24.8 s on the two sets with long merges.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 20:00:27 +02:00
leonarski_fandClaude Opus 5.5 a6bb16ce8c RotationIndexer: keep what RunIndexing computed, keyed by everything it reads
A rotation run asks the rotation indexer the same question many times: the
canonical pass's first pass repeats the rotation-scale probe at the stored
angles exactly (identical validation evidence on every set checked), and on
a crystal that does not index, pass 2 and every probe repeat pass 1's rescue
ladder rung for rung. Each such RunIndexing is an FFT search plus a serial
Ceres fixed-point chain, ~1-3 s on the GPU build.

RunIndexing is deterministic in its inputs, so its outcome - every member it
sets - is now kept process-wide under a key of all of them: the accumulated
spots (every field) and their angles, both geometries and the axis, the
experiment's indexing settings, cell and space group, and the settings of the
pool it indexes with (IndexerThreadPool::Settings). A RotationIndexer asking
with the same key takes the outcome. RUGNUX_VERIFY_FIRST_PASS_MEMO recomputes
and throws on a difference.

md5-identical output on four sets; GPU wall 28.9 -> 26.1 s, 43.9 -> 41.0 s,
30.3 -> 26.9 s and 143.6 -> 120.5 s (the set that does not index); the
verify mode found no difference on the last.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 19:52:46 +02:00
leonarski_fandClaude Opus 5.5 c7a8fabc7e BeamCenterFFTCPU: run the capture's independent 2-D transforms at the same time
PointSurfaces made three forward and four inverse transforms of the padded
detector (8748 x 8400 on a 16M) one after the other on the main thread -
~3.5 s of one core in a CPU-only run. The forwards do not depend on each
other, nor do the inverse products; each group now runs at the same time,
one workspace per transform. A workspace is built as the single one was -
its own std::vector buffers and its own FFTW_ESTIMATE plan for them - so the
same plan runs on the same data and the result is bit-identical. Costs about
2 GB more transient memory on a 16M detector.

CPU-only 16M rotation run: md5-identical, 216 -> 210 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 19:39:56 +02:00
leonarski_fandClaude Opus 5.5 59d92a7238 ParallelSort for the two large sorts on the merge path
Two whole-dataset sorts sat on the main thread at the end of a rotation run:
WilsonOutliers orders every full by resolution, and the unmerged MTZ is put
in H K L M/ISYM BATCH order by Mtz::sort(5) - together about 2 s of one
thread on a 1.4 M-observation set.

ParallelSort (common/ParallelFor.h) sorts one piece per worker and merges
them pairwise. It is only for comparators that are a strict total order,
where the sorted sequence is unique and the result is the serial sort's bit
for bit: WilsonOutliers already breaks ties on the index, and the unmerged
writer now sorts the rows itself on the five key columns and then the row
number - the order Mtz::sort's stable sort gives - and sets sort_order as it
did. An empty table still fails the way Mtz::sort does.

md5-identical p.hkl, p.mtz and p_unmerged.mtz; 44.4 -> 43.6 s on a 16M set.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 18:13:41 +02:00
leonarski_fandClaude Opus 5.5 c244d384b0 tNCS: evaluate the class-vector ladder's neighbours in parallel
Where a pseudo-translation is detected, the class vector is re-refined up a
ladder on the whole data by a greedy sweep over the 26 neighbours of the
current vector, each a cosine correlation over up to every acentric
reflection - seconds of one thread on the main path of the run (3.4 s and
1.5 s for the two calls on a P3_121 16M set).

The neighbours of the current vector are now evaluated together; the first
improvement in the sweep's own order is taken, and the neighbours after it
are evaluated again about the new vector - exactly the moves the serial
sweep makes. Each correlation is still one serial sum, so the result is
bit-identical. AnalyzeTranslationalNCS takes the thread count.

md5-identical output and report; 47.1 -> 44.4 s on that set.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 18:08:36 +02:00
leonarski_fandClaude Opus 5.5 50098d20ba CPU adaptive spot finder: second local pass only at the ring pixels
The adaptive finder keeps (local test & ring threshold), so the local test's
second pass is only wanted at the ring pixels. Its window there is the first
pass's window with the first pass's strong pixels taken out, and those are
few: the first pass now records its window sums at the ring pixels, and the
second pass subtracts the strong pixels in each window instead of sweeping
the whole image again. Integer sums and the same acceptance test (factored
into StrongInWindow), so the bits are exactly the dense pass's.

CPU-only rotation run of a 16M set: md5-identical output, 262 -> 242 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 17:30:44 +02:00
leonarski_fandClaude Opus 5.5 67ec3649a4 CPU Bragg integration: keep the full-frame scratch between images
RunImpl allocated the reflection mask (1 byte/pixel) and the owner map (4
bytes/pixel) afresh for every image - 90 MB at 18 Mpixel, above the malloc
mmap threshold, so each call paid an mmap, a zero-fill page fault per 4 kB and
an munmap (with its TLB shootdown across the other workers). On a CPU-only
rotation run of a 16M set that was 79 M page faults and 794 s of system time.

Both are now members of the engine, and each call clears the rectangles the
previous one wrote before it starts. Contents at every read are unchanged, so
the output is byte-identical; CPU-only wall 297 -> 262 s, system time
794 -> 214 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 17:23:12 +02:00
leonarski_fandClaude Opus 5.5 aac1c54fb7 GPU spot finder: warp-per-32-columns local test, second pass only at ring pixels
The local-box SNR kernel (analyze_pixel) was 55% of all GPU kernel time on a
rotation run and held the image loop GPU-bound. Three exact changes:

- analyze_pixel is rewritten as one warp per 32 output columns, each lane
  holding the vertical sums of two input columns in registers and the 31-wide
  horizontal window taken from warp prefix scans. No shared memory and no
  block-wide synchronisation; the sums are modular 64-bit integers, so the
  result is the same bits as before.
- The adaptive finder keeps only (local-test & ring threshold), and a pixel's
  second-pass result depends only on its own window, so the second pass is
  evaluated at the ring pixels alone (analyze_candidates) instead of densely
  followed by and_bits.
- The first pass is then only read within NBX of a ring pixel, so a warp whose
  tile no ring pixel can reach skips it; it also no longer reads an all-zero
  previous-pass buffer.

Checked bit for bit against the old kernel on every frame of a 16M rotation
run, and p.hkl / p_unmerged.mtz md5-identical on two inhouse EIGER2 16M sets.
Total kernel time 19.8 -> 12.1 s, image loop 2.69 -> 1.37 ms/image; wall
44.0 -> 37.2 s and 56.0 -> 48.7 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 17:16:06 +02:00
leonarski_fandClaude Opus 5.5 e9da892790 Rugnux report: low warning thresholds, worded as prompts to check; ice apart from powder
Build Packages / Create release (push) Successful in 15s
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 7m56s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 8m39s
Build Packages / build:viewer-tgz:cpu (push) Successful in 10m23s
Build Packages / build:viewer-tgz:cuda (push) Successful in 12m48s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 15m22s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 16m11s
Build Packages / build:windows:nocuda (push) Successful in 18m24s
Build Packages / build:windows:cuda (push) Successful in 20m42s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 24m38s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 19m11s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 19m24s
Build Packages / Generate python client (push) Successful in 45s
Build Packages / build:rugnux:windows (push) Successful in 11m0s
Build Packages / Build documentation (push) Successful in 1m22s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 20m36s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 19m52s
Build Packages / build:rpm (rocky8) (push) Successful in 18m15s
Build Packages / build:rpm (rocky9) (push) Successful in 18m10s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 14m2s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m29s
Build Packages / Unit tests (push) Successful in 1h14m50s
Owner decision: a warning is a prompt to check and must catch the real cases (9min) at the
cost of some spurious ones. The physically motivated corrections stay (<|L|> outside its
physical range is not twinning; a single sweep's indexing choice; NO_LATTICE on rotation;
the no-crystal report); thresholds raised only to cut noise come back down:

- SUPERCELL_POSSIBLE warns wherever the class measures and rocks (as before rc173's audit
  fix), worded as "check the cell", naming weak ordered intensity of a correct cell and spots
  of further lattice domains as the other readings. 9min (rock 4.2%) warns again.
- LATTICE_TRANSLATION warns on every admitted vector (>=75% of the origin); below 90% the
  wording names a very strong pseudo-translation as the other reading.
- PSEUDO_TRANSLATION warns on every detection; below a 20% peak it is worded as weak, check.
- SWEEP_GAPS warns where the degraded ranges cover at least 1% of the sweep (the 4 sets of 77
  below that had 1-2 frames, 0.4-0.6% of the sweep).
- Powder rings are split between hexagonal-ice positions and the rest (MeasurePowderRings,
  report-only fields); ICE_RINGS and POWDER_RINGS warn separately from 5% of the spots, and
  ICE_RINGS also where the merge's ice gate found ice.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 14:02:05 +02:00
leonarski_fandClaude Opus 5.5 8a8c992cca Rugnux report: warn only where the evidence supports the claim
Report-layer fixes from the warnings audit of the rc173 full battery. Processing is
unchanged (merged MTZ/HKL bit-identical on the sets checked).

- SUPERCELL_POSSIBLE stays a key and a summary line; a warning only where the rocking part
  reaches 20% (sub-cells that refine normally rocked at 2.1-16.2%, the accepted doubled cells
  at 4.2% and 29.1%). The summary says where further lattice domains may put spots on the
  half-integer nodes.
- LATTICE_TRANSLATION only where the Patterson at the vector reaches 90% of the origin
  (UNDECLARED_LATTICE_TRANSLATION_PCT, new); the two admitted vectors on the battery read 79%
  and 85% on crystals refining in the deposited cell, and are now reported as a very strong
  pseudo-translation. PSEUDO_TRANSLATION warns from a 20% peak (xtriage convention; detected
  peaks were 8.1% on one set, 24.8-63.2% on the rest).
- <|L|> outside 0.375-0.55 (on the adopted or the pre-search merge) is TWINNING_VERDICT=
  NOT_READABLE with no warning, instead of a twin or SYMMETRY_SUSPECT reading.
- NO_LATTICE no longer fires on a rotation run that indexed the sweep with no frame indexed
  on its own; INDEXING_RATE is --developer on rotation.
- A rotation run that finds no lattice writes a report (VERDICT= FAILED, NO_LATTICE) before
  exiting 1, instead of an input-parameter error and no report (NoLatticeFound).
- Powder and Ice summary lines, ICE_* keys, and a POWDER_RINGS flag from a 0.5 spot fraction.
- SWEEP_GAPS warns only where the degraded ranges cover 10% of the sweep (77 -> 48 sets).
- A single rotation sweep no longer carries an INDEXING_AMBIGUITY warning (one orientation
  matrix, consistent hand); INDEXING_AMBIGUITY_OPERATORS and an Indexing choice line instead.
- A DOMAIN more than 10 deg away is a second crystal; the MULTIPLE_LATTICES sum is documented.
- Mosaicity line and docs name XDS's suggested REFLECTING_RANGE_E.S.D. as the comparable
  number (0.90x median over 34 in-house sets; per-image SIGMAR runs ~1.35x higher).
- HARMONIC_CONTAMINATION, SCALING_NOT_CONVERGED, POWDER_RINGS in the documented vocabulary.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 14:02:05 +02:00
leonarski_fandClaude Opus 5.5 a1816e905e ShadowFinder: add beam-stop holder arms that let part of the beam through
After the existing mask is complete, a second step takes the pixels it left out that are
significantly dimmer than their ring, joins them through a 6 px bridge and across module gaps of
any width, and adds a piece whole when it holds >= 2000 dim pixels of which >= 200 are deep. The
existing mask is never touched, so a sweep with no such piece keeps its mask bit for bit.

Quick subset tests (battery --only, no model check) against the rc173 a0518abe6 full battery:
152 of 233 masks identical; 10 of 10 controls (including the sets where earlier shadow changes
regressed through marginal decisions) give byte-identical merges. On the 21 sets that gain a
piece the space group never changes, merge outlier rejections fall on 19, R_meas falls and ISa
rises on 17, and the shell-scaled model R improves on 16 of 18; a transmitting arm with an
over-subtracted background strip is recovered (R_meas 13.3 -> 12.0 %, ISa 13.4 -> 15.5).
Known costs: on one set with background bumps at ring radii the low-resolution agreement with
the model falls (CC 0.89 -> 0.85) while its internal statistics improve, and two sets cut
slightly coarser (1.50 -> 1.55 A, 1.72 -> 1.80 A) with a better model R.

Squashed from branch hq-beamstop (cfa080a59..60b6c6822), whose history records the variants
tried and dropped: a straight port of the earlier arm finder moved every mask, a two-shape
construction filled almost a whole sweep through hole filling, and a deep-fraction floor
rejected a real arm that is dim along its whole length.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 14:01:17 +02:00
leonarski_fandClaude Opus 5.5 5a9aca682b Rugnux: spherical-harmonic crystal-frame absorption as a candidate surface
A new correction surface, fitted after the time x detector surface and
before the goniometer-frame 8x8 grid: log A is a sum of real spherical
harmonics (l = 1..6, 48 terms) of the diffracted-beam direction de-rotated
into the crystal frame. The incident-beam path depends on phi alone and is
in the per-frame scale already.

It runs through ApplyCellSurface unchanged in everything but the update:
the cells are 32 x 64 equal-solid-angle direction bins, and each round the
per-cell sums (ref2, cross, the same damping) become one ridge-regularised
Gauss-Newton step on the coefficients (prior width 0.1/l per degree-l
coefficient) instead of independent per-cell steps. The half-set Fisher-z
gate adopts or refuses it exactly as it does the grids; where it is
refused, the grid after it sees what it saw before.

Why: the folded 8x8 grid (hemispheres share a cell) is the weak basis for
long-wavelength absorption. Offline, held out by unique reflection, this
basis lowered held-out scatter 4-8% on 6 of 11 long-wavelength sets where
no cell grid did, raised model-phased anomalous peaks 0.02-0.2 sigma, and
was neutral on hard-X-ray controls.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 13:43:59 +02:00
leonarski_fandClaude Opus 5.5 23f446acbb Rugnux: French-Wilson amplitudes as ctruncate gives them
Compared with ctruncate (CCP4 9, version 1.17.29) on rugnux's own merged
intensities of the open battery arm, three differences:

- ctruncate gives no amplitude to an intensity below -3.7 sigma (exactly
  that bound on every set checked; up to 2498 reflections on one set).
  Rugnux turned them into small, confident amplitudes (F/Fc ~0.3 on the
  sets whose background is over-subtracted on powder/ice rings). They now
  get F = NaN (missing in MTZ, '?' in mmCIF); IMEAN is kept. They are
  also left out of the shell mean that sets the Wilson prior.
- The switch to sqrt(I) at I/sigma = 4 left 4-6 sigma amplitudes 2-3%
  above ctruncate's on every set (1.019-1.028). The posterior now applies
  up to 20 sigma (emulated: 0.995-1.000).
- A shell whose mean intensity is not positive gave a prior at the 1e-10
  clamp and amplitudes of ~0 (one set's outer shell); it now takes the
  nearest lower-resolution shell's mean.

The remaining gap (weak amplitudes 2-5% below ctruncate's in the outer
shells) is ctruncate's anisotropy-corrected prior; not attempted. An
offline R-free ablation put the whole ctruncate conversion at -0.0006
median (15/18 sets better) and dropping its rejected negatives at -0.0009
mean (-0.009 on the worst set).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 13:43:59 +02:00
leonarski_fandClaude Opus 5.5 6b061880d8 Rugnux: bin merged reflections by the cell the completeness denominator uses
A merged reflection's d came from whichever frame first observed it, while the possible
reflections per shell are enumerated from the reference cell, so reflections near a shell edge
fell on the other side from their possible twin and shells read 100.1-100.9 % complete (17 of
210 battery sets). The group's d is now computed from the reference cell before the export, and
the per-hand table reads the group's d instead of an observation's.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 13:43:59 +02:00
leonarski_fandClaude Opus 5.5 02bc0b7e0e Rugnux: test angle promotions of the metric symmetry; score model frames at the data's best indexing
The lattice walk admits an angle up to 3 deg from 90 and the constrained refinement then holds it
there. The existing two-arm test (promoted vs demoted class, judged on the held-out residual) only
covered length equalities (a = b). It now also covers monoclinic/orthorhombic classes: pass 1 reads
how far the free refinement leaves a held angle from 90 (AngleEqualityDeparture_deg), turns it into
pixels at the far corner, and where that exceeds the integration disc runs the demoted arm on the
indexer's free (triclinic) cell. A primitive monoclinic crystal with beta 1.4 deg from 90 was
indexed and integrated as primitive orthorhombic on 60/60 validation frames, so the per-frame guard
never fired; demoted, the held-out residual halves (1.54e-6 -> 8.16e-7), ISa 5.5 -> 6.1, d_min
3.70 -> 3.37 A, CC1/2 at 3.9 A 0.54 -> 0.76. Unchanged on five controls (three oP/oI, two near-90
monoclinic).

Model validation scored alternative cell frames against the data in the merged indexing, before
the indexing-ambiguity probe. On a pseudo-orthorhombic P2 crystal whose data need h,-k,-l to match
the model, every frame read random (R 0.77-0.84) and noise picked an axis swap into P1121; each
frame is now scored at its best over the data's alternative indexings. CC_model 0.23 -> 0.45,
R-free 0.66 -> 0.44 on that set.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 13:43:58 +02:00
leonarski_fandClaude Opus 5.5 ae444437f4 rugnux: accept a PDB structure-factor mmCIF as the -z reference
-z (now also --reference; --reference-mtz still works) takes an SF-mmCIF
(e.g. a deposited -sf.cif, gzipped or not) as well as an MTZ. The format is
recognised by content: a file starting with "MTZ " is an MTZ, anything else
is parsed as CIF. The first merged reflection block with the requested
column (or, by default, one the auto choice accepts) is converted to a
gemmi::Mtz in memory with GEMMI's CifToMtz and then read by the unchanged
MTZ loader, so the in-memory reference is exactly what the MTZ path yields.
Unmerged (_diffrn_refln) and anomalous-only blocks are passed over; the log
names the block used.

The R-free set comes from _refln.status (f -> FreeR_flag 0, o -> 1: the
CCP4 convention the loader already reads) and is preferred to
_refln.pdbx_r_free_flag, whose convention varies by program; a status
column with no 'f' is ignored and pdbx_r_free_flag is used instead.

Checked on two open-arm sets in the deposited setting (one with F_meas_au
only, one with intensity_meas): reference loaded with the deposited cell and
group, the inherited free set agrees with the deposited status 'f' on every
common reflection, and the merge correlates at CC 0.99 with the deposition.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-25 20:17:03 +02:00
leonarski_fandClaude Opus 5.5 19f399a973 Rugnux: write the standard setting, or the user's; match a reference MTZ on its own axes
The files were written on the axes the space-group search named, which for
P2221/P21212 puts the unique axis wherever the a<b<c indexing put it (a 52.51
87.87 137.72 crystal came out P 2 21 21 where XDS writes 87.87 137.72 52.51
P 21 21 2), and a reference MTZ was matched in the data's frame: on permuted
axes its free-R flags landed on unrelated reflections while the log reported a
high matched count.

- New CrystalSetting (scale_merge): changes of basis between settings of one
  lattice (cell, group, index operator, basis matrix), the {-1,0,1} det +1
  candidates (CellMappingOperators, moved from ModelValidation), MetricViolation
  (moved from Rugnux), ChooseOutputSetting and SeatGroupByAbsences.
- Output setting: after every decision the merge, the integrated reflections
  (unmerged MTZ), the P1 cross-check, the lattice and the _process.h5 reindex
  matrix are relabelled into the ITA standard setting - or, in priority order,
  a reference MTZ's, a fitting model's, the -C axis order, a non-standard -S
  symbol's. Free-R flags are drawn again on the written axes. Reported as
  SETTING_OPERATOR / SETTING_SOURCE.
- Reference MTZ: the group is kept in its setting; after the merge every cell
  mapping onto the reference cell (times the twin laws) is scored by the
  reference CC, the best is re-seated and re-merged, and the free flags are
  inherited only where CC >= 0.5 over >= 50% of the reference range
  (REFERENCE_MISMATCH otherwise; --mode scale gates the same way).
  REFERENCE_OPERATOR / _CC / _MATCHED_FRACTION / _FREE_FLAGS_INHERITED.
- -S: a fixed group is put on the axes its absences name before merging
  (SeatGroupByAbsences), fixing -S 18 on a cell whose pure axis is not c.
- --model: a model in another setting is now a claim the null tests; where it
  fits, the data are written in its setting and the validation is remade on
  those axes (KeepModelVerdict carries the decisions over).

Tests: [setting] (synthetic #18/#17/I222/C222/c-unique P21/C2 beta/I2->C2/P1,
-C and -S order, absence seating, permuted reference with flags).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-25 20:16:42 +02:00
leonarski_fandClaude Opus 5.5 6c9205a670 Report the strongest index-2 superstructure class and the cell it doubles to
The report-only supercell probe now reaches the results report: SUPERCELL_CLASS (parity of the
primitive indices), its occupancy and Bragg-like (rocking) part against the lattice's own reflections,
its <I/sigma>, and SUPERCELL_DOUBLED_CELL, the Niggli-reduced cell to give with -C. No decision is
taken on it: correct cells whose half-integer class is diffuse, and cells whose depositor kept the
sub-cell, cannot be told from a real doubling by the data alone. On a crystal with a real doubled
axis the reported cell matched the deposited one, and processing on it with -C brought R_free
against the deposited model from 0.60 to 0.29.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-25 08:57:28 +02:00
leonarski_fandClaude Opus 5.5 9e510816e0 Merge hq-cure-merge into hq-pool3; R_MEAS stays the ordinary statistic, the weighted one is R_MEAS_WEIGHTED
R_MEAS is what users set beside XDS/AIMLESS, so its meaning does not change; the merge-weighted value
from hq-cure-merge is reported beside it (R_MEAS_WEIGHTED, REFRES_R_MEAS_WEIGHTED). The per-hand
table keeps the ordinary statistic.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-25 06:05:46 +02:00
leonarski_fandClaude Opus 5.5 2b476336ae rugnux: report-only supercell probe - the eight parity classes of the doubled cell
Every lattice decision counts spots or frames, and the sub-lattice's strong
reflections win every count, so a crystal whose cell is doubled by a weak
superstructure class (9min, 6z9g) is adopted at the half cell. This asks the
question in intensities instead. On 60 frames spread over the sweep, after
each frame's own integration, the 2a x 2b x 2c supercell of its primitive
lattice is predicted to 3 A with the frame's own refined orientation and
geometry and integrated on the same engine; the reflections are summed per
parity class in two shells (20-5, 5-3 A), with a fit of intensity against
partiality (I = a + b p) that separates what rocks like a Bragg reflection
from what sits at the node whatever the rocking. Only the tested frames are
predicted and nothing is retained, so memory is bounded (the previous
prototype predicted the whole run through the merge and ran out of GPU
memory). The probe's integrations are kept out of the engine's own counts,
which the two-pass stencil guard reads. Results are bit-identical with and
without it.

REPORT ONLY - it decides nothing, because on the battery it does not yet
separate a weak real class from what sits at the half-integer nodes of
crystals whose cell is right. Real classes: 6z9g class 101 at 24 % of the
lattice's intensity (29 % rocking), 9min 100 at 19 % (4 % rocking - its real
class does not rock like the lattice either). On correct cells the largest
classes reach 9-12 % raw (7n2s, 9i0a, 7os3) and 3-4 % rocking (7dkp, 7os3),
and 7mzt reads 40 % / 19 % on a class the deposition does not have. The log
line is the population a decision has to be calibrated on.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-25 06:02:50 +02:00
leonarski_fandClaude Opus 5.5 96b681e50d PostRefine: commit the geometry when the held-out residual falls by more than its noise
The joint crystal + detector fit was committed only when the held-out residual
fell by at least a fixed 2 %. On a crystal whose geometry is already right
that bound cuts through the noise: two builds read 1.9 % and 2.05 % for the
same 0.03 px / 0.01 % move (lyso_x06da_half_image), one committed and the
other did not, and everything downstream (smoothing window, stretch
segmentation) followed the coin.

The bar is now the residual's own noise - the standard errors of the two
held-out means combined - which is the bar a round of the geometry walk
already has to clear (HeldOutResidualFell). The log line carries the joint
residual and that noise. Noise on the battery is 2-7 % of the residual, so
the gate is about as strict as before but no longer at a fixed edge.

Checked on 57 open/in-house sets and 11 private ones against the pool
battery. The verdict changed on 8 open/in-house sets besides the passes that
follow the 7ris/8pqd lattice change: half_image, 6qaj, lalanine, 8xtf, 9ig7,
8sqt (commit -> reject of a 1-step move) and myob_x06da_split, 6cdl (reject
-> commit). Merge statistics are equal to the last digit everywhere except
half_image (REFRES R_meas 1.16 -> 0.95, CC1/2 0.79 -> 0.86), 9ig7 (5 fewer
rejected observations) and 8xtf, whose stretch disposition followed the
coin: multiplicity 20.8 -> 15.6, R_meas 0.91 -> 0.60, R_free 0.193 -> 0.189,
REFMAC R_free 0.179 -> 0.176. Private arm unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-25 06:02:50 +02:00
leonarski_fandClaude Opus 5.5 5c5eb00904 Merge statistics: R_meas weighted as the merge weights each observation
The delta-CC1/2 disposition keeps a weak stretch in the merge (downgraded)
wherever removing it would not raise CC1/2, and the merge then carries each
of its observations at the small 1/sigma^2 its scaled-up counting error gives
it. R_meas counted those observations at full weight, so the frames that add
next to nothing to the intensities set the number. hq-pool battery: 8xtf kept
115 weak frames (scale 0.1-0.2 of the run's) with the same CC1/2, <I/sigma>
and CC_model per shell as rc173 with 84 frames rejected, and R_meas was
1.3-1.5x higher in every shell; 9w3y (33 deg rejected -> 0, every model
metric better) and lyso_x10sa_strong read the
same way. The disposition itself is not segmentation-dependent: conviction
is on the batch grid, not the ledger ranges, and those sets' dispositions
changed because the corrected data changed.

R_meas now weights each observation by its merge weight v = 1/sigma^2 under
the error model (corrected_sigma on the host, ModelSigma on the GPU, with the
error model of the last MergeAccum), normalised per reflection to Kish's
effective count, so equal sigmas give the ordinary formula
(WeightedRmeasTerms). Where the proportional term dominates - strong
reflections - frames are weighted alike, as in the merge: 8xtf's lowest shell
reads 15.0%, the rc173 run with 84 frames rejected 15.2%. The per-hand table
is weighted the same way. R_MEAS_UNWEIGHTED / REFRES_R_MEAS_UNWEIGHTED keep
the XDS/AIMLESS convention and are what to set beside XDS; the battery scorer
records them. MULTIPLICITY stays a count.

Effect (weighted / unweighted): lyso_x06da_ref 0.0454 / 0.0479, 8xtf
0.225 / 0.907, insu_I_x06da_ref REFRES 0.072 / 0.232. GPU and CPU paths agree
to the last printed digit on lyso_x06da_ref.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-25 04:52:24 +02:00
leonarski_fandClaude Opus 5.5 94ed6aec04 Scaling: a frame below the credible floor does not hold the loop's step open
ComputeSmoothGWindow leaves a frame below MIN_CREDIBLE_SCALE_RATIO of the
median out of every window, and DropCollapsedScales drops it after the loop,
but RunScalingLoop still counted its step. Such a frame has lost its vote in
its own references and keeps moving, so the loop was stopped as "unsettled".
On 8xtf (hq-pool battery) the first-pass partials loop stopped after 16 rounds
at rms dlogG 1.8e-01 (base 9.5e-04); with the frame left out of the step it
settles in 23 rounds (8.0e-04), and every other loop of that run that had
stopped unsettled (pass 2 at 1.6e-03) now settles too. Every loop of that
sweep that stopped unsettled was one that went on to drop a frame. The final
8xtf merge is bit-identical; the false SCALING_NOT_CONVERGED warning goes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-25 04:44:50 +02:00
leonarski_f 5893a524a5 Merge branch 'hq-errmodel' into hq-pool 2026-09-24 14:36:50 +02:00
leonarski_fandClaude Opus 5.5 aa1f939347 Error model: calibrate each bin on the median NORMALISED deviation, binned by counting I/sigma
The rotation merge fitted var = a*s2 + b^2*<I>^2 from three separate medians (s2, I2, dev2) per
bin of I2. A median of dev2 over observations whose variances differ is not 0.455 times their mean
variance, so the ratio of medians read a too low and b too high: on the scaled fulls of 28 sets
(in-house, open and private) the core of the normalised deviations scattered at up to 1.8x its
stated variance in the weak and middle bins and at 0.1-0.7x in the strongest. Reproduced on
synthetic samples with a known model (a 1.3 read as 1.12, ISa 33 read as 31).

Now (ErrorModel.h/.cpp, host-only, so the GPU and CPU paths share it):
- bins are equal counts in counting I/sigma (I2/s2), where b is identified;
- each bin is calibrated on the median of dev2/var with var from the previous iteration, iterated
  to a fixed point - heterogeneity inside a bin no longer biases it, and the median keeps it robust
  to tails;
- s2 is the counting variance the merge actually applies the model to (rebuilt at the reflection's
  mean), not the observation's own sigma^2.
The separate 6-sigma misfit refit is gone: the median does not need it.

A mean-based fit (misfits cut at z^2 > 2 ln N) was tried first: it calibrates the total variance
best (median rms log chi2 over the 28 sets 0.14 vs 0.19 here) but on heavy-tailed data it sizes
the sigmas on the tails, the merge's outlier test widens with them, and CC1/2 fell 0.80 -> 0.71 on
a powder-contaminated set (0.83 with this fit). Rejected for that.

Offline, 28 sets: rms log chi2 of the median normalised deviation over (counting I/sigma x
resolution) 0.208 -> 0.139 (better on 23), of the mean 0.239 -> 0.191 (better on 22).
Battery, 45 of 48 sets against the 9b6736 run (3 lost to CUDA OOM from GPU contention): space group
unchanged on all; d_min unchanged except two poor multi-lattice sets (1.69 -> 1.56, 1.96 -> 1.90);
ISa x1.13 (median), in-house ISa/XDS 0.73 -> 0.93; ISa*R_meas_lo/0.8 0.93 -> 1.05 (XDS ~1.2);
CC1/2 over the XDS range +0.002 (mean; up 0.014-0.031 on the three poorest sets, else +-0.0001);
CC_model +0.0011, R_model_shell_scaled -0.0007 (mean over 17 open sets); CC_anom +0.003 (mean).
Six private sets: space group, d_min and CC1/2 unchanged, ISa up by 14-67% towards XDS's.

Remaining misfit, not addressed: the excess variance grows slower than <I>^2 (the effective
fractional error falls 2-2.5x from counting I/sigma 5 to 200), so the strongest reflections still
scatter below their sigma on open sets. A third, linear term (as in Aimless) fits it better on most
sets but leaves b unidentified on some (b -> 0 on 4 of 28); not landed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-24 14:32:13 +02:00
leonarski_f c0b7391159 Merge branch 'hq-rotrate2' into hq-pool 2026-09-24 14:06:24 +02:00
leonarski_fandClaude Opus 5.5 e7ef72ba40 Rotation prediction: place each partial at its slice of the rocking curve
A rotation frame records the part of a reflection's rocking curve inside its own oscillation, and
while the crystal turns through the curve the spot walks along its Debye ring. The predictor put
every partial at the exact diffracting condition, so a partial recorded on the curve's flank was
integrated pixels away from where its flux landed. The walk is largest where the reflection moves
nearly tangent to the Ewald sphere (low |zeta|), whose curves are widest.

The prediction now turns S about the beam by the rotation's component along the ring times the
flux-weighted centre of the frame's slice of the curve (a truncated-normal mean, RockingSlice.h),
on the CPU and the GPU predictor alike. 2theta is unchanged; a frame that straddles the condition
symmetrically gets no shift.

Measured (observed r1 centroid minus prediction, tangential, I/sig > 10): against the slice model
correlation 0.93-0.98 before, 0.0 after; residual rms 1.0-1.8 px -> 0.35-0.48 px at |zeta| < 0.3.
Emulated capture of low-|zeta| partials at high angle 0.50 -> 0.94 of a 12 px aperture. CPU runs
on four rotation sets (400-800 frames): R_meas -0.03..-0.09 pp, ISa +1.2..+1.8, rejected
observations roughly halved, space groups unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-24 14:06:24 +02:00
leonarski_fandClaude Opus 5.5 5c28f35f85 Rugnux: report the oscillation the run integrated at; test the rotation-scale fit on a simulated sweep
OSCILLATION_RANGE was read off the caller's experiment, so a run that adopted a goniometer rotation
scale still reported the file's oscillation; it is now scaled by the adopted k (the starting angle is
the stage's and stays). The post-refinement logs the range of k over the leave-a-fifth-out folds
instead of a ratio to k - 1, which is meaningless near k = 1.

PostRefine_RotationScale: partials of a 180 deg sweep simulated at a known stage rate are fitted
back to it (0.97 and 1.0, to 1e-3).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-24 11:37:38 +02:00
leonarski_fandClaude Opus 5.5 032bf9fe2b Rugnux: walk the goniometer rotation scale to its fixed point, decided on the whole sweep
The pass-1 post-refinement fits the rotation scale k only on the frames the stored angles still
track, and a rate error is exactly what stops them tracking the rest: on a sweep whose stage turned
~3 % slow the fit read 0.979 over 94 deg, failed its leave-a-fifth-out test and was thrown away,
leaving half the frames unscaled.

The pass no longer decides. Between the passes, at the pass-2 detector geometry, the lattice is
indexed (index-only probe) under the stored angles and under the fitted k and scored on the
validation frames of the whole sweep: share of the spots on the lattice beyond the wrong-spindle
null. k is adopted only where it scores higher by more than the binomial noise of the two
(ValidationEvidencePrefers - the test the beam-centre arms already used, now one function); the
run then integrates and post-refines at k (post-refine-only probe), fits again on top of it and
repeats until the next k no longer scores better (WalkRotationScale). The stored angles are the
first hypothesis. Measured: 1.000 25.1 %, 0.97874 58.9 %, 0.97041 90.0 %, 0.97006 90.7 % (not
significant) -> 0.97041 adopted.

Probes restore the experiment, the pass-2 geometry, pass-1 mosaicity and the beam-centre-search
flag; a probe opens no beam-centre search. Forced pass-1 results get their axis scaled; the header
revert drops the scale. GONIOMETER_ROTATION_SCALE reports the adopted k (SUSPECT = adopted). The
leave-a-fifth-out figure stays in the log only.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-24 11:34:06 +02:00
leonarski_fandClaude Opus 5.5 a2c2810551 Merge: a clipped observation is no witness in the Wilson test
The pair rule drops the improbable (higher) member of a discordant pair. If the
lower member lost its core to saturation or the mask, its profile estimate of
what remained may be low, and a genuinely strong reflection would be replaced
by the clipped value - the failure Aimless guards against.

Integration now marks a reflection whose signal disk was not fully readable
(Reflection::clipped, from the engines' existing full-disk test); the flag is
carried through the partials to the fulls (OR over an event's partials, CPU and
GPU combine alike), and the Wilson test never counts a clipped observation as
a probable witness against a larger one.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-24 07:26:42 +02:00
leonarski_fandClaude Opus 5.5 257686c6fd Merge: Wilson outlier test
The median outlier test needs three observations, so a reflection measured
once or twice (one good observation and one artefact, the usual case after
Friedel merging) was never tested, and where there were more, a precise-looking
artefact (a hot pixel, thousands of counts) could outweigh mates from weak
frames and become the weighted median itself.

Every observation of the written merge is now also judged against Wilson
statistics beyond 4 A: E^2 = I / (epsilon <I/epsilon>_shell), centric and
acentric laws, a bound set by alpha = 0.01 expected false rejections per
dataset (ln(2N/alpha); twice that for centrics), the lower confidence limit
I - z*sigma (own sigma, z from the same budget) required to exceed it, shells
whose <I> is not established at that significance not judged, and the bound
widened by the dataset's measured tail scale (peaks over threshold on the
well-measured shells; 1 for a Wilson crystal, larger under tNCS/anisotropy).
An improbable singleton is rejected; one with company only when most of its
reflection's other observations are probable and it disagrees with them. The
reflection is the group, or the Friedel pair under -A. Search merges are not
tested (epsilon is 1 in P1, and the screw rows are the reflections a wrong
epsilon would call improbable). The flags are handed to the device merge as
pre-rejections, so CPU and GPU agree.

Report: OBSERVATIONS_REJECTED_WILSON= (part of OBSERVATIONS_REJECTED=), and the
developer report lists each rejected observation (hkl, d, E^2, image, x, y).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-24 07:26:42 +02:00
leonarski_f cab91e2663 Merge branch 'hq-prescan' into hq-integ 2026-09-24 07:26:42 +02:00
leonarski_fandClaude Opus 5.5 23e1976e55 Surfaces: compare held-out half-set CC on Fisher's z; fit modulation, time, goniometer frame in that order
The correction-surface gate averaged the per-shell change of the held-out half-set CC over 20
equal-occupancy shells and adopted a surface on its sign. On the shells where a multiplicative error
shows, CC1/2 sits at 0.99-0.9996 and a large reduction of the error moves it in the fourth decimal;
the shells beyond the data's reach (up to 12 of 20 at CC ~ 0) each add +-0.01, so the sign was set
by noise (measured spread of the mean +-0.001-0.003 against effects of -0.0003..-0.0008). It refused
the 24x24 detector surface on four EIGER2 16M sweeps. The change is now averaged on atanh(CC) - the
candidate from the hq-bisect investigation.

The order of the three overlapping surfaces also decides what is adopted. With the gate fixed,
goniometer-frame first took a 5 keV insulin sweep from ISa 31.8 to 20.1 (the time surface then
refused; R_meas 7.8 -> 8.9%, CC to the model 0.8307 -> 0.8282) and cost a 0.1 deg-sliced thaumatin
sweep R_meas 18.2 -> 19.6%; time x detector first refused the detector surface on the 16 keV
thaumatin sweep (ISa 33 instead of 40). Modulation, time, goniometer frame avoids both.

Bare runs, --report-resolution at XDS's range, against rc173 (0252880e0): ISa / lowest-shell R_meas
thaumatin 16 keV 32.9/.030 -> 39.8/.028, lysozyme 90 deg sweeps 10.1/.059 -> 14.1/.048 and
10.1/.060 -> 13.8/.049, cytochrome c 19.2 -> 20.4 and 17.3 -> 19.3, thaumatin 3.8 keV 13.3 -> 14.8;
CC to a deposited model (standard-protein models on 14 in-house sets, the deposition on 13 open sets)
within +-0.0045, R_model_shell_scaled within +-0.002 except where d_min moved. 46 sets, one regression:
a half-image lysozyme test set, CC1/2 0.878 -> 0.862.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-24 02:13:33 +02:00