Commit Graph
1255 Commits
Author SHA1 Message Date
leonarski_fandClaude Opus 5.5 023afc4ee2 HotPixelFinder: read a ring's median and spread off a histogram of its counts
Each frame packed every ring's pixels together and selected in them twice - the median, then the
median absolute deviation - about 40% of the pass. A ring's background is a few counts, so both are
now read off a histogram of the values 0..1023, exactly, wherever the statistic lands inside it (and,
for the spread, no count is negative); otherwise the ring is packed and selected as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 23:03:21 +02:00
leonarski_fandClaude Opus 5.5 ca98ca8397 HotPixelFinder: keep the level and valid-frame sums per ring-sector, not per pixel
Every valid pixel of every sampled frame read and wrote its valid count and level sum, 20 of the 44
bytes the accumulation moved per pixel, on a pass that is bound by memory bandwidth. Both depend on
the pixel only through its ring-sector and its own error frames: they are now summed once per frame
per ring-sector, and a pixel keeps only what its error frames take out of them. Integers, so the same
sums.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 23:02:14 +02:00
leonarski_fandClaude Opus 5.5 d3afe0cb07 ImageSpotFinderCPU: the first pass slides its window only where its result is read
With candidates, the first pass already tested only the 32-column blocks within reach of a candidate,
but still slid the horizontal window across every column of every row. It now slides it only over
the runs of needed blocks, starting each run from the vertical sums it covers - integers, so the same
window sums - and sets the bits there directly; the rest stay 0 as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 23:01:21 +02:00
leonarski_fandClaude Opus 5.5 5b2b642e73 Rugnux: run the pre-scan's beam-centre capture beside the first pass
The capture (about 1.5 s on a 16M detector) sat at the end of the pre-scan, but nothing reads its
answer before the first pass's beam-centre check. It now runs on copies of the experiment, mask and
projection it measured, and is taken (JoinBeamCenterCapture) wherever background_center_ or
measured_beam_center_ is read. The first-pass memo key counts a capture still running as a centre
present, which it will be by the time the pass reads it. Its log lines now come from its own thread,
among the first pass's.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:58:42 +02:00
leonarski_fandClaude Opus 5.5 7e1be959a6 BraggIntegrationEngineCPU: the background clip walks the ring pixels pass A read
The sigma-clip pass recomputed the stencil distances over the whole box to find the same background
ring pixels pass A had just summed. Pass A now keeps their values and radial offsets in its own
order, and the clip runs over them - the same pixels, the same sums.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:56:29 +02:00
leonarski_fandClaude Opus 5.5 967d5691cc BraggIntegrationEngineCPU: gather the profile fit's pixels once, not on every reweighting
The four Kabsch reweighting iterations each walked the reflection's grid again with the same bounds,
validity and ownership tests the p_valid pass had just applied. That pass now keeps the profile value
and background-subtracted count of the pixels the fit reads, in grid order, and the iterations run over
them - the same terms summed in the same order (about 6 s of a CPU-only 16M run).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:55:48 +02:00
leonarski_fandClaude Opus 5.5 a6c27e9a04 ShadowFinder: the element-wise full-detector mask updates on all threads
Four per-pixel passes over char masks ran on the pre-scan's critical path on one thread; each
pixel's result depends on that pixel alone, so splitting them changes nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:55:17 +02:00
leonarski_fandClaude Opus 5.5 2746c50822 BeamCenterFromBackground: skip pixels clearly outside the band before sqrt and atan2
60-70% of a large detector lies outside the resolution band the walk bins, and every one paid a
square root and two arctangents on every iteration. A test on tan(2theta) = rho / lz with a 0.1%
margin, far above float rounding, drops them first; every pixel the exact test keeps still reaches it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:54:37 +02:00
leonarski_fandClaude Opus 5.5 02063960d6 Rugnux: answer an indexing probe from the first-pass memo before building the pass
A probe whose inputs a first pass has already run on returned that pass's evidence only after the
pass had built its azimuthal mapping, writer setup and indexer (0.15 s each on a 16M detector, twice
at the end of a run). The lookup now comes first; nothing in between changes what its key reads.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:54:01 +02:00
leonarski_fandClaude Opus 5.5 7adf8121a8 Rugnux: key the first-pass spot cache without the goniometer's scan
Spot finding reads whether there is a goniometer, never its axis or angles, and the rotation-scale
walk's probes change only those - so each probe found every frame's spots again (0.2-0.3 s on a 16M
sweep). RUGNUX_VERIFY_FIRST_PASS_MEMO still recomputes and compares.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:53:21 +02:00
leonarski_fandClaude Opus 5.5 b119b19707 BraggIntegrationEngineGPU: grow the per-reflection buffers to twice the record
Each growth is 11 pinned host and 19 device allocations under the driver's device-wide lock; at the
start of an image loop 16 workers growing 1.5x at a time spent about 0.3 s in them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:52:27 +02:00
leonarski_fandClaude Opus 5.5 cfe202e9a1 ScaledObservations: sample each image on its own thread
The ASU-key sample over every partial ran on one thread (0.2-1 s of the merge tail). Each image's
parts are now taken on the workers and joined in image order, so the sort sees the same parts in the
same order.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:52:09 +02:00
leonarski_fandClaude Opus 5.5 dd02e65693 PostRefine: rank the observation cap on I/sigma worked out once, not per comparison
nth_element compared through pointers into gigabytes of partials, a cache miss per comparison on one
thread (0.3 s on a 16M sweep). The ratio is now computed beside each pointer on all threads; the
comparisons and so the selection and its order are the same.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:51:31 +02:00
leonarski_fandClaude Opus 5.5 2c48ed3fcd RotationScaleMerge: measure the rocking-event span once, not twice per Run
At the smoothing window every Run has just restored corr to what Ingest built, and nothing else the
measurement reads changes after Ingest, so its answer is the same on every Run; it walked all the
partials twice per Run on the CPU path.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:50:11 +02:00
leonarski_fandClaude Opus 5.5 45a8db03d4 RotationScaleMerge: without a GPU, fold the partials' corr update into the rescale pass
RunScalingLoop rescales after every iteration before anything reads corr, so the iteration's
UpdateCorr and the rescale are one pass over the partials instead of two: the G and fitted frames
as the iteration left them, then the ratio on top - the same two roundings in the same order.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:49:32 +02:00
leonarski_fandClaude Opus 5.5 6df56e661b RotationScaleMerge: without a GPU, sum the partials' group means on all threads
ReduceGroupMeans ran once or twice per scaling iteration as one serial scatter over every partial
(about 8 s of a CPU-only run's main thread on a 16M sweep). ComputeAsuGroups already builds the
partials' group CSR - a stable counting sort, observation order within each group - for the GPU
reduction; it is now kept on the CPU path too, and each group is summed over it in the order the
serial loop added it, so the means are the same numbers.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:48:41 +02:00
leonarski_fandClaude Opus 5.5 698d06a0e2 BraggPredictionRot: reject solutions far from the frame before the rest of the arithmetic
zeta >= min_zeta, so the rocking-curve test rejects every solution with
min_zeta * (|phi| - half wedge) > multiplier * mosaicity. Tested right after phi, with a 1e-5 rad
margin (float rounding of the full test is below 1e-6), it skips the cross product, normalisation
and zeta of the 85-95% of solutions no frame keeps, and cannot reject one the full test would pass.
CPU build: on the order of 10 s of a 1.2 A sweep's prediction.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:47:12 +02:00
leonarski_fandClaude Opus 5.5 79bc90e512 RotationIndexer: keep 1024 indexing outcomes, keyed by a digest of the input
A run on data that index poorly asks over a hundred indexing questions - every rung of the
spot-budget ladder, in every pass and walk probe - and a later pass repeats a probe's ladder, which
the 32-entry memo had long evicted (about 20 s on one such sweep). The key holds every spot, so it is
now kept as two independent 64-bit hashes and its length instead of whole.
RUGNUX_VERIFY_FIRST_PASS_MEMO still recomputes and compares.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:46:20 +02:00
leonarski_fandClaude Opus 5.5 92877e5eb7 Rugnux: build the unmerged MTZ beside the P1 cross-check merge
The unmerged file reads the integration outcomes and the determined group, neither of which the P1
merge changes, except each image's mosaicity, which the batch headers carry and the merge rewrites.
UnmergedMtz builds the file without it on a second thread, and SetUnmergedMtzMosaicity fills it in
after the merge, so the file is the same bytes as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:45:48 +02:00
leonarski_fandClaude Opus 5.5 ab52c12137 Rugnux: measure CC1/2 before the surfaces only on the merge the quality guard reads
Every non-search merge of a rotation pass spent a whole extra merge on cc_half_before_corrections,
but only the pass's first merge (search_merge_cc_half) is read; the ProcessResult copy was never
read at all and is removed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 22:43:30 +02:00
leonarski_fandClaude Opus 5.5 b90b377b85 rugnux --mode scale: hold the merge cell by value, not a reference into a temporary
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 9m52s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 10m45s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 17m38s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 20m23s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m17s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 7m44s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 7m11s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 11m3s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m26s
Build Packages / Unit tests (push) Successful in 1h36m59s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 17m34s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 17m5s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 17m39s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 17m48s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 18m44s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 20m29s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m52s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 18m51s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 22m11s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 23m30s
Build Packages / Generate python client (push) Successful in 57s
Build Packages / Build documentation (push) Successful in 2m10s
Build Packages / Create release (push) Successful in 18s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 12m57s
GetUnitCell() returns std::optional by value, so a reference to its value() dangled at the end of
the statement (clang -Wdangling-gsl).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 21:54:58 +02:00
leonarski_fandClaude Opus 5.5 40d91072ba CUDAWrapper: spell out lock_guard's mutex type
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 21:49:41 +02:00
leonarski_f 13c03e1ef4 Merge branch 'perf-rotation' into rc173
# Conflicts:
#	common/ParallelFor.h
2026-09-26 21:48:30 +02:00
leonarski_fandClaude Opus 5.5 90be01a6cc CI: order jobs viewer, rugnux, jfjoch, release, each under a section header
Jobs moved as whole blocks; no job's content changes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 21:47:56 +02:00
leonarski_f bc9b5bb628 Merge branch 'fix-macos-build' into rc173 2026-09-26 21:46:49 +02:00
leonarski_fandClaude Opus 5.5 8a767d03c0 Rugnux: the P1 cross-check does not measure CC1/2 before its surfaces
Every rotation merge measured CC1/2 before the correction surfaces - a merge
of its own - and the P1 cross-check's value was never read. scale_and_merge
takes whether it is wanted, and the cross-check says no.

md5-identical p.hkl, p_P1.mtz, p_unmerged.mtz and report on four sets
(one log line fewer); ~0.3-0.6 s on the sets with large merges.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 21:23:04 +02:00
leonarski_fandClaude Opus 5.5 cd98787728 CPU analysis: take the azimuthal profile in the adaptive finder's ring pass
On the CPU path every image made a separate azimuthal-integration pass
(AzIntEngineCPU) over the 72 MB frame although the adaptive finder's first
ring pass reads the same pixels in the same order under the same rules
(skip the INT32_MIN/MAX sentinels, bins below the mapping's count). That
pass now also accumulates the corrected profile - the same statements as
AzIntEngineCPU, so the same float sums - and MXAnalysisWithoutFPGA takes the
profile from the finder instead of running the separate pass, as the fused
GPU engine already does. Only where the azimuthal engine would be the CPU
one; the finder the pre-scan uses does not accumulate it.

md5-identical p.hkl, p_unmerged.mtz, p_plot.txt and report; CPU-only
163 -> 149 s and 94 -> 91 s. GPU unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 21:18:04 +02:00
leonarski_fandClaude Opus 5.5 d0d1985ed0 Rugnux: run the leftover-lattice census beside the image loop
The census (report only: what the crystal's lattice leaves over on the
scheme and validation frames, up to three further lattices) ran on the
critical path of the canonical pass between the first pass and the image
loop - ~1.5 s on the GPU build for a set where it finds lattices.

Its spots are still found in place; the rest now runs in the background on
copies of what it reads (the experiment, those frames' spot lists, the
lattice, the validation settings) and is collected into the result before
the pass returns. It no longer runs on a pass that only post-refines, whose
result nothing reads.

md5-identical output and identical report on four sets; GPU 24.1 -> 22.8 s
on a set with leftover lattices.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 21:08:18 +02:00
leonarski_fandClaude Opus 5.5 c3c7e179c7 Rugnux: keep the first pass's spot lists across passes
Every first pass found the spots of its ~200 scheme and validation frames
afresh, although the canonical pass follows a probe on the same frames under
the same settings, and a pass repeating an earlier pass's rescue ladder asks
for the same frames again: on the CPU-only build that is several seconds per
first pass.

The spot lists are now also kept in a store shared by the run and its
copies, keyed by everything a frame's list depends on (SpotFindingKey: the
experiment key - geometry, goniometer, spot budget, settings - the rest of the
spot-finding settings including the measured rings, the ice-ring switch, and
the checksums of the pixel mask and of the spot mapping). A first pass takes
what an earlier one found under the same key. RUGNUX_VERIFY_FIRST_PASS_MEMO
finds the spots again and throws on any difference, field by field.

md5-identical output (four sets GPU, two CPU-only); CPU-only 169 -> 163 s and
96 -> 94 s, GPU 23.0 -> 22.4 s and 119 -> 116 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 21:01:54 +02:00
leonarski_fandClaude Opus 5.5 ef4090c830 CPU FFT indexer and beam-centre shortlist: less serial work
- FFTIndexerCPU::ExecuteFFT built the per-direction histograms and scanned the
  per-direction spectra for their most prominent peak in one thread; every
  direction has its own histogram and spectrum, so the directions are now
  split over the refinement threads, each filled and scanned in the same
  order as before.
- BeamCenterShortlist2D scanned the whole padded surface (73 M points on a 16M
  detector) once per candidate. It now keeps each row's maximum and the first
  index holding it and rescans only the rows a suppression touched; rows in
  order, first index within a row, is the same first maximum.

md5-identical output; CPU-only 177 -> 169 s and 101 -> 96 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 20:48:08 +02:00
leonarski_fandClaude Opus 5.5 87cfc879b6 CPU adaptive spot finder: sigma clips from a histogram, first pass only where read
Two whole-image passes per frame out of the CPU spot finder (the pre-scan's
finder on every build, and every image on the CPU-only build):

- AccumulateRings ran three passes over the frame - the plain ring statistics
  and two sigma clips. The plain pass now also counts each ring's valid
  values in a histogram (0..1023, the rest in a short list), and the clip
  passes sum over the distinct values: each meets the same float test its
  pixels would, and the sums are integers, so the totals are the same.
- The local test's first pass is read by DetectAt only inside a candidate's
  window. It now marks the row/32-column blocks those windows reach and
  keeps its sliding sums everywhere but skips the per-pixel test elsewhere;
  the bits it leaves unset are never read.

md5-identical output on four sets (GPU) and three (CPU-only). CPU-only
16M: 203 -> 177 s and 303 -> 263 s; GPU 16M 23.8 -> 22.8 s (the pre-scan's
finder).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 20:40:46 +02:00
leonarski_fandClaude Opus 5.5 7ce2993ae9 Pre-scan: measure the spots beside the shadow, defect and beam-centre steps
On a 16M rotation run the pre-scan was ~11 s on the critical path, the first
3 s of it the spot-width tiers (CPU spot finding on the sample frames),
which the shadow, the defective-pixel mask and the beam-centre capture
waited for although none of them reads what the tiers measure.

The sample is now read twice. The first pass builds the projection and
nothing else; the shadow, the diagnostic, the defective pixels and the
capture start as soon as it is in. The second finds the spots - the width
tiers exactly as before, on copies of the experiment and the pre-shadow mask
it always read - in the background beside them, and reads only the frames
it has something to measure on. What it measured (radii, bandwidth, powder
rings, spot quantiles, the spot-symmetry pool) is applied once both are done,
before anything that reads it. The diagnostic JPEG is rendered in the
background from copies too.

md5-identical p.hkl / p.mtz / p_P1.mtz / p_unmerged.mtz, report and
p_detector.jpg on four sets. GPU 26.2 -> 23.8 s and 38.3 -> 36.4 s on two
16M sets; CPU-only 210 -> 203 s and 114 -> 103 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 20:15:18 +02:00
leonarski_fandClaude Opus 5.5 94505d10c4 Merge tail: parallel delta-CC1/2 ranges, twin-immune zone evidence and P1 tNCS
Three serial stretches of the canonical pass's tail, each made parallel with
the same arithmetic in the same order:

- MeasureBatchDeltaCCHalf measured every batch of the curve, every open
  candidate of the rejection loop and every ledger range one after another.
  Each measurement is a pure function of its range and the fixed totals, so
  they now run side by side (measure_with, one scratch set per worker) and the
  decisions scan the results in the original order; the edge walk (locate)
  stays one at a time.
- The twin-immune zone evidence (CentricOverAcentric, a few hundred
  exponentials per reflection) is evaluated in parallel and summed in the
  original order.
- The P1 cross-check's AnalyzeTranslationalNCS was still called with one
  thread; it gets the run's thread count like the other two calls.

md5-identical p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz and report on four sets;
GPU wall 41.0 -> 38.3 s and 26.9 -> 24.8 s on the two sets with long merges.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 20:00:27 +02:00
leonarski_fandClaude Opus 5.5 a6bb16ce8c RotationIndexer: keep what RunIndexing computed, keyed by everything it reads
A rotation run asks the rotation indexer the same question many times: the
canonical pass's first pass repeats the rotation-scale probe at the stored
angles exactly (identical validation evidence on every set checked), and on
a crystal that does not index, pass 2 and every probe repeat pass 1's rescue
ladder rung for rung. Each such RunIndexing is an FFT search plus a serial
Ceres fixed-point chain, ~1-3 s on the GPU build.

RunIndexing is deterministic in its inputs, so its outcome - every member it
sets - is now kept process-wide under a key of all of them: the accumulated
spots (every field) and their angles, both geometries and the axis, the
experiment's indexing settings, cell and space group, and the settings of the
pool it indexes with (IndexerThreadPool::Settings). A RotationIndexer asking
with the same key takes the outcome. RUGNUX_VERIFY_FIRST_PASS_MEMO recomputes
and throws on a difference.

md5-identical output on four sets; GPU wall 28.9 -> 26.1 s, 43.9 -> 41.0 s,
30.3 -> 26.9 s and 143.6 -> 120.5 s (the set that does not index); the
verify mode found no difference on the last.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 19:52:46 +02:00
leonarski_fandClaude Opus 5.5 c7a8fabc7e BeamCenterFFTCPU: run the capture's independent 2-D transforms at the same time
PointSurfaces made three forward and four inverse transforms of the padded
detector (8748 x 8400 on a 16M) one after the other on the main thread -
~3.5 s of one core in a CPU-only run. The forwards do not depend on each
other, nor do the inverse products; each group now runs at the same time,
one workspace per transform. A workspace is built as the single one was -
its own std::vector buffers and its own FFTW_ESTIMATE plan for them - so the
same plan runs on the same data and the result is bit-identical. Costs about
2 GB more transient memory on a 16M detector.

CPU-only 16M rotation run: md5-identical, 216 -> 210 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 19:39:56 +02:00
Filip LeonarskiandClaude Opus 5.5 3c7efcd1e5 FFTW: build the NEON codelets on aarch64
Build Packages / Create release (push) Successful in 48s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 2m56s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m26s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 23m45s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 16m46s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 23m7s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 17m4s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 6m30s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 6m58s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 7m11s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 8m19s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 10m27s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 15m34s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 16m17s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 16m26s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 16m14s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 15m48s
Build Packages / Generate python client (push) Successful in 14s
Build Packages / Build documentation (push) Successful in 48s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 15m22s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 15m37s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 14m39s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 16m28s
Build Packages / Unit tests (push) Successful in 1h47m0s
FFTW's CMake build knows only SSE/SSE2/AVX/AVX2; NEON exists only in its
autotools build (--enable-neon). The ENABLE_NEON this file set was silently
ignored, so on Apple Silicon and the Linux aarch64 build FFTW ran its scalar
codelets (config.h: HAVE_NEON undefined, no simd/neon objects). The NEON
sources are now added to fftw3f the way its CMakeLists adds the SSE2/AVX
ones, with HAVE_NEON as a target definition (the config.h template carries
it only as a comment). No compiler flag: NEON is part of every aarch64 CPU.
FFTW 3.3.11 would not help - its CMake build has no NEON either, and its
Apple ARM cycle counter only matters to FFTW_MEASURE planning, while every
plan here is FFTW_ESTIMATE.

FFTW's runtime NEON probe on unix/linux executes .long 0xf2000150; on
aarch64 that decodes as ands x16, x10, #0x100000001 and runs without
SIGILL (checked on Apple Silicon), so the codelets are used on Linux too.

Measured on Apple Silicon against the scalar build: batched 1D r2c 1.8x
(power-of-two length) and 1.3x (730), 2D r2c 1024^2 1.9x, 3D c2r 128^3
2.6x, spectra equal to 1.6e-8 relative. rugnux -X FFTW on 360 frames of
the rotation test set: merged .hkl byte-identical, report identical; CPU
time unchanged (300 s), so FFTs are a small share of that run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-26 19:35:07 +02:00
Filip LeonarskiandClaude Opus 5.5 48ce0dd570 Docs: macOS artefacts - requirements, first open, building from source
RELEASE_CONTENTS gains the macOS .dmg and rugnux .tgz in the artefact
table, the CPU floor (any Apple Silicon Mac), the OS floor (macOS 13), the
CPU-only note in the CUDA/GPU tables, and a macOS section: Apple Silicon
only (Rosetta does not run arm64 code on Intel), drag-to-Applications,
notices inside the bundle, and how to open the not-yet-notarized release
(Open Anyway on macOS 15+, Control-click Open on 13/14, or xattr).
JFJOCH_VIEWER states the platforms and requirements, that the Mac build is
CPU-only, that D-Bus is Linux-only, and adds Building from source on macOS.
RUGNUX_INSTALL adds the macOS archive and the quarantine note - checked: a
browser-downloaded .tgz hands its quarantine flag to everything tar
extracts, Gatekeeper rejects rugnux, and xattr -dr clears it. DEPLOYMENT
points to the pre-built Windows/macOS viewers and the macOS rugnux archive.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-26 19:35:07 +02:00
leonarski_fandClaude Opus 5.5 4b1d7feebd GPU threads: block instead of spin, and stay on the GPU's NUMA node
- set_gpu_blocking_sync(): every device is put in
  cudaDeviceScheduleBlockingSync before its context exists, so a host thread
  waiting on the GPU sleeps instead of spinning on a core. On a 16M rotation
  run a fifth of all CPU time was that spinning; wall time unchanged within
  noise. Called first thing in rugnux.
- enable_gpu_numa_binding(): from then on pin_gpu() (and the new
  pin_gpu(dev), used by the first-pass spot workers that take a card by
  index) also keeps the thread on the CPUs of the NUMA node the card hangs
  off. The node and its CPUs come from /sys (no libnuma), intersected with
  the process's own mask; Linux only, and nothing happens on a machine with a
  single node. rugnux turns it on; the broker does not.
- A thread inherits its creator's affinity, so the shared ParallelFor pool
  would run every later pass on one socket if a pinned worker created it:
  its threads now reset to the mask the process started with
  (common/ThreadAffinity).

Byte-identical output. The NUMA part is a no-op on the single-node test box
and still has to be measured on a two-socket machine.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 19:30:01 +02:00
leonarski_fandClaude Opus 5.5 25ea9bf555 Rugnux: run independent probe passes beside each other, on copies of the run
The rotation two-pass runs several first-pass indexing probes one after the
other although some of them do not depend on each other:

- The rotation-scale walk always indexes at the stored angles and then at
  the pass-1 fit; both are known before it starts. The probe at the fit now
  runs at the same time, on a copy of the run as it stands before either.
  Its answer is taken only if the probe at the stored angles left behind
  none of the state the next pass reads (the spot-finding settings a
  first-pass ladder rung adopts; each probe puts the experiment back
  itself) - otherwise it is run again in sequence, as before.
- The geometry walk's indexing probe at the canonical pass's post-refined
  geometry depends on nothing that pass does after its post-refinement. It
  is now started there, on a copy, and runs beside the pass's scaling, merge
  and output; its evidence comes back through the first-pass memo (now a
  short list) and the walk's own probe takes it through the usual key check,
  so RUGNUX_VERIFY_FIRST_PASS_MEMO covers it too. It is started only for the
  passes whose post-refinement the walk probes with an indexing pass.

The copies are plain copies of Rugnux: the cancel flag becomes a shared
flag (so cancelling the run cancels them) and everything else was already
copyable. Probe passes run on a copy get no observer.

md5-identical output on two 16M sets; the walk's two probes take ~2.3 s
together instead of ~3 s, and the last probe of the geometry walk is hidden
behind the merge.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 19:30:01 +02:00
Filip LeonarskiandClaude Opus 5.5 f2985edcd8 macOS minimum 13.0; CI job names as build:<product>:<platform>:<variant>
Build Packages / Create release (push) Successful in 24s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m19s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m25s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 10m34s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 17m26s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 8m45s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 6m36s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 6m30s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 20m1s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 24m57s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 11m34s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 14m13s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 14m3s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 14m22s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 12m53s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 13m18s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m5s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 17m4s
Build Packages / Generate python client (push) Successful in 41s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 18m58s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 18m50s
Build Packages / Build documentation (push) Successful in 1m6s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 15m15s
Build Packages / Unit tests (push) Successful in 1h48m38s
CMAKE_OSX_DEPLOYMENT_TARGET 12.0 -> 13.0: the Qt 6.11 frameworks the .dmg
bundles are built for macOS 13, so the app could not start on 12 anyway, and
the linker warned about every framework ("building for macOS-12.0, but
linking with dylib ... built for newer version 13.0").

CI job names follow one scheme - build:viewer|rugnux|jfjoch, then the
platform (<os>-<arch> for the portable products, the distribution for the
packages), then cuda / nocuda (cuda-sls9 for the slsDetectorPackage 9
builds). The Linux viewer matrix's "cpu" variant is now "nocuda" like the
Windows one; the RPM/DEB matrix gains platform and variant beside the
distro key its upload step tests. Only display names change: job ids, and
so needs:, and artifact file names are as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-26 19:21:32 +02:00
Filip LeonarskiandClaude Opus 5.5 c516ac2c51 macOS packaging: working DMG, licences in the bundle, rugnux tgz, CI jobs
Build Packages / Create release (push) Successful in 15s
Build Packages / build:macos:viewer (arm64) (push) Successful in 3m5s
Build Packages / build:rugnux:macos (arm64) (push) Successful in 2m21s
Build Packages / build:rugnux:aarch64 (cross) (push) Successful in 8m8s
Build Packages / build:rugnux-tgz (x86_64) (push) Successful in 9m14s
Build Packages / build:viewer-tgz:cpu (push) Successful in 10m47s
Build Packages / build:viewer-tgz:cuda (push) Successful in 12m23s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 15m34s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 16m2s
Build Packages / build:windows:nocuda (push) Successful in 17m42s
Build Packages / build:windows:cuda (push) Successful in 20m14s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m17s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 18m28s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 18m48s
Build Packages / Generate python client (push) Successful in 40s
Build Packages / build:rugnux:windows (push) Successful in 11m3s
Build Packages / Build documentation (push) Successful in 1m17s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 19m27s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 19m53s
Build Packages / build:rpm (rocky8) (push) Successful in 18m21s
Build Packages / build:rpm (rocky9) (push) Successful in 18m25s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 18m49s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 19m36s
Build Packages / Unit tests (push) Successful in 2h11m55s
The .dmg cpack produced did not start: jfjoch_viewer links Qt as
@rpath/Qt*.framework, install strips the build tree's rpath, and no install
rpath was set, so dyld found none ("no LC_RPATH's found"); macdeployqt's
"Cannot resolve rpath" errors were the same gap. INSTALL_RPATH is now
@executable_path/../Frameworks. Info.plist had an empty identifier and
version; it now carries ch.psi.jfjoch.viewer, the numeric version (the rc
suffix in the long version string) and the name "JFJoch Viewer".

On macOS the notices go into jfjoch_viewer.app/Contents/Resources instead of
a share/ folder beside the app in the .dmg, which a user dragging the app
out would leave behind. They are installed before the Qt deploy script,
which signs the bundle - a file added afterwards invalidates the signature.
JFJOCH_NOTICE_FILES is set before the subdirectories for that.

Names: jfjoch-viewer-<version>-macos-<arch>.dmg (volume "Jungfraujoch
Viewer <version>"), and a rugnux-only build on macOS packs as
rugnux-<version>-macos-<arch>-cpu.tar.gz instead of claiming linux.

CI: build:macos:viewer and build:rugnux:macos on a host-mode runner
labelled macos-arm64. Each checks its artifact (the bundle finds its own
Qt, holds no path into the build machine and is validly signed; rugnux
links only system libraries and starts) and uploads it on a release build.
The Developer ID signing / notarytool / stapler steps are left as a comment
until the project has an Apple Developer account.

Verified locally: the DMG mounts, the app starts from it and loads every Qt
framework and plugin from inside the bundle; the rugnux tarball runs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-26 18:59:46 +02:00
leonarski_fandClaude Opus 5.5 59d92a7238 ParallelSort for the two large sorts on the merge path
Two whole-dataset sorts sat on the main thread at the end of a rotation run:
WilsonOutliers orders every full by resolution, and the unmerged MTZ is put
in H K L M/ISYM BATCH order by Mtz::sort(5) - together about 2 s of one
thread on a 1.4 M-observation set.

ParallelSort (common/ParallelFor.h) sorts one piece per worker and merges
them pairwise. It is only for comparators that are a strict total order,
where the sorted sequence is unique and the result is the serial sort's bit
for bit: WilsonOutliers already breaks ties on the index, and the unmerged
writer now sorts the rows itself on the five key columns and then the row
number - the order Mtz::sort's stable sort gives - and sets sort_order as it
did. An empty table still fails the way Mtz::sort does.

md5-identical p.hkl, p.mtz and p_unmerged.mtz; 44.4 -> 43.6 s on a 16M set.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 18:13:41 +02:00
leonarski_fandClaude Opus 5.5 c244d384b0 tNCS: evaluate the class-vector ladder's neighbours in parallel
Where a pseudo-translation is detected, the class vector is re-refined up a
ladder on the whole data by a greedy sweep over the 26 neighbours of the
current vector, each a cosine correlation over up to every acentric
reflection - seconds of one thread on the main path of the run (3.4 s and
1.5 s for the two calls on a P3_121 16M set).

The neighbours of the current vector are now evaluated together; the first
improvement in the sweep's own order is taken, and the neighbours after it
are evaluated again about the new vector - exactly the moves the serial
sweep makes. Each correlation is still one serial sum, so the result is
bit-identical. AnalyzeTranslationalNCS takes the thread count.

md5-identical output and report; 47.1 -> 44.4 s on that set.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 18:08:36 +02:00
Filip LeonarskiandClaude Opus 5.5 22d4e94bd8 Viewer: inspector width from its content, fold it away on a narrow window
The inspector's width bounds were 75-100 average character widths, which on
macOS (8 px) came to 600 px for content that needs 397. They are now
measured: at least the panel's minimum size hint plus the scroll bar, at
most a third more. CollapsibleSection reports its folded content's width
in minimumSizeHint, so the panel keeps its width when a section opens. On a
font change the measurement waits a turn, until the children have the new
font.

When shrinking the window would leave the image narrower than the
inspector, the inspector folds away, with the magnifier if it shares the
inspector's column; they come back once the image would still be at least
as wide as the inspector (plus 40 px of hysteresis). Only a shrink folds, a
dock shown by anyone else is no longer the window's to restore, and folded
docks are saved as open. Measured on a 1920 px screen: folds at 1140 px,
returns at 1260 px; at 960 px the image gets 620 px instead of ~145.

Dark theme: disabled text is set explicitly (#858AAE, ~4.8:1 on the panels)
instead of the derived colour, which was nearly the panel colour and made
greyed-out fields look empty. The splash text is black in either theme:
its message is rich text, which takes the palette's text colour rather
than showMessage's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-26 18:04:42 +02:00
Filip LeonarskiandClaude Opus 5.5 9f369567f1 Viewer: font size and theme buttons on the display toolbar, matching title bar
The right end of the display toolbar now carries three "A" buttons for the
100/125/150 % font sizes and a light/dark switch (a sun with a half-filled
disc), so both settings are visible without opening the View menu. The size
buttons and View > Font size stay in sync through JFJochViewerMenu's new
SetFontZoom / fontZoomChanged; the View > Theme tick follows ThemeNotifier,
so it moves when the toolbar switches the theme.

With Qt 6.8+ an explicit Light or Dark choice is also requested from the OS
(QStyleHints::setColorScheme), so the title bar macOS and Windows draw
matches the window, as apps with their own theme switch do. Follow system
withdraws the request first, since Qt otherwise reports the requested scheme
instead of the desktop's.

The image counter reads "1 / 1800" instead of "1/1800".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-26 17:51:50 +02:00
Filip LeonarskiandClaude Opus 5.5 0be8eb2bca Viewer: dark theme, switchable live from the View menu
View > Theme offers Follow system / Light / Dark, stored in the settings.
Follow system uses the desktop's colour scheme where Qt reports it (6.5+)
and is light otherwise. The dark theme is a deep indigo (#1A1D3A panels,
#12142B entry fields and charts).

The palettes live in the new ViewerTheme module. A switch applies at once:
the application stylesheet is set again after the palette so every widget
is re-polished (with a stylesheet active, polished widgets otherwise keep
their old palette); widgets that bake colours into stylesheets use
SetThemedStyleSheet, which re-applies them on ThemeNotifier::changed;
toolbar icons draw their glyph at paint time through an icon engine, so the
ink follows the theme (and stays sharp at any size); charts re-theme and
rebuild. Hard-coded white backgrounds on entry fields and tables are
removed in favour of the palette's base colour.

The idle detector-status badge in the status bar is now transparent instead
of an empty default-styled progress bar.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-26 17:51:50 +02:00
Filip LeonarskiandClaude Opus 5.5 f1a8c1d868 Viewer: force black text in the light theme on macOS
The viewer's light (salmon/white) theme was built on the system palette, so
under a dark macOS theme the text roles stayed white on light backgrounds; with
Fusion's standard palette macOS still supplies grey text. Start from Fusion's
standard palette and set window, entry, button and tooltip text to black for
the active and inactive groups, as on Linux and Windows. Disabled text keeps
Fusion's grey.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-26 17:51:50 +02:00
Filip LeonarskiandClaude Opus 5.5 ba43cb4319 Fix macOS build: rename Rugnux library, move in_worker TLS into a function
The Rugnux library and the rugnux executable differ only in case, so on a
case-insensitive filesystem (macOS, Windows) their CMakeFiles/<target>.dir
directories collide and the executable's build.make overwrites the library's
("No rule to make target rugnux/CMakeFiles/Rugnux.dir/depend"). The library is
renamed JFJochRugnux, in line with the other libraries.

Apple's linker rejects the TLS wrapper clang emits for the inline thread_local
static member WorkerPool::in_worker as a duplicate symbol once it is included
from several libraries. It is now a function-local thread_local.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-26 17:51:50 +02:00
leonarski_fandClaude Opus 5.5 ea867171fc Rugnux: run the short-axis indexing beside the standard one
The short-axis pass (the low-FFT-floor hypothesis for small-molecule cells)
ran after the standard first-pass indexing and its rescues, and it is itself
a full indexing of both schemes - about a second of mostly serial refinement,
paid in every first pass, probes included.

It is now started as soon as the standard schemes are fed, on a copy of the
experiment, and runs alongside them. At its own place it is taken only if what
it read is still what the run has there - the experiment and spot-finding
settings (ExperimentKey, split out of FirstPassInputKey), the spot list's
mapping and the per-image spot budget; a rescue that moved any of them makes
it run again there as before. Same inputs through the same code, so the
result is the same.

md5-identical output on two 16M rotation sets; 34.0 -> 29.5 s and
48.2 -> 47.1 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 17:50:58 +02:00
leonarski_fandClaude Opus 5.5 34a50507f6 Rugnux: reuse the canonical pass's first-pass evidence for a probe on the same inputs
The geometry walk after the canonical pass starts with an indexing probe at
the geometry in hand (probe_at_best), which is the canonical pass's own first
pass run over again - a full first-pass indexing whose answer is already
known (identical validation-spot counts in the logs).

The first pass is deterministic in its inputs, so each canonical or probe
first pass now leaves its validation evidence behind keyed by what it read
(geometry, goniometer, spot budget, cell, indexing and spot-finding settings,
and the run state it consults), and an indexing-only probe with the same key
takes it. Nothing is stored for the geometry pre-pass, for a pass that
changed its own inputs on the way (a rescue adopted), or for one that ran the
beam-centre ladder, which a probe never runs. RUGNUX_VERIFY_FIRST_PASS_MEMO
runs the probe anyway and fails the run on any difference.

md5-identical output; 37.2 -> 34.0 s on a 16M rotation set with a geometry
walk (one probe fewer).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 17:39:10 +02:00