jfjoch_broker: /start takes an optional `tokens` list (any number, all equivalent; writeOnly in the
API). While the current dataset has any, the endpoints that expose it - /statistics/data_collection,
/result/scan, /image_buffer/{start.cbor,image.cbor,image.jpeg,image.tiff}, /preview/plot{,.bin} -
answer 401 unless the request carries `Authorization: Bearer <token>`; /statistics keeps serving
the instrument view and only omits its `measurement` block. Enforcement is one pre-routing hook
over a named path set (the same set carries `bearerAuth` in jfjoch_api.yaml); tokens are compared
in constant time and never read back or logged. The tokens are replaced only by an accepted start,
and atomically with clearing the previous run's status, plots and image buffer, in this order:
clear, swap, import the new settings - so no moment serves the old run under the new tokens or the
new run's name under the old ones. Settings are validated on a copy first so a refused start
changes nothing.
jfjoch_viewer: a Token field (password echo) next to the http/https scheme in Open HTTP Connection,
JUNGFRAUJOCH_HTTP_TOKEN, and D-Bus LoadFile(..., token) / SetHttpToken; the dialog overrides the
others, nothing is persisted. A 401 clears the display, stops following and puts a line on the
status bar - no dialog, since a dataset changing hands is the normal cause.
Web frontend: key button in the top bar (token in sessionStorage, applied to the bearerAuth
operations by the generated client, and to the raw preview fetch), a tokens field in the start
form, and a "token required" hint on the plots. Python client: Configuration(access_token=...)
after regeneration.
Viewer: reference dataset accepts a structure-factor mmCIF (already read by content; the dialog
now offers it) and a model file can be chosen next to it (ProcessConfig::model_path, as
rugnux --model); the job also carries the reference's free-R flags, cell and setting, and the
copied command line states -z / --reference-column / --model. A grid scan keeps its own preferred
dataset-info plot and "Spots + background" means the spot count there. Dark theme: the navy hero
buttons, checked segments, warning texts and chart guide lines follow the theme instead of their
light-theme colours.
Docs: SECURITY.md section 3 is now the implemented scheme; a guided tour with four screenshots
in JFJOCH_VIEWER.md; broker, OpenAPI, Python client and frontend pages mention the token.
Tests: BearerTokens unit test and an HTTP round trip over the real broker HTTP layer (jfjoch_test
now compiles JFJochBrokerHttp.cpp and links httplib).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
One chain runs at a time there, after the parallel candidate solves, on 1-3.5 cores of an otherwise
idle machine. Ceres sums per-thread cost and gradient pieces, so the thread count may move a result at
rounding level; measured, it did not: identical outputs on the four reference sweeps and all 22 smoke
sets, and a sweep that indexes poorly (a chain after every spot-budget rung) went from 96 to 62 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
A pass's P1 merges - the search merge, the all-observations merge, the P1 cross-check - each ran the
partial scaling loop from scratch on identical inputs (about 5 s each on a 16M CPU-only run). The loop
restarts from corr_ingested, fits every frame from its own observations and the group means alone,
and writes G and corr only on the frames it fits; everything else it reads is fixed at Ingest. So
over the same ASU grouping and settings (space group, Friedel, resolution limits, partiality floor,
window, iteration cap) it ends where it ended before, and the frames, G and corr it left are kept for
the next Run to take. Frames it does not fit keep whatever G they had, as before.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Not bit-exact: the radial offset d.rad is ddx*ux + ddy*uy inlined at two sites, which the compiler may
contract to an FMA differently, so the clip's bin (lround(r0 + rad)) moved on a few pixels of one
sweep's CPU run.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The MTZ points at the experiment's space group, and the background build used a copy that died with
the task: the file was written through a dangling pointer (a verify-mode run failed with a garbled
space-group name; the ordinary runs happened to read the freed memory intact).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Wrong: every spot carries the goniometer angle of its frame (DiffractionSpot phi), so the rotation-
scale probes' spots differ from the stored ones - RUGNUX_VERIFY_FIRST_PASS_MEMO caught it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Each frame packed every ring's pixels together and selected in them twice - the median, then the
median absolute deviation - about 40% of the pass. A ring's background is a few counts, so both are
now read off a histogram of the values 0..1023, exactly, wherever the statistic lands inside it (and,
for the spread, no count is negative); otherwise the ring is packed and selected as before.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Every valid pixel of every sampled frame read and wrote its valid count and level sum, 20 of the 44
bytes the accumulation moved per pixel, on a pass that is bound by memory bandwidth. Both depend on
the pixel only through its ring-sector and its own error frames: they are now summed once per frame
per ring-sector, and a pixel keeps only what its error frames take out of them. Integers, so the same
sums.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
With candidates, the first pass already tested only the 32-column blocks within reach of a candidate,
but still slid the horizontal window across every column of every row. It now slides it only over
the runs of needed blocks, starting each run from the vertical sums it covers - integers, so the same
window sums - and sets the bits there directly; the rest stay 0 as before.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The capture (about 1.5 s on a 16M detector) sat at the end of the pre-scan, but nothing reads its
answer before the first pass's beam-centre check. It now runs on copies of the experiment, mask and
projection it measured, and is taken (JoinBeamCenterCapture) wherever background_center_ or
measured_beam_center_ is read. The first-pass memo key counts a capture still running as a centre
present, which it will be by the time the pass reads it. Its log lines now come from its own thread,
among the first pass's.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The sigma-clip pass recomputed the stencil distances over the whole box to find the same background
ring pixels pass A had just summed. Pass A now keeps their values and radial offsets in its own
order, and the clip runs over them - the same pixels, the same sums.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The four Kabsch reweighting iterations each walked the reflection's grid again with the same bounds,
validity and ownership tests the p_valid pass had just applied. That pass now keeps the profile value
and background-subtracted count of the pixels the fit reads, in grid order, and the iterations run over
them - the same terms summed in the same order (about 6 s of a CPU-only 16M run).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Four per-pixel passes over char masks ran on the pre-scan's critical path on one thread; each
pixel's result depends on that pixel alone, so splitting them changes nothing.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
60-70% of a large detector lies outside the resolution band the walk bins, and every one paid a
square root and two arctangents on every iteration. A test on tan(2theta) = rho / lz with a 0.1%
margin, far above float rounding, drops them first; every pixel the exact test keeps still reaches it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
A probe whose inputs a first pass has already run on returned that pass's evidence only after the
pass had built its azimuthal mapping, writer setup and indexer (0.15 s each on a 16M detector, twice
at the end of a run). The lookup now comes first; nothing in between changes what its key reads.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Spot finding reads whether there is a goniometer, never its axis or angles, and the rotation-scale
walk's probes change only those - so each probe found every frame's spots again (0.2-0.3 s on a 16M
sweep). RUGNUX_VERIFY_FIRST_PASS_MEMO still recomputes and compares.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Each growth is 11 pinned host and 19 device allocations under the driver's device-wide lock; at the
start of an image loop 16 workers growing 1.5x at a time spent about 0.3 s in them.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The ASU-key sample over every partial ran on one thread (0.2-1 s of the merge tail). Each image's
parts are now taken on the workers and joined in image order, so the sort sees the same parts in the
same order.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
nth_element compared through pointers into gigabytes of partials, a cache miss per comparison on one
thread (0.3 s on a 16M sweep). The ratio is now computed beside each pointer on all threads; the
comparisons and so the selection and its order are the same.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
At the smoothing window every Run has just restored corr to what Ingest built, and nothing else the
measurement reads changes after Ingest, so its answer is the same on every Run; it walked all the
partials twice per Run on the CPU path.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
RunScalingLoop rescales after every iteration before anything reads corr, so the iteration's
UpdateCorr and the rescale are one pass over the partials instead of two: the G and fitted frames
as the iteration left them, then the ratio on top - the same two roundings in the same order.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
ReduceGroupMeans ran once or twice per scaling iteration as one serial scatter over every partial
(about 8 s of a CPU-only run's main thread on a 16M sweep). ComputeAsuGroups already builds the
partials' group CSR - a stable counting sort, observation order within each group - for the GPU
reduction; it is now kept on the CPU path too, and each group is summed over it in the order the
serial loop added it, so the means are the same numbers.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
zeta >= min_zeta, so the rocking-curve test rejects every solution with
min_zeta * (|phi| - half wedge) > multiplier * mosaicity. Tested right after phi, with a 1e-5 rad
margin (float rounding of the full test is below 1e-6), it skips the cross product, normalisation
and zeta of the 85-95% of solutions no frame keeps, and cannot reject one the full test would pass.
CPU build: on the order of 10 s of a 1.2 A sweep's prediction.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
A run on data that index poorly asks over a hundred indexing questions - every rung of the
spot-budget ladder, in every pass and walk probe - and a later pass repeats a probe's ladder, which
the 32-entry memo had long evicted (about 20 s on one such sweep). The key holds every spot, so it is
now kept as two independent 64-bit hashes and its length instead of whole.
RUGNUX_VERIFY_FIRST_PASS_MEMO still recomputes and compares.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The unmerged file reads the integration outcomes and the determined group, neither of which the P1
merge changes, except each image's mosaicity, which the batch headers carry and the merge rewrites.
UnmergedMtz builds the file without it on a second thread, and SetUnmergedMtzMosaicity fills it in
after the merge, so the file is the same bytes as before.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Every non-search merge of a rotation pass spent a whole extra merge on cc_half_before_corrections,
but only the pass's first merge (search_merge_cc_half) is read; the ProcessResult copy was never
read at all and is removed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
macOS turns Shift + a vertical mouse wheel into horizontal scrolling, so
the notch arrives in angleDelta().x() and y() is 0; the contrast step
read only y and did nothing. The wheel handler now takes whichever axis
carries the notch, for zoom, foreground and background alike.
Help > Mouse Shortcuts names Cmd for Ctrl on macOS (Qt maps Command to
Ctrl), Ctrl-click for right click and Fn + arrows for Home/End/Page
Up/Down, and lists B held + wheel, which was missing. The magnifier's
tooltip says how to drive it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every rotation merge measured CC1/2 before the correction surfaces - a merge
of its own - and the P1 cross-check's value was never read. scale_and_merge
takes whether it is wanted, and the cross-check says no.
md5-identical p.hkl, p_P1.mtz, p_unmerged.mtz and report on four sets
(one log line fewer); ~0.3-0.6 s on the sets with large merges.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
On the CPU path every image made a separate azimuthal-integration pass
(AzIntEngineCPU) over the 72 MB frame although the adaptive finder's first
ring pass reads the same pixels in the same order under the same rules
(skip the INT32_MIN/MAX sentinels, bins below the mapping's count). That
pass now also accumulates the corrected profile - the same statements as
AzIntEngineCPU, so the same float sums - and MXAnalysisWithoutFPGA takes the
profile from the finder instead of running the separate pass, as the fused
GPU engine already does. Only where the azimuthal engine would be the CPU
one; the finder the pre-scan uses does not accumulate it.
md5-identical p.hkl, p_unmerged.mtz, p_plot.txt and report; CPU-only
163 -> 149 s and 94 -> 91 s. GPU unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The census (report only: what the crystal's lattice leaves over on the
scheme and validation frames, up to three further lattices) ran on the
critical path of the canonical pass between the first pass and the image
loop - ~1.5 s on the GPU build for a set where it finds lattices.
Its spots are still found in place; the rest now runs in the background on
copies of what it reads (the experiment, those frames' spot lists, the
lattice, the validation settings) and is collected into the result before
the pass returns. It no longer runs on a pass that only post-refines, whose
result nothing reads.
md5-identical output and identical report on four sets; GPU 24.1 -> 22.8 s
on a set with leftover lattices.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Every first pass found the spots of its ~200 scheme and validation frames
afresh, although the canonical pass follows a probe on the same frames under
the same settings, and a pass repeating an earlier pass's rescue ladder asks
for the same frames again: on the CPU-only build that is several seconds per
first pass.
The spot lists are now also kept in a store shared by the run and its
copies, keyed by everything a frame's list depends on (SpotFindingKey: the
experiment key - geometry, goniometer, spot budget, settings - the rest of the
spot-finding settings including the measured rings, the ice-ring switch, and
the checksums of the pixel mask and of the spot mapping). A first pass takes
what an earlier one found under the same key. RUGNUX_VERIFY_FIRST_PASS_MEMO
finds the spots again and throws on any difference, field by field.
md5-identical output (four sets GPU, two CPU-only); CPU-only 169 -> 163 s and
96 -> 94 s, GPU 23.0 -> 22.4 s and 119 -> 116 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
- FFTIndexerCPU::ExecuteFFT built the per-direction histograms and scanned the
per-direction spectra for their most prominent peak in one thread; every
direction has its own histogram and spectrum, so the directions are now
split over the refinement threads, each filled and scanned in the same
order as before.
- BeamCenterShortlist2D scanned the whole padded surface (73 M points on a 16M
detector) once per candidate. It now keeps each row's maximum and the first
index holding it and rescans only the rows a suppression touched; rows in
order, first index within a row, is the same first maximum.
md5-identical output; CPU-only 177 -> 169 s and 101 -> 96 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Two whole-image passes per frame out of the CPU spot finder (the pre-scan's
finder on every build, and every image on the CPU-only build):
- AccumulateRings ran three passes over the frame - the plain ring statistics
and two sigma clips. The plain pass now also counts each ring's valid
values in a histogram (0..1023, the rest in a short list), and the clip
passes sum over the distinct values: each meets the same float test its
pixels would, and the sums are integers, so the totals are the same.
- The local test's first pass is read by DetectAt only inside a candidate's
window. It now marks the row/32-column blocks those windows reach and
keeps its sliding sums everywhere but skips the per-pixel test elsewhere;
the bits it leaves unset are never read.
md5-identical output on four sets (GPU) and three (CPU-only). CPU-only
16M: 203 -> 177 s and 303 -> 263 s; GPU 16M 23.8 -> 22.8 s (the pre-scan's
finder).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
On a 16M rotation run the pre-scan was ~11 s on the critical path, the first
3 s of it the spot-width tiers (CPU spot finding on the sample frames),
which the shadow, the defective-pixel mask and the beam-centre capture
waited for although none of them reads what the tiers measure.
The sample is now read twice. The first pass builds the projection and
nothing else; the shadow, the diagnostic, the defective pixels and the
capture start as soon as it is in. The second finds the spots - the width
tiers exactly as before, on copies of the experiment and the pre-shadow mask
it always read - in the background beside them, and reads only the frames
it has something to measure on. What it measured (radii, bandwidth, powder
rings, spot quantiles, the spot-symmetry pool) is applied once both are done,
before anything that reads it. The diagnostic JPEG is rendered in the
background from copies too.
md5-identical p.hkl / p.mtz / p_P1.mtz / p_unmerged.mtz, report and
p_detector.jpg on four sets. GPU 26.2 -> 23.8 s and 38.3 -> 36.4 s on two
16M sets; CPU-only 210 -> 203 s and 114 -> 103 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Three serial stretches of the canonical pass's tail, each made parallel with
the same arithmetic in the same order:
- MeasureBatchDeltaCCHalf measured every batch of the curve, every open
candidate of the rejection loop and every ledger range one after another.
Each measurement is a pure function of its range and the fixed totals, so
they now run side by side (measure_with, one scratch set per worker) and the
decisions scan the results in the original order; the edge walk (locate)
stays one at a time.
- The twin-immune zone evidence (CentricOverAcentric, a few hundred
exponentials per reflection) is evaluated in parallel and summed in the
original order.
- The P1 cross-check's AnalyzeTranslationalNCS was still called with one
thread; it gets the run's thread count like the other two calls.
md5-identical p.hkl, p.mtz, p_P1.mtz, p_unmerged.mtz and report on four sets;
GPU wall 41.0 -> 38.3 s and 26.9 -> 24.8 s on the two sets with long merges.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
A rotation run asks the rotation indexer the same question many times: the
canonical pass's first pass repeats the rotation-scale probe at the stored
angles exactly (identical validation evidence on every set checked), and on
a crystal that does not index, pass 2 and every probe repeat pass 1's rescue
ladder rung for rung. Each such RunIndexing is an FFT search plus a serial
Ceres fixed-point chain, ~1-3 s on the GPU build.
RunIndexing is deterministic in its inputs, so its outcome - every member it
sets - is now kept process-wide under a key of all of them: the accumulated
spots (every field) and their angles, both geometries and the axis, the
experiment's indexing settings, cell and space group, and the settings of the
pool it indexes with (IndexerThreadPool::Settings). A RotationIndexer asking
with the same key takes the outcome. RUGNUX_VERIFY_FIRST_PASS_MEMO recomputes
and throws on a difference.
md5-identical output on four sets; GPU wall 28.9 -> 26.1 s, 43.9 -> 41.0 s,
30.3 -> 26.9 s and 143.6 -> 120.5 s (the set that does not index); the
verify mode found no difference on the last.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
PointSurfaces made three forward and four inverse transforms of the padded
detector (8748 x 8400 on a 16M) one after the other on the main thread -
~3.5 s of one core in a CPU-only run. The forwards do not depend on each
other, nor do the inverse products; each group now runs at the same time,
one workspace per transform. A workspace is built as the single one was -
its own std::vector buffers and its own FFTW_ESTIMATE plan for them - so the
same plan runs on the same data and the result is bit-identical. Costs about
2 GB more transient memory on a 16M detector.
CPU-only 16M rotation run: md5-identical, 216 -> 210 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
FFTW's CMake build knows only SSE/SSE2/AVX/AVX2; NEON exists only in its
autotools build (--enable-neon). The ENABLE_NEON this file set was silently
ignored, so on Apple Silicon and the Linux aarch64 build FFTW ran its scalar
codelets (config.h: HAVE_NEON undefined, no simd/neon objects). The NEON
sources are now added to fftw3f the way its CMakeLists adds the SSE2/AVX
ones, with HAVE_NEON as a target definition (the config.h template carries
it only as a comment). No compiler flag: NEON is part of every aarch64 CPU.
FFTW 3.3.11 would not help - its CMake build has no NEON either, and its
Apple ARM cycle counter only matters to FFTW_MEASURE planning, while every
plan here is FFTW_ESTIMATE.
FFTW's runtime NEON probe on unix/linux executes .long 0xf2000150; on
aarch64 that decodes as ands x16, x10, #0x100000001 and runs without
SIGILL (checked on Apple Silicon), so the codelets are used on Linux too.
Measured on Apple Silicon against the scalar build: batched 1D r2c 1.8x
(power-of-two length) and 1.3x (730), 2D r2c 1024^2 1.9x, 3D c2r 128^3
2.6x, spectra equal to 1.6e-8 relative. rugnux -X FFTW on 360 frames of
the rotation test set: merged .hkl byte-identical, report identical; CPU
time unchanged (300 s), so FFTs are a small share of that run.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
RELEASE_CONTENTS gains the macOS .dmg and rugnux .tgz in the artefact
table, the CPU floor (any Apple Silicon Mac), the OS floor (macOS 13), the
CPU-only note in the CUDA/GPU tables, and a macOS section: Apple Silicon
only (Rosetta does not run arm64 code on Intel), drag-to-Applications,
notices inside the bundle, and how to open the not-yet-notarized release
(Open Anyway on macOS 15+, Control-click Open on 13/14, or xattr).
JFJOCH_VIEWER states the platforms and requirements, that the Mac build is
CPU-only, that D-Bus is Linux-only, and adds Building from source on macOS.
RUGNUX_INSTALL adds the macOS archive and the quarantine note - checked: a
browser-downloaded .tgz hands its quarantine flag to everything tar
extracts, Gatekeeper rejects rugnux, and xattr -dr clears it. DEPLOYMENT
points to the pre-built Windows/macOS viewers and the macOS rugnux archive.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- set_gpu_blocking_sync(): every device is put in
cudaDeviceScheduleBlockingSync before its context exists, so a host thread
waiting on the GPU sleeps instead of spinning on a core. On a 16M rotation
run a fifth of all CPU time was that spinning; wall time unchanged within
noise. Called first thing in rugnux.
- enable_gpu_numa_binding(): from then on pin_gpu() (and the new
pin_gpu(dev), used by the first-pass spot workers that take a card by
index) also keeps the thread on the CPUs of the NUMA node the card hangs
off. The node and its CPUs come from /sys (no libnuma), intersected with
the process's own mask; Linux only, and nothing happens on a machine with a
single node. rugnux turns it on; the broker does not.
- A thread inherits its creator's affinity, so the shared ParallelFor pool
would run every later pass on one socket if a pinned worker created it:
its threads now reset to the mask the process started with
(common/ThreadAffinity).
Byte-identical output. The NUMA part is a no-op on the single-node test box
and still has to be measured on a two-socket machine.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
The rotation two-pass runs several first-pass indexing probes one after the
other although some of them do not depend on each other:
- The rotation-scale walk always indexes at the stored angles and then at
the pass-1 fit; both are known before it starts. The probe at the fit now
runs at the same time, on a copy of the run as it stands before either.
Its answer is taken only if the probe at the stored angles left behind
none of the state the next pass reads (the spot-finding settings a
first-pass ladder rung adopts; each probe puts the experiment back
itself) - otherwise it is run again in sequence, as before.
- The geometry walk's indexing probe at the canonical pass's post-refined
geometry depends on nothing that pass does after its post-refinement. It
is now started there, on a copy, and runs beside the pass's scaling, merge
and output; its evidence comes back through the first-pass memo (now a
short list) and the walk's own probe takes it through the usual key check,
so RUGNUX_VERIFY_FIRST_PASS_MEMO covers it too. It is started only for the
passes whose post-refinement the walk probes with an indexing pass.
The copies are plain copies of Rugnux: the cancel flag becomes a shared
flag (so cancelling the run cancels them) and everything else was already
copyable. Probe passes run on a copy get no observer.
md5-identical output on two 16M sets; the walk's two probes take ~2.3 s
together instead of ~3 s, and the last probe of the geometry walk is hidden
behind the merge.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C