Commit Graph
6 Commits
Author SHA1 Message Date
leonarski_fandClaude Opus 5.5 c20cd016dc Fixed-partition reductions: sums that do not depend on the thread count
ParallelChunks cuts a range into one chunk per worker, so a floating-point
sum folded per chunk and then added up rounds differently at another -N.
ParallelBlocks cuts it by n alone (n / min_per_block blocks, at most 256);
each block folds into its own slot and the slots are added in block order,
so the sum has the same bits at any thread count.

Converted:
- RotationScaleMerge::ApplyCellSurface: the per-cell cross/ref2 sums of the
  surface fit (modulation, absorption), previously per-thread partials cut by
  ThreadsForWork(idx_all) threads.
- PostRefine: the scale-scan cost grid (per-chunk slots cut by -N).
- PostRefine joint solve: Ceres at a fixed 16 threads (as the rotation
  indexer's chain) instead of -N. Ceres sums cost and gradient in
  4 * num_threads pieces, so the split no longer follows -N. The pieces are
  handed to its threads in scheduling order, which only one thread makes
  exact; one thread was measured at +1.0-1.4 s of post-refinement on myob and
  lyso (0.5 -> 1.9 s, 1.1 -> 2.1 s), so that channel is left.
- IndexAndRefine supercell probe: per-frame probes kept by image number and
  summed in frame order, instead of added under a mutex in completion order;
  the primitive is the highest probed frame's instead of the last finisher's.

Integer reductions and per-item passes are unchanged; the three ParallelSort
callers already break ties on the index.

Validation (myob, cytc, lyso, sparse; GPU and CPU builds; -N 8/16/32):
p.hkl, p.mtz, p_P1.mtz and p_unmerged.mtz are md5-identical across -N, and
identical to the base 351be7de0 at every -N. The base was already -N
invariant on these four sets (GPU -N 1..32, CPU lyso -N 2..32), so no output
moved: the converted sums differ from the old ones only in last bits that
never reached a written float.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-28 18:57:02 +02:00
leonarski_f 13c03e1ef4 Merge branch 'perf-rotation' into rc173
# Conflicts:
#	common/ParallelFor.h
2026-09-26 21:48:30 +02:00
leonarski_fandClaude Opus 5.5 4b1d7feebd GPU threads: block instead of spin, and stay on the GPU's NUMA node
- set_gpu_blocking_sync(): every device is put in
  cudaDeviceScheduleBlockingSync before its context exists, so a host thread
  waiting on the GPU sleeps instead of spinning on a core. On a 16M rotation
  run a fifth of all CPU time was that spinning; wall time unchanged within
  noise. Called first thing in rugnux.
- enable_gpu_numa_binding(): from then on pin_gpu() (and the new
  pin_gpu(dev), used by the first-pass spot workers that take a card by
  index) also keeps the thread on the CPUs of the NUMA node the card hangs
  off. The node and its CPUs come from /sys (no libnuma), intersected with
  the process's own mask; Linux only, and nothing happens on a machine with a
  single node. rugnux turns it on; the broker does not.
- A thread inherits its creator's affinity, so the shared ParallelFor pool
  would run every later pass on one socket if a pinned worker created it:
  its threads now reset to the mask the process started with
  (common/ThreadAffinity).

Byte-identical output. The NUMA part is a no-op on the single-node test box
and still has to be measured on a two-socket machine.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 19:30:01 +02:00
leonarski_fandClaude Opus 5.5 59d92a7238 ParallelSort for the two large sorts on the merge path
Two whole-dataset sorts sat on the main thread at the end of a rotation run:
WilsonOutliers orders every full by resolution, and the unmerged MTZ is put
in H K L M/ISYM BATCH order by Mtz::sort(5) - together about 2 s of one
thread on a 1.4 M-observation set.

ParallelSort (common/ParallelFor.h) sorts one piece per worker and merges
them pairwise. It is only for comparators that are a strict total order,
where the sorted sequence is unique and the result is the serial sort's bit
for bit: WilsonOutliers already breaks ties on the index, and the unmerged
writer now sorts the rows itself on the five key columns and then the row
number - the order Mtz::sort's stable sort gives - and sets sort_order as it
did. An empty table still fails the way Mtz::sort does.

md5-identical p.hkl, p.mtz and p_unmerged.mtz; 44.4 -> 43.6 s on a 16M set.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 18:13:41 +02:00
Filip LeonarskiandClaude Opus 5.5 ba43cb4319 Fix macOS build: rename Rugnux library, move in_worker TLS into a function
The Rugnux library and the rugnux executable differ only in case, so on a
case-insensitive filesystem (macOS, Windows) their CMakeFiles/<target>.dir
directories collide and the executable's build.make overwrites the library's
("No rule to make target rugnux/CMakeFiles/Rugnux.dir/depend"). The library is
renamed JFJochRugnux, in line with the other libraries.

Apple's linker rejects the TLS wrapper clang emits for the inline thread_local
static member WorkerPool::in_worker as a duplicate symbol once it is included
from several libraries. It is now a function-local thread_local.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-26 17:51:50 +02:00
leonarski_fandjungfrau 4dc2534dbf v1.0.0.rc-162 (#72)
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 18m57s
Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 16m55s
Build Packages / build:windows:cuda (push) Successful in 18m48s
Build Packages / build:viewer-tgz:cpu (push) Successful in 13m10s
Build Packages / build:viewer-tgz:cuda (push) Successful in 14m45s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 22m23s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 20m12s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 23m7s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 20m43s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 23m9s
Build Packages / XDS test (durin plugin) (push) Successful in 12m26s
Build Packages / build:rpm (rocky9) (push) Successful in 24m58s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 23m20s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (JFJoch plugin) (push) Successful in 12m37s
Build Packages / build:rpm (rocky8) (push) Successful in 27m58s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m38s
Build Packages / Build documentation (push) Successful in 59s
Build Packages / DIALS test (push) Successful in 23m16s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m38s
**Files written by Jungfraujoch now import correctly in DIALS, XDS and pyFAI.** A tilted detector, a grid scan, a still recorded at a goniometer position, and saturated or unreadable pixels were each described in a way that a third-party program acted on wrongly. If you process Jungfraujoch data outside Jungfraujoch, prefer this release to any earlier one.

* HDF5: the detector tilt (`rot1`/`rot2`/`rot3`) is exported correctly in the NXmx transformation chain; untilted geometries are unaffected.
* HDF5: a still recorded at a goniometer position is no longer read back as a single image, and a grid scan records a stationary spindle so a program that requires a rotation axis can open it.
* HDF5: the sample transformation chain is written in mounting order, with a Smargon head position told apart from the spindle, one entry per image, `module_offset` as a float unit vector, and `offset_units` on every offset.
* HDF5: saturated, underloaded and unreadable pixels are described so a downstream program masks them - `saturation_value`, `underload_value`, `error_value` and `bit_depth_readout` are written correctly, and a data file missing next to a VDS master reads as the error marker rather than as zero counts.
* HDF5: the rotation axis is read back under whatever name it carries, and `mirror_y` records whether the assembled image is mirrored in Y relative to the detector's raw readout.
* A grid scan and a goniometer axis can both be set; they are no longer alternatives.
* `images_per_file` is chosen from the acquisition when it is not given: a rotation sweep of at most 20000 images goes into a single data file, a grid scan splits on whole fast-axis rows, and stills and serial keep 1000.
* The writer refuses a stream whose start message declares a different pixel format than its images carry, and a DECTRIS detector sending signed images is no longer declared unsigned.
* The image stream can carry the sample transformation chain (`transformations`, in the END message); a producer that does not send it gets the same chain built by the writer.
* rugnux: fixing the space group with `-S` no longer prevents the lattice from being found - a lattice indexed in a different setting is reindexed into that group's own setting, and a run whose crystal does not have that group's lattice stops and names the cell it indexed as, rather than reporting statistics that cannot describe it.
* rugnux: the per-image resolution estimate now predicts the resolution the merged data reach rather than the highest-resolution spot found, and is reported as `SPOT_RESOLUTION_ESTIMATE`.
* rugnux: two runs of the same command on the same images produce the same merged intensities; the azimuthal profile written alongside them is not yet reproducible in the same way.
* rugnux: the offline lattice refinement is bounded by iterations rather than by a wall clock, so a loaded machine can no longer refine to a different lattice; a live acquisition keeps its real-time bound.
* rugnux: the detector-frame modulation correction is fitted on a grid spanning the detector, so whether it is applied no longer depends on how far integration reached.
* rugnux: the geometry pre-pass no longer writes `<prefix>_01.mtz`, `_01.cif`, `_01.hkl` and `_01_image.dat`; the refined second pass writes those files under `<prefix>`, and that is the result to use.
* rugnux: `_process.h5` describes the pixel format of the images it links to, and is written on a thread of its own.
* rugnux: the detector geometry is also logged in XDS's convention (`ORGX`/`ORGY`, detector axis vectors, rotation axis), so it can be compared with an XDS refinement.
* rugnux: an image integrated in pyFAI through the `.poni` file written by `--mode calibration` comes out with the correct azimuth, and the file declares pyFAI's `orientation`, which needs pyFAI 2024.01 or newer. Radial integration is unchanged.
* rugnux: a rotation run is substantially faster throughout - beam-stop detection, first-pass indexing, geometry refinement, integration, scaling and merging - and observations outside the scaling resolution range are dropped as they are ingested. The refined geometry, the space group chosen and the merged statistics are unchanged.
* Faster spot finding and indexing, on the broker as well as in rugnux; the spots found and the lattices indexed are unchanged.
* A run reserves substantially less GPU memory: nothing is allocated for buffers that are never read, and a worker builds only the engines it uses.
* rugnux: with `-N` left at its default the per-image loop of `--mode mx` uses at most 16 workers per GPU, rather than one per hardware thread; an explicit `-N` is obeyed as given.
* CUDA 12 builds now contain device code for Volta, so the RHEL 8 packages and the portable Linux `.tgz` run on a V100; the CUDA 13 artefacts (RHEL 9, Ubuntu, Windows) remain Turing and newer.
* The build resolves a single Eigen for the whole project, and refuses to configure if Ceres picks up a different one; a build that mixed two Eigen versions was undefined behaviour and crashed at -O2.
* Documentation: a security page, and the supported GPU generations and minimum NVIDIA driver version of every released artefact.

**Breaking change to OpenAPI** - regenerate the client (`jfjoch-client` 1.0.0-rc.162, `frontend/src/client`):
* `dataset_settings.images_per_file` is no longer `default: 1000` and no longer accepts `0`; it is optional, and its minimum is 1. A client sending `0` (previously "one file for the whole run") is now rejected - omit the field instead, which for a rotation sweep gives the same single file.
* `file_writer_format` now defaults to `NXmxVDS`, matching the server's own default and the layout recommended for DIALS, XDS and CrystFEL. A generated client that fills in schema defaults and does not set the format explicitly will write VDS masters where it previously wrote legacy ones; set `NXmxLegacy` explicitly to keep them.

---------

Co-authored-by: jungfrau <jungfrau@mx-aare-test.psi.ch>
Reviewed-on: #72
Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-08-25 08:21:39 +02:00