rc173
7
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2c1123c267 |
rugnux: calibration re-bin sums every image; clean out-of-host-memory failures
Calibration (--mode calibration): RebinAndRefit gave each worker one
AzimuthalIntegrationProfile that AzIntEngineCPU::Run clears on every
image, so the re-binned profile held only the last image each worker
read, added in completion order - on a 1800-frame LaB6 run the re-binned
fit changed with -N (197 / 181 / 205 ring points at rms ~7 px). Each
fixed block of images is now summed in image order into its own slot
(ParallelBlocks) and the slots added in block order, so the profile is
the sum over every image and the same at any thread count (rms 0.93 px,
identical at -N 1, 7 and 32). On the LaB6 runs with a sane header the
first pass still wins and the .poni is unchanged; with a header 100 px
off the re-binned fit is now adopted (218 vs 140 ring points). New test
Rugnux_CalibrationRebinsEveryImage spreads the rings over six frames
and fails on the old code (241 vs 141 points at -N 1 vs -N 4).
Host memory: out-of-memory now ends with "Processing failed: out of
host memory (...) - this data set needs more RAM than is available" and
exit 1. Tested with ulimit -v and an LD_PRELOAD allocator that refuses
allocations from a chosen phase (CPU and GPU builds, image loop through
merging and writing, calibration mode). Paths that crashed instead:
- PostIndexingRefinement ran candidate blocks on bare std::threads, so
a bad_alloc there called std::terminate (seen: "terminate called
recursively", SIGABRT); now std::async futures.
- FFTW aborts on its own failed allocation (CK(p) in kernel/alloc.c,
also inside buffered transforms: BeamCenterFFTCPU, FFTIndexerCPU);
rugnux now defines fftwf_assertion_failed to report it and exit 1.
- libjpeg's default error_exit calls exit() from the diagnostic-JPEG
thread ("Insufficient memory (case 12)"); WriteJPEGToMem now
longjmps back and throws (bad_alloc for JERR_OUT_OF_MEMORY).
- WorkerPool construction that fails to start a thread destroyed
joinable threads (terminate); it now joins them and rethrows.
Also: length_error counts as a fatal resource error, pinned host
allocation failure is MemAllocFailed, and the GPU scaling fail-fast
message says the CPU path needs a lot of host memory.
Docs: very large cells and what running out of host memory looks like
(including the OOM killer) in RUGNUX_INSTALL.md, pointer in RUGNUX.md.
Validation: myob, cytc, lyso, sparse md5-identical to
|
||
|
|
c20cd016dc |
Fixed-partition reductions: sums that do not depend on the thread count
ParallelChunks cuts a range into one chunk per worker, so a floating-point
sum folded per chunk and then added up rounds differently at another -N.
ParallelBlocks cuts it by n alone (n / min_per_block blocks, at most 256);
each block folds into its own slot and the slots are added in block order,
so the sum has the same bits at any thread count.
Converted:
- RotationScaleMerge::ApplyCellSurface: the per-cell cross/ref2 sums of the
surface fit (modulation, absorption), previously per-thread partials cut by
ThreadsForWork(idx_all) threads.
- PostRefine: the scale-scan cost grid (per-chunk slots cut by -N).
- PostRefine joint solve: Ceres at a fixed 16 threads (as the rotation
indexer's chain) instead of -N. Ceres sums cost and gradient in
4 * num_threads pieces, so the split no longer follows -N. The pieces are
handed to its threads in scheduling order, which only one thread makes
exact; one thread was measured at +1.0-1.4 s of post-refinement on myob and
lyso (0.5 -> 1.9 s, 1.1 -> 2.1 s), so that channel is left.
- IndexAndRefine supercell probe: per-frame probes kept by image number and
summed in frame order, instead of added under a mutex in completion order;
the primitive is the highest probed frame's instead of the last finisher's.
Integer reductions and per-item passes are unchanged; the three ParallelSort
callers already break ties on the index.
Validation (myob, cytc, lyso, sparse; GPU and CPU builds; -N 8/16/32):
p.hkl, p.mtz, p_P1.mtz and p_unmerged.mtz are md5-identical across -N, and
identical to the base
|
||
|
|
13c03e1ef4 |
Merge branch 'perf-rotation' into rc173
# Conflicts: # common/ParallelFor.h |
||
|
|
4b1d7feebd |
GPU threads: block instead of spin, and stay on the GPU's NUMA node
- set_gpu_blocking_sync(): every device is put in cudaDeviceScheduleBlockingSync before its context exists, so a host thread waiting on the GPU sleeps instead of spinning on a core. On a 16M rotation run a fifth of all CPU time was that spinning; wall time unchanged within noise. Called first thing in rugnux. - enable_gpu_numa_binding(): from then on pin_gpu() (and the new pin_gpu(dev), used by the first-pass spot workers that take a card by index) also keeps the thread on the CPUs of the NUMA node the card hangs off. The node and its CPUs come from /sys (no libnuma), intersected with the process's own mask; Linux only, and nothing happens on a machine with a single node. rugnux turns it on; the broker does not. - A thread inherits its creator's affinity, so the shared ParallelFor pool would run every later pass on one socket if a pinned worker created it: its threads now reset to the mask the process started with (common/ThreadAffinity). Byte-identical output. The NUMA part is a no-op on the single-node test box and still has to be measured on a two-socket machine. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C |
||
|
|
59d92a7238 |
ParallelSort for the two large sorts on the merge path
Two whole-dataset sorts sat on the main thread at the end of a rotation run: WilsonOutliers orders every full by resolution, and the unmerged MTZ is put in H K L M/ISYM BATCH order by Mtz::sort(5) - together about 2 s of one thread on a 1.4 M-observation set. ParallelSort (common/ParallelFor.h) sorts one piece per worker and merges them pairwise. It is only for comparators that are a strict total order, where the sorted sequence is unique and the result is the serial sort's bit for bit: WilsonOutliers already breaks ties on the index, and the unmerged writer now sorts the rows itself on the five key columns and then the row number - the order Mtz::sort's stable sort gives - and sets sort_order as it did. An empty table still fails the way Mtz::sort does. md5-identical p.hkl, p.mtz and p_unmerged.mtz; 44.4 -> 43.6 s on a 16M set. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C |
||
|
|
ba43cb4319 |
Fix macOS build: rename Rugnux library, move in_worker TLS into a function
The Rugnux library and the rugnux executable differ only in case, so on a
case-insensitive filesystem (macOS, Windows) their CMakeFiles/<target>.dir
directories collide and the executable's build.make overwrites the library's
("No rule to make target rugnux/CMakeFiles/Rugnux.dir/depend"). The library is
renamed JFJochRugnux, in line with the other libraries.
Apple's linker rejects the TLS wrapper clang emits for the inline thread_local
static member WorkerPool::in_worker as a duplicate symbol once it is included
from several libraries. It is now a function-local thread_local.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
4dc2534dbf |
v1.0.0.rc-162 (#72)
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 18m57s
Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 16m55s
Build Packages / build:windows:cuda (push) Successful in 18m48s
Build Packages / build:viewer-tgz:cpu (push) Successful in 13m10s
Build Packages / build:viewer-tgz:cuda (push) Successful in 14m45s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 22m23s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 20m12s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 23m7s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 20m43s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 23m9s
Build Packages / XDS test (durin plugin) (push) Successful in 12m26s
Build Packages / build:rpm (rocky9) (push) Successful in 24m58s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 23m20s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (JFJoch plugin) (push) Successful in 12m37s
Build Packages / build:rpm (rocky8) (push) Successful in 27m58s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m38s
Build Packages / Build documentation (push) Successful in 59s
Build Packages / DIALS test (push) Successful in 23m16s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m38s
**Files written by Jungfraujoch now import correctly in DIALS, XDS and pyFAI.** A tilted detector, a grid scan, a still recorded at a goniometer position, and saturated or unreadable pixels were each described in a way that a third-party program acted on wrongly. If you process Jungfraujoch data outside Jungfraujoch, prefer this release to any earlier one. * HDF5: the detector tilt (`rot1`/`rot2`/`rot3`) is exported correctly in the NXmx transformation chain; untilted geometries are unaffected. * HDF5: a still recorded at a goniometer position is no longer read back as a single image, and a grid scan records a stationary spindle so a program that requires a rotation axis can open it. * HDF5: the sample transformation chain is written in mounting order, with a Smargon head position told apart from the spindle, one entry per image, `module_offset` as a float unit vector, and `offset_units` on every offset. * HDF5: saturated, underloaded and unreadable pixels are described so a downstream program masks them - `saturation_value`, `underload_value`, `error_value` and `bit_depth_readout` are written correctly, and a data file missing next to a VDS master reads as the error marker rather than as zero counts. * HDF5: the rotation axis is read back under whatever name it carries, and `mirror_y` records whether the assembled image is mirrored in Y relative to the detector's raw readout. * A grid scan and a goniometer axis can both be set; they are no longer alternatives. * `images_per_file` is chosen from the acquisition when it is not given: a rotation sweep of at most 20000 images goes into a single data file, a grid scan splits on whole fast-axis rows, and stills and serial keep 1000. * The writer refuses a stream whose start message declares a different pixel format than its images carry, and a DECTRIS detector sending signed images is no longer declared unsigned. * The image stream can carry the sample transformation chain (`transformations`, in the END message); a producer that does not send it gets the same chain built by the writer. * rugnux: fixing the space group with `-S` no longer prevents the lattice from being found - a lattice indexed in a different setting is reindexed into that group's own setting, and a run whose crystal does not have that group's lattice stops and names the cell it indexed as, rather than reporting statistics that cannot describe it. * rugnux: the per-image resolution estimate now predicts the resolution the merged data reach rather than the highest-resolution spot found, and is reported as `SPOT_RESOLUTION_ESTIMATE`. * rugnux: two runs of the same command on the same images produce the same merged intensities; the azimuthal profile written alongside them is not yet reproducible in the same way. * rugnux: the offline lattice refinement is bounded by iterations rather than by a wall clock, so a loaded machine can no longer refine to a different lattice; a live acquisition keeps its real-time bound. * rugnux: the detector-frame modulation correction is fitted on a grid spanning the detector, so whether it is applied no longer depends on how far integration reached. * rugnux: the geometry pre-pass no longer writes `<prefix>_01.mtz`, `_01.cif`, `_01.hkl` and `_01_image.dat`; the refined second pass writes those files under `<prefix>`, and that is the result to use. * rugnux: `_process.h5` describes the pixel format of the images it links to, and is written on a thread of its own. * rugnux: the detector geometry is also logged in XDS's convention (`ORGX`/`ORGY`, detector axis vectors, rotation axis), so it can be compared with an XDS refinement. * rugnux: an image integrated in pyFAI through the `.poni` file written by `--mode calibration` comes out with the correct azimuth, and the file declares pyFAI's `orientation`, which needs pyFAI 2024.01 or newer. Radial integration is unchanged. * rugnux: a rotation run is substantially faster throughout - beam-stop detection, first-pass indexing, geometry refinement, integration, scaling and merging - and observations outside the scaling resolution range are dropped as they are ingested. The refined geometry, the space group chosen and the merged statistics are unchanged. * Faster spot finding and indexing, on the broker as well as in rugnux; the spots found and the lattices indexed are unchanged. * A run reserves substantially less GPU memory: nothing is allocated for buffers that are never read, and a worker builds only the engines it uses. * rugnux: with `-N` left at its default the per-image loop of `--mode mx` uses at most 16 workers per GPU, rather than one per hardware thread; an explicit `-N` is obeyed as given. * CUDA 12 builds now contain device code for Volta, so the RHEL 8 packages and the portable Linux `.tgz` run on a V100; the CUDA 13 artefacts (RHEL 9, Ubuntu, Windows) remain Turing and newer. * The build resolves a single Eigen for the whole project, and refuses to configure if Ceres picks up a different one; a build that mixed two Eigen versions was undefined behaviour and crashed at -O2. * Documentation: a security page, and the supported GPU generations and minimum NVIDIA driver version of every released artefact. **Breaking change to OpenAPI** - regenerate the client (`jfjoch-client` 1.0.0-rc.162, `frontend/src/client`): * `dataset_settings.images_per_file` is no longer `default: 1000` and no longer accepts `0`; it is optional, and its minimum is 1. A client sending `0` (previously "one file for the whole run") is now rejected - omit the field instead, which for a rotation sweep gives the same single file. * `file_writer_format` now defaults to `NXmxVDS`, matching the server's own default and the layout recommended for DIALS, XDS and CrystFEL. A generated client that fills in schema defaults and does not set the format explicitly will write VDS masters where it previously wrote legacy ones; set `NXmxLegacy` explicitly to keep them. --------- Co-authored-by: jungfrau <jungfrau@mx-aare-test.psi.ch> Reviewed-on: #72 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch> |