rc175
3
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b43844ae06 |
rugnux: the merging engines sit on their card's NUMA node while they drive it
The image-loop workers are kept on the CPUs of their card's NUMA node (pin_gpu), but the two
threads that drive the merging tail were not: the quality guard took its card with set_gpu(1),
and the main engine runs on the pipeline's own thread, which goes on to everything else. On a
multi-socket server their pageable copies to the card went through a remote socket.
RotationScaleMerge now holds a ScopedGpuPin (new, CUDAWrapper) for each IngestHost,
IngestDevice and Run: pin_gpu(gpu_device) for the call, and afterwards the thread's previous
device and CPU mask back (GetThreadCpus/SetThreadCpus, new in ThreadAffinity). So the binding
covers whichever thread drives an engine, and does not stay on the pipeline thread. The threads
the engine starts itself (ComputeAsuGroups, Combine, FitErrorModel, ApplyFrenchWilson use
std::async) call RestoreThreadAffinity first, as the pool's workers do, so they are not
squeezed onto one node. The guard thread, which does nothing but the guard's merge, uses
pin_gpu(1) instead of set_gpu(1). Pool workers are untouched: they restore the start mask
themselves; the new test checks both that and that a pinned caller keeps its own mask.
No flag, no environment variable. On one NUMA node (this workstation) only the device setting
happens, so nothing can move. Measured on the 5080 workstation (one card, one node), base =
rc175-next 41fab6156, --model dropped, under the GPU lock: p.mtz md5 identical on myob_x10sa,
cytc_x10sa, thau_x10sa_16keV, 8a1a, 11if, 7qis, 6h5t, kdp_x10sa_20keV at default -N (2 runs),
-N 8 and CUDA_VISIBLE_DEVICES=0. Tail ("scale, merge and symmetry") unchanged within noise,
e.g. 8a1a 30.06 -> 29.37 s (base runs 29.75/30.36), cytc 5.33 -> 5.35 s. The gain is for
multi-socket boxes (T4 server: guard copies 1.7-2.1 vs 3.5-4.3 GB/s) and is not measured here.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
|
||
|
|
50006eff88 |
Thread count from the CPUs the process may use, not the machine's
std::thread::hardware_concurrency() counts every CPU of the node, so a job given a CPU set (taskset, a cpuset, a Slurm allocation) or a container CPU quota started one thread per node CPU: the default -N, the shared worker pool and every "0 = all threads" default. AvailableCpus() (ThreadAffinity) counts the start affinity mask, capped by the cgroup CPU quota (v2 cpu.max, v1 cfs_quota/period, the process's own cgroup and then the mounted root), read once. Outside Linux it is hardware_concurrency(), so the viewer tree stays portable. It replaces every hardware_concurrency() call outside the tests and the vendored pocketfft. Only how many threads run changes; the passes split their work by n alone, so the results do not depend on it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi |
||
|
|
84228bf8be |
v1.0.0-rc.173 (#83)
Build Packages / Create release (push) Successful in 24s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m29s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m43s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 8m27s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 9m53s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 9m58s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 11m22s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 13m39s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 18m37s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 16m32s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 24m11s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m30s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 19m3s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 20m23s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 19m41s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 21m0s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m38s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 14m33s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 17m55s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 20m50s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 18m38s
Build Packages / Unit tests (push) Successful in 1h46m14s
* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports. * jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls. * Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results. * Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable. * Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate. * Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do. * Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence. * Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags. * Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check. * Rugnux: Clear error messages when a data set needs more GPU or host memory than is available. Reviewed-on: #83 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch> |