4b1d7feebde52fc093dccf6fcd080b53db2037d5
- set_gpu_blocking_sync(): every device is put in cudaDeviceScheduleBlockingSync before its context exists, so a host thread waiting on the GPU sleeps instead of spinning on a core. On a 16M rotation run a fifth of all CPU time was that spinning; wall time unchanged within noise. Called first thing in rugnux. - enable_gpu_numa_binding(): from then on pin_gpu() (and the new pin_gpu(dev), used by the first-pass spot workers that take a card by index) also keeps the thread on the CPUs of the NUMA node the card hangs off. The node and its CPUs come from /sys (no libnuma), intersected with the process's own mask; Linux only, and nothing happens on a machine with a single node. rugnux turns it on; the broker does not. - A thread inherits its creator's affinity, so the shared ParallelFor pool would run every later pass on one socket if a pinned worker created it: its threads now reset to the mask the process started with (common/ThreadAffinity). Byte-identical output. The NUMA part is a no-op on the single-node test box and still has to be measured on a two-socket machine. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
Jungfraujoch
Application to receive data from the PSI JUNGFRAU and EIGER detectors.
All documentation is now placed in docs/ subdirectory and for the current version hosted on Jungfraujoch Read The Docs page.
Languages
C++
78.1%
HTML
6%
C
4.8%
TypeScript
3.5%
Cuda
2.7%
Other
4.8%