Files
Jungfraujoch/common/ThreadAffinity.h
T
leonarski_fandClaude Opus 5.5 4b1d7feebd GPU threads: block instead of spin, and stay on the GPU's NUMA node
- set_gpu_blocking_sync(): every device is put in
  cudaDeviceScheduleBlockingSync before its context exists, so a host thread
  waiting on the GPU sleeps instead of spinning on a core. On a 16M rotation
  run a fifth of all CPU time was that spinning; wall time unchanged within
  noise. Called first thing in rugnux.
- enable_gpu_numa_binding(): from then on pin_gpu() (and the new
  pin_gpu(dev), used by the first-pass spot workers that take a card by
  index) also keeps the thread on the CPUs of the NUMA node the card hangs
  off. The node and its CPUs come from /sys (no libnuma), intersected with
  the process's own mask; Linux only, and nothing happens on a machine with a
  single node. rugnux turns it on; the broker does not.
- A thread inherits its creator's affinity, so the shared ParallelFor pool
  would run every later pass on one socket if a pinned worker created it:
  its threads now reset to the mask the process started with
  (common/ThreadAffinity).

Byte-identical output. The NUMA part is a no-op on the single-node test box
and still has to be measured on a two-socket machine.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 19:30:01 +02:00

25 lines
1.2 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <string>
// Keeping a GPU's worker threads on the socket the card hangs off, on a machine with more than one
// NUMA node, so the pinned host buffers they allocate and the copies to the card stay on that
// socket's memory controller and PCIe root. Read from /sys - no libnuma - and Linux only: elsewhere,
// and on a machine with a single node, every call is a no-op.
// The NUMA node of a PCI device, from its bus id as CUDA gives it ("0000:41:00.0"); -1 where it is
// not known or the machine has a single node.
int NumaNodeOfPciDevice(const std::string &pci_bus_id);
// Restrict the calling thread to the CPUs of `node` that the process may run on. Nothing happens for
// node < 0, or where that would leave the thread no CPU.
void PinThreadToNumaNode(int node);
// Give the calling thread back the CPUs the process started with. A thread inherits the affinity of
// the thread that creates it, so a thread pool first reached from a pinned worker would otherwise
// run every later parallel pass of the process on one socket.
void RestoreThreadAffinity();