Files
leonarski_fandClaude Opus 5.5 50006eff88 Thread count from the CPUs the process may use, not the machine's
std::thread::hardware_concurrency() counts every CPU of the node, so a job
given a CPU set (taskset, a cpuset, a Slurm allocation) or a container CPU
quota started one thread per node CPU: the default -N, the shared worker
pool and every "0 = all threads" default. AvailableCpus() (ThreadAffinity)
counts the start affinity mask, capped by the cgroup CPU quota (v2 cpu.max,
v1 cfs_quota/period, the process's own cgroup and then the mounted root),
read once. Outside Linux it is hardware_concurrency(), so the viewer tree
stays portable. It replaces every hardware_concurrency() call outside the
tests and the vendored pocketfft.

Only how many threads run changes; the passes split their work by n alone,
so the results do not depend on it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-08 11:15:29 +02:00

32 lines
1.7 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <string>
// Keeping a GPU's worker threads on the socket the card hangs off, on a machine with more than one
// NUMA node, so the pinned host buffers they allocate and the copies to the card stay on that
// socket's memory controller and PCIe root. Read from /sys - no libnuma - and Linux only: elsewhere,
// and on a machine with a single node, every call is a no-op.
// The NUMA node of a PCI device, from its bus id as CUDA gives it ("0000:41:00.0"); -1 where it is
// not known or the machine has a single node.
int NumaNodeOfPciDevice(const std::string &pci_bus_id);
// How many CPUs this process may use: the CPUs it was started on (its affinity mask - what taskset,
// a cpuset or a Slurm job gives it), and no more than its cgroup's CPU quota where one is set (a
// container started with --cpus). std::thread::hardware_concurrency() counts every CPU of the
// machine, so a job given 16 of a node's 192 would otherwise start 192 threads. Read once, at the
// first call. Outside Linux it is hardware_concurrency(). At least 1.
unsigned AvailableCpus();
// Restrict the calling thread to the CPUs of `node` that the process may run on. Nothing happens for
// node < 0, or where that would leave the thread no CPU.
void PinThreadToNumaNode(int node);
// Give the calling thread back the CPUs the process started with. A thread inherits the affinity of
// the thread that creates it, so a thread pool first reached from a pinned worker would otherwise
// run every later parallel pass of the process on one socket.
void RestoreThreadAffinity();