std::thread::hardware_concurrency() counts every CPU of the node, so a job given a CPU set (taskset, a cpuset, a Slurm allocation) or a container CPU quota started one thread per node CPU: the default -N, the shared worker pool and every "0 = all threads" default. AvailableCpus() (ThreadAffinity) counts the start affinity mask, capped by the cgroup CPU quota (v2 cpu.max, v1 cfs_quota/period, the process's own cgroup and then the mounted root), read once. Outside Linux it is hardware_concurrency(), so the viewer tree stays portable. It replaces every hardware_concurrency() call outside the tests and the vendored pocketfft. Only how many threads run changes; the passes split their work by n alone, so the results do not depend on it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
32 lines
1.7 KiB
C++
32 lines
1.7 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <string>
|
|
|
|
// Keeping a GPU's worker threads on the socket the card hangs off, on a machine with more than one
|
|
// NUMA node, so the pinned host buffers they allocate and the copies to the card stay on that
|
|
// socket's memory controller and PCIe root. Read from /sys - no libnuma - and Linux only: elsewhere,
|
|
// and on a machine with a single node, every call is a no-op.
|
|
|
|
// The NUMA node of a PCI device, from its bus id as CUDA gives it ("0000:41:00.0"); -1 where it is
|
|
// not known or the machine has a single node.
|
|
int NumaNodeOfPciDevice(const std::string &pci_bus_id);
|
|
|
|
// How many CPUs this process may use: the CPUs it was started on (its affinity mask - what taskset,
|
|
// a cpuset or a Slurm job gives it), and no more than its cgroup's CPU quota where one is set (a
|
|
// container started with --cpus). std::thread::hardware_concurrency() counts every CPU of the
|
|
// machine, so a job given 16 of a node's 192 would otherwise start 192 threads. Read once, at the
|
|
// first call. Outside Linux it is hardware_concurrency(). At least 1.
|
|
unsigned AvailableCpus();
|
|
|
|
// Restrict the calling thread to the CPUs of `node` that the process may run on. Nothing happens for
|
|
// node < 0, or where that would leave the thread no CPU.
|
|
void PinThreadToNumaNode(int node);
|
|
|
|
// Give the calling thread back the CPUs the process started with. A thread inherits the affinity of
|
|
// the thread that creates it, so a thread pool first reached from a pinned worker would otherwise
|
|
// run every later parallel pass of the process on one socket.
|
|
void RestoreThreadAffinity();
|