// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute // SPDX-License-Identifier: GPL-3.0-only #pragma once #include #include // Keeping a GPU's worker threads on the socket the card hangs off, on a machine with more than one // NUMA node, so the pinned host buffers they allocate and the copies to the card stay on that // socket's memory controller and PCIe root. Read from /sys - no libnuma - and Linux only: elsewhere, // and on a machine with a single node, every call is a no-op. // The NUMA node of a PCI device, from its bus id as CUDA gives it ("0000:41:00.0"); -1 where it is // not known or the machine has a single node. int NumaNodeOfPciDevice(const std::string &pci_bus_id); // How many CPUs this process may use: the CPUs it was started on (its affinity mask - what taskset, // a cpuset or a Slurm job gives it), and no more than its cgroup's CPU quota where one is set (a // container started with --cpus). std::thread::hardware_concurrency() counts every CPU of the // machine, so a job given 16 of a node's 192 would otherwise start 192 threads. Read once, at the // first call. Outside Linux it is hardware_concurrency(). At least 1. unsigned AvailableCpus(); // Restrict the calling thread to the CPUs of `node` that the process may run on. Nothing happens for // node < 0, or where that would leave the thread no CPU. void PinThreadToNumaNode(int node); // Give the calling thread back the CPUs the process started with. A thread inherits the affinity of // the thread that creates it, so a thread pool first reached from a pinned worker would otherwise // run every later parallel pass of the process on one socket. void RestoreThreadAffinity(); // The calling thread's CPUs now, for SetThreadCpus to give back after a while on one node. Empty // outside Linux. std::vector GetThreadCpus(); // Keep the calling thread on exactly these CPUs. Nothing happens for an empty list. void SetThreadCpus(const std::vector &cpus);