// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute // SPDX-License-Identifier: GPL-3.0-only #pragma once #include // Keeping a GPU's worker threads on the socket the card hangs off, on a machine with more than one // NUMA node, so the pinned host buffers they allocate and the copies to the card stay on that // socket's memory controller and PCIe root. Read from /sys - no libnuma - and Linux only: elsewhere, // and on a machine with a single node, every call is a no-op. // The NUMA node of a PCI device, from its bus id as CUDA gives it ("0000:41:00.0"); -1 where it is // not known or the machine has a single node. int NumaNodeOfPciDevice(const std::string &pci_bus_id); // Restrict the calling thread to the CPUs of `node` that the process may run on. Nothing happens for // node < 0, or where that would leave the thread no CPU. void PinThreadToNumaNode(int node); // Give the calling thread back the CPUs the process started with. A thread inherits the affinity of // the thread that creates it, so a thread pool first reached from a pinned worker would otherwise // run every later parallel pass of the process on one socket. void RestoreThreadAffinity();