Files
leonarski_fandClaude Opus 5.5 b43844ae06 rugnux: the merging engines sit on their card's NUMA node while they drive it
The image-loop workers are kept on the CPUs of their card's NUMA node (pin_gpu), but the two
threads that drive the merging tail were not: the quality guard took its card with set_gpu(1),
and the main engine runs on the pipeline's own thread, which goes on to everything else. On a
multi-socket server their pageable copies to the card went through a remote socket.

RotationScaleMerge now holds a ScopedGpuPin (new, CUDAWrapper) for each IngestHost,
IngestDevice and Run: pin_gpu(gpu_device) for the call, and afterwards the thread's previous
device and CPU mask back (GetThreadCpus/SetThreadCpus, new in ThreadAffinity). So the binding
covers whichever thread drives an engine, and does not stay on the pipeline thread. The threads
the engine starts itself (ComputeAsuGroups, Combine, FitErrorModel, ApplyFrenchWilson use
std::async) call RestoreThreadAffinity first, as the pool's workers do, so they are not
squeezed onto one node. The guard thread, which does nothing but the guard's merge, uses
pin_gpu(1) instead of set_gpu(1). Pool workers are untouched: they restore the start mask
themselves; the new test checks both that and that a pinned caller keeps its own mask.

No flag, no environment variable. On one NUMA node (this workstation) only the device setting
happens, so nothing can move. Measured on the 5080 workstation (one card, one node), base =
rc175-next 41fab6156, --model dropped, under the GPU lock: p.mtz md5 identical on myob_x10sa,
cytc_x10sa, thau_x10sa_16keV, 8a1a, 11if, 7qis, 6h5t, kdp_x10sa_20keV at default -N (2 runs),
-N 8 and CUDA_VISIBLE_DEVICES=0. Tail ("scale, merge and symmetry") unchanged within noise,
e.g. 8a1a 30.06 -> 29.37 s (base runs 29.75/30.36), cytc 5.33 -> 5.35 s. The gain is for
multi-socket boxes (T4 server: guard copies 1.7-2.1 vs 3.5-4.3 GB/s) and is not measured here.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi
2026-10-09 12:51:38 +02:00

40 lines
2.0 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <string>
#include <vector>
// Keeping a GPU's worker threads on the socket the card hangs off, on a machine with more than one
// NUMA node, so the pinned host buffers they allocate and the copies to the card stay on that
// socket's memory controller and PCIe root. Read from /sys - no libnuma - and Linux only: elsewhere,
// and on a machine with a single node, every call is a no-op.
// The NUMA node of a PCI device, from its bus id as CUDA gives it ("0000:41:00.0"); -1 where it is
// not known or the machine has a single node.
int NumaNodeOfPciDevice(const std::string &pci_bus_id);
// How many CPUs this process may use: the CPUs it was started on (its affinity mask - what taskset,
// a cpuset or a Slurm job gives it), and no more than its cgroup's CPU quota where one is set (a
// container started with --cpus). std::thread::hardware_concurrency() counts every CPU of the
// machine, so a job given 16 of a node's 192 would otherwise start 192 threads. Read once, at the
// first call. Outside Linux it is hardware_concurrency(). At least 1.
unsigned AvailableCpus();
// Restrict the calling thread to the CPUs of `node` that the process may run on. Nothing happens for
// node < 0, or where that would leave the thread no CPU.
void PinThreadToNumaNode(int node);
// Give the calling thread back the CPUs the process started with. A thread inherits the affinity of
// the thread that creates it, so a thread pool first reached from a pinned worker would otherwise
// run every later parallel pass of the process on one socket.
void RestoreThreadAffinity();
// The calling thread's CPUs now, for SetThreadCpus to give back after a while on one node. Empty
// outside Linux.
std::vector<int> GetThreadCpus();
// Keep the calling thread on exactly these CPUs. Nothing happens for an empty list.
void SetThreadCpus(const std::vector<int> &cpus);