std::thread::hardware_concurrency() counts every CPU of the node, so a job
given a CPU set (taskset, a cpuset, a Slurm allocation) or a container CPU
quota started one thread per node CPU: the default -N, the shared worker
pool and every "0 = all threads" default. AvailableCpus() (ThreadAffinity)
counts the start affinity mask, capped by the cgroup CPU quota (v2 cpu.max,
v1 cfs_quota/period, the process's own cgroup and then the mounted root),
read once. Outside Linux it is hardware_concurrency(), so the viewer tree
stays portable. It replaces every hardware_concurrency() call outside the
tests and the vendored pocketfft.
Only how many threads run changes; the passes split their work by n alone,
so the results do not depend on it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi