With two GPUs or more the two-pass quality guard's merge of the pre-pass
frames gets a scaling engine of its own on device 1 (RotationScaleMerge::
SetGpuDevice), and the main engine's device ingest and first merge no
longer wait for it to hand the card back. The first merge waits only for
the guard's ingest, the last time the guard reads the outcomes: that merge
writes per-frame fields (mosaicity among them) back into them.
Neither merge reads what the other writes, and a merge gives the same bits
on any card of one model, so only when the guard runs changes. With one
GPU, and in the CPU build, the order is the one it was.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi