The geometry walk re-integrates the whole sweep in every round, and on an
uncompressed SMV sweep that image loop is bound by memory traffic:
smv::ReadInto was the top function (33% of all user cycles on 2xfw), and
16 threads read 360 frames no faster than 1 thread. On Linux/macOS the
frame is now mapped and widened from the mapping, which drops the
read() copy of the whole file into a scratch buffer. Windows keeps the
ifstream read. Pixel values are unchanged; every pass benefits, not only
the walk.
Measured (RTX 5080 workstation, load 9-23 from other agents, 2 repeats,
base = rc175 0457cecaf, same flags):
2xfw WALL 48.15/48.33 -> 40.49/41.98 s, Geometry check 17.2/17.4 -> 13.6/13.8 s
6zqy WALL 26.10/.. -> 23.60/23.16 s, Geometry check 8.6 -> 7.1/7.2 s
2ygz WALL 32.05 -> 26.27 s
-N 8: 2xfw 50.09 -> 42.82 s, 6zqy 26.78 -> 22.97 s
peak RSS -0.5 GB (no scratch copy)
Microbenchmark, 360 2xfw frames, 16 threads: read 1.8-1.9 s, mmap 1.2-1.4 s.
p.hkl md5 identical to the base on every set run (2xfw, 6zqy, 2ygz, -N 8,
and the small SMV sets of the open arm).
Tried first and dropped: keeping the decoded frames in memory across the
walk's rounds. Exact, but filling 13.6 GB of fresh memory cost more than
the hits saved (2xfw round 1 +2.4 s, release +0.9 s, later rounds -0.8 s
each; CBF sets 5m17/7ph1 got slower), at +12.5 GB peak RSS.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SVmAWnzCmRKAXVUCdc4iNi