Files
Jungfraujoch/image_analysis/scale_merge
jungfrauandClaude Opus 5 27020d27e9 Take the merge's per-observation sweeps off one thread
Of the seventeen seconds a high-multiplicity crystal spends in scaling and merging, only three are
GPU work. The rest is the host, and most of it was running on one or two cores of forty-eight.

Eight of those passes are elementwise maps over the observation array - restoring the scaling
correction at the start of a pass, saving it before the pass filters, scattering it back from the
device, the zeta filter, the frame rejection, the two gathers that hand it to the device again, and
the collapsed-scale ratio. Each reads and writes an eighty-byte record per observation, each ran
serially, and each runs once per cycle with five cycles in a run. They are independent per element,
so chunking them changes nothing but the wall clock. The zeta filter's drop count is now one atomic
add per chunk rather than per observation, and it is an integer, so no arrival order can move it.

The download of the combined fulls did the same work twice over: `assign(nf, Obs{})` zeroed a
quarter of a gigabyte that the next loop overwrote completely, fifteen scratch vectors were
allocated and zeroed afresh every cycle, and the gather from them was a three-million-iteration
serial loop. The scratch is now kept between cycles and the gather is chunked.

The correction surfaces were the last of it. Their inner pass sums the reference intensity of every
usable full, thirty-nine times a run, and a comment asked for per-worker accumulators if it ever
mattered. It does now, but per-worker accumulators would re-associate the double sums. The fulls are
already grouped by a stable counting sort, so walking that grouping visits each group's members in
increasing index - the order the serial loop added them in - and the sums keep their exact sequence.
Copying the four fields the pass actually reads into a packed record first is what makes it pay:
what kept this serial was not the addition but the random read across 265 MB of fat structs, and 53
MB read in order is a different thing.

Ingest is parallel over frames now, which is safe because a frame's mean background is still summed
in that frame's own order by one thread - it is the incident-flux meter and it has to be exact. The
larger rewrite it deserves, sorting a narrow key first and building the fat record only for the ten
per cent that survive the resolution cut, is left alone.

Measured with the surrounding commits: a high-multiplicity set 35.6 s -> 32.4 s, a large-cell one
52.6 s -> 46.1 s, byte-identical merged output on both.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU
2026-08-23 12:59:35 -04:00
..
2026-08-13 17:03:10 +02:00
2026-08-13 17:03:10 +02:00
2026-08-13 17:03:10 +02:00
2026-08-13 17:03:10 +02:00
2026-08-13 17:03:10 +02:00
2026-08-13 17:03:10 +02:00
2026-08-13 17:03:10 +02:00
2026-08-13 17:03:10 +02:00
2026-07-19 09:39:28 +02:00
2026-07-12 19:42:29 +02:00
2026-07-12 19:42:29 +02:00
2026-07-13 13:54:03 +02:00
2026-07-13 13:54:03 +02:00
2026-08-13 17:03:10 +02:00
2026-08-13 17:03:10 +02:00
2026-08-13 17:03:10 +02:00
2026-08-13 17:03:10 +02:00
2026-08-13 17:03:10 +02:00
2026-08-13 17:03:10 +02:00
2026-08-13 17:03:10 +02:00
2026-08-13 17:03:10 +02:00
2026-08-13 17:03:10 +02:00