c29dd67b5fc860ab124fc2b924b73769e29198af
3
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
00c27584cb |
Spot width: stop doing obviously redundant work in the pixel loops
None of this is worth its speedup - the routine is 0.5-0.8% of the pre-scan's cycles on two independent profiles, and the whole change is 0.9 ms per data set. It is worth having because the code was doing work that has no reason to exist. The isolation grid was a vector<vector<uint32_t>> over ISOLATION_PX cells: on a 4148x4362 detector that is 23244 std::vector objects constructed, heap-allocated and destroyed per image for a structure that is read once. It is now a counting sort - one offset array, one index array. This was the largest single item and it is not a pixel loop. The encircled-flux curve added every pixel into every bin at or beyond its own radius, about 7.5 adds per pixel, which sums the same aperture R_MAX/2 times over. Each pixel now lands in the one bin its radius falls in and the curve is the running total. The bins are double where the running totals were float, so the result is more accurate, not merely faster. No square roots remain in the pixel loops. The isolation and beam gates compare squared distances, and the radial bin is a table lookup over the 197 squared distances the aperture can produce. That is the same bin, not an approximation: floor(sqrt(floor(y))) == floor(sqrt(y)) for every real y >= 0, because k*k is an integer, so the table recovers ceil(sqrt(rc2)) exactly once the perfect-square case is separated out. The disks are also walked as disks rather than as their bounding boxes - constexpr per-row half-widths - which takes the background pass from 1681 to 1257 pixel visits per spot. Verified over the rotation battery: r80 identical to nine significant figures on 37 of 37 measurable crystals, the 38th unmeasurable in both arms, and the chosen r1 identical on 38 of 38. The pass feeds nothing but r1 into the run, so merged output is unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CHMmeM1d489zvNFT7ZMN2P |
||
|
|
1bb4cecbd1 |
Spot width: measure on a growing sample and stop once it settles
The pre-scan measured the spot width on all 60 frames of its sample every
time. The radius that measurement feeds is a three-way choice - R1ForWidth
moves only where 2*r80 crosses 4.5 and 5.5 - and 29 of 37 measurable
crystals sit more than 0.29 px clear of both switches, so most of the
sample is spent confirming a bucket the first few frames already picked.
Measure instead on a growing share of it - every eighth frame, then every
fourth, every second, all of it - and stop at the first tier whose answer
has settled: within 0.40 px of what the smaller sample said AND 0.25 px
clear of both switches. Both conditions are load-bearing. Clearance alone
loses the crystal whose r80 is 0.07 px from a switch, because its small
samples read across it; the step test alone lets a sample settle on a
switch and stop there. Both bounds sit interior to a two-dimensional
region that is right on 38 of 38, and every candidate was re-scored at all
eight phase offsets of the tier ladder.
The tiers are sized in worker-rounds rather than frames: a first tier of
five frames occupies eight workers as long as one of eight does. A {12,4,
2,1} ladder measures 11% fewer frames than {8,4,2,1} and is 11% slower.
Over the rotation battery this reads 1215 of 2280 frames - 21 crystals
stop at 15, 4 at 30, 13 still run all 60 - and the chosen r1 is identical
on 38 of 38, as is the beam-stop mask the reordered loop also touches.
The pass feeds nothing else into the run, so every merged intensity is
unchanged by construction and no battery is required. Cost over the
battery goes 15.6 s to 10.1 s, median 0.18 s to 0.14 s per data set.
Note for anyone optimising this further: the measurement itself is 0.54%
of the pre-scan's cycles. The cost is the decode, preprocessing and spot
finding each frame needs before it - 64% of the pass - because the
beam-stop projection decodes on the GPU and never materialises the frame
on the host. Cutting frames is the only lever short of harvesting the
width from the two-pass run's first GPU pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CHMmeM1d489zvNFT7ZMN2P
|
||
|
|
48008e1447 |
rugnux: set the integration radius from the crystal's own spot width
On rotation data the signal radius is now r1 = clamp(round(2*r80), 4, 6),
where r80 is the 80% encircled-flux radius of the crystal's own spots. The
background ring keeps its area (r3 = sqrt(r2^2 + 133)), so r1 = 4 is the
shipped default bit for bit and 26 of the 38 battery crystals come out
byte-identical.
The width had to be measured somewhere new. rugnux already has one -
shell_sigma2[].tan - but it is a second moment taken inside the r1 disk it
would be setting, and it saturates at r1/2, so feeding it back measures the
cap and not the crystal. SpotWidth instead measures encircled flux over an
aperture fixed for the whole file (14 px, normalised at 8), in the pre-scan,
from spots the finder already produces on the frames the beam-stop projection
already reads. It touches no integrator output and runs before the first
integration pass, so there is no loop, it costs no extra frame reads, and
both passes - including the space-group search, which runs in pass 1 - see
the same radius. A default run pays a median 1.9 s.
k = 2 is not fitted. For a Gaussian r80 = 1.794 sigma, so r1 = 2*r80 is
3.59 sigma, where the truncated second moment recovers 0.990 of sigma^2. The
new test checks the estimator returns 1.794 sigma on a known Gaussian.
Battery, 38 crystals, both arms run twice: the space group is identical on
all 38 and 35 agree with the reference in both arms. Per shell on the 12
crystals the rule moves, 6 win and 4 tie, with mean per-shell <I/sigma> up
30.6, 24.6, 15.8, 9.1, 8.0 and 5.1 per cent and R_meas down as much as 23.8.
Runtime is neutral - 19m13s against 22m00s warm.
One crystal is a real cost and is named in docs/RUGNUX.md with its
workaround: an I222 case that is simultaneously the widest-spot and among the
highest-mosaicity in the set loses 28.5% of its observations at unchanged
completeness, because at r1 = 6 its predicted reflection density leaves the
background ring too few clean pixels. No cheap guard separates it - its
predicted spacing is mid-table, larger than five crystals that survive r1 =
12 - and the guard that would, on the measured drop rate out of pass 1, needs
a diagnostic channel out of both integration engines and is not yet
validated.
This depends on
|