c470bed93ac8d4da1dee373e6474b643ae01a0a1
1206
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c470bed93a |
Allocate off the driver lock, and give the loop the sixteen workers a card wants
Build Packages / build:windows:nocuda (push) Successful in 16m31s
Build Packages / build:windows:cuda (push) Successful in 19m15s
Build Packages / build:viewer-tgz:cpu (push) Successful in 19m58s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m51s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 24m5s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 25m3s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m51s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m55s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m35s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 21m19s
Build Packages / XDS test (durin plugin) (push) Successful in 11m17s
Build Packages / build:rpm (rocky9) (push) Successful in 23m30s
Build Packages / Generate python client (push) Successful in 33s
Build Packages / build:rpm (rocky8) (push) Successful in 25m35s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m37s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m40s
Build Packages / XDS test (neggia plugin) (push) Successful in 10m17s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 26m44s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 11m10s
Build Packages / DIALS test (push) Successful in 23m27s
Build Packages / Unit tests (push) Successful in 1h19m59s
Two changes to how the per-image loop is set up, neither of which touches what it computes. Device buffers were taken with cudaMalloc and returned with cudaFree, both of which are on CUDA's implicit-synchronisation list: each one synchronises the device across every stream. One analysis engine per worker, each making a few dozen of them, means the workers still constructing stall the workers already processing images, and the cost grows with the worker count. They are now stream-ordered allocations from the device's memory pool, with the synchronous pair kept as the fallback where no pool is available. Two deliberate limits on that. The pool's release threshold is one gibibyte rather than unbounded: holding the small per-worker buffers is the whole point, but the card also has to fit the merge afterwards, which asks for several gigabytes of its own. And the shared geometry tables keep the synchronous allocator, because their deleter runs on whichever thread drops the last reference, so an asynchronous free there would be ordered on a stream that says nothing about the engine streams whose kernels read the table; they are allocated once per card, so the pool bought them nothing. The loop's worker cap per card goes from eight to sixteen. The comment beside it already recorded where the measurement put the minimum - the loop's time falls to sixteen workers and then rises - and a later measurement on a single card agrees: at eight the loop waits on the queue rather than on the card. An explicit -N is still obeyed as given. Reflection files are byte-identical on four crystals with the worker count doubled, which is the property the frame-ordered mosaicity smoothing and the deterministic prediction order were built to give. Thirty consecutive runs of one crystal on the pooled allocator: no failure, every file identical to the first. Peak device memory over the whole rotation test set is 4.9 of 16 gibibytes. Four minutes seventeen to four minutes one over thirty-eight crystals, each binary repeating itself to within half a per cent; nineteen crystals faster, nineteen level, none slower, and every column of the comparison table identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EGpGdgmJ8MyY9pCGWjktyi |
||
|
|
1d16dcd2b9 |
Read a chunk without zeroing it first, and hold no frame the fused decoder never writes
Build Packages / build:windows:nocuda (push) Successful in 16m8s
Build Packages / build:windows:cuda (push) Successful in 18m58s
Build Packages / build:viewer-tgz:cpu (push) Successful in 20m35s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m31s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 25m9s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 25m6s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m57s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m58s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m58s
Build Packages / XDS test (durin plugin) (push) Successful in 12m3s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 22m24s
Build Packages / build:rpm (rocky9) (push) Successful in 21m45s
Build Packages / Generate python client (push) Successful in 53s
Build Packages / build:rpm (rocky8) (push) Successful in 26m9s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m37s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m34s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 22m0s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m53s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m29s
Build Packages / DIALS test (push) Successful in 23m40s
Build Packages / Unit tests (push) Successful in 1h20m1s
Three costs before and around the image loop. Every image allocated a fresh buffer for its compressed chunk and resized it, which value-initialises, and the read then overwrote every byte. At a few megabytes a chunk the allocation is large enough to be mapped rather than reused, so the zeroing was page-fault bound and cost more than the read it preceded - twenty gigabytes of it over a long sweep. The buffer now uses an allocator that does not construct, and the two HDF5 read paths are templated on the allocator so every existing caller compiles unchanged. The rebind is deliberate: without it the vector base rebinds to the default allocator and the zeroing quietly returns. The bitshuffle decoder allocated a whole uncompressed frame in its constructor - seventy megabytes a worker, five hundred and fifty across the loop - for the route that decodes the shuffled image separately. That route is taken only when a bitshuffle block is too large for the fused kernel, which neither writer this pipeline reads produces, so on a real frame the buffer is allocated, never touched, and freed. It is now allocated where it is used. The comment two lines below already warned against sizing a buffer from the uncompressed size; the line above it had not been given the same treatment. The first call into cuFFT pays the library's one-time initialisation, and it landed in the middle of the first pass with nothing to overlap it. It is now forced on a background thread at startup, alongside the file open and the mapping build, in the manner the shadow finder already uses. Finally, the detector mask was copied into the start message whether or not a file would carry it, which a merging run does not. It is filled where a writer is constructed - both places one is constructed, the second being the fallback that writes a process file when nothing indexed. Faster on eleven of thirty-eight crystals and slower on none; the whole rotation test set falls from four minutes thirty to four minutes seventeen, with each binary repeating itself to within half a per cent. Space groups thirty-five of thirty-eight and no failures throughout, and every column of the comparison table is identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EGpGdgmJ8MyY9pCGWjktyi |
||
|
|
641f890a40 |
Build the frame's constants once, and page-lock what the integration engine copies
The per-image geometry refinement is the largest stage of the image loop, and a third of it was arithmetic on numbers that never change. The residual derives the detector angles' sines and cosines, the goniometer's back-rotation - a three-argument hypot, a sine, a cosine and a division - and the reciprocal basis of the cell on every evaluation. On the rotation path the detector angles and the axis are held fixed and stored as plain doubles, so all of it is constant, not merely constant per block: there is one frame per image and one cell. Three solves an image, fifty iterations a solve and a thousand spots make it tens of thousands of repetitions of the same result. The frame's constants are now built once and handed in. The body they feed is the same body, split out rather than copied, so no expression is reassociated - in particular the reciprocal vector is still formed as the basis times the inverse volume, with the volume not folded into the basis. The spot confidence weights depend only on each spot's resolution and intensity, which no solver touches, and were recomputed identically for each of the three passes. They are computed once. The sort behind them ordered indices through a projection that chased a random eighty-byte-strided element per comparison; it now sorts a packed resolution and index, which makes the same comparisons in the same sequence and therefore the same permutation. The spot list itself was copied per image through an initializer list whose elements are const; it is passed as a view. The integration engine was the last one in the loop copying through pageable host memory - three transfers in and eight out per image, twenty-six bytes a reflection, while every other engine already page-locks its staging. A driver copy from pageable memory stages through its own pinned buffer on the calling thread, which is why an asynchronous copy was averaging a hundred and thirteen microseconds. Page-locked, the same seventeen thousand calls cost four hundred and thirty-two milliseconds instead of one and a half seconds, and the wait moves to the synchronisation point where it belongs. Two smaller ones: the reflections were copied into the per-image message for a process file that a merging run does not write, so the copy is made where a writer exists; and the intensity statistics and the Wilson estimate walked the same eighty-byte array twice to read twelve bytes, which is now one pass with each accumulation in its own order. Every reflection file is byte-identical on four crystals; the process file's reflections match dataset for dataset, and its azimuthal arrays differ no more between this build and the last than the last differs from itself. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EGpGdgmJ8MyY9pCGWjktyi |
||
|
|
9a50ae5a0f |
Take the asymptote where it is read, and count a shell's reflections on every core
Five costs in the merge, none of them arithmetic. The strong-reflection asymptote was estimated twice for every merge and the first estimate was read by nothing: no statement between the two touches the value, so wherever the resolution cutoff refits the error model the earlier one was thrown away. It is now taken once, at the point the number is reported. Its per-group scatter array was also built fresh on every call - fifty megabytes value-initialised on one thread and immediately overwritten - and now lives with the object. The merge accumulator did the same thing on a larger scale: ten arrays and a struct of accumulators, a quarter of a gigabyte in all, zeroed on one thread before the device wrote every element of them. The kernel is a grid-stride loop over all groups and writes all ten outputs unconditionally, so nothing was reading a zero it had put there. The arrays are kept and resized, and the unpack that follows runs over the cores instead of one; its only reduction is an integer count, which does not care in what order it is summed. The host path still clears, because it accumulates in place. Dropping collapsed frame scales walked every full at eighty bytes a record to read two fields of four. Both were already downloaded, so it reads those instead: a tenth of the traffic for the same answer. The error model's chi-square median was computed on every fit and printed once. It now keeps what the last fit used and takes the median where the line is written. Counting the reflections a resolution shell could hold walked the whole reciprocal box on one thread and built a vector of the survivors first. The walk is now split over the outer index with a per-thread tally summed in thread order, and the vector is gone. On the cells in the rotation test set that is four to eleven milliseconds a call against three calls a crystal; on a two hundred Angstrom cell it is ninety-three milliseconds down to five. A comment claiming the point-group pass is serial to fill the operator cache is no longer true and is corrected; the cache is filled by a parallel pass before it. Every reflection file is byte-identical on four crystals, under a pinned resolution limit and under the automatic cutoff - the latter being the configuration that actually exercises the moved asymptote, since a pinned limit never made the second call at all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EGpGdgmJ8MyY9pCGWjktyi |
||
|
|
0109676bdf |
Keep the constants out of the dual numbers, and clear only the boxes
Two costs in the per-image loop, each measured before it was touched. Geometry refinement is the largest item in that loop - about half to two thirds of its processor time on the datasets where the loop matters, and all of it on the host. It runs three solves per image, and each one spends four fifths of itself inside the solver at barely two iterations: the cost is not convergence, it is what every residual evaluation does. The residual carried the blocks it does not refine as dual numbers, so each evaluation recomputed the two detector rotations, the whole orthogonalisation matrix, three cross products and the cell volume - all of them constant for the image - through the derivative machinery, several million times per run. Split the observed and predicted sides so the un-refined blocks pass as plain doubles, evaluate the cell side once when the functor is built, and let the rotator take a point whose type differs from the angle's. A dual number times a double is a dual number times a dual number whose derivatives are zero, so the arithmetic is the same one with the zeros removed. Integration cleared the owner and mask images for the whole frame before every image. On a large detector that is more than three hundred megabytes of writes to reset pixels of which about one in twenty-five is ever marked, and it cost most of what the integration kernels themselves cost. The marking kernel gained an unmarking mode - one kernel, so the two cannot drift apart - and the engine clears whichever way is cheaper for the frame in front of it, with a flag to force the full clear the first time and after anything threw. The size test is not decoration: without it, clearing box by box is slower than the memset on a small detector with many predictions, which is what the measurement said before it was added. Faster on thirteen of thirteen matched pairs: refinement by a quarter to a third, whole-run wall by one to eight per cent depending on how much of the run is the loop. The two changes pay in opposite regimes - refinement where the loop is processor-bound, the clear where the detector is large enough for the card to be the constraint. Every reflection file over seven datasets is byte-identical, and the solver did not merely land in the same place: it took the same path, agreeing digit for digit on iteration, residual and Jacobian evaluation counts. A new test runs two mismatched frames through one engine and compares against a fresh one, which is what a mark left behind would break. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016NNnL26LAvruQ9eLUUWvrJ |
||
|
|
b074cca6ea |
Put the reflections in the fixed group's own setting, or say they cannot be
Fixing a space group told the merge which absences to apply but never told it which basis to apply them in. Where the indexed lattice was already conventional the group was simply stamped on it, so a primitive tetragonal cell asked to merge in a C-centred orthorhombic group had that centring rule evaluated in a frame the reflections were not in, and half of them were declared systematically absent. The reindex that would have fixed this existed but was reachable only from the triclinic arm. So ask the character table for the group's own class. The Bravais search grows an optional class filter - one continue that skips characters of the wrong class, one answer of "none fits" when the metric cannot carry it - and the reindexing that followed the triclinic arm is lifted out and offered to a fixed group whose centring is not the indexed lattice's, mapping a trigonal-P request onto the hexagonal-P setting it is described in. With no class asked for, both new statements are dead and the search is what it was. That splits the failing cases in two, and conflating them was what made this wrong in both directions. A lattice that HAS a setting carrying the group is reindexed into it: the tetragonal case above recovers every observation it had been discarding, and a centred monoclinic one that had been merging from a primitive cell without any reindex - which nothing had noticed - goes from an error model that could barely be fitted to a healthy one. A lattice that genuinely has no such setting - a triclinic metric several degrees from monoclinic-C, or an F-centred cubic one asked for hexagonal-P, whose hexagonal description is R-centred - has no basis to be put in, and every statistic computed from it is meaningless. Those now stop, naming the group, its centring and the cell that was actually indexed, and they stop only after the reindex has been tried, so a mistyped but reachable group is repaired rather than rejected. The second pass keeps its existing flag-and-decline instead. Separately, the geometry pre-pass predicted in the primitive lattice only when no group was fixed. With a centred group fixed it integrated half the events, moved the error model, and shifted the post-refined distance by more than a tenth of a millimetre - enough, in a loop this sensitive, to send the second pass down the other branch. It now predicts primitive there whatever the group, which is what it already did de novo and which its discarded intensities have no opinion about; the one dataset this cost its indexing rate recovers completely, and lands on the same answer it reaches with no group given. Thirty-one of thirty-eight pinned runs are bit-identical and none is worse. De novo nothing changes at all, by construction and on the whole rotation test set. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016NNnL26LAvruQ9eLUUWvrJ |
||
|
|
e5442e7a07 |
Read the cell surface's twenty bytes, and give each candidate group a thread
Build Packages / build:windows:nocuda (push) Successful in 17m17s
Build Packages / build:windows:cuda (push) Successful in 19m29s
Build Packages / build:viewer-tgz:cpu (push) Successful in 20m4s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m1s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 23m37s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 24m55s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 29m14s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 29m17s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 30m37s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 19m23s
Build Packages / XDS test (durin plugin) (push) Successful in 11m38s
Build Packages / build:rpm (rocky9) (push) Successful in 24m1s
Build Packages / build:rpm (rocky8) (push) Successful in 25m59s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m22s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m14s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 11m49s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 26m27s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m53s
Build Packages / DIALS test (push) Successful in 23m28s
Build Packages / Unit tests (push) Successful in 1h18m34s
Three unrelated costs in the tail, each measured before and after. The cell surface fits and scores by reaching into the fulls for sixteen bytes of an eighty-byte record, twenty-two times over. That is four times the traffic of the data it uses, and it was the whole of the cost: the arithmetic never was. One fused compaction of the intensity, sigma, correction, cell term and group, built in the passes that were already being made, and every later pass walks the compact array instead. Fit accumulation falls to a third, scoring to a third. What is left is the term build and the scatter, not the fit. The estimator is untouched. The space-group search hashed a reflection key per observation per candidate. The orbit representative is now interned to a dense index when the orbits are built, so the two tests that follow index an array. Both candidate loops also run a thread per candidate - the point-group loop, and the space-group loop, which was the larger of the two by far: a full pass over the merge with three absence tests and a map insert per reflection, once for every candidate group. The operator cache is filled by the serial pass that precedes them, so each candidate still sees the same operators, does its own arithmetic unchanged, and appends in the same order. The search runs in a third of the time. Writing the reflections was seventeen stream insertions per row for a quarter of a million rows. The rows are formatted in parallel blocks and written in order, in the same widths and precisions as before: the mmCIF in a tenth of the time, the hkl in a quarter. Faster on twelve of twelve matched pairs, eight per cent on the sum of minima, and region timers account for the wall clock to within five per cent. Every reflection file over the whole rotation test set is byte-identical, in both passes and all three formats. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016NNnL26LAvruQ9eLUUWvrJ |
||
|
|
9f26c5e8b1 |
Keep the per-observation corrections on the card that computed them
Run downloaded the correction factor for every partial, scattered it into an eight-and-a-half-million element array of eighty-byte records, kept a copy of it, filtered it twice on the host, gathered it back and uploaded it again. Five passes over six hundred and eighty megabytes, on the path where the data was already resident on the card. A comment above it warned that three host readers needed the scattered copy, and an earlier attempt read that as a reason to leave the whole thing alone. Taken one at a time the readers fall: both pass filters test quantities the card already holds - zeta and the frame index - so they become kernels over the correction array in place; the saved copy is now the download itself, thirty-four megabytes rather than a gather of the whole record; and the combine, the only genuine host reader, gets its own scatter immediately before it, which matters solely when observations are dumped. Zeta is compared in double on the device so the promotion matches the host's comparison exactly, and the drop count is an atomic add. The merge's own sweeps had the same shape: they walked the eighty-byte record to reach twenty bytes of it. They now build those twenty bytes once per merge, on all threads, and stream them. The reject median's first walk over every full goes entirely - the counts it was accumulating are the ones the error-model pass has already produced. The group histogram is one flat uninitialised buffer whose rows are cleared by the threads that use them, in place of a vector of vectors cleared twice, and the ingest no longer zeroes four hundred and fifty megabytes of staging that the following line overwrites. Faster on seventeen of seventeen matched pairs across two alternating sessions; eight to ten per cent of whole-run wall clock on the datasets where the tail dominates. The reflection files are byte-identical, including with a frame correlation cut and with an observation dump, which are what exercise the two filters and the host combine. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016NNnL26LAvruQ9eLUUWvrJ |
||
|
|
d19409f4a1 |
Resolve Eigen once, before Ceres, and refuse a build that mixes two
Ceres asks for Eigen with a version range - find_package(Eigen3 3.3.4...5 NO_MODULE) - and Eigen's own Eigen3ConfigVersion.cmake rejects any range whose endpoints differ in major version. A 3.4 Eigen therefore declares itself INCOMPATIBLE with Ceres' query, the search falls through, and Ceres binds whatever older Eigen comes next on the prefix path. Where a distro eigen3-devel 3.3.4 is installed alongside a 3.4 one, Ceres created Eigen3::Eigen first, in its own directory scope, pointing at the older headers, and exported it publicly; the project's own find_package then ran afterwards and made a second target pointing at the newer ones. Targets linking both - the geometry refinement and the scale/merge libraries - took the older Eigen first. The result was a binary holding Eigen 3.3.4 and 3.4.90 template instantiations at once. Identically mangled, they are merged at link time with disagreeing evaluator layouts, so the program is undefined: at -O2 it segfaulted inside an Eigen product under the lattice reduction, nine runs out of nine, and at -O3 it happened not to, which is luck rather than correctness. Resolve Eigen before Ceres is added. The first find_package to run creates the imported target and later ones leave it alone, so Ceres inherits ours. Then assert it: if Ceres ever creates an Eigen3::Eigen of its own, the configure fails with an explanation rather than producing a binary that is quietly ill-formed. The guard fires only on that condition, not merely because two Eigens are installed. After the change no translation unit sees the older headers - 0 of 227 flags files, against 30 before - and Ceres reports the Eigen it actually compiled against. The same nine runs that all crashed now all complete. Release output is unaffected: eight datasets give byte-identical reflection files and identical merging statistics either way, so nothing previously measured is invalidated. Eigen and ZLIB stay external find_package dependencies on every platform, and OVERRIDE_FIND_PACKAGE is not reintroduced. Where only one Eigen is installed - the Windows and macOS case - Ceres' non-range fallback honours the same Eigen3_DIR and the change is a no-op. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016NNnL26LAvruQ9eLUUWvrJ |
||
|
|
2f54a1189d |
Take the error-model split and the ASU grouping off one thread
Three regions of the merge tail, measured with instrumented timers and confirmed against a cycle profile. On a tail-heavy dataset the scale and merge tail is 70% of the run's wall clock at six of thirty-two logical cores busy, with the GPU idle 88% of the time, so this is where the CPU headroom is. fit_error_model ran a serial four-level nth_element cascade over the whole sample pool, twelve times per dataset. The two halves either side of a partition are disjoint and their contents are already fixed by the parent's nth_element, so the recursion can descend both at once; it now does while a range is worth a thread. The bins are unchanged. ComputeAsuGroups sorted indices with an indirect comparator, taking a cache miss per comparison into an array far larger than the last-level cache. It now sorts packed key-and-run pairs. Tie order does not matter because the packed key encodes h, k, l and the hand exactly, so every run in a tie reduces to the same reflection. The per-thread histogram prefix walked thirty-two separate histograms column-wise on one thread. It becomes a parallel per-group total, one sequential scan over two flat arrays, and a parallel hand-out of the bases - the same sums in the same order. Faster on 21 of 23 matched pairs in an alternating A/B, and on 15 of 15 in the quieter of the two sessions: 0.6% to 2.3% of whole-run wall clock depending on the dataset, around 1.8% in aggregate, and 3 to 4% of the time spent outside the image loop. The reflection files are byte-identical on every dataset tested. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016NNnL26LAvruQ9eLUUWvrJ |
||
|
|
54adcaafcc |
Estimate the resolution the merged data will reach, not the furthest spot found
The per-image resolution estimate was the 5th percentile of the spot d spacings - an extreme order statistic, so it measured where detection stops rather than how well the crystal diffracts. A large cell puts more reflections past the same threshold and scored better than a small cell that diffracts further; intensity was not used at all, so a weak crystal padded with spurious high-resolution detections ran away; and nothing clamped the answer to what the detector can deliver. Against the resolution the merged data actually reach it was 42% out in log-RMS, with 1 of 38 rotation datasets inside 0.2 A. Take instead the 1/d^2 beyond which 30% of the sum of sqrt(I) over the image's non-ice spots lies, report 1/(2.25 sqrt of it), clamp at the detector corner, and take the median over images. A quantile from the middle of the distribution measures the shape of the falloff - the crystal's own exp(-B/2d^2) - where an extreme one measures the threshold. sqrt(I) is the Poisson significance of a summed photon count, so a marginal high-resolution detection cannot carry the answer and neither can a handful of very strong low-resolution reflections. The 2.25 is the multiplicity gain: merging keeps measuring intensities a fixed factor in 1/d past the point where a single frame detects them. Spearman 0.881 -> 0.954, log-RMS 42% -> 8.9%, median error 0.79 -> 0.07 A, and 32 of 38 within 0.2 A. Both constants sit on a broad plateau, the scale is stable across dataset halves and across resolution ranges, and no second predictor survives leave-one-out. The residual is around 9%, set by multiplicity, symmetry and radiation damage - none of which a spot list can see. The estimate feeds only reporting: the image stream, HDF5, the plots, the scan result and the preview ring. It sets no cutoff and no search limit, and the scaling and merging output is byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016NNnL26LAvruQ9eLUUWvrJ |
||
|
|
baf302f813 |
Map an FFT candidate to its own Bravais lattice, not to the pinned group's
A user-fixed space group reached indexing in one place, and what it did there was relabel a cell rather than re-express it. build_sr took the conventional cell that LatticeSearch had reduced for whatever Bravais class the METRIC matched, then overwrote its system and centring with the pinned group's - without transforming the cell. The constrained refine then snapped that cell's real angles onto the pinned class's ideal ones. Measured: a C-centred orthorhombic cell relabelled primitive monoclinic indexed 1 of 60 validation frames, and an F-cubic one relabelled trigonal indexed 0 of 60, where the same frames index 36/60 and 51/60 with no group given. The pseudo-symmetry guard was withheld at the same time - has_tri required no group - so the unconstrained cell did not exist, which also disabled the false-promotion rescue in pick_best. The only remaining outcome for a bad constrained cell was the throw. That is what decided the two centred-monoclinic cases, where the relabelling is a no-op and the metric really is the pinned class: the constrained solve runs out to the length bound at a fraction of 0.005 while the unconstrained solve on the same candidate reaches 0.7. The group names the symmetry; it does not say which basis the candidate came back in. It is applied where it belongs, to the scaling and the merge. Pinning each rotation test dataset to its reference group: 5 hard failures of 38 become none, and no dataset is worse than before. The de-novo path is unchanged by construction - with no group the deleted branch never ran, and has_tri's condition reduces to its old form - and a ten-dataset de-novo control reproduces the baseline exactly. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016NNnL26LAvruQ9eLUUWvrJ |
||
|
|
bd2179b08f |
Let the parallel pass signal completion under the lock its waiter tests
RunTasks keeps its RunState - counter, mutex, condition variable - on its own stack frame and hands a pointer to the pool. The last worker to finish decremented the counter OUTSIDE the mutex and only then took it to notify, while the waiter's predicate was the counter itself. So the waiter could see zero the instant the decrement landed, find its predicate already true, never block, and return from RunTasks - popping the frame. The worker then locked a mutex and signalled a condition variable that no longer existed, writing pthread state into a frame the submitting thread had already reused. pthread_mutex_unlock writes owner and nusers as eight contiguous zero bytes. Land those on a live pointer and the next read of a member at offset 8 faults: the observed crash was fmt's buffer<char>::append with this == nullptr, in the log call immediately after a parallel pass, which is why it always appeared after the resolution-filter line - that line was simply the next thing to use the frame. Give the waiter a flag set under the same lock as the notify. Completion cannot then be observed until the notifier has released the mutex, i.e. after its last touch of the state. The counter keeps its lock-free fast path and decides only who notifies, so there is still exactly one lock per pass. Found independently by two investigations: a widened-window reproducer (2 crashes in 38 unfixed, 0 in 60 fixed; glibc's own "__owner == 0" assertion caught in the pool worker) and an isolated one that clobbered a freshly filled stack frame 732 times in 60000 and never after the fix. Growing RunState by eight bytes, changing nothing else, took the rate from 0/90 to 3 hard failures in 30. The dataset that failed about one run in twelve: 0 of 40. Space group and merged output unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cf4fbcf96a |
Reduce the anomalous split once per ASU group, not once per observation
ComputeAsuGroups states the rule for itself - "one ASU reduction per distinct raw hkl (not per observation)" - and the anomalous split then did a gemmi ASU reduction and an unordered_map lookup for every one of the millions of fulls. Both things it wants are properties of the observation's ASU GROUP rather than of the observation: group_h/k/l is the group's SIGNED representative, so the same reduction applied to it returns the Friedel-merged key and the hand together. Reduce once per group into a dense accumulator indexed from there. The hand only follows the group when the merge distinguishes the hands; a Friedel-merged run holds both in one group and still has to ask per observation. SigAno and the merged statistics are unchanged (2.96 over 53303 acentric pairs, merge table byte-identical). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
29eda90190 |
Report the time the whole run took, not the last pass of it
processing_time_s is set inside RunPipeline, so it measures one pass. A rotation two-pass run integrates everything twice and the beam-stop / beam-centre pre-scan happens before either pass, none of which the reported number saw: on an 18 Mpx dataset it printed 7.58 s for a run that took 23.46 s. Time the run in Run(), where every pass is inside, and report that. The last pass is still printed alongside it when there was more than one, because the gap between them is what the second pass costs. The frame rate and throughput stay per-pass: they say how fast rugnux moves through images, which running a second pass does not change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
65d0d6f238 |
Note the faster geometry refinement and merge in the changelog
Build Packages / build:viewer-tgz:cpu (push) Successful in 18m57s
Build Packages / build:viewer-tgz:cuda (push) Successful in 21m30s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 23m3s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 24m32s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m44s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m55s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 30m19s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 21m5s
Build Packages / XDS test (durin plugin) (push) Successful in 11m50s
Build Packages / build:rpm (rocky9) (push) Successful in 23m21s
Build Packages / Generate python client (push) Successful in 34s
Build Packages / Build documentation (push) Successful in 1m31s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (rocky8) (push) Successful in 27m15s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m43s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 22m10s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 11m12s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m39s
Build Packages / DIALS test (push) Successful in 23m39s
Build Packages / build:windows:nocuda (push) Successful in 13m15s
Build Packages / build:windows:cuda (push) Successful in 15m35s
Build Packages / Unit tests (push) Successful in 1h20m17s
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
be75001833 |
Build the merge's per-frame quantities once per frame, and count what a shell can hold only when it is reported
Three passes over the ingested observations were doing more than they needed. The smoothed-geometry pass rebuilt a CrystalLattice and its three reciprocal vectors for every observation, each one a cross product and a cell volume, for a value that depends only on which frame the observation came from. On a large sweep that is tens of millions of constructions against a couple of thousand distinct answers. The completeness column counts how many unique reflections a shell could hold. It is read off a merge that gets written out, never off the ones the space-group search runs on the way there - and those are the expensive ones to count, because the search merges in P1, where the list is the whole hemisphere rather than an asymmetric unit of it. The keep flags were written over the whole observation array as 1 and then immediately over it again as 0 whenever a resolution limit is set, which the default low-resolution limit always does. They are filled once now. The merged reflections also get their capacity up front rather than doubling their way to it several times per pass. Merged intensities are unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0d07141b5c |
Ask the geometry refinement for the derivatives it uses, and solve the normal equations
Four changes to the same least-squares fit, which the rotation first pass runs on every candidate lattice and the per-image path runs on every frame. The linear solver was DENSE_QR on a problem that is very tall and thin - thousands of spots against at most seventeen parameters. That is the shape QR handles worst: it copies the Jacobian out of Ceres' row-major storage into a column-major buffer on every solve, and Eigen's blocked Householder then degenerates to the unblocked path because its block size is the column count. Accumulating J^T J reads the Jacobian once instead. Both solve the same damped system, so the step is the same to round-off. Ceres sizes its dual numbers from the declared parameter blocks, not from which of them the caller then holds constant. Nothing outside a test set refine_distance_mm - the positional residual leaves the distance degenerate with the cell scale, which is why the rotation post-refinement fits it in a step of its own with the cell held fixed - so the block was declared only to be frozen, and every residual differentiated seventeen parameters to use sixteen. It is gone, along with the test that exercised distance recovery; that test seeded the distance off truth, which the cell would now absorb, so its seed moves to the true value. The post-refinement's own detector step held five of its seven blocks constant and now bakes them into the residual, leaving beam and distance. The predicted reciprocal vector was built by rotating all three direct columns and then crossing them. A rotation commutes with the cross product and leaves the triple product alone, so the same vector comes out of crossing the unrotated columns and turning the result once - three rotations become one, for every crystal system. The documentation described the arrangement before all this, and had drifted in a second way: the first-pass rotation indexing has been refining the detector tilt and the rotation axis by default, which the text said were held fixed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4b48d19064 |
Keep the threads the parallel passes run on, rather than making them per pass
ParallelChunks and ParallelFor started a thread per chunk with std::async and joined it again on every call. A thread costs tens of microseconds to create and join, and the analysis code repeats some of these passes thousands of times in a run - the merge alone has dozens of call sites, several of them inside iteration loops - so a short pass could spend more on its threads than on the work. Both now run on a pool made once and kept. The split is unchanged, so a pass whose per-element work is independent still gives the serial answer bit for bit. Two properties the futures gave for free had to be kept explicitly. A pass reached from inside a pool worker runs inline instead of queueing, since the workers are occupied by the outer pass and waiting for one of them could wait forever. And an exception from any task is held until every task has finished and then rethrown to the caller, so the others still run to completion. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f6983486c3 |
Put the test's mask in through LoadUserMask, not through a const_cast
The fused GPU preprocessor uploads a byte-per-pixel form of the mask that PixelMask derives when the mask is loaded. Writing the bitfield behind its back left that form stale, so the device worked from a mask the CPU reference did not have. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
aceaf8b714 |
Say what the program now does, in the places that still described what it used to
Build Packages / build:windows:nocuda (push) Successful in 13m21s
Build Packages / build:viewer-tgz:cpu (push) Successful in 26m8s
Build Packages / build:windows:cuda (push) Successful in 14m49s
Build Packages / build:viewer-tgz:cuda (push) Successful in 28m25s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 31m34s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 32m10s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 37m12s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 37m8s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 38m43s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 21m7s
Build Packages / XDS test (durin plugin) (push) Successful in 11m21s
Build Packages / build:rpm (rocky9) (push) Successful in 22m50s
Build Packages / build:rpm (rocky8) (push) Successful in 26m32s
Build Packages / Generate python client (push) Successful in 54s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m46s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m29s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 11m36s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 26m54s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m39s
Build Packages / DIALS test (push) Successful in 22m39s
Build Packages / Unit tests (push) Failing after 1h26m26s
A pass over the user-facing text against the code, after a day of changes. Almost all of it is deletion. The pre-pass files: RUGNUX.md promised them twice, CPU_DATA_ANALYSIS.md once, the usage line for --rotation-no-postrefine, the prose in the run report, a viewer tooltip and a comment in the anomalous script. They are not written any more, so the promises are gone and one is replaced by a sentence saying what the pre-pass is for and that it writes no merged files. CPU_DATA_ANALYSIS.md's account of decompression said two kernels do the work and, specifically, that the kernel stages nothing in shared memory and is therefore indifferent to the bitshuffle block size the file declares. Both halves stopped being true when the decode was fused into the preprocessing: it asks for the block as dynamic shared memory, and a block larger than 16 kB falls back to the old two-kernel route. That fallback is a behaviour a reader needs, so it is now described rather than contradicted. RUGNUX.md said spot finding, integration and scaling run on the CPU and scale with -N. Spot finding, preprocessing, azimuthal integration, prediction and Bragg integration are on the GPU, and rotation scaling and merging are GPU-resident too; that sentence now names what runs where. The -N row gains the per-GPU cap, and the usage line it is documenting gains it too - the usage message is the authority, so the two now agree. Deleted a promise of a per-iteration scaling file that nothing has written since the standalone scaling tool was removed. The changelog keeps four entries for today, cut back to what the house rules ask for: what changed, one line, no measurements or mechanism. The reproducibility entry now says MERGED intensities and says plainly that the azimuthal profile is not yet reproducible in the same way, which is what the code delivers - the corrected ring sums are still float atomics, and neither dropping the corrections (solid angle ramps 63% across the profile) nor widening them (the per-warp accumulators do not fit in shared memory at the bin counts in use) turned out to be a way to fix it. The review also turned up two defects in today's own code - the per-image cost line reported the thread count rather than the workers the loop actually ran, and the cap reached the azimuthal and calibration modes, whose worker preprocesses and integrates on the CPU and wants every thread it can have. Both are fixed in the preceding commit, which touches the same files. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
b380da2a91 |
Judge the refined pass against the pre-pass on the same footing
The two-pass guard rolls back to the header geometry when the refined pass looks worse than the pre-pass, and one of its three tests is a drop of more than 0.05 in CC1/2. Since the pre-pass stopped fitting its correction surfaces - it exists to pick a space group and post-refine the geometry, and its intensities are discarded - the two sides of that test were no longer measuring the same thing: the pre-pass's CC1/2 came out uncorrected and the refined pass's corrected. On one crystal here that flattered the refined pass by 0.008, and it is the wrong direction to be careless in, because it makes the guard slower to fire on a pass that really is bad. So measure the refined pass's CC1/2 before its surfaces are applied as well, and compare that. It cannot be had from the half-set accumulate alone, which was the cheap thing to hope for: CC1/2 correlates half-set means built on the error model's sigmas over the reflections the automatic resolution cutoff kept, so an accumulate on its own is a different quantity - and one biased low, which would make the guard fire too eagerly. It takes the same merge the pre-pass now does, without the statistics tail, before the surfaces run. That is one extra merge on the one pass that has surfaces, so a caller asks for it explicitly rather than paying for it by default: the two-pass driver does, --mode scale does not, because there is no other pass to compare against. Where the caller wants it but nothing was corrected anyway, the merge that already ran IS the uncorrected one and is reported as such; where nobody asked, the field stays absent rather than being filled in from the statistic that reads almost the same and is not. Also make the final in-symmetry merge unconditional, with P1 when no group was determined. The comment there has always said P1 stands in that case and the condition did the opposite, which would have written a search merge - zeta-filtered, ice-excluded, uncorrected - as the result. It turns out to be unreachable: with any reflections at all the search returns a group, because no symmetry leaves the identity point group whose representative is P1, a group with no screws or centering leaves a symmorphic candidate that has no absences to contradict and so is always eligible, and an empty merge throws in both engines before the search sees it. The two lines keep that promise here instead of resting on eligibility gates in another file that a later change could tighten without noticing what leaned on them; the reasoning is written at the site. Battery unchanged on all 24 crystals - same space group, reflection count and R_meas as the run before it - and the merged output is byte-identical on three crystals spanning the regimes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
c133fd89a8 |
Give the per-image loop as many workers as the cards can take, not as many as the machine has
-N defaults to every hardware thread and the per-image loop spawned one worker for each of them. Every worker submits its own kernels to a card, and a card runs out of room to accept them long before it runs out of work to do: measured on two GPUs, the loop's own time falls from 10.76 s at four workers to 9.69 s at sixteen and then climbs back to 10.49 s at forty-eight. Forty-eight workers is slower than eight. The same shape appears on a small detector, with the turn further out because a frame is a smaller piece of work. So cap the loop at eight workers per card when -N was left alone. Per card, because that is what the queue depth belongs to; eight, because that is where the curve turns on the hardware this was measured on. Everything outside the loop - the merge, the surfaces, post-refinement - still gets the whole machine, because none of it is waiting on a card. An explicit -N is obeyed exactly as given, and the cap says so in the log when it fires. A previous attempt at this overrode an explicit -N and applied to the azimuthal and calibration modes as well, which is why it was refused; this one is only about the default. It matters most where it cannot be measured here. A two-card production node with 192 threads runs ninety-six workers per card against a curve that turns at eight, while this box at -N 48 across four cards sits at twelve and looks fine. Even so, on four cards the battery goes 6m28s -> 6m10s, with every crystal's space group, reflection count and R_meas identical to the run before it - the cap changes no arithmetic at all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
7762d8bd30 |
Build an observation only if the resolution range keeps it, and let the pre-pass skip what it discards
Two changes to what the scaling stage does at all, rather than to how fast it does it. Ingest converted every integrated partial into the eighty-byte record the scaling works on, and then threw away whatever fell outside the requested resolution range. On a large cell that is sixty-three million records built and fifty-seven million discarded - five gigabytes written, most of it to be skipped by every consumer afterwards. It now emits a twenty-four byte key per observation, sorts and buckets those, decides from the runs which raw hkl the range keeps, and builds the full record only for the survivors. The flux meter still sums a frame's whole background in that frame's own order on one thread, because it is the number every merged intensity is divided by; the first usable d is still taken over the run rather than over the survivors; and the sort order was already total, so how the keys are filled cannot change it. The other is the pre-pass. It exists to choose a space group and post-refine the geometry, and its merged intensities are discarded - the second pass makes them again at the refined geometry. It was nonetheless fitting the decay, absorption and modulation surfaces, measuring radiation damage and sweep quality, assigning R-free flags, converting to amplitudes, walking the observations again for R_meas, splitting the anomalous pairs and analysing twinning, all for a result nobody reads. A flag threaded from the call site turns that off on the pre-pass, following the convention the anomalous split already used. What the pre-pass keeps is what is read later: the merge itself, the error model, the resolution cutoff, and the whole per-shell statistics block - because the second pass is judged against the first, and that guard needs the pre-pass's completeness and CC1/2. The flag that says the statistics exist is untouched and still set unconditionally; moving it is what disabled the guard entirely in an earlier attempt at this, and with it the completeness bound, the CC1/2 bound, the lattice-conflict test and the fall back to the header geometry. One consequence to be aware of: the pre-pass's CC1/2 is now measured without the correction surfaces while the second pass's is measured with them, so the two are no longer compared on quite the same footing. It moves the guard in the direction of firing less readily, never more, so it cannot roll back a good pass - but it is a small loss of sensitivity and the next commit removes it. Merged output byte-identical on a large-cell set, a high-multiplicity one and a small one; battery 6m28s against 7m55s, space group unchanged on all 24 crystals. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
ef0be2e52a |
Bound the offline lattice refinement by iterations, and stop rotating the same axis three times
Two things in the indexing path, one of them a reproducibility hole. XtalOptimizerData bounds a solve by iterations when it is told to and by WALL-CLOCK SECONDS when it is not, and its own header says why that matters: the same image refines to a different answer on a busier machine. The per-image refinement sets the iteration bound for exactly that reason. The rotation indexer never did, so its candidate-cell refinement ran under a one-second wall clock - three stages a candidate, up to eight candidates a scheme, twice a run. A run that has just been made reproducible from its prediction order to its accumulators was still free to pick a different lattice because the machine was loaded. It now takes the iteration bound offline and keeps the wall-clock one for a live acquisition, whose budget is real, which is the same split the per-image path already makes. The residual itself rotated the same axis three times over. It applies one orientation to three reciprocal-lattice vectors, and ceres::AngleAxisRotatePoint recomputes the angle, its sine, its cosine and the normalised axis on each call - and it does not inline at this optimisation level, so the compiler cannot notice. On a seventeen-parameter Jet each of those is a full dual-number evaluation. Computing the rotation once and applying it three times removes two hypots, two sines, two cosines, two divisions and six multiplies from every evaluation, which is about half the libm calls in it; hoisting a constant member's sine and cosine out of the same function takes two more. It runs everywhere the residual does - the indexer, the per-image refinement and the geometry refiner. Also lifts five SetParameterBlockConstant calls out of a per-observation loop in the detector solve, where they were executing once per observation to say the same thing. The rotation hoist was checked against the function it replaces on 200000 random dual numbers, including the small-angle branch, comparing the value and all seventeen derivative lanes: no difference in any component. Merged output is byte-identical on a 16 Mpx set and an ordinary one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
cf2336e523 |
Decode a compressed frame into shared memory and preprocess it there
The image loop on a 16 Mpx detector is two thirds of the run and both cards are busy for essentially all of it, so card time removed is wall time removed. Of the six milliseconds a frame costs, two and a half were spent decompressing it - and not because the card was short of bandwidth. The LZ4 pass moved 53 GB/s where the strong-pixel flagger, reading the same image and the same bin table, gets 276. It is latency, not bandwidth: the copy loop moves 32 bytes per warp iteration with a syncwarp after each one, and for a match copy the source and the destination both derive from the same pointer, so nothing pipelines. The warp spends its time waiting for global memory, one dependent round trip at a time. So decode where the waiting is cheap. One CUDA block now owns one bitshuffle block: its first warp decodes the payload into shared memory, and the whole block then un-transposes and preprocesses out of shared and writes finished pixels. A shared round trip is tens of cycles rather than hundreds, and the 72 MB shuffled intermediate never reaches DRAM at all - the pair of kernels moved about 238 MB a frame and the fused one moves 93. The parser is lifted into a device function that both kernels call over the same bytes, so the standalone path and the fused one cannot decode a chunk differently. The statistics reduction had to change with it: 48 bytes of static shared on top of a full bitshuffle block costs a whole resident block per multiprocessor, so the counts now reduce through a warp shuffle and one integer atomic per warp. Blocks larger than 16 kB keep the two-kernel path, and the beam stop's own decoder is untouched. What this costs is decoder parallelism: a block that holds 16 kB of shared is one of four resident per multiprocessor on this card, where the old kernel fitted thirty-two warps each decoding on its own. The trade is favourable here and should be better on the production cards, which have half again as much shared memory per multiprocessor. Measured on a 16 Mpx rotation set at the production GPU count, with the indexing work of the next commit: 37.2 s -> 32.9 s, and the merged output is byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
f09fe4e2c5 |
Report the per-image cost as something that can be attributed
Each stage timer is wall time inside one worker, so it counts whatever that worker spent blocked - on the GPU, above all - as well as its own work. Dividing the mean by the worker count, as this did, assumes every worker was busy for the whole loop. Measured occupancy is a third of the workers asked for on a large detector and less on a small one, so the number people tune against came out low by that factor, and it moved with the contention rather than with the work. Report the share instead. A stage's fraction of a worker's own per-image time is what that stage is responsible for whatever the contention was, and spending that fraction against the loop's wall time per image gives a figure that is attributable and that sums to the loop. The worker mean is printed at the end rather than divided away, because the gap between it and the wall is the waiting, and the size of that gap is worth seeing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
5ee0f22a61 |
Build the detector's lookup tables once, not once per worker
The image loop gives every worker its own analysis engine, so a run builds ninety-six of them. Each one derived, from scratch, tables that are the same in all of them: the byte-per-pixel mask, the resolution mask, the radial kernel, and the checksum that names the shared device tables. The checksum was the worst of it, because it is part of the cache KEY and so is computed before the lookup - a hit still hashed the whole table. On a 16 Mpx detector that is the bin table, the corrections and the mask, 126 MB an engine, about twelve gigabytes over a run, to answer a question whose answer had not changed. The header said it cost nothing measurable; a profile says otherwise, and says it is worst exactly during the ramp when the machine has nothing else to do. It cannot simply be remembered against the address, which is what it exists to catch: a buffer can be freed and another allocated where it was, and the cache would then hand back a device copy of something else. So the owner of the bytes computes it instead. The azimuthal mapping writes its two tables in its constructor and never again. The pixel mask re-derives its binary form and its checksum on every path that changes the mask, and all of those paths are now private to the class. The key therefore still describes the bytes as they are at the moment of the lookup. The resolution mask was two passes over every pixel - a float comparison into a vector<bool>, then a bit-by-bit repack - in each of the ninety-six. It is one pass now, writing the packed form directly, built once for the limits asked for and handed out as a shared pointer so a worker keeps the mask it was given. The radial kernel is cached on the six numbers it is derived from. Nothing computes a different value; only who computes it changes. Byte-identical merged output on a 16 Mpx set and on a small one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
27020d27e9 |
Take the merge's per-observation sweeps off one thread
Of the seventeen seconds a high-multiplicity crystal spends in scaling and merging, only three are
GPU work. The rest is the host, and most of it was running on one or two cores of forty-eight.
Eight of those passes are elementwise maps over the observation array - restoring the scaling
correction at the start of a pass, saving it before the pass filters, scattering it back from the
device, the zeta filter, the frame rejection, the two gathers that hand it to the device again, and
the collapsed-scale ratio. Each reads and writes an eighty-byte record per observation, each ran
serially, and each runs once per cycle with five cycles in a run. They are independent per element,
so chunking them changes nothing but the wall clock. The zeta filter's drop count is now one atomic
add per chunk rather than per observation, and it is an integer, so no arrival order can move it.
The download of the combined fulls did the same work twice over: `assign(nf, Obs{})` zeroed a
quarter of a gigabyte that the next loop overwrote completely, fifteen scratch vectors were
allocated and zeroed afresh every cycle, and the gather from them was a three-million-iteration
serial loop. The scratch is now kept between cycles and the gather is chunked.
The correction surfaces were the last of it. Their inner pass sums the reference intensity of every
usable full, thirty-nine times a run, and a comment asked for per-worker accumulators if it ever
mattered. It does now, but per-worker accumulators would re-associate the double sums. The fulls are
already grouped by a stable counting sort, so walking that grouping visits each group's members in
increasing index - the order the serial loop added them in - and the sums keep their exact sequence.
Copying the four fields the pass actually reads into a packed record first is what makes it pay:
what kept this serial was not the addition but the random read across 265 MB of fat structs, and 53
MB read in order is a different thing.
Ingest is parallel over frames now, which is safe because a frame's mean background is still summed
in that frame's own order by one thread - it is the incident-flux meter and it has to be exact. The
larger rewrite it deserves, sorting a narrow key first and building the fat record only for the ten
per cent that survive the resolution cut, is left alone.
Measured with the surrounding commits: a high-multiplicity set 35.6 s -> 32.4 s, a large-cell one
52.6 s -> 46.1 s, byte-identical merged output on both.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU
|
||
|
|
0fed94d75b |
Gather post-refinement's partials without touching four gigabytes twice
Post-refinement fits eight numbers, and it selects the twenty thousand best-recorded events to fit them from. Before it can select, it copies every integrated partial into an array of its own and sorts it. On a large cell that is 63 million of them, and the phase took 8.5 s of a 56 s run. Almost none of that was the sort. `std::vector<Partial> pts(n)` value-initialises: one thread writes 3.5 GB of zeroes, page by page, before the parallel fill overwrites every byte of it - and being the first touch, it also decides where the pages live, so the whole array lands on one NUMA node and every later pass over it runs at one node's bandwidth. The same again for the sorted copy. Allocate the storage without initialising it and let the parallel fill be the first touch. The record itself carried more than the sort reads. `angle_rad` is a function of the image number that the goniometer can give back on demand, and the two observed positions are wanted only by the distance step, and only for the twenty thousand it keeps. Storing what is read - and as the floats the fields already were, since widening a float to a double is exact - takes the record from 56 bytes to 32, which is a third off the fill and half off the sort's element moves. Then three passes that walked the whole array to no purpose. The h range is now taken in the count pass, which reads the same reflections anyway; the bucket histogram in the fill pass, which already has h in hand. The event split walked serially and grew its output by doubling - about a gigabyte of pure copying - although h is the leading sort key, so a rocking event never crosses an h bucket: count per bucket, prefix, fill in parallel, and the events come out in the order the serial walk produced them. And the copy of the whole event list, made only so that nth_element could destroy the original, is now an index array. Every one of these is the same arithmetic in the same order. Measured on a large-cell rotation set, with the two commits that follow: 52.6 s -> 46.1 s, and the merged .hkl, .mtz and .cif are byte-identical. `part_less` is deliberately left as it was, not a total order: what makes it reproducible is that each bucket reaches the sort in gather order, and the new chunking is still a contiguous span of that order. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
4a537dbfc2 |
Bin the error model's samples without sorting them
The (a, b) fit wants sixteen equal-count bins in I^2 and takes three medians out of each. It was getting them by sorting the whole pool - millions of 32-byte samples - and it did that fourteen times a run: the fit runs once per merge and twice where the resolution cutoff refits, the outlier refit doubles it again, and there are five merges. Each call also took its pool BY VALUE, so every one of those began by copying tens of megabytes, and each bin then built three more vectors by push_back to hand to a median. A bin only has to be the right SET. Put each boundary in place with nth_element instead, splitting the boundaries down the middle so every level halves the range it works on - four levels of linear work against n log n - and take the three medians straight off the bin's own span with the field wanted, which is what median_of was doing anyway: it returns the lower median, exactly the element nth_element leaves at that index. No copy is made at all, and the sixteen bins are disjoint so they divide over the cores. The comparator is now total. The sort it replaces was not stable, so which of two samples of equal I^2 landed in which bin was decided by the order the pool happened to arrive in - and the refit is handed a different order from the first fit. Ordering on the remaining fields, which are in the same cache line, makes the bin a property of the samples instead. This is why the merged intensities are not byte-identical to the previous release on about half a percent of reflections, at a median difference of zero and a worst case of 1.2e-2: those are the ties, whose old resolution was arbitrary. Every fitted (a, b, ISa, chi2) in the run agrees to four significant figures. The per-group outlier median goes the same way. It was building a vector per ASU group to hold a handful of floats - over a million allocations, their growth and their frees, five times a run - where the counts were already to hand from the pass above. One flat array with a per-group span gives the identical median, since a median does not care how the multiset was laid out. Measured together with the previous commit on a high-multiplicity rotation set: 38.3 s -> 35.1 s. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
c9fc46e6e2 |
Stop making the first pass write the files the second pass replaces
Pass 1 exists to choose the space group and post-refine the geometry. Its merged intensities are discarded - pass 2 remakes them seconds later at the refined geometry, and that is the answer anyone reads. It was nonetheless writing the full set of merged files at the end of every pass 1: a mmCIF of every unique reflection (22 MB on an ordinary crystal, 48 MB on a crowded one), an .hkl, an .mtz and the per-image scaling table, all through one thread. Measured on an ordinary rotation set: 0.60 s of a 15 s run, and pass 2's identical block right after it takes another 0.585 s to write the files that are kept. The pass-2 quality guard is untouched, which is what disqualified an earlier attempt at this: has_merge_statistics is set at the merge, well above the write, so pass 1 still reports the completeness and CC1/2 the guard compares against. Nothing numeric moves - the same run measures 38.5 s before and 36.6 s after with a byte-identical .hkl. Also hoist the pixel-mask accessor out of the preprocessor's per-pixel loop. It called .at() on every pixel of the detector - 18 million bounds checks per engine, and an engine is built per worker per pass - for a bound the loop already respects, which stopped it vectorising. The <prefix>_01.mtz/.cif/.hkl are documented output, so this is a deliberate behaviour change: the pre-pass result is no longer written. If it is wanted for comparison it should come back behind a flag rather than by default. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
16639e9de2 |
Add up the profile accumulators in an order the schedule cannot change
Every accumulator in this file that could be an integer already is one, and the spot finder's reduce_rings_shared says why: a preprocessed pixel is an exact int32, integer addition is associative, and a threshold that moves in its last bits between runs flips every pixel sitting on it. Four accumulators here were still floats, and they reach the intensity rather than a diagnostic. * The radial background curve. s_radv and rad_sum sum int32 pixel values, so int64 is not an approximation of the old sum, it IS the old sum - and the curve is subtracted from every reflection's background. * The learned profile grid and its second moments, which are sums of (px - bkg) / I over every strong reflection of the frame. There is no exact integer form, so these are fixed point at 2^20: a quantum of 1e-6 of one I-normalised pixel, far below the Poisson noise of the pixel it came from, and some five orders of headroom inside a signed 64-bit accumulator. * The normalisation total in build_profiles, which divides every cell of the profile - 128 lanes on one address, in arrival order. Now summed as integers, exactly, from the grid it normalises. * The fit's own reductions, s_num and s_den among them, which ARE the fitted intensity. These stay float, so fixed point would be a real precision trade over an unbounded range; instead each warp leaves its total in a slot of its own and every thread adds the slots up by warp index. WARP_ATOMIC_ADD is order-independent for the integer accumulators it was written for and not for these, which is what block_sum is for. With the prediction ordering of the previous commit, a run is now reproducible: the same command on the same images writes byte-identical .hkl and .mtz, at -N 1 and at -N 48, on a 16 Mpx rotation set and on a large-cell one. Before, all four differed. The battery is unchanged where it was ever stable: space group identical on all 24 crystals, reflection count on 20, R_meas on 22. The two that move are the two the battery has always seen move between runs of an unchanged binary - which is the point, since they stop moving now. Total 8m02s against 7m55s, inside the noise of per-crystal times quantised to a second. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
484a0e162a |
Give the predicted reflections an order of their own
The GPU predictors claim their output slot with atomicAdd(counter, 1), so a reflection's position in the array is whatever order the blocks happened to finish in. That position is not private to the predictor. BraggOwnerKey packs it into the owner map as the tie-break between two centres equidistant from a shared pixel - the map's atomicMin is order-independent, but the number it compares is not - and the ingest and post-refine bucket sorts, whose comparators are deliberately not total, resolve their ties by the order they are handed. So two runs of the same binary on the same images integrated a different set of reflections. Measured on a large-cell rotation dataset: 63301112 observations against 63301139, and 89% of the merged intensities differing by more than 1% of themselves, median 1.8%. Single-threaded as well as at -N 48, which is what ruled out thread ordering and pointed here. Order the downloaded list by (h, k, l, delta_phi) before TruncateToOutput, whose own pick is then reproducible as well. hkl is a property of the reflection rather than of the schedule, and delta_phi separates the two rocking solutions one hkl can have. The CPU predictors already emit in hkl order, so the two paths now agree on it. Sorting a 20-byte key and gathering once, rather than sorting the 88-byte reflections in place, keeps this off the clock: on a crystal predicting some 35000 reflections a frame the run measures 52.2 s against 52.3 s before. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
c3236eed44 |
Publish the underload, not the error marker, on the finalized-file socket
The finalized-file notification has carried j["underload"] = error_value since the two meant the same thing. They stopped meaning the same thing when error_value became the marker the pixels actually store: UINTx_MAX for an unsigned image, where it used to be GetUnderflow()'s -1, a value no unsigned pixel can hold and which therefore excluded nothing. So for an unsigned 16-bit run the key went from -1 to 65535. A facility that forwards it into an XDS UNDERLOAD or a DIALS trusted range - which is what a key called "underload" is for - would reject every pixel below 65535, i.e. all of them. Nothing in HDF5 is affected and none of the writer tests look at this socket, so it fails silently and outside the file. The start message already carries underload_value, the lowest valid value: 0 for an unsigned image and INTx_MIN+1 for a signed one. Send that. The signed case moves too, from the marker itself to one above it, which is what the key has always claimed to be. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
0cca6945e2 |
Find the first pass's spots on every worker
The first pass built one analysis engine and walked its candidate frames through it serially, while the main image loop had been giving an engine to each of its workers all along. On a 16 Mpx dataset that phase was 21% of the run on one thread and one card. A frame's spots are a pure function of that frame and the settings - the engine carries nothing from one image to the next, which is exactly why the main loop can hand one to every worker - so the search parallelises without changing anything it finds. Each worker keeps its engine for the whole pass and takes a card by index, so an engine always meets the card it was built on, and the engines are released before the main loop builds its own; the peak is no higher than the main loop already reaches. Results land in a slot indexed by position and are inserted into the cache afterwards by the owning thread, and the feed loop still walks ordinals in order and still stops on the same condition. So neither a frame's spots nor the set of frames the indexer sees depends on how the workers interleaved: the lattice picked is the same one, on the same frames. Ported from 2608-performance with its two unrelated riders - the flag_strong quad read and the beam-stop shard allocation - split into commits of their own. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
9b1ffbaa71 |
Allocate a beam-stop shard's accumulators when it is first used
SetShardCount allocated and zeroed three per-pixel accumulators for every shard up front. With a GPU present none of them is ever written - the frames are decoded and folded on the device - and on a 16 Mpx detector eight shards are 2.9 GB to allocate and clear, measured at 0.8 s of the pre-scan spent on memory nothing reads. A shard now allocates on the first frame that reaches it, and the fold skips shards that never got one. Two things that go with it, not in the version on 2608-performance. Reduce's single-shard fast path returns that shard directly, which is now an EMPTY projection if nothing was ever added to it, where before it was a zeroed full-size one - and both callers index it by pixel. The fast path therefore requires the shard to hold at least one frame; otherwise the general path builds the zeroed projection as before. Also drops a duplicate include of ParallelFor.h. Split out of "Find the first pass's spots on every worker", which carried it as an unrelated rider. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
0047d1065e |
Read four pixels at a time when flagging strong pixels
The ring reduction already reads its pixels four at a time; the pass that flags the strong ones still read them one at a time, over the same image. Give it the same quad read. The flag is a per-pixel comparison against a threshold the reduction has already fixed, so nothing is summed here and the result is unchanged pixel for pixel, including the tail the quad read does not cover and the masked pixels it skips. Split out of "Find the first pass's spots on every worker" on 2608-performance, which carried it as an unrelated rider. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
74e8b77a0d |
Stop scaling and merging what the resolution range excludes
A crystal integrated to the detector corner but merged well short of it carries observations through the whole merge that the merge then discards. On the heaviest dataset in the rotation test set that is 63.3 M partials of which 6.4 M are ever used: the other nine tenths are sorted, uploaded, scaled, combined and error-modelled before anything looks at their resolution. Ingest copied every one of them unconditionally, and the d_min limit was first applied far downstream, in the ASU grouping. They are now dropped at ingest, immediately after the one big sort: - WHOLE raw-hkl runs are dropped, on the same rawrun_d the ASU grouping already tests. A per-observation test is not equivalent - a run is in or out today by one member's d - and using a different rule here would put the two out of step. - The drop happens AFTER the flux meter, which takes each frame's mean background over every reflection on it, and after the sort, so neither changes. - The compaction runs in index order, so a frame's observations stay contiguous and keep their order, and every per-frame sum keeps its sequence of roundings. The incident-flux divide goes with it: it was reading one int and dividing one float across 5 GB in a pass of its own. The per-frame mean it needs is now accumulated by the ingest fill loop - one frame, one thread, same order, so bit-exact - and the divide rides on the finiteness pass that already touches that field. Ported from 2608-performance with two changes. The ingest fill loop there had been parallelised by an earlier commit that is not being taken, so the mean background is accumulated in the serial loop this branch still has; it is the same sum in the same order either way. And the post-refinement sampling that commit also introduced - thinning the fit to 8 M partials by a hash of the raw hkl - is NOT included. Every consumer of the dropped observations is gated on the ASU group, so dropping them is a no-op for the science; thinning post-refinement is not, its own measurement puts the cell scale breaking at 4 M against a pool of 8 to 16 M, and it makes the fit depend on how far integration ran. That belongs to its own decision. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
be73288748 |
Fit the modulation surface on a grid that spans the detector
The detector-frame modulation correction takes its 16x16 grid extent from a pass over every full, but the surface is fitted only on the fulls that belong to an ASU group. Those are two different populations, and the gap between them is whatever was integrated past the resolution the merge uses. That made the correction's fate depend on how far integration reached. Cut it back and the grid contracts onto the merged disc while the cell count stays the same, so each cell holds too few reflections, the surface over-fits, and cross-validation throws it away - correctly, on a surface that should never have been fitted at that scale. Varying only the integration limit on one rotation dataset, merged R_meas came out 28.4 / 33.1 / 29.0 / 32.8 / 31.7 %, and the four-point spread is entirely the correction switching on and off: every low value is a run where it was applied, every high value one where it was refused, with no exceptions. Nothing else moved. The grid now spans the detector. Cells with no observations in them keep a factor of 1 and cost nothing, and with integration running to the detector corner - the default - the grid is the one it always was, so the common case is unchanged. On a crystal carrying no resolution limit at all it takes merged R_meas from 39.7 % to 35.5 %. This is a correctness fix in its own right. It also has to come first: without it, any change that narrows the integrated resolution range trips the same over-fit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2eb9780fe2 |
Weight the corrected intensity, not its factors, in the shared reference
Folding the fit loop and the score loop into one reference() had to pick one of their two spellings, and it picked the fit loop's: w * I * corr * a, where the score loop had built Is = I * corr * a first and then summed w * Is. Those differ in the last place, and of the two callers it is the score that decides whether a surface is kept at all - so a gate sitting on the fence could go the other way for no reason but the order of three multiplications. Sum w * Is, which leaves the deciding path spelled as it was and matches how the rest of this file accumulates a weighted intensity. The fit's own reference moves by a last place instead; it is iterated to convergence and then scored, so that is the cheaper place to absorb it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
ca3ca7170e |
Spread the scaling corrections and the space-group search over the cores
Two thirds of a rotation run is one thread. The image loop is not the problem - on the heaviest crystal of the battery it is 1.8 s of 40 - and neither GPU nor CPU is saturated, because while the corrections and the space-group search run there is one core working and 47 idle. Mean occupancy over the whole run: 3.9 of 48. In the correction surfaces (absorption in the goniometer frame, detector-plane modulation, absorption against time and detector position - all one function): the per-cell accumulation, the score reduction and the final apply are now chunked, as are the three loops that assign a full to its cell, one of which spends a sine and a cosine per full de-rotating it into the crystal frame. Two full sorts of four million floats went with them: only the nine bin edges are wanted, so they are selected instead, each selection starting where the last one left off. The per-group pass is deliberately left serial. The terms of one group are spread all over the list, so the only way to give a thread groups of its own is to walk in group order, and that trades a near-sequential read of the fulls for a random one over a few hundred megabytes - the trade that already lost once in the combine kernel. The space-group search scores each candidate rotation by correlating I(h) against I(Rh) over the whole merge. Every operator it can ask about comes from a fixed list and none of them depend on each other, so they are scored up front, in parallel, and the search reads the cache. The scratch that stops a pair being counted twice is now per worker rather than shared. Worker counts are gated on how much work there is, not on how many cores the machine has (ThreadsForWork). Both parallel helpers start a thread per chunk, so a small dataset on a large node would otherwise pay for 48 thread starts to sum a few thousand terms - and this runs on 8-core laptops as well as on this node. Measured on the heaviest crystal, idle machine, two runs each, summed over both passes: those phases go 7.88 s -> 5.19 s. Whole-run wall time is the wrong ruler for it - it moves +-4 s between identical runs. Battery 9m45s -> 9m23s, space group 21/24, no failures; 16 of 24 crystals bit-identical to the previous run and the rest inside the noise floor of running one binary twice. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
01b619cfd0 |
Take the double precision out of the box integrator's inner loops
boxsum summed its ring background in double and compared each ring pixel against a double threshold. The pixels are integers: a sum of at most a thousand int32 values is exact in a 64-bit integer AND exact in a double, so the two agree bit for bit, and comparing an integer against the floor of the threshold accepts exactly the same pixels as comparing it against the threshold itself. Both loops now do integer arithmetic. That was 39% of the card's double-precision pipe on the development machine and about three quarters of it on the production one, where the double rate is unchanged from Turing while the single rate has doubled - so this is worth more there than here. Alongside it, three things in the combine kernel. rr_nusable was computed by a whole extra walk over every observation and then never downloaded or read by anything. sum_wb and sum_cwb have no F in them, so they are the same in all three reweights and only the last round's values are ever used - two thirds of them were two divisions each, discarded. And CombineParams was the one parameter struct in the file without __restrict__, so the compiler could not assume the observation arrays and the freshly allocated fulls arrays were distinct. Measured on a crystal with 66 million partial observations: boxsum 12.2 s -> 8.3 s, the combine kernel 8.0 s -> 7.6 s, whole crystal 1m17s -> 1m12s. Battery 15m32s -> 9m59s. Same space group on all 24 crystals, none failed. Two things measured and NOT kept, recorded so they are not tried again: sorting the raw-hkl runs by length so a warp holds runs of similar length - it trades away the locality of neighbouring runs in the permutation and came out slower (7.6 s -> 8.8 s); and page-locking the integrator's host staging arrays individually - eleven separate registrations of small heap allocations overlap on shared pages and the driver refuses them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a8cca3e5d4 |
Parallelise the incident-flux divide, drop a redundant sync
DivideOutIncidentFlux was still the last fully serial pass in Ingest: a sweep over every observation to take each frame's mean background, and another to divide every rlp by its frame's flux. Ten gigabytes of traffic on one thread. The per-frame means go a frame at a time rather than an observation at a time, so each frame's running sum stays in one thread and in the order it had - splitting by observation would cut a frame across two threads and the partial sums would have to be recombined, which is a different sequence of roundings. The divide is per-element and splits anywhere. The adaptive spot finder synchronised after flagging strong pixels. The extractor that reads those pixels runs on the same stream, so the ordering already guaranteed the flagging had finished; the wait only idled the host, once per image. Measured on a crystal with 66 million partial observations: Ingest 8.5 s and 7.7 s -> 7.1 s and 6.6 s, whole crystal 1m24s -> 1m17s. Merged statistics unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8ae53b7fdf |
Say what actually makes the post-refine bucket sort reproducible
The bucketing commit justified itself with "the partials order became total in an earlier commit", which is true of the scale/merge ingest and not of this sort: part_less ends at the image number, so two partials of one reflection on one image tie, exactly as they did before. Nothing is wrong with the result. The counting-sort prefix lays each bucket out chunk by chunk, and a chunk is a contiguous span of the gathered order, so every bucket arrives at std::sort in global gather order no matter how many threads scattered it - the order is reproducible run to run and identical across -N. Giving Partial a rank field to make the comparator total would settle those ties by index instead, at eight more bytes on an array that reaches tens of millions of elements, and would change nothing anyone can observe. So state the invariant where the comparator is, rather than leaving the next reader to trust a claim that does not hold for this half. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
01d16231b3 |
Sort the partials and the post-refine events in buckets, in parallel
Both were one std::sort on one thread over tens of millions of elements, and together they were a third of a crowded crystal's run. Bucketing by h first makes them parallel. h is the comparator's leading key, so the sorted array is exactly the buckets laid end to end, and each bucket sorts on its own thread. In Ingest the keys are built straight into their bucket slot, so this replaces the build pass rather than adding one and the packed-key array is never duplicated; the extra memory is a few hundred kilobytes of histograms. Buckets are taken largest first, because the tail of the phase is whichever bucket finishes last. The run split falls out of the same structure for free: a run of equal (h,k,l) never crosses an h boundary, so each bucket counts its own runs, a scan over the buckets gives the offsets, and the arrays are sized exactly - which also removes the repeated growth the push_backs were paying for. The h range comes from the finiteness pass, which already reads every observation. The partials order became total in an earlier commit, when the observation index was added as the last key. That is what makes this safe rather than merely fast: the permutation is uniquely determined, so a bucket sort produces the same one a single sort would. Measured on a crystal with 66 million partial observations: Ingest 15.2 s and 14.3 s -> 8.3 s and 7.4 s, the post-refine event sort out of the top ten gaps entirely, the whole crystal 2m22s -> 1m24s. Battery 15m32s -> 10m05s. Same space group on all 24 crystals, none failed, and no crystal's R_meas moved by more than 0.3 points. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c1b85c7e88 |
Reduce within the warp before the Bragg integration atomics
fit and boxsum were 80% of GPU time on a crowded crystal - 78 s of it. Neither was bandwidth- or occupancy-bound: both sat at about an eighth of the issue rate the card can sustain, stalled. What stalls them is the block-wide accumulations. Every one has all 128 lanes of the block adding into one shared address, and a shared-memory atomicAdd on a float or a 64-bit integer has no instruction on either Turing or Ada - it compiles to a compare-and-swap retry loop. So those 128 lanes serialise into 128 retries, eighteen times per thread in fit. Summing across the warp first and letting one lane do the atomic leaves four per block instead of 128. That is the whole story: the arithmetic below was worth 2%, the atomics 5.6x. The arithmetic is still worth having, and is what was expected to matter: - compute_shell ran on all 128 threads of a block for a value that belongs to the reflection. It is two software double-precision divisions, on a card whose double throughput is a thirty-second (a sixty-fourth on the production one) of its single. One thread does it now. - The Kabsch inner loop divided by the same weight three times; the compiler emits the whole correctly-rounded sequence each time. One reciprocal now. Likewise the two Gaussian widths and the profile normalisation, which are constant over a reflection's cells and were divided per cell. - boxsum read the pixel before deciding whether it wanted it. The window is the bounding box of an ellipse, so nearly half of it is neither the signal disk nor the background ring, and those slots were fetching a cache line for nothing. Measured: fit 50.8 s -> 9.1 s, boxsum 27.5 s -> 12.1 s. A crowded crystal 2m22s -> 1m58s, a 16M-pixel one 39.5 s -> 37.2 s, the whole battery 12m30s -> 11m35s. Same space group on all 24 crystals, none failed. The integer sums are unchanged - addition is associative. The float ones move in their last bits and become more reproducible, since a fixed shuffle tree replaces whatever order the atomics arrived in. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
52756273e1 |
Make the partials order total, and hoist 1/sigma out of the IRLS loop
The sort that orders every observation by (h,k,l,image_number) was not a total order: two observations can genuinely share all four. The predictor emits BOTH intersections of a reflection's rotation circle with the Ewald sphere, and near the blind region - where zeta is smallest - the two are close enough in angle that both are accepted on the same frame. Which of them came first was then whatever the sort happened to produce. That was observable. The combine takes on_ice from the FIRST member of a rocking event, so the order decided whether a full was flagged as ice at all, and its per-event sums are floating point, so it moved intensities in their last bits. The observation's own index is now the final key, which orders them by arrival - and, more usefully, makes the order unique, so it no longer depends on which algorithm sorted it. sigma never changes once it is uploaded, so 1/sigma is the same in all thirty IRLS iterations of all three scaling iterations of all five scaling passes. It was being recomputed every time: a 64-bit reciprocal is a hardware estimate plus five refinement steps, and the profile put the three divisions in that loop at 21 of its 31 double-precision instructions. It is computed once now, in the pass that already streams every observation. The CPU has always hoisted it; this is the GPU catching up. Same expression on the same operand, so the value is what the loop used to compute, bit for bit. Also: PrepScaleObsKernel is not a grid-stride loop, but the scale-fulls path capped its grid at 65535 blocks like the grid-stride kernels around it. Above 16.8 million fulls that silently left the tail of sco_coeff/sco_ok stale. No dataset here reaches it; the cap is simply wrong for that kernel. And the AoS-to-SoA staging that feeds the GPU - the widest pass in Ingest, reading an 80-byte struct and writing fourteen arrays out of it - ran on one thread. Full 24-crystal battery: same space group on all 24, none failed, one crystal moved R_meas by 0.8 points with CC unchanged (it moves by that much between runs of an identical binary). 15m32s -> 13m35s. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
594100accc |
Make the rotation-scale fit reproducible, and stop refining past a float
Two follow-ups to the closed-form fit. The five per-fifth sums were reduced under a mutex, so the order in which the chunks were added depended on which worker reached the lock first and the fitted scale moved in its last bits between runs of the same binary. Each chunk now folds into its own slot and the slots are summed in chunk order, which is the reduction pattern the rest of the analysis code uses. The split ParallelChunks makes is fixed, so the sum is now the same sequence every time. The golden section bracketed to 1e-9. The fit is narrowed to a float before it is applied, and a float's epsilon is 6e-8, so the last ten or so iterations - each a full parallel pass over every event - refined digits that are discarded on the next line. Bracket to 1e-7. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n8riB6X59oRjkrSHzNPAU |
||
|
|
c4c2d598d7 |
Fit the goniometer rotation scale in closed form
The fit has ONE parameter, and it was handed to Ceres as one residual block per rocking event - 8 million of them on a large crystal. Each block is a functor, an auto-diff cost function and a loss object on the heap, and the solver then factorises an 8-million-by-one Jacobian on every iteration. It cost 13.7 s. The residual is closed-form in k. A rotation preserves length, so |p_lab| is |e_mid| whatever k is and only the z component moves; Rodrigues gives it exactly: r(k) = C + A cos(a k) - B sin(a k) = C + R cos(a k + psi) C = lambda |e|^2 / 2 + u_z (u.e), A = e_z - u_z (u.e), B = (u x e)_z with a the event's angle from the sweep centre. That is the same function the functor computes - Ceres uses the exact Rodrigues form here, so there is no small-angle branch to disagree with - and it reduces the fit to minimising a smooth function of one variable over the interval the solver was bounded to. It is scanned on a grid and then closed in by golden section; the objective's curvature jumps wherever an event crosses the Huber knee, which is why this is not a Newton iteration. The coefficients are computed in double and stored narrowed. Their rounding moves the minimiser by ~1e-10, and k is carried downstream as a float, so the committed value is the same to far more digits than anything reads. One pass over the events yields the five per-fifth partial sums, so the all-data fit and the five leave-a-fifth-out folds share it. That matters because the jackknife only runs when the fit is big enough to act on, and on a crystal that trips it the old code paid for six full solves. The partials gather ahead of it counted first and then filled instead of growing one vector by push_back tens of millions of times, which copied the whole thing on every doubling. Measured: unchanged verdict and k to five decimals on the regression crystals. Full 24-crystal battery: same space group on all 24, none failed, 15m32s -> 13m35s together with the scale/merge changes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |