7eb8e93a0afc4daea730ff6efc19ea885bc4119a
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
baf2017c97 |
Defective pixels: mask a persistent patch, not only a lone pixel; sub-lattice ask reads the R contrast
Two defects that together turned a tetragonal small-molecule crystal (open-arm cuhf2, published
P4/nmm) into P222.
1. HotPixelFinder masked a persistent strong pixel only when it stood alone or in a pair ("a
larger patch is a feature of the scattering"). On cuhf2 a ring of ~20 pixels reads 1e5-3e5
counts on EVERY frame (dead centre) - stationary in the lab, so no reflection of the rotating
crystal. Unmasked, the reflections crossing it ((2,6,+-9)) merged to 3.5e6 / 6.7e5 against 4e4
for the strongest real reflection. The final merge's outlier test removes them (12
equivalents), but the space-group search's P1-like merge has no equivalents to judge by, and
the h<->k operator read CC 0.18 / R 0.28 instead of ~0.99 / 0.013. The persistence and strength
tests are unchanged; only the isolation requirement is dropped.
2. The sub-lattice ask kept every metric two-fold that passes min_operator_cc. On a pseudo-cubic
cell two false cubic two-folds still correlate at 0.31-0.33, LePage closes them with the
genuine ones into the full cubic metric, and the tetragonal class in between is never asked.
An operator now also has to pass the R contrast every promotion's added operators are held to
(min_operator_r_contrast against random_pairing_r / global_best_operator_r, read only where the
search reads it): the false ones sit at 0.2, the genuine ones at 1.0.
cuhf2: P222 -> P422, R_meas 2.9% -> 2.6%, ISa 45.5 -> 52.3 (overnight-cint had found P422 by the
sub-lattice path; rc174-all lost it). Prescan survey, masked pixels base -> new: unchanged on
lyso_x06da_ref/5keV/atten_wedge, thau, insu, cytc, myob_split, 9qw8, aspirin20, HEPES, YAG,
metformin, nidppe; cuhf2 3 -> 42, 5reo 2 -> 6, lcystine25 0 -> 9 (two clusters of noisy pixels at
3-7 counts/frame on a zero background). Targeted battery (sm2rest-fix vs rc174all-ctl / smt-all):
the 15 protein / sub-lattice sets (lyso x3, thau, insu, cytc, myob_split, 5reo, 9qw8, 6iu6, 6iu8,
6iu9, 6z8o, 8t7r, 9ea5) identical to the 4th digit; SM sets identical except cuhf2 (above) and
lcystine25, SHELXL R1 .112 -> .151: the nine masked pixels flip the fulls-only per-frame scale
smoother on that polycrystalline sweep from "settled after 29 iterations" to "stopped settling"
(the base run reproduces bit-identically) - a sensitivity of that smoother, reported to the
scaling work, not of the mask.
Test: HotPixelFinder_PersistentPixelNotBragg gains a 3x3 patch that must be masked. The
device-vs-host test now waits for each queued frame before overwriting the shared device buffer;
it remains flaky at ~4/25 on the base binary as well (a borderline pixel or two differ), which is
pre-existing and not addressed here.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
|
||
|
|
d4f4bab158 |
Pre-scan defective pixels: keep the GPU finder's work on the device
The hot-pixel step of the pre-scan (MaskDefectivePixels) on a GPU build:
- The device half is built once, before the workers start (HotPixelFinder::PrepareDevice),
instead of by the first worker's frame while the others waited. The unmasked pixels
grouped by key are sorted on the device (stable radix sort: the same order the host
fill gave) instead of scattered on the host and uploaded.
- No per-frame host round trip: the ring-sector levels and lit thresholds are made on
the device from the order statistics, and the per-key frame and level sums are kept
there too; frames queue on their workers' streams and the per-pixel accumulation is
ordered by an event instead of a host synchronisation.
- The mask: the chance rate's per-ring counts are summed on the device, and only the
pixels the tests can pass (error value on most frames, or lit on at least
min(max(2, k_chance), valid frames)) come back with their sums - not the five
per-pixel arrays (470 MB pageable on a 16 Mpx detector). The host tests run on them
unchanged.
The threshold is written as fma(nsigma, noise, level) + offset on the host - what GCC
already contracted it to - and the device takes the same two roundings, so the levels
are bit-identical (and no longer depend on whether a compiler contracts).
Exact: hot-pixel mask and p.mtz byte-identical to
|
||
|
|
24e36ae740 |
Pre-scan: parallel beam-stop mask, leaner background beam-centre fit
Exact: p.mtz and the pre-scan products (shadow mask, mean projection, defective-pixel mask, ring and capture centres, compared as hashes and hex floats) are bit-identical to rc174 on three in-house rotation sets, GPU and CPU builds. - ShadowFinder::GetMask: the serial parts run in parallel - connected components by row band joined with union-find (both the shadow and the transmitting-arm searches, and the hole fill), ring binning and the harmonic sector gather by blocks, gap bridging by line; ring pixel counts read off the ring offsets. Mean projection filled in parallel. - ShadowFinder host accumulation: one band-locked projection instead of a 20 B/px shard per pre-scan worker (2.7 GB zeroed and folded on a 16M detector); SetShardCount and the shard argument are gone. - FindBeamCenterFromBackground: the usable-pixel test is made once, the in-band pixels are kept in pixel order so the clipping rounds no longer sweep the whole detector, the 67 MB cell map is gone and the per-iteration block fold runs in parallel - same sums, same order. - HotPixelFinder::GetMask: the chance-rate counts in parallel (integers). Measured on a loaded box (load ~25 from other jobs), pre-scan window: GPU 5.9-6.5 s -> 3.2-3.4 s, CPU 8.4-9.0 s -> 6.1-7.4 s. The GPU-build pre-scan now ends with its background spot measurement (CPU spot finder on ~120 frames, ~13 core-s on 8 workers). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB |
||
|
|
84228bf8be |
v1.0.0-rc.173 (#83)
Build Packages / Create release (push) Successful in 24s
Build Packages / build:viewer:macos-arm64:nocuda (push) Successful in 3m29s
Build Packages / build:rugnux:macos-arm64:nocuda (push) Successful in 2m43s
Build Packages / build:rugnux:linux-aarch64:cuda (push) Successful in 8m27s
Build Packages / build:rugnux:linux-x86_64:cuda (push) Successful in 9m53s
Build Packages / build:viewer:linux-x86_64:nocuda (push) Successful in 9m58s
Build Packages / build:viewer:linux-x86_64:cuda (push) Successful in 11m22s
Build Packages / build:jfjoch:rocky8:nocuda (push) Successful in 13m39s
Build Packages / build:viewer:windows-x86_64:nocuda (push) Successful in 18m37s
Build Packages / build:jfjoch:rocky9:nocuda (push) Successful in 16m32s
Build Packages / build:viewer:windows-x86_64:cuda (push) Successful in 24m11s
Build Packages / HDF5 consumer tests (DIALS, XDS) (push) Successful in 25m30s
Build Packages / build:jfjoch:ubuntu2404:nocuda (push) Successful in 19m3s
Build Packages / build:jfjoch:ubuntu2204:nocuda (push) Successful in 20m23s
Build Packages / build:jfjoch:rocky8:cuda-sls9 (push) Successful in 19m41s
Build Packages / Generate python client (push) Successful in 50s
Build Packages / Build documentation (push) Successful in 1m16s
Build Packages / build:jfjoch:rocky9:cuda-sls9 (push) Successful in 21m0s
Build Packages / build:jfjoch:rocky8:cuda (push) Successful in 18m38s
Build Packages / build:rugnux:windows-x86_64:cuda (push) Successful in 14m33s
Build Packages / build:jfjoch:rocky9:cuda (push) Successful in 17m55s
Build Packages / build:jfjoch:ubuntu2204:cuda (push) Successful in 20m50s
Build Packages / build:jfjoch:ubuntu2404:cuda (push) Successful in 18m38s
Build Packages / Unit tests (push) Successful in 1h46m14s
* jfjoch_broker: Optional per-dataset authentication - statistics, images and plots can require a bearer token, which jfjoch_viewer supports. * jfjoch_viewer: Dark mode and a theme-matched colour scheme, a magnifier panel, and simpler contrast and background controls. * Rugnux: Multiple performance improvements on GPU and CPU (CPU-only processing up to 40% faster, faster image decoding on ARM), with unchanged results. * Rugnux: `--model` rigid-body refinement runs on the GPU, and the model-validation check is faster and more reliable. * Rugnux: Improved scaling and merging - error model, outlier rejection, absorption correction and French-Wilson amplitudes now agree more closely with XDS and ctruncate. * Rugnux: Improved integration - radial background on powder and ice rings, crowded rotation data keep their reflections, and CPU-only builds integrate large unit cells as GPU builds do. * Rugnux: More robust detector geometry - measured beam centre, X-ray bandwidth and goniometer rate, and geometry refinement accepted only on significant evidence. * Rugnux: Merged files are written in the standard setting, or in the setting of a reference MTZ, structure-factor mmCIF or model, with its free-R flags. * Rugnux: Richer report - ice and powder rings, further lattices, superstructure candidates and mosaicity, with warnings worded as prompts to check. * Rugnux: Clear error messages when a data set needs more GPU or host memory than is available. Reviewed-on: #83 Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch> |