8874a788e68054c4445694f424eca00a2c6af459
226
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1f2dad530c |
Space-group search: the merge with all the observations decides
A rotation run makes the space-group decision twice: once on a merge of only the well-measured observations (--search-min-zeta) and once on all of them, and the rule was to keep whichever search found more symmetry. Its justification was that the filtered arm can only ever LOSE an operator - discarding 40-80% of the observations starves the operator correlations - and never invent one. That premise was checked at one set of geometries, and there it holds exactly: over 1232 stored battery runs (38 geometries x 32 code variants) the two arms disagree 59 times and the all-observation arm is the higher one every single time. Away from those geometries it fails. Over 68 runs whose first pass was given a displaced beam centre, the filtered arm confirms an operator the full merge refuses 14 times, and the two arms never once both confirm the promotion. On one rhombohedral crystal the all-observation merge refuses a 3 -> 32 promotion at the correct beam centre, on the same twin-law statistic the search uses everywhere (1.78 against a bound of 1.70), and the filtered merge - which had thrown two thirds of the observations away - overrides it into the wrong space group. The rule was validated on the only data that cannot test it. So the filtered arm no longer promotes. Where the two disagree the answer is the one the merge with all the observations supports, as the systematic absences already were, and both are still reported. The filter keeps the job it was added for - stopping a near-tangential measurement from making a real operator look like a twin law - it simply cannot outvote the merge that has every observation in it. Costless on the 38-crystal battery, as predicted: the changed branch is never taken there (41 searches: 39 agree, 2 with the all-observation arm higher, none the other way), the space groups stay at 35/38 with the same three misses, and the crystals the second search does rescue - where the all-observation merge is the one finding the higher symmetry - are untouched. Injecting the post-refined beam centre into the crystal above now yields the right space group, 27900 unique reflections at R_meas 12.6% against 14078 at 21.3% before. |
||
|
|
7ce47bd4d1 |
Radiation damage: report a measurement, or nothing
The monitor fitted each batch's relative-B on SINGLE observations -
ln(I_ref/I_obs) regressed on s^2, weighted by (I_obs/sigma)^2, with the
logarithm requiring I_obs > 0. The observation therefore sits in the
response and in its own weight, and the positivity requirement keeps only
the upward half of the noise, so the estimate is biased downwards wherever
I/sigma approaches 1 and is unbounded in the limit. Simulated: on a batch
with no relative-B at all and <I/sigma> = 0.3 it reads -28 A^2; on a batch
whose true relative-B is +30 A^2 it reads -29. The bias grows with dose,
so it inverts the answer on exactly the data the number exists for.
That is not a corner case. Re-measured on stored integrated intensities,
a 360 deg sweep obstructed over a 60 deg wedge - whose honest curve is flat
for 130 deg, dips over the wedge and comes back - printed a per-batch curve
saturated at -31.4 A^2 for twelve consecutive batches (the +-50 A^2 clamp
less the low-dose anchor, "no data here" reported as a measurement) under a
headline of -19 A^2 of radiation damage. A deliberately dosed dataset
printed -39 A^2 where the honest measurement is about +34: the one crystal
with real damage got the sign wrong. Eight of thirty-eight datasets
reported |dB| > 5 A^2 and their curves oscillate by tens of A^2.
So pool the observations into ten equal-occupancy resolution shells per
batch before taking the logarithm, and fit slope AND intercept over the
shell means, weighting each shell by its own pooled (I/sigma)^2. A shell
mean is well determined where a single observation is not, it admits
negative intensities, and it carries the I/sigma that says whether the
batch can be measured at all. The same simulations then reproduce the
truth to under 1 A^2 at every signal level. The intercept keeps a batch
that is merely dimmer than the run - an attenuated beam, a mis-fitted frame
scale - out of the damage number: a batch mis-scaled by 2x read +12 A^2 of
"damage" without it and +0.05 with it.
A batch whose shells are too weak to fit is now absent from the curve,
printed as "-", instead of pinned to the clamp. And the clamp itself is
now an argument of the solve rather than one shared constant: it guards
against divergence, and the correction keeps the bound it was tuned with,
but with the estimator fixed a heavily dosed crystal's honest relative-B
runs past it - the monitor pinned thirteen consecutive batches at +49 A^2,
which is the same defect in the other direction. The monitor is given room
a real relative-B cannot reach and drops any batch that lands on it anyway.
The shells are laid inside the range the run actually diffracted to,
not across the whole merged range: a resolution limit taken from another
program or left generous spends most of an equal-occupancy grid on noise
and leaves a batch with too few shells to fit at all - on the battery that
silenced three crystals outright and cost two of them nineteen batches of
thirty-six. Where the merged range already sits inside the signal the grid
is unchanged and so is every number.
The first->last headline
is reported only where a straight line explains at least half of the
curve's variance, or where the curve is flat to within a couple of A^2 and
the answer is simply "no damage"; otherwise there is no headline and the
report says the loss was not dose and points at the sweep-quality section.
The three shapes separate cleanly - progressive damage R^2 0.97, the
obstructed sweep 0.24, the clean control flat at +0.35 A^2. And the label
now follows the sign: damage fades the high-resolution intensity, so only a
positive change is dose, where before any |dB| > 5 was called damage.
Report-only throughout - the monitor never touches corr, and the per-batch
curve's only consumer beyond the report is a sweep-quality field no reader
reads; classification runs on the per-frame scale and CC, and is unmoved.
The decay correction's global slope and the opt-in per-batch relative-B
share this estimator and are left alone here: they fold into the scale, so
Full 38-crystal rotation battery, twice (the second confirming the shell
placement), against a clean baseline at the same base:
space groups unchanged at 35/38
merge metrics move on three crystals only - the same three whose two-pass
lattice search takes a different branch on nearly every arm run
this session, one of which moves its own R_meas by 1.5 points on
thread count alone
A report-only change ought to be bit-identical and this is not quite, which is
worth saying plainly: the three crystals that move are the known unstable ones
and no space group moves, but "identical except where nothing is ever identical"
is a weaker statement than "identical", and the residue has not been chased to
ground.
The three validation cases behave as they must:
60 deg beam-obstructed wedge, no decay -23.02, labelled damage, twelve
batches printing the clamp
-> NOT_A_TREND, curve within 3 A^2, the
two unmeasurable batches absent, and a
pointer to the sweep-quality section
genuine progressive damage -39.46, sign inverted
-> +94.50, monotone, corroborated by a
per-image CC that falls 0.608 -> 0.159
and never recovers
clean control +0.18 -> +0.76, flat within 1 A^2
Across the battery the report now names four crystals as radiation-damaged
instead of ten; the other three are the two lowest-energy datasets and the
pink-beam one, each showing a monotone rise of about ten square Angstroms.
Sweep-quality classification is untouched, and the coupling that was assumed to
exist does not: rad_damage_b_batch reaches it through one field that is written
and never read. Ranges and reasons are identical on 35 of 38, the three that
differ by one to seven frames are the same unstable crystals, and the census of
stretches called radiation damage is one before and one after.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
2222a4085c |
Scaling: correct the absorption that changes as the crystal turns
RefineAbsorption indexes its surface by the diffracted direction de-rotated into the crystal frame, deliberately without a time axis, and RefineModulation indexes its by detector position, also without one. Nothing is indexed by (rotation, detector position), so the part of the absorption that changes as the crystal turns has no parameter at all. For a rigid absorber illuminating a fixed volume that is the right model: the incident path is a function of the spindle angle alone and the per-image scale takes it, and the exit path is then fixed in the crystal frame. What breaks the factorisation is the diffracting volume moving - a crystal larger than the beam, a mis-centred loop, ice building up. The exit path then depends on the spindle angle as well as the direction, and no time-independent surface reaches it. Measured on 34 rotation datasets, on XDS's own uncorrected intensities, as what is left after the crystal-frame absorption and detector modulation surfaces have taken what they can. The cross-validation gate lets the surface engage on 22 of them. Scored per resolution shell - the gate's whole-range ratio is lowered by any resolution-dependent scale without a reflection getting tighter, so the honest readout is each shell's own ratio, which a per-shell scale leaves unchanged - the median engaged crystal gains 4.1 %, the set gains 117 % summed against 19 % of damage, and 3 of the 22 are hurt. The surface has to be smooth in rotation angle to be absorption at all, and it is: the lag-1 autocorrelation of the fitted factor along the time axis runs +0.32 to +0.71 on the crystals it engages, against -0.08 for the same surface with its time bins shuffled. Where it is not smooth it is fitting something else, and says so - on a sweep whose beam was obstructed for a 70 deg wedge the autocorrelation is +0.16 and the profile is a cliff at the wedge, not a turn. Two null controls. Assign every observation a random cell and the gate refuses it (-1.2 % to -3.8 %). Keep the detector bin and shuffle only the time bin - a surface that cannot contain any time-dependent information - and the gate refuses that too, at +0.04 %, -0.60 % and +0.33 % on three crystals. Against the real surface's +3.7 % to +14.8 % on the same three. 12 time bins x a 10 x 10 detector grid = 1200 factors. On the per-shell score the median gain moves only between 3.1 % and 4.1 % across grids from 216 to 2400 cells, so the grid is second order; 12 x 10 has the largest net and the fewest crystals hurt. Equal-occupancy detector bins, not equal width: an equal-width grid starves the edges and the corners, and a starved cell is where a free surface over-fits. Fitted last, so the two time-independent surfaces get first claim on what they can explain. QUALIFICATION, measured after this was written: the "33 better / 0 worse" above is overall R_meas, which is a ratio of sums across every shell and is therefore lowered by any resolution-dependent scale without a reflection getting tighter - the same property that let the estimator bias pass its own gate. Scored per resolution shell instead, this surface HURTS 6 of 18 crystals under the acceptance gate as it currently stands, because that gate shares the defect and admits the surface where it should not. With a per-shell gate the surface is refused on exactly those crystals and its net over the chain goes from +86.3 to +159.5 per cent with none worse. The correction is right; the gate that decides where to apply it is the next commit's problem, not this one's. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Full 38-crystal rotation battery against its own matched baseline - the same binary with the corrected estimator and without this surface: R_meas better 33 / worse 0, -78.7 R_meas_lo better 24 / worse 2, -20.7 CC1/2 better 8 / worse 0, +9.6 ISa better 30 / worse 2, +99.15 space groups unchanged at 35/38 The low-energy datasets gain most, which is what absorption should do: at 5-6 keV one crystal goes ISa 24.68 -> 37.42 and another 14.21 -> 22.08, while the same protein measured at 13 keV moves 13.41 -> 14.83. Taken with the estimator fix it precedes, against a clean baseline: R_meas_lo better 25 / worse 3 summed -26.0, ISa better 31 / worse 2 summed +113.1, outer-shell CC1/2 +83.0, observations +25 880 on 36 crystals of 38, CC1/2 flat at -1.4 and no space group moved. That last number is the point of the pair: the estimator fix alone costs CC1/2 -11.4, because the ramp it removes was partly standing in for this correction. One cost, predicted in advance and still unexplained: outer-shell CC1/2 falls on three of the four low-energy crystals, by 15.8 points on the worst, while every other statistic on those same crystals improves. The fourth goes up. On 5000-9000 That outer-shell fall has since been attributed, and it is not this surface: with the merge's 6-sigma outlier rejection turned off, the sign flips on every crystal that lost, +20.5 and +20.7 where it read -15.8 and -8.0. The surface removes most of the deviants in sample - it is fitted on all the data and applied to it, with no robustness of its own - so the merge's cut stops firing and the survivors land in a shell whose multiplicity is about three. Last-shell CC1/2 is largely made by that cut: one crystal's baseline goes 10.4 to 92.7 purely by dropping 19 per cent of the shell. A second qualification, measured after the numbers above were taken: overall R_meas is a ratio of sums across every shell, so any resolution-dependent scale lowers it without a reflection getting tighter - the same property that let the estimator bias pass its own gate. Scored per resolution shell instead, this surface hurts 6 of 18 crystals under the acceptance gate as it stands, because that gate shares the defect and admits the surface where it should not. Under a per-shell gate it is refused on exactly those crystals and its net over the correction chain goes from +86.3 to +159.5 per cent with none worse. The correction is right; where to apply it is the gate's problem. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b8e7e9c8bf |
Scaling: fit a correction surface with the observation as the response, not the regressor
ApplyCellSurface fits one multiplicative factor per cell by least squares with the OBSERVATION on the regressor side - A = sum w Is Iref / sum w Is^2, the slope that carries Is onto Iref. A least-squares slope is attenuated by the noise in its own regressor, here by 1/(1 + (sigma/I)^2), and an observation carries all of a reflection's noise where the reference carries about 1/n of it. So every cell is pulled towards zero by an amount set by its own signal-to-noise - and on a detector that is a function of radius, which is to say of resolution. The gauge fix then spreads the ramp over the whole surface and the alternating rounds compound it. Nothing downstream catches it. The cross-validation splits by frame parity, and a bias that depends only on a cell's signal-to-noise is identical in both halves. And the held-out score is sum|Is - Iref| / sum Iref over the whole resolution range, which any resolution-dependent scale lowers without tightening a single reflection: applying a scale that is purely a function of d to a merged 360 deg sweep leaves every resolution shell's R_meas unchanged to 0.05 pp and takes the run's overall R_meas from 42.4% to 27.5%. The surface finds that manipulation because its own bias points exactly along it, and reports it as a 40% held-out gain. The cost is large wherever a sweep was taken past its signal. On one such run the fitted detector-plane "flat field" ran from 0.25 to 4.0 with 37% of its cells pinned at the low clamp - an 11x centre-to-edge ramp - and on the same combined fulls it moved the merge: low-resolution R_meas 8.8% -> 15.5%, <I/sigma> 33.2 -> 13.4, CC1/2 99.65 -> 98.87, error model b 5.0e-03 -> 4.3e-02, ISa 11.0 -> 3.5. The program's own --no-scaling-corrections run agrees on the same fulls and the same space group (b 5.4e-03 ISa 10.6 against b 4.1e-02 ISa 3.6). Regressing the other way round restores all of it - 8.6%, 33.6, 99.65, 4.9e-03, ISa 11.1 - and keeps the surface's real gain in the middle shells, where CC1/2 rises by 1-2 points. Where the data are well measured everywhere the two estimators are indistinguishable: on a control crystal the two surfaces agree to 0.02 pp in every shell and 0.02 in ISa, which is what a flat field should look like. Recovering a synthetic +-20% ripple imposed on the same fulls: 0.18 rms in log against 0.91 for the shipped form on the weak-outer-shell crystal, 0.059 against 0.082 on the well-measured one, with the spurious correlation between the fitted factor and detector radius falling from -0.30 to -0.01. The same estimator serves the absorption surface and any resolution-indexed surface fitted through this function, where the bias lands directly on the resolution axis: on a crystal the program itself reports as having no radiation damage (total dB 0.00 A^2), a batch x resolution-shell surface fitted the old way manufactures a monotone 12% falloff from low to high resolution out of nothing, and the new way gives 1%. The per-frame scale (FitPerFrameG) already regresses this way round. Full 38-crystal rotation battery against the same binary without it. Read it with the next commit, which supplies the correction this one stops faking; alone it is a partial state: observations better 34 / worse 4, +22 940 R_meas_lo better 4 / worse 8, -5.9 ISa better 20 / worse 14, +32.66 CC1/2 better 2 / worse 10, -11.4 R_meas better 1 / worse 30, +114.3 space groups unchanged at 35/38 Overall R_meas RISES, and that is the artefact leaving rather than arriving: it is a ratio of sums across every shell, so a resolution-dependent scale lowers it without one reflection getting tighter - measured, a scale that is purely f(d) leaves every shell's R_meas unchanged to 0.05 pp while moving the run's overall value from 42.4 to 27.5 per cent. That is precisely the shape of the bias, which is why the surface's own held-out score read it as a 40 per cent gain. CC1/2 falling on ten crystals is not covered by that argument and is the reason this commit is not defensible on its own: the ramp was partly standing in for a real time-dependent absorption that nothing else modelled. With that correction supplied by the following commit the same battery gives CC1/2 -1.4 and R_meas_lo -26.0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e4ce573075 |
Bragg integration: keep a reflection that lost a wing, not one that lost its peak
MINPK asks how MUCH of the expected profile is readable. It does not ask WHERE, and the two are not the same question. The renormalisation argument the rescue rests on - a fit over a subset of a normalised profile is unbiased - needs the pixels to go missing for reasons unrelated to the reflection. A gap, a mask or the edge of the sensor is such a reason: the loss is set by the detector, and the fit renormalises over what is left. A pixel invalidated BY THE FLUX IT SAW is not: it goes missing because the reflection was bright, and it is the peak. Measured on the combined fulls, against the mean of the complete observations of the same reflection, in the innermost resolution shell of the high-multiplicity control and of a weaker crystal: a rescued reflection whose unreadable pixel sits within a pixel of the predicted centre reads |I - <I>|/I of 0.50 and 0.53, against 0.073 and 0.212 for a complete observation - 6.8x and 2.5x - and carries several times the mean intensity of its shell. On the control that is 0.21% of the shell's observations supplying 1.77% of the R_meas numerator; on the weaker crystal 0.52% supplying 6.82%. Rescues that lost only rim pixels are unremarkable by the same measure, 1.19x and 0.88x. Dropping the peak-losers alone takes the shell's R_meas from 7.440% back to 7.315% (unrescued: 7.307%) and from 22.03% to 21.06% (unrescued: 21.28%) - which is the whole of the low-resolution R_meas the rescue cost, and on the second crystal rather more. Raw frames say what they are. The pattern is a dead-centre invalid pixel with 5878, 9875 and 27583 counts around it: the detector's per-frame invalid marker on the brightest reflections. MINPK cannot catch them because it cuts on profile MASS, and the peak of a broad spot is a few percent of the mass. So a second condition, in the loop that already measures the readable fraction: no unreadable pixel may carry more than 0.9 of the profile's own peak value. A fraction of the peak rather than a radius in pixels because the peak is as wide as the spot - for a Gaussian the cut is at sqrt(-2 ln f) sigma, 0.46 sigma here, which is the peak pixel alone where sigma is 0.8 px and the crest of the ridge where the profile is a bandwidth streak. Swept against the alternatives on two crystals: a fixed radius needs 1.0-1.5 px to do the same work and costs 3-9x more observations for it, and 0.5 px does not reach the peak of a sub-pixel-offset prediction at all; tightening the fraction to 0.5 or 0.2 buys nothing beyond 0.9 and costs 7x more. Six crystals, three detectors, against the rescue as it stands: the rule keeps 99.86-99.96% of the recovered observations and returns R_meas to its unrescued value or below (4.6 -> 4.5%, 6.7 -> 6.6%, 25.1 -> 25.0%), R_meas in the innermost shell likewise (2.7 -> 2.6%, 5.9 -> 5.3% against 5.4% unrescued, 16.5 -> 16.4%), <I/sigma> up or level everywhere, and every unique reflection the rescue won is kept. Raising --overlap-minpk to 0.90 instead reaches the same place on two of them and short of it on the third, while discarding 0.8% of the recovered observations rather than 0.05%. An elongated pink-beam profile on a 9M detector and an EIGER2 16M dataset are both untouched at 99.9%, so the crest protection does not over-reject a streak. One crystal is not improved: a dataset whose error model rugnux declines to fit for want of strong reflections, whose <I/sigma> is <= 0 in eight of its ten shells and whose R_meas is undefined in as many. There the rule costs about 3% of <I/sigma> in the one shell that has signal, reproducibly, on top of the 9% the rescue itself costs there - while its overall R_meas moves 1.5 points on nothing but the thread count. The parity test gains four sections. Unreadable pixels were only ever punched into empty sky, so neither the rescue nor this rule had any CPU/GPU coverage at all; they now go into the signal disks - the peak of every fifth reflection, ~1.1 sigma out of every seventh, the disk edge of every eleventh - for both profile modes, a box sum and an elongated stencil, with a check that the clipping actually costs reflections so the coverage cannot go quietly vacuous. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Full 38-crystal rotation battery against the rescue without this rule, both on the same base: ISa better 19 / worse 4, +0.73 CC1/2 better 3 / worse 1, +1.3 R_meas_lo better 4 / worse 3, -0.3 space groups unchanged for 17 770 observations, 0.09 % of the run total and under 2 % of what the rescue had won. The two crystals whose peak-loss population was measured beforehand land on their predicted values: a tetragonal reference goes R_meas_lo 2.7 -> 2.6 % and ISa 27.11 -> 27.42, a cubic insulin 5.9 -> 5.3 % and 20.34 -> 20.65. One crystal pays: a cubic case with 2381 unique reflections goes R_meas 8.8 -> 9.6 % and ISa 4.08 -> 3.49. It is the crystal in the battery with the fewest uniques, so its rescued population is small and its shell statistics are coarse, but the loss is real and not noise in the R_meas. The R_meas sum over the battery reads +1.3, of which +3.2 is one crystal whose R_meas moves 1.5 points on thread count alone; without it the sum is negative. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
12ae6c5228 |
Bragg integration: fit a reflection over the pixels it has, not only over all of them
A predicted reflection was discarded outright if ANY pixel of its signal disk was unreadable - masked, untrusted, in a detector gap, or overloaded. On a battery crystal that is 11.1% of all predictions, thrown away for a defect in one pixel of fifty, and the pixels concerned sit at fixed places on the detector, so the loss is systematic in reciprocal space rather than random. Neither XDS nor dials does that. Both estimate the missing part from the profile instead and keep the reflection while enough of it was seen: XDS's MINPK (default 75%, "the missing intensity is estimated from the learned profiles"), dials' integration.profile.valid_foreground_threshold (default 0.75). MOSFLM is the one program that rejects by default, and even it relaxes to 50% with PROFILE EDGE. We already had the argument and the machinery: a profile fit is the amplitude of a NORMALISED profile, so leaving pixels out renormalises the estimator by construction - it costs information, which sum P^2/v duly loses and sigma duly gains, and biases nothing. That is exactly why --overlap exclude drops a neighbour's pixels from the fit rather than the reflection. Unreadable pixels are the same case with a different reason, so they take the same treatment, cut on the same threshold, in the same place: the readable fraction of the expected profile, measured against the profile mass that lands on the detector at all so a reflection is judged on the pixels that exist. A box sum has no profile to renormalise with and keeps the all-or-nothing rule. Two consequences handled. The summation seed and its variance now count the pixels actually read, and the runaway guard scales the fit back to that same disk before comparing - both exactly as before wherever nothing is missing. (Its fallback then hands back that partial sum unrescaled, which would read low; the guard fires on 8 of 96 688 recovered reflections, and on none at all on a weak crystal, so it is not worth a branch.) And the profile, its resolution shells and their widths are learned from COMPLETE reflections only, as is the box-sum centroid post-refinement reads as an observed position: a disk with a hole gives a centroid pulled away from the hole, and the hole does not move between frames. That sigma gains what the missing pixels carried is the claim the whole change rests on, and it is measurable. Force the conventional CENTRED cell of a body-centred crystal in P1: the predictor then enumerates every lattice point, and the reflections the centring makes systematically absent have a true intensity of exactly zero, so their scatter about zero must equal their reported sigma. Over 7.1 M such observations, matched by resolution shell, the trimmed std(I)/rms(sigma) of the recovered reflections is 0.99 / 1.20 / 2.33 / 1.04 / 1.69 against 0.98 / 1.22 / 2.29 / 1.03 / 1.56 for the reflections that were complete - the same calibration to a few percent. The lever there is small, because the typical recovered reflection is missing only 5% of its disk. Lowering the threshold to 0.50 admits a band missing 25-50%, which is a real lever: there sigma comes out 8-43% larger than a complete reflection's in the same shell, and the scatter about zero tracks it, 0.97 / 1.09 / 1.92 / 0.99 / 1.37, at or below the complete population. Sigma grows, and by the amount it should. The threshold stays at XDS's and dials' 0.75, on that evidence and on quality. Below it the estimator starts to run out: on those same zero-intensity reflections the recovered ones read +0.8 counts high at 0.75 and +1.9 counts high in the 0.50-0.75 band, against a sigma of 12-17, and at 0.25 the fit degenerates outright, single reflections carrying sigma in the thousands. Above it there is nothing to buy: 0.90 leaves a fifth of the recoverable observations behind and measures no better for them. On the high-multiplicity control, R_rim over as-shipped / 0.90 / 0.75 / 0.50 runs 4.49% / 4.51% / 4.56% / 4.78% while <I/sigma> runs 33.47 / 34.02 / 33.89 / 33.43 - 0.50 is where the recovered observations stop paying for themselves. Probe against the previous commit, six crystals. The high-multiplicity control gains 4.2% more observations, 924 803 -> 963 946, which lands it on XDS's 961 379 from the same images, for <I/sigma> 33.47 -> 33.89, R_rim 4.49% -> 4.56% at 4.3% more multiplicity, CC1/2 unchanged at 0.9998 and ISa 27.80 -> 27.12. Five weaker crystals gain 3.3-4.8% of their observations and up to 1.0 point of completeness, for <I/sigma> +0.4 to +3.6%, R_rim between -8.1% and +5.8% relative, CC1/2 +6.6 / +0.3 / +0.2 / -0.0 / -1.2 points, and ISa between +0.3% and -3.4%. Some of that ISa is the point rather than the price: a reflection integrated over fewer pixels carries less information, and the absence test above says the sigma that reports so is honest. The GPU and CPU engines agree as before. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Full 38-crystal rotation battery, against the same binary without it: observations better 38 / worse 0, +937 100 unique refl better 30 / worse 0, +9 229 overall <I/sig> better 33 / worse 1, +7.00 CC1/2 better 5 / worse 1, +6.2 space groups unchanged at 35/38 Every crystal gains observations and not one loses a unique reflection. The two costs are small and both are understood. Low-resolution R_meas is worse on eight crystals, by +0.8 pp at most and +3.2 pp summed - a reflection whose own peak pixel is unreadable loses the part of the profile that carries most of the amplitude, and that population sits at low resolution; the following commit handles it. And ISa falls on 32 crystals, by 10.9 summed, which is what admitting 937 000 further observations does to the strong-reflection asymptote: R_meas excluding the one crystal whose thread-count noise is 1.5 pp is flat. |
||
|
|
677ece7b59 |
Post-refinement: correct a goniometer that turned further than it was told
The angles a rotation dataset stores are the COMMANDED ones, so a stage whose travel is miscalibrated leaves no trace in the header - every angle is self-consistently wrong. No existing parameter can absorb it either: the cell scale, the axis direction, the detector distance and the beam centre are all orthogonal to an error in rotation MAGNITUDE. So fit it as what it is - one scalar k, the ratio of the travel to the commanded angle - on the rocking events the geometry post-refinement already builds, after step A so the cell scale and the axis direction are fixed and k is the only free quantity. Two details decide whether the number means anything. The angle enters measured from the CENTRE of the sweep: the reference orientation was fitted against the commanded angles and has already absorbed their mean error, so measured from the goniometer's zero instead a constant missetting about the spindle leaks into k with a gain of <phi>/<phi^2>, which depends only on where the sweep happens to sit - on a short sweep starting near zero a 0.14 deg missetting fakes 1.4 % of k. Referred to the sweep centre that leak is identically zero at any width. And the robust loss is scaled to the scatter the events actually have, which varies by more than a decade between datasets, so any fixed constant is either inert or throws away real data. A stage fault is rare and a 1 % angle correction applied to a healthy dataset would damage it silently, so the correction is committed only when every test passes: at least 30 deg of sweep and 5000 events, |k-1| over 0.5 %, a misorientation of at least 0.5 deg at each end of the sweep, and the same k from every fifth of the sweep left out. The last test is not optional. A second lattice that dominates ONE END of a sweep - exactly what happens where the primary stops indexing - fakes a k that passes the other two, and the hkl-hash split used elsewhere in this file cannot see it, because both of its folds sit at the same angles and anything structured in phi survives in both. When it commits, the second pass re-integrates against the corrected angles. The pre-pass mosaicity is dropped with it: that is a width in degrees fitted against angles the second pass has just stopped using, and since the override can only ever raise the second pass's own estimate, carrying it over would hold the second pass at the rocking width the uncorrected angles produced - the correction half-applied. --rotation-scale asserts a known stage calibration by hand and overrides the fit. On the 38-crystal rotation battery the gate fires on exactly one dataset, at k = 1.01318 with 0.74 of that k surviving every fifth left out. The largest of the other 37 is 1.00211, which fails the end-error test; 34 of them sit below 1.0006. On the one that fires: R_meas 39.2 -> 23.9 % (XDS 37.1) CC1/2 86.5 -> 96.0 % (XDS 94.3) CC1/2 outer 1.4 -> 53.4 % (XDS 42.5) unique refl 40990 -> 41540 (XDS 41322) observations 74975 -> 103858 (XDS 129322) mosaicity 0.181 -> 0.159 deg which takes it from losing to XDS on R_meas, CC1/2 and outer-shell CC1/2 to beating it on all three, and the mosaicity drop is the inflation the uncorrected angles were producing. Its low-resolution R_meas is the one number that moves the wrong way, 12.0 -> 13.9 %, still well inside XDS's 18.3. No space group moves anywhere, and every other crystal's merge is unchanged beyond the two-pass loop's own jitter - measured here as the spread of the post-refined distance across arms that do not touch post-refinement at all, which is larger than anything this commit produces. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e5c0066129 |
rugnux: write a results report next to the reflections
Everything a run determines went to stdout and nowhere else. The space group and the evidence behind it, the error model, the post-refine commit-or-reject decisions and their held-out residuals, the two-pass adopt-or-roll-back, the resolution cut, the merging statistics - all of it scrolled past interleaved with progress lines and was gone. A user who was not watching had no record, and nothing could read it. `rugnux` had no log file at all; the `rugnux.log` in the regression harness is that harness capturing stdout. Write `<prefix>_report.txt` alongside the .cif/.mtz/.hkl, always, with no option to ask for it. It holds what the run DETERMINED; timing, rates, per-image progress and engine chatter stay on stdout, where they belong. Every line rugnux logs was classified result-or-process against the regression corpus to decide what crosses over. The format follows XDS's CORRECT.LP, which has been read by people and parsed by other programs for twenty years: `KEY= value` assignment lines a script greps one at a time, fixed-width tables with stable headers and a total row, `WARNING:` sentences in plain English, section banners. REPORT_VERSION says when that interface last changed. It is assembled from results the pipeline already computed, so an unconditional file costs nothing, and a failure to write it is logged and swallowed - a run that produced good reflections must not be lost to a side file. One thing CORRECT.LP does not have to solve: a rotation run integrates twice and writes both passes, so every report says which pass it describes and why that pass was adopted. `--no-merge` gets a report too, saying MERGE= NOT_PERFORMED rather than leaving a reader to infer it from absent sections. An empty output prefix still writes nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b9078d9a59 |
Scaling: drop a frame whose scale collapses, do not merge it unscaled
Two guards catch a per-frame scale far below the run median. Both then invented a value for it - one substituted the run median, the other set corr = 1 and merged the frame "unscaled". For a frame whose scale really is 1/17402 of its neighbours', asserting 1 is worse than asserting nothing, and it is the assertion that does the damage: those observations enter the merge at full weight carrying an intensity scale that is wrong by four orders of magnitude. It surfaced when rotation started integrating every frame the sweep's lattice explains, but it is not caused by that change - six crystals in the battery already tripped these guards before it. What the extra frames did was find a crystal where the collapsed population is large enough to dominate: R_meas 19.1 -> 90.1%, ISa 25.60 -> 4.58, from 17% more observations. The frames are not sparse and the fit is not running away. A per-frame dump shows 3394 observations on the median collapsed frame against 3469 on live ones - the scale is over-determined 3400:1 for one parameter - and 113 frames fit exactly zero. They form one contiguous arc of about 68 degrees once the sweep's wrap is accounted for, over which the per-frame correlation to the merge is 0.035 against 0.85 elsewhere, while the flux measured from the background varies by only 1.55x. So the fitted zero is a well-determined measurement that the frame holds no diffraction from this lattice, not a failure to measure. The frames are empty, not under-determined. That is also why the smooth or shrunk alternatives do not apply, and both were built and measured rather than argued away: giving a collapsed frame the geometric mean of its credible neighbours is worse than the baseline (R_meas 115.3%), because it merges noise at the weight of a good frame, and a dead region 112 and 232 frames wide has no local neighbourhood to borrow from in any case. Dropping them: R_meas 90.1 -> 38.3%, low-resolution R_meas 26.9 -> 10.2% (past XDS's 14.3), ISa 4.58 -> 22.00, CC1/2 99.0 -> 99.9, with 5.4% more observations retained than before frames were integrated at all. Over the full battery, against the same binary without either change, ISa moves from -22.5 to -3.4 summed, CC1/2 +24.6, and 216247 more observations. The crystal that motivated the integration change is untouched by this one, bit for bit. The detection and the MIN_CREDIBLE_SCALE_RATIO threshold are unchanged. Note that threshold is now marginal: its own comment records 0.070 as the smallest legitimate ratio seen, and one crystal here has a legitimate live frame at 0.026, so it cannot be raised to catch the partly-dead transition frames at the edges of an arc without risking real data. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9885d1fc28 |
Rotation: integrate every frame the sweep's lattice explains
A rotation dataset has ONE lattice. Once the first pass has found it and the goniometer gives each frame its orientation, every frame of the sweep is a frame of that crystal - yet integration was gated on each frame re-indexing on its own, a test that carries an absolute floor of 9 indexed spots. A weakly diffracting crystal shows a handful of spots per image while the geometry still puts ~1500 reflections on the detector, so the floor threw away whole frames that had nothing wrong with them. Measured on a 360-degree battery crystal: 1484 of its 1800 frames failed that gate, all of them on the spot-count floor alone and none on the consistency test - the median failing frame had 4 spots and the lattice indexed all 4. Integration therefore ran on 17.7% of the sweep and the merge came out 35.7% complete at multiplicity 1.1, against XDS's 97.7% at 2.81 from the same images. XDS's own INTEGRATE.LP shows why the floor is the wrong test there: 964 of its frames have fewer than 9 strong spots and it predicts ~1483 reflections near the Ewald sphere on every one of them, because INTEGRATE works from the global orientation and has no per-frame indexing gate at all. Neither does dials.integrate. Split the one verdict into the two questions it was answering. "Does this frame index?" - what the indexing rate reports and what the first pass scores candidate lattices on - keeps the floor, because a handful of spots sit on almost any lattice by chance. "Is this frame worth integrating?" keeps only the consistency part, and only where the lattice does not come from this frame. A frame whose spots largely MISS the lattice is still refused: on another battery crystal that is 35% of the sweep, and integrating those collapsed the space group to P1 - the floor had been shielding the merge from frames the model does not describe, which is a different defect and not one to paper over here. Two consequences had to be handled. A frame that is too sparse to index is also too sparse to fit its own rocking width, and the placeholder it used to predict with was being reported onward as if measured, into the frame-order average that recomputes every partiality; report nothing instead, and fill the gaps in that average with the run's median rather than a fixed default. Probe (XDS in brackets): the crystal above goes 9 700 -> 81 956 observations, 8 618 -> 23 960 unique [23 576], 35.7% -> 99.4% complete [97.7%], R_meas 21.2% -> 68.6% [76.7%], CC1/2 96.0% -> 86.4% [81.1%], low-shell R_meas 7.2% -> 14.3% [20.6%], ISa unmeasurable -> 13.8 [10.4] - better than XDS on every statistic, where before it was merging a third of the data. A second crystal gains 41% more observations with R_meas 12.6% -> 8.5% and ISa 3.3 -> 3.7. The high-multiplicity control is unchanged to 2 observations in 924 782, and four further crystals move within recompilation noise. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d8029524e7 |
Scaling: never let an observation's own fluctuation set its weight
A weighted mean is only unbiased while the weights are independent of the values
being averaged. The IUCr's own nomenclature report (Schwarzenbach et al., Acta
Cryst A45 (1989) 63-75) puts it directly: weights in averaging "should not be
based on the counting statistics of the individual observations whose estimated
variances are biased and result in larger weights for accidentally low
intensities". Two places in the rotation pipeline were doing exactly that, and
between them they drove whole resolution shells of merged intensity negative.
1. The profile fit computed its non-signal variance as
var_bkg = max(0, 1/den - max(0, I) + bkg-estimate term)
The point of a separate var_bkg is that it does NOT move with the
reflection's own fluctuation, and 1/den - I is the quantity that does not:
1/den is the fit variance taken at the fitted intensity and grows with it
roughly one for one. Clamping the subtrahend at zero left a down-fluctuated
reflection's own deflated variance standing as its background variance.
Measured over 6.9 M partials of one weak rotation dataset, var_bkg/bkg came
out at 3.7-5.4 for observations with I < 0 against 11.4-13.7 for I > 0 - the
down-fluctuated half of every reflection carried a variance ~2.7x too small
and was weighted up by the same factor, first in the 3D combine and then
again in the merge. Removing the clamp makes var_bkg flat in I (~13 x bkg
across the whole range).
2. The merge then weighted each combined full by 1/sigma_full^2, and sigma_full
is by construction a function of the full's own answer: the combine's
variance carries a corr*max(0, F) signal term, so every full with F <= 0 got
the smallest variance the model allows while the strongest quartile got
2.26x more. The merge now rebuilds that variance at the reflection's mean
instead, from a linear model var(I) = var_bkg + var_per_I * I that the
combine measures and stores on the full. This mirrors
MergeOnTheFly::CorrectedSigma, whose comment already claimed to mirror the
rotation combine.
Verified against an estimator that cannot see the fluctuation - summing the
partials and dividing by the summed partiality, the classical construction every
other program uses (Greenhough & Suddath, J. Appl. Cryst. 19 (1986) 400-409, via
Leslie, Acta Cryst D55 (1999) 1696-1702: profile fitting biases the individual
partials but not their sum). Reproducing the merge on dumped observations, the
shipped weighting sat ~1.9 sigma below that reference in the noise shells; the
two changes recover most of it, and every intensity-independent weighting
scheme agrees with the reference once (1) is in.
Four-crystal probe, XDS resolution limits, branch fingerprint identical on all
four (so none of these is a two-pass branch flip):
weak cubic case last shell <I/sig> -1.6 -> +0.2 (XDS +0.10), last shell
R_meas 478% -> 250% (XDS 246%), overall <I/sig> 6.1 -> 7.5
(XDS 7.18), R_meas 18.3% -> 18.1%, CC1/2_hi 38.2% -> 43.7%
tetragonal case outer shells <I/sig> -0.4/-0.8/-0.9/-1.0 -> +1.8/+1.2/
+0.9/+0.4, R_meas 184%/595%/7614%/nan -> 95%/119%/135%/232%
(the nan was the shell mean crossing zero), R_meas 33.3% ->
32.9%, CC1/2_hi 38.3% -> 56.5%
trigonal case R_meas 13.0% -> 12.5%, CC1/2_hi 14.4% -> 16.5%
strong control unchanged to every printed digit but ISa
Cost: ISa falls (17.2 -> 14.0 and 16.7 -> 14.9 on the two mid-strength cases,
28.3 -> 27.8 on the control). Strong reflections are untouched by (1) - their
partials are all positive, so var_bkg is bit-identical - but the joint a/b fit
redistributes: honest weak sigmas lower a, and b rises to keep the strong bins
fitted. The median reduced chi^2 improves (1.25 -> 1.14, 1.35 -> 1.28) so the
new split describes the scatter better, but ISa is the one headline metric that
moves the wrong way and it should be watched over the full battery.
The integrator change is shared, so the stills merge sees it too; there it feeds
GetExpectedVarianceMerge, which had been handed the same contaminated var_bkg.
That path is untested here.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
43627e22dc |
docs: a rule for licences and academic credit, and apply it
Several methods adopted recently came from other crystallographic packages - the screw-absence test from POINTLESS, MINPK and the profile-fit reweighting from XDS/Otwinowski, the CC1/2 cutoff and merge outlier rejection from DIALS, the per-frame indexing gate from CrystFEL - and nothing in the repository said where such a debt is recorded. The licence side was already worked out (licences beside the vendored code, verbatim texts in licenses/ collected by COLLECT.sh, a row in THIRD_PARTY_NOTICES.md, all installed under share/doc/jfjoch); the credit side was ad hoc. Write the rule into CLAUDE.md. It states the distinction that matters: vendoring or linking someone's CODE creates a LICENCE obligation, discharged in licenses/ and THIRD_PARTY_NOTICES.md; reimplementing an algorithm from a PAPER creates none of that but creates an obligation of academic CREDIT, discharged in docs/ACKNOWLEDGEMENT.md and in a comment at the algorithm. Neither substitutes for the other, and taking both source and paper incurs both. It also fixes the citation form (authors, title, year, journal, volume, pages, verified DOI), and says in-source credit goes at the algorithm, not the file header, in the one-line style the code already uses. Then bring the repository into compliance for the works concerned: docs/ACKNOWLEDGEMENT.md gains a section acknowledging XDS, DIALS, POINTLESS/CCP4, MOSFLM, CrystFEL, GEMMI, the Kabsch/Otwinowski profile fit, the Diederichs & Karplus statistics and the IUCr nomenclature reports, each with a DOI checked against Crossref; docs/CPU_DATA_ANALYSIS.md's reference list gains the ones it was missing; and four algorithms gain a line naming their source where no adjacent comment carried one. No licence change. licenses/ and THIRD_PARTY_NOTICES.md are untouched. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6d39a4e1ab |
Space-group search: decide a screw from the evidence, not from a count of absences
A screw's predicted-absent class was required to hold min_absent_observed = 8 reflections before the screw could be claimed. That count is the wrong measure of evidence, and it is wrong in both directions. A screw extinguishes one row of reciprocal space, and that row is often the one a rotation sweep records least: it lies near the spindle, where the blind cusp maps onto itself and symmetry cannot fill it in. Counting it measures the geometry of the sweep. A monoclinic crystal whose 2-fold sits 7.6 deg from the spindle contributed six 0k0-odd reflections, every one of them measured between -0.013 and 4e-5 of the shell mean with zero violations, against a 0k0 row averaging 1.44x the shell mean - and was refused its 2_1 for being six rather than eight. XDS's own integration of the same images finds seventeen of those reflections and every one of them is likewise dead. The count is equally wrong the other way: a uniformly weak axial row produces no violations at all, so with enough reflections on it a screw is claimed from no evidence whatsoever. The second new test section demonstrates exactly that on the old gate. Judge the class by how unlikely it would be if the screw did not exist. Under "no screw" the absent class and the rest of its row are both Wilson-distributed with the same mean, so with each absent intensity taken in units of its row's control mean, sum_u/(sum_u + n_control) follows Beta(n_absent, n_control) exactly; the reported evidence is -log of that lower tail. The row's own strength cancels, which is the property the count lacks, and the scale is set by the number of reflections, so few-but-decisive and many-but-marginal are told apart. It is sigma-free by design: the merged sigma carries the error model's intensity-proportional term and so shrinks with I, reading much the same on an absent reflection as on a present one. This follows POINTLESS (Evans, Acta Cryst D67, 282-292 (2011), Appendix A3), which likewise scores an absence against the rest of its own axial row rather than against a global mean or a fixed cut, and likewise lets confidence fall away with the number of axial reflections instead of refusing outright below a count. POINTLESS calibrates its null width from control transforms of non-axial reflections; the Beta tail here is an analytic null in its place. XDS is not a reference for this: it "deliberately avoids any test for the presence of screw axes as these tests would depend strongly on the completeness of the data" (Kabsch, Acta Cryst D66, 133-144 (2010), section 6), so a screw axis in a CORRECT.LP was supplied to it, not determined by it. Measured over five probe crystals, genuine screw conditions read 34-800 nats and false ones - the 4_1/4_3 conditions of a cubic crystal that has no screw, whose predicted-absent class is STRONGER than its control row - read -7 to -8.5. The bound is set at 20, in the gap, at p <= 2e-9: three well-measured dead axial reflections clear it and two do not. min_absent_observed keeps its job for CENTERING, where a count is a fair measure - that class is a third to a half of every reflection in the data set and the bound is never binding on a centering that exists. The candidate table now prints the screw-absent count and this evidence in place of the two E^2 medians that were its raw ingredients, so a refusal can be read off the log. Measured on the five probes: the monoclinic crystal above returns to P2_1 with every merge statistic unchanged (R_meas 58.8 -> 58.7%, CC1/2 49.1 -> 49.3%, ISa 6.61 -> 6.59 - P2 and P2_1 share a point group, so only the symbol and the absent reflections differ). The other four are untouched, space group included, and the two-pass branch fingerprint (indexed frames, distance, mosaicity) is identical on all five. The full battery has not been run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7bbf072ad2 |
Scaling: do not report an ISa that was never measured
The error model's systematic term b is identified only by the spread of I^2/sigma^2 across the intensity bins the fit uses, and those bins hold equal COUNTS. So when fewer reflections are strong than one bin holds - a sixteenth of the pool - the top bin's median sits at an intensity where b cannot be measured at all, and the fit hands it the bins' own noise-selection slope instead: sorting noise by its group mean squared makes dev2 rise with I2 even when the true b is zero, and with no strong bin to out-vote it that slope becomes b. The result is not a small error. On the battery's weakest crystal, 2.2% of whose fulls reach I/sigma 2, the fit returns b = 5.6 - sigma -> 2*I at the strong end - and since corrected_sigma applies b at the GROUP MEAN, sigma^2 = a*sigma^2 + (b*mean)^2 is a per-group constant that caps merged |I/sigma| at sqrt(n)/b. The cap lands at 2.3, so 98.8% of merged reflections come out below 3 and the reported ISa is 0.50, on data whose CC1/2 is 99.3% at multiplicity 18.7. XDS fits 6.13 from the same images. Feeding XDS's own scaled observations through this estimator returns 0.84, so it is the estimator and not the data; synthetic data built with b = 0 and 1.8% strong reproduces a = 0.51 and ISa 0.50 to two digits, and recovers the truth as soon as the strong fraction passes one bin. So refuse to report what was not measured: when the strongest bin's own (I/sigma)^2 is below 4, fit a alone, hold b at zero and warn that ISa is unmeasured. The threshold is not delicate - the two crystals it fires on sit at 0.22 and 0.84 while the next crystal in the battery is at 31.7 and a healthy one at 342, so anything from 4 to 25 selects the same two. Full 38-crystal rotation battery: it fires on those two crystals and no others, and space groups are unchanged at 35/38. Dropping the spurious term also fixes the merge weights it had been distorting - on the worse of the two, R_meas 19.2 -> 13.6%, low-resolution R_meas 13.8 -> 6.9% against XDS's 14.1%, CC1/2 98.7 -> 100.0%, with chi2 1.11 on the one-parameter model. Two further crystals move slightly; the guard never fires on either, and they are marginal crystals of the kind whose two-pass branch any recompilation can shift. This reports the parameter as unmeasured rather than clamping it to something plausible, because the honest statement is that the data do not reach far enough for a systematic error to be seen - not that there is none. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b40abe31cf |
Scaling: one exact-Bragg angle per rocking event
Every partial's delta_phi was solved from its OWN frame's lattice - by the predictor, and again by SmoothGeometry. Per-frame geometry is re-refined against that frame's spots alone, so what is left of its jitter entered each frame of an event independently and the frames of one rocking event stopped sitting exactly one oscillation apart on the curve. Their partialities then no longer tile it, and because a broad rocking curve spans more frames, the error grows as 1/zeta - which is how it has been showing up: a zeta-graded systematic that nothing in the integrator could reach. For the frames of one event the geometry is exact. Each frame has already turned one oscillation further, so delta_phi is linear in frame number with slope minus the increment; the sign is checked against the data rather than derived, the measured mean frame-to-frame slope being -0.19996 deg/frame at an increment of 0.20000. Fit the one free number, the offset, over the event and lay its partials back on that line. The rms departure removed is 0.25 deg - larger than the oscillation itself, because a small orientation wobble is amplified by 1/zeta. The raw-hkl runs the merge already builds give the grouping, so this costs one pass over the partials and no extra sort. Full 38-crystal rotation battery against the same binary without it, on unchanged data (observations +0.20%, unique reflections +0.03%, so none of this is selection): R_meas_lo better 22 / worse 5, summed -41.0 pp; excess against XDS -46.9 -> -87.9 R_meas better 22 / worse 5, summed -24.7 CC1/2 better 17 / worse 2, summed +36.8 ISa better 15 / worse 22, summed +6.87; shortfall against XDS 28.1 -> 21.2 space groups unchanged at 35/38 The low-resolution R_meas gains land on the crystals that have carried this gap: 19.3 -> 10.8, 20.7 -> 13.8 (now past XDS), 25.6 -> 19.7, 17.5 -> 12.4 per cent. Exactly one crystal shows any change in the two-pass branch fingerprint, so unlike most changes on this path the result is not confounded by that bistability. ISa falls on more crystals than it rises, and that is the estimator becoming honest rather than the data getting worse: every crystal whose ISa dropped materially was over-optimistic against its own R_meas_lo and moved toward consistency, and the median ratio of reported ISa to the value its own R_meas_lo implies goes 1.09 -> 1.01, against 1.19 for XDS. The one real loss is a crystal going 1.11 -> 0.95 on that ratio. High-shell CC1/2 is worse on 22 crystals, by about 1.2 points each. It is the one metric that dissents, and it is also the one that has failed as an arbiter repeatedly on this data, while overall CC1/2, R_meas, R_meas_lo and reflection count all improve on an unchanged number of observations. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6f7b136ec2 |
Bragg integration: a shared signal pixel belongs to the nearer reflection
Nothing kept a neighbour's flux out of a reflection's own signal disk. The union mask keeps neighbour cores out of the BACKGROUND ring, but the r1 disk was read whole, so on a dense pattern a crowded reflection measures part of its neighbour as its own. Ownership is decided once per image into a per-pixel (quantised distance, reflection) key written with an atomic minimum, so the nearest predicted centre wins whatever order the writes arrive in and the lowest index breaks a tie. `--overlap exclude`, now the default, drops the pixels a nearer neighbour owns from the profile fit. A profile fit is the amplitude of a normalised profile, so leaving pixels out renormalises the estimator by construction and the reflection stays unbiased rather than being discarded; the summation-fallback guard is scaled back to the disk the box-sum seed actually read, so it still compares like with like. `--overlap reject` is the XDS MINPK alternative - drop the reflection when less than `--overlap-minpk` of its expected profile is cleanly its own. A box sum has no profile to renormalise with, so `exclude` is a no-op there and only `reject` acts on it. Widening the split - keeping a pixel only where no other centre is within its distance PLUS a margin - was built and measured, and it is worse monotonically: the residual bias of the pixels that were kept grows from +0.072 to +0.209 in ln intensity at 0 to 3 px of margin. What the margin removes is the reflection's own profile, not the neighbour's tail, so the plain nearest-centre split is the rule. Measured on the full 38-crystal rotation battery against the same binary with the treatment off: ISa better 15 / worse 8, summed shortfall against XDS 39.7 -> 28.1. Three of the losses are the two-pass loop taking its other branch - their median mosaicity moves between the two known attractors - rather than the change under test; excluding those it is better 15 / worse 5 and the shortfall goes 31.3 -> 14.4. The two crowded crystals gain 38% and 52% of their ISa, one of them passing XDS. High-shell CC1/2 over the 35 crystals that neither flipped branch nor carry a collapsed error model is better 7 / worse 7. Space groups unchanged at 35/38. The owner map is built only when a treatment is asked for and costs 1.1% of the battery's wall clock - 23% on a genuinely crowded crystal, nothing where no two predictions touch. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
06a5a118bf |
Bragg integration: widen the background ring to r3 = 13
The background is estimated from the r2..r3 ring and then subtracted from every pixel of the r1 disk, so the ring mean's own error enters the intensity n_inner times over: var(I) carries n_inner^2 * bkg / n_B. That term is first-order in sigma, and it is set by how many pixels the ring holds - not by anything about the reflection. At r3 = 10 the ring holds about 200 px against the disk's 50. Widening it to 13 roughly doubles that. The signal disk is untouched, and the pixels gained lie further from the reflection rather than nearer, so nothing is traded for them. The effect is not subtle once looked for. Matched observation by observation on one crystal, halving the ring's pixel count leaves the intensity alone and inflates sigma by 4.7%, and the inflation rank-orders with the ring collapse across the battery. Over the whole rotation battery, against the same binary at r3 = 10: ISa better on 14 crystals and worse on 4, the summed shortfall against the reference 164.7 -> 155.9, the summed low-resolution R_meas excess 69.7 -> 59.1 percentage points, and one more crystal reaching the reference space group (33/37 -> 34/37, a trigonal case that was over-promoting). Largest gains where the ring was starved worst; the four losses are 0.25 to 2.16 in ISa and none of them changes a space group. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
27d0b1db74 |
docs: state the reflection-file conventions
The mmCIF carries eleven items rugnux invents, and the rule they follow - a jfjoch_ prefix inside whichever standard category the quantity belongs to - was nowhere written down, so the only way to learn what was in a merged .cif was to read WriteReflections.cpp. Tabulate them, with the values a reader needs in order to interpret each one (the untwinned and perfect-twin values for the L test and the second moment, the sign convention for the radiation-damage B). Two of them need more than a name. The compatibility note records that jfjoch_diffrn_ISa changed meaning and that a file carries no marker saying which. And the HKLF-4 .hkl has two properties that are invisible in the file and change what a comparison means: Bijvoet mates are separate records, and the intensities carry a single global rescale so the largest fits F8.2 - harmless to SHELXC and ANODE, which use ratios, but not something to compare magnitudes across. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
adf87e8675 |
Merging: export the XDS-comparable ISa under jfjoch_diffrn_ISa
The mmCIF's _reflns.jfjoch_diffrn_ISa carried the strong-reflection asymptote, a tier XDS has no equivalent of, while the name invites comparison with XDS's ISa - which is the whole-range 1/sqrt(a*b). rugnux_vs_xds.py reads that item for the battery's ISa column, so the comparison that column exists to make was between two different quantities, flattering rugnux by the difference between the tiers. Write the whole-range value there, move the asymptote to _reflns.jfjoch_diffrn_ISa_asymptotic, and add _reflns.jfjoch_error_model_a and _b in XDS's convention so the number can be re-derived from the file rather than taken on trust. On a broadband rotation dataset the battery column now reads 13.25 against XDS's 21.18 where it read 15.6 before, and the two error models can be compared term by term for the first time: a 1.538 vs 1.249 and b 3.71e-03 vs 1.78e-03, so the gap is in BOTH the counting and the systematic term (1.23x and 2.08x, and sqrt(1.23*2.08) = 1.60 = 21.18/13.25). This is a deliberate redefinition of an exported item, not an addition: a file written by an earlier version carries the asymptote under the old name and there is no version marker to tell them apart. Noted in the changelog and in docs/CPU_DATA_ANALYSIS.md. Nothing reads the item back into the pipeline - it is written and never parsed by rugnux itself - so no stored file is reinterpreted in a way that changes a result. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ac06b5c64f |
Merging: report the error model in XDS's convention
rugnux fits sigma^2 = a*sigma0^2 + (b*<I>)^2, so its `b` is a fraction of the intensity. XDS fits
sigma^2 = a*(sigma0^2 + b*I^2) and prints ISa = 1/sqrt(a*b). The two `a` are the same number, but the
two `b` are not - b_xds = b^2/a - so the pair rugnux printed could not be read against a CORRECT.LP,
which is the only reason anyone looks at it.
Convert at the report. The fit, the merge weights and both engines' variance expressions are
untouched, so this is a re-expression and not a change: on a rotation dataset the merged intensities
move strictly less between before and after than they do between two runs of the SAME binary (99.9%
identical, max |dI/I| 9.1e-4 against the run-to-run control's 7.5e-3), with the same reflection set.
The rotation path also printed the wrong ISa for the comparison it invites. What it calls ISa is the
strong-reflection asymptote, a tier XDS has no equivalent of and which can only ever be the more
optimistic of the two; XDS's ISa is the whole-range 1/sqrt(a*b), which in rugnux units is exactly
1/b. Print both, labelled. On a broadband rotation dataset that is 13.2 (whole range) and 15.6
(asymptote) against XDS's 21.18 - so the number previously compared was flattering rugnux by 2.4.
A third, unrelated `b` lives in the space-group search: fitted with the sigma^2 coefficient held at 1,
with gate constants calibrated in that convention, and a ratio bound does not survive the mapping
(1.90 would have to become 3.61) while the absolute floor has no correct value at all, there being no
`a`. It is now commented as such, since making the three consistent is the obvious wrong move.
Also corrects three comments and two doc passages that still described a merged-sigma systematic
floor deleted in
|
||
|
|
61d24db59f |
Bragg integration: elongate the background ring per reflection
The signal disk and the r2..r3 background ring were fixed pixel circles, identical for every reflection at every resolution. A reflection is not round: a finite bandwidth streaks it radially by bw_sigma*Rpx, so at high resolution the ring sits within 1.3-2.2 sigma of the reflection's own profile and measures its tails as background. --integration-stencil <k> makes the RING an ellipse, elongated along the beam->reflection direction by k times that streak, capped at 2*r3. The tangential half-widths stay r2 and r3, and the r1 signal disk stays a circle: r1 drives the all-or-nothing n_inner_valid == n_inner gate, so growing it rejects any reflection carrying one bad pixel along a long streak, and the flux a circular r1 loses is a function of resolution alone, which the per-shell scale absorbs. The geometry lives in one shared header compiled by both the host compiler and nvcc, so the seven pixel-classification sites - the CPU mask/main/clip loops and the GPU mark_mask/main/trim/clip kernels - cannot drift apart. Rather than evaluate an ellipse, each pixel's squared distance has its radial part scaled down, d2 - q*rad^2 against r2^2/r3^2 with q = 1 - (r/(r+grow))^2, so grow = 0 gives q = 0 and both tests collapse onto d2 exactly in floating point. The width is the bandwidth streak alone, not the profile's full radial variance, which also carries the sensor parallax and weak-spot capture terms. Deriving the growth from those was implemented first and measured on the rotation battery: at k=1 it took Thau_9's high-shell CC1/2 from 75.8 to 27.9 and Benas_3's from 14.1 to 6.0, against cytC_10 +1.2 and lyso_ref flat. On a monochromatic beam they are the only terms there are, and C_CAPTURE is 64% of them. Keeping only the streak also makes the option exactly inert without a bandwidth, rather than merely small. Default 0. Measured on broadband rotation data with the bandwidth set to its spectroscopic value, matched resolution limits: high-shell CC1/2 30.6 -> 46.4 at k=4, and better in EVERY shell in both CC1/2 and R_meas (top shell R_meas 194.7% -> 138.7%), with completeness, multiplicity and space group unchanged and 28 of 98833 unique reflections lost. Anomalous peak height over 18 sites +0.107 +- 0.039 sigma (p = 0.013). The full 38-crystal rotation battery is unchanged to every reported digit, base against k=3. Two consequences of an elongated ring are handled rather than inherited. The neighbour exclusion marks the inner ELLIPSE in each neighbour's own frame, or an elongated neighbour leaks its tails into this reflection's ring. And the radial-background curvature kernel becomes a small table indexed by the growth, because its azimuthal average makes one kernel serve every reflection only while their stencils are identical; the GPU's radial window, previously a fixed 32 bins, is now sized on the host from the widest ring on the detector. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
52ea727650 |
Reader: reconstruct the background variance a legacy file does not store
A _process.h5 written before background_variance existed was read with var_bkg = 0, on the reasoning that zero leaves the combine with the signal term alone, "which is what it had before". It does not. Before, the combine back-derived the non-signal variance from sigma itself, and on a weak reflection that is essentially the whole of sigma^2; zero deletes the dominant term and weights the reflection by roughly 1/I instead of 1/sigma^2. Recover it from the integrator's own identity, sigma^2 = I + var_bkg, when the dataset is absent. Measured by re-scaling a stills _process.h5 with the dataset deleted, against the same file with it intact: the automatic resolution cutoff was reading 1.66 A where the intact file reads 1.81, with 14731 unique reflections against 11508 - i.e. the zeroed file looked good enough to merge 0.15 A past its own limit. Reconstructed, it reads 1.78 A and 11926, within 0.03 A of the intact file. The residue is the profile-fit path, where var_bkg is not exactly sigma^2 - I and only the box-sum identity is exact; that is recoverable to a closer approximation only by storing it, which is what files written from now on do. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a7c5d7a89b |
CBOR: carry the reflection's background variance
var_bkg was added to Reflection and to the HDF5 writer but never to the CBOR reflection map, so it survived only where rugnux drives the writer in-process. Everything that reaches the writer over the wire - i.e. every acquisition the broker records - wrote /entry/reflections/*/background_variance as an array of zeros, presented as a measured quantity, and re-processing such a file fed the merge a non-signal variance of zero. Encode and decode it. The key is optional on both sides, so a stream from an older version still reads and one from this version still reads on an older client. The round-trip test only checked h, k, l, the predicted position and d - which is why a missing float was invisible. It now gives every field a distinct value and checks all of them, so the next field added to Reflection and forgotten here fails immediately. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bc1c4c6800 |
Rotation: land the rest of the bandwidth term
|
||
|
|
1239c49731 |
Bragg integration: separate the three things a bandwidth used to switch
Setting a bandwidth flipped three unrelated switches at once: it changed the profile's radial capture term, it moved the width measurement from the signal disk to the whole fit grid, and it silently overrode the background clip and trim, so --background-clip under --bandwidth was ignored - the two runs were bit-identical. The width measurement was the damaging one. The fit grid is an azimuthally averaged stack, so its second moment is sigma_r^2 + sigma_t^2 and the radial smear of a bandwidth leaked into the tangential model - a tangential width of 3.04 px against a 1.06 px truth, inflating the effective background pixel count where the weak signal is. The result was a step rather than a slope: on genuinely monochromatic data, declaring a 0.2% bandwidth cost ISa 28.4 -> 22.2. Measure the two widths separately, accumulated in each spot's own radial/tangential frame over the signal disk, from the signed profile cells - away from the peak a learned cell is background noise centred on zero, so the signed sum is unbiased, while clamping it at zero turns that noise into a pedestal the r^2 weight reads as width. The radial term is then the measured excess or the analytic floor, whichever is larger. With the two widths separated there is nothing left for the broadband switch to select, so it is gone - which is the proof the three were independent. The background clip and trim now come from the settings in every case; the tuned 3-sigma broadband default moves to the rugnux front end, which is the only place that knows whether the user gave a value. Monochromatic data: declaring a 0.2% bandwidth now costs ISa 28.4 -> 27.9 rather than 22.2, and forcing the old 3-sigma clip in the new build reproduces the good result, so none of the step came from the clip. On large-bandwidth data CC1/2 improves in 8 of 10 shells. Across 12 monochromatic crystals the space groups are unchanged and CC1/2 moves by at most 0.2 points. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a29c36600f |
Beam-stop shadow detection, and a low-resolution limit for scaling
rugnux finds the beam stop and its holder in a projection of 60 images and marks them in the pixel mask as bit 9 (--detect-beam-stop[=N|off], on by default). Reflections behind the stop are attenuated but not flagged, so they integrate low with a plausible sigma and nothing downstream catches them: the signal-box gate requires 100% valid pixels and shadow pixels are valid, the background clip is high-side only, and the |zeta| cut applies only to the space-group search merge. The detection compares each pixel's background against the typical background at the same radius on two channels. An azimuthal one (the ring median) finds the holder arm, which is a minority of its ring; a radial one (the background just outside) finds the disk, which the ring median cannot see because inside a fully blocked ring the median is the shadow itself. Pixels are pooled over a 5x5 box and tested only where the background has actually been counted, so low-background data no longer masks the whole detector. Recorded reflections are carved back out - a beam stop cannot block a reflection that was measured. Bit 9 belongs to the run that found it, not to the dataset: it is cleared when a run starts, so a mask read back from a file that carries one starts clear. The user mask (bit 8) is left alone. Scaling and merging gain a low-resolution limit, default 50 A (--scaling-low-resolution <num>, 0 removes it), applied per observation before scaling so it also protects the per-frame scale fit and the space-group search. 50 A is the value XDS configurations use; rugnux_vs_xds.py now matches both of XDS's resolution limits instead of only the high one, so the lowest shell is the same shell in the two programs. The viewer draws the detected shadow in coral with a "Show beam stop" switch in the side panel, exposes the low-resolution limit in the settings dock, and offers detection in its processing jobs. Adding an image marker meant giving the reader a MIN_REAL_PXL_VALUE, because several places classify a pixel by range rather than by equality and would otherwise read the new marker as a very negative intensity. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
df9a9c2a2c |
Fix the defects found reviewing the branch before merge
Build Packages / build:viewer-tgz:cpu (push) Successful in 19m32s
Build Packages / build:windows:nocuda (push) Successful in 19m57s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m45s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m38s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m24s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m8s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m9s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m18s
Build Packages / XDS test (durin plugin) (push) Successful in 11m10s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 20m21s
Build Packages / build:windows:cuda (push) Successful in 22m5s
Build Packages / build:rpm (rocky9) (push) Successful in 20m57s
Build Packages / Generate python client (push) Successful in 34s
Build Packages / Build documentation (push) Successful in 1m29s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m41s
Build Packages / DIALS test (push) Successful in 21m19s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m34s
Build Packages / build:rpm (rocky8) (push) Successful in 27m4s
Build Packages / XDS test (neggia plugin) (push) Successful in 10m19s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m58s
Build Packages / Unit tests (push) Successful in 1h17m36s
Image buffer: the per-image CBOR metadata headroom had been re-derived from the online reflection cap alone, which cut it from 4 MiB to 2.55 MB while the measured worst case - reflections plus the capped spot list plus the three azimuthal arrays - is 2.9 MB, so the receiver dropped the frames with the most to say. Restore it and give it a name that both the code and its guard test read: written down twice, the two had drifted and the test kept passing against the value the code had left. Spot finding: an unset low_resolution_limit means no limit at that end, as an unset high_resolution_limit already did. An optional rather than a zero sentinel, because zero is not a natural "no limit" here - every pixel lies above it, so the plain comparison masked the whole image instead of none of it, and nothing validated the zero. The API field is no longer required; a zero is folded into the unset case at the boundary, where older clients still send it, so one spelling reaches the analysis code. The FPGA takes its fixed-point ceiling instead, since ap_ufixed<16,9> wraps above 512 A and would have masked everything. image_preprocessing: check the CUDA calls on the fused decode path - the one new GPU file with none, and the path fed by bytes we did not produce. An unchecked synchronise returned the host-written sentinel as if it were a measurement, so the decode looked successful and the fallback to the host decoder never fired. rugnux: --stride no longer writes one past the end of the per-image arrays, whose count floored where the worker loop ceils, and the written process file links the images actually processed rather than the first N - each frame's picture now sits next to its own analysis. Powder calibration: the face-centred calibrants no longer list their systematically absent rings, so the distance fit starts from a reflection that exists rather than an extinct one; the triclinic calibrant covers both signs of h and k instead of a single octant, which is only valid for a diagonal metric. The test asserted the old behaviour - one ring formula for every cubic standard - and is rewritten. CBOR: skip an unknown tagged value in the end block, as the other four blocks already do. One advance lands on the tagged item rather than past it, so an older reader fed a newer end message threw and never finalized its file. Viewer: a settings value the setter rejects no longer escapes as an uncaught throw from a worker slot, and the field offers only what the setter accepts. Space-group search: judge stage B on the same "present" cut stage A already computes. Merged sigma is floored so no reflection reads above ISa, so on a low-ISa merge the fixed cut left both stage B tests unsatisfiable - every screw axis passed unchallenged and the centering rescue switched itself off on exactly the weak data it exists for. Where the fixed cut is the smaller of the two they are equal and this is inert: over the 37-crystal rotation battery every crystal reports the identical space group and identical merge statistics, so it is a no-op there and the low-ISa case it targets remains unmeasured. rugnux: --polarization reaches --mode azint, which parsed the flag and then dropped it; that mode also applies the same polarization default as every other mode. Acknowledge the ACTS/traccc project, whose sparse connected-component labelling both spot extractors take their algorithm from, with its citation and its license. The rc.161 change list is brought back to one line per entry, and the user-visible changes that were missing from it added. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5830f78d57 |
Revert the azimuthal-integration sigma clip
Removes azim_int_settings.sigma_clip / rugnux --azim-sigma-clip and the clipping
machinery in AzIntEngine. This is a partial revert of
|
||
|
|
a8d289e7cf |
Powder calibration: write Poni1/Poni2 in pyFAI's frame, not ours
The same frame mismatch as the rot2/rot3 fix, in the other two fields. Our pixel coordinates are pixel-centred - 948.0 is the CENTRE of pixel 948 - while pyFAI measures from the edge of the sensor and puts the centre of pixel i at (i + 0.5) * pixel size. Poni1/Poni2 went out as beam * pixel size, so anything reading the file placed the pattern half a pixel (37.5 um at 75 um pixels) off ours. The previous commit's "Poni1/Poni2 need no such change" was right about the axis directions and wrong about the origin. The proof was already in the tree. The pyFAI reference values in DiffractionGeometryTest were computed for a .poni with Poni2: 0.150 and a 75 um pixel, which the tests translate to beam_x = 2000 - but pyFAI's numbers are reproduced only at 1999.5. At 2000 every one of them is out by 2.6e-3 nm^-1, which the 1e-2 tolerance hid. The tests now use the beam centre those headers actually mean, and agree with pyFAI to 1e-6 - float precision - across untilted q, azimuth, rot1, rot1+rot2, rot3, rot1+rot2+rot3 and the solid-angle correction. Tolerances drop to 1e-4 (1e-5 for solid angle): ~100x the observed float noise, and 26x tighter than the half pixel they were blind to. The viewer's calibration window printed "PONI x = ... mm" from the un-offset value beside the path of the file it disagreed with; it now matches the file. Also moves the viewer's beam-centre cross half a pixel down and right, where the spot, prediction, top-pixel and saturation markers already are. Our coordinates are pixel-centred and the Qt scene's are pixel-cornered, so the map between them is +0.5, and DrawBeamCenter was the one overlay missing it. The convention itself is now written down in docs/DETECTOR_GEOMETRY.md, with the conversions to XDS ORGX/ORGY and to the edge-of-sensor programs, this being the second bug to come out of it. Only exported and displayed values change; the fitted geometry, spot positions and integration were always self-consistent. A .poni written by an earlier build is half a pixel off. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6194fe6fbf |
viewer: calibrate the whole dataset from "Analyze dataset"
Build Packages / build:viewer-tgz:cpu (push) Successful in 11m45s
Build Packages / build:viewer-tgz:cuda (push) Successful in 17m11s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 18m59s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 21m9s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 24m51s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 25m1s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 21m32s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 17m49s
Build Packages / build:rpm (rocky8) (push) Successful in 23m8s
Build Packages / build:rpm (rocky9) (push) Successful in 20m8s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m4s
Build Packages / XDS test (durin plugin) (push) Successful in 10m53s
Build Packages / Generate python client (push) Successful in 28s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m30s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 23m22s
Build Packages / DIALS test (push) Successful in 18m9s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m55s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m56s
Build Packages / Unit tests (push) Successful in 1h54m52s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
The powder panel could only calibrate the image on screen. Calibration is now a third page beside MX and AzInt, so the dataset button runs it over every image the same way it runs the other two - which is the point, since a powder ring is measured far better by summing a run than by one frame. The page carries the calibrant and the method (rings or spots); the interactive Guess/Refine buttons stay where they were and now share the one calibrant selection, so there is no second combo to drift. analyzeDataset() carries the ProcessMode rather than a bool: a third state was coming, and two bools would have had one combination that cannot be valid. The calibrant list gains ICE, which it could not offer before: the widget worked in unit cells, and hexagonal ice has none that generates its rings correctly (P6_3/mmc would include systematically absent ones). FindCenter now takes the ring list its first line used to derive, so the interactive path gets ice as well. The result window leads with the residual rms rather than the fitted sigma. The sigma is a formal scatter estimate and understates a bad fit badly - measured on ice, 0.215 px reported against a 1.70 px residual - while the rms separates a usable fit from one that has locked onto the wrong thing. A rings run needs the profile binned in azimuth; below four sectors it returns nothing at all. The viewer raises the count to 32 exactly as the CLI does, and says so in the panel and in the job dialog rather than doing it silently. Also fixes a CLI inconsistency this comparison exposed: rugnux's calibration branch never applied the standard offline analysis defaults, so it measured the rings in a profile built with the file's polarization factor while every other mode - and the viewer - uses 0.99. Found because the two disagreed by 0.005 px in PONI x, and confirmed by reproducing the viewer exactly with --polarization 0.99. With it applied the CLI and the viewer write byte-identical .poni files on LaB6 by rings, LaB6 by spots, and an iced dataset over 1800 images. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6468dd13be |
rugnux: --mode, and detector calibration from powder rings
Build Packages / Unit tests (push) Failing after 6m24s
Build Packages / build:rpm (rocky9_nocuda) (push) Failing after 14m5s
Build Packages / build:viewer-tgz:cpu (push) Failing after 14m34s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Failing after 14m54s
Build Packages / build:viewer-tgz:cuda (push) Failing after 16m14s
Build Packages / build:rpm (rocky8_nocuda) (push) Failing after 16m24s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Failing after 18m53s
Build Packages / build:rpm (rocky9_sls9) (push) Failing after 13m2s
Build Packages / build:rpm (rocky8_sls9) (push) Failing after 19m34s
Build Packages / build:rpm (rocky9) (push) Failing after 14m54s
Build Packages / Generate python client (push) Successful in 42s
Build Packages / build:rpm (ubuntu2404) (push) Failing after 14m10s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (durin plugin) (push) Successful in 12m14s
Build Packages / XDS test (neggia plugin) (push) Successful in 11m52s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 12m15s
Build Packages / Build documentation (push) Successful in 2m10s
Build Packages / build:rpm (rocky8) (push) Failing after 18m29s
Build Packages / build:rpm (ubuntu2204) (push) Failing after 17m55s
Build Packages / DIALS test (push) Successful in 17m4s
Build Packages / build:windows:nocuda (push) Canceled after 0s
Build Packages / build:windows:cuda (push) Canceled after 0s
--azint-only and --scale are replaced by --mode mx|azint|scale|calibration, with mx the default. The old flags are removed rather than aliased. Calibration mode fits the detector geometry - PONI x/y, the two tilts and the distance - to a calibrant's powder rings and writes a pyFAI .poni alongside a report of how far each parameter moved from the header. Bragg data constrain the beam centre worst, because it is gauge-coupled to the crystal orientation; a powder ring has no orientation to couple to. --calibrant takes lab6, agbh, ceo2, si or ice. A calibrant is a list of ring positions rather than a unit cell, because hexagonal ice is P6_3/mmc: rings enumerated from its cell would include systematically absent ones. So the crystalline standards generate their rings from a cell and ice carries the measured list, and RingsFromAzimuthalProfile, GuessGeometry and OptimizeGeometry all take ring q. The calibrant table is shared with the viewer's powder panel, which previously carried its own copy. --calibration picks how the rings are measured: rings (default) sums the (q x azimuth) profile over every processed image and fits the arcs in it; spots pools the found spots and fits those. Both use the whole run, with -s/-e/-t selecting images. rings defaults --azim-phi-bins to 32, since a profile with one azimuthal bin has averaged the ring over every direction and cannot locate it. Two fixes this exposed: The extraction window is capped at half the gap to the neighbouring ring. The background under a peak is taken from the ends of its window, so a window wider than half that gap measures the next ring's flank as this ring's background - and hexagonal ice has three rings within 0.06 1/A. Ice calibration was 3.5 px out before this and 0.29 px after; LaB6 is unaffected. RingOptimizer holds rot1/rot2 fixed when only one ring is present. A tilt and a centre offset both move a ring as cos(phi) and are separated only by the tilt's amplitude growing as the ring radius squared, so on a single ring they are exactly degenerate. Measured. LaB6 at five distances: the fitted direct beam is within 0.36 px of an independent implementation out to 300 mm, and D = -0.046 + 1.000788 dtz with an rms of 0.011 mm. At 500 mm one ring is fully on the detector and a second only clips the corners, which is not enough to constrain a tilt - restricting the q range to the resolved ring recovers 0.06 px. Ice: 5.53 -> 0.29 px on one crystal and 4.71 -> 0.80 px on another, against XDS's refined direct beam. On an ice-free crystal the fit is worse than the header, which is the correct outcome. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d81d2e4696 |
docs: state the metric-symmetry rule rather than how it was arrived at
CPU_DATA_ANALYSIS describes how the pipeline works; the account of which bar was tried first belongs in the commit that changed it. Same facts, no narrative. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cfd3697ddb |
docs: cover the azimuthal sigma clip, the indexing flag, and the metric-symmetry check
Build Packages / build:windows:nocuda (push) Successful in 17m17s
Build Packages / build:windows:cuda (push) Successful in 21m20s
Build Packages / build:viewer-tgz:cpu (push) Successful in 11m54s
Build Packages / build:viewer-tgz:cuda (push) Successful in 12m55s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 16m59s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 20m19s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 19m55s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 16m47s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 19m57s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 16m31s
Build Packages / build:rpm (rocky9) (push) Successful in 17m38s
Build Packages / build:rpm (rocky8) (push) Successful in 19m29s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 19m16s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 17m38s
Build Packages / Generate python client (push) Successful in 11s
Build Packages / DIALS test (push) Successful in 15m9s
Build Packages / Build documentation (push) Successful in 42s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (durin plugin) (push) Successful in 9m32s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m24s
Build Packages / XDS test (neggia plugin) (push) Successful in 6m20s
Build Packages / Unit tests (push) Successful in 1h19m27s
Three things had reached the code without reaching the documentation. The azimuthal sigma clip had a RUGNUX.md row and a CPU_DATA_ANALYSIS section but no changelog entry - and the only "sigma clip" the changelog mentioned was the background ring's, which is a different thing at a different stage. --index-ice-rings was in the options table but nowhere in the changelog, so the entry describing the ice gate still implied that whether indexing uses the ice-band spots is decided per run, which it no longer is. CPU_DATA_ANALYSIS section 6 still described the Bravais class as simply "the highest-symmetry class that matches within tolerances", which is the behaviour that lost a crystal outright. It now records that the class is chosen from the UNREFINED candidate against a fixed angular tolerance, that a pseudo-symmetric lattice therefore gets promoted a class too far, and that the first pass settles it on validation-frame counts with a clear-majority bar - including why the bar is a majority rather than a margin, since that distinction is the whole reason the check is safe. It also records that the first pass finds its own spots rather than reading the acquisition's, which was not written down anywhere. Also a build note: M_PI is not standard C++ and MSVC does not define it, so the Bragg integrator's use of it broke the Windows viewer build. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1b5e2e85fd |
Regenerate the API documentation from the spec
Build Packages / build:windows:nocuda (push) Successful in 14m15s
Build Packages / build:windows:cuda (push) Successful in 20m17s
Build Packages / build:viewer-tgz:cpu (push) Successful in 16m9s
Build Packages / build:viewer-tgz:cuda (push) Successful in 17m41s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 18m44s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 20m49s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 24m20s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 22m56s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 19m52s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 22m35s
Build Packages / build:rpm (rocky9) (push) Successful in 19m44s
Build Packages / build:rpm (rocky8) (push) Successful in 25m16s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m48s
Build Packages / Generate python client (push) Successful in 43s
Build Packages / Build documentation (push) Successful in 1m5s
Build Packages / Create release (push) Skipped
Build Packages / XDS test (durin plugin) (push) Successful in 9m52s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 24m7s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 8m36s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m3s
Build Packages / DIALS test (push) Successful in 15m45s
Build Packages / Unit tests (push) Successful in 1h54m32s
update_version.sh at 1.0.0-rc.161. The only substantive change is the one that had drifted: the spot-finding ice-ring half-width was still documented as 0.02 in the generated Python client and its docs while broker/jfjoch_api.yaml has said 0.03 since the band was widened to the measured ring FWHM. Anyone reading the client docs - or relying on the client's default when omitting the field - got a band two-thirds the width the pipeline actually uses. The TypeScript frontend client regenerates identically (the spec itself did not move), and python-client/ is not tracked here. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1df9556ec1 |
Bragg integration: default the radial background correction off again
Auto rode in with the ice work rather than on its own evidence, and measured over the 37-crystal rotation battery it does not carry itself yet. It TARGETS correctly - it fires on ten crystals and every one is ice-positive, no failures, no space-group changes - but it costs 1.35x the wall clock (median +3 s per crystal, worst +29 s) and on the merge statistics it is the familiar sign-mixed trade: high-shell CC1/2 worse on three of the four crystals that move materially, mean -0.76. The case for it is real but rests on agreement with a fixed external model - 43 % of the ice bands' excess amplitude removed on smooth ice, the effect 7x stronger inside the bands than outside - which is the better arbiter and also the narrower one. That deserves settling on its own, not riding along with a set of ice defaults. `--background-radial=auto` keeps it a flag away. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a6be35ccdb |
Azimuthal integration: optional sigma clipping of the reported profile
The profile is the MEAN of each bin, so a few strong reflections landing in a bin lift it exactly as a smooth powder ring does. That is the wrong quantity whenever the profile is wanted as a background rather than as a measurement of what is in the bin - the ice score being the case in point, where reading a plain profile INVERTED the metric: over 37 rotation crystals the two highest-scoring crystals had no ice at all. The adaptive spot finder already computes the right thing, a sigma-clipped per-resolution-ring background, as a byproduct of its own threshold. Where it runs, the ice score uses that. Where it does not - --no-adaptive-spots, --azint-only, and anything reading the profile the broker wrote - there was no way to get it. This adds one: azim_int_settings.sigma_clip (rugnux --azim-sigma-clip), 0 = off, minimum 2 because a tighter clip rejects a large part of a clean Gaussian bin and biases the estimate low rather than removing outliers. Two clip passes follow the plain one, matching the finder's recipe - the first pass's standard deviation is itself inflated by the peaks being removed, so one pass leaves a threshold that is still too generous. A bin with fewer than eight pixels is left alone: at the detector edge and behind the beam stop there is no spread to clip on. Both engines do it. On the GPU the accept range is computed by a small kernel and stays resident, so a clip pass is one more read of the same pixels and no round trip; the two accumulation kernels take the range as a pointer that is null on the plain pass. Measured on a JUNGFRAU rotation dataset, non-adaptive path: azimuthal integration 0.02 -> 0.06 ms per image, exactly the 3x the extra passes predict, against a 0.34 ms per-image total. Note what the result IS: the smooth background under the peaks, not the bin mean. It should not be switched on where a ring's integrated intensity is wanted - the powder-ring geometry fit reads ring peaks, and those are what a clip is designed to remove. Off by default, so nothing changes unless it is asked for. Not exposed over the REST API - that needs the generated model regenerated, which is a separate step. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3d4209d803 |
docs: the first-pass ice measurement, the promotion fix, and powder-ring geometry
Changelog entries for the three changes above, and a new CPU_DATA_ANALYSIS section on determining detector geometry from powder rings: why a ring is an independent constraint on the beam centre (it has no crystal orientation to be gauge-coupled to, unlike everything else that fits geometry here), what a ring can and cannot determine, and how the ring points are obtained. The section states the harmonics correctly, which is worth writing down because the intuitive version is wrong: a detector tilt shows up as cos(phi), the same harmonic as a beam-centre error, and the two are separated by the amplitude growing as the ring radius SQUARED - so it takes at least two rings, and on one ring they are exactly degenerate. The genuine cos(2 phi) term is hundredths of a pixel. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b5f5879a1d |
rugnux: measure the ice in the first pass, and always find its own spots
Ice handling was gated on a measurement the run only made AFTER the images had been processed, so the per-image pass could not use it. The flagging therefore ran unconditionally: ice-band spots were ordered last in the --max-spots budget and held out of the indexer seed and the geometry refinement on every crystal, iced or not. The eleven bands are fixed geometry holding 16-26 % of the unique reflections whether or not there is ice, so on a clean crystal that discards a fifth of the spots - the strongest first - for nothing. Measured on a crystal whose gate never fires, that moved the merged data by a mean of 0.85 sigma against a run-to-run floor of 9.3e-5. Measure it in the first pass instead. That pass already looks at ~100 images spread over the sweep, and it already stops at the spot finder, so it sees the azimuthal profile for the smooth channel and the unfiltered connected components for the spot channel. Both counts SpotAnalyze takes are pre-filter, so pooling them there is the run's own verdict, reached before anything has been discarded and in time for the pass that acts on it. Where the sample sees no ice, the run indexes on the ice-band spots too. It has to be the whole sample: the spot channel is a ratio pooled over images, because one frame carries a handful of control spots. A per-image gate is not an alternative - two of the crystals whose indexing this rescues fire on that channel alone, at profile scores of 1.12 and 1.22, so gating per image on the profile score would drop exactly the cases that matter. This also removes the first-pass spot reuse, and with it --redo-rotation-spots and the reuse path. Finding the ~100 first-pass spots costs little, and reusing was actively wrong here: the stored spots were found online at the acquisition's threshold and have already had their ice-band entries ordered last and dropped by its spot budget, so counting ice from them under-reads it by construction, and the lattice search never saw the spot-finding settings at all. It also removes the need for the machinery that re-found spots whenever a spot-finding option was named, which made those options impossible to A/B. IndexAndRefine cached index_ice_rings at construction, which happens before the first pass; it holds a reference to the experiment, so it now reads the setting where it uses it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f0cdb027e1 |
Ice: default the merge mask off, gate the radial background on smooth ice, and pick detection by geometry
Three defaults, each settled by measurement rather than by argument. The arbiter throughout is structure-referenced - anomalous peak height where a crystal can carry it, and otherwise the agreement of the ice bands with a fixed external model against resolution-matched DECOY bands carrying no ice. The band-versus-decoy contrast is used because R-free here tracks completeness, and every one of these switches moves completeness. The damage is real and it localizes: over the rotation battery the ice bands' excess amplitude reaches +9.6% on a smooth-ice crystal and +35% on the worst, while a clean control sits at +0.6% (z +0.45). On the worst crystal, nine of the ten largest excess peaks in a q scan land on hexagonal ring positions. Turning ice handling off leaves the contrast unchanged and forcing it on a clean crystal does not create one, so it is the ice and not the machinery. MERGE-TIME RING MASK -> OFF. It deletes reflections, which no other program does by default - AIMLESS, DIALS, xia2, XDS and CrystFEL all keep ice-band reflections in the merge and exclude them only from the model fit; autoPROC is the sole exception. On the one battery crystal where the mask fires and an anomalous arbiter can score it, dropping the band moved the mean peak height at the known sites by -0.001 +- 0.018 sigma, 2% of the site height, while removing 1149 unique reflections whose mean I/sigma was 3.62 against the dataset's own 3.05 - better than average data - and costing 17 completeness points in that shell. It fires on 5 of 37 crystals, changes no space group, and those 5 disagree in sign: it clearly helps the two most heavily iced, is a wash on two and costs a third. So it stays as a switch, worth setting by hand on a badly iced crystal where it shows in the high shell, but it is not a default. RADIAL BACKGROUND -> AUTO, gated per image. The correction models the background as a function of radius alone, and that is exactly when it works. On a crystal with pure smooth powder ice it removes 43% of the bands' excess amplitude, with the improvement 7x larger inside the bands than outside; on a crystal whose ice is discrete crystallite spots - no smooth ring to model - the excess amplitude GREW by half; on clean data it is inert to four decimals. The two ice channels already separate those morphologies, so --background-radial takes on|off|auto and auto applies it to an image when that image's peak-excluded score reaches --ice-min-score. Auto never engages without such a score, because the plain profile carries the Bragg peaks and cannot support an absolute threshold. Per image rather than per run, and that was tested rather than assumed: the gate fires on 100% and 94% of frames on the two crystals that want it, and on 1.5% of frames - 32 blocks, 23 of them single frames - on the textured-ice crystal. A seam statistic against off + f*(on - off) is null on both mixed runs, every merge statistic is bracketed by the pure arms, and the textured crystal's auto arm lands on `off` rather than on `on`'s harm. A run-level gate would need the score before the pass that integrates, i.e. rotation-only plumbing, and buys nothing measurable. The kernel was already built unconditionally, so flipping the flag per image is free - except on the GPU, where the launches were gated on a construction-time n_rad. That is why the buffers are now allocated whenever the correction could run, and Run() decides per image. DETECTION -> the geometry's default when the file is silent: on for rotation, off for stills, with the command line and then the file taking precedence. A rotation sweep sits on the same rings for the whole run, so ice there is a coherent systematic and the presence gate keeps it inert on a clean crystal; a serial stills run has too few spots per image to spend any on flagging. The master file's key is kept as written rather than collapsed to a bool, so "the file said nothing" is distinguishable from "the file said no" - it used to fall silently to off, taking the exclusion from the scale fit with it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
61a7c91b90 |
Ice: detect it on two channels, and only handle it when it is there
The per-image ice score was read off the PLAIN azimuthal profile. That profile is a per-ring mean, so a few strong Bragg reflections landing in a ring's q bin lift it exactly as ice would. Measured over 37 rotation crystals, that did not merely add noise - it INVERTED the metric: the two highest-scoring crystals had no ice at all (4.23 and 4.06), while a clean control read 1.57. A decoy null - the identical statistic evaluated at q positions where hexagonal ice cannot be - reaches 1.51 at its 99th percentile and 2.70 at its maximum, so that metric cannot support any absolute threshold whatsoever. The adaptive spot finder already computes the right input for its own threshold: a sigma-clipped per-resolution-ring background, in the same bins. A powder ring is azimuthally smooth and survives the clip; Bragg peaks do not. On the clipped profile the clean population tightens to 1.00-1.22 and the crystals with confirmed ice sit at 2.08-2.37, against a decoy null that never exceeds 1.29. That channel is blind to one thing: ice in large crystallites diffracts as DISCRETE spots and leaves the radial profile flat. So a second channel counts found spots on the rings against the same q width of ice-free flanks beside them. The two barely overlap - the smooth-ice crystals read 2.1-2.4 / ~1.0 and the textured ones ~1.1 / 3.8-17.6, while a clean crystal reads 1.04 on both. Both are then used as a GATE (--ice-min-score 1.5, --ice-min-spot-ratio 2.0, both calibrated on the battery, 0 disables): the eleven fixed hexagonal bands cover 16-26 % of the unique reflections at typical resolutions whether or not the crystal has ice, so flagging, the exclusion from the scale fit and the merge-time CC1/2 ring mask are now all skipped when neither channel sees any. The gate is applied in the full pipeline and in --scale, which reads the stored per-image values back out of the _process.h5. Also fixes the merge-time mask's control: the shoulder now excludes reflections that are themselves on an ice ring. The rings are not evenly spaced - 1.947/1.916/1.882 A sit 0.05-0.06 apart in q - so for those three the [w,3w) shoulder landed squarely on the neighbours and the test compared ice against ice. Measured, that is the only thing this changes: it removes firings on those three rings and leaves every other firing's CC pair identical to three decimals. And the online ice half-width, which was 0.02 in the API against 0.03 offline, so the same data got a narrower band online than the measured ~0.06 ring FWHM justifies. Battery (37 rotation crystals, against the previous behaviour): space groups 34/37 in both and NO crystal's space group changes; 6 crystals gain unique reflections, 1 loses. Best of them gains 7082 unique reflections with R_meas 16.0 -> 14.3, CC1/2 95.9 -> 97.3 and ISa 13.7 -> 19.0; another goes R_meas 54.9 -> 42.9, CC1/2 84.0 -> 90.4, ISa 3.9 -> 5.5; a third reaches CC1/2 99.4 from 95.7 at an unchanged reflection count. The one crystal that loses reflections improves on both R_meas and CC1/2. Not done here: the ScanResult/API/plot-type/frontend/viewer layers for the new spot_count_ice_control (they need the OpenAPI regeneration). Message, CBOR, HDF5 write/read and the receiver plots are. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
227f1bf1b4 |
Add rugnux_anomalous.py: judge partiality changes by anomalous peak height
A change that touches partiality - a mosaicity estimator, a rocking-curve model, a background change - cannot be judged by the statistics we normally reach for, and this was learned the expensive way. ISa is anti-correlated with external accuracy and is the largest mover of any statistic; last-shell R_meas moves with its denominator, i.e. the wrong way by construction; `--model` R-free tracks its own zero-information floor, which shifts ~22x more than R-free itself over the same sweep; and per-shell agreement with XDS is biased, because XDS never divides by partiality, so "divide less" moves us toward it mechanically - measured to put the optimum ~1.4x too low. Anomalous difference density at known scatterer sites has none of those problems. It is read in units of the map's own sigma, so the uniform intensity rescale a partiality change produces cancels exactly, and it is referenced to the structure rather than to another program's partiality model. The script runs SHELXC + ANODE per arm against a model placed ONCE and then held fixed, and reports the mean site height, the off-site noise floor, and the paired per-site change between arms. Numeric arm labels turn a set of arms into a curve with a per-dataset optimum. The dataset table lives outside the repository, as rugnux_vs_xds.py already does for the battery, because dataset and sample identities are not committed. It reproduces the measurements it was built from: all nine points of three pooled curves, every per-crystal optimum, the site heights, the paired t statistics, and the adversarial control in which a model refined against the worst arm reproduces the curves to <=0.005 and the same optimum on 4/4. Four things the ad-hoc scripts it replaces got wrong, all now handled: * Keying sites on the ANODE atom label silently drops an alternate conformation sharing that label - one dataset class has 18 sulfur sites, not 17, and the uncorrected mean read 15.03 against a true 14.49. * The off-site floor skipped any peak within 1.0 A of ANY atom, so a ripple sitting on a light atom was not counted as background; requiring 1.5 A from an anomalous scatterer raises one floor from 7.12 to 9.48 sigma. * Special-position peaks are Fourier ripples, not background. Excluding them is load-bearing on 3 of 9 datasets and they are now reported in their own column rather than dropped silently. * Enantiomorph care turned out to be unnecessary - passing the merged file's screw label to ANODE while the model sits in the other hand gives byte identical peaks. What does matter is the pair whose absences are identical, I23 vs I2_13, which phaser's automatic hand test does not cover. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bb7df09086 |
rugnux: separate the merge-time ice-ring mask from ice detection
--detect-ice-rings did two unrelated jobs at once: flagging ice spots so indexing de-prioritises them and keeping ice reflections out of the scale fit, AND gating the merge-time mask that drops a decorrelated ice ring and re-merges. Turning it off to de-confound a merge-stage experiment therefore also changed how the data were indexed - measured, that breaks indexing outright on two of the 37 rotation battery crystals - while leaving it on lets the mask land differently between two arms of an experiment and contaminate the comparison (measured on up to 19 of 37 crystals in response to a small intensity change). Add --ice-ring-mask[=on|off], default on, gating only the merge-time mask. Verified with =off: ice-spot flagging and the scaling exclusion still log and still apply, no mask line, no second merge, and the first error model is bit-identical to the =on arm. The full pipeline and the offline --scale path reach the same verdict on the same data, as they must. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
09fb8e0306 |
Bragg integration: clip the background ring high side instead of trimming it
The r2..r3 background ring was averaged with a 10% SYMMETRIC trimmed mean. A symmetric trim is not a consistent estimator of the mean of a right-skewed (Poisson) sample: on a clean Poisson ring it sits ~0.1 ct/px BELOW the true mean at every level, and with ~50 signal pixels in the r1 disk that under-subtraction adds ~5 counts to every partial on every frame. Measured two independent ways on four rotation datasets - stored background_mean against a plain ring mean over the same pixels on reflection-free frames, and directly on apertures that provably hold no reflection. Empty-aperture pedestal, counts: plain mean -0.03..-0.20, 10% symmetric trim +5.05..+6.34, 4 sigma clip +0.02..+0.54. Replace it with a high-side-only sigma clip at mean + n*sqrt(mean), n = 4 for monochromatic data. It rejects the same one-sided contamination the trim was there for - better, in fact: a 40 px neighbour core at +100 ct shifts the trim by +10.1 ct/px, because a symmetric trim collapses once contamination exceeds ~10% of the ring, versus +0.009 ct/px at 4 sigma. False rejection on a clean ring is 0.04-0.39%. Broadband data keep their tuned 3 sigma clip unchanged. The trim stays reachable with --background-trim for back compatibility; setting either estimator clears the other, so they can never stack. --integrator boxsum does not take the clip (matching what the shipped clip already did), so it now uses the plain ring mean unless --background-trim is given. The intensities get measurably more accurate: per-shell agreement with an independent processing of the same images improves on 14 of 16 crystals (weighted -0.0347, outermost shell 12/4), the outermost-shell R_meas NUMERATOR - absolute scatter, not a denominator effect - falls 13.5% median on 16/5, and CC1/2 in the outer shell improves on 14/7. EXPECT <I/sigma> TO FALL AND EDGE R_meas TO RISE. Both are inflated by information-free counts, so both get worse when the bias is removed; neither is evidence against this change. That fingerprint is exactly how the trimmed mean was accepted in the first place. Known cost: over the 37-crystal rotation battery the de-novo space-group count goes 34 OK / 3 DIFF to 33 / 4. The single regression is a two-lattice crystal whose merge fails the absolute-sanity gate under either background (R_meas 63.5%, CC1/2 72.2%) and which carries an unresolved indexing ambiguity on the very operator being scored, so its operator CC is diluted by construction. No other crystal changes space group, and twin protection is not weakened - the H-ratio veto that refuses genuinely twinned crystals gets MORE decisive (1.63 -> 1.84, 2.83 -> 3.99). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
fb55645b81 |
Revert "rugnux: fit the profile radius from the strongest spots too"
Build Packages / build:viewer-tgz:cpu (push) Successful in 16m51s
Build Packages / build:viewer-tgz:cuda (push) Successful in 18m44s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 20m42s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 21m37s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 24m37s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 24m48s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 20m13s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 25m19s
Build Packages / build:rpm (rocky9) (push) Successful in 23m23s
Build Packages / DIALS test (push) Successful in 21m35s
Build Packages / Generate python client (push) Successful in 40s
Build Packages / build:rpm (rocky8) (push) Successful in 29m16s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (ubuntu2404) (push) Successful in 23m17s
Build Packages / XDS test (durin plugin) (push) Successful in 11m5s
Build Packages / Build documentation (push) Successful in 1m14s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 27m29s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m21s
Build Packages / XDS test (neggia plugin) (push) Successful in 7m24s
Build Packages / build:windows:nocuda (push) Successful in 13m58s
Build Packages / build:windows:cuda (push) Successful in 16m6s
Build Packages / Unit tests (push) Successful in 1h18m59s
Reverts the profile-radius part of 457b1bfd1; the comparison-script and mosaicity-column changes from that commit are kept. The cap was validated on the rotation battery, which cannot test it: the profile radius feeds `ewald_dist_cutoff` in IndexAndRefine, and that is read only by the STILLS predictors (BraggPrediction/BraggPredictionGPU). The rotation predictors gate on the mosaicity window instead and never look at it. So "no space-group changes, 36 of 37 crystals bit-identical" showed the quantity is inert for rotation, not that capping it is safe - and the one regime where it does act was never exercised. Validating it needs the serial-stills battery, which is a much larger exercise. Until then the arbitrary constant is not worth carrying in a code path nobody measured. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5eb386e333 |
docs: describe the per-frame geometry smoothing
Build Packages / build:viewer-tgz:cpu (push) Successful in 19m9s
Build Packages / build:viewer-tgz:cuda (push) Successful in 21m0s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 23m12s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m58s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m17s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m53s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m59s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 18m55s
Build Packages / XDS test (durin plugin) (push) Successful in 11m20s
Build Packages / build:rpm (rocky9) (push) Successful in 21m37s
Build Packages / Generate python client (push) Successful in 38s
Build Packages / build:rpm (rocky8) (push) Successful in 24m58s
Build Packages / Create release (push) Skipped
Build Packages / Build documentation (push) Successful in 1m29s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 11m23s
Build Packages / DIALS test (push) Successful in 20m30s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m7s
Build Packages / XDS test (neggia plugin) (push) Successful in 9m26s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 26m1s
Build Packages / Unit tests (push) Successful in 1h27m56s
Build Packages / build:windows:nocuda (push) Successful in 16m37s
Build Packages / build:windows:cuda (push) Successful in 17m39s
Goes in §10.3 next to the per-frame scale and mosaicity smoothing, since it is the same mechanism applied for the same reason, and trims the changelog line to one sentence. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
017f64690c |
rugnux: smooth the per-frame geometry before scaling
Geometry is re-refined independently on every frame, against that frame's spots alone - as few as a dozen on a sparse crystal, where XDS fits its equivalent to about sixty times more data. Measured over ten datasets the per-frame orientation carries two components: a slow drift that is real, with rugnux and XDS agreeing to R^2 0.83-0.88 on the two crystals that genuinely slip by 1.5 and 0.54 degrees, and a fast jitter that is fit noise, scaling with spots-per-frame at exponent -0.79 where counting noise alone would give -0.5. The jitter is worth 1-8% on merged intensities, 24% on the sparsest crystal. It cannot be fixed by refining less. Turning per-image refinement off entirely loses six space groups and a whole crystal, and even a 624-spot-per-frame crystal collapses; dropping the beam-centre terms holds the space groups but is worse on 31 of 37 crystals. The freedom is earning its keep, so keep it and suppress only the band that cannot be physical - a crystal does not re-orient and snap back from one frame to the next. So smooth the orientation in frame order after integration and recompute each partial's delta_phi, and hence its partiality, from the smoothed lattice. Batching at integration time was not an option: frames are processed independently and the online path depends on that. This runs before the GPU upload, so the device path picks it up with no separate kernel. The window is chosen per dataset by leave-one-out cross-validation, because the two components vary far too much for one number - drift spans 0.018 to 1.288 degrees and jitter 0.005 to 0.221, so any fixed window over-smooths one crystal while under-smoothing another. Chosen windows range from +-1 to +-20 frames. It is capped: cross-validation scores how well neighbours predict a frame's orientation, which on a barely-drifting crystal keeps improving with width, but the per-frame fit is also absorbing a real per-frame systematic and smoothing too wide destroys it - uncapped, one crystal chose +-60 and lost 16% of its ISa. Battery over 37 crystals: space groups unchanged at 34 matching XDS, R_meas better on 31 and worse on 6, low-resolution R_meas 30/7, ISa 26/10, high-resolution CC1/2 23/12. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
457b1bfd1d |
rugnux: fit the profile radius from the strongest spots too
Build Packages / build:viewer-tgz:cpu (push) Successful in 18m20s
Build Packages / build:viewer-tgz:cuda (push) Successful in 20m23s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m35s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 23m47s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m24s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m35s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 29m9s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 19m23s
Build Packages / XDS test (durin plugin) (push) Successful in 10m17s
Build Packages / build:rpm (rocky9) (push) Successful in 20m45s
Build Packages / Generate python client (push) Successful in 33s
Build Packages / Build documentation (push) Successful in 1m7s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (rocky8) (push) Successful in 26m5s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 11m15s
Build Packages / DIALS test (push) Successful in 20m23s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 20m50s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m36s
Build Packages / XDS test (neggia plugin) (push) Successful in 10m3s
Build Packages / Unit tests (push) Successful in 1h17m43s
Build Packages / build:windows:nocuda (push) Successful in 16m24s
Build Packages / build:windows:cuda (push) Successful in 17m50s
Same defect as the mosaicity in
|
||
|
|
2c94f3013e |
rugnux: fit the mosaicity from the strongest spots only
Build Packages / build:viewer-tgz:cpu (push) Successful in 21m50s
Build Packages / build:viewer-tgz:cuda (push) Successful in 22m21s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 22m34s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 24m7s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 28m34s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 28m30s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 28m37s
Build Packages / XDS test (durin plugin) (push) Successful in 11m30s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 22m6s
Build Packages / build:rpm (rocky9) (push) Successful in 21m35s
Build Packages / Generate python client (push) Successful in 43s
Build Packages / Build documentation (push) Successful in 1m17s
Build Packages / Create release (push) Skipped
Build Packages / DIALS test (push) Successful in 20m20s
Build Packages / build:rpm (rocky8) (push) Successful in 27m13s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 25m40s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 21m19s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 10m37s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m5s
Build Packages / Unit tests (push) Successful in 1h19m31s
Build Packages / build:windows:nocuda (push) Successful in 19m12s
Build Packages / build:windows:cuda (push) Successful in 22m29s
The per-image mosaicity MLE ran over the whole indexed spot list, so it rode on --max-spots, which is an indexing budget. A spot is detected when I_full * R(tau) clears the finder threshold, so selecting by intensity censors on R(tau): a deeper list holds proportionally more large-|tau| partially recorded spots and the fit widens with it. Raising the budget 250 -> 1000 widened sigma_M 0.059 -> 0.075 deg on a rotation dataset whose measured rocking width says 0.054. That is not cosmetic. An over-wide mosaicity mis-states every partiality in scaling: forcing the mosaicity across that range moved the merge error model from b 0.039 / ISa 26 to b 0.167 / ISa 6, and the space-group search lost a genuine 422 with it, merging the crystal in 222 instead. Cap the fit at the strongest 250 spots. FilterSpotsByCount leaves the list strongest-first, so this selects exactly the spots a smaller --max-spots would, and the mosaicity becomes invariant: 0.0538 deg at 250, 500, 1000 and 2000 spots, with the correct space group at each. Trimming or down-weighting the tau tail does not work - the censoring is multiplicative in R(tau), so it widens the whole distribution rather than adding a tail. Battery over 37 crystals: exactly one change, the demoted crystal repaired (33 space groups matching XDS -> 34). 23 of 37 are bit-identical, never reaching 250 spots. Unaffected elsewhere: the default spot count is 250, and stills have no goniometer so they return before the fit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f25fea7024 |
rugnux: keep 1000 spots per image instead of 250
Offline reprocessing is not bound by the online spot budget, and the cap is applied at the end of SpotAnalyze, so it is exactly the spot list the indexer and the per-image refinement see. jfjoch_viewer already sends 1000, so the two front ends now agree on the same file. Measured as a paired A/B over the 37-crystal rotation battery, de novo, with the resolution and Friedel setting matched to the XDS reference, both arms from the same binary bar this constant: R_meas low shell 16 better 0 worse 19 unchanged R_meas 14 better 4 worse 17 unchanged ISa 14 better 6 worse 15 unchanged CC1/2 6 better 3 worse 26 unchanged Low-resolution R_meas is a clean sweep. Around half the battery is bit-identical: those frames never reach 250 spots, so the cap never bound. Wall clock is unchanged (10m00s vs 10m44s, uncontrolled for page cache). Known cost, and the reason this is its own commit: one crystal in the battery reproducibly loses symmetry, tetragonal 422 -> orthorhombic 222, doubling its asymmetric unit. Its R_meas and ISa "improve" there, but that is what merging in too low a symmetry always does, and the lower symmetry then admits a merohedral indexing ambiguity. An intermediate cap of 500 demotes it too, so it buys none of the safety. This is the known point-group-decision-moves-with-data-amount fragility of the space-group search rather than an argument for starving the indexer of spots - the search is the thing to fix. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2473e03cf7 |
rugnux: default --spot-sigma to 4.0, the value the viewer already uses
The two front ends disagreed on the fixed-threshold spot finder: rugnux started from 3.0, jfjoch_viewer from the SpotFindingSettings default of 4.0, so the same file processed either way could give different spots. Inert on the default path - the adaptive finder derives its threshold from each image's own per-resolution-ring noise and never reads signal_to_noise_threshold (only ImageSpotFinderCPU/GPU and DetModuleSpotFinder do). It changes behaviour only under --no-adaptive-spots, and there it now matches the viewer. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |