The rotation merge fitted var = a*s2 + b^2*<I>^2 from three separate medians (s2, I2, dev2) per bin of I2. A median of dev2 over observations whose variances differ is not 0.455 times their mean variance, so the ratio of medians read a too low and b too high: on the scaled fulls of 28 sets (in-house, open and private) the core of the normalised deviations scattered at up to 1.8x its stated variance in the weak and middle bins and at 0.1-0.7x in the strongest. Reproduced on synthetic samples with a known model (a 1.3 read as 1.12, ISa 33 read as 31). Now (ErrorModel.h/.cpp, host-only, so the GPU and CPU paths share it): - bins are equal counts in counting I/sigma (I2/s2), where b is identified; - each bin is calibrated on the median of dev2/var with var from the previous iteration, iterated to a fixed point - heterogeneity inside a bin no longer biases it, and the median keeps it robust to tails; - s2 is the counting variance the merge actually applies the model to (rebuilt at the reflection's mean), not the observation's own sigma^2. The separate 6-sigma misfit refit is gone: the median does not need it. A mean-based fit (misfits cut at z^2 > 2 ln N) was tried first: it calibrates the total variance best (median rms log chi2 over the 28 sets 0.14 vs 0.19 here) but on heavy-tailed data it sizes the sigmas on the tails, the merge's outlier test widens with them, and CC1/2 fell 0.80 -> 0.71 on a powder-contaminated set (0.83 with this fit). Rejected for that. Offline, 28 sets: rms log chi2 of the median normalised deviation over (counting I/sigma x resolution) 0.208 -> 0.139 (better on 23), of the mean 0.239 -> 0.191 (better on 22). Battery, 45 of 48 sets against the 9b6736 run (3 lost to CUDA OOM from GPU contention): space group unchanged on all; d_min unchanged except two poor multi-lattice sets (1.69 -> 1.56, 1.96 -> 1.90); ISa x1.13 (median), in-house ISa/XDS 0.73 -> 0.93; ISa*R_meas_lo/0.8 0.93 -> 1.05 (XDS ~1.2); CC1/2 over the XDS range +0.002 (mean; up 0.014-0.031 on the three poorest sets, else +-0.0001); CC_model +0.0011, R_model_shell_scaled -0.0007 (mean over 17 open sets); CC_anom +0.003 (mean). Six private sets: space group, d_min and CC1/2 unchanged, ISa up by 14-67% towards XDS's. Remaining misfit, not addressed: the excess variance grows slower than <I>^2 (the effective fractional error falls 2-2.5x from counting I/sigma 5 to 200), so the strongest reflections still scatter below their sigma on open sets. A third, linear term (as in Aimless) fits it better on most sets but leaves b unidentified on some (b -> 0 on 4 of 28); not landed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
49 lines
2.5 KiB
C++
49 lines
2.5 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <cstddef>
|
|
#include <vector>
|
|
|
|
// The error model of the rotation merge: var = a * s2 + b^2 * <I>^2, fitted so that the scatter of
|
|
// symmetry equivalents matches it - in each bin of counting I/sigma, the median of dev2 / var is the
|
|
// median of a chi^2 with one degree of freedom.
|
|
//
|
|
// One sample per observation: s2 is its counting variance as the merge rebuilds it at the reflection's
|
|
// mean intensity, I2 that mean squared, dev2 its leverage-corrected squared deviation from the mean.
|
|
//
|
|
// Two choices, each against a measured failure of the fit this replaces, which took the medians of s2,
|
|
// I2 and dev2 separately over bins of I2:
|
|
// * Bins are equal counts in COUNTING I/sigma (I2 / s2), not in I2. b is identified by how far the
|
|
// bins reach into the regime where b*<I> dominates the counting term, and that is what this ranks
|
|
// on; ranking on I2 mixes weak high-resolution reflections into the strong bins.
|
|
// * The median is taken of the NORMALISED deviation, each sample divided by the variance the previous
|
|
// iteration gave it. A median of dev2 over observations whose variances differ is not 0.455 times
|
|
// their mean variance, and the ratio of three medians read a too low and b too high: on most of 28
|
|
// sets the core of the normalised deviations then scattered at up to 1.8 times its stated variance
|
|
// in the weak and middle bins and at 0.1-0.7 times it in the strongest.
|
|
// A median, not a mean: on heavy-tailed data (split and powder-contaminated crystals) a mean calibrates
|
|
// the sigmas on the tails, the merge's outlier test widens with them, and CC1/2 fell 0.80 -> 0.71 on one
|
|
// such set where this fit raises it to 0.83.
|
|
struct ErrorModelSample {
|
|
double s2, I2, dev2;
|
|
float d;
|
|
};
|
|
|
|
struct ErrorModelFit {
|
|
bool active = false; // the pool was large enough to fit
|
|
double a = 1.0, b2 = 0.0;
|
|
bool b_measured = true; // false: no bin reached where b could be seen; b is held at 0
|
|
bool b_resolved = true; // b^2 stood clear of two of its standard errors
|
|
};
|
|
|
|
// Scratch: the partitioned copy of the pool (reused between calls to save the allocation), with the rank
|
|
// key and the current normalised deviation.
|
|
struct ErrorModelBinned {
|
|
double snr2, s2, I2, dev2, z2;
|
|
};
|
|
|
|
ErrorModelFit FitErrorModel(const std::vector<ErrorModelSample> &pool, std::vector<ErrorModelBinned> &scratch,
|
|
size_t nthreads);
|