Bragg integration: keep a reflection that lost a wing, not one that lost its peak
MINPK asks how MUCH of the expected profile is readable. It does not ask WHERE, and the two are not the same question. The renormalisation argument the rescue rests on - a fit over a subset of a normalised profile is unbiased - needs the pixels to go missing for reasons unrelated to the reflection. A gap, a mask or the edge of the sensor is such a reason: the loss is set by the detector, and the fit renormalises over what is left. A pixel invalidated BY THE FLUX IT SAW is not: it goes missing because the reflection was bright, and it is the peak. Measured on the combined fulls, against the mean of the complete observations of the same reflection, in the innermost resolution shell of the high-multiplicity control and of a weaker crystal: a rescued reflection whose unreadable pixel sits within a pixel of the predicted centre reads |I - <I>|/I of 0.50 and 0.53, against 0.073 and 0.212 for a complete observation - 6.8x and 2.5x - and carries several times the mean intensity of its shell. On the control that is 0.21% of the shell's observations supplying 1.77% of the R_meas numerator; on the weaker crystal 0.52% supplying 6.82%. Rescues that lost only rim pixels are unremarkable by the same measure, 1.19x and 0.88x. Dropping the peak-losers alone takes the shell's R_meas from 7.440% back to 7.315% (unrescued: 7.307%) and from 22.03% to 21.06% (unrescued: 21.28%) - which is the whole of the low-resolution R_meas the rescue cost, and on the second crystal rather more. Raw frames say what they are. The pattern is a dead-centre invalid pixel with 5878, 9875 and 27583 counts around it: the detector's per-frame invalid marker on the brightest reflections. MINPK cannot catch them because it cuts on profile MASS, and the peak of a broad spot is a few percent of the mass. So a second condition, in the loop that already measures the readable fraction: no unreadable pixel may carry more than 0.9 of the profile's own peak value. A fraction of the peak rather than a radius in pixels because the peak is as wide as the spot - for a Gaussian the cut is at sqrt(-2 ln f) sigma, 0.46 sigma here, which is the peak pixel alone where sigma is 0.8 px and the crest of the ridge where the profile is a bandwidth streak. Swept against the alternatives on two crystals: a fixed radius needs 1.0-1.5 px to do the same work and costs 3-9x more observations for it, and 0.5 px does not reach the peak of a sub-pixel-offset prediction at all; tightening the fraction to 0.5 or 0.2 buys nothing beyond 0.9 and costs 7x more. Six crystals, three detectors, against the rescue as it stands: the rule keeps 99.86-99.96% of the recovered observations and returns R_meas to its unrescued value or below (4.6 -> 4.5%, 6.7 -> 6.6%, 25.1 -> 25.0%), R_meas in the innermost shell likewise (2.7 -> 2.6%, 5.9 -> 5.3% against 5.4% unrescued, 16.5 -> 16.4%), <I/sigma> up or level everywhere, and every unique reflection the rescue won is kept. Raising --overlap-minpk to 0.90 instead reaches the same place on two of them and short of it on the third, while discarding 0.8% of the recovered observations rather than 0.05%. An elongated pink-beam profile on a 9M detector and an EIGER2 16M dataset are both untouched at 99.9%, so the crest protection does not over-reject a streak. One crystal is not improved: a dataset whose error model rugnux declines to fit for want of strong reflections, whose <I/sigma> is <= 0 in eight of its ten shells and whose R_meas is undefined in as many. There the rule costs about 3% of <I/sigma> in the one shell that has signal, reproducibly, on top of the 9% the rescue itself costs there - while its overall R_meas moves 1.5 points on nothing but the thread count. The parity test gains four sections. Unreadable pixels were only ever punched into empty sky, so neither the rescue nor this rule had any CPU/GPU coverage at all; they now go into the signal disks - the peak of every fifth reflection, ~1.1 sigma out of every seventh, the disk edge of every eleventh - for both profile modes, a box sum and an elongated stencil, with a check that the clipping actually costs reflections so the coverage cannot go quietly vacuous. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Full 38-crystal rotation battery against the rescue without this rule, both on the same base: ISa better 19 / worse 4, +0.73 CC1/2 better 3 / worse 1, +1.3 R_meas_lo better 4 / worse 3, -0.3 space groups unchanged for 17 770 observations, 0.09 % of the run total and under 2 % of what the rescue had won. The two crystals whose peak-loss population was measured beforehand land on their predicted values: a tetragonal reference goes R_meas_lo 2.7 -> 2.6 % and ISa 27.11 -> 27.42, a cubic insulin 5.9 -> 5.3 % and 20.34 -> 20.65. One crystal pays: a cubic case with 2381 unique reflections goes R_meas 8.8 -> 9.6 % and ISa 4.08 -> 3.49. It is the crystal in the battery with the fewest uniques, so its rescued population is small and its shell statistics are coarse, but the loss is real and not noise in the R_meas. The R_meas sum over the battery reads +1.3, of which +3.2 is one crystal whose R_meas moves 1.5 points on thread count alone; without it the sum is negative. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+1
-1
@@ -9,7 +9,7 @@ This is an UNSTABLE release. It includes many experimental features, as well as
|
||||
* Bragg integration: the profile fit's `background_variance` now takes the fitted intensity itself out of the fit variance instead of `max(0, I)`, so a reflection that fluctuated below zero no longer reports a background variance two to three times too small and is no longer weighted up for it.
|
||||
* Scaling: the rotation merge weights each combined full by its variance rebuilt at the reflection's mean intensity rather than by the full's own sigma, as the stills merge already did.
|
||||
* Scaling: a rotation frame whose fitted scale collapses - it recorded no diffraction from the indexed lattice - is now **dropped from the merge** instead of being merged unscaled, which had asserted a scale of 1 for a frame demonstrably nowhere near it.
|
||||
* Bragg integration: a reflection whose signal disk is cut by a mask, an untrusted region, a detector gap or an overload is now profile-fitted over the pixels that remain instead of being discarded, as long as at least `--overlap-minpk` of its expected profile is readable (XDS's MINPK); `--integrator boxsum` still discards it.
|
||||
* Bragg integration: a reflection whose signal disk is cut by a mask, an untrusted region, a detector gap or an overload is now profile-fitted over the pixels that remain instead of being discarded, as long as at least `--overlap-minpk` of its expected profile is readable (XDS's MINPK) **and the unreadable part does not take the profile's peak**; `--integrator boxsum` still discards it.
|
||||
* rugnux: Rotation data are integrated on **every frame whose spots the sweep's lattice explains**, instead of only on frames that would also index on their own; the reported indexing rate still counts the latter.
|
||||
* Scaling: a rotation frame too sparse to fit a rocking width of its own now takes the run's median instead of a fixed default.
|
||||
* rugnux: The error-model **a** and **b** are reported in XDS's convention, and `_reflns.jfjoch_diffrn_ISa` now carries the whole-range `1/sqrt(a*b)` that XDS's ISa denotes; the strong-reflection asymptote moves to `_reflns.jfjoch_diffrn_ISa_asymptotic`. **A file written by an earlier version carries the asymptote under the old name.**
|
||||
|
||||
@@ -640,6 +640,8 @@ where $c$ is the pixel value and the de-biased variance $v$ (background plus mod
|
||||
|
||||
**Pixels the fit cannot use (MINPK).** A profile fit is the amplitude of a *normalised* profile, so a pixel left out of the sum renormalises the estimator by construction: it costs information — $\sum P^2/v$ shrinks and $\sigma$ grows — but biases nothing. That is what keeps a reflection whose signal disk is cut by a mask, an untrusted region, a detector gap or an overload: those pixels are simply not read, and the fit is taken over the rest, exactly as the shared pixels of a crowded reflection are (`--overlap exclude`). The reflection is kept only while enough of the expected profile survives — at least `--overlap-minpk` of the profile mass that falls on the detector at all, default 0.75, which is XDS's `MINPK` and dials' `valid_foreground_threshold`. The complete reflections alone teach the profile, its resolution shells and their widths. `--integrator boxsum` has no profile to renormalise with and keeps the all-or-nothing rule of §9.2.
|
||||
|
||||
"Biases nothing" holds while the pixels go missing for reasons that have nothing to do with the reflection, which is true of a gap, a mask or the edge of the sensor — where the loss is set by the detector, and the same hole recurs on every frame of the rocking curve because the spot does not move off it. It is not true of a pixel invalidated *by the flux it saw*: that pixel goes missing **because** the reflection was bright, and it is the peak. The fit then has only the wings to set the amplitude from and reads low — measured at $-50\%$ against the symmetry mates, on reflections carrying several times the mean intensity of their shell, which are the largest terms of $R_\mathrm{meas}$. MINPK cannot separate the two cases, because it cuts on profile *mass* and the peak of a broad spot is a few percent of the mass. So a second condition applies alongside it: **no unreadable pixel may carry more than 0.9 of the profile's own peak value**. As a fraction of the peak rather than a radius in pixels, that scales with the spot — for a Gaussian it is a cut at $\sqrt{-2\ln f}\,\sigma = 0.46\sigma$, the peak pixel alone where $\sigma$ is 0.8 px and the crest of the ridge where the profile is a bandwidth streak — and it needs nothing the fit does not already compute. It costs 0.05–0.07 % of the recovered observations.
|
||||
|
||||
The integrator is selected by `--integrator boxsum|gaussian|empirical` (default `gaussian`).
|
||||
|
||||
### 9.4 Lorentz–polarization factor handling
|
||||
|
||||
@@ -84,6 +84,22 @@ constexpr float MAX_STENCIL_GROW_OVER_R3 = 2.0f;
|
||||
// summation (box-sum) intensity when the profile result disagrees with the summation seed by more than
|
||||
// this many box-sum sigmas (a real fit agrees within counting noise, so the margin is generous).
|
||||
constexpr double PROFILE_SUMMATION_MAX_NSIGMA = 10.0;
|
||||
|
||||
// MINPK keeps a reflection while enough of its expected profile is readable. It says nothing about
|
||||
// WHERE the unreadable part is, and the two are not the same question. A pixel lost to a gap, a mask
|
||||
// or the edge of the sensor is lost for reasons that have nothing to do with this reflection, and the
|
||||
// fit renormalises over what is left with no bias. A pixel lost because the flux it saw put it over
|
||||
// the detector's range is lost BECAUSE the reflection was bright, and it is the peak: the fit then has
|
||||
// only the wings to set the amplitude from, and reads low - measured at -50% on the strongest
|
||||
// low-resolution reflections, which are also the largest terms of R_meas. So no unreadable pixel may
|
||||
// carry more than this fraction of the profile's own peak value.
|
||||
//
|
||||
// A fraction of the peak rather than a radius in pixels, because the peak is as wide as the spot: for
|
||||
// a Gaussian the cut sits at sqrt(-2 ln f) sigma, i.e. 0.46 sigma here, which is the peak pixel alone
|
||||
// where sigma is 0.8 px and the crest of the ridge where it is 2.4 px or a bandwidth streak. It also
|
||||
// needs nothing the fit does not already have, so it costs one max-reduction in the loop that
|
||||
// measures the readable fraction, and it applies to the learned empirical profile unchanged.
|
||||
constexpr double MINPK_MAX_MISSING_PEAK = 0.9;
|
||||
} // namespace bragg_engine
|
||||
|
||||
// One reflection's extracted intensity, produced by the derived engine and turned into a
|
||||
|
||||
@@ -479,6 +479,7 @@ std::vector<Reflection> BraggIntegrationEngineCPU::RunImpl(const Sampler &img,
|
||||
// summation seed the runaway guard compares against actually summed; with nothing missing
|
||||
// and nothing excluded it is 1 and the guard is untouched. ---
|
||||
double p_grid = 0.0, p_valid = 0.0, p_own = 0.0, m_all = 0.0, m_read = 0.0;
|
||||
double p_peak = 0.0, p_lost_peak = 0.0;
|
||||
for (int dy = -Rf; dy <= Rf; ++dy)
|
||||
for (int dx = -Rf; dx <= Rf; ++dx) {
|
||||
const double Pp = (*Pvec)[(dy + Rf) * Gf + (dx + Rf)];
|
||||
@@ -487,14 +488,23 @@ std::vector<Reflection> BraggIntegrationEngineCPU::RunImpl(const Sampler &img,
|
||||
if (x < 0 || y < 0 || x >= W || y >= H) continue;
|
||||
const bool in_disk = dx * dx + dy * dy < r1_sq;
|
||||
p_grid += Pp;
|
||||
p_peak = std::max(p_peak, Pp);
|
||||
if (in_disk) m_all += Pp;
|
||||
if (!valid(img[y * W + x])) continue;
|
||||
if (!valid(img[y * W + x])) {
|
||||
p_lost_peak = std::max(p_lost_peak, Pp);
|
||||
continue;
|
||||
}
|
||||
p_valid += Pp;
|
||||
const bool own = clean(x, y, i);
|
||||
if (own) p_own += Pp;
|
||||
if (in_disk && (own || !exclude)) m_read += Pp;
|
||||
}
|
||||
if (p_valid < overlap_min_peak * p_grid) continue;
|
||||
// A hole in the profile's PEAK is a different defect from a hole in its wings, and the mass
|
||||
// fraction above cannot tell them apart - the peak of a broad spot is a few percent of the
|
||||
// mass, so MINPK passes a reflection that has lost the one part of the profile its amplitude
|
||||
// is determined by. See the header.
|
||||
if (p_lost_peak > MINPK_MAX_MISSING_PEAK * p_peak) continue;
|
||||
if (overlap == OverlapMode::Reject && p_own < overlap_min_peak) continue;
|
||||
|
||||
const double B = std::max(rh.bkg, PIXEL_VARIANCE_FLOOR);
|
||||
|
||||
@@ -34,6 +34,7 @@ struct BraggGpuParams {
|
||||
float claim_sq, inv_claim; // how far a reflection claims pixels in the owner map
|
||||
float minpk; // least readable/clean profile fraction that is kept (XDS MINPK)
|
||||
int partial_ok; // keep a signal disk with unreadable pixels: anything but a box sum
|
||||
float peak_frac; // most of the profile's peak value an unreadable pixel may carry
|
||||
};
|
||||
|
||||
__device__ inline bool valid(int32_t v) { return v != INT32_MIN && v != INT32_MAX; }
|
||||
@@ -467,6 +468,9 @@ __global__ void fit(const int32_t *img, const uint32_t *owner, const float *px_x
|
||||
extern __shared__ float Pbuf[];
|
||||
__shared__ float s_gs, s_num, s_den, s_I, s_wsum;
|
||||
__shared__ float s_pgrid, s_pvalid, s_pown, s_mall, s_mread;
|
||||
// Max reductions. The profile is non-negative, and for non-negative floats the IEEE bit pattern
|
||||
// orders exactly as the value does, so an integer atomicMax on that pattern is an exact float max.
|
||||
__shared__ int s_ppeak_i, s_plost_i;
|
||||
__shared__ int s_Rf, s_Gf;
|
||||
|
||||
if (!ok_a[i]) { if (threadIdx.x == 0) ok_o[i] = 0; return; }
|
||||
@@ -526,9 +530,11 @@ __global__ void fit(const int32_t *img, const uint32_t *owner, const float *px_x
|
||||
// CPU engine.
|
||||
if (threadIdx.x == 0) {
|
||||
s_pgrid = 0.0f; s_pvalid = 0.0f; s_pown = 0.0f; s_mall = 0.0f; s_mread = 0.0f;
|
||||
s_ppeak_i = 0; s_plost_i = 0;
|
||||
}
|
||||
__syncthreads();
|
||||
float l_pgrid = 0.0f, l_pvalid = 0.0f, l_pown = 0.0f, l_mall = 0.0f, l_mread = 0.0f;
|
||||
float l_ppeak = 0.0f, l_plost = 0.0f;
|
||||
for (int k = threadIdx.x; k < GfGf; k += blockDim.x) {
|
||||
const float Pp = Pbuf[k];
|
||||
if (Pp <= 0.0f) continue;
|
||||
@@ -537,8 +543,12 @@ __global__ void fit(const int32_t *img, const uint32_t *owner, const float *px_x
|
||||
if (x < 0 || y < 0 || x >= p.W || y >= p.H) continue;
|
||||
const bool in_disk = (float) (dx * dx + dy * dy) < p.r1_sq;
|
||||
l_pgrid += Pp;
|
||||
l_ppeak = fmaxf(l_ppeak, Pp);
|
||||
if (in_disk) l_mall += Pp;
|
||||
if (!valid(img[y * p.W + x])) continue;
|
||||
if (!valid(img[y * p.W + x])) {
|
||||
l_plost = fmaxf(l_plost, Pp);
|
||||
continue;
|
||||
}
|
||||
l_pvalid += Pp;
|
||||
const bool own = !p.overlap || BraggOwnedBy(owner[y * p.W + x], i);
|
||||
// Zeroing the profile here is how Exclude drops the pixel: the fit skips any cell with
|
||||
@@ -550,11 +560,17 @@ __global__ void fit(const int32_t *img, const uint32_t *owner, const float *px_x
|
||||
}
|
||||
atomicAdd(&s_pgrid, l_pgrid); atomicAdd(&s_pvalid, l_pvalid); atomicAdd(&s_pown, l_pown);
|
||||
atomicAdd(&s_mall, l_mall); atomicAdd(&s_mread, l_mread);
|
||||
atomicMax(&s_ppeak_i, __float_as_int(l_ppeak)); atomicMax(&s_plost_i, __float_as_int(l_plost));
|
||||
__syncthreads();
|
||||
if (s_pvalid < p.minpk * s_pgrid) {
|
||||
if (threadIdx.x == 0) ok_o[i] = 0;
|
||||
return;
|
||||
}
|
||||
// A hole in the profile's PEAK is a different defect from a hole in its wings. See the CPU engine.
|
||||
if (__int_as_float(s_plost_i) > p.peak_frac * __int_as_float(s_ppeak_i)) {
|
||||
if (threadIdx.x == 0) ok_o[i] = 0;
|
||||
return;
|
||||
}
|
||||
if (p.overlap == 1 && s_pown < p.minpk) {
|
||||
if (threadIdx.x == 0) ok_o[i] = 0;
|
||||
return;
|
||||
@@ -745,6 +761,7 @@ std::vector<Reflection> BraggIntegrationEngineGPU::Run(const ImagePreprocessorBu
|
||||
.exclude_px = (overlap == OverlapMode::Exclude && mode != IntegratorMode::BoxSum) ? 1 : 0,
|
||||
.claim_sq = claim * claim, .inv_claim = inv_claim, .minpk = overlap_min_peak,
|
||||
.partial_ok = mode != IntegratorMode::BoxSum ? 1 : 0,
|
||||
.peak_frac = static_cast<float>(MINPK_MAX_MISSING_PEAK),
|
||||
};
|
||||
|
||||
// Whether the radial correction runs for THIS image. n_rad only says the buffers exist - under
|
||||
|
||||
@@ -41,7 +41,12 @@ Reflection MakeReflection(float x, float y, float d, int hkl) {
|
||||
// companion_dx > 0 puts a second spot that many pixels beside every grid spot, so their r1 signal
|
||||
// disks share pixels while the background rings still see clean sky - which is what a dense pattern
|
||||
// actually looks like (crowded along one reciprocal axis, sparse across it).
|
||||
Scene BuildScene(size_t width, size_t height, int spacing = 60, float companion_dx = 0.0f) {
|
||||
// clip_spots punches unreadable pixels into the spots themselves rather than into empty sky: the
|
||||
// centre of every 5th, a mid-profile pixel of every 7th and a disk-edge pixel of every 11th. That is
|
||||
// the MINPK rescue's own case - a reflection kept and fitted over the pixels it has - and with it the
|
||||
// peak-loss rule, which has to fire on the same reflections in both engines.
|
||||
Scene BuildScene(size_t width, size_t height, int spacing = 60, float companion_dx = 0.0f,
|
||||
bool clip_spots = false) {
|
||||
Scene s;
|
||||
s.width = width;
|
||||
s.height = height;
|
||||
@@ -91,6 +96,19 @@ Scene BuildScene(size_t width, size_t height, int spacing = 60, float companion_
|
||||
const size_t idx = (static_cast<size_t>(k) * 2654435761u) % s.image.size();
|
||||
s.image[idx] = (k % 2) ? INT32_MIN : INT32_MAX;
|
||||
}
|
||||
|
||||
if (clip_spots)
|
||||
for (size_t n = 0; n < s.predicted.size(); ++n) {
|
||||
int dx = 0, dy = 0;
|
||||
if (n % 5 == 0) { dx = 0; dy = 0; } // the peak itself: the rule must reject
|
||||
else if (n % 7 == 0) { dx = 1; dy = 1; } // ~1.1 sigma out: near the rule's boundary
|
||||
else if (n % 11 == 0) { dx = 3; dy = -2; } // disk edge: MINPK keeps it, the rule does not fire
|
||||
else continue;
|
||||
const int x = static_cast<int>(std::lround(s.predicted[n].predicted_x)) + dx;
|
||||
const int y = static_cast<int>(std::lround(s.predicted[n].predicted_y)) + dy;
|
||||
if (x < 0 || y < 0 || x >= static_cast<int>(width) || y >= static_cast<int>(height)) continue;
|
||||
s.image[y * width + x] = (n % 2) ? INT32_MAX : INT32_MIN;
|
||||
}
|
||||
return s;
|
||||
}
|
||||
|
||||
@@ -126,7 +144,8 @@ void CompareCpuVsGpu(IntegratorMode mode, std::optional<float> bandwidth_fwhm,
|
||||
float clip_nsigma = 4.0f, bool radial = false, int spacing = 60,
|
||||
float stencil_k = 0.0f,
|
||||
float r1 = 0.0f, float r2 = 0.0f, float r3 = 0.0f,
|
||||
OverlapMode overlap = OverlapMode::Off, float companion_dx = 0.0f) {
|
||||
OverlapMode overlap = OverlapMode::Off, float companion_dx = 0.0f,
|
||||
bool clip_spots = false) {
|
||||
const DiffractionExperiment experiment =
|
||||
MakeExperiment(mode, bandwidth_fwhm, clip_nsigma, radial, DetJF(2), stencil_k, r1, r2, r3,
|
||||
overlap);
|
||||
@@ -135,7 +154,7 @@ void CompareCpuVsGpu(IntegratorMode mode, std::optional<float> bandwidth_fwhm,
|
||||
const size_t npixel = experiment.GetPixelsNum();
|
||||
REQUIRE(npixel == width * height);
|
||||
|
||||
const Scene scene = BuildScene(width, height, spacing, companion_dx);
|
||||
const Scene scene = BuildScene(width, height, spacing, companion_dx, clip_spots);
|
||||
REQUIRE(scene.image.size() == npixel);
|
||||
REQUIRE(scene.predicted.size() > 60);
|
||||
|
||||
@@ -161,6 +180,18 @@ void CompareCpuVsGpu(IntegratorMode mode, std::optional<float> bandwidth_fwhm,
|
||||
// atomic summation of the learned profile, so compare up to a small tolerance.
|
||||
REQUIRE(out_gpu.size() == out_cpu.size());
|
||||
REQUIRE(out_cpu.size() > 40);
|
||||
if (clip_spots) {
|
||||
// Guard against the coverage going vacuous: the punched pixels have to actually cost some
|
||||
// reflections, or the two engines are being compared on a case neither of them meets.
|
||||
const Scene clean_scene = BuildScene(width, height, spacing, companion_dx, false);
|
||||
ImagePreprocessorBuffer clean_image(npixel);
|
||||
for (size_t i = 0; i < npixel; ++i)
|
||||
clean_image[i] = clean_scene.image[i];
|
||||
BraggIntegrationEngineCPU clean_cpu(experiment);
|
||||
const auto out_clean = clean_cpu.Run(clean_image, clean_scene.predicted,
|
||||
clean_scene.predicted.size(), 5);
|
||||
CHECK(out_cpu.size() < out_clean.size());
|
||||
}
|
||||
for (size_t i = 0; i < out_cpu.size(); ++i) {
|
||||
INFO("mode " << static_cast<int>(mode) << " reflection " << i << " hkl " << out_cpu[i].h);
|
||||
CHECK(out_gpu[i].h == out_cpu[i].h);
|
||||
@@ -239,6 +270,29 @@ TEST_CASE("BraggIntegrationEngineGPU_MatchesCPU") {
|
||||
0.0f, 0.0f, 0.0f, OverlapMode::Exclude);
|
||||
}
|
||||
SECTION("ProfileGaussian mono trim") { CompareCpuVsGpu(IntegratorMode::ProfileGaussian, std::nullopt, 0.0f); }
|
||||
// Unreadable pixels inside the signal disks themselves: the MINPK rescue keeps the reflection and
|
||||
// fits it over what is left, and the peak-loss rule throws back the ones that lost the profile's
|
||||
// maximum. Both decisions are per-reflection cuts on a reduction over the profile grid, computed
|
||||
// independently in the two engines (serial max vs an atomicMax on the float bit pattern), so they
|
||||
// have to reject exactly the same reflections - a mismatch shows up as a size mismatch here.
|
||||
SECTION("ProfileGaussian clipped disks") {
|
||||
CompareCpuVsGpu(IntegratorMode::ProfileGaussian, std::nullopt, 4.0f, false, 60, 0.0f,
|
||||
0.0f, 0.0f, 0.0f, OverlapMode::Off, 0.0f, true);
|
||||
}
|
||||
SECTION("ProfileEmpirical clipped disks") {
|
||||
CompareCpuVsGpu(IntegratorMode::ProfileEmpirical, std::nullopt, 4.0f, false, 60, 0.0f,
|
||||
0.0f, 0.0f, 0.0f, OverlapMode::Off, 0.0f, true);
|
||||
}
|
||||
// The same, with an elongated profile: the peak is then a ridge, so the fraction-of-peak test has
|
||||
// to protect a crest rather than one pixel, and the grid it reduces over is reflection-dependent.
|
||||
SECTION("ProfileGaussian clipped disks stencil") {
|
||||
CompareCpuVsGpu(IntegratorMode::ProfileGaussian, 0.005f, 4.0f, false, 120, 3.0f,
|
||||
0.0f, 0.0f, 0.0f, OverlapMode::Off, 0.0f, true);
|
||||
}
|
||||
SECTION("BoxSum clipped disks") {
|
||||
CompareCpuVsGpu(IntegratorMode::BoxSum, std::nullopt, 4.0f, false, 60, 0.0f,
|
||||
0.0f, 0.0f, 0.0f, OverlapMode::Off, 0.0f, true);
|
||||
}
|
||||
// The radial background curvature correction is computed independently in the two engines
|
||||
// (host loop vs radial_correct kernel), so it needs its own parity coverage.
|
||||
SECTION("BoxSum radial") { CompareCpuVsGpu(IntegratorMode::BoxSum, std::nullopt, 4.0f, true); }
|
||||
|
||||
Reference in New Issue
Block a user