Digging into the selection logic showed the caps were not deciding the science - the two-pass geometry post-refinement was, and the caps only fed it randomness. Caps. The prediction buffer now grows to whatever a frame predicts instead of keeping an arbitrary subset of it, and the per-image reflection limit is raised to 65536, with the image-buffer transport headroom derived from the same constant so the two cannot drift. Measured: bit-identical output on five battery crystals, because a normal cell never approached the old limits - only a large cell (~2.8e6 A^3, ~30000-44000 predictions per frame) ever did. Pass-2 guard. The refined pass is normally the better answer, which is why it is the canonical output, but it was adopted whatever it produced. On that same crystal it merged more unique reflections than its own cell can hold - completeness "117%", which is arithmetically impossible - while the header- geometry pass sat at 92.6% and CC1/2 0.98. Compare the two and, when the refined pass is not credible, go back to the header geometry and re-run so the canonical files are the ones that are kept. Both bounds are set where only a failure reaches them. Together on that crystal: 111639 unique against XDS's 118730 (was 88000-99000 and different every run), CC1/2 98.0% (was 96.9-97.7%), ISa 8.54, and two runs now agree bit for bit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
55 lines
1.3 KiB
C++
55 lines
1.3 KiB
C++
// SPDX-FileCopyrightText: 2025 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <vector>
|
|
|
|
#include "BraggPrediction.h"
|
|
|
|
#include "../indexing/CUDAMemHelpers.h"
|
|
|
|
struct KernelConsts {
|
|
float det_width_pxl;
|
|
float det_height_pxl;
|
|
float beam_x;
|
|
float beam_y;
|
|
float coeff_const;
|
|
float one_over_wavelength;
|
|
float one_over_dmax_sq;
|
|
float ewald_cutoff;
|
|
float bandwidth_sigma; // relative Δλ/λ (sigma); 0 = monochromatic
|
|
Coord Astar, Bstar, Cstar, S0;
|
|
float rot[9];
|
|
char centering;
|
|
};
|
|
|
|
class BraggPredictionGPU : public BraggPrediction {
|
|
CudaRegisteredVector<Reflection> reg_out; // requires stable storage
|
|
CudaDevicePtr<Reflection> d_out;
|
|
|
|
// Dedicated stream
|
|
CudaStream stream;
|
|
|
|
// Device allocations via helpers
|
|
CudaDevicePtr<KernelConsts> dK;
|
|
CudaDevicePtr<int> d_count;
|
|
|
|
// Host pinned buffer for async download (optional but faster)
|
|
CudaHostPtr<int> h_count;
|
|
|
|
public:
|
|
explicit BraggPredictionGPU(int max_reflections = kPredictionCapacity);
|
|
|
|
protected:
|
|
void GrowCapacity(int count) override;
|
|
|
|
public:
|
|
|
|
int Calc(const DiffractionExperiment &experiment,
|
|
const CrystalLattice &lattice,
|
|
const BraggPredictionSettings &settings) override;
|
|
};
|
|
|
|
|