Three related changes to the FFT candidate path, batteried together because they touch the same function. A COPLANAR CANDIDATE REACHED REFINEMENT. ReduceResults filtered triples on lengths and angles only - the 30-150 degree bound admits any flat combination - and there was no volume test. On one dataset 41 of 5535 candidates had |V|/abc below 0.05, with a clean decade gap to the next, and three of them reached the optimizer. UnitCell is float, and for a cell that flat the metric determinant is around 1.5e-7, so float32 gets its sign wrong 19% of the time where float64 never does. The guard against a negative argument to sqrt then CREATES the singularity it was meant to prevent: it puts c in the a-b plane, the reciprocal volume is 1/0, and the residual is 0 times infinity. Ceres reported a not-a-number Jacobian and wrote several hundred lines of solver output per failed solve. VolumeFraction() is |V|/(|a||b||c|), rejected below 0.02 - about 1.1 degrees off flat, ten times below the flattest real candidate observed and a thousand times above where float loses the sign. It is enforced at the producer and at the two optimizer entry points. Note the existing sanity checks use ABSOLUTE volume, which a 320 cubic-angstrom flat cell passes. The same reciprocal-volume division is now guarded at the two remaining sites that share the pattern. A SHORTLIST CONFINED TO ONE PLANE cannot close a cell, and the row it is missing is the plane normal. That is detected from the scatter-matrix eigenvalue ratio - measured, degenerate clouds score 2e-5 to 3.3e-4 against 0.026 or more for every non-degenerate one, a factor of eighty - and one further transform is spent with the same direction count inside a three-degree cap about the normal, so the plan and buffers are untouched. More directions cannot substitute: at the exact true direction the long axis ranks 1422 of 16384 by prominence while the shortlist cut is four times higher. Ranking, not sampling, is the obstacle. A four-fold denser grid was measured and rejected - it reaches the same answer to three decimal places and takes a run from 2.5 to 8 GB of device memory. fft_min_unit_cell_A is reachable as --fft-min-unit-cell and is lowered automatically by -C, mirroring how the maximum is already raised. The default of 10 is unchanged: a lower floor admits spurious sub-cells on protein data, and over 73 protein runs the floor was never lowered while the sibling maximum did fire twice, so the path is live and correctly inert. Corpus of 93 datasets, both arms, one build: 72 bit-identical on report content and p.hkl checksum, 13 failing identically, and the count of working datasets rises by one. The volume guard fires on 58 of 93 and 47 of those stay bit-identical - it fires constantly and almost never changes an answer, which is what it should do. Solver chatter falls from 919 lines across three datasets to none. The cap fires on 4 of 93, none of them in the in-house or private arms. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Lc5JG6kJqZoCWaoZ43JGTW
53 lines
1.5 KiB
C++
53 lines
1.5 KiB
C++
// SPDX-FileCopyrightText: 2025 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
// This include should be only included in sections of the code, where it is certain that CUDA is present
|
|
// so with JFJOCH_USE_CUDA preprocessor definition, given this file is included in the source only in this case
|
|
|
|
#include <vector>
|
|
#include <mutex>
|
|
#include <optional>
|
|
#include "CUDAMemHelpers.h"
|
|
#include "../../common/Coord.h"
|
|
#include "../../common/CrystalLattice.h"
|
|
#include "FFTIndexer.h"
|
|
#include "../common/IndexingSettings.h"
|
|
#include "FFTResult.h"
|
|
|
|
class FFTIndexerGPU : public FFTIndexer {
|
|
CudaDevicePtr<float> d_dir_x;
|
|
CudaDevicePtr<float> d_dir_y;
|
|
CudaDevicePtr<float> d_dir_z;
|
|
|
|
CudaDevicePtr<float> d_spot_x;
|
|
CudaDevicePtr<float> d_spot_y;
|
|
CudaDevicePtr<float> d_spot_z;
|
|
|
|
CudaHostPtr<float> spot_x;
|
|
CudaHostPtr<float> spot_y;
|
|
CudaHostPtr<float> spot_z;
|
|
|
|
CudaDevicePtr<float> d_input_fft;
|
|
CudaDevicePtr<cufftComplex> d_output_fft;
|
|
CudaDevicePtr<FFTResult> d_result_fft;
|
|
|
|
CudaRegisteredVector<FFTResult> result_fft_reg;
|
|
|
|
CudaFFTPlan plan;
|
|
|
|
CudaStream stream;
|
|
|
|
void ExecuteFFT(const std::vector<Coord> &coord, size_t nspots) override;
|
|
void DirectionsChanged() override;
|
|
public:
|
|
explicit FFTIndexerGPU(const IndexingSettings& settings);
|
|
FFTIndexerGPU(const FFTIndexerGPU &i) = delete;
|
|
const FFTIndexerGPU &operator=(const FFTIndexerGPU &i) = delete;
|
|
~FFTIndexerGPU() override = default;
|
|
};
|
|
|
|
|
|
|