The crystal refinement (XtalOptimizer, both the seven-block and the reduced beam+orientation form, and XtalOptimizerRotationOnly) no longer builds a ceres::Problem. XtalRefine holds the problem as data and solves it with LMSolver, which follows Ceres' trust-region LM step for step - Jacobi scaling, damping and radius updates, stopping rules, box projection, the projected Armijo line search with cubic interpolation on bounded problems, the SphereManifold for the spindle - but takes J^T J and J^T r directly instead of a Jacobian. The residual is the same XtalResidual code, now Ceres-free and evaluated on a forward-mode Dual (Dual.h); everything that depends on parameters alone (detector-angle trig, per-frame back-rotation, reciprocal basis, orientation rotation) is worked out once per evaluation, and the observed and predicted halves carry 6 and 9 derivative lanes rather than 16. The sums are cut into blocks that depend on the residual count alone, so the answer does not depend on the thread count. Because the line-search trial point is the candidate point, a bounded iteration costs one evaluation instead of Ceres' three. Validation (rc174 + this, -march=x86-64-v3): - p.mtz md5 identical to the Ceres build on myob/cytc/thau x10sa, GPU and CPU builds, and on the lyso8 stills reference. - Solve corpus (every 16-parameter solve and every 10th per-image solve of the three sets, 8.3k problems, inputs and Ceres results dumped from a run that reproduced the md5s): usable/failed agree on all, iteration counts identical on all, parameters agree to <2e-11 (in px / rad / 0.01 A units), costs to 1e-13. - Same process, same threads: 7-9x faster per solve than Ceres. - In-run (GPU, loaded box): xtal 16-parameter solves myob 22.2 -> 6.4 core-s, cytc 88 -> 26 core-s; per-image solves 5.4 -> 1.2 core-s (myob); cytc first pass indexing windows 1.9 -> 0.85 s, myob 1.1 -> 0.45 s; solver share of the whole cytc run 17% -> 4% of CPU samples. New tests compare the solver with Ceres on synthetic rotation problems (full/weighted/reduced) and check thread-count independence. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
79 lines
3.7 KiB
C++
79 lines
3.7 KiB
C++
// SPDX-FileCopyrightText: 2025 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#pragma once
|
|
|
|
#include <optional>
|
|
#include <span>
|
|
|
|
#include "../common/GoniometerAxis.h"
|
|
#include "../common/CrystalLattice.h"
|
|
#include "../common/DiffractionGeometry.h"
|
|
#include "../common/SpotToSave.h"
|
|
#include "gemmi/symmetry.hpp"
|
|
|
|
struct XtalOptimizerData {
|
|
DiffractionGeometry geom;
|
|
CrystalLattice latt;
|
|
gemmi::CrystalSystem crystal_system = gemmi::CrystalSystem::Triclinic;
|
|
int64_t min_spots = 8;
|
|
|
|
float min_length_A = 5.0;
|
|
float max_length_A = 500.0;
|
|
float min_angle_deg = 60.0f;
|
|
float max_angle_deg = 120.0f;
|
|
|
|
bool refine_beam_center = true;
|
|
bool refine_detector_angles = false;
|
|
bool refine_unit_cell = true; // This refines unit cell size + angles - orientation is always refined
|
|
bool refine_rotation_axis = false;
|
|
|
|
bool index_ice_rings = true;
|
|
|
|
// Weight each spot by how strong it is for its resolution, so that low-confidence spots contribute
|
|
// without driving the fit (see SpotConfidenceWeights). Off by default: the indexers call this with a
|
|
// spot list they have already selected, it is the per-image refinement that gets the raw list.
|
|
bool weight_spots_by_confidence = false;
|
|
|
|
// Stopping rule. max_iterations > 0 bounds the solver by ITERATIONS, which is reproducible;
|
|
// otherwise it is bounded by max_time, wall-clock seconds, which is not - the same image refines
|
|
// to a different answer on a busier machine. Online acquisition needs the wall-clock bound because
|
|
// its budget is real; offline reprocessing wants the reproducible one.
|
|
float max_time = 1.0;
|
|
int max_iterations = 0;
|
|
|
|
std::optional<GoniometerAxis> axis;
|
|
|
|
// The rocking geometry, for the acceptance gate's dead zone alone - NOT for back-rotation. A
|
|
// refinement that holds a single frame does not back-rotate (the frame's angle is a gauge its
|
|
// orientation block absorbs) and so passes no `axis`, but the spots on that frame still
|
|
// diffracted at different angles spanning the exposure, and the gate still has to be told it
|
|
// does not know which. Left unset - or a wedge below the coarse-slicing trigger - the gate is
|
|
// the plain fractional-index test.
|
|
std::optional<Coord> rocking_spindle;
|
|
float rocking_wedge_deg = 0.0f;
|
|
|
|
// output
|
|
std::optional<double> beam_corr_x;
|
|
std::optional<double> beam_corr_y;
|
|
|
|
// For rotation only optimizer
|
|
std::optional<double> angle_corr;
|
|
std::optional<Coord> angle_axis;
|
|
};
|
|
|
|
// num_threads sets the thread count of the internal least-squares refine (the answer does not depend on it). It defaults
|
|
// to 1 because XtalOptimizer is usually called from many threads at once; raise it only when a caller
|
|
// runs a small number of refinements concurrently and wants each to use several cores.
|
|
bool XtalOptimizer(XtalOptimizerData &data, std::span<const std::vector<SpotToSave>> spots,
|
|
int num_threads = 1);
|
|
// Single frame. Not the same as passing {spots} to the overload above: a braced list copies the spot
|
|
// list, its elements being const, which on the per-image path is the whole list once per image.
|
|
bool XtalOptimizer(XtalOptimizerData &data, const std::vector<SpotToSave> &spots, int num_threads = 1);
|
|
bool XtalOptimizerRotationOnly(XtalOptimizerData &data, const std::vector<SpotToSave> &spots, float tolerance);
|
|
|
|
// The widest of the three fractional-index gates XtalOptimizer fits through. Its first pass selects
|
|
// spots on this one and its last pass on 0.1, so the spots between the two are carried into the fit
|
|
// and then dropped from it: this is the population a converged solve does NOT optimise.
|
|
constexpr float XTAL_OPTIMIZER_WIDE_TOLERANCE = 0.3f;
|