Files
Jungfraujoch/image_analysis/geom_refinement/XtalOptimizer.h
T
leonarski_fandClaude Opus 4.8 b1cd49c6b5 rugnux: parallelize two-pass rotation first phase
The first pass of two-pass rotation indexing was the slow "getting there" phase.
Profiling RunIndexing showed the cost is the serial Ceres candidate refinement
(up to 4 FFT candidate cells x 2 XtalOptimizer solves each), not spot finding or
the cuFFT kernel.

- XtalOptimizer: add a num_threads arg (default 1) so a caller running a few
  refinements concurrently can give each several cores.
- RotationIndexer::RunIndexing: seed the candidates serially, run their up-to-8
  XtalOptimizer solves concurrently (4 Ceres threads each), select serially in
  candidate order (identical result). ~10-13x on the candidate loop.
- Rugnux: feed both first-pass schemes (spread + wedge) serially, then run their
  RunIndexing() passes in parallel - overlapping one scheme's cuFFT with the
  other's Ceres. Size the FFT indexer pool to 2 for the rotation path (it fires
  only twice) instead of the default 4, avoiding wasted cuFFT-plan init.

First-pass wall 2.1-6.5x faster across a 12-crystal battery; whole lysoC run
17.7->10.5s. All 12 crystals bit-identical (scheme, validation count, space group).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-11 22:35:54 +02:00

54 lines
1.7 KiB
C++

// SPDX-FileCopyrightText: 2025 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <optional>
#include "../common/GoniometerAxis.h"
#include "../common/CrystalLattice.h"
#include "../common/DiffractionGeometry.h"
#include "../common/SpotToSave.h"
#include "gemmi/symmetry.hpp"
struct XtalOptimizerData {
DiffractionGeometry geom;
CrystalLattice latt;
gemmi::CrystalSystem crystal_system = gemmi::CrystalSystem::Triclinic;
int64_t min_spots = 8;
float min_length_A = 5.0;
float max_length_A = 500.0;
float min_angle_deg = 60.0f;
float max_angle_deg = 120.0f;
bool refine_beam_center = true;
bool refine_distance_mm = false;
bool refine_detector_angles = false;
bool refine_unit_cell = true; // This refines unit cell size + angles - orientation is always refined
bool refine_rotation_axis = false;
bool index_ice_rings = true;
float max_time = 1.0;
std::optional<GoniometerAxis> axis;
// output
std::optional<double> beam_corr_x;
std::optional<double> beam_corr_y;
// For rotation only optimizer
std::optional<double> angle_corr;
std::optional<Coord> angle_axis;
};
// num_threads sets the Ceres solver thread count for the internal least-squares refine. It defaults
// to 1 because XtalOptimizer is usually called from many threads at once; raise it only when a caller
// runs a small number of refinements concurrently and wants each to use several cores.
bool XtalOptimizer(XtalOptimizerData &data, const std::vector<std::vector<SpotToSave>> &spots,
int num_threads = 1);
bool XtalOptimizerRotationOnly(XtalOptimizerData &data, const std::vector<SpotToSave> &spots, float tolerance);