Files
Jungfraujoch/image_analysis/indexing/FFTIndexerCPU.cpp
T
leonarski_f 67dca388bd
Build Packages / Unit tests (push) Skipped
Build Packages / build:windows:cuda (push) Successful in 18m44s
Build Packages / build:viewer-tgz:cpu (push) Successful in 6m11s
Build Packages / build:viewer-tgz:cuda (push) Successful in 6m54s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 9m40s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 10m41s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 10m10s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 10m4s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 11m5s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 12m23s
Build Packages / build:rpm (rocky8) (push) Successful in 11m30s
Build Packages / build:rpm (rocky9) (push) Successful in 12m51s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 12m8s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 11m21s
Build Packages / DIALS test (push) Successful in 13m22s
Build Packages / XDS test (durin plugin) (push) Successful in 9m2s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 7m55s
Build Packages / XDS test (neggia plugin) (push) Successful in 5m57s
Build Packages / Generate python client (push) Successful in 23s
Build Packages / Build documentation (push) Successful in 57s
Build Packages / Create release (push) Skipped
Build Packages / build:windows:nocuda (push) Successful in 10m24s
v1.0.0-rc.160 (#70)
This is an UNSTABLE release. It includes many experimental features, as well as many AI generated fixes. We recommend using rc.152 for production use.

* rugnux: Add `--model model.pdb` - score the merged data against an atomic model and compute initial maps. It reports R-work/R-free (scaling the model to the observed amplitudes with an overall scale, an anisotropic B and a flat bulk solvent - the standard few-parameter model, so a batch of maps stays directly comparable) and writes 2Fo-Fc / Fo-Fc electron-density maps (CCP4) plus a map-coefficient MTZ. The structure itself is not refined; the model is only re-fractionalised into the data cell.
* rugnux: The merged reflection output now carries French-Wilson amplitudes (|F| and its sigma) next to the intensities - MTZ `F`/`SIGF`, mmCIF `_refln.F_meas_au`, and the text HKL - computed with the correct centric/acentric Wilson prior and epsilon multiplicity, so a downstream program (e.g. phenix.refine) can refine against amplitudes. The intensity columns are unchanged.
* rugnux: R-free test-set flags are now assigned deterministically and consistently across symmetry - a Bijvoet pair I(+)/I(-) is never split between the work and free sets, and the assignment is a reproducible per-hkl hash that depends only on the reflection index, so every dataset of one crystal form gets the same ~5% free set (what a multi-dataset campaign such as PanDDA needs). On small data the fraction is floored so the test set stays large enough for a stable R-free (~500 reflections, capped at 10%); it stays flat at 5% on ordinary data. When a reference MTZ carries a `FreeR_flag` column its test set is imported instead, letting a whole campaign inherit one shared free set.
* rugnux: A reference MTZ (`--reference-mtz`) can now fix the space group and cell for rotation data too (previously rejected), without being used to scale - the rotation merge stays self-consistent. When the crystal has an indexing (merohedral) ambiguity - a lattice symmetry higher than its Laue symmetry, e.g. P3/P4/P6/C2 - the reference also resolves it: each candidate reindexing (identity plus the twin-law cosets of the metric symmetry) is scored by its intensity correlation against the reference and the data are re-merged in the best-correlating one. This is a metric-preserving relabelling of hkl (the cell is unchanged) and a no-op for a holohedral crystal such as lysozyme.
* rugnux: `--model` validation now aligns the data to the model before scoring - the observed reflections are reindexed into the model's enantiomorph when the two differ only by hand (indistinguishable from merged intensities). A merohedral indexing ambiguity is resolved against the reference MTZ when one is given (so a whole campaign shares one indexing convention); only with a model and no reference does validation fall back to fitting each candidate reindexing and keeping the lowest R-free.
* rugnux: De-novo symmetry - recover a genuine high-symmetry group whose data are imperfectly scaled. Such a merge's within-orbit chi² lands just past the self-consistency bound (each real symmetry step adds a little systematic scatter), right where a merohedral twin also lands, so the chi² ratio alone cannot separate them. The candidate is now rescued when the extra intensity-proportional systematic error it invokes stays small relative to the confirmed subgroup - a genuine symmetry step gains multiplicity without inflating the merge error model's b, whereas a twin forces non-equivalent reflections together and b balloons. Fixes cubic insulin (I23 instead of I222) with no change to any other crystal in the test battery, including the twins that must stay in their lower symmetry.
* Docs: Document the French-Wilson amplitude estimation, R-free flagging, reference-based space-group/ambiguity resolution, and model-based validation/maps in CPU_DATA_ANALYSIS.md.
* Frontend: The status-bar pill now shows a progress bar during detector calibration (previously only during measurement), and the calibration state and its button are labelled "Calibration"/"CALIBRATE" (the internal `Pedestal` state name is unchanged for back-compatibility).Reviewed-on: #70

Co-authored-by: Filip Leonarski <filip.leonarski@psi.ch>
2026-07-19 09:39:28 +02:00

131 lines
4.7 KiB
C++

// SPDX-FileCopyrightText: 2025 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#include "FFTIndexerCPU.h"
#include <cmath>
#include <algorithm>
#include <stdexcept>
#include <cassert>
#include <mutex>
#include <vector>
static std::mutex fftw_plan_mutex;
static inline double dot_abs(const Coord& a, const Coord& b) {
return std::fabs(a.x * b.x + a.y * b.y + a.z * b.z);
}
FFTIndexerCPU::FFTIndexerCPU(const IndexingSettings& settings)
: FFTIndexer(settings) {
// Allocate host buffers
h_input_fft.resize(input_size, 0.0);
h_output_fft.resize(output_size);
// Validate allocations vs. expected sizes
if (h_input_fft.size() != input_size)
throw std::runtime_error("FFTWIndexer: input buffer size mismatch");
if (h_output_fft.size() != output_size)
throw std::runtime_error("FFTWIndexer: output buffer size mismatch");
const int H = static_cast<int>(histogram_size);
const int out_len = (H / 2) + 1;
int n[1] = { H };
{
std::unique_lock ul(fftw_plan_mutex);
plan = fftwf_plan_many_dft_r2c(
1, n, nDirections,
h_input_fft.data(), nullptr, 1, H,
reinterpret_cast<fftwf_complex*>(h_output_fft.data()), nullptr, 1, out_len,
FFTW_ESTIMATE);
}
if (!plan)
throw std::runtime_error("fftw_plan_many_dft_r2c failed");
}
FFTIndexerCPU::~FFTIndexerCPU() {
std::unique_lock ul(fftw_plan_mutex);
fftwf_destroy_plan(plan);
}
void FFTIndexerCPU::ExecuteFFT(const std::vector<Coord> &coord, size_t nspots) {
// Build histograms: one per direction
const int H = static_cast<int>(histogram_size);
const int D = nDirections;
const int out_len = (H / 2) + 1;
std::fill(h_input_fft.begin(), h_input_fft.end(), 0.0);
for (int d = 0; d < D; ++d) {
float* hist = h_input_fft.data() + static_cast<size_t>(d) * H;
for (size_t i = 0; i < nspots; i++) {
const auto& r = coord[i];
double dot = dot_abs(direction_vectors[d], r);
long long bin = static_cast<long long>(dot / histogram_spacing);
if (bin >= 0 && bin < H) {
hist[bin] += 1.0;
}
}
}
// Plan and execute batched R2C FFT with FFTW
fftwf_execute(plan);
// Post-process: pick the peak past min_length_A by PROMINENCE above a local background.
// The projected histogram has a broad low-frequency ENVELOPE (spots cluster near the
// origin) whose magnitude can exceed the true lattice peaks; a plain argmax|spec| then
// returns a short envelope vector (~10A) on weak/pink-beam frames and the real axes are
// lost. Subtracting a running-mean background of half-width BG_HALF bins removes that
// smooth envelope (it cancels to ~0) while sharp lattice peaks - fundamentals AND
// harmonics - keep their height. The prominence is also reported as the magnitude so
// FilterFFTResults ranks directions by real-peak strength, not by envelope.
const double len_coeff = 2.0 * static_cast<double>(max_length_A) / static_cast<double>(H);
// Background half-window ~15 A (in length, so it is independent of the histogram sizing);
// wide enough to span the envelope yet narrow enough not to smooth real peaks. Validated
// optimum on serial-still jet data (de-novo indexing improved markedly vs FFBIDX); a second dataset unchanged.
constexpr double BG_HALF_WIDTH_A = 15.0;
const int bg_half = std::max(1, static_cast<int>(std::lround(BG_HALF_WIDTH_A / len_coeff)));
std::vector<double> mag(out_len), pref(out_len + 1);
for (int d = 0; d < D; ++d) {
const auto* spec = h_output_fft.data() + static_cast<size_t>(d) * out_len;
pref[0] = 0.0;
for (int j = 0; j < out_len; ++j) {
mag[j] = std::hypot(spec[j][0], spec[j][1]);
pref[j + 1] = pref[j] + mag[j];
}
double best_prom = 0.0;
double best_len = -1.0;
for (int j = 0; j < out_len; ++j) {
double len = len_coeff * static_cast<double>(j);
if (len <= static_cast<double>(min_length_A)) continue;
const int lo = std::max(0, j - bg_half);
const int hi = std::min(out_len, j + bg_half + 1);
const double background = (pref[hi] - pref[lo]) / static_cast<double>(hi - lo);
const double prominence = mag[j] - background;
if (prominence > best_prom) {
best_prom = prominence;
best_len = len;
}
}
result_fft[d] = FFTResult{
.magnitude = static_cast<float>(best_prom),
.direction = d,
.length = static_cast<float>(best_len)
};
}
}