Files
Jungfraujoch/rugnux/RugnuxCommandLine.cpp
T
leonarski_fandClaude Opus 5 09fb8e0306 Bragg integration: clip the background ring high side instead of trimming it
The r2..r3 background ring was averaged with a 10% SYMMETRIC trimmed mean. A
symmetric trim is not a consistent estimator of the mean of a right-skewed
(Poisson) sample: on a clean Poisson ring it sits ~0.1 ct/px BELOW the true
mean at every level, and with ~50 signal pixels in the r1 disk that
under-subtraction adds ~5 counts to every partial on every frame. Measured two
independent ways on four rotation datasets - stored background_mean against a
plain ring mean over the same pixels on reflection-free frames, and directly on
apertures that provably hold no reflection. Empty-aperture pedestal, counts:
plain mean -0.03..-0.20, 10% symmetric trim +5.05..+6.34, 4 sigma clip
+0.02..+0.54.

Replace it with a high-side-only sigma clip at mean + n*sqrt(mean), n = 4 for
monochromatic data. It rejects the same one-sided contamination the trim was
there for - better, in fact: a 40 px neighbour core at +100 ct shifts the trim
by +10.1 ct/px, because a symmetric trim collapses once contamination exceeds
~10% of the ring, versus +0.009 ct/px at 4 sigma. False rejection on a clean
ring is 0.04-0.39%. Broadband data keep their tuned 3 sigma clip unchanged. The
trim stays reachable with --background-trim for back compatibility; setting
either estimator clears the other, so they can never stack. --integrator boxsum
does not take the clip (matching what the shipped clip already did), so it now
uses the plain ring mean unless --background-trim is given.

The intensities get measurably more accurate: per-shell agreement with an
independent processing of the same images improves on 14 of 16 crystals
(weighted -0.0347, outermost shell 12/4), the outermost-shell R_meas NUMERATOR
- absolute scatter, not a denominator effect - falls 13.5% median on 16/5, and
CC1/2 in the outer shell improves on 14/7.

EXPECT <I/sigma> TO FALL AND EDGE R_meas TO RISE. Both are inflated by
information-free counts, so both get worse when the bias is removed; neither is
evidence against this change. That fingerprint is exactly how the trimmed mean
was accepted in the first place.

Known cost: over the 37-crystal rotation battery the de-novo space-group count
goes 34 OK / 3 DIFF to 33 / 4. The single regression is a two-lattice crystal
whose merge fails the absolute-sanity gate under either background (R_meas
63.5%, CC1/2 72.2%) and which carries an unresolved indexing ambiguity on the
very operator being scored, so its operator CC is diluted by construction. No
other crystal changes space group, and twin protection is not weakened - the
H-ratio veto that refuses genuinely twinned crystals gets MORE decisive
(1.63 -> 1.84, 2.83 -> 3.99).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 18:03:11 +02:00

195 lines
9.2 KiB
C++

// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
// SPDX-License-Identifier: GPL-3.0-only
#include "RugnuxCommandLine.h"
#include "../common/DiffractionExperiment.h"
#include <sstream>
#include <vector>
namespace {
std::string quote_if_needed(const std::string &s) {
if (s.find_first_of(" \t\"'") == std::string::npos)
return s;
std::string out = "\"";
for (char c: s) {
if (c == '"' || c == '\\')
out += '\\';
out += c;
}
out += '"';
return out;
}
const char *indexing_alg_flag(IndexingAlgorithmEnum a) {
switch (a) {
case IndexingAlgorithmEnum::FFBIDX: return "ffbidx";
case IndexingAlgorithmEnum::FFT: return "fft";
case IndexingAlgorithmEnum::FFTW: return "fftw";
case IndexingAlgorithmEnum::None: return "none";
case IndexingAlgorithmEnum::Auto:
default: return "auto";
}
}
const char *refine_flag(GeomRefinementAlgorithmEnum r) {
switch (r) {
case GeomRefinementAlgorithmEnum::None: return "none";
case GeomRefinementAlgorithmEnum::OrientationOnly: return "orientation";
case GeomRefinementAlgorithmEnum::Flex: return "flex";
case GeomRefinementAlgorithmEnum::BeamCenter:
default: return "beam_and_lattice";
}
}
std::string num(double v) {
std::ostringstream o;
o << v;
return o.str();
}
}
std::string RugnuxCommandLine(const ProcessConfig &config,
const DiffractionExperiment &experiment,
const std::string &input_file) {
std::vector<std::string> args;
const bool azint = (config.mode == ProcessMode::AzimuthalIntegration);
args.emplace_back("rugnux");
if (azint)
args.emplace_back("--azint-only");
auto add = [&](const std::string &flag, const std::string &val) {
args.push_back(flag);
args.push_back(val);
};
if (!config.output_prefix.empty())
add("-o", config.output_prefix);
add("-N", std::to_string(config.nthreads));
if (config.start_image != 0)
add("-s", std::to_string(config.start_image));
if (config.end_image >= 0)
add("-e", std::to_string(config.end_image));
if (config.stride != 1)
add("-t", std::to_string(config.stride));
if (azint) {
const auto a = experiment.GetAzimuthalIntegrationSettings();
add("--azim-min-q", num(a.GetLowQ_recipA()));
// An unset maximum Q means "to the detector edge"; emitting the resolved number would pin it
// to this run's geometry, so leave the flag out and let it resolve again.
if (const auto high_q = a.GetRequestedHighQ_recipA())
add("--azim-max-q", num(*high_q));
add("--azim-q-spacing", num(a.GetQSpacing_recipA()));
add("--azim-phi-bins", std::to_string(a.GetAzimuthalBinCount()));
add("--polarization-correction", a.IsPolarizationCorrection() ? "on" : "off");
add("--solid-angle-correction", a.IsSolidAngleCorrection() ? "on" : "off");
} else {
const auto &sf = config.spot_finding;
add("--spot-sigma", num(sf.signal_to_noise_threshold));
add("--spot-threshold", std::to_string(sf.photon_count_threshold));
// Emit the explicit flag rather than relying on the default, so the command line reproduces
// this run even if the default changes.
args.emplace_back(sf.adaptive_threshold ? "--adaptive-spots" : "--no-adaptive-spots");
if (sf.adaptive_threshold)
add("--spot-false-pixels", num(sf.false_pixels_per_frame));
// min-pix is chosen per image unless an explicit value is given, so emit --min-pix-per-spot only
// when a fixed min-pix was selected; its absence selects the adaptive per-image path.
if (sf.min_pix_per_spot.has_value())
add("--min-pix-per-spot", std::to_string(*sf.min_pix_per_spot));
// Same for the spot-finding limit: absent means "as far as the detector reaches".
if (sf.high_resolution_limit.has_value())
add("--spot-high-resolution", num(*sf.high_resolution_limit));
add("--max-spots", std::to_string(experiment.GetMaxSpotCount()));
const auto idx = experiment.GetIndexingSettings();
add("-X", indexing_alg_flag(idx.GetAlgorithm()));
add("-r", refine_flag(idx.GetGeomRefinementAlgorithm()));
if (const auto sg = experiment.GetSpaceGroupNumber())
add("-S", std::to_string(*sg));
if (const auto uc = experiment.GetUnitCell()) {
std::ostringstream o;
o << uc->a << "," << uc->b << "," << uc->c << "," << uc->alpha << "," << uc->beta << "," << uc->gamma;
add("-C", o.str());
}
if (const auto bw = experiment.GetBandwidthFWHM())
add("--bandwidth", num(*bw));
const auto bragg = experiment.GetBraggIntegrationSettings();
std::ostringstream radii;
radii << bragg.GetR1() << "," << bragg.GetR2() << "," << bragg.GetR3();
add("--integration-radius", radii.str());
// Background ring: the CLI defaults to the 4 sigma high-side clip, so emit a flag only when the
// GUI chose otherwise. The trim is the alternative estimator (it clears the clip), and a clip of
// 0 with no trim is the plain ring mean, which needs the flag to be reproduced.
if (bragg.GetBackgroundTrimFraction() > 0.0f)
add("--background-trim", num(bragg.GetBackgroundTrimFraction()));
else if (bragg.GetBackgroundClipNSigma() != BraggIntegrationSettings().GetBackgroundClipNSigma())
add("--background-clip", num(bragg.GetBackgroundClipNSigma()));
if (const auto max_hkl = bragg.GetMaxHKL())
add("--max-hkl", std::to_string(*max_hkl));
// Unset means "to the detector edge"; emitting the resolved number would pin it to this run's
// geometry, so leave the flag out and let it resolve again.
if (const auto d_min = bragg.GetDMinLimit_A())
add("--integration-high-resolution", num(*d_min));
if (config.rotation_indexing) {
if (config.two_pass_rotation)
// -R takes an optional argument, which getopt only accepts attached (-R100), never as a
// separate token - so emit it joined or the copied command line will not re-parse.
args.push_back("-R" + std::to_string(config.rotation_indexing_image_count));
else
args.emplace_back("--single-pass-rotation");
if (!config.reuse_rotation_spots)
args.emplace_back("--redo-rotation-spots");
} else if (experiment.GetGoniometer().has_value()) {
// rotation dataset processed as stills -> the user overrode the default with --force-still
args.emplace_back("--force-still");
}
// Stills geometry-refinement two-pass (--refine-geometry). getopt takes its optional argument
// only when attached (=N), never as a separate token, so emit it joined. The CLI defaults it ON
// for a stills-with-cell run, so emit =off when the GUI turned it off in that same case, to
// reproduce the GUI's choice in the copied command line.
const bool stills_with_cell = !config.rotation_indexing && experiment.GetUnitCell().has_value();
if (config.refine_geometry.has_value())
args.push_back("--refine-geometry=" + std::to_string(*config.refine_geometry));
else if (stills_with_cell)
args.emplace_back("--refine-geometry=off");
// Rotation two-pass geometry post-refine. The CLI defaults it ON for a rotation run, so emit the
// disable flag only when the GUI turned it off on a rotation dataset (to reproduce that choice).
if (config.rotation_indexing && !config.rotation_postrefine_geometry)
args.emplace_back("--rotation-no-postrefine");
// Merging is on by default; emit --no-merge only when it was turned off.
if (config.run_scaling) {
const auto sc = experiment.GetScalingSettings();
if (!sc.GetMergeFriedel())
args.emplace_back("-A");
if (!sc.GetStillsPartialityRefine())
args.emplace_back("--simple-stills");
if (!sc.GetExpectedVarianceMerge())
args.emplace_back("--no-expected-variance-merge");
// When merging, the CLI skips the large _process.h5 unless asked; emit the flag when it is
// wanted so a copied command matches the GUI's "Save _process.h5" choice. (write_merged has
// no CLI equivalent - the CLI always writes the .mtz/.cif when merging.)
if (config.write_process_h5)
args.emplace_back("--write-process-h5");
} else {
args.emplace_back("--no-merge");
}
}
args.push_back(input_file);
std::ostringstream cmd;
for (size_t i = 0; i < args.size(); i++) {
if (i)
cmd << ' ';
cmd << quote_if_needed(args[i]);
}
return cmd.str();
}