Files
Jungfraujoch/image_analysis/WriteReflections.h
T
leonarski_fandClaude Opus 5.5 59d92a7238 ParallelSort for the two large sorts on the merge path
Two whole-dataset sorts sat on the main thread at the end of a rotation run:
WilsonOutliers orders every full by resolution, and the unmerged MTZ is put
in H K L M/ISYM BATCH order by Mtz::sort(5) - together about 2 s of one
thread on a 1.4 M-observation set.

ParallelSort (common/ParallelFor.h) sorts one piece per worker and merges
them pairwise. It is only for comparators that are a strict total order,
where the sorted sequence is unique and the result is the serial sort's bit
for bit: WilsonOutliers already breaks ties on the index, and the unmerged
writer now sorts the rows itself on the five key columns and then the row
number - the order Mtz::sort's stable sort gives - and sets sort_order as it
did. An empty table still fails the way Mtz::sort does.

md5-identical p.hkl, p.mtz and p_unmerged.mtz; 44.4 -> 43.6 s on a 16M set.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1G8gJVAy6gp1K5Dz3NE5C
2026-09-26 18:13:41 +02:00

78 lines
4.1 KiB
C++

// SPDX-FileCopyrightText: 2025 Paul Scherrer Institute
// SPDX-License-Identifier: GPL-3.0-only
#pragma once
#include <string>
#include <vector>
#include "../common/Reflection.h"
#include "../common/UnitCell.h"
#include "../common/DiffractionExperiment.h"
#include "IntegrationOutcome.h"
struct MergeStatistics;
struct TwinningAnalysisResult;
// The error model as it is reported, already formatted. `isa` is the whole-range 1/sqrt(a*b), the
// same quantity XDS's ISa denotes, so a file written here is directly comparable with a CORRECT.LP;
// `isa_asymptotic` is the strong-reflection tier, which only the rotation path has. `a` and `b` are
// in XDS's convention, sigma^2 = a*(sigma0^2 + b*I^2). Empty strings are written as unknown.
struct ErrorModelReport {
std::string isa;
std::string isa_asymptotic;
std::string a;
std::string b;
};
// nthreads: workers for the per-reflection row formatting, which is the bulk of the file.
void WriteMmcifReflections(const std::vector<MergedReflection> &reflections,
const UnitCell &unitCell,
const DiffractionExperiment &experiment,
const MergeStatistics &statistics,
const ErrorModelReport &error_model,
const TwinningAnalysisResult &twinning,
const std::string &filename,
size_t nthreads);
void WriteMtzReflections(const std::vector<MergedReflection> &reflections,
const UnitCell &unitCell,
const DiffractionExperiment &experiment,
const std::string &filename);
// SHELX HKLF-4 text file (h k l I sigma(I), Bijvoet mates separate) for SHELXC / ANODE.
// nthreads: workers for the per-reflection row formatting, as for the mmCIF.
void WriteShelxHklReflections(const std::vector<MergedReflection> &reflections,
const DiffractionExperiment &experiment,
const std::string &filename,
size_t nthreads);
// Unmerged observations in the column and batch-header layout POINTLESS writes: aimless, pointless,
// careless and iotbx.merging_statistics all read that layout. H K L are the ASU indices and M/ISYM
// recovers the index the reflection was measured at (which is what careless needs to see the crystal
// frame) and says whether the observation is a partial.
// sum_partials: add the partials of each rocking event into one full, written at the batch of the
// event's centroid with the summed rocking-curve fraction in FRACTIONCALC - the plain sum every
// rotation program writes, over the events the combine would assemble (--min-partiality). False
// writes one row per integrated box instead, flagged as partials for the reading program to sum.
// Stills have no rocking events and are unaffected either way.
// The intensities carry the deterministic per-reflection corrections and nothing else - Lorentz
// and polarization in LP, the sensor's angle-dependent efficiency in QE, and the attenuation of the
// flight-path medium in FLIGHT - so raw counts are I / LP * QE * FLIGHT. The partiality and the
// per-image scale are left for the reading program to fit, since every program this file is for
// fits a scale model of its own.
void WriteUnmergedMtzReflections(const std::vector<IntegrationOutcome> &outcomes,
const UnitCell &unitCell,
const DiffractionExperiment &experiment,
bool sum_partials,
const std::string &filename,
size_t nthreads = 1);
void WriteReflections(const std::vector<MergedReflection> &reflections,
const UnitCell &unitCell,
const DiffractionExperiment &experiment,
const MergeStatistics &statistics,
const ErrorModelReport &error_model,
const TwinningAnalysisResult &twinning,
const std::string &filename,
size_t nthreads);