rugnux: make the tail's two P1 merges on the run's own engine again

The all-observation arm of the space-group search and the P1 cross-check
were made on a second RotationScaleMerge engine beside the run's own
(871347b7a). That engine holds a second device copy of every observation,
and on the largest sets the two no longer fit a 16 GB card: 8a1a, 8qaw and
8tyy (55-125 M partial observations) stopped with an out-of-memory error
in scaling. Measured on a quiet box the engine bought 0.7 s (cytc) and
1.0 s (thau) of tail and nothing on myob, which does not justify a memory
budget, so it is removed and both merges run on rsm in sequence, as
before 871347b7a.

Kept from 871347b7a: the GPU scaling's own non-blocking stream, the
cross-check not writing per-frame G/CC/mosaicity back (the per-image table
still describes the merge that was written), the restored scaling
iteration counts, and the anisotropy analysis beside the other analyses.

p.mtz byte-identical to before on myob/cytc/thau, GPU and CPU builds;
p_plot.txt unchanged apart from the GPU bkg column.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K5K8jvPPbmCrbqnWkddTuB
This commit is contained in:
2026-10-03 20:32:59 +02:00
co-authored by Claude Opus 5.5
parent 581e1c8ca2
commit 0f728ecabf
3 changed files with 5 additions and 101 deletions
@@ -121,8 +121,7 @@ public:
// Whether Run() writes the per-frame G / CC / mosaicity back onto the outcomes (on by default). Off
// for a merge that is not the run's answer - the P1 cross-check - so the per-image table and the
// unmerged MTZ describe the merge that was written, and so an engine merging beside another one
// does not write the outcomes they share.
// unmerged MTZ describe the merge that was written.
void SetWriteBackPerFrameScale(bool on) { write_back_per_frame_scale = on; }
// Override the high-resolution cut for the next Run() - used to gate the de-novo P1 search pass at
@@ -19,9 +19,8 @@ namespace {
constexpr int MIN_REFLECTIONS = 20;
// Every kernel and copy here is queued on the instance's own stream (Impl::stream), not on the
// legacy NULL stream: two merges made side by side - the run's and the one it makes ahead of time -
// and the image analysis of a probe pass beside them would otherwise wait for each other's work at
// every launch and every synchronisation. As in BeamCenterFFTGPU no buffer comes from the pool. A
// legacy NULL stream: the merge and the image analysis of a probe pass beside it would otherwise
// wait for each other's work at every launch and every synchronisation. As in BeamCenterFFTGPU no buffer comes from the pool. A
// pooled buffer is freed with cudaFreeAsync on the thread's allocation stream, which is not
// ordered after this one. Each entry point below waits for its own work before it returns, so no
// free here has yet overtaken a read, but that holds only by that convention, and not at all for a
+2 -96
View File
@@ -6674,28 +6674,6 @@ ProcessResult Rugnux::RunPipeline(RugnuxObserver *observer, bool write_output, b
const auto &rot_ss = experiment_.GetScalingSettings();
const bool is_rotation = experiment_.IsRotationIndexing(); // rotation indexing -> rotation scaling/merge
std::optional<RotationScaleMerge> rsm;
// How many times rsm has been ingested on this pass: a re-ingest follows a reindex or a cell
// change, after which a merge made from the first ingest is a merge of other indices.
int rsm_ingests = 0;
// Two P1 merges of this pass's integration made beside the rest of the tail, on an engine of
// their own: the all-observation arm of the space-group search (see there) and the P1
// cross-check (see where it is written). Neither reads anything the search, the in-symmetry
// merge or the analyses decide, so neither has to wait for them. The engine is ingested together
// with rsm, before any merge writes per-frame values back onto the outcomes, so the two start
// from the same state, and it writes nothing back itself. Used only while rsm was ingested once
// (rsm_ingest); otherwise both merges are made on rsm, as before.
struct P1MergesAhead {
DiffractionExperiment x;
Logger log = Logger::Buffered(); // what the engine logs
Logger all_observations_log = Logger::Buffered(); // ...up to the all-observation merge
std::optional<RotationScaleMerge> engine;
std::future<void> ingested;
std::future<RotationScaleMerge::Result> all_observations;
std::future<RotationScaleMerge::Result> crosscheck;
bool crosscheck_made = false;
int rsm_ingest = 0;
};
std::unique_ptr<P1MergesAhead> p1_ahead;
std::optional<PostRefineObservations> prepass_postrefine_obs;
// The rotation geometry post-refinement (see its call sites below for what it is for).
const auto post_refine_geometry = [&] {
@@ -6826,60 +6804,6 @@ ProcessResult Rugnux::RunPipeline(RugnuxObserver *observer, bool write_output, b
throw JFJochException(JFJochExceptionCategory::InputParameterInvalid,
"Rotation scaling/merging (RotationScaleMerge) does not support "
"wedge refinement");
// The conditions of p1_crosscheck and of the all-observation arm below that are known here;
// the rest (a pass that turns out superseded, a search that finds no point group) only means
// a merge is made and not used.
const bool all_observations_ahead = !geometry_prepass && !experiment_.GetGemmiSpaceGroup().has_value()
&& rot_ss.GetSearchMinZeta() > 0.0;
const bool crosscheck_ahead = !geometry_prepass && write_files && config_.write_merged
&& config_.write_p1_crosscheck && result.consensus_cell
&& (!experiment_.GetGemmiSpaceGroup().has_value() || indexer->GetPredictionCentring() == 'P');
if ((all_observations_ahead || crosscheck_ahead) && config_.observation_dump_path.empty()) {
p1_ahead = std::make_unique<P1MergesAhead>();
p1_ahead->x = experiment_;
p1_ahead->rsm_ingest = 1;
p1_ahead->crosscheck_made = crosscheck_ahead;
std::promise<void> ingested;
std::promise<RotationScaleMerge::Result> all_observations;
p1_ahead->ingested = ingested.get_future();
if (all_observations_ahead)
p1_ahead->all_observations = all_observations.get_future();
p1_ahead->crosscheck = std::async(std::launch::async,
[&a = *p1_ahead, &outcomes = indexer->GetIntegrationOutcome(),
cell = result.consensus_cell, iter = static_cast<int>(config_.scaling_iter),
nthreads = config_.nthreads, ingested = std::move(ingested),
all_observations = std::move(all_observations), all_observations_ahead,
crosscheck_ahead]() mutable -> RotationScaleMerge::Result {
try {
a.engine.emplace(a.x, outcomes, cell, iter, nthreads, a.log);
a.engine->SetWriteBackPerFrameScale(false);
a.engine->Ingest();
} catch (...) {
ingested.set_exception(std::current_exception());
throw;
}
ingested.set_value();
// Ingested in the group rsm was, merged in P1 - as both merges on rsm are.
a.x.SpaceGroupNumber(1);
if (all_observations_ahead) {
try {
a.engine->SetSearchMinZeta(0.0);
auto merged = a.engine->Run(/*for_search=*/true, /*full_stats=*/true,
/*measure_cc_before_corrections=*/false);
a.all_observations_log = a.log;
a.log = Logger::Buffered();
all_observations.set_value(std::move(merged));
} catch (...) {
all_observations.set_exception(std::current_exception());
throw;
}
}
if (!crosscheck_ahead)
return {};
return a.engine->Run(/*for_search=*/false, /*full_stats=*/true,
/*measure_cc_before_corrections=*/false);
});
}
// A reference MTZ is allowed for rotation: it fixes the space group / cell (on the CLI) and
// resolves the indexing ambiguity (below), but is NOT used to scale - the rotation merge stays
// self-consistent.
@@ -6887,9 +6811,6 @@ ProcessResult Rugnux::RunPipeline(RugnuxObserver *observer, bool write_output, b
static_cast<int>(config_.scaling_iter),
config_.nthreads, logger, config_.observation_dump_path);
rsm->Ingest();
++rsm_ingests;
if (p1_ahead)
p1_ahead->ingested.get();
}
// The geometry pre-pass reads the per-image reflections exactly twice more: they were just
@@ -7224,14 +7145,9 @@ ProcessResult Rugnux::RunPipeline(RugnuxObserver *observer, bool write_output, b
// all-observation arm on both twin gates and promoted to R32 by the filtered one. So the
// losing arm's refusal is logged and carried to the report below; the rule is unchanged.
if (rsm && rsm->GetSearchMinZeta() > 0.0 && !sg_search.point_group_hm.empty()) {
std::optional<RotationScaleMerge::Result> ahead;
if (p1_ahead && p1_ahead->all_observations.valid() && p1_ahead->rsm_ingest == rsm_ingests) {
ahead = p1_ahead->all_observations.get();
p1_ahead->all_observations_log.ReplayInto(logger);
}
const double zeta = rsm->GetSearchMinZeta();
rsm->SetSearchMinZeta(0.0);
auto sm_all = scale_and_merge("P1, all observations", true, false, std::move(ahead));
auto sm_all = scale_and_merge("P1, all observations", true);
rsm->SetSearchMinZeta(zeta);
merged_filtered_isa = sg_opts.merge_isa; // the filtered arm's, before it is replaced
sg_opts.merge_isa = result.error_model_isa; // this arm's own error model
@@ -7910,7 +7826,6 @@ ProcessResult Rugnux::RunPipeline(RugnuxObserver *observer, bool write_output, b
static_cast<int>(config_.scaling_iter),
config_.nthreads, logger, config_.observation_dump_path);
rsm->Ingest();
++rsm_ingests;
}
const auto &uc = *result.consensus_cell;
logger.Info("{} names a cell of {}x the volume of its reference setting: the "
@@ -8106,7 +8021,6 @@ ProcessResult Rugnux::RunPipeline(RugnuxObserver *observer, bool write_output, b
result.consensus_cell, static_cast<int>(config_.scaling_iter),
config_.nthreads, logger, config_.observation_dump_path);
rsm->Ingest();
++rsm_ingests;
}
}
}
@@ -8134,7 +8048,6 @@ ProcessResult Rugnux::RunPipeline(RugnuxObserver *observer, bool write_output, b
static_cast<int>(config_.scaling_iter),
config_.nthreads, logger, config_.observation_dump_path);
rsm->Ingest();
++rsm_ingests;
}
}
experiment_.SetSpaceGroup(sg);
@@ -8468,7 +8381,6 @@ ProcessResult Rugnux::RunPipeline(RugnuxObserver *observer, bool write_output, b
static_cast<int>(config_.scaling_iter),
config_.nthreads, logger, config_.observation_dump_path);
rsm->Ingest();
++rsm_ingests;
phase("Re-merging in the reference's frame");
sm = scale_and_merge(moved->short_name(), false);
const auto &uc = *result.consensus_cell;
@@ -8964,8 +8876,6 @@ ProcessResult Rugnux::RunPipeline(RugnuxObserver *observer, bool write_output, b
&& !geometry_prepass && !superseded && config_.write_p1_crosscheck
&& p1_integration_complete && is_rotation;
std::optional<RotationScaleMerge::Result> p1_merged_early;
const bool p1_ahead_usable = p1_crosscheck && p1_ahead && p1_ahead->crosscheck_made
&& p1_ahead->rsm_ingest == rsm_ingests;
// Model validation runs BEFORE the reflection files are written, because it is what settles the
// frame they are written in: the enantiomorph, which merged intensities cannot choose, and -
@@ -9000,7 +8910,7 @@ ProcessResult Rugnux::RunPipeline(RugnuxObserver *observer, bool write_output, b
// P1 merge afterwards, through merge_to_written, as it did when it was made below. The
// validation's log lines are held and printed as one block once it is done.
ModelValidationResult validation;
if (p1_crosscheck && rsm && !p1_ahead_usable) {
if (p1_crosscheck && rsm) {
Logger held = Logger::Buffered();
auto pending = std::async(std::launch::async, validate, std::ref(held));
experiment_.SpaceGroupNumber(1);
@@ -9202,10 +9112,6 @@ ProcessResult Rugnux::RunPipeline(RugnuxObserver *observer, bool write_output, b
// Both the merge and the MTZ read the group from the experiment, so it is set for
// the whole of it and restored after.
experiment_.SpaceGroupNumber(1);
if (p1_ahead_usable) {
p1_merged_early = p1_ahead->crosscheck.get();
p1_ahead->log.ReplayInto(logger);
}
if (!p1_merged_early)
rsm->SetWriteBackPerFrameScale(false);
auto p1 = scale_and_merge("P1 cross-check", false, false, std::move(p1_merged_early));