Build Packages / build:viewer-tgz:cpu (push) Successful in 7m31s
Build Packages / build:viewer-tgz:cuda (push) Successful in 8m41s
Build Packages / build:rpm (ubuntu2404_nocuda) (push) Successful in 13m27s
Build Packages / build:rpm (rocky8_nocuda) (push) Successful in 13m46s
Build Packages / build:rpm (ubuntu2204_nocuda) (push) Successful in 13m49s
Build Packages / build:rpm (rocky8_sls9) (push) Successful in 14m4s
Build Packages / build:rpm (rocky9_nocuda) (push) Successful in 14m10s
Build Packages / build:windows:nocuda (push) Successful in 15m55s
Build Packages / build:windows:cuda (push) Successful in 18m8s
Build Packages / build:rpm (rocky8) (push) Successful in 11m29s
Build Packages / build:rpm (rocky9_sls9) (push) Successful in 12m56s
Build Packages / XDS test (durin plugin) (push) Successful in 7m50s
Build Packages / Generate python client (push) Successful in 33s
Build Packages / Build documentation (push) Successful in 1m6s
Build Packages / Create release (push) Skipped
Build Packages / build:rpm (rocky9) (push) Successful in 13m4s
Build Packages / build:rpm (ubuntu2404) (push) Successful in 12m52s
Build Packages / build:rpm (ubuntu2204) (push) Successful in 13m56s
Build Packages / DIALS test (push) Successful in 14m2s
Build Packages / XDS test (neggia plugin) (push) Successful in 8m4s
Build Packages / XDS test (JFJoch plugin) (push) Successful in 9m14s
Build Packages / Unit tests (push) Successful in 1h37m26s
AssignRfreeFlags gains a min_free_reflections floor (default 500): on small data, where the 5% fraction would give too few test reflections for a stable R-free (Brunger's ~500-2000 rule), the fraction is lifted toward ~500 free reflections, capped at 10% so a large test set never steals working data. An explicit rfree_fraction is still honoured; pass 0 to disable. For ordinary data (5% already clears the floor) this is inactive and the fraction stays flat, so the cross-dataset-identical property is preserved; the floor only lifts the fraction on genuinely small datasets, where per-dataset R-free stability outweighs cross-dataset identity (a shared reference set keeps exact identity there). Tests updated to isolate the pure-hash guarantees with the floor off, plus a new small-data floor + cap test. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
80 lines
3.7 KiB
C++
80 lines
3.7 KiB
C++
// SPDX-FileCopyrightText: 2026 Filip Leonarski, Paul Scherrer Institute <filip.leonarski@psi.ch>
|
|
// SPDX-License-Identifier: GPL-3.0-only
|
|
|
|
#include "RfreeFlags.h"
|
|
|
|
#include <algorithm>
|
|
#include <unordered_map>
|
|
#include <unordered_set>
|
|
|
|
#include "HKLKey.h"
|
|
|
|
namespace {
|
|
// splitmix64 bit-mix of a key -> uniform double in [0, 1). Same key -> same value, so all
|
|
// mates of a reflection (which share the Laue-ASU key) get the same draw. Same idiom as the
|
|
// CC1/2 half-set split (HalfForImage in Merge.cpp).
|
|
double UniformFromKey(uint64_t key) {
|
|
uint64_t z = key + 0x9e3779b97f4a7c15ULL;
|
|
z = (z ^ (z >> 30)) * 0xbf58476d1ce4e5b9ULL;
|
|
z = (z ^ (z >> 27)) * 0x94d049bb133111ebULL;
|
|
z = z ^ (z >> 31);
|
|
return static_cast<double>(z >> 11) * (1.0 / 9007199254740992.0);
|
|
}
|
|
}
|
|
|
|
void AssignRfreeFlags(std::vector<MergedReflection> &merged, int32_t space_group_number,
|
|
double rfree_fraction, int min_free_reflections) {
|
|
for (auto &r : merged)
|
|
r.rfree_flag = false;
|
|
if (rfree_fraction <= 0.0 || merged.empty())
|
|
return;
|
|
|
|
// The flag is a pure function of the Friedel-merged (Laue) ASU key: symmetry- and Friedel-
|
|
// equivalent reflections collapse to one key and so share a flag (a Bijvoet pair I(+)/I(-) is
|
|
// never split across the work and free sets), and the draw depends only on the reflection index
|
|
// - not on this dataset's resolution range or which reflections it happens to contain. So every
|
|
// dataset of one crystal form gets the SAME free set, which is what a multi-dataset campaign
|
|
// (ensemble refinement, PanDDA) needs. A uniform hash draws ~rfree_fraction of the distinct
|
|
// reflections free; a stratified per-shell draw would be tied to the dataset and break that.
|
|
const HKLKeyGenerator laue_key(/*merge_friedel=*/true, space_group_number);
|
|
|
|
// Count the distinct test-eligible reflections (distinct Laue-ASU keys; mates collapse to one) so
|
|
// the fraction can be floored to a usable test-set size on small data.
|
|
std::unordered_set<uint64_t> distinct;
|
|
distinct.reserve(merged.size());
|
|
for (const auto &r : merged)
|
|
distinct.insert(laue_key(r).pack());
|
|
|
|
// Effective fraction: at least rfree_fraction, lifted toward min_free_reflections/N on small data
|
|
// (so R-free is not sampling-noise dominated), but the floor's lift is capped at MAX_FRACTION so a
|
|
// large test set never steals working data. An explicitly large rfree_fraction is always honoured.
|
|
constexpr double MAX_FRACTION = 0.10;
|
|
const double floor_fraction =
|
|
std::min(min_free_reflections / static_cast<double>(distinct.size()), MAX_FRACTION);
|
|
const double eff_fraction = std::max(rfree_fraction, floor_fraction);
|
|
|
|
for (auto &r : merged)
|
|
r.rfree_flag = UniformFromKey(laue_key(r).pack()) < eff_fraction;
|
|
}
|
|
|
|
size_t ApplyReferenceFreeFlags(std::vector<MergedReflection> &merged, int32_t space_group_number,
|
|
const std::vector<MergedReflection> &reference) {
|
|
// Reference free/work partition keyed by the Friedel-merged (Laue) ASU index, so it transfers
|
|
// regardless of which Bijvoet mate / symmetry equivalent each dataset happens to have measured.
|
|
const HKLKeyGenerator laue_key(/*merge_friedel=*/true, space_group_number);
|
|
std::unordered_map<uint64_t, bool> ref_flag;
|
|
ref_flag.reserve(reference.size());
|
|
for (const auto &r : reference)
|
|
ref_flag[laue_key(r).pack()] = r.rfree_flag;
|
|
|
|
size_t matched = 0;
|
|
for (auto &r : merged) {
|
|
const auto it = ref_flag.find(laue_key(r).pack());
|
|
if (it != ref_flag.end()) { // reflections absent from the reference keep their hash flag
|
|
r.rfree_flag = it->second;
|
|
++matched;
|
|
}
|
|
}
|
|
return matched;
|
|
}
|