From 342ef5f66ae6e546b3bd6d0a057bebf66ba767be Mon Sep 17 00:00:00 2001 From: Filip Leonarski Date: Sat, 19 Sep 2026 18:32:15 +0200 Subject: [PATCH] tools/battery: merge quality is CC1/2 against XDS only The owner's scoring decision: no R_meas criterion, and merge quality judged only by CC1/2 relative to XDS. - The global rule (R_meas > 60% or CC1/2 < 0.5 fails) is gone. - XDS arms: a set fails `merge` when (1/CC1/2 - 1) of the reference-range table is more than twice XDS's, XDS's CC1/2 taken at the bottom of its 0.1% rounding. CC1/2 = S/(S+E), so 1/CC1/2 - 1 = E/S at any CC1/2, and E goes as 1/observations: 2x is XDS's merge with half its observations. Stored as cc_half_noise_ratio. No REFRES table or no XDS CC1/2: no criterion. - Open arm: no merge criterion; CC1/2 and R_meas stay reported numbers. - Low-resolution R_meas is reported, not scored: the lowest shell of the own table, the reference-range table and XDS's (lowres_*), with lowres_r_meas_ratio in the like-for-like table and a ratio plot. - Every schema-3 run is re-scored when it is read (report, compare, baseline delta) with today's scorer and today's manifest rows, so both sides of a comparison are scored alike. results.json keeps the verdicts as scored at run time; rows without a lattice keep them. Re-scoring the rc172 reference run: inhouse 29 -> 27 pass (three new merge fails, one R_meas fail lifted), open 137 -> 138 (one R_meas fail lifted). Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_013nW6FNRP1bBJJ8pfHiByAT --- tools/battery/README.md | 50 +++++++++++++++++++++++------ tools/battery/battery.py | 2 ++ tools/battery/report.py | 65 +++++++++++++++++++++++++++++++++---- tools/battery/score.py | 69 ++++++++++++++++++++++++++++++++-------- 4 files changed, 156 insertions(+), 30 deletions(-) diff --git a/tools/battery/README.md b/tools/battery/README.md index 65d6c2005..80d33c4b0 100644 --- a/tools/battery/README.md +++ b/tools/battery/README.md @@ -15,14 +15,35 @@ with each other. | **inhouse** | standard test crystals measured at the SLS (lysozyme, thaumatin, insulin, cytochrome C, myoglobin), plus no-crystal controls | XDS, from the `CORRECT.LP` beside each dataset | `inhouse.json` (committed) | | **private** | user data | XDS, like inhouse | outside the repository; the local site config gives its path | -Scoring checks four things in order, and the first one that fails decides the verdict: did it -run, is the lattice right, is the symmetry right, is the merge usable (R_meas at most 60%, CC1/2 at -least 0.5). The lattice test compares Niggli-reduced primitive cells: the primitive volume ratio +Scoring checks these things in order, and the first one that fails decides the verdict: did it +run, is the lattice right, is the symmetry right, and on the XDS arms, is the merge as good as +XDS's. The lattice test compares Niggli-reduced primitive cells: the primitive volume ratio must be 0.95-1.05 and the reduced edges must agree within 2%. On the open arm the space group is scored with `sgequiv`. On the XDS arms only the point group is scored, because XDS never tests a screw axis. A no-crystal control (`"expect": "no_lattice"`) passes only if no lattice is reported. `score.py` has the details. +**Merge quality is judged only against XDS, and only by CC1/2.** No R_meas is a pass/fail +criterion, and no absolute CC1/2 either. On the XDS arms the set fails `merge` when + + (1/CC1/2_rugnux - 1) / (1/(CC1/2_XDS - 0.0005) - 1) > 2 + +with rugnux's CC1/2 from the reference-range table (`REFRES_CC_HALF`, XDS's own range) and XDS's +from the `CORRECT.LP` total. CC1/2 = S / (S + E), with S the signal variance and E the half-set +error variance, so 1/CC1/2 - 1 = E/S at any CC1/2, and E goes as 1/(observations): a ratio above 2 +is a merge noisier than XDS's would be with half of its observations. XDS prints CC1/2 to 0.1%, so +its value is taken at the bottom of that rounding (-0.0005), and a printed 100.0 does not demand a +perfect CC1/2. The ratio is stored as `cc_half_noise_ratio`. A set without the reference-range +table or without an XDS CC1/2 (the controls, a hand-measured reference) has no merge criterion. +**The open arm has no merge criterion at all**: it is scored on lattice and symmetry, and its +resolution, CC1/2 and R_meas are reported, not judged. + +**Low-resolution R_meas** is reported, not scored: the lowest-resolution shell of rugnux's own +table (`lowres_r_meas`), of the reference-range table (`refres_lowres_r_meas`) and of XDS's +`CORRECT.LP` (`lowres_r_meas_ref`, stored in the manifest as `r_meas_low`), with the shells' +high-resolution limits. The reference-range table uses XDS's shell limits, so its lowest shell and +XDS's cover the same reflections; `lowres_r_meas_ratio` is the one over the other, plotted per set. + ## One command per arm Every set runs **once**, with one command per arm: @@ -52,9 +73,9 @@ the reference: - **`--report-resolution`** (XDS arms): a second statistics table from the same merge, binned over XDS's own range (`dmin_xds` and XDS's d_max as given, 999 included, so it covers exactly what XDS's CORRECT.LP totals cover), and `-A` where XDS merged with `FRIEDEL'S_LAW=FALSE` (`"anomalous": true` in the - manifest), so the two tables count the same way. Its REFRES_* numbers are scored against XDS's - CORRECT.LP totals: completeness, multiplicity, R_meas, CC1/2 and ISa (I/sigma is shown; the - manifest holds no XDS value for it). Where rugnux's own cut is coarser than the reference + manifest), so the two tables count the same way. Its REFRES_* numbers are shown beside XDS's + CORRECT.LP totals: completeness, multiplicity, R_meas, the lowest shell's R_meas, CC1/2 and ISa + (I/sigma is shown; the manifest holds no XDS value for it). Only CC1/2 is scored (above). Where rugnux's own cut is coarser than the reference (`REFRES_SHELLS_PAST_LIMIT` > 0), the shells past it are not merged, so the completeness is read as **coverage** of the reference range, not as a quality regression, and the other numbers are over the shells rugnux reached. @@ -211,6 +232,13 @@ This is a flat list with one object per set, and every arm uses the same keys. M `null`. `manifest.json` records the schema version as `results_schema`, and each arm's command as `command_doc`. +**Verdicts are scored again whenever a run is read** (`report`, `compare`, the baseline delta): +every row whose `p_report.txt` has a lattice (or that is a control) goes through today's +`score.judge` with today's manifest row (reference, `unscored`, `expect`), so the two sides of a +comparison are always scored by the same rules against the same references. A row whose report +has no lattice keeps its stored verdict, because why it failed came from the run. `results.json` +keeps the verdicts as they were scored when the run finished. + Schema 3 is one row per set again: `variant` and `first_read` are gone, and `refres_*`, `isa_ratio`, `r_meas_ratio`, `multiplicity_ref`, `cc_model`, `sg_label`, `model` and `model_note` are new. `sgno`/`sg` are now the data's own space group where `--model` relabelled the hand. @@ -228,7 +256,7 @@ was run `--unforced`), so its forced XDS-arm rows are left out the same way. |---|---| | `set`, `arm`, `tags`, `input`, `cmd` | the set's id, arm, population tags, input file, and the exact command that ran | | `verdict` | `pass`, `fail`, `unscored` (no reference) or `not_run` (input missing) | -| `cause` | why it failed: `crash`, `reader`, `indexing`, `timeout`, `lattice_halved`, `lattice_doubled`, `lattice_other`, `sym_under`, `sym_over`, `sym_screw`, `sym_other`, `merge`, `false_lattice`; or `no_reference` / `no_input` | +| `cause` | why it failed: `crash`, `reader`, `indexing`, `timeout`, `lattice_halved`, `lattice_doubled`, `lattice_other`, `sym_under`, `sym_over`, `sym_screw`, `sym_other`, `merge` (XDS arms: CC1/2 noise ratio above 2), `false_lattice`; or `no_reference` / `no_input` / `reference_problem` | | `reason` | one line for a human | | `sgno`, `sg`, `pg` / `sgno_ref`, `sg_ref`, `pg_ref` | space group number and name, and point group: ours (the data's own determination) / the reference's | | `sg_label` | the space group rugnux reported where `--model` put the model's enantiomorph on it, else null | @@ -240,6 +268,8 @@ was run `--unforced`), so its forced XDS-arm rows are left out the same way. | `r_meas_ref`, `cc_half_ref`, `isa_ref`, `completeness_ref`, `multiplicity_ref` | XDS's overall values, over XDS's own range | | `refres_range`, `refres_shells_past_limit`, `refres_unique_reflections`, `refres_completeness`, `refres_multiplicity`, `refres_i_over_sigma`, `refres_r_meas`, `refres_cc_half`, `refres_isa` | XDS arms: the REFRES_* keys, the same merge over the reference range | | `isa_ratio`, `r_meas_ratio` | refres_isa / isa_ref and refres_r_meas / r_meas_ref | +| `cc_half_noise_ratio` | XDS arms: (1/refres_cc_half - 1) / (1/(cc_half_ref - 0.0005) - 1), the merge criterion | +| `lowres_d`, `lowres_r_meas`, `refres_lowres_d`, `refres_lowres_r_meas`, `lowres_d_ref`, `lowres_r_meas_ref`, `lowres_r_meas_ratio` | the lowest-resolution shell's high-resolution limit and R_meas: rugnux's own table, the reference-range table, XDS's; and refres_lowres_r_meas / lowres_r_meas_ref | | `model`, `model_note` | open arm: the deposited coordinates given to `--model`, or why there were none | | `rfree`, `rwork`, `cc_model`, `model_fit`, `rfree_deposited`, `rfree_ratio` | open arm with a model: R_FREE, R_WORK, CC_MODEL_OVERALL and MODEL_FIT as rugnux reports them (placement-only, own free set: trend fields), the published R-free, and rfree / rfree_deposited | | `refmac_rfree`, `refmac_rwork`, `refmac_rfree_depflags`, `refmac_rfree_depdata`, `refmac_rfree_ratio`, `refmac_status`, `refmac_reason` | the REFMAC check (`--model-check`, open arm) | @@ -256,8 +286,8 @@ SVG, no external files): - a table **per population** (tags such as `cubic`, `cbf`, `lysozyme`) and the distributions of the main metrics; - plots per set (HTML only; hovering a point shows the set): d_min(rugnux) / d_min(reference) - for each arm; on the XDS arms the reference-range ISa and R_meas over XDS's; on the open arm - R_free / published R_free; + for each arm; on the XDS arms the reference-range ISa, R_meas, lowest-shell R_meas and CC1/2 + noise over XDS's; on the open arm R_free / published R_free; - on the XDS arms the **like-for-like table**: the reference-range numbers beside XDS's, with the rows where rugnux's cut is coarser marked as coverage; - the **failures**, and **one row per set** with all the numbers; @@ -286,7 +316,7 @@ $B compare RUN_A RUN_B --rerun-changed # rerun the changed sets wit `compare` pairs rows by arm and set id, following the manifests' `aliases` across renames. A run from before schema 3 is compared through its `bare` rows (see the schema above). It lists every row whose verdict, space group, lattice, d_min, ISa, R_meas, CC1/2, completeness, cell, R-free, -reference-range R_meas or ISa (where both runs have them), or time moved by more than the noise +reference-range R_meas, lowest-shell R_meas or ISa (where both runs have them), or time moved by more than the noise thresholds in `report.py` (`NOISE`). Those thresholds are a first guess. `--rerun-changed` measures the noise directly: it reruns the changed sets with A's saved binary and options. A set that moves again under the same binary is noise. A set that diff --git a/tools/battery/battery.py b/tools/battery/battery.py index 27ae2c150..72ed37663 100644 --- a/tools/battery/battery.py +++ b/tools/battery/battery.py @@ -662,6 +662,8 @@ def main(): a = ap.parse_args() site = load_site(a.site) report.ALIASES = load_aliases(site) + report.MANIFESTS = {arm: {e["id"]: e for e in load_sets(site, arm)} + for arm, cfg in site["arms"].items() if os.path.exists(cfg["manifest"])} {"run": cmd_run, "list": cmd_list, "compare": cmd_compare, "report": cmd_report, "abort": cmd_abort, "discover": cmd_discover, "refs": cmd_refs, "remap": cmd_remap}[a.cmd](a, site) diff --git a/tools/battery/report.py b/tools/battery/report.py index df86c2e13..23ea0089a 100644 --- a/tools/battery/report.py +++ b/tools/battery/report.py @@ -5,6 +5,7 @@ import socket import statistics import sys +import score from render import Doc, ratio_dots, verdict_bars # A change smaller than these is treated as run-to-run noise. They are a starting point, to be @@ -19,6 +20,7 @@ NOISE = { "rfree": ("abs", 0.005), "refres_r_meas": ("rel", 0.05), "refres_isa": ("rel", 0.05), + "refres_lowres_r_meas": ("rel", 0.05), "wall_s": ("rel", 0.25), # and at least 10 s - see beyond_noise() } @@ -53,9 +55,27 @@ def load_run(run_dir, allow_incomplete=False): schema1(man, res) if schema < 3: res[:] = schema2(res) + else: + rescore(run_dir, man, res) return man, res +def rescore(run_dir, man, res): + """Score every row again from its saved report, with today's scorer and today's manifest rows + (references, `unscored`, `expect`), so a report or a compare never mixes two scorings or two + references. results.json keeps the verdicts as they were scored when the run finished. + + A row whose report has no lattice keeps its verdict: why it failed (crash, reader, timeout) + came from the run, not from the report.""" + own = {(e["arm"], e["id"]): e for e in man["sets"]} + for r in res: + e = MANIFESTS.get(r["arm"], {}).get(key(r)[1]) or own.get((r["arm"], r["set"])) + rep = score.read_report(os.path.join(run_dir, "work", r["arm"], r["set"], "p_report.txt")) + lattice = rep.get("SPACE_GROUP_NUMBER") and rep.get("UNIT_CELL_CONSTANTS") + if e and (lattice or e.get("expect") == "no_lattice"): + r.update(score.judge(dict(e, arm=r["arm"]), rep, "")) + + def schema1(man, res): """Read a run from before the variants (results schema 1) as one: it ran each set once, plain on the open arm and with XDS's settings on the XDS arms - unless --unforced, which left the @@ -114,6 +134,9 @@ RESULTS_SCHEMA = 3 # dataset was renamed still pairs with one made after ALIASES = {} +# {arm: {set id: manifest row}} of today's manifests (battery.py sets it); rescore() scores with them +MANIFESTS = {} + def key(r): return (r["arm"], ALIASES.get(r["arm"], {}).get(r["set"], r["set"])) @@ -166,7 +189,7 @@ def changes(ra, rb): if beyond_noise(name, ra.get(name), rb.get(name)): out.append(name) # the like-for-like table only where both runs have one (a pre-schema-3 run has none) - for name in ("refres_r_meas", "refres_isa"): + for name in ("refres_r_meas", "refres_isa", "refres_lowres_r_meas"): if ra.get(name) is not None and rb.get(name) is not None and beyond_noise(name, ra[name], rb[name]): out.append(name) if beyond_noise("wall_s", ra.get("wall_s"), rb.get("wall_s")): @@ -192,7 +215,7 @@ def arrow(x, y, fmt=lambda v: str(v)): def compare_table(doc, rows, only_changed=True): head = ["set", "arm", "verdict", "space group", "d_min", "ISa", "R_meas", "CC1/2", - "cell dev %", "R_free", "ref-range R_meas", "time s", "beyond noise"] + "cell dev %", "R_free", "ref-range R_meas", "ref-range low-res R_meas", "time s", "beyond noise"] out = [] for (arm, s), ra, rb, ch in rows: if only_changed and not ch: @@ -207,6 +230,8 @@ def compare_table(doc, rows, only_changed=True): arrow(ra.get("cell_dev_pct"), rb.get("cell_dev_pct"), lambda v: f"{v:.2f}"), arrow(ra.get("rfree"), rb.get("rfree"), lambda v: f"{v:.3f}"), arrow(ra.get("refres_r_meas"), rb.get("refres_r_meas"), lambda v: f"{100 * v:.1f}%"), + arrow(ra.get("refres_lowres_r_meas"), rb.get("refres_lowres_r_meas"), + lambda v: f"{100 * v:.1f}%"), arrow(ra.get("wall_s"), rb.get("wall_s"), lambda v: f"{v:.0f}"), ", ".join(ch)]) doc.table(head, out, row_class=lambda r: "fail" if "verdict" in r[-1] else "") @@ -301,6 +326,9 @@ def build(run_dir, baseline=None, allow_incomplete=False): ("cc_half", "CC1/2", "{:.3f}"), ("res_gain_pct", "res. gain %", "{:+.1f}"), ("isa_ratio", "ref-range ISa / XDS", "{:.3f}"), ("r_meas_ratio", "ref-range R_meas / XDS", "{:.3f}"), + ("lowres_r_meas", "low-res shell R_meas", None), + ("lowres_r_meas_ratio", "ref-range low-res R_meas / XDS", "{:.3f}"), + ("cc_half_noise_ratio", "ref-range CC1/2 noise / XDS", "{:.2f}"), ("rfree", "R_free (placement)", "{:.3f}"), ("rfree_ratio", "R_free ratio", "{:.3f}"), ("wall_s", "time s", "{:.0f}")): q = quartiles(r.get(name) for r in rs) @@ -322,7 +350,13 @@ def build(run_dir, baseline=None, allow_incomplete=False): ("isa_ratio", "ISa", "ISa of the reference-range table (the error model refitted on it) / " "XDS's ISa; above 1 = rugnux's error model is the better one"), ("r_meas_ratio", "R_meas", "R_meas over the reference range / XDS's R_meas; below 1 = " - "rugnux's merge is the more consistent one")): + "rugnux's merge is the more consistent one"), + ("lowres_r_meas_ratio", "low-res R_meas", "R_meas of the lowest-resolution shell of the " + "reference-range table / of XDS's table; below 1 = " + "rugnux's strong reflections agree better"), + ("cc_half_noise_ratio", "CC1/2 noise", "half-set noise over the reference range read off CC1/2 " + "(1 / CC1/2 - 1) / XDS's; above 2 fails (the merge is as " + "noisy as XDS's with half its observations)")): for arm in arms: pts = [(r["set"], r[name], r["verdict"], f"{what} ratio, {refres_note(r)}") for r in res if r["arm"] == arm and r.get(name)] @@ -363,11 +397,21 @@ def build(run_dir, baseline=None, allow_incomplete=False): f(r.get("refres_multiplicity"), "{:.1f}") + " / " + f(r.get("multiplicity_ref"), "{:.1f}"), f(r.get("refres_i_over_sigma"), "{:.1f}"), pct(r.get("refres_r_meas")) + " / " + pct(r.get("r_meas_ref")), - f(r.get("refres_cc_half"), "{:.3f}") + " / " + f(r.get("cc_half_ref"), "{:.3f}"), + lowres_cell(r), + f(r.get("refres_cc_half"), "{:.4f}") + " / " + f(r.get("cc_half_ref"), "{:.3f}"), + f(r.get("cc_half_noise_ratio"), "{:.2f}"), f(r.get("refres_isa"), "{:.1f}") + " / " + f(r.get("isa_ref"), "{:.1f}"), refres_note(r)]) doc.table(["set", "arm", "range A", "own d_min", "compl % (rugnux / XDS)", "mult", "I/sigma", - "R_meas", "CC1/2", "ISa", "reading"], rows, num=(3, 6)) + "R_meas", "low-res R_meas (to d A)", "CC1/2", "CC1/2 noise ratio", "ISa", "reading"], + rows, num=(3, 6, 10)) + doc.p("Merge quality is scored here and only here, by CC1/2 against XDS's over the same range. " + "CC1/2 = S / (S + E) (signal and half-set error variances), so 1 / CC1/2 - 1 = E / S is the " + "half-set noise; the noise ratio is rugnux's over XDS's, with XDS's CC1/2 taken at the " + "bottom of its printed rounding (-0.0005). A set fails `merge` when the ratio is above " + f"{score.CC_HALF_NOISE_LIMIT:g}: noisier than XDS's merge would be with half its " + "observations. The low-resolution R_meas is the lowest shell of each table (the shells " + "are XDS's, so both cover the same reflections); it is reported, not scored.") derived = {r["set"]: r for r in res if (r.get("d_min_ref_rule") or "xds_range") != "xds_range"} if derived: @@ -394,7 +438,7 @@ def build(run_dir, baseline=None, allow_incomplete=False): model += [("REFMAC R_free", "refmac_rfree_depflags"), ("REFMAC dep data", "refmac_rfree_depdata"), ("REFMAC ratio", "refmac_rfree_ratio")] if any(r.get("refmac_status") for r in res) else [] head = ["set", "arm", "verdict", "space group", "ref", "cell dev %", "V ratio", "d_min", "ref d_min", - "gain", "compl %", "mult", "R_meas", "CC1/2", "ISa", "ref ISa", "idx"] + \ + "gain", "compl %", "mult", "R_meas", "low-res R_meas", "CC1/2", "ISa", "ref ISa", "idx"] + \ [h for h, _ in model] + (["model fit"] if model else []) + ["time s", "note"] rows = [] for r in sorted(res, key=lambda r: (order.get(r["arm"], 9), r["set"])): @@ -404,7 +448,8 @@ def build(run_dir, baseline=None, allow_incomplete=False): f(r.get("cell_dev_pct")), f(r.get("volume_ratio"), "{:.3f}"), f(r.get("d_min")), f(r.get("d_min_ref")), f(r.get("res_gain_pct"), "{:+.1f}%"), f(r.get("completeness"), "{:.1f}"), f(r.get("multiplicity"), "{:.1f}"), - pct(r.get("r_meas")), f(r.get("cc_half"), "{:.3f}"), f(r.get("isa"), "{:.1f}"), + pct(r.get("r_meas")), pct(r.get("lowres_r_meas")), f(r.get("cc_half"), "{:.3f}"), + f(r.get("isa"), "{:.1f}"), f(r.get("isa_ref"), "{:.1f}"), pct(r.get("indexing_rate"), 0)] + [f(r.get(k), "{:.3f}") for _, k in model] + ([r.get("model_fit") or "-"] if model else []) @@ -430,6 +475,12 @@ def build(run_dir, baseline=None, allow_incomplete=False): return doc +def lowres_cell(r): + """rugnux / XDS R_meas of the lowest-resolution shell, with the shells' high-resolution limits.""" + return (pct(r.get("refres_lowres_r_meas")) + " / " + pct(r.get("lowres_r_meas_ref")) + + f" ({f(r.get('refres_lowres_d'))} / {f(r.get('lowres_d_ref'))})") + + def refres_note(r): """How to read a row of the reference-range table.""" if r.get("refres_shells_past_limit"): diff --git a/tools/battery/score.py b/tools/battery/score.py index 331722d7d..865daf681 100644 --- a/tools/battery/score.py +++ b/tools/battery/score.py @@ -1,9 +1,9 @@ """Scoring one rugnux run against its reference. One dataset, one verdict, one cause, decided in this order: did it run -> is the lattice right -> -is the symmetry right -> is the merge usable. The order matters: a halved axis also makes the -screw along it unobservable, and counting that row twice would hide the defect that cost the -reflections. +is the symmetry right -> (XDS arms only) is the merge as good as XDS's by CC1/2. The order matters: +a halved axis also makes the screw along it unobservable, and counting that row twice would hide +the defect that cost the reflections. """ import itertools import os @@ -17,13 +17,28 @@ REPORT_LINE = re.compile(r"^([A-Z_0-9]+)=\s*(.*)$") def read_report(path): - """KEY= value lines of a rugnux _report.txt.""" + """KEY= value lines of a rugnux _report.txt, plus the lowest-resolution shell of its two shell + tables: LOWRES_D / LOWRES_R_MEAS from the table over its own range, REFRES_LOWRES_D / + REFRES_LOWRES_R_MEAS from the one over the reference range (it follows REFRES_RANGE=).""" out = {} - if os.path.exists(path): - for line in open(path, errors="replace"): - m = REPORT_LINE.match(line.rstrip("\n")) - if m: - out[m.group(1)] = m.group(2).strip() + if not os.path.exists(path): + return out + header = None + for line in open(path, errors="replace"): + line = line.rstrip("\n") + m = REPORT_LINE.match(line) + if m: + out[m.group(1)] = m.group(2).strip() + continue + f = line.lstrip("#").split() + if f[:2] == ["d_min", "N_obs"]: + header = f + elif header and f and re.match(r"^\d+\.\d+$", f[0]): + prefix = "REFRES_" if "REFRES_RANGE" in out else "" + out[prefix + "LOWRES_D"] = f[0] + r_meas = f[header.index("R_meas")].rstrip("%") # "3.7%", or "-" for none + out[prefix + "LOWRES_R_MEAS"] = "" if r_meas == "-" else str(round(float(r_meas) / 100, 4)) + header = None # only the first row of each table return out @@ -74,6 +89,24 @@ def cell_dev_pct(cell, ref): for p in itertools.permutations(cell[:3])) +# XDS prints CC1/2 in % with one decimal; its value is taken at the bottom of that rounding, so a +# printed 100.0 does not demand a perfect CC1/2 of rugnux +XDS_CC_HALF_ROUNDING = 0.0005 +CC_HALF_NOISE_LIMIT = 2.0 + + +def cc_half_noise_ratio(cc, cc_xds): + """rugnux's half-set noise over XDS's, both read off CC1/2 over XDS's range; None without both. + + CC1/2 = S / (S + E), with S the variance of the signal and E that of the half-set error, so + E / S = 1 / CC1/2 - 1 at any CC1/2. E goes as 1 / (observations): a ratio of 2 + (CC_HALF_NOISE_LIMIT) is a merge as noisy as XDS's would be with half of its observations. + A CC1/2 at or below 0 is no signal at all, counted as 0.001.""" + if cc is None or cc_xds is None: + return None + return round((1 / max(cc, 0.001) - 1) / (1 / (cc_xds - XDS_CC_HALF_ROUNDING) - 1), 3) + + def point_group(sgno): return gemmi.find_spacegroup_by_number(int(sgno)).point_group_hm() @@ -115,6 +148,15 @@ def judge(entry, rep, run_note): r["isa_ratio"] = round(r["refres_isa"] / r["isa_ref"], 4) if r["refres_isa"] and r["isa_ref"] else None r["r_meas_ratio"] = (round(r["refres_r_meas"] / r["r_meas_ref"], 4) if r["refres_r_meas"] and r["r_meas_ref"] else None) + # R_meas of the lowest-resolution shell: rugnux's own table, the reference-range table and XDS's + # (reported, not scored) + r.update(lowres_d=fnum(rep.get("LOWRES_D")), lowres_r_meas=fnum(rep.get("LOWRES_R_MEAS")), + refres_lowres_d=fnum(rep.get("REFRES_LOWRES_D")), + refres_lowres_r_meas=fnum(rep.get("REFRES_LOWRES_R_MEAS")), + lowres_d_ref=ref.get("dmin_low"), lowres_r_meas_ref=ref.get("r_meas_low")) + r["lowres_r_meas_ratio"] = (round(r["refres_lowres_r_meas"] / r["lowres_r_meas_ref"], 4) + if r["refres_lowres_r_meas"] and r["lowres_r_meas_ref"] else None) + r["cc_half_noise_ratio"] = cc_half_noise_ratio(r["refres_cc_half"], r["cc_half_ref"]) if rep.get("SPACE_GROUP_NUMBER"): r["sgno"] = int(fnum(rep["SPACE_GROUP_NUMBER"])) r["sg"] = rep.get("SPACE_GROUP_NAME") or sg_name(r["sgno"]) @@ -191,10 +233,11 @@ def judge(entry, rep, run_note): "sym_over" if order > order_ref else "sym_other") return verdict("fail", cause, f"{r['sg']} vs reference {r['sg_ref']}") - if r["r_meas"] is not None and r["r_meas"] > 0.6: - return verdict("fail", "merge", f"R_meas {r['r_meas']:.0%} overall") - if r["cc_half"] is not None and r["cc_half"] < 0.5: - return verdict("fail", "merge", f"CC1/2 {r['cc_half']:.2f} overall") + # Merge quality is judged only against XDS, by CC1/2 over XDS's own range (cc_half_noise_ratio). + # The open arm has no merge criterion. + if r["cc_half_noise_ratio"] is not None and r["cc_half_noise_ratio"] > CC_HALF_NOISE_LIMIT: + return verdict("fail", "merge", f"CC1/2 {r['refres_cc_half']:.4f} vs XDS {r['cc_half_ref']:.3f} " + f"over XDS's range: {r['cc_half_noise_ratio']:.1f}x XDS's half-set noise") note = "" if r["sgno"] != ref["sgno"]: note = f"{r['sg']} vs reference {r['sg_ref']} ({r.get('sg_relation') or 'screws not judged against XDS'})"