Files
Jungfraujoch/docs/CPU_DATA_ANALYSIS_DECISIONS.md
T
leonarski_fandClaude Opus 5 20f869c0b8 docs: split the data-analysis reference into four parts along the pipeline
CPU_DATA_ANALYSIS.md becomes a short landing page (scope, part map,
references) over four parts in pipeline order - images to spots (0-3),
indexing and geometry (4-7), integration/scaling/merging (8-12), space group
and validation (13-14). Pure moves: the section numbering is continuous and
unchanged, since the rest of the documentation and the source cite sections
by number. Inbound topical links now land on the right part; the build has
zero warnings and the rendered-HTML anchor check finds no dead link.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EFEJG6WBQv8th4UJFNe53N
2026-09-02 09:19:16 +02:00

29 KiB
Raw Blame History

Data analysis: space group and validation (§13–§14)

Part of the CPU/GPU data-analysis reference; the section numbers are continuous across its four parts.

:local:
:depth: 2

13. Space-group determination and merge-level decisions

13.1 Space-group determination

When no space group is supplied, a POINTLESS-like search scores Laue-group symmetry (CC of E^2(h) vs E^2(Rh) — the intensities normalised by the mean of their own resolution shell — plus merge self-consistency) and detects screw/centering absences from the $P1$-merged intensities. Three tests gate a promotion to higher symmetry, all aimed at the merohedral twin, whose twin law forces non-equivalent reflections together and so mimics symmetry:

  1. Merge self-consistency (\chi^2 under the candidate group, relative to the confirmed subgroup). On its own this is not sufficient: it is a ratio to an error model that moves with the amount of data — the parent's systematic term grows as \sigma shrinks with 1/\sqrt{N}, while a twin's is already saturated — so its verdict depends on how much data the search saw.
  2. Error-model $b$ (the intensity-proportional systematic). A genuine symmetry step gains multiplicity without inflating b; merging a twin law's extra operator inflates it. A $\chi^2$-passing promotion is vetoed when b rises past a bound relative to the confirmed subgroup.
  3. Operator disagreement, a sigma-free statistic H=\mathrm{median}\,|I_1-I_2|/(I_1+I_2), formed as the ratio of the operators a promotion adds to the parent's own, measured on the same reflections. Normalising against the parent divides out the systematic floor that symmetry mates carry on real data, which varies by crystal and by operator; a median is used because a twin perturbs every pair whereas a badly-measured minority perturbs only the tail. Where a candidate has several parents of the same order, it is judged against the worst of them, since a rival subgroup can itself contain the twin laws.

The correlation is on resolution-normalised intensity E^2 = I/\langle I\rangle(\text{shell}), normalised over exactly the reflections the correlation pairs. Both members of a symmetry pair lie at the same |s|, so on raw I the resolution fall-off is variance shared perfectly between the two arms and appears as a positive correlation for any pairing at all: a shell-matched random pairing — the exact null for a metrically-allowed false operator — scores a median 0.31 across the rotation battery, and on one crystal 0.53 — above the bound the correlation is tested against. That floor varies more from crystal to crystal (spread 0.46) than the whole true/false gap is wide (0.38), so an absolute bound on the raw statistic is a different test on every crystal; and it moves with the search resolution cut, which is what made that cut a symmetry-deciding parameter. Normalised, the floor has a median of 0.015, never exceeds 0.06, and barely moves with the cut.

Both the correlation stage and the absence tests need to know whether a reflection is genuinely present, and that question is asked of its counting significance, not of the merged I/\sigma. A merged \sigma carries the error model's intensity-proportional term, \sigma^2 = a\,\sigma_0^2 + (b\,I)^2, so merged I/\sigma saturates — at \mathrm{ISa}\sqrt{n} for a reflection observed n times, and at \mathrm{ISa} exactly for one observed once. Above that knee it stops rising with the intensity: on the weakest search merge of the rotation battery the I/\sigma of every decile of E^2 reads 1.581.65 against an \mathrm{ISa} of 1.70, so reflections an order of magnitude apart in real intensity report the same number. A single constant applied there demands anywhere between 2.2 and 11.9 in counting significance depending on the crystal.

The nominal cut T (default 3.0) is therefore converted once, using the ISa of the merge being searched. For a reflection observed once \sigma_\text{counting}^2 = \sigma^2 - (b\,I)^2, so I/\sigma_\text{counting} \ge T is exactly

\frac{I}{\sigma} \;\ge\; \frac{T}{\sqrt{1 + (T/\mathrm{ISa})^2}}

The converted cut lies strictly below \mathrm{ISa} for every \mathrm{ISa}, so it is always reachable by the reflection the ceiling binds hardest, and it is within 1 % of T on any merge with \mathrm{ISa} \ge 21 — a healthy merge is left exactly where it was. Multiplicity is taken as 1 deliberately rather than estimated, for the same reason. Where the merge reports no ISa (or b = 0) the cut is used as it stands.

The candidates are enumerated in every setting the refined cell can host, not only in the settings the International Tables call the reference one. A setting is a statement about direction — P112_1 puts its 2_1 on c where P12_11 puts it on b, and both are space group 4 — so a search restricted to reference settings can only ever put a screw or a centring on the axis the convention chose, whatever the data say. Two things bound the widened set. A candidate is offered only if the cell's own metric admits the rotations its setting names (\alpha=\beta=90^\circ for a $2$-fold on c), and a candidate predicting exactly the absences another candidate already predicts is dropped as the same hypothesis under a second name. A non-reference setting is additionally refused when its centring class holds no reflection in this merge, since there is then nothing to confirm it with. Because the candidates of a point group all share its rotations, this cannot raise the symmetry: it decides which axes carry the screws and the centring, never how many operators there are. Where a determined group is not the reference setting, the report and the reflection files name it by its full HermannMauguin symbol, and SPACE_GROUP_NAME in the run report carries it beside the number, which alone would be ambiguous.

Several space groups may share an absence pattern exactly. Where they do, the search scores them identically and all of them are named in the result rather than one being reported as the answer: some are enantiomorph pairs, which merged intensities cannot distinguish in principle, and others differ only by a screw condition that the centering condition already implies, so the screw has no observable signature at all. The representative reported first is the lowest space-group number, which is a convention and not a measurement.

The Lorentz factor \zeta (§8.3) governs how well a reflection can be measured, so when the spindle lies in a plane of the lattice, an operator permuting the two in-plane axes samples a different mixture of measurement qualities than one that only flips signs. The search is therefore run a second time on a merge of only the well-measured observations (--search-min-zeta, rotation default 0.85), both answers are reported, and where they disagree the merge of all the observations decides. The filter discards 4080 % of the observations, which can starve an operator correlation the full merge confirms and can equally leave an operator confirmed that the full merge refuses, so the decision — the point group as well as the absences, which live in the weak reflections the filter removes — rests on the arm with every observation behind it. A tie (same order, different symmetry) is reported with both candidates named, for trying in molecular replacement.

Both search merges also drop the frames whose fitted per-frame scale came out below a tenth of the run median — the stretches where the crystal was barely in the beam. The scale enters as 1/G, so such a frame's intensities arrive amplified tenfold or more, with their \sigma amplified by the identical factor and the scale's own error nowhere in it; the production merge survives that because a reflection in the determined symmetry is measured ten or twenty times and --reject-outliers removes the amplified observation, but the search merges in P1, where a reflection has two or three observations and no majority exists to call any of them an outlier. On a crystal that repeatedly left the beam over a full turn this cost the 422 point group outright, its operators reading $0.08$0.38 on a merge that gives \mathrm{CC}_{1/2} = 96\,\% in that same point group once it is assumed. Like the \zeta filter, this is a filter of the search passes alone — the production merge keeps every frame.

Screw axes are scored per axial zone, not pooled over them. The conditions on h00, 0k0 and 00l are independent, so their log-likelihoods add, and a zone that was never measured is allowed to abstain rather than to argue: pooling let one unmeasured row veto a confirmed one, and it read three genuine screws as weaker than two whenever the third row was shallow.

Centering is accepted when the systematically-absent class is weak relative to the present one by either of two floor-independent tests: its mean signed I/\sigma well below the present mean, or its rate of individually-significant reflections well below the present class's own significant rate. The second test covers weak and low-energy data, where a positive intensity floor (background and profile leakage) lifts the absent class's mean I/\sigma well above zero and, when the present class is itself weak, carries the plain mean ratio past its bound; a false centering fails both tests, its absent class being as strong as the present one. When several centerings pass they are ranked by a Beta-tail likelihood rather than by a count of net absences: a count lets a centering that extinguishes many more reflections out-rank one that violates far fewer, and the likelihood weighs the violations against the class each belongs to instead. The acceptance bounds above are unchanged by that ranking.

13.2 Twinning check

A PadillaYeates $L$-test (\langle|L|\rangle, \langle L^2\rangle — 0.500 and 0.333 untwinned, 0.375 and 0.200 for a perfect twin) and the second moment \langle I^2\rangle/\langle I\rangle^2 (2.0 for untwinned acentric data, 1.5 for a perfect twin; taken per resolution shell with noise-only shells skipped and Wilson outliers rejected, so a single strong reflection in a collapsed-mean shell cannot skew it) are written to the merged mmCIF as a twinning diagnostic. A merohedral twin law exists only where the Laue class is a proper subgroup of the lattice holohedry, so in the high-symmetry holohedral classes (4/mmm, 6/mmm, m\bar{3}m, and \bar{3}m on a rhombohedral lattice) twinning is never flagged, and a low \langle|L|\rangle there is reported as a statistical artefact instead. The low-symmetry holohedral classes (\bar{1}, 2/m, mmm) also admit no strictly merohedral law, but they stay eligible for the flag on purpose: pseudo-merohedral twinning through an accidentally special metric cannot be ruled out from the symmetry alone, and those are the classes it happens in.

13.3 Outlier rejection

Merging applies an optional per-observation median-based N\sigma cut (--reject-outliers, default 6σ for rot3d, off otherwise). The same N\sigma cut is fed back into the error model: after an initial a,b fit the parameters are re-fit once on the reflections that survive rejection (dropping any whose squared deviation exceeds N^2\,[a\,\sigma^2 + (b\,\langle I\rangle)^2]), so the calibrated errors describe the reflections that actually enter the merge rather than the pre-rejection pool.

13.4 Automatic resolution cutoff

By default the reported/written high-resolution limit is trimmed where \mathrm{CC}_{1/2} falls off: a logistic is fitted to \mathrm{CC}_{1/2}(s), and the limit is set one reported-shell width past the point where the fit crosses 0.30 — deliberately "one shell too far", so weak-but-real data below the crossing are kept rather than discarded. The extension is measured over the range that is actually kept, not the full measured range, so a detector reaching far past where the crystal diffracts cannot inflate it. --scaling-high-resolution overrides the limit and --resolution-cutoff off disables it.

13.5 Diffraction anisotropy

How fast the intensity falls off with resolution can depend on direction. rugnux measures that, reports it, and does nothing else with it: no intensity is corrected, no reflection is removed on a directional criterion, and the merged data and the written files do not depend on direction at all.

The tensor. A deviatoric anisotropic displacement tensor is fitted to the merged intensities as

\ln \langle I(\mathbf{s})\rangle = c(\text{shell}) - \tfrac{1}{2}\,\mathbf{s}^\mathsf{T} B\, \mathbf{s},\qquad \mathbf{s} = \text{reciprocal-space vector},\ |\mathbf{s}| = 1/d

(§10.6 and §14.2 write s for \sin\theta/\lambda = 1/2d, so their s^2 = 1/4d^2; the two conventions give the same exponent -(B/2)(1/d^2), and B is the same B.)

with one free constant per resolution shell, so every isotropic feature — the Wilson curve, an ice ring, a noise floor, a scaling error — is absorbed exactly and only the \ell = 2 angular part drives the tensor. For an isotropic B this reduces to the ordinary Wilson plot, so B here is the ordinary crystallographic (B = 8\pi^2 U) B, directly comparable with phenix.xtriage's B_cart, ctruncate's anisotropic B eigenvalues and AIMLESS's anisotropic \Delta B. Only the deviatoric part is fitted: the isotropic part is degenerate with the overall scale. The tensor is constrained to the directions the Laue class allows — five free deviatoric parameters in triclinic, three in monoclinic, two in orthorhombic, one in tetragonal, trigonal and hexagonal, and none at all in cubic, where symmetry forces \Delta B to be exactly zero.

The fit is on intensities, with no positivity cut. Fitting amplitudes, or dropping non-positive intensities as an amplitude-based tool must, loses roughly 40% of the signal: in a direction that has died half the merged intensities are negative, so a positivity cut keeps only the positive noise excursions and flattens the fall-off exactly where the anisotropy is largest.

Two different quantities are reported, and they are not interchangeable. \Delta B (the range of the principal components) is a rate; the diffraction limit along each principal direction — where \langle I/\sigma(I)\rangle in a 20° cone about that direction falls through 2, read by interpolation in s^2 over equal-count shells — is where the signal actually runs out. A crystal can have a large \Delta B and almost no spread in directional limit, or the reverse. Where \langle I/\sigma(I)\rangle never falls through 2 in a direction, the limit returned is the edge of the measured data rather than the crystal's own; such a direction is marked — with a < in the report, a 1 in ANISOTROPY_D_MIN_CENSORED, and a note in the mmCIF — so the spread is not read as a measurement when it is a lower bound. Both constants are AIMLESS's defaults (cone half-angle 20°, \langle I/\sigma\rangle = 2). The 2 is a per-direction diagnostic level only: the dataset-wide resolution cut (§13.4) is CC$_{1/2}$-based, and the two criteria are not interchangeable.

The resolution signature. A genuine DebyeWaller B makes the directional deficit a straight line through the origin in s^2. The per-shell \ell = 2 amplitude is therefore fitted against s^2 and the curve is classified: linear (a real B), flat (a deficit that does not follow \exp(-\tfrac12 \mathbf{s}^\mathsf{T} B \mathbf{s}) at all, so the fitted \Delta B describes the data with the wrong functional form and may be an under-estimate), or convex (a deficit that grows faster than s^2, which a B cannot do). The verdict is re-derived at 8 and at 16 shells, and reported as undetermined if it moves.

The verdict, and what it is measured against. Whether an anisotropy is real is not decided against a counting-statistics error bar. Real data carry systematic error far larger than counting error, and gating on the latter reports anisotropy on datasets that have none. Instead the data set measures its own systematic error: in the tensor directions the Laue class forbids, the true tensor is exactly zero whatever the crystal is, so whatever is measured there is systematic. That measurement needs the unmerged observations — a merge has exact Laue symmetry by construction, and the forbidden directions are identically zero in it — so it is made on the scaled, rocking-curve-assembled observations. The counting part is subtracted, the counting error of the directions actually being tested is added back, and the result is the floor. What is tested against it is \Delta B_\text{linear} — the part of the fall-off that actually follows \exp(-\tfrac12\mathbf{s}^\mathsf{T}B\mathbf{s}), clamped at zero — and not the headline \Delta B; the report names which of the two it is quoting. The ratio is banded: below 2 not established, 23.5 marginal, above 3.5 established, above 5 strong. The bands are calibrated against known ground truth — merging cubic crystals in proper subgroups of their own Laue class, where the true anisotropy is exactly zero — which puts the false-positive rate at 24 % at 2.0, 10 % at 3.5 and 5 % at 5.0.

The report says NOT DETECTED, DETECTED, or CANNOT DETERMINE, and the third is a real answer rather than an evasion. It is returned when the Laue class is triclinic (no forbidden direction exists, so there is no internal measurement of the systematic error and no substitute for it), when the observed rotation range is under about 90° (a lab-fixed systematic then reaches several tensor directions instead of one), when the merged data are at the noise floor, when the scale model carried no dose term (an uncorrected dose ramp manufactures anisotropy that no significance test can see through), or when no unmerged observations were available. The smallest \Delta B that could have been established on the data set is reported with the verdict; it is set by the systematic error rather than by counting, so it does not improve with more reflections or a longer exposure.

A too-high space-group assignment is the one failure mode that is silent: real anisotropy is then pushed into the directions used to measure the systematic error, which inflates the floor and biases the answer towards reporting none. A caution says so, but only where that is actually indicated — a single free direction, no detection, and a forbidden-direction measurement far above its own counting noise — rather than on every tetragonal, trigonal and hexagonal data set.

Everything lands in <prefix>_report.txt section 9 (ANISOTROPY_* keys), in the printed statistics, and in the merged mmCIF: the eigen-decomposition of the tensor as the standard _reflns.pdbx_aniso_B_tensor_* items (relative to the weakest direction, since only the deviatoric part is determined), and the directional limits, the shape and the verdict under the _reflns.jfjoch_aniso_* local prefix.

13.6 Practical notes and limitations

  • Bragg integration is profile-fitted by default (per-shell Gaussian profile, Kabsch extraction; §9.3), with plain box summation available as a fallback (--integrator boxsum). The profiles are built per frame from that frame's strong spots, which suits fast-feedback and serial/streaming use; a profile shared across many frames (as in full offline workflows) is not currently formed.
  • Space-group symmetry beyond centering absences is not enforced during prediction/integration unless the space group is supplied and used downstream.
  • Resolution masking is controllable, and so is every stage of ice-ring handling (§3.3, §10.10). None of it runs unless the crystal is measured to have ice, because the fixed bands are a fixed cost in unique reflections whether it does or not.
  • Rotation vs still modes differ substantially in prediction and scaling: partiality is angle-driven in rotation data, while stills are predicted within an excitation-error window and get their partiality from the default-on per-crystal tilt post-refinement (§10.2) — or unit partiality with --simple-stills.
  • Amplitudes and intensities. The merged output carries both intensities (mmCIF intensity_meas, MTZ IMEAN/SIGIMEAN) and FrenchWilson amplitudes (mmCIF F_meas_au, MTZ F/SIGF; §10.8), so a downstream program can refine against either.

14. Model-based validation: R-free against a model and electron-density maps

Offline (rugnux --model model.pdb) the merged data can be scored against a supplied atomic model and initial electron-density maps computed — enough to confirm that a model fits the data and to inspect the density, not a substitute for refinement. The structure itself is not refined; the model is only re-fractionalized into the data unit cell (a rigid cell adjustment, so a deposited model with a slightly different cell still lines up), and the observed amplitudes are the FrenchWilson |F| from §10.8, so the R-free and the maps use exactly the same amplitudes as the written reflection file. The model, structure-factor, bulk-solvent and FFT machinery is provided by GEMMI.

14.1 Model structure factors

The model electron density is sampled on a grid (IT92 X-ray form factors, with a Refmac-compatible Gaussian blur chosen for the grid spacing) and Fourier-transformed to structure factors F_\mathrm{calc}(hkl) up to the data resolution.

14.2 Bulk solvent and scaling

A flat bulk-solvent mask around the model is transformed to F_\mathrm{mask}, and the model is scaled to the observed amplitudes by an overall least-squares fit of a scale k, an anisotropic B, and the flat-solvent parameters k_\mathrm{sol}, B_\mathrm{sol}:

$ F_\mathrm{model} = k,e^{-\mathbf{h}^\top \mathbf{B},\mathbf{h}/4}\left(F_\mathrm{calc} + k_\mathrm{sol},e^{-B_\mathrm{sol},s^2},F_\mathrm{mask}\right),\quad s^2 = 1/4d^2. $

This is the standard, few-parameter scaling model used by refinement programs. No free-form per-resolution-shell rescale is applied: such a rescale is dataset-specific and reshapes each map's radial amplitude profile differently, which would make maps from a multi-dataset campaign no longer directly comparable.

14.3 R-work and R-free

Crystallographic R-factors are reported over the work and free sets (the §10.7 flags):

$ R = \frac{\sum \big|,|F_o| - |F_\mathrm{model}|,\big|}{\sum |F_o|}, $

with R-free the same sum restricted to the free set. Note that the scaling of §14.2 is fitted over all reflections, work and free alike — its few parameters (k, an anisotropic B, k_\mathrm{sol}, B_\mathrm{sol}) are far too few to absorb individual reflections, but R-free here is strictly "free of refinement", not free of the scaling fit.

14.4 Electron-density maps

Two maps are formed with the model phases \varphi_\mathrm{model}: a 2F_o-F_c map, coefficients (2|F_o|-|F_\mathrm{model}|)\,e^{i\varphi_\mathrm{model}}, and an F_o-F_c difference map, (|F_o|-|F_\mathrm{model}|)\,e^{i\varphi_\mathrm{model}}, each inverse-Fourier-transformed to a real-space CCP4 map (<prefix>_2fofc.ccp4, <prefix>_fofc.ccp4). A map-coefficient MTZ (<prefix>_maps.mtz: FP, FC, PHIC, FWT/PHWT, DELFWT/PHDELWT, FREE) is written alongside so the maps can be reopened or rebuilt in Coot / PyMOL. These are unweighted difference coefficients (no \sigma_A / figure-of-merit weighting), which is why they are described as initial maps.

14.5 Aligning the data to the model: enantiomorph and indexing ambiguity

The model fixes a definite hand and indexing, but the merged data need not share them, so before comparison the observed reflections are brought into the model's frame.

  • Enantiomorph / screw. When the data space group is the enantiomorph of the model's (e.g. data P4_12_12, model P4_32_12; or P3_1/P3_2), the two are indistinguishable from merged intensities|F_\mathrm{calc}| is invariant under the change of hand, so R-free cannot choose between them and probing would be meaningless. The model's group is therefore adopted as the label the reflections are written under, and the reflections themselves are left untouched. The two groups of an enantiomorphic pair differ only in the translations of their operations: their rotations are identical, so they transform hkl identically, share a reciprocal ASU, and assign the Bijvoet hands identically. The label carries no handedness, and there is nothing about it to undo. Reindexing by the change-of-hand operator — which is the inversion — would instead swap I(+) with $I(-)$, flipping every anomalous difference on the strength of a label the space-group search itself reports as undetermined; where the model is genuinely the wrong enantiomorph for the crystal, it would manufacture agreement rather than reveal the mismatch. What does carry the hand is the indexing the data already have, from the diffraction geometry, and the anomalous differences that come with it. §14.6 is what tests them against the model.
  • Indexing (merohedral) ambiguity. When the crystal has a merohedral ambiguity (§10.9), the observed intensities do differ between indexings, and the right one is chosen against the best available reference. If a reference MTZ was supplied, the data were already reindexed to agree with it (§10.9 — by the reference-intensity correlation, at the merge stage for rotation data or per image in stills scaling), and model validation keeps that authoritative choice. Only with a model and no reference does validation resolve the ambiguity itself, as a fallback: the scaled model is fit to each reindexing of the data (identity plus the twin-law cosets) and the one giving the lowest R-free is kept. This matters for a multi-dataset campaign — a single shared reference fixes one indexing convention for every dataset, whereas an independent per-dataset lowest-R-free choice could send borderline datasets to different conventions. A no-op either way for a holohedral crystal (no twin laws). The two decisions then reach the written output differently, because only one of them moves reflections. The ambiguity choice is applied to the merged reflections themselves — and to the integrated observations behind the unmerged export — which are written after this step, so the reflection file, the R-factors and the maps describe one indexing. The change of hand changes only the space group the files are written under (the model's enantiomorph): no reflection moves, exactly as the first bullet says, so I(+) and I(-) stay as measured and §14.6's hand check remains a genuine test rather than an agreement manufactured by reindexing. The ambiguity choice is reported with the R-free of the winner and of the runner-up, since the margin between them is what says whether the data decided or the two came out within noise of each other.

14.6 Anomalous difference map and the sites it names

Where the merge kept the Bijvoet split (§10.5) — which a rotation merge does by default, whether or not the mates were averaged — an anomalous difference map is computed as well, with coefficients

$ \big(|F(+)| - |F(-)|\big), e^{i(\varphi_\mathrm{model} - \pi/2)}, $

over the acentric reflections that have both hands (a centric reflection has no anomalous difference, only noise). Turning the model phase back by 90° is what makes the anomalous scattering, which is 90° out of phase with the normal scattering, add up in the real part: the map's peaks then sit on the anomalous scatterers. It is written as <prefix>_anom.ccp4. Its hand is the one §14.5 settled: with the mates the wrong way round every peak becomes a trough, so a map of clean peaks is itself a check that the frame is right.

Rather than searching the map for blobs and leaving a list of coordinates, the map is read at the model's own atom centres (hydrogens excluded — they scatter no anomalous signal), and the ten highest, in units of the map's r.m.s., are reported in the log and as ANOMALOUS_SITE_01ANOMALOUS_SITE_10 in the results report. Each site is therefore named — the atom, residue and chain it belongs to — which is what says what carries the signal, not just where it is. The reading is cubic, not linear: the map is sampled every d_\mathrm{min}/3, and a peak that sharp read by trilinear interpolation comes out up to a quarter low — unevenly enough to reorder the sites. This reading is the one ANODE reports.

Because the map is built on the model's phases and the data's own indexing, it is also the only test of whether the two agree about the hand (§14.5). A model and a dataset in opposite hands turn every anomalous peak into a trough, so a map whose deepest hole at an atom is both deeper than 5\sigma and deeper than its highest peak says so, and the run reports it as a warning naming that atom. It is not repaired by reindexing: that would make the two agree by construction and destroy the evidence for which of the model and the data is in the wrong hand. Note that R-free cannot see this at all — a mirrored model gives R-free to four decimal places unchanged, and an exactly inverted anomalous map.

The list is always ten entries long, so it is their height that carries the information: on a sulfur-SAD dataset the sulfurs fill the top of the list and are followed by a clear drop to the couple of sigma that is the map's noise, while a dataset with no anomalous signal has no such separation and lists ten unrelated atoms at noise level. A scatterer the model does not contain — a bound ion, a soaked heavy atom — is by construction invisible in the list, and is what the map file is for.