Reading the numbers¶
A finished fit hands back a RefinementResult. This chapter is what is on it,
and how to read each part: the agreement indices, the two structure R-factors,
the distances and angles, what any restraints did, and the two counts that say
whether the pattern could support the model at all. The refined parameter values
themselves are the one part covered elsewhere, in The parameter table, because they are
a view of the parameter table.
The guidelines rank that evidence, and the ranking is not the one a table of R values suggests. The two most important criteria for judging a refinement are the fit of the calculated pattern to the observed one and “the chemical sense of the structure” [McCusker et al., 1999]. Rwp is half of the first.
Where a section here names a formula, Part 2 carries it as a numbered equation and this chapter links to it.
Fit statistics¶
RefinementResult.statistics is a Statistics object. The definitions follow
Toby [Toby, 2006], and Part 2 gives them as equation
(8.2).
Field |
Is |
Reads as |
|---|---|---|
|
the weighted profile R-factor |
how well the calculated pattern matches the measured one, weighted by the counting statistics. A fraction, so 0.0933 is 9.3 %. |
|
the unweighted profile R-factor |
the same comparison with every point weighted equally, so strong peaks dominate it. |
|
the expected R-factor |
the Rwp that perfect counting statistics alone would produce. It is a property of the data and the parameter count, not of the model. |
|
the reduced χ², Σw δ²/(N − P) |
how far the fit is from statistical perfection. 1.0 means the residual is the size of the noise. |
|
Rwp / Rexp |
the square root of |
|
Rwp with the background removed from both patterns |
the more meaningful figure when the background carries most of the raw intensity. |
|
the serial-correlation statistic |
≈ 2 means neighbouring residuals are independent. Far from 2 means the misfit is structured, whatever Rwp says. |
|
the Bérar-Lelann factor |
the amount the reported esds were multiplied by to account for that serial correlation. It has already been applied. |
|
N and P |
what makes the rest interpretable. |
|
max |Δθ|/esd over the last accepted step, external units both sides |
the convergence quantity (McCusker 1999 §7: converged when ≤ 0.1, a band quoted and never tuned). A converged fit satisfies it a fortiori; on a stage that stopped on its iteration budget it says how far the solve was still moving. |
|
the report’s identifiability sentence, verbatim: the exchange finding with the swap experiment and its license, and/or the softest notable mode |
the same sentence the report’s summary carries, delivered beside the numbers so a consumer that reads only the statistics block still gets it. The evidence behind it is |
Statistics.chi2 is the reduced χ², not Σw δ². The two differ by a factor of
N − P, which on a real pattern is several thousand.
Every figure in the block, and RefinementResult.y_calc beside it, is measured
on a compile at the values the result returns. Two results are not: a Pawley
result, whose intensity block belongs to the solve and is returned as the
frozen compile’s, and a joint fit, which rietx.multi builds itself. Both
report the frozen compile’s figures, and neither carries
FROZEN_COMPILE_STALE. The last stage solved on a
compile whose peak windows and quadrature node counts were frozen at the
values it started from, and when the two disagree by more than 1 % of χ² the
result says so as FROZEN_COMPILE_STALE, both figures in its message. The
esds are not re-measured: they are the last solve’s, read off its Jacobian.
To reproduce the number yourself, compile Refinement.fitted_structure and
Refinement.fitted_instrument with moving_paths set to the paths in
result.parameters. A compile claiming nothing moves sizes the peak
quadrature for the values alone, and on the Si SRM 640c acceptance fixture
that alone moves χ² by 1.4 %.
The literature is not consistent about which of the two is called χ². The IUCr
guidelines [McCusker et al., 1999] write χ² = Rwp/Rexp and say it should approach 1;
that quantity is Statistics.gof here, and Statistics.chi2 is its square. Both
conventions say the same thing about a fit, and the naming here follows Toby
[Toby, 2006], but a number copied from a paper needs the convention copied
with it.
Statistics.esd_inflation is conservative by construction. Perfectly white
residuals still land near 1.51, because same-sign runs happen by chance, and lab
data with unmodelled profile detail typically lands at 2 to 4. Treat it as an
upper bound on the serial-correlation damage rather than as a measurement of it.
Structure agreement indices¶
Every field above measures the pattern. Two more measure the structure, and a
journal will ask for at least one of them [McCusker et al., 1999]. They live on
RefinementResult.phase_agreement, one PhaseAgreement row per phase, keyed by
PhaseAgreement.name, the same name the phase carries everywhere else. Part 2
defines both as equation (8.3).
Field |
Is |
Reads as |
|---|---|---|
|
R_B, the Bragg-intensity R-factor |
how well the model reproduces the integrated intensities I = m·F², with m the multiplicity, rather than the profile. Written to a CIF as |
|
R_F, the structure-factor R-factor |
the same comparison on the amplitude rather than the intensity, so it is the number a single-crystal R is directly comparable with. Written as |
|
how many reflections were summed |
smaller than the phase’s reflection list whenever one falls off the Ewald sphere. |
A joint multi-histogram fit reports them per histogram, on
HistogramResult.phase_agreement: the partition is of one pattern’s counts, so
there is no pooled value.
Both are biased towards the model you are testing. A powder pattern does not measure individual reflection intensities, because overlapped peaks share their counts, so the “observed” intensity of each reflection is the observed pattern partitioned in proportion to the calculated one. A wrong model therefore receives the intensity it predicted, and both indices flatter it. That is the paper’s own warning, and it fixes what they are for: watching an R_B fall as you improve a model, rather than judging a model in isolation.
A trace phase’s R_B is not comparable with the major phase’s. Neither index is weighted (the sums count every reflection alike, whatever its counting statistics), so a reflection the fit barely constrains weighs as much as one that dominates it, and a minor phase’s windows sit under the major phase’s peaks, where the counts the major phase failed to describe are handed out too. Measured on the 11-BM NAC pattern with its 1.35 wt % CaF₂ impurity: NAC reads R_B 0.052 and the impurity 0.385, with the whole of the impurity’s misfit in four reflections at I(obs)/I(calc) ≈ 2.2, every one of them under a strong NAC peak with a large positive residual. Read R_B beside the phase’s weight fraction, and treat a trace phase’s value as a question rather than a measurement.
Both are absent, an empty list, for a Le Bail or Pawley fit. There the intensities are the fit, so the partition would be compared against itself.
The bonding geometry¶
The other thing a journal asks for is the distances and angles, and the
guidelines rank them above every R value: the two most important criteria for
judging a refinement are the profile fit and the chemical sense of the structure
[McCusker et al., 1999]. RefinementResult.geometry is a
GeometryTable carrying them.
Field |
Is |
Reads as |
|---|---|---|
|
one |
every asymmetric-unit atom’s whole environment, so the number of rows naming an atom is its coordination number. A bond between two sites is therefore in the list twice, once from each end. |
|
the two halves of that list |
split by |
|
one |
at every vertex, over that vertex’s complete bonded environment. Contacts are not arms. |
|
where coverage stopped |
a per-atom contact cap or a phase too large to search. Empty means nothing was bounded. |
A row’s value is GeometryDistance.distance in ångströms, or
GeometryAngle.angle in degrees. Each names its atoms twice. Once by the labels
the structure carries, as GeometryDistance.atom_1 and
GeometryDistance.atom_2, with GeometryAngle.atom_1, GeometryAngle.atom_2
and GeometryAngle.atom_3 putting the vertex in the middle. Once by index, as
GeometryDistance.atom_index_1, GeometryDistance.atom_index_2,
GeometryAngle.atom_index_1, GeometryAngle.atom_index_2 and
GeometryAngle.atom_index_3, which index Phase.atoms directly.
GeometryDistance.phase_index and GeometryAngle.phase_index say which phase,
in a multi-phase fit.
GeometryDistance.symmetry_2 is the CIF n_klm code of the image the second
atom sits at, and resolves against the operation list the exported CIF writes
beside it. GeometryDistance.symmetry_1 is always ., the published position;
so is GeometryAngle.symmetry_2, the vertex, while GeometryAngle.symmetry_1
and GeometryAngle.symmetry_3 code the two arms.
The esd on a distance¶
GeometryDistance.stderr is propagated through the whole parameter covariance,
as the guidelines require of any derived quantity: “the whole correlation
matrix, not just the diagonal elements, should be included in the calculation”
[McCusker et al., 1999]. Part 2 gives the propagation as equation
(8.6). GeometryDistance.stderr_diagonal is the same propagation
with the refined parameters’ correlations zeroed, the number a reader combining
the printed parameter esds in quadrature would get. It is carried so the
difference is visible, and it is never the answer.
GeometryAngle.stderr and GeometryAngle.stderr_diagonal are the same pair, in
degrees.
Dropping the correlations is not a conservative approximation, and both numbers
are reported for that reason. The figure is the 88 distances of an
11-BM NAC refinement under mccusker_structural (Rwp 0.0819), each plotted as
its diagonal-only esd over its full one: the ratio runs from 0.86 to 1.41, on
both sides of equality, so the cheap number is sometimes too small and
sometimes 40 % too large, with nothing in the row to say which.
The spread appears only once the coordinates refine. Under a plan that frees the cell and nothing structural, a cubic phase’s distances depend on one free parameter, the quadratic form has one term, and the two esds agree to the last digit. Correlation between coordinates is what the guideline is about, which is also why a table quoted from a profile-only refinement has nothing to say here.
An esd is None when it cannot be measured, which is absence rather than
σ = 0. Four routes lead there, and they mean the same thing. A result with no
covariance behind it (any evaluate-only pass, a replay of a history node
included) has distances and no esds at all. So does a row nothing free
reaches. A row whose value is fixed by symmetry has no variance to report: a
fluorite Ca–F distance with the cell held, or a rutile O–Ti–O angle, which stays
at exactly 90° however the one free coordinate degree of freedom moves. And an
angle at 0° or 180° is a stationary point, where the linearisation behind
(8.6) does not hold at all. The distance or angle itself is exact in
every one of the four; it is the uncertainty that is withheld.
Geometry is Rietveld-only. RefinementResult.geometry is None in Le Bail and
Pawley mode, where the dummy atom the mode requires is not a structure to
measure, and it is computed at the close of the fit rather than on demand: the
covariance it needs is read off the final Jacobian, which is never stored.
The size and the strain, in physical units¶
A refined profile width is a number of degrees, and a number of degrees is not
transferable: the same 0.09° of 1/cosθ broadening is a 400 Å coherent domain on
Mo Kα and a 233 Å one on Cu Kα. RefinementResult.microstructure is one
PhaseMicrostructure per phase, in phase order, carrying the four sample
coefficients read as the two quantities behind them. Part 2 states the two
conversions as equations (6.3) and (6.5).
Field |
Is |
Reads as |
|---|---|---|
|
which phase this block is about |
the index into |
|
one |
four rows ( |
|
|
which conversion produced |
|
the refined coefficient itself |
in its stored units, deg 2θ for the Lorentzian pair and deg² for the Gaussian variances, so a row can be checked against |
|
the reading and its uncertainty |
|
|
the parameter path it came from |
|
|
the two constants the size reading used |
the source’s longest declared line, and the Scherrer K. Carried so a size can be rescaled to another convention without knowing which build produced it. |
|
the Gaussian reading divided by the Lorentzian one |
1 means the two independent columns describe one specimen. |
|
whether size and strain are distinguishable over the range measured |
read below before quoting either number. |
FitReport.microstructure is the same list, carried onto the report so a reader
working from Refinement.report() need not go back to the result. The report
adds one thing to it, the separability caveat below.
A coherent domain size is the length over which the lattice diffracts in phase.
It is not a particle size and must not be reported as one.
Phase.particle_radius_um is a different quantity for a different correction,
the absorption path through a grain, and a particle is an aggregate of domains.
Nor is a domain size a two-figure number. K moves 10–20 % with crystallite
shape, so SCHERRER_K = 0.9 is one convention among several. GSAS-II reads its
size off the same FWHM with K = 1, and FullProf off the integral breadth, which
for a pure Lorentzian is 1.571× rietx’s answer. Quote the size as an order of
magnitude, with the K the row carries.
Separating size from strain¶
Over a short 2θ range 1/cosθ and tanθ are nearly the same curve, so a fit can
trade a size against a strain with no cost in Rwp. That is the Williamson–Hall
problem, and it is the reason a size is reported with PhaseMicrostructure.separable
attached rather than on its own. The verdict is the width trend’s own
(TrendAnalysis.separable on the "width" observable, with
PhaseMicrostructure.size_strain_collinearity carrying the correlation it was
decided on), so the report holds one opinion about it and not two. It is None,
meaning no claim made, when Layer 1 abstained or no width trend was fitted.
Where it says False, quote neither number alone. Refining one of the pair and
holding the other is not a workaround: the answer then depends on which one was
held. Measuring over a wider range is.
The esds, and the four absences¶
Each reading is a function of exactly one refined coefficient, so its esd is
that coefficient’s scaled by the derivative and nothing is dropped: a size’s
relative esd equals its coefficient’s, and so does a strain’s. A combined
Gaussian-plus-Lorentzian size would be a function of two correlated columns, and
is deliberately not reported. The two are read separately and compared through
PhaseMicrostructure.size_agreement instead.
MicrostructureTerm.unavailable names why a number is missing, and the four
answers mean different things.
|
What happened |
|---|---|
|
the coefficient is at its off state, where a size is infinite and a strain is a perfect lattice. True, and not a number to report. This is what a phase whose widths were never refined says, in all four rows. |
|
the source declares no emission line, so no size can be read. Never reaches a strain, which has no wavelength in it. |
|
the coefficient carries no esd, so |
|
nothing is missing. |
Size in a joint fit¶
A joint fit shares the crystallite size and not the degrees it is read in. In a
multi-histogram fit MultiHistogramRefinement normalises the size
coefficients by wavelength, because a size coefficient is proportional to λ
while a strain coefficient is not. The shared column is the coefficient at the
first histogram’s wavelength and every histogram’s own structure copy carries
the one it needs, so the crystallite behind them is a single number and the
SIZE_NORMALISED_ACROSS_WAVELENGTHS diagnostic says so. Read a shared
phases.*.lor_size as histogram 0’s coefficient, never as the number your
second pattern shows.
How many observations there are¶
Statistics.n_points is the N the least-squares algorithm uses, and it is not
the number of observations the pattern holds. Only the integrated intensities of
individual reflections are unique observations [McCusker et al., 1999]; the profile
steps across one peak are repeated measurements of the same number. The
consequence is that the algorithm will refine far more parameters than the data
support, without complaining, because its N runs into the thousands.
RefinementResult.data_support is a DataSupport object carrying the count
that answers this. Part 2 states the correction it applies as equation
(8.4).
Field |
Is |
Reads as |
|---|---|---|
|
reflections this pattern measured, summed over phases |
one symmetry orbit is one reflection, and a Kα doublet’s second line is the same reflection measured again, not a second observation. A reflection counts when a fitted channel lies within half its own FWHM of its position, so an excluded region removes what sits under it and a peak half-measured at a range end still counts. |
|
the same count corrected for overlap |
each reflection contributes the fraction of its own area on which no overlapping reflection stands higher, so an isolated line is worth 1 and an exactly coincident pair is worth 1 between the two of them. A float, because a partly resolved pair is worth more than one and less than two. |
|
the free parameters the ratio is about |
the atomic ones: coordinate DOFs, occupancies, Biso, ADP components. The cell, zero, profile, background, scale, preferred orientation and extinction are excluded, because peak positions and shape determine those rather than the intensities being counted. |
|
the raw count divided by the parameters |
the upper bound on the ratio. |
|
the effective count divided by the parameters |
the number the guideline is about: at least three and preferably five [McCusker et al., 1999]. |
The complement of DataSupport.n_structural_parameters is
Statistics.n_free_parameters minus it: the profile, background and cell
parameters, which the same fit refined against the same pattern but which the
peak positions pay for.
DataSupport.n_unique_reflections over-counts, on purpose. Two reflections at
the same 2θ are one observation, and both are counted here. In a cubic cell that
pair is common: (300) and (221) coincide exactly. The raw count is therefore an
upper bound on the information, and the ratio built from it is optimistic.
DataSupport.n_effective_observations is the corrected number, by the method of
Altomare et al. [Altomare et al., 1995], and the gap between the two is the
pattern’s overlap.
The figure is one Cu Kα LaB6 pattern over 15-140°, holding the cell and the reflection list fixed and widening the peaks with Lorentzian size broadening alone, so the overlap is the only thing that moves. Twenty-six reflections throughout; 22.0 effective observations while the lines are sharp, and 3.7 by the time the pattern is one hump. The raw count cannot see any of that, so it is the wrong denominator for a parameter ratio.
The estimate is not a theorem. The paper says so, and the IUCr guidelines repeat it: the approach “may not have a rigorous basis”, and what it gives is a reasonable estimate of how many parameters the data will support. Two numbers a reader should have with it: the interval each reflection is judged over is ±2 FWHM, and the paper’s own check at ±4 FWHM lands 6.5 % lower on average, so the reported figure is a little generous.
Nothing here refuses anything. The number is evidence, read beside the fit
rather than as a gate on it, and a ratio below three is a reason to hold
parameters rather than a reason the fit is wrong. The sharper question is which
parameter is unsupported, and that is the next section. A ratio below five
raises the DATA_SUPPORT_LOW diagnostic, as a warning below three and as
information between three and five.
Which parameters the data could not separate¶
RefinementResult.identifiability is an Identifiability, and it exists because
these statistics cannot be recovered afterwards. They are read off the final
Jacobian, an N × P array nothing serializes (a history node stores state, not
curves), so they are screened at fit time or lost. Rwp and the residual are the
contrast, since any consumer can recompute those from the arrays already on the
result.
Field |
Is |
Reads as |
|---|---|---|
|
dot-path → the block projection R² of that structural column onto the background column span |
how much of a parameter the background could imitate. Every screened path is here, not only the ones over the guard threshold: a fired/not-fired bit is a verdict, 0.46 against 0.08 is evidence, and these are the same numbers the |
|
the worst-|ρ| pairs, each a |
which two refined parameters moved together |
|
the softest directions of the scale-normalised normal matrix, each a |
the same problem where it involves three parameters or more, which no pairwise list can show |
|
one |
whether a parameter you held could have been absorbed by the ones you refined |
Field |
Is |
|---|---|
|
the two dot-paths |
|
their signed correlation, from the same final Jacobian the |
|
of ĴᵀĴ with every column normalised to unit length, so it is dimensionless: for a single pair with correlation ρ it is 1 − |ρ|, and 0 is exact degeneracy |
|
dot-path → component of the unit eigenvector, kept above 0.1 and signed so the largest is positive |
|
the held parameter’s dot-path |
|
the block projection R² of its column onto the span of the free ones, evaluated at the converged values with the candidate freed on a copy of the vary set, never refined |
|
dot-path → signed loading of the reconstruction, kept above 0.05: which fitted values absorb the held parameter’s signature |
The softest modes are carried whatever their eigenvalues, because the number is
the evidence and deciding where comment starts is the report’s job. Read
ExchangeRow.r2 alone with care: it is a property of the design matrix over the
sampled range and fires on clean fits too, measured at 0.999945 on both a
degenerate fixture and its clean reference. The discriminating half, whether
anything significant is riding the exchange, is in the report, whose
ExchangeFinding carries the fitted partner’s value and esd beside the same R².
That chapter’s FitReport.identifiability is a different type from this one, and
the warning there says how they differ.
The other half of “how flexible was the background” is not in that table. The
absorption column says what the background could imitate;
RefinementResult.n_background_components says with how many explicit
humps it was allowed to do it. N humps are 3N parameters with
unconstrained positions, which a reader comparing two Rwp values has to be able
to see.
There are two counts because there are two kinds of component.
RefinementResult.n_extra_components is the length of
Instrument.extra_components: everything you declared, humps and
declared sharp peaks alike.
n_background_components is the humps alone, and it is the one this paragraph
is about. A declared peak is freedom you granted, and it is not background
freedom; background_absorption excludes it for the same reason. Declare no
peaks and the two numbers are equal, which is every pattern written before v1.4.
It sits on the result rather than in the table above because it is declared and
not measured. The four fields above are read off the final Jacobian and exist
only where a solve measured them, while this is a count of what the instrument
carried, so it is there on a replayed node too, where identifiability is
None. 0 means none was
declared, exactly off; None means nothing counted, which is a joint
multi-histogram fit (one count per histogram) or a result recorded before the
field existed.
How finely the peaks were sampled¶
The other half of the question is about the experiment rather than the model. There should be at least five steps across the top of each peak, and generally not more than ten [McCusker et al., 1999]. Below five, the integrated intensity of that reflection was never measured, and no refinement afterwards recovers it.
PatternDiagnostics.steps_per_fwhm is that number: the median, over the
pattern’s resolved peaks, of how many channels span the peak’s half-height
width. PatternDiagnostics.n_peaks_measured says how many peaks the median was
taken over. Both come from diagnose(data), which needs no model, so this is a
question to ask of a file before refining it, and the answer does not change
afterwards:
d = rx.diagnose(data)
print(d.steps_per_fwhm, "steps per FWHM over", d.n_peaks_measured, "peaks")
A refinement reports the same measurement on its fitted channels, as the
PATTERN_UNDERSAMPLED diagnostic, and only when it falls below five. There is
no code for the upper end of the band: oversampling costs beam time, not
validity.
Everything else diagnose measures¶
The object that call returns is a PatternDiagnostics, and the two sampling
fields above are two of its seventeen. The rest describe the pattern itself, with
no model involved, and they are what the background chapter’s defaults are chosen
from.
Field |
Is |
Reads as |
|---|---|---|
|
channels in the file |
|
|
the range, in degrees |
|
|
the median channel noise |
the scale a peak height is significant against |
|
share of channels more than 3σ above the background envelope |
how much of the pattern is peak rather than background |
|
resolved peaks found |
|
|
those per degree 2θ |
above roughly 2/deg the pattern is dense, which favours a stiff baseline and a low background order |
|
near-maximum net signal over the median background level |
|
|
how much of the envelope’s cubic-fit residual variance a 1/(2θ) column explains |
a nested-model test for the low-angle air-scatter rise; a high value is what turns the 1/x background term on |
|
RMS of the envelope residual after both the cubic and the 1/x term, over the median level |
what is left is genuinely broad non-polynomial structure (amorphous content, capillary glass), and calls for a more flexible background |
|
the arPLS stiffness the whiteness rule picked for this pattern |
|
|
the sampling pair above |
null when no peak was measurable |
|
the lines of a Kβ or W Lα contamination finding, each a |
empty is the usual answer; see the warning below |
|
the bulk pattern’s σ²/max(y, 1), median over the middle half of the range |
1.0 is pure Poisson counting; anything else says the file’s σ is something else (merged detectors, a monitor normalisation). Null means σ was not measured, so nothing was checked |
|
stretches whose σ carries more variance per count than that plateau, each a |
the pattern’s statistical weight is not uniform across its range; see below |
|
ends of the range where the level collapsed and stayed down, each a |
read this one first; see below |
|
short runs, inside the range or at either end, that measure nothing and outvote the pattern while doing it, each a |
empty also means not checkable: the test needs the file’s own σ. See below |
A contamination is one finding, and each flag is one line of it. The first four fields are that line’s, the last three the finding’s, and the last three repeat across every flag of one finding.
Field |
Is |
|---|---|
|
which ghost line the finding is of |
|
this weak peak’s position |
|
the strong peak it is a ghost of |
|
this peak’s intensity over that parent’s |
|
the ratio fitted across every supporting parent |
|
strong parents that supported the finding |
|
strong parents whose predicted ghost position fell inside the range |
Whether Kβ reaches the detector is a property of the optics, so a leak puts a
line at the predicted position of every strong reflection, all at one ratio.
Reading each match on its own could not tell that apart from a coincidence, and
a coincidence is much the commoner event: five of the eight strongest parents
must now carry a candidate at a common ratio before anything is reported at
all. Read leak_ratio to decide whether the beam is contaminated, and
intensity_ratio only to see how one line contributed.
Warning
An empty PatternDiagnostics.contamination means “nothing was flagged, or
nothing was checked”. The Kβ position is anode-specific, so the screen needs
wavelength=, and a wavelength matching no tabulated Kα1 is skipped silently.
Measured on the 11-BM pattern, diagnose(data) and
diagnose(data, wavelength=0.4139090) both return an empty list: the first
because nothing was asked, the second because a synchrotron wavelength has no
anode.
The joint bar costs sensitivity, deliberately. A Kβ image injected into six round-robin patterns is found on every one of them from about 10 % of its parent upwards. Below that it depends on how many peaks the pattern has: zincite is caught at 1 %, magnetite, which yields 22 usable peaks, not until 10 %. An unfiltered tube sits at 14 % (Hölzer et al. 1997), which is the case the screen is for. A residual leak past a working filter is below the floor and comes back as nothing.
The region below the first reflection¶
air_scatter_gain above tests the envelope’s shape over the whole fitted range,
so a beam or air-scatter tail a few tens of channels wide out of several hundred
barely moves it. Measured on a real pattern it reads 0.019 against the 0.3
trigger, because the tail carried 0.3 % of the whole-range residual the nested
cubic-against-cubic+1/x fit compares. A fitted result carries a second, narrower
check for exactly that case. LOW_ANGLE_UNMODELLED reads the fit’s own residual
over [two_theta_min, first_tick − 2·FWHM), taking every phase, every
emission line and every declared peak so that a peak’s own low-angle flank is
never counted, and fires when that region’s mean weighted-squared residual
exceeds three times the whole-pattern reduced χ² (Diagnostic.value).
first_tick is the lowest tick with intensity behind it: an image, on any
emission line, of a reflection whose strongest calculated point reaches 1σ of
the noise on some line, which is the per-phase threshold PHASE_UNCONSTRAINED
reads applied reflection by reflection. A Kα2 or λ/2 image of a reflection the
data sees therefore keeps its place however weak it is on its own. A tick with
nothing behind it (a phase held at a vanishing scale, a declared peak at zero
area, a reflection whose structure factor is zero) does not move the boundary.
It is silent when the region holds fewer than ten channels: no reflections
carrying intensity at all, or the first one sits at or near the low edge. The message reports the ratio and leaves the choice of remedy open.
Raising the pattern’s lower limit to where the residual falls under 2σ and adding
the background’s air-scatter term are both available, and only whoever is looking
at the pattern can tell whether the region is genuinely outside the beam or the
background model is locally wrong there.
A declined air-scatter term¶
auto_background declares the P-spline’s 1/(2θ) term only when
air_scatter_gain crosses 0.3, and a declined term is absent, so no plan can
free it. On a line-dense synchrotron capillary pattern with a broad hump that
envelope test is blind. Its 3° windows miss a rise confined to the first few
degrees and sit far above the true background where the lines crowd. On the
synthetic NAC pattern of tests/test_air_scatter_undeclared.py (λ = 0.4133 Å,
0.5–50°, 0.005° steps) it reads 0.0042 with a strong rise and 0.044 without one.
AIR_SCATTER_UNDECLARED asks the fit instead. It fires on a P-spline with no
air term when re-fitting the background with a 1/(2θ) column, the Bragg model
held, would remove at least 0.5 % of χ² and at least 25 times the reduced χ²
(Diagnostic.value is the fraction). It is the one-column score test on the
fit’s own residual, so it costs no refit. It is one-sided because the term is
bounded at zero: a residual leaning the other way is a direction the term
cannot take. On that pattern it reads 0.13 with the rise, where declaring the
term cuts χ² by 24 %, and nothing on the control, where declaring it cuts
nothing. A weaker rise that LOW_ANGLE_UNMODELLED misses, because it runs on
under the first lines, still reads 0.011. Where the spline can draw 1/(2θ) on its
own, the projected column is small and the code is silent. The remedy is in
the suggestion: declare the term and refit.
How many observations stand behind each channel¶
PatternDiagnostics.coverage_regions is the one measurement here that reads the
σ column rather than the intensities. The statistic is the variance
inflation σ²/max(y, 1), which is 1 for pure Poisson counting and proportional to
1/n_eff for a channel averaged over n_eff independent observations; a region is a
run of channels carrying more of it than the bulk of the scan does
(rietx.background.counting_coverage).
On an instrument with a bank of detectors on a circle, the number contributing to a given 2θ falls off at both ends of the range, and this counts them. Measured on two NIST BT-1 constant-wavelength neutron patterns (3.00–166.25° at 0.05°, σ from the file), both show the same ladder: about 5× below 8°, 2.2–2.6× out to 15°, tapering to 1× by 55°, then a step back to 2.2× inside a single channel at 161.30° and held to the end of the scan. The levels are quantised because detectors are integers, so the ends read as steps rather than as a gradual falloff.
Field |
Is |
|---|---|
|
where the region runs, in degrees |
|
the median σ²/max(y, 1) over the region, divided by the plateau: how many times the bulk’s variance per count it carries. A median over the whole span, so any bridged sub-threshold channels (below) are in it |
|
the region’s width in channels. Runs closer together than the smoothing window are merged into one, so this is the region’s span, bridged in-region dips back below the threshold included, rather than a count of channels each individually over it |
|
|
Warning
This is descriptive, and it triggers nothing. If the file’s σ is right then weighted least squares already gives those channels the weight they deserve, σ being for exactly that, so a region is no argument for trimming the range. What it tells you is that the pattern’s statistical weight is not uniform across it. That is a fact about how many detectors saw each angle and not about the specimen. Where it disagrees with a hand-chosen fit range, that is information about the range, in either direction and with no recommendation attached.
Read CoverageRegion.edge as a location rather than a diagnosis. It says
whether the run reached an end of the range or sat inside it, and no more. The
causes it leaves unseparated are several. An end region can be a detector bank’s
coverage running out (a quantised falloff), or a synchrotron analyser bank
tapering out smoothly at high angle; an interior one can be a dead or excluded
detector, two scans stitched together, a variable counting-time or slit
schedule, or threshold-crossing chatter a channel or two wide. The table below
is where those readings live.
On the bundled patterns, three with a measured σ fire, and none of them is the plain detector falloff the BT-1 provenance shows:
pattern |
fires |
how to read it |
|---|---|---|
|
high 46.0–60.0°, 2.80× |
a synchrotron analyser bank tapering out: the ratio rises smoothly across the last third, where a detector count would make a quantised step |
|
high 53.0–67.0°, 2.82× |
the same taper |
|
low 20.3–62.5°, 5.50×; interior 66.7–75.0°, 2.74×; interior 140.5–147.5°, 1.58× |
a ladder non-monotonic in both directions, which reads as a variable counting-time or slit schedule rather than a detector count; the two interior regions are threshold-crossing chatter |
It is silent on the one other σ-bearing bundled pattern
(panalytical_attenuator.xrdml) and returns null on the two with no σ column.
The σ column is genuinely structured on all three, so the reading to take is
that edge locates the structure and this table names its cause.
One thing it does settle: a σ column that shows this structure was measured. The
Poisson fallback σ = √max(y, 1) makes the ratio identically 1, so a pattern
without a σ column reports null and an empty list, meaning not checkable,
which is a different answer from an empty list beside a plateau. The same ratio
also rises wherever the counts collapse to nothing with σ staying finite, a
different phenomenon wearing the same statistic.
Where the instrument stopped seeing the sample¶
PatternDiagnostics.signal_cutoffs is the phenomenon that last paragraph names,
measured from the intensities where it belongs, and it is the one field here to
read before the others. Each entry is a SignalCutoff: an end of the range
where the level collapsed to a small fraction of the interior level and stayed
there. The cause can be a detector’s active area running out, a beamstop or
sample-environment shielding. Whichever it is, those channels carry no
diffraction information and belong outside the fit range.
Field |
Is |
|---|---|
|
|
|
the boundary, meaning where the collapse began rather than where it reached its floor, because that is what a fit range should stop at. Still a live channel, so a trim keeps it |
|
the region’s median level over the interior level: how dead it is |
|
channels outside the boundary, i.e. exactly what a trim there would drop |
|
the implied precision penalty: the region’s median σ/y over the interior’s. Null when σ was not measured |
Measured on a private constant-wavelength neutron PSD scan (σ from the file),
the two ends are different shapes and carry different arguments. The trailing
end is a cliff: from the interior level to a floor of a few per cent within a
couple of degrees, then flat for the rest of the range. That is a factor of
tens in a couple of degrees, and past it nothing at all. The leading end is a
graded degradation: a direct-beam shoulder, a shadowed floor of a few per cent,
a broad bump at a modest fraction of the interior level, then a climb that
reaches the interior level only tens of degrees in. The bump sits at d ≈ 30 Å,
so it is not sample diffraction. There is structure at the low end, and it is a
beamstop halo and air scatter rather than the specimen, so fitting a background
through it means describing non-specimen structure at several times the
interior’s fractional error. signal_cutoffs reports a boundary at each end:
the leading one within half a degree of the window the data owner’s own TOPAS
refinements of that file declare, the trailing one a little further out.
Warning
Nothing is trimmed for you and nothing else is re-measured. diagnose reports
these and stops, because a fit range is a protocol decision and the numbers above
do not settle it on their own.
That is also why the ordering matters: every other field of a
PatternDiagnostics is measured over the whole range it was handed. On that
same file, full range against the data owner’s TOPAS window,
amorphous_hump_score is inflated nearly twofold by the dead tail.
air_scatter_gain is understated by more than an order of magnitude, hiding the
real low-angle rise entirely, and baseline_lambda moves by two decades. If you
decide to trim, call diagnose again on the trimmed pattern. The first answer
described the range you gave it.
SignalCutoff.relative_error_ratio is derived rather than a second observation.
Where σ²/y is constant, as it is across that whole pattern straight through both
transitions, σ/y is 1/√y up to a constant. So the ratio is 1/√floor_fraction
and says nothing the level did not. It is reported because it is the number an
experimenter reads. That σ²/y stays flat is also what separates this from
PatternDiagnostics.coverage_regions: the file’s σ there is honest, and those
channels are empty rather than thinly covered.
A channel that measures nothing and outvotes the pattern¶
PatternDiagnostics.dead_channels is the short companion to the section
above: a run shorter than a cutoff, judged wherever it sits, in the interior
or touching either end of the range. A run spanning the whole pattern is not
judged, having no live neighbour to be weighed against. A dead or masked
detector cell, a gap between banks, a channel the electronics dropped: its
intensity falls to nothing and its esd falls with it.
Weights are 1/σ², so it does not merely contribute nothing. It outvotes its
neighbours, and the background model is pulled down to meet it.
Field |
Is |
|---|---|
|
the interval you would exclude |
|
how many channels it spans |
|
the run’s median intensity over the local background level |
|
how many live channels one of these outvotes. Always present, because the census answers nothing without a measured σ |
Measured on an ILL D1B constant-wavelength neutron scan of Co₃O₄ (λ = 2.52 Å),
two cells read 3 and 5 counts at σ = 1.000 and 1.414, beside live channels at
36 503 and σ = 55.1. Each therefore carries about 3 000 times the weight of a
live channel. Fitted with those two channels inside the range, a 12-term
Chebyshev background is dragged through zero to reach them and Rwp comes back
at 0.087; excluded, the same refinement gives 0.0074. The fit with them in
emits fourteen bound hits, a BACKGROUND_ABSORPTION for every phase and a
DATA_SUPPORT_LOW. None of those is the cause.
Warning
This one needs the file’s own σ column and answers nothing without it. Under
the Poisson fallback σ = √max(y, 1) a dead cell and a channel that honestly
counted zero are the same two numbers, so there is nothing to tell apart. An
empty list on a pattern whose coverage_plateau is null means not checked.
That is a different answer from checked and clean.
Nothing is excluded for you, here or anywhere else: excluded_regions on the
project is where a fit range is declared, and it is a protocol decision.
What the restraints did¶
If the fit carried soft restraints, RefinementResult.restraints is a
RestraintReport, and on a restrained refinement it is the first thing to read,
before Rwp. How a refinement works covers declaring them and scheduling their weight;
this is the object that comes back.
Field |
Is |
Reads as |
|---|---|---|
|
one |
|
|
Σ weight·(deviation/σ)² |
the penalty term S_G of (9.12), at unit weight scale. |
|
the c_w the stage ran at |
so the penalty the fit actually minimised is |
|
how many rows there were |
the rows are in the covariance but out of Rwp, so this is not part of N. |
The deviations themselves are always reported unscaled, because “is this
restraint satisfied?” is a question about the geometry and c_w is a choice about
how hard to insist on the answer. A row past three σ raises RESTRAINT_TENSION:
the restraint and the data are pulling against each other, and one of them is
wrong. A joint fit reports the same object per histogram on
HistogramResult.restraints.
What the statistics cannot tell you¶
They measure agreement, not correctness. A background flexible enough to imitate the peaks biases displacement parameters up and phase fractions down while Rwp improves. Of the eight corrections shipped in v0.5, two provably cannot move Rwp, one moves it the wrong way when it is right, and the two largest accuracy wins are invisible in it.
That is what the report is for, and why a correction in this package ships with a record field or a diagnostic saying what it changed, rather than with an Rwp comparison as its evidence.
What the absorption correction did¶
RefinementResult.absorption is the worked example of that rule.
Specimen absorption is one seam with three geometries behind it, and one of them
provably cannot move Rwp: the capillary factor is exactly a constant times
exp(c·sin²θ), so applying it is an exact reparameterisation of the phase scale
and the displacement parameters. A comparison of fits would show nothing. The
AbsorptionCorrection record is where the correction says what it changed.
It is present only for a Rietveld fit that carried a specimen dimension, and
None otherwise, including for a fit that declared none, which is not the same
as a specimen of no thickness. Patterns, structures and instruments covers declaring it.
Field |
Is |
Reads as |
|---|---|---|
|
|
which geometry’s expression ran |
|
the dimensionless µ·(length) |
the capillary radius for the cylinder, the specimen thickness for the two plates |
|
|
estimated means composition × packing × that length, so it inherits their uncertainty |
|
the λ it was computed at |
µ is wavelength-dependent, so the number is not portable between sources |
|
the Biso bias, in Ų, that refining without this correction would have absorbed |
positive means add this to recover the unbiased value. For the cylinder this is the entire content of the correction |
|
the share of ln A that a free scale and a free Biso cannot reproduce |
zero for the cylinder to rounding, a few to tens of per cent for a plate, and hence how far to trust the ΔBiso above |
|
the same measure applied to ∂lnA/∂µt |
the number behind the decision not to make the thickness refinable |
|
µt·exp(1 − µt), transmission only |
the counts this specimen delivered as a fraction of the best it could have. A specimen-preparation number no fit statistic can express: a badly chosen thickness costs counting statistics, not accuracy |
|
whether µR left the expression’s validated range |
|
|
why no correction ran, or null |
absence with a reason attached |
The flat-plate cases are not exactly reparameterisable, which makes their
AbsorptionCorrection.equivalent_delta_biso a lower bound rather than the
answer. The projection behind it is unweighted while a refinement finds a
weighted compromise, and measured against synthetic refits the bias a fit really
absorbs runs 1.06 to 1.5× the predicted one, tracking
AbsorptionCorrection.unabsorbed_fraction.
Diagnostics¶
RefinementResult.diagnostics is the channel for “your answer is wrong although
Rwp is fine”. Each entry is a Diagnostic with a Diagnostic.code, a
Diagnostic.level, a human-readable Diagnostic.message, a
Diagnostic.suggestion, Diagnostic.where naming the parameter paths
involved, and Diagnostic.value carrying the headline number where the
diagnostic has one (a correlation’s ρ, an absorption’s block R²). None
means “no single number”, not zero.
The codes are an open vocabulary, deliberately: a new correction ships with the diagnostic that states what it changed. Read them before the statistics, every time.
A persistently correlated pair is deduplicated across a plan’s stages (the
worst |ρ| kept, every stage it fired in named in the message) and the whole
list is capped at ten HIGH_CORRELATION findings, worst first. Past the cap a
HIGH_CORRELATION_OMITTED entry gives the omitted count and points at
result.identifiability for the rest, so the list is never silently
truncated.
Printing a result¶
print(result), which is str(result), is the view a bare RefinementResult
can answer on its own: per-stage status and the last stage’s convergence number
(McCusker et al. 1999 §7’s max|Δθ|/esd), every diagnostic, provenance, and the
agreement indices last. It needs no compiled model, so it works on a result read
back from a project or handed across a process boundary, and it stays under 80
lines on a five-stage fit.
Refinement.summary(), called after fit() on the session that ran it, prints
more, because a session holds the compiled model: the same per-stage and
diagnostics rows, then the Layer 0/1 misfit summary (the worst regions by χ²
share, the unmatched-peak count, the serial correlation read out as
Durbin-Watson with the esd inflation it costs), the next parameter worth
freeing, the protocol actually run (the plan, every held path grouped by why,
the excluded ranges, N points against N reflections, the σ source), and finally
the agreement indices and a named visual check:
print(ref.summary())
The next: line is a Refinement.suggest probe on the channels the fit ran on,
read as ΔBIC rather than as the Δχ² that ranks. It prints next: free instrument.profile.w, predicted ΔBIC +48.2 at N_eff 612 (Δχ² 1.2e+03), or
next: nothing ΔBIC admits when the leading candidate’s gain does not pay for
its parameter. N_eff is the count the ΔBIC was charged at, the channel count
over the squared Bérar–Lelann factor, and not the raw channel count. It costs one Jacobian build and no solve; The fit report has
the field and the two questions it separates.
deliverable= adds the rows one purpose actually decides on, because the report
will not infer your purpose for you. They are "phase_id" (unmatched observed
peaks, the Le Bail gap ratio), "qpa" (background.absorption’s worst entry,
the weight fractions with esds) and "structure"
(identifiability.exchangeability). How much of each phase, and “Structure agreement
indices” above, say what each row means:
print(ref.summary(deliverable="qpa"))
A figure takes a title only when you ask for one. RefinementResult.plot
(rietx.viz.plots.plot_result), PatternData.plot and SeriesResult.plot
(rietx.viz.plots.plot_trajectory) accept title=, which draws the caller’s
words above the figure in the panel’s own type. Without it no title is drawn
and the figure is byte for byte the one it was before the keyword existed. For
a paper the caption is the title, and the fit statistics stay a corner
annotation either way. A notebook cell or a batch of thirty patterns has no
caption, and that is where the keyword belongs:
result.plot(path="p03.png", title="pattern 3 of 30, 250 °C")
plot_for_vlm refuses title= by name. Its panel titles are the evidence the
vision model reads, and a caller’s words prefixed to them would push those
numbers further along the line.
plot= writes plot_for_vlm to that path and names it on the last line. It
draws the same regions the text names, with their numbers in the panel titles,
so the picture confirms the text rather than standing in for it:
print(ref.summary(plot="fit.png"))
report= takes the FitReport that ref.report() already built, so a caller
that wants the text and the structured report builds one report instead of two.
The report is the expensive half of the call, and in a batch the doubling is the
dominant cost. Pass the report or plan=, never both: a report was built under
its own plan.
report = ref.report()
print(ref.summary(deliverable="qpa", report=report))
A series prints its own view. SeriesResult.summary(), which is
str(series_result), is the trajectory table (first and last few entries, with
the count between them), the SEQUENTIAL_* series-level diagnostics, and each
shown entry’s status and Rwp. See Refining many patterns for what the series-level
diagnostics mean.
The fourth deliverable is the series’ own, because no single pattern carries a
SEQUENTIAL_* row: series.summary(deliverable="series") adds the rows that
decide a parameter measured against the series axis: whether the ordering was
checked at all (direction="both", or a plain statement that it was not), the
persistent findings, each step with what an independent cold pair reproduced,
the phases held for want of support, and the two things no diagnostic can
supply, a stated 2θ-scale anchor and the precision/accuracy split. The other
three purposes are refused there by name, and Refinement.summary (deliverable="series") answers with where they live:
print(series.summary(deliverable="series"))
Progress¶
A staged fit or a long series can take longer than a caller wants to wait on
silently. progress= takes a text stream or a path, on fit() and on
refine_sequential()/SequentialRefinement.fit(), and writes one line per
stage boundary (one per pattern under a series, each stamped with its place in
it):
result = ref.fit(data, progress=sys.stdout)
# stage scale_bkg converged Rwp 0.0933 2s
# stage zero converged Rwp 0.0894 1s
# ...
series = rx.refine_sequential(patterns, structure, instrument,
labels=["250C", "300C"], progress="run.log")
# [series 1/2] 250C stage scale_bkg converged Rwp 0.0812 3s
# [series 2/2] 300C stage warm_refit converged Rwp 0.0805 1s
It is implemented as an events= consumer rather than a second telemetry
channel. Every number on the line is read straight off the same event
rietx watch tails, so the two never disagree about what happened. Pass both
freely: progress= adds a subscriber, and does not replace a callback or path
already given as events=.