7. Background

7.1. Additive background models

The background is part of the model in (1.1), and is never subtracted from the data. An estimated baseline is held additively under the refined polynomial, or co-refined under a smoothness penalty. Three models refine: Chebyshev polynomials, a fixed estimated baseline plus Chebyshev, and a P-spline [Eilers and Marx, 1996]. The P-spline is a B-spline basis whose coefficients \(c\) are disciplined by second-difference penalty rows appended to the residual (1.5):

(7.1)\[r_{\mathrm{pen}} \;=\; \sqrt{\lambda\, m}\;\frac{D_2\, c}{\bar\sigma},\]

Source: rietx.background.models.pspline_penalty_scale

with \(D_2\) the \((n-2) \times n\) second-difference matrix, \(\bar\sigma\) the median \(\sigma\) over the fitted channels and \(m = N/n\) the fitted channels per coefficient, both frozen at stage compile. The coefficients carry the intensity’s unit and every data row is divided by its \(\sigma\), so without \(\bar\sigma\) the penalty’s weight would move by \(k^2\) when the intensities and their \(\sigma\) are multiplied by \(k\). Dividing by \(\bar\sigma\) takes the unit out of \(\lambda\), and \(m\) takes out the sampling density, because the data’s own weight on one coefficient grows as \(m/\bar\sigma^2\). So \(\lambda\) is a pure number, the penalty’s weight against the data’s. A background feature of width \(W\) then costs about \(\lambda (h/W)^4\) of the fit it buys, with \(h\) the knot spacing. That is derived, not measured, and its constant of order one is not computed. Setting lambda_units to "intensity" drops both factors, which gives the rows \(\sqrt{\lambda}\, D_2 c\) that every fit used before and keeps them as the bit-identical way back.

The rows land in \(J^\top J\), so the covariance is regularised, and they are excluded from \(R_{wp}\) and the serial-correlation statistics. They are soft observations rather than data.

7.1.1. What the pure number leaves to the counts

A pure λ fixes the penalty’s weight against the data’s, so it fixes the relative bias a too-stiff spline leaves in a broad feature. The χ² that bias costs grows with the feature’s height in σ, so the same λ is stiffer, in what it costs, on a better-counted pattern. Measured on the synthetic NAC capillary pattern of tests/test_air_scatter_undeclared.py (0.005° steps, auto_background’s 3° knots, a hump of FWHM about 8° and 900 counts over a 300-count level, no air rise), with every count multiplied by k:

k

hump height / σ

χ²_red at λ = 1

at λ = 0.1

at λ = 0.01

0.01

2.6

1.137

1.111

1.105

0.1

8.2

1.301

1.079

1.039

1

26

3.764

1.478

1.060

10

82

28.44

5.489

1.293

The excess over the λ = 0.01 column grows about tenfold per decade of k, as \((h/\sigma)^2\) would. FitReport.background.worst_absorption reads 0.02 in every cell and BACKGROUND_ABSORPTION fires in none, so on this pattern the lower λ buys the hump without reaching the lines. The rows are invariant to the step: the k = 1 cell at λ = 1 reads 3.761 at 0.001° steps. The default stays 1. On a well-counted pattern with a broad hump, the signature of a λ stiffer than the data supports is the one increase_background_flexibility reports: an off-region χ²_red well above 1 at a Durbin–Watson d well under 2 (here 4.04 and 0.48 at λ = 1, 1.08 and 1.79 at λ = 0.01). Lower lambda_smooth a decade at a time, and read worst_absorption and BACKGROUND_ABSORPTION at each step before keeping one.

7.2. A measured background

The fixed curve can be a measurement rather than an estimate. An empty vanadium can, a blank capillary, a matrix-only scan and an empty furnace are all scans of everything the specimen is not. Such a curve enters as \(f\), sampled onto the pattern grid, with one refinable multiplier on it:

(7.2)\[y_{\mathrm{bkg}}(2\theta) \;=\; \sum_n c_n T_n(x) \;+\; s\, f(2\theta).\]

Source: rietx.model.forward.CompiledModel.background

Both forms are TOPAS’s bkg_file [Coelho, 2018], which takes the scale as an argument or leaves it at one.

A blank is never on the specimen’s scale. The two scans differ in monitor normalisation and in counting time, and the specimen attenuates the container’s own scattering. The polynomial on top cannot repair that. It is additive, so it moves the level and never rescales the shape, while the error in a fixed curve is multiplicative. Quantitative phase analysis is the workflow least tolerant of the difference, because what a wrong background level biases is the phase scales and hence the weight fractions, while \(R_{wp}\) improves.

The model is linear in \(s\), so its Jacobian column is the sampled curve itself and it joins the background’s linear block. It is also largely parallel to the constant Chebyshev term wherever the curve is not strongly featured. Large esds on \(s\) and on \(c_0\) are then the correct report, and HIGH_CORRELATION is where that surfaces.

The curve’s own counting statistics, where it has them, enter the channel weight:

(7.3)\[\sigma_i^2 \;=\; \sigma_{y,i}^2 \;+\; s^2\, \sigma_{f,i}^2 .\]

Source: rietx.model.forward.compile_model

That is what makes a short blank scan honestly worse than a long one. The \(s\) in it is the value the stage compiled at, frozen for the stage like every other discrete choice. Two limits stand outside the formula. Interpolating a noisy curve onto the pattern grid correlates neighbouring channels’ errors, so the propagated \(\sigma\) is a lower bound unless the curve is smoothed first, and smoothing is a modelling choice that belongs in the record. One blank reused across a series carries error that is fully correlated rather than error that averages down, so a refined scale that trends along a ramp is partly an artefact of the single measurement.

A scalar is the leading-order correction rather than the exact one. The container’s scattering reaches the detector through the specimen in the sample scan and through nothing in the blank scan, so the exact multiplier follows the specimen transmission of (5.5). A single scale is therefore right per pattern, and not as one constant across a temperature or atmosphere series.

7.3. Explicit humps

The three models above are global. A Chebyshev term and a spline coefficient each act over a stretch of the pattern, so describing a feature confined to a few degrees means making the whole curve flexible. The alternative is an additive Gaussian on the angle axis, one per declared feature, summed on top of whichever model is in use:

(7.4)\[y_{\mathrm{peak}}(2\theta) \;=\; h \, \exp\!\left[-4\ln 2 \left(\frac{2\theta - 2\theta_0}{\Gamma}\right)^{\!2}\right] \qquad [\text{counts}],\]

Source: rietx.background.models.hump_curve

This is an empirical basis function rather than a peak shape, and no physical derivation is claimed for it. Genuinely amorphous scattering is a Debye or radial-distribution term in \(Q\). The citation is for the practice: GSAS-II exposes an explicit broad background peak as a hump [Toby and Von Dreele, 2013], and TOPAS exposes one as a cell-less “peaks phase” [Coelho, 2018].

Two consequences follow from the form. It is nonlinear in \(2\theta_0\) and \(\Gamma\), where every term in (7.1)’s block is linear, so it stays out of the linear design matrix and its Jacobian columns are finite differences of the whole model. And it is not identifiable on its own. \(h\), \(\Gamma\) and the low-order polynomial terms underneath it all raise the same region of the curve, so the width is what makes the term a background term. With \(\Gamma_{\mathrm{inst}}(2\theta)\) the resolution function of (3.1)-(3.3), the admissible régime is

(7.5)\[\Gamma \;\gtrsim\; m\,\Gamma_{\mathrm{inst}}(2\theta_0),\]

Source: rietx.strategy.staged.HUMP_MIN_WIDTH_MULT

with \(m =\) 4.0. Below that width the term is a reflection with no cell and no structure factor behind it. The condition depends on the refined resolution parameters and on \(2\theta_0\) itself, so it cannot be a box constraint. It is carried as a reported guard, like the Stephens strain cone, and a firing means the peak parameters are not quotable.

7.4. Model-free estimation

Baseline estimators serve the fixed-plus-Chebyshev model and the automatic pipeline. The Whittaker smoother [Eilers, 2003] solves the banded (pentadiagonal) system

(7.6)\[(W + \lambda D_2^\top D_2)\, z \;=\; W y,\]

Source: rietx.background.estimators.whittaker_solve

and arPLS [Baek et al., 2015] iterates it with asymmetric reweighting so peaks are progressively excluded from the baseline. SNIP [Ryan et al., 1988] is available as an independent alternative.

7.5. Choosing the flexibility

Two knobs are selected with the same two ingredients: the Chebyshev order (or P-spline λ), and the estimator’s λ.

  • BIC on peak-masked channels [Schwarz, 1978]. Background flexibility has to be justified by the background channels alone, so Bragg-peak channels (net

    3σ above a robust baseline) are masked out of \(\mathrm{BIC} = m\ln(\mathrm{RSS}/m) + k\ln m\).

  • Durbin-Watson whiteness stopping [Durbin and Watson, 1950; Hill and Flack, 1987]. \(d = \sum(\Delta_i - \Delta_{i-1})^2 / \sum\Delta_i^2\) rises toward 2 as the background stops leaving serially correlated structure. Past the stopping threshold, extra flexibility chases noise. Masked channels are treated as contiguous, which makes the test slightly conservative, in the safe direction.

Source: rietx.background.select

7.6. Background flexibility and bias

A background able to imitate the peaks biases ADPs up, and scales (hence QPA fractions) down, while \(R_{wp}\) improves. The right measure is the block projection of a structural Jacobian column \(j_i\) onto the span \(B\) of the background columns:

(7.7)\[R^2_i \;=\; 1 - \frac{\lVert j_i - P_B\, j_i \rVert^2}{\lVert j_i \rVert^2},\]

Source: rietx.optimize.statistics.background_absorption

\(R^2_i\) is the fraction of the parameter’s effect the background can reproduce. Pairwise correlation is the wrong statistic here. With ~100 spline coefficients each individual \(|\rho|\) stays small (~0.2), while the block collectively absorbs ~50 % of the parameter (measured). The projection has to include the penalty rows of (7.1): they stiffen the background against imitating a peak, and dropping them overstates the risk by ~5×.