Files and projects¶
This chapter is the map of what the package reads off disk and what it writes back. Patterns, structures and instruments is what the objects it hands you contain.
graph LR
P["pattern file<br/><i>.xye .fxye .raw .cif …</i>"] --> RP["read_pattern"]
C["structure<br/><i>.cif</i>"] --> SC["Structure.from_cif"]
I["instrument profile<br/><i>.json</i>"] --> LP["load_instrument_profile"]
G["GSAS-I instrument file<br/><i>.prm</i>"] --> GP["read_gsas_prm"]
M["another program's refinement<br/><i>.inp .pcr .EXP</i>"] --> PM["read_project_model"]
RP --> REF(["refinement"])
SC --> REF
LP --> REF
GP --> REF
PM --> REF
subgraph rex ["my_sample.rex/"]
PJ["project.json<br/><i>settings</i>"]
PC["the pattern file<br/><i>copied byte for byte</i>"]
H["history.jsonl<br/><i>model state</i>"]
LV["live/<br/><i>one directory per run</i>"]
EX["exports/<br/><i>CIF, tables</i>"]
end
REF --> PJ
REF --> PC
REF --> H
REF --> LV
REF --> EX
REF -.-> RUNS[".rietx/runs/<br/><i>one directory per fit,<br/>with no project</i>"]
Every fit writes the dotted arrow whether you asked for it or not. A refinement
that came from a project records into that project’s live/ instead, and
.rietx/ is what a fit outside a project leaves in the working directory. Both
are below.
Pattern files¶
read_pattern opens a pattern and returns a PatternData, whose fields are in
Patterns, structures and instruments. It identifies the format from the file itself rather than from the
extension:
from rietx import capabilities
caps = capabilities()
[(fmt.name, fmt.sigma) for fmt in caps.reader_formats]
[(opt.name, opt.help) for opt in caps.reader_options]
Ask capabilities() rather than trusting a list in prose. The format list went
from five to ten in two days once. ReaderCapability.sniff says how each format
is recognised, ReaderCapability.sigma says where the uncertainties come from,
and ReaderCapability.refuses says what the reader declines and why.
Four properties of the readers reach a caller, and each of them can change a number you quote.
A multi-range file holds scans, and the reader selects one. Pass scan= to
choose. The ranges are never concatenated, because two ranges are usually two
weighting regimes, and joining them silently mixes them.
A pdCIF holds blocks instead of scans, and takes block=. A file with
a _meas block and a _calc block is a different pattern depending on which
you ask for. read_pdcif reads one directly.
A reader may repair a file only where it can say that it did. Pass a list as
diagnostics= and the repairs come back in it:
import rietx as rx
notes = []
data = rx.read_pattern("my_sample.raw", diagnostics=notes)
for note in notes:
print(note.code, note.message)
The intensities and σ need not be the file’s numbers. Vendors disagree about
whether an attenuator factor is already applied (four formats, three answers),
so the reader applies it or not by measured convention, and σ goes through the
same transformation either way. Where the scale cannot be established the reader
withholds σ and says so with PATTERN_INTENSITY_SCALED, because the Poisson
fallback is wrong by √t on a rate.
The scanned axis is never assumed. Most vendor files are not powder scans at all, so a file whose axis is something other than 2θ is refused by name, and an axis the reader cannot identify says so.
Weights follow from all this. The package uses the file’s esd column when the
file has one, and Poisson σ = √max(y, 1) only as the fallback. It never
subtracts an estimated background: hold the background additively
(BackgroundFixedPlusChebyshev) or co-refine it under a smoothness penalty
(BackgroundPSpline).
Structures from CIF¶
Structure.from_cif reads a crystal structure. It takes phase_name= to pick
one block from a multi-phase file, aniso=True to read an anisotropic
displacement loop, and the same diagnostics= channel:
notes = []
structure = rx.Structure.from_cif("my_phase.cif", aniso=True, diagnostics=notes)
aniso is opt-in on purpose. Several CIFs carry an anisotropic loop, and
reading a file must not silently change which parameters a plan will free.
Two repairs happen at read, and both are recorded rather than assumed. A species
label that is not a recognised scatterer is normalised
(CIF_SPECIES_NORMALISED), and a cell angle that disagrees with its space group
by a small amount is snapped to the symmetry value
(CIF_CELL_ANGLE_CORRECTED): a β of 90.002(3) under P m m m is an
experimenter quoting a refined number. Past that threshold the symbol and the
angle contradict each other, one of the two is wrong, and choosing between them
is yours: the value is left byte for byte and the read raises.
A B_iso outside the 0–25 Ų starting bounds is kept, and its bound is widened
to hold it (BISO_BOUND_WIDENED). Published structures carry such values. The
methylammonium C in COD 4335638 sits at 26.8 Ų, and a light atom is sometimes
refined slightly negative. The number is the file author’s model, so refine it
or set a physical start before a stage holds it. A fit that keeps a negative B
reports BISO_NEGATIVE.
A further note is a report rather than a repair. A site can sit within 1e-4 of a
special position without being on it, as a file quoting five decimals often
leaves one. Such a site has its orbit expanded at that position, so its
multiplicity is the special one, and SITE_SNAPPED_TO_SPECIAL_POSITION names
the site, the shift and the multiplicity. The stored coordinates are unchanged
and the fit is unaffected. What the multiplicity decides is how many atoms the
site puts in the cell, and so ZMV and every weight fraction; compare it against
the file’s own _atom_site_symmetry_multiplicity.
Building a phase by hand, with no file behind it, has one more way to go quiet.
A bare Hermann-Mauguin symbol resolves to the first setting the tables hold, and
40 symbols hold more than one: the :1/:2 origin choices and the rhombohedral
:H/:R axes. Site multiplicities differ between settings, so
spinel’s origin-2 coordinates under a bare F d -3 m build Mg₂AlO₄ where
MgAl₂O₄ was meant, with Rwp unmoved. SPACE_GROUP_SETTING_ASSUMED quotes the
cell contents each setting implies, which is the part you can recognise:
import rietx as rx
P = rx.Parameter
spinel = rx.Phase(
name="spinel",
space_group="F d -3 m:2", # the setting, and not the bare symbol
cell=rx.Cell.cubic(8.0806),
atoms=[
rx.Atom(label="Mg", species="Mg", x=P(value=0.125),
y=P(value=0.125), z=P(value=0.125)),
rx.Atom(label="Al", species="Al", x=P(value=0.5),
y=P(value=0.5), z=P(value=0.5)),
rx.Atom(label="O", species="O", x=P(value=0.2624),
y=P(value=0.2624), z=P(value=0.2624)),
],
)
Naming the setting silences it. Files carry theirs, so a CIF or a TOPAS .inp
never raises it.
A magnetic structure: reading and writing a magCIF¶
A magnetic structure reaches you as a magCIF, and Structure.from_cif reads
one. MAGNDATA, ISODISTORT, k-SUBGROUPSMAG and Bilbao’s own tools all emit that
form, and what it carries is exactly what the model stores: the operator list
with its time-reversal signs, the moments in crystal-axis components and the
BNS symbol.
notes = []
structure = rx.Structure.from_cif("0.642.mcif", moment_ions={"Mn1": "Mn3+"},
diagnostics=notes)
phase = structure.phases[0]
print(phase.space_group, phase.magnetic_symmetry.bns_number)
print(phase.atoms[1].moment.values(), phase.atoms[1].moment.ion)
Four things about that read are worth knowing before you trust the answer.
The nuclear space group’s label comes from _parent_space_group.name_H-M_alt
(a magCIF states no ordinary _space_group_name_H-M_alt at all, since the
magnetic block replaces it). The IUCr underscore screw-axis spelling
(P 2_1/c) is normalised to the plain one gemmi’s lookup table accepts (P 21/c) before the symbol is resolved, and the file’s own spelling stays
unchanged in every diagnostic. The symbol names the type, and the setting is
checked against the file’s own magnetic operators with time reversal dropped,
because gemmi’s tabulated operator table for a Hermann-Mauguin symbol can sit
at a different origin or axis choice than a non-standard setting’s own
(measured on Ima2, space group 46), and a symbol lookup on its own would attach
the wrong coordinate frame to a right-looking name. The atom positions refine under
the highest tabulated group the file’s atoms satisfy, and the moments always
under the file’s magnetic group. The parent is kept when it contains those
operators and gives every listed site the same number of images as they do:
that is a nuclear phase under the parent carrying moments under a magnetic
group that may be a subgroup of it, Cr₂WO₆’s Pn'nm in P4₂/mnm, and
CIF_MAGNETIC_NUCLEAR_GROUP says so (“positions refine under ‘P 42/m n m’,
index 2 over the file’s own ‘P n n m’; moments under the file’s magnetic
group”). Otherwise the phase is built under the tabulated setting whose
operations are exactly the file’s own, and CIF_MAGNETIC_NUCLEAR_SETTING says
which and why. Structure.from_cif(..., nuclear_group="file") always takes the
file’s own group, and nuclear_group="parent" always takes the parent and is
refused by name where the parent cannot carry the file’s atoms; the default,
"auto", is the rule above.
That covers a symbol tabulated at another origin, and a symmetry-lowering
magnetic group whose asymmetric unit lists one parent orbit as two sites,
which under the parent would be one orbit counted twice. A file whose own
group is in a setting no tabulated symbol names, an origin or axis choice
gemmi’s table does not hold, is read with that group as its operation list:
Phase.symmetry_operations holds the file’s operations with time reversal
dropped, in the file’s order, Phase.space_group holds a bracketed label
naming the closest type ('P -1 [unnamed in this cell]'), and
CIF_MAGNETIC_NUCLEAR_SETTING says so. Every symmetry consumer reads the list,
never the label, and Structure.to_cif writes the list back. A writer whose
format states a group only as a symbol refuses such a phase by name. A group
whose rotations are in no tabulated orientation at all, a two-fold along a face
diagonal for instance, is refused by name: an operation list takes its crystal
system and cell ties from the tabulated group with the same rotations, and
there is none.
A _parent_space_group.transform_Pp_abc
that is not the identity is a record of which setting the symbol was
tabulated in, one every non-standard-setting MAGNDATA entry carries, and is
not by itself a reason to refuse the file. What is refused is a genuine
supercell, judged by the child transform’s determinant and by k, never by the
parent transform. A file with no space group this reader can resolve at all,
neither an ordinary tag nor a magCIF parent symbol, is refused by name.
The operators are stored verbatim, and transform_BNS_Pp_abc is not applied.
A published structure is usually written in its own setting rather than the BNS
standard one (Cr₂WO₆’s record carries 'b,-a,c;0,0,0'), and that setting is
the one its cell and coordinates are in. So the list is kept as written, the
transform goes into MagneticSymmetry.setting as the record it is, and applying
it would move the group away from the atoms beside it.
MagneticSymmetry.uni_number is derived from the operators through spglib, so
it is the same on both sides of a round trip. A stated number_BNS or
number_OG is checked against the same identification: a number labels the
group’s type in every setting, so one the operators are not is a contradiction.
The operators are what is refined, so their numbers are stored, the stated
name_BNS is dropped with the wrong number, and CIF_MAGNETIC_NUMBER_MISMATCH
names both. The reader cannot tell which side is wrong: a forgotten
anti-centring in the centring loop makes the operators a smaller group while
the number stays right.
The magnetic ion is not in the file. The magnetic dictionary has no item for
it, and MAGNDATA writes a bare Mn for a site that is chemically Mn³⁺, so
without moment_ions= the neutral atom’s ⟨j₀⟩ is used and
CIF_MAGNETIC_ION_UNCHARGED says so. The two curves differ, and a moment
refined under the wrong one absorbs the difference into its magnitude. A bare
lanthanide or actinide symbol has no neutral-atom row at all, so there the
fallback goes one step further, to the element’s majority oxidation state
that is magnetic (Ln³⁺, except Eu²⁺; U⁴⁺, Np⁴⁺, Pu³⁺), reporting
MAGNETIC_ION_ASSUMED instead. moment_g= is the same shape for the Landé g
a 4f moment needs. Where the ion itself was assumed this way, its free-ion
Hund’s-rule g_J is assumed too (magCIF has no item for g any more than it has
one for the ion, so refusing on a quantity the file could never have carried
either way would be the wrong-shaped fix), reporting LANDE_G_ASSUMED. An ion
stated explicitly (moment_ions=) still needs moment_g= stated alongside
it, unchanged: that gap is the caller’s, unrelated to this default. Both
arguments are keyed by _atom_site_moment.label, and a key that names no row
of that loop is refused with the labels the file does carry, since a
misspelled label would leave its site on the default with nothing said.
All three moment forms are read. The dictionary defines the moment in
crystal axes, in spherical coordinates and in Cartesian ones; a reader taking
only the first would silently drop the other two. The spherical and Cartesian
frames are the one where x ∥ a and z ∥ c*, which is the frame the ADP
code already uses, so no second convention enters. Where a file states more than
one form they must agree, and a file stating _atom_site_moment.magnitude has
it cross-checked against the components under the cosine metric the
dictionary defines: |m| = √(mᵀ·G·m) and not √(Σmᵢ²), which on hexagonal axes is
29 % out and is the commonest way that line goes wrong. The check compares
|magnitude| against the (always non-negative) derived norm, since a
negative stated magnitude is a sign convention rather than a mismatch, at a
tolerance drawn from the file’s own stated precision (magnitude_su, an
inline value(su), or half the last printed digit when neither is given)
rather than a fixed band, widened by each component’s own imprecision
propagated through the metric, since the derived norm is itself only as
precise as the components it came from. A magnitude rounded to fewer
figures than the components is read, not refused. A component written ?
or . beside numbers in the same form is refused, naming the site and the
item. In CIF ? means unknown and . inapplicable, the dictionary gives a
moment component no default, and reading either as 0 μ_B would state a moment
the file did not. A form whose cells in a row are all ? or . is one that
row does not use, and the other form is read.
Standard uncertainties follow the form the moment is taken from. A
crystal-axis row keeps its own. A Cartesian row’s go through the conversion,
which is linear, so each crystal-axis esd is the Cartesian ones propagated
through it, assuming no correlation between components, since a CIF row
states none. A spherical row’s are not propagated, because that conversion
is nonlinear and singular on the pole, and CIF_MAGNETIC_SPHERICAL_ESD_NOT_STORED
names each su that was dropped, so a component’s stderr=None there is not
read as “not refined”.
Two constructs are refused by name rather than half-read. A modulated
structure (any _atom_site_moment_Fourier loop, special function, cell wave
vector or superspace group), because reading its k = 0 amplitude alone would
return a collinear structure for a modulated one. And a structure stated in a
supercell of its parent, which is the generic commensurate k ≠ 0 record: the
asymmetric unit of the magnetic group in that cell splits one nuclear orbit
into several sites carrying different moments, and the model stores one moment
per nuclear site. Both refusals name what would be needed. The k ≠ 0 test
reduces every component modulo 1 first, because a propagation vector is only
ever meaningful modulo the reciprocal lattice: an all-integer k such as
(1, 1, 1) is Γ identically and reads exactly like k = 0, never as a
supercell. The parent k is read for that test and nothing else, so it is not
stored and not written back.
Structure.to_cif and Refinement.write_cif write the same form back. The
operator and centring loops go out verbatim, the BNS/OG metadata with them
(a MagneticSymmetry whose number is not its operators’ is refused by name),
and the moments in crystal-axis components with the dictionary’s own
_su items beside them, not in the value(su) notation every other number
in the file uses, and for a measured reason: a refined moment’s esd is routinely
0.05-0.9 μ_B, so two significant figures of su would quote the value to one
decimal and the number you refined would not come back. Two items have no
magnetic-dictionary tag at all, the form-factor ion and the Landé g, so they go
out under a namespaced _rietx_atom_site_moment. category; without them a round
trip would lose which form factor the refinement used. A refinement CIF also
carries _atom_site_moment.magnitude_su, the esd of the modulus, which is
where a moment’s uncertainty lives. _space_group_magn.transform_BNS_Pp_abc is
written only when MagneticSymmetry.setting is a transform in the dictionary’s
form, the (P,p) from the current setting to the BNS one ('a,b,c;0,0,0'), and
is left out otherwise. A group built from its number is in the BNS standard
setting, so its setting is 'a,b,c;0,0,0'. A magnetic_supercell phase has
none: the parent-to-child transform it knows is a different map, and it stays
on the statement and in the phase name.
A magnetic phase from a TOPAS .inp¶
A TOPAS .inp states a magnetic phase as mag_space_group on a str, with
mlx mly mlz (and optionally mg) on each magnetic site. Its magnetic
examples, and ISODISTORT’s TOPAS export, often state no nuclear space_group
at all. rx.read_topas_inp reads such a str as a phase. The build then takes
the nuclear group from the magnetic one: its operators with time reversal
dropped, under the magCIF reader’s tier-2 rule above. TOPAS itself generates
the atoms from those operators.
A mag_space_group written as a BNS number (62.448, 1.1) or an OG number
is used as the group directly. It resolves to the operators of that number’s
standard setting, and TOPAS_MAGNETIC_GROUP_READ says so. A Shubnikov
symbol is only kept as metadata, because nothing here parses one. The build
refuses a phase that states moments under a symbol until you pass
to_structure(magnetic_symmetry=...).
mlx mly mlz are components in the fractional basis,
m = mlx·a + mly·b + mlz·c with the edges in Å. The stored
crystal-axis moment is therefore (mlx·|a|, mly·|b|, mlz·|c|). This is a
per-axis scale, never a rotation. It is the TOPAS Technical Reference’s own
reading, § 13: Fmagc = L·Fmag, and its MM_CrystalAxis_Display macro gives
mxc = mlx·a. A file’s MM_CrystalAxis_Display line is the number to compare
with. On the Durham LaMnO₃ tutorial it agrees to the printed digits. It is
also measured against TOPAS’s own calculated intensities. On a monoclinic
cell with β = 115° and a moment on all three axes, TOPAS 6’s magnetic
intensity ratios between reflections of equal d match this reading on every
pair to four digits, and its absolute magnetic intensity gives the same |m|.
A crystal-axis reading misses the median pair by about 58 %, and a Cartesian
one by about 80 %. With the moment along b alone, where the three readings
coincide, all three match, which shows the comparison can tell them apart.
mg is the site’s Landé g, which enters only the ⟨j₂⟩ weight of the
magnetic form factor. TOPAS can refine it (Technical Reference § 13), but
Moment.g is a fixed number. So an mg the file refines (mg @ 1.9, or a
named parameter) is read as held at the file’s value, and
TOPAS_MOMENT_G_HELD names each such site.
mag_only and mag_only_for_mag_sites switch a site’s nuclear scattering off.
The file that uses them typically restates the magnetic sites in a second
str. Building that without the switch would count those sites’ nuclear
scattering twice, so it is refused by name rather than dropped.
rx.write_topas_inp writes a magnetic phase back in the same form. The str
states mag_space_group with the BNS number and no space_group, and each
moment-bearing site states mlx mly mlz, the stored μ_B divided by the edge,
plus mg where a Landé g is set. Read back, each component is within one ulp
of the one written. On 180 random moments over an orthorhombic, a monoclinic and
an oblique triclinic cell, most come back bit for bit. Whatever the str
cannot state is refused by name:
a group in a setting other than its number’s standard one, since TOPAS generates the standard operators;
a nuclear group larger than the magnetic group’s own family group, since TOPAS would generate only part of each parent orbit;
a magnetic group on a phase with no moment;
a moment ion that differs from the site’s species, since TOPAS takes both from the one
occ;a propagation vector, either the parent k of a supercell or a phase’s own. A magCIF drops both as well (a supercell is written as k = 0 in its own cell), so the
Structure’s own JSON is the export that keeps them.
The group’s symbol rides as a comment and is not read back.
A magnetic phase from a FullProf .pcr¶
FullProf states a magnetic structure as a separate pure magnetic phase
(Jbt = 1, or -1 with the moment as (M, φ, θ)). The phase lists only the
magnetic atoms and has its own Scale. FullProf’s manual requires that phase’s
scale and structural parameters to be constrained to those of its
crystallographic counterpart. So to_structure reads it as the moments of that
counterpart, one Jbt = 0 phase of the same file, and only when all of the
following hold:
the magnetic phase is stated in the counterpart’s own cell, with k = 0 (the magnetic cell, where the manual says the Fourier term is the moment);
its symmetry is
Isy = -1SYMM/MSYM pairs, each moment matrix ±det(R)·R with phase 0, so each pair is a Shubnikov operation;each magnetic atom sits on one of the counterpart’s sites, with the same Biso and the same number of images;
the two phases’
Scale × f²agree, f being theOccfactor, because one scale is all the built phase has.
The moments are μ_B on unit vectors along the cell axes. That is magCIF’s
basis, so the numbers carry over as written. FULLPROF_MAGNETIC_PHASE_READ
reports the merge. A moment carries one refine flag, so a site whose file
frees only some of its three moment columns (Jbt = -1 with only M coded,
say) comes in with the whole moment free, and the message names the site.
Every other magnetic phase is refused by name, with its reason:
the Fourier-component form, meaning a k other than 0, an imaginary part, or a
MagPhor MSYM phase;basis functions of an irreducible representation (
Isy = -2);a moment matrix that is not an axial action;
a counterpart stated in a different cell, as when the nuclear phase is in the chemical cell and the magnetic one in the supercell, or one whose space group does not build;
a magnetic phase with no sites;
Jbt = -1on a cell that is not orthogonal, where the manual leaves the in-plane frame of φ open.
rietx.io.projects.fullprof.magnetic_reading(model, phase) returns that
reason as a string, so you can ask before building. nuclear_only=True still
omits every magnetic phase, and says so.
Instrument profiles¶
A calibrated instrument is a file. save_instrument_profile writes one and
load_instrument_profile reads it back with every parameter vary=False, which
is the second half of the two-step lab workflow:
result = rx.refine(standard_data, standard, instrument, plan="lab_calibrate")
rx.save_instrument_profile(ref.fitted_instrument, "cu_ka_10mm.json")
instrument = rx.load_instrument_profile("cu_ka_10mm.json")
result = rx.refine(data, structure, instrument, plan="lab_sample_refine")
Calibrate on a standard with its certified cell held fixed. That is what
decorrelates the zero shift from the sample displacement from the cell, and it
is why lab_sample_refine is the only plan whose size and strain numbers mean
what they say.
A profile saved by an earlier release can carry a parameter whose declared
range that release dropped, and reading one restores it. That is the repair the
history log gets (§ The history log below), and it reports on the same
diagnostics= list, which load_instrument_profile takes too. A profile
stamps the schema version it was written under, and one with no stamp is read
as the oldest.
The GUI writes and reads the same file from the Model panel, with Save profile… and Load profile…. Saving lands it in the project’s exports/
directory. It needs a model and not a fit, unlike everything else written
there: the other exports describe a refinement result, while a profile
describes the instrument as it stands.
Reading a GSAS-I .prm instrument file¶
read_gsas_prm reads a GSAS-I instrument-parameter file (Larson & Von Dreele,
GSAS, LAUR 86-748) straight into an Instrument, with the same vary=False
contract as load_instrument_profile. It is the text file an APS 11-BM mail-in
ships beside its pattern:
instrument = rx.read_gsas_prm("beamline.prm")
It reads the dominant case the format ships: one bank, HTYPE PXCR
(constant-wavelength X-ray), GSAS profile function 3. Its neutron twin, HTYPE PNCR with profile function 3, reads too, onto
Instrument.constant_wavelength_neutron at the file’s LAM1: the profile
record is defined by profile function rather than by radiation, so the same
coefficients land in the same places. A neutron source has one wavelength and
no polarization, so such a file’s POLA and KRATIO are read and not applied,
and a non-zero LAM2 is refused. It converts GU/GV/GW
from centidegrees² and LX/LY from centidegrees into the degrees² and degrees
ProfileTCHZ uses. A file stating a Kα1/Kα2 doublet comes back with two
emission lines, the second weighted by the KRATIO field, which is the Kα2/Kα1
intensity ratio and so is what EmissionLine.weight means. A doublet with no
KRATIO is refused: the conventional 0.5 would be the reader’s number rather
than the file’s, and the polarization two fields earlier is conventionally 0.5
as well. A neutron time-of-flight file (HTYPE PNTR) and every other
GSAS profile function are refused by name rather than approximated, and the
refusal says which reason applies to which. A time-of-flight file puts something
onto the axis that ProfileTCHZ’s constant-wavelength Caglioti/TCH law cannot
express. A profile function other than 3 is refused under PXCR and PNCR
alike, naming the type the file carries, because each function has its own
coefficient layout and no verified one exists here for the others. A GSAS .LST refinement output has no reader and is transcribed
by hand. A GSAS .EXP, a TOPAS .inp and a FullProf .pcr are whole
refinements rather than instrument files, and have their own readers in the
next section.
Two things this reader chooses rather than reads, both on the diagnostics=
channel the sections above use. A .prm states no geometry at all, and
HTYPE PXCR spans Bragg-Brentano and Debye-Scherrer, so the Instrument comes
back debye_scherrer, which is the 11-BM capillary the corpus this reader was
built against. A flat-plate calibration must have its geometry set by the
caller, and doing that takes one care, because the geometry is not the only
thing assumed. S/L and H/L (PRCF positions 7-8) are read from the file,
and they land on geometry.axial_sl/geometry.axial_hl rather than on the
profile. A replacement Geometry built from scratch starts at 0 for both, which
is an instrument with no axial divergence at all, so copy those two across. That
matters beyond bookkeeping: Geometry.kind selects the position correction and
its suggested action, and the two geometries’ absorption corrections have
different off states. The file’s fields that the Instrument cannot carry are
dropped at their identity value and reported the same way, one per record that
carried them:
notes = []
instrument = rx.read_gsas_prm("beamline.prm", diagnostics=notes)
[(d.code, d.message) for d in notes]
# GSAS_PRM_FIELD_DROPPED: ICONS's IPOLA and refine controls, PRCF's GP …
# GSAS_PRM_GEOMETRY_ASSUMED: the geometry was not read from the file
Writing a GSAS-I .prm back¶
write_gsas_prm is that reader’s inverse, and the fourth of the
foreign-format writers. It takes a calibrated Instrument and states it as
GSAS states one: the wavelengths, the polarization, profile.u/v/w as
GU/GV/GW and profile.x/y as LX/LY multiplied back into
centidegrees, and geometry.axial_sl/axial_hl as S/L and H/L:
notes = []
rx.write_gsas_prm(instrument, "beamline.prm", header="LaB6, March",
diagnostics=notes)
Unlike the three structure writers, the refine flags are deliberately not the
payload. An instrument-parameter file is a beamline calibration rather than a
starting guess, which is why both readers here hand one back frozen, so the
PRCF header’s flag columns are left blank as a real calibration file leaves
them. GSAS_PRM_FIELD_NOT_WRITTEN names what this instrument carries that the
format cannot state, the geometry always among it. GSAS_PRM_VALUE_NARROWED
names what did not fit: a .prm is read by column and a column is the budget,
so a converged calibration’s full-precision numbers are written to what the
field holds.
One value is refused rather than reported. ICONS’ ZERO field is the one
number here whose unit this package has not established, and read_gsas_prm
refuses a non-zero one on the way in for that reason, so writing a guess would
make a file this package will not read back and a ZERO wrong by 100× puts
every peak in the wrong place. Set zero_shift to 0 and let the receiving
program refine it; a zero shift belongs to the mount rather than to the
goniometer. A third emission line and a second line whose weight is not a
KRATIO are refused for the same reason: each would make a file this
package’s own reader declines. A neutron source is refused for a different
one. The reader takes a type-3 HTYPE PNCR file, but what such a file’s POLA
and KRATIO should hold, for a source that has neither, is spelled by one real
file only, and that is a reading rather than a convention to write by.
Pass no list and the read is silent and identical, so the channel is opt-in rather than a behaviour change.
Reading a GSAS-II .instprm file¶
GSAS-II keeps a project in the binary .gpx the next section opens, and ships
a calibration as a small text file instead. read_gsas2_instprm reads one into
a frozen Instrument, on the same contract as the two readers above:
notes = []
instrument = rx.read_gsas2_instprm("beamline.instprm", diagnostics=notes)
It reads the constant-wavelength types, PXC (X-ray) and PNC (neutron), and
a neutron file comes back with a NeutronSource. U, V and W are
centidegrees squared and X and Y centidegrees, converted through the same
factors the .gpx reader uses. Zero is already in degrees, and it means what
instrument.zero_shift means: a constant added to the calculated 2θ. SH/L is
GSAS-II’s combined (S+H)/L, so it is split evenly into axial_sl and
axial_hl, which is the symmetric Finger-Cox-Jephcoat reading and the only one
a single number admits.
Several banks in one file are a selection rather than a default: pass
bank=, as #Bank n: states it, or the read refuses and names the numbers it
found. Refused by name as well: a time-of-flight or energy-dispersive Type,
a non-zero Z (this package’s profile is u, v, w, x, y exactly, and Z is
zero in all 95 constant-wavelength histograms of GSAS-II’s own tutorial
corpus), a non-zero Azimuth, which changes what Polariz. means, and a
negative U, W, X or Y. That last one is the ordinary case rather than a
corner. GSAS-II bounds none of its width coefficients, this package’s are
bounded at zero through a softplus, and two of the four constant-wavelength
files in that corpus converged to a negative X.
The geometry comes from Diff-type where the file states one. Where it does
not, the reader falls back the way GSAS-II’s own does (a stated doublet is
Bragg-Brentano and anything else Debye-Scherrer) and says so with
GSAS2_INSTPRM_GEOMETRY_ASSUMED, because the choice was not read from the
file.
Writing a GSAS-II .instprm back¶
write_gsas2_instprm is that reader’s inverse and the machine half of what
GSAS-II imports. The structure half is a CIF, below.
notes = []
rx.write_gsas2_instprm(instrument, "beamline.instprm", diagnostics=notes)
The refine flags are dropped here as they are from a .prm, an .instprm
stating none at all. Two values are worth knowing about before you hand the
file to GSAS-II. axial_sl and axial_hl are written as their sum, GSAS-II
modelling one number where this package holds two, and an uneven pair is named
GSAS2_INSTPRM_VALUE_MERGED rather than quietly halved on the way back. And
GSAS-II floors SH/L at 0.002 when it evaluates a profile, so a smaller sum is
written faithfully and still modelled as the floor by the program the file is
for; that is GSAS2_INSTPRM_VALUE_FLOORED, and no round trip through this
package can show it.
Refinement files another program wrote¶
A TOPAS .inp, a FullProf .pcr and their kind are neither patterns nor
structures. They state someone’s whole refinement: the phases, the instrument,
and which parameters were free. That last part is the one nobody can reconstruct
from a CIF plus a pattern. Transcribing one by hand is the failure mode this
reader exists to remove, because a mistyped coordinate stays symmetry-valid and
fails silently, so what comes out is a plausible wrong answer rather than an
error.
Provisional
The foreign-refinement readers are under active development, so the names in
this section are documented and not frozen. read_project_model,
identify_project_format, read_topas_inp, read_fullprof_pcr,
read_gsas_exp, read_gsas2_gpx, write_topas_inp, write_fullprof_pcr,
write_gsas_exp, write_gsas2_phase_cif and
the per-format models they answer with (rietx.io.projects) may change in a
1.x release: the registry has four formats and one more queued, each of which
is evidence about its shape, and all four can now be written. A format’s own
model mirrors that format, so its fields move when the reader’s coverage does.
Provisional by declaration has the promise in full.
Opening one¶
read_project_model takes a file and works out which program wrote it:
model = rx.read_project_model("refined.inp")
model.format.name # 'topas_inp'
model.stated.r_wp # the run's own figure, as the file states it
structure = model.to_structure()
Dispatch is on content, never on the extension. .inp is written by several
unrelated programs, and claiming one by name is how a reader returns a plausible
wrong model for a finite-element deck. A file nothing here reads is refused by a
message naming what this build does open, and pointing at the three neighbouring
readers: rx.read_pattern for a pattern, rx.read_gsas_prm for a GSAS-I
instrument file, rx.read_recipe for a PowderLine recipe. Those are what such a
file usually turns out to be.
Call rx.read_topas_inp, rx.read_fullprof_pcr, rx.read_gsas_exp or
rx.read_gsas2_gpx directly when you already know what you have; rx.identify_project_format answers which format claims a
file without parsing it, reading only enough of the head to decide.
Writing one back¶
rx.write_topas_inp(structure, path), rx.write_fullprof_pcr(structure, path) and rx.write_gsas_exp(structure, path) are each format’s inverse of
its own reader +
ProjectModel.to_structure. All three write a file that states the same
phases, the same cell and the same atoms, plus the part a CIF cannot carry:
the same refine flags, each Parameter.vary written in the target format’s
own grammar (TOPAS’s @/!; FullProf’s codeword, one free parameter to one
codeword number, never a shared tie; GSAS’s Y on the cell record and its
X/U/F letters on a site’s).
rx.write_topas_inp(structure, "exported.inp")
back = rx.read_topas_inp("exported.inp").to_structure()
rx.write_fullprof_pcr(structure, "exported.pcr")
back = rx.read_fullprof_pcr("exported.pcr").to_structure()
rx.write_gsas_exp(structure, "exported.EXP", title="my experiment")
back = rx.read_gsas_exp("exported.EXP")
All three write space groups from get_spacegroup(...).xhm(), never a
phase’s own stored spelling, so a setting this build already resolved is not
laundered back into an ambiguous symbol. The TOPAS writer spells an ion
sign-first (Zr4+ → occ Zr+4, O2- → occ O-2), because TOPAS stops on
the IUCr order rietx stores. An isotope keeps its leading mass number, which is
TOPAS’s order too (7Li, and 7Li1+ → occ 7Li+1). It refuses an ion with a sign and no magnitude
(Cu+): rietx reads that as the neutral atom, and TOPAS reads Cu+1 as the
ion. The FullProf writer spells a species by the radiation, because FullProf
looks the two tables up differently: X-rays key the form factor on element,
sign and magnitude (Zr4+ → ZR+4, O2- → O-2), and neutrons key b on
the element alone, so the charge is not written (O2- → O; FullProf stops on
O-2 there). An isotope is written only on a neutron file, as a LINE-12 user
scattering length (7Li → Typ LI7 with LI7 -0.222 0.0 0), because a bare
LI7 runs at natural abundance with no message. An isotope on an X-ray file,
Cu+ on an X-ray file (a neutron file writes it as CU, Cu’s b in both
programs), and an isotope whose name would not fit LINE 12’s four characters
(157Gd) are refused by name. FullProf’s grammar has no origin or
axis suffix at all, though, so it can only state a setting its own
bare-symbol convention already prefers (root CLAUDE.md’s “an R lattice on
rhombohedral axes” and “choice 2 wherever the bare symbol lands on choice
1”). A phase whose resolved setting disagrees, most commonly origin choice 1,
is refused by name instead of being silently written as the other setting.
TOPAS and GSAS can both spell every setting, so neither owes that refusal. An
anisotropic site is refused by the FullProf and GSAS writers, because
to_structure refuses to assume a displacement-tensor convention on the way
in and writing one out would assume the very thing the reader declines to
read back. A magnetic phase, one carrying magnetic_symmetry or a site
moment, is written by write_topas_inp (see
A magnetic phase from a TOPAS .inp) and refused by name by
the other three writers, write_gsas2_phase_cif included, because none of
their target programs would read the moments back (GSAS-II 5.6.3’s CIF import,
measured, drops the magCIF loops without a warning); Structure.to_cif writes
the magCIF. All four writers refuse a phase carrying a propagation vector
(Phase.propagation_vector), since none of their files would read back with
its k. No CIF this package writes carries a nuclear phase’s k either, so
the Structure’s own JSON (Structure.model_dump_json) is the export for one.
A .EXP is the one target read by column rather than by token, and two
things follow from that. A number is worth as many characters as its field
has, so a value whose shortest exact decimal fits ten columns crosses
untouched and one that does not is written to the precision the field holds
and named, GSAS_EXP_VALUE_NARROWED, once for the file. The displacement is
always in that second class: GSAS stores Uiso and rietx stores Biso, so
the 8π² between them turns a two-character Biso into a Uiso no ten-column
field can hold exactly. Ten columns keep it to about seven significant
figures, five orders below the esd a refinement reports, and two figures
better than the six decimals GSAS itself writes. The other thing is that GSAS
states one refine flag for a whole cell and one for a site’s three
coordinates, both meaning “refine as symmetry permits”. A structure whose six
or three disagree is written free and the group is named,
GSAS_EXP_REFINE_FLAG_MERGED, so the file still says the cell refined and
you learn which held parameter it frees. Pass diagnostics=[] to collect
both. A species is written as the manual’s atom type, aasv_nnn: the symbol
in upper case, the valence sign-first with its number, then _ and the isotope
number (Zr4+ → ZR+4, 7Li1+ → LI+1_7, D → H_2); Cu+ is refused.
That spelling is from the manual only, since no GSAS run has checked it. The
reader takes GSAS’s NI+2_58 back as 58Ni2+, the ⁵⁸Ni isotope of the ion.
Some things do not travel, because no to_structure builds them from its
file. Common to all three: the emission profile and instrument geometry
(Instrument is not part of what any of these readers returns), and cell and
site bound windows. FullProf-specific: the fitted 2θ range, the background and
every control/output switch on a .pcr are protocol to_structure never
reads into a Structure. write_fullprof_pcr fills them with safe, inert
placeholders purely to keep the file complete, since a .pcr is positional
and every line the reader expects has to exist even where a Structure
carries nothing for it. The widths and the radiation are the exception,
because FullProf will not run a file without them: pass
write_fullprof_pcr(structure, path, instrument=instrument) and the file
states the instrument’s wavelengths (Job = 1 for a neutron source) and its
widths as FullProf’s TCH pseudo-Voigt (Npr = 7), each phase’s own sample
broadening added and FullProf’s X/Y letters swapped onto rietx’s strain/size
terms. Without an instrument it writes Cu Kα1/Kα2 and a default
ProfileTCHZ()’s widths; zero widths are a FullProf hard stop. An X-ray file
also states rietx’s own anomalous dispersion, one LINE-12 nam f′ f″ 2 per
Typ (an override or dispersion=None, which this module’s reader would refuse, is refused at write): without it FullProf applies its own f′/f″,
measured against FullProf 8.20 to differ from rietx’s away from Cu Kα. A
profile term FullProf’s file would not state the same way is refused by name:
an exact Voigt shape, a Stephens strain block, axial divergence, harmonics, a
third emission line, a first line of weight 0, a negative second weight, and a
non-finite wavelength, weight or width. So is every term the file writes only
at its identity, wherever it is away from it: a zero shift, sample
displacement or transparency, a capillary offset, absorption, surface
roughness, a polarisation other than K = 0.5, a declared hump or peak,
preferred orientation (r ≠ 1), extinction, and a partly occupied site. A written .EXP states no
histograms at all, which is what a GSAS experiment file looks like before any
data is loaded rather than an omission.
GSAS-II is the fourth target and the one with no project file to write.
It imports a phase from a CIF and a machine from an .instprm, so the pair is
what rx.write_gsas2_phase_cif(structure, path) and rx.write_gsas2_instprm
produce:
rx.write_gsas2_phase_cif(structure, "exported.cif")
back = rx.Structure.from_cif("exported.cif")
Its atom tags are the ones GSAS-II’s own importer reads, so the file is the
package’s ordinary structure block with the symmetry stated three times. That
is not belt and braces. GSAS-II resolves a bare two-origin symbol such as
F d -3 m to origin choice 2, and calls choice 1 a setting not compatible with
it; gemmi, and so this package, resolves the same string to choice 1. No single
symbol satisfies both, so each tag carries the spelling its own reader takes:
_symmetry_space_group_name_H-M the bare symbol, which is the tag GSAS-II
reads first and the only grammar it accepts, _space_group_name_H-M_alt the
resolved xhm(), which gemmi prefers when both are present, and
_space_group_symop_operation_xyz the operations themselves, which need no
convention at all. GSAS-II checks its own reading of the symbol against those
operations and offers to transform a structure that disagrees. A phase whose
symbol is ambiguous is named GSAS2_CIF_SETTING_IN_OPERATORS. What a CIF
cannot state (the refine flags, the phase scale, the sample broadening) is
named GSAS2_CIF_FIELD_NOT_WRITTEN. An isotope other than deuterium is
refused by name, and so is Cu+. GSAS-II keeps an isotope as a per-type
choice in the phase’s General data, which no CIF tag reaches. Its importer
reads 7Li as hydrogen and Cu+ as carbon, and says so only on stdout. Write
the element, then choose the isotope in GSAS-II. 2H is written D, the one
isotope label GSAS-II takes, and every site written D is named
GSAS2_CIF_DEUTERIUM_AS_D: GSAS-II’s D has b = 6.681 fm, against the
6.671 fm (Sears) rietx uses.
What comes back¶
The answer is a ProjectModel. It names the format and hands on that format’s
own model untouched. It is deliberately not one shape with blanks where a file
was silent, because a blank reads as an answer and a caller cannot tell “this
file carried none” from “the reader found none”.
Field |
What it holds |
|---|---|
|
the |
|
the file that was read |
|
what the file stated, in that format’s own model |
|
what the read repaired or assumed, kept whether or not you asked |
|
build a |
ProjectModel.to_structure passes its keywords through to the format’s own
conversion, so dataset= reaches a multi-dataset .inp and nuclear_only=
reaches a .pcr with a magnetic phase. They are not flattened into one
vocabulary: two formats’ options that happened to share a name would not share
a meaning.
What each format states¶
A model is what the file said, seeded with nothing. A value the file omitted
arrives as None and never as a default. The two differ because the formats
do.
read_topas_inp returns a TopasModel:
Field |
What it holds |
|---|---|
|
the file it was read from |
|
the phases the file states, with their sites and refine flags |
|
the run’s own figures, where the grammar settles which is the file’s |
|
every such value in file order, since a multi-dataset file states one per dataset |
|
how many |
|
the pattern files it points at |
|
the emission profile, as stated or as a macro named it |
|
the diffractometer, where the file says |
|
how many background coefficients were refined |
|
phase blocks that stated no name, or neither a |
|
what the reader met and did not carry; see below |
|
the datasets that are neutron time of flight, each with the constructs that said so; |
read_fullprof_pcr returns a FullProfModel:
Field |
What it holds |
|---|---|
|
the file, and the names it gives itself |
|
every phase, nuclear and magnetic |
|
the same, split; a magnetic phase always reads, and builds only as the moments of the nuclear phase it restates |
|
the converged figure, from the comments FullProf rewrites each cycle |
|
the Job/Npr/Nph control line, field by field |
|
which diffraction experiment the file declares |
|
the wavelengths, and which slot stated them |
|
the zero correction, with its refine codeword decoded |
|
the background the file declares |
|
what the run fitted and what it left out |
|
a neutron file’s LINE-12 user scattering lengths, each read as the isotope it is the Sears b of |
|
an X-ray file’s LINE-12 f′/f″ per element as stated; a pair differing from rietx’s Cromer-Liberman value at the file’s wavelength is refused at read |
|
the pattern it points at |
|
how the run was driven, and which codeword numbered which parameter |
|
the output options the file sets |
read_gsas_exp returns a GsasModel. A GSAS experiment file holds the whole
experiment rather than one phase, so the model is a tree: phases, histograms,
and one entry per phase-and-histogram pair, which is where GSAS keeps the
profile coefficients.
Field |
What it holds |
|---|---|
|
the file, its |
|
every phase the file states |
|
every histogram, with its machine and the pattern it points at |
|
the phase-and-histogram entries, keyed |
|
the run’s own agreement indices |
|
GSAS’s |
|
how many parameters the run refined, and over how many channels |
|
one line per thing the read could not carry, kept whether or not you asked |
|
pick one histogram, or one pair’s refined profile; both refuse to choose when the file carries several and you named none |
A phase is GsasPhase:
Field |
What it holds |
|---|---|
|
its |
|
the symbol, as written |
|
the lattice, its esds and its one refine flag |
|
the sites |
|
GSAS’s phase type, and whether it is one of the two magnetic ones |
|
the |
Its lattice is a GsasCell:
Field |
What it holds |
|---|---|
|
the lattice |
|
their esds, where the run produced them |
|
the cell volume |
|
GSAS refines the cell as metric tensor elements under one flag, so there is one here rather than six |
Each of its sites is a GsasAtom:
Field |
What it holds |
|---|---|
|
the name you chose and the scattering species, in the file’s own upper case |
|
the site |
|
the isotropic displacement in Ų, or |
|
the six anisotropic values, or |
|
the |
|
the per-group damping codes |
A histogram is GsasHistogram:
Field |
What it holds |
|---|---|
|
its ordinal and GSAS’s |
|
that type, read for you |
|
the pattern and |
|
the detector bank, and the anode |
|
the emission lines, Kα1 first |
|
the zero correction and its flag |
|
the polarization fraction and its type code |
|
the Kα2/Kα1 ratio, or |
|
what the run left out, GSAS’s spacer pair dropped |
|
the range the data covers |
|
how many points were fitted, of how many recorded |
|
the histogram scale and its flag |
|
the background function and its coefficients |
|
the profile sets the instrument file supplied, which is where the histogram started |
|
this histogram’s own indices |
A phase-and-histogram pair is a GsasHap:
Field |
What it holds |
|---|---|
|
which pair this is |
|
the phase fraction and its flag |
|
the refined profile for this pair |
|
the March-Dollase rows, each with its axis and flag |
|
the powder extinction coefficient |
Its profile is a GsasProfile of GsasProfileTerm rows:
Field |
What it holds |
|---|---|
|
which GSAS profile function, over how many coefficients |
|
the coefficients, named |
|
the same, keyed by GSAS’s name for each |
|
the ones the run was free to move, which is the protocol |
|
GSAS’s name for the coefficient and its 1-based position |
|
the file’s own number, and its flag |
|
the same quantity in degrees, or |
The background is a GsasBackground:
Field |
What it holds |
|---|---|
|
which of GSAS’s background functions, over how many terms |
|
the terms themselves |
|
whether the run refined them |
Three things about a .EXP decide how you read one.
Coefficient names depend on the profile function¶
GSAS’s constant-wavelength
function 2 has LX fourth and function 3 has GP fourth, so the names come from
the function the file declares. A function this build cannot name is refused
rather than read positionally, because a mis-assigned width is a plausible wrong
answer rather than an error.
The coefficients are in centidegrees¶
GsasProfileTerm.degrees is the
conversion, and it is None for the L11…L23 anisotropic terms, whose unit
also involves d². None there means no conversion was made, not that the
value is zero.
A blank field is not a zero¶
GsasHistogram.ka2_ratio is the case that
catches people: both a polarization fraction and a Kα2/Kα1 ratio are
conventionally 0.5, and they sit in adjacent fields on the same record. Where
the file leaves the ratio blank this reads None, and choosing a ratio is then
yours to do explicitly.
The converse holds too. A field that is written but is not a number is not
blank, and the read is refused rather than returning None. The message names
the file, the record key, the card columns and the text. A record like that is
misaligned or corrupt, so nothing on it can be trusted by column. The .prm
reader shares the record grammar and refuses the same way. A field of nothing
but asterisks is the exception with a known cause: it is Fortran’s output for a
value too wide for its field, so the writer had the number and ran out of
columns. In a required field, such as a cell edge or a .prm’s LAM1, the
read is refused and the message says overflow. In an optional one, such as an
esd or the history Rwp, the field reads as None and the reader reports
GSAS_FIELD_OVERFLOW on its diagnostics= list, naming the record and the
columns. One case no
per-field rule can catch: a number that spills across a field boundary and
leaves two readable halves, such as a left-aligned 46.60 read as 46 and
.60.
read_gsas2_gpx returns a Gsas2Model. A GSAS-II project is the same shape of
tree as a .EXP: phases, histograms, and one entry per phase-and-histogram
pair. It carries two things the older format has no room for. One is the set of
constraints the refinement ran under. The other is the list of variables it
actually refined.
Field |
What it holds |
|---|---|
|
the file you opened, where GSAS-II last saved it, and the revision that wrote it |
|
every phase, and every powder histogram |
|
the phase-and-histogram entries, keyed |
|
the variables the last refinement refined, in GSAS-II’s |
|
the equivalences, holds and constraint equations, with their variables resolved into names |
|
the run’s own agreement indices. |
|
how many channels, how many parameters, and whether GSAS-II called it converged. |
|
one line per thing the read could not carry, kept whether or not you asked |
|
pick one by name or by number; both refuse to choose when the file carries several and you named none |
A phase is a Gsas2Phase:
Field |
What it holds |
|---|---|
|
the name GSAS-II gives it, and its |
|
the symbol, as written |
|
the setting the project’s own operators describe, and the symbol to build under. The second is the first where the file settled one, and |
|
the six lattice parameters and the cell volume |
|
GSAS-II refines a cell under one flag, so there is one here rather than six |
|
the sites |
|
the isotope chosen per atom type, GSAS-II’s |
|
GSAS-II’s phase type, and whether it is a magnetic one |
|
GSAS-II’s |
|
whether the phase was fitted by Pawley extraction, in which case its refined quantities are intensities rather than sites |
Each of its sites is a Gsas2Atom:
Field |
What it holds |
|---|---|
|
the name you chose and the scattering species, in the file’s own spelling. GSAS-II writes an ionic charge sign first, as |
|
the site |
|
what GSAS-II worked out about the position |
|
|
|
the isotropic displacement in Ų, and the six anisotropic values |
|
the letters the file carries, such as |
|
those letters read one at a time |
A histogram is a Gsas2Histogram:
Field |
What it holds |
|---|---|
|
its |
|
GSAS-II’s type code, such as |
|
the instrument coefficients, and the same keyed by GSAS-II’s name for each |
|
the emission lines, primary first |
|
the range the refinement used, and the range the file was loaded with. Two numbers rather than one, because that pair is what makes a narrowed range visible |
|
ranges masked out of the fit |
|
the background function and its coefficients |
|
the refinable sample parameters with their flags, keyed by GSAS-II’s own names ( |
|
|
|
how many channels the file carries |
An instrument coefficient is a Gsas2Term:
Field |
What it holds |
|---|---|
|
GSAS-II’s own name, such as |
|
the refined number and the one it started from, which the format keeps side by side |
|
whether the run was free to move it |
|
the same quantity in degrees for the six centidegree coefficients, and |
The background is a Gsas2Background:
Field |
What it holds |
|---|---|
|
GSAS-II’s own name for it, such as |
|
the terms, and whether the run refined them |
|
how many diffuse terms and fitted background peaks sat beside them |
A phase-and-histogram pair is a Gsas2Hap:
Field |
What it holds |
|---|---|
|
which pair this is |
|
the phase’s scale in this histogram, and whether the pair was in the fit at all |
|
the sample broadening, which GSAS-II refines per pair |
|
the hydrostatic-strain terms |
|
the model, its value, its flag and its axis |
|
the extinction coefficient |
|
whether this pair’s intensities were extracted rather than calculated |
Size and microstrain are each a Gsas2Broadening:
Field |
What it holds |
|---|---|
|
|
|
the direct numbers with their flags: a size in µm, a microstrain in units of 1e-6 |
|
the unique axis a uniaxial model uses |
|
the hkl-dependent coefficients. The format writes this row whatever the model, so on an isotropic one it is the unused default rather than a second opinion |
A constraint is a Gsas2Constraint:
Field |
What it holds |
|---|---|
|
the format’s own letter: |
|
which variables it constrains: |
|
|
|
the right-hand side of a constraint equation, and whether a new variable is refined |
|
the one shape |
Anything with a flag beside it is a Gsas2Value, which is Gsas2Value.value
and Gsas2Value.refined and nothing else.
Three things about a .gpx decide how you read one.
It is a pickle, so the reader refuses names rather than running them¶
A .gpx is a sequence of python pickles, and loading a pickle runs whatever it
names. So this reader resolves an allow-list and refuses every other name,
naming it, without reading the file at all. That covers GSAS-II’s own two
classes too. They are admitted as inert stand-ins that can hold data and do
nothing, which is why a project with constraints reads at all. If you are handed
a .gpx from somewhere you do not trust, this is the property to know about:
the refusal message names the global, and a file that opens has named nothing
outside the list.
The profile coefficients are centidegrees and the zero correction is not¶
Gsas2Term.degrees is the conversion, and it exists for U, V, W, X, Y
and Z only. Zero is already in degrees, which is where the two GSAS
generations differ from each other, so a None there means “no conversion was
needed” rather than “no conversion was made”.
What to_structure refuses, it refuses by name¶
A magnetic phase, a macromolecular one, a Pawley phase, a phase with
anisotropic sites, a phase with no sites at all, and a phase whose stored Uiso
is negative. The last two are what a corpus of real projects contains and a
specification never mentions: GSAS-II fits a phase with no sites by extracting
its intensities, and a refinement really can end with a negative displacement
parameter. Each refusal names the phase and leaves the numbers on the model,
because deciding what a negative Uiso should have been is yours. An isotope choice that is not a mass number, or
names a nucleus the Sears table rietx scatters from does not carry, is refused
the same way, since reading it as natural abundance would change the site’s
scattering length without saying so.
Which formats this build reads¶
rx.capabilities() publishes the registry, so a client asks rather than
guesses:
import rietx as rx
caps = rx.capabilities()
[(f.name, f.extensions, f.reports_at) for f in caps.project_formats]
Field |
What it holds |
|---|---|
|
one |
|
the registry key, and what a person calls it |
|
the conventional suffixes, which are informational and never the dispatch |
|
how the format is recognised, in words |
|
what the file holds beyond a structure; read this before reading a model you have no common shape for |
|
|
|
the top-level verb that writes this format, or |
|
set when the build recognises a format in order to decline it, carrying why |
reports_at is a real difference and not bookkeeping. A .inp’s repairs, such
as a rewritten species spelling or a translated origin suffix, happen while
parsing, so its channel is read_project_model(..., diagnostics=notes). A
.pcr’s all happen while codewords become a Structure, so its channel is
ProjectModel.to_structure(diagnostics=notes). A .EXP repairs at both ends
and declares "both". It reports the histograms it could not carry while
parsing, and the species it normalises while building. Passing a list to both
calls collects any of the three without your having to know which.
Passing your list to the wrong call still does not lose the read’s half.
Whatever the read reported is on ProjectModel.diagnostics as well, filled list
or no list. That matters more than it sounds. An empty list reads as “this file
needed no repairs”, and a caller who had passed it to the other call would
believe it.
writes is a name rather than a flag, so a client learns what to call and not
only that something is possible. It is None where this build has no writer,
which is .gpx today, and None rather than False on purpose: a False
would be an answer about a format nobody had wired.
The same facts are in the registry the arm is built from. ProjectFormat.name,
ProjectFormat.title, ProjectFormat.extensions, ProjectFormat.sniff,
ProjectFormat.carries, ProjectFormat.reports_at and ProjectFormat.refuses
carry the declarations, while ProjectFormat.matches, ProjectFormat.read,
ProjectFormat.to_structure and ProjectFormat.write are the callables the
dispatch and the writers use.
What a reader will not guess¶
These formats hold constructs this package has no model for, and dropping one
changes the model rather than its presentation. So every construct the TOPAS
reader knows about is a Feature carrying a declared Stance, a keyword with
no stance fails a test instead of vanishing, and what a particular file turned
out to contain comes back as a Coverage of Hit rows. A construct written
as a library macro (TCHZ_Peak_Type(…), CS_L(…),
Specimen_Displacement(…)) is met as well as one written in keywords, inside
a phase or at dataset level:
from rietx.io.projects import coverage
Name |
What it is |
|---|---|
|
carried into the model |
|
deliberately not carried, and harmless |
|
met, not carried, and named in the diagnostics |
|
met, and the file is refused rather than returned half-built |
|
one construct, what it is and the argument for its stance |
|
the stance, and the format keywords it governs |
|
the format’s library macros that state the same construct, met wherever the file invokes one |
|
one construct actually met in a file, and where |
|
the hits of each kind |
|
those hits in a sentence |
A refused construct raises, naming the file and the line. A reported one arrives on the diagnostics channel, and without a list the read is silent and identical, on the same opt-in terms as every other reader here.
The .rex project directory¶
A project is the one durable thing a session can point at. It is a directory rather than an archive:
my_sample.rex/
project.json settings and the data reference
11BM_NAC.fxye the pattern file, byte for byte as measured
history.jsonl the refinement DAG, append-only
live/ one directory per run, for `rietx watch`
20260915-092640-80245/
20260915-101302-80245/
exports/ CIFs, reflection tables, QPA tables
live/ holds one directory per run. Two writers on one project, a GUI session
and an agent fitting the same directory, would otherwise interleave their events
into a single file.
A directory because the history log’s crash safety is append-only writes by one writer. Zipping would force a rewrite on every save and lose exactly the property that makes a JSONL log recoverable and tailable while a fit runs.
project = rx.Project.create("my_sample.rex", pattern="my_sample.xye",
structure=structure, instrument=instrument)
project.fit()
project.save()
project = rx.Project.open("my_sample.rex")
Project.create builds the directory around a pattern file and a model.
Project.open reads one back and resumes at the history head. It re-checks
every binding on the way rather than assuming it: the pattern file is still
there, its bytes still hash to the recorded digest, this build still parses those
bytes to the recorded numbers, and the history was recorded against that same
pattern. Each of the four raises with its own message, because each has a
different cause and a different fix.
A project with no structure¶
structure= may be left out. What you get is a pattern-only project: zero
phases, for a pattern whose phase you do not know yet.
project = rx.Project.create("unknown.rex", pattern="unknown.xye",
instrument=instrument)
Peak picking and indexing work over it, the parameter table holds the instrument
and the background, and the routes out of it both end in a phase: index the
peaks and adopt a candidate cell (Indexing an unknown cell), or build a Le Bail scaffold
from a space-group symbol and the cell parameters that setting leaves free
(rietx.schemas.structure.lebail_scaffold). In the GUI it is the third answer
to the wizard’s structure step.
What it cannot do is refine. A phase reaches the pattern only through
scale × |F|² × profile, so with no phase there is nothing but the background
to fit, and a plan run over one converges on the background and reports
success. fit, run_stage, refine_multi and refine_sequential therefore
raise NoPhasesError before they start. NoPhasesError.code is the string
NO_PHASES on every surface: the agent envelope’s fourth error code, and the
GUI’s reason for a disabled Run button.
import rietx as rx
print(rx.NoPhasesError.code)
NO_PHASES
Refinement.predict still works, because evaluating the background as it stands
is not a refinement.
An open project holds the session as seven attributes.
Attribute |
Holds |
|---|---|
|
the project directory |
|
the |
|
the |
|
the |
|
that refinement’s |
|
what the reader repaired or assumed on the last read |
|
what reading the history log repaired, on the last open |
Project.data_diagnostics is held in memory and is not a project.json field.
The repairs are a function of the bytes, the reader and its options, and the
data reference below already records all three. Project.history_diagnostics
is held the same way, for the same reason: it is a function of the log and of
the release reading it. The history log below says what it can hold.
Every verb that changes a project writes into its directory as it runs, and
Project.open appends a line to the log before any verb is called. There is
no read-only way to open one. To look at a project without changing it, open a
copy: rietx gui PROJECT.rex --scratch copies the directory to a temporary
one and opens that, printing where it went (The command line). The copy is
byte-for-byte, so it opens exactly as the original does.
Each fact has one authority. project.json holds the settings: the selected
plan and mode, the 2θ limits, the excluded regions, and the GUI’s own ui keys.
history.jsonl holds the model state, and its head is the working state. No
parameter value is written in both places.
Saving persists settings, and is not what makes the work durable. Every verb
that changes the model commits a history node the moment it runs, so the work is
on disk whether or not anyone calls Project.save. What save persists is the
half of a session that nothing else owns.
ProjectDoc is that half, field by field.
Field |
Holds |
|---|---|
|
the data references, one per pattern |
|
the plan the next run will use, as a |
|
the intensity mode the next run will use |
|
the range the next run will fit |
|
2θ regions masked out of the residual |
|
the next indexing run’s controls |
|
the log’s filename inside the directory |
|
the version of the |
|
the version of rietx that wrote it |
|
when it was created, and last saved |
|
keys a front end persists, untyped |
The three settings after the plan are what fit and run_stage will be called
with. A history node records a mode and limits too, and that is a different
fact: the node says what a past run used, the document says what the next one
will use. Before the first run there is no node to ask.
ProjectDoc.patterns is a list because stacking several patterns into one joint
residual is a later milestone’s work. A project holds one today, and
Project.open refuses a document carrying more rather than opening the first
and looking like it worked.
ProjectDoc.ui is untyped deliberately. A front end owns those keys, and the
container only stores them, so a layout change is not a schema change.
The pattern is copied verbatim rather than re-serialised, because the bytes are
the contract: the reader takes σ from the file’s own column and never overrides
it. Project.data_ref returns the DataRef that makes those bytes trustworthy
on re-open. It carries DataRef.sha256 of the file, DataRef.fingerprint of
the parsed arrays, and DataRef.reader with DataRef.options. The reader
call itself is part of the reference, because a pdCIF is a different pattern
depending on the block. Agreeing bytes with a disagreeing fingerprint mean the
reader changed, not that the project is corrupt.
Four more fields say what the pattern is: DataRef.filename names it inside the
directory, DataRef.n_points and DataRef.two_theta_range describe it, and
DataRef.has_sigma records whether σ was measured or fell back to Poisson. That
last one is a correctness property of every fit in the project and is invisible
once the data are read, so it is written down.
Project.set_excluded_regions records regions to leave out of the fit. They
live in the document rather than in a history node because they are protocol
that is in neither the file nor the model: a node cannot say what was excluded
when it ran. Project.fitted_mask is the one authority for which channels the
next run fits. An inverted or empty interval is refused rather than reordered.
The two settings compose. On the 11-BM pattern of A first refinement, limits of 2–24° leave 22 003 of 59 498 channels in the residual, and excluding 7.4–7.6° as well leaves 21 803.
Project.exports_dir and Project.live_dir are where the last two directories
live, and Project.parameters, Project.fit and Project.run_stage are the
session verbs, with the same meaning they have on Refinement.
Run directories¶
A fit that did not come from a project records under .rietx/runs/ in the
working directory, one directory per fit, named for the date, the time and the
process id:
.rietx/runs/
.gitignore contains `*`, so runs stay out of `git status`
20260915-092640-80245/
events.jsonl one line per event
snapshot.json the stage's curves, ticks and statistics
summary.txt the termination view
meta.json label, start time, version, cwd, command
status.json state, pid, host, heartbeat, stage, Rwp, gof
run.lock held by the writing process for its life
Two directories are spelled .rietx and they have nothing to do with each
other. $HOME/.rietx is the GUI’s per-user state, the recent list and the theme,
and --state-dir or RIETX_STATE_DIR moves it. The one above is the runs root,
it is relative to wherever the process is running, and no environment variable
moves it. To put runs elsewhere, name a root per call with telemetry=
(Running a refinement).
A fit you run inside somebody else’s package leaves .rietx/ in their working
directory. That is the working directory doing its job rather than a bug, and
RIETX_TELEMETRY=0 is how a process declines.
What a run directory contains¶
Every free parameter’s value at every recorded evaluation, in events.jsonl,
alongside the stage boundaries and the statistics at each one. snapshot.json
holds the observed, calculated and difference curves for the stage, decimated
for drawing.
No pattern bytes are ever copied into a run. The curves in a snapshot are decimated to a drawing budget of 4000 points, the same budget the GUI and the comparison UI use, so a long pattern reaches a run as a picture of itself. A short one arrives nearly whole: a 4200-point pattern kept 4143.
The parameter trajectory is the part to weigh before writing runs to a shared
filesystem, and RIETX_TELEMETRY=0 is how you decline (Running a refinement).
The history log¶
history.jsonl is an append-only record of the refinement DAG, one JSON object
per line. RefinementTree.save and RefinementTree.load are the file
interface, RefinementTree.records is what gets written, and
RefinementTree.summary prints the tree. The refinement history is the DAG itself: what
a node holds, and the verbs that restore, fork and merge one.
One repair is made when a log is read. A release before a class inherited
its fields’ declared ranges (The parameter table) wrote a Parameter you had supplied
with the bounds it then carried: (−inf, inf), no unit, the identity
transform. Such a log would reopen unbounded, because every key it stores is
present and nothing is inherited from a key that is there.
RefinementTree.load restores the declared range in every node that stored
one, where the node was written before that class’s inheritance. Each node
carries the schema version it was written under (HistoryNode.schema_version),
and one written before that field is read at its log header’s version. Each
restored parameter is reported once, as a DECLARED_RANGE_RESTORED
diagnostic naming its dot-path, the atom where there is one, and both ranges.
Pass diagnostics=[] to collect them; Project.open passes one for you.
No value moves. Where any node’s value lies outside the declared range, the
parameter is left as stored in every node and reported as
DECLARED_RANGE_NOT_RESTORED, naming the node and the value. It still refines
unbounded. The declared box is not physics for every instrument: a coarse
neutron line’s Caglioti u sits well outside the X-ray one. A range that
refused the record would lose the one value it holds. A range that is merely
wider than the declared one is a choice and is left alone, and so is an
unbounded one written after the class inherited.
Each line is a HistoryRecord, a tagged union of the three things the log
carries. The tag is what keeps the file append-only: a reader branches on it
rather than on the file’s shape, so a new line never invalidates the ones before
it.
Field |
Is |
Reads as |
|---|---|---|
|
|
the tag. Branch on this and read the matching field |
|
a |
the tree’s identity and the fingerprint pinning it to one pattern |
|
a |
one state, appended when a stage or an edit commits |
|
an |
a tag or a note attached to a node afterwards, so it is a separate line rather than a field on the node |
The other three fields are null on any given line. A tag applied to a node that was written an hour ago appends an annotation line; it never rewrites the node.
A node stores state and not curves. A node is about 10 kB; embedding the calculated pattern would make it 1.24 MB. Le Bail extracted intensities are the exception, because they live outside the parameter vector and are path-dependent, so they are serialized per node.
The tree branches. A stage adds a node under the head, a model edit adds one
too, and Refinement.checkout moves the head back so the next fit forks
instead of continuing:
graph TD
root["root<br/>initial model"] --> n1["fit: lebail<br/>Rwp 0.113"]
n1 --> n2["edit: add CaF₂ phase"]
n2 --> n3["fit: rietveld<br/>Rwp 0.093"]
n1 --> n4["fit: rietveld<br/>no impurity<br/>Rwp 0.141"]
RefinementTree.to_mermaid prints that diagram for a real tree, in the same
syntax, so you can paste it into any markdown that renders mermaid:
print(ref.history.to_mermaid())
Exports¶
Three writers turn a result into a file someone else can read:
Call |
Writes |
|---|---|
|
the refined structure and fit as a CIF |
|
one row per (emission line, reflection): hkl, d, 2θ, |F|², intensity |
|
the quantitative phase analysis |
Exports and derived tables has the rows those last two are made of, and the same three
writers as methods on Refinement.
The CIF is the one a journal asks for, so it carries more than the coordinates.
Per phase it writes the agreement indices of
Reading the numbers as _refine_ls_R_I_factor (R_B), _refine_ls_R_factor_all (R_F)
and _refine_ls_number_reflns, and the geometry as _geom_bond_,
_geom_contact_ and _geom_angle loops with esds in su notation. Each
_geom_*_site_symmetry_* code indexes a _space_group_symop_ loop written
into the same block, because a code that pointed at whatever order the reader’s
own library generated would name a different atom. The loops list each bond
once, unlike GeometryTable, whose audience is a chemist counting neighbours
rather than a parser. Take structure from Refinement.fitted_structure, which
is where the refined values and their esds are.
viz.html.write_html writes the interactive page, and RefinementResult.plot
writes the static figure. The page needs nothing beyond a base install. The
figure needs the viz extra.
For agents
Capabilities.project_format_version versions the .rex directory, and
Capabilities.textdoc_format_version versions .rxt, the line-oriented text
rendering of a project that the GUI edits. Both move independently of the
package version. See Calling rietx from a program.