API Reference
- dendros.open_outputs(path, output_root='Outputs')[source]
Open a Galacticus output collection.
- Parameters:
path (str | Path | List[str | Path] | Mapping[str, str | Path | List[str | Path]]) –
One of:
A single filename – e.g.
"galacticus.hdf5". If sibling MPI-rank files (*_MPI:????) exist they are included automatically.A glob string – e.g.
"run*/galacticus*.hdf5".An explicit list of filenames.
A dict
{label: path-or-paths}to open several models for side-by-side comparison. Equivalent toopen_models(path); returns aModelCollection.
output_root (str) – Top-level HDF5 group containing the
Output*groups. Defaults to"Outputs". Pass"Lightcone"for lightcone runs or any other custom group name as needed.
- Returns:
Collection – For the single-, glob-, or list-of-files forms (one model, possibly MPI-split).
ModelCollection – For the dict form (one entry per model, suitable for multi-model comparison plots).
- Raises:
FileNotFoundError – If no files are found matching path.
TypeError – If path is not a
str,pathlib.Path,list, ordict.
- Return type:
Examples
Open a single file:
c = open_outputs("galacticus.hdf5")
Auto-detect MPI-split files (given any one rank’s file):
c = open_outputs("galacticus_MPI:0000.hdf5")
Open via glob:
c = open_outputs("run001/galacticus*.hdf5")
Open an explicit list:
c = open_outputs(["file_a.hdf5", "file_b.hdf5"])
Open several models for comparison plots (see also
open_models()):m = open_outputs({"Fiducial": "fid.hdf5", "Variant": "var.hdf5"}) figs = m.plot_analyses()
Lightcone mode:
c = open_outputs("lightcone.hdf5", output_root="Lightcone")
- dendros.open_models(models, output_root='Outputs')[source]
Open several Galacticus runs as a labelled
ModelCollection.Each entry is opened with
open_outputs(), so a single filename auto-detects its*_MPI:????peers and a list-of-files entry is accepted as-is. The returnedModelCollectioncan be passed directly toplot_analyses()to overlay analyses from every model on a single figure.- Parameters:
models (Mapping[str, str | Path | List[str | Path]] | Sequence[str | Path | List[str | Path]]) – Either a dict
{label: path-or-paths}— keys become legend labels — or a list ofpath-or-paths, in which case default labels are derived from each model’s primary file stem (with any:MPIxxxxsuffix stripped).output_root (str) – Forwarded to
open_outputs().
- Return type:
- Raises:
ValueError – If a list form produces duplicate default labels. Pass a dict with explicit labels to disambiguate.
Examples
Compare two models with explicit labels:
with open_models({"Fiducial": "fid.hdf5", "Variant": "var.hdf5"}) as m: figs = dendros.plot_analyses(m)
Or with labels derived from filenames:
models = open_models(["fid.hdf5", "var.hdf5"])
- class dendros.Collection(files, output_root='Outputs')[source]
Bases:
objectA collection of one or more Galacticus HDF5 output files.
Prefer constructing instances through
open_outputs()rather than calling this constructor directly.- Parameters:
Examples
>>> from dendros import open_outputs >>> with open_outputs("galacticus.hdf5") as c: ... c.validate_completion() ... print(c.list_outputs()) ... data = c.read("Output1", ["nodeData/haloMass"])
- validate_completion(mode='error')[source]
Check that all files report a successful completion status.
Galacticus writes a
statusCompletionattribute to the root of the HDF5 file when it finishes. This method verifies that the attribute equals"complete"for every file in the collection.- Parameters:
mode (str) –
What to do when an incomplete file is found:
"error"(default) – raiseRuntimeError."warn"– emit aUserWarningand continue."ignore"– do nothing.
- Raises:
ValueError – If mode is not one of the accepted values.
RuntimeError – If mode is
"error"and at least one file is incomplete.
- Return type:
None
- property outputs: OutputIndex
An
OutputIndexfor this collection.The index is scanned lazily on first access and then cached.
- list_outputs(format='astropy')[source]
Return a table of available outputs.
Scans all
Output*groups inside/{output_root}/and extractsoutputTimeandoutputExpansionFactorattributes. Redshift is computed as z = 1/a − 1.- Parameters:
format (str) –
"astropy"(default) returns anastropy.table.Table;"pandas"returns apandas.DataFrame;"tabulate"returns astrformatted using thetabulatelibrary.- Return type:
astropy.table.Table, pandas.DataFrame, or tabulate-formatted string
- list_properties(output, format='astropy')[source]
Return a table of datasets available in the
nodeDatagroup.The table has columns
name,dtype,shape,description(the dataset’scomment), andunits(the human-readable units description, e.g."Solar masses"; blank for dimensionless datasets).- Parameters:
- Return type:
astropy.table.Table, pandas.DataFrame, or tabulate-formatted string
- read(output, datasets, where=None, as_quantity=True)[source]
Read one or more datasets from an output group.
For multi-file collections, arrays from all files are concatenated along axis 0 before any selection is applied.
- Parameters:
output (str | int) – Output name (e.g.
"Output1") or 1-based integer index.datasets (List[str] | Dict[str, str]) – Either a list of relative dataset paths under the output group (e.g.
["nodeData/haloMass"]), in which case the same strings are used as dict keys in the return value; or adictmapping user-chosen labels to relative paths.where –
Nonereads all rows. A boolean mask array of length N_total or an integer index array selects a subset.as_quantity (bool) – When
True(default), datasets that carry aquantitystring in theirunitsattribute are returned asastropy.units.Quantityobjects. Dimensionless datasets (emptyquantity) and datasets without units metadata are always returned as plainnumpy.ndarrayobjects. Set toFalseto return plain arrays for every dataset.
- Returns:
Mapping from dataset name / label to
numpy.ndarrayorastropy.units.Quantity.- Return type:
- trace_history(ids, properties, outputs=None, *, id_dataset='nodeData/nodeUniqueIDBranchTip', on_duplicate_file_match='error', int_sentinel=-1)[source]
Trace the history of specified galaxies across outputs.
Convenience wrapper around
dendros.trace_galaxy_history(). See that function for full parameter and return-value documentation.
- list_analyses(format='astropy')[source]
Return a table of
function1Danalyses in/analyses.Convenience wrapper around
dendros.list_analyses().- Parameters:
format (str)
- plot_analyses(name=None, output_directory=None, *, show_target=True, figsize=(7.0, 5.0), dpi=120, file_format='pdf')[source]
Plot
function1Danalyses from/analyses.Convenience wrapper around
dendros.plot_analyses().
- class dendros.ModelCollection[source]
Bases:
dictA
dictmapping label →Collection, one entry per model.Returned by
open_models()and accepted byplot_analyses()so analyses from several Galacticus runs can be overlaid on a single figure for comparison.Acts as a regular dict but also supports the context-manager protocol —
__exit__closes every contained Collection.- close()[source]
Close every contained Collection.
Each close is attempted independently; if one fails the others are still closed and a
UserWarningis emitted naming the failing model so the problem is visible.- Return type:
None
- list_analyses(format='astropy')[source]
Return the union of
function1Danalyses across all models.Each row carries an extra
modelscolumn listing the labels of the models that contain the analysis (sorted, comma-separated). Per-row metadata (description, axis labels, log flags, target presence) is taken from the first model that supplies the analysis.- Parameters:
format (str)
- plot_analyses(name=None, output_directory=None, *, show_target=True, figsize=(7.0, 5.0), dpi=120, file_format='pdf')[source]
Plot
function1Danalyses overlaid across every model.Convenience wrapper around
dendros.plot_analyses()applied to thisModelCollection. Labels come from this dict’s keys; the target overlay is drawn once.
- class dendros.OutputIndex(collection)[source]
Bases:
objectIndex of all
Output*groups found in aCollection.Instances are obtained via
outputs.- Parameters:
collection (Collection) – The parent
Collection.
- table(format='astropy')[source]
Return a table of output metadata.
- Parameters:
format (str) –
"astropy"(default) returns anastropy.table.Table;"pandas"returns apandas.DataFrame.- Return type:
astropy.table.Table or pandas.DataFrame
- class dendros.OutputMeta(name, path, index, time, scale_factor, redshift, output_type)[source]
Bases:
objectMetadata for a single Galacticus output snapshot.
- Parameters:
- dendros.sfh_collapse_metallicities(dataset)[source]
Collapse a formation history over metallicity.
Collapses (sums) star formation histories over the metallicity axis.
The return type depends on how the history was tabulated:
If a shared
timeattribute is present (fixed tabulation times common to every galaxy), a fixed-length 2Dnumpy.ndarrayof shape(n_galaxies, n_times)is returned, and any empty entries are filled with zeros.If the history was tabulated with the
fixedAgesmethod (typically used for lightcone outputs), each galaxy is tabulated at a fixed set of ages relative to its lightcone-crossing time. Ages that precede the Big Bang are dropped, so galaxies crossing earlier retain fewer bins. Since the ages form a fixed, nested set, the collapsed histories are right-aligned (the crossing-time bin is last) into a non-ragged 2Dnumpy.ndarrayof shape(n_galaxies, n_ages), front-padded with zeros where bins were dropped. Columnjcorresponds to the same tabulation age across all galaxies; usesfh_times()to recover the per-galaxy times for each column.Otherwise (variable per-galaxy times with no fixed-age structure), a list of 1D
numpy.ndarrayobjects is returned.
- Parameters:
dataset (DatasetProxy) – The dataset containing the star formation history data.
- dendros.sfh_times(dataset)[source]
Return times associated with a star formation history.
Returns None if no times are associated with this star formation history.
The return type depends on how the history was tabulated:
If a shared
timeattribute is present, a 1Dnumpy.ndarrayof the tabulation times (common to every galaxy) is returned.If the history was tabulated with the
fixedAgesmethod (typically used for lightcone outputs), the times differ from galaxy to galaxy and are stored in a companion...Timesdataset. These per-galaxy times are right-aligned (the crossing-time bin is last) into a non-ragged 2Dnumpy.ndarrayof shape(n_galaxies, n_ages), front-padded withNaNwhere bins were dropped. This matches the alignment of the array returned bysfh_collapse_metallicities(), so columnjof both arrays refer to the same tabulation bin. Note that, because each galaxy crosses the lightcone at a different cosmic time, a given column holds a different absolute time for each galaxy (but the same lookback age relative to crossing).Otherwise,
Noneis returned.
- Parameters:
dataset (DatasetProxy) – The dataset containing the star formation history data.
- dendros.trace_galaxy_history(collection, ids, properties, outputs=None, *, id_dataset='nodeData/nodeUniqueIDBranchTip', on_duplicate_file_match='error', int_sentinel=-1)[source]
Extract per-galaxy property histories across Galacticus outputs.
Galaxies are traced across
Output*groups via an integer branch-tip identifier (usuallynodeUniqueIDBranchTip) that is constant over time for a given object and unique within a single HDF5 file. For each requested property and each chosen output, this function locates every requested ID in every file of the collection, assembles the per-galaxy slice, and stacks the results along a trailing “output” axis.Slots where a galaxy is absent at a given output are filled with
numpy.nan(floating-point properties, and thetime,redshiftandexpansion_factormetadata arrays), withint_sentinel(integer properties), or withFalse(boolean properties). The returnedpresentmask is the canonical indicator of presence/absence and should be preferred to sentinel checks.- Parameters:
collection (Collection) – An open
Collection.ids – Array-like of integer
nodeUniqueIDBranchTipvalues to trace. Coerced tonumpy.ndarrayofint64. Input order is preserved along the first axis of every returned array.properties (Union[List[str], Dict[str, str]]) – Either a list of relative dataset paths under each
Output*group (e.g.["nodeData/basicMass"]), matchingCollection.read(), or adictmapping user-chosen labels to relative paths.outputs (Optional[Sequence[Union[int, str]]]) – Optional iterable selecting a subset of outputs to include. Each element may be a 1-based integer (e.g.
3) or a group name (e.g."Output3"). Arangeis accepted. Defaults to all outputs in the collection, in temporal order.id_dataset (str) – Relative path of the tracing ID dataset under each
Output*group. Defaults to"nodeData/nodeUniqueIDBranchTip".on_duplicate_file_match (str) –
What to do if the same ID is found in more than one file at the same output (IDs are only unique within a file in multi-file collections):
"error"(default) – raiseValueError."warn"– emit aUserWarningand keep the first file’s match."first"– silently keep the first file’s match.
int_sentinel (int) – Missing-slot value used for integer-typed properties. Defaults to
-1.
- Returns:
Contains:
one entry per property –
numpy.ndarrayof shape(n_galaxies,) + per_galaxy_tail + (n_outputs,). A 1-D source dataset yields a 2-D(n_galaxies, n_outputs)array; a 2-D dataset of shape(N, W)yields a 3-D(n_galaxies, W, n_outputs)array; and so on."time"– float array(n_galaxies, n_outputs)of output times, NaN where the galaxy is absent."redshift"– float array(n_galaxies, n_outputs)of redshifts, NaN where the galaxy is absent."expansion_factor"– float array(n_galaxies, n_outputs)of expansion factors, NaN where the galaxy is absent."present"– bool array(n_galaxies, n_outputs)that isTrueexactly where the galaxy was located."output_names"– 1-D object array of output group names in temporal order."ids"– 1-Dint64array of normalized input IDs.
- Return type:
- Raises:
KeyError – If
id_datasetis not present in any chosen output of any file (e.g. the Galacticus run did not emitnodeUniqueIDBranchTip), or if a requested property is missing from a chosen output.ValueError – If
propertiescontains a reserved label, ifoutputsis empty, if any chosen output has anoutputTypeof"node"or"lightcone"(for which tracing is undefined), if the tail shape of a property differs between outputs, or (by default) if an ID appears in more than one file at the same output.NotImplementedError – If a property has a dtype other than integer, floating, or boolean.
Notes
A galaxy need not be present at every output (it may have formed later or merged earlier); ragged histories are expected. Requesting IDs that are never found anywhere produces a
UserWarningrather than an error, since exploratory workflows often probe IDs of uncertain provenance.
- dendros.list_analyses(collection, format='astropy')[source]
Return a table of
function1Danalyses available in the collection.- Parameters:
collection (Collection) – A
Collection. Only the primary file is consulted — for MPI runs, the/analysesdata has been reduced over all ranks and is identical in every file.format (str) –
"astropy"(default),"pandas", or"tabulate".
- Return type:
astropy.table.Table, pandas.DataFrame, or tabulate-formatted string
- Raises:
KeyError – If the file has no top-level
/analysesgroup.
- dendros.plot_analyses(collection, name=None, output_directory=None, *, labels=None, show_target=True, figsize=(7.0, 5.0), dpi=120, file_format='pdf')[source]
Plot one, several, or all
function1Danalyses.A single
Collectionproduces one model curve per figure (legacy behaviour). A list, dict, orModelCollectionof Collections overlays one curve per model on each figure, plotting the target/observational overlay once (since it is shared across models). The union of analyses discovered across models is plotted — figures whose analysis is absent from a given model simply do not include its curve.- Parameters:
collection (_MultiInput) – A
Collection; a sequence of Collections; or a mapping{label: Collection}(e.g. one returned byopen_models()).name (Union[None, str, List[str]]) –
None(default) plots everyfunction1Danalysis discovered across all models. A single name (str) or list of names plots only those.output_directory (Union[None, str, 'Path']) – If given, each figure is also saved as
<output_directory>/<safe_name>.<file_format>. The directory is created if it does not exist.labels (Optional[Sequence[str]]) – Optional sequence of legend labels, one per Collection, used only when collection is a list/tuple of Collections. When omitted, each model is labelled by its primary file’s stem (with any
:MPIxxxxsuffix stripped). Cannot be combined with a dict input.show_target (bool) – If
True(default), overlay target/observational data when present. For multi-model plots the target is plotted only once, from the first model that has it.dpi (int) – Forwarded to matplotlib.
file_format (str) – Forwarded to matplotlib.
- Returns:
Mapping from analysis name to
matplotlib.figure.Figure.- Return type:
- Raises:
KeyError – If a model has no
/analysesgroup, or if a requested name is missing from every model.ImportError – If matplotlib is not installed; install with
pip install 'dendros[plot]'.
MCMC
Entry point
- dendros.open_mcmc(config_path, *, max_steps=None)[source]
Open an MCMC run by parsing its config XML.
- Parameters:
config_path (str | Path) – Path to the Galacticus MCMC
<parameters>XML file.max_steps (int | None) – When given, retain only the last
max_stepsrecorded steps of each chain. Intended for run-time monitoring of a live chain, where only the recent window is of interest: it bounds the per-row parse cost and speeds up every downstream diagnostic (R-hat, ESS, autocorrelation).None(the default) reads the full history — use that for definitive, post-hoc analysis (corner plots, final convergence, posterior sampling) where truncation would bias the result.
- Return type:
Examples
>>> from dendros import open_mcmc >>> with open_mcmc("mcmcConfig.xml") as run: ... print(run.parameters) ... chains = run.chains
Fast run-time monitoring on the last 1000 steps:
>>> with open_mcmc("mcmcConfig.xml", max_steps=1000) as run: ... rhat = run.gelman_rubin()
- class dendros.MCMCRun(config, *, max_steps=None)[source]
Bases:
objectAn MCMC run, parsed from its config file and lazily backed by chain data.
Construct via
open_mcmc(). The chain files are not read untilchainsis first accessed; subsequent accesses return the cachedChainSet.- Parameters:
config (MCMCConfig) – Parsed
MCMCConfig.max_steps (Optional[int]) – When given, retain only the last
max_stepsrecorded steps of each chain when the chain data are first read. Intended for run-time monitoring of a live chain: it bounds the per-row parse cost and speeds up every downstream diagnostic, at the cost of discarding older history.None(the default) reads the entire history. Assigning to themax_stepsproperty later invalidates any cached chains.
- property config: MCMCConfig
The parsed
MCMCConfig.
- property parameters: Tuple[ModelParameter, ...]
Active model parameters, in chain-file column order.
- property max_steps: int | None
Retain only the last this-many steps per chain when reading (
None= all).Assigning a new value invalidates the cached
chainsso the next access re-reads with the new window.
- property chains: ChainSet
Lazily-loaded
ChainSetfor this run.Honors
max_steps: when set, only the lastmax_stepssteps of each chain are read and cached.
- gelman_rubin(*, drop_chains=(), step_grid=None, n_grid=200, min_steps=10, alpha_interval=0.15)[source]
Convenience wrapper around
dendros.gelman_rubin().
- convergence_step(*, threshold=1.1, sustained_for=1, drop_chains=(), step_grid=None, n_grid=200, min_steps=10)[source]
First simulation-step count at which max-Rhat is sustained below threshold.
Computes a Gelman-Rubin trace via
gelman_rubin()and returns theRhatResult.stepsvalue at which convergence is first declared. ReturnsNoneif convergence is never reached on the chosen grid.
- geweke(*, first=0.1, last=0.5)[source]
Convenience wrapper around
dendros.geweke().
- ensemble_drift(*, drop_chains=(), burn=0, first=0.5, last=0.5)[source]
Convenience wrapper around
dendros.ensemble_drift().The effect-size stationarity check appropriate for the
differentialEvolutionensemble: gate onresult.max_drift() < thresholdrather than on Gelman-Rubin / Geweke, which over-reject for large ensembles.
- outlier_chains(*, alpha=0.05, max_outliers=10, parameters=None)[source]
Convenience wrapper around
dendros.outlier_chains().
- acceptance_rate(*, post_burn=None)[source]
Convenience wrapper around
dendros.acceptance_rate().
- acceptance_rate_trace(*, window=30, post_burn=0)[source]
Convenience wrapper around
dendros.acceptance_rate_trace().
- autocorrelation_time(*, post_burn=None, c=5.0)[source]
Convenience wrapper around
dendros.autocorrelation_time().
- effective_sample_size(*, post_burn=None, c=5.0)[source]
Convenience wrapper around
dendros.effective_sample_size().
- maximum_posterior(*, drop_chains=())[source]
Convenience wrapper around
dendros.maximum_posterior().
- maximum_likelihood(*, drop_chains=())[source]
Convenience wrapper around
dendros.maximum_likelihood().
- posterior_samples(n, *, post_burn=None, drop_chains=(), rng=None, replace=None)[source]
Convenience wrapper around
dendros.posterior_samples().
- projection_pursuit(*, post_burn=None, drop_chains=())[source]
Convenience wrapper around
dendros.projection_pursuit().- Parameters:
- Return type:
- multivariate_normal_fit(*, post_burn=None, drop_chains=())[source]
Convenience wrapper around
dendros.multivariate_normal_fit().
- write_parameter_file(state, out_path, *, likelihood_index=0)[source]
Emit a single Galacticus parameter file for one likelihood leaf.
Reads
leaves[likelihood_index].base_parameters_file, applies the leaf’sparameter_map(or the full state when no map is set), and writes the result to out_path.- Parameters:
state –
(n_params,)state vector in physical (model) space. Galacticus stores chain rows in physical space, so a row frommaximum_posterior()/posterior_samples()can be passed directly.out_path – Output path. Parent directory is created if missing.
likelihood_index (int) – Which leaf of the likelihood tree to use. Defaults to
0.
- Returns:
Resolved output path.
- Return type:
- corner_plot(*, parameters=None, post_burn=None, drop_chains=(), labels=None, **corner_kwargs)[source]
Convenience wrapper around
dendros.corner_plot().
- write_parameter_files(state, out_dir, *, name_format=None)[source]
Emit one parameter file per likelihood leaf into out_dir.
For
independentLikelihoodsconfigs each leaf has its ownbaseParametersFileNameandparameterMap; this writes one file per leaf, with each file’s filename derived from the base file’s stem.- Parameters:
state –
(n_params,)state vector in physical space.out_dir – Output directory (created if missing).
name_format (str | None) – Output-filename format string accepting
leaf_indexandstem. Defaults to"{stem}.xml"for a single leaf and"{leaf_index:02d}_{stem}.xml"for multiple.
- Returns:
One entry per leaf, in document order.
- Return type:
list of (leaf_index, pathlib.Path)
Configuration
- dendros.parse_mcmc_config(path)[source]
Parse a Galacticus MCMC
<parameters>XML file.- Parameters:
- Return type:
- Raises:
FileNotFoundError – If path does not exist.
ValueError – If the file’s root element is not
<parameters>, or if the requiredposteriorSampleSimulation/logFileRootelements are missing.
- class dendros.MCMCConfig(config_path, log_file_root, simulation_kind, parameters, likelihood)[source]
Parsed Galacticus MCMC configuration.
- Parameters:
config_path (Path)
log_file_root (Path)
simulation_kind (str)
parameters (Tuple[ModelParameter, ...])
likelihood (Likelihood | None)
- config_path
Absolute path to the parsed XML file.
- Type:
- log_file_root
Resolved chain log-file root (relative paths resolved against
config_path’s directory). Per-rank chain files are atf"{log_file_root}_{rank:04d}.log".- Type:
- simulation_kind
Value attribute of
posteriorSampleSimulation, e.g."differentialEvolution"or"particleSwarm". Determines whether chain rows carry trailing per-particle velocity columns.- Type:
- parameters
Tuple of active
ModelParameterentries in document order. This is the canonical ordering used by chain-file columns.- Type:
Tuple[dendros._mcmc._config.ModelParameter, …]
- likelihood
Root of the
posteriorSampleLikelihoodtree, orNoneif the config lacks a likelihood block.- Type:
- state_indices_for(leaf)[source]
Indices of the global state vector applicable to leaf.
For a leaf inside an
independentLikelihoodssubtree, returns the positions inparametersnamed in the leaf’sparameter_map. For other leaves, returns(0, 1, ..., n_params - 1)(identity).- Parameters:
leaf (Likelihood) – A
LikelihoodfromLikelihood.leaves().- Raises:
KeyError – If a name in
parameter_mapisn’t among the active parameters.- Return type:
- class dendros.ModelParameter(name, label=None, prior=None, mapper='identity', perturber=None)[source]
A single
<modelParameter value="active">entry from the config.- Parameters:
name (str)
label (str | None)
prior (PriorSpec | None)
mapper (str)
perturber (PerturberSpec | None)
- label
Optional LaTeX label for plotting.
Nonewhen the config omits the<label>sub-element. Usedisplay_labelto obtain a plottable string regardless.- Type:
str | None
- prior
Parsed
distributionFunction1DPriorblock, if present.- Type:
- perturber
Parsed
distributionFunction1DPerturberblock, if present.- Type:
- class dendros.Likelihood(kind, base_parameters_file=None, parameter_map=None, children=<factory>, options=<factory>)[source]
A node in the
posteriorSampleLikelihoodtree.- Parameters:
- base_parameters_file
Resolved path to the
baseParametersFileNameelement’s value when present.Nonefor non-leaf nodes (e.g.independentLikelihoodswithout a base file of its own).- Type:
pathlib.Path | None
- parameter_map
For children of
posteriorSampleLikelihoodIndependentLikelihoods, the parsed<parameterMap value="space separated names"/>for this child. Each entry is a parameter name from the active model parameters.Noneoutside of anindependentLikelihoodscontext, in which case identity mapping (all active parameters) is implied.- Type:
Tuple[str, …] | None
- children
Tuple of child
Likelihoodinstances. Empty for leaves.- Type:
Tuple[dendros._mcmc._config.Likelihood, …]
- options
The remaining scalar sub-parameters of this
posteriorSampleLikelihoodelement, as a{tag: value}mapping of raw strings. Each likelihood class defines its own options (fileNames,redshifts,massRangeMinimum, … forhaloMassFunction), so they are kept uninterpreted here; read them viaoption(),option_float(),option_bool()oroption_words().baseParametersFileNameandparameterMapare excluded, having their own fields.
- option_words(tag)[source]
Return whitespace-separated option tag split into a tuple.
Empty when the option is absent. Galacticus uses whitespace-separated lists for vector-valued parameters such as
fileNamesandredshifts.
- leaves()[source]
Flatten the tree to its leaf likelihoods (in document order).
- Return type:
Tuple[Likelihood, …]
Chains
- dendros.read_chains(config, *, max_steps=None, log_file_root=None, on_header_mismatch='raise')[source]
Discover and read all per-rank chain files for config.
- Parameters:
config (MCMCConfig) – Parsed
MCMCConfig.max_steps (int | None) – When given, retain only the last
max_stepsrecorded steps of each chain (see_read_chain_file()). Speeds up reading and every downstream diagnostic for run-time monitoring, where only the recent window matters.None(the default) reads the entire history.log_file_root (str | Path | None) – Override for
config.log_file_root. A run copied from the machine it executed on will have an absolute path in its config that does not exist locally; pass the local root here.on_header_mismatch (str) – What to do when a chain file’s header parameter names disagree with the config:
"raise"(default),"warn"or"ignore". Headers are written once at file creation, so a resumed run — or one analysed against a regenerated config — can carry stale names with correct columns.
- Return type:
- Raises:
FileNotFoundError – If no chain files are found.
- class dendros.Chain(chain_index, path, step, eval_time, converged, log_posterior, log_likelihood, state, velocity=None)[source]
One MPI rank’s MCMC chain.
- Parameters:
- path
Source log-file path.
- Type:
- step
Integer simulation-step index, one per row.
- Type:
- eval_time
Wall-clock evaluation time per step, in seconds.
- Type:
- converged
Boolean flag indicating whether the simulation had declared convergence at this step.
- Type:
- log_posterior
Log posterior probability per step.
- Type:
- log_likelihood
Log likelihood per step.
- Type:
- state
(n_steps, n_params)array of parameter values, inMCMCConfig.parametersorder. Values are in physical (model) space — Galacticus applies the inverse ofoperatorUnaryMapperbefore writing.- Type:
- velocity
(n_steps, n_params)array of per-parameter particle velocities forparticleSwarmsimulations;Nonefor differential-evolution and other state-only simulations.- Type:
numpy.ndarray | None
- class dendros.ChainSet(config, chains)[source]
An ordered collection of
Chainobjects from one MCMC run.Iteration yields chains in MPI-rank order.
- Parameters:
config (MCMCConfig) – The parsed
MCMCConfigthe chains correspond to.chains (Sequence[Chain]) – The per-rank chains.
Convergence
- dendros.gelman_rubin(chains, *, drop_chains=(), step_grid=None, n_grid=200, min_steps=10, alpha_interval=0.15)[source]
Brooks-Gelman corrected Rhat as a function of simulation step.
For each chosen truncation point
sthe firstsrows of every surviving chain are used to compute the standard between-chain (B) and within-chain (W) variances and the Brooks-Gelman corrected potential-scale reduction factor \(\hat{R}_c\). The non-parametric interval-length ratio \(R_{\rm interval}\) (Brooks & Gelman 1998 section 1.3) is also computed at the same evaluation points.- Parameters:
chains (ChainSet) –
ChainSetto evaluate. Must contain at least two non-dropped chains and at leastmin_stepsrows per chain.drop_chains (Sequence[int]) – Iterable of
chain_indexvalues to exclude before computing. Use this with the indices returned byoutlier_chains().step_grid (Sequence[int] | None) – Optional explicit 1-D iterable of truncation step counts (1-based). When given,
n_gridandmin_stepsare ignored.n_grid (int) – Number of evenly-spaced evaluation points to use when
step_gridisNone. Capped at the shortest surviving chain length minusmin_steps+ 1.min_steps (int) – Smallest truncation step count to evaluate. Must be
>= 2.alpha_interval (float) – Two-sided significance level for
R_interval(default 0.15, i.e. 85 % credible intervals — matches the Galacticus Perl reference).
- Return type:
- Raises:
ValueError – If fewer than two chains survive
drop_chainsormin_stepsis too small.
- class dendros.RhatResult(steps, Rhat_c, R_interval, parameter_names, alpha_interval, chains_used)[source]
Result of
gelman_rubin().- Parameters:
- steps
(n_eval,)1-D array of truncation step counts at which Rhat was computed (i.e. each entrysmeans “use the firstsrows of every chain”). These are 1-based step counts so the smallest value is the chosenmin_steps.- Type:
- Rhat_c
(n_eval, n_params)array of Brooks-Gelman corrected potential-scale reduction factors.- Type:
- R_interval
(n_eval, n_params)array of non-parametric interval-length ratios (mixed-chain credible interval / mean per-chain credible interval) at the chosenalpha_interval.- Type:
- chains_used
chain_indexvalues of the chains that contributed (afterdrop_chainswas applied).- Type:
Tuple[int, …]
- Rhat_c_max:
Per-step max-over-parameters of
Rhat_c, useful as the input toconvergence_step().
- dendros.convergence_step(rhat_max, *, threshold=1.1, sustained_for=1)[source]
Index into the Rhat grid at which convergence is first declared.
Searches for the smallest index
isuch that every entry ofrhat_max[i : i + sustained_for]is at or belowthreshold.- Parameters:
rhat_max (ndarray) – 1-D array of (max-over-parameters) Rhat values, e.g.
RhatResult.Rhat_c_max().threshold (float) – Convergence threshold. Defaults to
1.1.sustained_for (int) – Number of consecutive grid points that must all be below the threshold before convergence is declared. Defaults to
1(strict first crossing).
- Returns:
Grid index at which convergence is first sustained, or
Noneif the threshold is never met.- Return type:
int or None
Notes
Use
RhatResult.stepsto translate the returned grid index to a simulation-step count.
- dendros.geweke(chains, *, first=0.1, last=0.5)[source]
Per-chain Geweke z-scores comparing the means of two chain segments.
For each chain and each parameter, returns
\[z = \frac{\bar{x}_1 - \bar{x}_2}{\sqrt{s^2_1/n_1 + s^2_2/n_2}}\]where segment 1 covers the first
firstfraction of the chain and segment 2 covers the lastlastfraction. Large|z|for any parameter suggests the chain has not yet reached a stationary distribution — useful when the chains were started from an under-dispersed state (which makes Gelman-Rubin uninformative).- Parameters:
- Returns:
(n_chains, n_params)z-score array. Chains shorter than 4 rows in either segment yieldNaN.- Return type:
np.ndarray
Notes
The variance estimator used here is the simple sample variance, which treats each draw as independent. Autocorrelated chains will produce artificially-large
|z|; once a proper integrated-autocorrelation-time estimator lands (Phase 3) this can be inflated by the ACL to recover the classical spectral-density-at-zero variant.
- dendros.ensemble_drift(chains, *, drop_chains=(), burn=0, first=0.5, last=0.5)[source]
Standardized drift of the ensemble mean between an early and late window.
Pools all surviving walkers to form the ensemble at each step, then compares the pooled mean over an early window against a late window and reports the shift in units of the posterior width — an effect-size stationarity check appropriate for interacting ensemble samplers, where per-walker Gelman-Rubin / Geweke are inflated by the long integrated autocorrelation time and, being significance tests, reject on negligible drift once the ensemble is large (see
EnsembleDriftResultnotes).For each parameter,
\[\mathrm{drift} = \frac{|\bar{x}_{\rm late} - \bar{x}_{\rm early}|} {\sigma_{\rm post}}\]where the means pool every surviving walker and every step in each window, and \(\sigma_{\rm post}\) is the pooled standard deviation over the whole (burned) analysis span.
- Parameters:
chains (ChainSet) –
ChainSet. All surviving chains are truncated to the shortest length before windowing.drop_chains (Sequence[int]) – Iterable of
chain_indexvalues to exclude (e.g. the output ofoutlier_chains()).burn (int) – Number of leading rows to drop from every chain before windowing, to exclude burn-in. Compare two late windows by burning first: the full-run first half otherwise carries the approach-to-stationarity and inflates the drift.
first (float) – Fractions in
(0, 1]giving the lengths of the early and late windows within the burned span. Must satisfyfirst + last <= 1so the windows do not overlap. Default0.5/0.5(compare the two halves of the burned span).last (float) – Fractions in
(0, 1]giving the lengths of the early and late windows within the burned span. Must satisfyfirst + last <= 1so the windows do not overlap. Default0.5/0.5(compare the two halves of the burned span).
- Return type:
- Raises:
ValueError – If no chains survive
drop_chains, the fractions are out of range or overlap, orburnleaves too few rows to form both windows.
- class dendros.EnsembleDriftResult(drift, delta_mean, sigma_post, parameter_names, early_steps, late_steps, chains_used)[source]
Result of
ensemble_drift().- Parameters:
- drift
(n_params,)standardized drift of the ensemble mean between the early and late windows:|mean_late - mean_early| / sigma_post. A dimensionless effect size — the shift of the pooled ensemble mean expressed in units of the posterior width.- Type:
- delta_mean
(n_params,)signedmean_late - mean_early(model units).- Type:
- sigma_post
(n_params,)pooled ensemble standard deviation over the analysis span (all surviving walkers × steps afterburn), i.e. the posterior width used to standardizedrift.- Type:
- early_steps, late_steps
(start, stop)row-index half-open ranges (into the burned span) of the two windows that were compared.
Notes
This is deliberately an effect-size statistic, not a significance test. For an interacting ensemble sampler (e.g. Galacticus
differentialEvolution) with many walkers, the ensemble mean is estimated so precisely that a significance test (Gelman-Rubin, Geweke, a pooled z-test) rejects stationarity on a negligible drift — it has too much power. The drift effect size answers the question that actually matters for calibration: “has the ensemble mean stopped moving, relative to the posterior width?” Gate ondrift.max() < threshold(e.g.0.1), with the threshold a human choice. Pair witheffective_sample_size()/autocorrelation_time()for the independent-sample count.
- dendros.outlier_chains(chains, *, alpha=0.05, max_outliers=10, parameters=None)[source]
Iterative two-sided Grubbs test on each chain’s final state.
Each chain contributes its last row (the most recent state) as a single multivariate point. The Grubbs test is applied iteratively over the active chains, dropping the chain whose maximum per-parameter deviation exceeds the critical value at each step, until none exceed it or
max_outlierschains have been removed.- Parameters:
chains (ChainSet) –
ChainSet. Must contain at least three chains.alpha (float) – Two-sided significance level. Defaults to
0.05to match the Galacticus Perl reference’s hard-coded value.max_outliers (int) – Maximum number of chains to declare as outliers.
parameters (Iterable[str] | None) – Optional iterable of parameter names to restrict the test to a subset. Unknown names raise
KeyError.
- Returns:
chain_indexvalues of the chains flagged as outliers, in the order they were removed.- Return type:
Mixing diagnostics
- dendros.autocorrelation_function(chains, *, post_burn=0, max_lag=None)[source]
Per-chain, per-parameter normalized autocorrelation function.
- Parameters:
- Returns:
Array of shape
(n_chains, max_lag + 1, n_params). Chains are truncated to a common length (the shortest post-burn chain) so this is rectangular.- Return type:
np.ndarray
- dendros.autocorrelation_time(chains, *, post_burn=None, c=5.0)[source]
Integrated autocorrelation time per parameter, in steps.
Implements the standard Sokal automatic-windowing estimator over the chain-averaged autocovariance. For each parameter, the per-chain autocovariances are averaged before integrating, which is more stable than averaging per-chain
τ_intestimates.- Parameters:
chains (ChainSet) –
ChainSet. All chains are truncated to the shortest post-burn length.post_burn (int | None) – Number of leading rows to skip from each chain.
Nonetriggers automatic detection viagelman_rubin()/convergence_step(); if convergence is not reached aUserWarningis emitted and0is used.c (float) – Sokal window constant. Defaults to
5.0.
- Returns:
(n_params,)array of integrated autocorrelation times in steps.- Return type:
np.ndarray
- dendros.effective_sample_size(chains, *, post_burn=None, c=5.0)[source]
Effective sample size per parameter.
Defined as
N_total / τ_intwhereN_totalis the total number of post-burn samples summed across all chains andτ_intis the chain- averaged integrated autocorrelation time fromautocorrelation_time().- Parameters:
post_burn (int | None) – See
autocorrelation_time().c (float) – Sokal window constant.
- Returns:
(n_params,)array of effective sample sizes.- Return type:
np.ndarray
- dendros.acceptance_rate(chains, *, post_burn=None)[source]
Per-chain post-burn acceptance rate.
A step is “accepted” iff any parameter component differs from the previous step. Galacticus emits the same row when a proposal is rejected, so this is the canonical acceptance count.
- Parameters:
post_burn (int | None) – Number of leading rows to skip in each chain.
Nonetriggers automatic detection viagelman_rubin()/convergence_step(); if convergence is not reached aUserWarningis emitted and0is used.
- Returns:
(n_chains,)array.NaNfor any chain with fewer than two post-burn rows.- Return type:
np.ndarray
- dendros.acceptance_rate_trace(chains, *, window=30, post_burn=0)[source]
Sliding-window acceptance rate as a function of step.
For each chain, returns a 1-D array whose
i-th entry is the fraction of the previous window transitions that were accepted (i.e. changed at least one parameter). The firstwindowentries are filled withnumpy.nanbecause the window is not yet full.- Parameters:
- Returns:
One 1-D array per chain, each of length
n_steps_post_burn. Returned as a list because chains may have different post-burn lengths.- Return type:
list of np.ndarray
Posterior analyses
- dendros.maximum_posterior(chains, *, drop_chains=())[source]
State vector at the maximum log posterior across all surviving chains.
- dendros.maximum_likelihood(chains, *, drop_chains=())[source]
State vector at the maximum log likelihood across all surviving chains.
- class dendros.MaxResult(state, log_posterior, log_likelihood, chain_index, step, parameter_names)[source]
Result of
maximum_posterior()ormaximum_likelihood().- Parameters:
- state
(n_params,)parameter vector at the maximizing step.- Type:
- step
Simulation-step value of the maximizing row (1-based, matching
Chain.step).- Type:
- dendros.posterior_samples(chains, n, *, post_burn=None, drop_chains=(), rng=None, replace=None)[source]
Draw n uniformly-random rows from the post-burn concatenated chain.
- Parameters:
n (int) – Number of samples to draw. Must be positive.
post_burn (int | None) – Number of leading rows to skip in each chain.
Nonetriggers automatic detection viagelman_rubin()/convergence_step(); if convergence is not reached aUserWarningis emitted and0is used.drop_chains (Sequence[int]) – Iterable of
chain_indexvalues to exclude.rng (Generator | None) –
numpy.random.Generator. Defaults tonumpy.random.default_rng(), which seeds from system entropy. Pass an explicit generator for reproducibility.replace (bool | None) – Whether to sample with replacement.
None(default) means “with replacement only when n exceeds the available pool”, matching the common case where a smallnis desired and identical rows would be misleading.
- Return type:
- Raises:
ValueError – If n is non-positive, or if all chains are dropped, or if
replace=Falseand n exceeds the pool size.
- class dendros.PosteriorSamples(state, log_posterior, log_likelihood, chain_index, step, parameter_names)[source]
Sampled rows from the post-burn concatenated chain.
- Parameters:
- state
(n_samples, n_params)parameter values.- Type:
- log_posterior
(n_samples,)log posterior at each sample.- Type:
- log_likelihood
(n_samples,)log likelihood at each sample.- Type:
- chain_index
(n_samples,)source chain indices.- Type:
- step
(n_samples,)source simulation-step values.- Type:
Notes
Adjacent steps in an MCMC chain are correlated — N draws here represent significantly fewer independent samples from the posterior. Use
effective_sample_size()to estimate the equivalent count, and thin by the integrated autocorrelation time if independence is required.
- dendros.projection_pursuit(chains, *, post_burn=None, drop_chains=())[source]
Find the linear combinations of parameters best constrained by the data.
Each post-burn parameter column is mapped via its
operatorUnaryMapper, normalised bysqrt(prior variance), and mean-centred. The covariance matrix of the resulting samples is eigendecomposed, and the eigenvalues/eigenvectors are returned sorted by ascending eigenvalue — soeigenvectors[:, 0]is the linear combination most tightly constrained relative to the prior.- Parameters:
- Return type:
- Raises:
NotImplementedError – If any active parameter uses a prior or mapper not yet supported by
projection_pursuit()(currently uniform/normal priors andidentitymapper only).ValueError – If the post-burn pool is empty or contains fewer than two rows.
- class dendros.ProjectionPursuitResult(eigenvalues, eigenvectors, parameter_names, parameter_labels, prior_sigma)[source]
Result of
projection_pursuit().- Parameters:
- eigenvalues
(n_params,)ascending eigenvalues of the rescaled-sample covariance matrix. Smaller is “better constrained”.- Type:
- eigenvectors
(n_params, n_params)matrix whose[:, k]column is the eigenvector foreigenvalues[k], expressed in rescaled-mapped space (i.e. in the same coordinates used for the eigendecomposition).- Type:
- parameter_labels
ModelParameter.display_labelstrings parallel toparameter_names.- Type:
Tuple[str, …]
- prior_sigma
(n_params,)square root of the prior variance for each parameter, the rescaling that was applied before eigendecomposition.- Type:
- direction:
Components of one eigenvector that exceed a contribution threshold.
- latex_summary:
LaTeX-rendered summary line for a chosen direction.
- dendros.multivariate_normal_fit(chains, *, post_burn=None, drop_chains=())[source]
Fit a multivariate normal to the post-burn concatenated chain.
- Parameters:
post_burn (int | None) – Number of leading rows to skip per chain.
Nonetriggers automatic detection viagelman_rubin()/convergence_step().drop_chains (Sequence[int]) –
chain_indexvalues to exclude.
- Return type:
- Raises:
ValueError – If fewer than
n_params + 1post-burn samples remain (so that the sample covariance is rank-deficient).np.linalg.LinAlgError – If the sample covariance is not positive-definite (which can happen for parameters that are degenerate post-burn). Drop the offending parameter or supply more samples.
- class dendros.MVNFit(mean, covariance, cholesky, parameter_names)[source]
Multivariate-normal fit to post-burn samples.
- Parameters:
- mean
(n_params,)sample mean.- Type:
- covariance
(n_params, n_params)sample covariance, symmetrised.- Type:
- cholesky
(n_params, n_params)lower-triangular Cholesky factor ofcovariance. SatisfiesL @ L.T == covariance.- Type:
- write_reparameterization_config:
Emit a Galacticus-style XML config that re-parameterizes the active parameters in terms of independent unit-normal meta parameters.
- write_reparameterization_config(out_path, *, n_sigma=5.0, perturber_scale=1e-05)[source]
Write a Galacticus reparameterization XML config.
For an n-parameter MVN fit with mean \(\mu\) and Cholesky factor \(L\), the emitted config declares n active
metaParameter{i}parameters with truncated unit-normal priors (limits \(\pm n_\sigma\)), and n derived parameters expressing the original active parameters as\[x_i = \mu_i + \sum_j L_{ij} \, m_j .\]Re-running the MCMC against this config samples in coordinates where the posterior is approximately spherical.
- Parameters:
- Returns:
Resolved path of the written file.
- Return type:
Parameter-file emission
- dendros.read_parameter_file(path)[source]
Parse a Galacticus parameter XML file.
- Parameters:
- Return type:
- dendros.resolve_parameter_path(root, path)[source]
Locate the XML element identified by a Galacticus parameter path.
- Parameters:
root (Element) – Root
xml.etree.ElementTree.Elementto search.path (str) – Slash- or
::-separated parameter path. Each segment is an element name, optionally followed by a[N]integer (1-based) instance selector or a[@value='x']attribute filter — matching the Galacticus parameter-file convention.
- Return type:
- Raises:
KeyError – If any segment does not match an element under the current node.
ValueError – If a path segment is malformed.
- dendros.apply_state(tree, parameters, state, *, parameter_map=None)[source]
Set
value=attributes in tree from a state vector.- Parameters:
tree (ElementTree) – Parsed XML tree to modify in place.
parameters (Sequence[ModelParameter]) – Active model parameters (the same ordering used by chain columns).
state (ndarray) –
(n_params,)state vector aligned with parameters.parameter_map (Iterable[str] | None) – Optional iterable of parameter names to apply.
Noneapplies every entry of parameters; non-Noneapplies only those named and is the typical case for anindependentLikelihoodsleaf, whose base parameter file mentions only that leaf’s parameters.
- Raises:
KeyError – If a parameter’s path does not resolve in tree, or if a name in parameter_map is not among parameters.
ValueError – If state doesn’t match the length of parameters.
- Return type:
None
- dendros.emit_parameter_files(state, config, out_dir, *, name_format=None)[source]
Write one Galacticus parameter file per leaf of
config.likelihood.Reads each leaf’s
base_parameters_file, applies the subset of state selected by that leaf’sparameter_map(or the full state when no map is set), and writes the modified XML into out_dir.- Parameters:
state (ndarray) –
(n_params,)state vector (in physical / model space, as stored in the chain log file — no mapper inversion is applied).config –
MCMCConfig.out_dir (str | Path) – Output directory (created if missing).
name_format (str | None) – Format string for output filenames; receives
leaf_indexandstem(the base file’s stem). Defaults to"{stem}.xml"for a single leaf and"{leaf_index:02d}_{stem}.xml"for multiple, so per-leaf files don’t collide when several leaves share a base stem.
- Returns:
One tuple per leaf, in document order.
- Return type:
list of (leaf_index, written_path)
- Raises:
ValueError – If
config.likelihoodisNone, or if any leaf lacks abase_parameters_file.KeyError – If a parameter’s path does not resolve in the corresponding base file, or a
parameter_mapreferences an unknown parameter.
Corner plots
- dendros.corner_plot(chains, *, parameters=None, post_burn=None, drop_chains=(), labels=None, **corner_kwargs)[source]
Render a corner plot of post-burn chain samples.
- Parameters:
chains (ChainSet) –
ChainSetwhose post-burn samples will be plotted.parameters (Iterable[str] | None) – Optional iterable of parameter names to restrict the plot to a subset (in the order given).
Noneplots every active parameter.post_burn (int | None) – Number of leading rows to skip per chain.
Nonetriggers automatic detection viagelman_rubin()/convergence_step().drop_chains (Sequence[int]) – Iterable of
chain_indexvalues to exclude.labels (Sequence[str] | None) – Optional axis labels.
Noneuses each parameter’s LaTeXModelParameter.display_label, wrapped in$...$socornerrenders them in math mode.**corner_kwargs – Additional keyword arguments forwarded to
corner.corner().
- Return type:
matplotlib.figure.Figure
- Raises:
ImportError – If the optional
cornerpackage is not installed. Install viapip install 'dendros[mcmc]'.KeyError – If a name in parameters is not among the active parameters.
ValueError – If the post-burn pool is empty.
Model prediction vectors
- dendros.index_prediction_files(samples_dir)[source]
Scan samples_dir once, grouping sample files by constraint label.
A production run writes one file per constraint per MPI rank, so this directory routinely holds hundreds of thousands of entries. Scan it once and pass the result to
read_predictions()viafiles: globbing per-label instead re-walks every entry for each constraint, which dominates the run time and scales quadratically in the number of constraints.
- dendros.discover_prediction_labels(samples_dir)[source]
Return the sorted constraint labels present in samples_dir.
A label is a sample filename with its
_NNNN.txtrank suffix removed.
- dendros.read_predictions(samples_dir, label, *, ranks=None, files=None)[source]
Read every rank’s prediction file for constraint label.
- Parameters:
samples_dir (str | Path) – Directory holding the sample files. Pass this explicitly rather than taking it from the config: a run analysed after being copied from the machine it ran on will have a stale
pathSamples.label (str) – Constraint label, as returned by
discover_prediction_labels().ranks (Sequence[int] | None) – When given, read only these MPI ranks.
files (Sequence[Tuple[int, Path]] | None) – Pre-resolved
[(rank, path), ...]for this label, as produced byindex_prediction_files(). Supplying it skips the directory scan — strongly preferred when reading many constraints from one directory.
- Return type:
- Raises:
FileNotFoundError – If no files match label.
- class dendros.PredictionSet(label, series, *, abscissa=None, abscissa_name=None)[source]
All ranks’ prediction records for one constraint.
- Parameters:
label (str)
series (Sequence[PredictionSeries])
abscissa (Optional[np.ndarray])
abscissa_name (Optional[str])
- records_by_chain()[source]
Return
{chain_index: (step, prediction)}, pooled across files.When records carry their own chain index they are attributed by it, so the result is correct even under load balancing, where a chain’s model may have been evaluated and written by any process. Otherwise each file is attributed wholesale to the process that wrote it, which is only valid if load balancing was off.
- step_offset(chains)[source]
Detect which step-labelling convention this run used.
Returns whichever of
SAMPLE_STEP_OFFSETSaccounts for more of the chains’ accepted steps. Every acceptance was necessarily evaluated and so must have a record, which normally makes this decisive: proposals rejected on the prior never reach the likelihood, so records cover only a fraction of steps and the wrong offset leaves a conspicuous shortfall.When records happen to cover every step the two conventions are indistinguishable from step indices alone — both account for all accepted steps, while pairing each state with a different record. That case warns and returns
0; passstep_offsetexplicitly topaired()if the run predates the labelling fix.
- paired(chains, *, burn=0, drop_chains=(), step_offset=None, proposals=None)[source]
Join records to the parameter states that produced them.
Without proposals only accepted steps can be paired, since a rejected proposal’s parameters appear nowhere in the chain log. Supplying a
ProposalSetpairs every evaluated record instead, which for a typical acceptance rate is several times as much data and covers a wider region of parameter space.- Parameters:
chains (ChainSet) – The
ChainSetfor the same run. Matched to series bychain_index.burn (int) – Discard chain steps at or below the
burn-th recorded step of each chain before joining.drop_chains (Sequence[int]) –
chain_indexvalues to exclude entirely.step_offset (int | None) – Added to each record’s step index to obtain the chain step. Detected via
step_offset()when omitted, which is normally what you want: the convention changed between Galacticus versions.proposals (ProposalSet | None) – Proposed states from
read_proposals(), written whenlogProposalsis set on the simulation. When given, rejected-step records are paired too, andPairedPredictions.multiplicityis zero for them — they carry no posterior weight even though they are perfectly good samples of the model’s response. Passweights=Nonetojacobian_from_samples()to use them; weighting by multiplicity would discard exactly the extra coverage they provide.
- Return type:
- class dendros.PredictionSeries(chain_index, path, step, prediction, record_chain=None)[source]
One rank’s model prediction records for a single constraint.
- Parameters:
- path
Source file path.
- Type:
- step
(n_records,)integer step index exactly as recorded in the file. It is not necessarily comparable withdendros.Chain.step— see the module docstring — so usePredictionSet.paired()to align them.- Type:
- record_chain
(n_records,)chain index each record belongs to, when the file records one;Nonefor files written before that column was added. Under load balancing this differs fromchain_index, which is only the process that wrote the file.- Type:
numpy.ndarray | None
- prediction
(n_records, n_bins)model prediction vectors.- Type:
- class dendros.PairedPredictions(label, abscissa, abscissa_name, parameter_names, state, prediction, log_likelihood, chain_index, step, multiplicity)[source]
Model predictions paired with the chain states that produced them.
Contains only records at accepted steps, pooled across ranks; see the module docstring for why rejected-step records cannot be paired.
- Parameters:
- abscissa
(n_bins,)abscissa read from the file header, orNoneif absent.- Type:
numpy.ndarray | None
- state
(n_pairs, n_params)parameter states.- Type:
- log_likelihood
(n_pairs,)total log likelihood logged for each state. This is the sum over all constraints, not just this one.- Type:
- chain_index, step
(n_pairs,)provenance of each pair.
- multiplicity
(n_pairs,)number of consecutive chain steps for which each accepted state was retained — its posterior weight. Use as sample weights when a posterior-averaged quantity is wanted; the raw rows over-represent frequently-rejected regions.- Type:
- dendros.SAMPLE_STEP_OFFSETS = (0, 1)
Built-in immutable sequence.
If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable’s items.
If the argument is a tuple, the return value is the same object.
- dendros.LEGACY_SAMPLE_STEP_OFFSET = 1
int([x]) -> integer int(x, base=10) -> integer
Convert a number or string to an integer, or return 0 if no arguments are given. If x is a number, return x.__int__(). For floating-point numbers, this truncates towards zero.
If x is not a number or if base is given, then x must be a string, bytes, or bytearray instance representing an integer literal in the given base. The literal can be preceded by ‘+’ or ‘-’ and be surrounded by whitespace. The base defaults to 10. Valid bases are 0 and 2-36. Base 0 means to interpret the base from the string as an integer literal. >>> int(‘0b100’, base=0) 4
Proposed states
- dendros.read_proposals(config, *, log_file_root=None)[source]
Read every rank’s proposal log for config.
- Parameters:
config (MCMCConfig) – Parsed
MCMCConfig.log_file_root (str | Path | None) – Override for
config.log_file_root, for a run analysed away from the machine it executed on.
- Return type:
- Raises:
FileNotFoundError – If no proposal logs are found. A run without
logProposalsset writes none; in that case only accepted steps can be paired with predictions.
- dendros.discover_proposal_files(log_file_root)[source]
Return all per-rank proposal logs matching
<root>Proposals_NNNN.log.
- class dendros.ProposalSet(config, series)[source]
An ordered collection of
ProposalSeries, one per MPI rank.- Parameters:
config (MCMCConfig)
series (Sequence[ProposalSeries])
- class dendros.ProposalSeries(chain_index, path, step, accepted, log_posterior, log_likelihood, state)[source]
One rank’s proposed states.
- Parameters:
- path
Source log-file path.
- Type:
- step
(n_proposals,)simulation step at which each proposal was evaluated. Directly comparable withdendros.Chain.step.- Type:
- accepted
(n_proposals,)whether the proposal was accepted.- Type:
- log_posterior, log_likelihood
(n_proposals,)values of the proposal, not of the retained state.
- state
(n_proposals, n_params)proposed parameter vectors, inMCMCConfig.parametersorder and in physical (unmapped) space.- Type:
Count-data likelihoods
- dendros.log_likelihood_bins(count, mu, variance_fractional=0.0)[source]
Per-bin log-likelihood of observing count given model mean mu.
Poisson when
variance_fractional <= 0, else negative binomial with variancemu + variance_fractional * mu**2.Bins with
mu <= 0contribute0where the observed count is also zero, andLOG_IMPROBABLEwhere it is not — matching Galacticus, which treats a positive count against a zero prediction as impossible.- Parameters:
- Returns:
Per-bin log-likelihood.
- Return type:
- dendros.count_variance(mu, variance_fractional=0.0)[source]
Return the per-bin count variance
mu + f mu**2.
- dendros.fisher_weights(mu, variance_fractional=0.0, *, count=None, observed=False)[source]
Return per-bin Fisher weights for the model prediction.
These are the diagonal of the weight matrix
WinJ^T W J, letting the Gaussian propagation indendros._mcmc._fisherbe applied to a count likelihood.- Parameters:
mu (ndarray) – Model-predicted mean counts per bin.
variance_fractional (float) – Fractional model-discrepancy variance
f.count (ndarray | None) – Observed counts. Required when
observed=True.observed (bool) – When
False(the default) return the expected information1 / (mu + f mu^2). WhenTruereturn the observed information-d2 lnL / d mu2evaluated at the data, which tracks the local curvature of the realized likelihood more closely but is noisier and can go negative in poorly-fit bins.
- Returns:
Per-bin weights. Zero where
mu <= 0, those bins carrying no information about the model.- Return type:
- dendros.negative_binomial_shape(variance_fractional)[source]
Return the negative-binomial shape
rfor a fractional model variance.Galacticus parameterizes over-dispersion by the fractional variance
fadded in quadrature to the Poisson term, giving variancemu + f mu^2; the corresponding negative-binomial shape isr = 1 / f. Returnsinfforf <= 0, the Poisson limit.
- dendros.LOG_IMPROBABLE = -1e+30
Convert a string or number to a floating-point number, if possible.
Linearized error propagation
- dendros.jacobian_from_samples(state, prediction, *, weights=None, active=None, degree=1, center=None, rcond=None)[source]
Fit
m(theta)as a polynomial inthetaby least squares over samples.Regressing the recorded predictions on the recorded states gives the model’s response over the posterior volume without any new model evaluations.
Prefer
degree=2unless the model is known to be near-linear. A model that curves appreciably over the posterior volume will have its linear-fit slope biased, because the fit trades curvature against the slopes of correlated parameters; a quadratic fit separates the two and its linear coefficients are the derivative atJacobianFit.center. CompareJacobianFit.nonlinearitybetween degrees to see whether it matters.- Parameters:
state (ndarray) –
(n_samples, n_params)parameter states.prediction (ndarray) –
(n_samples, n_bins)model predictions, row-matched to state.weights (ndarray | None) –
(n_samples,)non-negative sample weights, e.g. the posterior multiplicity of each accepted state. Uniform when omitted.active (Sequence[int] | None) – Column indices of state to fit against. Columns outside this set get a zero Jacobian column. Use it for parameters a constraint cannot see (a Galacticus
parameterMapsubset), which are otherwise fitted to pure noise.degree (int) – Polynomial degree, 1 or 2. Degree 2 adds all cross- and square terms in the active columns, costing
n_active (n_active + 3) / 2coefficients.center (ndarray | None) –
(n_params,)point about which to expand. Defaults to the weight-weighted mean of state. Withdegree=2the reported Jacobian is the derivative at this point, so set it explicitly when the samples are not centred where the derivative is wanted — fitting over proposals (accepted and rejected alike) while expanding about the posterior mean, for instance.rcond (float | None) – Cutoff passed to
numpy.linalg.lstsq().
- Return type:
- class dendros.JacobianFit(jacobian, intercept, center, residual_rms, prediction_rms, n_samples, degree=1)[source]
A linear model of the prediction vector over the posterior volume.
- Parameters:
- jacobian
(n_bins, n_params)matrix ofdm/dtheta.- Type:
- center
(n_params,)parameter values the expansion is about.- Type:
- residual_rms
(n_bins,)root-mean-square residual of the linear fit per bin.- Type:
- prediction_rms
(n_bins,)root-mean-square variation of the prediction about its mean, per bin. The ratio toresidual_rmsmeasures how well the model is described as linear over this volume.- Type:
- dendros.parameter_covariances(jacobian, weights_assumed, covariance_true, *, prior_precision=None, parameter_names=None, rcond=1e-12)[source]
Propagate a data covariance through a linearized model.
- Parameters:
jacobian (ndarray) –
(n_data, n_params)model Jacobian.weights_assumed (ndarray) – The weight matrix the fit used:
(n_data,)for a diagonal weighting (e.g.dendros._mcmc._counts.fisher_weights()) or(n_data, n_data)for a general one.covariance_true (ndarray) –
(n_data, n_data)true data covariance.prior_precision (ndarray | None) –
(n_params, n_params)prior precision to add to each Fisher matrix. Omit for flat priors.parameter_names (Sequence[str] | None) – Optional names, carried through to the result.
rcond (float) – Relative eigenvalue cutoff for the pseudo-inverses. Fisher matrices are singular whenever a parameter is unconstrained by the data — which is normal for a single constraint of a larger fit — so inverses here are Moore–Penrose throughout.
- Return type:
- class dendros.InflationResult(assumed, optimal, sandwich, parameter_names=None)[source]
Parameter covariances under the assumed and true data covariances.
- Parameters:
- assumed
(n_params, n_params)covariance implied by the weighting the fit used. Compare against the chain’s own covariance to validate.- Type:
- optimal
Covariance a correctly-weighted fit would have given.
- Type:
- sandwich
True covariance of the mis-weighted estimator that was used.
- Type:
- dendros.fisher_matrix(jacobian, weights)[source]
Return
J^T W Jfor diagonal (1-D) or full (2-D) weights.
Rank constraints by how much they tighten the posterior.
For each constraint, the fraction by which dropping it would inflate the parameter volume —
(det Sigma_without / det Sigma_all)^(1/2p) - 1. A constraint scoring near zero cannot move the result no matter how its data covariance is treated, which is what makes this the cheap triage step before any covariance work: it is computed from the diagonal weights already in hand.- Parameters:
fishers (Mapping[str, ndarray]) –
{label: (n_params, n_params) Fisher matrix}, additive over independent constraints.prior_precision (ndarray | None) – Added to the total and to each leave-one-out total. Strongly recommended: without a prior, dropping a constraint can leave a singular Fisher matrix and an undefined determinant.
parameters (Sequence[int] | None) – Restrict the comparison to this subset of parameter indices, to ask about the constraints on particular parameters rather than all of them.
rcond (float) – Relative eigenvalue cutoff for the pseudo-inverses.
- Returns:
{label: share}, descending by share.- Return type:
Covariance utilities
- dendros.to_correlation(covariance)[source]
Return the correlation matrix of covariance.
Rows and columns with non-positive variance are returned as zero off the diagonal and one on it, rather than producing NaNs.
- dendros.rescale_correlation(correlation, sigma)[source]
Combine a correlation matrix with a separate set of standard deviations.
C = diag(sigma) R diag(sigma). Use this when the correlation structure and the variances come from different sources — for example a correlation matrix measured from a sub-volume or a subsample, applied at the variances the likelihood actually used. It sidesteps any amplitude mismatch in the estimate, and keeps the resulting covariance’s diagonal consistent with the fit being corrected, which is what makes the validity check indendros._mcmc._fishermeaningful.
- dendros.shrink_to_diagonal(covariance, alpha)[source]
Interpolate between the diagonal of covariance and the full matrix.
alpha=0returns the diagonal (no correlations),alpha=1the input unchanged, and intermediate values scale the off-diagonal terms. Values above 1 strengthen the correlations, which is useful for bracketing: if a parameter-error inflation is flat across a range of alpha, the precise correlation estimate does not matter and a better one is not worth obtaining. Large alpha can leave the matrix indefinite, so check the eigenvalues.
- dendros.nearest_positive_definite(covariance, *, floor=1e-10)[source]
Clip eigenvalues of a symmetric matrix to make it positive definite.
Bootstrap and jackknife covariance estimates are commonly rank-deficient (more bins than independent resamplings, or exactly degenerate bins), leaving zero or slightly negative eigenvalues that break a Cholesky factorization.
- dendros.hartlap_factor(n_realizations, n_data)[source]
Return the Hartlap debiasing factor for an inverse sample covariance.
The inverse of a covariance estimated from
n_realizationsindependent realizations of ann_data-dimensional vector is biased high; multiplying it by(n_realizations - n_data - 2) / (n_realizations - 1)removes the bias. Raises when there are too few realizations for the inverse to exist.
- dendros.dodelson_schneider_factor(n_realizations, n_data, n_params)[source]
Return the variance inflation of parameter errors from a noisy covariance.
Using a covariance estimated from a finite number of realizations inflates parameter variances by
1 + (n_data - n_params) / (n_realizations - n_data - 2)beyond the errors a perfectly known covariance would give. Take the square root for the 1-sigma inflation.
Internal helpers
- class dendros._collection.GroupProxy(collection, path)[source]
Read-only h5py-like proxy for an HDF5 group.
- Parameters:
collection (Collection) – Parent
Collection.path (str) – HDF5 path to the group within the file.
- class dendros._collection.DatasetProxy(collection, path)[source]
Read-only h5py-like proxy for an HDF5 dataset.
For multi-file
Collectioninstances,read()concatenates data from all files along axis 0.- Parameters:
collection (Collection) – Parent
Collection.path (str) – HDF5 path to the dataset within the file.
- property dtype
NumPy dtype of the dataset.
- read(where=None)[source]
Read the dataset into a
numpy.ndarray.For multi-file collections the arrays from all files are concatenated along axis 0 before the optional where selection is applied.
- Parameters:
where –
Nonereads everything. A boolean mask or integer index array is applied after concatenation.- Return type: