TEXAS#

TetraEther indeX of Ammonia oxidizerS — Bayesian proxy system model for TEX86 paleothermometry

License PyPI Zenodo

TEXAS-PSM is a Bayesian proxy system model (PSM) for TEX86-based sea surface temperature (SST) reconstruction, built around the TEXAS sensor model. It fits hierarchical generalized-logistic Stan models to isoGDGT Ring Index data — with optional non-thermal corrections for GDGT-2/3 ratio (AOA ecology) and NO₃ (nutrient effect) — and reconstructs paleotemperatures with full posterior uncertainty.

The result is a posterior distribution of temperature for each downcore sample, not just a point estimate with a fixed RMSE.


How it works#

TEXAS uses a two-stage workflow:

Stage 1 — Forward calibration fits a hierarchical Bayesian generalized logistic curve to modern culture, mesocosm, and coretop Ring Index–temperature data. The output is a posterior distribution of calibration parameters saved as a .nc file. Pre-computed posteriors are available on Zenodo — most users can skip this stage entirely.

Stage 2 — Inverse reconstruction passes your downcore Scaled RI measurements through the forward posterior, marginalizing over all calibration parameter uncertainty, and returns a full temperature posterior per sample.


Quickstart#

Install#

pip install texas-psm
# or, with uv:  uv add texas-psm

Or open the interactive notebook in Google Colab — no installation needed:

Open in Colab

For Docker, conda-lock, uv, and development installs see Installation.


Step 1 — Compute Scaled Ring Index#

Before prediction you need Scaled Ring Index (RI₀₋₃) values. Pass raw LC/MS peak areas or fractional abundances — the formula normalises by the six-GDGT total, so either works:

import pandas as pd
from TEXAS import compute_scaledRI

df = pd.read_csv("my_gdgt_data.csv")

df["scaledRI_cren3"] = compute_scaledRI(
    df["GDGT-0"], df["GDGT-1"], df["GDGT-2"], df["GDGT-3"],
    df["cren"],   df["cren_prime"],   # cren_weight=3 by default → RI₀₋₃
)

!!! note “Which Ring Index convention?” The canonical TEXAS posteriors are calibrated against RI₀₋₃ (cren_weight=3, crenarchaeol and its regioisomer weighted 3). Pass cren_weight=4 to reproduce the RI₀₋₄ convention of Zhang et al. (2016), but the canonical posteriors were not calibrated against that convention — the weight also sets the scale the index is expressed on, so the two are not interchangeable inputs.

`cren_weight` was called `cren_rings` before 2026-08-21. The old keyword still works and warns.


Step 2 — Choose a calibration posterior#

You can skip this step. The full multivariate T₀-shift calibration ships inside the package, and is used whenever you do not name another one, so a reconstruction needs no download and no network access:

from TEXAS import predict_T_from_proxyObs

result = predict_T_from_proxyObs(          # fwd_posterior omitted →
    proxyObs=df["scaledRI_cren3"].values,  # tx.GHEB.sst.sri03.G23-N1p0
    prior_mu_t=15.0, prior_sigma_t=10.0,
    gdgt23ratio=df["gdgt23ratio"].values,
    no3=df["no3"].values,
)

Pass temptype="thermoT" to get the thermocline-integrated calibration instead; it ships too. Any other posterior is fetched once from Zenodo and cached:

import TEXAS

# Univariate SST — the temperature-only calibration (<1 MB)
TEXAS.download_posteriors(["tx.GHPU.sst.sri03.p0"])

# Everything in the current Zenodo record (~490 MB; the six
# full multivariate posteriors are ~78-81 MB each)
TEXAS.download_all()

# What you already have — bundled and downloaded alike
TEXAS.list_posteriors()

Available forward posteriors:

Name (no .nc)

Model

Temperature

Size

tx.GHEB.sst.sri03.G23-N1p0

T₀-shift multivariate coretop (default)

SST

bundled

tx.GHEB.thm.sri03.G23-N1p0

T₀-shift multivariate coretop (default)

Thermo T

bundled

tx.GCDU.cul.sri03.p0

Culture + mesocosm

Culture T

<1 MB

tx.GHPU.sst.sri03.p0

Univariate coretop

SST

<1 MB

tx.GHPU.thm.sri03.p0

Univariate coretop

Thermo T

<1 MB

tx.GHEA.sst.sri03.G23-N1p0

Response-offset multivariate coretop (preprint)

SST

~78 MB

tx.GHEA.thm.sri03.G23-N1p0

Response-offset multivariate coretop (preprint)

Thermo T

~78 MB

The bundled pair is the same posterior as the archived one with the error-in-variables model’s per-site latent variables removed — everything the forward and inverse models read is present and unmodified, which is what takes the file from ~80 MB to 0.37 MB. The complete versions, latents included, are archived on Zenodo with the release record that the paper cites.

From v0.3.0 the Zenodo record’s flat files are named by case id (tx.GHPU.sst.sri03.p0.fwd.nc). Version 0.2.0 of the record — the initial submission’s additive fits — carries the older names instead (e.g. gen_logi_fixed_hier_crtp_univ_priorApprox_SST_scaledRI_cren3.nc), and download_posteriors() still resolves those, pinned to that version, so preprint-era notebooks keep working. load_posterior() accepts either spelling, and downloads are unpacked into the case-directory layout shown below.

How posterior files are named#

Posteriors use CESM-style case ids: fixed dot-delimited positions instead of ever-growing concatenated descriptions.

tx . GHEB . sst . sri03 . G23-N1p0 . fwd.nc
│    │      │     │       │         └ role: fwd = forward calibration;
│    │      │     │       │           inv.<site>.<tags> = a reconstruction
│    │      │     │       └ predictors: G23, N1p0 (NO₃, cutoff 1.0 — `p`
│    │      │     │         is the decimal point), p0 = none
│    │      │     └ proxy: sri03 = scaledRI_cren3, sri = scaledRI, tex = TEX86
│    │      └ target temperature: sst, thm (Thermo-T), cul (culture T)
│    └ compset — the model recipe, one letter per axis (see table)
└ project prefix

The four compset letters encode, in order:

Axis

Letters

Curve

G generalized logistic · L logistic · N linear

Training set

H hierarchical coretop · C culture+mesocosm · J culture+mesocosm+coretop · T coretop-only

Estimator

P priorApprox · E priorApprox + errors-in-variables · D full hierarchical

Predictor structure

U univariate · A additive (β on the response, unbounded) · B bounded-by-construction (the T₀-shift structure: γ on T₀)

So tx.GHEB.sst.sri03.G23-N1p0 is the published T₀-shift multivariate SST calibration, and tx.GHPU.sst.sri03.p0 the univariate SST one. A date/run token may be appended to distinguish refits of the same configuration. Every function that takes a posterior name resolves case ids and legacy long names alike.

Note

Two formulations of the non-thermal terms exist, and they are not interchangeable. The current calibration — the bundled GHEB pair — applies G₂/₃ and NO₃ inside the logistic, as a shift of the curve location T₀ in °C (gamma_G23_crtp / gamma_NO3_crtp), so the predicted Scaled RI stays inside its bounds for any predictor value. The GHEA posteriors on Zenodo are the preprint’s response-offset formulation (beta_G23_crtp / beta_NO3_crtp, in Scaled-RI units per predictor unit); they are kept for reproducing the preprint and are documented on the preprint archive page.

You do not have to choose between them by hand: TEXAS reads which formulation a posterior uses from the coefficients it carries and selects the matching inverse model. The univariate and culture/mesocosm posteriors are unaffected either way — they have no non-thermal predictors.


Step 3 — Forward prediction (temperature → proxy)#

Useful for plotting the calibration curve and its uncertainty envelope:

import numpy as np
from TEXAS import predict_proxy_from_T

result = predict_proxy_from_T(
    temperatures=np.linspace(5, 35, 100),
    posterior="tx.GHPU.sst.sri03.p0",
)
result["p50"]   # median Scaled RI (numpy array, length 100)
result["p5"]    # 5th percentile
result["p95"]   # 95th percentile

Step 4 — Inverse reconstruction (proxy → temperature)#

Note

temptype= is optional here. It does not change the reconstruction, which follows whichever calibration you supply — the target is read from that posterior’s own attributes, and a temptype that contradicts it warns. It matters in exactly one case: when fwd_posterior is omitted, it picks which bundled calibration is used, "SST" (the default) or "thermoT". On the forward side, get_posterior(..., temptype=...) is a real argument — that is where the target is decided.

=== “Univariate”

```python
from TEXAS import predict_T_from_proxyObs

result = predict_T_from_proxyObs(
    proxyObs=df["scaledRI_cren3"].values,
    prior_mu_t=15.0,        # prior mean temperature (°C) — geological estimate
    prior_sigma_t=10.0,     # prior uncertainty (°C) — use wide prior if unsure
    fwd_posterior="tx.GHPU.sst.sri03.p0",   # temperature-only calibration
)
result["p50"]   # median SST (°C), one value per sample
result["p5"]    # 5th percentile
result["p95"]   # 95th percentile
```

=== “Multivariate (GDGT-2/3 + NO₃)”

```python
result = predict_T_from_proxyObs(
    proxyObs=df["scaledRI_cren3"].values,
    prior_mu_t=15.0,
    prior_sigma_t=10.0,
    # fwd_posterior omitted → the bundled tx.GHEB.sst.sri03.G23-N1p0
    gdgt23ratio=df["gdgt23ratio"].values,
    no3=df["no3"].values,   # µmol/L; scalar or per-sample array
)
```

=== “NO₃ from WOA23 climatology”

```python
import xarray as xr

ocean_ds = xr.load_dataset("ocean_prop_ds.nc")  # WOA23-derived, from SI_code1

result = predict_T_from_proxyObs(
    proxyObs=df["scaledRI_cren3"].values,
    prior_mu_t=15.0, prior_sigma_t=10.0,
    gdgt23ratio=df["gdgt23ratio"].values,
    site_lat=15.3, site_lon=-23.7,   # modern drill-site coordinates
    no3_dataset=ocean_ds,
)
# Prints: WOA23 NO₃ lookup: lat=15.3, lon=-23.7 → 0.42 µmol/L
```

=== “Load from disk / Google Drive”

If you have a posterior `.nc` file locally or on Google Drive, pass it directly — no cache lookup, no download:

```python
import xarray as xr
from TEXAS import predict_T_from_proxyObs

# Colab: mount Google Drive first, then load
# (files downloaded straight from Zenodo keep the archived legacy name)
ds = xr.load_dataset("/content/drive/MyDrive/posteriors/gen_logi_fixed_hier_crtp_univ_priorApprox_SST_scaledRI_cren3.nc")

result = predict_T_from_proxyObs(
    proxyObs=df["scaledRI_cren3"].values,
    prior_mu_t=15.0, prior_sigma_t=10.0,
    fwd_posterior=ds,    # xr.Dataset — skips all file I/O
)
```

Saving results#

By default predict_T_from_proxyObs returns a dict in memory and writes nothing to disk. Pass save_results=True to persist:

result = predict_T_from_proxyObs(
    ...,
    save_results=True,            # writes quantile .nc + .npz
    save_draws=True,              # also saves raw MCMC draws as _draws.nc
    cache_dir="/your/output/",    # default: ~/.texas/cache/TEXAS_invT_posterior_cache/
)

Running forward calibration from scratch#

Only needed if you want to re-fit the model to your own data or reproduce the published calibration. Requires CmdStan and the GDGT training database (TEXAS.download_training_data()).

from TEXAS import build_fwd_data, get_posterior, save_posterior

data = build_fwd_data(
    t_cul=cul_df["SST"].values,       proxy_cul=cul_df["scaledRI"].values,
    t_meso=meso_df["SST"].values,     proxy_meso=meso_df["scaledRI"].values,
    t_crtp=crtp_df["SST"].values,     proxy_crtp=crtp_df["scaledRI"].values,
    gdgt23ratio_crtp=crtp_df["gdgt23ratio"].values,
    no3_crtp=crtp_df["no3"].values,   # no3_cutoff auto-calculated via Spearman if omitted
)

posterior, diagnostics = get_posterior(
    data,
    stan_file="gen_logi_fixed_hier_crtp_multiv_priorApprox_eiv_t0shift",
    temptype="SST",
    proxy_name="scaledRI_cren3",
)
save_posterior(posterior)

..._eiv_t0shift is the published calibration. The non-thermal predictors enter inside the logistic, shifting the curve’s location parameter T₀ rather than adding an offset to the response:

T₀_eff = T₀ + γ_{G₂/₃}·G₂/₃ + γ_{NO₃}·log₁₀(NO₃)
Scaled RI = b + (1 − b) / (1 + exp(−k·(T − T₀_eff)))^(1/ν)

so the fitted coefficients are gamma_G23_crtp and gamma_NO3_crtp, in °C per predictor unit, and the predicted Scaled RI is confined to (b, 1) by construction. R2_thermal must be supplied to this model — compute it from a thermal-only coretop fit with gen_logi_fixed_hier_crtp_univ_priorApprox.


Citation#

If you use TEXAS in published work, please cite:

Rattanasriampaipong, R. et al. (in prep). TEXAS: A proxy system model for TEX86 paleothermometry. AGU Paleoceanography and Paleoclimatology.

See CITATION.cff for machine-readable metadata.