Differences from statsmodels#
This library follows Durbin and Koopman, Time Series Analysis by State Space Methods
(2nd ed., 2012). Where statsmodels’ tsa.statespace makes a different choice,
we follow DK. This page lists those choices and the features only one of the two
libraries has.
The comparison is against statsmodels 0.15.0 (the version in .venv). Section
numbers refer to DK. Full citations are in references.md.
Methods in both libraries#
Filter and smoother output#
This library |
statsmodels |
|
|---|---|---|
v, F, F∞ in diffuse and univariate periods |
In the original coordinates, the same as in conventional periods. The per-element values of the univariate recursion are also stored, in the |
In the transformed (element-by-element) coordinates |
K and F⁻¹ in univariate periods |
NaN, because the univariate recursion computes neither |
The values from the transformed coordinates |
F∞ for a missing element |
Z P∞ Z’ like any other element |
0 |
Smoothed ε in diffuse and univariate periods |
Full matrices, from the exact identity ε_o = y_o − d_o − Z_o α_t |
Diagonal variances only, in the transformed coordinates |
Smoothed ε for a missing element |
The true conditional moments: mean H u, variance H − H D H (DK 4.5). EM needs these |
Mean 0 and variance H, the unconditional moments |
Standardized residuals where undefined (missing, or F = 0 in the diffuse period) |
NaN |
0 |
Smoothing weights (4.8.3) in diffuse periods |
Computed, including the weights on c and a₁ |
NaN |
The two libraries agree everywhere else. Across the fixtures, the largest relative difference is about 1e-10.
Filter variants#
Exact diffuse filter (5.2). statsmodels always uses the univariate form. Ours also uses it by default.
DIFFUSE_MULTIVARIATEswitches to DK’s multivariate form, which falls back to the univariate form in periods where F∞ is singular but nonzero.fres%method(t)records which recursion each period used.Univariate treatment with correlated H (6.4.3). Ours works with H_t that is non-diagonal, time-varying and partly missing: each period’s observed block gets its own LDL’ factorization.
Diffuse tolerances. These match statsmodels: 1e-10 on F∞ and on ‖P∞‖²_F. Elements with F* ≤ 1e-10 are skipped in the diffuse period, as in statsmodels.
Steady state (DK 4.3.4). Both conventional filters stop updating P, F and K once P_t converges in a time-invariant model, but they test convergence differently.
statsmodels switches when ‖P_{t+1} − P_t‖²_F < 1e-19. That threshold is absolute, so whether it is reached depends on the scale of the data.
Ours switches when ‖P_{t+1} − P_t‖F ≤
tol_steady· ‖P{t+1}‖_F, withtol_steady= 1e-15 by default. Setting it negative turns the switch off, andfilter_result_t%t_steadyreports the first period that used it.In ours, a missing observation restarts the full recursion. The ~1e-10 differences in some time-invariant multivariate fixtures come from where the two filters switch.
Default initialization of the built-in models#
This library |
statsmodels UC and SARIMAX |
|
|---|---|---|
Nonstationary states |
Exact diffuse (5.2) |
Approximate diffuse, with variance 1e6 (1e10 for SARIMAX with state regression). |
Damped cycle |
Stationary, as in DK 3.2.4 |
Diffuse, like the other UC states |
Log likelihood |
Every observation contributes |
With the approximate diffuse default, |
Parameter transforms#
Parameter |
This library |
statsmodels |
|---|---|---|
Variances |
σ² = exp(2ψ) (DK 7.3.2) |
σ² = x² |
Full covariance matrices |
Σ = LL’, with L’s diagonal exp(ψ) |
UC has no full covariances |
Cycle frequency λ in (λ_lo, λ_hi) |
mid + half · ψ/√(1+ψ²) |
A logistic between the bounds |
Cycle damping ρ in (0, 1) |
ψ/√(1+ψ²), mapped onto (0, 1) |
A logistic |
AR and MA coefficients |
Monahan’s transform; MA = −(the AR transform) |
The same, so the values agree to 1e-14 |
The constrained parameters mean the same thing in both libraries, so fitted values and log likelihoods agree. The unconstrained values, and the optimizer path, differ.
The cycle period bounds default to 2 up to the number of observations. statsmodels uses 1.5 to 12 years when the data’s frequency is known, and (2, ∞) otherwise.
Estimation#
Optimizer. Both use L-BFGS-B: ours is the Fortran
jacobwilliams/lbfgsb, and statsmodels uses SciPy’s. The defaults are statsmodels’: factr = 1e7, pgtol = 1e-5, m = 10, minimizing −llf/n.Gradient.
Ours defaults to DK’s analytic score (7.14)/(7.16), which comes from one smoother pass through the Fisher identity. It covers parameters in H, R, Q and a stationary P*.
Parameters that move Z or T (AR/MA coefficients, cycle frequency and damping, loadings) get central differences, as DK recommend. This happens per parameter, so the variances in the same model keep their analytic score.
statsmodels defaults to complex-step differentiation. Its analytic score is
_score_harvey(Harvey 1989), which differentiates the filter recursions.
Standard errors. Ours come from a finite-difference Hessian in the unconstrained parameters ψ (DK 7.3.6), mapped to the constrained parameters by the delta method. Every finite-difference step is then a valid parameter, even for a variance near zero. statsmodels defaults to the outer product of gradients (
cov_type="opg"). Itscov_type="approx"is closer to ours, but differentiates in the constrained parameters, by complex step.Prediction error variance (DK 7.4). Ours is the steady-state F̄ (DK 2.11, 4.3.4), from
steady_state. statsmodels has no such function; its conventional filter’s steady-state switch is internal to the filter.AIC/BIC. We follow statsmodels: k = k_params + k_diffuse (+1 with a concentrated scale), and n = nobs − burn. DK (7.4) also divide by n; neither library does.
Built-in models#
Unobserved components (structural_model_t vs UnobservedComponents)
Ours is multivariate: SUTSE (DK 3.3) with diagonal, full or no disturbance covariance per component, and signal loadings for common levels, latent risk and dynamic factors. statsmodels’ UC is univariate.
Ours has the Harrison–Stevens seasonal (3.2.2) alongside the dummy and trigonometric forms. statsmodels’
freq_seasonalcan keep fewer than s/2 harmonics; ours always uses all of them.Interventions (step, pulse, slope) and regression effects are components in ours. Regression coefficients are diffuse states, either fixed or random walks. statsmodels handles these through
exog, estimated either by MLE or as states.statsmodels’
autoregressivecomponent corresponds to adding anarima_tcomponent in ours.
ARIMA (arima_t vs SARIMAX)
The state layout is the same, DK 3.4 with the differences held as states, so the states agree one by one.
Ours allows seasonal differencing up to D = 1; statsmodels allows any D.
statsmodels has several options ours lacks:
trendpolynomials,simple_differencing,hamilton_representationandmeasurement_error. In ours, a trend comes from aregression_tcomponent and measurement error from anirregular_tcomponent.SARIMAX with
time_varying_regressioncorresponds toregression_twithrandom_walk.
Forecasting#
Our forecast supports only time-invariant models. For a time-varying model, append
NaNs to y and filter. statsmodels’ get_forecast handles time-varying models, given
future exog.
Only in this library#
Feature |
DK |
|---|---|
Augmented Kalman filter and smoother, with δ̂ and Var(δ̂). statsmodels defines |
5.7 |
Square-root filter and smoother, by Householder QR. statsmodels defines |
6.3 |
Multivariate exact initial filter and smoother |
5.2–5.3 |
General initialization from any (a₁, P*, P∞), as well as block-wise mixtures |
5.1 |
Fast, two-filter and Whittle smoothers |
4.6.2–4.6.4 |
Updating smoothed estimates, and fixed-point and fixed-lag smoothers |
4.4.5–4.4.6 |
Filtering weights |
4.8.2 |
de Jong–Shephard disturbance simulation smoother, which also draws α₁ |
4.9.3 |
Dense matrix form of the likelihood and smoother, used as a test oracle |
4.13 |
Linear restrictions on the states |
6.6 |
Likelihood with elements of α₁ fixed but unknown |
7.2.4 |
Marginal likelihood for regression effects (Francke, Koopman and de Vos 2010) |
7.2.6 |
EM for H and Q in any model. statsmodels has EM only in |
7.3.4 |
Effect of parameter estimation errors on smoothed states (bias estimate, antithetic draws) |
7.3.7 |
Auxiliary residuals, element-wise and vector-standardized; de Jong–Penzer statistics |
7.5 |
Least squares residuals |
6.2.4 |
R²_D and prediction error variance |
7.4 |
Harrison–Stevens seasonal; multivariate structural models (SUTSE, common levels, latent risk) |
3.2–3.3 |
Continuous-time local level and smooth trend, and weighted irregulars for unequal spacing |
3.8 |
Discrete and continuous smoothing splines |
3.9 |
Steady-state P̄ by structure-preserving doubling ( |
4.3.4 |
|
Only in statsmodels#
Feature |
Notes |
|---|---|
Chandrasekhar recursions ( |
Not in DK |
Memory-conservation options, filter timing, choice of matrix inversion method, forced symmetry |
Ours always uses Cholesky for F and stores every output |
Chan–Jeliazkov (CFA) simulation smoother |
Not in DK |
News and revisions ( |
|
|
|
Forecasting for time-varying models |
See Forecasting |
VARMAX, DynamicFactorMQ (mixed frequency), ETS |
Our DK 3.5 exponential smoothing is only tested as equivalent to the local level filter. Recursive residuals are the innovations of a regression model |
OPG (the default) and robust covariance types |
Ours has only the numerical Hessian |
pandas indexes and plots; prediction objects for time-varying models, news, impulse responses, |
Out of scope. |
Comparing results#
Two conventions change numbers that look comparable:
loglikelihood_burn. UC and SARIMAX leave the first observations out ofllfby default, and a fitted SARIMAX’sllfleaves out the first d. Usellf_obsto compare.Diffuse damped cycle. statsmodels starts a damped (stationary) cycle as diffuse. DK 3.2.4 give its stationary distribution, which changes the log likelihood; in one fixture, −65.72 against statsmodels’ −63.60.