# Differences from statsmodels This library follows Durbin and Koopman, *Time Series Analysis by State Space Methods* (2nd ed., 2012). Where statsmodels' `tsa.statespace` makes a different choice, we follow DK. This page lists those choices and the features only one of the two libraries has. The comparison is against **statsmodels 0.15.0** (the version in `.venv`). Section numbers refer to DK. Full citations are in [references.md](references.md). ## Methods in both libraries ### Filter and smoother output | | This library | statsmodels | |---|---|---| | v, F, F∞ in diffuse and univariate periods | In the original coordinates, the same as in conventional periods. The per-element values of the univariate recursion are also stored, in the `uv_*` arrays | In the transformed (element-by-element) coordinates | | K and F⁻¹ in univariate periods | NaN, because the univariate recursion computes neither | The values from the transformed coordinates | | F∞ for a missing element | Z P∞ Z' like any other element | 0 | | Smoothed ε in diffuse and univariate periods | Full matrices, from the exact identity ε_o = y_o − d_o − Z_o α_t | Diagonal variances only, in the transformed coordinates | | Smoothed ε for a missing element | The true conditional moments: mean H u, variance H − H D H (DK 4.5). EM needs these | Mean 0 and variance H, the unconditional moments | | Standardized residuals where undefined (missing, or F = 0 in the diffuse period) | NaN | 0 | | Smoothing weights (4.8.3) in diffuse periods | Computed, including the weights on c and a₁ | NaN | The two libraries agree everywhere else. Across the fixtures, the largest relative difference is about 1e-10. ### Filter variants - **Exact diffuse filter (5.2).** statsmodels always uses the univariate form. Ours also uses it by default. `DIFFUSE_MULTIVARIATE` switches to DK's multivariate form, which falls back to the univariate form in periods where F∞ is singular but nonzero. `fres%method(t)` records which recursion each period used. - **Univariate treatment with correlated H (6.4.3).** Ours works with H_t that is non-diagonal, time-varying and partly missing: each period's observed block gets its own LDL' factorization. - **Diffuse tolerances.** These match statsmodels: 1e-10 on F∞ and on ‖P∞‖²_F. Elements with F\* ≤ 1e-10 are skipped in the diffuse period, as in statsmodels. - **Steady state (DK 4.3.4).** Both conventional filters stop updating P, F and K once P_t converges in a time-invariant model, but they test convergence differently. - statsmodels switches when ‖P_{t+1} − P_t‖²_F < 1e-19. That threshold is absolute, so whether it is reached depends on the scale of the data. - Ours switches when ‖P_{t+1} − P_t‖_F ≤ `tol_steady` · ‖P_{t+1}‖_F, with `tol_steady` = 1e-15 by default. Setting it negative turns the switch off, and `filter_result_t%t_steady` reports the first period that used it. - In ours, a missing observation restarts the full recursion. The ~1e-10 differences in some time-invariant multivariate fixtures come from where the two filters switch. ### Default initialization of the built-in models | | This library | statsmodels UC and SARIMAX | |---|---|---| | Nonstationary states | Exact diffuse (5.2) | Approximate diffuse, with variance 1e6 (1e10 for SARIMAX with state regression). `use_exact_diffuse=True` selects exact diffuse | | Damped cycle | Stationary, as in DK 3.2.4 | Diffuse, like the other UC states | | Log likelihood | Every observation contributes | With the approximate diffuse default, `loglikelihood_burn` equals the number of diffuse states and those observations are left out of `llf`. `llf_obs` still has every period, so the tests compare `llf_obs` | ### Parameter transforms | Parameter | This library | statsmodels | |---|---|---| | Variances | σ² = exp(2ψ) (DK 7.3.2) | σ² = x² | | Full covariance matrices | Σ = LL', with L's diagonal exp(ψ) | UC has no full covariances | | Cycle frequency λ in (λ_lo, λ_hi) | mid + half · ψ/√(1+ψ²) | A logistic between the bounds | | Cycle damping ρ in (0, 1) | ψ/√(1+ψ²), mapped onto (0, 1) | A logistic | | AR and MA coefficients | Monahan's transform; MA = −(the AR transform) | The same, so the values agree to 1e-14 | The constrained parameters mean the same thing in both libraries, so fitted values and log likelihoods agree. The unconstrained values, and the optimizer path, differ. The cycle period bounds default to 2 up to the number of observations. statsmodels uses 1.5 to 12 years when the data's frequency is known, and (2, ∞) otherwise. ### Estimation - **Optimizer.** Both use L-BFGS-B: ours is the Fortran `jacobwilliams/lbfgsb`, and statsmodels uses SciPy's. The defaults are statsmodels': factr = 1e7, pgtol = 1e-5, m = 10, minimizing −llf/n. - **Gradient.** - Ours defaults to DK's analytic score (7.14)/(7.16), which comes from one smoother pass through the Fisher identity. It covers parameters in H, R, Q and a stationary P\*. - Parameters that move Z or T (AR/MA coefficients, cycle frequency and damping, loadings) get central differences, as DK recommend. This happens per parameter, so the variances in the same model keep their analytic score. - statsmodels defaults to complex-step differentiation. Its analytic score is `_score_harvey` (Harvey 1989), which differentiates the filter recursions. - **Standard errors.** Ours come from a finite-difference Hessian in the unconstrained parameters ψ (DK 7.3.6), mapped to the constrained parameters by the delta method. Every finite-difference step is then a valid parameter, even for a variance near zero. statsmodels defaults to the outer product of gradients (`cov_type="opg"`). Its `cov_type="approx"` is closer to ours, but differentiates in the constrained parameters, by complex step. - **Prediction error variance (DK 7.4).** Ours is the steady-state F̄ (DK 2.11, 4.3.4), from `steady_state`. statsmodels has no such function; its conventional filter's steady-state switch is internal to the filter. - **AIC/BIC.** We follow statsmodels: k = k_params + k_diffuse (+1 with a concentrated scale), and n = nobs − burn. DK (7.4) also divide by n; neither library does. ### Built-in models **Unobserved components** (`structural_model_t` vs `UnobservedComponents`) - Ours is multivariate: SUTSE (DK 3.3) with diagonal, full or no disturbance covariance per component, and signal loadings for common levels, latent risk and dynamic factors. statsmodels' UC is univariate. - Ours has the Harrison–Stevens seasonal (3.2.2) alongside the dummy and trigonometric forms. statsmodels' `freq_seasonal` can keep fewer than s/2 harmonics; ours always uses all of them. - Interventions (step, pulse, slope) and regression effects are components in ours. Regression coefficients are diffuse states, either fixed or random walks. statsmodels handles these through `exog`, estimated either by MLE or as states. - statsmodels' `autoregressive` component corresponds to adding an `arima_t` component in ours. **ARIMA** (`arima_t` vs `SARIMAX`) - The state layout is the same, DK 3.4 with the differences held as states, so the states agree one by one. - Ours allows seasonal differencing up to D = 1; statsmodels allows any D. - statsmodels has several options ours lacks: `trend` polynomials, `simple_differencing`, `hamilton_representation` and `measurement_error`. In ours, a trend comes from a `regression_t` component and measurement error from an `irregular_t` component. - SARIMAX with `time_varying_regression` corresponds to `regression_t` with `random_walk`. ### Forecasting Our `forecast` supports only time-invariant models. For a time-varying model, append NaNs to y and filter. statsmodels' `get_forecast` handles time-varying models, given future exog. ## Only in this library | Feature | DK | |---|---| | Augmented Kalman filter and smoother, with δ̂ and Var(δ̂). statsmodels defines `FILTER_AUGMENTED` but doesn't implement it | 5.7 | | Square-root filter and smoother, by Householder QR. statsmodels defines `FILTER_SQUARE_ROOT` but doesn't implement it | 6.3 | | Multivariate exact initial filter and smoother | 5.2–5.3 | | General initialization from any (a₁, P\*, P∞), as well as block-wise mixtures | 5.1 | | Fast, two-filter and Whittle smoothers | 4.6.2–4.6.4 | | Updating smoothed estimates, and fixed-point and fixed-lag smoothers | 4.4.5–4.4.6 | | Filtering weights | 4.8.2 | | de Jong–Shephard disturbance simulation smoother, which also draws α₁ | 4.9.3 | | Dense matrix form of the likelihood and smoother, used as a test oracle | 4.13 | | Linear restrictions on the states | 6.6 | | Likelihood with elements of α₁ fixed but unknown | 7.2.4 | | Marginal likelihood for regression effects (Francke, Koopman and de Vos 2010) | 7.2.6 | | EM for H and Q in any model. statsmodels has EM only in `DynamicFactorMQ` | 7.3.4 | | Effect of parameter estimation errors on smoothed states (bias estimate, antithetic draws) | 7.3.7 | | Auxiliary residuals, element-wise and vector-standardized; de Jong–Penzer statistics | 7.5 | | Least squares residuals | 6.2.4 | | R²_D and prediction error variance | 7.4 | | Harrison–Stevens seasonal; multivariate structural models (SUTSE, common levels, latent risk) | 3.2–3.3 | | Continuous-time local level and smooth trend, and weighted irregulars for unequal spacing | 3.8 | | Discrete and continuous smoothing splines | 3.9 | | Steady-state P̄ by structure-preserving doubling (`steady_state`), independent of the filter | 4.3.4 | | `fit_many`: independent fits in parallel with OpenMP | | ## Only in statsmodels | Feature | Notes | |---|---| | Chandrasekhar recursions (`FILTER_CHANDRASEKHAR`) | Not in DK | | Memory-conservation options, filter timing, choice of matrix inversion method, forced symmetry | Ours always uses Cholesky for F and stores every output | | Chan–Jeliazkov (CFA) simulation smoother | Not in DK | | News and revisions (`news`), smoothed-state decomposition and gains, impulse responses | | | `append`/`extend`/`apply` for new data, `fix_params`, `fit_constrained` | | | Forecasting for time-varying models | See [Forecasting](#forecasting) | | VARMAX, DynamicFactorMQ (mixed frequency), ETS `ExponentialSmoothing`, `RecursiveLS` with CUSUM tests | Our DK 3.5 exponential smoothing is only tested as equivalent to the local level filter. Recursive residuals are the innovations of a regression model | | OPG (the default) and robust covariance types | Ours has only the numerical Hessian | | pandas indexes and plots; prediction objects for time-varying models, news, impulse responses, `append`/`extend` | Out of scope. `ssfortran` returns numpy arrays; it has fitted values, residuals, forecasts with intervals (time-invariant models), component decompositions and a summary with residual tests | ## Comparing results Two conventions change numbers that look comparable: - **`loglikelihood_burn`.** UC and SARIMAX leave the first observations out of `llf` by default, and a fitted SARIMAX's `llf` leaves out the first d. Use `llf_obs` to compare. - **Diffuse damped cycle.** statsmodels starts a damped (stationary) cycle as diffuse. DK 3.2.4 give its stationary distribution, which changes the log likelihood; in one fixture, −65.72 against statsmodels' −63.60.