Diagnostic checking#

Residuals#

The standardized one-step prediction errors \(e_t = L_t^{-1} v_t\), \(F_t = L_t L_t'\), are independent standard normal when the model is correct (DK §2.12, §7.5). They are NaN in the diffuse periods and where observations are missing.

>>> y = np.loadtxt("data/nile.csv", delimiter=",", skiprows=1)[:, 1]
>>> res = ss.StructuralModel(y, [ss.Irregular(), ss.Level()]).fit()
>>> e = res.standardized_residuals
>>> bool(np.isnan(e[0]))
True

Tests#

diagnostics() applies three tests to the residuals after the diffuse periods (DK §2.12.1): Ljung-Box for serial correlation, Jarque-Bera for normality and H for heteroskedasticity. They are also available on any series in ssfortran.diagnostics, and summary() reports them:

>>> print(res.summary())
============...
Model:          StructuralModel         Log likelihood:       -633.4646
Observations:   100                     AIC:                  1272.9291
Diffuse periods:1                       BIC:                  1280.7446
============...

Auxiliary residuals#

The smoothed disturbances, standardized by their variances (auxiliary_residuals()), point to outliers (observation disturbances) and structural breaks (state disturbances) (DK §2.12.2, §7.5). The statistics of de Jong and Penzer (1998) (de_jong_penzer()) do the same for every state. In the Nile series the auxiliary level residual is largest in 1899, when the Aswan dam was built:

>>> eps, eta = res.model.representation(res.params).auxiliary_residuals()
>>> int(1871 + np.nanargmax(np.abs(eta[0])))
1898

The state disturbance of 1898 moves the level of 1899.

Goodness of fit#

r2_diffuse() compares the one-step prediction errors with those of a random walk with drift (Harvey 1989); steady_state() gives the prediction error variance of DK §7.4.