# Three Underrated Statsmodels Methods Every Time Series Analyst Should Know
Fitted statistical models in Python carry more information than most practitioners ever extract from them. When you call a forecasting method, the returned object has already done the hard work of estimating uncertainty, handling seasonal structures, and storing parameter estimates. All too often, analysts manually rebuild these components — introducing errors, duplicating effort, and wasting compute time in the process.
This guide walks through three powerful methods that are already sitting inside the statsmodels results object, waiting to be used correctly. Each example works on the same monthly dataset and the same fitted model, so you can clearly see how switching the method call changes the output — nothing else.
The code has been verified against statsmodels version 0.15.0.
## Getting Started
Install statsmodels and its dependencies if you haven’t already:
“`python
pip install statsmodels
“`
Load a monthly CO₂ concentration dataset, resample it to the start of each month, and split it into a training set and a holdout set of the last twelve observations:
“`python
import statsmodels.api as sm
from statsmodels.tsa.arima.model import ARIMA
co2 = sm.datasets.co2.load_pandas().data[“co2”]
co2 = co2.resample(“MS”).mean().ffill()
train, recent = co2[:-12], co2[-12:]
res = ARIMA(train, order=(1, 1, 1), seasonal_order=(1, 1, 1, 12)).fit()
“`
Everything that follows builds on this single fitted model.
—
## Method 1: Retrieving Confidence Intervals Alongside Predictions
The simplest call, `res.forecast(12)`, returns only a flat array of point predictions for the next twelve months. If you want the uncertainty bands that the model already computed, you need a different accessor.
Replace `forecast` with `get_forecast`, and the returned object is a `PredictionResults` instance. This object carries the standard error, the confidence interval bounds, and more — all from the same internal computation that generated the point forecast. No extra work is being done behind the scenes; the short-cut method simply discards this information when you use the default call.
“`python
forecast_df = res.get_forecast(12).summary_frame()
print(forecast_df.head())
“`
The output gives you four columns at once: the predicted mean, its standard error, and the lower and upper confidence bounds.
| Date | Mean | Std Err | CI Lower | CI Upper |
|————|———-|———|———–|———–|
| 2001-01-01 | 370.52 | 0.32 | 369.89 | 371.16 |
| 2001-02-01 | 371.25 | 0.39 | 370.49 | 372.01 |
| 2001-03-01 | 372.20 | 0.43 | 371.36 | 373.04 |
| 2001-04-01 | 373.47 | 0.46 | 372.56 | 374.38 |
| 2001-05-01 | 373.86 | 0.49 | 372.89 | 374.83 |
A forecast published without uncertainty intervals is an incomplete forecast. The intervals are not an optional extra — they are part of the same model output.
If instead of a forward-looking forecast you want fitted values across the historical training period, use `get_prediction`. It works on any range of observations, including in-sample dates, and returns the same rich structure with prediction intervals for each point.
—
## Method 2: Incorporating New Observations Without Retraining
When twelve months of fresh data arrive, the instinctive response is to append them to the original training set and run `.fit()` again from scratch. This re-estimates every parameter, consuming time and potentially destabilizing a model that was already well-calibrated.
The `append` method offers a smarter path. It constructs a new results object over the combined historical and new data, and when called with `refit=False` — the default — it preserves every coefficient from the original fit:
“`python
updated = res.append(recent, refit=False)
print(updated.get_forecast(6).summary_frame().head())
“`
| Date | Mean | Std Err | CI Lower | CI Upper |
|————|———-|———|———–|———–|
| 2002-01-01 | 371.97 | 0.32 | 371.34 | 372.60 |
| 2002-02-01 | 372.75 | 0.39 | 371.99 | 373.51 |
| 2002-03-01 | 373.65 | 0.43 | 372.81 | 374.50 |
| 2002-04-01 | 374.83 | 0.46 | 373.93 | 375.74 |
| 2002-05-01 | 375.33 | 0.49 | 374.36 | 376.30 |
Use `refit=True` only when enough new data has accumulated that you genuinely want the parameters recomputed.
Statsmodels provides three closely related methods, each suited to a different scenario:
– **`append`** — Re-runs the Kalman filter across both the original and new observations. Best when the history is moderate in length.
– **`extend`** — Filters only the new observations, skipping a full pass over the existing history. Ideal when the training set is large and the new data is a small addition.
– **`apply`** — Designed for a completely different dataset, not a continuation of the same series.
Choosing the right variant saves computation and avoids unnecessary refitting.
—
## Method 3: Delegating Seasonality Handling to STLForecast
Handling seasonality manually is a three-step chore: decompose the series, forecast the seasonally adjusted portion, and then re-add the seasonal component to the forecast. The third step is where sign mistakes, index misalignments, and off-by-one errors creep in — especially when the decomposition and the forecasting model are not synchronized properly.
`STLForecast` collapses all three steps into a single object. Under the hood, it uses Seasonal-Trend decomposition via Loess to remove the seasonal pattern, fits the specified time-series model to the deseasonalized data, and then reconstructs the forecast by adding the seasonal component back. Everything is handled consistently and automatically.
“`python
from statsmodels.tsa.forecasting.stl import STLForecast
stlf = STLForecast(train, ARIMA, model_kwargs={“order”: (1, 1, 1), “trend”: “t”})
print(stlf.fit().forecast(12).head())
“`
The output:
“`
2001-01-01 370.53
2001-02-01 370.96
2001-03-01 371.92
2001-04-01 373.14
2001-05-01 373.14
Freq: MS, dtype: float64
“`
There is one API detail that catches most newcomers off guard: you pass the **model class** itself — `ARIMA` — not a pre-fitted instance. The model arguments are supplied separately through `model_kwargs`. Passing `ARIMA(…)` with parentheses is the most common mistake made with this class.
—
## Bringing It All Together
Every technique described above is a method that already exists on objects you have already constructed. Writing custom logic to replicate what these methods do is longer, slower, and prone to subtle bugs. The reason most people end up reimplementing these steps is that they never take the time to read what the fitted results object actually contains.
Explore the object returned by `.fit()`. Browse its attributes and methods. You will almost always find that the tool you need already exists — and using it correctly is faster and more reliable than building something from scratch.
—
## Frequently Asked Questions
**Q: Why does `forecast()` return fewer values than `get_forecast()`?**
A: `forecast()` returns only the point predictions. `get_forecast()` returns a rich object that includes the same point predictions plus the standard error and confidence interval. The underlying computation is identical; `forecast()` simply strips away the extra information.
**Q: Can I use `append` with models other than ARIMA?**
A: Yes. Any model that produces a results object compatible with statsmodels’ state-space or time-series framework can generally be used with `append`, `extend`, and `apply`. The key requirement is that the underlying model supports filtering over combined data.
**Q: What is the difference between `get_prediction` and `get_forecast`?**
A: `get_prediction` can generate prediction intervals for any range of observations — including those already seen in the training data (in-sample). `get_forecast` is specifically for out-of-sample periods that lie beyond the end of the fitted data.
**Q: When should I switch `refit` from `False` to `True` in `append`?**
A: Keep `refit=False` when the new data is a small addition relative to the existing dataset and you want to preserve the stability of your original estimates. Switch to `refit=True` when the new data volume is large enough that the parameters should be reconsidered — for example, after a structural shift in the data-generating process.
**Q: Does `STLForecast` handle trends automatically?**
A: Yes. The STL decomposition step handles both seasonal and trend components. You control the trend smoothing through the `trend` parameter and can pass any additional ARIMA configuration via `model_kwargs`.
**Q: Can I change the confidence level for the intervals returned by `summary_frame()`?**
A: Absolutely. Pass an `alpha` parameter to `get_forecast()` or `get_prediction()` — for example, `res.get_forecast(12, alpha=0.1)` produces 90% confidence intervals instead of the default 95%.
—
## Conclusion
Statsmodels is more than a collection of fitting functions — it is a framework where fitted models carry a wealth of derived information. By using the right accessor methods, analysts can extract prediction intervals, update models with new data efficiently, and delegate seasonal decomposition to a single unified class. These practices reduce code complexity, minimize the chance of errors, and produce more informative outputs. The best tool for the job is often the one already built into the library; the key is knowing it exists and understanding how to call it.
Thank you for reading



