AR Filtering or Lag Controls? What the Frisch-Waugh-Lovell Theorem Really Says

When studying geopolitical risk, should we first filter GPR through an AR(12) model and use the residual as the shock, or should we enter current GPR and its twelve lags directly in the outcome equation? In a linear OLS setting, the two approaches can be coefficient-equivalent. But this equivalence has nothing to do with structural exogeneity.

Researchers working with monthly geopolitical risk often face a seemingly simple choice. One approach is to estimate an autoregressive model for GPR, retain the residual, and label that residual a geopolitical-risk innovation. Another is to use current GPR directly while controlling for its previous twelve monthly values.

At first sight, these procedures look different. The first is presented as a two-step shock-construction exercise; the second is an ordinary distributed-lag regression. Yet the Frisch-Waugh-Lovell theorem shows that, under carefully matched linear specifications, they use the same variation in current GPR and produce the same coefficient.

This result is useful, but it is also easy to overinterpret. AR filtering removes the component of GPR that is linearly predictable from the specified information set. It does not establish that the remaining innovation is independent of the omitted determinants of the economic outcome. In short:

Unpredictability is not exogeneity.

This post develops these two points in a panel setting and then applies them to panel local projections.

1. Two ways to handle persistent GPR

Let \(g_{i,t}\) denote the country-specific GPR index for country \(i\) in month \(t\). Let \(y_{i,t+h}\) be the economic or financial outcome at horizon \(h\). For example, \(y\) could be a sovereign yield, investment, industrial production, inflation, or an exchange rate.

Monthly GPR can be persistent. A geopolitical confrontation reported this month may continue to dominate the news during subsequent months. A researcher may therefore want to distinguish the new movement in GPR from the component already predictable from its history.

Approach A: include current GPR and its twelve lags

In a panel local projection, the direct specification is:

\[ y_{i,t+h} = \alpha_{i,h}+\lambda_{t,h} +\beta_h g_{i,t} +\sum_{j=1}^{12}\delta_{h,j}g_{i,t-j} +\Gamma_h’X_{i,t-1} +\varepsilon_{i,t+h}. \]

Here, \(\alpha_{i,h}\) denotes country fixed effects, \(\lambda_{t,h}\) denotes time fixed effects, and \(X_{i,t-1}\) contains any additional controls. Because the twelve GPR lags are held constant, \(\beta_h\) is identified from movements in current GPR that are not explained by those lags and the other included regressors.

Approach B: construct an AR(12) innovation first

The alternative is to estimate an auxiliary AR(12) projection:

\[ g_{i,t} =a_i+\sum_{j=1}^{12}\rho_jg_{i,t-j}+u_{i,t}, \]

and define the unstandardized innovation as:

\[ \widehat{u}_{i,t} =g_{i,t} -\widehat{E}\!\left( g_{i,t}\mid g_{i,t-1},\ldots,g_{i,t-12},a_i \right). \]

The second-stage local projection replaces current GPR with this innovation:

\[ y_{i,t+h} =\alpha_{i,h}+\lambda_{t,h} +\beta_h^{u}\widehat{u}_{i,t} +\sum_{j=1}^{12}\widetilde{\delta}_{h,j}g_{i,t-j} +\Gamma_h’X_{i,t-1} +e_{i,t+h}. \]

If the specifications are aligned correctly, then:

\[ \widehat{\beta}_h^{u}=\widehat{\beta}_h. \]

The separate filtering step has not created a new source of identification. It has only reparametrized the same linear regression.

2. Why the coefficients are identical

The intuition follows directly from the AR decomposition:

\[ g_{i,t} =\widehat{g}_{i,t}+\widehat{u}_{i,t}, \]

where \(\widehat{g}_{i,t}\) is a linear combination of the country effect and the twelve lagged GPR values. Substituting this decomposition into the direct outcome equation gives:

\[ \begin{aligned} y_{i,t+h} ={}&\alpha_{i,h}+\lambda_{t,h} +\beta_h\left(\widehat{g}_{i,t}+\widehat{u}_{i,t}\right)\\ &+\sum_{j=1}^{12}\delta_{h,j}g_{i,t-j} +\Gamma_h’X_{i,t-1} +\varepsilon_{i,t+h}. \end{aligned} \]

Because \(\widehat{g}_{i,t}\) is itself constructed from regressors already included in the equation, its contribution can be absorbed into their coefficients. The coefficient multiplying \(\widehat{u}_{i,t}\) remains \(\beta_h\). Thus, the two specifications span the same regressor space.

This is the regression anatomy formalized by Frisch and Waugh (1933) and later associated with Lovell (1963).

3. The Frisch-Waugh-Lovell interpretation

Write the outcome equation compactly as:

\[ y=\beta g+Z\gamma+\varepsilon, \]

where \(Z\) collects all nuisance regressors: the twelve GPR lags, fixed effects, lagged outcomes, macroeconomic controls, trends, and any other variables included in the specification.

Define the residual-maker matrix:

\[ M_Z=I-Z(Z’Z)^{-1}Z’. \]

Partialling \(Z\) out of current GPR gives:

\[ \widetilde{g}=M_Zg. \]

Partialling the same \(Z\) out of the outcome gives:

\[ \widetilde{y}=M_Zy. \]

The Frisch-Waugh-Lovell theorem states that the coefficient on current GPR in the complete regression is exactly the coefficient obtained from:

\[ \widetilde{y}=\beta\widetilde{g}+\widetilde{\varepsilon}. \]

Therefore, the coefficient on current GPR conditional on its lags is estimated from the component of current GPR left after those lags have been partialled out. This is precisely the sense in which including the lags directly performs AR filtering inside the outcome regression.

Key result. In a linear OLS panel regression, including current GPR and its twelve lags or replacing current GPR with the corresponding unstandardized AR(12) innovation produces the same coefficient when the parameterization is matched and the filtering variables are retained in the outcome equation. The same result follows from a pure FWL regression when both GPR and the outcome are partialled out using the same information set and estimation sample.

4. What “the same information set” requires

The equivalence is exact, not approximate, but only under the corresponding linear specifications. Several details matter.

  • The same lags: an AR(12) residual is not equivalent to a regression controlling for only two GPR lags.
  • The same parameterization: a pooled panel AR with common lag coefficients is not the same as estimating a separate AR(12) for every country. A direct regression with common lag coefficients corresponds to the pooled filter.
  • The same fixed effects: country and time effects included in the filtering projection must either be retained in the outcome regression or included in the common residualization set.
  • The same controls: if GPR is residualized only against its lags, those lags should remain in the outcome regression. Alternatively, both GPR and the outcome can be residualized against the complete set of lags, fixed effects, and controls.
  • The relevant estimation sample: if the filtering variables remain in the outcome equation, coefficient equivalence follows from the linear reparametrization even when a fixed AR projection was estimated on a broader sample. If those variables are dropped and one instead uses a residual-on-residual FWL regression, both residualizations must be performed on the exact outcome-equation sample. This distinction matters in local projections because the usable sample changes with \(h\).
  • The same scale: standardizing the AR residual changes the coefficient’s units. If \(z_{i,t}=\widehat{u}_{i,t}/s_u\), then the response to a one-standard-deviation innovation is \(s_u\widehat{\beta}_h\), not \(\widehat{\beta}_h\).

There is also a distinction between coefficient equivalence and a stripped-down two-step inference procedure. When the filtered regression retains the complete conditioning set, it is simply a linear reparametrization: fitted values and residuals coincide, and the covariance estimate for the corresponding coefficient also coincides up to numerical precision. If the researcher instead runs only the residualized outcome on residualized GPR, the coefficient is still reproduced by FWL, but degrees-of-freedom and covariance calculations must be adjusted consistently. For applied work, the safest procedure is usually to estimate the complete outcome equation directly and use the desired clustered or heteroskedasticity-robust covariance estimator there.

Does monthly frequency automatically imply twelve lags?

No. AR(12) is a convenient way to condition on one year of monthly GPR history, but monthly frequency does not mechanically justify twelve lags. The lag length should reflect the economic application and should be assessed using information criteria, residual-serial-correlation diagnostics, and robustness to shorter alternatives such as AR(1), AR(3), or AR(6). The FWL equivalence applies to whichever lag set is chosen; it does not determine which information set is substantively appropriate.

5. A simple numerical intuition

Suppose the previous twelve months predict a country GPR value of 80, but the observed value is 110:

\[ \widehat{g}_{i,t}=80,\qquad g_{i,t}=110. \]

The AR innovation is:

\[ \widehat{u}_{i,t}=110-80=30. \]

Now consider another country whose GPR is 150 but whose past values predict 140. Its innovation is only 10. Although the second country has the higher GPR level, the first country experiences the larger movement relative to its own recent history.

A regression containing current GPR and the twelve lags uses exactly this conditional comparison. Once past GPR is held constant, the coefficient on current GPR is identified by deviations such as 30 and 10—not simply by the observed levels of 110 and 150.

6. Application to panel local projections

Jordà’s (2005) local-projection method estimates a separate regression for every horizon:

\[ y_{i,t+h} =\alpha_{i,h}+\lambda_{t,h} +\beta_hg_{i,t}+Z_{i,t}’\gamma_h +\varepsilon_{i,t+h}, \qquad h=0,1,\ldots,H. \]

The FWL theorem applies separately at every \(h\). At each horizon, \(\beta_h\) is estimated from current GPR after the controls relevant to that regression have been partialled out.

This has an important practical implication. A single AR-filtered GPR series constructed before the LP exercise remains coefficient-equivalent if the filtering variables are retained in every horizon-specific equation. However, if the lags are subsequently dropped and the researcher relies on residualization alone, the filtering must be aligned with each horizon-specific sample. In either case, including current GPR and its lags directly makes the conditioning set more transparent.

For country-specific GPR, full time fixed effects can also be included because GPR varies across countries within a given month. These effects absorb global shocks common to all countries. For a purely global GPR series that varies only over time, full monthly time effects would instead absorb the shock itself.

7. Unpredictability is not exogeneity

The most important conceptual point is that AR residualization addresses predictability—not causal identification.

After estimating the auxiliary projection, OLS ensures the sample orthogonality conditions:

\[ \sum_{i,t}\widehat{u}_{i,t}g_{i,t-j}=0, \qquad j=1,\ldots,12. \]

The residual is therefore linearly orthogonal, in the estimation sample, to the regressors used in the AR filter. With a correctly specified conditional-mean model, one may write the stronger population condition:

\[ E\!\left(u_{i,t}\mid g_{i,t-1},\ldots,g_{i,t-12}\right)=0. \]

Neither statement implies the exogeneity condition required for a causal interpretation of the outcome equation:

\[ E\!\left(u_{i,t}\varepsilon_{i,t+h}\mid Z_{i,t}\right)=0. \]

The two conditions concern different relationships:

  • Unpredictability: can current GPR be forecast from the variables included in the filtering model?
  • Exogeneity: is the GPR innovation independent of the omitted causes of the outcome?

An innovation can satisfy the first condition while violating the second.

8. The converse case: exogenous but predictable

The converse is equally important: a variable need not be unpredictable to be exogenous. Predictability concerns its conditional mean, whereas exogeneity concerns its relationship with the error in the outcome equation.

Consider a persistent country GPR process:

\[ g_{i,t}=\rho g_{i,t-1}+\eta_{i,t}, \qquad |\rho|<1. \]

Because

\[ \mathbb{E}\!\left(g_{i,t}\mid g_{i,t-1}\right) =\rho g_{i,t-1}, \]

current GPR is predictable whenever \(\rho\neq0\). For example, if \(\rho=0.8\) and last month’s GPR equals 100, the forecast for this month is 80. If \(\eta_{i,t}=5\), observed GPR is 85: 80 was predictable and only 5 represents new information. Neither component is necessarily endogenous.

Now consider the outcome equation:

\[ y_{i,t+h} =\alpha_{i,h}+\lambda_{t,h} +\beta_h g_{i,t} +\delta_h g_{i,t-1} +X_{i,t-1}’\gamma_h +\varepsilon_{i,t+h}. \]

Suppose that, conditional on past GPR and the controls, neither the predictable component nor the new GPR innovation is correlated with the outcome-equation error:

\[ \begin{aligned} \mathbb{E}\!\left( \varepsilon_{i,t+h} \mid g_{i,t-1},X_{i,t-1} \right)&=0,\\ \mathbb{E}\!\left( \eta_{i,t}\varepsilon_{i,t+h} \mid g_{i,t-1},X_{i,t-1} \right)&=0. \end{aligned} \]

It follows that current GPR is conditionally exogenous even though it is predictable:

\[ \begin{aligned} &\mathbb{E}\!\left( g_{i,t}\varepsilon_{i,t+h} \mid g_{i,t-1},X_{i,t-1} \right)\\ &\quad= \rho g_{i,t-1} \mathbb{E}\!\left( \varepsilon_{i,t+h} \mid g_{i,t-1},X_{i,t-1} \right) +\mathbb{E}\!\left( \eta_{i,t}\varepsilon_{i,t+h} \mid g_{i,t-1},X_{i,t-1} \right) =0. \end{aligned} \]

A stylized GPR example is a preannounced sequence of foreign military exercises or sanctions stages whose calendar is determined outside a small economy. Country GPR may rise in a persistent and partly forecastable way around that sequence. If short-run innovations in the small economy cannot affect the foreign timetable, and relevant common determinants are properly controlled for, this GPR exposure can be exogenous to the domestic outcome error despite being predictable from its past.

Anticipation still matters for the interpretation. Markets may respond when the timetable is announced rather than when each stage occurs. The exposure may therefore be exogenous, while a local projection dated at the implementation month fails to capture the complete dynamic effect. Exogeneity does not imply surprise, and it does not determine the appropriate event date.

Strictly speaking, the predictable component should be called an exogenous exposure or exogenous treatment, not an exogenous structural shock. A structural shock is normally defined as an innovation relative to an information set. The logical distinction can be summarized as:

\[ \text{unpredictable}\ \not\!\Rightarrow\ \text{exogenous}, \qquad \text{predictable}\ \not\!\Rightarrow\ \text{endogenous}. \]

9. A counterexample: an unpredictable but endogenous innovation

Let \(q_t\) be an unexpected global disturbance. Suppose it affects both GPR coverage and the economic outcome:

\[ g_{i,t}=\pi’L_{i,t-1}+\kappa q_t+\nu_{i,t}, \]

where \(L_{i,t-1}\) contains the twelve GPR lags. If both \(q_t\) and \(\nu_{i,t}\) are unpredictable from those lags, the AR residual is:

\[ u_{i,t}=\kappa q_t+\nu_{i,t}. \]

Now suppose the outcome is:

\[ y_{i,t+h}=\beta_hu_{i,t}+\theta_hq_t+\eta_{i,t+h}. \]

If \(q_t\) is omitted, it becomes part of the outcome-equation error. Therefore:

\[ \operatorname{Cov}(u_{i,t},\varepsilon_{i,t+h}) =\kappa\theta_h\operatorname{Var}(q_t)\neq0. \]

The AR residual remains unpredictable from past GPR, but it is endogenous because it contains the omitted disturbance \(q_t\).

In a GPR application, the substantive defense of exogeneity must come from the nature and timing of geopolitical events, not from the AR filter. Researchers may argue that wars, terrorist attacks, military escalations, or diplomatic confrontations are not caused by short-run movements in the economic outcome. They may reinforce this argument using predetermined controls, time fixed effects when country-specific variation is available, lead tests, narrative timing, high-frequency identification, recursive restrictions, or external instruments. The influential GPR analysis of Caldara and Iacoviello (2022), for example, combines the nature of geopolitical events with explicit dynamic identification and validation exercises.

But none of these identifying arguments is supplied mechanically by residualization.

10. Reduced-form innovation versus structural shock

The terminology matters. An AR residual is appropriately described as:

  • a forecast error relative to a specified linear model;
  • an innovation relative to the chosen information set; or
  • the component of GPR linearly unpredictable from the included lags.

None of these descriptions is equivalent to calling the residual an exogenous structural shock. The distinction is familiar from VAR analysis. Consider the reduced-form VAR:

\[ \mathbf{x}_t =A_1\mathbf{x}_{t-1} +\cdots +A_p\mathbf{x}_{t-p} +\mathbf{u}_t, \qquad \mathbb{E}\!\left( \mathbf{u}_t\mid\mathcal{I}_{t-1} \right)=\mathbf{0}. \]

The residual vector \(\mathbf{u}_t\) is unpredictable from the variables contained in the lagged information set \(\mathcal{I}_{t-1}\). Nevertheless, its elements will generally be contemporaneously correlated:

\[ \Sigma_u =\mathbb{E}\!\left( \mathbf{u}_t\mathbf{u}_t’ \right). \]

This means that the reduced-form residual associated with one variable may combine several underlying economic disturbances. Structural identification introduces a mapping between the reduced-form innovations and a vector of mutually orthogonal structural shocks:

\[ \underbrace{\mathbf{u}_t}_{\text{reduced-form innovations}} = \underbrace{B}_{\text{contemporaneous impact matrix}} \underbrace{\boldsymbol{\varepsilon}_t}_{\text{structural shocks}}, \qquad \mathbb{E}\!\left( \boldsymbol{\varepsilon}_t\boldsymbol{\varepsilon}_t’ \right)=I. \]

It follows that:

\[ \Sigma_u=BB’. \]

Estimating the reduced-form VAR identifies \(\Sigma_u\), but it does not uniquely identify \(B\). For a system containing \(n\) variables, the symmetric covariance matrix provides \(n(n+1)/2\) distinct moments, whereas \(B\) contains \(n^2\) elements. Consequently, \(n(n-1)/2\) additional restrictions are required to recover the structural shocks. These may come from a recursive Cholesky ordering, sign restrictions, narrative restrictions, long-run restrictions, heteroskedasticity, or an external instrument.

The same identification problem exists even when GPR is filtered through a univariate AR model. A scalar GPR residual may still combine several underlying disturbances:

\[ u_t^{GPR} = b_G\varepsilon_t^{G} +b_F\varepsilon_t^{F} +b_O\varepsilon_t^{O} +b_M\varepsilon_t^{M}, \]

where \(\varepsilon_t^{G}\) denotes a genuine geopolitical disturbance, while the other terms may represent unexpected financial stress, oil-market developments, or changes in media attention. All these components can be unpredictable from past GPR. The AR filter cannot determine which of them generated the observed innovation.

Calling \(u_t^{GPR}\) an exogenous geopolitical shock therefore requires an additional identification argument. Researchers may defend a short-run exogeneity assumption using the nature and timing of wars, terrorist attacks, military escalations, or diplomatic confrontations. Alternatively, they may use contemporaneous recursive restrictions, narrative information, high-frequency identification, or external instruments. These arguments may be convincing, but they come from the economic research design—not from the AR residualization itself.

Once the structural shock has been identified, the same shock can be used in a VAR, a Bayesian VAR, or a local projection. The estimators differ in how they trace the subsequent dynamics, not necessarily in how the shock is identified. I discuss this distinction in more detail in the EconMacro post Can Local Projections Have Short-Run Restrictions?

AR filtering identifies a forecast error. Structural identification determines what economic disturbance produced that forecast error.

11. A Stata illustration

The following schematic example uses monthly country data. Replace the variable names with those in the application.

* Panel structure
xtset country month

capture drop outcome_h sample_h gpr_ar12 z_gpr_ar12

* Example: outcome at horizon h
local h = 6
generate double outcome_h = F`h'.outcome

* Additional controls
global controls L(1/2).inflation L(1/2).growth

* ------------------------------------------------------------
* A. Direct specification: current GPR plus 12 GPR lags
* ------------------------------------------------------------
reghdfe outcome_h gpr L(1/12).gpr $controls, ///
    absorb(country month) ///
    vce(cluster country)

scalar b_direct = _b[gpr]
generate byte sample_h = e(sample)

* ------------------------------------------------------------
* B. AR(12) filtering, followed by the equivalent regression
* ------------------------------------------------------------
reghdfe gpr L(1/12).gpr if sample_h, ///
    absorb(country month) ///
    residuals(gpr_ar12)

reghdfe outcome_h gpr_ar12 L(1/12).gpr $controls if sample_h, ///
    absorb(country month) ///
    vce(cluster country)

scalar b_filtered = _b[gpr_ar12]

display "Direct coefficient  = " %12.10f b_direct
display "Filtered coefficient = " %12.10f b_filtered

Up to numerical precision, the coefficients should coincide because \(gpr\) and \(gpr\_ar12\) differ only by a linear combination of regressors retained in the second equation.

If the residual is standardized, the coefficients will no longer be numerically identical:

summarize gpr_ar12 if sample_h, meanonly
generate double z_gpr_ar12 = gpr_ar12 / r(sd)

reghdfe outcome_h z_gpr_ar12 L(1/12).gpr $controls if sample_h, ///
    absorb(country month) ///
    vce(cluster country)

The new coefficient measures the response to a one-standard-deviation GPR innovation. It equals the unstandardized coefficient multiplied by the residual standard deviation.

12. Common mistakes

  1. Filtering GPR and then dropping the conditioning variables. Regressing the outcome only on an AR residual is not generally equivalent to the complete distributed-lag model. The filtering variables must be retained, or the complete conditioning set must be partialled out consistently.
  2. Ignoring horizon-specific samples after dropping the lags. If the filtering variables are retained, a fixed linear reparametrization remains valid on the LP subsample. If they are omitted and the analysis relies on residualization alone, GPR and the outcome must be partialled out on the same horizon-specific sample.
  3. Mixing pooled and country-specific dynamics. Separate country AR models allow twelve country-specific slope coefficients; a pooled panel regression usually does not.
  4. Ignoring standardization. Standardization preserves the underlying variation but changes the coefficient’s scale.
  5. Calling every residual a structural shock. Orthogonality to included lags is not orthogonality to the outcome-equation error.
  6. Treating FWL as a free shortcut in a VAR or BVAR. Univariate filtering of GPR does not reproduce the cross-variable lag structure of a VAR. In a BVAR, shrinkage priors also make the exercise different from an OLS reparametrization. Parameter proliferation should be addressed with a justified lag choice or shrinkage—not by silently discarding the relevant information set.

13. Conclusion

The Frisch-Waugh-Lovell theorem provides a clean way to understand AR filtering in panel regressions and panel local projections. Including current GPR together with its twelve lags already identifies the coefficient from the component of GPR not explained by those lags. Constructing an AR(12) residual first can be an equivalent reparametrization, but it is not an additional identification strategy.

The concepts are logically distinct in both directions. An unpredictable GPR innovation can still be endogenous, while a persistent and forecastable GPR exposure can be conditionally exogenous. The crucial distinction is therefore:

AR filtering asks whether a GPR movement was predictable from the specified past information. Exogeneity asks whether that movement is independent of the omitted determinants of the outcome. Only the second condition supports a causal interpretation.

For applied work, my preference is to include current GPR and its lags directly in the outcome equation. This keeps the information set visible, respects the horizon-specific sample in local projections, facilitates lag-length robustness, and avoids giving an ordinary reduced-form residual a stronger structural interpretation than the research design can support.

References

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.