When studying geopolitical risk, should we first filter GPR through an AR(12) model and use the residual as the shock, or should we enter current GPR and its twelve lags directly in the outcome equation? In a linear OLS setting, the two approaches can be coefficient-equivalent. But this equivalence has nothing to do with structural exogeneity.
Researchers working with monthly geopolitical risk often face a seemingly simple choice. One approach is to estimate an autoregressive model for GPR, retain the residual, and label that residual a geopolitical-risk innovation. Another is to use current GPR directly while controlling for its previous twelve monthly values.
At first sight, these procedures look different. The first is presented as a two-step shock-construction exercise; the second is an ordinary distributed-lag regression. Yet the Frisch-Waugh-Lovell theorem shows that, under carefully matched linear specifications, they use the same variation in current GPR and produce the same coefficient.
This result is useful, but it is also easy to overinterpret. AR filtering removes the component of GPR that is linearly predictable from the specified information set. It does not establish that the remaining innovation is independent of the omitted determinants of the economic outcome. In short:
Unpredictability is not exogeneity.
This post develops these two points in a panel setting and then applies them to panel local projections.
1. Two ways to handle persistent GPR
Let \(g_{i,t}\) denote the country-specific GPR index for country \(i\) in month \(t\). Let \(y_{i,t+h}\) be the economic or financial outcome at horizon \(h\). For example, \(y\) could be a sovereign yield, investment, industrial production, inflation, or an exchange rate.
Monthly GPR can be persistent. A geopolitical confrontation reported this month may continue to dominate the news during subsequent months. A researcher may therefore want to distinguish the new movement in GPR from the component already predictable from its history.
Approach A: include current GPR and its twelve lags
In a panel local projection, the direct specification is:
Here, \(\alpha_{i,h}\) denotes country fixed effects, \(\lambda_{t,h}\) denotes time fixed effects, and \(X_{i,t-1}\) contains any additional controls. Because the twelve GPR lags are held constant, \(\beta_h\) is identified from movements in current GPR that are not explained by those lags and the other included regressors.
Approach B: construct an AR(12) innovation first
The alternative is to estimate an auxiliary AR(12) projection:
and define the unstandardized innovation as:
The second-stage local projection replaces current GPR with this innovation:
If the specifications are aligned correctly, then:
The separate filtering step has not created a new source of identification. It has only reparametrized the same linear regression.
2. Why the coefficients are identical
The intuition follows directly from the AR decomposition:
where \(\widehat{g}_{i,t}\) is a linear combination of the country effect and the twelve lagged GPR values. Substituting this decomposition into the direct outcome equation gives:
Because \(\widehat{g}_{i,t}\) is itself constructed from regressors already included in the equation, its contribution can be absorbed into their coefficients. The coefficient multiplying \(\widehat{u}_{i,t}\) remains \(\beta_h\). Thus, the two specifications span the same regressor space.
This is the regression anatomy formalized by Frisch and Waugh (1933) and later associated with Lovell (1963).
3. The Frisch-Waugh-Lovell interpretation
Write the outcome equation compactly as:
where \(Z\) collects all nuisance regressors: the twelve GPR lags, fixed effects, lagged outcomes, macroeconomic controls, trends, and any other variables included in the specification.
Define the residual-maker matrix:
Partialling \(Z\) out of current GPR gives:
Partialling the same \(Z\) out of the outcome gives:
The Frisch-Waugh-Lovell theorem states that the coefficient on current GPR in the complete regression is exactly the coefficient obtained from:
Therefore, the coefficient on current GPR conditional on its lags is estimated from the component of current GPR left after those lags have been partialled out. This is precisely the sense in which including the lags directly performs AR filtering inside the outcome regression.
Key result. In a linear OLS panel regression, including current GPR and its twelve lags or replacing current GPR with the corresponding unstandardized AR(12) innovation produces the same coefficient when the parameterization is matched and the filtering variables are retained in the outcome equation. The same result follows from a pure FWL regression when both GPR and the outcome are partialled out using the same information set and estimation sample.
4. What “the same information set” requires
The equivalence is exact, not approximate, but only under the corresponding linear specifications. Several details matter.
- The same lags: an AR(12) residual is not equivalent to a regression controlling for only two GPR lags.
- The same parameterization: a pooled panel AR with common lag coefficients is not the same as estimating a separate AR(12) for every country. A direct regression with common lag coefficients corresponds to the pooled filter.
- The same fixed effects: country and time effects included in the filtering projection must either be retained in the outcome regression or included in the common residualization set.
- The same controls: if GPR is residualized only against its lags, those lags should remain in the outcome regression. Alternatively, both GPR and the outcome can be residualized against the complete set of lags, fixed effects, and controls.
- The relevant estimation sample: if the filtering variables remain in the outcome equation, coefficient equivalence follows from the linear reparametrization even when a fixed AR projection was estimated on a broader sample. If those variables are dropped and one instead uses a residual-on-residual FWL regression, both residualizations must be performed on the exact outcome-equation sample. This distinction matters in local projections because the usable sample changes with \(h\).
- The same scale: standardizing the AR residual changes the coefficient’s units. If \(z_{i,t}=\widehat{u}_{i,t}/s_u\), then the response to a one-standard-deviation innovation is \(s_u\widehat{\beta}_h\), not \(\widehat{\beta}_h\).
There is also a distinction between coefficient equivalence and a stripped-down two-step inference procedure. When the filtered regression retains the complete conditioning set, it is simply a linear reparametrization: fitted values and residuals coincide, and the covariance estimate for the corresponding coefficient also coincides up to numerical precision. If the researcher instead runs only the residualized outcome on residualized GPR, the coefficient is still reproduced by FWL, but degrees-of-freedom and covariance calculations must be adjusted consistently. For applied work, the safest procedure is usually to estimate the complete outcome equation directly and use the desired clustered or heteroskedasticity-robust covariance estimator there.
Does monthly frequency automatically imply twelve lags?
No. AR(12) is a convenient way to condition on one year of monthly GPR history, but monthly frequency does not mechanically justify twelve lags. The lag length should reflect the economic application and should be assessed using information criteria, residual-serial-correlation diagnostics, and robustness to shorter alternatives such as AR(1), AR(3), or AR(6). The FWL equivalence applies to whichever lag set is chosen; it does not determine which information set is substantively appropriate.
5. A simple numerical intuition
Suppose the previous twelve months predict a country GPR value of 80, but the observed value is 110:
The AR innovation is:
Now consider another country whose GPR is 150 but whose past values predict 140. Its innovation is only 10. Although the second country has the higher GPR level, the first country experiences the larger movement relative to its own recent history.
A regression containing current GPR and the twelve lags uses exactly this conditional comparison. Once past GPR is held constant, the coefficient on current GPR is identified by deviations such as 30 and 10—not simply by the observed levels of 110 and 150.
6. Application to panel local projections
Jordà’s (2005) local-projection method estimates a separate regression for every horizon:
The FWL theorem applies separately at every \(h\). At each horizon, \(\beta_h\) is estimated from current GPR after the controls relevant to that regression have been partialled out.
This has an important practical implication. A single AR-filtered GPR series constructed before the LP exercise remains coefficient-equivalent if the filtering variables are retained in every horizon-specific equation. However, if the lags are subsequently dropped and the researcher relies on residualization alone, the filtering must be aligned with each horizon-specific sample. In either case, including current GPR and its lags directly makes the conditioning set more transparent.
For country-specific GPR, full time fixed effects can also be included because GPR varies across countries within a given month. These effects absorb global shocks common to all countries. For a purely global GPR series that varies only over time, full monthly time effects would instead absorb the shock itself.
7. Unpredictability is not exogeneity
The most important conceptual point is that AR residualization addresses predictability—not causal identification.
After estimating the auxiliary projection, OLS ensures the sample orthogonality conditions:
The residual is therefore linearly orthogonal, in the estimation sample, to the regressors used in the AR filter. With a correctly specified conditional-mean model, one may write the stronger population condition:
Neither statement implies the exogeneity condition required for a causal interpretation of the outcome equation:
The two conditions concern different relationships:
- Unpredictability: can current GPR be forecast from the variables included in the filtering model?
- Exogeneity: is the GPR innovation independent of the omitted causes of the outcome?
An innovation can satisfy the first condition while violating the second.
8. The converse case: exogenous but predictable
The converse is equally important: a variable need not be unpredictable to be exogenous. Predictability concerns its conditional mean, whereas exogeneity concerns its relationship with the error in the outcome equation.
Consider a persistent country GPR process:
Because
current GPR is predictable whenever \(\rho\neq0\). For example, if \(\rho=0.8\) and last month’s GPR equals 100, the forecast for this month is 80. If \(\eta_{i,t}=5\), observed GPR is 85: 80 was predictable and only 5 represents new information. Neither component is necessarily endogenous.
Now consider the outcome equation:
Suppose that, conditional on past GPR and the controls, neither the predictable component nor the new GPR innovation is correlated with the outcome-equation error:
It follows that current GPR is conditionally exogenous even though it is predictable:
A stylized GPR example is a preannounced sequence of foreign military exercises or sanctions stages whose calendar is determined outside a small economy. Country GPR may rise in a persistent and partly forecastable way around that sequence. If short-run innovations in the small economy cannot affect the foreign timetable, and relevant common determinants are properly controlled for, this GPR exposure can be exogenous to the domestic outcome error despite being predictable from its past.
Anticipation still matters for the interpretation. Markets may respond when the timetable is announced rather than when each stage occurs. The exposure may therefore be exogenous, while a local projection dated at the implementation month fails to capture the complete dynamic effect. Exogeneity does not imply surprise, and it does not determine the appropriate event date.
Strictly speaking, the predictable component should be called an exogenous exposure or exogenous treatment, not an exogenous structural shock. A structural shock is normally defined as an innovation relative to an information set. The logical distinction can be summarized as:
9. A counterexample: an unpredictable but endogenous innovation
Let \(q_t\) be an unexpected global disturbance. Suppose it affects both GPR coverage and the economic outcome:
where \(L_{i,t-1}\) contains the twelve GPR lags. If both \(q_t\) and \(\nu_{i,t}\) are unpredictable from those lags, the AR residual is:
Now suppose the outcome is:
If \(q_t\) is omitted, it becomes part of the outcome-equation error. Therefore:
The AR residual remains unpredictable from past GPR, but it is endogenous because it contains the omitted disturbance \(q_t\).
In a GPR application, the substantive defense of exogeneity must come from the nature and timing of geopolitical events, not from the AR filter. Researchers may argue that wars, terrorist attacks, military escalations, or diplomatic confrontations are not caused by short-run movements in the economic outcome. They may reinforce this argument using predetermined controls, time fixed effects when country-specific variation is available, lead tests, narrative timing, high-frequency identification, recursive restrictions, or external instruments. The influential GPR analysis of Caldara and Iacoviello (2022), for example, combines the nature of geopolitical events with explicit dynamic identification and validation exercises.
But none of these identifying arguments is supplied mechanically by residualization.
10. Reduced-form innovation versus structural shock
The terminology matters. An AR residual is appropriately described as:
- a forecast error relative to a specified linear model;
- an innovation relative to the chosen information set; or
- the component of GPR linearly unpredictable from the included lags.
None of these descriptions is equivalent to calling the residual an exogenous structural shock. The distinction is familiar from VAR analysis. Consider the reduced-form VAR:
The residual vector \(\mathbf{u}_t\) is unpredictable from the variables contained in the lagged information set \(\mathcal{I}_{t-1}\). Nevertheless, its elements will generally be contemporaneously correlated:
This means that the reduced-form residual associated with one variable may combine several underlying economic disturbances. Structural identification introduces a mapping between the reduced-form innovations and a vector of mutually orthogonal structural shocks:
It follows that:
Estimating the reduced-form VAR identifies \(\Sigma_u\), but it does not uniquely identify \(B\). For a system containing \(n\) variables, the symmetric covariance matrix provides \(n(n+1)/2\) distinct moments, whereas \(B\) contains \(n^2\) elements. Consequently, \(n(n-1)/2\) additional restrictions are required to recover the structural shocks. These may come from a recursive Cholesky ordering, sign restrictions, narrative restrictions, long-run restrictions, heteroskedasticity, or an external instrument.
The same identification problem exists even when GPR is filtered through a univariate AR model. A scalar GPR residual may still combine several underlying disturbances:
where \(\varepsilon_t^{G}\) denotes a genuine geopolitical disturbance, while the other terms may represent unexpected financial stress, oil-market developments, or changes in media attention. All these components can be unpredictable from past GPR. The AR filter cannot determine which of them generated the observed innovation.
Calling \(u_t^{GPR}\) an exogenous geopolitical shock therefore requires an additional identification argument. Researchers may defend a short-run exogeneity assumption using the nature and timing of wars, terrorist attacks, military escalations, or diplomatic confrontations. Alternatively, they may use contemporaneous recursive restrictions, narrative information, high-frequency identification, or external instruments. These arguments may be convincing, but they come from the economic research design—not from the AR residualization itself.
Once the structural shock has been identified, the same shock can be used in a VAR, a Bayesian VAR, or a local projection. The estimators differ in how they trace the subsequent dynamics, not necessarily in how the shock is identified. I discuss this distinction in more detail in the EconMacro post Can Local Projections Have Short-Run Restrictions?
AR filtering identifies a forecast error. Structural identification determines what economic disturbance produced that forecast error.
11. A Stata illustration
The following schematic example uses monthly country data. Replace the variable names with those in the application.
* Panel structure
xtset country month
capture drop outcome_h sample_h gpr_ar12 z_gpr_ar12
* Example: outcome at horizon h
local h = 6
generate double outcome_h = F`h'.outcome
* Additional controls
global controls L(1/2).inflation L(1/2).growth
* ------------------------------------------------------------
* A. Direct specification: current GPR plus 12 GPR lags
* ------------------------------------------------------------
reghdfe outcome_h gpr L(1/12).gpr $controls, ///
absorb(country month) ///
vce(cluster country)
scalar b_direct = _b[gpr]
generate byte sample_h = e(sample)
* ------------------------------------------------------------
* B. AR(12) filtering, followed by the equivalent regression
* ------------------------------------------------------------
reghdfe gpr L(1/12).gpr if sample_h, ///
absorb(country month) ///
residuals(gpr_ar12)
reghdfe outcome_h gpr_ar12 L(1/12).gpr $controls if sample_h, ///
absorb(country month) ///
vce(cluster country)
scalar b_filtered = _b[gpr_ar12]
display "Direct coefficient = " %12.10f b_direct
display "Filtered coefficient = " %12.10f b_filtered
Up to numerical precision, the coefficients should coincide because \(gpr\) and \(gpr\_ar12\) differ only by a linear combination of regressors retained in the second equation.
If the residual is standardized, the coefficients will no longer be numerically identical:
summarize gpr_ar12 if sample_h, meanonly
generate double z_gpr_ar12 = gpr_ar12 / r(sd)
reghdfe outcome_h z_gpr_ar12 L(1/12).gpr $controls if sample_h, ///
absorb(country month) ///
vce(cluster country)
The new coefficient measures the response to a one-standard-deviation GPR innovation. It equals the unstandardized coefficient multiplied by the residual standard deviation.
12. Common mistakes
- Filtering GPR and then dropping the conditioning variables. Regressing the outcome only on an AR residual is not generally equivalent to the complete distributed-lag model. The filtering variables must be retained, or the complete conditioning set must be partialled out consistently.
- Ignoring horizon-specific samples after dropping the lags. If the filtering variables are retained, a fixed linear reparametrization remains valid on the LP subsample. If they are omitted and the analysis relies on residualization alone, GPR and the outcome must be partialled out on the same horizon-specific sample.
- Mixing pooled and country-specific dynamics. Separate country AR models allow twelve country-specific slope coefficients; a pooled panel regression usually does not.
- Ignoring standardization. Standardization preserves the underlying variation but changes the coefficient’s scale.
- Calling every residual a structural shock. Orthogonality to included lags is not orthogonality to the outcome-equation error.
- Treating FWL as a free shortcut in a VAR or BVAR. Univariate filtering of GPR does not reproduce the cross-variable lag structure of a VAR. In a BVAR, shrinkage priors also make the exercise different from an OLS reparametrization. Parameter proliferation should be addressed with a justified lag choice or shrinkage—not by silently discarding the relevant information set.
13. Conclusion
The Frisch-Waugh-Lovell theorem provides a clean way to understand AR filtering in panel regressions and panel local projections. Including current GPR together with its twelve lags already identifies the coefficient from the component of GPR not explained by those lags. Constructing an AR(12) residual first can be an equivalent reparametrization, but it is not an additional identification strategy.
The concepts are logically distinct in both directions. An unpredictable GPR innovation can still be endogenous, while a persistent and forecastable GPR exposure can be conditionally exogenous. The crucial distinction is therefore:
AR filtering asks whether a GPR movement was predictable from the specified past information. Exogeneity asks whether that movement is independent of the omitted determinants of the outcome. Only the second condition supports a causal interpretation.
For applied work, my preference is to include current GPR and its lags directly in the outcome equation. This keeps the information set visible, respects the horizon-specific sample in local projections, facilitates lag-length robustness, and avoids giving an ordinary reduced-form residual a stronger structural interpretation than the research design can support.
References
- Caldara, D., & Iacoviello, M. (2022). Measuring geopolitical risk. American Economic Review, 112(4), 1194–1225.
- Frisch, R., & Waugh, F. V. (1933). Partial time regressions as compared with individual trends. Econometrica, 1(4), 387–401.
- Jordà, Ò. (2005). Estimation and inference of impulse responses by local projections. American Economic Review, 95(1), 161–182.
- Lovell, M. C. (1963). Seasonal adjustment of economic time series and multiple regression analysis. Journal of the American Statistical Association, 58(304), 993–1010.
Appendix A. The Frisch–Waugh–Lovell theorem, step by step
This appendix derives the Frisch–Waugh–Lovell (FWL) result used in the main text. The purpose is to show exactly why including current GPR together with its lags uses the same conditional variation as an appropriately matched residualization procedure. The derivation also clarifies the slightly different—but closely related—argument behind replacing current GPR with an AR residual while retaining the AR regressors.
The result to be established. In a linear model estimated by OLS, the coefficient on current GPR can be obtained either from the complete regression or from a regression using the parts of the outcome and current GPR that remain after the nuisance regressors have been partialled out. Replacing current GPR by an unstandardized AR residual also leaves this coefficient unchanged when all variables used in the AR filter remain in the outcome equation.
A.1. Stack the panel local projection
At horizon \(h\), let \(\mathcal S_h\) denote the set of usable country–month observations and let \(n_h=|\mathcal S_h|\). Stack the horizon-specific local projection as:
Here, \(\mathbf y_h\) stacks the outcomes \(y_{i,t+h}\), \(\mathbf g_h\) stacks current GPR \(g_{i,t}\), and \(\mathbf Z_h\) collects every nuisance regressor: the twelve GPR lags, additional controls, and the fixed effects.
Suppose that the horizon-\(h\) sample contains \(I_h\) countries and \(T_h\) origin dates, while \(K\) denotes the number of additional control columns. Using an intercept, \(I_h-1\) country dummies, and \(T_h-1\) time dummies, the fixed-effect block has \(I_h+T_h-1\) linearly independent columns.
| Object | Content | Dimension |
|---|---|---|
| \(\mathbf y_h\) | Outcome at horizon \(h\) | \(n_h\times1\) |
| \(\mathbf g_h\) | Current GPR | \(n_h\times1\) |
| \(\mathbf L_h\) | Twelve GPR lags | \(n_h\times12\) |
| \(\mathbf C_h\) | Additional controls | \(n_h\times K\) |
| \(\mathbf F_h\) | Intercept and two-way fixed effects | \(n_h\times(I_h+T_h-1)\) |
| \(\mathbf Z_h=[\mathbf L_h\;\mathbf C_h\;\mathbf F_h]\) | Complete nuisance matrix | \(n_h\times k_h\) |
The number of nuisance regressors is therefore:
The complete design matrix adds current GPR as one further column:
What does \(\mathbb R^{n_h\times(k_h+1)}\) mean?
The symbol \(\mathbb R\) denotes the set of real numbers. The expression \(\mathbb R^{n_h\times(k_h+1)}\) denotes the set of all real-valued matrices with \(n_h\) rows and \(k_h+1\) columns. It therefore describes the shape of the complete design matrix:
- the \(n_h\) rows correspond to the usable country–month observations at horizon \(h\);
- the first column is current GPR, \(\mathbf g_h\);
- the remaining \(k_h\) columns are the nuisance regressors collected in \(\mathbf Z_h\).
Two uses of the word dimension should be distinguished. The matrix size is \(n_h\times(k_h+1)\). As a vector space, \(\mathbb R^{n_h\times(k_h+1)}\) has dimension \(n_h(k_h+1)\), because a matrix of this size contains that many independently selectable real entries. For example, if \(n_h=1{,}000\) and \(k_h=20\), then \(\mathbf X_h\in\mathbb R^{1000\times21}\): it has 1,000 rows, 21 columns, and 21,000 entries. In the regression, the relevant description is normally its matrix size, \(1000\times21\).
Provided that the design has full column rank, the residual degrees of freedom are:
The subtraction of one in \(I_h+T_h-1\) reflects the fixed-effect normalization. One valid coding uses an intercept, \(I_h-1\) country dummies, and \(T_h-1\) time dummies. Another uses all \(I_h\) country dummies, \(T_h-1\) time dummies, and no separate intercept. If all country and time dummies are included without an intercept, one linear dependence must be removed; adding an intercept to all of those dummies would create two dependencies. Commands that absorb fixed effects avoid constructing these columns explicitly, but the underlying rank is still \(I_h+T_h-1\).
Rank conditions
Writing an ordinary inverse such as \((\mathbf Z_h^{\prime}\mathbf Z_h)^{-1}\) requires the nuisance matrix to have full column rank:
Identifying the coefficient on current GPR also requires current GPR not to be a linear combination of the nuisance regressors:
Equivalently, current GPR must retain positive residual variation after conditioning on \(\mathbf Z_h\):
Why full rank is equivalent to positive residual variation
Let \(\mathbf M_{Z,h}=\mathbf I_{n_h}-\mathbf P_{Z,h}\) denote the residual-maker for \(\mathbf Z_h\), as derived formally below, and define the residualized GPR vector:
Because a residual-maker is symmetric and idempotent, \(\mathbf M_{Z,h}^{\prime}=\mathbf M_{Z,h}\) and \(\mathbf M_{Z,h}^2=\mathbf M_{Z,h}\). It follows that:
The quadratic form is therefore exactly the residual sum of squares from projecting current GPR on \(\mathbf Z_h\). It cannot be negative. It is strictly positive if and only if at least one residualized GPR observation is nonzero:
Now, \(\mathbf M_{Z,h}\mathbf g_h=\mathbf0\) holds if and only if current GPR lies entirely in the column space of \(\mathbf Z_h\). In that case \(\mathbf g_h=\mathbf P_{Z,h}\mathbf g_h\), so some coefficient vector \(\mathbf a\) satisfies \(\mathbf g_h=\mathbf Z_h\mathbf a\). Adding \(\mathbf g_h\) to \(\mathbf Z_h\) then adds no new direction: the complete design matrix is not full column rank. Conversely, if \(\mathbf M_{Z,h}\mathbf g_h\neq\mathbf0\), current GPR is not a linear combination of the nuisance regressors and adds one independent column. Provided \(\mathbf Z_h\) already has full column rank:
Two languages for the same requirement. In rank language, current GPR must not be an exact linear combination of the nuisance regressors. In residual language, current GPR must not disappear completely after those regressors have been partialled out.
Global versus country-specific GPR with monthly fixed effects
The distinction is especially important in a panel with complete monthly fixed effects. Consider two countries observed in three months. If GPR is global, its value is identical across countries within each month:
| Month | Country A | Country B |
|---|---|---|
| 1 | 10 | 10 |
| 2 | 20 | 20 |
| 3 | 15 | 15 |
With observations stacked month by month, the GPR vector is:
where \(\mathbf D_t\) is the dummy vector for month \(t\). Thus, global GPR is an exact linear combination of the monthly dummies. Monthly fixed effects reproduce it perfectly, implying:
There is no variation left with which to distinguish the GPR coefficient from the monthly effects. By contrast, suppose GPR is country-specific:
| Month | Country A | Country B |
|---|---|---|
| 1 | 8 | 12 |
| 2 | 18 | 22 |
| 3 | 12 | 18 |
The monthly means are 10, 20, and 15. Removing the monthly fixed effects leaves:
Why does the quadratic form equal 34?
Define the residualized GPR vector by:
Although the expression \(\mathbf g^{\prime}\mathbf M_Z\mathbf g\) contains the original vector \(\mathbf g\) on the left, it is exactly the squared length of the residualized vector. This follows because the residual-maker is symmetric and idempotent:
Consequently:
The inner product of a vector with itself is the sum of the squares of its elements. Hence:
The same result can be verified by multiplying the original GPR vector by the residualized GPR vector. With the observations ordered by month and then by country:
Therefore:
Why does multiplying the original vector by the residual vector give the residual sum of squares? Decompose current GPR into its fitted and residual components:
The fitted monthly means are:
OLS fitted values are orthogonal to OLS residuals, so:
It follows that:
Interpretation. The inequality \(34>0\) does not mean that every residual is positive: some residuals are negative and some are positive. It means that the residual sum of squares is positive. This occurs because country-specific GPR is not completely explained by the monthly fixed effects. The quadratic form would equal zero only if every residual were zero—that is, only if GPR had no cross-country variation within any month.
The identifying variation is now the within-month difference across countries. Country-specific GPR therefore avoids this particular mechanical collinearity, provided that some cross-country variation remains after all other controls are also removed.
How the horizon determines the sample size
In a balanced panel with \(I\) countries, \(T\) monthly observations per country, twelve required GPR lags, no other missing values, and an outcome dated \(t+h\), a simple count is:
The first twelve origin dates are unavailable because the lag matrix is incomplete, while the final \(h\) origin dates are unavailable because \(y_{i,t+h}\) lies beyond the sample. In an unbalanced panel, missing outcomes, missing controls, singleton removal, and gaps in the time series make \(n_h\) horizon-specific in a less mechanical way.
What fixed-effect absorption does
Let \(\mathbf F_h\) denote a full-rank basis for the fixed effects, and place the lags and observed controls in \(\mathbf W_h=[\mathbf L_h\;\mathbf C_h]\). Define:
The associated fixed-effect projection matrix is \(\mathbf P_{F,h}=\mathbf F_h(\mathbf F_h^{\prime}\mathbf F_h)^{-1}\mathbf F_h^{\prime}\), so \(\mathbf M_{F,h}=\mathbf I_{n_h}-\mathbf P_{F,h}\). The projection extracts the component explained by the fixed effects; the residual-maker subtracts that component.
Why \(\mathbf M_{F,h}\mathbf F_h=\mathbf0\)
Every column of \(\mathbf F_h\) already lies in the space generated by the fixed effects. Projecting the fixed-effect dummies on themselves therefore reproduces them exactly. Algebraically:
The same result applies to every linear combination of the fixed effects:
This is why multiplication by \(\mathbf M_{F,h}\) is described as absorbing the fixed effects: their entire fitted contribution becomes zero. With one-way country effects, \(\mathbf M_{F,h}\mathbf v_h\) subtracts the appropriate country mean from every observation of any variable \(\mathbf v_h\). With multiple high-dimensional effects, the geometric operation is the same even when it is computed iteratively rather than by explicitly constructing the matrix.
Deriving the combined residual-maker step by step
Absorbing the fixed effects alone gives \(\mathbf M_{F,h}\). The longer formula below arises because the remaining regressors \(\mathbf W_h\) must subsequently be partialled out as well. To see this, take any horizon-specific variable \(\mathbf v_h\)—for example, the outcome or current GPR—and consider its auxiliary regression on both sets of nuisance regressors:
Apply \(\mathbf M_{F,h}\) to every term. Since \(\mathbf M_{F,h}\mathbf F_h=\mathbf0\), the fixed effects disappear:
At the OLS solution, \(\mathbf r_h\) is orthogonal to every column of \(\mathbf F_h\). It is therefore already in the residual space of the fixed effects, which implies \(\mathbf M_{F,h}\mathbf r_h=\mathbf r_h\).
Define the within-transformed variable and controls:
The coefficient on \(\mathbf W_h\) conditional on the fixed effects is the OLS coefficient from regressing \(\widetilde{\mathbf v}_h^{F}\) on \(\widetilde{\mathbf W}_h^{F}\):
The second equality uses symmetry and idempotence: \(\mathbf M_{F,h}^{\prime}=\mathbf M_{F,h}\) and \(\mathbf M_{F,h}^2=\mathbf M_{F,h}\). For example:
After removing the fixed effects, subtract the part of \(\widetilde{\mathbf v}_h^{F}\) explained by the within-transformed controls:
Because the expression in brackets maps any variable \(\mathbf v_h\) into its residual after projecting it on both \(\mathbf F_h\) and \(\mathbf W_h\), it is the combined residual-maker:
The dimensions confirm that the second term is an \(n_h\times n_h\) matrix. If \(\mathbf W_h\) contains \(q\) columns, then:
We can verify directly that the combined matrix removes both blocks. First, it removes the fixed effects because \(\mathbf M_{F,h}\mathbf F_h=\mathbf0\). Second:
Thus, absorbing fixed effects and then applying FWL to the within-transformed remaining regressors is algebraically equivalent to including a full-rank set of fixed-effect dummies in one large regression.
Why not simply use \(\mathbf M_W\mathbf M_F\)? After fixed-effect absorption, the relevant controls are \(\mathbf M_F\mathbf W\), not the raw controls \(\mathbf W\). In general, projecting the fixed-effect-adjusted variable on unadjusted controls does not produce a residual orthogonal to both blocks. The correct sequential operation projects on the within-transformed controls, which is exactly what the combined formula implements.
Software such as reghdfe computes this operation efficiently without forming the potentially enormous \(n_h\times n_h\) matrix explicitly. The inverse in the displayed formula also requires \(\mathbf W_h^{\prime}\mathbf M_{F,h}\mathbf W_h\) to be nonsingular: after fixed-effect absorption, the retained columns of \(\mathbf W_h\) must still contain linearly independent variation.
A.2. The OLS objective and its first-order conditions
To reduce notation, suppress the horizon subscript. Write:
where \(\mathbf y\) and \(\mathbf g\) are \(n\times1\), \(\mathbf Z\) is \(n\times k\), \(\beta\) is a scalar, and \(\boldsymbol\gamma\) is \(k\times1\). Define the residual vector:
OLS chooses \(\beta\) and \(\boldsymbol\gamma\) to minimize the residual sum of squares:
Expanding this scalar objective function gives:
Differentiating with respect to the scalar \(\beta\) gives:
Differentiating with respect to the \(k\times1\) vector \(\boldsymbol\gamma\) gives a \(k\times1\) gradient:
Any OLS minimizer must make both derivatives equal zero. After dividing by two and rearranging, the normal equations are:
Stacking the scalar equation for \(\widehat\beta\) and the \(k\) equations for \(\widehat{\boldsymbol\gamma}\) produces:
Geometric interpretation. The normal equations are equivalent to \(\mathbf g^{\prime}\widehat{\mathbf e}=0\) and \(\mathbf Z^{\prime}\widehat{\mathbf e}=\mathbf0\). The OLS residual is orthogonal, in sample, to current GPR and to every nuisance regressor.
Why the first-order conditions give a minimum
Let \(\boldsymbol\theta=[\beta\;\boldsymbol\gamma^{\prime}]^{\prime}\) and \(\mathbf X=[\mathbf g\;\mathbf Z]\). The objective can be written compactly as:
If \(\mathbf X\) is \(n\times p\), where \(p=k+1\), then \(\boldsymbol\theta\) is \(p\times1\). Expanding the objective gives:
where the second equality uses the fact that \(\mathbf y^{\prime}\mathbf X\boldsymbol\theta\) is a scalar and therefore equals its transpose, \(\boldsymbol\theta^{\prime}\mathbf X^{\prime}\mathbf y\). Its gradient and Hessian are:
The gradient is a \(p\times1\) vector of first derivatives. Setting it equal to zero gives the normal equations. A zero gradient identifies a stationary point, but in a general optimization problem it does not by itself tell us whether that point is a minimum, maximum, or saddle point. The Hessian—a \(p\times p\) matrix of second derivatives—provides the required curvature information.
Why full column rank makes the Hessian positive definite
A symmetric matrix \(\mathbf H\) is positive definite when \(\mathbf a^{\prime}\mathbf H\mathbf a>0\) for every nonzero conformable vector \(\mathbf a\). For the OLS Hessian, take any nonzero \(p\times1\) vector \(\mathbf a\). Then:
A squared norm is always nonnegative. If \(\mathbf X\) has full column rank, \(\mathbf X\mathbf a=\mathbf0\) is possible only when \(\mathbf a=\mathbf0\). Hence, for every nonzero \(\mathbf a\):
The Hessian is therefore positive definite. The sum-of-squares surface curves upward in every possible coefficient direction. In one dimension this is the familiar upward-opening parabola; positive definiteness is its multivariate counterpart.
A direct proof of the unique global minimum
The result can also be shown without relying only on the vocabulary of convexity. Let \(\widehat{\boldsymbol\theta}\) solve the normal equations, let \(\widehat{\mathbf e}=\mathbf y-\mathbf X\widehat{\boldsymbol\theta}\), and consider any alternative coefficient vector:
The alternative residual is:
Its sum of squares is:
The normal equations imply \(\mathbf X^{\prime}\widehat{\mathbf e}=\mathbf0\), so the cross term vanishes:
If \(\mathbf X\) has full column rank and \(\mathbf d\neq\mathbf0\), then \(\lVert\mathbf X\mathbf d\rVert^2>0\). Every coefficient vector different from \(\widehat{\boldsymbol\theta}\) therefore produces a strictly larger residual sum of squares. This proves that the normal-equation solution is the unique global minimum.
If \(\mathbf X\) is rank deficient, a nonzero \(\mathbf d\) may satisfy \(\mathbf X\mathbf d=\mathbf0\). Then \(\widehat{\boldsymbol\theta}\) and \(\widehat{\boldsymbol\theta}+\mathbf d\) generate the same fitted values and the same residual sum of squares. The objective remains convex, but the coefficients are not uniquely determined: geometrically, the minimum is a flat valley rather than a single point.
A.3. Eliminate the nuisance coefficients
The parameter of interest is \(\beta\). The vector \(\boldsymbol\gamma\) is called a nuisance parameter because it is required for conditioning but is not the object whose coefficient we want to interpret. Use the second normal equation to solve for it:
Substitute this expression into the first normal equation:
Distribute the matrix product:
Collect the two terms containing \(\widehat\beta\) on the left and move the remaining term to the right:
Each side is a scalar. For example:
Introduce the shorthand \(\mathbf P_Z=\mathbf Z(\mathbf Z^{\prime}\mathbf Z)^{-1}\mathbf Z^{\prime}\). Then the preceding equation is exactly:
Where does this equation come from? It is not an additional assumption. It is the first normal equation after the second normal equation has been solved for \(\widehat{\boldsymbol\gamma}\) and that expression has been substituted back. This substitution eliminates the nuisance coefficients while preserving the equation that determines \(\widehat\beta\).
A.4. Projection and residual-maker matrices
Define the projection matrix onto the column space of \(\mathbf Z\):
For any \(n\times1\) vector \(\mathbf v\), \(\mathbf P_Z\mathbf v\) contains the fitted values from the OLS projection of \(\mathbf v\) on \(\mathbf Z\). Define the associated residual-maker matrix:
Then \(\mathbf M_Z\mathbf v=\mathbf v-\mathbf P_Z\mathbf v\) is the residual vector from that projection. The two matrices satisfy:
Why these matrix properties hold
The matrix \(\mathbf Z^{\prime}\mathbf Z\) is symmetric, and so is its inverse. Therefore:
Moreover:
Idempotence means that projecting twice has exactly the same effect as projecting once. The corresponding results for \(\mathbf M_Z=\mathbf I_n-\mathbf P_Z\) follow immediately:
Finally:
The orthogonal decomposition
Every vector \(\mathbf v\in\mathbb R^{n\times1}\) can be decomposed uniquely as:
The two components are orthogonal because:
FWL uses only the second component of current GPR: the part lying outside the space spanned by its lags, controls, and fixed effects.
How the identity matrix reveals the residual-maker
Return to the boxed equation obtained after eliminating \(\widehat{\boldsymbol\gamma}\):
The identity matrix leaves any \(n\times1\) vector unchanged:
Consequently:
Inserting \(\mathbf I_n\) changes no numerical value. It is only an algebraic device that allows the two terms on each side to be factored. After insertion, the equation is:
Use distributivity to factor out \(\mathbf g^{\prime}\) on the left of each expression and \(\mathbf g\) or \(\mathbf y\) on the right:
Since \(\mathbf M_Z=\mathbf I_n-\mathbf P_Z\), the term in parentheses is precisely the residual-maker. We therefore obtain:
Provided that \(\mathbf g^{\prime}\mathbf M_Z\mathbf g>0\), so that current GPR retains variation after conditioning on \(\mathbf Z\), the coefficient is:
A.5. The FWL residual-on-residual regression
Define the residualized regressor and residualized outcome:
Both are \(n\times1\). Because \(\mathbf M_Z\) is symmetric and idempotent:
Therefore:
This is the OLS slope from regressing the part of the outcome unexplained by \(\mathbf Z\) on the part of current GPR unexplained by the same \(\mathbf Z\). The residual-on-residual regression is run without an intercept: when an intercept is included in \(\mathbf Z\), both residualized variables already have sample mean zero. It is exactly the coefficient on current GPR in the complete regression.
FWL in words. First remove the twelve GPR lags, controls, and fixed effects from current GPR. Remove the same variables from the outcome. The slope relating the two remaining components is exactly the coefficient on current GPR in the complete OLS regression.
Fitted values and residuals
Since \(\widetilde{\mathbf g}=\mathbf M_Z\mathbf g\) is orthogonal to every column of \(\mathbf Z\), the complete regressor space can be written as the orthogonal sum of \(\mathcal C(\mathbf Z)\) and the one-dimensional space generated by \(\widetilde{\mathbf g}\). The complete fitted values are therefore:
The residual from the residual-on-residual regression is:
This is exactly the residual vector from the complete regression. The residualized regression does not display the fitted contribution of \(\mathbf Z\), but it reproduces both the coefficient of interest and the complete model’s OLS residual vector.
A.6. A fully worked numerical example
Consider four observations, one regressor of interest \(\mathbf g\), and a nuisance matrix \(\mathbf Z\) containing an intercept and a binary group indicator:
The first two observations belong to group zero and the final two to group one. Before applying FWL to \(\mathbf g\) and \(\mathbf y\), it is useful to construct the relevant projection matrix one multiplication at a time.
Step 1: represent the two groups
The matrix \(\mathbf Z\) uses an intercept and a group-one dummy. An equivalent coding uses one indicator column for each group:
Each row represents an observation. The first column indicates membership in group zero and the second indicates membership in group one. Its transpose is:
Step 2: calculate \(\mathbf F^{\prime}\mathbf F\)
Write the two columns as \(\mathbf F=[\mathbf f_0\;\mathbf f_1]\), where:
Then:
Calculate the four dot products:
The diagonal entries count the observations in each group. The off-diagonal entries are zero because no observation belongs to both groups. Hence:
The dimensions are \((2\times4)(4\times2)=2\times2\). More generally, with \(n_0\) and \(n_1\) observations in the two groups, \(\mathbf F^{\prime}\mathbf F=\operatorname{diag}(n_0,n_1)\).
Step 3: invert the group-count matrix
The inverse of a nonsingular diagonal matrix is obtained by taking the reciprocal of each diagonal entry:
because:
The factors \(1/2\) turn each group total into a group average.
Step 4: calculate \(\mathbf F(\mathbf F^{\prime}\mathbf F)^{-1}\)
The dimensions are \((4\times2)(2\times2)=4\times2\). Each group indicator has effectively been divided by its group size.
Step 5: multiply by \(\mathbf F^{\prime}\)
For example, the first row of the product is:
The two nonzero entries average observations 1 and 2. The zeros prevent observations from the other group from entering that average. In general:
Step 6: see the group means directly
For an arbitrary vector \(\mathbf v=[v_1\;v_2\;v_3\;v_4]^{\prime}\):
Projection therefore replaces each observation by its group mean. If \(\mathbf v=[2\;6\;10\;14]^{\prime}\), for example, then \(\mathbf P_F\mathbf v=[4\;4\;12\;12]^{\prime}\).
Step 7: why \(\mathbf P_F=\mathbf P_Z\)
The matrices \(\mathbf F\) and \(\mathbf Z\) use different dummy coding, but they span the same space. In fact:
The matrix \(\mathbf A\) is invertible, so every column of \(\mathbf Z\) is a linear combination of the columns of \(\mathbf F\), and vice versa:
An orthogonal projection depends on the space being projected onto, not on the particular basis used to describe that space. Therefore:
Using two group dummies or using an intercept plus one group dummy therefore produces exactly the same fitted group means.
Step 8: construct the residual-maker
Subtract the projection matrix from the \(4\times4\) identity matrix:
The diagonal entries are \(1-1/2=1/2\), the within-group off-diagonal entries are \(0-1/2=-1/2\), and the between-group entries remain zero. Thus, for any \(\mathbf v\):
Applying \(\mathbf M_Z\) subtracts the relevant group mean from every observation:
The FWL numerator is:
The denominator is:
Consequently:
Running the complete regression of \(\mathbf y\) on \(\mathbf g\), an intercept, and the group indicator gives exactly the same coefficient, \(\widehat\beta=3\). The example makes the partialling-out operation concrete: FWL compares deviations from the group-specific means rather than the raw levels of \(\mathbf g\) and \(\mathbf y\).
A second numerical example: fixed effects followed by a control
The first example contains only group fixed effects. We now add one observed control in order to illustrate numerically the sequential residual-maker derived in Section A.1. Consider two countries, A and B, with two observations each:
| Observation | Country | Current GPR \(g\) | Control \(w\) | Outcome \(y\) |
|---|---|---|---|---|
| 1 | A | 1 | 1 | 15 |
| 2 | A | 3 | 3 | 25 |
| 3 | B | 2 | 2 | 25 |
| 4 | B | 4 | 6 | 39 |
These observations satisfy:
The full regression must therefore recover a coefficient of 3 on current GPR. We will obtain it by successively removing the country fixed effects and the control.
Use an intercept and a country-B indicator as a full-rank fixed-effect basis:
The fixed-effect residual-maker subtracts the country-specific mean:
Apply \(\mathbf M_F\) to current GPR, the control, and the outcome:
For instance, country A has GPR mean \((1+3)/2=2\), so its transformed GPR values are \(1-2=-1\) and \(3-2=1\). Country B has control mean \((2+6)/2=4\), so its transformed control values are \(2-4=-2\) and \(6-4=2\).
Because there is only one control column:
The combined residual-maker is consequently:
This single matrix removes both the country fixed effects and the within-country component explained by \(w\). Applying it gives:
The FWL numerator and denominator are:
Therefore:
The direct regression gives an intercept of 10, a country-B effect of 5, a GPR coefficient of 3, and a control coefficient of 2. The sequential calculation gives exactly the same GPR coefficient. Numerically, the procedure is:
- remove the country means from \(g\), \(w\), and \(y\);
- remove from the transformed \(g\) and \(y\) the components explained by transformed \(w\);
- regress the remaining outcome variation on the remaining GPR variation.
A.7. Construct the AR(12) residual
Now distinguish the complete FWL residual from a narrower AR residual. Let:
contain the variables used in the AR filter. If the data have already been demeaned and the filter contains only twelve lags, then \(r=12\). With an intercept, \(r=13\). If the filter also contains a full-rank basis for fixed effects, \(r\) increases accordingly.
For a lag-only AR(12), the rows of \(\mathbf A_h\) have the form:
The auxiliary AR projection is:
The OLS coefficient vector and the unstandardized residual are:
Thus, current GPR is decomposed into a component predicted by the AR regressors and an innovation relative to that auxiliary model:
A.8. Why retaining the AR regressors preserves the coefficient
Let \(\mathbf W_h\in\mathbb R^{n_h\times s}\) contain the remaining nuisance regressors that were not used in the AR filter. The complete conditioning matrix is:
Every column of \(\mathbf A_h\) is therefore contained in the column space of \(\mathbf Z_h\). Projecting \(\mathbf A_h\) on \(\mathbf Z_h\) reproduces \(\mathbf A_h\) exactly:
Consequently, the residual-maker annihilates the AR regressors:
Now partial out the complete nuisance set from the AR residual:
After conditioning on \(\mathbf Z_h\), current GPR and its AR residual therefore contain exactly the same remaining variation. Their difference, \(\mathbf A_h\widehat{\boldsymbol\rho}\), is constructed entirely from regressors that the outcome equation already holds constant.
The same equivalence can be seen immediately by substitution. Start from:
Because \(\mathbf g_h=\widehat{\mathbf u}_h+\mathbf A_h\widehat{\boldsymbol\rho}\), substitution gives:
The coefficient on the AR residual remains \(\beta_h\). Only the coefficients attached to the retained AR regressors are reparameterized.
The same result as an invertible matrix transformation
Define the two complete design matrices:
Because \(\mathbf g_h=\widehat{\mathbf u}_h+\mathbf A_h\widehat{\boldsymbol\rho}\), the direct design is obtained from the filtered design through:
The matrix \(\mathbf T_h\) is square, block triangular, and has ones on its diagonal. Hence \(\det(\mathbf T_h)=1\), so it is invertible. Therefore:
where \(\mathcal C(\cdot)\) denotes a matrix’s column space. OLS depends on the regressor space, not on the particular basis used to describe that space. The transformation changes the numerical coefficients on the AR regressors, but it does not change the fitted vector, the residual vector, or the coefficient multiplying the first column.
Exact coefficient equivalence.
The two complete regressions span the same regressor space. They consequently produce the same fitted values, OLS residuals, residual sum of squares, and—when the same covariance estimator and finite-sample convention are used—the same conventional, heteroskedasticity-robust, or clustered standard error for the corresponding coefficient.
A.9. The complete FWL residual is not necessarily the AR residual
This distinction is important. The complete FWL residual is:
where \(\mathbf Z_h\) includes all lags, controls, and fixed effects. The narrower AR residual is:
where \(\mathbf A_h\) contains only the variables used in the AR filter. Unless the two projections use the same regressor space, the residual vectors need not be numerically identical:
There is no contradiction. Once the remaining controls are also partialled out in the outcome equation:
The coefficient equivalence therefore requires a statement about the complete regression, not a claim that a simple AR residual is always numerically identical to the full FWL residual.
A useful refinement. When the AR regressors remain in the outcome equation, the coefficient equality is an exact change of regressor coordinates. Indeed, subtracting any fixed linear combination of retained regressors from current GPR would leave its conditional coefficient unchanged. Estimating the combination by AR OLS is what gives the resulting variable its forecast-error interpretation; it is not what creates the coefficient equivalence.
A.10. What changes when the AR regressors are dropped?
Suppose the researcher constructs \(\widehat{\mathbf u}_h=\mathbf M_{A,h}\mathbf g_h\) and then estimates an outcome equation containing only \(\widehat{\mathbf u}_h\), while dropping \(\mathbf A_h\) and the remaining controls. That regression is generally not equivalent to the complete specification.
There are two safe implementations:
- Retain the complete conditioning set. Regress the outcome on the unstandardized AR residual, the AR regressors, the remaining controls, and the same fixed effects. This is an exact reparameterization of the direct regression.
- Apply pure FWL residualization. On the exact estimation sample, residualize both current GPR and the outcome against the complete nuisance matrix \(\mathbf Z_h\), then regress \(\mathbf M_{Z,h}\mathbf y_h\) on \(\mathbf M_{Z,h}\mathbf g_h\) without an intercept. Degrees-of-freedom and robust or clustered covariance adjustments must still reflect the complete model.
For applied work, estimating the complete outcome equation directly is usually the clearest approach because the conditioning information remains visible and the requested inference can be computed in a single regression.
The parameterization must match. A pooled AR(12) with common lag coefficients corresponds to twelve common lag columns. Separate country-specific AR coefficients correspond to country-by-lag interaction columns. The latter filter is coefficient-equivalent only if those interactions are also retained in the outcome equation. The same caution applies to state-dependent or smooth-transition specifications: replacing \(g_{i,t}S_{i,t}\) with \(\widehat u_{i,t}S_{i,t}\) generally requires retaining the corresponding interactions between the AR regressors and the state variable.
A.11. Why the local-projection horizon matters
Every LP horizon has its own usable sample:
Consequently, \(n_h\), \(\mathbf Z_h\), and the residual-maker \(\mathbf M_{Z,h}\) can change with \(h\). The horizon-specific FWL formula is:
If the AR regressors remain in every horizon-specific outcome equation, a previously constructed residual remains an exact linear reparameterization on each subsample because \(\mathbf g_h=\widehat{\mathbf u}_h+\mathbf A_h\widehat{\boldsymbol\rho}\) continues to hold observation by observation. If the AR regressors are dropped and the analysis relies only on residualization, the projection must instead be aligned with each horizon-specific sample.
Nothing in the theorem changes when the LP is lag-augmented or the response is estimated in levels. Additional lags of the outcome, the shock, or other predetermined controls simply add columns to \(\mathbf Z_h\). FWL then defines the identifying GPR variation relative to that enlarged conditioning set.
A.12. Coefficient equivalence and statistical inference
FWL first establishes equality of the coefficient estimates. Inference requires an additional accounting point. Let:
Under conditionally homoskedastic errors, the conditional variance of the GPR coefficient is:
The complete regression estimates \(\sigma^2\) using all the estimated parameters:
A software regression containing only \(\widetilde{\mathbf y}\) and \(\widetilde{\mathbf g}\) may instead use \(n-1\) in the denominator because it does not know that \(k\) nuisance coefficients were estimated during residualization. The coefficient remains correct, but the default finite-sample variance correction need not be.
Heteroskedasticity-robust inference
Using the residual vector from the complete model, the uncorrected HC0 variance for the scalar coefficient can be written as:
HC1 adds a degrees-of-freedom correction. HC2 and HC3 additionally depend on the leverage values from the complete design. A stripped-down residual regression can therefore report different corrected standard errors unless the full-model parameter count and leverage are supplied.
Clustered inference
Suppose the panel contains \(G\) clusters and let \(\widetilde{\mathbf g}_c\) and \(\widehat{\boldsymbol\varepsilon}_c\) denote the vectors belonging to cluster \(c\). The uncorrected cluster-robust variance is:
Cluster finite-sample corrections likewise depend on the number of estimated parameters and sometimes on cluster-level leverage. This is why the most reliable implementation is to estimate the complete outcome equation directly, or to retain the complete conditioning set when replacing current GPR with its AR residual.
No separate generated-regressor correction in the exact reparameterization. When the AR residual and every AR regressor are included together, the filtered design is an invertible transformation of the direct design. The two regressions have the same hat matrix, fitted values, residuals, and leverage. Under the same covariance estimator and finite-sample convention, inference for the corresponding GPR coefficient is therefore identical.
A.13. Standardization changes the coefficient’s scale
The coefficient equality above uses the unstandardized AR residual. Let \(s_u\) denote its sample standard deviation and define:
Since \(\widehat{\mathbf u}_h=s_u\mathbf z_h^{u}\), the coefficient on the standardized innovation is:
This is not a failure of FWL. The underlying variation is the same, but the coefficient now measures the response to a one-standard-deviation GPR innovation rather than to a one-unit movement in GPR.
A.14. What the algebra does—and does not—establish
The FWL theorem is an algebraic result about OLS coefficients. It establishes which variation in current GPR estimates \(\beta_h\). It does not establish the causal orthogonality condition:
An AR residual is orthogonal in sample to the regressors used in its auxiliary projection. It may nevertheless contain unexpected financial news, commodity-market disturbances, changes in media attention, measurement error, or other contemporaneous innovations that also affect the outcome.
The relevant causal requirement is conditional orthogonality—or an appropriate conditional zero-mean restriction—not full statistical independence. Neither condition follows merely from the OLS identity \(\mathbf A_h^{\prime}\widehat{\mathbf u}_h=\mathbf0\).
| Claim | Correct conclusion |
|---|---|
| Complete regression versus residual-on-residual FWL | The coefficient on current GPR is identical. |
| Current GPR versus its unstandardized AR residual, with AR regressors retained | The corresponding coefficient, fitted values, and OLS residuals are identical; the coefficients on the retained AR regressors are reparameterized. |
| Plain AR residual versus complete FWL residual | The two residual series need not be identical when the full outcome equation contains additional controls. |
| AR residual standardized by its standard deviation | The underlying variation is unchanged, but the coefficient and its standard error are multiplied by that standard deviation. |
| AR residual orthogonal to the included lags | This establishes in-sample forecast orthogonality, not exogeneity with respect to the outcome error. |
| Residual-only regression with controls discarded | Coefficient equivalence requires matched FWL residualization on the complete, horizon-specific estimation sample; default inference may still need correction. |
Finally, the derivation above is an OLS result. Weighted least squares has an analogous theorem using weighted projection matrices. IV and GMM require their own projection and moment-condition arguments; the ordinary OLS residual-maker \(\mathbf M_Z\) cannot simply be carried over without modification.
Final takeaway.
AR filtering answers a forecasting question. Structural identification answers a causal question. The former does not mechanically supply the latter.