Why can an estimator become less credible when it uses more instruments? Because additional moment conditions are not necessarily additional identifying information.
Generalized Method of Moments was a genuine econometric breakthrough. It allowed economists to estimate models from a limited set of theoretically meaningful restrictions without specifying a complete likelihood. The difficulty came later, especially in dynamic-panel applications, when mechanically generated internal instruments began to substitute for an identification argument. This note explains, step by step, why numerous weak and highly correlated moments can add little genuine information while increasing finite-sample bias, overfitting and weighting-matrix instability.
\[ \boxed{\text{GMM can combine identifying information; it cannot create it.}} \]1. What I actually dislike
The title is deliberately provocative. GMM itself has not disappeared, and I do not object to its mathematics. OLS, instrumental variables and two-stage least squares can all be represented inside the broad GMM framework. Moment-based estimation remains fundamental in econometrics.
My criticism concerns a particular applied practice: using a large set of internally generated instruments—especially in difference or system GMM—as if the estimator itself solved endogeneity.
I distrust GMM when multiplying moment conditions substitutes for an identification argument.
Endogeneity is not fundamentally an optimization problem. It is an information problem. If a regressor is correlated with the structural error, we need a credible source of exogenous variation. No weighting matrix can manufacture that variation.
The correct hierarchy is
\[ \boxed{ \text{research design} \longrightarrow \text{credible exogenous variation} \longrightarrow \text{moment conditions} \longrightarrow \text{estimation} }. \]Mechanical applications often reverse this order: choose system GMM, generate many lags, report a Hansen test, and then claim that endogeneity has been addressed. That is precisely the practice I find unconvincing.
2. Two rises and a credibility correction
The expression “rise and fall” should not be interpreted as the disappearance of GMM. The historically defensible story is that GMM rose in two waves and subsequently underwent a credibility correction.
\[ \boxed{ \text{innovation} \rightarrow \text{diffusion with simultaneous warnings} \rightarrow \text{dynamic-panel boom} \rightarrow \text{credibility correction} } \]Before 1982 — Method of moments, IV, minimum distance and overidentification were already established. Sargan (1958) is a central precursor. The basic idea of translating economic assumptions into orthogonality restrictions predates the name “GMM.”
1982 — Hansen unified estimation, efficient weighting and overidentification testing. Hansen and Singleton applied the framework to nonlinear rational-expectations models. This was the first rise: economists could estimate theoretically meaningful restrictions without specifying the complete data-generating process.
1981–1998 — Dynamic-panel IV developed through Anderson–Hsiao, Arellano–Bond, Arellano–Bover and Blundell–Bond. This was the second rise: internal lags appeared to offer a practical answer to fixed effects, dynamics and the absence of external instruments.
1986–1990s — Tauchen, Bekker, Hansen–Heaton–Yaron, Altonji–Segal and Ziliak documented finite-sample, weak-instrument, many-instrument and estimated-weighting problems. The warnings began during the expansion—not after it.
Late 1990s–2000s — Difference and system GMM diffused rapidly through firm, household and cross-country panels, helped by automated software. This was the applied boom: sophisticated-looking estimates and diagnostics could be produced even without external instruments.
2000–2010 — Stock–Wright, Bowsher, Windmeijer, Roodman, Newey–Windmeijer and Bun–Windmeijer made weak identification and instrument proliferation central applied concerns. This was a highly visible credibility correction, not a sudden end to GMM usage.
Since 2010 — More attention has been paid to restricted or collapsed instrument sets, lag-window sensitivity, corrected inference and alternative panel estimators. This is better understood as selective survival and methodological maturation.
The chronology is important. Tauchen (1986) was already documenting a bias–variance trade-off from expanding lag instruments. Yet in 2009, Roodman still described difference and system GMM as growing in popularity. There was no clean “enthusiasm first, criticism later” sequence. Adoption and criticism overlapped.
Wooldridge (2001) offered a particularly balanced assessment during the boom: sophisticated GMM can be indispensable in complicated models, yet it is often unlikely to improve convincingly on OLS or 2SLS in typical applied settings.
3. From OLS to a moment condition
Consider the simple linear model
\[ y_i=\beta_0x_i+u_i. \]For clarity, suppose the variables have been centered, so an intercept is unnecessary. OLS minimizes the sum of squared residuals. Its first-order condition is
\[ \frac{1}{n}\sum_{i=1}^{n}x_i(y_i-\beta x_i)=0. \]Solving gives
\[ \widehat\beta_{\mathrm{OLS}} = \frac{\sum_i x_i y_i}{\sum_i x_i^2}. \]Substitute \(y_i=\beta_0x_i+u_i\):
\[ \widehat\beta_{\mathrm{OLS}} = \beta_0+ \frac{\sum_i x_i u_i}{\sum_i x_i^2}. \]Consequently,
\[ \operatorname{plim}\widehat\beta_{\mathrm{OLS}} = \beta_0+ \frac{E[x_i u_i]}{E[x_i^2]}. \]If
\[ E[x_i u_i]\neq0, \]OLS is inconsistent. The estimator has not failed to perform its calculation. The population condition underlying that calculation is false.
First lesson: an estimator does not eliminate endogeneity. It works only when the identifying restrictions behind it are credible.
4. What an instrument must actually provide
Suppose we find an instrument \(z_i\). It must satisfy two distinct conditions:
\[ \underbrace{E[z_i u_i]=0}_{\text{validity}}, \qquad \underbrace{E[z_i x_i]\neq0}_{\text{relevance}}. \]The IV estimator solves
\[ \frac1n\sum_i z_i(y_i-\beta x_i)=0, \]which gives
\[ \widehat\beta_{\mathrm{IV}} = \frac{\sum_i z_i y_i}{\sum_i z_i x_i}. \]Substituting the structural equation yields the most useful IV formula:
\[ \boxed{ \widehat\beta_{\mathrm{IV}}-\beta_0 = \frac{\frac1n\sum_i z_i u_i} {\frac1n\sum_i z_i x_i} }. \]The numerator concerns validity. The denominator concerns relevance. A perfectly exogenous instrument can nevertheless be nearly useless if the denominator is close to zero.
A simple analogy helps. A valid instrument is an honest witness: it is not systematically related to the unobserved forces contained in \(u_i\). A relevant instrument is a witness who actually saw the event: it contains substantial information about \(x_i\). An honest witness who saw almost nothing cannot identify what happened.
In short, exogeneity makes an instrument honest; relevance makes it informative. We need both.
Validity says that the population moment \(E[z_i u_i]\) is zero. It does not force the sample moment \(\overline{zu}\) to equal zero exactly. When \(\overline{zx}\) is close to zero, an ordinary sampling discrepancy in the numerator is divided by almost no relevance and is therefore greatly magnified.
Suppose the sample moment in the numerator is only \(0.01\). Then:
When \(\overline{zx}=0.50\), the coefficient error is \(0.01/0.50=0.02\).
When \(\overline{zx}=0.05\), the coefficient error is \(0.01/0.05=0.20\).
When \(\overline{zx}=0.01\), the coefficient error is \(0.01/0.01=1.00\).
The same small sample cross-moment with the error produces a coefficient error fifty times larger when relevance falls from \(0.50\) to \(0.01\).
Second lesson:
\[ \boxed{\text{Validity does not imply relevance.}} \]A weak instrument generates a flat statistical problem: ordinary sampling noise is divided by very little identifying information.
5. What GMM really adds
Now suppose there are \(K\) instruments collected in the vector
\[ z_i= \begin{pmatrix} z_{1i}\\ z_{2i}\\ \vdots\\ z_{Ki} \end{pmatrix}. \]They generate \(K\) moment conditions:
\[ E\!\left[z_{ji}(y_i-\beta_0x_i)\right]=0, \qquad j=1,\ldots,K. \]Define
\[ g_i(\beta)=z_i(y_i-\beta x_i), \qquad \overline g_n(\beta)=\frac1n\sum_{i=1}^{n}g_i(\beta). \]GMM chooses
\[ \widehat\beta_{\mathrm{GMM}} = \arg\min_{\beta} \overline g_n(\beta)’W_n\overline g_n(\beta), \]where \(W_n\) is a positive-definite weighting matrix.
If the number of moments equals the number of parameters, the model is exactly identified. The sample moments can be set to zero, and the choice of \(W_n\) does not affect the estimate.
If there are more moments than parameters, the model is overidentified. Sampling variation generally makes it impossible to set every sample moment to zero. GMM must choose a weighted compromise.
GMM as an average of instrument-specific estimates
For each instrument, define the just-identified estimate
\[ \widehat\beta_j = \frac{\overline{z_jy}}{\overline{z_jx}}, \qquad a_j=\overline{z_jx}. \]If \(W_n\) is diagonal, GMM becomes
\[ \boxed{ \widehat\beta_{\mathrm{GMM}} = \frac{\sum_{j=1}^{K}w_j a_j^2\widehat\beta_j} {\sum_{j=1}^{K}w_j a_j^2} }. \]It is a weighted average of the estimates preferred by the different moments.
Suppose one moment favours \(\beta=1\), while another favours \(\beta=3\). Then
\[ Q(\beta)=w_1(\beta-1)^2+w_2(\beta-3)^2, \]and
\[ \widehat\beta=\frac{w_1+3w_2}{w_1+w_2}. \]Thus,
\[ (w_1,w_2)=(9,1) \quad\Rightarrow\quad \widehat\beta=1.2, \]whereas
\[ (w_1,w_2)=(1,9) \quad\Rightarrow\quad \widehat\beta=2.8. \]If both moments are valid and strong, their preferred estimates converge to the same population parameter. If they remain incompatible because one is invalid or both are weak, the weighting matrix decides which moment to believe.
Why this appeared so attractive
Under independent sampling—or when \(i\) indexes independent clusters—let
\[ S=E[g_i(\theta_0)g_i(\theta_0)’] \]be the covariance matrix of the moments. With serially dependent moment contributions, \(S\) instead denotes their long-run covariance. Also let
\[ G=E\!\left[\frac{\partial g_i(\theta_0)}{\partial\theta’}\right] \]measure how strongly the moments respond to the parameters. The general asymptotic variance is
\[ \operatorname{Avar}\!\left[\sqrt n(\widehat\theta-\theta_0)\right] = (G’WG)^{-1}G’WSWG(G’WG)^{-1}. \]The efficient choice is
\[ W=S^{-1}, \]which gives
\[ \operatorname{Avar}\!\left[\sqrt n(\widehat\theta-\theta_0)\right] = (G’S^{-1}G)^{-1}. \]Under valid moments, sufficient identification, fixed \(K\), conventional regularity conditions and a large sample, adding informative moments can improve efficiency. This was—and remains—a genuine virtue of GMM.
The qualifications matter:
\[ \text{valid moments} + \text{relevant moments} + \text{fixed }K + \text{large effective sample}. \]Applied instrument-heavy GMM can depart substantially from this ideal experiment.
6. The master formula for GMM sensitivity
A first-order expansion of the GMM conditions gives
\[ \boxed{ \widehat\theta-\theta_0 \approx -(G’WG)^{-1}G’W\overline g_n(\theta_0) }. \]This formula contains the entire intuition:
- \(\overline g_n(\theta_0)\): sampling noise—or systematic moment invalidity.
- \(W\): which moments GMM chooses to trust.
- \(G\): how strongly the moments respond to the parameters.
- \((G’WG)^{-1}\): how strongly noise is amplified when identification is weak.
For scalar linear GMM, define
\[ a_n=\frac{Z’x}{n}, \qquad e_n=\frac{Z’u}{n}. \]Then, for a fixed final-step weighting matrix,
\[ \boxed{ \widehat\beta-\beta_0 = \frac{a_n’W_ne_n}{a_n’W_na_n} }. \]The numerator \(a_n’W_ne_n\) measures the weighted noise that GMM combines. The denominator \(a_n’W_na_n\) measures weighted identifying strength. The matrix \(W_n\) changes how both are combined, but it cannot create relevance that is absent from the moments.
Equivalently, the same idea can be remembered as
\[ \boxed{ \text{estimation error} = \frac{\text{weighted moment noise}} {\text{identifying strength}} }. \]Weak moments make the criterion flat
The curvature of the scalar linear-GMM criterion is
\[ \frac{\partial^2 Q(\beta)}{\partial\beta^2} =2a_n’W_na_n. \]If the instruments barely predict the endogenous regressor, \(a_n’W_na_n\) is small. The objective has a flat valley. A small change in the sample moments can then move its minimum considerably, even when numerical optimization is flawless.
Slight invalidity changes the population target
Let
\[ a=E[z_i x_i], \qquad \delta=E[z_i u_i]. \]If some moments are invalid, \(\delta\neq0\), and the pseudo-true GMM value satisfies
\[ \boxed{ \beta^*(W)-\beta_0 = \frac{a’W\delta}{a’Wa} }. \]Different instrument sets and different weighting matrices can then imply genuinely different probability limits. Sensitivity is no longer merely a small-sample inconvenience; it reflects incompatible identifying assumptions.
Correlated moments are echoes, not independent information
Suppose \(K\) moments have equal relevance \(a\), equal variance, and pairwise correlation \(r\). In a simplified equal-weight setting,
\[ \widehat\beta-\beta_0 = \frac{1}{Ka}\sum_{j=1}^{K}e_j. \]If
\[ \operatorname{Var}(e_j)=\frac{\sigma_e^2}{n}, \qquad \operatorname{Corr}(e_j,e_\ell)=r, \]then
\[ \operatorname{Var}(\widehat\beta-\beta_0) = \frac{\sigma_e^2}{na^2} \frac{1+(K-1)r}{K}. \]If \(r=0\), variance falls as \(1/K\). As \(r\) approaches one, the variance reduction disappears. At \(r=1\), the moments are exact duplicates and their covariance matrix is singular: ten copies of the same moment are not ten independent sources of identification.
Ten highly correlated moment conditions are closer to ten microphones recording the same voice than to ten independent witnesses. The data contain ten recordings, but mostly one underlying signal. In the same way, numerous lags of one persistent variable can generate many columns in the instrument matrix while contributing very little independent information.
The key distinction is:
\[ \boxed{ \text{number of moments} \neq \text{amount of independent identifying information} }. \]7. How many weak moments aggravate the problem
Several individually weak moments can jointly be informative. The problem is not “more moments” by itself. The problem arises when the number of moments grows faster than genuine identifying information because the additional instruments are weak, redundant, highly correlated or slightly invalid.
Many instruments can confuse fit with identification
Two-stage least squares, a special case of linear GMM, can be written as
\[ \widehat\beta_{\mathrm{2SLS}} = (X’P_ZX)^{-1}X’P_Zy, \]where
\[ P_Z=Z(Z’Z)^{-1}Z’ \]projects variables onto the instrument space.
Write the population linear projection of the endogenous regressor on the instruments as
\[ X=Z\pi+v. \]Here \(\pi\) is the population projection coefficient, defined so that \(E[z_i v_i]=0\).
The fitted regressor is
\[ P_ZX = \underbrace{Z\pi}_{\text{population signal predicted by }Z} + \underbrace{P_Zv}_{\text{in-sample fit to }v}. \]The first component is the population variation predicted by \(Z\). It is useful for identification only if the instruments are valid and sufficiently relevant. The second component is sample-specific fit to the first-stage residual. When \(v\) contains the endogenous part of \(X\), a flexible instrument set can accidentally fit some of that endogenous variation.
With a small, economically justified instrument set, we hope the population signal dominates. As the number of instruments increases, the projection becomes more flexible and can fit more of \(v\) inside the estimation sample. The first stage may then fit the observed \(X\) extremely well without isolating much genuinely exogenous variation.
In the extreme full-rank thought experiment, if \(K=n\), then
\[ P_Z=I_n. \]Consequently,
\[ \widehat\beta_{\mathrm{2SLS}} = (X’X)^{-1}X’y = \widehat\beta_{\mathrm{OLS}}. \]The first stage now fits \(X\) perfectly, but that perfect fit is completely unhelpful: it reproduces every component of \(X\), including the endogenous component.
A perfect first-stage fit is not necessarily perfect identification. With enough instruments, the first stage may simply have learned the estimation sample too well.
\[ \boxed{ \text{More instruments can improve in-sample fit without adding exogenous information.} } \]The same mechanism can be seen algebraically. Let \(\sigma_{vu}\) denote the conditional covariance between the first-stage residual \(v_i\) and the structural error \(u_i\). If \(E[vu’\mid Z]=\sigma_{vu}I_n\) and \(Z\) has rank \(K\), then
\[ E[v’P_Zu\mid Z] = \sigma_{vu}\operatorname{tr}(P_Z) = K\sigma_{vu}. \]This is not a universal formula for GMM bias. It isolates the conventional many-instrument mechanism: expanding the instrument space increases its in-sample capacity to fit the part of the first-stage residual that is correlated with the structural error. This helps explain the familiar finite-sample tendency:
\[ \boxed{ \text{a highly saturated instrument space can pull IV toward endogenous OLS} }. \]The estimated weighting matrix is another noisy object
Efficient two-step GMM uses
\[ \widehat W=\widehat S^{-1}, \qquad \widehat S = \frac1N\sum_{i=1}^{N} g_i(\widehat\theta_1)g_i(\widehat\theta_1)’. \]In panel GMM, \(N\) is typically the number of independent individuals or clusters. Thus, the relevant comparison is often \(K\) relative to \(N\), not relative to \(NT\).
Because \(\widehat S\) is a sum of only \(N\) outer products,
\[ \operatorname{rank}(\widehat S)\leq\min(K,N). \]If \(K>N\), the matrix cannot have full rank; it can become poorly conditioned well before that boundary.
A symmetric \(K\times K\) matrix contains
\[ \frac{K(K+1)}{2} \]distinct covariance elements. Therefore,
\[ K=10\Rightarrow55, \qquad K=50\Rightarrow1{,}275. \]Technical aside: why inverse weights can become unstable
When moments are nearly duplicates, their covariance matrix is nearly singular. Inverting that matrix can amplify small sampling differences in the direction where the moments almost cancel.
For the simple covariance matrix
\[ S= \begin{pmatrix} 1 & r\\ r & 1 \end{pmatrix}, \qquad \lambda_{\min}(S)=1-r. \]When \(r=0.99\), the smallest eigenvalue is only \(0.01\). The matrix is therefore almost singular. The perturbation identity
\[ \boxed{ d(S^{-1})=-S^{-1}(dS)S^{-1} }. \]shows why estimation noise in \(\widehat S\) can be magnified through inversion. This does not mean that inverse weighting is intrinsically wrong under the ideal fixed-moment model. It means that finite-sample noise or slight misspecification can matter greatly when the moments are nearly redundant.
This is why changing from one-step to two-step GMM, changing the lag window, or collapsing the instrument matrix may affect the point estimate—not merely its reported standard error.
8. Why dynamic-panel GMM is especially fragile
Consider the dynamic panel
\[ y_{it}=\rho y_{i,t-1}+\alpha_i+\varepsilon_{it}, \]where \(\alpha_i\) is an individual effect. First differencing removes it:
\[ \Delta y_{it} = \rho\Delta y_{i,t-1} + \Delta\varepsilon_{it}. \]But the differenced lag is endogenous. Indeed,
\[ \Delta y_{i,t-1}=y_{i,t-1}-y_{i,t-2} \]contains \(\varepsilon_{i,t-1}\), while
\[ \Delta\varepsilon_{it} = \varepsilon_{it}-\varepsilon_{i,t-1}. \]Therefore,
\[ \operatorname{Cov}(\Delta y_{i,t-1},\Delta\varepsilon_{it})\neq0. \]Difference GMM
Under the usual assumptions that sufficiently old outcomes are uncorrelated with subsequent innovations and that the idiosyncratic errors have no serial correlation, those old levels satisfy
\[ E[y_{i,t-s}\Delta\varepsilon_{it}]=0, \qquad s\geq2. \]The available instruments accumulate:
At \(t=3\), the available level instrument is \(y_{i1}\).
At \(t=4\), the available level instruments are \(y_{i1},y_{i2}\).
At \(t=5\), the available level instruments are \(y_{i1},y_{i2},y_{i3}\), and the sequence continues in this way.
In a balanced panel with one endogenous lagged variable, all available level instruments and an uncollapsed difference-GMM matrix, the count is
\[ K_D = 1+2+\cdots+(T-2) = \frac{(T-1)(T-2)}{2}. \]Hence,
\[ T=10\Rightarrow K_D=36, \qquad T=20\Rightarrow K_D=171. \]The count grows approximately with \(T^2\). A modest increase in the time dimension can therefore generate a surprisingly large instrument set.
These counts concern this specific uncollapsed setup. Collapsing, restricting the lag range, adding other endogenous variables, including time indicators, or adding the system-GMM level equations changes the count.
Persistence predicts levels but not changes
In a stripped-down centered AR(1),
\[ y_{t-1}=\rho y_{t-2}+\varepsilon_{t-1}. \]Therefore,
\[ \Delta y_{t-1} = (\rho-1)y_{t-2}+\varepsilon_{t-1}. \]In this stripped-down single-instrument first stage, the population coefficient on the level instrument is
\[ \boxed{\pi=\rho-1}. \]This result resolves an apparent paradox. When \(\rho\) is close to one, \(y_{t-2}\) predicts the next level, \(y_{t-1}\), extremely well. But difference GMM needs it to predict the change, \(\Delta y_{t-1}\). Persistence preserves information about levels while leaving very little predictable movement in the change.
For a near-random walk,
\[ y_{t-1}\approx y_{t-2}+\varepsilon_{t-1} \qquad\Longrightarrow\qquad \Delta y_{t-1}\approx\varepsilon_{t-1}. \]The change is essentially new information. Under the maintained assumption that the innovation is orthogonal to past information, an older level cannot predict it. At the exact random-walk limit, \(\rho=1\), the approximation becomes an equality.
Thus,
\[ \rho=0.95\Rightarrow\pi=-0.05, \qquad \rho=0.99\Rightarrow\pi=-0.01. \]Under stationarity and \(0<\rho<1\), the corresponding correlation is
\[ \operatorname{Corr}(y_{t-2},\Delta y_{t-1}) = -\sqrt{\frac{1-\rho}{2}} \longrightarrow0. \]Technical precision: it is the first-stage slope and correlation that reveal weak relevance. The raw covariance need not converge to zero because \(\operatorname{Var}(y_t)\) changes with \(\rho\). Under stationary normalization,
\[ \operatorname{Cov}(y_{t-2},\Delta y_{t-1}) =-\frac{\sigma_\varepsilon^2}{1+\rho}. \]Dynamic-panel GMM can consequently produce the particularly uncomfortable configuration
\[ \boxed{ K\uparrow \qquad\text{while}\qquad \text{instrument relevance}\downarrow }. \]Dynamic-panel GMM may therefore create many versions of the same weak signal and mistake their number for strength.
System GMM: stronger instruments through stronger assumptions
System GMM adds moment conditions for the equation in levels, such as
\[ E\!\left[\Delta y_{i,t-1}(\alpha_i+\varepsilon_{it})\right]=0. \]These additional moments can improve relevance and finite-sample performance, but they require restrictions on the initial-condition or mean-stationarity process. They are not generated by algebra alone.
This is the central trade-off:
\[ \boxed{ \text{potentially stronger instruments} \quad\Longleftrightarrow\quad \text{stronger identifying assumptions} }. \]Moreover, system GMM does not guarantee strong identification. Bun and Windmeijer (2010) show that the level equation can itself suffer from weak instruments.
9. Why specification tests may not rescue the design
The Hansen statistic is
\[ J = N\overline g_N(\widehat\theta)’ \widehat S^{-1} \overline g_N(\widehat\theta). \]Under correct specification, adequate identification, fixed \(K\) and conventional asymptotics,
\[ J\overset{a}{\sim}\chi^2_{K-p}. \]A rejection tells us that the joint restrictions and model are incompatible with the data. A non-rejection does not prove that the instruments are valid.
With many instruments:
- the endogenous regressors can be overfitted;
- the high-dimensional covariance matrix is difficult to estimate;
- the fixed-\(K\) chi-square approximation can become unreliable;
- the test may have very little power against invalid moments.
Bowsher (2002) showed that too many dynamic-panel moments can make overidentification tests severely undersized and extremely low-powered. In Roodman’s (2009) simulations, an invalid full-instrument specification at \(T=20\) produced an average Hansen \(p\)-value of \(1.000\), whereas sharply reducing the instrument set made the violation readily detectable.
A Hansen \(p\)-value close to one is therefore a warning sign in an instrument-heavy specification, not proof of validity. Difference-in-Hansen tests inherit many of the same limitations.
The Arellano–Bond AR(2) test is also narrower than sometimes believed. Passing it supports a serial-correlation implication needed for particular lag instruments. It does not validate every exclusion restriction or establish that the instruments are strong.
What the Windmeijer correction actually fixes
Conventional two-step GMM standard errors neglect finite-sample variation created by estimating \(\widehat W\) from first-step residuals. Windmeijer (2005) provides a correction for this estimated variance.
\[ \boxed{ \text{Windmeijer correction} \Rightarrow \text{improved two-step standard errors} }. \]It does not:
- debias the coefficient estimate;
- strengthen weak instruments;
- make invalid moments valid;
- reduce instrument proliferation;
- restore the power of the Hansen test.
10. What I would require from a convincing GMM application
The conclusion is not that every GMM application is uninformative. A small set of economically defensible, sufficiently strong moments can be entirely convincing. But the burden of proof belongs to the research design, not to the estimator.
I would want to see:
- An economic justification for each family of moments. Why should the proposed instruments be orthogonal to the structural error?
- A relevance argument. Why should the instruments meaningfully predict the endogenous variables after the chosen transformation?
- A deliberately restricted instrument set. Every available lag should not be included merely because software permits it.
- Sensitivity across lag windows and collapsing choices. Large coefficient movements reveal that the moment set is doing substantial identifying work.
- One-step and corrected two-step results. A major movement in the point estimate can indicate sensitivity to the estimated weighting matrix.
- Difference versus system comparisons. The additional level moments should be defended, not treated as automatically valid.
- Weak-identification-aware analysis. Conventional standard errors and Wald tests can be unreliable when moments are weak.
- Comparison with credible alternatives. Depending on the application, these may include external IV, bias-corrected fixed effects, split-panel jackknife methods, or a more modest non-causal interpretation.
There is no universal instrument-count rule. Keeping \(K<N\) is not a theorem guaranteeing validity or strength. A small instrument set can still be weak or invalid. Instrument reduction is necessary in many applications, but it cannot replace an economic argument.
11. Conclusion: identification first, estimation last
That is why I dislike GMM—or, more accurately, why I distrust what applied economists sometimes ask it to accomplish.
Hansen’s GMM was a breakthrough because it provided a disciplined way to combine credible economic restrictions without specifying a complete likelihood. The difficulty arose when this logic was reversed: because software could generate hundreds of internal instruments, the existence of many moments began to be treated as evidence of strong identification.
But weighting cannot manufacture either exogeneity or relevance. When moments are weak, redundant, highly correlated, numerous or slightly invalid, GMM may multiply sampling noise faster than it accumulates genuine information. More instruments can improve in-sample fit without adding exogenous information, making weak identification appear stronger without actually making it stronger.
A single strong and economically credible instrument may be worth more than a forest of weak internal instruments.
The final message is therefore not that “adding moments is always bad.” Under the ideal fixed-\(K\) theory, valid and independently informative moments can improve efficiency. The correct conclusion is more precise:
\[ \boxed{ \text{More moment conditions mean more identification only when they add} } \] \[ \boxed{ \text{valid, relevant and genuinely independent information.} } \] \[ \boxed{ \text{The solution to endogeneity is credible exogenous variation—not GMM itself.} } \]References and further reading
Foundations and early applications
- Sargan, J. D. (1958). “The Estimation of Economic Relationships Using Instrumental Variables.” Econometrica, 26(3), 393–415. https://doi.org/10.2307/1907619.
- Hansen, L. P. (1982). “Large Sample Properties of Generalized Method of Moments Estimators.” Econometrica, 50(4), 1029–1054. Author’s page and paper.
- Hansen, L. P., and Singleton, K. J. (1982). “Generalized Instrumental Variables Estimation of Nonlinear Rational Expectations Models.” Econometrica, 50(5), 1269–1286. Author’s page and paper.
- Newey, W. K., and West, K. D. (1987). “A Simple, Positive Semi-definite, Heteroskedasticity and Autocorrelation Consistent Covariance Matrix.” Econometrica, 55(3), 703–708. NBER version.
- Wooldridge, J. M. (2001). “Applications of Generalized Method of Moments Estimation.” Journal of Economic Perspectives, 15(4), 87–100. https://doi.org/10.1257/jep.15.4.87.
Dynamic-panel GMM
- Anderson, T. W., and Hsiao, C. (1981). “Estimation of Dynamic Models with Error Components.” Journal of the American Statistical Association, 76(375), 598–606. https://doi.org/10.1080/01621459.1981.10477691.
- Holtz-Eakin, D., Newey, W., and Rosen, H. S. (1988). “Estimating Vector Autoregressions with Panel Data.” Econometrica, 56(6), 1371–1395. NBER version.
- Arellano, M., and Bond, S. (1991). “Some Tests of Specification for Panel Data: Monte Carlo Evidence and an Application to Employment Equations.” Review of Economic Studies, 58(2), 277–297. Journal page.
- Arellano, M., and Bover, O. (1995). “Another Look at the Instrumental Variable Estimation of Error-Components Models.” Journal of Econometrics, 68(1), 29–51. https://doi.org/10.1016/0304-4076(94)01642-D.
- Blundell, R., and Bond, S. (1998). “Initial Conditions and Moment Restrictions in Dynamic Panel Data Models.” Journal of Econometrics, 87(1), 115–143. IFS journal page.
Weak moments, finite samples and instrument proliferation
- Tauchen, G. E. (1986). “Statistical Properties of Generalized Method-of-Moments Estimators of Structural Parameters Obtained from Financial Market Data.” Journal of Business & Economic Statistics, 4(4), 397–416. Duke research page.
- Bekker, P. A. (1994). “Alternative Approximations to the Distributions of Instrumental Variable Estimators.” Econometrica, 62(3), 657–681. https://doi.org/10.2307/2951662.
- Hansen, L. P., Heaton, J., and Yaron, A. (1996). “Finite-Sample Properties of Some Alternative GMM Estimators.” Journal of Business & Economic Statistics, 14(3), 262–280. Author’s page and paper.
- Altonji, J. G., and Segal, L. M. (1996). “Small-Sample Bias in GMM Estimation of Covariance Structures.” Journal of Business & Economic Statistics, 14(3), 353–366. NBER version.
- Ziliak, J. P. (1997). “Efficient Estimation with Panel Data When Instruments Are Predetermined: An Empirical Comparison of Moment-Condition Estimators.” Journal of Business & Economic Statistics, 15(4), 419–431. https://doi.org/10.1080/07350015.1997.10524720.
- Stock, J. H., and Wright, J. H. (2000). “GMM with Weak Identification.” Econometrica, 68(5), 1055–1096. https://doi.org/10.1111/1468-0262.00151.
- Bowsher, C. G. (2002). “On Testing Overidentifying Restrictions in Dynamic Panel Data Models.” Economics Letters, 77(2), 211–220. https://doi.org/10.1016/S0165-1765(02)00130-1.
- Windmeijer, F. (2005). “A Finite Sample Correction for the Variance of Linear Efficient Two-Step GMM Estimators.” Journal of Econometrics, 126(1), 25–51. IFS journal page.
- Roodman, D. (2009). “A Note on the Theme of Too Many Instruments.” Oxford Bulletin of Economics and Statistics, 71(1), 135–158. https://doi.org/10.1111/j.1468-0084.2008.00542.x.
- Newey, W. K., and Windmeijer, F. (2009). “Generalized Method of Moments with Many Weak Moment Conditions.” Econometrica, 77(3), 687–719. https://doi.org/10.3982/ECTA6224.
- Bun, M. J. G., and Windmeijer, F. (2010). “The Weak Instrument Problem of the System GMM Estimator in Dynamic Panel Data Models.” The Econometrics Journal, 13(1), 95–126. https://doi.org/10.1111/j.1368-423X.2009.00299.x.
Technical appendix: additional detailed derivations
This appendix provides the mathematical steps behind additional results used in the main text. The objective is to show exactly where each formula comes from and what it means economically.
A.1. From OLS to a moment condition
Begin with the linear model
\[ y_i=\alpha_0+\beta_0x_i+u_i, \]where \(y_i\) is the outcome, \(x_i\) is the explanatory variable, \(\beta_0\) is the true slope, and \(u_i\) contains all the other determinants of \(y_i\).
To simplify the notation, subtract the sample means from \(x_i\) and \(y_i\). Equivalently, we can say that the intercept has already been removed. The model then becomes
\[ y_i=\beta_0x_i+u_i. \]OLS chooses a value of \(\beta\) that minimizes the average squared residual:
\[ SSR(\beta) = \frac{1}{n}\sum_{i=1}^{n}(y_i-\beta x_i)^2. \]The residual associated with observation \(i\) is
\[ e_i(\beta)=y_i-\beta x_i. \]To find the minimum, differentiate the objective function with respect to \(\beta\). For one observation, the chain rule gives
\[ \frac{\partial}{\partial\beta}(y_i-\beta x_i)^2 = 2(y_i-\beta x_i)(-x_i). \]Therefore,
\[ \frac{\partial SSR(\beta)}{\partial\beta} = -\frac{2}{n} \sum_{i=1}^{n} x_i(y_i-\beta x_i). \]At an interior minimum, this derivative must equal zero:
\[ -\frac{2}{n} \sum_{i=1}^{n} x_i(y_i-\widehat\beta_{\mathrm{OLS}}x_i) = 0. \]Dividing by \(-2\) gives
\[ \frac{1}{n} \sum_{i=1}^{n} x_i(y_i-\widehat\beta_{\mathrm{OLS}}x_i) = 0. \]This is the OLS sample moment condition. It says that the regressor is orthogonal to the estimated residual:
\[ \overline{x\widehat u}=0. \]Expanding the moment condition gives
\[ \frac{1}{n}\sum_{i=1}^{n}x_iy_i – \widehat\beta_{\mathrm{OLS}} \frac{1}{n}\sum_{i=1}^{n}x_i^2 = 0. \]Consequently,
\[ \widehat\beta_{\mathrm{OLS}} = \frac{\frac{1}{n}\sum_{i=1}^{n}x_iy_i} {\frac{1}{n}\sum_{i=1}^{n}x_i^2} = \frac{\sum_{i=1}^{n}x_iy_i} {\sum_{i=1}^{n}x_i^2}. \]Now substitute the true structural equation
\[ y_i=\beta_0x_i+u_i \]into the OLS formula:
\[ \widehat\beta_{\mathrm{OLS}} = \frac{\sum_i x_i(\beta_0x_i+u_i)} {\sum_i x_i^2}. \]Distributing \(x_i\) in the numerator gives
\[ \widehat\beta_{\mathrm{OLS}} = \frac{ \beta_0\sum_i x_i^2+\sum_i x_iu_i }{ \sum_i x_i^2 }. \]Separating the two terms produces
\[ \widehat\beta_{\mathrm{OLS}} = \beta_0 + \frac{\sum_i x_iu_i} {\sum_i x_i^2}. \]Therefore, the OLS estimation error is
\[ \boxed{ \widehat\beta_{\mathrm{OLS}}-\beta_0 = \frac{\frac{1}{n}\sum_i x_iu_i} {\frac{1}{n}\sum_i x_i^2} }. \]The numerator measures the sample relationship between the regressor and the structural error. The denominator measures the amount of variation in the regressor.
Under the law of large numbers, and provided the relevant expectations exist,
\[ \frac{1}{n}\sum_i x_iu_i \overset{p}{\longrightarrow} E[x_iu_i] \]and
\[ \frac{1}{n}\sum_i x_i^2 \overset{p}{\longrightarrow} E[x_i^2]. \]Hence,
\[ \operatorname{plim} \widehat\beta_{\mathrm{OLS}} = \beta_0 + \frac{E[x_iu_i]}{E[x_i^2]}. \]Thus, if
\[ E[x_iu_i]=0 \]and
\[ E[x_i^2]>0, \]then
\[ \operatorname{plim} \widehat\beta_{\mathrm{OLS}} = \beta_0. \]OLS can therefore be viewed as a method-of-moments estimator based on the population condition
\[ E[x_i(y_i-\beta_0x_i)]=0. \]The sample estimator is obtained by replacing the population expectation with its sample analogue and setting it equal to zero.
A.2. Derivation and interpretation of the IV error formula
Suppose that \(x_i\) is endogenous:
\[ E[x_iu_i]\neq 0. \]OLS is then inconsistent because the regressor is systematically related to the structural error.
Assume that we have an instrument \(z_i\). A valid instrument must satisfy the population moment condition
\[ E[z_iu_i]=0. \]Because
\[ u_i=y_i-\beta_0x_i, \]the IV moment condition can also be written as
\[ E[z_i(y_i-\beta_0x_i)]=0. \]The sample analogue is
\[ \frac{1}{n} \sum_{i=1}^{n} z_i(y_i-\widehat\beta_{\mathrm{IV}}x_i) = 0. \]Expanding this expression gives
\[ \frac{1}{n}\sum_i z_iy_i – \widehat\beta_{\mathrm{IV}} \frac{1}{n}\sum_i z_ix_i = 0. \]Solving for the IV estimator produces
\[ \widehat\beta_{\mathrm{IV}} = \frac{ \frac{1}{n}\sum_i z_iy_i }{ \frac{1}{n}\sum_i z_ix_i }. \]Now substitute the structural equation
\[ y_i=\beta_0x_i+u_i. \]We obtain
\[ \widehat\beta_{\mathrm{IV}} = \frac{ \frac{1}{n}\sum_i z_i(\beta_0x_i+u_i) }{ \frac{1}{n}\sum_i z_ix_i }. \]Expanding the numerator gives
\[ \widehat\beta_{\mathrm{IV}} = \frac{ \beta_0\frac{1}{n}\sum_i z_ix_i + \frac{1}{n}\sum_i z_iu_i }{ \frac{1}{n}\sum_i z_ix_i }. \]Separating the two terms yields
\[ \widehat\beta_{\mathrm{IV}} = \beta_0 + \frac{ \frac{1}{n}\sum_i z_iu_i }{ \frac{1}{n}\sum_i z_ix_i }. \]Consequently,
\[ \boxed{ \widehat\beta_{\mathrm{IV}}-\beta_0 = \frac{\overline{zu}} {\overline{zx}} = \frac{ \frac{1}{n}\sum_i z_iu_i }{ \frac{1}{n}\sum_i z_ix_i } }. \]The numerator and denominator have distinct economic meanings.
The numerator
\[ \overline{zu} = \frac{1}{n}\sum_i z_iu_i \]measures the sample relationship between the instrument and the structural error. Instrument validity requires
\[ E[z_iu_i]=0. \]This does not mean that \(\overline{zu}\) will equal zero exactly in every finite sample. Even a perfectly valid instrument normally has some accidental sample correlation with the error.
The denominator
\[ \overline{zx} = \frac{1}{n}\sum_i z_ix_i \]measures instrument relevance. It tells us how strongly the instrument is related to the endogenous regressor.
If the instrument is strong, ordinary sampling noise in the numerator is divided by a reasonably large number. If the instrument is weak, the same noise is divided by a very small number and can generate a very large estimation error.
For example, suppose that
\[ \overline{zu}=0.01. \]If
\[ \overline{zx}=0.50, \]then
\[ \widehat\beta_{\mathrm{IV}}-\beta_0 = \frac{0.01}{0.50} = 0.02. \]But if
\[ \overline{zx}=0.05, \]then
\[ \widehat\beta_{\mathrm{IV}}-\beta_0 = \frac{0.01}{0.05} = 0.20. \]Finally, if
\[ \overline{zx}=0.01, \]then
\[ \widehat\beta_{\mathrm{IV}}-\beta_0 = \frac{0.01}{0.01} = 1. \]The sample relationship between the instrument and the error has not changed. Only instrument relevance has changed. Yet the estimation error has increased from \(0.02\) to \(1\).
Exogeneity and relevance therefore perform different functions:
\[ \text{exogeneity controls the numerator,} \] \[ \text{relevance prevents the denominator from becoming too small.} \]An intuitive analogy is that a valid instrument is an honest witness, while a relevant instrument is a witness who actually observed the event. An honest witness who saw almost nothing cannot provide much information.
Thus:
\[ \boxed{ \text{Exogeneity makes an instrument honest; relevance makes it informative.} } \]A.3. Why GMM is a weighted average of instrument-specific estimates
Suppose that there is one parameter \(\beta\) and \(K\) instruments. Instrument \(j\) generates the sample moment
\[ \overline g_j(\beta) = \frac{1}{n} \sum_{i=1}^{n} z_{ji}(y_i-\beta x_i). \]Define
\[ b_j = \frac{1}{n}\sum_i z_{ji}y_i \]and
\[ a_j = \frac{1}{n}\sum_i z_{ji}x_i. \]The moment can then be written as
\[ \overline g_j(\beta) = b_j-a_j\beta. \]If instrument \(j\) were used alone, its just-identified IV estimate would solve
\[ b_j-a_j\widehat\beta_j=0. \]Therefore, provided \(a_j\neq0\),
\[ \widehat\beta_j = \frac{b_j}{a_j} = \frac{\frac{1}{n}\sum_i z_{ji}y_i} {\frac{1}{n}\sum_i z_{ji}x_i}. \]Because \(b_j=a_j\widehat\beta_j\), the corresponding moment discrepancy is
\[ \overline g_j(\beta) = a_j(\widehat\beta_j-\beta). \]This equation has an important interpretation. Each moment condition has its own preferred estimate, \(\widehat\beta_j\). The moment is zero when \(\beta=\widehat\beta_j\).
Suppose that the GMM weighting matrix is diagonal:
\[ W = \operatorname{diag}(w_1,\ldots,w_K). \]The GMM criterion is
\[ Q_n(\beta) = \sum_{j=1}^{K} w_j\overline g_j(\beta)^2. \]Substituting
\[ \overline g_j(\beta) = a_j(\widehat\beta_j-\beta) \]gives
\[ Q_n(\beta) = \sum_{j=1}^{K} w_ja_j^2 (\widehat\beta_j-\beta)^2. \]Differentiate with respect to \(\beta\):
\[ \frac{\partial Q_n(\beta)}{\partial\beta} = 2\sum_{j=1}^{K} w_ja_j^2 (\beta-\widehat\beta_j). \]At the minimum,
\[ \sum_{j=1}^{K} w_ja_j^2 (\widehat\beta_{\mathrm{GMM}}-\widehat\beta_j) = 0. \]Rearranging,
\[ \widehat\beta_{\mathrm{GMM}} \sum_{j=1}^{K}w_ja_j^2 = \sum_{j=1}^{K} w_ja_j^2\widehat\beta_j. \]Therefore,
\[ \boxed{ \widehat\beta_{\mathrm{GMM}} = \frac{ \sum_{j=1}^{K} w_ja_j^2\widehat\beta_j }{ \sum_{j=1}^{K} w_ja_j^2 } }. \]Define the effective weight attached to instrument \(j\) as
\[ \lambda_j = \frac{w_ja_j^2} {\sum_{\ell=1}^{K}w_\ell a_\ell^2}. \]These effective weights satisfy
\[ \lambda_j\geq0 \]and
\[ \sum_{j=1}^{K}\lambda_j=1. \]Thus,
\[ \widehat\beta_{\mathrm{GMM}} = \sum_{j=1}^{K} \lambda_j\widehat\beta_j. \]With a diagonal weighting matrix, GMM is therefore a weighted average of the estimates preferred by the individual instruments.
The effective importance of an instrument depends on two elements:
\[ \text{effective importance} = \text{GMM weight} \times \text{squared sample relevance}. \]This follows from the term \(w_ja_j^2\). A weak instrument has a small \(a_j\), so its direct contribution is reduced. But a collection of numerous weak and correlated instruments can still create serious finite-sample problems through overfitting and unstable estimation of the weighting matrix.
A.3.1. A two-moment illustration
Suppose that one moment condition is minimized at
\[ \beta=1, \]while another is minimized at
\[ \beta=3. \]For simplicity, absorb the relevance terms into the weights. The GMM criterion becomes
\[ Q(\beta) = w_1(\beta-1)^2 + w_2(\beta-3)^2. \]The first term penalizes distance from \(1\). The second penalizes distance from \(3\).
Differentiate the criterion:
\[ Q'(\beta) = 2w_1(\beta-1) + 2w_2(\beta-3). \]At the minimum,
\[ 2w_1(\widehat\beta-1) + 2w_2(\widehat\beta-3) = 0. \]Dividing by \(2\) and expanding gives
\[ w_1\widehat\beta-w_1 + w_2\widehat\beta-3w_2 = 0. \]Collecting the terms containing \(\widehat\beta\),
\[ (w_1+w_2)\widehat\beta = w_1+3w_2. \]Hence,
\[ \boxed{ \widehat\beta = \frac{w_1+3w_2}{w_1+w_2} }. \]If the first moment receives nine times as much weight as the second,
\[ (w_1,w_2)=(9,1), \]then
\[ \widehat\beta = \frac{9+3}{10} = 1.2. \]If the weights are reversed,
\[ (w_1,w_2)=(1,9), \]then
\[ \widehat\beta = \frac{1+27}{10} = 2.8. \]The result can also be understood as a balance of opposing forces. At the minimum,
\[ w_1(\widehat\beta-1) = w_2(3-\widehat\beta). \]With \((w_1,w_2)=(9,1)\),
\[ 9(1.2-1) = 1(3-1.2) = 1.8. \]The first moment is only \(0.2\) away from its preferred value, but its weight is nine. The second is \(1.8\) away from its preferred value, but its weight is only one. The weighted pressures are equal.
This example reveals what GMM does when the sample moments disagree:
\[ \boxed{ \text{GMM does not determine which moment is correct; it chooses a weighted compromise.} } \]If all moments are valid and strongly identifying, their preferred estimates should converge to the same true parameter as the sample grows. Their finite-sample disagreement is then ordinary sampling noise.
If some moments are invalid, however, their disagreement can persist even in very large samples. In that case, changing the weighting matrix changes the population value toward which GMM converges.
With a non-diagonal weighting matrix, the estimator need not be a simple convex average. Correlations among moments can generate complicated implicit weights, which may even be negative.
A.4. Deriving the GMM sandwich variance and efficient weighting matrix
Let \(\theta_0\) be a \(p\times1\) vector of true parameters and let \(g_i(\theta)\) be a \(K\times1\) vector of moment conditions satisfying
\[ E[g_i(\theta_0)]=0. \]The sample average of the moments is
\[ \overline g_n(\theta) = \frac{1}{n} \sum_{i=1}^{n} g_i(\theta). \]GMM chooses \(\widehat\theta\) to minimize
\[ Q_n(\theta) = \overline g_n(\theta)’ W \overline g_n(\theta), \]where \(W\) is a symmetric positive-definite \(K\times K\) weighting matrix.
A.4.1. The meanings of \(S\) and \(G\)
Under independent sampling, define
\[ S = \operatorname{Var}[g_i(\theta_0)]. \]Because the population moments are zero,
\[ E[g_i(\theta_0)]=0, \]we have
\[ S = E[g_i(\theta_0)g_i(\theta_0)’]. \]The diagonal elements of \(S\) measure the variance of each moment. The off-diagonal elements measure how the moments move together.
With serially dependent observations, \(S\) is instead the long-run covariance matrix:
\[ S = \sum_{h=-\infty}^{\infty} \operatorname{Cov} \left( g_t(\theta_0), g_{t-h}(\theta_0) \right). \]Now define the \(K\times p\) Jacobian matrix
\[ G = E\left[ \frac{\partial g_i(\theta_0)} {\partial\theta’} \right]. \]The matrix \(G\) measures how strongly the moments respond when the parameters change.
For one parameter and one moment, \(G\) is simply a slope. If \(G\) is close to zero, changing the parameter barely changes the moment condition. The moment then contains little identifying information.
With several parameters, identification requires \(G\) to have full column rank. If its columns are nearly linearly dependent, some combinations of the parameters are only weakly identified.
A.4.2. Linearizing the GMM estimator
Differentiate the GMM criterion. Treating the final-step weighting matrix as fixed asymptotically, the first-order condition is
\[ 2\widehat G_n(\widehat\theta)’ W \overline g_n(\widehat\theta) = 0, \]where
\[ \widehat G_n(\theta) = \frac{\partial\overline g_n(\theta)} {\partial\theta’}. \]For a large sample, \(\widehat G_n(\widehat\theta)\) is close to its population value \(G\). The first-order condition can therefore be approximated by
\[ G’W\overline g_n(\widehat\theta) \approx 0. \]Next, use a first-order Taylor expansion of the sample moments around the true parameter:
\[ \overline g_n(\widehat\theta) \approx \overline g_n(\theta_0) + G(\widehat\theta-\theta_0). \]Substituting this expansion into the first-order condition gives
\[ G’W \left[ \overline g_n(\theta_0) + G(\widehat\theta-\theta_0) \right] \approx 0. \]Expanding,
\[ G’W\overline g_n(\theta_0) + G’WG(\widehat\theta-\theta_0) \approx 0. \]Move the first term to the other side:
\[ G’WG(\widehat\theta-\theta_0) \approx – G’W\overline g_n(\theta_0). \]Provided \(G’WG\) is invertible,
\[ \widehat\theta-\theta_0 \approx – (G’WG)^{-1} G’W \overline g_n(\theta_0). \]Multiplying by \(\sqrt n\) gives the fundamental GMM approximation:
\[ \boxed{ \sqrt n(\widehat\theta-\theta_0) \approx – (G’WG)^{-1} G’W \sqrt n\,\overline g_n(\theta_0) }. \]This equation says that estimation error is transformed moment noise. The sample moments fluctuate around zero, and the matrix
\[ (G’WG)^{-1}G’W \]translates those fluctuations into parameter fluctuations.
If \(G’WG\) has a very small eigenvalue, its inverse contains a very large eigenvalue. Small moment disturbances can then generate large changes in the estimated parameters. This is the multivariate form of the weak-denominator problem.
A.4.3. Obtaining the sandwich variance
Under a central limit theorem,
\[ \sqrt n\,\overline g_n(\theta_0) \overset{d}{\longrightarrow} N(0,S). \]Recall the elementary covariance rule
\[ \operatorname{Var}(Ax) = A\operatorname{Var}(x)A’. \]Define
\[ A = – (G’WG)^{-1}G’W. \]Then
\[ \sqrt n(\widehat\theta-\theta_0) \approx A\sqrt n\,\overline g_n(\theta_0). \]Its asymptotic variance is therefore
\[ \operatorname{Avar} \left[ \sqrt n(\widehat\theta-\theta_0) \right] = ASA’. \]Because \(W\) is symmetric,
\[ A’ = – WG(G’WG)^{-1}. \]Substituting \(A\), \(S\), and \(A’\) gives
\[ \boxed{ \operatorname{Avar} \left[ \sqrt n(\widehat\theta-\theta_0) \right] = (G’WG)^{-1} G’WSWG (G’WG)^{-1} }. \]This is called a sandwich variance because the matrix \(G’WSWG\) is placed between two copies of \((G’WG)^{-1}\).
Since this expression is the variance of \(\sqrt n(\widehat\theta-\theta_0)\), the approximate variance of the estimator itself is
\[ \operatorname{Var}(\widehat\theta) \approx \frac{1}{n} (G’WG)^{-1} G’WSWG (G’WG)^{-1}. \]A.4.4. Why \(W=S^{-1}\) is efficient
The matrix \(S\) describes the noise and correlation structure of the moment conditions. The efficient GMM weighting matrix is
\[ W=S^{-1}. \]This choice gives less influence to noisy combinations of moments and accounts for correlations between them.
One way to understand this result is to define standardized moments
\[ h_i(\theta) = S^{-1/2}g_i(\theta). \]The covariance matrix of these transformed moments is
\[ \operatorname{Var}[h_i(\theta_0)] = S^{-1/2}SS^{-1/2} = I. \]Thus, the transformation removes differences in scale and correlation. The efficient GMM criterion can be written as
\[ \overline g_n(\theta)’S^{-1}\overline g_n(\theta) = \overline h_n(\theta)’\overline h_n(\theta). \]Efficient GMM therefore minimizes moment discrepancies after measuring them in standardized, correlation-adjusted units.
Substitute \(W=S^{-1}\) into the general sandwich formula:
\[ \operatorname{Avar} \left[ \sqrt n(\widehat\theta-\theta_0) \right] = (G’S^{-1}G)^{-1} G’S^{-1}SS^{-1}G (G’S^{-1}G)^{-1}. \]Because
\[ S^{-1}SS^{-1}=S^{-1}, \]the middle part becomes
\[ G’S^{-1}G. \]Consequently,
\[ \operatorname{Avar} \left[ \sqrt n(\widehat\theta-\theta_0) \right] = (G’S^{-1}G)^{-1} (G’S^{-1}G) (G’S^{-1}G)^{-1}. \]Using the identity
\[ A^{-1}AA^{-1}=A^{-1}, \]we obtain
\[ \boxed{ \operatorname{Avar} \left[ \sqrt n(\widehat\theta-\theta_0) \right] = (G’S^{-1}G)^{-1} }. \]The matrix
\[ G’S^{-1}G \]can be interpreted as the identifying information contained in the moments after accounting for their noise and correlation. The asymptotic variance is its inverse.
A.4.5. Connection with the weak-instrument formula
Consider one parameter and one IV moment:
\[ g_i(\beta) = z_i(y_i-x_i\beta). \]At the true parameter,
\[ g_i(\beta_0)=z_iu_i. \]Therefore,
\[ S = E[z_i^2u_i^2]. \]The derivative of the moment with respect to \(\beta\) is
\[ \frac{\partial g_i(\beta)}{\partial\beta} = -z_ix_i, \]so
\[ G=-E[z_ix_i]. \]The efficient asymptotic variance becomes
\[ \operatorname{Avar} \left[ \sqrt n(\widehat\beta-\beta_0) \right] = \frac{S}{G^2}. \]Substituting \(S\) and \(G\),
\[ \boxed{ \operatorname{Avar} \left[ \sqrt n(\widehat\beta-\beta_0) \right] = \frac{ E[z_i^2u_i^2] }{ \left(E[z_ix_i]\right)^2 } }. \]The numerator measures moment noise. The denominator measures instrument relevance, squared.
If
\[ E[z_ix_i]\approx0, \]the denominator is very small and the variance becomes very large. No weighting scheme can eliminate this basic lack of identifying information.
This is the central connection between weak IV and GMM:
\[ \boxed{ \text{GMM can combine identifying information efficiently, but it cannot create identifying information.} } \]The efficiency result for \(W=S^{-1}\) is therefore conditional on valid moment conditions, sufficiently strong identification, a fixed and manageable number of moments, a reliably estimated covariance matrix, and a sufficiently large sample. When there are numerous weak and highly correlated moments, these conditions can provide a poor description of the finite sample.
A.5. Why \(P_ZX\) is the first-stage fitted regressor
Consider a model with one endogenous regressor:
\[ y=X\beta_0+u, \]where \(X\) and \(y\) are \(n\times1\) vectors. Let \(Z\) be an \(n\times K\) matrix of instruments with full column rank \(K\).
The projection matrix onto the column space of \(Z\) is
\[ P_Z=Z(Z’Z)^{-1}Z’. \]The column space of \(Z\) contains every linear combination of the instruments across the \(n\) observations. Multiplying a vector by \(P_Z\) returns the part of that vector lying in this instrument space.
A.5.1. Obtaining the first-stage fitted values
The sample first-stage regression is
\[ X=Z\widehat\pi+\widehat v. \]Ordinary least squares gives
\[ \widehat\pi=(Z’Z)^{-1}Z’X. \]The corresponding fitted values are therefore
\[ \begin{aligned} \widehat X &=Z\widehat\pi\\ &=Z(Z’Z)^{-1}Z’X\\ &=P_ZX. \end{aligned} \]Hence,
\[ \boxed{\widehat X=P_ZX}. \]The vector \(P_ZX\) contains the value of the endogenous regressor predicted by the instruments for every observation. This is the regressor used in the familiar two-stage interpretation of 2SLS.
A.5.2. Decomposing fitted values into population signal and sample fit
Now write the population linear projection of \(X\) on \(Z\) as
\[ X=Z\pi+v, \]where \(\pi\) is the population projection coefficient and \(v\) is the population projection error. The coefficient \(\pi\) is defined by the orthogonality condition
\[ E[z_iv_i]=0. \]Apply \(P_Z\) to both sides of the population projection:
\[ P_ZX=P_Z(Z\pi+v). \]By linearity,
\[ P_ZX=P_ZZ\pi+P_Zv. \]Because
\[ \begin{aligned} P_ZZ &=Z(Z’Z)^{-1}Z’Z\\ &=Z, \end{aligned} \]we obtain
\[ \boxed{ P_ZX=Z\pi+P_Zv }. \]The two terms have different meanings:
\[ \underbrace{Z\pi}_{\text{population signal predicted by }Z} \qquad+\qquad \underbrace{P_Zv}_{\text{population residual fitted in this sample}}. \]The first term, \(Z\pi\), is genuine population-level predictive content. It represents the component of \(X\) that the instruments systematically explain.
The second term, \(P_Zv\), is different. It is the component of the population projection error that happens to align with the instrument space in the observed sample. It contributes to the sample first-stage fit even though it is not part of the population signal.
The same decomposition follows directly from the sample first-stage coefficient:
\[ \begin{aligned} \widehat\pi &=(Z’Z)^{-1}Z'(Z\pi+v)\\ &=\pi+(Z’Z)^{-1}Z’v. \end{aligned} \]Multiplying by \(Z\) gives
\[ \begin{aligned} Z\widehat\pi &=Z\pi+Z(Z’Z)^{-1}Z’v\\ &=Z\pi+P_Zv. \end{aligned} \]Thus, the difference between the sample fitted regressor and the population signal is
\[ \boxed{ \widehat X-Z\pi=P_Zv }. \]A.5.3. Why \(P_Zv\) need not be zero
The population orthogonality condition
\[ E[z_iv_i]=0 \]does not imply that its finite-sample counterpart is exactly zero:
\[ Z’v\neq0 \]in a typical sample. Sampling variation can create an accidental empirical relationship between the instruments and the population projection errors. Consequently,
\[ P_Zv=Z(Z’Z)^{-1}Z’v \]will generally not equal zero.
This may initially seem to contradict a familiar property of OLS: first-stage residuals are orthogonal to the instruments. The resolution is that the population residual \(v\) and the sample residual \(\widehat v\) are not the same object.
The sample residual is
\[ \widehat v=X-Z\widehat\pi=(I-P_Z)X. \]It satisfies the sample normal equations
\[ Z’\widehat v=0. \]Therefore,
\[ P_Z\widehat v = Z(Z’Z)^{-1}Z’\widehat v = 0. \]Indeed, substituting \(X=Z\pi+v\) gives
\[ \begin{aligned} \widehat v &=X-P_ZX\\ &=Z\pi+v-\left(Z\pi+P_Zv\right)\\ &=(I-P_Z)v. \end{aligned} \]The sample regression mechanically divides the population residual into two orthogonal components:
\[ \boxed{ v=P_Zv+(I-P_Z)v }. \]The first component, \(P_Zv\), is absorbed into the fitted values. Only the second component, \((I-P_Z)v=\widehat v\), remains in the reported first-stage residual. Sample orthogonality is therefore achieved partly by incorporating accidental noise into \(\widehat X\).
A.5.4. Why many instruments can confuse fit with identification
Adding instruments enlarges the column space of \(Z\). A larger projection space can capture more of the genuine signal in \(X\), but it can also capture more of the population residual \(v\).
Consequently, an improvement in the sample first-stage fit can come from either
\[ \text{more genuine predictive signal} \]or
\[ \text{more accidental fitting of }v. \]The fitted regressor alone does not distinguish between them.
In the extreme case in which \(K=n\) and \(Z\) has full rank, its columns span the entire \(n\)-dimensional observation space. Then
\[ P_Z=I_n, \]so
\[ P_ZX=X. \]The first stage fits \(X\) perfectly, but this perfect in-sample fit does not constitute meaningful identification. It occurs because the instrument space is large enough to reproduce every sample vector.
This yields the central distinction:
\[ \boxed{ \text{First-stage fit measures what }Z\text{ predicts in the sample;} } \] \[ \boxed{ \text{identification requires stable population variation in }Z\pi. } \]Numerous instruments can make \(P_ZX\) look highly informative by enlarging \(P_Zv\), even when the genuine population signal \(Z\pi\) remains weak.
A.6. Why fitting more first-stage noise creates many-instrument bias
The preceding decomposition shows that
\[ P_ZX=Z\pi+P_Zv. \]We now quantify why the second term is problematic when the endogenous part of \(X\) is correlated with the structural error.
Let the structural and first-stage equations be
\[ y=X\beta_0+u \]and
\[ X=Z\pi+v. \]The covariance between \(v_i\) and \(u_i\) captures the endogeneity remaining after projecting \(X\) on the instruments. Denote it by
\[ \sigma_{vu} = \operatorname{Cov}(v_i,u_i\mid Z). \]For clarity, suppose that the errors are conditionally centered and that their conditional cross-covariance matrix takes the simple form
\[ E[vu’\mid Z] = \sigma_{vu}I_n. \]Element by element, this assumption means
\[ E[v_i u_j\mid Z] = \begin{cases} \sigma_{vu}, & i=j,\\ 0, & i\neq j. \end{cases} \]Thus, \(v_i\) and \(u_i\) may be correlated within an observation, but errors from different observations are conditionally uncorrelated.
A.6.1. Deriving the expected contamination
Because \(P_Z\) is symmetric,
\[ v’P_Zu=(P_Zv)’u. \]This scalar measures the alignment between the component of the first-stage residual fitted by the instruments, \(P_Zv\), and the structural error \(u\).
Let \(p_{ij}\) denote element \((i,j)\) of \(P_Z\). Expanding the expression gives
\[ v’P_Zu = \sum_{i=1}^{n} \sum_{j=1}^{n} v_i p_{ij}u_j. \]Conditional on \(Z\), the projection matrix is fixed. Therefore,
\[ E[v’P_Zu\mid Z] = \sum_{i=1}^{n} \sum_{j=1}^{n} p_{ij}E[v_i u_j\mid Z]. \]All terms for which \(i\neq j\) vanish under the maintained covariance assumption. Only the diagonal terms remain:
\[ E[v’P_Zu\mid Z] = \sigma_{vu} \sum_{i=1}^{n}p_{ii}. \]The sum of a matrix’s diagonal elements is its trace, so
\[ E[v’P_Zu\mid Z] = \sigma_{vu}\operatorname{tr}(P_Z). \]A.6.2. Why \(\operatorname{tr}(P_Z)=K\)
Using the cyclic property of the trace,
\[ \begin{aligned} \operatorname{tr}(P_Z) &= \operatorname{tr} \left[ Z(Z’Z)^{-1}Z’ \right]\\ &= \operatorname{tr} \left[ (Z’Z)^{-1}Z’Z \right]\\ &= \operatorname{tr}(I_K)\\ &=K. \end{aligned} \]The same result has a geometric interpretation. A projection matrix has an eigenvalue of one for every dimension onto which it projects and an eigenvalue of zero for every orthogonal direction. Since the column space of \(Z\) has dimension \(K\),
\[ \operatorname{rank}(P_Z)=K \]and
\[ \operatorname{tr}(P_Z)=K. \]Consequently,
\[ \boxed{ E[v’P_Zu\mid Z] = K\sigma_{vu} }. \]This result quantifies the endogenous noise-fitting mechanism. Each additional linearly independent instrument increases the dimension of the fitted space and, under these simplifying assumptions, adds \(\sigma_{vu}\) to the expected alignment between fitted first-stage noise and the structural error.
Dividing by sample size gives
\[ \boxed{ E\left[ \frac{1}{n}v’P_Zu \;\middle|\; Z \right] = \frac{K}{n}\sigma_{vu} }. \]The instrument-to-sample-size ratio \(K/n\) is therefore central.
A.6.3. How the contamination enters 2SLS
For one endogenous regressor, the 2SLS estimator is
\[ \widehat\beta_{\mathrm{2SLS}} = (X’P_ZX)^{-1}X’P_Zy. \]Substituting \(y=X\beta_0+u\) gives
\[ \widehat\beta_{\mathrm{2SLS}}-\beta_0 = \frac{X’P_Zu}{X’P_ZX}. \]Now use the population first stage \(X=Z\pi+v\):
\[ \begin{aligned} X’P_Zu &=(Z\pi+v)’P_Zu\\ &=\pi’Z’u+v’P_Zu, \end{aligned} \]where the identity \(Z’P_Z=Z’\) has been used.
The first component,
\[ \pi’Z’u, \]is the interaction between the genuine population first-stage signal and the structural error. Under instrument validity, for example under
\[ E[u\mid Z]=0, \]its conditional expectation is zero:
\[ E[\pi’Z’u\mid Z]=0. \]The second component is different:
\[ E[v’P_Zu\mid Z] = K\sigma_{vu}. \]Thus, even when the instruments are exogenous, the estimated first stage can fit some of the endogenous component \(v\). Because \(v\) is correlated with \(u\), this fitted noise contaminates the second-stage regression.
The resulting numerator can be written as
\[ \boxed{ X’P_Zu = \underbrace{\pi’Z’u}_{\text{valid-signal sampling variation}} + \underbrace{v’P_Zu}_{\text{fitted endogenous noise}} }. \]A.6.4. Why the problem grows with the number of instruments
If \(K\) remains fixed while \(n\) grows,
\[ \frac{K}{n}\longrightarrow0. \]The normalized endogenous-noise term then disappears asymptotically. This is the conventional fixed-\(K\) environment in which standard 2SLS asymptotics are developed.
If \(K\) is large relative to \(n\), or grows proportionally with \(n\), then
\[ \frac{K}{n} \]need not be small. The fitted endogenous component can remain important even in a sample that appears large in absolute terms.
The identity
\[ E[v’P_Zu\mid Z]=K\sigma_{vu} \]does not by itself equal the exact bias of 2SLS. The estimator is a ratio of random variables, so in general
\[ E\left[ \frac{X’P_Zu}{X’P_ZX} \right] \neq \frac{E[X’P_Zu]}{E[X’P_ZX]}. \]Nevertheless, the identity isolates the mechanism behind many-instrument finite-sample bias. When the denominator is of order \(n\), the systematic contamination in the numerator is approximately of order \(K\), making the corresponding bias component roughly of order
\[ \frac{K}{n}, \]with its exact magnitude depending on first-stage strength and the joint distribution of the errors.
The extreme case makes the direction of the problem transparent. If \(K=n\) and \(Z\) has full rank, then
\[ P_Z=I_n. \]Consequently,
\[ \widehat\beta_{\mathrm{2SLS}} = (X’X)^{-1}X’y = \widehat\beta_{\mathrm{OLS}}. \]The instruments fit the endogenous regressor perfectly, but 2SLS has completely lost its separation from OLS. As the instrument space becomes excessively rich, the first stage increasingly reproduces \(X\), including the endogenous component that IV was intended to remove.
Under more general heteroskedasticity or dependence, define
\[ \Omega_{uv} = E[uv’\mid Z]. \]The corresponding result is
\[ E[v’P_Zu\mid Z] = \operatorname{tr}(P_Z\Omega_{uv}). \]It reduces to \(K\sigma_{vu}\) when \(\Omega_{uv}=\sigma_{vu}I_n\). The precise expression may therefore change outside the simple homoskedastic setting, but the underlying geometry remains the same: enlarging the instrument space allows the first stage to fit more directions in the endogenous residual.
The central lesson is
\[ \boxed{ \text{More fitted variation is not necessarily more identifying variation.} } \]Numerous instruments can increase the apparent first-stage fit by increasing \(P_Zv\). When \(v\) is correlated with the structural error, this converts sample overfitting into finite-sample bias rather than credible identification.
A.7. Why inverse GMM weights can become unstable
Efficient GMM uses the inverse of the covariance matrix of the moment conditions. Let
\[ \overline g_n(\theta) = \frac{1}{n}\sum_{i=1}^{n}g_i(\theta) \]denote the vector of sample moments, and let \(S\) denote the asymptotic covariance matrix of \(\sqrt n\,\overline g_n(\theta_0)\). The GMM criterion is
\[ Q_n(\theta) = \overline g_n(\theta)’W\overline g_n(\theta), \]with the efficient population weighting matrix
\[ W=S^{-1}. \]In practice, \(S\) is unknown. Two-step GMM therefore estimates it from first-step residuals and uses
\[ \widehat W=\widehat S^{-1}. \]This procedure is asymptotically efficient under valid moments, sufficiently strong identification, a fixed and manageable number of moments, and regularity conditions. The finite-sample difficulty is that \(\widehat S^{-1}\) can be very unstable when some moments are almost redundant.
A.7.1. Two nearly duplicate moments
Consider two standardized moment conditions whose asymptotic covariance matrix is
\[ S = \begin{pmatrix} 1&r\\ r&1 \end{pmatrix}, \]where each moment has unit asymptotic variance and \(r\) is their correlation. When \(r\) is close to one, the two moments move almost identically.
The eigenvalues of \(S\) solve
\[ \det(S-\lambda I) = \det \begin{pmatrix} 1-\lambda&r\\ r&1-\lambda \end{pmatrix} = (1-\lambda)^2-r^2 = 0. \]Hence,
\[ \lambda_{+}=1+r, \qquad \lambda_{-}=1-r. \]For \(0\leq r<1\), the smallest eigenvalue is therefore
\[ \boxed{ \lambda_{\min}(S)=1-r }. \]The corresponding normalized eigenvectors are
\[ q_{+} = \frac{1}{\sqrt{2}} \begin{pmatrix} 1\\ 1 \end{pmatrix}, \qquad q_{-} = \frac{1}{\sqrt{2}} \begin{pmatrix} 1\\ -1 \end{pmatrix}. \]These eigenvectors define two economically useful combinations of the moments:
\[ \overline g_{+} = q_{+}’\overline g = \frac{\overline g_1+\overline g_2}{\sqrt{2}}, \]which measures their common movement, and
\[ \overline g_{-} = q_{-}’\overline g = \frac{\overline g_1-\overline g_2}{\sqrt{2}}, \]which measures their difference. The variances represented by \(S\) along these two directions are
\[ \operatorname{Avar}(\sqrt n\,\overline g_{+})=1+r, \qquad \operatorname{Avar}(\sqrt n\,\overline g_{-})=1-r. \]When \(r\) is close to one, the moments contain almost the same information. Their common component has substantial variance, but their difference has almost no variance. At \(r=1\), the moments are exact duplicates:
\[ \det(S)=1-r^2=0, \]so \(S\) has rank one and cannot be inverted.
For example, if
\[ r=0.99, \]then
\[ \lambda_{+}=1.99, \qquad \lambda_{-}=0.01. \]The covariance matrix is therefore very close to singular.
A.7.2. Inversion amplifies the direction in which the moments almost cancel
The inverse of \(S\) is
\[ S^{-1} = \frac{1}{1-r^2} \begin{pmatrix} 1&-r\\ -r&1 \end{pmatrix}. \]When \(r=0.99\), this becomes approximately
\[ S^{-1} \approx \begin{pmatrix} 50.251&-49.749\\ -49.749&50.251 \end{pmatrix}. \]The large positive diagonal entries and large negative off-diagonal entries indicate that inverse weighting places considerable importance on the difference between the two nearly identical moments.
The eigenvalues of \(S^{-1}\) are the reciprocals of the eigenvalues of \(S\):
\[ \lambda_{+}(S^{-1}) = \frac{1}{1+r}, \qquad \lambda_{-}(S^{-1}) = \frac{1}{1-r}. \]Consequently, the GMM criterion can be written in the rotated moment directions as
\[ \boxed{ \overline g’S^{-1}\overline g = \frac{\overline g_{+}^{\,2}}{1+r} + \frac{\overline g_{-}^{\,2}}{1-r} }. \]For \(r=0.99\),
\[ \overline g’S^{-1}\overline g \approx 0.5025\,\overline g_{+}^{\,2} + 100\,\overline g_{-}^{\,2}. \]Thus, the common movement of the moments receives a weight of approximately \(0.5\), whereas their small difference receives a weight of \(100\).
In the ideal population model, this logic is sensible. If a particular combination of valid moments truly has very little variance, it is potentially a highly precise source of information. In a finite sample, however, the small difference
\[ \overline g_1-\overline g_2 \]may be driven disproportionately by sampling noise, numerical error, or slight misspecification. Inverse weighting then treats this fragile difference as exceptionally informative.
A.7.3. The perturbation identity
The sensitivity of an inverse matrix can be seen directly. Begin with the identity
\[ SS^{-1}=I. \]Take the differential of both sides. By the matrix product rule,
\[ (dS)S^{-1} + S\,d(S^{-1}) = 0. \]Premultiplying by \(S^{-1}\) gives
\[ S^{-1}(dS)S^{-1} + d(S^{-1}) = 0. \]Therefore,
\[ \boxed{ d(S^{-1}) = -S^{-1}(dS)S^{-1} }. \]This is the matrix analogue of the scalar derivative
\[ d\left(\frac{1}{s}\right) = -\frac{1}{s^2}\,ds. \]The two appearances of \(S^{-1}\) show how a small perturbation \(dS\) can be magnified through inversion. In the spectral norm, the first-order change satisfies
\[ \left\|d(S^{-1})\right\| \leq \left\|S^{-1}\right\|^2 \left\|dS\right\|. \]For the covariance matrix considered above,
\[ \left\|S^{-1}\right\| = \frac{1}{\lambda_{\min}(S)} = \frac{1}{1-r}. \]The possible amplification therefore grows at the rate
\[ \frac{1}{(1-r)^2}. \]The sensitivity can also be seen from the inverse weight attached to the difference direction:
\[ w_{-} = \frac{1}{1-r}. \]Small changes in \(r\) produce very large changes in this weight near \(r=1\):
\[ r=0.98 \quad\Longrightarrow\quad w_{-}=50, \] \[ r=0.99 \quad\Longrightarrow\quad w_{-}=100, \] \[ r=0.995 \quad\Longrightarrow\quad w_{-}=200. \]A rise of only \(0.005\), from \(r=0.99\) to \(r=0.995\), doubles the weight attached to the difference between the moments. A decline of \(0.01\), from \(r=0.99\) to \(r=0.98\), halves it.
The same problem is summarized by the spectral condition number
\[ \kappa(S) = \frac{\lambda_{\max}(S)} {\lambda_{\min}(S)} = \frac{1+r}{1-r}. \]At \(r=0.99\),
\[ \kappa(S) = \frac{1.99}{0.01} = 199. \]A large condition number means that inversion is sensitive to small perturbations in the estimated covariance matrix.
A.7.4. Why the point estimate can change
The GMM weighting matrix does not determine only the reported standard errors. In an overidentified model, it also enters directly into the point estimate.
For illustration, suppose that the sample moments are linear in a \(p\times1\) parameter vector:
\[ \overline g(\theta) = a-B\theta, \]where \(a\) is \(K\times1\) and \(B\) is \(K\times p\). Minimizing
\[ Q(\theta) = (a-B\theta)’W(a-B\theta) \]gives the first-order condition
\[ B’W(a-B\widehat\theta_W)=0. \]Provided \(B’WB\) is invertible,
\[ \boxed{ \widehat\theta_W = (B’WB)^{-1}B’Wa }. \]The weighting matrix \(W\) therefore appears explicitly in the formula for the coefficient estimate.
When \(K>p\), the model is overidentified. Because of sampling variation, the different moments will generally not all equal zero at the same parameter value. GMM must choose a weighted compromise among them. If the estimated weighting matrix changes sharply, that compromise can also change sharply.
This explains why the following choices may affect the point estimate:
one-step GMM uses an initial weighting matrix, whereas two-step GMM uses an estimated inverse covariance matrix \(\widehat S^{-1}\);
changing the lag window changes both the number and the composition of the moment conditions;
collapsing the instrument matrix removes or combines moment conditions and changes their covariance structure;
adding numerous highly correlated instruments can create very small eigenvalues in \(\widehat S\), producing large and unstable implicit weights.
Dynamic-panel GMM is particularly exposed to this issue because instruments such as
\[ y_{i,t-2}, \quad y_{i,t-3}, \quad y_{i,t-4}, \quad \ldots \]can be highly correlated. The associated moments may then be close substitutes rather than genuinely distinct sources of identifying information.
There is an important exception. Under exact identification, \(K=p\), and if \(B\) is nonsingular, the moments can be set exactly equal to zero:
\[ a-B\widehat\theta=0. \]Hence,
\[ \widehat\theta=B^{-1}a, \]which does not depend on \(W\). Weighting affects the point estimate only when GMM must arbitrate among more moment conditions than parameters.
Under the ideal fixed-moment asymptotic model, different positive-definite weighting matrices remain consistent for the same true parameter, while \(S^{-1}\) delivers the smallest asymptotic variance. The concern is therefore not that inverse weighting is intrinsically incorrect. It is that, with numerous nearly redundant moments, finite-sample estimation noise or slight misspecification in \(\widehat S\) can receive disproportionate influence.
The main lesson is:
\[ \boxed{ \text{When moments are nearly redundant, } \widehat S^{-1} \text{ can transform small sampling differences into large GMM weights.} } \]This is why large movements between one-step and two-step estimates, alternative lag windows, or collapsed and uncollapsed instrument sets are substantively informative. They may reveal that the estimated result depends heavily on a fragile weighting of nearly duplicate moments, not merely that its reported standard error has changed.
A.8. Why lagged levels become weak instruments near a random walk
Consider the demeaned stationary AR(1) process
\[ y_t=\rho y_{t-1}+\varepsilon_t, \qquad 0<\rho<1, \]where
\[ E[\varepsilon_t\mid\mathcal F_{t-1}]=0, \qquad \operatorname{Var}(\varepsilon_t)=\sigma_\varepsilon^2. \]The innovation \(\varepsilon_t\) is therefore orthogonal to all information dated \(t-1\) or earlier. In difference GMM, an older level such as \(y_{t-2}\) may be used as an instrument for the differenced variable \(\Delta y_{t-1}\). This appendix derives the strength of that relationship as the process approaches a random walk.
A.8.1. The population first stage
Write the AR(1) equation one period earlier:
\[ y_{t-1}=\rho y_{t-2}+\varepsilon_{t-1}. \]Subtracting \(y_{t-2}\) from both sides gives
\[ \begin{aligned} \Delta y_{t-1} &=y_{t-1}-y_{t-2}\\ &=(\rho-1)y_{t-2}+\varepsilon_{t-1}. \end{aligned} \]This is the population first-stage equation. Because \(y_{t-2}\) belongs to the information set available before \(\varepsilon_{t-1}\) is realized,
\[ E[y_{t-2}\varepsilon_{t-1}]=0. \]The population first-stage slope is therefore
\[ \boxed{\pi=\rho-1}. \]For example,
\[ \rho=0.95 \quad\Longrightarrow\quad \pi=-0.05, \]whereas
\[ \rho=0.99 \quad\Longrightarrow\quad \pi=-0.01. \]As \(\rho\) approaches one, the first-stage slope approaches zero:
\[ \rho\longrightarrow1^{-} \quad\Longrightarrow\quad \pi=\rho-1\longrightarrow0. \]The intuition is immediate. Near the random-walk limit,
\[ y_{t-1}\approx y_{t-2}+\varepsilon_{t-1}, \]and hence
\[ \Delta y_{t-1}\approx\varepsilon_{t-1}. \]The change is then almost entirely new information. An older level cannot predict an innovation that is orthogonal to past information.
A.8.2. Variances of the level and the change
Let
\[ V_y=\operatorname{Var}(y_t). \]Stationarity implies that \(\operatorname{Var}(y_t)=\operatorname{Var}(y_{t-1})\). Taking variances in the AR(1) equation gives
\[ V_y=\rho^2V_y+\sigma_\varepsilon^2. \]Therefore,
\[ (1-\rho^2)V_y=\sigma_\varepsilon^2, \]so the stationary variance of the level is
\[ \boxed{ V_y = \operatorname{Var}(y_t) = \frac{\sigma_\varepsilon^2}{1-\rho^2} }. \]To obtain the variance of the change, use
\[ \Delta y_{t-1}=y_{t-1}-y_{t-2}. \]For a stationary AR(1),
\[ \operatorname{Cov}(y_{t-1},y_{t-2}) = \rho V_y. \]Consequently,
\[ \begin{aligned} \operatorname{Var}(\Delta y_{t-1}) &= \operatorname{Var}(y_{t-1}) + \operatorname{Var}(y_{t-2})\\ &\quad -2\operatorname{Cov}(y_{t-1},y_{t-2})\\ &= V_y+V_y-2\rho V_y\\ &= 2(1-\rho)V_y. \end{aligned} \]Substituting the stationary variance of \(y_t\),
\[ \begin{aligned} \operatorname{Var}(\Delta y_{t-1}) &= 2(1-\rho) \frac{\sigma_\varepsilon^2} {(1-\rho)(1+\rho)}\\ &= \boxed{ \frac{2\sigma_\varepsilon^2}{1+\rho} }. \end{aligned} \]A.8.3. Covariance between the older level and the change
Starting from the first-stage equation,
\[ \Delta y_{t-1} = (\rho-1)y_{t-2} + \varepsilon_{t-1}, \]take the covariance with \(y_{t-2}\):
\[ \begin{aligned} \operatorname{Cov}(y_{t-2},\Delta y_{t-1}) &= (\rho-1)\operatorname{Var}(y_{t-2})\\ &\quad+ \operatorname{Cov}(y_{t-2},\varepsilon_{t-1}). \end{aligned} \]Innovation orthogonality implies
\[ \operatorname{Cov}(y_{t-2},\varepsilon_{t-1})=0. \]It follows that
\[ \operatorname{Cov}(y_{t-2},\Delta y_{t-1}) = (\rho-1)V_y. \]Substituting
\[ V_y=\frac{\sigma_\varepsilon^2}{1-\rho^2} \]gives
\[ \begin{aligned} \operatorname{Cov}(y_{t-2},\Delta y_{t-1}) &= (\rho-1) \frac{\sigma_\varepsilon^2} {(1-\rho)(1+\rho)}\\ &= -(1-\rho) \frac{\sigma_\varepsilon^2} {(1-\rho)(1+\rho)}\\ &= \boxed{ -\frac{\sigma_\varepsilon^2}{1+\rho} }. \end{aligned} \]The negative sign reflects mean reversion. When \(y_{t-2}\) is unusually high and \(0<\rho<1\), the subsequent change tends to be negative. The same result can be obtained directly from
\[ \begin{aligned} \operatorname{Cov}(y_{t-2},\Delta y_{t-1}) &= \operatorname{Cov}(y_{t-2},y_{t-1}-y_{t-2})\\ &= \rho V_y-V_y\\ &= (\rho-1)V_y. \end{aligned} \]A.8.4. From covariance to correlation
Correlation standardizes covariance by the standard deviations of the two variables:
\[ \operatorname{Corr}(X,Y) = \frac{ \operatorname{Cov}(X,Y) }{ \sqrt{\operatorname{Var}(X)} \sqrt{\operatorname{Var}(Y)} }. \]Therefore,
\[ \operatorname{Corr}(y_{t-2},\Delta y_{t-1}) = \frac{ (\rho-1)V_y }{ \sqrt{V_y} \sqrt{2(1-\rho)V_y} }. \]The denominator simplifies as follows:
\[ \sqrt{V_y} \sqrt{2(1-\rho)V_y} = V_y\sqrt{2(1-\rho)}. \]Hence,
\[ \begin{aligned} \operatorname{Corr}(y_{t-2},\Delta y_{t-1}) &= \frac{\rho-1}{\sqrt{2(1-\rho)}}\\ &= -\frac{1-\rho}{\sqrt{2(1-\rho)}}\\ &= \boxed{ -\sqrt{\frac{1-\rho}{2}} }. \end{aligned} \]The correlation approaches zero as persistence approaches the random-walk limit:
\[ \rho\longrightarrow1^{-} \quad\Longrightarrow\quad \operatorname{Corr}(y_{t-2},\Delta y_{t-1}) \longrightarrow0. \]Numerically, when \(\rho=0.95\),
\[ \operatorname{Corr}(y_{t-2},\Delta y_{t-1}) = -\sqrt{\frac{0.05}{2}} \approx-0.158. \]When \(\rho=0.99\),
\[ \operatorname{Corr}(y_{t-2},\Delta y_{t-1}) = -\sqrt{\frac{0.01}{2}} \approx-0.071. \]In a simple first-stage regression containing one instrument and an intercept—or, equivalently, using demeaned variables—the population \(R^2\) is the squared correlation:
\[ \boxed{ R^2 = \operatorname{Corr}(y_{t-2},\Delta y_{t-1})^2 = \frac{1-\rho}{2} }. \]Thus,
\[ \rho=0.95 \quad\Longrightarrow\quad R^2=0.025, \]so the older level explains only \(2.5\%\) of the variation in the change. Similarly,
\[ \rho=0.99 \quad\Longrightarrow\quad R^2=0.005, \]so it explains only \(0.5\%\). The precise weak-instrument diagnosis also depends on sample size and the concentration parameter, but these calculations clearly show why relevance deteriorates as \(\rho\) approaches one.
A.8.5. Why the raw covariance does not approach zero
At first sight, there appears to be a contradiction. The first-stage slope and correlation both approach zero, but, holding the innovation variance fixed,
\[ \operatorname{Cov}(y_{t-2},\Delta y_{t-1}) = -\frac{\sigma_\varepsilon^2}{1+\rho} \longrightarrow -\frac{\sigma_\varepsilon^2}{2}. \]The raw covariance therefore does not converge to zero. The reason is that covariance depends on scale. As \(\rho\) approaches one under the stationary family with fixed \(\sigma_\varepsilon^2\),
\[ \operatorname{Var}(y_{t-2}) = \frac{\sigma_\varepsilon^2}{1-\rho^2} \longrightarrow\infty. \]Recall that
\[ \operatorname{Cov}(y_{t-2},\Delta y_{t-1}) = \underbrace{(\rho-1)}_{\longrightarrow0} \underbrace{\operatorname{Var}(y_{t-2})}_{\longrightarrow\infty}. \]The vanishing first-stage slope is multiplied by an exploding variance. More precisely,
\[ \begin{aligned} (\rho-1)\operatorname{Var}(y_{t-2}) &= -(1-\rho) \frac{\sigma_\varepsilon^2} {(1-\rho)(1+\rho)}\\ &= -\frac{\sigma_\varepsilon^2}{1+\rho}. \end{aligned} \]The two effects offset one another, leaving a finite nonzero covariance. This does not indicate strong relevance. Dividing the covariance by the variance of the instrument recovers the first-stage slope:
\[ \begin{aligned} \pi &= \frac{ \operatorname{Cov}(y_{t-2},\Delta y_{t-1}) }{ \operatorname{Var}(y_{t-2}) }\\ &= \frac{ -\sigma_\varepsilon^2/(1+\rho) }{ \sigma_\varepsilon^2/[(1-\rho)(1+\rho)] }\\ &= -(1-\rho) = \rho-1 \longrightarrow0. \end{aligned} \]An alternative normalization makes the scale issue transparent. If the unconditional variance of \(y_t\) is held equal to one as \(\rho\) changes, then the innovation variance must satisfy
\[ \sigma_\varepsilon^2=1-\rho^2. \]Under this normalization,
\[ \operatorname{Cov}(y_{t-2},\Delta y_{t-1}) = \rho-1 \longrightarrow0. \]The correlation is unchanged by this rescaling and remains
\[ -\sqrt{\frac{1-\rho}{2}}. \]This is why the first-stage slope and correlation, rather than an unstandardized covariance considered in isolation, reveal the deterioration in relevance.
A.8.6. The exact unit-root boundary
The preceding formulas were derived under stationarity and therefore apply only when \(\rho<1\). At exactly \(\rho=1\), the process becomes
\[ y_t=y_{t-1}+\varepsilon_t, \]so
\[ \Delta y_{t-1}=\varepsilon_{t-1}. \]With a finite initial condition and innovations orthogonal to past information, \(y_{t-2}\) is composed entirely of information dated \(t-2\) or earlier. Therefore,
\[ \operatorname{Cov}(y_{t-2},\Delta y_{t-1}) = \operatorname{Cov}(y_{t-2},\varepsilon_{t-1}) = 0. \]This may seem inconsistent with the stationary limit
\[ \lim_{\rho\to1^{-}} \operatorname{Cov}(y_{t-2},\Delta y_{t-1}) = -\frac{\sigma_\varepsilon^2}{2}. \]There is no contradiction because the two calculations use different initializations. In the stationary calculation, the process is initialized from its stationary distribution for every \(\rho<1\). The variance of that distribution is
\[ \frac{\sigma_\varepsilon^2}{1-\rho^2}, \]which diverges as \(\rho\) approaches one. At the exact unit root, no finite-variance stationary distribution exists. The random walk must instead be defined from a finite initial condition or another nonstationary initialization.
The two operations therefore do not commute:
\[ \boxed{ \text{take the stationary limit as }\rho\to1^{-} \neq \text{set }\rho=1\text{ first}. } \]The economically relevant conclusion is nevertheless the same in both cases. Near the unit-root boundary, the lagged level has very little standardized predictive power for the subsequent change; at the exact random walk, it has none:
\[ \boxed{ \text{high persistence makes lagged levels weak instruments for first differences.} } \]