Why can an estimator become less credible when it uses more instruments? Because additional moment conditions are not necessarily additional identifying information.
Generalized Method of Moments was a genuine econometric breakthrough. It allowed economists to estimate models from a limited set of theoretically meaningful restrictions without specifying a complete likelihood. The difficulty came later, especially in dynamic-panel applications, when mechanically generated internal instruments began to substitute for an identification argument. This note explains, step by step, why numerous weak and highly correlated moments can add little genuine information while increasing finite-sample bias, overfitting and weighting-matrix instability.
\[ \boxed{\text{GMM can combine identifying information; it cannot create it.}} \]1. What I actually dislike
The title is deliberately provocative. GMM itself has not disappeared, and I do not object to its mathematics. OLS, instrumental variables and two-stage least squares can all be represented inside the broad GMM framework. Moment-based estimation remains fundamental in econometrics.
My criticism concerns a particular applied practice: using a large set of internally generated instruments—especially in difference or system GMM—as if the estimator itself solved endogeneity.
I distrust GMM when multiplying moment conditions substitutes for an identification argument.
Endogeneity is not fundamentally an optimization problem. It is an information problem. If a regressor is correlated with the structural error, we need a credible source of exogenous variation. No weighting matrix can manufacture that variation.
The correct hierarchy is
\[ \boxed{ \text{research design} \longrightarrow \text{credible exogenous variation} \longrightarrow \text{moment conditions} \longrightarrow \text{estimation} }. \]Mechanical applications often reverse this order: choose system GMM, generate many lags, report a Hansen test, and then claim that endogeneity has been addressed. That is precisely the practice I find unconvincing.
2. Two rises and a credibility correction
The expression “rise and fall” should not be interpreted as the disappearance of GMM. The historically defensible story is that GMM rose in two waves and subsequently underwent a credibility correction.
\[ \boxed{ \text{innovation} \rightarrow \text{diffusion with simultaneous warnings} \rightarrow \text{dynamic-panel boom} \rightarrow \text{credibility correction} } \]Before 1982 — Method of moments, IV, minimum distance and overidentification were already established. Sargan (1958) is a central precursor. The basic idea of translating economic assumptions into orthogonality restrictions predates the name “GMM.”
1982 — Hansen unified estimation, efficient weighting and overidentification testing. Hansen and Singleton applied the framework to nonlinear rational-expectations models. This was the first rise: economists could estimate theoretically meaningful restrictions without specifying the complete data-generating process.
1981–1998 — Dynamic-panel IV developed through Anderson–Hsiao, Arellano–Bond, Arellano–Bover and Blundell–Bond. This was the second rise: internal lags appeared to offer a practical answer to fixed effects, dynamics and the absence of external instruments.
1986–1990s — Tauchen, Bekker, Hansen–Heaton–Yaron, Altonji–Segal and Ziliak documented finite-sample, weak-instrument, many-instrument and estimated-weighting problems. The warnings began during the expansion—not after it.
Late 1990s–2000s — Difference and system GMM diffused rapidly through firm, household and cross-country panels, helped by automated software. This was the applied boom: sophisticated-looking estimates and diagnostics could be produced even without external instruments.
2000–2010 — Stock–Wright, Bowsher, Windmeijer, Roodman, Newey–Windmeijer and Bun–Windmeijer made weak identification and instrument proliferation central applied concerns. This was a highly visible credibility correction, not a sudden end to GMM usage.
Since 2010 — More attention has been paid to restricted or collapsed instrument sets, lag-window sensitivity, corrected inference and alternative panel estimators. This is better understood as selective survival and methodological maturation.
The chronology is important. Tauchen (1986) was already documenting a bias–variance trade-off from expanding lag instruments. Yet in 2009, Roodman still described difference and system GMM as growing in popularity. There was no clean “enthusiasm first, criticism later” sequence. Adoption and criticism overlapped.
Wooldridge (2001) offered a particularly balanced assessment during the boom: sophisticated GMM can be indispensable in complicated models, yet it is often unlikely to improve convincingly on OLS or 2SLS in typical applied settings.
3. From OLS to a moment condition
Consider the simple linear model
\[ y_i=\beta_0x_i+u_i. \]For clarity, suppose the variables have been centered, so an intercept is unnecessary. OLS minimizes the sum of squared residuals. Its first-order condition is
\[ \frac{1}{n}\sum_{i=1}^{n}x_i(y_i-\beta x_i)=0. \]Solving gives
\[ \widehat\beta_{\mathrm{OLS}} = \frac{\sum_i x_i y_i}{\sum_i x_i^2}. \]Substitute \(y_i=\beta_0x_i+u_i\):
\[ \widehat\beta_{\mathrm{OLS}} = \beta_0+ \frac{\sum_i x_i u_i}{\sum_i x_i^2}. \]Consequently,
\[ \operatorname{plim}\widehat\beta_{\mathrm{OLS}} = \beta_0+ \frac{E[x_i u_i]}{E[x_i^2]}. \]If
\[ E[x_i u_i]\neq0, \]OLS is inconsistent. The estimator has not failed to perform its calculation. The population condition underlying that calculation is false.
First lesson: an estimator does not eliminate endogeneity. It works only when the identifying restrictions behind it are credible.
4. What an instrument must actually provide
Suppose we find an instrument \(z_i\). It must satisfy two distinct conditions:
\[ \underbrace{E[z_i u_i]=0}_{\text{validity}}, \qquad \underbrace{E[z_i x_i]\neq0}_{\text{relevance}}. \]The IV estimator solves
\[ \frac1n\sum_i z_i(y_i-\beta x_i)=0, \]which gives
\[ \widehat\beta_{\mathrm{IV}} = \frac{\sum_i z_i y_i}{\sum_i z_i x_i}. \]Substituting the structural equation yields the most useful IV formula:
\[ \boxed{ \widehat\beta_{\mathrm{IV}}-\beta_0 = \frac{\frac1n\sum_i z_i u_i} {\frac1n\sum_i z_i x_i} }. \]The numerator concerns validity. The denominator concerns relevance. A perfectly exogenous instrument can nevertheless be nearly useless if the denominator is close to zero.
A simple analogy helps. A valid instrument is an honest witness: it is not systematically related to the unobserved forces contained in \(u_i\). A relevant instrument is a witness who actually saw the event: it contains substantial information about \(x_i\). An honest witness who saw almost nothing cannot identify what happened.
In short, exogeneity makes an instrument honest; relevance makes it informative. We need both.
Validity says that the population moment \(E[z_i u_i]\) is zero. It does not force the sample moment \(\overline{zu}\) to equal zero exactly. When \(\overline{zx}\) is close to zero, an ordinary sampling discrepancy in the numerator is divided by almost no relevance and is therefore greatly magnified.
Suppose the sample moment in the numerator is only \(0.01\). Then:
When \(\overline{zx}=0.50\), the coefficient error is \(0.01/0.50=0.02\).
When \(\overline{zx}=0.05\), the coefficient error is \(0.01/0.05=0.20\).
When \(\overline{zx}=0.01\), the coefficient error is \(0.01/0.01=1.00\).
The same small sample cross-moment with the error produces a coefficient error fifty times larger when relevance falls from \(0.50\) to \(0.01\).
Second lesson:
\[ \boxed{\text{Validity does not imply relevance.}} \]A weak instrument generates a flat statistical problem: ordinary sampling noise is divided by very little identifying information.
5. What GMM really adds
Now suppose there are \(K\) instruments collected in the vector
\[ z_i= \begin{pmatrix} z_{1i}\\ z_{2i}\\ \vdots\\ z_{Ki} \end{pmatrix}. \]They generate \(K\) moment conditions:
\[ E\!\left[z_{ji}(y_i-\beta_0x_i)\right]=0, \qquad j=1,\ldots,K. \]Define
\[ g_i(\beta)=z_i(y_i-\beta x_i), \qquad \overline g_n(\beta)=\frac1n\sum_{i=1}^{n}g_i(\beta). \]GMM chooses
\[ \widehat\beta_{\mathrm{GMM}} = \arg\min_{\beta} \overline g_n(\beta)’W_n\overline g_n(\beta), \]where \(W_n\) is a positive-definite weighting matrix.
If the number of moments equals the number of parameters, the model is exactly identified. The sample moments can be set to zero, and the choice of \(W_n\) does not affect the estimate.
If there are more moments than parameters, the model is overidentified. Sampling variation generally makes it impossible to set every sample moment to zero. GMM must choose a weighted compromise.
GMM as an average of instrument-specific estimates
For each instrument, define the just-identified estimate
\[ \widehat\beta_j = \frac{\overline{z_jy}}{\overline{z_jx}}, \qquad a_j=\overline{z_jx}. \]If \(W_n\) is diagonal, GMM becomes
\[ \boxed{ \widehat\beta_{\mathrm{GMM}} = \frac{\sum_{j=1}^{K}w_j a_j^2\widehat\beta_j} {\sum_{j=1}^{K}w_j a_j^2} }. \]It is a weighted average of the estimates preferred by the different moments.
Suppose one moment favours \(\beta=1\), while another favours \(\beta=3\). Then
\[ Q(\beta)=w_1(\beta-1)^2+w_2(\beta-3)^2, \]and
\[ \widehat\beta=\frac{w_1+3w_2}{w_1+w_2}. \]Thus,
\[ (w_1,w_2)=(9,1) \quad\Rightarrow\quad \widehat\beta=1.2, \]whereas
\[ (w_1,w_2)=(1,9) \quad\Rightarrow\quad \widehat\beta=2.8. \]If both moments are valid and strong, their preferred estimates converge to the same population parameter. If they remain incompatible because one is invalid or both are weak, the weighting matrix decides which moment to believe.
Why this appeared so attractive
Under independent sampling—or when \(i\) indexes independent clusters—let
\[ S=E[g_i(\theta_0)g_i(\theta_0)’] \]be the covariance matrix of the moments. With serially dependent moment contributions, \(S\) instead denotes their long-run covariance. Also let
\[ G=E\!\left[\frac{\partial g_i(\theta_0)}{\partial\theta’}\right] \]measure how strongly the moments respond to the parameters. The general asymptotic variance is
\[ \operatorname{Avar}\!\left[\sqrt n(\widehat\theta-\theta_0)\right] = (G’WG)^{-1}G’WSWG(G’WG)^{-1}. \]The efficient choice is
\[ W=S^{-1}, \]which gives
\[ \operatorname{Avar}\!\left[\sqrt n(\widehat\theta-\theta_0)\right] = (G’S^{-1}G)^{-1}. \]Under valid moments, sufficient identification, fixed \(K\), conventional regularity conditions and a large sample, adding informative moments can improve efficiency. This was—and remains—a genuine virtue of GMM.
The qualifications matter:
\[ \text{valid moments} + \text{relevant moments} + \text{fixed }K + \text{large effective sample}. \]Applied instrument-heavy GMM can depart substantially from this ideal experiment.
6. The master formula for GMM sensitivity
A first-order expansion of the GMM conditions gives
\[ \boxed{ \widehat\theta-\theta_0 \approx -(G’WG)^{-1}G’W\overline g_n(\theta_0) }. \]This formula contains the entire intuition:
- \(\overline g_n(\theta_0)\): sampling noise—or systematic moment invalidity.
- \(W\): which moments GMM chooses to trust.
- \(G\): how strongly the moments respond to the parameters.
- \((G’WG)^{-1}\): how strongly noise is amplified when identification is weak.
For scalar linear GMM, define
\[ a_n=\frac{Z’x}{n}, \qquad e_n=\frac{Z’u}{n}. \]Then, for a fixed final-step weighting matrix,
\[ \boxed{ \widehat\beta-\beta_0 = \frac{a_n’W_ne_n}{a_n’W_na_n} }. \]The numerator \(a_n’W_ne_n\) measures the weighted noise that GMM combines. The denominator \(a_n’W_na_n\) measures weighted identifying strength. The matrix \(W_n\) changes how both are combined, but it cannot create relevance that is absent from the moments.
Equivalently, the same idea can be remembered as
\[ \boxed{ \text{estimation error} = \frac{\text{weighted moment noise}} {\text{identifying strength}} }. \]Weak moments make the criterion flat
The curvature of the scalar linear-GMM criterion is
\[ \frac{\partial^2 Q(\beta)}{\partial\beta^2} =2a_n’W_na_n. \]If the instruments barely predict the endogenous regressor, \(a_n’W_na_n\) is small. The objective has a flat valley. A small change in the sample moments can then move its minimum considerably, even when numerical optimization is flawless.
Slight invalidity changes the population target
Let
\[ a=E[z_i x_i], \qquad \delta=E[z_i u_i]. \]If some moments are invalid, \(\delta\neq0\), and the pseudo-true GMM value satisfies
\[ \boxed{ \beta^*(W)-\beta_0 = \frac{a’W\delta}{a’Wa} }. \]Different instrument sets and different weighting matrices can then imply genuinely different probability limits. Sensitivity is no longer merely a small-sample inconvenience; it reflects incompatible identifying assumptions.
Correlated moments are echoes, not independent information
Suppose \(K\) moments have equal relevance \(a\), equal variance, and pairwise correlation \(r\). In a simplified equal-weight setting,
\[ \widehat\beta-\beta_0 = \frac{1}{Ka}\sum_{j=1}^{K}e_j. \]If
\[ \operatorname{Var}(e_j)=\frac{\sigma_e^2}{n}, \qquad \operatorname{Corr}(e_j,e_\ell)=r, \]then
\[ \operatorname{Var}(\widehat\beta-\beta_0) = \frac{\sigma_e^2}{na^2} \frac{1+(K-1)r}{K}. \]If \(r=0\), variance falls as \(1/K\). As \(r\) approaches one, the variance reduction disappears. At \(r=1\), the moments are exact duplicates and their covariance matrix is singular: ten copies of the same moment are not ten independent sources of identification.
Ten highly correlated moment conditions are closer to ten microphones recording the same voice than to ten independent witnesses. The data contain ten recordings, but mostly one underlying signal. In the same way, numerous lags of one persistent variable can generate many columns in the instrument matrix while contributing very little independent information.
The key distinction is:
\[ \boxed{ \text{number of moments} \neq \text{amount of independent identifying information} }. \]7. How many weak moments aggravate the problem
Several individually weak moments can jointly be informative. The problem is not “more moments” by itself. The problem arises when the number of moments grows faster than genuine identifying information because the additional instruments are weak, redundant, highly correlated or slightly invalid.
Many instruments can confuse fit with identification
Two-stage least squares, a special case of linear GMM, can be written as
\[ \widehat\beta_{\mathrm{2SLS}} = (X’P_ZX)^{-1}X’P_Zy, \]where
\[ P_Z=Z(Z’Z)^{-1}Z’ \]projects variables onto the instrument space.
Write the population linear projection of the endogenous regressor on the instruments as
\[ X=Z\pi+v. \]Here \(\pi\) is the population projection coefficient, defined so that \(E[z_i v_i]=0\).
The fitted regressor is
\[ P_ZX = \underbrace{Z\pi}_{\text{population signal predicted by }Z} + \underbrace{P_Zv}_{\text{in-sample fit to }v}. \]The first component is the population variation predicted by \(Z\). It is useful for identification only if the instruments are valid and sufficiently relevant. The second component is sample-specific fit to the first-stage residual. When \(v\) contains the endogenous part of \(X\), a flexible instrument set can accidentally fit some of that endogenous variation.
With a small, economically justified instrument set, we hope the population signal dominates. As the number of instruments increases, the projection becomes more flexible and can fit more of \(v\) inside the estimation sample. The first stage may then fit the observed \(X\) extremely well without isolating much genuinely exogenous variation.
In the extreme full-rank thought experiment, if \(K=n\), then
\[ P_Z=I_n. \]Consequently,
\[ \widehat\beta_{\mathrm{2SLS}} = (X’X)^{-1}X’y = \widehat\beta_{\mathrm{OLS}}. \]The first stage now fits \(X\) perfectly, but that perfect fit is completely unhelpful: it reproduces every component of \(X\), including the endogenous component.
A perfect first-stage fit is not necessarily perfect identification. With enough instruments, the first stage may simply have learned the estimation sample too well.
\[ \boxed{ \text{More instruments can improve in-sample fit without adding exogenous information.} } \]The same mechanism can be seen algebraically. Let \(\sigma_{vu}\) denote the conditional covariance between the first-stage residual \(v_i\) and the structural error \(u_i\). If \(E[vu’\mid Z]=\sigma_{vu}I_n\) and \(Z\) has rank \(K\), then
\[ E[v’P_Zu\mid Z] = \sigma_{vu}\operatorname{tr}(P_Z) = K\sigma_{vu}. \]This is not a universal formula for GMM bias. It isolates the conventional many-instrument mechanism: expanding the instrument space increases its in-sample capacity to fit the part of the first-stage residual that is correlated with the structural error. This helps explain the familiar finite-sample tendency:
\[ \boxed{ \text{a highly saturated instrument space can pull IV toward endogenous OLS} }. \]The estimated weighting matrix is another noisy object
Efficient two-step GMM uses
\[ \widehat W=\widehat S^{-1}, \qquad \widehat S = \frac1N\sum_{i=1}^{N} g_i(\widehat\theta_1)g_i(\widehat\theta_1)’. \]In panel GMM, \(N\) is typically the number of independent individuals or clusters. Thus, the relevant comparison is often \(K\) relative to \(N\), not relative to \(NT\).
Because \(\widehat S\) is a sum of only \(N\) outer products,
\[ \operatorname{rank}(\widehat S)\leq\min(K,N). \]If \(K>N\), the matrix cannot have full rank; it can become poorly conditioned well before that boundary.
A symmetric \(K\times K\) matrix contains
\[ \frac{K(K+1)}{2} \]distinct covariance elements. Therefore,
\[ K=10\Rightarrow55, \qquad K=50\Rightarrow1{,}275. \]Technical aside: why inverse weights can become unstable
When moments are nearly duplicates, their covariance matrix is nearly singular. Inverting that matrix can amplify small sampling differences in the direction where the moments almost cancel.
For the simple covariance matrix
\[ S= \begin{pmatrix} 1 & r\\ r & 1 \end{pmatrix}, \qquad \lambda_{\min}(S)=1-r. \]When \(r=0.99\), the smallest eigenvalue is only \(0.01\). The matrix is therefore almost singular. The perturbation identity
\[ \boxed{ d(S^{-1})=-S^{-1}(dS)S^{-1} }. \]shows why estimation noise in \(\widehat S\) can be magnified through inversion. This does not mean that inverse weighting is intrinsically wrong under the ideal fixed-moment model. It means that finite-sample noise or slight misspecification can matter greatly when the moments are nearly redundant.
This is why changing from one-step to two-step GMM, changing the lag window, or collapsing the instrument matrix may affect the point estimate—not merely its reported standard error.
8. Why dynamic-panel GMM is especially fragile
Consider the dynamic panel
\[ y_{it}=\rho y_{i,t-1}+\alpha_i+\varepsilon_{it}, \]where \(\alpha_i\) is an individual effect. First differencing removes it:
\[ \Delta y_{it} = \rho\Delta y_{i,t-1} + \Delta\varepsilon_{it}. \]But the differenced lag is endogenous. Indeed,
\[ \Delta y_{i,t-1}=y_{i,t-1}-y_{i,t-2} \]contains \(\varepsilon_{i,t-1}\), while
\[ \Delta\varepsilon_{it} = \varepsilon_{it}-\varepsilon_{i,t-1}. \]Therefore,
\[ \operatorname{Cov}(\Delta y_{i,t-1},\Delta\varepsilon_{it})\neq0. \]Difference GMM
Under the usual assumptions that sufficiently old outcomes are uncorrelated with subsequent innovations and that the idiosyncratic errors have no serial correlation, those old levels satisfy
\[ E[y_{i,t-s}\Delta\varepsilon_{it}]=0, \qquad s\geq2. \]The available instruments accumulate:
At \(t=3\), the available level instrument is \(y_{i1}\).
At \(t=4\), the available level instruments are \(y_{i1},y_{i2}\).
At \(t=5\), the available level instruments are \(y_{i1},y_{i2},y_{i3}\), and the sequence continues in this way.
In a balanced panel with one endogenous lagged variable, all available level instruments and an uncollapsed difference-GMM matrix, the count is
\[ K_D = 1+2+\cdots+(T-2) = \frac{(T-1)(T-2)}{2}. \]Hence,
\[ T=10\Rightarrow K_D=36, \qquad T=20\Rightarrow K_D=171. \]The count grows approximately with \(T^2\). A modest increase in the time dimension can therefore generate a surprisingly large instrument set.
These counts concern this specific uncollapsed setup. Collapsing, restricting the lag range, adding other endogenous variables, including time indicators, or adding the system-GMM level equations changes the count.
Persistence predicts levels but not changes
In a stripped-down centered AR(1),
\[ y_{t-1}=\rho y_{t-2}+\varepsilon_{t-1}. \]Therefore,
\[ \Delta y_{t-1} = (\rho-1)y_{t-2}+\varepsilon_{t-1}. \]In this stripped-down single-instrument first stage, the population coefficient on the level instrument is
\[ \boxed{\pi=\rho-1}. \]This result resolves an apparent paradox. When \(\rho\) is close to one, \(y_{t-2}\) predicts the next level, \(y_{t-1}\), extremely well. But difference GMM needs it to predict the change, \(\Delta y_{t-1}\). Persistence preserves information about levels while leaving very little predictable movement in the change.
For a near-random walk,
\[ y_{t-1}\approx y_{t-2}+\varepsilon_{t-1} \qquad\Longrightarrow\qquad \Delta y_{t-1}\approx\varepsilon_{t-1}. \]The change is essentially new information. Under the maintained assumption that the innovation is orthogonal to past information, an older level cannot predict it. At the exact random-walk limit, \(\rho=1\), the approximation becomes an equality.
Thus,
\[ \rho=0.95\Rightarrow\pi=-0.05, \qquad \rho=0.99\Rightarrow\pi=-0.01. \]Under stationarity and \(0<\rho<1\), the corresponding correlation is
\[ \operatorname{Corr}(y_{t-2},\Delta y_{t-1}) = -\sqrt{\frac{1-\rho}{2}} \longrightarrow0. \]Technical precision: it is the first-stage slope and correlation that reveal weak relevance. The raw covariance need not converge to zero because \(\operatorname{Var}(y_t)\) changes with \(\rho\). Under stationary normalization,
\[ \operatorname{Cov}(y_{t-2},\Delta y_{t-1}) =-\frac{\sigma_\varepsilon^2}{1+\rho}. \]Dynamic-panel GMM can consequently produce the particularly uncomfortable configuration
\[ \boxed{ K\uparrow \qquad\text{while}\qquad \text{instrument relevance}\downarrow }. \]Dynamic-panel GMM may therefore create many versions of the same weak signal and mistake their number for strength.
System GMM: stronger instruments through stronger assumptions
System GMM adds moment conditions for the equation in levels, such as
\[ E\!\left[\Delta y_{i,t-1}(\alpha_i+\varepsilon_{it})\right]=0. \]These additional moments can improve relevance and finite-sample performance, but they require restrictions on the initial-condition or mean-stationarity process. They are not generated by algebra alone.
This is the central trade-off:
\[ \boxed{ \text{potentially stronger instruments} \quad\Longleftrightarrow\quad \text{stronger identifying assumptions} }. \]Moreover, system GMM does not guarantee strong identification. Bun and Windmeijer (2010) show that the level equation can itself suffer from weak instruments.
9. Why specification tests may not rescue the design
The Hansen statistic is
\[ J = N\overline g_N(\widehat\theta)’ \widehat S^{-1} \overline g_N(\widehat\theta). \]Under correct specification, adequate identification, fixed \(K\) and conventional asymptotics,
\[ J\overset{a}{\sim}\chi^2_{K-p}. \]A rejection tells us that the joint restrictions and model are incompatible with the data. A non-rejection does not prove that the instruments are valid.
With many instruments:
- the endogenous regressors can be overfitted;
- the high-dimensional covariance matrix is difficult to estimate;
- the fixed-\(K\) chi-square approximation can become unreliable;
- the test may have very little power against invalid moments.
Bowsher (2002) showed that too many dynamic-panel moments can make overidentification tests severely undersized and extremely low-powered. In Roodman’s (2009) simulations, an invalid full-instrument specification at \(T=20\) produced an average Hansen \(p\)-value of \(1.000\), whereas sharply reducing the instrument set made the violation readily detectable.
A Hansen \(p\)-value close to one is therefore a warning sign in an instrument-heavy specification, not proof of validity. Difference-in-Hansen tests inherit many of the same limitations.
The Arellano–Bond AR(2) test is also narrower than sometimes believed. Passing it supports a serial-correlation implication needed for particular lag instruments. It does not validate every exclusion restriction or establish that the instruments are strong.
What the Windmeijer correction actually fixes
Conventional two-step GMM standard errors neglect finite-sample variation created by estimating \(\widehat W\) from first-step residuals. Windmeijer (2005) provides a correction for this estimated variance.
\[ \boxed{ \text{Windmeijer correction} \Rightarrow \text{improved two-step standard errors} }. \]It does not:
- debias the coefficient estimate;
- strengthen weak instruments;
- make invalid moments valid;
- reduce instrument proliferation;
- restore the power of the Hansen test.
10. What I would require from a convincing GMM application
The conclusion is not that every GMM application is uninformative. A small set of economically defensible, sufficiently strong moments can be entirely convincing. But the burden of proof belongs to the research design, not to the estimator.
I would want to see:
- An economic justification for each family of moments. Why should the proposed instruments be orthogonal to the structural error?
- A relevance argument. Why should the instruments meaningfully predict the endogenous variables after the chosen transformation?
- A deliberately restricted instrument set. Every available lag should not be included merely because software permits it.
- Sensitivity across lag windows and collapsing choices. Large coefficient movements reveal that the moment set is doing substantial identifying work.
- One-step and corrected two-step results. A major movement in the point estimate can indicate sensitivity to the estimated weighting matrix.
- Difference versus system comparisons. The additional level moments should be defended, not treated as automatically valid.
- Weak-identification-aware analysis. Conventional standard errors and Wald tests can be unreliable when moments are weak.
- Comparison with credible alternatives. Depending on the application, these may include external IV, bias-corrected fixed effects, split-panel jackknife methods, or a more modest non-causal interpretation.
There is no universal instrument-count rule. Keeping \(K<N\) is not a theorem guaranteeing validity or strength. A small instrument set can still be weak or invalid. Instrument reduction is necessary in many applications, but it cannot replace an economic argument.
11. Conclusion: identification first, estimation last
That is why I dislike GMM—or, more accurately, why I distrust what applied economists sometimes ask it to accomplish.
Hansen’s GMM was a breakthrough because it provided a disciplined way to combine credible economic restrictions without specifying a complete likelihood. The difficulty arose when this logic was reversed: because software could generate hundreds of internal instruments, the existence of many moments began to be treated as evidence of strong identification.
But weighting cannot manufacture either exogeneity or relevance. When moments are weak, redundant, highly correlated, numerous or slightly invalid, GMM may multiply sampling noise faster than it accumulates genuine information. More instruments can improve in-sample fit without adding exogenous information, making weak identification appear stronger without actually making it stronger.
A single strong and economically credible instrument may be worth more than a forest of weak internal instruments.
The final message is therefore not that “adding moments is always bad.” Under the ideal fixed-\(K\) theory, valid and independently informative moments can improve efficiency. The correct conclusion is more precise:
\[ \boxed{ \text{More moment conditions mean more identification only when they add} } \] \[ \boxed{ \text{valid, relevant and genuinely independent information.} } \] \[ \boxed{ \text{The solution to endogeneity is credible exogenous variation—not GMM itself.} } \]References and further reading
Foundations and early applications
- Sargan, J. D. (1958). “The Estimation of Economic Relationships Using Instrumental Variables.” Econometrica, 26(3), 393–415. https://doi.org/10.2307/1907619.
- Hansen, L. P. (1982). “Large Sample Properties of Generalized Method of Moments Estimators.” Econometrica, 50(4), 1029–1054. Author’s page and paper.
- Hansen, L. P., and Singleton, K. J. (1982). “Generalized Instrumental Variables Estimation of Nonlinear Rational Expectations Models.” Econometrica, 50(5), 1269–1286. Author’s page and paper.
- Newey, W. K., and West, K. D. (1987). “A Simple, Positive Semi-definite, Heteroskedasticity and Autocorrelation Consistent Covariance Matrix.” Econometrica, 55(3), 703–708. NBER version.
- Wooldridge, J. M. (2001). “Applications of Generalized Method of Moments Estimation.” Journal of Economic Perspectives, 15(4), 87–100. https://doi.org/10.1257/jep.15.4.87.
Dynamic-panel GMM
- Anderson, T. W., and Hsiao, C. (1981). “Estimation of Dynamic Models with Error Components.” Journal of the American Statistical Association, 76(375), 598–606. https://doi.org/10.1080/01621459.1981.10477691.
- Holtz-Eakin, D., Newey, W., and Rosen, H. S. (1988). “Estimating Vector Autoregressions with Panel Data.” Econometrica, 56(6), 1371–1395. NBER version.
- Arellano, M., and Bond, S. (1991). “Some Tests of Specification for Panel Data: Monte Carlo Evidence and an Application to Employment Equations.” Review of Economic Studies, 58(2), 277–297. Journal page.
- Arellano, M., and Bover, O. (1995). “Another Look at the Instrumental Variable Estimation of Error-Components Models.” Journal of Econometrics, 68(1), 29–51. https://doi.org/10.1016/0304-4076(94)01642-D.
- Blundell, R., and Bond, S. (1998). “Initial Conditions and Moment Restrictions in Dynamic Panel Data Models.” Journal of Econometrics, 87(1), 115–143. IFS journal page.
Weak moments, finite samples and instrument proliferation
- Tauchen, G. E. (1986). “Statistical Properties of Generalized Method-of-Moments Estimators of Structural Parameters Obtained from Financial Market Data.” Journal of Business & Economic Statistics, 4(4), 397–416. Duke research page.
- Bekker, P. A. (1994). “Alternative Approximations to the Distributions of Instrumental Variable Estimators.” Econometrica, 62(3), 657–681. https://doi.org/10.2307/2951662.
- Hansen, L. P., Heaton, J., and Yaron, A. (1996). “Finite-Sample Properties of Some Alternative GMM Estimators.” Journal of Business & Economic Statistics, 14(3), 262–280. Author’s page and paper.
- Altonji, J. G., and Segal, L. M. (1996). “Small-Sample Bias in GMM Estimation of Covariance Structures.” Journal of Business & Economic Statistics, 14(3), 353–366. NBER version.
- Ziliak, J. P. (1997). “Efficient Estimation with Panel Data When Instruments Are Predetermined: An Empirical Comparison of Moment-Condition Estimators.” Journal of Business & Economic Statistics, 15(4), 419–431. https://doi.org/10.1080/07350015.1997.10524720.
- Stock, J. H., and Wright, J. H. (2000). “GMM with Weak Identification.” Econometrica, 68(5), 1055–1096. https://doi.org/10.1111/1468-0262.00151.
- Bowsher, C. G. (2002). “On Testing Overidentifying Restrictions in Dynamic Panel Data Models.” Economics Letters, 77(2), 211–220. https://doi.org/10.1016/S0165-1765(02)00130-1.
- Windmeijer, F. (2005). “A Finite Sample Correction for the Variance of Linear Efficient Two-Step GMM Estimators.” Journal of Econometrics, 126(1), 25–51. IFS journal page.
- Roodman, D. (2009). “A Note on the Theme of Too Many Instruments.” Oxford Bulletin of Economics and Statistics, 71(1), 135–158. https://doi.org/10.1111/j.1468-0084.2008.00542.x.
- Newey, W. K., and Windmeijer, F. (2009). “Generalized Method of Moments with Many Weak Moment Conditions.” Econometrica, 77(3), 687–719. https://doi.org/10.3982/ECTA6224.
- Bun, M. J. G., and Windmeijer, F. (2010). “The Weak Instrument Problem of the System GMM Estimator in Dynamic Panel Data Models.” The Econometrics Journal, 13(1), 95–126. https://doi.org/10.1111/j.1368-423X.2009.00299.x.
Technical appendix: Four detailed derivations
This appendix provides the mathematical steps behind four results used in the main text. The objective is to show exactly where each formula comes from and what it means economically.
A.1. From OLS to a moment condition
Begin with the linear model
\[ y_i=\alpha_0+\beta_0x_i+u_i, \]where \(y_i\) is the outcome, \(x_i\) is the explanatory variable, \(\beta_0\) is the true slope, and \(u_i\) contains all the other determinants of \(y_i\).
To simplify the notation, subtract the sample means from \(x_i\) and \(y_i\). Equivalently, we can say that the intercept has already been removed. The model then becomes
\[ y_i=\beta_0x_i+u_i. \]OLS chooses a value of \(\beta\) that minimizes the average squared residual:
\[ SSR(\beta) = \frac{1}{n}\sum_{i=1}^{n}(y_i-\beta x_i)^2. \]The residual associated with observation \(i\) is
\[ e_i(\beta)=y_i-\beta x_i. \]To find the minimum, differentiate the objective function with respect to \(\beta\). For one observation, the chain rule gives
\[ \frac{\partial}{\partial\beta}(y_i-\beta x_i)^2 = 2(y_i-\beta x_i)(-x_i). \]Therefore,
\[ \frac{\partial SSR(\beta)}{\partial\beta} = -\frac{2}{n} \sum_{i=1}^{n} x_i(y_i-\beta x_i). \]At an interior minimum, this derivative must equal zero:
\[ -\frac{2}{n} \sum_{i=1}^{n} x_i(y_i-\widehat\beta_{\mathrm{OLS}}x_i) = 0. \]Dividing by \(-2\) gives
\[ \frac{1}{n} \sum_{i=1}^{n} x_i(y_i-\widehat\beta_{\mathrm{OLS}}x_i) = 0. \]This is the OLS sample moment condition. It says that the regressor is orthogonal to the estimated residual:
\[ \overline{x\widehat u}=0. \]Expanding the moment condition gives
\[ \frac{1}{n}\sum_{i=1}^{n}x_iy_i – \widehat\beta_{\mathrm{OLS}} \frac{1}{n}\sum_{i=1}^{n}x_i^2 = 0. \]Consequently,
\[ \widehat\beta_{\mathrm{OLS}} = \frac{\frac{1}{n}\sum_{i=1}^{n}x_iy_i} {\frac{1}{n}\sum_{i=1}^{n}x_i^2} = \frac{\sum_{i=1}^{n}x_iy_i} {\sum_{i=1}^{n}x_i^2}. \]Now substitute the true structural equation
\[ y_i=\beta_0x_i+u_i \]into the OLS formula:
\[ \widehat\beta_{\mathrm{OLS}} = \frac{\sum_i x_i(\beta_0x_i+u_i)} {\sum_i x_i^2}. \]Distributing \(x_i\) in the numerator gives
\[ \widehat\beta_{\mathrm{OLS}} = \frac{ \beta_0\sum_i x_i^2+\sum_i x_iu_i }{ \sum_i x_i^2 }. \]Separating the two terms produces
\[ \widehat\beta_{\mathrm{OLS}} = \beta_0 + \frac{\sum_i x_iu_i} {\sum_i x_i^2}. \]Therefore, the OLS estimation error is
\[ \boxed{ \widehat\beta_{\mathrm{OLS}}-\beta_0 = \frac{\frac{1}{n}\sum_i x_iu_i} {\frac{1}{n}\sum_i x_i^2} }. \]The numerator measures the sample relationship between the regressor and the structural error. The denominator measures the amount of variation in the regressor.
Under the law of large numbers, and provided the relevant expectations exist,
\[ \frac{1}{n}\sum_i x_iu_i \overset{p}{\longrightarrow} E[x_iu_i] \]and
\[ \frac{1}{n}\sum_i x_i^2 \overset{p}{\longrightarrow} E[x_i^2]. \]Hence,
\[ \operatorname{plim} \widehat\beta_{\mathrm{OLS}} = \beta_0 + \frac{E[x_iu_i]}{E[x_i^2]}. \]Thus, if
\[ E[x_iu_i]=0 \]and
\[ E[x_i^2]>0, \]then
\[ \operatorname{plim} \widehat\beta_{\mathrm{OLS}} = \beta_0. \]OLS can therefore be viewed as a method-of-moments estimator based on the population condition
\[ E[x_i(y_i-\beta_0x_i)]=0. \]The sample estimator is obtained by replacing the population expectation with its sample analogue and setting it equal to zero.
A.2. Derivation and interpretation of the IV error formula
Suppose that \(x_i\) is endogenous:
\[ E[x_iu_i]\neq 0. \]OLS is then inconsistent because the regressor is systematically related to the structural error.
Assume that we have an instrument \(z_i\). A valid instrument must satisfy the population moment condition
\[ E[z_iu_i]=0. \]Because
\[ u_i=y_i-\beta_0x_i, \]the IV moment condition can also be written as
\[ E[z_i(y_i-\beta_0x_i)]=0. \]The sample analogue is
\[ \frac{1}{n} \sum_{i=1}^{n} z_i(y_i-\widehat\beta_{\mathrm{IV}}x_i) = 0. \]Expanding this expression gives
\[ \frac{1}{n}\sum_i z_iy_i – \widehat\beta_{\mathrm{IV}} \frac{1}{n}\sum_i z_ix_i = 0. \]Solving for the IV estimator produces
\[ \widehat\beta_{\mathrm{IV}} = \frac{ \frac{1}{n}\sum_i z_iy_i }{ \frac{1}{n}\sum_i z_ix_i }. \]Now substitute the structural equation
\[ y_i=\beta_0x_i+u_i. \]We obtain
\[ \widehat\beta_{\mathrm{IV}} = \frac{ \frac{1}{n}\sum_i z_i(\beta_0x_i+u_i) }{ \frac{1}{n}\sum_i z_ix_i }. \]Expanding the numerator gives
\[ \widehat\beta_{\mathrm{IV}} = \frac{ \beta_0\frac{1}{n}\sum_i z_ix_i + \frac{1}{n}\sum_i z_iu_i }{ \frac{1}{n}\sum_i z_ix_i }. \]Separating the two terms yields
\[ \widehat\beta_{\mathrm{IV}} = \beta_0 + \frac{ \frac{1}{n}\sum_i z_iu_i }{ \frac{1}{n}\sum_i z_ix_i }. \]Consequently,
\[ \boxed{ \widehat\beta_{\mathrm{IV}}-\beta_0 = \frac{\overline{zu}} {\overline{zx}} = \frac{ \frac{1}{n}\sum_i z_iu_i }{ \frac{1}{n}\sum_i z_ix_i } }. \]The numerator and denominator have distinct economic meanings.
The numerator
\[ \overline{zu} = \frac{1}{n}\sum_i z_iu_i \]measures the sample relationship between the instrument and the structural error. Instrument validity requires
\[ E[z_iu_i]=0. \]This does not mean that \(\overline{zu}\) will equal zero exactly in every finite sample. Even a perfectly valid instrument normally has some accidental sample correlation with the error.
The denominator
\[ \overline{zx} = \frac{1}{n}\sum_i z_ix_i \]measures instrument relevance. It tells us how strongly the instrument is related to the endogenous regressor.
If the instrument is strong, ordinary sampling noise in the numerator is divided by a reasonably large number. If the instrument is weak, the same noise is divided by a very small number and can generate a very large estimation error.
For example, suppose that
\[ \overline{zu}=0.01. \]If
\[ \overline{zx}=0.50, \]then
\[ \widehat\beta_{\mathrm{IV}}-\beta_0 = \frac{0.01}{0.50} = 0.02. \]But if
\[ \overline{zx}=0.05, \]then
\[ \widehat\beta_{\mathrm{IV}}-\beta_0 = \frac{0.01}{0.05} = 0.20. \]Finally, if
\[ \overline{zx}=0.01, \]then
\[ \widehat\beta_{\mathrm{IV}}-\beta_0 = \frac{0.01}{0.01} = 1. \]The sample relationship between the instrument and the error has not changed. Only instrument relevance has changed. Yet the estimation error has increased from \(0.02\) to \(1\).
Exogeneity and relevance therefore perform different functions:
\[ \text{exogeneity controls the numerator,} \] \[ \text{relevance prevents the denominator from becoming too small.} \]An intuitive analogy is that a valid instrument is an honest witness, while a relevant instrument is a witness who actually observed the event. An honest witness who saw almost nothing cannot provide much information.
Thus:
\[ \boxed{ \text{Exogeneity makes an instrument honest; relevance makes it informative.} } \]A.3. Why GMM is a weighted average of instrument-specific estimates
Suppose that there is one parameter \(\beta\) and \(K\) instruments. Instrument \(j\) generates the sample moment
\[ \overline g_j(\beta) = \frac{1}{n} \sum_{i=1}^{n} z_{ji}(y_i-\beta x_i). \]Define
\[ b_j = \frac{1}{n}\sum_i z_{ji}y_i \]and
\[ a_j = \frac{1}{n}\sum_i z_{ji}x_i. \]The moment can then be written as
\[ \overline g_j(\beta) = b_j-a_j\beta. \]If instrument \(j\) were used alone, its just-identified IV estimate would solve
\[ b_j-a_j\widehat\beta_j=0. \]Therefore, provided \(a_j\neq0\),
\[ \widehat\beta_j = \frac{b_j}{a_j} = \frac{\frac{1}{n}\sum_i z_{ji}y_i} {\frac{1}{n}\sum_i z_{ji}x_i}. \]Because \(b_j=a_j\widehat\beta_j\), the corresponding moment discrepancy is
\[ \overline g_j(\beta) = a_j(\widehat\beta_j-\beta). \]This equation has an important interpretation. Each moment condition has its own preferred estimate, \(\widehat\beta_j\). The moment is zero when \(\beta=\widehat\beta_j\).
Suppose that the GMM weighting matrix is diagonal:
\[ W = \operatorname{diag}(w_1,\ldots,w_K). \]The GMM criterion is
\[ Q_n(\beta) = \sum_{j=1}^{K} w_j\overline g_j(\beta)^2. \]Substituting
\[ \overline g_j(\beta) = a_j(\widehat\beta_j-\beta) \]gives
\[ Q_n(\beta) = \sum_{j=1}^{K} w_ja_j^2 (\widehat\beta_j-\beta)^2. \]Differentiate with respect to \(\beta\):
\[ \frac{\partial Q_n(\beta)}{\partial\beta} = 2\sum_{j=1}^{K} w_ja_j^2 (\beta-\widehat\beta_j). \]At the minimum,
\[ \sum_{j=1}^{K} w_ja_j^2 (\widehat\beta_{\mathrm{GMM}}-\widehat\beta_j) = 0. \]Rearranging,
\[ \widehat\beta_{\mathrm{GMM}} \sum_{j=1}^{K}w_ja_j^2 = \sum_{j=1}^{K} w_ja_j^2\widehat\beta_j. \]Therefore,
\[ \boxed{ \widehat\beta_{\mathrm{GMM}} = \frac{ \sum_{j=1}^{K} w_ja_j^2\widehat\beta_j }{ \sum_{j=1}^{K} w_ja_j^2 } }. \]Define the effective weight attached to instrument \(j\) as
\[ \lambda_j = \frac{w_ja_j^2} {\sum_{\ell=1}^{K}w_\ell a_\ell^2}. \]These effective weights satisfy
\[ \lambda_j\geq0 \]and
\[ \sum_{j=1}^{K}\lambda_j=1. \]Thus,
\[ \widehat\beta_{\mathrm{GMM}} = \sum_{j=1}^{K} \lambda_j\widehat\beta_j. \]With a diagonal weighting matrix, GMM is therefore a weighted average of the estimates preferred by the individual instruments.
The effective importance of an instrument depends on two elements:
\[ \text{effective importance} = \text{GMM weight} \times \text{squared sample relevance}. \]This follows from the term \(w_ja_j^2\). A weak instrument has a small \(a_j\), so its direct contribution is reduced. But a collection of numerous weak and correlated instruments can still create serious finite-sample problems through overfitting and unstable estimation of the weighting matrix.
A.3.1. A two-moment illustration
Suppose that one moment condition is minimized at
\[ \beta=1, \]while another is minimized at
\[ \beta=3. \]For simplicity, absorb the relevance terms into the weights. The GMM criterion becomes
\[ Q(\beta) = w_1(\beta-1)^2 + w_2(\beta-3)^2. \]The first term penalizes distance from \(1\). The second penalizes distance from \(3\).
Differentiate the criterion:
\[ Q'(\beta) = 2w_1(\beta-1) + 2w_2(\beta-3). \]At the minimum,
\[ 2w_1(\widehat\beta-1) + 2w_2(\widehat\beta-3) = 0. \]Dividing by \(2\) and expanding gives
\[ w_1\widehat\beta-w_1 + w_2\widehat\beta-3w_2 = 0. \]Collecting the terms containing \(\widehat\beta\),
\[ (w_1+w_2)\widehat\beta = w_1+3w_2. \]Hence,
\[ \boxed{ \widehat\beta = \frac{w_1+3w_2}{w_1+w_2} }. \]If the first moment receives nine times as much weight as the second,
\[ (w_1,w_2)=(9,1), \]then
\[ \widehat\beta = \frac{9+3}{10} = 1.2. \]If the weights are reversed,
\[ (w_1,w_2)=(1,9), \]then
\[ \widehat\beta = \frac{1+27}{10} = 2.8. \]The result can also be understood as a balance of opposing forces. At the minimum,
\[ w_1(\widehat\beta-1) = w_2(3-\widehat\beta). \]With \((w_1,w_2)=(9,1)\),
\[ 9(1.2-1) = 1(3-1.2) = 1.8. \]The first moment is only \(0.2\) away from its preferred value, but its weight is nine. The second is \(1.8\) away from its preferred value, but its weight is only one. The weighted pressures are equal.
This example reveals what GMM does when the sample moments disagree:
\[ \boxed{ \text{GMM does not determine which moment is correct; it chooses a weighted compromise.} } \]If all moments are valid and strongly identifying, their preferred estimates should converge to the same true parameter as the sample grows. Their finite-sample disagreement is then ordinary sampling noise.
If some moments are invalid, however, their disagreement can persist even in very large samples. In that case, changing the weighting matrix changes the population value toward which GMM converges.
With a non-diagonal weighting matrix, the estimator need not be a simple convex average. Correlations among moments can generate complicated implicit weights, which may even be negative.
A.4. Deriving the GMM sandwich variance and efficient weighting matrix
Let \(\theta_0\) be a \(p\times1\) vector of true parameters and let \(g_i(\theta)\) be a \(K\times1\) vector of moment conditions satisfying
\[ E[g_i(\theta_0)]=0. \]The sample average of the moments is
\[ \overline g_n(\theta) = \frac{1}{n} \sum_{i=1}^{n} g_i(\theta). \]GMM chooses \(\widehat\theta\) to minimize
\[ Q_n(\theta) = \overline g_n(\theta)’ W \overline g_n(\theta), \]where \(W\) is a symmetric positive-definite \(K\times K\) weighting matrix.
A.4.1. The meanings of \(S\) and \(G\)
Under independent sampling, define
\[ S = \operatorname{Var}[g_i(\theta_0)]. \]Because the population moments are zero,
\[ E[g_i(\theta_0)]=0, \]we have
\[ S = E[g_i(\theta_0)g_i(\theta_0)’]. \]The diagonal elements of \(S\) measure the variance of each moment. The off-diagonal elements measure how the moments move together.
With serially dependent observations, \(S\) is instead the long-run covariance matrix:
\[ S = \sum_{h=-\infty}^{\infty} \operatorname{Cov} \left( g_t(\theta_0), g_{t-h}(\theta_0) \right). \]Now define the \(K\times p\) Jacobian matrix
\[ G = E\left[ \frac{\partial g_i(\theta_0)} {\partial\theta’} \right]. \]The matrix \(G\) measures how strongly the moments respond when the parameters change.
For one parameter and one moment, \(G\) is simply a slope. If \(G\) is close to zero, changing the parameter barely changes the moment condition. The moment then contains little identifying information.
With several parameters, identification requires \(G\) to have full column rank. If its columns are nearly linearly dependent, some combinations of the parameters are only weakly identified.
A.4.2. Linearizing the GMM estimator
Differentiate the GMM criterion. Treating the final-step weighting matrix as fixed asymptotically, the first-order condition is
\[ 2\widehat G_n(\widehat\theta)’ W \overline g_n(\widehat\theta) = 0, \]where
\[ \widehat G_n(\theta) = \frac{\partial\overline g_n(\theta)} {\partial\theta’}. \]For a large sample, \(\widehat G_n(\widehat\theta)\) is close to its population value \(G\). The first-order condition can therefore be approximated by
\[ G’W\overline g_n(\widehat\theta) \approx 0. \]Next, use a first-order Taylor expansion of the sample moments around the true parameter:
\[ \overline g_n(\widehat\theta) \approx \overline g_n(\theta_0) + G(\widehat\theta-\theta_0). \]Substituting this expansion into the first-order condition gives
\[ G’W \left[ \overline g_n(\theta_0) + G(\widehat\theta-\theta_0) \right] \approx 0. \]Expanding,
\[ G’W\overline g_n(\theta_0) + G’WG(\widehat\theta-\theta_0) \approx 0. \]Move the first term to the other side:
\[ G’WG(\widehat\theta-\theta_0) \approx – G’W\overline g_n(\theta_0). \]Provided \(G’WG\) is invertible,
\[ \widehat\theta-\theta_0 \approx – (G’WG)^{-1} G’W \overline g_n(\theta_0). \]Multiplying by \(\sqrt n\) gives the fundamental GMM approximation:
\[ \boxed{ \sqrt n(\widehat\theta-\theta_0) \approx – (G’WG)^{-1} G’W \sqrt n\,\overline g_n(\theta_0) }. \]This equation says that estimation error is transformed moment noise. The sample moments fluctuate around zero, and the matrix
\[ (G’WG)^{-1}G’W \]translates those fluctuations into parameter fluctuations.
If \(G’WG\) has a very small eigenvalue, its inverse contains a very large eigenvalue. Small moment disturbances can then generate large changes in the estimated parameters. This is the multivariate form of the weak-denominator problem.
A.4.3. Obtaining the sandwich variance
Under a central limit theorem,
\[ \sqrt n\,\overline g_n(\theta_0) \overset{d}{\longrightarrow} N(0,S). \]Recall the elementary covariance rule
\[ \operatorname{Var}(Ax) = A\operatorname{Var}(x)A’. \]Define
\[ A = – (G’WG)^{-1}G’W. \]Then
\[ \sqrt n(\widehat\theta-\theta_0) \approx A\sqrt n\,\overline g_n(\theta_0). \]Its asymptotic variance is therefore
\[ \operatorname{Avar} \left[ \sqrt n(\widehat\theta-\theta_0) \right] = ASA’. \]Because \(W\) is symmetric,
\[ A’ = – WG(G’WG)^{-1}. \]Substituting \(A\), \(S\), and \(A’\) gives
\[ \boxed{ \operatorname{Avar} \left[ \sqrt n(\widehat\theta-\theta_0) \right] = (G’WG)^{-1} G’WSWG (G’WG)^{-1} }. \]This is called a sandwich variance because the matrix \(G’WSWG\) is placed between two copies of \((G’WG)^{-1}\).
Since this expression is the variance of \(\sqrt n(\widehat\theta-\theta_0)\), the approximate variance of the estimator itself is
\[ \operatorname{Var}(\widehat\theta) \approx \frac{1}{n} (G’WG)^{-1} G’WSWG (G’WG)^{-1}. \]A.4.4. Why \(W=S^{-1}\) is efficient
The matrix \(S\) describes the noise and correlation structure of the moment conditions. The efficient GMM weighting matrix is
\[ W=S^{-1}. \]This choice gives less influence to noisy combinations of moments and accounts for correlations between them.
One way to understand this result is to define standardized moments
\[ h_i(\theta) = S^{-1/2}g_i(\theta). \]The covariance matrix of these transformed moments is
\[ \operatorname{Var}[h_i(\theta_0)] = S^{-1/2}SS^{-1/2} = I. \]Thus, the transformation removes differences in scale and correlation. The efficient GMM criterion can be written as
\[ \overline g_n(\theta)’S^{-1}\overline g_n(\theta) = \overline h_n(\theta)’\overline h_n(\theta). \]Efficient GMM therefore minimizes moment discrepancies after measuring them in standardized, correlation-adjusted units.
Substitute \(W=S^{-1}\) into the general sandwich formula:
\[ \operatorname{Avar} \left[ \sqrt n(\widehat\theta-\theta_0) \right] = (G’S^{-1}G)^{-1} G’S^{-1}SS^{-1}G (G’S^{-1}G)^{-1}. \]Because
\[ S^{-1}SS^{-1}=S^{-1}, \]the middle part becomes
\[ G’S^{-1}G. \]Consequently,
\[ \operatorname{Avar} \left[ \sqrt n(\widehat\theta-\theta_0) \right] = (G’S^{-1}G)^{-1} (G’S^{-1}G) (G’S^{-1}G)^{-1}. \]Using the identity
\[ A^{-1}AA^{-1}=A^{-1}, \]we obtain
\[ \boxed{ \operatorname{Avar} \left[ \sqrt n(\widehat\theta-\theta_0) \right] = (G’S^{-1}G)^{-1} }. \]The matrix
\[ G’S^{-1}G \]can be interpreted as the identifying information contained in the moments after accounting for their noise and correlation. The asymptotic variance is its inverse.
A.4.5. Connection with the weak-instrument formula
Consider one parameter and one IV moment:
\[ g_i(\beta) = z_i(y_i-x_i\beta). \]At the true parameter,
\[ g_i(\beta_0)=z_iu_i. \]Therefore,
\[ S = E[z_i^2u_i^2]. \]The derivative of the moment with respect to \(\beta\) is
\[ \frac{\partial g_i(\beta)}{\partial\beta} = -z_ix_i, \]so
\[ G=-E[z_ix_i]. \]The efficient asymptotic variance becomes
\[ \operatorname{Avar} \left[ \sqrt n(\widehat\beta-\beta_0) \right] = \frac{S}{G^2}. \]Substituting \(S\) and \(G\),
\[ \boxed{ \operatorname{Avar} \left[ \sqrt n(\widehat\beta-\beta_0) \right] = \frac{ E[z_i^2u_i^2] }{ \left(E[z_ix_i]\right)^2 } }. \]The numerator measures moment noise. The denominator measures instrument relevance, squared.
If
\[ E[z_ix_i]\approx0, \]the denominator is very small and the variance becomes very large. No weighting scheme can eliminate this basic lack of identifying information.
This is the central connection between weak IV and GMM:
\[ \boxed{ \text{GMM can combine identifying information efficiently, but it cannot create identifying information.} } \]The efficiency result for \(W=S^{-1}\) is therefore conditional on valid moment conditions, sufficiently strong identification, a fixed and manageable number of moments, a reliably estimated covariance matrix, and a sufficiently large sample. When there are numerous weak and highly correlated moments, these conditions can provide a poor description of the finite sample.