Joint Inference in LP: Wald Tests and Simultaneous Confidence Bands

Joint Inference in LP: Wald Tests and Simultaneous Confidence Bands

A response curve can have a 95% simultaneous confidence band that includes zero at every horizon and still reject an all-zero path in a joint Wald test. This is precisely the pattern in additional calculations for the diplomatic-tension response of Brent oil prices in my work with Will Ginn. The apparent discrepancy is a useful starting point for understanding joint inference.

An impulse response is a collection of estimates. We may want to know whether a particular coefficient differs from zero, whether there is any response over the entire horizon, or where we can identify a positive or negative response after examining the whole curve. These questions require care. Wald tests and simultaneous bands offer complementary ways to evaluate a response path. When used to test an all-zero path, they can address the same null hypothesis with different rejection rules.

When the main question is whether the response is zero at every horizon in a prespecified range, I recommend using a valid full-path Wald test as the primary test of that hypothesis. Simultaneous bands help identify the months in which the response is statistically different from zero. Testing many months creates more opportunities to find an apparent effect by chance. A 95% simultaneous band adjusts for this multiplicity: under the maintained assumptions, it aims to limit the probability of making even one false declaration of a nonzero response across the selected months to 5%.

1. From one horizon to an entire response

Suppose we estimate the response of an outcome from impact through month 48. The estimated path contains 49 coefficients:

\[ \widehat{\boldsymbol\beta} = (\widehat\beta_0,\widehat\beta_1,\ldots,\widehat\beta_{48})’. \]

At one prespecified horizon, a conventional 95% pointwise interval is approximately

\[ \widehat\beta_h\pm1.96\,\widehat s_h, \]

where \(\widehat s_h\) is the coefficient’s standard error. Its coverage statement concerns that horizon alone.

Reading the entire chart introduces a different problem. Even when the true response is zero at every horizon, sampling variation may place some estimates far from zero. As a deliberately simplified illustration, if 49 tests were independent and each had a 5% rejection probability under the null, the probability of at least one rejection would be

\[ 1-0.95^{49}\approx 0.919. \]

This is not the rejection probability for our local projections: their estimates are correlated across horizons. It illustrates why 49 separate 95% statements cannot simply be read as one 95% statement about the curve.

2. The Wald test

The full-path Wald test starts with an explicit hypothesis:

\[ H_0:\beta_0=\beta_1=\cdots=\beta_{48}=0. \]

Let \(\widehat V\) be the estimated covariance matrix of the coefficient estimates, including all cross-horizon covariances. The statistic is

\[ W=\widehat{\boldsymbol\beta}’\widehat V^{-1} \widehat{\boldsymbol\beta}. \]

The statistic measures the distance of the estimated vector from zero, taking account of uncertainty and correlation. If two estimates tend to move together because they share sampling noise, the test takes that common movement into account. It also evaluates differences between estimates relative to the precision with which those differences are measured.

Under the null, joint asymptotic normality and a consistent, nonsingular covariance estimator give the usual approximation

\[ W\overset{a}{\sim}\chi^2_{49},\qquad p_{\mathrm{Wald}}=1-F_{\chi^2_{49}}(W). \]

The 49 degrees of freedom count the restrictions. They are not the residual degrees of freedom of an individual regression. This is conventional Wald inference within the large-sample framework used for moment estimators; Hansen (1982) provides the GMM foundation.

A rejection means that the data provide evidence against an entirely zero response. It does not identify every significant horizon or establish that the response has the same sign throughout.

3. Simultaneous sup-t bands

A simultaneous confidence band aims to cover the entire coefficient vector at a stated confidence level. Instead of using 1.96 at every horizon, it chooses a common cutoff \(c_{0.95}\) calibrated to the largest absolute standardized estimation error:

\[ \mathcal B_h= \left[\widehat\beta_h-c_{0.95}\widehat s_h, \widehat\beta_h+c_{0.95}\widehat s_h\right]. \]

One convenient construction uses correlated Gaussian draws with the estimated covariance. For each draw \(b\), compute

\[ \boldsymbol g^{(b)}\sim N(\boldsymbol0,\widehat V),\qquad M^{(b)}=\max_h\left|\frac{g_h^{(b)}}{\widehat s_h}\right|. \]

The empirical 95th percentile of the simulated maxima supplies \(c_{0.95}\). This is the Gaussian plug-in sup-t approach discussed by Montiel Olea and Plagborg-Møller (2019) and reviewed by Jordà and Taylor (2025). Its coverage is asymptotic and depends on a valid distributional and covariance approximation.

The band can also test the all-zero null. Reject if zero falls outside the band at any horizon. Equivalently, reject when

\[ M_{\mathrm{obs}}=\max_h\left|\frac{\widehat\beta_h}{\widehat s_h}\right| >c_{0.95}. \]

Both methods use cross-horizon dependence. It enters the Wald statistic through the inverse covariance matrix and the sup-t cutoff through the joint distribution of the simulated errors.

4. Why the conclusions can differ

With the same covariance matrix, the Wald test and the sup-t test can still disagree. Their acceptance regions have different shapes: the Wald region is an ellipsoid, while the sup-t region is a rectangle. Neither shape generally contains the other.

Consider a hypothetical two-horizon example. Both standard errors equal one, and the estimation errors have correlation 0.8:

\[ V=\begin{pmatrix}1&0.8\\0.8&1\end{pmatrix},\qquad W=\frac{\widehat\beta_1^2-1.6\widehat\beta_1\widehat\beta_2+ \widehat\beta_2^2}{0.36}. \]

Assume a Gaussian sampling distribution with known covariance for this illustration. At 5%, the Wald cutoff is approximately 5.9915 and the sup-t cutoff is approximately 2.1524. Both tests then have the same nominal size, but they reject different vectors.

Wald and sup-t tests have different acceptance regionsFor two normal estimates with standard errors one and correlation 0.8, the Wald ellipse and sup-t square are both centered at the zero null. Point A at 1.8, 0.3 is outside the ellipse but inside the square. Point B at 2.2, 2.2 is inside the ellipse but outside the square. 2026-09-22T18:39:33.811502 image/svg+xml Matplotlib v3.10.8, https://matplotlib.org/ −3 −2 −1 0 1 2 3 Estimate at the first horizon / standard error −3 −2 −1 0 1 2 3 Estimate at the second horizon / standard error A (1.8, 0.3) Wald rejects only B (2.2, 2.2) Sup-t rejects only Wald acceptance region Sup-t acceptance region
Figure 1. Two 95% acceptance regions under the all-zero null. The blue ellipse belongs to the Wald test and the dashed gold square to the sup-t test. A point outside a region rejects under that procedure. These are hypothetical standardized estimates, not Brent results; the covariance is known, with correlation 0.8.
Two examples using exactly the same covariance matrix
CaseEstimatesWald statisticWald p-valueMaximum |t|Decision at 5%
A(1.8, 0.3)6.8500.03251.800Wald rejects; sup-t does not
B(2.2, 2.2)5.3780.06802.200Sup-t rejects; Wald does not

In case A, neither individual estimate reaches the pointwise 1.96 threshold. Yet the estimates differ substantially. With strongly positively correlated errors, such a gap is relatively unusual under the all-zero null. The Wald test detects it.

In case B, both estimates exceed the simultaneous cutoff. But they move together along a direction with relatively high sampling variability, so the Wald statistic remains below its threshold.

The example explains why a failure to reject with one procedure does not imply failure with the other. It also shows why “Wald is more powerful” is too broad a claim. The relative performance depends on the alternative and the covariance structure, consistent with the discussion in Montiel Olea and Plagborg-Møller (2019). Simulation error is not needed to produce the disagreement.

5. A simple Python implementation

The following function implements both procedures using NumPy and SciPy. The inputs are an estimated coefficient vector, beta, and its full covariance matrix, V. Importantly, V means \(\widehat{\operatorname{Var}}(\widehat{\boldsymbol\beta})\). If software instead reports the asymptotic covariance of \(\sqrt n(\widehat{\boldsymbol\beta}-\boldsymbol\beta)\), divide that matrix by \(n\) first.

Install the two packages if needed:

python -m pip install numpy scipy
import numpy as np
from scipy.stats import chi2


def joint_inference(beta, V, alpha=0.05, draws=50_000, seed=2026):
    """Test the all-zero path and construct simultaneous confidence bands."""
    beta = np.asarray(beta, dtype=float)
    V = np.asarray(V, dtype=float)
    k = beta.size
    if beta.ndim != 1 or k == 0 or V.shape != (k, k):
        raise ValueError("Supply a coefficient vector and matching covariance.")
    if not (np.isfinite(beta).all() and np.isfinite(V).all()):
        raise ValueError("Inputs must be finite.")
    if not 0 < alpha < 1 or not isinstance(draws, int) or draws < 1:
        raise ValueError("Require 0 < alpha < 1 and a positive integer draws.")
    if not np.allclose(V, V.T):
        raise ValueError("The covariance matrix must be symmetric.")
    V = (V + V.T) / 2  # Remove numerical asymmetry only.
    L = np.linalg.cholesky(V)  # Requires positive definite covariance.
    se = np.sqrt(np.diag(V))

    # 1. Wald test of beta_0 = ... = beta_H = 0.
    W = float(beta @ np.linalg.solve(V, beta))
    p_wald = chi2.sf(W, df=k)

    # 2. Correlated Gaussian errors for the sup-t critical value.
    rng = np.random.default_rng(seed)
    errors = rng.standard_normal((draws, k)) @ L.T
    simulated_max = np.max(np.abs(errors / se), axis=1)
    cutoff = np.quantile(simulated_max, 1 - alpha)
    observed_max = np.max(np.abs(beta / se))
    p_sup = (1 + np.sum(simulated_max >= observed_max)) / (draws + 1)

    return dict(W=W, df=k, p_wald=p_wald, p_sup=p_sup,
                max_t=observed_max, cutoff=cutoff,
                lower=beta - cutoff * se, upper=beta + cutoff * se)

The Cholesky factor turns independent standard-normal draws into correlated errors with covariance V. The same covariance is used for the Wald statistic and the simultaneous bands. The function requires a positive definite matrix; it deliberately stops if that condition fails. A pseudoinverse with unchanged degrees of freedom is not a general solution to a singular covariance problem.

We can reproduce the two hypothetical examples as follows:

V = np.array([[1.0, 0.8], [0.8, 1.0]])
for name, beta in {"A": [1.8, 0.3], "B": [2.2, 2.2]}.items():
    r = joint_inference(beta, V)
    excludes_zero = (r["lower"] > 0) | (r["upper"] < 0)
    print(f"{name}: W={r['W']:.3f}, Wald p={r['p_wald']:.4f}, "
          f"sup-t p={r['p_sup']:.4f}, cutoff={r['cutoff']:.3f}")
    print("   Band excludes zero:", excludes_zero.tolist())

With the displayed seed and 50,000 draws, the output is approximately:

A: W=6.850, Wald p=0.0325, sup-t p=0.1097, cutoff=2.148
   Band excludes zero: [False, False]
B: W=5.378, Wald p=0.0680, sup-t p=0.0434, cutoff=2.148
   Band excludes zero: [True, True]

The simulated cutoff is about 2.148, close to the 2.1524 value obtained by numerical integration for Figure 1. Changing the seed changes the simulation output slightly; increasing the number of draws reduces this Monte Carlo variability. Neither choice resolves an incorrectly estimated covariance matrix. The simulated sup-t p-value uses a plus-one correction and remains an approximation; that correction does not make inference with an estimated covariance exact in finite samples.

For a 49-horizon application, supply 49 coefficients and their 49-by-49 covariance matrix, in the same horizon order. Separate standard errors are insufficient: they provide the diagonal but omit the cross-horizon covariances. Setting those off-diagonal entries to zero would impose an independence assumption.

To obtain 90% bands, set alpha=0.10. The returned lower and upper arrays can be plotted around the estimated response. A horizon excludes zero when its lower endpoint is positive or its upper endpoint is negative.

6. An application to geopolitical turning points

Additional calculations prepared for a future revision of my working paper with Will Ginn on geopolitical turning points and oil prices illustrate the distinction. The joint IV local projections include military conflict, diplomatic tensions, and nuclear threats. The following results concern real Brent prices and use the calendar-aligned cross-horizon HC3 covariance from the revision calculations.

Joint-zero Wald tests for real Brent response paths
CategoryHorizonsWald statisticDegrees of freedomp-value
Diplomatic tensions0, 6, 12, 24, 36, 4814.78760.0220
Diplomatic tensionsEvery month from 0 to 4870.042490.0259
Nuclear threats0, 6, 12, 24, 36, 4812.41160.0534
Nuclear threatsEvery month from 0 to 4876.172490.0077

Source: Brent joint-inference revision calculations, September 21, 2026. Each row tests one category separately. Probability values use the asymptotic chi-square distribution, conditional on the maintained IV and covariance assumptions. No correction across categories or alternative tests is applied. The 49-horizon tests are additional revision diagnostics.

The selected six-horizon nuclear test omits months 18–22, where the delayed positive response is especially clear. Testing all 49 plotted horizons gives a p-value of 0.0077. This change is not a general rule that more horizons imply greater significance: the statistic and the number of restrictions both change. The complete horizon set should be motivated by the economic question and reported transparently.

For diplomatic tensions, the independent-monthly 95% simultaneous band includes zero throughout, whereas the full-path Wald test rejects at 5%. The geometry above explains how such findings can coexist. At 90%, the diplomatic band excludes zero negatively at months 23–26 under all three multiplier schemes. At 95%, the six- and twelve-month block schemes exclude zero negatively at months 22–28. For nuclear threats, all three 95% schemes exclude zero positively at months 18–22.

The empirical bands use Rademacher influence multipliers, with independent monthly weights or weights shared within six- and twelve-month blocks. The Python example illustrates a Gaussian plug-in implementation; it is not an exact replication of those multiplier calculations. The block schemes change the treatment of temporal dependence. This is separate from finite simulation variability and from the distinction between Wald and sup-t statistics.

7. What to report and what to check

Jordà and Taylor (2025) offer a useful applied framework. Their review discusses system-based joint inference and graphical uncertainty measures; Figure 6 of its August 2024 working-paper version reports a 49-restriction LP-IV joint-zero test alongside pointwise intervals. This provides a precedent for reporting a formal path test with the response plot. The article appeared in the Journal of Economic Literature.

For a study whose central question concerns the existence of a response somewhere along the declared path, I recommend giving the full-path Wald test priority in the reporting of statistical evidence. Report the tested horizon set, the covariance estimator, the Wald statistic, its degrees of freedom, and its p-value. Simultaneous bands provide complementary evidence about the timing and sign of individual responses. Pointwise intervals remain useful for showing marginal precision.

Exclusion of zero by a simultaneous band should not be imposed as an additional requirement for rejecting the all-zero path with a valid Wald test. A band containing zero at every horizon means that the sup-t procedure does not reject that null at the chosen level; it does not invalidate the Wald rejection. Both results should be reported with their respective interpretations. This priority is a recommendation about how to answer the economic question, not a claim that Wald tests are uniformly more powerful.

The primary hypothesis, horizon range, and testing procedure should be chosen before inspecting the results. The 49-horizon Brent tests reported above remain additional revision diagnostics; the recommendation here does not imply that they were prespecified in the original analysis.

The covariance estimator is central to both procedures. An HC3 adjustment addresses leverage within a heteroskedasticity-consistent calculation, following MacKinnon and White (1985). It does not by itself account for arbitrary serial correlation in IV scores. Validity may require a long-run covariance estimator or an appropriately justified resampling procedure. Lag augmentation needs a justification for the particular model and inferential target. Conventional Wald inference also requires sufficiently strong identification; it is not automatically robust to weak instruments.

The recommended reporting order is therefore explicit: use the full-path Wald test as the primary assessment of the all-zero response hypothesis, and use simultaneous bands to qualify statements about particular horizons. This separates evidence that some response exists from evidence about exactly when it occurs and what sign it takes. Both interpretations remain conditional on the identification assumptions and the validity of the covariance and distributional approximations.

References

Hansen, L. P. (1982). Large Sample Properties of Generalized Method of Moments Estimators. Econometrica, 50(4), 1029–1054. https://doi.org/10.2307/1912775.

Jordà, Ò., and Taylor, A. M. (2025). Local Projections. Journal of Economic Literature, 63(1), 59–110. https://doi.org/10.1257/jel.20241521. See also the August 2024 working-paper version, Section 7 and Figure 6.

MacKinnon, J. G., and White, H. (1985). Some heteroskedasticity-consistent covariance matrix estimators with improved finite sample properties. Journal of Econometrics, 29(3), 305–325. https://doi.org/10.1016/0304-4076(85)90158-7.

Montiel Olea, J. L., and Plagborg-Møller, M. (2019). Simultaneous Confidence Bands: Theory, Implementation, and an Application to SVARs. Journal of Applied Econometrics, 34(1), 1–17. https://doi.org/10.1002/jae.2656. The author version explains the sup-t construction in Section 2 and discusses the absence of uniform power dominance in footnote 6.

Python documentation: NumPy Cholesky decomposition and SciPy’s chi-square distribution.

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.