You run a strategy, look at the returns, and spot a whiff of momentum — yesterday’s up day seems to lead to today’s. Is that a real edge or a coincidence? A single lag of autocorrelation is anecdote: you looked at one of many possible lags and one of them was bound to look interesting. The Ljung-Box test replaces the anecdote with evidence. It rolls the whole autocorrelation function up to lag h into one number and asks a single question: is there structure here, or is this white noise?
1. The portmanteau idea
Instead of testing one lag, a portmanteau test squares the sample autocorrelations from lag 1 to lag h and sums them. Under the null that the series is serially uncorrelated, each autocorrelation should hover near zero, so the sum should be small; a large sum is evidence of memory somewhere in the first h lags. Ljung & Box (1978) scale that sum so it follows a χ² distribution with h degrees of freedom:
Q = n(n+2) · Σₖ ₋₁ᵢᵎ ρ̂ᵢ² / (n − k) ~ χ²(h)
A small p-value rejects “no autocorrelation.” The older Box-Pierce (1970) statistic drops the (n+2)/(n−k) weighting; Ljung-Box adds it as a small-sample correction and is the version worth reporting.
2. Point it at three different things
The same test answers three different questions depending on what you feed it. On returns, it tests for predictability — momentum or mean reversion in the mean. On squared or absolute returns, it tests for volatility clustering: calm and stormy periods bunching together, the fingerprint of ARCH effects, even when the returns themselves look random. On model residuals, it is a specification check — if your model captured the structure, its residuals should be white noise; leftover autocorrelation means something is missing (drop the fitted parameters from the degrees of freedom).
3. Reading it honestly
Rejecting the null means there is autocorrelation; it does not mean there is a trade. The predictable component can be far smaller than the spread and fees you would pay to harvest it — statistically real, economically dead. And because the test aggregates across lags, it tells you structure exists somewhere in the first h, not where. Treat a rejection as a prompt to look closer, not a signal by itself.
4. Computing it
Our open-source orderflow-metrics library ships both statistics, dependency-free, in TypeScript and Python. The χ² tail is computed exactly for any number of lags:
import { ljungBox } from "orderflow-metrics";
// returns with runs of same-sign moves (trending microstructure)
const returns = [/* 0.010, 0.013, 0.009, -0.008, -0.011, ... */];
ljungBox(returns, 5);
// { statistic: 50.72, degreesOfFreedom: 5, pValue: 9.9e-10 }
// p ≈ 0 — the returns are serially correlated, not white noise
// point it at squared returns to test for volatility clustering:
ljungBox(returns.map(r => r * r), 5);
// { statistic: 13.21, degreesOfFreedom: 5, pValue: 0.021 } — clustering present
The Python distribution exposes the same function (ljung_box, plus box_pierce). It lifts the single-lag autocorrelation into a joint test, the natural companion to the variance-ratio test and the mean-reversion diagnostics in the wider market-microstructure toolkit.
5. Conclusion
Before you trust an edge — or a model — ask whether what you are seeing survives the whole autocorrelation function, not one flattering lag. Ljung-Box gives a single number and a confidence level: on returns it flags predictability, on squared returns volatility clustering, on residuals misspecification. It is the difference between “looks autocorrelated” and “is.” Explore the rest of the toolkit in our quantitative research library, or read the implementation on our open-source page.