Here is the pairs trade that looks perfect and quietly bankrupts you. You plot the spread between two correlated assets; it wiggles around a level, up and back, down and back — textbook mean reversion. So you fade the extremes and wait for the pull back to the mean. For a while it works. Then the spread pushes to a new extreme and doesn’t come back. You add, because the strategy says so. It widens again. The “mean” you were trading around had been drifting the whole time — and now it’s gone, along with your margin. The problem was never your entries. You never checked whether the spread was mean-reverting in the first place.
1. A random walk wiggles too
This is the trap: a random walk also meanders around a level for long stretches. Over a finite window a genuine mean-reverting series and a pure random walk can look almost identical, and no correlation will tell them apart. The difference is structural. In a stationary series, shocks decay — push it away and a force pulls it back. In a random walk (a unit root), shocks are permanent — every move is baked into the new level and the series is free to wander anywhere. One comes back. The other doesn’t. Trading the first is a strategy; trading the second is a slow-motion blow-up.
2. The test that decides
The Augmented Dickey-Fuller test settles it. It regresses today’s change on yesterday’s level, plus a few lagged changes to mop up short-run autocorrelation:
Δyₜ = α + γ·yₜ₋₁ + Σ δᵢ·Δyₜ₋ᵢ + εₜ
and reads the t-statistic on γ. If the level has no pull (γ = 0), the series is a random walk. If γ is reliably negative — high levels get pushed down, low levels pulled up — it is stationary and it reverts. The lagged differences are the “augmentation” that keeps the test valid on real, autocorrelated data.
3. Why the t-statistic isn’t normal
The twist that trips people up: under the unit-root null, that t-statistic does not follow a normal or Student distribution. You cannot compare it to ±1.96. It has its own (Dickey-Fuller) distribution, skewed to the negative side, which is why the test carries its own critical values — more negative than the usual ones. A statistic below the critical value rejects the unit root; the p-value uses MacKinnon’s response-surface approximation.
4. Reading it honestly
Two caveats worth their weight. Failing to reject is not proof of a random walk — ADF has famously low power near the unit root, so on a short window a truly reverting spread can slip past. Don’t confuse “couldn’t reject” with “it’s a random walk.” And passing the test is a filter, not a thesis: a spread that is stationary this year can break when the economic link behind it breaks. Pair the test with a reason the two assets should move together, and a stop that respects being wrong.
5. Computing it
Our open-source orderflow-metrics library ships the ADF test, dependency-free, in TypeScript and Python — the statistic, MacKinnon p-value, and critical values in one call:
import { augmentedDickeyFuller } from "orderflow-metrics";
// a candidate spread between two correlated assets
const spread = [/* 2.0, 1.4, 0.9, 1.2, 0.5, -0.1, ... */];
augmentedDickeyFuller(spread, 1, "c");
// { statistic: -3.874, pValue: 0.0022, usedLag: 1, nobs: 38,
// criticalValues: { "1%": -3.616, "5%": -2.941, "10%": -2.609 } }
// the statistic is below the 1% critical value — reject the unit root → stationary
The statistic clears the 1% hurdle, so this spread genuinely mean-reverts — the reversion is real, not a random walk that got lucky. The Python distribution exposes the same function (augmented_dickey_fuller), with a constant ("c") or constant-plus-trend ("ct") specification. It sits beside the mean-reversion and variance-ratio diagnostics; we framed the same idea for a trading audience in a short note on spreads that wander.
6. Conclusion
Before you bet that a spread comes back, prove that it does. The ADF test asks the one question mean-reversion trading depends on — stationary, or unit root? — and answers it with a statistic and a set of critical values built for the job. That is the whole discipline: test for stationarity first, size the reversion second. Explore the rest of the toolkit in our quantitative research library, or read the implementation on our open-source page.