Back to Research
Statistical Arbitrage September 21, 2026 • 10 min read

Is Your Spread Really Mean-Reverting? The Augmented Dickey-Fuller Test

A spread that wiggles around a level on a chart looks mean-reverting — but a random walk wiggles too, and trading the reversion of one that never comes back is how pairs trades blow up. The Augmented Dickey-Fuller test tells them apart before you put on the trade.

Here is the pairs trade that looks perfect and quietly bankrupts you. You plot the spread between two correlated assets; it wiggles around a level, up and back, down and back — textbook mean reversion. So you fade the extremes and wait for the pull back to the mean. For a while it works. Then the spread pushes to a new extreme and doesn’t come back. You add, because the strategy says so. It widens again. The “mean” you were trading around had been drifting the whole time — and now it’s gone, along with your margin. The problem was never your entries. You never checked whether the spread was mean-reverting in the first place.

1. A random walk wiggles too

This is the trap: a random walk also meanders around a level for long stretches. Over a finite window a genuine mean-reverting series and a pure random walk can look almost identical, and no correlation will tell them apart. The difference is structural. In a stationary series, shocks decay — push it away and a force pulls it back. In a random walk (a unit root), shocks are permanent — every move is baked into the new level and the series is free to wander anywhere. One comes back. The other doesn’t. Trading the first is a strategy; trading the second is a slow-motion blow-up.

2. The test that decides

The Augmented Dickey-Fuller test settles it. It regresses today’s change on yesterday’s level, plus a few lagged changes to mop up short-run autocorrelation:

Δyₜ = α + γ·yₜ₋₁ + Σ δᵢ·Δyₜ₋ᵢ + εₜ

and reads the t-statistic on γ. If the level has no pull (γ = 0), the series is a random walk. If γ is reliably negative — high levels get pushed down, low levels pulled up — it is stationary and it reverts. The lagged differences are the “augmentation” that keeps the test valid on real, autocorrelated data.

3. Why the t-statistic isn’t normal

The twist that trips people up: under the unit-root null, that t-statistic does not follow a normal or Student distribution. You cannot compare it to ±1.96. It has its own (Dickey-Fuller) distribution, skewed to the negative side, which is why the test carries its own critical values — more negative than the usual ones. A statistic below the critical value rejects the unit root; the p-value uses MacKinnon’s response-surface approximation.

4. Reading it honestly

Two caveats worth their weight. Failing to reject is not proof of a random walk — ADF has famously low power near the unit root, so on a short window a truly reverting spread can slip past. Don’t confuse “couldn’t reject” with “it’s a random walk.” And passing the test is a filter, not a thesis: a spread that is stationary this year can break when the economic link behind it breaks. Pair the test with a reason the two assets should move together, and a stop that respects being wrong.

5. Computing it

Our open-source orderflow-metrics library ships the ADF test, dependency-free, in TypeScript and Python — the statistic, MacKinnon p-value, and critical values in one call:

import { augmentedDickeyFuller } from "orderflow-metrics";

// a candidate spread between two correlated assets
const spread = [/* 2.0, 1.4, 0.9, 1.2, 0.5, -0.1, ... */];

augmentedDickeyFuller(spread, 1, "c");
// { statistic: -3.874, pValue: 0.0022, usedLag: 1, nobs: 38,
//   criticalValues: { "1%": -3.616, "5%": -2.941, "10%": -2.609 } }
// the statistic is below the 1% critical value — reject the unit root → stationary

The statistic clears the 1% hurdle, so this spread genuinely mean-reverts — the reversion is real, not a random walk that got lucky. The Python distribution exposes the same function (augmented_dickey_fuller), with a constant ("c") or constant-plus-trend ("ct") specification. It sits beside the mean-reversion and variance-ratio diagnostics; we framed the same idea for a trading audience in a short note on spreads that wander.

6. Conclusion

Before you bet that a spread comes back, prove that it does. The ADF test asks the one question mean-reversion trading depends on — stationary, or unit root? — and answers it with a statistic and a set of critical values built for the job. That is the whole discipline: test for stationarity first, size the reversion second. Explore the rest of the toolkit in our quantitative research library, or read the implementation on our open-source page.

For more on market efficiency, regime diagnostics, and open-source tooling, visit our official resources:

🧩 Open Source 💻 orderflow-metrics on GitHub 📚 More Research