Back to Research
Risk Management September 18, 2026 • 10 min read

Does Your VaR Model Actually Work? Backtesting with Kupiec & Christoffersen

A 99% VaR should be breached one day in a hundred — no more, and never in a cluster. Backtesting is how you check that a risk model earned its number, and it is exactly what regulators require. Kupiec and Christoffersen turn “did the model hold up?” into a hypothesis test with a p-value.

A Value-at-Risk number is a promise: over the next day, the book will not lose more than this, except on 5% of days. It is one of the most quoted figures in risk management — and one of the least checked. A forecast nobody verifies is decoration on a risk report. Backtesting is the verification, and since 1996 it has not been optional: the Basel framework makes institutions count their VaR breaches and penalises models that fail. The same discipline belongs in any book that sizes on VaR.

1. The count you have to earn

Start with the obvious check. A 95% VaR should be breached — the realised loss exceeding the forecast — about 5% of the time. Over 250 trading days that is roughly twelve or thirteen exceptions. Far more, and the model understates risk; far fewer, and it overstates it, quietly wasting capital on margin you never needed. Kupiec's (1995) proportion-of-failures test turns “is the breach rate right?” into a likelihood-ratio statistic, distributed χ² with one degree of freedom, so you get a p-value instead of a shrug.

LR uc = 2·ln[ (1−π̂)ⁿ⁻ˣ·π̂ˣ / ((1−p)ⁿ⁻ˣ·pˣ) ]  ~  χ²(1)

2. The count is not enough

A model can breach the right number of times and still be dangerous. Picture twelve exceptions in a year — a perfect count — but all twelve falling in the same turbulent fortnight. That is not a model that occasionally slips; it is one that goes blind exactly when volatility rises and losses cluster. Christoffersen's (1998) independence test catches it. It reads the sequence of 0/1 breach flags and asks whether a breach today makes a breach tomorrow more likely — serial dependence in the exceptions — again as a χ²(1) likelihood-ratio test.

3. The joint test: conditional coverage

The one to report combines both. Conditional coverage adds the two likelihood ratios and tests them together against χ² with two degrees of freedom: the model must get the number of breaches and their spacing right. A VaR that passes Kupiec on the count can still be rejected here for clustering — and that rejection is the more useful warning, because clustered breaches are how a book takes several limit-exceeding losses in a row.

LR cc = LR uc + LR ind  ~  χ²(2)

4. Reading it honestly

Two cautions. A one-year window is short for tail testing: at 99% you expect only two or three breaches in 250 days, so the tests have limited power and a weak model can survive a quiet year — which is why regulators lean on multi-year records and a graduated “traffic-light” response. And failing to reject is not proof the model is right; it is the absence of evidence that it is wrong. Backtesting narrows the set of models you can trust, it does not certify one.

5. Computing it

Our open-source orderflow-metrics library ships both tests, dependency-free, in TypeScript and Python. Feed a 0/1 breach series and the target rate:

import { kupiecPOF, christoffersenConditionalCoverage } from "orderflow-metrics";

// 250-day breach flags (1 = the day’s loss exceeded the 95% VaR); 18 exceptions, some clustered
const breaches = [/* 0, 0, 1, 0, 1, 1, 1, ... */];

kupiecPOF(breaches, 0.05);
// { exceptions: 18, observations: 250, statistic: 2.256, pValue: 0.133 }
// 18 breaches vs ~12.5 expected — on the count alone, not enough to reject (p = 0.13)

christoffersenConditionalCoverage(breaches, 0.05);
// { statistic: 25.54, pValue: 0.0000028 }
// the joint test rejects hard — the breaches cluster, so the model misses volatility regimes

Here the count would have passed, and the joint test still fails because the exceptions bunch together — the whole reason to backtest past the headline number. The Python distribution exposes the same functions (kupiec_pof, christoffersen_conditional_coverage). They sit directly beside the Value-at-Risk & Expected Shortfall tooling that produces the forecast: measure the risk, then prove the measure.

6. Conclusion

A VaR you do not backtest is a number you hope is true. Kupiec asks whether the breaches arrive at the right rate; Christoffersen asks whether they arrive independently; conditional coverage insists on both. Run them on a rolling window and a failing model announces itself before the losses do. Explore the rest of the toolkit in our quantitative research library, or read the implementation on our open-source page (npm and PyPI, MIT-licensed).

For more on market efficiency, regime diagnostics, and open-source tooling, visit our official resources:

🧩 Open Source 💻 orderflow-metrics on GitHub 📚 More Research