Back to Research
Market Microstructure August 9, 2026 • 18 min read

Information-Driven Bars: Tick, Volume & Dollar Sampling for Market Microstructure

Almost every trader learns the market on time bars — 1-minute, 5-minute, daily candles. Yet the clock is arguably the worst possible way to sample a market. This article explains why, and how sampling by activity — a bar every N trades, N units of volume, or N dollars traded — produces returns that are dramatically better behaved for the statistical models that sit downstream.

1. The Hidden Cost of the Clock

A time bar aggregates every trade that prints within a fixed wall-clock interval. It is simple, universally supported, and — for quantitative purposes — deeply flawed. Information does not arrive at a constant rate. Markets are frantic in the minutes after a macro release and nearly silent at 3 a.m. When you sample on a fixed clock, you attach the same weight to a minute that carried a single lonely print as to a minute that absorbed ten thousand trades and a full percent of price movement.

The consequences are not cosmetic. Time bars oversample low-activity periods and undersample high-activity periods. The returns they produce exhibit serial correlation, heteroskedasticity, and fat tails far in excess of what most models assume. Any downstream estimator — a volatility forecast, a variance ratio test, a signal regression — inherits that distortion. As we argue in why most alpha models fail in production, a surprising share of "alpha decay" is nothing more than a sampling artifact that never survived contact with a properly clocked market.

The fix, popularized by Marcos López de Prado in Advances in Financial Machine Learning (2018), is to stop sampling on time and start sampling on information. Instead of "every 60 seconds," we sample "every time the market has done a fixed amount of work." These are information-driven bars.

2. Sampling by Activity: The Three Bars

All three variants share the same idea — accumulate a running counter over the raw trade stream and close a bar the moment the counter crosses a threshold. They differ only in what they count.

Tick bars

A tick bar closes after a fixed number of trades, regardless of their size. If the threshold is T trades, the bar boundaries fall at cumulative trade counts T, 2T, 3T, …. Tick bars normalize for the frequency of activity: a bar always represents the same amount of matching-engine "events."

close a bar when  count(trades) ≥ T

Volume bars

A volume bar closes once the cumulative traded quantity reaches a threshold V. Because a single large print can carry more size than a thousand small ones, volume bars weight the market by how much was actually transacted — a far better proxy for information flow than raw trade count.

close a bar when  Σ sizei ≥ V

Dollar bars

A dollar bar closes once the cumulative traded value — price times size, summed — reaches a threshold D. Dollar bars normalize for both activity and price level, which makes them uniquely robust across regimes: when an asset doubles in price, a fixed volume threshold silently doubles the economic content of each bar, whereas a dollar threshold does not.

close a bar when  Σ (pricei × sizei) ≥ D

In every case the trade that crosses the threshold is included whole — never split — so the realized volume or value of a bar may slightly overshoot its target. A trailing partial bar (still below threshold when the stream ends) is dropped, exactly as an incomplete volume bucket is dropped when computing VPIN.

3. Why It Matters: A Side-by-Side View

The table below summarizes what each sampling scheme normalizes for, and where it breaks down.

Bar type Closes on Normalizes for Main weakness
Time Wall-clock interval Nothing Over/undersamples with activity; non-IID returns
Tick Trade count Frequency of events Sensitive to order fragmentation & quote stuffing
Volume Traded quantity Transacted size Economic content drifts with price level
Dollar Traded value Size and price level Threshold needs periodic recalibration as the market grows

Empirically, information-driven bars — dollar bars in particular — yield return series that are closer to independent and identically distributed, with distributions much nearer to Gaussian and far less serial correlation. That is precisely the property every classical estimator quietly assumes. Better-behaved inputs mean a variance ratio test that actually measures mean reversion rather than sampling noise, and volatility forecasts whose residuals behave.

4. Why Dollar Bars Are Usually Preferred

Suppose a token trades at $10 and you set a volume threshold of 100,000 units — every bar carries roughly $1M of traded value. Six months later the token trades at $40. The same 100,000-unit threshold now packs $4M into each bar: the "resolution" of your series has quietly quadrupled, and any model calibrated on the old bars is now looking at a different object. A dollar threshold sidesteps this entirely — $1M is $1M whether the price is $10 or $40.

Dollar bars also degrade gracefully around corporate actions, supply changes, and the extreme price ranges common in the crypto and digital-asset markets we cover. For most research pipelines they are the sensible default; volume and tick bars remain useful as diagnostics and for venues where a stable notional is hard to define.

5. From Bars to Alpha: Feeding the Rest of the Stack

Information-driven bars are not an end in themselves — they are the sampling layer the rest of a microstructure stack should sit on. Once the raw feed has been normalized and, where needed, the book reconstructed from incremental updates, bars become the natural clock for everything that follows:

  • Order Flow Imbalance — aggregating signed size per bar (each bar already carries buy and sell volume separately) turns OFI into a clean, activity-clocked predictor of short-horizon price drift.
  • VPIN & flow toxicity — volume bars are conceptually the same construct VPIN uses for its buckets, so the two compose directly.
  • Realized volatility — computing realized variance on dollar-bar returns removes most of the intraday seasonality that plagues fixed-interval estimates.
  • Impact & execution — bar VWAP and signed flow are the raw material for calibrating impact models and optimal execution schedules.

This is the philosophy behind our broader market microstructure analytics: get the sampling right first, and every metric built on top inherits the benefit.

6. Reference Implementation

We maintain an open-source, dependency-free TypeScript implementation of these bars — and the metrics that consume them — in our public orderflow-metrics library. The API is deliberately small: pass a trade stream and a threshold, receive an array of OHLCV bars carrying open/high/low/close, volume, traded value, VWAP, tick count, and signed buy/sell volume.

import { dollarBars, realizedVolatility } from "orderflow-metrics";

// Raw prints from the exchange feed
const trades = [
  { ts: 1_700_000_000_000, price: 42_150.5, size: 0.8, side: "buy"  },
  { ts: 1_700_000_000_120, price: 42_151.0, size: 1.2, side: "buy"  },
  { ts: 1_700_000_000_340, price: 42_149.5, size: 0.5, side: "sell" },
  // ...thousands more
];

// Sample the stream into $250k dollar bars
const bars = dollarBars(trades, 250_000);

for (const bar of bars) {
  const signedFlow = bar.buyVolume - bar.sellVolume;
  console.log(bar.close, bar.vwap.toFixed(2), signedFlow);
}

// Downstream metrics now run on activity-clocked bars,
// not on the distorted wall clock.
const barReturns = bars
  .slice(1)
  .map((b, i) => Math.log(b.close / bars[i].close));

realizedVolatility(barReturns);

Swapping dollarBars for volumeBars or tickBars changes nothing else in the pipeline — the returned Bar shape is identical, so you can benchmark all three sampling schemes against the same downstream model with a one-line change.

7. Practical Considerations

Threshold selection. A useful starting point is to target a desired average number of bars per day: estimate daily traded value (or volume, or trade count) and divide by the number of bars you want. Because activity itself trends over weeks and months, treat the threshold as something to recalibrate periodically rather than a constant set once.

Partial bars. Dropping the trailing incomplete bar keeps every emitted bar comparable, but it means the most recent (still-forming) bar is not yet visible. Live systems typically track the in-progress accumulator separately and only "commit" a bar on threshold crossing.

Beyond the basics. The same accumulator idea extends to imbalance and run bars, which close when signed order flow — not raw activity — crosses a threshold, sampling the market most finely exactly when informed trading clusters. These are a natural next step once tick, volume, and dollar bars are in place, and they connect directly back to order flow imbalance as a triggering signal.

8. Conclusion

The clock is a convenience, not a law of markets. Sampling by activity — tick, volume, and above all dollar bars — realigns your data with the way information actually arrives, producing returns that are closer to IID and far friendlier to every estimator downstream. It is one of the cheapest, highest-leverage changes a quantitative pipeline can make: no new data, no new model, just a better way to read the tape. For the rest of the toolkit that builds on this foundation, browse the full quantitative research library or explore the infrastructure that produces these feeds in the first place.

For more on low-latency infrastructure and open-source microstructure tooling, visit our official resources:

🧩 Open Source 💻 orderflow-metrics on GitHub 📚 More Research