Skip to content

Check data quality with audit

ob-analytics audit <source> runs the pipeline, prints a per-run data-quality summary, and exits non-zero when a check fails — so it works both as a thing to read and as a gate in a script. Point it at the same source you would pass to process:

ob-analytics audit orders.csv
ob-analytics audit data/ --source lobster --trading-date 2012-06-21
ob-analytics audit orders.csv --json           # machine-readable, for CI
ob-analytics audit results/ --from-parquet     # a saved 'process' output
ob-analytics audit orders.csv --strict         # warnings fail too

validate is the old name for this verb and still works.

Text output looks like this (bundled Bitstamp sample):

Data quality summary
  feed type             : diff_feed
  events / orders       : 314,057 / 156,902
  trades                : 284
  crossed resting book  : 91.61% of session (7238 episode(s)) [diff feed, but 2 stale resting order(s) stay in the book — see stale resting orders]
  stale resting orders  : 2 (worst: ask 2002347646152704 at 78,333 held the ask touch for 27.4 min after a trade printed through it)
  unmatched trades      : 0.70% [maker and taker]
  duplicate event ids   : 0
  duplicate created ids : 0
  pre-existing orders   : 13
  orphan orders         : 13 (13 event(s), no created row)
  impossible values     : 49 non-positive price(s) / 0 negative volume(s)
  clock order           : 0 venue-after-receive / 11 reordered
  venue sequence        : 0 missing / 0 out-of-order (0 row(s) numbered)
Checks: 0 error(s), 4 warning(s)
  WARNING orphan_orders: 13 order(s) are changed or deleted with no created event ...
  WARNING stale_orders: 2 resting order(s) a trade printed through and the venue did not report again within 1 s ...
  WARNING nonpositive_price: 49 row(s) are priced at or below zero: not a tradeable level
  WARNING exchange_time_reordered: 11 message(s) arrived out of venue order ...

Reading the metrics

Field Read it as
feed type matched_book (LOBSTER/MBO) or diff_feed (Bitstamp) — sets expectations for the next line
crossed resting book Share of session time with best_bid > best_ask. ~0% for a matched book; can be high and faithful for a diff feed, unless stale resting orders cause it
stale resting orders Resting orders a trade printed through that the venue did not report again within 1 s. The worst is named with its side, price and how long it held the touch
unmatched trades Trades whose maker or taker order could not be found among the order events. Only the orders the feed can show are looked for: the note in brackets says which
duplicate event ids / created ids Should be 0; anything else is a feed defect worth chasing
pre-existing orders Orders already resting when the capture began (no created row) — structurally unclassifiable, not errors
orphan orders Orders changed or deleted with no created row at all. The opening book is the honest source of these; a rise mid-session is the stream losing messages
impossible values Levels priced at or below zero, and negative volumes or fills
clock order Rows the venue stamped after we received them, and messages that reached the capture out of venue order
venue sequence Skipped and non-advancing sequence numbers: dropped and reordered messages (gap detection)

A high crossed resting book number on a diff_feed can be expected — see Data quality: matched book vs diff feed for why, and for the uncross= option that cleans the book up for display without touching the data you analyse. On a matched_book, a non-zero figure is a red flag.

Read it together with stale resting orders. On the bundled sample one stale order holds the ask touch for 27 minutes, and it causes almost all of the 91.61%. When the run has stale orders, the crossing note says so instead of calling the crossing normal.

What fails a run

Every metric above is scored by a named check carrying a severity, and the exit code follows the severities rather than the numbers:

Severity Meaning Exit code
error The data contradicts something that must hold non-zero
warning Worth reading, but a sound capture can show it 0, or non-zero with --strict
info Context; never fails a run 0

Errors: duplicate_event_ids, duplicate_created_ids, sequence_gaps, sequence_out_of_order, negative_volume, exchange_time_after_receive, and crossed_book on a matched book only.

Warnings: orphan_orders, stale_orders, nonpositive_price, exchange_time_reordered, unmatched_trades (above 5%), and crossed_book when no feed type was declared.

Two of these are judgement calls worth stating plainly:

  • A crossed book is scored by feed type, not by size. A crossed book is always a defect in a matched book, but can be a faithful replay of a diff feed. Only the source's declared FeedType can tell them apart. Size alone says little: on the bundled sample, almost all of the 92% comes from two stale orders, which stale_orders reports. With --from-parquet and no --source, the feed type is undeclared and crossing drops to a warning rather than being guessed.
  • A dropped created message cannot be told apart from an order that was already resting when the capture began — both leave an order that is only ever changed or deleted. So orphan_orders is a warning, and the hard evidence for dropped messages is sequence_gaps, which needs a feed that carries a venue sequence. audit always loads with sequence tracking on. A skipped number is a dropped message only when the venue adds one per message. A ccxt capture records in meta.json that its sequence only rises, and audit then checks only that it never goes back, and prints gaps not checked.

Most sources carry no sequence that can prove a message was lost. On those, audit prints 0 row(s) numbered for the venue sequence, and a dropped message shows only indirectly, as orphan or stale orders:

Source Venue sequence What audit can check
bitstamp (files and live capture) none: the feed has a timestamp only nothing
lobster none nothing
cryptofeed one number per message, on venues that publish one gaps and order
ccxt, including Binance rises, but skips on its own order only
ccxt for Kalshi and Polymarket none nothing
databento rises, but skips on its own order only

Databento numbers every message on the venue's channel. A file usually holds one instrument of that channel, and its trade and fill records become trades rather than book events, so the numbers left in the events skip even when nothing was lost. The source declares this, and audit then prints gaps not checked. Audit a saved Databento output with --from-parquet --source databento: without --source, nothing says what the numbers promise, and every skip counts as a lost message. - Unmatched trades count only the orders the feed can show. Every trade has a maker, the order that was resting, and a taker, the order that arrived and traded against it. Only a feed that reports every order shows the taker. Each source declares which it can show, as its TradeAttribution:

Source Shows What unmatched_trades counts
bitstamp maker and taker trades missing either
lobster, databento, cryptofeed at L3 maker only: a taker trades on arrival and never rests trades missing the maker
ccxt, depth_csv, cryptofeed at L2 neither: price levels have no order identity nothing

LOBSTER's trades do carry a taker, but it is a guess: Nasdaq ITCH names only the resting order of an execution, and the reader picks the most recent new order on the other side that could have traded. So the check does not count it. - A capture is checked against the source that made it. A live capture records its source and that source's declarations in meta.json, and ob-analytics process copies meta.json into its output. audit uses the record even when --source names another source. --source also says how to read the files, and a cryptofeed L3 capture can only be read as bitstamp, whose feed shows more. The log says when the record overrides --source. - A capture of several segments is audited one segment at a time. Given a capture directory (or the process output made from one), audit prints a report for each segment and then the capture's own checks from manifest.json: capture_gaps (time no segment covered), unfinished_segments (segments a dead capture process left open), and dropped_messages. All three are warnings, so --strict fails on them. See Running for days. - A stale order is reported, not removed. A trade above a resting ask (or below a resting bid) shows the order has gone, because a matching engine fills the better price first. The venue normally reports that order within milliseconds, so stale_orders waits one second before it counts one. The test needs to know what the feed said after the trade, which a live capture cannot know in time, so order_book() keeps the order. detect_stale_orders runs the same test from Python.

In CI

ob-analytics audit orders.csv --json > quality.json || exit 1

--json writes the whole summary, including ok and every check, so a build can read the verdict instead of parsing the text block:

{
  "feed_type": "diff_feed",
  "orphan_orders": 13,
  "ok": true,
  "checks": [
    {
      "name": "duplicate_event_ids",
      "passed": true,
      "severity": "error",
      "detail": "0 event_id value(s) occur more than once; ..."
    }
  ]
}

From Python

from ob_analytics import Pipeline, BitstampSource, FeedType, data_quality_summary

result = Pipeline().run("orders.csv")
summary = data_quality_summary(
    result.events, result.trades,
    feed_type=BitstampSource().feed_type,   # or getattr(source, "feed_type", FeedType.UNKNOWN)
    depth=result.depth,                      # faithful depth; not depth_summary
    tick_size=result.config.tick_size,       # stale-order prices in the quote currency
)
print(summary.render())
summary.ok            # False when an error-severity check failed
summary.errors        # the failed error checks, each with a one-line detail
summary.warnings      # the failed warning checks
summary.stale_orders  # StaleOrder records, worst first
summary.to_dict()     # JSON-serialisable, including every check

The sequence checks need the venue's numbers, which a run keeps only with PipelineConfig(track_sequence=True); audit turns this on for you. Pass the source's sequence_kind too, or a source whose numbers skip on their own, such as databento or ccxt, is read as losing messages:

from ob_analytics import DatabentoSource, Pipeline, PipelineConfig, data_quality_summary
from ob_analytics.protocols import sequence_kind_of, trade_attribution_of

source = DatabentoSource()
result = Pipeline(PipelineConfig(track_sequence=True), source=source).run("aapl.mbo.dbn.zst")
summary = data_quality_summary(
    result.events, result.trades,
    feed_type=source.feed_type,
    depth=result.depth,
    sequence_kind=sequence_kind_of(source),
    trade_attribution=trade_attribution_of(source),
)

Pass depth, not depth_summary

Crossing is measured from the faithful resting book. depth_summary is already uncrossed by the depth engine, so passing it would always report ~0%. Omit depth and it is recomputed from events.