Analytics¶
Format-agnostic post-processing analytics. These functions work with the output of any format's pipeline run (Bitstamp, LOBSTER, or custom).
Trade Analysis¶
order_aggressiveness ¶
Calculate order aggressiveness with respect to the best bid or ask in BPS.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
events
|
DataFrame
|
The events DataFrame (must contain |
required |
depth_summary
|
DataFrame
|
The order book summary statistics DataFrame (must contain |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
The events DataFrame with an added |
trade_impacts ¶
Generate a DataFrame containing order book impact summaries.
Aggregates trade records by taker order ID to summarise how each aggressive order swept through the book (price range, number of fills, total volume, VWAP, duration).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
trades
|
DataFrame
|
The trades DataFrame (must contain |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
A DataFrame summarising market order impacts with columns:
|
Order Type Classification¶
set_order_types ¶
Determine limit order types.
Classifies each order as one of: market, resting-limit, flashed-limit, or market-limit, based on how the order interacts with the book over its lifetime.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
events
|
DataFrame
|
The limit order events DataFrame. |
required |
trades
|
DataFrame
|
The executions DataFrame. |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
The events DataFrame with an updated 'type' column indicating order types. |
Order Book Reconstruction¶
The reconstructions themselves live in the order-book engine; these are their frame-level faces.
order_book ¶
order_book(
events: DataFrame,
tp: datetime | None = None,
max_levels: int | None = None,
bps_range: int = 0,
min_bid: float = 0,
max_ask: float = np.inf,
uncross: bool = False,
) -> dict[str, datetime | pd.Timestamp | pd.DataFrame]
Reconstruct the order book at a specific point in time.
The reconstruction itself is :func:ob_analytics.engine.book_state; this
function is its frame adapter, and owns the display window (max_levels,
bps_range) the engine has no opinion about.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
events
|
DataFrame
|
DataFrame containing order events. |
required |
tp
|
datetime or Timestamp
|
The point in time at which to evaluate the order book. If None, uses the latest event timestamp in the data. |
None
|
max_levels
|
int
|
The maximum number of price levels to include for bids and asks. |
None
|
bps_range
|
int
|
Basis points range to filter the bids and asks. Default is 0. |
0
|
min_bid
|
float
|
Minimum bid price. Default is 0. |
0
|
max_ask
|
float
|
Maximum ask price. Default is infinity. |
inf
|
uncross
|
bool
|
When |
False
|
Returns:
| Type | Description |
|---|---|
dict[str, datetime or DataFrame]
|
A dictionary containing: - 'timestamp': The evaluation timestamp. - 'asks': DataFrame of active ask orders. - 'bids': DataFrame of active bid orders. |
order_lifecycles ¶
Collapse events into one row per order: placement → outcome.
The canonical lifecycle table (one derivation, shared by the L3 faces
and the order-book reconstruction). A thin frame wrapper over
:func:ob_analytics.engine.order_lifecycles, which relies on the
schemas.py volume contract: volume is the outstanding size after each
event and fill the executed delta, so an order is terminated when a
deleted row arrives or its outstanding size reaches zero — the
latter is how fully-executed LOBSTER orders end, which never emit a
deleted event.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
events
|
DataFrame
|
Events satisfying the schemas.py contract. Orders without a
|
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
One row per order id:
|
uncross_book_sides ¶
Evict crossed levels from two reconstructed book sides for display.
The frame-level counterpart of order_book(..., uncross=True) for
callers that already hold per-order book sides — e.g. the book_snapshot
/ depth_chart visualization prepares. Both frames are returned
best-first (bids by descending price, asks by ascending price) with the
crossed best-end orders removed so best_bid < best_ask; liquidity
is recomputed when present and every other column is preserved.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
bids
|
DataFrame
|
Per-order book sides carrying at least |
required |
asks
|
DataFrame
|
Per-order book sides carrying at least |
required |
Returns:
| Type | Description |
|---|---|
tuple of (pandas.DataFrame, pandas.DataFrame)
|
The uncrossed |
Data Quality¶
See Data quality: matched book vs diff feed for the
concepts and the audit how-to for the CLI.
data_quality_summary ¶
data_quality_summary(
events: DataFrame,
trades: DataFrame,
*,
feed_type: FeedType = FeedType.UNKNOWN,
depth: DataFrame | None = None,
) -> DataQualitySummary
Summarise the data quality of one reconstructed session.
Surfaces the health signals that matter before trusting a feed — most
importantly how crossed the resting book is, which distinguishes a matched
book from a diff feed (see :class:~ob_analytics.protocols.FeedType).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
events
|
DataFrame
|
Classified events (must carry the canonical columns and the
|
required |
trades
|
DataFrame
|
The trades frame, with |
required |
feed_type
|
FeedType
|
The source's declared feed type, recorded on the summary and used to
interpret |
UNKNOWN
|
depth
|
DataFrame
|
A faithful price-level-volume frame (e.g. |
None
|
Returns:
| Type | Description |
|---|---|
DataQualitySummary
|
|
DataQualitySummary
dataclass
¶
DataQualitySummary(
feed_type: FeedType,
n_events: int,
n_orders: int,
n_trades: int,
crossed_pct: float,
crossed_episodes: int,
unmatched_trades_pct: float,
duplicate_event_ids: int,
duplicate_created_ids: int,
pre_existing_orders: int,
events_with_sequence: int = 0,
sequence_gaps: int = 0,
sequence_out_of_order: int = 0,
orphan_orders: int = 0,
orphan_events: int = 0,
nonpositive_price_rows: int = 0,
negative_volume_rows: int = 0,
exchange_time_after_receive: int = 0,
exchange_time_reordered: int = 0,
)
Per-run data-quality metrics for a reconstructed session.
Built by :func:data_quality_summary. All percentages are 0–100 floats.
The fields are the measurements; :attr:checks turns them into pass/fail
verdicts with a :class:Severity each, and :attr:ok is the one-line
answer to "is this feed trustworthy?" that ob-analytics audit exits on.
Attributes:
| Name | Type | Description |
|---|---|---|
feed_type |
FeedType
|
The source's declared crossing invariant (see
:class: |
n_events, n_orders, n_trades |
int
|
Row / distinct-order / trade counts. |
crossed_pct |
float
|
Percentage of session time the faithful book is crossed
( |
crossed_episodes |
int
|
Number of distinct crossed intervals. |
unmatched_trades_pct |
float
|
Percentage of trades missing a resolved |
duplicate_event_ids |
int
|
Count of |
duplicate_created_ids |
int
|
Count of order ids with more than one |
pre_existing_orders |
int
|
Distinct orders resting before the capture window (classifier label
|
events_with_sequence |
int
|
Rows carrying a non-null venue |
sequence_gaps |
int
|
Dropped-message count: skipped venue sequence numbers. |
sequence_out_of_order |
int
|
Reordered or duplicated messages: sequence steps that did not advance. |
orphan_orders |
int
|
Distinct order ids with a |
orphan_events |
int
|
Rows belonging to those orphan orders. |
nonpositive_price_rows |
int
|
Rows priced at or below zero. Legal in the schema (prices are signed integer ticks) but not a tradeable level. |
negative_volume_rows |
int
|
Rows with a negative |
exchange_time_after_receive |
int
|
Rows whose venue clock ( |
exchange_time_reordered |
int
|
Steps where the venue clock goes backwards while the receive clock moves forward: messages that reached the capture out of order. |
checks
property
¶
Every check this run was scored against, errors first.
The crossing check reads its severity off :attr:feed_type: a crossed
resting book is a defect in a matched book and a faithful property of a
diff feed, so the same number means opposite things and only the
declared feed type can tell them apart.
errors
property
¶
Failed checks whose severity is :attr:Severity.ERROR.
warnings
property
¶
Failed checks whose severity is :attr:Severity.WARNING.
QualityCheck
dataclass
¶
One named data-quality check and how it read on this run.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
Stable identifier, matching the summary field it reads
(e.g. |
passed |
bool
|
Whether the data satisfied the check. |
severity |
Severity
|
What a failure means (see :class: |
detail |
str
|
One line saying what was found and how to read it. |
Severity ¶
Bases: str, Enum
How much a failed data-quality check matters.
A check carries its severity so the policy — what fails a run — lives with
the measurement rather than in each caller. The enum mixes in str
(Severity.ERROR == "error"), which keeps CLI and JSON output plain.
Attributes:
| Name | Type | Description |
|---|---|---|
ERROR |
The data contradicts something that must hold (a duplicate
|
|
WARNING |
A signal worth reading before trusting the feed, but one a sound
capture can legitimately show (orders resting before the capture
began, zero-priced levels, messages reordered in transit). Fails the
run only under |
|
INFO |
Reported for context; never fails a run. |
detect_sequence_gaps ¶
detect_sequence_gaps(
frame: DataFrame,
*,
sequence_col: str = SEQUENCE_COLUMN,
order_col: str = INGEST_SEQ_COLUMN,
group_cols: Sequence[str] = _SEQUENCE_GROUP_COLUMNS,
) -> SequenceGapReport
Report missing or out-of-order venue sequence numbers in frame.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
frame
|
DataFrame
|
Events or depth rows. A missing sequence_col (a source with no venue sequence) yields an empty, clean report — the column is optional. |
required |
sequence_col
|
str
|
Column holding the venue's per-event sequence (nullable integer). |
SEQUENCE_COLUMN
|
order_col
|
str
|
Column defining ingest order; when absent, the frame's current row order is used instead. |
INGEST_SEQ_COLUMN
|
group_cols
|
sequence of str
|
Columns that identify one channel (instrument / venue). Only those present in frame are used; with none present the whole frame is one channel. Sequences from different channels are not comparable, so each group is scored on its own. |
_SEQUENCE_GROUP_COLUMNS
|
Returns:
| Type | Description |
|---|---|
SequenceGapReport
|
|
Notes
Consecutive equal sequence values are collapsed before comparison, so a single venue update that emits several rows (e.g. every changed level of one book diff) counts once. A non-consecutive repeat still shows as out-of-order.
SequenceGapReport
dataclass
¶
SequenceGapReport(
n_sequenced: int,
n_updates: int,
n_missing: int,
n_out_of_order: int,
max_gap: int,
first_break_seq: int | None,
)
Result of scanning a frame's venue sequence column for breaks.
Built by :func:detect_sequence_gaps. A feed numbers each message on a
channel with a sequence that should rise by exactly one; a skip is a
dropped message and a step that does not rise is a reordered (or repeated)
one. Rows are read in ingest order (ingest_seq when present, else the
frame's row order) and grouped per instrument / venue when those columns
exist, so interleaved channels are scored independently.
Attributes:
| Name | Type | Description |
|---|---|---|
n_sequenced |
int
|
Rows carrying a non-null venue |
n_updates |
int
|
Distinct consecutive sequence values seen — one per venue message (a book update that emits several rows shares one sequence, counted once). |
n_missing |
int
|
Total count of skipped sequence numbers: the dropped-message count. |
n_out_of_order |
int
|
Steps where the sequence did not advance (a repeat or a decrease) — a reordered or duplicated message. |
max_gap |
int
|
Largest single run of consecutive missing numbers ( |
first_break_seq |
int or None
|
The last in-order sequence value before the first break, a hint for
where to resync; |