Trade-Sign Classification¶
Infer the aggressor side of each trade when the feed doesn't label it.
L3 crypto (Bitstamp) ships buy_order_id / sell_order_id, so the trades
frame carries a real direction; L2 / aggregated feeds and many CCXT
sources don't, so the signed-flow metrics
(compute_vpin,
order_flow_imbalance)
have nothing to work with. These classifiers fill that gap and are wired in
as an automatic fallback.
Functions¶
classify_trade_sign ¶
classify_trade_sign(
trades: DataFrame,
method: str = "lee_ready",
quotes: DataFrame | None = None,
) -> pd.Series
Classify the aggressor side of each trade.
A drop-in source of the direction column for feeds that don't label
the aggressor. Sorts trades chronologically, applies the chosen
per-trade classifier, and returns the result realigned to the original
index.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
trades
|
DataFrame
|
Trades with at least |
required |
method
|
str
|
|
'lee_ready'
|
quotes
|
DataFrame
|
Required for |
None
|
Returns:
| Type | Description |
|---|---|
Series
|
Categorical |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If method is unknown, if |
ObAnalyticsError
|
If trades is empty. |
tick_rule ¶
Classify trade signs by the tick rule.
Signs each trade from the sign of its price change relative to the
previous trade: an uptick is buyer-initiated (+1), a downtick
seller-initiated (-1). A zero tick (unchanged price) inherits
the last non-zero sign — the classic Lee–Ready convention.
prices must already be in trade order (chronological). A leading run
of zero ticks (before the first price move) is back-filled from the
first determinable sign; a perfectly flat series has no information and
defaults to +1.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
prices
|
array - like
|
Trade prices in chronological order. Anything
:func: |
required |
Returns:
| Type | Description |
|---|---|
ndarray
|
|
lee_ready ¶
Classify trade signs by the Lee–Ready quote-midpoint test.
A trade above the prevailing mid is buyer-initiated (+1), below it
seller-initiated (-1). A trade at the mid — or one with no
prevailing quote (mid is NaN) — falls back to the
:func:tick_rule.
prices and mid must be equal-length and in chronological trade
order; mid is the quote midpoint prevailing at (or just before) each
trade — see :func:classify_trade_sign for aligning quotes to trades.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
prices
|
array - like
|
Trade prices in chronological order. |
required |
mid
|
array - like
|
Prevailing quote midpoint per trade ( |
required |
Returns:
| Type | Description |
|---|---|
ndarray
|
|
bulk_volume_classification ¶
bulk_volume_classification(
trades: DataFrame,
bucket_volume: float,
*,
sigma: float | None = None,
) -> pd.DataFrame
Split volume bars into buy / sell fractions (Easley–LdP–O'Hara BVC).
Partitions cumulative trade volume into equal-sized buckets (the same
volume bars :func:~ob_analytics.flow_toxicity.compute_vpin uses; a
trade straddling a boundary is split proportionally) and estimates each
bucket's buy fraction as
buy_fraction = Φ(ΔP / σ)
where ΔP is the bucket's close-to-close price change and σ the
standard deviation of those changes. Unlike a per-trade classifier this
labels volume, so it needs neither the aggressor side nor quotes — the
VPIN-native method for feeds that carry only trade prints.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
trades
|
DataFrame
|
Trades with |
required |
bucket_volume
|
float
|
Total volume per bucket (instrument-specific). |
required |
sigma
|
float
|
Standard deviation of bucketed price changes. Estimated from the data (sample std of the bucket ΔP series) when omitted. |
None
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
One row per completed bucket with columns |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If required columns are missing. |
ObAnalyticsError
|
If trades is empty. |
ValueError
|
If bucket_volume is not positive, or sigma is not positive. |
Shared helpers¶
The signed-flow analytics reach the quotes and the aggressor side through these two, rather than each re-deriving them.
resolve_direction ¶
resolve_direction(
trades: DataFrame,
sign_method: str | None,
quotes: DataFrame | None,
context: str,
) -> pd.DataFrame
Return trades guaranteed to carry a buy/sell direction.
Signed-flow analytics need the taker's aggressor side. L3 feeds provide
it natively; L2 / aggregated feeds don't, so synthesize it with a
trade-sign classifier (:func:classify_trade_sign).
sign_method=None— keep a nativedirectionif present, filling any row whose value is neither"buy"nor"sell"with the classifier below; otherwise classify every row with Lee–Ready when quotes are supplied, else the tick rule.sign_method="tick"/"lee_ready"— always (re)classify with that method, overriding any existingdirection.
The frame is only copied when a direction column is written, so a feed
that already labels every trade is passed straight through.
The returned column is guaranteed to hold only "buy" and "sell".
That matters because every consumer reads it as == "buy" and treats
everything else as a sell: an unlabelled trade left in place is not
dropped by them, it is counted on the wrong side.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
trades
|
DataFrame
|
Trades with at least |
required |
sign_method
|
str or None
|
|
required |
quotes
|
DataFrame or None
|
Quote frame for Lee–Ready (e.g. a pipeline |
required |
context
|
str
|
Caller name, used in the error message. |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
trades with a |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If sign_method is |
prevailing_mid ¶
prevailing_mid(
timestamps: ndarray,
quotes: DataFrame,
context: str = "classify_trade_sign",
*,
allow_exact: bool = True,
skip_crossed: bool = False,
mid_column: str | None = None,
require_covered: bool = False,
) -> np.ndarray
Midpoint prevailing at or before each (sorted) timestamp.
A backward as-of join of timestamps against quotes: each instant gets
the midpoint of the last quote published at or before it. Instants before
the first quote get NaN rather than the first quote's mid, so a caller
can tell "no quote yet" from a real number (Lee–Ready falls back to the
tick rule there; the cost metrics leave the row unmeasured).
Any of the accepted quote-column spellings works — a mid column (mid /
midprice / mid_price) or a bid/ask pair (best_bid_price /
best_ask_price, best_bid / best_ask, or bid / ask) — so
a pipeline depth_summary can be passed straight in.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
timestamps
|
ndarray
|
Instants to price, sorted ascending ( |
required |
quotes
|
DataFrame
|
Quote frame with |
required |
context
|
str
|
Caller name, used in the error message. |
'classify_trade_sign'
|
mid_column
|
str
|
Column to read the midpoint from, e.g. |
None
|
allow_exact
|
bool
|
Whether a quote stamped at exactly the same instant counts as
prevailing. |
True
|
skip_crossed
|
bool
|
Whether to drop crossed quotes — best bid above best ask — from the
reference series, so an instant standing on one reaches back to the
last quote that was not crossed. Default |
False
|
require_covered
|
bool
|
Whether an instant past the newest usable quote is |
False
|
Returns:
| Type | Description |
|---|---|
ndarray
|
The prevailing midpoint per instant, |
Raises:
| Type | Description |
|---|---|
ConfigError
|
If quotes lacks |