Skip to content

Feature table

One tidy table: a point in time on each row, a microstructure feature in each column. Two decisions make it, and they are separate — where the rows fall, and what each column measures.

Where the rows fall is a bar rule, so features() takes the same sampling arguments bars() does and the same arguments give the same cut in both. What each column measures is a feature, and ten ship with the package:

Feature Columns Reads
price open, high, low, close, vwap the bar's trades
returns log_return, realized_vol the bar's trades
flow volume, turnover, n_trades, signed_volume, trade_imbalance the bar's trades
spread spread, spread_bps the book at the close
mid_price mid_price the book at the close
micro_price micro_price, micro_price_offset the book at the close
imbalance obi, obi_depth the book at the close
depth best_bid_vol, best_ask_vol, bid_depth, ask_depth the book at the close
vpin vpin the last 20 bars
kyle_lambda kyle_lambda the last 20 bars

Each row is stated as of the close of its bar. The trade columns hold what happened inside the bar and the book columns hold the book as it stood when it closed, so nothing from later reaches the row and the table carries no look-ahead. It carries no target either: a target looks forward, and building one is a shift the caller makes.

See the "Build a feature table" how-to for a worked example and a baseline model, and Extending for writing a feature of your own.

Functions

features

features(
    trades: DataFrame,
    quotes: DataFrame | None = None,
    rule: str = "time",
    threshold: Any = None,
    *,
    target_bars: int = 50,
    include: Sequence[str] | None = None,
    sign_method: str | None = None,
) -> pd.DataFrame

Build a feature table from trades and the book in quotes.

The trades are cut into bars by rule and threshold — the same cut :func:~ob_analytics.bars.bars makes from the same arguments — and each bar becomes one row. Every registered feature that the inputs support then writes its columns onto that row.

Parameters:

Name Type Description Default
trades DataFrame

Trades with at least timestamp, price and volume — a pipeline result's trades frame, or any frame shaped like it. A direction column ("buy" / "sell", the taker's side) is used when present; otherwise the aggressor side is classified — see sign_method.

required
quotes DataFrame

Book snapshots supplying the book columns: timestamp plus the touch (best_bid_price, best_bid_vol, best_ask_price, best_ask_vol) and any depth bins. A pipeline depth_summary satisfies this. None (the default) leaves the book features out and measures the trades alone.

None
rule str

Registered bar rule deciding where the rows fall: "time" (the default), "tick", "volume", "dollar" or "imbalance". See :func:~ob_analytics.bars.list_bar_rules.

'time'
threshold optional

How much of the rule's own quantity closes a bar. None (the default) asks the rule for a threshold that yields about target_bars rows.

None
target_bars int

How many rows the default threshold aims at. Ignored when threshold is given.

50
include sequence of str

The features to measure, in the order their columns are written. None (the default) measures every registered feature the inputs support: the ten in :data:DEFAULT_FEATURES first, then any further registration, sorted by name. A feature named here whose inputs are missing raises; one selected by default is skipped instead.

None
sign_method str or None

How to classify the aggressor side: None (the default) keeps a direction column if the trades have one and otherwise classifies, "tick" or "lee_ready" always classify. See :func:~ob_analytics.trade_sign.resolve_direction.

None

Returns:

Type Description
DataFrame

One row per bar, in time order. The first three columns are :data:INDEX_COLUMNS — bar, timestamp_start and timestamp — and the rest are the features' own, in the order they were measured.

timestamp is the close of the bar, and the instant the whole row is stated as of: the trade columns hold what happened between timestamp_start and it, and the book columns hold the book as it stood at it. Nothing later reaches the row, whichever rule cut the bars, so the table carries no look-ahead.

A row with no readable quote behind it has no book to read, and a trailing window with too little history behind it has nothing to measure; both are NaN rather than a filled-in value. Two quote states are not readable as a book and are skipped rather than taken at face value — an empty side, and a crossed one. See :func:readable_quotes.

The frame's attrs carry bar_rule and bar_threshold, the cut the rows were made on; features, the names measured; and features_skipped, the names left out for want of their inputs.

Raises:

Type Description
ConfigError

If required columns are missing, if the threshold does not suit the rule, if include is a bare string or names a feature twice, if a feature named in include cannot read what it needs, or if two features would write the same column.

KeyError

If rule, or a name in include, is not registered; the message lists the registered names.

ObAnalyticsError

If trades is empty.

Notes

Prices and sizes pass through in the units they arrive in. A pipeline result holds prices as whole ticks and sizes as whole lots, so spread is in ticks and vwap is a tick count; spread_bps, obi and trade_imbalance are ratios and do not depend on the unit.

The table is features only. A model also needs a target, and a target looks forward: build one from the table with a negative shift, which is the one place look-ahead belongs.

No row reads data from after its own close, whatever the threshold. The choice of threshold is another matter: left to default it is worked out from the whole trades frame, so where the boundaries fall depends on the whole capture. Pass a threshold when the cut itself has to be something the rows could have been given at the time.

Examples:

>>> from ob_analytics import Pipeline, features, sample_csv_path
>>> result = Pipeline().run(sample_csv_path())
>>> table = features(
...     result.trades, result.depth_summary, "volume", 100
... )
>>> table[["timestamp", "close", "spread_bps", "obi", "vpin"]]

register_feature

register_feature(feature: Feature) -> None

Register feature under its own :attr:~ob_analytics.protocols.Feature.name.

Case-insensitive; overwriting an existing registration is allowed, so a feature of your own may deliberately shadow a built-in one.

list_features

list_features() -> list[str]

Return a sorted list of registered feature names.

get_feature

get_feature(name: str) -> Feature

Return the feature registered under name (case-insensitive).

Raises:

Type Description
KeyError

If no feature is registered under name; the message lists the registered names.

readable_quotes

readable_quotes(quotes: DataFrame) -> pd.DataFrame

Return the rows of quotes whose book can be read as a price.

Two states get through a depth summary that are not books anything could have traded against, and both would otherwise reach a row as ordinary numbers:

A side with nothing resting on it. The depth engine writes a price and a volume of 0 for an empty side, which is a marker and not a price. Taken at face value it makes a spread the width of the whole instrument, a mid at half the other side, and a micro-price of zero — three finite numbers, none of them true, and none of them marked.

A crossed book, where the best bid is above the best ask. A diff feed can hold genuinely crossed resting orders, so this is an expected state on such a feed rather than a fault, but its midpoint is not a price and its spread is negative. The test is bid > ask, the same one :func:~ob_analytics.trade_sign.prevailing_mid applies: a locked book, bid equal to ask, is a real state at a spread of zero and is kept.

Dropping these from the reference series is what makes a bar reach back to the last quote that could be read, the way it reaches back over any other instant with no quote of its own. A frame carrying no best_bid_price / best_ask_price pair cannot be tested and is returned unchanged.

Parameters:

Name Type Description Default
quotes DataFrame

Book snapshots — a pipeline depth_summary, or any frame shaped like one.

required

Returns:

Type Description
DataFrame

The readable rows, in their original order.

The built-in features

PriceFeature

The bar's four prices and its volume-weighted average price.

open / high / low / close Trade prices within the bar, in the units of the trades frame. vwap Turnover divided by volume.

compute

compute(frame: DataFrame) -> dict[str, npt.ArrayLike]

Return the bar's price columns unchanged.

ReturnsFeature dataclass

ReturnsFeature(
    name: str = "returns", window: int = DEFAULT_WINDOW
)

The bar's return, and how much returns have been moving lately.

log_return log(close / previous close). The first bar has no previous close, so it is NaN; so is a bar whose close is not a positive price. realized_vol Standard deviation of log_return over the last :attr:window bars, the row's own included. It is per bar and not annualized: only clock bars span equal amounts of time, so there is no one factor that would scale it to a year.

Attributes:

Name Type Description
name str

Registered feature name.

window int

How many bars realized_vol looks back over.

compute

compute(frame: DataFrame) -> dict[str, npt.ArrayLike]

Return the log return per bar and its trailing standard deviation.

FlowFeature

How much traded in the bar, and how one-sided it was.

volume / turnover / n_trades Size traded, price × size traded, and the number of trades. signed_volume Buyer-initiated volume minus seller-initiated volume. trade_imbalance signed_volume / volume — the signed share of the bar's volume, from -1 (every trade a sell) to +1 (every trade a buy).

compute

compute(frame: DataFrame) -> dict[str, npt.ArrayLike]

Return the bar's flow columns and the imbalance they imply.

SpreadFeature

What it cost to cross the book at the bar's close.

spread Best ask price minus best bid price, in the units of the quotes frame. spread_bps The same as a share of the mid-price, in basis points.

compute

compute(frame: DataFrame) -> dict[str, npt.ArrayLike]

Return the spread in price units and in basis points.

MidPriceFeature

The mid-price at the bar's close: the average of the two best prices.

compute

compute(frame: DataFrame) -> dict[str, npt.ArrayLike]

Return the plain mid-price per bar.

MicroPriceFeature

The mid-price weighted by the size resting on the opposite side.

micro_price The size-weighted mid from :func:~ob_analytics.depth.micro_price. It leans toward the side carrying the heavier opposite book, which is the direction price is more likely to move. micro_price_offset micro_price - mid_price: how far it leans, in price units. This is the part a model wants — the micro-price itself tracks the mid so closely that the two say almost the same thing.

compute

compute(frame: DataFrame) -> dict[str, npt.ArrayLike]

Return the micro-price and its distance from the mid.

ImbalanceFeature dataclass

ImbalanceFeature(
    name: str = "imbalance", depth_levels: int = 5
)

The signed share of resting volume on the bid side.

obi Book imbalance at the touch — :func:~ob_analytics.depth.book_imbalance with levels=1. obi_depth The same measured over :attr:depth_levels, so it reads the pressure behind the touch as well as on it. The level count is clamped to the depth bins the quotes frame actually carries, so a summary with fewer bins gives a shallower reading rather than an error.

Attributes:

Name Type Description
name str

Registered feature name.

depth_levels int

Depth for obi_depth, counted the way :func:~ob_analytics.depth.book_imbalance counts it: 1 is the touch alone and each further level adds one depth bin.

compute

compute(frame: DataFrame) -> dict[str, npt.ArrayLike]

Return the touch imbalance and the deeper one.

DepthFeature

How much size was resting at the bar's close.

best_bid_vol / best_ask_vol Size at the touch on each side. bid_depth / ask_depth Size within every depth bin the quotes frame carries — the innermost bin already includes the touch. A quotes frame with no depth bins reports the touch size, which is all the depth it knows about.

compute

compute(frame: DataFrame) -> dict[str, npt.ArrayLike]

Return the touch size and the measured depth on each side.

VpinFeature dataclass

VpinFeature(
    name: str = "vpin", window: int = DEFAULT_WINDOW
)

How one-sided the flow has been over the last :attr:window bars.

vpin The mean of the absolute trade imbalance over the window. It runs from 0 (every bar balanced) to 1 (every bar one-sided), and a high reading says the flow has been persistently directional, which is what a market maker is exposed to.

Sampled on volume bars this is VPIN, the volume-synchronized probability of informed trading of Easley, López de Prado and O'Hara, whose buckets are equal amounts of traded volume. On another rule it is the same measure read on that rule's clock. Either way the imbalance is divided by the bar's own volume, so a bar that traded less is not read as a quieter one.

The reading depends on how many trades a bar holds. A bar of one or two trades is almost always all buying or all selling, so its imbalance is close to 1 whatever the flow is doing, and a window of such bars reads close to 1 too. The original measure uses buckets of about a fiftieth of a day's volume, each holding many trades. Choose bars that each hold many trades before reading this column as toxicity; n_trades in the same table shows how many each bar holds.

:func:~ob_analytics.flow_toxicity.compute_vpin computes VPIN without a bar table, and cuts its buckets slightly differently: a trade that straddles a bucket boundary is split between the two, where a volume bar keeps the trade whole and closes a little past its threshold. The two readings agree only when the bars and the buckets are about the same size.

Attributes:

Name Type Description
name str

Registered feature name.

window int

How many bars the mean covers — the bucket count of the original measure, which the paper puts at 50.

columns tuple of str

The one column, named after :attr:name, so a second window registered under another name writes a column of its own rather than overwriting this one.

compute

compute(frame: DataFrame) -> dict[str, npt.ArrayLike]

Return the trailing mean absolute trade imbalance.

KyleLambdaFeature dataclass

KyleLambdaFeature(
    name: str = "kyle_lambda",
    window: int = DEFAULT_WINDOW,
    min_bars: int = 3,
)

How far price moved per unit of net order flow, lately.

kyle_lambda The slope of close - open on signed_volume, fitted over the last :attr:window bars. A high value says a given imbalance pushed the price a long way, which is a book that was expensive to trade against. It is in the price units of the trades frame per unit of volume: on a pipeline result that is ticks per lot, so multiply by config.tick_size to read it in the quote currency.

This is Kyle's λ, estimated on a trailing window rather than over a whole run. :func:~ob_analytics.flow_toxicity.compute_kyle_lambda fits the same regression once across every window of a run, and reports the t-statistic and R² that say whether the fit means anything; a rolling slope reports neither, so read it as a level that moves, not as a test.

Attributes:

Name Type Description
name str

Registered feature name.

window int

How many bars each fit covers.

min_bars int

The fewest bars that produce a slope. Earlier rows are NaN.

columns tuple of str

The one column, named after :attr:name, so a second window registered under another name writes a column of its own rather than overwriting this one.

compute

compute(frame: DataFrame) -> dict[str, npt.ArrayLike]

Return the trailing regression slope per bar.