Build a feature table¶
A model or a study wants one tidy table: a point in time on each row and a
microstructure feature in each column. The library measures all of those
already, but each in its own table on its own clock.
features() does the join once.
from ob_analytics import Pipeline, features, sample_csv_path
from ob_analytics.visualization import display_result
result = Pipeline().run(sample_csv_path())
# Pipeline prices are whole ticks and sizes whole lots; display_result
# converts both to the quote currency and the base asset.
shown = display_result(result)
trades, quotes = shown.trades, shown.depth_summary
table = features(trades, quotes, "volume", 0.05)
print(table[["timestamp", "close", "spread_bps", "obi", "trade_imbalance", "vpin"]])
timestamp close spread_bps obi trade_imbalance vpin
2026-05-02 02:36:23.889000+00:00 78319.0 0.127684 0.772174 1.0 1.0
2026-05-02 02:36:23.918000+00:00 78319.0 1.915085 -0.216647 1.0 1.0
2026-05-02 02:36:23.924000+00:00 78319.0 1.915085 -0.216647 1.0 1.0
2026-05-02 02:36:23.936000+00:00 78320.0 1.915085 -0.216647 1.0 1.0
2026-05-02 02:36:23.936000+00:00 78321.0 1.915085 -0.216647 1.0 1.0
Two decisions make the table, and they are separate: where the rows fall, and what each column measures.
Where the rows fall¶
The rows are bars, and features() takes the same sampling
arguments bars() does — the rule, the threshold, and target_bars when you
leave the threshold out. The same arguments give the same cut in both.
features(trades, quotes, "time", "1min") # a time grid
features(trades, quotes, "tick", 100) # every 100 trades
features(trades, quotes, "volume", 0.5) # every 0.5 BTC traded
features(trades, quotes, "dollar", 50_000) # every $50,000 of turnover
features(trades, quotes, "imbalance", 0.5) # every 0.5 BTC of net flow
features(trades, quotes, target_bars=200) # let the rule choose
Activity rules are usually the better sampling for a model: a quiet hour and a busy minute produce the same number of rows, so the rows come much closer to being independent and identically distributed. The bars how-to explains the five rules.
The cut is recorded, so a table built with a defaulted threshold still says what it was built on:
table.attrs["bar_rule"] # 'volume'
table.attrs["bar_threshold"] # 0.05
table.attrs["features"] # the feature names measured
What each column measures¶
Each column comes from a feature. Ten ship with the package:
| Feature | Columns | Reads |
|---|---|---|
price |
open, high, low, close, vwap |
trades |
returns |
log_return, realized_vol |
trades |
flow |
volume, turnover, n_trades, signed_volume, trade_imbalance |
trades |
spread |
spread, spread_bps |
book |
mid_price |
mid_price |
book |
micro_price |
micro_price, micro_price_offset |
book |
imbalance |
obi, obi_depth |
book |
depth |
best_bid_vol, best_ask_vol, bid_depth, ask_depth |
book |
vpin |
vpin |
trades |
kyle_lambda |
kyle_lambda |
trades |
Name the ones you want, in the order you want their columns:
features(trades, quotes, "volume", 0.5, include=["spread", "imbalance"])
# bar, timestamp_start, timestamp, spread, spread_bps, obi, obi_depth
Leave include out and every registered feature the inputs support is
measured. Without a quotes frame the five book features are skipped and the
table holds the trade features alone:
table = features(trades, None, "volume", 0.5)
table.attrs["features_skipped"]
# ['spread', 'mid_price', 'micro_price', 'imbalance', 'depth']
Naming a book feature without quotes raises instead: an explicit request that cannot be met is an error, while a default that cannot be met is a smaller table.
VPIN needs bars that hold many trades¶
vpin is the average, over the last 20 bars, of how one-sided each bar's
trades were. A bar that holds one or two trades is nearly always all buying or
all selling, so it reads close to 1 whatever the flow is doing. The example
above cuts a bar every 0.05 BTC, which on the sample is about two trades a bar,
and its vpin column sits near 1 on every row. That says the bars are small,
not that the flow is toxic.
On the bundled sample:
| Rule | Trades per bar | Median vpin |
|---|---|---|
volume, 0.05 BTC |
2.2 | 1.00 |
time, 1 minute |
9.5 | 0.96 |
volume, 0.5 BTC |
12.3 | 0.87 |
tick, 20 trades |
18.9 | 0.78 |
volume, 2 BTC |
40.6 | 0.72 |
The value falls as the bars grow, so compare vpin only between tables cut
the same way. The original measure uses buckets of about a fiftieth of a day's
volume, which on a liquid market holds many trades each. Check n_trades
before you read vpin as toxicity. compute_vpin
behaves the same way: with buckets the size of these bars, its readings on
the sample are within 0.06 of the column's.
No look-ahead¶
Each row is stated as of the close of its bar, which is the timestamp
column. The trade columns hold what happened between timestamp_start and
that instant, and the book columns hold the book as it stood at it — a
backward as-of join, so a row takes the last quote published at or before its
close. Nothing from later reaches the row.
The trailing-window features — realized_vol, vpin, kyle_lambda — look
back over the last 20 rows, the row's own included, and are NaN until there
is enough history behind them. A row with no readable quote behind it has no
book to read, and is NaN too.
Quotes that are not books¶
Two states reach a depth summary that nothing could have traded against, and both would otherwise arrive in the table as ordinary numbers:
- A side with nothing resting on it. The depth engine writes a price and a
volume of
0for an empty side. Read as a price that gives a spread the width of the instrument, a mid at half the other side, and a micro-price of zero — three finite numbers, none of them true. - A crossed book, where the best bid is above the best ask. A diff feed can hold genuinely crossed resting orders, so on such a feed this is an expected state rather than a fault; but its spread is negative and its mid is not a price.
Both are dropped from the reference series, so a row reaches back to the last
quote that could be read — the same thing it does for any other instant with
no quote of its own. A locked book, bid equal to ask, is a real state at a
spread of zero and is kept. The same test is what
transaction_costs() applies before measuring against
a mid.
One thing is not settled by the table: the choice of threshold. Left to default it is worked out from the whole trades frame, so where the boundaries fall depends on the whole capture. Pass a threshold when the cut itself has to be something the rows could have been given at the time.
Build a target¶
The table is features only. A target looks forward, and building one is the one place look-ahead belongs — so it is yours to write, deliberately:
The last row then has no target, and rows whose trailing windows are still
filling have no features. Drop both together with .dropna().
A baseline model¶
Enough to see the shape. Six features, ordinary least squares, and a chronological split — never a random one, because shuffling rows lets the model learn from the future:
import numpy as np
table = features(trades, quotes, "volume", 0.05)
table["target"] = table["log_return"].shift(-1)
columns = [
"obi",
"obi_depth",
"micro_price_offset",
"trade_imbalance",
"spread_bps",
"vpin",
]
fit = table[[*columns, "target"]].dropna()
split = int(len(fit) * 0.7)
train, test = fit.iloc[:split], fit.iloc[split:]
def design(frame):
return np.column_stack([np.ones(len(frame)), frame[columns].to_numpy(float)])
beta, *_ = np.linalg.lstsq(design(train), train["target"].to_numpy(float), rcond=None)
def r_squared(frame):
actual = frame["target"].to_numpy(float)
predicted = design(frame) @ beta
return 1 - ((actual - predicted) ** 2).sum() / ((actual - actual.mean()) ** 2).sum()
print(f"in sample R² {r_squared(train):+.3f} ({len(train)} rows)")
print(f"out of sample R² {r_squared(test):+.3f} ({len(test)} rows)")
Read that as a demonstration of the shape, not as a finding. The bundled sample is ten minutes of one instrument, so 38 test rows can say almost anything; and the split is one split, not a walk forward. What it does show is the workflow: sample on activity, measure as of each row's close, shift a target back by one row, and split in time.
A feature of your own¶
A feature says what one column measures and nothing else. Register one and
features() puts it in the table — see
Extending.
Somewhere else¶
The table is a plain pandas frame, so it goes wherever one does:
Polars is not a dependency of ob-analytics; install it yourself to run that second line. See the schema page for what the package guarantees about frames and files.