Skip to content

Transaction cost and liquidity

What trading cost, and what liquidity was there to trade against. The quoted spread says what the book advertised; these say what a taker actually paid, and split that payment into the part the liquidity provider kept and the part the market moved:

effective spread = realized spread + price impact

Two of the measures need quotes and an aggressor side. Two read the trade prices alone, so they work on a bare tape.

See Measure transaction costs for a worked example, and Glossary: transaction cost for what each measure means.

Per-trade costs

transaction_costs

transaction_costs(
    trades: DataFrame,
    quotes: DataFrame,
    *,
    horizon: str = "1min",
    sign_method: str | None = None,
    mid_column: str | None = None,
) -> pd.DataFrame

Measure what each trade cost the taker, and where that cost went.

For a trade at price :math:p_t with aggressor side :math:D_t (+1 for a buy, -1 for a sell), against the mid-price :math:m_t prevailing when it printed and the mid-price :math:m_{t+\Delta} prevailing one horizon later:

  • effective spread :math:= 2 D_t (p_t - m_t) — the round-trip cost of crossing, as the taker experienced it. Doubled so it is comparable with a quoted spread, which also spans both sides of the mid.
  • realized spread :math:= 2 D_t (p_t - m_{t+\Delta}) — the part the liquidity provider kept, measured once the trade's information has had horizon to reach the price.
  • price impact :math:= 2 D_t (m_{t+\Delta} - m_t) — the rest: how far the trade moved the market. The two parts add back to the effective spread exactly, row by row.

The horizon matters and there is no neutral choice. Too short and the price has not finished reacting, so impact is understated; too long and unrelated moves are counted as this trade's impact. Five minutes is the equity convention; the default here is one minute, because captures of a fast crypto tape are usually measured in minutes rather than hours. Report the horizon alongside the number.

A trade in the last horizon of the quote frame has no future mid to be measured against. Rather than reuse the final quote — which would report a shrinking horizon as though it were the full one — those rows get NaN realized spread and impact, and keep their effective spread.

:math:m_t is the mid of the last quote strictly before the trade. When the quote frame is a depth_summary built from the same event stream, the row stamped at the trade's own instant is the book after the trade took the touch, and measuring against it would charge the taker nothing for a move they caused. Venue timestamps are coarse, so a quote a few events older is the closest honest reference available. Crossed quotes are skipped for the same reason: a book whose best bid is above its best ask has no midpoint, so the last uncrossed quote is used instead.

Every number here inherits the quality of the book it is measured against. On a diff feed a trade can print through resting orders the venue never withdrew, which reads as a negative effective spread — the taker apparently paying less than the mid. That is a finding about the capture, not about the market: run :func:~ob_analytics.analytics.detect_stale_orders (or the audit command) before reading the costs as execution quality.

Parameters:

Name Type Description Default
trades DataFrame

Trades with timestamp, price and volume. A direction column ("buy" / "sell", the taker side) is used when present; otherwise it is inferred — see sign_method.

required
quotes DataFrame

Quote frame supplying the mid-price: timestamp plus either a mid column (mid / midprice / mid_price) or a bid/ask pair (best_bid_price / best_ask_price, best_bid / best_ask, or bid / ask). A pipeline depth_summary satisfies this.

required
horizon str

Pandas offset string for :math:\Delta, the wait before the realized spread is read. Default "1min".

'1min'
sign_method str

How to obtain the aggressor side when there is no native direction. None (default) keeps a native direction and otherwise classifies with Lee-Ready against quotes. "tick" / "lee_ready" force that classifier, overriding a native direction. See :func:~ob_analytics.trade_sign.classify_trade_sign.

None
mid_column str

Column of quotes to take the reference price from. None (default) uses the plain mid. Pass "micro_price" to measure against the size-weighted mid instead (:func:~ob_analytics.depth.micro_price, present on a frame from :func:~ob_analytics.depth.depth_signals): the micro-price is the better forecast of where the price is going, so an effective spread measured against it charges the taker for crossing but not for the move the book was already leaning toward. The two answer different questions, so the choice is the caller's.

None

Returns:

Type Description
DataFrame

One row per trade, chronological, with the columns in :data:COST_COLUMNS. mid_price and future_mid_price are the two mids the measures were taken against, kept so a number can be traced back to the quotes that produced it. The horizon is recorded in frame.attrs["horizon"].

Raises:

Type Description
ConfigError

If a required column is missing, or sign_method is "bvc", which labels volume bars rather than trades.

ObAnalyticsError

If trades is empty.

Examples:

>>> from ob_analytics import Pipeline, sample_csv_path, transaction_costs
>>> result = Pipeline().run(sample_csv_path())
>>> costs = transaction_costs(result.trades, result.depth_summary, horizon="5s")
>>> costs["effective_spread_bps"].median()
0.1276...

cost_summary

cost_summary(costs: DataFrame) -> CostSummary

Reduce a :func:transaction_costs frame to one volume-weighted figure each.

Each measure is averaged over the trades that actually carry it, weighted by trade size: the effective spread over every trade with a mid, the realized spread and the impact over the smaller set that also had a mid a horizon later. A measure with no usable trade is NaN rather than an error, so a summary of a capture shorter than its own horizon still returns and says so through :attr:~CostSummary.n_realized.

Parameters:

Name Type Description Default
costs DataFrame

The frame :func:transaction_costs returned. The horizon is read from costs.attrs["horizon"] when present.

required

Returns:

Type Description
CostSummary

Raises:

Type Description
ConfigError

If costs is missing a column :func:transaction_costs writes.

ObAnalyticsError

If costs is empty.

Price-only liquidity

amihud

amihud(
    trades: DataFrame, *, window: str | None = None
) -> pd.DataFrame

Amihud (2002) illiquidity: price move per unit of turnover.

::

amihud = |return| / turnover

Read it as the price response a unit of trading buys. A market where a large turnover barely moves the price is deep, and scores low; one where a small turnover swings it is thin, and scores high. It needs no quotes and no aggressor side, so it is the one liquidity measure available on a tape that carries nothing but prices, sizes and times.

The return is taken within each window, from its first trade price to its last, and turnover is the price times size summed over the same trades. Amihud's original is a daily close-to-close return over that day's volume; a capture is not a series of trading days, and taking each window's own first and last print keeps every window usable — including the first, which a close-to-close definition cannot measure — and makes the whole-session case (window None) the same calculation over one window rather than a different one.

The result carries a unit: one over the turnover unit of the input, which on a pipeline frame is ticks times lots. It is comparable across windows of one run, and across runs of one instrument, but not across instruments without rescaling. Published figures are multiplied by a power of ten to reach readable digits; this returns the raw ratio and leaves that choice to the caller.

Parameters:

Name Type Description Default
trades DataFrame

Trades with timestamp, price and volume.

required
window str or None

Pandas offset string for the window, e.g. "5min". None (default) measures the whole session as a single window and returns one row.

None

Returns:

Type Description
DataFrame

One row per window with timestamp (the window's start), first_price, last_price, abs_return, turnover, n_trades and amihud. Windows with no trade are dropped; a window whose turnover is zero has NaN for amihud.

Raises:

Type Description
ConfigError

If a required column is missing.

ObAnalyticsError

If trades is empty.

roll_spread

roll_spread(
    trades: DataFrame, *, window: str | None = None
) -> pd.DataFrame

Roll (1984): the spread implied by bid-ask bounce in the price series.

A tape with no quotes still shows the spread, because consecutive trades alternate between hitting the bid and lifting the ask, and that bounce makes successive price changes negatively correlated. Roll turns the size of that correlation back into a spread::

roll_spread = 2 * sqrt(-cov(dp_t, dp_{t-1}))

where dp is the change in trade price.

The model assumes an efficient price plus a bounce of constant half-spread, with order flow that carries no information. Under it the bounce is the only source of price change, so the lag-1 autocorrelation of the changes is exactly -0.5. That is the number to check: autocorrelation is returned beside the estimate, and the further it sits from -0.5, the less of the price movement the bounce explains.

Two things push it away. Real flow carries information, which raises the autocovariance and makes Roll read low against a measured effective spread. More decisively, the efficient price moves between trades, and on a sparse tape in a volatile instrument it moves much further than half a spread — the bounce then accounts for a few per cent of the variance and the autocovariance comes out positive, where the formula has no real root. The estimate is NaN there rather than a number the model does not support, and the two diagnostic columns say why. Sampling more often than the tape trades will not help: the fix is a denser tape or a wider spread, and where quotes exist :func:transaction_costs measures the spread directly instead of inferring it.

Parameters:

Name Type Description Default
trades DataFrame

Trades with timestamp and price, chronological or not.

required
window str or None

Pandas offset string for the window, e.g. "5min". None (default) estimates over the whole session and returns one row.

None

Returns:

Type Description
DataFrame

One row per window with timestamp (the window's start), n_trades, mean_price, autocovariance, autocorrelation, roll_spread and roll_spread_bps. autocorrelation is the lag-1 autocorrelation of the price changes, which Roll's model puts at -0.5; a window whose value is far from that is one the model does not describe. A window with fewer than four trades — too few for two overlapping price changes — or a non-negative autocovariance has NaN for both spread columns.

Raises:

Type Description
ConfigError

If a required column is missing.

ObAnalyticsError

If trades is empty.

Models

CostSummary dataclass

CostSummary(
    effective_spread: float,
    realized_spread: float,
    price_impact: float,
    effective_spread_bps: float,
    realized_spread_bps: float,
    price_impact_bps: float,
    horizon: str,
    n_trades: int,
    n_realized: int,
    volume: float,
)

One run's transaction costs, volume-weighted.

Each figure weights a trade by its size, so it reports what the average unit traded cost, not what the average trade cost. That is the convention the execution-cost literature reports (Bessembinder 2003), and it is the number a taker sizing an order needs: one large trade at a wide spread costs more than ten small ones at a narrow spread, and an unweighted mean would say the opposite.

The realized-spread and price-impact figures are averaged over the trades that have a future mid to measure against, which is fewer than the effective-spread figure has — the last horizon of the capture has no future left in it. :attr:n_realized says how many that was, so a summary computed on a short capture cannot quietly look like a full one.

Attributes:

Name Type Description
effective_spread float

Volume-weighted effective spread, in the price unit of the trades.

realized_spread float

Volume-weighted realized spread, same unit.

price_impact float

Volume-weighted price impact, same unit. Equal to effective_spread - realized_spread only when both averages cover the same trades, which is why it is measured rather than subtracted.

effective_spread_bps, realized_spread_bps, price_impact_bps float

The same three in basis points of the mid-price.

horizon str

The realized-spread horizon the costs were measured at.

n_trades int

Trades with a mid-price to measure the effective spread against.

n_realized int

Trades that also had a mid-price horizon later, so the realized spread and the impact could be measured.

volume float

Total size of the :attr:n_trades trades — the weight behind the effective-spread figure.