Skip to content

Databento

Support for Databento DBN files in the market-by-order (MBO) schema: a per-order feed with an order_id on every record, so it replays through the full L3 reconstruction.

Use via Pipeline(source=DatabentoSource(...)) or Pipeline.from_source("databento"). See the how-to for the workflow, including how to size a query before downloading it.

Key differences from LOBSTER:

  • Fixed-point prices — raw prices are integers where one unit is 1e-9 of the quote currency (price_divisor=1_000_000_000).
  • Two clocks per record — ts_recv (Databento's receive time, monotonic) becomes timestamp; ts_event (the venue's own) becomes exchange_timestamp.
  • Fills are separate records — an execution is an F record that does not change the book, followed by the C or M that takes the size off it. The loader pairs the two, so fill tells an execution apart from a cancel.
  • No date to supply — DBN carries absolute UTC timestamps, so required_context() is empty.

databento is an optional dependency:

pip install "ob-analytics[databento]"

DatabentoLoader

DatabentoLoader(
    config: PipelineConfig | None = None,
    *,
    settings: DatabentoSettings | None = None,
    venue: str | None = None,
    symbol: str | None = None,
)

Load Databento MBO records into the canonical per-order events frame.

Satisfies the :class:~ob_analytics.protocols.EventLoader protocol.

The fills the loader pairs with their book events are kept on :attr:trade_records for :class:DatabentoTradeReader to project into the trades frame — the same arrangement :class:~ob_analytics.lobster.LobsterSource uses for the companion orderbook file, and for the same reason: the second table is a by-product of reading the first, and reading the file twice to get it would be wasted work on a multi-million-record window.

Parameters:

Name Type Description Default
config PipelineConfig

Pipeline configuration. price_divisor is Databento's fixed-point scale (:data:DBN_PRICE_DIVISOR) and lot_size is normally 1, both supplied by :meth:DatabentoSource.config_defaults.

None
settings DatabentoSettings

Which instrument and publisher of a multi-instrument file to read.

None
venue str

Optional instrument identity (issue #147). When either is supplied the loaded frame gains per-row venue / symbol columns; venue falls back to "databento".

None
symbol str

Optional instrument identity (issue #147). When either is supplied the loaded frame gains per-row venue / symbol columns; venue falls back to "databento".

None

load

load(source: Any) -> pd.DataFrame

Load MBO records from source and return the events frame.

Parameters:

Name Type Description Default
source str, Path, DBNStore, or pandas.DataFrame

Anything :func:read_mbo_frame accepts.

required

Returns:

Type Description
DataFrame

Canonical events: id, timestamp, exchange_timestamp, price (integer ticks), volume and fill (integer lots), action, direction, event_id, original_number, raw_event_type (the DBN action letter) and raw_size (the record's own size field).

DatabentoTradeReader

DatabentoTradeReader(
    config: PipelineConfig | None = None,
    *,
    loader: DatabentoLoader,
    settings: DatabentoSettings | None = None,
)

Build trades from the fill and trade records of a Databento MBO file.

Satisfies the :class:~ob_analytics.protocols.TradeSource protocol.

Which record makes a trade is :attr:DatabentoSettings.trades_from:

  • Fills (F), the default. A fill names the resting order, so each one becomes a trade row with a maker and a maker_event_id, one row per resting order a sweep took out — the same per-maker-leg shape the LOBSTER and Bitstamp readers produce. A file with no fills falls back to its prints. The catch: a trade the publisher sent no fill for — an auction, a trade against a non-displayed order, an off-exchange print — has nothing to become, so it is not in the frame. The reader warns with the volume left out.
  • Trade prints (T). The venue's whole tape, one row per print, but no maker on any row.

The two are not mixed. A print and the fills behind it describe one execution, and nothing in a DBN record ties them together reliably — a fill and the modify it causes can carry different receive times — so joining them would risk counting the same volume twice.

Where the venue states the aggressor, direction is its answer rather than an estimate. Where it does not, the pipeline classifies it against the reconstructed quotes.

The taker's own order is not identified. A DBN trade record does not reliably carry the aggressing order's id, so taker and taker_event_id are NA. That is worth knowing before reading :func:~ob_analytics.analytics.set_order_types: with no taker ids it labels executed resting orders resting-limit and never market or market-limit.

Parameters:

Name Type Description Default
config PipelineConfig

Pipeline configuration.

None
loader DatabentoLoader

The loader that read the events frame. The fills come from it, so it must be the same instance, already used for the run (:class:DatabentoSource wires this up).

required
settings DatabentoSettings

Supplies :attr:~DatabentoSettings.trades_from.

None

load

load(events: DataFrame, source: Any) -> pd.DataFrame

Build the trades DataFrame for the run.

Parameters:

Name Type Description Default
events DataFrame

The loaded events frame, used to map a maker event back to its original_number.

required
source Any

Unused; the trade records came off the file with the events.

required

Returns:

Type Description
DataFrame

DatabentoWriter

DatabentoWriter(
    config: PipelineConfig | None = None,
    *,
    settings: DatabentoSettings | None = None,
)

Write an events frame back out as a DBN file of MBO records.

Satisfies the :class:~ob_analytics.protocols.DataWriter protocol, and inverts :class:DatabentoLoader record for record: a created row becomes A, a changed row M, a deleted row C removing the outstanding size, and a non-zero fill becomes an F record immediately before the row that carries it.

The point is a round trip — reading a window, working on it, and writing something another DBN reader can open — not re-creating a vendor file byte for byte. The metadata says OB.ANALYTICS (or whatever :attr:DatabentoSettings.dataset says) rather than claiming to be Databento's own data.

Parameters:

Name Type Description Default
config PipelineConfig

Pipeline configuration; its price_divisor, tick_size and lot_size invert the canonical ticks and lots.

None
settings DatabentoSettings

dataset, instrument_id and publisher_id to stamp on the records (defaulting to 1 and 1 when unset).

None

write

write(
    data: dict[str, DataFrame],
    dest: str | Path,
    *,
    symbol: str = "SYMBOL",
    **kwargs: Any,
) -> Path

Write data["events"] to dest as a DBN file.

Parameters:

Name Type Description Default
data dict of str to DataFrame

Must contain "events".

required
dest str or Path

Output file. A directory is filled with events.dbn.

required
symbol str

The raw symbol to record in the file's metadata.

'SYMBOL'

Returns:

Type Description
Path

The file written.

DatabentoSource dataclass

DatabentoSource(
    settings: SourceSettings = DatabentoSettings(),
)

The Databento source: offline replay of DBN market-by-order files (L3).

Offline only — Databento's Live client is not wired up here — so it satisfies :class:~ob_analytics.protocols.OfflineSource and has no live capability. Use it as::

from ob_analytics import Pipeline
from ob_analytics.databento import DatabentoSource, DatabentoSettings

source = DatabentoSource(settings=DatabentoSettings(raw_symbol="AAPL"))
result = Pipeline(source=source).run("xnas-itch-20240403.mbo.dbn.zst")

The defaults suit a US equity: a one-cent tick and whole shares. Another instrument needs its own, e.g. PipelineConfig(tick_size=0.25, price_decimals=2) for the E-mini S&P 500 future.

dbn_settings

dbn_settings() -> DatabentoSettings

Return :attr:settings as :class:DatabentoSettings.

Raises:

Type Description
ConfigError

If the source was built with another source's settings.

DatabentoSettings

Bases: SourceSettings

Typed settings for :class:DatabentoSource.

A DBN file can hold several instruments and, for a consolidated dataset such as DBEQ.BASIC, several publishers of the same instrument. Each (instrument, publisher) pair is a book of its own, so a run covers exactly one of them: name it here, or the loader raises and lists what the file holds. A file with one instrument and one publisher needs no setting.

Attributes:

Name Type Description
instrument_id (int, optional)

Keep only records with this instrument_id.

publisher_id (int, optional)

Keep only records from this publisher (one venue of a consolidated dataset).

raw_symbol (str, optional)

Keep only records whose mapped symbol is this. Works only on a file that carries Databento's symbol mapping; use instrument_id otherwise.

trades_from {'fills', 'prints'}

Which records the trades frame is built from. "fills" (the default) uses the F records, one row per resting order a trade hit, each naming that order — which is what order classification and queue analysis read. A file with no fills at all falls back to its prints. "prints" uses the T records: the venue's whole tape, including the auction, non-displayed and off-exchange trades that come with no fill, but with no maker named on any row. Pick "prints" when the question is about the tape — VWAP, bars, flow toxicity, costs — and the loader has warned that fills leave part of it out.

dataset str

The dataset id :class:DatabentoWriter stamps into the DBN metadata it writes. Read only on the way out; the loader takes the file's own.

read_mbo_frame

read_mbo_frame(source: Any) -> pd.DataFrame

Return the raw MBO records of source as a DataFrame, in file order.

Parameters:

Name Type Description Default
source str, Path, DBNStore, or pandas.DataFrame

A .dbn / .dbn.zst path, an already-open databento.DBNStore, or a frame of MBO records the caller read itself (store.to_df(price_type="fixed")). A frame is taken as is, which is how a caller who already sliced or filtered the records hands them straight to the loader.

required

Returns:

Type Description
DataFrame

One row per record with at least :data:_REQUIRED_MBO_COLUMNS. Prices are Databento's raw fixed-point integers, not decimals.

Raises:

Type Description
ConfigError

If the file holds a schema other than mbo, or a frame is missing required columns.

ImportError

If a path or store was given and databento is not installed.