Changelog¶
All notable changes to ob-analytics are documented in this file.
The format is based on Keep a Changelog.
[Unreleased]¶
Added¶
-
A live capture can run for days (#150).
ob-analytics capturenow writes a directory of segments with amanifest.json. A lost connection ends the segment; the capture waits (1 s, doubling to 60 s) and starts a new one from a fresh snapshot.--roll-minutesand--roll-mbstart a new segment by time or size, and the new segment is streaming before the old one stops, so a roll loses nothing. Running the same command again continues the capture, and closes a segment a crashed process left open; a lock on the directory stops a second process from writing to it at once. The manifest records every segment, why it ended, and every gap with its cause.processandauditread the whole capture, andauditadds the checkscapture_gaps,unfinished_segmentsanddropped_messages. From Python:ob_analytics.live.run_captureandread_manifest. -
Each source declares which orders of a trade it can name (#284). The new
Source.trade_attribution(TradeAttribution.BOTH,MAKER_ONLYorNONE) says whether a feed's order events show the trade's maker, its taker, or neither.bitstampnames both.lobster,databentoandcryptofeedat L3 name the maker only: their feeds show resting orders, and a taker trades on arrival. The L2 sources name neither. Theunmatched_tradescheck now counts only the orders the feed can name, and says which. Before, a Databento run could only report every trade as unmatched, and LOBSTER's guessed takers counted as matches. A source that does not declare it is read asBOTH, the old check. - A capture records what its source declares, and
auditholds it to that (#284).meta.jsonnow recordssource,feed_typeandtrade_attribution.ob-analytics processcopiesmeta.jsoninto its output.auditchecks a capture against the declarations of the source that made it, even when--sourcenames another: a cryptofeed L3 capture is read with--source bitstamp, whose feed shows more. Read the record withrecorded_source,recorded_feed_typeandrecorded_trade_attribution.
Changed¶
- The capture directory layout (#150). A capture's files are now in
seg-0001/,seg-0002/, ... under--out, next tomanifest.json. Read one segment as before, from itsorders.csvordepth.csv; giveprocessandauditthe capture directory to read them all.--outmust be new, empty, or a capture of the same venue and pair. - Live sources no longer reconnect by themselves (#150). The Bitstamp
source reconnected without a new snapshot, so an order deleted while it was
disconnected stayed on the book to the end of the run, and only a
reconnectscount (now removed) recorded it. The ccxt source stopped its book or trade loop on a network error and ran on with half its feed. Both now end the stream with the error, and the capture starts a new segment.
Fixed¶
- A Bitstamp capture no longer keeps trades from before its snapshot
(#301). It already skipped order messages from before the REST snapshot, but
kept the trades from the same time. The orders those trades filled are not in
the capture, so
auditcounted the trades as unmatched: up to 52% of a segment's trades with 30-second segments. The trades are now skipped and counted aspre_snapshot_trades_skippedinmeta.json. At a roll, the previous segment already has these trades. - On Python 3.11, a capture stops each segment when asked (#296). The
live sources waited for messages with
asyncio.wait_for, which on Python 3.11 can drop a cancel that arrives with a message. A roll then left the old segment streaming next to the new one to the end of the capture, with no further rolls, and SIGTERM or Ctrl-C did nothing. The sources now useasyncio.timeout, and ruff bansasyncio.wait_forin this repository. The runner cancels a stream again while it keeps yielding items, so a plug-in source with the same pattern cannot block a stop; a stream that is closing its connection is left to finish. Python 3.12 and later were not affected. - A segment that does not stop cannot hold up the capture (#296). A segment asked to stop has 20 seconds to close. After that it is cancelled and closed from its files, like a segment a crash left open. The manifest keeps why it was stopped and records the delay as its error, not as a gap. If closing its files fails, the error says so and the capture carries on. A segment covers the market only until it is asked to stop, and one asked to stop before its first live event covers nothing, so it no longer adds a second gap next to the real one. At the end of a capture all running segments are stopped together, so SIGTERM ends a capture in under a minute even when a source hangs.
-
auditno longer fails a complete Databento file (#298). Databento numbers every message on the venue's channel, so the numbers in one instrument's events skip the other instruments' messages and the trade and fill records, which become trades.auditread every skip as a lost message and failed a sound file with asequence_gapserror.DatabentoSourcenow declaressequence_kind = SequenceKind.MONOTONIC, soauditchecks only that the numbers never go back.auditreads a source's declaredsequence_kindwhen no capturemeta.jsonrecords one; read it with the newsequence_kind_of, and pass it torecorded_sequence_kindas its newdefault. -
A Bitstamp L3 capture through cryptofeed passes
audit(#284). cryptofeed's Bitstamp L3 channel isdetail_order_book: a picture of the top 100 bids and top 100 asks about 10 times a second, not every order. Four defects followed from that: - An order that dropped past the 100th place was recorded as deleted, and as created again when it came back. It now stays tracked at its last size until the book shows its price again. This removes the duplicate created ids, and the clock errors the Bitstamp reader made from them.
- Book rows carried the venue's time in both clock columns.
timestampis now the time the capture received the message, like the trades. - Trades dropped the maker and taker order ids that Bitstamp sends. They are now read from the raw message.
- A fill showed only as a size change between two pictures, so a fully filled order read as cancelled. A trade that names a tracked order now reports its fill as it arrives, and a filled order is deleted at size 0, as the native Bitstamp feed reports it. A picture that shows an order gone before its trade arrives holds the delete for up to 2 seconds for it.
On a 90-second capture, makers now link on 40 of 42 trades, up from 16. For
analysis, the native bitstamp source still shows more: every order,
takers included.
- A capture that fails now exits non-zero and says why (#283).
ob-analytics captureused to exit 0 and write"errors": 0tometa.jsonwhen the stream raised, or when the source's optional extra was missing. A source is now checked before the capture starts: a missing extra, or an unknown ccxt or cryptofeed venue, stops the run with status 1 and creates no output directory. An error in the snapshot, the stream or the shutdown events keeps the rows already written, is recorded inmeta.jsonascapture_errorandcapture_error_phaseand counted inerrors, and makes the command exit 1. A source can take part in the early check by adding apreflight()method (SupportsPreflight). - A Binance capture no longer loses price levels near the top of the book
(#101). ccxt deletes the levels of a Binance book that fall past the depth
it is given, and Binance sends a level again only when it changes. After
the price moved away and back, the captured book had holes in its top 20
levels. The ccxt source now asks for the whole Binance book (see #275
below for what it does with that book) and never gives ccxt less than
1,000 levels a side to track, so a level rarely falls out of what ccxt
itself knows. ccxt's opening snapshot is now as deep as
--depth-limit(never less than 1,000 levels), so a Binance capture can record up to the 5,000 levels a side that Binance sends, about 1% from the price. A deeper--depth-limitis refused. auditno longer fails every ccxt capture with missing sequence numbers (#101). The ccxtnonceonly rises: a Binance diff covers a range of update IDs, and ccxt can apply several diffs before it returns a book. A capture now records"sequence_kind": "monotonic"inmeta.json, andauditthen checks only that the number never goes back. Captures made before this change record no kind and are still read as contiguous.- ccxt book rows are stamped with the time they arrived (#101). The
timestampof a ccxtdepth.csvrow was the venue's book time, while its trades used the receive time. ccxt stamps its first Binance book with its own snapshot's time and then applies older diffs, so sorting on that time swapped two updates and left a stale level:auditreported the book crossed for 90% of a session.timestampis now the receive time, and the venue's time is kept in a newexchange_timestampcolumn. - A ccxt book that loses sync is fetched again (#101). When ccxt finds a
missing Binance diff it drops its book and raises an error, which used to
end the book for the rest of the capture while trades went on. The capture
now asks for the book again (up to 10 times) and counts it as
book_resyncsinmeta.json. - A venue that refuses your location gives a one-line error. Binance
answers HTTP 451 from some countries.
capture ccxtnow says so and names--exchange binanceusand--market-data-mirror, instead of printing a traceback. - A level that only left
--depth-limit, not the book, is no longer recorded as cancelled (#275). Every ccxt capture used to cropdepth.csv/raw.jsonlto the top--depth-limitlevels a side, so a level that was still resting just outside that crop read as a0row when the price moved, the same as a real cancel — 28% of the removals in a five-minute Binance BTC/USDT capture. The capture now records whatever ccxt reports, uncropped: Coinbase, Bitstamp and OKX ignore--depth-limitand always hand ccxt their whole book (over 20,000 levels a side on Coinbase); Binance and its family are asked for their whole book too (see #101 above). Kraken is unaffected, because it subscribes at exactly--depth-limitlevels and drops a level from its own book once the price moves it out of that window — a0row there can still be either a real cancel or Kraken's own window exit, which is inherent to how Kraken reports its book.docs/howto/ccxt.mdnow says what a0row means, per venue. compute_vpinnow warns when its default bucket size is too big for the capture (#274). The defaultbucket_volume(average daily volume ÷ 50) scales a short capture up to a full day, so the bundled sample fills only one bucket and the trailingvpin_avgnever covers a full window. That was already recorded inattrs["diagnostics"], easy to miss on a frame that otherwise looks fine —compute_vpinnow also raises aUserWarningwhen there are fewer complete buckets thann_buckets, naming a smallerbucket_volumeorn_bucketsas the fix. The flow-toxicity how-to gains a short-captures section with a working example on the sample.
Changed¶
-
The price-level depth follows an order that moves or grows.
price_level_volumeused to count every one of an order's rows on the price of its first row, and ignored a rise in its size. It now reads achangedrow that reports no execution and carries a new price as a move: the order's volume leaves the old level and joins the new one. Achangedrow that reports no execution and a larger size adds the difference where the order rests. Deletes and rows that report an execution are still taken off where the order rests, whatever price they carry, so the fix for Bitstamp's deletes at the wrong price is kept. The price-level rebuild now agrees with the per-order rebuild (book_state) on these orders. This matters for Databento, whose modify can move an order or make it bigger. The Bitstamp and LOBSTER outputs are unchanged: neither feed moves or grows an order this way. -
A trade the venue left unlabelled is now classified on the L3 path too.
Pipeline.runlabels any trade with no aggressor side against the reconstructed quotes, filling one subset at a time instead of all or nothing, and never overwriting a side the venue did state. This was already what the price-level path did; it now also covers a per-order feed that states the aggressor on most trades but not all, which is what Databento does for auctions, non-displayed orders and off-exchange prints. No change for a feed that labels every trade, which is every other source in the package.
Added¶
-
A Bokeh plot backend (#123).
result.plot(concept, backend="bokeh")renders the core concepts —trade_tape,depth_heatmap,book_snapshot,depth_chart— as interactive Bokeh figures, alongside the static Matplotlib default and the Plotly backend. Suited to Bokeh / Panel server dashboards and streaming views. Ships in the newbokehextra:pip install "ob-analytics[bokeh]". -
Binance venue notes and a market-data mirror (#101). A new how-to page covers capturing Binance spot through ccxt: the location block, the
binanceusalternative, the depth a 100-level book reaches, and trade sides.capture ccxt --market-data-mirror(CcxtSettings.market_data_mirror) reads Binance spot from Binance's market-data-only endpoints. NewSequenceKindandrecorded_sequence_kind(), akindargument ondetect_sequence_gaps(), and asequence_kindargument ondata_quality_summary(). -
Find hidden liquidity (#111).
detect_icebergs(events, trades)finds iceberg orders from their refills: a resting order filled out, then a new order at the same side and price within one millisecond. It chains refills into one suspected iceberg and gives each aconfidence.hidden_trades(events, trades, depth_summary)returns the trades that printed strictly inside the visible spread. On one day of LOBSTER AAPL it finds 85% of the type-5 hidden executions, with no false ones. The synthetic generator now labels its iceberg slices inSynthSession.icebergs, so the detector can be scored exactly. See the hidden liquidity how-to. -
Iceberg refills and hidden trades drawn on the depth heatmap and order activity map (#272). Both L3 faces now overlay
detect_icebergsandhidden_trades: a diamond marks each refill, joined by a line per iceberg (opacity =confidence); a star marks a hidden trade, with a thin line to the standing best bid and best ask so the print reads as inside that spread. A filled star is a confirmed hidden order; an open star is a trade to check, where the maker order was actually visible or its maker identity did not resolve at all — the diff-feed case the how-to guide describes. Both overlays clip to the gallery's zoom window, and are absent without error when a run has neither. Newprepare.hidden_liquidity_overlaybuilds the same overlay for a custom plot.hidden_tradesnow returnsbest_bid_price/best_ask_priceindepth_summary's own dtype instead of always casting toint64, so a caller already holding display-unit floats gets floats back rather than a silently truncated value. -
A feature table for models (#149).
features(trades, quotes)returns one tidy table: a point in time on each row and a microstructure feature in each column — the shape a model or a study wants. It replaces a join per measurement.
Two decisions make the table, and they are separate. Where the rows fall is
a bar rule, so features() takes the same sampling arguments bars() does
and the same arguments give the same cut in both; a time grid is the time
rule. What each column measures is a Feature, and ten ship: price,
returns, flow, spread, mid_price, micro_price, imbalance,
depth, vpin and kyle_lambda, writing 20 columns between them.
register_feature adds one of your own, usable by name with no edit to the
package.
Every row is stated as of the close of its bar. The trade columns hold what happened inside the bar, and the book columns hold the book as it stood at the close, a backward as-of join. Nothing from after that instant reaches the row, so the table carries no look-ahead — which is tested by truncating the inputs and checking that the rows that survive are unchanged, for every rule and every feature. The table holds no target either: a target looks forward, and building one is a shift the caller makes deliberately.
Two quote states are not books anything could have traded against, and both
would otherwise arrive as ordinary numbers: a side with nothing resting on
it, which the depth engine marks with a price of 0, and a crossed book,
which a diff feed can genuinely hold. readable_quotes() drops them from
the reference series, so a row reaches back to the last quote it could read
— the same test transaction_costs already applied before measuring against
a mid. A locked book, bid equal to ask, is a real state at a spread of zero
and is kept.
Without a quotes frame the five book features are skipped and the table
holds the trade features alone; naming one explicitly raises instead. Two
features that would write the same column are an error rather than a silent
overwrite, and so is a name listed twice in include; a trailing window a
feature cannot use is refused when the feature is built. See the "Build a feature table"
how-to,
which ends in a baseline model.
- Databento market-by-order files (#100).
DatabentoSourcereads Databento's DBN files in the MBO schema, which is a per-order feed: every record carries an order id, so a file replays through the full L3 path with order lifetimes, queue position and order classification. It reaches many venues that no other source in the package does, US equities and futures among them.
Databento reports an execution as a fill record that does not change the
book, followed by the cancel or modify that takes the size off it. The loader
pairs the two, so fill tells an execution apart from a cancel the trader
asked for, and each trade names the resting order it hit. The aggressor's
side comes from the venue rather than a classifier. A book clear deletes the
orders still resting, and the record's two clocks are kept apart: Databento's
receive time orders the events, the venue's own becomes
exchange_timestamp.
A modify that moves an order to another price or makes it bigger is recorded
as a changed event with the new price and size, and the depth follows it
(see Changed). The loss of queue priority is not modelled. The one modify the
depth still cannot follow is one that carries a fill and also moves the order
or changes its size by more than the fill; the loader says how many rows the
depth will be off by.
A feed the loader does not understand is refused: a publisher that only sends top-of-book or price-level data, because its order ids mean nothing; a price-level schema; a file covering more than one book; an action outside DBN's own alphabet; an order id too big for the schema's signed 64-bit id. A malformed record inside a feed it does understand — no price, or no side on a book action — is dropped and counted in a warning, because refusing a whole session over a handful of them would be worse.
Trades are built from the fill records by default, so each names the resting
order it hit. A trade the publisher sent no fill for — an auction, a trade
against a non-displayed order, an off-exchange print — is then left out, and
the loader warns with the volume. DatabentoSettings(trades_from="prints")
builds the trades from the whole tape instead, without makers; use it for
VWAP, bars, flow toxicity and costs. DatabentoWriter
writes an events frame back out as DBN. scripts/databento_window.py sizes a
query against the in-memory envelope before downloading it, then runs one
window at a time. databento is an optional extra
(pip install "ob-analytics[databento]"). See the "Process Databento MBO
files"
how-to.
- Flow-toxicity results say when they rest on too little data (#119). VPIN
and Kyle's λ were designed for markets that trade thousands of times a
minute, and on a thin tape they used to return a number with no sign that it
meant little.
KyleLambdaResultnow hassignificantanddiagnostics: λ is flagged when there are fewer than 30 regression windows (KYLE_MIN_WINDOWS), when|t|is below 2 (KYLE_MIN_T_STAT), or when the fit is undefined. It also carries a confidence interval,ci_low/ci_high, from a block bootstrap over the windows (1000 resamples by default, under a millisecond on the bundled sample, seeded withseed=0so the same trades give the same interval;n_boot=0skips it).
compute_vpin records how it ran in the frame's attrs: bucket_volume,
bucket_volume_rule, n_buckets, and diagnostics, which flags a result
with fewer complete buckets than n_buckets. bucket_volume is now
optional. Left out, it is picked by the new vpin_bucket_volume(trades),
which applies the common rule of average daily volume ÷ 50; a session
shorter than a day is scaled up to a day at the rate it traded
(trading_day="24h" by default). Existing calls are unchanged.
On the bundled sample, λ at 5-minute windows is flagged on both counts (7 windows, t = 1.45) and its interval spans zero; VPIN with the default bucket fills one bucket and says so. Tutorial chapter 6 and the flow-toxicity how-to now show these checks in place of the hand-written caveats.
- Bars: the trade stream resampled into OHLCV rows (#148).
bars(trades, rule, threshold)cuts a trades frame into bars and returns one row each with open, high, low, close, volume, turnover, VWAP, and the buy/sell split of that volume. Five rules ship with it:time(a fixed span of the clock),tick(a fixed number of trades),volumeanddollar(a fixed amount of size, or of price × size), andimbalance(signed size drifting a set amount from where the bar opened). Leave the threshold out and the rule picks one aiming at about 50 bars.
What differs between bar types is only where the boundaries fall, so that is
all a rule decides: BarRule states a name, how it reads its threshold,
and which bar each trade belongs to. register_bar_rule adds one of your
own, usable by name with no edit to the package. Feeds that don't label the
aggressor are classified the same way the flow-toxicity metrics do.
Bars draw as a "bars" plot face on both backends — candles over a strip of
volume coloured by the net aggressor — and bars_panel() puts them in a
gallery. The demos show a clock cut and a volume cut side by side. See the
"Build bars from trades"
how-to.
- Transaction cost and price impact (#110). A new
costmodule answers what trading cost, rather than what the book advertised.transaction_costs(trades, quotes)returns one row per trade with the effective spread — what the taker paid to cross — split into the realized spread the liquidity provider kept and the price impact the trade caused, in price units and in basis points. The three add up exactly, trade by trade.cost_summary()reduces that to volume-weighted session figures and reports how many trades each one could be measured on.
amihud() and roll_spread() read liquidity from the trade prices alone,
so they run on a tape with no quotes and no aggressor side: the price move a
unit of turnover buys, and the spread implied by bid-ask bounce. Roll also
returns the lag-1 autocorrelation of the price changes, which its model
puts at exactly -0.5; how far the number sits from that is how little of
the price movement the bounce explains. On the bundled capture it is
+0.197 and the estimate has no real root, so it is NaN rather than a
number the model does not support. The diagnostic matters in the other
direction too: when the autocovariance lands negative by chance Roll returns
a spread that is not there, and the autocorrelation is what catches it.
The mid a trade is measured against is the last quote strictly before it,
skipping crossed quotes: on a frame built from the same event stream, the
quote sharing a trade's instant is the book after that trade took the touch,
and a crossed book has no midpoint at all. A trade in the last horizon of
the capture has no future mid, so its realized spread is NaN rather than
the final quote reused.
transaction_costs takes mid_column to measure against a reference other
than the plain mid — "micro_price" for the size-weighted mid, which on the
bundled capture reads 1.24 bps against the plain mid's 1.45.
The decomposition draws as a level-less transaction_costs face on both
backends — two lines with the impact as the band between them — and
transaction_costs_panel() puts it in a gallery. Both demos now include it.
See the "Measure transaction costs"
how-to.
- Polymarket prediction markets, through the ccxt source (#103).
ob-analytics capture ccxt --exchange polymarket --pair <token id>streams one outcome's order book and trades over Polymarket's public websocket, with no account or API key.--pairis Polymarket's token id for the outcome, which the Gamma API lists asclobTokenIds. Each outcome is its own book, and its trades are priced in that outcome.
Polymarket makes a market's tick finer as the price nears 0 or 1, so a ccxt
capture now makes its recorded tick size finer when a price arrives between
two ticks, and counts each change in tick_size_changes in meta.json. The
replay then reads every price exactly. See the "Capture Polymarket
prediction markets"
how-to.
- Kalshi prediction markets, through the ccxt source (#102).
ob-analytics capture ccxt --exchange kalshi --pair <market ticker>records a Kalshi market's order book and trades from Kalshi's public API, with no account or API key, andob-analytics processreplays it through the L2 path. The ccxt source now looks up CCXT's prediction markets (ccxt.prediction: Kalshi, Polymarket and others) as well as its crypto exchanges; before,--exchange kalshifailed with "Unknown CCXT exchange".binanceandhyperliquidare in both lists, so the plain id keeps meaning the crypto exchange andprediction/<id>picks the prediction market.
The captured book is the market's Yes book: a bid to buy No at p is
recorded as an offer to sell Yes at 1 - p, and a trade is priced in Yes
and signed from the Yes side. A ccxt capture also records its market's tick
size in meta.json, which process and audit use. See the "Capture
Kalshi prediction markets"
how-to.
- A capture records which rows came from its opening snapshot (#237).
orders.csvanddepth.csvgain anorigincolumn:snapshotfor the opening book,streamfor a live message,shutdownfor a synthetic close-out. The capture runner fills it in, so every live source gets it without a change, and the loaders carry it through toevents. Before this, the only way to tell a snapshot row from a live one was to compare itsexchange_timestampwithsnapshot_microtimestampinmeta.json.
meta.json also reports n_snapshot_unconfirmed: how many orders in the
opening book no later order event or trade mentioned. The bundled Bitstamp
sample has 6,294 of 6,512. Almost all of them sit far from the touch and did
not trade, but two stale asks among them held the best ask for most of the
session. See "Capture live
data".
-
auditnames stale resting orders (#234). A trade above a resting ask, or below a resting bid, shows that the order has gone. An order the venue then does not report again within one second is now reported as astale_orderswarning, and the worst one is named with its id, side, price and how long it held the touch. On the bundled Bitstamp sample this names ask2002347646152704, which held the ask touch for 27 minutes and causes almost all of the 91.6% crossed time. The crossing note no longer calls a diff feed's crossing normal when the run has stale orders. Nothing is removed:order_book()still replays what the feed said. New public names:detect_stale_orders,StaleOrder,DataQualitySummary.stale_orders, and atick_size=argument ondata_quality_summary. -
A metric registry, so a user metric runs and plots with no core edit (#140). A metric is a plain object with a
name, atitle, thelevelsit applies to,compute(result)andprepare(frame)— no base class to inherit, the same structural typing sources and writers use. Register it withregister_metric(metric), or ship it in your own package under theob_analytics.metricsentry-point group andload_metric_plugins()finds it atimport ob_analytics.
A registered metric is a level-less plot concept under its own name, so a
renderer at (name, None, backend) is its face. It then appears in
available_concepts(result), renders through result.plot(name), and gets
its own gallery card with no extra_panels=. Metrics run when asked for, not
during Pipeline.run: result.metric(name) computes one and
result.metrics() computes every metric whose levels include the run's
resolution — so an L3-only metric is skipped on an L2 run instead of failing
on its empty events table, and a metric that raises is logged and its card
dropped, so one broken metric cannot stop the gallery being built. New public
names: Metric,
register_metric, list_metrics, get_metric, load_metric_plugins,
PipelineResult.metric / .metrics. See the "A new metric"
how-to.
ob-analytics audit, a data-quality gate (#108). The oldvalidateverb is nowaudit(the old name still works), it scores the run against named checks, and it exits non-zero when one fails — so a script can stop before trusting a feed. Five checks are new: orphan orders (changed or deleted with nocreatedrow), non-positive prices, negative volumes or fills, and the two clock-order defects — a venue timestamp later than the receive timestamp, and messages that arrived out of venue order.auditalso loads withtrack_sequenceon, so the dropped-message check (#146) reads a venue sequence whenever the feed carries one.
Each check carries a Severity: an error fails the run, a warning
fails it only under --strict, and info never does. A crossed resting
book is scored by feed type, not by size — an error on a matched book, a
faithful replay on a diff feed. --json emits every check plus an ok
verdict; --from-parquet audits a saved process output without re-running
the pipeline. New public names: Severity, QualityCheck, and
DataQualitySummary.ok / .errors / .warnings / .checks. See the
"Check data quality with audit" how-to.
PipelineResult.to_arrow()andPipelineResult.to_polars()(#104). Both return the run's four tables —events,trades,depth,depth_summary— keyed by name, with the same keys on every run: on an L2 runeventsis an empty table, not a missing key. The Arrow tables carry the schema version and tick size in their metadata, the same key-value metadata the Parquet files carry, so a reader handed tables in memory is no worse off than one reading files. Polars is not a dependency and is not installed;to_polars()raisesImportErrorwith an install hint when it is missing, and Polars keeps no schema metadata, so the version and tick size do not survive that conversion.- The frame-type contract is written down in
Frame types: pandas in, pandas out:
public functions take and return pandas, plug-ins are handed pandas, and the
versioned Parquet is how other tools read the output. The reasoning is in
adr/0002-dataframe-library.md. - cryptofeed source for live L2 and L3 capture (
ob-analytics capture cryptofeed --exchange <venue> --pair <symbol>). The per-order complement to the CCXT source: venues publishing an order-by-order book recordorders.csvwith the venue's own ids and replay through the full reconstruction pipeline; the rest recorddepth.csvfor the L2 path. The level is discovered from the venue's declared channels rather than a hardcoded list, and--levelforces it — except L3 on a venue that publishes none, which raises. Ships as the optional[cryptofeed]extra, imported lazily. See the new "Capture cryptofeed venues" how-to. -
sequenceis now written toorders.csv. The capture sink dropped the venue sequence on the L3 path, sodetect_sequence_gapshad nothing to read; L2 already kept it. cryptofeed captures also reportsequence_gaps/sequence_missinginmeta.json. -
L2 (price-level) depth-native ingestion path. Price-level feeds (Binance, Kalshi, Polymarket, most CCXT sources) publish
[price, quantity]levels and diffs with no order IDs; ob-analytics now ingests them as a first-class L2 resolution instead of faking per-order state. A format declares itsresolution(Level.L2/Level.L3, exposed asob_analytics.Level); an L2 format's loader is aDepthSourcethat yields the depth frame directly, andPipeline.runtakes the price-level path — depth metrics / spread and trade-sign classification run, while the per-order stages (set_order_types,order_aggressiveness, queue reconstruction) are skipped.PipelineResultgains aresolutionfield and, on an L2 run, returns an empty (schema-valid)eventsframe. Ships thedepth_csvformat (L2DepthLoader,L2TradeReader,DepthCsvWriter,DepthCsvFormat) for the canonical L2 CSV schema, atoy_l2_depth()/toy_l2_trades()synthetic snapshot+delta fixture, anddata_quality_summary+ the gallery degrade gracefully (L3-only faces skipped, not errored).ob-analytics process|validate --format depth_csvworks from the CLI. Unblocks the aggregated venue connectors. Documented in a new "Process L2 feeds" how-to. - Trade-sign classification (
ob_analytics.trade_sign) for feeds that don't label the aggressor side.tick_rule(last-price-change sign),lee_ready(quote-midpoint test with a tick-rule fallback), andbulk_volume_classification(BVC — the buy fraction of a volume bar via the standardized-price-change normal CDF).classify_trade_sign(trades, method=..., quotes=...)is the per-trade entry point.compute_vpinandorder_flow_imbalancegainsign_method/quotesarguments and now synthesizedirectionautomatically when the trades frame has none — so VPIN and OFI run on L2 / aggregated captures, not just L3. A nativedirectionis still honored unchanged (sign_method=None). On the bundled Bitstamp L3 sample the classifiers agree with the true maker/taker side ~0.83 (tick) / ~0.79 (Lee–Ready) — validated by a test harness. - Feed classification. Every format declares a
FeedType(matched_bookvsdiff_feed) through afeed_typeattribute —BitstampFormat→diff_feed,LobsterFormat→matched_book— so downstream code reasons about crossed books by coordinate, not by format name. Exposed asob_analytics.FeedType. order_book(..., uncross=True)evicts crossed resting orders for display, mirroring the depth engine's crossed-level eviction. The default stays faithful, so a diff feed's genuinely crossed resting orders are replayed as-is. Threaded throughprepare.book_snapshot(..., uncross=True)(also drivesdepth_chart) and available frame-level asanalytics.uncross_book_sides.- Per-run data-quality summary.
data_quality_summary()and the newob-analytics validate <source>CLI verb report the crossed-resting %, unmatched-trades %, duplicate ids, and pre-existing-order count. A new "Data quality: matched book vs diff feed" explanation page and avalidatehow-to document the distinction.
Fixed¶
- The VPIN chart draws an empty panel when no bucket is complete. A
capture with less volume than one bucket gives
compute_vpinzero rows, andplot("vpin", ...)then raised: aTypeErrorwith matplotlib, aKeyErrorwith both backends when the frame had no columns. Both backends now draw the axes with "(no complete buckets)" in the title.compute_vpinnow returns its usual columns and dtypes when it has no rows, where before it returned a frame with no columns (or, withsign_method="bvc", columns of object dtype). -
A price-level file no longer has its prices rounded to the tick size.
L2DepthLoaderandL2TradeReaderconverted each price to the nearest whole number of ticks, so a price finer thantick_sizemoved without a warning: a Kalshi price of 0.036 loaded as 0.04 at the default 0.01 tick, and the most traded Kalshi markets quote in tenths of a cent. Both now raiseConfigErrorwhen a price is not a whole number of ticks, and say which tick size to set. A ccxt capture records its market's tick size inmeta.json, andob-analytics processandob-analytics auditread it from there (recorded_tick_size), so a CLI replay needs no extra option. -
A streamed ccxt capture no longer writes a trade twice. ccxt's Polymarket websocket handed back a trade it had already delivered, together with the next new one, and only a polled capture skipped repeats. Every ccxt capture now skips a trade identical to one it has written: same id, time, price, size and side. Two fills that share a Polymarket id (the settling transaction) are both kept.
meta.jsoncounts the skipped repeats induplicate_trades. -
A polled ccxt capture no longer records trades from before it started. On a venue without websockets, the first poll of the trade tape returns the venue's recent history, which on Kalshi reached back nine hours. Those trades were written with the capture's receive time, as if they had just happened. The capture now drops trades older than its opening book.
-
A Bitstamp capture no longer starts from a snapshot older than its stream (#237). The capturer subscribes to the WebSocket, then fetches the REST book. It assumed the stream already covered the moment the book describes, but it often does not: in a live test the first order message came 0.7 s after the snapshot's
microtimestamp, and the bundled sample shows the same 0.74 s gap. An order deleted in that gap stayed in the capture until the syntheticdeletedat shutdown. In the bundled sample, one such ask was the best ask for 89% of the session.
The capturer now fetches the book again, a second apart and up to 10 times,
until some buffered order message is at or before the snapshot's
microtimestamp. meta.json gains snapshot_fetches and
snapshot_overlap. In the live test, the second fetch no longer listed any
of the 15 orders that were gone. Eight of those were orders that trades
printed through.
- A price level now empties when the order resting on it goes away.
price_level_volumeadded an order's volume at the price on itscreatedrow and subtracted it at the price on whichever later row removed it. Those two prices are not always the same: Bitstamp reports adeletedcarrying a price the order never rested at for 1.3% of orders, and the subtraction then landed on a level the volume was never added to, leaving the created level holding it for the rest of the session. Every later row now subtracts at the order's created price, so+vand-valways cancel on one level.
On the bundled Bitstamp sample this removed 104 price levels holding 29.95
BTC that no order was resting on. They were the reported touch on both sides
— best bid $78,495.00 against a real best bid of $78,350.00, and best ask
$78,324.00 against a real best ask of $78,333.00 — so best_bid_price moves
on 35.7% of depth_summary rows and best_ask_vol on 68.9%. The per-order
rebuild (engine.book_state) tracks orders by id and never had this problem;
the two rebuilds now agree on how long that book is crossed.
aggressiveness_bpsis NaN, not an infinity, against a zero touch. The depth engine reports a zero price for an empty side, and the Bitstamp sample also carries orders priced at zero (auditreports these asnonpositive_price). Dividing by that produced a signed infinity that travelled through every downstream mean. A distance from a price that is not tradeable has no value, so it is now NaN.
Changed¶
-
depth.bin_volume_columns()is public. It returns the per-bps depth-bin volume columns a depth summary carries, ordered from the touch outward, and it was already the answerbook_imbalanceanddepth_signalsneeded. The feature table needs the same answer, and so does anyone writing a depth feature of their own, so it is no longer private. Behaviour is unchanged. -
trade_sign.resolve_direction()now makes the guarantee its docstring already claimed: thedirectioncolumn it returns holds only"buy"and"sell". A native column was previously passed back untouched however it was filled, and every consumer reads it as== "buy"and takes the rest as a sell — so a partly-labelled feed did not lose its unlabelled trades, it counted them on the wrong side.compute_vpin,order_flow_imbalance,barsand the cost metrics were all affected. Rows that are neither side are now inferred the same way a wholly unlabelled feed is, with a warning saying how many. A feed that labels every trade is passed through unchanged.compute_kyle_lambdareaches the same guarantee: it still requires adirectioncolumn rather than inferring one, but a column being present no longer means every row in it is trusted. -
trade_sign.prevailing_mid()is now public, and takesallow_exact,skip_crossed,mid_columnandrequire_covered— the last quote strictly before an instant, crossed books skipped, a named reference column such asmicro_price, andNaNrather than the final quote reused once the quotes stop reaching. All four default to the previous behaviour, soclassify_trade_signis unchanged. An empty quote frame now returns allNaNinstead of raising a pandasMergeError. -
Sizes are integer lots plus a
lot_size, not floats (issue #226). Breaking: the on-disk schema goes 3.0 → 4.0. Everyvolumeandfillcolumn is now a whole number of lots (int64) instead of adoublein the base asset. The base-asset size islots * lot_size, wherelot_sizeis the instrument's minimum size increment (PipelineConfig.lot_size, default1e-8; LOBSTER sets1, whole shares). This is the size half of the integer-tick decision (issue #155) and it fixes a real defect rather than only re-expressing the data.
A price level is a running sum of adds, cancels and fills. A float sum does
not return to exactly zero when the last order leaves, so a level landed on
residue such as 5.55e-17, stayed live, and was reported as the best bid or
ask ahead of the real one. On the bundled Bitstamp sample that corrupted the
reported best bid on 25,611 of 313,565 rows (8.2%) and the best ask on 30,096
(9.6%) — the spread on about one row in eleven. Integer lots cancel exactly,
so a level empties or it does not, and those counts are now zero.
It was found by the new cross-check against hftbacktest (issue #224), and
that is what confirms the fix: replaying an exported session through
hftbacktest's own L3 reconstruction now agrees with depth_summary on the
best bid and ask for every row across five synthetic seeds, and Nautilus'
book agrees too. Before the fix the two disagreed on up to 78 rows a seed.
The change reaches every size-valued column — depth_summary's per-bin
volumes, placed_vol and filled_vol, the book snapshot's liquidity, and
the queue's ahead_volume and remaining — so their sums are exact as well.
Three float-era workarounds went with it: the Kahan compensation behind
filled_vol, the simulator's _vol_eps exhaustion tolerance, and the
LOBSTER book replay's 1e-12 level cutoff. Loaders convert on the way in;
the plots and the round-trip and export writers convert back, so what a user
sees and what another tool reads are unchanged. lot_size travels in each
Parquet file's key-value metadata under ob_analytics_lot_size, next to
ob_analytics_tick_size, and load_data surfaces it as
df.attrs["lot_size"]. Files written at 1.0–3.0 still read, as the
float-size frames they are. Golden outputs were re-baselined on purpose.
-
The export writers leave out orders that never rested (issue #224). A marketable order is recorded as a transient add on its own side at the touch, then the fill, then a delete;
ob_analytics.depthhas always excluded these from the book, but the hftbacktest and Nautilus writers were sending them. A backtesting engine reads an add as real liquidity, so its book crossed at the touch and dropped the resting level the order traded against — its reconstruction drifted permanently thinner than ours. Both writers now exclude them, which is what makes the two books agree. -
The order-book engine is its own module (issue #136). The rebuild (
order_book), the per-order lifecycles, and the FIFO queue reconstruction moved out ofanalytics.py/queue.pyintoob_analytics/engine/, behind one input and one output: order events in, book states and order lifecycles out, and nothing else. The engine imports no pandas — everything crosses its interface as NumPy arrays, with the shared schema (issue #112) as the input, timestamps as int64 UTC nanoseconds (issue #154) and prices as integer ticks (issue #155). Results carry a row index back into the caller's event arrays instead of copying columns out, so adding a column to the schema does not widen the interface and the engine never learns a vocabulary — order types, venue names — belonging to the layer above.ob_analytics/_engine_frames.pyis the one place pandas and the engine meet;analytics.order_book,analytics.order_lifecycles, and theob_analytics.queuefunctions are now its frame adapters and keep their exact signatures, dtypes, column order, and index behaviour. Output is unchanged byte for byte — the golden-output gates from issue #143 pass on their recorded fingerprints. Two things did move: the display window (max_levels,bps_range) and the queue sampling window are set by the frame adapters rather than the engine, which reconstructs the whole book and replays to the instants it is given. A new import test (tests/test_engine_boundary.py) keeps the engine free of pandas and of every layer above it. This is what lets the inside be replaced with a faster implementation (#138) or fed one event at a time (#139) without touching anything else.Direction,Action, andOutcomeareIntEnumcode vocabularies that derive their schema strings from their own member names, so a code and its label cannot drift apart. Two details of the frame code are reproduced deliberately rather than rewritten: an order's executed total is accumulated with compensated (Kahan) summation, as the pandas aggregation it replaced did, and placement values are taken per column as the first non-null among an order'screatedrows. The lifecycle table is now covered bytests/test_golden_synth.py, which it was not before. -
BitstampTradeReaderno longer requires integer order ids. It keyed its maker/taker lookup onint(order_id), which crashed on a public trade tape carrying no ids (int(NaN)) and on venues publishing UUIDs. Integer ids behave exactly as before; other ids match on their string form, and a missing id resolves toNaNinstead of raising. -
One
Sourceshape for every data source, file or live (issue #137; settled #145 as "optional extras plus entry-point plug-ins"). File loaders and live capturers were two separate designs with two registries; they are now oneSourceprotocol with two capability refinements —OfflineSource(replay stored files: the loader / trade-source / writer / depth factories) andLiveSource(capture a venue:snapshot/stream/shutdown_synthetic_events). A source states itslevel(L2/L3) andfeed_type, carries typedsettings, and registers in the singleSOURCESregistry viaregister_source. A source can be both:BitstampSourcenow covers offline replay and live capture in one descriptor. This is a breaking API change with no back-compat shims:Format→OfflineSource;LiveCapturer→LiveSource;BitstampFormat/LobsterFormat/DepthCsvFormat→BitstampSource/LobsterSource/DepthCsvSource;CcxtCapturer→CcxtSource.Pipeline(format=...)→Pipeline(source=...);Pipeline.from_format→Pipeline.from_source.- The
FORMATS/CAPTURERSregistries and theirregister_format/register_capturer/list_formats/list_capturers/get_capturerhelpers are replaced bySOURCES/register_source/list_sources/get_source(inob_analytics.sources). PipelineResult.resolution→PipelineResult.level(one coordinate name across the codebase;Source.level, matching the visualization layer).CaptureConfig.extras(the untyped settings dict) is removed. Per-source settings are now typedSourceSettingson the source itself, e.g.CcxtSource(settings=CcxtSettings(exchange="binance", depth_limit=100)).- CLI:
process/validatetake--source(was--format), and theformatsverb is nowsources(it also shows each source's capability and required context).
- Third-party sources load through entry points. A source can ship in its
own package and advertise itself under the
ob_analytics.sourcesentry-point group;ob_analytics.sources.load_source_plugins()discovers and registers it at import time, with no edit to ob-analytics. The built-in sources (bitstamp, lobster, depth_csv, ccxt) self-register on import and stay behind today's[live]/[ccxt]extras. - Prices are now integer ticks, not floats (issue #155). Every
pricecolumn — events, trades, depth, depth_summary, book snapshot, and order lifecycles — is a whole number of ticks (int64); the quote-currency price isticks * tick_size, wheretick_sizeis the instrument's minimum price increment (PipelineConfig.tick_size, default0.01). Loaders convert a raw price to ticks on load; the plots and the round-trip writers convert back for display, so figures and CSV output are unchanged. Storing the exact integer removes the float rounding that made small-tick and 0-1 instruments show crossed levels that were not real, and the depth engine now bins and compares levels on exact integers instead of multiplying and rounding each event — LOBSTER'sprice_divisoris now just the raw-feed encoding scale.tick_sizeis written to each Parquet file'sob_analytics_tick_sizekey-value metadata (a JSON map keyed by instrument, ready for per-(venue, symbol)ticks in #147) and surfaced onload_dataframes'attrs. Breaking: the dtype of everypricecolumn changed fromdoubletoint64and the stored numbers changed (prices re-expressed as ticks; price-valued analytics such astrade_impactsVWAP and Kyle's λ are now in tick units — multiply bytick_sizefor the quote currency; scale-free metrics such as bps depth and order-book imbalance are unchanged). The canonical Parquet schema version is now3.0(a1.0/2.0file still reads — Parquet is self-describing — as the float-price frame it stored, whose prices are not directly comparable to a3.0file's ticks; re-save it to move it onto the tick model). Golden-output baselines were re-recorded behind the correctness gate (#143). - Timestamps are now tz-aware UTC nanoseconds (
timestamp[ns, tz=UTC]) on both clocks —timestamp(receive) andexchange_timestamp(matching engine) — across every table, loader, the synthetic generator, and the toy datasets (issue #154). Before, they were tz-naive and in each venue's native clock (millisecond-resolution UTC for Bitstamp, US/Eastern for LOBSTER), and frames from different venues were declared not comparable. Now every frame sits on one UTC clock, so cross-venue frames can be joined or concatenated directly. LOBSTER's seconds-after-midnight are converted to UTC from the session date and a venue time zone (RunContext(session_tz=...), defaultAmerica/New_York); Bitstamp / CCXT keep their wall-clock instants and only gain the zone and the nanosecond unit, so their values do not move. The schema also documents a same-instant total order —timestamp, thensequence, thenevent_id, theningest_seq(ob_analytics.schemas.time_order_keys), which the per-order reconstructions sort by. Breaking: the dtype of every timestamp column changed, so the canonical Parquet schema version is now2.0(a1.0file still reads — Parquet is self-describing — as the tz-naive frame it stored; re-save it to move it onto the UTC clock). Consumers that compared pipeline timestamps against tz-naivepandas.Timestamps must now use tz-aware (UTC) ones.
Fixed¶
- Order lifecycles read every filled order as cancelled when sizes were
floats (#226 regression).
order_lifecyclessummed each order's fills and cast the total toint64. On integer lots that is exact, but the function also accepts base-asset floats, and it is handed them on every gallery run:display_resultconverts a whole result to display units before any face builds. Base-asset sizes are mostly below 1, so a 0.121 BTC fill truncated to0, the order read as never executed, and the three lifecycle-derived L3 faces — Order Activity, Order Outcome and Queue Position — drew a book of nothing but cancellations. On the bundled Bitstamp sample the Order Activity face lost 224 of its 226 filled spans. The sum now keeps the units it was given, integer lots summing exactly and base-asset floats with the compensation that was dropped as part of #226.
LOBSTER was never affected: its lot size is 1, so a truncated size equals the size. Every LOBSTER face is pixel-identical across the change.
- LOBSTER's
fillcolumn wasfloat64, not integer lots (#226). A0.0literal in the expression that built it widened the whole column, so a schema-4.0 LOBSTER run wrote base-asset-looking floats that were really lot counts. Nothing raised; the values only differ from the correct ones once the lot size is not 1.
[0.1.0] - 2026-06-26¶
First public release (PyPI). The Python port of the R obAnalytics package, reworked into a pipeline API with pluggable formats, flow-toxicity metrics, L2/L3 visualization, and Matplotlib/Plotly backends — plus the packaging, documentation, and distribution that make it installable. The sections below also record how the API was deliberately de-bloated and unified during the port (the pipeline's numeric output is unchanged — the regression fingerprints pass; only the shape of the public API moved). See Extending ob-analytics.
Packaging & distribution¶
- The bundled Bitstamp sample ships gzip-compressed (
orders.csv.gz, ~23 MB → ~2.9 MB installed);sample_csv_path()returns the.gzpath, read transparently by pandas. No API change. - Published documentation site (GitHub Pages),
CITATION.cff, an explicit GPL-2.0-or-later license section, and a "Scale envelope" doc. - PyPI release workflow (
release.yml, trusted publishing), package classifiers and project URLs, andob_analytics.__version__. - Fixed quickstart and API-reference documentation drift.
Breaking¶
- Pydantic models removed.
ob_analytics.models(OrderEvent,Trade,DepthLevel,OrderBookSnapshot) deleted; the data contract is now column-list constants +validate_events_df/validate_trades_df/validate_depth_dfinob_analytics.schemas. metrics/package removed.ToxicityMetric,Vpin,Ofi,KyleLambda,register_metric, andlist_metricsare gone. Callcompute_vpin,compute_kyle_lambda, andorder_flow_imbalanceonresult.tradesdirectly.Pipeline(metrics=...)removed. Metrics are no longer a pipeline stage — compute them after the run.PipelineConfig.vpin_bucket_volumeremoved — passbucket_volume=tocompute_vpin.PipelineResultslimmed to exactlyevents,trades,depth,depth_summary, andconfig. Thevpin,ofi,metrics,metadata, andextrasattributes are gone.- The thirteen
plot_*wrappers removed → oneplot(name, *, backend="matplotlib", ax=None, **data)dispatcher keyed by(plot_name, backend); renderers self-register intoRENDERERS. - Global theme state removed.
set_plot_theme/get_plot_theme/_current_themedeleted; passtheme=PlotTheme(...)toplot(). - Exception hierarchy collapsed to
ObAnalyticsError+ConfigError.InvalidDataError,MatchingError,InsufficientDataError, andConfigurationErrorare removed. - Top-level
__all__trimmed to ~22 orchestration names. Low-level helpers now import from their submodules —ob_analytics.bitstamp,ob_analytics.lobster,ob_analytics.analytics,ob_analytics.depth,ob_analytics.data,ob_analytics.visualization,ob_analytics.flow_toxicity. Formatis now atyping.Protocol— there is no base class to inherit; any conforming object is recognised structurally.- Low-level helpers no longer re-exported from the package root (e.g.
depth_metricsis nowfrom ob_analytics.depth import depth_metrics). RunContext.extrasandFormat.collect_extrasremoved. LOBSTER trading halts are read fromLobsterLoader.trading_haltsand composed into the gallery viaextra_panels=.DepthMetricsEngine.update()removed → the public hot-path method isupdate_side(price, volume, side, out).
Added¶
ob_analytics.schemas— the single data contract: column-list constants (EVENT_COLUMNS,TRADE_COLUMNS,DEPTH_COLUMNS) plus thevalidate_*functions, run at the pipeline's Protocol boundaries. Replaces the Pydantic model layer.- One generic
Registry[K, V](ob_analytics._registry) backs the format, writer, capturer, and renderer registries. Register through the public helpersregister_format,register_writer,register_capturer, andRENDERERS.register/register_plot_backend. - Unified
plot()dispatcher +RENDERERSregistry keyed by(plot_name, backend), so new plots and backends plug in without a wrapper function. The HTML gallery composes custom panels viaextra_panels=. ob_analytics.live— optional sub-package for live order-book capture: theLiveCapturerprotocol (with an optionalSupportsDiagnosticscapability),CaptureConfig,CaptureResult,CaptureSink, and a generic asyncio runner. Capture output drops straight into the pipeline (orders.csvschema unchanged). Install withpip install "ob-analytics[live]".ob-analytics capture <venue>CLI verb with a built-inbitstampcapturer (ob_analytics/live/bitstamp.py);--listshows registered capturers.scripts/collect_bitstamp_btcusd.pyis now a thin wrapper around it.TradeSourceprotocol andBitstampTradeReader— read an authoritative companiontrades.csvand join it to events via thefillcolumn.RunContextdataclass (ob_analytics.protocols, re-exported at the top level) for per-run parameters such as LOBSTERtrading_datethat don't belong on long-livedFormatinstances.- Docs —
docs/extending.md(add a data source / writer / plot / metric / capturer). - Tests —
test_bitstamp.py,test_cli.py(subprocess smoke tests for all CLI subcommands),test_exceptions.py,test_data_registry.py, a regression snapshot suite pinning demo Parquet hashes + the Kyle-λ baseline, andob_analytics/__main__.py(python -m ob_analytics).
Changed¶
- Bundled sample —
ob_analytics/_sample_data/now shipsorders.csvandtrades.csvfrom a modern BTC/USD live capture (replaces the legacy 2015 orders-only slice). - Demos consolidated into
ob_analytics._demos;scripts/bitstamp_demo.py,scripts/lobster_demo.py, and thebitstamp-demo/lobster-demoCLI subcommands are now thin argparse wrappers. Behaviour unchanged. - Performance — the LOBSTER book is maintained as a
SortedDict(no per-event re-sort), Bitstamp trade→event resolution is indexed, LOBSTER depth uses a single strategy, the Plotly import is memoised, and depth metrics sum active levels into bps bins. Numeric output is unchanged (pinned by the regression snapshots). compute_kyle_lambdacomputes its OLS vianp.linalg.lstsq(was hand-rolled; agrees with the prior implementation tortol=1e-10).- Internal modules reorganized (renames from the 0.x line): e.g.
event_processing.py→bitstamp.py, validation/time helpers →_utils.py, and the visualization modules split into avisualization/subpackage. - Type checking is Astral's
ty(not mypy); lint and format are Ruff.
Removed¶
pacmanorder type. A legacy artifact of the 2015 Bitstamp HTTP API, where a singleorder_idcould appear at multiple prices over its lifetime. Modern Bitstamp WS v2 and LOBSTER do not produce this pattern (price-modifies become cancel + new id). ThetypeCategorical no longer includes"pacman", the set-subtraction classification path is gone, andLobsterLoaderno longer renumbers hidden-execution ids (raw type 5 now retains the native LOBSTERid=0).- Bitstamp trade inference. A companion
trades.csvnext toorders.csvis now required. Removed: Needleman–Wunsch matching,BitstampMatcher,BitstampTradeInferrer, theMatchingEngine/TradeInferrerprotocols,NeedlemanWunschMatcher, and thematch_cutoff_ms/price_jump_thresholdfields onPipelineConfig. - Zombie detection —
get_zombie_idsand thezombie_offset_seconds/skip_zombie_detectionconfig fields. - LOBSTER
LobsterMatcher— removed;LobsterTradeInferrerrenamed toLobsterTradeReaderwithload(events, source). - Legacy Bitstamp-only wrappers
load_event_data,event_match,match_trades,process_data, andplot_price_levels_faster. - 12 unused runtime dependencies (scikit-learn, scipy, jupyter, bokeh, …) and stale dev dependencies (black, flake8 + plugins, darglint).
Fixed¶
depth_metricsno longer overflows for prices > $9,999.99 — dynamicdict[int, int]state replaces the fixed array.best_bid/best_askare tracked correctly from the first event (were initialised with dataset-wide max/min).datetime_to_epochuses.astype("int64")instead of the deprecated.view("int64").- All
print()replaced withlogurulogging; all bareassertstatements replaced with raised exceptions;plt.show()removed from plot functions (callers control display).