On-Demand Webinar

US Book Depth: Market Microstructure Analysis with OneTick Cloud

This workshop, hosted by OneTick and KX, demonstrates how you can use OneTickCloud to access on-demand Level 3 Market Data sourced from US Equity Exchange PCAPs to analyze both market structure across venues and microstructure within venues.

The Latest from OneTick Cloud | AI-Ready Market Data as a Service

In this workshop, Peter Simpson demonstrates how you can use OneTick Cloud to access on-demand Level-3 Market Data sourced from US Equity Exchange PCAPs, to analyze both market structure across venues and microstructure within venues.

This new standardized US Level 3 book depth offering allows quants and analysts to analyze unified US equities and options market data with nanosecond-precision, in a research-ready format, without the heavy burden of in-house engineering. See every individual order, execution, modification, and cancellation on the order book.

About OneTick Cloud:

OneTick Cloud is a high-quality, on-demand managed time-series data and analytics platform that provides instant, global, AI-ready market data, seamlessly fueling compute and analytics engines. It offers a single vendor end-to-end solution, significantly reducing the time to value.

KX and OneTick are the only vendor delivering AI-ready, hydrated, temporal market data as a managed service. Our data is pre-normalized across 250+ venues and 30+ years of history, point-in-time with no look-ahead bias, machine-readable from day one, and fed natively into Python, SQL, and KDB-X.

Save time, money, and resources by letting the KX OneTick team clean feeds, map symbols, and align timestamps so your quants, analysts, and AI models can do their work.

Webinar Agenda:

OneTick Cloud Datasets

  • L1 through to L3
  • Real Time, Intraday and Historic
  • Global Equities, Futures, Spreads & Options

Calculating Market Structure Metrics

  • Calculating US Trade Volume Types, by Venue
  • Comparing Exchange Trades against NBBO
  • Comparing Exchange Quotes against NBBO

Calculating & Visualizing Books

  • Single and Consolidated Books Across Time
  • Single & Consolidated Books at a Point in Time
  • Replaying Composite Markets

Calculating Market Microstructure Metrics

  • Sweeps to Book Level
  • Sweeps to Price Skew
  • Sweeps to Accumulated Quantity
  • Order & Order Msg Metrics
  • Trade & Cancel Durations

This session is designed for financial professionals who are interested in learning about both available Book Depth datasets and analysis techniques. The workshop will focus on SQL and Python market structure and microstructure analysis.

Speaker:

  • Peter Simpson, OneTick Product Owner, KX

Watch the Recording:

 

Webinar Transcript:

Today's webinar is focused on US book depth, and it's a workshop because we're going through a series of examples. So to set the scene, I'll introduce myself, the platform, how it's supported, and then the types of book analysis that are available. Now I am Peter Simpson, and I'm responsible for the OneTick cloud platform. And I've been doing this for at least the last seven years, and I have a history in financial markets, especially analyzing capital markets data. Now we provide OneTick cloud, which is a market data on demand service where you can both access and analyze historic market data.

This covers real time streaming, intraday, and, of course, historic market data across global equities, futures, spreads, and options.

And as well as the collected data, we create AI feature sets from the base tick data and also create composites for regions with fragmented liquidity. Now our coverage includes around six million symbols, a big chunk being options with the data going back to nineteen ninety three in the case of the US SIP. Different markets go back different time periods. Now we're storing around two terabytes of data a day. People can access the data and importantly run their analytics as close as possible to data.

So it's both data access and data analytics. Now OneTick Cloud is supported by a set of teams.

To manage this market data on demand service, we have teams covering the whole data onboarding process and then ongoing operations.

You should make sure that data is available on time, both the source tick data and also all the derived AI feature sets. Now a new customer can get up and running in minutes rather than wasting time and resources for replicating this whole environment and this whole set of resources themselves.

Now I'll be using the OneTick Cloud market data on demand service to both query, analyze historic l three order book data for US markets.

We'll be using SQL and Python to query the data and want it dashboards to display the data visually so we can replay through market activity identifying anomalies.

Doing this, both working per venue and working across venues where we'll create a consolidated order book. Now analysis we'll be performing would include pretty simple data aggregation and filtering, Then we'll be building order method metrics, order duration metrics. So by order duration, how long a particular order takes to execute or be cancelled at a particular level in the book.

We'll look at book reconstruction at a particular point in time, then booked reconstruction across time. That's per venue. We'll also look at consolidated order book reconstruction across all of the venues that we select, how to filter output books by different mechanisms, how to build metrics either at a book level, by price skew, or by size, and then potentially by spread.

Now we have a series of displays which are primarily used for reference, and they include listing our cloud coverage, schemas, enumerations, etcetera, how to explore the data visually, data retrieval methods for the data could be just REST or Parquet.

Now if you're using REST, you have access to our analytics side where we provide a large and ever growing set of examples and documentation which are fed through our MCP server. So from your IDE, you can ask questions and easily generate query syntax in both SQL and Python. And then you can run your queries through our server and get back your results. Now before we jump into SQL, let's look at the data. Now I've already logged in to the service, which I came to by going to wantic dot com and clicking the red buttons. So I'm on my profile page where I can see data assets, our universe, MCP connection instructions, examples of how to query via SQL, how to query via Python, how to access our rest endpoints where you have a limited analytics and data retrieval, and how to access Parquet virus three. So let's first go to data assets.

This will list all of our datasets, and we can see they divide into a couple of categories. I'm just going to filter on United States. If you scroll up to the top Now, historically, we provided the US SIP, so the consolidated tape, which provided level one data for US markets across venues, and that's provided through the US comp databases, whether we're looking at tick data, one minute bars, daily metrics, latest pricing if we're subscribing to real time data, or market share metrics where comparing trades and quotes to the MBBO. But we can see now there's a whole set of equity markets.

These are the level three markets providing kind of market by order data from NICE, NASDAQ, and CBOE, and other US venues will follow shortly. So let's look at ARCA. We can see we have these symbols coming in, and we have have a set of tables. So both trades, quotes, PRR full will be the market by order data.

And we have indicative pricing around the auctions, end of day records where we're splitting volume across different categories of data or different categories of traits. Let's look at our data explorer, and let's look at NICE.

Let's go Arca. Okay. So if I pick a symbol, let's say Alcoa, we have a set of tables as we saw before, but now we're seeing the data. So we can see quote, trade, PRL full, end for the auctions, market for the changes in market phases, day, which is the end of day record, and then stat, which will be our kind of static data for each symbol.

Now quotes are included in the top of the quotes with prices, sizes, and you can see the order counts on both the bid and the ask side. Trades include every trade, including both the aggressor side, whether it's an odd lot. And if we scroll over to the right, should see the executed order ID and common fields across all databases, which would include the trade period and the book type. And those common fields allow us to query across different venues across different regions of the world.

The market includes the changes in phase or session. Again, we have a standardized field called OMD status. Stat holds the static data for the instrument, so name, MIC, lot size, etcetera. Ind has the auction and balance data including the price and the imbalance volume on the imbalance side.

And finally, we get to p r l full. This holds the book depth data, and this includes every lit passive order message that impacts the book.

So an IOC that just cancels wouldn't be here. It's more than the published market data, but every lit passive order message that impacts the book would be. So we have we scroll over to the left.

We have price. This was representing the limit price. We have size, which is representing the lit quantity for the order. Then we have the trade ID if that particular order message is executing as a trade. The update type, which we scroll here.

This is telling us the type of message. So this is gonna be a for a new order, d for a cancel, p for a partial fill, f for a fill, and m for a place. And we have the order ID.

And if we keep on scrolling, we see buy sell flag.

This is zero for buys and one for sells, or zero for bids, one for asks.

Now each database has a time priority field, which we'll use in ranking within a given book level. So we can see within a good book level what's the priority of each order. We have the fill price and fill size for a particular order message. Then we have old price and old size. So as the order goes through its life cycle, we'll see what the order looked like before.

We can also see old order ID and original order ID. Now this occurs for orders that change their order ID as they go through a replace, and this occurs on the nice and Nasdaq markets.

For the CBO markets, we have a consistent order ID across time. Now we can also see the price level and the old price level. So its order position within the book and its old order position. And then there's size ahead. So that's the total order size ahead of the order in the order key. Now each of these fields is documented in our reference data, and you can go to the data assets listing as we saw before and step through that.

Let's go to the start of the page. We can see this record type zed and record type r. R is a general update record. We've got a new message.

Zed is a specific record type. It's indicating that there's a clear book.

So we're clearing the book at the start of each trading day, then we're adding messages throughout that day.

Now OneTick is also storing data periodically so that we can guarantee fast query times and fast book rebuild. So now we understand the data.

Let's kind of go through the analysis we want to run. So starting with data aggregation and filtering, then order message metrics, which are kind of effectively order counts across time depending on the type of message. Order duration metrics gets a bit more interesting.

When we're calculating the time difference between order entry and order execution or order cancellation depending on when it entered into the book or what level it entered at. Book reconstruction at time, so we're gonna rebuild the book at a particular point in time.

Then we're gonna rebuild the book across time, either continuously or periodically.

Then we're gonna be filtering the output books down to our criteria.

That would also allow us to build up our own metrics based on that filtering criteria, which could be by book level, by price queue, by accumulated size, or by spread.

Okay. Let's close data explorer, and let's go into our SQL examples. Now you can run a SQL query just via REST. Just call the REST call and embed the SQL query, or you can use the Python example that we pre prepared.

Or if you use our Python wheel, which you can pip install, then you'll get data directly back into a data frame, or you can use our flight and get data directly back into a data frame. The advantage of using our Python wheel is we can go and query SQL, but we can also get into query with more of a pandas like syntax, which we'll see slightly later. So there are around two hundred and fifty examples.

We can see this AI query assistant, and all of these examples go into our training sets for our MCP server, which you can pull in to your IDE and your assistant of choice.

So let's start writing queries because we can use this interface to just test inquiries. So let's say I want to query trades from NICE for Cisco from last week for a day.

You can see I'm setting my time zone. I apply, and I get my data back. In this case, the first ten trades. The trade times look a bit strange because I'm looking at my output data set in UTC.

The set's the New York. We can see trades during the kind of early auction, and then we are into continuous trading. And as we saw through the data explorer, we can see our fields. So the odd lot field, aggressive side, trade period, book type.

So if you just swap to quotes, I just need to change the table to quote. Now we're getting our bid and ask price, bid size, ask size, and our bid and ask size well, number of orders on the bid and on the ask. And see where the bid is empty, it's returned as an empty value. Now if we just swap this to look at book depth, now we're seeing our price and size.

So I'll limit price on our remaining lit quantity in the book, the symbol, the side of the book, the update type, the order message, so order ID, the time priority, etcetera. So we've seen individual tables. Let's do something a bit more. We'll count the number of messages by order ID.

Apply.

I can see a lot of orders have two messages. As we've got a large number of orders, let's just look at the average order message count and the total message counts.

So we had a hundred thousand orders, and the orders average one point three messages per order. And our total order message count was two hundred thousand. Now in IC, order IDs change and replace, so let's do this slightly differently.

Now we'll look at a particular order ID or original order ID and apply that.

Now we can see an ad where it comes in as this particular order ID, and then there's a modify, and the order ID changes. Then there's another modify, and it changes again. We can see the price has changed.

Then we get to a loss modification, price has changed again, and then we have a partial fill and a fill. And if we move across time, we're gonna see each time we get our modification or replace coming through, the time priority is also changed.

And we can also see as we go through our order ID, we have our old order ID. So this value is referring to this value. This value is referring to this one. Now we're consistent again. And if we keep on going, we'll see our original order ID being represented here throughout the book.

Now we can see it's actually being at the top of book, and then when it completely fills, it's removed from the book. So price level is zero. So based on original order IDs and orders changing, a better query than we saw previously would look like this. So now we're grouping by original order ID. So now we have ninety one thousand original order IDs, two hundred thousand order messages, and it's around two point two messages per order. So here we're just looking at order counts. But if we extend this a bit further, and let's do this in steps.

So we're putting back some data. In this case, we're saying, is there a trait? So price size and is there a trait. Now we'll go through and we'll aggregate up by original order ID. And rather than just counting the messages, this time we'll look at the first size. So what's the quantity of an order when it comes in to the book?

What was its original order price?

From that, we'll calculate its order value. We'll also sum the trade count and the trade size and also the trade value. So we've applied that. We can now see for each order of that message, how many messages, its price and value, and its trade size and value. And we'll see some that execute.

So now let's wrap this all up with another layer of aggregation. So rather than going by original order ID, now we'll look across all orders.

And if we apply that, now we're just seeing one row returned, which is our order message count, our trade count, our trade size, our trade value, our order count, our order size, and our order value.

And from here, we can then calculate our order to trade ratios, whether we're looking at account of orders or account of messages, or we're looking at by volume or trade size and order size. We're looking by notional, so trade value and order value. And we can keep on building more metrics by aggregating and filtering.

And we could be filtering for certain types of order or certain types of trade, for example. So just show me odd lots. Now we can go further and divide this into buy and sell counts, and then we can have ratios on the buy and sell side. So next step, let's look at order durations.

Now you want to look at average time duration to trade a cancel to trade or cancel an order based on the order level it exit entered the book at because that's one of the fields that's available. So let's look at this in a few steps. Firstly, we want to retrieve the orders that have a price level less than five. So let's have a look.

So we'll bring our order ID. I want the first time the order came in and the first price level of that order. I'm gonna only return the records where the price level is less than five. I'm gonna group by original order ID again.

So now I can see my price level and my original order ID. Let's look the other way. So rather than order came into the book where it's the first time, let's look an order when it finishes. So an order will finish in a fill or a cancel.

So this time, we're gonna pull back the original order ID again. But rather than the first time, I'm gonna take the last time and just verify the last update type. I mean, see a series of deletes or cancels, then that's a fill. So now we have both when an order enters the book and when an order leaves the book.

Let's try and combine those together. So now we're going to join when orders start. So enter the book when orders end the book. And now we're going to as we have the start time and the end time for each order, we can calculate in milliseconds by using the state diff function, the order duration. And as we have the level that it entered into the book as and how it ended up. We end up with for each order ID let's look at this. So this order started at this time, ended at this time, which is a difference of five thousand one hundred and seventeen milliseconds or just over five seconds.

It came in at level two, and its update type was cancel. I will see some here that are fills. So this took thirteen and a half seconds to fill.

That came in at level one. So if we just aggregate this data up slightly further, now we can see the average order duration, the minimum order duration, and the maximum. Let's look at our standard deviation. Just count the number of orders, and we're grouping by the level that the order came into the book in and the update type.

Now we have a series of metrics coming through. Now we're seeing for average durations, these numbers look quite high. We are we're looking across an hour period, so from ten to eleven. So this is what the data looks like.

And we're querying, in this case, nicely. If we run the same analysis for Nasdaq or for Arca or Edge or Bats, we're gonna come back with a different set of results. So so far, we've just been looking at the order messages.

Let's now move on to reconstructing the order book. Let's just clear this. Now in SQL, we can reconstruct the order book. We can also do that in Python. For SQL, we can do that for individual books, but we need to use our Python API for access to a consolidated book.

Now OneTick supports this kind of fast reconstruction of book depth at a specific point in time or across time because of how we store the data. So let's see how we can do this in SQL. I'm gonna ask for the book at a particular point in time, so twelve midday. Asking for NICE for Alcoa again.

I'm using this OB snapshot, and I'm setting my table p r l full, and I come back with an output. Now there are a series of different outputs of how we want to output the book.

This default way, so OB snapshot, for each row, each record returns, we're defining a price level, in this case, level one, and a side of the book. So we have the limit price and the lit visible quantity in the book. But let's say we wanted both the bid NAF side on the same row. Now we can replace OB snapshot with OB snapshot wide.

That will now return one row for both the bid and the ask for a particular level. And I won't specify any filters here, so my price level keeps going down. And we can see as this is full book depth, the price level keeps on going and keeps on going. There's a final method but we do need to apply a filter.

We're saying the maximum number of levels we want to return. This is called OB snapshot and flat. You can see again this time stamp is saying, what's the book at this point in time?

We apply this.

We now see a single record, and that record represents the whole book down to a certain number of levels.

So we have the bid price, when that bid was last updated, its size, same for ask.

So level one first, level two, level three, level four, level five.

So this output format has a dynamic number of columns, which is based on your filter clause. In this case, I'm setting ten columns, so now I keep scrolling. Let's say I want that to be twenty five and apply, and I get more columns back.

So if we go back to our basic view where we had each level in the book or each row in the resulting records representing a level in the book and a side, so one for ask, zero for bid. Let's say rather than returning individual price levels, we want to return the number of orders at each price level. We just add this in, show number of orders at level. So that's true and apply, and I get another field, number of orders. If I'm doing this with wide, then I have bid number of orders and ask number of orders.

Let's put that back.

So so far, we're still at market by level. We just have a bit more information on the number of orders at that level. I want to see each individual order. So now I can specify instead of showing the number of orders, I'm gonna say show full detail is true and apply that. Now we see the price and size. We'll see the level again. But now we have the update type and the order ID and basically the full order message.

It's the original order ID here. So this is returning at a particular point in time. Let's update that limit. The book is representing one thousand eight hundred and eleven individual order messages are in the book exactly at twelve noon. So I'd like to filter this output down to something that I find more useful.

So I could filter by book level, and we saw this briefly when we were looking at one of the output structures. I'm just including into my OBI snapshot function.

Max levels is five. Apply that.

Now we only have nineteen records. We've got our individual price levels and individual orders. So we have free orders on the ask side at level one. Three orders at level two. Two orders at level three, etcetera.

If I don't want to filter by maximum levels, I could equally filter by accumulated depth. So if I wanted to trade ten thousand shares, now we have a few more rows coming down.

First, on the ask side, and all of these sizes when accumulated together will add up to a thousand. Once they're complete, then we'll go on to the bid side of the book, which is a much smaller set of records, and that's because their sizes are a lot bigger. Or rather than going by depth, we can say, what if we want to only trade up to a certain SKU from best?

We can filter on map depth for price, and this is as a percentage. Let's apply that. Now we're getting our price SKU. Or finally, we can have spread.

Now this is an absolute spread. So if I wanted to trade a spread of one dollar, now we're seeing our spread where we have our worst ask. It's at forty seven dot nine seven, under worst bid, forty six dot nine, covering our spread of one. So each of these filters, we can either use independently or combine them to retrieve our book.

Now at the moment, we've been looking at a particular point in time, but we don't have to.

We can return a book across time. In this case, still op snapshot. I can say max levels five.

I'm using this it's running aggregate equals true. I'm looking across a thirty minute period. So the time stamp is greater or equal to our start time and less than our end time. And this will return every book update. In each update, we'll see the same time stamp during that buck book update, and then we'll move on to the next one. And we can see our time stamps are kind of a nanosecond accuracy because that's the accuracy that we're being provided.

So this is giving us every single book update, which across a day and across a liquid symbol can be a lot.

So we could say we want to output the book periodically. Now rather than using this is running aggregate, we use bucket interval equals sixty. In this case, every sixty seconds. Now every minute, we'll see our output book.

Let me just do this the right way. Sorry. K. Every minute, we're seeing our output book.

As we're seeing it to five levels, see first five levels of the book on the ask, then on the bid, then we're going to the next minute. If I wanted it every second, I just change that. I can change that down to every millisecond where I'm outputting the book periodically every millisecond. If I want to go below a millisecond, then I can go back to my is running aggregate.

It's true and output each individual book update.

So we can see the individual fields being returned.

We saw the book at full detail. But if I don't want to output the book, but instead I want to return statistics on a field to view of the book, I then instead of using OB snapshot, I can use OB summary. Let's look at this. OB summary bucket interval sixty.

So every sixty seconds return our output metrics where we're going to trade a thousand shares for Alcoa and apply. This comes back with a series of fields. Inside the results, we see bid size and ask size, which is set to a thousand. This is saying that the liquidity is there.

If I change from NICEE to Amex and apply, we'll see our liquidity is a lot lower. So I can't trade a thousand shares on Amex. They're just the book isn't big enough.

Let's go.

Nasdaq. Yeah. The liquidity is here.

Nasdaq.

Let's go back to NIC. So in the return schema as well as is the liquidity there, we see the ask VWAP and the bid VWAP.

So what price would I actually get if I traded the liquidity down at each of the price levels to get to my thousand shares?

There's also the best bid price and the best ask price and the number of levels in the book I need to eat through to execute that quantity. So based on my fields where I have my bid and ask VWAP, if I just wrap the statement up a bit, I can calculate my bid skew and my ask skew and calculate my effective spread. If we're looking during the day, we can see our values every minute. And if we look at our charts, you can see our effective spread as we go across the day with just above zero.

Let's try this again. Now we can see our effective spread changing throughout the day. And we can repeat this exercise across different venues and see how different venues, both in terms of the effective spread and their bid mask skew change. Now we may may be more focused on returning statistics based on skew, based on spread, based on levels, all the shares that we just saw.

And instead of calculating them every minute, I could calculate them on every book update using this is running aggregate equals true instead of bucket interval.

So let's try that. So same kind of data with running aggregate is true.

Let's limit a thousand. So we can see the liquidity here and our values.

And we can see there's lots of rows because it's every single book update. So now we can calculate the book across a time period of interest because we can then TWAP our results.

So TW average for a time weighted average for our bid and ask price or effective spread and our SKUs. Now we get one row back. Now we can repeat this exercise for every individual venue, and we can keep on going building more complex queries. Now we can see the queries can, in SQL, have multiple levels of nesting, and the SQL can get a bit unwieldy. So everything we've seen so far, we can do in Python, and we can go further in Python. Also, if we look here on the SQL side, we can see there's a selection of examples and book depth. If we go to Python, well, we're following a panda style where we define a data source that looks like a panda's data frame, and then we define operations that look like operations on a data frame.

What instead we're doing is building a query that's getting sent to our servers. We execute where everything's running in parallel at c plus plus speeds, and we're sending back the results as a data frame. And, again, here, there are a whole set of examples on working with book depth. So on the Python side, everything we've been doing, we can do in Python, and I'll very briefly go through this.

We can also go further and do a consolidated book reconstruction. So let's look at the same. So if we're retrieving OB snapshot again, now we're first defining our data source. In this case, I'm selecting NICI, picking our table PRL full, doing our OB snapshot.

We're then running that query. And here, our start time and end time are the same.

In Python, we always have to spec specify start and end time even if we're looking at a point in time. So I've just set these two values to be the same, and I'm asking for our color again.

And that will return our results, which will look as we saw in SQL. So if we wanted the wide version where the bid and ask were on the same output row rather than on different rows, we'll just append sort of OB snapshot. We'll have OB snapshot white. And if you want the variation where the output is a single row per book update, where we have lots of columns based on our filter for the maximum number of levels, then we can use this syntax.

So instead of wide or none, we're now saying flat. And again, start and end time are the same. And as we saw in SQL, if we wanted full book depth, we're specifying show full detail equals true. And if we want to full book depth down to a certain number of levels, show full book depth is true and maximum number of levels set to five.

Now if you want full book detail, we have to be using OB snapshot because each of the other output variations is designed for market by level. You can't have the bid and ask orders on the same time, on the same row because they could be representing there could be multiple levels in the book. Same goes if we want to combine a single row for an order book. It doesn't make sense for full book depth.

So full book depth is always gonna be a b snapshot or we're creating metrics.

And our filters could be based on maximum number of levels, maximum accumulated size into the book, so how much we want to trade, for example. If I want to trade a thousand shares, that will give me what the book looks like at that point in time.

We wanted to trade down to a price skew, so a percentage away from the best. And this skew is from best rather than from mid. What specifies the price skew, or we can specify an absolute spread.

And as we saw before, this is showing us at a point in time, so start and end are the same.

We will look across time to output the book every sixty seconds down to five levels. So that could be every second, every five seconds, every fifteen minutes, or every millisecond. We're just choosing how often you want to output the book.

And if you wanted to output the book continuously, so you have a very, very liquid book, you don't want it at every millisecond, you want it every time a book changes, then we'll use running as true. We can also output rather than the full book at every change just the changes in the book with another argument. So we saw all of these in SQL, and we saw them executing. Let's look at statistics.

So just as I used, let's go back into here and let's put depth at quant you can see OB summary here in our examples. If we look at our SQL examples and all the examples, In this case, we're using our sample LSC dataset, which again is full book depth. So full book depth isn't just for the US. There are a whole set of markets, both across equities and across futures and options and spreads.

And there's a previous webinar that I ran on CME book depth, for example. Okay. So we can see in SQL, OV summary, where I'm specifying our bucket interval and max depth in shares. I am my equivalent here.

I'm asking from nine thirty until four PM. So we had our example where we calculated our bid and ask queue. And that example in SQL looked like this. We're running our function OB summary, then we're counting our SKU and effective spread.

In Python, pretty similar. First, we do our OB summary, output every seconds down to a maximum depth of a thousand shares or a maximum accumulated bit size. Then that's gonna return our VWAP for bid and ask. We want to calculate our effective spread from that, and our ask SKU and bid SKU is gonna be our VWAPs compared to our best prices, and we run that across our time range.

Then the last example, once we've calculated our SKUs and effective spread, then we use our aggregation function and time weighted aggregate to come back and pass in a dictionary of aggregates, and then we'll get our results across the time window. In this case, one record for the output.

Now each of these queries is looking for a per symbol or or per venue basis, but we can also combine books together. So in this case, I have a symbol list. I'm just looking at four venues, ARCA, Amex, Nasdaq, and ICEE. I'm choosing which venues go into my analysis. Now I'm retrieving the book down to three levels, and I'm running the book.

Now I have another example where I'm pulling back OB summary.

So retrieving my stats down to a thousand shares for a particular point in time, and now every sixty seconds across the trading day. Let's run this one. So first query, we were pulling back the book at a point in time to three levels. Second query, we're seeing if we trade to a thousand, how it would look.

Third query, We're doing this every minute. Okay. So now we're getting the book being retrieved. I'm just seeing six rows, three for either side.

I'm not seeing an output book per venue because this is now consolidated across all of the venues that I've selected. Similarly, when I'm trading a thousand shares, this is my bid VWAP and ask VWAP if I traded across all of the markets, picking out the best prices. And similarly here, I'm doing the same per minute. So I can see how my bid view app and ask view app changes across time.

So the only difference here is that I'm passing in, so not asking for a single symbol where I'm specifying my venue up front. Now I'm specifying a set of symbols across different venues. Now I can run my query across consolidated book.

Okay. We saw lots of examples, so let's just look at data visually now to finish off with. So we can use our displays and running exactly the same queries under the surface to look at Cassandra did replay of order messages across time and at time. So to look at the book visually, and then we can look at our metrics as graphs again across time. So rather than looking at our Seahawall display, let's look at book replay. Let's look at Home Depot. That's coming back.

On the graph, I'm showing the MBBO, and I'm showing my bid and ask across today. This is the US regulatory MBBO, and I'm seeing each individual message order message or at least the first fifty thousand order messages. So you can see the spread is very wide in the early and evening sessions, then it narrows during the main session.

Let's look at Apple.

Apple's spread is actually pretty consistent throughout the day, both including the early and evening sessions, but most of the volume is during the main session.

Let's go back to Home Depot. Let's zoom in to a time window. We're now seeing more of the bid and last. Can we see the interplay between IC edge and Nasdaq? And as I select a particular order message, I can see what the book looks like.

And the book is either showing us bids and asks, or we can look by source where I can see bats, nicey, NASDAQ.

So we've got Nasdaq on the ask top of the ask side, a nicey on the top of the bid, along with a little bit of Nasdaq, and then we have bats again. Now here we're looking across all venues.

I can filter that to, let's say, just show me Nicey or just show me Nicey and Nasdaq together. Now we get our consolidated book on the right hand side or bring back everything. Let's zoom in again.

Let's go from let's drag this a bit to the left.

Okay. Let's go ten twenty to ten forty. Now we've zoomed in far enough, we're starting to see trades highlighted. Either the trades are occurring on the BBO, or we can see trades inside the BBO. And we have a good set of order messages coming from different venues. So if I just click the forward message, forward button, I'm just stepping through each message in turn. And as I step through the messages, the book updates.

So it's by default showing each individual order. The size of the box is representing the order size, and the position is representing its time priority in terms of its stack and on the y axis in terms of its price. So greens for bids, reds for asks, and we can either have the size proportion to the order size or just the order count. So you can see we have a lots of orders here, less prominent. Now the y axis can be just straight numeric, so we can see the gap of the spread.

We can hide the gaps, and then we're seeing each individual price level. We can also increase how much the book we see. So rather than top ten levels, we can filter that to the top twenty five levels.

Also, as we have the level here each time in the order book messages, we can filter the messages down to a certain number of levels away from the best so we're not having to step through messages that aren't impacting the book. So let's go from ten fifty five to twelve o five. So let's zoom out. Let's zoom in again. Let's look at this ten minute period, and we can see all of the trade inside and the book changing.

Now during this selected time period, I also have a set of stats so I can see what's been trading across each individual venue. And I've divided those stats based in volume, based in trade counts, and based in notional trade value. And as we have where the trades are coming through, which side is aggressive, in this case, most of the aggressive trades on the ARCA side have been buys, while on Nasdaq, most of the aggressive trades have been sells. So here we've been looking at replaying the book message by message, filtering as we would like. Let's look at the book visually, not just at a point in time, but across time.

Okay. And let's change this from five levels to twenty five levels. Now on the bottom of the screen, I see my size up to twenty five levels. We can see that expands dramatically in the kind of evening session, although our spreads are probably much wider.

And then if I zoom in to the main session, we'll see our book across the day. Let's look at the same kind of time period. Let's say eleven fifty five to twelve o five. Now we're replaying the book where on the right hand side we're seeing each individual order message either by side or by the source venue.

On the bottom, we have both the accumulated size across all of the order books that I've selected in terms of the bid and the ask and then the imbalance.

And on the top, we're seeing each individual price level where the intensity of the color is representing the size left.

So if we zoom into here, we'll see that there's a large set of orders on the bid side. Also, on the r side, we have various levels.

Let's just zoom in to this time window. Let's remove the gaps. So now we can see we have bright colors here indicating we've got this two set of orders.

We have a bright color here indicating this kind of big volume and going out. Now in this display, we see lots of kind of horizontal bands, and those are kind of price levels without a lot of movement.

So let's see how this looks per venue. So if I look at Amex, there's just not pretty much there. Let's look at Arca. Now the top of this book is active, but it's rare around the top of book.

I can see there are some messages that kind of trade around it. Let's take that down to ten levels. There's a bit of activity around. BYX, just straight horizontal bands.

BZX. Okay. That's more interesting. We can see that it's an asymmetrical book. It's jumping up to the bid and then the the ask and then the bid side.

Edgar, fixed bands again, Ajax, more bands away from the bid and the ask. Let's look at Nasdaq. Well, that looks look like a real book, and I can see there are kind of price levels moving away. PSX is gonna be fixed bands. Nasdaq, Texas, not much.

Let's look at NICE. Now here we're seeing the kind of the pricing at the top of book, and then we see some other levels with fixed bands, some other levels which seem to mirror the book. So different venues, depending on the instrument, will have different profiles, Whether they're looking at the book and they have fixed bands or whether they're moving and have price levels that are outside or whether they have price levels outside of the best which are moving consistently with the best about a couple of spreads away. Okay.

So finally, as I've been using Home Depot today, I can pick my day and see across my time window what I want to look at. Let's zoom in. Okay. I'm seeing a series of markets across my time window.

Only show those markets are showing me my size based on how much I wanted to trade, which is here. And based on the venues that I wanted to select, what's the liquidity there, what's my effective spread, as an absolute value and as a percentage, and my price skews. So I'm using exactly the same OB summary queries to pull back the data. But rather than doing this for a single venue, now I'm doing this across the venues that I have access to.

And I can look at this across the day or across particular time windows or look across the average day to see is there correlation there. Okay. Thank you for watching. Hopefully, this gave you introduction to the analytic capabilities of OneTick and OneTick Cloud and how you can query book depth effectively without having to write loads of code within a couple of lines.

You can rebuild the book, pull back statistics, and perform your analysis.

Thank you, Peter.

At this point, we will answer any last questions. Peter, do you have anything more you'd like to add to those that have come in?

Okay. I would reemphasize that although we're showing US book depth today, the same applies for other markets. So in Canada, we store all of the Canadian markets at kind of full book depth. We do the same for most European markets, and we're gradually kind of expanding our book depth in Asian markets. The same goes across derivatives whether you're looking at kind of CME, ICE, UREX or smaller derivatives venues.

You can, in the same method that I showed today, query the book or create a consolidation of different venues and query that. And that could be around surveillance where we're kind of replaying through activity as we saw in the book replay, or it could equally be around investigation microstructure where we're drilling into how the book evolves within a particular venue or across correlated venues, which we've seen today with the US.

There was one last question. So he was asking about modifications to a quote. Do they generate new order IDs?

Generally, if an order is being reduced in its size, it would just be a modification rather than a replace. So it won't generate a new order ID.

But that's the exact rules of whether order IDs get replaced or not to very much venue specific, and that becomes more complex when you're looking at time priority where each venue will have a slightly different definition of when time priority will be updated.

NICE has different rules to Nasdaq, for example.

We are kind of sitting above that because we're taking the direct output of the matching engine. So the Exchange is publishing this data in their feeds, which is being stored as PCAPs. We're processing the PCAPs from each of the venues and providing the representation of the data so you can step through each individual order message that impacts the book.

Now what is in the book or in the data that we have is everything the exchange is publishing, which would be the lit passive order book. If there are, let's say, IOC orders that immediately get canceled and don't impact the book, they wouldn't be shown in this data. So typically, when we're combining the order book data with the customer's own order flow, we'd match based on order ID.

But equally, we'd add in their own IOC orders which won't match because they didn't trade. Same goes with for locale orders.

Great. One last question I'm seeing here. How much delay is there for a real time L3 feed?

Okay. Currently, we provide in real time level one data.

We do not provide level three in terms of our cloud offering.

What we've seen with customers that want to deploy level three data in real time is it can be very region specific.

So if I want to pick up US markets, that's pretty easy. But if I want to pick up Canadian markets at a level three level, then I may be picking it up from Toronto.

if we're looking at level one data, our cloud is hosted in US East. So if we're looking at US equity data in real time, our latency is around fifteen milliseconds.

And that's both the collection time, the difference between the exchange time and when we collect it, and then the loading into our in memory database for querying.

When we're looking at European data, that latency is probably around a hundred milliseconds because we're taking the data across the Atlantic, and then Asian data would be bigger. So latency is significantly driven by where the collection point is and where the exchange is geographically.

Okay. That'll do it. Thank you, Peter, so much for this presentation, and thank you all for attending today. The recording for this workshop will be made available on our website and via email.

Thank you.