On-Demand Recording | Hosted by A-Team Group

Navigating the Build vs. Buy Dilemma: Cloud Strategies for Accelerating Quantitative Research

Access Global Market Data and Run Analytics Quickly with OneTick Cloud

This workshop, hosted by the A-Team Group, examines how cloud-delivered, managed market data platforms are reshaping the build-versus-buy decision for quantitative research infrastructure. OneTick Senior Cloud Architect Mick Hittesdorf and panelists explore the practical trade-offs, architectural considerations and operational implications of moving from self-built data pipelines to on-demand, pre-normalized data environments.

For many quantitative trading firms and asset managers, building a self-provisioned historical market data environment remains one of the most time-consuming and resource-intensive steps in establishing a new research capability. Sourcing data, normalising symbologies, handling corporate actions and maintaining infrastructure can take months and absorb significant budget before a single model is tested.

At the same time, the expectations placed on quant teams are accelerating. Firms need faster time-to-market for new strategies, broader asset class coverage, and the ability to work with granular tick-level and depth-of-book data across hundreds of global venues, all without scaling headcount or infrastructure linearly.

This webinar examines how cloud-delivered, managed market data platforms are reshaping the build-versus-buy decision for quantitative research infrastructure. Our panellists will explore the practical trade-offs, architectural considerations and operational implications of moving from self-built data pipelines to on-demand, pre-normalised data environments.

Discussion topics will include:

  • The build-versus-buy calculus for market data infrastructure: Why sourcing, cleaning and maintaining historical data in-house remains costly and slow, and where managed alternatives are gaining traction
  • Accelerating quant research workflows: How pre-normalised, API-accessible data environments are enabling researchers to move from onboarding to back testing in days rather than months
  • Cloud-native compute for intensive workloads: The operational case for elastic, vendor-managed infrastructure to support burst workloads such as multi-year back tests and stress testing
  • Data depth and coverage: The growing demand for granular market data, including full depth of book, market-by-order and consolidated liquidity across global venues, and the infrastructure implications of sourcing and normalising it
  • Integration and flexibility: Practical considerations around Python, SQL and REST API access, and how managed platforms fit alongside existing notebooks, pipelines and ML toolchains

This webinar is designed for senior professionals in quantitative research, data engineering, trading technology and infrastructure, including CTOs, heads of quant research, data scientists, portfolio managers, and technology leaders responsible for market data strategy and research platform architecture.

Speakers:

  • Chris Kelliher, Portfolio Manager, Multi-Asset Systematic Strategies at Fidelity Investments

  • Renato Guerrieri, Head of Quantitative Strategy - Liquid Alternatives at Downing

  • Mick Hittesdorf, OneTick Senior Cloud Architect at OneTick / KX

  • Peter Simpson, OneTick Product Owner at OneTick / KX

  • Moderator: Mike O’Hara, Editor at TradingTech Insight

 

Watch the Recording:

 

Webinar Transcript:

Well, good afternoon, good morning, good evening, wherever you are, and welcome to this webinar from the A-Team Group entitled Navigating the Build versus Buy Dilemma: Cloud Strategies for Accelerating Quantitative Research.

I'm Mike O'Hara. I'm an editor at A-Team Group, and I'll be moderating today's webinar. And our panelists for today are Renato Guarieri from Downing, Chris Kelleher of Fidelity Investments, and from OneMarketData KX who are kindly sponsoring today's webinar, Peter Simpson and Mick Hittesdorf. 

So good afternoon, everyone. I'll ask you each to introduce yourselves in a bit more detail once we get through a few of these logistics details for the audience of today's webinar. So we'll actually be running two polls, two audience polls today, and we'll be using Slido for this and for audience questions.

So, members of the audience, if you'd like to take part in those polls and to, ask any questions of our speakers, you can either scan this QR code that's on your screen now, or you can go to slido dot com and enter the code t t I webinars, that's all one word, in the, in the box. We'll also place the Slido web browser link in the chat bar, so you'll see that in the chat.

If you do have any difficulties with Slido, then please type your questions into the chat, and we'll make sure that the panelists see them.

So we do have some upcoming dates for your diaries, if we can bring those on the, on the screen. On the twenty first of May, so that's tomorrow, we have another webinar, and that's Agility is Alpha, How Trading Infrastructure Determines Who Wins in Volatile Markets. And then a couple we've got an event coming up in New York City on the eleventh of June, is the Trading Tech Summit, and another webinar on the thirtieth of June, optimizing cloud marketplaces and managed data services. So please go ahead and register for any of those if you're interested.

So let's crack on with the webinar. I'd like to ask each of our panelists to introduce themselves, and, I'll start with you, Renato.

Yes. Good afternoon, good evening, good morning, everyone. My name is Renato Guadieri. I am the head of quantitative strategy for liquid alternatives at Downing. So my work in a nutshell sits between quantitative research and portfolio implementation.

So systematic strategy design, portfolio analytics, regime analysis, derivative pricing a lot, and stress testing.

Perspective, therefore, on a building versus buy is going to be very practical today. And today, I care less about whether, you know, the infrastructure is fashionable, for example, and a lot more about whether it helps researchers to test ideas, validate them properly, and turn them into, I would define, robust portfolio decisions. For me, the question is, where does the firm's edge really come from? And it is in the plumbing or is the investment judgment built around it?

Thank you. Chris?

Yeah. Sure. So I'm Chris Kelleher. I'm a portfolio manager in the multi asset systematic strategies team at Fidelity.

I've been at Fidelity for about seven years. What we do on our team is we try to build and run systematic strategies in the multi asset class space. So things like multi strategy products, manage futures, some option overlay stuff, and alternative risk premia are sort of all under the the guides of our team. Before I joined Fidelity, I was in the hedge fund world, you know, for about fifteen years.

Over the course of my career, I've had a bunch of different roles, mostly in the quant research and portfolio management space. I also have some experience in academia, so I've taught in math finance programs for about seven years and published a book on quant finance with Python.

Great. Thank you, Chris. Mick?

Hi, everybody. Good day. Thanks for joining. Mick Hittestorf, with OneTick, recently acquired by KX. My background is in finance and technology.

Only recently joined about a year and a half ago, OneTick as a senior cloud architect. But prior to that, I was working as a head of data engineering at an options market making firm here in Chicago where I led the development of a cloud based market data platform that serviced the entire organization and all the quant trading, reporting, and compliance use cases. So I'm gonna bring a very practical and recent experience of actually building a market data platform, having intimately, you know, been familiar with doing that and spent, almost three years, doing that.

And, now I have the vendor perspective. So I've been with OneTick for about a year and a half, so I think I'll be able to lend a very practical and firsthand, perspective on the pros and cons of building versus buying.

Yeah. Some good insight that I'm expecting there, Mick. Thanks. And, finally, Peter.

Hi. I've worked in financial markets for twenty plus years. I'll leave it at that. Mostly sell side and exchanges. And for the last ten years in fintech companies supporting stream processing and tick analytics.

For at OneTick, historically, we've sold tick analytics software to customers for on premise deployments for storing and analyzing their market data.

Now we very much focus on market data on demand for cloud delivery, and we've seen that migration from customers from on premise do it yourself to more cloud focused delivery.

Great. Thank you, Peter. So, quite a wide cross section of perspectives, I think, we'll have on this, on this webinar. So, let's start with an audience poll, if we can bring up the first audience poll.

And it's about market data. So which of the following is the biggest barrier your firm faces in building and maintaining a historical market data environment in house. Now that doesn't necessarily mean, you know, building an entire kind of tick database, in house, but a market data environment could mean many things. But I'm just interested to see what the audience sees as the biggest barrier.

So, while the audience is responding to that poll, we'll start with the first question. And I'm gonna address this one to you first of all, Chris. What do you see as being the main driver of the kind of renewed focus on build versus buy for quant research infrastructure right now? Is it mostly around cost?

Is it a speed to market conversation? Is it something more fundamental?

Where do you think that driver is?

So I think it depends a little bit on what type of data we're talking about. I think if it's data that's very close to the alpha generation process, then there absolutely, I think, is the potential for it to be something fundamental.

If you're a portfolio manager or a researcher trying to, you know, articulate an edge, then if you're getting the data externally, it creates a challenge to understand the edge and also understand how that edge will persist.

If it's more sort of traditional datasets that you're sort of not using as part of your or transforming before you turn them into alpha, then I think the discussion becomes a little bit different.

And in that case, I think it can be more open, you know, for those types of things to be done externally. Another thing that I would sort of flag in my experience is that there's this sort of, like, trade off between the time to market. Right? So, initially, often, we'll do something that we want to get to market very fast.

And to do that, you know, we'll use the data that we have available. And then as we sort of and that makes sense as we're prototyping. But then as we move further and further along, we want to institutionalize that dataset. And as we do that, the cost of that becomes higher and higher as we sort of move further along that path.

And so if you go too far, you've almost missed the chance to externalize that dataset. So I think that's another factor that's going into it as well.

Great. I'll be interested to get Renato's thoughts on this question as well. I mean, are you seeing the same drivers, Renato? Are you seeing anything different?

I totally agree with what Chris said, and definitely the point is very relevant. Depends really much on what type of data you are dealing with, what is your strategy.

And that is one hundred percent impactful.

I would add on I think that it is also partly cost and partly, how do you call it, speed to market. But I also believe that the deeper issues are competitively edged between firms. So firms are asking a more honest question now compared with the recent past, which is which parts of the research stock are generally differentiating and which parts are just expensive pumping because we have seen this happening ten or fifteen years ago. And for most firms, I believe that maintaining another version of a standard historical market data is not where the edge actually sits. Nobody's going to find much alpha just for, you know, historical prices. So the edge is usually in how the data is used as Chris was saying. From my perspective, I can see that signal generation, portfolio construction, risk management, you know, interpretation, and how quickly the weak ideas are going to be rejected.

So in a nutshell, I don't see the builder versus buy as a religious dogmatic debate. Build what is proprietary and buy what is a commodity. The mistake, in my opinion, is just building everything for control or buying everything and then outsourcing your brain.

Yeah. Thank you. That's some good insights there. So I'd like to bring up the results of the first poll if we can do that. So the biggest barrier that our audience sees in building, maintaining a historical market data environment is ongoing maintenance and data operations burden with opportunity cost fairly low down there.

Any surprises there?

Mick, you've done some of this stuff too, yourself. 

I mean yeah. Yeah. I'm actually quite, I guess, gratified to see that and and kind of surprised a little bit as well because that's the one that people notoriously overlook when when thinking about this problem is, oh, I you know, a lot a lot of times the thinking is I just once I get this platform built, I get my data. I'm done. But that's actually really when the hard work starts, especially since the data is being created every day in exchanges around the world. And so there is a requirement to ingest, collect, monitor, clean, verify, validate, distribute, that data, on an ongoing basis, forever, essentially. You know? So that is a burden that a lot of organizations require specialized staff who understand marketing protocols.

Yeah. So hundred percent. I'm glad to see the audience recognizes that as a significant, you know, input into the build by decision. Yeah.

Thank you. And that leads us quite nicely into the second question actually, which is, you know, when firms do try to quantify the real cost, the true cost of a self built historical market based environment, what is it they typically underestimate? I mean, you alluded to the fact there, Mick, that they they you know, what often happens is they underestimate the kind of maintenance side of things.

So is that what people underestimate, Peter? Is it more the kind of the upfront build cost or is there something else? I mean, what do they typically underestimate?

I would still go with the kind of ongoing maintenance. I was actually surprised that that wasn't higher in the poll.

People kind of bake in that volumes will rise.

They are less relevant than data changes. And sometimes it's changes that have been flagged for months or years. So let's say the US SIP adopted Odd Lot quotes.

And then there's the plan to go from Odd Lot quotes, which are level one down to book depth, which keeps getting pushed down.

But there are other changes that we have days or hours of notice. So, for example, the National Stock Exchange of India yesterday, sorry, end of last week, pulled its delivery of level three data.

Because of the regulator changing, they had to limit what they're doing. So then they I did a completely new spec of data on Monday, and that continuously happens. So it's having to always understand each of the venues that you're subscribing and pulling down and understanding that they will change.

Yeah. Yeah. Because those exchange driven changes are constant, aren't they? And there's, you know, so many of them.

Renato, I'd be interested to get your thoughts on this as well in terms of the, you know, again, what users underestimate. I mean, that is ongoing maintenance and do you think users are generally aware of how many exchange driven changes that there are on an ongoing basis?

So I believe there are two orders of impact, at least in our front office, a quantitative research department. So the upfront build is visible when you get approved that you have a budget, more or less, you have an idea from quotes how much it's going to cost.

Even a minimum running order is something that you already know upfront, and that is it's too easy to, at that point, to underestimate what happens afterwards because this cost is what happens immediately once you have it in place. So in my experience, it has been that firms often underestimate all the boring problems. Identifiers, corporate actions, missing data, correction, entitlements, in time logics a lot, versioning, monitoring, you know, all these small breaks that quietly damage your research quality.

And the other because is the researcher's time. So this now is really practical. If your quantum is constantly fixing data pumping, they are not testing ideas, validating signals, or improving portfolio.

So the biggest risk is just a false confidence, in my opinion. You can get a clean looking back test from a fragile data foundation, and that is very dangerous, especially when you are going to allocate capital, not just doing, you know, designing strategies because it does not look obviously wrong. It just makes you believe the wrong thing with confidence.

Yeah. Yeah. Chris, anything you'd add to that?

Yeah. No. I would echo a lot of what the other panel has said. I think that the ongoing data maintenance is a significant part. So where I sit in liquid alternatives, the data itself is quite complex. Think about, like, options.

And so, you know, the management of that can be its own big problem. As a professor, I always used to tell my students, you know, eighty percent of your time should be working on the data, and only twenty twenty percent of your time should be working on the modeling. And I think that's true not just in the upfront build, but also in a steady state where you're sort of monitoring the data on a day to day basis.

I do think the upfront cost can also have sort of a slippery slope as well. And in particular, if you're doing something internally, there can be scope creep where you try to sort of keep everyone happy and manage every requirement.

And then then, of course, there's the trade off between that and generalization. So I think that can also be a factor, and you need to be diligent about that. But I definitely would agree with my other panelists that this sort of ongoing maintenance is sort of the biggest potential issue.

Great. Thank you very much. So, Peter, I'd like to come to you here. If we look at, you know, a managed market data platform where, you know, everything is kind of, bought in or outsourced, that's obviously not a universal fit for everybody. But I'd be interested to hear from you, you know, what types of firm or what types of strategy or you use case are best suited to this fully managed approach?

And where does it still make sense for firms to build and maintain their own environments?

Okay.

I would say it's less around the type of firm and more around their coverage, so the asset classes they're looking at.

So best suited would be a firm that's looking at exchange based data, and ideally exchange based data across a large collection of exchanges.

They may have a kind of lack of knowledge or just resource limitations on the exchange data itself, the reference data landscape, the whole kind of storage loading, the kind of the dead DevOps cloud ops around all of this.

So as the coverage gets larger, there are more data questions to resolve. And if you don't have that level of understanding, you can run into problems pretty quickly. The biggest risks I've seen at customer adoptions have been the people doing the build process don't understand the data.

They may understand infrastructure really well, but they don't understand market data.

So they receive requirements from the quant teams, and they build something based on their understanding of the requirements, and it goes wrong pretty quickly. You end up with abstraction layers, which don't do what people need it to do and end up being thrown away.

If we look at the other side where you are building your own environment, I would say that's best suited for kind of either kind of OTC datasets. Let's say you're receiving FX pricing, the FX pricing is unique to you.

Or you have limited market coverage. You're looking at a single venue which you can focus on over a very small number of venues.

And you have export resources both to manage the IT side of that environment, but also the data side as changes come in.

Well, that all makes sense.

Any of the other panelists that agree or disagree with that analysis?

I guess that means you all agree with Peter, the way Peter sees this. Great. Okay. So just moving on then. How does, and I'll address this one to you, Chris, initially. How does the choice of data infrastructure actually affect the day to day work of a quant researcher?

Does moving to a kind of managed and API accessible environment change the kinds of questions that the researchers can practically ask, or is it mainly about doing the same work faster?

So I think it's definitely about doing the same work faster. Right? So you can definitely get productivity gains from that. I would say even more fundamentally, it provides consistency.

Right? So you know if two people are asking the same research question and using the same data, they're gonna get the same answer. Right? And so I think that provides a value in and of itself.

I also think that doing it, you know, having this sort of streamlined dataset can help you to be able to test things that wouldn't be possible if things are coming in disjoint places. So I think, potentially, it can be a little bit of both. And it can also make sort of things more accessible to a broader range of users. Right?

So if somebody who's a little bit less technical, is maybe gonna be able to interact with the sort of the API in a way that they wouldn't be able to interact with the raw data. So I think across the board, it sort of does provide those benefits.

Peter, your thoughts?

I think it all depends on the API that's provided. If the API is too basic or too limiting, then the quant's just going to pull the data down locally. It's just gonna be a straight data retrieval and then build everything he wants locally, and you've kind of then repeated the problem further down. So that API has to provide a level of value where you accelerate.

There's more complexity there. So when we say faster, we could say, can I run my analysis faster?

There's also the side of that, can I actually develop my analysis faster? Is it quicker to code this than before?

And that has connotations as we move it more into kind of an AI assisted world.

And the final side is faster data retrieval because I'm not necessarily sitting in the same region where the kind of cloud data center sits.

And I may be on a different platform. So it's different connotations and is faster, and it's just making sure that what you're trying to do fits into the platform and the API that's provided to you.

Yeah. And, Renato, how do you see the choice of data infrastructure impacting the day to day work of the quant researcher?

So I agree with what is already being said. From my end, I can add something maybe it's a little more subtle. But when you have a team, in the end, you'll see it happening. And that is it.

Basically, I don't think it's only about, you know, doing the same work faster. I mean, DBI faster definitely helps. But it also changes what questions the researchers are willing to ask. And now I still have in my memory when I was just a researcher.

If data access is low, messy, painful, the researchers will naturally avoid certain questions. I mean, this is human nature. They stay closer to what is easy to test.

That in the end is going to limit the research agenda. So if the environment is API accessible, governed properly, and it's fast, you can test more hypotheses, reject weak ideas faster, spend more time on interpretation rather than, you know, simple low level tasks like data repairing.

So but speed alone is, you know, is not enough. If the research process is weaker internally, faster infrastructure just helps you, again, reach the wrong answer faster. So the real objective, at least for me, is not just the speed by itself. It is time to validate it inside. So I will take a more organic view, not just give to the researchers faster for the sake of it.

Understood. Thank you. I'd like to bring the next poll up, actually, because we've got another audience poll that’d be good to run. So if we can bring that up on screen.

And this just gives us a better idea of where our audience stands at the moment. Where is your firm currently positioned on the build versus buy spectrum for quantitative research data infrastructure? And you can see there are various choices there from fully self built to primary primarily using a managed platform with various options in between.

So again, if the audience can vote on that, we'll go to our next question. So this one, I'll stick with you, Renato.

We think of multi year back tests, stress tests, regime change simulations, they all place very uneven demands on compute. Right? So how is cloud native elastic infrastructure changing the economics of running these workloads, and where does it genuinely outperform traditional kind of on premise approaches?

Sure. So premise is this is unique to every firm, to every process, so it's not a dogma. What I've seen in different firms is that these workflows are naturally uneven. So a normal search may not need much compute.

Then suddenly, once you have, you know, defined your hypothesis, you want to run a multi year backtest, for example, or scenario grids, stat stress, parameters of WIP or regime change simulation across many assumptions at once. And that is where, you know, the Elastic Cloud infrastructure can actually make sense. You don't necessarily want it to own permanently infrastructure for big workloads that you only need, you know, occasionally in the end. In my area, for example, if you are testing portfolio behavior across different regimes and when you are, for example, in macro strategies, at least it's twenty years.

We start from twenty years.

Or if you are doing repricing scenarios for derivatives or running stress grids, the ability to scale pure temporary can change the sales cycle massively impact. So you can ask every question without waiting for days or permanently overbuilding your local setup.

But at the same time, I will say that, you know, it’s not automatically cheaper and it's not automatically better.

You still need the cost control or reproducibility, data locality, governance especially, and a clear understanding of what you are actually running. Otherwise, cloud, you know, it becomes it just converts technical inefficiency into a larger invoice. And so my view is that cloud works best when the workload is adversity, parallelizable, and controllable.

Without control, it's just another expensive flexibility.

Yeah. Yeah. Indeed. Can we bring up the poll results? Please have a look at those and get the panelists' thoughts on this.

So, most firms, according to our audience, are currently running a kinda hybrid mix of self built and managed components. I guess there are no real surprises there. Anyone like to comment, Mick?

So it’s about what I would expect. Also, I think it's actually very healthy.

I think the firms that we see that are making effective use of data analytics and data and marketing infrastructure are doing just that, perhaps using the cloud for, its elasticity, it the the time to value, the fact that you can dip your toes into, you know, into the water and experiment with a new dataset or new API, a new idea, get access to data that is, you know, too expensive to bring in house. Maybe you're contemplating moving into a new, international exchange like Brazil or something and just need access to the data to do some analysis.

That's all great for the cloud, but perhaps maybe you have a more latency sensitive strategy. Maybe you're more sell side market making, something that makes or you wanna bring the data, you know, in house on-prem and collect that data yourself and maybe do real time, signal generation, something like that. So that may be a better fit, for, you know, in house infrastructure. The two are very complementary. They're not mutually exclusive. So, yeah, I would expect to see that. And I actually think, again, that's quite quite healthy, that strategy.

Yeah. Yeah. Understood.

So, Peter, I'd like to come to you now with our next question, which I can just bring up here.

So I figured about granular market microstructure data. So if we think about full depth of book, market by order, consolidated liquidity liquidity across, you know, hundreds of venues potentially.

The need for that kind of data seems to be growing across both systematic and discretionary firms. So again, if you talk to us a little bit about what's driving that and where the real challenges are in sourcing, normalizing, and working with data at this kind of scale?

I would say it's not necessarily just level three. It's also venue specific data.

Let's say there's a history of firms receiving data from a market data provider where it's been over standardized, and that over standardization has mean you've lost nuances. Nuances. So you've lost fields that that particular venue has been providing.

So making that venue specific information available adds value to the quant that's investigating or analyzing that particular venue. And then when you're looking at book depth, book depth impacts price performance.

So if you're just looking at the kind of level one BBO, you're missing what's happening below that.

When you're receiving the level three depth, you have the choice of do you receive that depth from an aggregator, which is, standardizing, so you're losing some of the benefits, or you're moving more into a PCAP world where you're taking exactly what the exchange provides and then pulling that data down and loading into formats that you can analyze.

I think there's always this trade off with how much you standardize.

So there's a consistent schema across different venues and how much you add on the venue specific information so you're not losing it.

Now, those datasets are big.

When we look at the data that we're storing, which is around four petabytes, three quarters to eighty percent of that is booked up.

So it's just it's adding to your volumes. It's adding to complexity.

And also, it's stateful where you're looking at level one data. If you miss a quote, yes, you've missed that quote, but the rest of the day is consistent.

Well, if you miss, let's say, one second of an order book, the rest of the day's order book is complete garbage.

So there's a much heavier demand around data quality analysis when you're putting that data in.

And as we're definitely in North America and Europe in a very much fragmented liquidity world around equities at least, you're having to both retrieve and store and manage high quality level three datasets for each venue and then combine them into a consolidated book.

So all of that is big data questions, but also big management questions of that data.

Again, it's I guess coming back to where we started, it's the ongoing maintenance of all of that is, yeah, I think it's hard to say each venue has nuances.

It's understanding those nuances and understanding how those nuances change across time.

Yeah. How you can look from going from an individual venue to a country level or in the case of Europe, a pan European level.

Just looking at some of the questions that are coming through from the audience, there's one that is kind of, you know, relevant to this, although it's a little bit, tangential, but it is very relevant. And that the question is, how well are managed platforms handling deep historical data quality? So things like corporate actions, symbology changes, survivorship, bad tics. Is this solved, or is it still a source of friction? And I'd like to maybe put that to either Chris or Renato as one of the kind of market practitioners as opposed to our vendor gentleman here. So, Chris or Renato, how well do you think managed platforms handle deep historical data quality?

You're both on mute, by the way.

So I think in my view, it depends on the platform. I think some and I'll let Renato chime in as well. But I think some platforms are pretty good about handling this, and other platforms are less good. And so you have to maybe have some of your own data quality checks even on top of the platform.

And, generally, the sort of better data quality sources maybe have a little bit of a higher cost has been my observation.

Would you agree with that, Renato? Anything you'd like to add?

Chris is right, but I believe this question is rabbit hole because here we start the conversation which we are never going to end. For my experience, at least because I'm multi asset, what I've seen is that you are not going to find anyone which is good at everything. So if you have derivatives and you have problems because the derivatives maybe are a OTC or is direct contractor. I have exotic options, for example, is direct contracts with with the banks. If an institution doesn't have a good linkage with another institution, and you cannot expect them to have all of the ten major banks, for example.

But maybe they are great with managing corporate actions in small caps. So it's really difficult to have someone really good at everything. Even if you take the best of the best of the best for data management like Bloomberg, it's not that great with OTC contracts, and you need to babysit Bloomberg in order to give you the right ticker. So I agree with Chris, but here, I'm not able to point... I mean, we cannot expect anybody to be absolutely able to cover one hundred percent of all asset classes. And I just mentioned two asset classes.

Yeah.

Just finding the right platform for what they do.

Yeah. For what they do. Yeah.

You. So, Mick, coming to you. One concern that firms do raise about moving to a managed platform is vendor dependency. Yeah. Become coming too dependent on a vendor.

And, like, what might happen to the research capability if that relationship changes or if the platform evolves in directions that the firm might not want. So how do you think firms should think about mitigating that risk?

You know, both contractually and architecturally?

Yeah. One, I think it's good to get to know your vendor upfront, right, to do your due diligence. You know, does your vendor have a strong reputation, a long history, you know, has established relationships, the trust of, you know, leading financial institutions, all the stuff you would do, when you're really contemplating any, you know, significant purchase or making a commitment. Right? So that's kind of maybe already understood.

Beyond that, you know, things still do change. So as an architect, right, I always look to, you know, draw on good basic architectural practices. Right? So lean into interoperability, you know, try to only take advantage of proprietary features when they provide distinct and differentiating value.

So things like open data formats, Parquet, and, you know, Iceberg and and some of these new standards are designed to allow you to, you know, have greater ownership of your data and to use common tools and APIs. So for example, pandas. Right? If you have a research team using Python, which most, you know, research teams do, you know, do have that skill set. If your vendor offers a pandas type style API, again, you're gonna be able to leverage, you know, the skills you have, the experience you have, and also take, you know, take that knowledge to, you know, the next research set or API or vendor.

And then there are just some basic good architectural practices, you know, around, you know, encapsulation of certain things. If you feel like, you know, that's really proprietary and something you wanna protect, you can still use a vendor's API and potentially, you know, build it behind your own internal API. Again, I wouldn't emphasize that as a strategy because there's a cost to encapsulating as well.

But, again, those are some, I think, strategies that are both proven and and and also still very applicable to today's world.

Do any of the other panelists have any thoughts on vendor dependency? And it can be a tricky one to address, particularly if you're a vendor.

But, Chris or Renato, have you had to deal with these risks?

I have had this conversation a couple of times in my career, but I to answer, I would zoom out a little bit because I would think that, you know, about about vendor dependency, both I like to think of it both contractually and also architecturally. We focus too much on the contractually.

Contractually, firms need clarity on access, experts route, continuity, sir service levels, like, maybe it was specifying data usage, of course. But also architecturally, I see a lot of importance.

They need a modular research code, open interfaces, lineage, you know, the ability to move or reproduce a critical workflow if needed.

Convenience is valuable definitely, but it should not become a captivity. This is the only limitation that at some point for governance as the business grows and scale is going to face. And if it wasn't clear since the beginning, some friction may happen, or at least this is what I've seen.

Thank you. Yeah. Go on. Sorry.

Yep. I can add on the contractual side, bake into your initial contract not just what you're planning to do immediately, but what your plans are longer term. So, for example, if you're starting looking at a single venue, let's say you're looking at kind of US derivative markets, but you know you're planning to expand to European derivative markets, lockdown what the cost of that changes. And the same goes with compute and storage if you're planning to do some kind of elastic back test at some point in the future, how much is that gonna cost you so you don't get the surprise?

And then on just a broader form, we talked about kind of proprietary standards interoperability. There's also proprietary data. So if a vendor's providing you a proprietary symbology, and let's say you're looking at European equities, and they're providing you proprietary trade classifications.

Why would you want to go down that path when you've got standardized MiFID classifications you could use, and you have standard symbols you can use, whether that's kind of Figgy, I Sin, Seadle, QSIP, etcetera?

Just try and avoid proprietary lock-in where possible because the price might be very attractive in year one, but, you know, due to year three, it's going to change.

So and and ensure your vendor kind of adheres to those standards, I guess.

If I can add something on that. I'm not a portfolio manager. Not yet, at least. And I totally agree with Peter.

What I can say why that happens is the incentives sometimes. And, for example, when you are building your marketing or you are building a narrative, some portfolio managers, on the discretionary side more than on the systematic side, they will need some tailored classifications because it fits their logic and what they are trying to implement in portfolio, which you couldn't represent even with simple statistics if you follow, let's say, the standardized criteria of classification. It's a pain. I one hundred percent agree.

Sometimes it's just incentives and the business structuring. We cannot avoid it. This is why I've seen it happening. Sometimes that weird classification, I agree, but sometimes it's a business need. So from the quant side, I need to be flexible. Also from the, you know, from the provider side, it's handy to have this type of flexibility already in place.

Thank you. So we haven't really talked about AI at all so far during this conversation. We've got forty five minutes in without really mentioning it. So I'd like to maybe spend a few minutes on AI now.

And really, the way I'd like to approach this is to look at, you know, how is AI changing the build versus buy equation? Because if you think of it as a tool for not only as a tool for accelerating integration work, but maybe as a new class of component that needs to be plugged into trading workflows and think of things like MCP and so on.

And the growing use of vibe coding, I mean, that must be changing the build by equation as well. So I'll come to you first on this, Mick, but I'm sure other panelists have got some thoughts on this as well.

Sure. I mean, AI is a force multiplier across the industry. Yes. End users are able to do more with less and take advantage of AIs to build faster, but vendors are leveling up as well. So I see the whole game just moving to a new level. And so there's a temptation to think that as an end user now, you can build it all yourself. Just point Claude code at it and let Claude build it.

There's a lot of domain knowledge, however, that is specialized knowledge with respect to market data that AIs just don't have and, I can't predict the future, but I don't see it coming for a long time. At the same time, like, as I said, your vendors are taking advantage of that same capability and delivering, you know, more software faster, better data faster.

So it's making everyone who is adopting AI more productive and building, you know, higher quality outputs.

So, yeah, I wouldn't, I wouldn't say necessarily.

It may change the bill versus buy equation for some small projects where the code is highly commoditized.

But for specialized market data infrastructure and analytics, it certainly does help. But, you know, your vendors are also aggressively pursuing that and are still probably in the best place to deliver, you know, highly specialized tools and data.

That's my point of view.

Thank you. Chris Chris, have you got thoughts on this–on how AI is changing again?

Yeah. I would totally agree with that. So first of all, I would also add that, like, the evolution is sort of happening rapidly. Right?

So AI does seem to be getting better at certain things. And so I think, you know, it's hard to project that out, but I think that will continue to evolve. I definitely agree that it's a productivity enhancer. On the internal side, it makes it quicker to test research ideas, quicker to code, things like that.

But I would say, like, the bigger benefit comes from people who are sort of already experts. Right? So if you use Cloud Code to write a piece of code for you, if you're an expert on what you're doing and coding, you can get it all the way done faster.

And so I think that's an important thing. On sort of the third party vendor side, I think it has the ability to maybe change how interaction with data happens.

So, like, an API could be friendly, whereas instead of accessing the data in a raw way, maybe a vendor is able to sort of use an LLM to sort of have a prompt-based framework. So I think that can be a powerful concept as well in terms of, like, just just making the whole process more seamless and more productive end to end sort of like Mick was saying.

Yeah. And we're seeing more and more: vendors coming out with announcements about MCP servers and so on, which is interesting.

Renato, your thoughts?

I totally agree with what has been said.

Also, Chris was very sharp, but I would sharpen even a bit more, which is that it's definitely helping on the workflows, but not on the model. I mean, let's be honest.

Maybe because I come from a machine learning engineer background, but we have not invented the new models.

Even if we use and now we made it very available, very easy to do, you know, earning transcript, filings, research notes, all this stuff.

When I started in 2017 and doing all this stuff, we had massive books on machine learning with unreadable academic papers in order to do those exercises, and they were difficult. Now anybody can do it, but it's still the same thing. I mean, whatever chat, chatbot, or agent that you are using, it will call or it will implement the same model that was available already ten years ago. So there is no invention there. It's just made the workflow more accessible and easier.

The other thing is where instead is adding something new or at least spinning up massively is the agent driven research, is real and is it it is today. It is already happening.

But it's not, how to put it, is not an agent… what I mean is that they can help, but they can extract, test, document, monitor. That is useful, and they do it great. But in test, autonomous signal discovery and generation, that one is still limited by the context, especially evaluation, hallucination risk, which in production, you cannot put a factor, hallucination risk, reproducibility as well as, you know, portfolio constraints. So the agents need rails. And yesterday, we were at the conference, and I got a similar question. And at the end, I concluded by saying, you know, the agents need the rails because without the rails, they are basically expensive interns with the supernatural.

So I repeat that.

I like that.

Yeah.

The agents are expensive interns.

Yeah. Nice. I like that.

I'd like to add to that because we see the opposite side where we have we're monitoring all the compute that's occurring in our platform and the queries coming in. And we've seen the amount of queries ramping up significantly, and we can see that ever growing as more of the analysis coming in is agentically driven.

But the inexperienced intern is driving the need, and the requirement for Rails is driving the need for feature sets.

We don't want a customer's AI agents to run thousands of queries to build up tick by tick sets of metrics.

Because that's expensive from a compute perspective, and it takes a long time as you're going back years to generate this data.

So we need to provide more feature sets that are relevant to our customers' needs. And then the agents need to and we need to provide the information to the AI assistant so that they know to look at the feature set and pull back the relevant context, not go and start querying every single quote NBBO and trade from the last five years for a whole market.

And I think you've just answered one of the audience questions there, Peter. If we can just, bring up the audience questions again. I mean, there's one, that's pretty much just asked what you've responded there. Where does AI fit in? Are managed platforms likely to start bundling feature engineering or model ready datasets? And, I guess the answer is yes. I mean, is there anything you'd like to add to what you're just just saying there?

I think definitely the bundling and the feature sets, and that continues to grow based on customer bond because there will always be a point when there's something very niche to a particular customer, and they want it to be niche just to them. So then you have public feature sets, which everyone should know about, and then private feature sets that are specific to a customer.

And then it's exposing the MCP tools in such a way that whoever the customer is gets their own nuanced view of the data assets that are available.

Yeah. Yeah. While we've got the questions up here, just like to maybe ask the panel the question that hasn't been addressed yet, which is the middle one there.

For firms migrating from a self built environment to a managed platform, how long does it realistically take to get researchers productive? And I guess this is one of those how long is a piece of string questions, but I would be interested to get the response.

I don't think this is for you, Mick, do you think?

Sure.

So it can be, you know, hours, you know, with a cloud based platform.

Our OneTick Cloud product allows you to subscribe, self register, install software, you know, in minutes, and begin using the API and pulling down data. So, you know, that's a very… It's an experience, you know, that allows a researcher to become productive again within hours or days, you know, and decline the learning curve, you know, very, very quickly. I think that's one of the primary benefits of a fully managed outsourced platform is that, you know, highly accelerated time to value, which you really can't or doesn't really compare to any other, you know, approach.

Is that… Chris and Renato, does that sound realistic? Up and running in a few hours, getting research productive in a few hours? From the practitioner point of view?

If I can take this one from my perspective, I see it like a circle. If you are first self built, you are basically building around your data, which is at the core.

If you are going to outsource a piece of it, basically, from yourself, it goes to the outside of the managed platform in order to come back to your researcher.

The question says that the research is productive.

I experienced a lot of difference between having the researchers just with the feed and having the researchers productive. So if your researchers have been used to to a different setup for ages, even if technically Mick has done a fantastic job in five minutes, the guys are not going to be very productive. And everything they used to call it or how they used to find it is completely… The engineering is completely changed. So what I've served and I experienced this one in person once is that all your, you know, code base is is not needed to be rewritten, but surely there is extensive work.

So you need to put, you know, you need to program at some times in which you are not going to produce a new feature, so the signals are not produced as smoothly as they used to. AI may do it faster. At the time, there wasn't this option. We needed to re-write a lot of stuff.

Today is going to be faster, but getting them productive, I don't know. It takes a bit. It depends very much on your previous diagrams and, you know, it's very one by one case. I couldn't generalize too much, but I don't see it's extremely smooth.

Yeah. It depends. I think the big question there is what are you trying to accomplish. Right? What's the use case?

You know, how tightly constrained and scoped is that, yeah, if the goal is to migrate a lot of existing models or research to a new platform, you know, whether that's a cloud based platform or something built in house, it's gonna take some time. 

But if you have a very specific use case in mind, let's say you wanna do level three order book analysis of, you know, Deutsche Borse or something like that, in a particular instrument. Right?

You can do that very quickly, again, assuming you have some knowledge of that exchange and maybe some of its, you know, microstructure dynamics, and get access to the data very quickly. If the API is, let's say, based on something you understand like pandas, then yeah. I mean, realistically, within hours to days, right, can start getting some meaningful insights. And so I I think it depends on productive doing what and what the scope is, and that's and that's I think that's where there is some, you know, some a range, right, of time it might take to get fully productive. But there's certainly a lot that can be done very quickly, using cloud based and API based, know, approaches.

Okay. We're, we're coming towards the end of the hour, and I've just got one more question, which is kind of forward looking, which I'll put to each of you. So, the question is, looking eighteen to twenty four months out, what's the single most important shift you expect to see in how quant firms source, manage, and work with market data? And what should firms currently reassessing their infrastructure be paying closest attention to?

And I'll start with you, Chris.

Yeah. So I think, similar to how we talked about AI, I think the potential for the way we engage with data to change is, you know, very possible. Right? So I think that you know, if you look back historically, what we meant by data is very different than maybe potentially what we mean by data today.

Right? So it used to be when we talked about data, we meant numerical values, and now we've had the ability to to parse text and all these things, and we continue to see these innovations. And I think that trend will continue. Right?

So I think we'll continue to see new datasets, and I think our ability to interact with them as potentially helped by LLMs will get stronger and stronger.

And so I think that's, like, a big thing and sort of having adaptable data sources that can handle this type of evolution is probably the most important thing.

Thanks. And same question to you, Renato.

Yes. I'll go fast. I think that the biggest shift will be toward the decision ready research infrastructure. Need firms need you know, they don't just need more data, more datasets, more compute, or more tools. We need a cleaner workflow from data to research, to portfolio decision with lineage, reproducibility, monitoring, governance built in. And the winning setup, in my opinion, will be the hybrid.

Managing infrastructure with scale, coverage, and maintenance are just commodities, and internal ownership where research logic, portfolio construction, and investment judgments are proprietary and are paramount.

For firms that are currently reassessing infrastructure, I would ask one simple question, and I will put it this way. Does this setup reduce the time to validate it inside?

It does, it is helping the investment process. If it is only at the complexity layer, it is probably just another layer of infrastructure there. This is how I call it.

So Mick and Peter, quickly, Mick, you first.

Well, I think that was very, very well said.

I agree one hundred percent with the hybrid thesis, and it can also be a continuum where a firm or a new desk or a new pod can start, in the cloud and be able to very quickly research and eliminate, you know, research ideas and and and strategies.

And, perhaps once something has proven, you know, to have value, it can be brought more in house, and datasets, for example, can be procured and delivered and into an internal infrastructure.

And that can be a a continuum where you move from experimentation, research, or time to insight is and and fail fast is really the goal to then moving to more of a, you know, hybrid or or even in in house infrastructure where there's real edge and the ideas have been proven, and you can, you know, invest in something you know that's going to produce real returns. So we see that a lot in our customer base, that dynamic. And, again, I think that's a very healthy strategy.

Right. Thanks. And last word to you, Peter?

I won't repeat what we've already said on AI, so I'll focus on changing markets.

We know over the next twenty four months, we're gonna have volume growth, then there's an increasing move to twenty four hour trading or twenty four by seven trading depending on the venue, move to more fractional trading, more move to increased tokenization.

There'll be further liquidity fragmentation and where people are looking at data. Here's my real time data coming into my trading system. Here's my historical data.

They'll merge more at level three level. So we see a significant amount of data change coming.

Right. And that integrates into the whole AI framework.

Yeah. Okay. Well, I think it's, we've got interesting times ahead. This has been a really fascinating discussion, gents, so thank you very much, for that.

Just before we finish, just a quick word, about upcoming events again. We've got on the twenty first of May, our webinar, that's tomorrow, Agility is alpha. You can see on the screen now the eleventh of June in New York City, our trading tech summit. So please do register for that.

And on the tenth of September, we have our New York alternative data conference. So we just had our London one yesterday, which Renato was speaking at. It was a very good event. So if you're interested in alternative data, then please and if you're in New York, then please come along to go to that event.

So all that remains for me is to thank my panelists, Chris, Mick, Peter, and Renato. Great discussion. Thanks very much for your time. I really appreciate it.

And to thank the audience for sticking with us throughout the hour and for coming up with some good questions and responding to the polls. And to OneMarketData KX for sponsoring the webinar. Thank you very much. So that's it from me.

Thank you, everyone.

Thank you, everybody. Thank you all.

Thank you.

Bye now.

More About OneTick Cloud:

OneTick Cloud provides real-time, intra-day and historic data and analytics on trading activity leading to actionable insights for sales, trading, and surveillance. With hundreds of customers, you know that OneTick is a recognized platform you can trust. We will cover these many powerful advantages to OneTick, and show you how, together, we can meet your business needs.