Moving to one master: Why multi-strategy funds are rebuilding reference data for real-time markets
Reference data used to be a back-office concern. In markets that reprice on a headline, it has become the foundation on which every trading and research decision stands.
Greg Georges Senior Solutions Architect Published October 5, 2026 Read time 9 min read
Across the top tier of multi-strategy hedge funds, reference data management is being rebuilt. Not tuned, not extended. Rebuilt.
Aerospike has been working with several of these firms, and the same two forces show up every time: growth and consolidation. Together, they explain why firms are replacing infrastructure.
Growth: These funds are expanding into new markets and launching new products inside them. Every expansion brings new identifiers, corporate action types, trading calendars, and cross-references to maintain. Each one stresses an architecture that was not intended to handle it.
Consolidation: Most funds do not run one reference master, but several. Typically, one is built for the trading organization, and another for research, and research itself often is divided further, with the equity analysts on one system and the commodities analysts on another. No one decided to run reference data in four places. It accumulated, one reasonable decision at a time, until one security could be defined four different ways across one firm. Consolidation means combining them into one, so every desk reads the same definition of the same instrument.
Firms are already taking on this project. In July 2023, Insider published an inside account of Citadel, the hedge fund, rebuilding its reference data platform from the ground up1, which is just one example of work happening across the industry. The platform being replaced was roughly two decades old, built when Citadel was largely an equity and fixed income shop rather than the multi-strategy operation it became. It served well for years, and it eventually could not handle the business’ complexity. The rollout reached every desk: equities, fixed income, quantitative (quant) research, commodities, credit converts, and equity volatility. That reach is itself the argument for consolidation: When one rebuild has to serve trading and research across every asset class at once, it shows that reference data cannot stay split between a trading master and a research master. It is one shared dependency, and it has to be built as one.
The stated goal for Citadel's rebuild was speed into new asset classes. "When new opportunities emerge, we need to react quickly,” Robert Tan, lead engineer, told Insider. Citadel built the platform so branching into a new strategy or instrument type could take days rather than quarters. Its engineers also noted that desks increasingly trade across one another's territory, which makes shared, precise definitions between businesses more important. Two reference masters cannot deliver that. One can.
When one of the world's largest multi-strategy managers decides the risk of rebuilding is lower than the risk of standing still, that is worth attention.
Aerospike real-time database architecture
Unlock the secrets behind Aerospike’s real-time database architecture, where zero downtime, ultra-low latency, and 90% smaller server footprint redefine scale. Discover how you can deliver high availability, strong consistency, and dramatic cost savings.
The split was a concession, and it is now a liability
To be fair to the architecture that produced this, the split was a reasonable answer to a real constraint. Trading needed real-time access at low latency and high throughput, because traders backtest machine learning (ML) scenarios across many candidate trades and need answers inside the decision window. Research historically did not carry the same urgency. Put both workloads on one system, and the trading desk inherits the latency profile of the research workload, which is not acceptable.
So the industry settled on a familiar solution: a cache, such as Redis, sitting in front of a persistent store, such as Postgres, or its managed cloud form, Amazon Aurora. The cache handles the hot reads. The store holds the durable data. For a long time, that worked.
It is working less well every quarter, for four reasons.
Markets now move faster than the cache. Tariffs are announced. Conflict in the Middle East changes energy assumptions overnight. The Fed signals a change in interest rates. Each of these events reprices instruments, changes corporate actions, and alters the relationships between securities. Reference data that was accurate an hour ago describes a market that no longer exists. An hour is nowhere near good enough. Increasingly, neither is a minute.
Geography multiplies the problem. A fund this size runs trading desks in New York, Chicago, London, and the major Asian hubs, with research distributed across the same footprint. Every desk needs the same reference data, kept current and available at local speed. The conventional answer is to replicate the master into every region. That costs more and is more work, creates more opportunities for drift, and makes it inevitable that, at some point, two desks will be trading against different data.
Expansion becomes a repeated project. When reference data is stored in three or four places, a new asset class is onboarded three or four times, then reconciled.
Data gets throttled. When the platform cannot handle the load, the response is to ration it. Quant researchers get capped on parallel threads. Write jobs queue for hours before the next research run starts. The demand was always there. The architecture simply refused to serve it, and over time people stop asking for what they know they cannot get. The most revealing number in any assessment we run is not what the legacy system processes. It is how much the legacy system would have processed unthrottled.
What a unified reference master changes
This is where reference data stops being a storage problem and becomes an operational one. Intelligence is about the quality of reasoning, while operational intelligence is about how well that reasoning incorporates the enterprise’s current state and context when a decision needs to be made, said Aerospike CEO Don Dama. Every quant model, every backtest, and every trade reasons against reference data in capital markets, and the quality of the decision depends on whether it’s using current data.
Here is that loop as it runs inside a top ten multi-strategy fund.
Figure 1: The Aerospike operating loop, running against reference data at a top ten multi-strategy hedge fund.
State. The reference data sitting behind every quant model, such as securities, identifiers, corporate actions, entity hierarchies, curves, calendars.
Context. The slice of that state relevant to the decision at hand. Every quant portfolio manager (PM) gets the same low-latency, point-in-time dataset, under massively parallel access, rather than a regional copy that has drifted.
Decision. Quant PMs run models, simulations, and backtests across thousands of parallel threads.
Action. Research and trading workloads execute without throttling, sustaining hundreds of thousands of parallel reads per second.
Consequence. Throttled quant capacity is eliminated, along with write waits that used to run for hours.
Updated state. New reference data and completed trades write back continuously, so the next research run starts from current reality rather than yesterday's.
The last step is the one legacy architecture cannot support.
The center of the operational loop is consequence. An action changes enterprise reality, and the next decision must begin from the reality the last one created. A completed trade changes a position. A new corporate action changes an instrument's definition. When that consequence lands in the persistent store, but the cache has not caught up, the next decision begins from outdated information, and the system produces confident answers about a market that has moved on. This is the difference between a database that happens to serve AI workloads and one built to keep the enterprise current as decisions are made against it.
The scale sets the bar. A core reference system of this kind processes well over a trillion records in a month while serving tens of millions of requests. That understates demand, because the legacy platform was throttled. On the same workload, Aerospike benchmarked at several times the throughput of managed Postgres and better than an order of magnitude over vanilla Postgres, with p99 latency holding in the low single-digit milliseconds.
A unified master built on Aerospike closes the loop. One system serves trading latency and research breadth at the same time, which removes the reason for the split. It also resolves the geography problem, because a cross-datacenter replication approach, such as Aerospike's Cross Datacenter Replication (XDR), keeps regional copies synchronized while letting each desk read locally. The goal is not simply to distribute data, but to preserve a consistent reference point across a global trading operation without forcing every request back to a distant master. That means New York, London, and Asia desks reprice from the same current state rather than from four regional copies drifting apart.
Redis benchmark
Aerospike consistently delivers lower latency and higher throughput than Redis at multi-terabyte scale. It also reduces infrastructure cost per transaction by up to 9.5x under real-world workloads. Download the benchmark report to see how Aerospike compares to Redis in production-level tests.
Quant capacity stops being rationed. Researchers run the thread counts the work calls for, not the number the platform will handle.
AI and ML backtests run against the same reference data the desk trades on, so research results and live behavior stop diverging.
A new market or a new product is onboarded once, not once per system.
Research and trading debate the thesis instead of debating whose data is correct.
When a headline hits, every seat in every office reprices from the same current state, inside the same decision window.
The largest multi-strategy firms are already rebuilding publicly, and they are moving now because the cost of stale, fragmented reference data has become visible in returns.
Where to start
If you are in the middle of this rebuild, or about to be, the first question is not which database to choose. Ask how old your reference data is when a trader acts on it, and whether you can answer that the same way in every office. Then ask your quant team what they would run if nothing throttled them. That’s the real size of the opportunity.
We spend a lot of time on this problem with funds working through these decisions. Want to compare notes on how others have approached the consolidation? We are happy to have that conversation.
Try Aerospike Cloud
Break through barriers with the lightning-fast, scalable, yet affordable Aerospike distributed NoSQL database. With this fully managed DBaaS, you can go from start to scale in minutes.
Bianca Chan, "Ken Griffin's Citadel Is Rebuilding One of Its Most Critical Platforms. Here's an Inside Look at the Project," Insider (Business Insider), July 6, 2023, reprint PDF, https://www.citadel.com/wp-content/uploads/2023/07/Insider-Reference-Data-Team-Feature.pdf.
Additional resources
For a deeper understanding and more insights, explore these additional resources.