Blog
Choosing the best distributed NoSQL database in 2026
Choosing a NoSQL database depends on workload, consistency needs, reliability, and cost, not feature lists. Read the full decision framework.
Blog
Choosing a NoSQL database depends on workload, consistency needs, reliability, and cost, not feature lists. Read the full decision framework.
Database decisions aren’t made on feature lists. A feature list sorts NoSQL databases into document versus key-value versus wide-column versus graph, then compares consistency models and query languages. That sorts out what exists, not what might go wrong for a given team.
The right answer for a specific team depends on the total cost of ownership: the workload, how reliably the database must keep serving through node, disk, and network failures, and how much it costs in production.
Ask engineers who have lived through a NoSQL rollout what drove the decision, and the answer is rarely a feature. It is who gets paged at 3 a.m., and whether the monthly bill matches what finance approved. A database with a richer query language can still be the wrong choice if it adds too much work for the team.
Operational load, cost mechanics, and potential problems are just as important as data model questions as decision criteria, not as an afterthought once the feature comparison is done. That has to start with a shared, accurate picture of what these databases are, because the earliest mistakes come from misunderstanding that.
"NoSQL" is often described as schemaless, but it’s more than that. Every dataset has a schema, whether the database enforces it or the application does. NoSQL does not eliminate schema design; it typically moves more of that responsibility from a database-enforced schema toward application-level conventions and the types of access a team decides to support. How much moves, and how much the engine still enforces on its own, varies by database.
The four models that make up the NoSQL category illustrate that range.
Key-value stores generally do not enforce the structure of the value at all.
Document databases often permit heterogeneous documents within the same collection, though several also support optional schema validation.
Wide-column stores carry some schema elements, such as declared column types, while still allowing rows within a table to differ in which columns they populate.
Graph databases range from weak to optional schema enforcement, depending on the implementation.
None of these four models removes the need for a data model; each one draws the line between "enforced by the engine" and "left to application convention" in a different place, and a team has to find that line before writing code.
Teams that skip that step tend to model NoSQL the way they modeled a relational database, by normalizing the data. That causes problems because many NoSQL systems require access patterns to be considered during data modeling, rather than left for query time. That means treating a lighter enforced schema as a responsibility handed over, not a feature that removes work.
That responsibility changes how a team should approach data modeling: from the first design conversation.
Relational design starts from the data: normalize it, then let SQL join tables at query time. Most NoSQL engines cannot do that cheaply, so the design process has to run in the opposite direction. Before creating a collection or table, write down every query the application needs to serve, including the ones planned for six months out, and let those queries determine how data gets grouped and duplicated.
This is the idea behind single-table design in DynamoDB, popularized through a sequence of AWS re:Invent talks by Rick Houlihan1, then an AWS practice manager. The technique packs multiple entity types into one table, using composite keys and secondary indexes to serve different types of access from one physical structure, trading storage duplication for simpler queries and less work.
The idea spread far enough that people started doing it automatically regardless of whether it fit the workload. Houlihan himself has since had to respond publicly to criticism that teams were force-fitting everything into one table2 because it had become "how you're supposed to do it," in cases where a simpler multi-table layout would have made monitoring, backups, and debugging easier. The technique is a means of matching physical layout to data access, not a requirement to use one table regardless of what that access looks like.
The solution? List the ways applications gain access to the data, before picking a schema, revisit that list whenever a new feature adds a query the original design didn’t include, and treat any model, single-table or otherwise, as a means to that end rather than an end in itself.
Every distributed database determines which node stores a given piece of data by hashing something, but what gets hashed, and how much choice a developer has over it, differs by engine. MongoDB calls this a shard key, and a developer picks fields from the document. Cassandra and DynamoDB call it a partition key, again developer-chosen. Aerospike hashes the record key into one of 4,096 logical partitions per namespace, rather than requiring the application to select a partition-key field. The hash determines the partition, and Aerospike distributes those partitions across cluster nodes to balance data and traffic.
The bigger problem is that workloads can become uneven when the mechanism used to distribute data or requests concentrates traffic on a small number of partitions or records. In Aerospike, the hashing scheme is designed to distribute records uniformly, but an application can still create a hot-key problem if a few records receive most of the requests.
Before MongoDB 5.0, changing a shard key meant exporting a collection and reloading it under a new key, an offline operation that could take days on a large collection. Percona has documented a case of a customer running a five-shard, 50-terabyte cluster with a suboptimal key, where the fix required weeks of custom migration scripts and additional hardware3. It became so much of a problem that MongoDB added live resharding in the 5.0 release4, letting a shard key be changed without downtime.
Not all NoSQL engines support live resharding. Even when they do, changing the distribution key is more work than a routine schema change.
Regardless of which database is on the table, and whether the specific mechanism is a chosen key or a whole-record hash:
Favor fields with many unique values and no natural hot value in whatever gets distributed
Model for the types of access you use the most, with particular attention to write distribution and hot-key risk because those problems tend to appear only at production scale, which makes them easy to overlook in a development environment
Fix data access types that concentrate on one value, one popular user, a trending item, or one date before they become a problem
In 2000, Eric Brewer described a tradeoff among three properties a distributed data system might need:
Consistency: Operations behave as though there were one up-to-date copy of the data, so every request sees the same value regardless of which replica handles it.
Availability: Every request to a node that has not failed eventually receives a response.
Partition tolerance: The system keeps operating despite communication failures between nodes.
Two years later, Seth Gilbert and Nancy Lynch published a formal proof5 that a system cannot guarantee all three at once, turning Brewer's conjecture into what is now called the CAP theorem.
The popular shorthand, pick two of three, oversimplifies what was proved. Partition tolerance is not really a choice because networks always fail. The system has to determine during a partition whether to keep answering with data that might be stale or to hold off answering until it can guarantee correctness. That CP-or-AP choice is fixed by the system's design or configuration, not switched at runtime, but it only comes into play while a partition lasts. Because partitions are rare and usually brief, the tradeoff shapes normal operation far less than the pick-two framing suggests, as Brewer himself noted twelve years later6.
CAP also leaves out another tradeoff that shows up on every request, not only during a partition. Daniel Abadi called it PACELC: if there is a Partition (P), the system trades Availability (A) against Consistency (C); Else (E), during normal operation, it trades Latency (L) against Consistency (C). Even when a system is operating normally, with no partition in sight, it still has to choose between latency and consistency7. In practice, that usually means choosing between accepting some latency to confirm a write everywhere, or responding immediately and accepting that a read elsewhere might not reflect that write yet. Wojciech Golab later gave this second tradeoff its own formal treatment8. How large that latency cost is, however, depends on the replication design; it is not a fixed penalty.
This isn’t academic. A vendor claiming strong consistency, high availability, and full partition tolerance with no tradeoffs is impossible, and it’s worth asking about when a vendor glosses over it.
The reverse claim deserves the same scrutiny: strong consistency does not have to be slow. In Aerospike, for example, a strong-consistency namespace with a replication factor of 2 performs about the same as an availability-mode namespace when no partition is occurring. Ask vendors to measure their strongest setting in steady state rather than assuming it must cost latency.
So what does that mean for your organization? Consistency in a NoSQL context is a spectrum.
Linearizability: in a system that provides linearizable reads, a successful read always receives the most recently completed write, no matter which replica answers or the type of network splits occurring. Some vendors use "strong consistency" to mean this most of the time, but the term is not standardized, so it is worth confirming what a specific vendor means by it rather than assuming.
Eventual consistency: Replicas may disagree briefly and will converge given enough time, with no fixed guarantee of how long that takes.
Tunable consistency (offered by systems including Cassandra and DynamoDB): Lets an application choose per request how many replicas must agree before a read or write counts as successful, trading latency against certainty on a request-by-request basis.
Ask vendors:
Is "strong consistency" here linearizable, or something looser?
Is that guarantee scoped to one key, or multiple keys in the same partition, or does it extend across partitions?
Can consistency be configured per table, or per individual read or write?
What happens to reads and writes during an active network partition?
What is the latency cost of the strongest setting available?
Whether eventual consistency is safe depends on what a stale read costs. A social feed that shows a post ten seconds late typically isn’t a problem. An account balance read at the wrong instant during a transfer, briefly disagreeing across two replicas, spends money that wasn’t available.
The same technology choice that suits one part of an application can be the wrong choice for another part of the same application, which is why several databases let consistency be set per table, or even per individual read or write, instead of once for the whole system. Some NoSQL databases support transactions and strong consistency guarantees, within a defined scope, depending on the vendor:
Aerospike's strong consistency mode supports session consistency by default, with linearizable reads selectable on a per-read basis. An independent Jepsen audit9 found that after a round of fixes, it delivered linearizable single-key operations through network partitions and node crashes, with caveats around clock-skew tolerance and multi-key operations that fell outside the guarantee at the time of that test. Aerospike built on that foundation in 2025, when Database 8 added distributed multi-record ACID transactions with strict serializability.
Cassandra's QUORUM and LOCAL_QUORUM levels give an application a tunable tradeoff between latency, availability, and how many replicas must agree, but those levels should not be treated as equivalent to linearizable consistency; they solve a different, more relaxed problem.
Cosmos DB offers configurable consistency levels up to strong.
MongoDB has multi-document ACID transactions.
DynamoDB offers strongly consistent reads per request, but they consume twice the read capacity of eventually consistent reads, and transactional reads and writes also cost double.
Redis groups commands with MULTI/EXEC, but in a cluster every key in a transaction must live in the same hash slot, and asynchronous replication means it cannot guarantee consistency through a failover.
More generally, cross-partition transactions carry more coordination overhead and latency cost in a distributed system than transactions confined to one partition, though this is a matter of degree rather than a hard boundary: some distributed databases, including several distributed SQL systems, do provide full distributed transaction support, at a performance cost that scales with how many partitions a transaction uses.
Before launch, walk every write path in the application and ask what happens if two users see different answers to the same question for a few hundred milliseconds.
Where the answer is “nothing important,” eventual consistency typically provides better performance.
Where the answer involves money, inventory, or anything else that cannot be reconciled later, it’s worth the latency cost of a stronger guarantee.
The next decision step is figuring out who’s going to run it and whether the team can handle the work involved. Even two databases with near-identical feature sets may require different amounts of maintenance.
A workable sequence runs through four layers:
Type of workload. Does the workload involve heavy, time-ordered writes, frequent point lookups, or multi-hop relationship traversal? In addition, what are the expected data volume and working-set size, read/write ratio, request concurrency, and the latency and availability targets the application needs to reach? This eliminates many options, because wide-column stores, key-value stores, and graph databases are built around different types of data access.
Consistency, durability, and reliability. How much data loss is tolerable if a node or a region goes down, and how quickly does service need to recover (sometimes called recovery point and recovery time objectives)? How does the database behave when a node, disk, rack, or region fails: does it keep serving without manual intervention, and how quickly does it recover and rebalance data afterward?
Operational tolerance. Does the team have, or want to develop, the expertise to run compaction tuning, repair schedules, and capacity planning, or does the organization need to pay a managed service to handle it?
Cost model. This comes last, not first, because it doesn’t save money if you can’t run it safely.
Consider team expertise and existing stack uniformity. For example, an organization that already runs Cassandra at scale for one service has a reason to prefer Cassandra again for a new service, because it already knows how to run it without the operational cost of learning something new. That said, it’s possible to go too far with that reasoning.
Running this sequence before comparing features means no longer looking for the database with the most features, but the smallest number of engines that could plausibly satisfy the first three filters.
At that point, cost becomes a useful tiebreaker, but there’s more to it than the price. What matters more is how a provider meters use, and how much human time a given architecture needs to keep running safely.
For example, DynamoDB meters reads in 4-kilobyte increments and writes in 1-kilobyte increments, rounding up on every request. An item just over a boundary bills as if it were a full increment larger, and on-demand pricing can run several times more expensive than provisioned capacity at full utilization10, though most real workloads run well below that ceiling.
One common and avoidable problem comes from batch operations. A BatchWriteItem call uses the same write capacity per item whether those items are sent in one batch or one at a time11, so a batch of 25 items bills as 25 write units, not one, a detail easy to miss.
Node-based systems such as Cassandra and ScyllaDB have a different issue. While the server bill looks predictable, they require a lot of work to keep the cluster healthy, including tuning compaction, running repairs on schedule, and responding when a node needs replacing.
Part of this is based on the operating model, so take that into consideration. Cassandra and ScyllaDB are both available as fully managed cloud services in addition to self-hosted deployment, while DynamoDB's low operational burden is partly because it’s a managed service.
Independent verification is hard to come by in a category this competitive, so when reading vendor benchmarks, check the workload, replication factor, and hardware configuration, because all three can be tuned toward whichever system ran the test.
When calculating a cost estimate, include how many engineer-hours per month the architecture realistically requires, priced at what this team's engineers cost.
Support requests that cause outages and support tickets are never the average ones. Instead of asking a vendor about average latency, ask for the 99th or 99.9th percentile: p99 is the latency value below which about 99% of requests complete, and p99.9 is the latency value below which about 99.9% of requests complete. Even though it’s a small percentage, databases have so many requests that it still adds up to a lot of unhappy users.
In one case, migrating trillions of stored messages from Cassandra to ScyllaDB cut historical-message read latency at the 99th percentile from a range of 40 to 125 milliseconds down to a steady 15 milliseconds, while shrinking the cluster from 177 nodes to 7212. On the Cassandra cluster, the team also dealt with JVM garbage collection (GC) pauses severe enough that an operator sometimes had to intervene manually to bring a node back to health. The two problems, unpredictable tail latency and GC-driven node instability, compounded each other rather than one causing the other, and together they justified a multi-year migration.
Similarly, LexisNexis Risk Solutions also saw infrastructure and latency improvement moving off Cassandra onto Aerospike, measured as end-to-end application request latency rather than database read latency: average latency down from 120 milliseconds to 30 milliseconds, and server count down from 96 to 28, while the data volume being served grew roughly two and a half times over the same period.
The two cases aren’t identical; the first figure is a p99 database read measurement and LexisNexis' is an average end-to-end application latency, but the result is the same: a Cassandra migration driven by unpredictable performance, resolved with a smaller, steadier footprint.
Different architectures try to reduce tail latency from different angles.
A shard-per-core design, where each CPU core owns a fixed slice of data and avoids contention with the others, removes one common source of latency spikes: threads waiting on each other.
A hybrid memory design that keeps indexes in RAM while values sit on fast local storage removes a different source: unpredictable storage read time on the hot path.
The two approaches are not mutually exclusive, and a database can combine them; which one matters more depends on which bottleneck a given workload encounters.
When architecture alone is not enough, request hedging offers a client-side mitigation: issue a second, identical request after a short delay if the first has not returned, and accept whichever response arrives first. Global Payments documented a 30% reduction in tail latency13 on a DynamoDB-backed credit card authorization platform using this technique.
Hedging is not free, though: it increases total request volume and load on the backend, because some requests now get sent twice. It also adds development and maintenance effort, and extra database cost if the database charges per read. Moreover, it is only safe to apply to reads and to writes that are either naturally idempotent or made idempotent, because issuing a non-idempotent write twice may duplicate the effect rather than just the request. That extra load is also worth considering if a system ever tips into sustained overload.
Tail latency also compounds. The more database calls it takes to serve one user request, the more likely that request hits at least one slow response: with 100 calls per request, about 63% of requests hit at least one p99 response and about 10% hit a p99.9 one. AI applications make this worse, because agents and retrieval pipelines often issue many database calls to answer even a simple question.
Ask for p99 and p99.9 numbers under a realistic workload, ask what happens to those numbers once the dataset outgrows what fits in cache, and treat any latency claim built only on an average as incomplete.
Latency numbers describe steady-state behavior. What happens when steady state breaks is a different question. Some problems never show up in a proof of concept, because they show up only under sustained production load.
Here are three common ones.
Cassandra and its variants handle deletes by writing a marker called a tombstone rather than removing data immediately, and that marker has to survive on disk for a configurable grace period, ten days by default14, so every replica gets a chance to become informed about the deletion before it is permanently forgotten. Skip a repair cycle for longer than that window, and a node that missed the deletion can reintroduce deleted data once it reconnects, because nothing remains to prove the deletion happened. Excessive tombstones also require more reads, because a query reading from an affected partition has to scan past accumulated markers to reach live data. This follows from how deletion works in a leaderless, replicated system, and it is a maintenance task that needs scheduling and monitoring rather than being left to run itself.
Hot partitions cause a different problem: a cluster looks adequately provisioned in aggregate while one specific key, one viral post, one popular product, or one badly chosen timestamp bucket handles most of the traffic and bogs the cluster down. This is why data partitioning is important.
A newer and less widely understood problem is the metastable failure, described formally by a group of distributed-systems researchers15 with hyperscale production experience, and examined further in a follow-up field study16. A brief trigger – a deploy, a network blip, or a traffic spike overloads a system, and a self-sustaining feedback loop, usually retries, keeps it there even after the original trigger disappears. That’s because optimizing a system for high steady-state utilization means it can’t handle a sudden increase in load or latency, which makes a feedback loop, once triggered, harder to recover from on its own.
These three problems are not reasons to avoid distributed NoSQL databases. They are reasons to ask these questions before going live, and weigh the answers.
How does this database handle deletion at scale, and what's the mechanism behind it?
How does it detect and mitigate an overloaded partition?
What happens to retry behavior under sustained backpressure?
There are business questions as well as operational ones, and they don’t always show up at the beginning.
There are business questions as well as operational ones, and they don’t always show up at the beginning.
Portability: DynamoDB runs only on AWS, BigTable only runs on GCP, which works for teams already committed to that ecosystem, but can be a problem when someone wants to move clouds later. Increasingly, organizations are adopting a multi-cloud strategy for risk-mitigation compliance.
Managed service limits: "Managed" does not mean the same thing as "no responsibility remains." A fully managed, serverless database removes server patching and hardware failover from a team's plate, which is real and valuable, but it doesn’t handle:
Schema and access-pattern design
Partition key selection
Index management
Cost-model tuning
Application-level retry behavior
Observability
Backup and restore testing
Teams that treat a managed service as fully hands-off tend to discover this the hard way.
Feature-tier gating. A proof of concept that runs on a free or open-source tier can miss a capability a production deployment eventually needs, and that capability's cost is worth pricing into an evaluation from the beginning rather than being surprised later.
Database | Gated to a paid or Enterprise tier |
|---|---|
Aerospike | Strong consistency mode and Cross Datacenter Replication (XDR), confirmed in Aerospike's edition comparison |
Apache Cassandra | None: The Apache Software Foundation project has no paid tier, though commercial vendors built on top of it, DataStax among them, sell their own enterprise layers |
MongoDB | Kerberos and LDAP authentication, audit logging, and an in-memory storage engine17 |
Redis |
Raising these questions during evaluation saves money and time in the long run.
With those questions answered, now you’re ready to look at features, including their limitations.
Aerospike’s hybrid memory architecture keeps indexes in RAM and data on SSDs, so it can serve datasets far larger than available RAM; namespaces that need even lower latency can instead run entirely in memory. Consequently, Aerospike customers commonly run production workloads on as little as 20% of the server count a comparable NoSQL deployment would need. Its lowest-latency access pattern, with plain key-value reads and writes, is most directly comparable to Redis, and reports lower tail latency and fewer required nodes than ScyllaDB at multi-terabyte scale.
No engine on this list wins across every dimension. The right choice is the one whose specific limitation a given team can live with, not the one with the longest feature list. Pulling all of this into something usable is the framework’s job.
Database | Data model | Consistency options | Deployment model | Where it wins | Key limitation |
|---|---|---|---|---|---|
Aerospike | Multi-model: key-value, document, and graph | Session consistency or full linearizability (Strong Consistency mode); distributed ACID transactions | Self-managed or Aerospike Cloud | Low-latency workloads at large scale, with as little as 20% of the server count of comparable deployments | Strong consistency, XDR, and multi-site clustering are Enterprise-only |
Apache Cassandra | Wide-column | Tunable (QUORUM, LOCAL_QUORUM); not linearizable | Self-managed, or managed through third parties such as DataStax | Write-heavy, always-on workloads at large scale | Tombstones, compaction, and ongoing repair scheduling |
DynamoDB | Key-value and document | Eventually consistent by default; strongly consistent reads available per request | Fully managed only, AWS exclusive | Predictable single-digit-millisecond latency without managing servers directly | No joins, boundary provisioning cost model, single cloud |
MongoDB | Document, with native full-text and vector search | Tunable read/write concern; multi-document ACID transactions | Self-managed or MongoDB Atlas | Flexible, developer-friendly document modeling | Steep learning curve on aggregation pipelines and sharding at scale |
Redis | Key-value | Primarily asynchronous replication; stronger replication/durability can be requested with WAIT/WAITAOF, but it cannot guarantee strong consistency in every failure scenario | Self-managed or managed cloud; standalone, replicated or clustered | Caching, sessions, and ultra-low-latency state | In-memory design limits dataset size to available RAM; while Redis 8 added more features19, it does not support joins across keys or data structures |
| Wide-column, Cassandra-compatible | Tunable, Cassandra-compatible | Self-managed or ScyllaDB Cloud | High-throughput, Cassandra-compatible wide-column workloads | Smaller community than Cassandra's; hardware-tuning curve |
Workload, consistency requirements, reliability, operational tolerance, and cost are still factors.
The “best” database depends on a workload, a team, and a budget.
Start with the workload: write-heavy and time-ordered, read-heavy and point-lookup, or relationship-driven, and let that narrow the field before talking about vendors.
Next, what consistency guarantee does the workload need, tested against what a stale read would cost, not what sounds safest in the abstract.
Then reliability: how the database behaves through node, disk, and region failures, and how much manual intervention recovery takes.
Add operational tolerance: what a team can run today, and what it is realistically willing to learn.
Bring in total cost of ownership only after that, priced as invoice plus engineer-hours rather than sticker price alone.
Only then, compare the few engines standing on their specific strengths and limitations.
That sequence turns an open-ended search into a short, defensible list, and turns a vendor conversation into "confirm this fits the five things already known to matter."
Rick Houlihan (unverified), "AWS re:Invent 2019: [REPEAT 1] Amazon DynamoDB Deep Dive: Advanced Design Patterns (DAT403-R1)," YouTube, https://www.youtube.com/watch?v=6yqfmXiZTlM.
madhead, "On DynamoDB's Single Table Design," madhead, February 23, 2024, https://madhead.me/posts/std/.
Corrado Pandiani, "Resharding in MongoDB 5.0," Percona Blog, August 24, 2021, https://www.percona.com/blog/resharding-in-mongodb-5-0/.
Mat Keep, Cesar Rojas, and Garaudy Etienne, "Scale Out Without Fear or Friction: Live Resharding in MongoDB," MongoDB Blog, January 26, 2022, updated January 31, 2023, https://www.mongodb.com/resources/products/capabilities/scale-out-without-fear-friction-live-resharding-mongodb.
Seth Gilbert and Nancy A. Lynch, "Perspectives on the CAP Theorem," IEEE Computer, 2012, https://groups.csail.mit.edu/tds/papers/Gilbert/Brewer2.pdf.
Eric Brewer, "CAP Twelve Years Later: How the 'Rules' Have Changed," InfoQ, 2012, https://www.infoq.com/articles/cap-twelve-years-later-how-the-rules-have-changed/.
Daniel Abadi, "Problems with CAP, and Yahoo's Little Known NoSQL System," DBMS Musings, April 23, 2010, http://dbmsmusings.blogspot.com/2010/04/problems-with-cap-and-yahoos-little.html.
Wojciech Golab, "Proving PACELC," ACM SIGACT News 49, no. 1 (March 2018): 73–81, https://dl.acm.org/doi/10.1145/3197406.3197420.
Kyle Kingsbury, "Jepsen: Aerospike 3.99.0.3," Jepsen, March 7, 2018, https://jepsen.io/analyses/aerospike-3-99-0-3.
Alex DeBrie, "How You Should Think About DynamoDB Costs," DeBrie Advisory, May 8, 2023, https://www.alexdebrie.com/posts/dynamodb-costs/.
"BatchWriteItem," Amazon DynamoDB API Reference, AWS Documentation, accessed October 6, 2026, https://docs.aws.amazon.com/amazondynamodb/latest/APIReference/API_BatchWriteItem.html.
Bo Ingram, "How Discord Stores Trillions of Messages," Discord Blog, March 6, 2023, https://discord.com/blog/how-discord-stores-trillions-of-messages.
Esteban Serna Parra, Don Anton Dehipitiarachchi, and Krishna Chaitanya Sarvepalli, "How Global Payments Inc. Improved Their Tail Latency Using Request Hedging with Amazon DynamoDB," AWS Database Blog, August 26, 2025, https://aws.amazon.com/blogs/database/how-global-payments-inc-improved-their-tail-latency-using-request-hedging-with-amazon-dynamodb.
"Tombstones," Apache Cassandra Documentation (v5.0), accessed October 6, 2026, https://cassandra.apache.org/doc/latest/cassandra/managing/operating/compaction/tombstones.html.
Nathan Bronson, Abutalib Aghayev, Aleksey Charapko, and Timothy Zhu, "Metastable Failures in Distributed Systems," in Proceedings of HotOS '21, Ann Arbor, MI, May 31–June 2, 2021, https://sigops.org/s/conferences/hotos/2021/papers/hotos21-s11-bronson.pdf.
Lexiang Huang, Matthew Magnusson, Abishek Bangalore Muralikrishna, Salman Estyak, Rebecca Isaacs, Abutalib Aghayev, Timothy Zhu, and Aleksey Charapko, "Metastable Failures in the Wild," ;login: Online, USENIX, June 20, 2022, https://www.usenix.org/publications/loginonline/metastable-failures-wild.
"Upgrade MongoDB Community to MongoDB Enterprise," MongoDB Database Manual, accessed October 6, 2026, https://www.mongodb.com/docs/manual/administration/upgrade-community-to-enterprise/.
"Active-Active Geo-Distribution," Redis, accessed October 6, 2026, https://redis.io/active-active/.
"Redis 8.0," Redis Docs, last modified July 30, 2026, https://redis.io/docs/latest/develop/whats-new/8-0/.
For a deeper understanding and more insights, explore these additional resources.
See more