Skip to content

Fast restart

For the complete documentation index see: llms.txt

All documentation pages available in markdown.

This page describes the fast restart feature which enables Aerospike nodes to re-join their clusters quickly.

Fast restart is available in Aerospike Database Enterprise Edition (EE) and Standard Edition (SE). It is not available in Community Edition (CE).

Fast Restart (AKA warm restart, warmstart) occurs after a clean shutdown of the Aerospike daemon (asd), during which time asd persists its indexes to their storage. When a new asd process restarts, it does not need to read records from namespace data storage. The server reattaches to the indexes, scans them to rebuild statistics, then rejoins the cluster after all its namespaces have been restarted.

The speed of warm restart is relative to the media the indexes are stored in - shared memory will be the fastest, followed closely by Intel Optane Persistent Memory (PMem), then indexes on flash SSDs.

The default behavior of Aerospike is to warm restart:

Terminal window
sudo systemctl start aerospike
# or
service aerospike start

You can verify in the logs whether a namespace is going through warm restart.

INFO (namespace): (namespace_ee.c:361) {test} beginning warm restart

Namespaces restart independently. Some may warm restart and some may cold restart, depending if the previously described conditions are met. The Aerospike server node only joins the cluster after all its namespaces have restarted.

When does warm restart NOT happen?

There are several situations in which the server cannot warm restart, and will switch to a cold restart. If a cluster node does switch to cold restart, its log file will mention it, and should indicate the reason for the switch.

Warm restart when shared memory is gone

Warm restart relies on the shared memory segments surviving the restart of asd. A host reboot or a container replacement wipes them, so the node cold restarts instead. Two Enterprise Edition features carry that shared memory across the boundary, including the primary index, the secondary index when it is held there, and, for a storage-engine memory namespace with no shadow, the records themselves:

  • The index checkpoint writes a durable copy of a namespace’s shared memory from inside asd, on an operator-issued checkpoint-save, and hydrates it automatically on the next start. Because the copy happens inside asd, this is the only option that covers a container replacement, where no external tool can run. Available in Database 8.2.0 and later as a preview feature.
  • The Aerospike Shared Memory Tool (ASMT) does a similar job from outside asd, for host reboots on bare metal and virtual machines. Use asmt after the asd process stops to copy its shared memory segments to persistent storage. After the host machine or VM reboots, use asmt to rehydrate those shared memory segments ahead of warm-restarting asd.
  • A namespace with index-type pmem or index-type flash keeps its primary index outside shared memory, so it warm restarts without either feature whenever its index storage follows the node.

Can I monitor system shared memory used by Aerospike?

You can see the system’s shared memory blocks, using the command:

Terminal window
sudo ipcs -m

All blocks listed that have keys starting with “0xae”, “0xa2”, or “0xad” are Aerospike shared memory blocks. With 6.1.0 and above, the shared memory blocks used for secondary indexes have keys that start with “0xa2”. With 7.0.0 and above, the shared memory blocks used for record data have keys that start with “0xad”.

‘ae’ (as in Aerospike) are Aerospike shared memory blocks. An instance of Aerospike EE will always have “0xae” keys for the primary index. If EE and 6.1+ and secondary indices are defined then there will be “0xa2” keys. If EE and 7.0+ and any namespace is configured with storage-engine memory then there will be “0xad” segments.

Terminal window
[root@da38772fdefc ~]# ipcs
------ Message Queues --------
key msqid owner perms used-bytes messages
------ Shared Memory Segments --------
key shmid owner perms bytes nattch status
0xae001100 0 root 666 1073741824 1
0xae002100 1 root 666 1073741824 1
0xad001000 2 root 666 536870912 1
0xad001001 3 root 666 536870912 1
0xad001002 4 root 666 536870912 1
0xad001003 5 root 666 536870912 1
0xad001004 6 root 666 536870912 1
0xad001005 7 root 666 536870912 1
0xad001006 8 root 666 536870912 1
0xad001007 9 root 666 536870912 1
0xad002000 10 root 666 536870912 1
0xad002001 11 root 666 536870912 1
0xad002002 12 root 666 536870912 1
0xad002003 13 root 666 536870912 1
0xad002004 14 root 666 536870912 1
0xad002005 15 root 666 536870912 1
0xad002006 16 root 666 536870912 1
0xad002007 17 root 666 536870912 1
------ Semaphore Arrays --------
key semid owner perms nsems