Checkpoint a node
For the complete documentation index see: llms.txt
All documentation pages available in markdown.
This page describes how to checkpoint one node before planned maintenance, so that it warm restarts with its indexes and its in-memory data intact instead of cold starting and rebuilding them.
See Index checkpoint for a description of the index checkpoint, how to enable and configure it, and how to size and secure the checkpoint directory.
How it works
The checkpoint-save command triggers a graceful node shutdown while writing a durable copy of its shared memory to disk.
When a node completes this process, it enters a parked state, meaning the process remains alive and out of the cluster, serving status commands(such as checkpoint-status or a re-issued checkpoint-save) until manually stopped or until the park timeout elapses.
Shutdown sequence
-
Cluster departure: The node leaves the cluster to isolate its state.
-
State copying: The node copies shared memory segments for each checkpointed namespace to
index-checkpoint-path. If a node has nothing to checkpoint, it skips this step. -
Parking: The node enters the parked state.
Saving a checkpoint is an integrated phase of the shutdown process. A node without data to checkpoint still accepts the command, shuts down gracefully, and parks.
The node is out of the cluster before the first byte is written, so size the maintenance window against the copy plus the park rather than the park alone. The copy scales with the size of what the node holds in shared memory. See What the checkpoint captures for more details.
On the next start, asd hydrates from the checkpoint if the node’s shared memory segments are gone,
then deletes the checkpoint as the index goes live, but only the checkpoint it actually hydrated
from. Hydrating protects one restart. A checkpoint the node did not consume, such as one left
behind by a same-host restart that reattached its surviving shared memory, stays on disk as a
fallback until the next checkpoint-save replaces it.
Prerequisites
- Aerospike Database Enterprise Edition 8.2.0 or later. The index checkpoint is not available in Community Edition.
- The target node has
index-checkpoint-pathconfigured in itsservicestanza and the--preview index-checkpointflag in its startup arguments. See Enable the index checkpoint. - Enough free space on the checkpoint volume for the shared memory footprint of every namespace the
node checkpoints, measured uncompressed.
asdwill start with a short volume, but the copy fails, and by then the node has already left the cluster. See Size the checkpoint volume. - A user holding the
sys-adminrole, whichcheckpoint-saveandcheckpoint-statusrequire. Aerospike enforces that role only on a cluster with asecuritystanza configured; without one, anyone who can reach the info port can take a node out of service. See Secure the checkpoint directory. asinfoandasadm, available at the management host. Themanage checkpointcommands require Tools 13.1.0 (asadm5.1.0) or later. The following examples target a four-node Enterprise Edition cluster with namespacetest, checkpointingnode2at10.0.3.2.
Checkpoint a node
-
Verify target namespaces.
Run
checkpoint-statuson the running node before your maintenance window to confirm which namespaces will be checkpointed.Shell asinfo -h 10.0.3.2 -p 3000 -v 'checkpoint-status'test:state=none:files_completed=0:files_total=0:is_parked=false:park_ms=0A namespace reporting
state=noneindicates a healthy, steady state prior to checkpointing. -
Confirm node checkpoint readiness.
Evaluate the response to ensure the node is configured to write checkpoints.
When
index-checkpoint-pathis unset, the feature is off for the node and the command returns:ERROR:4:'index-checkpoint-path' is not configuredWhen the path is set but no namespace is left to checkpoint, the command does not error. It returns a single node-global entry with no namespace name:
is_parked=false:park_ms=0A node in this state still accepts
checkpoint-saveto shut down gracefully and park, but it writes no checkpoint. Treat that as a warning that no index data will be saved on restart. -
Confirm the checkpoint volume has room for the copy.
Shell df -h /opt/aerospike/checkpoint -
Verify the cluster is stable.
Confirm that no migrations are in progress and that all nodes report the same cluster key. See Verify the cluster is stable.
asadm Admin> info network~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~Network Information (2026-04-13 23:57:26 UTC)~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~Node| Node ID| IP| Build|Migrations|~~~~~~~~~~~~~~~~~~Cluster~~~~~~~~~~~~~~~~~~|Client| Uptime| | | | |Size| Key|Integrity| Principal| Conns|node1:3000 | BB94B7AEB45DB52|10.0.3.1:3000|E-8.2.0.0| 0.000 | 4|AA5AF50552AF|True |BB9D787F6BAF3D6| 5|00:20:34node2:3000 |*BB9D787F6BAF3D6|10.0.3.2:3000|E-8.2.0.0| 0.000 | 4|AA5AF50552AF|True |BB9D787F6BAF3D6| 5|00:20:34node3:3000 | BB989A1BF1D8116|10.0.3.3:3000|E-8.2.0.0| 0.000 | 4|AA5AF50552AF|True |BB9D787F6BAF3D6| 5|00:20:34node4:3000 | BB9D48DF5A70CEE|10.0.3.4:3000|E-8.2.0.0| 0.000 | 4|AA5AF50552AF|True |BB9D787F6BAF3D6| 5|00:20:34Number of rows: 4 -
Quiesce the node.
Quiescing is not required to issue
checkpoint-save, but without it clients hold references to a node that is about to depart, which causes timeouts that quiescing avoids. The node drops to the end of every succession list and hands off master ownership. See Quiesce a node.asadm Admin> enableAdmin+> manage quiesce with node2:3000Admin+> manage reclusterConfirm the quiesce handoff is complete before you continue. Ops/sec on the quiesced node must drop to zero and proxy counters must stop incrementing.
-
Save the checkpoint and watch it complete.
manage checkpoint saveissues the save and then polls the node until every namespace finishes, so this one command covers both. It requires Tools 13.1.0 (asadm5.1.0) or later. It prompts for confirmation first, because a checkpoint cannot be undone. Type the challenge string to proceed.asadm Admin+> manage checkpoint save with node2:3000You are about to checkpoint node(s): node2:3000. Each leaves the cluster, saves its shared-memory segments to durable storage, then parks until it is stopped or the park timeout elapses (300s). This cannot be undone.Confirm that you want to proceed by typing a1b2c3, or cancel by typing anything else.If the node is not quiesced,
asadmwarns before the prompt but does not stop you:WARNING: Not quiesced: node2:3000 (test). Departure will trigger migrations. Recommended: 'manage quiesce with node2:3000', then 'manage recluster', and checkpoint once migrations finish.The node returns
ok, departs the cluster, writes its checkpoint, and parks.manage checkpoint savekeeps polling it in the same session after it leaves the cluster, and prints a table per poll:~~Checkpoint Save~~Node|Responsenode2:3000|okNumber of rows: 1~~~~~Shared-Memory Checkpoint (2026-09-24 21:25:07 UTC)~~~~Namespace| Node| State|Files|Progress|Parked| Park| | | | | | Timetest |node2:3000|copying|3/11 | 27.27 %|False |00:00:00Number of rows: 1~~~~Shared-Memory Checkpoint (2026-09-24 21:25:13 UTC)~~~Namespace| Node|State|Files|Progress|Parked| Park| | | | | | Timetest |node2:3000|done |11/11| 100.0 %|True |00:00:01Number of rows: 1Checkpoint complete on node2:3000.The park timeout defaults to 300 seconds. It bounds only the park, the window in which you stop the process, and does not bound the copy, which begins earlier. To hold the park longer, pass
--timeout, from 1 to 3600 seconds.--poll-intervalsets the seconds between polls, default 2. Flags go beforewith:asadm Admin+> manage checkpoint save --timeout 1200 --poll-interval 5 with node2:3000For the other flags, including
--no-warnto skip the prompt in a script, seemanage checkpoint. -
Confirm every namespace reached
done.manage checkpoint savereturns once every namespace isdoneorfailed, and reportsCheckpoint completeor anERRORline naming the failed namespaces. To read the status again afterwards, in the same session:asadm Admin+> manage checkpoint status with node2:3000~~~~Shared-Memory Checkpoint (2026-09-24 21:25:15 UTC)~~~Namespace| Node|State|Files|Progress|Parked| Park| | | | | | Timetest |node2:3000|done |11/11| 100.0 %|True |00:00:03Number of rows: 1If you lose the session while the node is still parked, start a new one seeded at the parked node. A session seeded at another node does not reach it, because it has left the cluster.
Shell asadm -h 10.0.3.2 --enable -e 'manage checkpoint status'The State column reads
nonebefore any checkpoint has run,copyingwhile one is in progress,donewhen it is complete, andfailedif it did not finish. Files and Progress count the shared memory segment files copied against the total to copy. Parked readsTrueonce the node has entered the park, and Park Time shows how long it has been there. -
Stop
asdand check its exit status.Shell sudo systemctl stop aerospikeasdexits0only when every namespace’s checkpoint succeeded and storage shut down cleanly, and non-zero if either failed. Read that status from the unit. The exit code ofsystemctl stopis the result of the stop job, not the exit status ofasd, and reports0even for a node that exited non-zero on a failed checkpoint:Shell systemctl show aerospike --property=ExecMainStatusExecMainStatus=0In a container, let the orchestrator deliver the
SIGTERM. Do not useSIGKILL, which bypasses the park. Oncecheckpoint-statusreportsdonefor every namespace, letting the park time out is also safe: the checkpoint is already durable, and the process exits on its own with the same status. -
Perform the maintenance, then start
asd.Reboot the host, upgrade the OS, replace hardware, or replace the container. Then start the node, with the
--preview index-checkpointflag still present in its startup arguments.Shell sudo systemctl start aerospikeThe node hydrates from the checkpoint, deletes it as the index goes live, and rejoins the cluster. If it halts instead of rejoining, see If the node halts on the next start. To confirm how each namespace actually started, see Verify how each namespace started.
-
Verify the node rejoined, then wait for migrations.
asadm Admin> info networkConfirm the node is back at the expected cluster size, then wait for migrations to complete. All nodes converge on the same cluster key. See Wait for migrations.
asadm Admin+> watch 10 30 asinfo -v 'cluster-stable:size=4;ignore-migrations=false'
Repeat Checkpoint a node for each remaining node, starting again at the quiesce.
The checkpoint is consumed by the restart that hydrates from it, so a node you need to restart a
second time needs a new checkpoint-save first. Restarting again without one is a cold start.
If the node halts on the next start
For an in-memory namespace without storage-backed persistence, the checkpoint or the stripes still in
shared memory can be the only copy of that node’s records. Rather than start empty over a copy that
might be good, asd halts before it reads, deletes, or changes anything. A halted node has lost
nothing, and it stays recoverable for as long as you leave it down.
Work through these in order. Each log message names the namespace and ends with the action to take.
-
Restart the node once. A transient storage condition may have cleared. Nothing is consumed by a halted start, so this costs nothing to try.
-
Fix the cause, then restart. Restore the checkpoint directory’s ownership and permissions, remove a foreign entry or symlink at the path, or set
data-sizeback to its previous value. This is the only recovery that keeps the data. -
If the cause cannot be fixed, discard and heal. Remove the namespace’s checkpoint directory, or start the node with
--cold-start. Both bring the namespace up empty and let the other nodes refill it by migration.
For the full list of conditions that halt a start, see When the node halts and waits for you.
Parked node behavior
The node refuses most info commands from the moment it accepts checkpoint-save until the process
exits. That covers the copy as well as the park, and the copy is the longer of the two. Every command
other than checkpoint-status and a re-issued checkpoint-save returns:
ERROR:22:checkpoint-save in progress - only checkpoint-status and checkpoint-save are availableWhat this means in practice:
- A cluster-wide
asadmview drops the node, because it has departed.manage checkpoint saveandmanage checkpoint statusreach it anyway: the session that issued the save keeps polling it, and a new session reaches it when seeded at the parked node. When every node a session connects to is parked,asadmstarts and warns that onlymanage checkpoint statusandmanage checkpoint savewill answer. See Checking status from a script for what not to use. checkpoint-statusanswering while everything else refuses is how you tell an intentionally parked node from a dead one. Use it as the health check for the duration.- Client transactions are unaffected when the quiesce handoff completed first, because the node owns
no partitions by then. If you skipped the quiesce, transactions that land on the node during the
save fail with a retryable
UNAVAILABLErather than hanging, and clients retry them elsewhere. - Monitoring that scrapes info commands shows the node as down for the whole save. Socket-level
probes still pass, because
asdkeeps accepting TCP connections throughout, so a TCP health check reports a healthy node while every info-level check fails.
Checking status from a script
A script can run manage checkpoint save with asadm -e. Pass --no-warn, because otherwise
the command waits for the confirmation prompt. The command exits with status 2 when the save does
not start, a namespace fails, or polling stops before the node finishes, so a stop command chained
with && runs only after a successful checkpoint:
asadm -h 10.0.3.2 --enable -e 'manage checkpoint save --no-warn with 10.0.3.2:3000' && sudo systemctl stop aerospikeA script or an orchestrator can also issue the info commands directly with asinfo:
asinfo -h 10.0.3.2 -p 3000 -v 'checkpoint-save'asinfo -h 10.0.3.2 -p 3000 -v 'checkpoint-status'checkpoint-status returns a single line, a semicolon-separated list with one record per
checkpointed namespace, each of the form
NAMESPACE:state=STATE:files_completed=N:files_total=N:is_parked=BOOL:park_ms=N. The semicolon
separates records and does not terminate the last one:
test:state=done:files_completed=11:files_total=11:is_parked=true:park_ms=1453;analytics:state=done:files_completed=6:files_total=6:is_parked=true:park_ms=1453is_parked turns true once every namespace has finished and the node has entered the park, and
park_ms then climbs on each poll. Together they distinguish a node that is about to be reaped from
one that has been parked far longer than expected.
Re-issuing checkpoint-save is safe. It reports state and never re-runs the copy, returning
checkpoint-save already in progress or checkpoint-save already complete as plain success
strings, with no ERROR: prefix, so a retrying script cannot corrupt the committed copy.
Poll every few seconds rather than continuously. On a cluster with security configured, each
asinfo call authenticates from scratch, so a tight loop spends more on logging in than on the
answer.
Next steps
- Index checkpoint. Configure the feature, size and secure the checkpoint directory, and read the full restart-outcome matrix.
- Quiesce a node. Drain client traffic before checkpointing.
- Managing strong consistency. Planned maintenance procedure for strong-consistency namespaces.
- Fast restart. Warm restart and shared-memory index persistence.