Skip to content

Checkpoint a node

For the complete documentation index see: llms.txt

All documentation pages available in markdown.

This page describes how to checkpoint one node before planned maintenance, so that it warm restarts with its indexes and its in-memory data intact instead of cold starting and rebuilding them.

See Index checkpoint for a description of the index checkpoint, how to enable and configure it, and how to size and secure the checkpoint directory.

How it works

The checkpoint-save command triggers a graceful node shutdown while writing a durable copy of its shared memory to disk. When a node completes this process, it enters a parked state, meaning the process remains alive and out of the cluster, serving status commands(such as checkpoint-status or a re-issued checkpoint-save) until manually stopped or until the park timeout elapses.

Shutdown sequence

  • Cluster departure: The node leaves the cluster to isolate its state.

  • State copying: The node copies shared memory segments for each checkpointed namespace to index-checkpoint-path. If a node has nothing to checkpoint, it skips this step.

  • Parking: The node enters the parked state.

Saving a checkpoint is an integrated phase of the shutdown process. A node without data to checkpoint still accepts the command, shuts down gracefully, and parks.

The node is out of the cluster before the first byte is written, so size the maintenance window against the copy plus the park rather than the park alone. The copy scales with the size of what the node holds in shared memory. See What the checkpoint captures for more details.

On the next start, asd hydrates from the checkpoint if the node’s shared memory segments are gone, then deletes the checkpoint as the index goes live, but only the checkpoint it actually hydrated from. Hydrating protects one restart. A checkpoint the node did not consume, such as one left behind by a same-host restart that reattached its surviving shared memory, stays on disk as a fallback until the next checkpoint-save replaces it.

Prerequisites

  • Aerospike Database Enterprise Edition 8.2.0 or later. The index checkpoint is not available in Community Edition.
  • The target node has index-checkpoint-path configured in its service stanza and the --preview index-checkpoint flag in its startup arguments. See Enable the index checkpoint.
  • Enough free space on the checkpoint volume for the shared memory footprint of every namespace the node checkpoints, measured uncompressed. asd will start with a short volume, but the copy fails, and by then the node has already left the cluster. See Size the checkpoint volume.
  • A user holding the sys-admin role, which checkpoint-save and checkpoint-status require. Aerospike enforces that role only on a cluster with a security stanza configured; without one, anyone who can reach the info port can take a node out of service. See Secure the checkpoint directory.
  • asinfo and asadm, available at the management host. The manage checkpoint commands require Tools 13.1.0 (asadm 5.1.0) or later. The following examples target a four-node Enterprise Edition cluster with namespace test, checkpointing node2 at 10.0.3.2.

Checkpoint a node

  1. Verify target namespaces.

    Run checkpoint-status on the running node before your maintenance window to confirm which namespaces will be checkpointed.

    Shell
    asinfo -h 10.0.3.2 -p 3000 -v 'checkpoint-status'
    test:state=none:files_completed=0:files_total=0:is_parked=false:park_ms=0

    A namespace reporting state=none indicates a healthy, steady state prior to checkpointing.

  2. Confirm node checkpoint readiness.

    Evaluate the response to ensure the node is configured to write checkpoints.

    When index-checkpoint-path is unset, the feature is off for the node and the command returns:

    ERROR:4:'index-checkpoint-path' is not configured

    When the path is set but no namespace is left to checkpoint, the command does not error. It returns a single node-global entry with no namespace name:

    is_parked=false:park_ms=0

    A node in this state still accepts checkpoint-save to shut down gracefully and park, but it writes no checkpoint. Treat that as a warning that no index data will be saved on restart.

  3. Confirm the checkpoint volume has room for the copy.

    Shell
    df -h /opt/aerospike/checkpoint
  4. Verify the cluster is stable.

    Confirm that no migrations are in progress and that all nodes report the same cluster key. See Verify the cluster is stable.

    asadm
    Admin> info network
    ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~Network Information (2026-04-13 23:57:26 UTC)~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
    Node| Node ID| IP| Build|Migrations|~~~~~~~~~~~~~~~~~~Cluster~~~~~~~~~~~~~~~~~~|Client| Uptime
    | | | | |Size| Key|Integrity| Principal| Conns|
    node1:3000 | BB94B7AEB45DB52|10.0.3.1:3000|E-8.2.0.0| 0.000 | 4|AA5AF50552AF|True |BB9D787F6BAF3D6| 5|00:20:34
    node2:3000 |*BB9D787F6BAF3D6|10.0.3.2:3000|E-8.2.0.0| 0.000 | 4|AA5AF50552AF|True |BB9D787F6BAF3D6| 5|00:20:34
    node3:3000 | BB989A1BF1D8116|10.0.3.3:3000|E-8.2.0.0| 0.000 | 4|AA5AF50552AF|True |BB9D787F6BAF3D6| 5|00:20:34
    node4:3000 | BB9D48DF5A70CEE|10.0.3.4:3000|E-8.2.0.0| 0.000 | 4|AA5AF50552AF|True |BB9D787F6BAF3D6| 5|00:20:34
    Number of rows: 4
  5. Quiesce the node.

    Quiescing is not required to issue checkpoint-save, but without it clients hold references to a node that is about to depart, which causes timeouts that quiescing avoids. The node drops to the end of every succession list and hands off master ownership. See Quiesce a node.

    asadm
    Admin> enable
    Admin+> manage quiesce with node2:3000
    Admin+> manage recluster

    Confirm the quiesce handoff is complete before you continue. Ops/sec on the quiesced node must drop to zero and proxy counters must stop incrementing.

  6. Save the checkpoint and watch it complete.

    manage checkpoint save issues the save and then polls the node until every namespace finishes, so this one command covers both. It requires Tools 13.1.0 (asadm 5.1.0) or later. It prompts for confirmation first, because a checkpoint cannot be undone. Type the challenge string to proceed.

    asadm
    Admin+> manage checkpoint save with node2:3000
    You are about to checkpoint node(s): node2:3000. Each leaves the cluster, saves its shared-memory segments to durable storage, then parks until it is stopped or the park timeout elapses (300s). This cannot be undone.
    Confirm that you want to proceed by typing a1b2c3, or cancel by typing anything else.

    If the node is not quiesced, asadm warns before the prompt but does not stop you:

    WARNING: Not quiesced: node2:3000 (test). Departure will trigger migrations. Recommended: 'manage quiesce with node2:3000', then 'manage recluster', and checkpoint once migrations finish.

    The node returns ok, departs the cluster, writes its checkpoint, and parks. manage checkpoint save keeps polling it in the same session after it leaves the cluster, and prints a table per poll:

    ~~Checkpoint Save~~
    Node|Response
    node2:3000|ok
    Number of rows: 1
    ~~~~~Shared-Memory Checkpoint (2026-09-24 21:25:07 UTC)~~~~
    Namespace| Node| State|Files|Progress|Parked| Park
    | | | | | | Time
    test |node2:3000|copying|3/11 | 27.27 %|False |00:00:00
    Number of rows: 1
    ~~~~Shared-Memory Checkpoint (2026-09-24 21:25:13 UTC)~~~
    Namespace| Node|State|Files|Progress|Parked| Park
    | | | | | | Time
    test |node2:3000|done |11/11| 100.0 %|True |00:00:01
    Number of rows: 1
    Checkpoint complete on node2:3000.

    The park timeout defaults to 300 seconds. It bounds only the park, the window in which you stop the process, and does not bound the copy, which begins earlier. To hold the park longer, pass --timeout, from 1 to 3600 seconds. --poll-interval sets the seconds between polls, default 2. Flags go before with:

    asadm
    Admin+> manage checkpoint save --timeout 1200 --poll-interval 5 with node2:3000

    For the other flags, including --no-warn to skip the prompt in a script, see manage checkpoint.

  7. Confirm every namespace reached done.

    manage checkpoint save returns once every namespace is done or failed, and reports Checkpoint complete or an ERROR line naming the failed namespaces. To read the status again afterwards, in the same session:

    asadm
    Admin+> manage checkpoint status with node2:3000
    ~~~~Shared-Memory Checkpoint (2026-09-24 21:25:15 UTC)~~~
    Namespace| Node|State|Files|Progress|Parked| Park
    | | | | | | Time
    test |node2:3000|done |11/11| 100.0 %|True |00:00:03
    Number of rows: 1

    If you lose the session while the node is still parked, start a new one seeded at the parked node. A session seeded at another node does not reach it, because it has left the cluster.

    Shell
    asadm -h 10.0.3.2 --enable -e 'manage checkpoint status'

    The State column reads none before any checkpoint has run, copying while one is in progress, done when it is complete, and failed if it did not finish. Files and Progress count the shared memory segment files copied against the total to copy. Parked reads True once the node has entered the park, and Park Time shows how long it has been there.

  8. Stop asd and check its exit status.

    Shell
    sudo systemctl stop aerospike

    asd exits 0 only when every namespace’s checkpoint succeeded and storage shut down cleanly, and non-zero if either failed. Read that status from the unit. The exit code of systemctl stop is the result of the stop job, not the exit status of asd, and reports 0 even for a node that exited non-zero on a failed checkpoint:

    Shell
    systemctl show aerospike --property=ExecMainStatus
    ExecMainStatus=0

    In a container, let the orchestrator deliver the SIGTERM. Do not use SIGKILL, which bypasses the park. Once checkpoint-status reports done for every namespace, letting the park time out is also safe: the checkpoint is already durable, and the process exits on its own with the same status.

  9. Perform the maintenance, then start asd.

    Reboot the host, upgrade the OS, replace hardware, or replace the container. Then start the node, with the --preview index-checkpoint flag still present in its startup arguments.

    Shell
    sudo systemctl start aerospike

    The node hydrates from the checkpoint, deletes it as the index goes live, and rejoins the cluster. If it halts instead of rejoining, see If the node halts on the next start. To confirm how each namespace actually started, see Verify how each namespace started.

  10. Verify the node rejoined, then wait for migrations.

    asadm
    Admin> info network

    Confirm the node is back at the expected cluster size, then wait for migrations to complete. All nodes converge on the same cluster key. See Wait for migrations.

    asadm
    Admin+> watch 10 30 asinfo -v 'cluster-stable:size=4;ignore-migrations=false'

Repeat Checkpoint a node for each remaining node, starting again at the quiesce.

The checkpoint is consumed by the restart that hydrates from it, so a node you need to restart a second time needs a new checkpoint-save first. Restarting again without one is a cold start.

If the node halts on the next start

For an in-memory namespace without storage-backed persistence, the checkpoint or the stripes still in shared memory can be the only copy of that node’s records. Rather than start empty over a copy that might be good, asd halts before it reads, deletes, or changes anything. A halted node has lost nothing, and it stays recoverable for as long as you leave it down.

Work through these in order. Each log message names the namespace and ends with the action to take.

  1. Restart the node once. A transient storage condition may have cleared. Nothing is consumed by a halted start, so this costs nothing to try.

  2. Fix the cause, then restart. Restore the checkpoint directory’s ownership and permissions, remove a foreign entry or symlink at the path, or set data-size back to its previous value. This is the only recovery that keeps the data.

  3. If the cause cannot be fixed, discard and heal. Remove the namespace’s checkpoint directory, or start the node with --cold-start. Both bring the namespace up empty and let the other nodes refill it by migration.

For the full list of conditions that halt a start, see When the node halts and waits for you.

Parked node behavior

The node refuses most info commands from the moment it accepts checkpoint-save until the process exits. That covers the copy as well as the park, and the copy is the longer of the two. Every command other than checkpoint-status and a re-issued checkpoint-save returns:

ERROR:22:checkpoint-save in progress - only checkpoint-status and checkpoint-save are available

What this means in practice:

  • A cluster-wide asadm view drops the node, because it has departed. manage checkpoint save and manage checkpoint status reach it anyway: the session that issued the save keeps polling it, and a new session reaches it when seeded at the parked node. When every node a session connects to is parked, asadm starts and warns that only manage checkpoint status and manage checkpoint save will answer. See Checking status from a script for what not to use.
  • checkpoint-status answering while everything else refuses is how you tell an intentionally parked node from a dead one. Use it as the health check for the duration.
  • Client transactions are unaffected when the quiesce handoff completed first, because the node owns no partitions by then. If you skipped the quiesce, transactions that land on the node during the save fail with a retryable UNAVAILABLE rather than hanging, and clients retry them elsewhere.
  • Monitoring that scrapes info commands shows the node as down for the whole save. Socket-level probes still pass, because asd keeps accepting TCP connections throughout, so a TCP health check reports a healthy node while every info-level check fails.

Checking status from a script

A script can run manage checkpoint save with asadm -e. Pass --no-warn, because otherwise the command waits for the confirmation prompt. The command exits with status 2 when the save does not start, a namespace fails, or polling stops before the node finishes, so a stop command chained with && runs only after a successful checkpoint:

Shell
asadm -h 10.0.3.2 --enable -e 'manage checkpoint save --no-warn with 10.0.3.2:3000' && sudo systemctl stop aerospike

A script or an orchestrator can also issue the info commands directly with asinfo:

Shell
asinfo -h 10.0.3.2 -p 3000 -v 'checkpoint-save'
asinfo -h 10.0.3.2 -p 3000 -v 'checkpoint-status'

checkpoint-status returns a single line, a semicolon-separated list with one record per checkpointed namespace, each of the form NAMESPACE:state=STATE:files_completed=N:files_total=N:is_parked=BOOL:park_ms=N. The semicolon separates records and does not terminate the last one:

test:state=done:files_completed=11:files_total=11:is_parked=true:park_ms=1453;analytics:state=done:files_completed=6:files_total=6:is_parked=true:park_ms=1453

is_parked turns true once every namespace has finished and the node has entered the park, and park_ms then climbs on each poll. Together they distinguish a node that is about to be reaped from one that has been parked far longer than expected.

Re-issuing checkpoint-save is safe. It reports state and never re-runs the copy, returning checkpoint-save already in progress or checkpoint-save already complete as plain success strings, with no ERROR: prefix, so a retrying script cannot corrupt the committed copy.

Poll every few seconds rather than continuously. On a cluster with security configured, each asinfo call authenticates from scratch, so a tight loop spends more on logging in than on the answer.

Next steps