---
title: "Checkpoint a node"
description: "Use checkpoint-save and checkpoint-status to preserve a node's shared memory before a host reboot or pod replacement, so the node warm restarts instead of cold starting."
---

# Checkpoint a node

> For the complete documentation index see: [llms.txt](https://aerospike.com/docs/llms.txt)
> 
> All documentation pages available in markdown.

This page describes how to checkpoint one node before planned maintenance, so that it warm restarts with its indexes and its in-memory data intact instead of cold starting and rebuilding them.

See [Index checkpoint](https://aerospike.com/docs/database/manage/database/index-checkpoint) for a description of the index checkpoint, how to enable and configure it, and how to size and secure the checkpoint directory.

::: preview feature
The index checkpoint requires the `--preview index-checkpoint` flag in the `asd` startup arguments on every start. Without it, `asd` refuses to start when `index-checkpoint-path` is configured. Set the flag in the service unit or the configuration file so it persists through package upgrades.
:::

## How it works

The [`checkpoint-save`](https://aerospike.com/docs/database/reference/info#checkpoint-save) command triggers a graceful node shutdown while writing a durable copy of its shared memory to disk. When a node completes this process, it enters a _parked_ state, meaning the process remains alive and out of the cluster, serving status commands(such as `checkpoint-status` or a re-issued `checkpoint-save`) until manually stopped or until the park timeout elapses.

### Shutdown sequence

-   **Cluster departure**: The node leaves the cluster to isolate its state.
    
-   **State copying**: The node copies shared memory segments for each checkpointed namespace to `index-checkpoint-path`. If a node has nothing to checkpoint, it skips this step.
    
-   **Parking**: The node enters the parked state.
    

Saving a checkpoint is an integrated phase of the shutdown process. A node without data to checkpoint still accepts the command, shuts down gracefully, and parks.

The node is out of the cluster before the first byte is written, so size the maintenance window against the copy plus the park rather than the park alone. The copy scales with the size of what the node holds in shared memory. See [What the checkpoint captures](https://aerospike.com/docs/database/manage/database/index-checkpoint#what-the-checkpoint-captures) for more details.

On the next start, `asd` hydrates from the checkpoint if the node’s shared memory segments are gone, then deletes the checkpoint as the index goes live, but only the checkpoint it actually hydrated from. Hydrating protects one restart. A checkpoint the node did not consume, such as one left behind by a same-host restart that reattached its surviving shared memory, stays on disk as a fallback until the next `checkpoint-save` replaces it.

::: checkpoint-save cannot be undone
No command checkpoints a running node, nothing cancels a checkpoint once issued, and there is no `unquiesce`\-style reversal. Once you issue `checkpoint-save`, the node is going down, and the only way forward is to stop the process and start it again. Confirm you intend to take this node out of service before you run it.
:::
::: the node stops answering monitoring for the whole save
From the moment the node accepts `checkpoint-save` until the process exits, throughout both copying and parking, it responds only to `checkpoint-status` and a re-issued `checkpoint-save`. All other info commands such as `statistics`, `build`, `partitions`, `replicas`, and `peers`, returns:

```text
ERROR:22:checkpoint-save in progress - only checkpoint-status and checkpoint-save are available
```

Suppress alerting for this node, and disable or extend any probe that could restart, evict, or kill it, before you issue `checkpoint-save`. A probe-driven restart during the copy destroys the checkpoint in progress, and there is no earlier copy to fall back on, so the node comes back cold. See [Parked node behavior](#parked-node-behavior).
:::

## Prerequisites

-   Aerospike Database Enterprise Edition 8.2.0 or later. The index checkpoint is not available in Community Edition.
-   The target node has [`index-checkpoint-path`](https://aerospike.com/docs/database/reference/config#service__index-checkpoint-path) configured in its `service` stanza and the `--preview index-checkpoint` flag in its startup arguments. See [Enable the index checkpoint](https://aerospike.com/docs/database/manage/database/index-checkpoint#enable-the-index-checkpoint).
-   Enough free space on the checkpoint volume for the shared memory footprint of every namespace the node checkpoints, measured uncompressed. `asd` will start with a short volume, but the copy fails, and by then the node has already left the cluster. See [Size the checkpoint volume](https://aerospike.com/docs/database/manage/database/index-checkpoint#size-the-checkpoint-volume).
-   A user holding the `sys-admin` role, which `checkpoint-save` and `checkpoint-status` require. Aerospike enforces that role only on a cluster with a `security` stanza configured; without one, anyone who can reach the info port can take a node out of service. See [Secure the checkpoint directory](https://aerospike.com/docs/database/manage/database/index-checkpoint#secure-the-checkpoint-directory).
-   [`asinfo`](https://aerospike.com/docs/database/tools/asinfo) and [`asadm`](https://aerospike.com/docs/database/tools/asadm), available at the management host. The `manage checkpoint` commands require Tools 13.1.0 (`asadm` 5.1.0) or later. The following examples target a four-node Enterprise Edition cluster with namespace `test`, checkpointing `node2` at `10.0.3.2`.

::: restart one node at a time
`checkpoint-save` shuts the node down. Checkpoint one node at a time and let migrations finish before you move to the next. For namespaces using [`strong-consistency`](https://aerospike.com/docs/database/reference/config#namespace__strong-consistency), follow the [planned maintenance procedure for strong consistency](https://aerospike.com/docs/database/manage/cluster/consistency#planned-maintenance). Restarting more than one node without it can drop some partitions below write quorum, blocking writes to those partitions until the node returns.
:::

## Checkpoint a node

1.  Verify target namespaces.
    
    Run `checkpoint-status` on the running node before your maintenance window to confirm which namespaces will be checkpointed.
    
    Shell
    
    ```bash
    asinfo -h 10.0.3.2 -p 3000 -v 'checkpoint-status'
    ```
    
    ```text
    test:state=none:files_completed=0:files_total=0:is_parked=false:park_ms=0
    ```
    
    A namespace reporting `state=none` indicates a healthy, steady state prior to checkpointing.
    
    ::: verify missing namespaces before proceeding
    A namespace is omitted from `checkpoint-status` when it sets [`skip-checkpoint`](https://aerospike.com/docs/database/reference/config#namespace__skip-checkpoint), or when it has no volatile data left to checkpoint. Absence from the output is not success. Confirm that every namespace you expect to survive the restart is listed, because `checkpoint-save` returns `ok` at the node level even when individual namespaces are skipped.
    :::
    
2.  Confirm node checkpoint readiness.
    
    Evaluate the response to ensure the node is configured to write checkpoints.
    
    When `index-checkpoint-path` is unset, the feature is off for the node and the command returns:
    
    ```text
    ERROR:4:'index-checkpoint-path' is not configured
    ```
    
    When the path is set but no namespace is left to checkpoint, the command does not error. It returns a single node-global entry with no namespace name:
    
    ```text
    is_parked=false:park_ms=0
    ```
    
    A node in this state still accepts `checkpoint-save` to shut down gracefully and park, but it writes no checkpoint. Treat that as a warning that no index data will be saved on restart.
    
3.  Confirm the checkpoint volume has room for the copy.
    
    Shell
    
    ```bash
    df -h /opt/aerospike/checkpoint
    ```
    
4.  Verify the cluster is stable.
    
    Confirm that no migrations are in progress and that all nodes report the same cluster key. See [Verify the cluster is stable](https://aerospike.com/docs/database/manage/cluster/quiesce-node#verify-the-cluster-is-stable).
    
    asadm
    
    ```bash
    Admin> info network
    
    ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~Network Information (2026-04-13 23:57:26 UTC)~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
    
             Node|         Node ID|          IP|    Build|Migrations|~~~~~~~~~~~~~~~~~~Cluster~~~~~~~~~~~~~~~~~~|Client|  Uptime
    
                 |                |            |         |          |Size|         Key|Integrity|      Principal| Conns|
    
    node1:3000   | BB94B7AEB45DB52|10.0.3.1:3000|E-8.2.0.0|   0.000  |   4|AA5AF50552AF|True     |BB9D787F6BAF3D6|     5|00:20:34
    
    node2:3000   |*BB9D787F6BAF3D6|10.0.3.2:3000|E-8.2.0.0|   0.000  |   4|AA5AF50552AF|True     |BB9D787F6BAF3D6|     5|00:20:34
    
    node3:3000   | BB989A1BF1D8116|10.0.3.3:3000|E-8.2.0.0|   0.000  |   4|AA5AF50552AF|True     |BB9D787F6BAF3D6|     5|00:20:34
    
    node4:3000   | BB9D48DF5A70CEE|10.0.3.4:3000|E-8.2.0.0|   0.000  |   4|AA5AF50552AF|True     |BB9D787F6BAF3D6|     5|00:20:34
    
    Number of rows: 4
    ```
    
5.  Quiesce the node.
    
    Quiescing is not required to issue `checkpoint-save`, but without it clients hold references to a node that is about to depart, which causes timeouts that quiescing avoids. The node drops to the end of every succession list and hands off master ownership. See [Quiesce a node](https://aerospike.com/docs/database/manage/cluster/quiesce-node).
    
    asadm
    
    ```bash
    Admin> enable
    
    Admin+> manage quiesce with node2:3000
    
    Admin+> manage recluster
    ```
    
    Confirm the [quiesce handoff](https://aerospike.com/docs/database/manage/cluster/quiesce-node#verify-quiesce-handoff) is complete before you continue. Ops/sec on the quiesced node must drop to zero and proxy counters must stop incrementing.
    
6.  Save the checkpoint and watch it complete.
    
    `manage checkpoint save` issues the save and then polls the node until every namespace finishes, so this one command covers both. It requires Tools 13.1.0 (`asadm` 5.1.0) or later. It prompts for confirmation first, because a checkpoint cannot be undone. Type the challenge string to proceed.
    
    asadm
    
    ```bash
    Admin+> manage checkpoint save with node2:3000
    
    You are about to checkpoint node(s): node2:3000. Each leaves the cluster, saves its shared-memory segments to durable storage, then parks until it is stopped or the park timeout elapses (300s). This cannot be undone.
    
    Confirm that you want to proceed by typing a1b2c3, or cancel by typing anything else.
    ```
    
    ::: include save
    `manage checkpoint` without `save` only displays status and does not save a checkpoint.
    :::
    
    If the node is not quiesced, `asadm` warns before the prompt but does not stop you:
    
    ```text
    WARNING: Not quiesced: node2:3000 (test). Departure will trigger migrations. Recommended: 'manage quiesce with node2:3000', then 'manage recluster', and checkpoint once migrations finish.
    ```
    
    The node returns `ok`, departs the cluster, writes its checkpoint, and parks. `manage checkpoint save` keeps polling it in the same session after it leaves the cluster, and prints a table per poll:
    
    ```text
    ~~Checkpoint Save~~
    
          Node|Response
    
    node2:3000|ok
    
    Number of rows: 1
    
    ~~~~~Shared-Memory Checkpoint (2026-09-24 21:25:07 UTC)~~~~
    
    Namespace|      Node|  State|Files|Progress|Parked|    Park
    
             |          |       |     |        |      |    Time
    
    test     |node2:3000|copying|3/11 | 27.27 %|False |00:00:00
    
    Number of rows: 1
    
    ~~~~Shared-Memory Checkpoint (2026-09-24 21:25:13 UTC)~~~
    
    Namespace|      Node|State|Files|Progress|Parked|    Park
    
             |          |     |     |        |      |    Time
    
    test     |node2:3000|done |11/11| 100.0 %|True  |00:00:01
    
    Number of rows: 1
    
    Checkpoint complete on node2:3000.
    ```
    
    The park timeout defaults to 300 seconds. It bounds only the park, the window in which you stop the process, and does not bound the copy, which begins earlier. To hold the park longer, pass `--timeout`, from 1 to 3600 seconds. `--poll-interval` sets the seconds between polls, default 2. Flags go before `with`:
    
    asadm
    
    ```bash
    Admin+> manage checkpoint save --timeout 1200 --poll-interval 5 with node2:3000
    ```
    
    For the other flags, including `--no-warn` to skip the prompt in a script, see [`manage checkpoint`](https://aerospike.com/docs/database/tools/asadm/live-mode#checkpoint).
    
    ::: raise the shutdown timeout, not the park timeout, for a slow copy
    What bounds the copy is the shutdown timeout of whatever stops the node: `TimeoutSec` in the systemd unit, or `terminationGracePeriodSeconds` in Kubernetes. If either cuts the copy short, the process is killed mid-copy and the checkpoint is lost. Raise that setting, then retry with a higher [`index-checkpoint-threads`](https://aerospike.com/docs/database/reference/config#namespace__index-checkpoint-threads) or faster checkpoint storage. Raising the park timeout does not help a copy that runs long.
    :::
    
7.  Confirm every namespace reached `done`.
    
    `manage checkpoint save` returns once every namespace is `done` or `failed`, and reports `Checkpoint complete` or an `ERROR` line naming the failed namespaces. To read the status again afterwards, in the same session:
    
    asadm
    
    ```bash
    Admin+> manage checkpoint status with node2:3000
    ```
    
    ```text
    ~~~~Shared-Memory Checkpoint (2026-09-24 21:25:15 UTC)~~~
    
    Namespace|      Node|State|Files|Progress|Parked|    Park
    
             |          |     |     |        |      |    Time
    
    test     |node2:3000|done |11/11| 100.0 %|True  |00:00:03
    
    Number of rows: 1
    ```
    
    If you lose the session while the node is still parked, start a new one seeded at the parked node. A session seeded at another node does not reach it, because it has left the cluster.
    
    Shell
    
    ```bash
    asadm -h 10.0.3.2 --enable -e 'manage checkpoint status'
    ```
    
    The State column reads `none` before any checkpoint has run, `copying` while one is in progress, `done` when it is complete, and `failed` if it did not finish. Files and Progress count the shared memory segment files copied against the total to copy. Parked reads `True` once the node has entered the park, and Park Time shows how long it has been there.
    
    ::: do not reboot or replace the container on failed state
    A `failed` state means that namespace has no checkpoint at all once the copy has begun. Retention is single-copy: `checkpoint-save` deletes the previous checkpoint just before it starts writing the new one, so a save that fails during the copy leaves nothing to fall back on. A save rejected by an earlier check, before the copy starts, leaves the previous checkpoint intact. The node parks on `failed` as well as on `done`, so a node that has parked is not evidence that a checkpoint was written.
    
    The node is already out of the cluster and cannot be returned to service in place. Stop the process and start it again normally, so the node rejoins and re-replicates. Then read the log for the reason the save failed, and fix it before you retry the maintenance. Rebooting or replacing the container instead restarts the node with no checkpoint, and an in-memory namespace without storage-backed persistence comes up empty.
    :::
    
8.  Stop `asd` and check its exit status.
    
    Shell
    
    ```bash
    sudo systemctl stop aerospike
    ```
    
    `asd` exits `0` only when every namespace’s checkpoint succeeded and storage shut down cleanly, and non-zero if either failed. Read that status from the unit. The exit code of `systemctl stop` is the result of the stop job, not the exit status of `asd`, and reports `0` even for a node that exited non-zero on a failed checkpoint:
    
    Shell
    
    ```bash
    systemctl show aerospike --property=ExecMainStatus
    ```
    
    ```text
    ExecMainStatus=0
    ```
    
    In a container, let the orchestrator deliver the `SIGTERM`. Do not use `SIGKILL`, which bypasses the park. Once `checkpoint-status` reports `done` for every namespace, letting the park time out is also safe: the checkpoint is already durable, and the process exits on its own with the same status.
    
    ::: a non-zero exit means the node will not come up warm
    A failed save leaves no checkpoint. Do not replace the container or reboot the host on a non-zero exit. For an in-memory namespace without storage-backed persistence, restarting now loses every record this node holds that is not on a surviving replica.
    :::
    
9.  Perform the maintenance, then start `asd`.
    
    Reboot the host, upgrade the OS, replace hardware, or replace the container. Then start the node, with the `--preview index-checkpoint` flag still present in its startup arguments.
    
    Shell
    
    ```bash
    sudo systemctl start aerospike
    ```
    
    The node hydrates from the checkpoint, deletes it as the index goes live, and rejoins the cluster. If it halts instead of rejoining, see [If the node halts on the next start](#if-the-node-halts-on-the-next-start). To confirm how each namespace actually started, see [Verify how each namespace started](https://aerospike.com/docs/database/manage/database/index-checkpoint#verify-how-each-namespace-started).
    
10.  Verify the node rejoined, then wait for migrations.
     
     asadm
     
     ```bash
     Admin> info network
     ```
     
     Confirm the node is back at the expected cluster size, then wait for migrations to complete. All nodes converge on the same cluster key. See [Wait for migrations](https://aerospike.com/docs/database/manage/cluster/quiesce-node#wait-for-migrations).
     
     asadm
     
     ```bash
     Admin+> watch 10 30 asinfo -v 'cluster-stable:size=4;ignore-migrations=false'
     ```
     

Repeat [Checkpoint a node](#checkpoint-a-node) for each remaining node, starting again at the quiesce.

The checkpoint is consumed by the restart that hydrates from it, so a node you need to restart a second time needs a new `checkpoint-save` first. Restarting again without one is a cold start.

## If the node halts on the next start

For an in-memory namespace without storage-backed persistence, the checkpoint or the stripes still in shared memory can be the only copy of that node’s records. Rather than start empty over a copy that might be good, `asd` halts before it reads, deletes, or changes anything. A halted node has lost nothing, and it stays recoverable for as long as you leave it down.

Work through these in order. Each log message names the namespace and ends with the action to take.

1.  Restart the node once. A transient storage condition may have cleared. Nothing is consumed by a halted start, so this costs nothing to try.
    
2.  Fix the cause, then restart. Restore the checkpoint directory’s ownership and permissions, remove a foreign entry or symlink at the path, or set [`data-size`](https://aerospike.com/docs/database/reference/config#namespace__data-size) back to its previous value. This is the only recovery that keeps the data.
    
3.  If the cause cannot be fixed, discard and heal. Remove the namespace’s checkpoint directory, or start the node with `--cold-start`. Both bring the namespace up empty and let the other nodes refill it by migration.
    

::: try the fix before you discard
Discarding is not a repair. The namespace comes up empty and recovers only because other nodes still hold the records, so it depends on the rest of the cluster being healthy and current. At `replication-factor 1`, or where the surviving stripes are the only copy, discarding loses those records permanently.
:::

For the full list of conditions that halt a start, see [When the node halts and waits for you](https://aerospike.com/docs/database/manage/database/index-checkpoint#when-the-node-halts-and-waits-for-you).

## Parked node behavior

The node refuses most info commands from the moment it accepts `checkpoint-save` until the process exits. That covers the copy as well as the park, and the copy is the longer of the two. Every command other than `checkpoint-status` and a re-issued `checkpoint-save` returns:

```text
ERROR:22:checkpoint-save in progress - only checkpoint-status and checkpoint-save are available
```

What this means in practice:

-   A cluster-wide `asadm` view drops the node, because it has departed. `manage checkpoint save` and `manage checkpoint status` reach it anyway: the session that issued the save keeps polling it, and a new session reaches it when seeded at the parked node. When every node a session connects to is parked, `asadm` starts and warns that only `manage checkpoint status` and `manage checkpoint save` will answer. See [Checking status from a script](#checking-status-from-a-script) for what not to use.
-   `checkpoint-status` answering while everything else refuses is how you tell an intentionally parked node from a dead one. Use it as the health check for the duration.
-   Client transactions are unaffected when the quiesce handoff completed first, because the node owns no partitions by then. If you skipped the quiesce, transactions that land on the node during the save fail with a retryable `UNAVAILABLE` rather than hanging, and clients retry them elsewhere.
-   Monitoring that scrapes info commands shows the node as down for the whole save. Socket-level probes still pass, because `asd` keeps accepting TCP connections throughout, so a TCP health check reports a healthy node while every info-level check fails.

::: kubernetes
An exec or HTTP probe that issues an info command fails for the whole save. Suspend it for the duration or scope it to `checkpoint-status`, which answers throughout, and confirm that Pod Disruption Budgets tolerate a node reporting as unresponsive. A probe-driven restart during the copy destroys the checkpoint. Under the Aerospike Kubernetes Operator (AKO), also confirm that `terminationGracePeriodSeconds` covers the copy plus the window in which you confirm the checkpoint and send `SIGTERM`.
:::

### Checking status from a script

A script can run `manage checkpoint save` with `asadm -e`. Pass `--no-warn`, because otherwise the command waits for the confirmation prompt. The command exits with status `2` when the save does not start, a namespace fails, or polling stops before the node finishes, so a stop command chained with `&&` runs only after a successful checkpoint:

Shell

```bash
asadm -h 10.0.3.2 --enable -e 'manage checkpoint save --no-warn with 10.0.3.2:3000' && sudo systemctl stop aerospike
```

A script or an orchestrator can also issue the info commands directly with `asinfo`:

Shell

```bash
asinfo -h 10.0.3.2 -p 3000 -v 'checkpoint-save'

asinfo -h 10.0.3.2 -p 3000 -v 'checkpoint-status'
```

`checkpoint-status` returns a single line, a semicolon-separated list with one record per checkpointed namespace, each of the form _`NAMESPACE:state=STATE:files_completed=N:files_total=N:is_parked=BOOL:park_ms=N`_. The semicolon separates records and does not terminate the last one:

```text
test:state=done:files_completed=11:files_total=11:is_parked=true:park_ms=1453;analytics:state=done:files_completed=6:files_total=6:is_parked=true:park_ms=1453
```

`is_parked` turns `true` once every namespace has finished and the node has entered the park, and `park_ms` then climbs on each poll. Together they distinguish a node that is about to be reaped from one that has been parked far longer than expected.

Re-issuing `checkpoint-save` is safe. It reports state and never re-runs the copy, returning `checkpoint-save already in progress` or `checkpoint-save already complete` as plain success strings, with no `ERROR:` prefix, so a retrying script cannot corrupt the committed copy.

Poll every few seconds rather than continuously. On a cluster with security configured, each `asinfo` call authenticates from scratch, so a tight loop spends more on logging in than on the answer.

::: do not poll with asinfo run inside asadm
Use `manage checkpoint status`, or `asinfo` from the shell. Do not poll a checkpointing node by running the info command through `asadm`:

Shell

```bash
# Does not work for polling a checkpointing node

asadm -h 10.0.3.2 --enable -e 'asinfo -v "checkpoint-status"'
```

That form re-checks the cluster before each command, and a node that has departed for a checkpoint fails the check and is dropped. It can return a poll or two before that happens, so a single successful result is not evidence the next one will work. `manage checkpoint status` recognizes a parked node and reaches it directly, so it is unaffected.
:::

## Next steps

-   [Index checkpoint](https://aerospike.com/docs/database/manage/database/index-checkpoint). Configure the feature, size and secure the checkpoint directory, and read the full restart-outcome matrix.
-   [Quiesce a node](https://aerospike.com/docs/database/manage/cluster/quiesce-node). Drain client traffic before checkpointing.
-   [Managing strong consistency](https://aerospike.com/docs/database/manage/cluster/consistency#planned-maintenance). Planned maintenance procedure for strong-consistency namespaces.
-   [Fast restart](https://aerospike.com/docs/database/manage/database/fast-start). Warm restart and shared-memory index persistence.