System metadata readiness on node join
For the complete documentation index see: llms.txt
All documentation pages available in markdown.
This page describes how a new or returning node waits for initial system metadata (SMD) sync before it answers on the service or admin ports, and before it receives incoming record migrations. The node still joins the cluster without delay.
The wait keeps client work from running against incomplete metadata, such as missing secondary index definitions, incomplete security users and roles, or missing user-defined functions (UDFs).
Applies to
- Product: Aerospike Database 8.2.0 and later.
- Audience: operators, site reliability engineers (SREs), and developers whose clients connect during node join, restart, or rolling upgrade.
- Configuration: none. There is nothing to set, tune, or disable.
The wait is identical whether or not security is configured.
What a joining node does
Cluster formation, exchange, and partition balance are not delayed. Other nodes list the joining node in the cluster while the wait is in progress.
Until initial SMD sync completes, the joining node:
- Completes a TCP handshake on the service and admin ports, then sends no protocol response.
- Defers incoming record migrations. Source nodes retry until the joining node is ready.
- Continues outgoing migrations, heartbeat, and fabric traffic. Those paths are not gated on SMD.
After the wait ends, a short, pre-existing window can remain in which data commands return a retryable unavailable error until initial partition balance resolves. That window is unrelated to SMD.
Copying SMD files onto a new node remains the recommended practice. See Initialize new cluster nodes with SMD. This wait does not replace that practice.
What clients and tools see
A connection attempt during the wait succeeds at the TCP layer and then receives no response. The client fails with a timeout. Existing client retry policy handles that failure. No client library change is required.
asadm, asinfo, and other tools that use the service or admin port against the joining node also time out until SMD sync completes. Use another node in the cluster, or the joining node’s log file, while the wait is in progress.
Rolling upgrade
The wait is active only when every node in the cluster is Database 8.2.0 or later. While any earlier node is present, a joining node does not wait for SMD sync.
To see whether the cluster has crossed that threshold, read cluster_min_compatibility_id from an upgraded node that is already answering:
Admin> show statistics like cluster_min_compatibility_idThe value is 16 or higher when every node is Database 8.2.0 or later.
Verify the wait completes
This behavior does not add statistics or ticker fields. The joining node does not answer info commands during the wait, so you verify an in-progress wait from that node’s logs and from its peers. The service ready log line appears before the wait ends. Use the initial SMD sync done line as the signal that the node is answering connections.
After the node answers, smd-info reports initial_sync_done and mixed_cluster in the global smd: section. Those fields are unreachable during the wait. initial_sync_done=true means the node has cleared its startup SMD check; mixed_cluster shows the cluster’s current version mix. They do not record whether the node waited at startup. Use the log lines for that.
asinfo -v 'smd-info'The two fields appear before the per-module records. See Configure SMD wire compression for an example output line.
| Field | Meaning |
|---|---|
mixed_cluster | true while any node runs a build earlier than Database 8.2.0. The readiness wait is not active on this node. |
initial_sync_done | true after the node has cleared its startup SMD check. |
Log lines on the joining node
All four lines use the smd logging context.
| Level | Message | When |
|---|---|---|
| INFO | waiting for initial SMD sync | Once, when the node begins waiting |
| WARNING | still waiting for initial SMD sync - elapsed ELAPSED_US us | About every 10 seconds while the wait continues |
| INFO | initial SMD sync done for cluster key CLUSTER_KEY | Once, when the sync completes |
| INFO | initial SMD sync done - elapsed ELAPSED_US us | Once, when the wait ends |
Elapsed time is in microseconds.
If initial SMD sync has already finished before the node starts waiting, neither the waiting for initial SMD sync line nor the initial SMD sync done - elapsed line appears. Only initial SMD sync done for cluster key is logged.
The repeating WARNING is expected. A few occurrences during a normal node start are not a fault.
The line exists because the node does not answer info commands during the wait. Logs are the only way to distinguish a node that is still syncing from one that has stopped making progress.
If the WARNING continues without a matching initial SMD sync done line, SMD sync is not completing. See Troubleshoot a stuck wait.
grep 'initial SMD sync' /var/log/aerospike/aerospike.logExpected output for a wait that completes:
INFO (smd): waiting for initial SMD syncWARNING (smd): still waiting for initial SMD sync - elapsed 10001234 usINFO (smd): initial SMD sync done for cluster key 7a54b93fdb52INFO (smd): initial SMD sync done - elapsed 10004567 usSignals on other nodes
While the joining node defers incoming migrations:
migrate_tx_partitions_remainingstays above zero on nodes that are sending data to the joining node.cluster-stablereports an unstable cluster while any node still has migrations outstanding to the joining node.
Admin> show statistics namespace like migrate_tx_partitions_remainingAdmin> asinfo -v 'cluster-stable:size=CLUSTER_SIZE;ignore-migrations=false'Replace CLUSTER_SIZE with the expected number of nodes, including the joining node.
Troubleshoot a stuck wait
If the WARNING repeats without a matching initial SMD sync done line appearing, check the following:
- Verify fabric and mesh network reachability between the joining node and its peers. SMD sync travels over the fabric link.
- Look for WARNING-level lines in the
smdlog context that indicate a sync failure, other than the repeatingstill waitingline. CRITICAL lines indicate a crash. - Restart the node as a last resort to retry the sync.