Skip to content

System metadata readiness on node join

For the complete documentation index see: llms.txt

All documentation pages available in markdown.

This page describes how a new or returning node waits for initial system metadata (SMD) sync before it answers on the service or admin ports, and before it receives incoming record migrations. The node still joins the cluster without delay.

The wait keeps client work from running against incomplete metadata, such as missing secondary index definitions, incomplete security users and roles, or missing user-defined functions (UDFs).

Applies to

  • Product: Aerospike Database 8.2.0 and later.
  • Audience: operators, site reliability engineers (SREs), and developers whose clients connect during node join, restart, or rolling upgrade.
  • Configuration: none. There is nothing to set, tune, or disable.

The wait is identical whether or not security is configured.

What a joining node does

Cluster formation, exchange, and partition balance are not delayed. Other nodes list the joining node in the cluster while the wait is in progress.

Until initial SMD sync completes, the joining node:

  • Completes a TCP handshake on the service and admin ports, then sends no protocol response.
  • Defers incoming record migrations. Source nodes retry until the joining node is ready.
  • Continues outgoing migrations, heartbeat, and fabric traffic. Those paths are not gated on SMD.

After the wait ends, a short, pre-existing window can remain in which data commands return a retryable unavailable error until initial partition balance resolves. That window is unrelated to SMD.

Copying SMD files onto a new node remains the recommended practice. See Initialize new cluster nodes with SMD. This wait does not replace that practice.

What clients and tools see

A connection attempt during the wait succeeds at the TCP layer and then receives no response. The client fails with a timeout. Existing client retry policy handles that failure. No client library change is required.

asadm, asinfo, and other tools that use the service or admin port against the joining node also time out until SMD sync completes. Use another node in the cluster, or the joining node’s log file, while the wait is in progress.

Rolling upgrade

The wait is active only when every node in the cluster is Database 8.2.0 or later. While any earlier node is present, a joining node does not wait for SMD sync.

To see whether the cluster has crossed that threshold, read cluster_min_compatibility_id from an upgraded node that is already answering:

asadm
Admin> show statistics like cluster_min_compatibility_id

The value is 16 or higher when every node is Database 8.2.0 or later.

Verify the wait completes

This behavior does not add statistics or ticker fields. The joining node does not answer info commands during the wait, so you verify an in-progress wait from that node’s logs and from its peers. The service ready log line appears before the wait ends. Use the initial SMD sync done line as the signal that the node is answering connections.

After the node answers, smd-info reports initial_sync_done and mixed_cluster in the global smd: section. Those fields are unreachable during the wait. initial_sync_done=true means the node has cleared its startup SMD check; mixed_cluster shows the cluster’s current version mix. They do not record whether the node waited at startup. Use the log lines for that.

asinfo
asinfo -v 'smd-info'

The two fields appear before the per-module records. See Configure SMD wire compression for an example output line.

FieldMeaning
mixed_clustertrue while any node runs a build earlier than Database 8.2.0. The readiness wait is not active on this node.
initial_sync_donetrue after the node has cleared its startup SMD check.

Log lines on the joining node

All four lines use the smd logging context.

LevelMessageWhen
INFOwaiting for initial SMD syncOnce, when the node begins waiting
WARNINGstill waiting for initial SMD sync - elapsed ELAPSED_US usAbout every 10 seconds while the wait continues
INFOinitial SMD sync done for cluster key CLUSTER_KEYOnce, when the sync completes
INFOinitial SMD sync done - elapsed ELAPSED_US usOnce, when the wait ends

Elapsed time is in microseconds.

If initial SMD sync has already finished before the node starts waiting, neither the waiting for initial SMD sync line nor the initial SMD sync done - elapsed line appears. Only initial SMD sync done for cluster key is logged.

The repeating WARNING is expected. A few occurrences during a normal node start are not a fault.

The line exists because the node does not answer info commands during the wait. Logs are the only way to distinguish a node that is still syncing from one that has stopped making progress.

If the WARNING continues without a matching initial SMD sync done line, SMD sync is not completing. See Troubleshoot a stuck wait.

Shell
grep 'initial SMD sync' /var/log/aerospike/aerospike.log

Expected output for a wait that completes:

INFO (smd): waiting for initial SMD sync
WARNING (smd): still waiting for initial SMD sync - elapsed 10001234 us
INFO (smd): initial SMD sync done for cluster key 7a54b93fdb52
INFO (smd): initial SMD sync done - elapsed 10004567 us

Signals on other nodes

While the joining node defers incoming migrations:

asadm
Admin> show statistics namespace like migrate_tx_partitions_remaining
Admin> asinfo -v 'cluster-stable:size=CLUSTER_SIZE;ignore-migrations=false'

Replace CLUSTER_SIZE with the expected number of nodes, including the joining node.

Troubleshoot a stuck wait

If the WARNING repeats without a matching initial SMD sync done line appearing, check the following:

  • Verify fabric and mesh network reachability between the joining node and its peers. SMD sync travels over the fabric link.
  • Look for WARNING-level lines in the smd log context that indicate a sync failure, other than the repeating still waiting line. CRITICAL lines indicate a crash.
  • Restart the node as a last resort to retry the sync.