Skip to content

Key metrics to monitor

For the complete documentation index see: llms.txt

All documentation pages available in markdown.

Aerospike recommends that you monitor your system with the metrics listed on this page. For a complete list of metrics, see the Metric reference.

Feature-specific counters are covered on the page for the feature. For intra-cluster bandwidth, see Configure wire compression and Configure SMD wire compression.

Operating system and server health

Finding total namespace memory

In Aerospike Database 7.0.0 and later, calculate memory use separately for each namespace on each node. When you created the namespace, you allocated storage for it and set a stop-writes threshold with stop-writes-sys-memory-pct that tells Aerospike when system memory is full enough to stop writes.

Run info namespace or its usage variant in Aerospike Admin (asadm) to identify storage placement and current use. The output tables include Primary Index, Secondary Index, and Storage Engine columns with type and used bytes. Use the following table to choose which metrics to sum for each component. This sum is the used record-data and index bytes for RAM-backed components. It is not the namespace’s allocated or resident footprint, which also includes unused index stages, unused in-memory data capacity, and write caches.

ComponentInclude whenMetric
Set indexesAlwaysset_index_used_bytes
Record datastorage-engine is memory.data_used_bytes
Primary indexindex-type is not flash or pmem.index_used_bytes
Secondary indexessindex-type is not flash or pmem.sindex_used_bytes

When record data is stored on SSD or persistent memory, data_used_bytes reports storage use. For example, in a hybrid memory architecture namespace with record data on SSD and indexes in shared memory, sum index_used_bytes, set_index_used_bytes, and sindex_used_bytes.

Index memory monitoring

Monitor current index use and remaining index capacity according to where each index is stored:

Index storageCurrent useCapacity
RAM and shared-memory indexes (primary and set)Use index_used_bytes and set_index_used_bytes.Use indexes_memory_used_pct only when indexes-memory-budget is greater than 0.
RAM and shared-memory secondary indexesUse sindex_used_bytes when sindex-type is not flash or pmem.Use indexes_memory_used_pct only when indexes-memory-budget is greater than 0.
Primary index on flash or persistent memoryUse index_used_bytes.Use index_mounts_used_pct.
Secondary indexes on flash or persistent memoryUse sindex_used_bytes.Use sindex_mounts_used_pct.

Apply the rows independently because primary and secondary indexes can use different storage types. Set indexes are always stored in memory.

Memory budgets and stage allocation

indexes-memory-budget is an optional stop-writes threshold for the combined used bytes of RAM and shared-memory primary, secondary, and set indexes. The default value of 0 disables the index-specific cap. When the budget is 0, track the applicable *_used_bytes metrics and monitor overall memory pressure with stop-writes-sys-memory-pct. In containers, enable cgroup-mem-tracking so that the threshold reflects the control group (cgroup) limit.

Indexes allocate space in stages as they grow, as configured by index-stage-size and sindex-stage-size. The *_used_bytes metrics report memory holding live index elements. Allocated stage footprint is larger. For a primary index on flash, index_flash_alloc_bytes reports allocated space. Database 8.2.0 adds the following metrics for allocated stages and the untouched space in the newest stage:

Eviction and stop-writes act on used bytes, so allocated bytes can sit above indexes-memory-budget while those thresholds remain unbreached. See Monitor index memory allocation for how to interpret allocated, used, and tail bytes and reconcile them with process and control group memory.

Namespace capacity and write protection

MetricDescription
stop_writestrue when the namespace is not allowing client-originated writes.
hwm_breachedtrue when a high-water disk or memory percentage threshold is breached for the namespace.
data_avail_pctMinimum available write blocks as a percentage across all storage files in a namespace (Database 7.0.0+).
clock_skew_stop_writestrue when clock skew is outside tolerance and the namespace stops accepting writes or NSUP is blocked.
index_shmem_tail_bytesUntouched space in the newest shared-memory primary-index stage in Database 8.2.0 and later. A value of 0 means that the next insert might claim another index-stage-size of memory.

Strong consistency

MetricDescription
dead_partitionsNumber of partitions unavailable when all roster nodes are present.
unavailable_partitionsNumber of partitions unavailable when roster nodes are missing.

Cluster and connections

MetricDescription
cluster_sizeSize of the cluster. Compare across all nodes.
client_connectionsNumber of active client connections to this node.
client_connections_openedNumber of client connections created since the node started.
fabric_connections_openedNumber of fabric connections created since the node started.
heartbeat_connections_openedNumber of heartbeat connections created since the node started.

System memory

MetricDescription
system_free_mem_kbytes / system_free_mem_pctFree memory on the node. Control-group-aware when cgroup-mem-tracking is enabled.
host_free_mem_kbytes / host_free_mem_pctHost-level free memory (Database 8.1.2+), regardless of cgroup-mem-tracking.
process_rss_bytesResident memory of the asd process (Database 8.2.0 and later). In a container, compare cgroup_memory_used_bytes with cgroup_memory_limit_bytes instead; the control group’s total can exceed asd’s RSS.

Wire compression

MetricDescription
repl_wire_compression_delta_reject_applyA sustained rise points at the replica’s stored base records, not at compression settings.
repl_wire_compression_delta_reject_otherInvestigate when the other three delta_reject_* counters stay near zero.

Metrics to track

Track these metrics over time. Compare error counters to their corresponding success or complete counters.

Client and query errors

MetricWhat it measures
client_read_errorFailed client read commands, such as invalid set name, unavailable partition, device I/O error, or key busy.
client_write_errorFailed client write commands, including generation conflicts, key busy, record too big, and XDR-forbidden errors.
client_delete_errorFailed client delete commands.
client_udf_errorFailed client-originated UDF commands (excludes timeouts).
batch_index_errorBatch index requests that completed with an error, often after client timeout or send delays.
pi_query_short_basic_errorFailed basic short primary index queries.
pi_query_long_basic_errorFailed basic long primary index queries.
pi_query_aggr_errorFailed primary index query aggregations.
pi_query_ops_bg_errorFailed ops background primary index queries.
pi_query_udf_bg_errorFailed UDF background primary index queries.

Storage and node runtime

MetricWhat it measures
storage-engine.device[ix].defrag_qNumber of write blocks queued for defragmentation on a device.
storage-engine.file[ix].write_qNumber of write blocks queued for write to a storage file.
index_flash_alloc_pctPercentage of mount space allocated for the primary index when index-type is flash (Enterprise Edition).
heap_efficiency_pctRatio of allocated to active jemalloc heap memory; lower values indicate higher fragmentation.
rw_in_progressRead-write transactions parked in the rw hash while waiting on replicas or duplicate resolution.

Wire compression

MetricWhat it measures
migrate_wire_compression_bytes_savedCumulative migration payload bytes kept off the wire.
migrate_wire_compression_fallbacksMigration records sent uncompressed because the codec failed on the sender.
repl_wire_compression_bytes_savedCumulative replication payload bytes kept off the wire.
repl_wire_compression_delta_apply_device_readsIncrements on every delta apply for storage-engine device and storage-engine pmem, before the prior-record load and without a cache check. It reads 0 on storage-engine memory. With cache-replica-writes enabled, some loads might not be physical device I/O.
repl_wire_compression_delta_hit_pctSender-side lifetime ratio that freezes at its last value, often 100.000, if deltas stop being attempted. Alert on the rate of repl_wire_compression_delta_attempts against the namespace write rate instead.
repl_wire_compression_fallbacksReplication writes that the replica rejected and the partition master retransmitted uncompressed.
wire_comp_cpu_pctCombined CPU percentage for all wire-compression codec operations. Reads 0.00 unless enable-benchmarks-wire-compression is enabled.

Cross-datacenter replication

MetricWhat it measures
lagReports the number of seconds that records wait at the source before shipping.
successNumber of records successfully shipped to a remote datacenter.
abandonedRecords abandoned after permanent destination failures; destination configuration must change to resume.
recoveriesPartitions recovered by reading from the primary index when the in-memory queue is full or incomplete.
recoveries_pendingRecoveries currently in progress; non-zero while catch-up runs.
retry_conn_resetRetries after connection reset to the destination (timeouts, network issues, destination restarts).
retry_destRetries after temporary destination errors such as key busy or device overload.
retry_no_nodeRetries when XDR cannot determine the destination master node.
lap_usMicroseconds to process one XDR lap (diagnostic; datacenter level only).
latency_msAverage network latency for successful shipments (datacenter level only).