Key metrics to monitor
For the complete documentation index see: llms.txt
All documentation pages available in markdown.
Aerospike recommends that you monitor your system with the metrics listed on this page. For a complete list of metrics, see the Metric reference.
Feature-specific counters are covered on the page for the feature. For intra-cluster bandwidth, see Configure wire compression and Configure SMD wire compression.
Operating system and server health
- Monitor system metrics with the Prometheus Node Exporter or an OS-specific tool.
- Use the Aerospike Health check to detect health outliers.
Enable
enable-health-checkon all nodes, then queryhealth-statsfor moving averages orhealth-outliersfor flagged nodes and devices.
Finding total namespace memory
In Aerospike Database 7.0.0 and later, calculate memory use separately for each namespace on each node.
When you created the namespace, you allocated storage for it and set a stop-writes threshold with stop-writes-sys-memory-pct that tells Aerospike when system memory is full enough to stop writes.
Run info namespace or its usage variant in Aerospike Admin (asadm) to identify storage placement and current use.
The output tables include Primary Index, Secondary Index, and Storage Engine columns with type and used bytes.
Use the following table to choose which metrics to sum for each component.
This sum is the used record-data and index bytes for RAM-backed components. It is not the namespace’s allocated or resident footprint, which also includes unused index stages, unused in-memory data capacity, and write caches.
| Component | Include when | Metric |
|---|---|---|
| Set indexes | Always | set_index_used_bytes |
| Record data | storage-engine is memory. | data_used_bytes |
| Primary index | index-type is not flash or pmem. | index_used_bytes |
| Secondary indexes | sindex-type is not flash or pmem. | sindex_used_bytes |
When record data is stored on SSD or persistent memory, data_used_bytes reports storage use.
For example, in a hybrid memory architecture namespace with record data on SSD and indexes in shared memory, sum index_used_bytes, set_index_used_bytes, and sindex_used_bytes.
Index memory monitoring
Monitor current index use and remaining index capacity according to where each index is stored:
| Index storage | Current use | Capacity |
|---|---|---|
| RAM and shared-memory indexes (primary and set) | Use index_used_bytes and set_index_used_bytes. | Use indexes_memory_used_pct only when indexes-memory-budget is greater than 0. |
| RAM and shared-memory secondary indexes | Use sindex_used_bytes when sindex-type is not flash or pmem. | Use indexes_memory_used_pct only when indexes-memory-budget is greater than 0. |
| Primary index on flash or persistent memory | Use index_used_bytes. | Use index_mounts_used_pct. |
| Secondary indexes on flash or persistent memory | Use sindex_used_bytes. | Use sindex_mounts_used_pct. |
Apply the rows independently because primary and secondary indexes can use different storage types. Set indexes are always stored in memory.
Memory budgets and stage allocation
indexes-memory-budget is an optional stop-writes threshold for the combined used bytes of RAM and shared-memory primary, secondary, and set indexes.
The default value of 0 disables the index-specific cap.
When the budget is 0, track the applicable *_used_bytes metrics and monitor overall memory pressure with stop-writes-sys-memory-pct.
In containers, enable cgroup-mem-tracking so that the threshold reflects the control group (cgroup) limit.
Indexes allocate space in stages as they grow, as configured by index-stage-size and sindex-stage-size.
The *_used_bytes metrics report memory holding live index elements.
Allocated stage footprint is larger.
For a primary index on flash, index_flash_alloc_bytes reports allocated space.
Database 8.2.0 adds the following metrics for allocated stages and the untouched space in the newest stage:
- Primary index:
index_{shmem,pmem}_alloc_bytesandindex_{shmem,pmem}_tail_bytes. - Secondary index:
sindex_{shmem,pmem,flash}_alloc_bytesandsindex_{shmem,pmem,flash}_tail_bytes. - Set index:
set_index_alloc_bytes. - Node-level footprint and container limits:
process_rss_bytes,cgroup_memory_used_bytes, andcgroup_memory_limit_bytes.
Eviction and stop-writes act on used bytes, so allocated bytes can sit above indexes-memory-budget while those thresholds remain unbreached.
See Monitor index memory allocation for how to interpret allocated, used, and tail bytes and reconcile them with process and control group memory.
Recommended alert metrics
Namespace capacity and write protection
| Metric | Description |
|---|---|
stop_writes | true when the namespace is not allowing client-originated writes. |
hwm_breached | true when a high-water disk or memory percentage threshold is breached for the namespace. |
data_avail_pct | Minimum available write blocks as a percentage across all storage files in a namespace (Database 7.0.0+). |
clock_skew_stop_writes | true when clock skew is outside tolerance and the namespace stops accepting writes or NSUP is blocked. |
index_shmem_tail_bytes | Untouched space in the newest shared-memory primary-index stage in Database 8.2.0 and later. A value of 0 means that the next insert might claim another index-stage-size of memory. |
Strong consistency
| Metric | Description |
|---|---|
dead_partitions | Number of partitions unavailable when all roster nodes are present. |
unavailable_partitions | Number of partitions unavailable when roster nodes are missing. |
Cluster and connections
| Metric | Description |
|---|---|
cluster_size | Size of the cluster. Compare across all nodes. |
client_connections | Number of active client connections to this node. |
client_connections_opened | Number of client connections created since the node started. |
fabric_connections_opened | Number of fabric connections created since the node started. |
heartbeat_connections_opened | Number of heartbeat connections created since the node started. |
System memory
| Metric | Description |
|---|---|
system_free_mem_kbytes / system_free_mem_pct | Free memory on the node. Control-group-aware when cgroup-mem-tracking is enabled. |
host_free_mem_kbytes / host_free_mem_pct | Host-level free memory (Database 8.1.2+), regardless of cgroup-mem-tracking. |
process_rss_bytes | Resident memory of the asd process (Database 8.2.0 and later). In a container, compare cgroup_memory_used_bytes with cgroup_memory_limit_bytes instead; the control group’s total can exceed asd’s RSS. |
Wire compression
| Metric | Description |
|---|---|
repl_wire_compression_delta_reject_apply | A sustained rise points at the replica’s stored base records, not at compression settings. |
repl_wire_compression_delta_reject_other | Investigate when the other three delta_reject_* counters stay near zero. |
Metrics to track
Track these metrics over time. Compare error counters to their corresponding success or complete counters.
Client and query errors
| Metric | What it measures |
|---|---|
client_read_error | Failed client read commands, such as invalid set name, unavailable partition, device I/O error, or key busy. |
client_write_error | Failed client write commands, including generation conflicts, key busy, record too big, and XDR-forbidden errors. |
client_delete_error | Failed client delete commands. |
client_udf_error | Failed client-originated UDF commands (excludes timeouts). |
batch_index_error | Batch index requests that completed with an error, often after client timeout or send delays. |
pi_query_short_basic_error | Failed basic short primary index queries. |
pi_query_long_basic_error | Failed basic long primary index queries. |
pi_query_aggr_error | Failed primary index query aggregations. |
pi_query_ops_bg_error | Failed ops background primary index queries. |
pi_query_udf_bg_error | Failed UDF background primary index queries. |
Storage and node runtime
| Metric | What it measures |
|---|---|
storage-engine.device[ix].defrag_q | Number of write blocks queued for defragmentation on a device. |
storage-engine.file[ix].write_q | Number of write blocks queued for write to a storage file. |
index_flash_alloc_pct | Percentage of mount space allocated for the primary index when index-type is flash (Enterprise Edition). |
heap_efficiency_pct | Ratio of allocated to active jemalloc heap memory; lower values indicate higher fragmentation. |
rw_in_progress | Read-write transactions parked in the rw hash while waiting on replicas or duplicate resolution. |
Wire compression
| Metric | What it measures |
|---|---|
migrate_wire_compression_bytes_saved | Cumulative migration payload bytes kept off the wire. |
migrate_wire_compression_fallbacks | Migration records sent uncompressed because the codec failed on the sender. |
repl_wire_compression_bytes_saved | Cumulative replication payload bytes kept off the wire. |
repl_wire_compression_delta_apply_device_reads | Increments on every delta apply for storage-engine device and storage-engine pmem, before the prior-record load and without a cache check. It reads 0 on storage-engine memory. With cache-replica-writes enabled, some loads might not be physical device I/O. |
repl_wire_compression_delta_hit_pct | Sender-side lifetime ratio that freezes at its last value, often 100.000, if deltas stop being attempted. Alert on the rate of repl_wire_compression_delta_attempts against the namespace write rate instead. |
repl_wire_compression_fallbacks | Replication writes that the replica rejected and the partition master retransmitted uncompressed. |
wire_comp_cpu_pct | Combined CPU percentage for all wire-compression codec operations. Reads 0.00 unless enable-benchmarks-wire-compression is enabled. |
Cross-datacenter replication
| Metric | What it measures |
|---|---|
lag | Reports the number of seconds that records wait at the source before shipping. |
success | Number of records successfully shipped to a remote datacenter. |
abandoned | Records abandoned after permanent destination failures; destination configuration must change to resume. |
recoveries | Partitions recovered by reading from the primary index when the in-memory queue is full or incomplete. |
recoveries_pending | Recoveries currently in progress; non-zero while catch-up runs. |
retry_conn_reset | Retries after connection reset to the destination (timeouts, network issues, destination restarts). |
retry_dest | Retries after temporary destination errors such as key busy or device overload. |
retry_no_node | Retries when XDR cannot determine the destination master node. |
lap_us | Microseconds to process one XDR lap (diagnostic; datacenter level only). |
latency_ms | Average network latency for successful shipments (datacenter level only). |