Monitor client performance
For the complete documentation index see: llms.txt
All documentation pages available in markdown.
Applies to
- Aerospike Developer SDK (Java 21+ and Python 3.11+)
- Aerospike Database 6.0 or later unless a section states otherwise
Client-side performance monitoring helps you detect regressions, troubleshoot timeouts and connection issues, and correlate application behavior with cluster health. Typical signals include operation volume, error and retry rates, latency distributions, and connection pool usage, aggregated by cluster, node, or namespace when your observability stack supports those dimensions.
The Python SDK has a built-in, cluster-scoped metrics collector: call cluster.enable_metrics() and poll await cluster.metrics() for a structured snapshot of latency histograms, byte counts, connection lifecycle, and retry/timeout counters. It’s off by default. The Java SDK does not currently embed an equivalent collector or periodic metrics log like the classic Aerospike clients; for Java, instrument your application around SDK calls and export measurements to the tools your team already uses (for example Prometheus, OpenTelemetry, Datadog, or structured logs). The sections below cover the Python built-in collector, followed by common metric categories and manual instrumentation patterns for both languages. For complementary request-level detail during development, see Configure logging.
Except where noted, snippets on this page use the imports below. A snippet lists additional import lines only when it needs a type not shown here. When this page includes a Complete example section, that block is fully self-contained with every import required to run it.
import com.aerospike.client.sdk.Cluster;import com.aerospike.client.sdk.ClusterDefinition;import time
from aerospike_sdk import Behavior, ClusterDefinition, DataSetAvailable metrics
Standard metrics
| Metric | Type | Description |
|---|---|---|
aerospike.operations.total | Counter | Total operations executed |
aerospike.operations.errors | Counter | Failed operations |
aerospike.latency.read | Histogram | Read operation latency |
aerospike.latency.write | Histogram | Write operation latency |
aerospike.connections.active | Gauge | Current open connections |
aerospike.connections.pool | Gauge | Connections in pool |
Extended metrics
| Metric | Type | Description |
|---|---|---|
aerospike.batch.size | Histogram | Batch operation sizes |
aerospike.retries | Counter | Operation retry count |
aerospike.timeouts | Counter | Timeout events |
aerospike.cluster.nodes | Gauge | Nodes in cluster |
Enable metrics programmatically
// Additional imports for this example:import io.micrometer.core.instrument.MeterRegistry;import io.micrometer.core.instrument.Timer;import io.micrometer.prometheus.PrometheusConfig;import io.micrometer.prometheus.PrometheusMeterRegistry;
// Application-owned registry (not provided by the Aerospike JAR)MeterRegistry registry = new PrometheusMeterRegistry(PrometheusConfig.DEFAULT);
try (Cluster cluster = new ClusterDefinition("localhost", 3000).connect()) { Timer.Sample sample = Timer.start(registry); try { // session.query(...).execute(); etc. } finally { sample.stop(registry.timer("aerospike.sdk.requests")); }}📖 API reference:
ClusterDefinition(String,int)|ClusterDefinition.connect()|Cluster.close()|ChainableQueryBuilder.execute()
async def main(): async with await ClusterDefinition("localhost", 3000).connect() as cluster: session = cluster.create_session(Behavior.DEFAULT) users = DataSet.of("test", "users")
start = time.perf_counter() stream = await session.query(users.id("user-1")).execute() await stream.first() stream.close() duration_ms = (time.perf_counter() - start) * 1000 print(f"aerospike.read.latency_ms={duration_ms:.2f}")📖 API reference:
ClusterDefinition|Cluster.create_session()|DataSet.of()|DataSet.id()|Behavior.DEFAULT|Session.query()|RecordStream.first()|RecordStream.close()|QueryBuilder.execute()
Prometheus integration
# Additional imports for this example:from prometheus_client import CollectorRegistry, generate_latestfrom prometheus_client import Counter, Histogram
registry = CollectorRegistry()reads_total = Counter("aerospike_reads_total", "Read operations", registry=registry)read_latency = Histogram("aerospike_read_latency_ms", "Read latency ms", registry=registry)
async def main(): async with await ClusterDefinition("localhost", 3000).connect() as cluster: session = cluster.create_session(Behavior.DEFAULT) users = DataSet.of("test", "users") with read_latency.time(): stream = await session.query(users.id("user-1")).execute() await stream.first() stream.close() reads_total.inc()
# Expose metrics endpoint@app.route('/metrics')def metrics(): return generate_latest(registry)📖 API reference:
ClusterDefinition|Cluster.create_session()|DataSet.of()|DataSet.id()|Behavior.DEFAULT|Session.query()|RecordStream.first()|RecordStream.close()|QueryBuilder.execute()
Access metrics snapshots
Get current metric values programmatically:
// The Java client does not expose cluster.getMetrics() / MetricsSnapshot.// Export Micrometer/Prometheus from your app or scrape logs for connection events.📖 API reference:
com.aerospike.client.sdk
from aerospike_sdk.metrics import CommandType, LatencyType
cluster.enable_metrics()
# ... application traffic runs here ...
snapshot = await cluster.metrics()reads = snapshot.latency(LatencyType.READ)print(f"{reads.count} reads, avg {reads.average:.1f} ms")
# Per-command, cluster-wide detail:agg = snapshot.cluster_aggregatedgets = agg.command_histogram(CommandType.GET)print(f"{gets.count} GET commands, {gets.average:.0f}ms avg")Snapshot values are cumulative since metrics were enabled, not deltas since the last poll. Poll on an interval (for example every 30 seconds) rather than per operation.
📖 PSDK Guide: Client metrics
Metrics log files
Enable periodic metrics logging to files:
// No built-in periodic metrics log in the Java client; use your logging framework// and rotate files under /var/log/... as needed (see Configure logging).📖 API reference:
com.aerospike.client.sdk
# Additional imports for this example:import logging
logger = logging.getLogger("aerospike_app_metrics")logger.setLevel(logging.INFO)logger.info("aerospike.read.latency_ms=12.7")📖 PSDK Guide: Logging
Client identification labels
Add labels to identify metrics by application or environment:
// Apply labels in your metrics backend (Prometheus relabel, Datadog tags, etc.), not on ClusterDefinition.📖 API reference:
com.aerospike.client.sdk
# Additional imports for this example:from dataclasses import dataclass
@dataclass(frozen=True)class MetricTags: app: str = "my-service" env: str = "production" region: str = "us-west-2"
tags = MetricTags()print(tags)Integration with monitoring systems
Prometheus + Grafana
- Export metrics via HTTP endpoint
- Configure Prometheus to scrape your application
- Import the Aerospike client dashboard in Grafana
Datadog
// Additional imports for this example:import io.micrometer.datadog.DatadogConfig;import io.micrometer.datadog.DatadogMeterRegistry;import java.time.Clock;
DatadogMeterRegistry registry = new DatadogMeterRegistry( DatadogConfig.DEFAULT, Clock.SYSTEM);// Register timers/counters around your own Aerospike calls; the SDK does not wire this for you.📖 API reference:
com.aerospike.client.sdk
# Additional imports for this example:from datadog import statsd
async def main(): async with await ClusterDefinition("localhost", 3000).connect() as cluster: session = cluster.create_session(Behavior.DEFAULT) users = DataSet.of("test", "users") stream = await session.query(users.id("user-1")).execute() row = await stream.first() stream.close() statsd.increment("aerospike.read.total", tags=["app:my-service"]) statsd.gauge("aerospike.read.found", 1 if row is not None else 0)📖 API reference:
ClusterDefinition|Cluster.create_session()|DataSet.of()|DataSet.id()|Behavior.DEFAULT|Session.query()|RecordStream.first()|RecordStream.close()|QueryBuilder.execute()
Performance impact
Metrics collection has minimal overhead:
- Counter/gauge operations: ~100ns
- Histogram recording: ~500ns
- Total impact: <1% CPU overhead
For latency-critical applications, you can disable histograms:
// Configure Micrometer histogram buckets on your own meters; unrelated to the Aerospike JAR.📖 API reference:
com.aerospike.client.sdk
# Configure histogram behavior in your metrics backend, not on ClusterDefinition.# Example: use a Counter-only approach in your instrumentation path.Next steps
Configure Logging
Set up client-side logging for debugging.
Tune Performance
Optimize client configuration with Behaviors.