Queries
For the complete documentation index see: llms.txt
All documentation pages available in markdown.
This page describes foreground and background queries against Aerospike namespaces and sets, including primary and secondary indexes, filter expressions, projection, pagination, and the query optimizer.
A query searches a namespace or set using a primary index (PI) or a secondary index (SI). Clients send the query to cluster nodes, which run it against the target index.
The Java SDK and Python SDK tabs use Developer SDK AEL text where it can express the example; the other tabs build the same expression with their client’s Exp builder. See the AEL reference and Query records in the Developer SDK.
Foreground queries
Similar to the SELECT statement in a relational database, a foreground query is read-only.
As the query iterates through the namespace data partitions it streams the current version of each record to the client.
Characteristics
- Selection of records:
- Queries, regardless of the index involved, can use filter expressions
such as
last_update()orset_name(). The filter acts as aWHEREclause. - For a foreground read query, the query optimizer selects a secondary index by reading the AEL the query carries. A query that carries no AEL must name its index with an explicit secondary index filter.
- Otherwise, the query targets the primary index. In the Developer SDKs, a query whose AEL no index can serve fails with error 201 (index not found) unless it allows the scan with
allowScansWithWhere; the AEL then filters the primary-index query. A PI query against a specified set name automatically leverages a set index, if one was created on this set.
- Queries, regardless of the index involved, can use filter expressions
such as
- Projection of record data:
- A list of named bins, also known as bin projection.
- operation projection of read-only bin-read operations, CDT read operations, and read expressions (which return their result in a computed bin). Available in Database 8.1.2 and later, matching the capability of single-record and batched
operate. QueryPolicy.includeBinDatacontrols whether to only return record metadata (digest, generation and TTL).
- Query by data partitions: If no partition IDs are specified, the client automatically runs the query against all 4096. This capability can be leveraged to horizontally scale the processing of results from a large dataset in parallel between multiple clients.
- Pagination: Return a specified number of records with the ability to continue the query from that point. The client uses a partition filter object to track progress across multiple partitions.
- Migration-tolerant queries: clients compatible with Database 6.0.0 and later ensure that queries handle the automatic data rebalancing (migration) that happens after a permanent cluster size changes.
Example: query with a filter expression
DataSet profiles = DataSet.of("test", "profiles");
RecordStream stream = session.query(profiles) .where("$.region == 'NA' and $.age >= 21") .withHint(hint -> hint.allowScansWithWhere()) .execute();stream.forEach(result -> { Record rec = result.recordOrThrow(); System.out.printf("uid=%s%n", rec.getString("uid"));});stream.close();from aerospike_sdk import DataSet, QueryHint
profiles = DataSet.of("test", "profiles")
stream = ( session.query(profiles) .where("$.region == 'NA' and $.age >= 21") .with_hint(QueryHint(allow_scans_with_where=True)) .execute())for row in stream: rec = row.record_or_raise() print(f"uid={rec.bins['uid']}")stream.close()use aerospike::expressions::{and, eq, ge, int_bin, int_val, string_bin, string_val};use aerospike::{Bins, PartitionFilter, QueryPolicy, Statement};use futures::StreamExt;
let stmt = Statement::new("test", "profiles", Bins::All);
let mut qp = QueryPolicy::default();qp.base_policy.filter_expression = Some(and(vec![ eq(string_bin("region".into()), string_val("NA".into())), ge(int_bin("age".into()), int_val(21)),]));
let rs = client.query(&qp, PartitionFilter::all(), stmt).await?;let mut stream = rs.into_stream();while let Some(result) = stream.next().await { let rec = result?; println!("uid={}", rec.bins["uid"]);}Statement stmt = new();stmt.SetNamespace("test");stmt.SetSetName("profiles");
QueryPolicy queryPolicy = new(){ filterExp = Exp.Build( Exp.And( Exp.EQ(Exp.StringBin("region"), Exp.Val("NA")), Exp.GE(Exp.IntBin("age"), Exp.Val(21))))};
using RecordSet rs = client.Query(queryPolicy, stmt);while (rs.Next()){ Record rec = rs.Record; Console.WriteLine($"uid={rec.GetString("uid")}");}// Requires: import as "github.com/aerospike/aerospike-client-go/v8"stmt := as.NewStatement("test", "profiles")
qp := as.NewQueryPolicy()qp.FilterExpression = as.ExpAnd( as.ExpEq(as.ExpStringBin("region"), as.ExpStringVal("NA")), as.ExpGreaterEq(as.ExpIntBin("age"), as.ExpIntVal(21)))
rs, err := client.Query(qp, stmt)if err != nil { panic(err)}for rec := range rs.Results() { if rec.Err != nil { panic(rec.Err) } fmt.Printf("uid=%v\n", rec.Record.Bins["uid"])}const query = client.query("test", "profiles");
const queryPolicy = new Aerospike.QueryPolicy({ filterExpression: exp.and( exp.eq(exp.binStr("region"), exp.str("NA")), exp.ge(exp.binInt("age"), exp.int(21)) ),});
const stream = query.foreach(queryPolicy);stream.on("data", (rec) => { console.log(`uid=${rec.bins.uid}`);});await new Promise((resolve, reject) => { stream.on("error", reject); stream.on("end", resolve);});bool query_cb(const as_val* val, void* udata) { if (!val) return false; as_record* rec = as_record_fromval(val); printf("uid=%s\n", as_record_get_str(rec, "uid")); return true;}
as_query q;as_query_init(&q, "test", "profiles");
as_exp_build(filter, as_exp_and( as_exp_cmp_eq(as_exp_bin_str("region"), as_exp_str("NA")), as_exp_cmp_ge(as_exp_bin_int("age"), as_exp_int(21))));
as_policy_query p;as_policy_query_init(&p);p.base.filter_exp = filter;
aerospike_query_foreach(&as, &err, &p, &q, query_cb, NULL);as_exp_destroy(filter);as_query_destroy(&q);Statement stmt = new Statement();stmt.setNamespace("test");stmt.setSetName("profiles");
QueryPolicy queryPolicy = new QueryPolicy();queryPolicy.filterExp = Exp.build( Exp.and( Exp.eq(Exp.stringBin("region"), Exp.val("NA")), Exp.ge(Exp.intBin("age"), Exp.val(21))));
try (RecordSet rs = client.query(queryPolicy, stmt)) { while (rs.next()) { Record rec = rs.getRecord(); System.out.printf("uid=%s%n", rec.getString("uid")); }}from aerospike_helpers.expressions import base as exp
query = client.query("test", "profiles")
expr = exp.And( exp.Eq(exp.StrBin("region"), exp.Val("NA")), exp.GE(exp.IntBin("age"), exp.Val(21)),).compile()
policy = {"expressions": expr}
for _, _, bins in query.results(policy): print(f"uid={bins['uid']}")No secondary index covers region or age, so this runs as a primary-index query that reads every record in the profiles set. The Developer SDK tabs allow that with allowScansWithWhere.
For a longer example with metadata filters and operation projection, see Query with a filter expression.
Background queries
A client application can issue an asynchronous background query to
modify records in place on the server. This is similar to an UPDATE
statement in a relational database.
Background queries apply either multiple native bin operations or a user-defined function (UDF) written in Lua to the records matched by the query. Using bin operations, also known as background ops, is more efficient and higher performing than using a Lua UDF, also known as a background UDF.
Characteristics
- Selection of records:
- Background queries, regardless of the index involved, can use filter expressions
such as
last_update()orset_name(). The filter acts as aWHEREclause. - The query optimizer applies to foreground read queries only. Background queries still require an explicit secondary index filter.
- Otherwise, the query targets the primary index. A PI query against a specified set name automatically leverages a set index, if one was created on this set.
- Background queries, regardless of the index involved, can use filter expressions
such as
- Clients can poll for progress and completion of a background query.
Background queries are not migration-tolerant and might miss records when a partition is migrating during data rebalancing.
Limiting query speed
Each individual query can be capped to run at a specified records per-second limit.
SREs can enforce that the totality of a user’s commands, including queries, are limited to a record per-second rate quota.
Query runtime optimization
A query’s runtime can be optimized by specifying its expected duration.
-
A query that is expected to run for a relatively long period of time is termed long query, and is the default expected duration. The runtime of a query depends on the complexity of its filter expression, the size of the dataset, indexing strategies, and the IOPS (input/output operations per second) capacity of the cluster nodes. Long queries include primary index queries running against a large dataset, or secondary index queries that return a large number of records.
-
A query that is expected to consistently run for a short duration and return a small number of records is termed short query. Explicitly designating a query’s expected duration as short allows Aerospike to optimize it for lower latency and a higher queries per-second (QPS) throughput.
-
The expected duration optimization works with queries running against both strong consistency (SC) and available and partition-tolerant (AP) namespaces.
Differences between short and long queries
The following table shows how queries are optimized to execute differently based on their expected duration.
| Short Query | Long Query |
|---|---|
| Short queries have a default 1s timeout (on the socket) | Long queries do not time out once they’re running |
| Short query execution cannot be throttled | Long query execution can be throttled by setting an RPS (records per second) cap |
| Short queries cannot be aborted | Long queries can be aborted |
| Short queries are not tracked | Active long queries are tracked and statistics on the most recent completed queries are kept in a queue (of size query-max-done) |
| Short queries are measured by query latency and record count histograms | Long queries do not provide latency and record count histograms |
| A short query runs in a single query thread, enabling a higher number of QPS | A long query runs in the number of query threads defined by single-query-threads |
| Short queries can be inlined to run in service threads | Long queries only run in query threads |
Long queries
In a long query there are typically lots of records to read in each data partition. The client begins by reserving and querying full partitions, and defers other partitions (for examples ones that are migrating) to a subsequent round, by which time the migration of these partitions to another cluster node is expected to have completed. If a deferred partition still isn’t full at the time the client tries to reserve it, it will retry several (by default 5) more times. Setting the query policy with adequate sleep between retries is more insurance that migrations do not time out and return an error code 11 PARTITION_UNAVAILABLE.
Long queries against an AP namespace that don’t return many records might error if they fail to reserve a migrating partition, because even moving it to the end of the querying sequence might still catch it while it’s migrating. For this reason, Database 7.1.0 moved from a boolean QueryPolicy.shortQuery to a QueryPolicy.expectedDuration with three options - short, long, and “relaxed AP long query”.
Following proper operating procedure, such as waiting for migrations to complete during a rolling upgrade, both avoids query errors and missing records. Under this approach, a relaxed long query against an AP namespace doesn’t miss any records, same as a long query against an SC namespace. Both settings avoid partition unavailable errors during properly executed rolling upgrade.
Short queries
Short queries don’t run for long enough to take advantage of the strategy governing long queries, where the client defers querying migrating partitions to the end. For this reason, short queries against an AP namespace are always relaxed, allowing the client to reserve a partition that is not full.
Short queries against SC namespaces will fail if they can’t reserve all the (full) partitions.