Skip to content

Secondary indexes

For the complete documentation index see: llms.txt

All documentation pages available in markdown.

This page describes how to create and query secondary indexes with the Java and Python Developer SDKs, including index types, AEL predicates, complete examples, and performance considerations. For server-side query types, index filters, and the query optimizer, see Secondary index queries.

What is a secondary index?

By default, Aerospike retrieves records by their primary key (digest). A secondary index (SI) lets you query records by the value of a specific bin.

Query typeBehavior
Single primary key lookupkey → record (O(1), always fast)
Multiple primary key lookupskeys → records (done in parallel across servers, usually fast)
Secondary index querybin value → records (requires index)
Primary index query(queries all records in parallel, O(N), can be slow)

Index types

TypeBin data typeQuery operations
IndexType.STRING (Java) / .string() (Python)StringEquality (==)
IndexType.INTEGER (Java) / .integer() (Python)IntegerEquality and range comparisons (<, <=, >, >=)
IndexType.GEO2DSPHERE (Java) / .geo2dsphere() (Python)GeoJSONWithin radius, within polygon (geoCompare in AEL)

When to use secondary indexes

✅ Good use cases:

  • Filtering by status, category, or type fields
  • Range queries on timestamps or numeric IDs
  • Geospatial queries (find nearby)

❌ Avoid when:

  • High cardinality (millions of unique values)
  • Zero or one result will be returned (for example, an alternate ID for a record)
  • Frequently updated bins
  • Can use primary key lookup instead

Creating an index

Note: Creating indexes in code in production environments is often an anti-pattern due to resource consumption. Use tools like asadm to manage indexes in higher environments instead.

import com.aerospike.client.sdk.DataSet;
import com.aerospike.client.sdk.query.IndexCollectionType;
import com.aerospike.client.sdk.query.IndexType;
DataSet users = DataSet.of("test", "users");
// Create an integer index on the "age" bin
session.createIndex(users, "age_idx", "age", IndexType.INTEGER, IndexCollectionType.DEFAULT)
.waitTillComplete();
// Create a string index on the "status" bin
session.createIndex(users, "status_idx", "status", IndexType.STRING, IndexCollectionType.DEFAULT)
.waitTillComplete();
// Create a GEO2DSPHERE index on a GeoJSON bin
DataSet places = DataSet.of("test", "places");
session.createIndex(places, "places_loc_idx", "loc", IndexType.GEO2DSPHERE, IndexCollectionType.DEFAULT)
.waitTillComplete();

📖 API reference: DataSet.of(...) | Session.createIndex(...)

Querying with indexes

Once an index exists, the query optimizer can select it: it reads your AEL and picks a readable secondary index (SI) that can serve the filter. If no readable index matches, the query fails with AS_ERR_SINDEX_NOT_FOUND (201), because allowScansWithWhere is false by default. Set it to true to let the query run as a PI query and read every record in the set instead. Always verify that indexes exist and are RW before production. See Query optimizer.

import com.aerospike.client.sdk.Record;
import com.aerospike.client.sdk.RecordStream;
// Server optimizer may select age_idx for $.age > 21
try (RecordStream stream = session.query(users)
.where("$.age > 21")
.execute()) {
stream.forEach(result -> {
Record row = result.recordOrThrow();
// Process row (for example, row.getString("name"))
});
}

📖 API reference: Session.query(DataSet) | ChainableQueryBuilder.where(...) | ChainableQueryBuilder.execute() | RecordStream.forEach(...) | RecordResult.recordOrThrow() | Record.getString(...)

Complete example

import com.aerospike.client.sdk.AerospikeException;
import com.aerospike.client.sdk.Cluster;
import com.aerospike.client.sdk.ClusterDefinition;
import com.aerospike.client.sdk.DataSet;
import com.aerospike.client.sdk.Record;
import com.aerospike.client.sdk.RecordStream;
import com.aerospike.client.sdk.Session;
import com.aerospike.client.sdk.policy.Behavior;
import com.aerospike.client.sdk.query.IndexCollectionType;
import com.aerospike.client.sdk.query.IndexType;
public class SecondaryIndexExample {
public static void main(String[] args) {
try (Cluster cluster = new ClusterDefinition("localhost", 3000).connect()) {
Session session = cluster.createSession(Behavior.DEFAULT);
DataSet users = DataSet.of("test", "users");
String k1 = "sidx-example-1";
String k2 = "sidx-example-2";
String k3 = "sidx-example-3";
String ageIndex = "age_idx_demo";
String statusIndex = "status_idx_demo";
// Cleanup so the example is repeatable.
session.delete(users.ids(k1, k2, k3)).execute().close();
try {
session.dropIndex(users, ageIndex).waitTillComplete();
} catch (AerospikeException ignored) {
// Index may not exist yet.
}
try {
session.dropIndex(users, statusIndex).waitTillComplete();
} catch (AerospikeException ignored) {
// Index may not exist yet.
}
// Seed sample data.
session.insert(users)
.bins("name", "age", "status")
.id(k1).values("Alice", 28, "sidx_active")
.id(k2).values("Bob", 42, "sidx_active")
.id(k3).values("Carol", 19, "sidx_inactive")
.execute();
// Create secondary indexes.
session.createIndex(users, ageIndex, "age", IndexType.INTEGER, IndexCollectionType.DEFAULT)
.waitTillComplete();
session.createIndex(users, statusIndex, "status", IndexType.STRING, IndexCollectionType.DEFAULT)
.waitTillComplete();
// Compound predicate; server optimizer selects the most selective readable index
System.out.println("sidx_active users age >= 30:");
try (RecordStream ageQuery = session.query(users)
.where("$.status == 'sidx_active' and $.age >= 30")
.readingOnlyBins("name", "age")
.execute()) {
ageQuery.forEach(result -> {
Record user = result.recordOrThrow();
System.out.println(" - " + user.getString("name") + " (" + user.getInt("age") + ")");
});
}
// Server optimizer selects an SI for $.status == 'sidx_active'
System.out.println("\nsidx_active users:");
try (RecordStream statusQuery = session.query(users)
.where("$.status == 'sidx_active'")
.readingOnlyBins("name", "status")
.execute()) {
statusQuery.forEach(result -> {
Record user = result.recordOrThrow();
System.out.println(" - " + user.getString("name"));
});
}
}
}
}

📖 API reference: ClusterDefinition(String,int) | ClusterDefinition.connect() | Cluster.createSession(Behavior) | Cluster.close() | DataSet.of(...) | DataSet.ids(...) | Session.insert(DataSet) | Session.delete(List) | Session.delete(Key) | Session.createIndex(...) | Session.dropIndex(...) | Session.query(DataSet) | OperationObjectBuilder.bins(...) | IdValuesBuilder.id(...) | IdValuesRowBuilder.values(...) | ChainableQueryBuilder.where(...) | ChainableQueryBuilder.readingOnlyBins(...) | ChainableQueryBuilder.execute() | ChainableNoBinsBuilder.execute() | RecordStream.forEach(...) | RecordStream.close() | RecordResult.recordOrThrow() | Record.getString(...) | Record.getInt(...) | AerospikeException

Performance considerations

  • Indexes consume memory on every node
  • Index updates add write latency
  • Queries without matching indexes run as PI queries (slow)
  • A missing or dropped index sends the query to the primary index, where it reads every record. Monitor show jobs queries alongside PI metrics to catch this.
  • Monitor index memory usage in production

Next steps