Regular expression syntax
For the complete documentation index see: llms.txt
All documentation pages available in markdown.
Aerospike’s string regex surfaces (Database 8.2.0 and later) all use
International Components for Unicode (ICU) regular expression syntax
(the java.util.regex-modeled dialect): the
regex_compare and
regex_replace operations, their
string_regex_compare and
string_regex_replace expression forms, and the
Aerospike expression language (AEL) =~ match operator and regexReplace() function. That guide is
the full syntax reference for character classes, anchors, quantifiers, groups, backreferences,
Unicode properties, and more.
Writing the pattern
The pattern is a string argument, so it passes through the host language’s string literal before the server sees it. Languages with a raw-literal form show the pattern as written; the rest escape every backslash:
| Language | Raw form | \d+ written as |
|---|---|---|
| Python | r"…" | r"\d+" |
| Rust | r"…" | r"\d+" |
| Go | `…` | `\d+` |
| C# | @"…" | @"\d+" |
| Java, C | none | "\\d+" |
"\\d+" and r"\d+" reach the server identically, so this is only a question of which form a reader can check against the ICU grammar at a glance. Use the raw form where the language has one and the pattern contains a backslash; a pattern with no backslash needs neither.
This applies to a regex inside AEL text as well, where the literal sits in a larger string — in Python, r"$.text:STRING.regexReplace(pattern: /\d+/g, replace: '')".
Flags
Regex flags are a bitmask integer on the operation and expression APIs. Combine them with bitwise OR. The client constants are RegexFlags in Python, StringRegexFlags in Java, and as_string_regex_flags in C.
| Flag | Bit value | AEL | Description |
|---|---|---|---|
NONE / DEFAULT | 0 | — | Case-sensitive; . does not match line terminators; ^ and $ match only at the start and end of the string. |
CASE_INSENSITIVE | 1 | i | Case-insensitive matching. |
MULTILINE | 2 | m | ^ and $ match at line boundaries in addition to the start and end of the string. |
DOTALL | 4 | s | . matches any character, including \n and \r. |
UNIX_LINES | 8 | — | Only \n is recognized as a line terminator for ^, $, and .. |
GLOBAL | 16 | g | Replace all non-overlapping matches. Replace operations only (regex_replace, string_regex_replace, AEL regexReplace()); without it, only the first match is replaced. |
CASE_INSENSITIVE, MULTILINE, DOTALL, and UNIX_LINES are ICU match flags, and Aerospike supports this subset of them. GLOBAL is not a match flag: it selects replace-all rather than changing how the pattern matches.
The server validates the flags word rather than masking it, so any bit outside this set fails the operation at parse time with AS_ERR_PARAMETER, naming the unrecognized bits, instead of being ignored. Passing GLOBAL to a read operation such as regex_compare is the same parse-time AS_ERR_PARAMETER; when GLOBAL is among the rejected bits, the message names only GLOBAL.
Where the flags go
Where a flag goes depends on which API you are writing:
| Surface | Flags go | Example |
|---|---|---|
| Operate API | In their own argument | regex_compare("email", "^ana", CASE_INSENSITIVE) |
| String expressions, client builders | In their own argument | StringExp.regexCompare(Exp.val("^ana"), CASE_INSENSITIVE, bin) |
| String expressions, AEL text | In the pattern | $.email:STRING =~ /^ana/i |
The first two take the integer bit field above. AEL has no flags argument at all — the flags are letters written onto the end of the regex literal, after its closing /, as in /pattern/ims. AEL takes a subset: there is no letter for UNIX_LINES. g is accepted only on the pattern: operand of regexReplace(), and using it with =~ is a parse error; without g, regexReplace() replaces the first match only.
Reads and writes take different flags
regex_compare reads, so it takes the match flags only. regex_replace writes, so it takes the match flags plus GLOBAL, and a write policy — CREATE_ONLY, UPDATE_ONLY, NO_FAIL — as a separate argument.
Do not mix the two. Both arguments are integers, so passing a write flag where a regex flag belongs is accepted rather than rejected, and the operation silently does something other than what was asked. Use the regex flag constants for the regex argument and the write flag constants for the policy. The same applies to string_regex_replace.
Replacement string syntax
In replace operations the replacement string is ICU’s replacement dialect, not a server-parsed format. It accepts:
- Plain text: any literal characters.
$0: the whole match.$N: a capture group by number. ICU consumes digits greedily up to the pattern’s capture-group count, so$12is group 12 when the pattern has at least 12 groups, and group 1 followed by a literal2when it does not.${name}: a named capture group, the partner of(?<name>...).\escapes: the backslash is consumed and the next character emitted literally, so\$is a literal$and\\a literal backslash.\1is therefore a literal1, not a backreference.\uXXXXand\UXXXXXXXXare Unicode escapes.
A $ that does not resolve to an existing capture group fails the operation with AS_ERR_OP_NOT_APPLICABLE and no detail message on any record the pattern matches. That covers $$, a trailing bare $, $x, ${1}, and any group number the pattern does not have — which is where a replacement typo lands. On a record the pattern does not match, the replacement is never evaluated: the operation succeeds and the value is unchanged.
ICU readings that differ from PCRE
The following constructs are valid ICU syntax and are accepted. Their ICU semantics differ from what the same construct means in Perl Compatible Regular Expressions (PCRE).
| Construct | ICU reading | PCRE reading |
|---|---|---|
\K | Literal K | Resets the start of the match (“keep”) |
\o{…} | Literal o followed by a quantifier | Octal code point |
\g{N} (all-digit N) | Literal g followed by a quantifier | Backreference to group N |
\g<n> / \g'n' | All literal characters | Backreference to group n |
\N{name} | Named Unicode character (such as \N{LATIN SMALL LETTER A}) | Named Unicode character; same behavior |
Non-ICU constructs rejected at parse time
Several constructs valid in PCRE or Python are not valid ICU syntax. The server rejects these when it parses the pattern, returning AS_ERR_PARAMETER with a message that names the construct and its ICU equivalent, rather than passing the pattern to the regex engine for a generic compile failure:
string_regex_compare: (?P<name>...) is not valid ICU regex syntax - use (?<name>...)The check runs before the pattern touches any record, so the outcome never depends on the data and NO_FAIL cannot mask it.
| Rejected spelling | Use instead |
|---|---|
(?P<name>...) | (?<name>...) |
(?P=name) | \k<name> |
(?'name'...) | (?<name>...) |
\g{name} | \k<name> |
{,n} | {0,n} |
\p or \P without braces | \p{...}, so \pL becomes \p{L} |
\p{^...} | \P{...}, or \p{...} for \P{^...} |
(?|...) | Branch reset has no ICU equivalent |
(?R) | Pattern recursion has no ICU equivalent |
(?n) | Subroutine calls have no ICU equivalent |
(?(...)...) | Conditionals have no ICU equivalent |
(*...) | Backtracking control verbs have no ICU equivalent |
\N | No ICU equivalent outside the named form \N{name} |
After an escaping backslash, or inside \Q...\E, these characters are literals and do not trigger a rejection. A character class is not an exception to the syntax: [\pL] escapes the guided rejection above but still fails to compile, returning AS_ERR_PARAMETER with subcode AS_SUB_PARAM_STRING_REGEX_INVALID and a generic message rather than the guided one. Write \p{L} everywhere, including inside a character class.
A spelling that ICU accepts but reads differently is not rejected at all — it silently matches something else. Those are in ICU readings that differ from PCRE.
Examples
The examples below use ICU syntax. Combine flag bits with bitwise OR; CASE_INSENSITIVE | MULTILINE is 1 | 2, which equals 3.
The match examples read a bin applog holding two lines:
boot okError: disk fullThe replace examples start from a bin sku holding "A-1024".
Match with regex_compare (Operate API)
try (RecordStream rs = session.query(key) .bin("applog").regexCompare("^error", StringRegexFlags.CASE_INSENSITIVE | StringRegexFlags.MULTILINE) .execute()) { boolean matched = rs.next().recordOrThrow().operationResult(0).getBoolean(); // matched: true}stream = session.query(key).bin("applog").str_regex_compare( r"^error", StringRegexFlags.CASE_INSENSITIVE | StringRegexFlags.MULTILINE).execute()matched = stream.first_or_raise().record_or_raise().operation_result(0)# matched: True// Requires: use aerospike::operations::string as str_op;let rec = client.operate(&WritePolicy::default(), &key, &[str_op::regex_compare_with_flags("applog", r"^error", StringRegexFlags::CASE_INSENSITIVE | StringRegexFlags::MULTILINE)]).await?;let matched = rec.bins.get("applog");// matched: trueRecord rec = client.Operate(null, key, StringOperation.RegexCompare("applog", "^error", StringRegexFlags.CASE_INSENSITIVE | StringRegexFlags.MULTILINE));bool matched = (bool)rec.GetValue("applog");// matched: true// Requires: import as "github.com/aerospike/aerospike-client-go/v8"rec, err := client.Operate(nil, key, as.StrRegexCompareWithFlagsOp("applog", `^error`, as.StringRegexCaseInsensitive|as.StringRegexMultiline))matched := rec.Bins["applog"].(bool)// matched: trueconst Aerospike = require('aerospike')const strings = Aerospike.stringsconst { CASE_INSENSITIVE, MULTILINE } = strings.regexFlags
const record = await client.operate(key, [ strings.regexCompareFlags('applog', '^error', CASE_INSENSITIVE | MULTILINE)])const matched = record.bins.applog// matched: trueas_operations ops;as_operations_init(&ops, 1);as_operations_string_regex_compare_flags(&ops, "applog", NULL, "^error", AS_STRING_REGEX_FLAGS_CASE_INSENSITIVE | AS_STRING_REGEX_FLAGS_MULTILINE);
as_record* rec = NULL;if (aerospike_key_operate(&as, &err, NULL, &key, &ops, &rec) == AEROSPIKE_OK) { bool matched = as_record_get_bool(rec, "applog"); // matched: true as_record_destroy(rec);}as_operations_destroy(&ops);Record record = client.operate(null, key, StringOperation.regexCompare("applog", "^error", StringRegexFlags.CASE_INSENSITIVE | StringRegexFlags.MULTILINE));boolean matched = record.getBoolean("applog");// matched: truefrom aerospike_helpers.operations import string_operations as sofrom aerospike_helpers.string_helpers import RegexFlags
_, _, bins = client.operate(key, [so.regex_compare( "applog", r"^error", RegexFlags.CASE_INSENSITIVE | RegexFlags.MULTILINE)])matched = bins["applog"]# matched: TrueResult: true.
^error matches the second line, not the value as a whole. MULTILINE is what lets ^
anchor at each line boundary rather than only at the start of the bin, and
CASE_INSENSITIVE is what lets the lowercase pattern match the capitalised Error. Drop
either flag and the same call returns false.
Filter with Aerospike expression language (AEL) =~
$.applog:STRING =~ /^error/imResult: true — the same outcome as the previous example, from the same stored value.
The i and m suffix letters map to CASE_INSENSITIVE and MULTILINE, so /^error/im
is the bitmask 3 written in AEL.
Replace with string_regex_replace (expression)
// sku holds "A-1024"String exp = "$.sku:STRING.regexReplace(pattern: /(\\d+)/, replace: 'id=$1')";// exp evaluates to "A-id=1024"; the bin is unchanged# sku holds "A-1024"exp = r"$.sku:STRING.regexReplace(pattern: /(\d+)/, replace: 'id=$1')"# exp evaluates to "A-id=1024"; the bin is unchanged// sku holds "A-1024"use aerospike::expressions::{string as str_exp, string_bin, string_val};use aerospike::operations::string::{StringPolicy, StringRegexFlags};
let exp = str_exp::regex_replace(&StringPolicy::default(), string_bin("sku".into()), string_val(r"(\d+)".into()), string_val("id=$1".into()), StringRegexFlags::DEFAULT);// exp evaluates to "A-id=1024"; the bin is unchanged// sku holds "A-1024"Expression exp = Exp.Build( StringExp.RegexReplace(StringPolicy.Default, Exp.Val(@"(\d+)"), Exp.Val("id=$1"), StringRegexFlags.DEFAULT, Exp.StringBin("sku")));// exp evaluates to "A-id=1024"; the bin is unchanged// Requires: import as "github.com/aerospike/aerospike-client-go/v8"// sku holds "A-1024"exp := as.ExpStringRegexReplace(as.DefaultStringPolicy, as.ExpStringBin("sku"), as.ExpStringVal(`(\d+)`), as.ExpStringVal("id=$1"), as.StringRegexDefault)// exp evaluates to "A-id=1024"; the bin is unchangedconst Aerospike = require('aerospike')const exp = Aerospike.exp
// sku holds "A-1024"const expression = exp.string.regexReplace(null, '(\\d+)', 'id=$1', Aerospike.strings.regexFlags.NONE, exp.binStr('sku'))// expression evaluates to "A-id=1024"; the bin is unchanged// sku holds "A-1024"as_exp_build(predexp, as_exp_string_regex_replace(NULL, "(\\d+)", "id=$1", AS_STRING_REGEX_FLAGS_NONE, as_exp_bin_str("sku")));// exp evaluates to "A-id=1024"; the bin is unchanged// sku holds "A-1024"Expression exp = Exp.build( StringExp.regexReplace(StringPolicy.Default, Exp.val("(\\d+)"), Exp.val("id=$1"), StringRegexFlags.DEFAULT, Exp.stringBin("sku")));// exp evaluates to "A-id=1024"; the bin is unchanged# sku holds "A-1024"from aerospike_helpers.expressions import string as str_exprfrom aerospike_helpers.string_helpers import StringPolicy, RegexFlags
exp = str_expr.RegexReplace( StringPolicy(), pattern=r"(\d+)", replacement="id=$1", regex_flags=RegexFlags.DEFAULT, bin="sku",).compile()# exp evaluates to "A-id=1024"; the bin is unchangedResult: the expression evaluates to "A-id=1024". The bin is unchanged — an
expression produces a value, and storing it takes a write.
(\d+) captures 1024, and $1 in the replacement puts it back after the literal id=.
Without GLOBAL only the first match is replaced, which is the only match here. See
Replacement string syntax.
Write with regex_replace (Operate API)
// sku holds "A-1024"try (RecordStream rs = session.upsert(key) .bin("sku").regexReplace("(\\d+)", "id=$1", StringRegexFlags.DEFAULT) .execute()) { rs.next().recordOrThrow();}// sku now holds "A-id=1024"# sku holds "A-1024"session.upsert(key).bin("sku").str_regex_replace( r"(\d+)", "id=$1", StringRegexFlags.DEFAULT).execute()# sku now holds "A-id=1024"// Requires: use aerospike::operations::string as str_op;// sku holds "A-1024"client.operate(&WritePolicy::default(), &key, &[str_op::regex_replace(&StringPolicy::default(), "sku", r"(\d+)", "id=$1", StringRegexFlags::DEFAULT)]).await?;// sku now holds "A-id=1024"// sku holds "A-1024"client.Operate(null, key, StringOperation.RegexReplace(StringPolicy.Default, "sku", @"(\d+)", "id=$1", StringRegexFlags.DEFAULT));// sku now holds "A-id=1024"// Requires: import as "github.com/aerospike/aerospike-client-go/v8"// sku holds "A-1024"client.Operate(nil, key, as.StrRegexReplaceOp(as.DefaultStringPolicy, "sku", `(\d+)`, "id=$1", as.StringRegexDefault))// sku now holds "A-id=1024"const Aerospike = require('aerospike')const strings = Aerospike.strings
// sku holds "A-1024"await client.operate(key, [ strings.regexReplace('sku', '(\\d+)', 'id=$1', strings.regexFlags.NONE)])// sku now holds "A-id=1024"// sku holds "A-1024"as_operations ops;as_operations_init(&ops, 1);as_operations_string_regex_replace(&ops, "sku", NULL, NULL, "(\\d+)", "id=$1", AS_STRING_REGEX_FLAGS_NONE);
aerospike_key_operate(&as, &err, NULL, &key, &ops, NULL);as_operations_destroy(&ops);// sku now holds "A-id=1024"// sku holds "A-1024"client.operate(null, key, StringOperation.regexReplace(StringPolicy.Default, "sku", "(\\d+)", "id=$1", StringRegexFlags.DEFAULT));// sku now holds "A-id=1024"# sku holds "A-1024"from aerospike_helpers.operations import string_operations as sofrom aerospike_helpers.string_helpers import RegexFlags
client.operate(key, [so.regex_replace("sku", r"(\d+)", "id=$1", RegexFlags.DEFAULT)])# sku now holds "A-id=1024"Result: sku now holds "A-id=1024". The operation itself returns no value.
That is the difference from the expression above: this one writes the result back to the
bin in place. To see the new value, add a read of the same bin to the same operate()
call. See regex_replace.
Resource limits
A pattern that exhausts the regex engine’s time or stack budget for a given input returns an
AS_ERR_OP_NOT_APPLICABLE (26) error with subcode
AS_SUB_OPNOT_STRING_REGEX_LIMIT_EXCEEDED, rather than running indefinitely. This is a
different outcome from a non-match: the pattern was too expensive to evaluate against that
value, not absent from it. Handle it as a signal to simplify the pattern rather than as a
negative result.