Skip to content

Regular expression syntax

For the complete documentation index see: llms.txt

All documentation pages available in markdown.

Aerospike’s string regex surfaces (Database 8.2.0 and later) all use International Components for Unicode (ICU) regular expression syntax (the java.util.regex-modeled dialect): the regex_compare and regex_replace operations, their string_regex_compare and string_regex_replace expression forms, and the Aerospike expression language (AEL) =~ match operator and regexReplace() function. That guide is the full syntax reference for character classes, anchors, quantifiers, groups, backreferences, Unicode properties, and more.

Writing the pattern

The pattern is a string argument, so it passes through the host language’s string literal before the server sees it. Languages with a raw-literal form show the pattern as written; the rest escape every backslash:

LanguageRaw form\d+ written as
Pythonr"…"r"\d+"
Rustr"…"r"\d+"
Go`…``\d+`
C#@"…"@"\d+"
Java, Cnone"\\d+"

"\\d+" and r"\d+" reach the server identically, so this is only a question of which form a reader can check against the ICU grammar at a glance. Use the raw form where the language has one and the pattern contains a backslash; a pattern with no backslash needs neither.

This applies to a regex inside AEL text as well, where the literal sits in a larger string — in Python, r"$.text:STRING.regexReplace(pattern: /\d+/g, replace: '')".

Flags

Regex flags are a bitmask integer on the operation and expression APIs. Combine them with bitwise OR. The client constants are RegexFlags in Python, StringRegexFlags in Java, and as_string_regex_flags in C.

FlagBit valueAELDescription
NONE / DEFAULT0—Case-sensitive; . does not match line terminators; ^ and $ match only at the start and end of the string.
CASE_INSENSITIVE1iCase-insensitive matching.
MULTILINE2m^ and $ match at line boundaries in addition to the start and end of the string.
DOTALL4s. matches any character, including \n and \r.
UNIX_LINES8—Only \n is recognized as a line terminator for ^, $, and ..
GLOBAL16gReplace all non-overlapping matches. Replace operations only (regex_replace, string_regex_replace, AEL regexReplace()); without it, only the first match is replaced.

CASE_INSENSITIVE, MULTILINE, DOTALL, and UNIX_LINES are ICU match flags, and Aerospike supports this subset of them. GLOBAL is not a match flag: it selects replace-all rather than changing how the pattern matches.

The server validates the flags word rather than masking it, so any bit outside this set fails the operation at parse time with AS_ERR_PARAMETER, naming the unrecognized bits, instead of being ignored. Passing GLOBAL to a read operation such as regex_compare is the same parse-time AS_ERR_PARAMETER; when GLOBAL is among the rejected bits, the message names only GLOBAL.

Where the flags go

Where a flag goes depends on which API you are writing:

SurfaceFlags goExample
Operate APIIn their own argumentregex_compare("email", "^ana", CASE_INSENSITIVE)
String expressions, client buildersIn their own argumentStringExp.regexCompare(Exp.val("^ana"), CASE_INSENSITIVE, bin)
String expressions, AEL textIn the pattern$.email:STRING =~ /^ana/i

The first two take the integer bit field above. AEL has no flags argument at all — the flags are letters written onto the end of the regex literal, after its closing /, as in /pattern/ims. AEL takes a subset: there is no letter for UNIX_LINES. g is accepted only on the pattern: operand of regexReplace(), and using it with =~ is a parse error; without g, regexReplace() replaces the first match only.

Reads and writes take different flags

regex_compare reads, so it takes the match flags only. regex_replace writes, so it takes the match flags plus GLOBAL, and a write policy — CREATE_ONLY, UPDATE_ONLY, NO_FAIL — as a separate argument.

Do not mix the two. Both arguments are integers, so passing a write flag where a regex flag belongs is accepted rather than rejected, and the operation silently does something other than what was asked. Use the regex flag constants for the regex argument and the write flag constants for the policy. The same applies to string_regex_replace.

Replacement string syntax

In replace operations the replacement string is ICU’s replacement dialect, not a server-parsed format. It accepts:

  • Plain text: any literal characters.
  • $0: the whole match.
  • $N: a capture group by number. ICU consumes digits greedily up to the pattern’s capture-group count, so $12 is group 12 when the pattern has at least 12 groups, and group 1 followed by a literal 2 when it does not.
  • ${name}: a named capture group, the partner of (?<name>...).
  • \ escapes: the backslash is consumed and the next character emitted literally, so \$ is a literal $ and \\ a literal backslash. \1 is therefore a literal 1, not a backreference. \uXXXX and \UXXXXXXXX are Unicode escapes.

A $ that does not resolve to an existing capture group fails the operation with AS_ERR_OP_NOT_APPLICABLE and no detail message on any record the pattern matches. That covers $$, a trailing bare $, $x, ${1}, and any group number the pattern does not have — which is where a replacement typo lands. On a record the pattern does not match, the replacement is never evaluated: the operation succeeds and the value is unchanged.

ICU readings that differ from PCRE

The following constructs are valid ICU syntax and are accepted. Their ICU semantics differ from what the same construct means in Perl Compatible Regular Expressions (PCRE).

ConstructICU readingPCRE reading
\KLiteral KResets the start of the match (“keep”)
\o{…}Literal o followed by a quantifierOctal code point
\g{N} (all-digit N)Literal g followed by a quantifierBackreference to group N
\g<n> / \g'n'All literal charactersBackreference to group n
\N{name}Named Unicode character (such as \N{LATIN SMALL LETTER A})Named Unicode character; same behavior

Non-ICU constructs rejected at parse time

Several constructs valid in PCRE or Python are not valid ICU syntax. The server rejects these when it parses the pattern, returning AS_ERR_PARAMETER with a message that names the construct and its ICU equivalent, rather than passing the pattern to the regex engine for a generic compile failure:

string_regex_compare: (?P<name>...) is not valid ICU regex syntax - use (?<name>...)

The check runs before the pattern touches any record, so the outcome never depends on the data and NO_FAIL cannot mask it.

Rejected spellingUse instead
(?P<name>...)(?<name>...)
(?P=name)\k<name>
(?'name'...)(?<name>...)
\g{name}\k<name>
{,n}{0,n}
\p or \P without braces\p{...}, so \pL becomes \p{L}
\p{^...}\P{...}, or \p{...} for \P{^...}
(?|...)Branch reset has no ICU equivalent
(?R)Pattern recursion has no ICU equivalent
(?n)Subroutine calls have no ICU equivalent
(?(...)...)Conditionals have no ICU equivalent
(*...)Backtracking control verbs have no ICU equivalent
\NNo ICU equivalent outside the named form \N{name}

After an escaping backslash, or inside \Q...\E, these characters are literals and do not trigger a rejection. A character class is not an exception to the syntax: [\pL] escapes the guided rejection above but still fails to compile, returning AS_ERR_PARAMETER with subcode AS_SUB_PARAM_STRING_REGEX_INVALID and a generic message rather than the guided one. Write \p{L} everywhere, including inside a character class.

A spelling that ICU accepts but reads differently is not rejected at all — it silently matches something else. Those are in ICU readings that differ from PCRE.

Examples

The examples below use ICU syntax. Combine flag bits with bitwise OR; CASE_INSENSITIVE | MULTILINE is 1 | 2, which equals 3.

The match examples read a bin applog holding two lines:

boot ok
Error: disk full

The replace examples start from a bin sku holding "A-1024".

Match with regex_compare (Operate API)

try (RecordStream rs = session.query(key)
.bin("applog").regexCompare("^error",
StringRegexFlags.CASE_INSENSITIVE | StringRegexFlags.MULTILINE)
.execute()) {
boolean matched = rs.next().recordOrThrow().operationResult(0).getBoolean();
// matched: true
}

Result: true.

^error matches the second line, not the value as a whole. MULTILINE is what lets ^ anchor at each line boundary rather than only at the start of the bin, and CASE_INSENSITIVE is what lets the lowercase pattern match the capitalised Error. Drop either flag and the same call returns false.

Filter with Aerospike expression language (AEL) =~

$.applog:STRING =~ /^error/im

Result: true — the same outcome as the previous example, from the same stored value.

The i and m suffix letters map to CASE_INSENSITIVE and MULTILINE, so /^error/im is the bitmask 3 written in AEL.

Replace with string_regex_replace (expression)

// sku holds "A-1024"
String exp = "$.sku:STRING.regexReplace(pattern: /(\\d+)/, replace: 'id=$1')";
// exp evaluates to "A-id=1024"; the bin is unchanged

Result: the expression evaluates to "A-id=1024". The bin is unchanged — an expression produces a value, and storing it takes a write.

(\d+) captures 1024, and $1 in the replacement puts it back after the literal id=. Without GLOBAL only the first match is replaced, which is the only match here. See Replacement string syntax.

Write with regex_replace (Operate API)

// sku holds "A-1024"
try (RecordStream rs = session.upsert(key)
.bin("sku").regexReplace("(\\d+)", "id=$1", StringRegexFlags.DEFAULT)
.execute()) {
rs.next().recordOrThrow();
}
// sku now holds "A-id=1024"

Result: sku now holds "A-id=1024". The operation itself returns no value.

That is the difference from the expression above: this one writes the result back to the bin in place. To see the new value, add a read of the same bin to the same operate() call. See regex_replace.

Resource limits

A pattern that exhausts the regex engine’s time or stack budget for a given input returns an AS_ERR_OP_NOT_APPLICABLE (26) error with subcode AS_SUB_OPNOT_STRING_REGEX_LIMIT_EXCEEDED, rather than running indefinitely. This is a different outcome from a non-match: the pattern was too expensive to evaluate against that value, not absent from it. Handle it as a signal to simplify the pattern rather than as a negative result.