Skip to content

Compare String values

For the complete documentation index see: llms.txt

All documentation pages available in markdown.

The expression comparison operators — eq, ne, gt, ge, lt, and le, available on String values since Database 5.2 — order String values by UTF-8 bytes. String search operations and normalize_nfc (Database 8.2.0 and later) treat canonically equivalent spellings as the same text. Those two rules disagree on mixed Unicode forms of the same word.

Byte order

Expression comparison operators order String values by UTF-8 bytes, which matches Unicode code-point order. All ASCII sorts before any accented letter, and every uppercase ASCII letter sorts before every lowercase ASCII letter. When one value is a prefix of the other, the shorter value sorts first.

Consequences of that order:

  • "Zebra" < "apple" is true, because Z sorts before a.
  • "Ålesund" > "Zurich" is true, because Å is a multi-byte UTF-8 letter and sorts after every ASCII letter.

This is the same byte order a List or Map uses to order String elements, which Order and compare collection elements covers for every data type.

Worked example

Seven stored city names, with Málaga stored twice: once with a precomposed á (U+00E1), and once as a plus a combining acute accent (U+0301). The two Málaga spellings look the same on screen.

#Byte order (expression eq, ne, gt, ge, lt, le)Typical dictionary order
1Málaga (a + combining acute)Ålesund
2Málaga (precomposed á)bergen
3São PauloMálaga
4ZurichMálaga
5bergenÖstersund
6ÅlesundSão Paulo
7ÖstersundZurich

A letter-range filter follows byte order, not dictionary order. Over the list above, city >= "A" AND city < "B" matches no record: Ålesund is the value a dictionary range would return, and it sorts after Z.

Equality and search disagree

Canonical equivalence means two Unicode spellings represent the same text: a precomposed é (U+00E9) and an e followed by a combining acute accent (U+0301) are equivalent. Unicode Normalization Form C (NFC) stores the precomposed letter. Normalization Form D (NFD) stores the letter plus the combining mark.

For a bin holding NFC café (precomposed é), compared with NFD café (e + U+0301):

The six search operations treat the two spellings as the same text. eq compares UTF-8 bytes, so the two spellings are not equal. split, regex_compare, and regex_replace do not use canonical equivalence either; they match on exact code points.

// city holds NFC "café" (precomposed é, U+00E9)
String eqExp = "$.city:STRING == 'cafe\u0301'";
// eqExp evaluates to false — eq compares UTF-8 bytes
String containsExp = "$.city:STRING.contains(needle: 'cafe\u0301')";
// containsExp evaluates to true — search uses canonical equivalence

Search matching is case-sensitive. "Café" and "café" do not match.

Normalize before you compare

normalize_nfc rewrites a String bin to NFC. The expression form string_normalize_nfc (Aerospike Expression Language (AEL) normalizeNFC()) returns that NFC value without writing the bin.

Normalize to NFC on write so stored values share one form. If mixed forms are already stored, normalize both sides of a comparison, or normalize the bin before comparing it with an NFC literal. Until the stored bytes share one form, comparator results are UTF-8 byte results.

// city holds NFD "café" (e + combining acute)
try (RecordStream rs = session.upsert(key)
.bin("city").normalizeNfc()
.execute()) {
rs.next().recordOrThrow();
}
// city now holds NFC "café" (precomposed é)
String exp = "$.city:STRING.normalizeNFC() == 'caf\u00e9'";
// true whichever form city is stored in

Invalid UTF-8 is a different problem from mixed NFC and NFD. Both NFC and NFD are valid UTF-8, so they pass encoding checks and still compare as different under eq. For bytes that are not valid UTF-8, see String operations and UTF-8 validation.

Next steps