Collator Class Reference

#include <coll.h>

Detailed Description

The Collator class performs locale-sensitive string comparison.
You use this class to build searching and sorting routines for natural language text.
Important: The ICU collation service has been reimplemented in order to achieve better performance and UCA compliance. For details, see the collation design document.

Collator is an abstract base class. Subclasses implement specific collation strategies. One subclass, RuleBasedCollator, is currently provided and is applicable to a wide set of languages. Other subclasses may be created to handle more specialized needs.

Like other locale-sensitive classes, you can use the static factory method, createInstance, to obtain the appropriate Collator object for a given locale. You will only need to look at the subclasses of Collator if you need to understand the details of a particular collation strategy or if you need to modify that strategy.

The following example shows how to compare two strings using the Collator for the default locale.

 // Compare two strings in the default locale
 UErrorCode success = U_ZERO_ERROR;
 Collator* myCollator = Collator::createInstance(success);
 if (myCollator->compare("abc", "ABC") < 0)
   cout << "abc is less than ABC" << endl;
   cout << "abc is greater than or equal to ABC" << endl;

You can set a Collator's strength property to determine the level of difference considered significant in comparisons. Five strengths are provided: PRIMARY, SECONDARY, TERTIARY, QUATERNARY and IDENTICAL. The exact assignment of strengths to language features is locale dependant. For example, in Czech, "e" and "f" are considered primary differences, while "e" and "\u00EA" are secondary differences, "e" and "E" are tertiary differences and "e" and "e" are identical. The following shows how both case and accents could be ignored for US English.

 //Get the Collator for US English and set its strength to PRIMARY
 UErrorCode success = U_ZERO_ERROR;
 Collator* usCollator = Collator::createInstance(Locale::US, success);
 if (usCollator->compare("abc", "ABC") == 0)
     cout << "'abc' and 'ABC' strings are equivalent with strength PRIMARY" << endl;

For comparing strings exactly once, the compare method provides the best performance. When sorting a list of strings however, it is generally necessary to compare each string multiple times. In this case, sort keys provide better performance. The getSortKey methods convert a string to a series of bytes that can be compared bitwise against other sort keys using strcmp(). Sort keys are written as zero-terminated byte strings. They consist of several substrings, one for each collation strength level, that are delimited by 0x01 bytes. If the string code points are appended for UCOL_IDENTICAL, then they are processed for correct code point order comparison and may contain 0x01 bytes but not zero bytes.

An older set of APIs returns a CollationKey object that wraps the sort key bytes instead of returning the bytes themselves. Its use is deprecated, but it is still available for compatibility with Java.

Note: Collators with different Locale, and CollationStrength settings will return different sort orders for the same set of strings. Locales have specific collation rules, and the way in which secondary and tertiary differences are taken into account, for example, will result in a different sorting order for same strings.

Public Types

enum  ECollationStrength {
enum  EComparisonResult { LESS = -1, EQUAL = 0, GREATER = 1 }

Public Member Functions

virtual Collatorclone (void) const =0
virtual UCollationResult compare (UCharIterator &sIter, UCharIterator &tIter, UErrorCode &status) const
virtual UCollationResult compare (const UChar *source, int32_t sourceLength, const UChar *target, int32_t targetLength, UErrorCode &status) const =0
virtual EComparisonResult compare (const UChar *source, int32_t sourceLength, const UChar *target, int32_t targetLength) const
virtual UCollationResult compare (const UnicodeString &source, const UnicodeString &target, int32_t length, UErrorCode &status) const =0
virtual EComparisonResult compare (const UnicodeString &source, const UnicodeString &target, int32_t length) const
virtual UCollationResult compare (const UnicodeString &source, const UnicodeString &target, UErrorCode &status) const =0
virtual EComparisonResult compare (const UnicodeString &source, const UnicodeString &target) const
virtual UCollationResult compareUTF8 (const StringPiece &source, const StringPiece &target, UErrorCode &status) const
UBool equals (const UnicodeString &source, const UnicodeString &target) const
virtual UColAttributeValue getAttribute (UColAttribute attr, UErrorCode &status)=0
virtual CollationKeygetCollationKey (const UChar *source, int32_t sourceLength, CollationKey &key, UErrorCode &status) const =0
virtual CollationKeygetCollationKey (const UnicodeString &source, CollationKey &key, UErrorCode &status) const =0
virtual UClassID getDynamicClassID (void) const =0
virtual const Locale getLocale (ULocDataLocaleType type, UErrorCode &status) const =0
virtual int32_t getSortKey (const UChar *source, int32_t sourceLength, uint8_t *result, int32_t resultLength) const =0
virtual int32_t getSortKey (const UnicodeString &source, uint8_t *result, int32_t resultLength) const =0
virtual ECollationStrength getStrength (void) const =0
virtual UnicodeSetgetTailoredSet (UErrorCode &status) const
virtual uint32_t getVariableTop (UErrorCode &status) const =0
virtual void getVersion (UVersionInfo info) const =0
UBool greater (const UnicodeString &source, const UnicodeString &target) const
UBool greaterOrEqual (const UnicodeString &source, const UnicodeString &target) const
virtual int32_t hashCode (void) const =0
virtual UBool operator!= (const Collator &other) const
virtual UBool operator== (const Collator &other) const
virtual CollatorsafeClone (void)=0
virtual void setAttribute (UColAttribute attr, UColAttributeValue value, UErrorCode &status)=0
virtual void setStrength (ECollationStrength newStrength)=0
virtual void setVariableTop (const uint32_t varTop, UErrorCode &status)=0
virtual uint32_t setVariableTop (const UnicodeString varTop, UErrorCode &status)=0
virtual uint32_t setVariableTop (const UChar *varTop, int32_t len, UErrorCode &status)=0
virtual ~Collator ()

Static Public Member Functions

static Collator *U_EXPORT2 createInstance (const Locale &loc, UErrorCode &err)
static Collator *U_EXPORT2 createInstance (UErrorCode &err)
static UCollatorcreateUCollator (const char *loc, UErrorCode *status)
static StringEnumeration *U_EXPORT2 getAvailableLocales (void)
static const Locale *U_EXPORT2 getAvailableLocales (int32_t &count)
static int32_t U_EXPORT2 getBound (const uint8_t *source, int32_t sourceLength, UColBoundMode boundType, uint32_t noOfLevels, uint8_t *result, int32_t resultLength, UErrorCode &status)
static UnicodeString &U_EXPORT2 getDisplayName (const Locale &objectLocale, UnicodeString &name)
static UnicodeString &U_EXPORT2 getDisplayName (const Locale &objectLocale, const Locale &displayLocale, UnicodeString &name)
static Locale U_EXPORT2 getFunctionalEquivalent (const char *keyword, const Locale &locale, UBool &isAvailable, UErrorCode &status)
static StringEnumeration *U_EXPORT2 getKeywords (UErrorCode &status)
static StringEnumeration *U_EXPORT2 getKeywordValues (const char *keyword, UErrorCode &status)
static StringEnumeration *U_EXPORT2 getKeywordValuesForLocale (const char *keyword, const Locale &locale, UBool commonlyUsed, UErrorCode &status)
static void U_EXPORT2 operator delete (void *, void *) U_NO_THROW
static void U_EXPORT2 operator delete (void *p) U_NO_THROW
static void U_EXPORT2 operator delete[] (void *p) U_NO_THROW
static void *U_EXPORT2 operator new (size_t, void *ptr) U_NO_THROW
static void *U_EXPORT2 operator new (size_t size) U_NO_THROW
static void *U_EXPORT2 operator new[] (size_t size) U_NO_THROW
static URegistryKey U_EXPORT2 registerFactory (CollatorFactory *toAdopt, UErrorCode &status)
static URegistryKey U_EXPORT2 registerInstance (Collator *toAdopt, const Locale &locale, UErrorCode &status)
static UBool U_EXPORT2 unregister (URegistryKey key, UErrorCode &status)

Protected Member Functions

 Collator (const Collator &other)
 Collator (UCollationStrength collationStrength, UNormalizationMode decompositionMode)
 Collator ()
virtual void setLocales (const Locale &requestedLocale, const Locale &validLocale, const Locale &actualLocale)

Private Member Functions

Collatoroperator= (const Collator &other)

Static Private Member Functions

static CollatormakeInstance (const Locale &desiredLocale, UErrorCode &status)


class CFactory
class ICUCollatorFactory
class ICUCollatorService
class SimpleCFactory

