Logo Search packages:      
Sourcecode: icu version File versions

Collator Class Reference

#include <coll.h>

Inheritance diagram for Collator:

RuleBasedCollator

List of all members.


Detailed Description

The Collator class performs locale-sensitive string comparison.
You use this class to build searching and sorting routines for natural language text.
Important: The ICU collation service has been reimplemented in order to achieve better performance and UCA compliance. For details, see the collation design document.

Collator is an abstract base class. Subclasses implement specific collation strategies. One subclass, RuleBasedCollator, is currently provided and is applicable to a wide set of languages. Other subclasses may be created to handle more specialized needs.

Like other locale-sensitive classes, you can use the static factory method, createInstance, to obtain the appropriate Collator object for a given locale. You will only need to look at the subclasses of Collator if you need to understand the details of a particular collation strategy or if you need to modify that strategy.

The following example shows how to compare two strings using the Collator for the default locale. <blockquote>

 
 // Compare two strings in the default locale
 UErrorCode success = U_ZERO_ERROR;
 Collator* myCollator = Collator::createInstance(success);
 if (myCollator->compare("abc", "ABC") < 0)
   cout << "abc is less than ABC" << endl;
 else
   cout << "abc is greater than or equal to ABC" << endl;
</blockquote>

You can set a Collator's strength property to determine the level of difference considered significant in comparisons. Five strengths are provided: PRIMARY, SECONDARY, TERTIARY, QUATERNARY and IDENTICAL. The exact assignment of strengths to language features is locale dependant. For example, in Czech, "e" and "f" are considered primary differences, while "e" and "\u00EA" are secondary differences, "e" and "E" are tertiary differences and "e" and "e" are identical. The following shows how both case and accents could be ignored for US English. <blockquote>

 
 //Get the Collator for US English and set its strength to PRIMARY 
 UErrorCode success = U_ZERO_ERROR;
 Collator* usCollator = 
                            Collator::createInstance(Locale::US, success);
 usCollator->setStrength(Collator::PRIMARY);
 if (usCollator->compare("abc", "ABC") == 0)
   cout << 
 "'abc' and 'ABC' strings are equivalent with strength PRIMARY" << 
 endl;
</blockquote>

For comparing strings exactly once, the compare method provides the best performance. When sorting a list of strings however, it is generally necessary to compare each string multiple times. In this case, sort keys provide better performance. The getSortKey methods convert a string to a series of bytes that can be compared bitwise against other sort keys using strcmp(). Sort keys are written as zero-terminated byte strings. They consist of several substrings, one for each collation strength level, that are delimited by 0x01 bytes. If the string code points are appended for UCOL_IDENTICAL, then they are processed for correct code point order comparison and may contain 0x01 bytes but not zero bytes.

An older set of APIs returns a CollationKey object that wraps the sort key bytes instead of returning the bytes themselves. Its use is deprecated, but it is still available for compatibility with Java.

Note: Collators with different Locale, and CollationStrength settings will return different sort orders for the same set of strings. Locales have specific collation rules, and the way in which secondary and tertiary differences are taken into account, for example, will result in a different sorting order for same strings.

See also:
RuleBasedCollator

CollationKey

CollationElementIterator

Locale

Normalizer

Version:
2.0 11/15/01

Definition at line 154 of file coll.h.


Public Types

enum  ECollationStrength {
  PRIMARY = 0, SECONDARY = 1, TERTIARY = 2, QUATERNARY = 3,
  IDENTICAL = 15
}
enum  EComparisonResult { LESS = -1, EQUAL = 0, GREATER = 1 }

Public Member Functions

virtual Collatorclone (void) const =0
virtual EComparisonResult compare (const UChar *source, int32_t sourceLength, const UChar *target, int32_t targetLength) const =0
virtual EComparisonResult compare (const UnicodeString &source, const UnicodeString &target, int32_t length) const =0
virtual EComparisonResult compare (const UnicodeString &source, const UnicodeString &target) const =0
UBool equals (const UnicodeString &source, const UnicodeString &target) const
virtual UColAttributeValue getAttribute (UColAttribute attr, UErrorCode &status)=0
virtual CollationKeygetCollationKey (const UChar *source, int32_t sourceLength, CollationKey &key, UErrorCode &status) const =0
virtual CollationKeygetCollationKey (const UnicodeString &source, CollationKey &key, UErrorCode &status) const =0
virtual Normalizer::EMode getDecomposition (void) const =0
virtual UClassID getDynamicClassID (void) const =0
virtual const Locale getLocale (ULocDataLocaleType type, UErrorCode &status) const =0
virtual int32_t getSortKey (const UChar *source, int32_t sourceLength, uint8_t *result, int32_t resultLength) const =0
virtual int32_t getSortKey (const UnicodeString &source, uint8_t *result, int32_t resultLength) const =0
virtual ECollationStrength getStrength (void) const =0
virtual uint32_t getVariableTop (UErrorCode &status) const =0
virtual void getVersion (UVersionInfo info) const =0
UBool greater (const UnicodeString &source, const UnicodeString &target) const
UBool greaterOrEqual (const UnicodeString &source, const UnicodeString &target) const
virtual int32_t hashCode (void) const =0
virtual UBool operator!= (const Collator &other) const
virtual UBool operator== (const Collator &other) const
virtual CollatorsafeClone (void)=0
virtual void setAttribute (UColAttribute attr, UColAttributeValue value, UErrorCode &status)=0
virtual void setDecomposition (Normalizer::EMode mode)=0
virtual void setStrength (ECollationStrength newStrength)=0
virtual void setVariableTop (const uint32_t varTop, UErrorCode &status)=0
virtual uint32_t setVariableTop (const UnicodeString varTop, UErrorCode &status)=0
virtual uint32_t setVariableTop (const UChar *varTop, int32_t len, UErrorCode &status)=0
virtual ~Collator ()

Static Public Member Functions

static CollatorcreateInstance (const Locale &loc, UVersionInfo version, UErrorCode &err)
static CollatorcreateInstance (const Locale &loc, UErrorCode &err)
static CollatorcreateInstance (UErrorCode &err)
static const Locale * getAvailableLocales (int32_t &count)
static int32_t getBound (const uint8_t *source, int32_t sourceLength, UColBoundMode boundType, uint32_t noOfLevels, uint8_t *result, int32_t resultLength, UErrorCode &status)
static UnicodeStringgetDisplayName (const Locale &objectLocale, UnicodeString &name)
static UnicodeStringgetDisplayName (const Locale &objectLocale, const Locale &displayLocale, UnicodeString &name)

Protected Member Functions

 Collator (const Collator &other)
 Collator (UCollationStrength collationStrength, UNormalizationMode decompositionMode)
 Collator ()

The documentation for this class was generated from the following files:

Generated by  Doxygen 1.6.0   Back to index