Unicode Properties
Unicode properties are standardized metadata attributes that define each character's behavior, rendering, and semantics for precise text processing.
- Long Name
- Age
Version Added is the Unicode version in which a character was first encoded, used to track historical introduction and stability.
- Long Name
- General_Category
General Category is a classification system that assigns each character to one of several groups, such as letters, numbers, punctuation, or symbols.
- Long Name
- Canonical_Combining_Class
Canonical Combining Class is a numeric value that orders diacritical marks for reordering, ensuring consistent text rendering and normalization.
- Long Name
- Bidi_Class
Bidirectional Class is a character property that defines how text aligns and orders within mixed directional content, such as Arabic and Latin.
- Long Name
- Bidi_Mirrored
Mirrored is a binary character property indicating that a glyph should be horizontally reversed in bidirectional text, like parentheses or arrows.
- Long Name
- Bidi_Control
Bidirectional Control is a character property that determines how text direction is managed, enabling explicit overrides for correct display of mixed scripts.
- Long Name
- Bidi_Paired_Bracket_Type
Bidirectional Paired Bracket Type is a character attribute that classifies brackets as opening, closing, or none, aiding proper mirroring in bidirectional text.
- Long Name
- Decomposition_Type
Decomposition Type is a tag categorizing how a precomposed character splits into its component parts, like font, compat, or circle.
- Long Name
- Composition_Exclusion
Composition Exclusion is a flag marking precomposed characters that should not be decomposed during normalization, preserving their distinct identity.
- Long Name
- Full_Composition_Exclusion
Full Composition Exclusion is a list of precomposed characters that must not be decomposed, then recomposed, during normalization.
- Long Name
- NFC_Quick_Check
NFC Quick Check is a binary flag that instantly determines if a string is already in Normalization Form C, avoiding full normalization.
- Long Name
- NFD_Quick_Check
NFD Quick Check is a binary flag indicating whether a string is already in Normalization Form D, enabling fast rejection of non-normalized text.
- Long Name
- NFKC_Quick_Check
NFKC Quick Check is a binary value indicating whether a string is already in NFKC form, needs normalization, or may need it after further processing.
- Long Name
- NFKD_Quick_Check
NFKD Quick Check is a binary flag indicating whether a string is already in NFKD normal form, enabling fast validation without full decomposition.
- Long Name
- Numeric_Type
Numeric Type is a classification that groups characters by their numerical value, including integers, fractions, and other numeric forms.
- Long Name
- Joining_Type
Joining Type is a character classification that defines how letters connect in cursive scripts, like Arabic, across joining, dual joining, or nonjoining forms.
- Long Name
- Joining_Group
Joining Group is a property that classifies Arabic letters by their contextual joining behavior, like Beh or Seen, shaping cursive text rendering.
- Long Name
- Join_Control
Join Control is a property that marks formatting characters which prevent or enable cursive joining between adjacent letters, affecting text shaping.
- Long Name
- Line_Break
Line Break is a text segmentation rule defining where lines can wrap, based on character classes like spaces, punctuation, and ideographs.
- Long Name
- East_Asian_Width
East Asian Width is a classification that assigns each character to one of six width categories, aiding in text layout for East Asian scripts.
- Long Name
- Uppercase
Uppercase is a binary character property that defines whether a letter has a capital form, like A, B, or Ç.
- Long Name
- Lowercase
Lowercase is a binary character property that marks letters with lowercase semantics, enabling case mapping, caseless matching, and text normalization.
- Long Name
- Other_Uppercase
Other Uppercase is a character property identifying letters that are uppercase but lack a direct lowercase mapping, like Roman numerals or circled capitals.
- Long Name
- Other_Lowercase
Other Lowercase is a character property identifying lowercase letters that lack a corresponding uppercase form, such as modifier letters and phonetic symbols.
- Long Name
- Case_Ignorable
Case Ignorable is a binary character attribute marking letters that can be omitted when determining case folding or case-insensitive matching.
- Long Name
- Cased
Cased is a binary character property that marks letters with distinct uppercase and lowercase forms, enabling case mapping and case-insensitive matching.
- Long Name
- Changes_When_Casefolded
Changes When Casefolded is a binary flag that marks characters whose case folding, used for case-insensitive matching, alters their code point sequence.
- Long Name
- Changes_When_Casemapped
Changes When Casemapped is a binary flag indicating whether a character’s case mapping, like uppercase to lowercase, alters its code point value.
- Long Name
- Changes_When_Lowercased
Changes When Lowercased is a binary flag indicating whether a character’s lowercase mapping differs from its original code point.
- Long Name
- Changes_When_NFKC_Casefolded
Changes When NFKC Casefolded is a derived flag indicating if a character’s casefolded form changes after NFKC normalization.
- Long Name
- Changes_When_Titlecased
Changes When Titlecased is a binary attribute that marks characters whose titlecase mapping differs from their uppercase mapping, like digraphs or ligatures.
- Long Name
- Changes_When_Uppercased
Changes When Uppercased is a binary flag indicating whether a character’s uppercase mapping differs from its original code point.
- Long Name
- Script
Script is a character property that groups codepoints by writing system, enabling multilingual text rendering, sorting, and detection.
- Long Name
- Hangul_Syllable_Type
Hangul Syllable Type is a classification of each Korean character as leading, vowel, trailing, or syllable, enabling correct text processing and rendering.
- Long Name
- Indic_Syllabic_Category
Indic Syllabic Category is a character classification that defines the phonetic and orthographic role of letters and signs in Indic scripts.
- Long Name
- Indic_Matra_Category
Indic Matra Category is a character classification that groups vowel signs by their placement, like above, below, or to the side.
- Long Name
- Indic_Positional_Category
Indic Positional Category is a value that classifies how a consonant's vowel sign attaches to its base letter, using positions like left, right, top, or bottom.
- Long Name
- Indic_Conjunct_Break
Indic Conjunct Break is a character property that defines where ligatures can form or break between consonants in Indic scripts.
- Long Name
- ID_Start
ID Start is the set of characters that can begin an identifier in programming languages, including letters, ideographs, and certain other symbols.
- Long Name
- Other_ID_Start
Other ID Start is a character property that marks code points allowed to begin identifiers, like U+1885, beyond standard ID Start.
- Long Name
- XID_Start
XID Start is a character class that permits the first character of an identifier, excluding those with Pattern_Syntax or Pattern_White_Space.
- Long Name
- ID_Continue
ID Continue is a character property indicating which letters, digits, underscores, and combining marks can appear after the first character in an identifier.
- Long Name
- Other_ID_Continue
Other ID Continue is a character property that permits certain symbols in identifiers beyond standard ID_Continue rules.
- Long Name
- XID_Continue
XID Continue is a character property that allows letters, digits, and certain symbols to be used within identifiers after the first character.
- Long Name
- ID_Compat_Math_Start
ID Compat Math Start is a character property that identifies mathematical symbols allowed to begin an identifier under Unicode's identifier compatibility rules.
- Long Name
- ID_Compat_Math_Continue
ID Compat Math Continue is a character property marking symbols that may continue a mathematical identifier under compatibility rules.
- Long Name
- Pattern_Syntax
Pattern Syntax is a character classification that marks punctuation and symbols safe to use in regular expression patterns without escaping.
- Long Name
- Pattern_White_Space
Pattern White Space is a binary character property that flags whitespace characters ignored during line breaking and pattern matching in regex.
- Long Name
- Dash
Dash is a binary character property marking whether a grapheme acts as a hyphen, minus, or other dash punctuation, aiding line breaking and text analysis.
- Long Name
- Quotation_Mark
Quotation Mark is a binary character property that identifies punctuation used to enclose direct speech or quotations, including paired and single forms.
- Long Name
- Terminal_Punctuation
Terminal Punctuation is a character property marking punctuation that ends a sentence, like periods, question marks, and exclamation points.
- Long Name
- Sentence_Terminal
Sentence Terminal is a binary character property that marks punctuation typically ending a sentence, useful for text segmentation and processing.
- Long Name
- Diacritic
Diacritic is a binary character property marking letters that combine with base characters to modify their sound or meaning, like accents or cedillas.
- Long Name
- Extender
Extender is a character property marking letters or symbols that continue a word, like the macron in Māori or the Arabic tatweel.
- Long Name
- Prepended_Concatenation_Mark
Prepended Concatenation Mark is a character class that attaches to a following symbol, like Arabic number signs, without breaking word boundaries in text.
- Long Name
- Modifier_Combining_Mark
Modifier Combining Mark is a character class that attaches to a base letter, altering its pronunciation or tone without changing its core identity.
- Long Name
- Soft_Dotted
Soft Dotted is a character property marking letters like i and j whose dot disappears in diacritic combinations.
- Long Name
- Alphabetic
Alphabetic is a binary character property that identifies letters and letter-like symbols, including those with derived alphabetic status across scripts.
- Long Name
- Other_Alphabetic
Other Alphabetic is the set of characters used in words but not classed as letters, including digits, marks, and some symbols, aiding text segmentation.
- Long Name
- Math
Math is a binary character property marking symbols used in mathematical notation, enabling their distinct handling in text processing and rendering.
- Long Name
- Other_Math
Other Math is a character property marking symbols used in mathematical notation, excluding those already classified as math operators or separators.
- Long Name
- Hex_Digit
Hex Digit is a binary character property that identifies the sixteen symbols, 0 through 9 and A through F, used in hexadecimal notation.
- Long Name
- ASCII_Hex_Digit
ASCII Hex Digit is a binary character property that marks the 22 characters used in hexadecimal notation, covering digits 0–9 and letters A–F in both cases.
- Long Name
- Default_Ignorable_Code_Point
Default Ignorable Code Point is a character that should be hidden unless explicitly needed, like soft hyphens or zero-width spaces.
- Long Name
- Other_Default_Ignorable_Code_Point
Other Default Ignorable Code Point is a character property for invisible or formatting characters that should be ignored in rendering unless explicitly needed.
- Long Name
- Logical_Order_Exception
Logical Order Exception is a rule where certain combining marks are stored before the base character despite being visually placed after it.
- Long Name
- White_Space
White Space is a binary character property that marks characters like spaces, tabs, and line breaks for separation, not content, in text processing.
- Long Name
- Vertical_Orientation
Vertical Orientation is a text property that defines whether a character rotates, remains upright, or transforms in vertical writing systems.
- Long Name
- Regional_Indicator
Regional Indicator is a character category for two letter symbols representing country codes, used in pairs to form flag emojis.
- Long Name
- Grapheme_Base
Grapheme Base is a character class that marks where combining marks or spacing can attach in a grapheme cluster.
- Long Name
- Grapheme_Extend
Grapheme Extend is a character class that includes marks and other signs which combine with a base character without breaking text into separate graphemes.
- Long Name
- Other_Grapheme_Extend
Other Grapheme Extend is a character property marking symbols that combine with or modify preceding characters without breaking text into separate units.
- Long Name
- Grapheme_Cluster_Break
Grapheme Cluster Break is a rule that defines where text segments between user-perceived characters, like emoji with modifiers, can split.
- Long Name
- Word_Break
Word Break is a categorization of characters into classes that defines where words can split across lines, aiding text segmentation and wrapping.
- Long Name
- Sentence_Break
Sentence Break is a character classification that determines where sentences start and end, aiding in text segmentation for processing and display.
- Long Name
- Ideographic
Ideographic is a binary character property marking codepoints used in logographic writing systems like Chinese, Japanese, and Korean.
- Long Name
- Unified_Ideograph
Unified Ideograph is a binary character property that marks codepoints designated as CJK unified ideographs, used for East Asian text processing.
- Long Name
- IDS_Binary_Operator
IDS Binary Operator is a character property flagging symbols used in ideographic descriptions to link components into a new complex ideograph.
- Long Name
- IDS_Trinary_Operator
IDS Trinary Operator is a boolean value that indicates whether a character can act as a trinary operator in ideographic description sequences.
- Long Name
- IDS_Unary_Operator
IDS Unary Operator is a binary operator in ideographic descriptions, combining two characters to form a third, used in CJK character composition.
- Long Name
- Radical
Radical is a binary character property that marks code points used as CJK radical components in dictionary indexes and Kangxi radicals.
- Long Name
- Deprecated
Deprecated is a binary character property flagging code points that are discouraged from use, though they remain valid for interchange and processing.
- Long Name
- Variation_Selector
Variation Selector is a value that alters a preceding character's glyph, enabling textual variants without changing its semantic meaning.
- Long Name
- Noncharacter_Code_Point
Noncharacter Code Point is a code point permanently reserved for internal use, never assigned to any character, ensuring it is ignored in text interchange.
- Long Name
- Emoji
Emoji is a binary character property that marks code points intended for emoji display, including sequences and modifiers, to enable consistent rendering.
- Long Name
- Emoji_Presentation
Emoji Presentation is a binary property that determines whether a character defaults to a colorful emoji glyph or a monochrome text symbol.
- Long Name
- Emoji_Modifier
Emoji Modifier is a set of code points that alter skin tone for certain emoji, enabling diverse representation in text.
- Long Name
- Emoji_Modifier_Base
Emoji Modifier Base is a character property marking symbols like skin tone options, enabling subsequent emoji modifiers to alter their appearance.
- Long Name
- Emoji_Component
Emoji Component is a property marking characters that serve as building blocks for emoji sequences, like skin tones, hair, and keycaps.
- Long Name
- Extended_Pictographic
Extended Pictographic is a binary character property identifying emoji and similar symbols, aiding in text segmentation and emoji detection.