Unicode Properties

Unicode Version 18.0

Unicode properties are standardized metadata attributes that define each character's behavior, rendering, and semantics for precise text processing.

Version Added
Long Name
Age

Version Added is the Unicode version in which a character was first encoded, used to track historical introduction and stability.

General Category
Long Name
General_Category

General Category is a classification system that assigns each character to one of several groups, such as letters, numbers, punctuation, or symbols.

Canonical Combining Class
Long Name
Canonical_Combining_Class

Canonical Combining Class is a numeric value that orders diacritical marks for reordering, ensuring consistent text rendering and normalization.

Bidirectional Class
Long Name
Bidi_Class

Bidirectional Class is a character property that defines how text aligns and orders within mixed directional content, such as Arabic and Latin.

Mirrored
Long Name
Bidi_Mirrored

Mirrored is a binary character property indicating that a glyph should be horizontally reversed in bidirectional text, like parentheses or arrows.

Bidirectional Control
Long Name
Bidi_Control

Bidirectional Control is a character property that determines how text direction is managed, enabling explicit overrides for correct display of mixed scripts.

Bidirectional Paired Bracket Type
Long Name
Bidi_Paired_Bracket_Type

Bidirectional Paired Bracket Type is a character attribute that classifies brackets as opening, closing, or none, aiding proper mirroring in bidirectional text.

Decomposition Type
Long Name
Decomposition_Type

Decomposition Type is a tag categorizing how a precomposed character splits into its component parts, like font, compat, or circle.

Composition Exclusion
Long Name
Composition_Exclusion

Composition Exclusion is a flag marking precomposed characters that should not be decomposed during normalization, preserving their distinct identity.

Full Composition Exclusion
Long Name
Full_Composition_Exclusion

Full Composition Exclusion is a list of precomposed characters that must not be decomposed, then recomposed, during normalization.

NFC Quick Check
Long Name
NFC_Quick_Check

NFC Quick Check is a binary flag that instantly determines if a string is already in Normalization Form C, avoiding full normalization.

NFD Quick Check
Long Name
NFD_Quick_Check

NFD Quick Check is a binary flag indicating whether a string is already in Normalization Form D, enabling fast rejection of non-normalized text.

NFKC Quick Check
Long Name
NFKC_Quick_Check

NFKC Quick Check is a binary value indicating whether a string is already in NFKC form, needs normalization, or may need it after further processing.

NFKD Quick Check
Long Name
NFKD_Quick_Check

NFKD Quick Check is a binary flag indicating whether a string is already in NFKD normal form, enabling fast validation without full decomposition.

Numeric Type
Long Name
Numeric_Type

Numeric Type is a classification that groups characters by their numerical value, including integers, fractions, and other numeric forms.

Joining Type
Long Name
Joining_Type

Joining Type is a character classification that defines how letters connect in cursive scripts, like Arabic, across joining, dual joining, or nonjoining forms.

Joining Group
Long Name
Joining_Group

Joining Group is a property that classifies Arabic letters by their contextual joining behavior, like Beh or Seen, shaping cursive text rendering.

Join Control
Long Name
Join_Control

Join Control is a property that marks formatting characters which prevent or enable cursive joining between adjacent letters, affecting text shaping.

Line Break
Long Name
Line_Break

Line Break is a text segmentation rule defining where lines can wrap, based on character classes like spaces, punctuation, and ideographs.

East Asian Width
Long Name
East_Asian_Width

East Asian Width is a classification that assigns each character to one of six width categories, aiding in text layout for East Asian scripts.

Uppercase
Long Name
Uppercase

Uppercase is a binary character property that defines whether a letter has a capital form, like A, B, or Ç.

Lowercase
Long Name
Lowercase

Lowercase is a binary character property that marks letters with lowercase semantics, enabling case mapping, caseless matching, and text normalization.

Other Uppercase
Long Name
Other_Uppercase

Other Uppercase is a character property identifying letters that are uppercase but lack a direct lowercase mapping, like Roman numerals or circled capitals.

Other Lowercase
Long Name
Other_Lowercase

Other Lowercase is a character property identifying lowercase letters that lack a corresponding uppercase form, such as modifier letters and phonetic symbols.

Case Ignorable
Long Name
Case_Ignorable

Case Ignorable is a binary character attribute marking letters that can be omitted when determining case folding or case-insensitive matching.

Cased
Long Name
Cased

Cased is a binary character property that marks letters with distinct uppercase and lowercase forms, enabling case mapping and case-insensitive matching.

Changes When Casefolded
Long Name
Changes_When_Casefolded

Changes When Casefolded is a binary flag that marks characters whose case folding, used for case-insensitive matching, alters their code point sequence.

Changes When Casemapped
Long Name
Changes_When_Casemapped

Changes When Casemapped is a binary flag indicating whether a character’s case mapping, like uppercase to lowercase, alters its code point value.

Changes When Lowercased
Long Name
Changes_When_Lowercased

Changes When Lowercased is a binary flag indicating whether a character’s lowercase mapping differs from its original code point.

Changes When NFKC Casefolded
Long Name
Changes_When_NFKC_Casefolded

Changes When NFKC Casefolded is a derived flag indicating if a character’s casefolded form changes after NFKC normalization.

Changes When Titlecased
Long Name
Changes_When_Titlecased

Changes When Titlecased is a binary attribute that marks characters whose titlecase mapping differs from their uppercase mapping, like digraphs or ligatures.

Changes When Uppercased
Long Name
Changes_When_Uppercased

Changes When Uppercased is a binary flag indicating whether a character’s uppercase mapping differs from its original code point.

Script
Long Name
Script

Script is a character property that groups codepoints by writing system, enabling multilingual text rendering, sorting, and detection.

Hangul Syllable Type
Long Name
Hangul_Syllable_Type

Hangul Syllable Type is a classification of each Korean character as leading, vowel, trailing, or syllable, enabling correct text processing and rendering.

Indic Syllabic Category
Long Name
Indic_Syllabic_Category

Indic Syllabic Category is a character classification that defines the phonetic and orthographic role of letters and signs in Indic scripts.

Indic Matra Category
Long Name
Indic_Matra_Category

Indic Matra Category is a character classification that groups vowel signs by their placement, like above, below, or to the side.

Indic Positional Category
Long Name
Indic_Positional_Category

Indic Positional Category is a value that classifies how a consonant's vowel sign attaches to its base letter, using positions like left, right, top, or bottom.

Indic Conjunct Break
Long Name
Indic_Conjunct_Break

Indic Conjunct Break is a character property that defines where ligatures can form or break between consonants in Indic scripts.

ID Start
Long Name
ID_Start

ID Start is the set of characters that can begin an identifier in programming languages, including letters, ideographs, and certain other symbols.

Other ID Start
Long Name
Other_ID_Start

Other ID Start is a character property that marks code points allowed to begin identifiers, like U+1885, beyond standard ID Start.

XID Start
Long Name
XID_Start

XID Start is a character class that permits the first character of an identifier, excluding those with Pattern_Syntax or Pattern_White_Space.

ID Continue
Long Name
ID_Continue

ID Continue is a character property indicating which letters, digits, underscores, and combining marks can appear after the first character in an identifier.

Other ID Continue
Long Name
Other_ID_Continue

Other ID Continue is a character property that permits certain symbols in identifiers beyond standard ID_Continue rules.

XID Continue
Long Name
XID_Continue

XID Continue is a character property that allows letters, digits, and certain symbols to be used within identifiers after the first character.

ID Compat Math Start
Long Name
ID_Compat_Math_Start

ID Compat Math Start is a character property that identifies mathematical symbols allowed to begin an identifier under Unicode's identifier compatibility rules.

ID Compat Math Continue
Long Name
ID_Compat_Math_Continue

ID Compat Math Continue is a character property marking symbols that may continue a mathematical identifier under compatibility rules.

Pattern Syntax
Long Name
Pattern_Syntax

Pattern Syntax is a character classification that marks punctuation and symbols safe to use in regular expression patterns without escaping.

Pattern White Space
Long Name
Pattern_White_Space

Pattern White Space is a binary character property that flags whitespace characters ignored during line breaking and pattern matching in regex.

Dash
Long Name
Dash

Dash is a binary character property marking whether a grapheme acts as a hyphen, minus, or other dash punctuation, aiding line breaking and text analysis.

Quotation Mark
Long Name
Quotation_Mark

Quotation Mark is a binary character property that identifies punctuation used to enclose direct speech or quotations, including paired and single forms.

Terminal Punctuation
Long Name
Terminal_Punctuation

Terminal Punctuation is a character property marking punctuation that ends a sentence, like periods, question marks, and exclamation points.

Sentence Terminal
Long Name
Sentence_Terminal

Sentence Terminal is a binary character property that marks punctuation typically ending a sentence, useful for text segmentation and processing.

Diacritic
Long Name
Diacritic

Diacritic is a binary character property marking letters that combine with base characters to modify their sound or meaning, like accents or cedillas.

Extender
Long Name
Extender

Extender is a character property marking letters or symbols that continue a word, like the macron in Māori or the Arabic tatweel.

Prepended Concatenation Mark
Long Name
Prepended_Concatenation_Mark

Prepended Concatenation Mark is a character class that attaches to a following symbol, like Arabic number signs, without breaking word boundaries in text.

Modifier Combining Mark
Long Name
Modifier_Combining_Mark

Modifier Combining Mark is a character class that attaches to a base letter, altering its pronunciation or tone without changing its core identity.

Soft Dotted
Long Name
Soft_Dotted

Soft Dotted is a character property marking letters like i and j whose dot disappears in diacritic combinations.

Alphabetic
Long Name
Alphabetic

Alphabetic is a binary character property that identifies letters and letter-like symbols, including those with derived alphabetic status across scripts.

Other Alphabetic
Long Name
Other_Alphabetic

Other Alphabetic is the set of characters used in words but not classed as letters, including digits, marks, and some symbols, aiding text segmentation.

Math
Long Name
Math

Math is a binary character property marking symbols used in mathematical notation, enabling their distinct handling in text processing and rendering.

Other Math
Long Name
Other_Math

Other Math is a character property marking symbols used in mathematical notation, excluding those already classified as math operators or separators.

Hex Digit
Long Name
Hex_Digit

Hex Digit is a binary character property that identifies the sixteen symbols, 0 through 9 and A through F, used in hexadecimal notation.

ASCII Hex Digit
Long Name
ASCII_Hex_Digit

ASCII Hex Digit is a binary character property that marks the 22 characters used in hexadecimal notation, covering digits 0–9 and letters A–F in both cases.

Default Ignorable Code Point
Long Name
Default_Ignorable_Code_Point

Default Ignorable Code Point is a character that should be hidden unless explicitly needed, like soft hyphens or zero-width spaces.

Other Default Ignorable Code Point
Long Name
Other_Default_Ignorable_Code_Point

Other Default Ignorable Code Point is a character property for invisible or formatting characters that should be ignored in rendering unless explicitly needed.

Logical Order Exception
Long Name
Logical_Order_Exception

Logical Order Exception is a rule where certain combining marks are stored before the base character despite being visually placed after it.

White Space
Long Name
White_Space

White Space is a binary character property that marks characters like spaces, tabs, and line breaks for separation, not content, in text processing.

Vertical Orientation
Long Name
Vertical_Orientation

Vertical Orientation is a text property that defines whether a character rotates, remains upright, or transforms in vertical writing systems.

Regional Indicator
Long Name
Regional_Indicator

Regional Indicator is a character category for two letter symbols representing country codes, used in pairs to form flag emojis.

Grapheme Base
Long Name
Grapheme_Base

Grapheme Base is a character class that marks where combining marks or spacing can attach in a grapheme cluster.

Grapheme Extend
Long Name
Grapheme_Extend

Grapheme Extend is a character class that includes marks and other signs which combine with a base character without breaking text into separate graphemes.

Other Grapheme Extend
Long Name
Other_Grapheme_Extend

Other Grapheme Extend is a character property marking symbols that combine with or modify preceding characters without breaking text into separate units.

Grapheme Cluster Break
Long Name
Grapheme_Cluster_Break

Grapheme Cluster Break is a rule that defines where text segments between user-perceived characters, like emoji with modifiers, can split.

Word Break
Long Name
Word_Break

Word Break is a categorization of characters into classes that defines where words can split across lines, aiding text segmentation and wrapping.

Sentence Break
Long Name
Sentence_Break

Sentence Break is a character classification that determines where sentences start and end, aiding in text segmentation for processing and display.

Ideographic
Long Name
Ideographic

Ideographic is a binary character property marking codepoints used in logographic writing systems like Chinese, Japanese, and Korean.

Unified Ideograph
Long Name
Unified_Ideograph

Unified Ideograph is a binary character property that marks codepoints designated as CJK unified ideographs, used for East Asian text processing.

IDS Binary Operator
Long Name
IDS_Binary_Operator

IDS Binary Operator is a character property flagging symbols used in ideographic descriptions to link components into a new complex ideograph.

IDS Trinary Operator
Long Name
IDS_Trinary_Operator

IDS Trinary Operator is a boolean value that indicates whether a character can act as a trinary operator in ideographic description sequences.

IDS Unary Operator
Long Name
IDS_Unary_Operator

IDS Unary Operator is a binary operator in ideographic descriptions, combining two characters to form a third, used in CJK character composition.

Radical
Long Name
Radical

Radical is a binary character property that marks code points used as CJK radical components in dictionary indexes and Kangxi radicals.

Deprecated
Long Name
Deprecated

Deprecated is a binary character property flagging code points that are discouraged from use, though they remain valid for interchange and processing.

Variation Selector
Long Name
Variation_Selector

Variation Selector is a value that alters a preceding character's glyph, enabling textual variants without changing its semantic meaning.

Noncharacter Code Point
Long Name
Noncharacter_Code_Point

Noncharacter Code Point is a code point permanently reserved for internal use, never assigned to any character, ensuring it is ignored in text interchange.

Emoji
Long Name
Emoji

Emoji is a binary character property that marks code points intended for emoji display, including sequences and modifiers, to enable consistent rendering.

Emoji Presentation
Long Name
Emoji_Presentation

Emoji Presentation is a binary property that determines whether a character defaults to a colorful emoji glyph or a monochrome text symbol.

Emoji Modifier
Long Name
Emoji_Modifier

Emoji Modifier is a set of code points that alter skin tone for certain emoji, enabling diverse representation in text.

Emoji Modifier Base
Long Name
Emoji_Modifier_Base

Emoji Modifier Base is a character property marking symbols like skin tone options, enabling subsequent emoji modifiers to alter their appearance.

Emoji Component
Long Name
Emoji_Component

Emoji Component is a property marking characters that serve as building blocks for emoji sequences, like skin tones, hair, and keycaps.

Extended Pictographic
Long Name
Extended_Pictographic

Extended Pictographic is a binary character property identifying emoji and similar symbols, aiding in text segmentation and emoji detection.