Canonical Combining Class

Unicode Version 18.0

Canonical Combining Class is a numeric value assigned to each character that determines its position relative to a preceding base character during text normalization, specifically when combining marks are applied. It ranges from 0 to 254, where 0 means the character does not combine with others, and higher values dictate the order in which multiple diacritical marks stack above or below a base glyph. This ordering follows a fixed schema based on the mark’s type, such as above, below, or attached, ensuring that visually identical sequences of characters and marks produce a canonically equivalent normalized form, regardless of the order in which they were originally typed. This prevents ambiguity in rendering and searching across different text systems.

Not Reordered
Value
0
Abbr
NR
Count
309339

Not Reordered is the default combining class for base characters that do not combine with preceding marks.

Overlay
Value
1
Abbr
OV
Count
34

Overlay is a spacing mark that combines with a base character by drawing a line or stroke through it, altering its appearance.

Han Reading
Value
6
Abbr
HANR
Count
2

Han Reading is a method for determining how CJK characters combine with diacritics, aiding in the proper ordering and rendering of annotated Han text.

Nukta
Value
7
Abbr
NK
Count
27

Nukta is a combining class used for the diacritical mark that alters consonant pronunciation in Indic scripts, with a fixed numeric value.

Kana Voicing
Value
8
Abbr
KV
Count
2

Kana Voicing is a canonical combining class that marks the voicing mark in kana, affecting diacritic ordering in text.

Virama
Value
9
Abbr
VR
Count
69

Virama is a canonical combining class value of 9, indicating a spacing character that kills the inherent vowel of a preceding consonant.

CCC10
Value
10
Abbr
CCC10
Count
2

CCC10 is a placeholder value for a canonical combining class that is not formally assigned, indicating no specific class applies.

CCC11
Value
11
Abbr
CCC11
Count
1

CCC11 is assigned to Hebrew points, marking them as non-spacing marks that reorder with base characters.

CCC12
Value
12
Abbr
CCC12
Count
1

CCC12 is a fixed spacing value used to indicate that a character attaches above its base letter in certain diacritic systems.

CCC13
Value
13
Abbr
CCC13
Count
1

CCC13 is a combining class that indicates characters are positioned above the base letter, used for diacritical marks like acute accents.

CCC14
Value
14
Abbr
CCC14
Count
1

CCC14 is assigned to marks that attach above the base character, specifically for subscript or superscript positioning, with a combining class of 14.

CCC15
Value
15
Abbr
CCC15
Count
1

CCC15 is a fixed value assigned to certain combining marks, indicating they attach above or below their base character without reordering.

CCC16
Value
16
Abbr
CCC16
Count
1

CCC16 is a fixed value indicating that two adjacent characters combine without reordering, typically used for marks attached to a base character.

CCC17
Value
17
Abbr
CCC17
Count
1

CCC17 is a value assigned to characters like Greek ypogegrammeni, indicating they attach above the preceding base character in canonical decomposition.

CCC18
Value
18
Abbr
CCC18
Count
2

CCC18 is a canonical combining class value indicating a subscript, assigned to characters like certain Arabic and Hebrew marks.

CCC19
Value
19
Abbr
CCC19
Count
2

CCC19 is a canonical combining class value indicating a spacing split, meaning the mark is positioned between two base characters.

CCC20
Value
20
Abbr
CCC20
Count
1

CCC20 is the Canonical Combining Class value assigned to Hebrew points, indicating they combine with base characters without reordering.

CCC21
Value
21
Abbr
CCC21
Count
2

CCC21 is a fixed value for the canonical combining class, indicating the character is not reordered with others.

CCC22
Value
22
Abbr
CCC22
Count
1

CCC22 is a numeric value indicating a character combines with preceding marks in a specific way, often for Hebrew or Arabic pointing.

CCC23
Value
23
Abbr
CCC23
Count
1

CCC23 is a canonical combining class value indicating that a character is positioned above its base character, with a combining class of 23.

CCC24
Value
24
Abbr
CCC24
Count
1

CCC24 is a numeric value used to define the canonical ordering of combining marks, specifically for marks with a fixed position in text.

CCC25
Value
25
Abbr
CCC25
Count
1

CCC25 is a classification for characters that attach to a preceding base character in a specific, non-spacing way with a fixed combining order.

CCC26
Value
26
Abbr
CCC26
Count
1

CCC26 is a canonical combining class value indicating that a character is attached above the preceding base character.

CCC27
Value
27
Abbr
CCC27
Count
2

CCC27 is a combining class value indicating characters that form a distinct, indivisible unit in Arabic script, like medial forms, for canonical ordering.

CCC28
Value
28
Abbr
CCC28
Count
2

CCC28 is a value indicating that a character reorders with adjacent marks as a distinct class, typically used for Musical Symbol Combining flags.

CCC29
Value
29
Abbr
CCC29
Count
2

CCC29 is a canonical combining class indicating a non-spacing mark that attaches to the base character with a fixed, non-reordering priority.

CCC30
Value
30
Abbr
CCC30
Count
2

CCC30 is a fixed combining class for characters that must not reorder, such as Hebrew points and Arabic vowel marks.

CCC31
Value
31
Abbr
CCC31
Count
2

CCC31 is a fixed value indicating the character is a left-facing Hebrew vowel mark, with no reordering relative to adjacent marks.

CCC32
Value
32
Abbr
CCC32
Count
2

CCC32 is a numeric value indicating that a character attaches directly above the preceding base character in a stacked diacritical mark.

CCC33
Value
33
Abbr
CCC33
Count
1

CCC33 is a fixed value indicating that a character combines with the preceding base character using the “Above” attachment position.

CCC34
Value
34
Abbr
CCC34
Count
1

CCC34 is a specific Canonical Combining Class value indicating that a character is attached above with 222 degrees, used for diacritical mark ordering.

CCC35
Value
35
Abbr
CCC35
Count
1

CCC35 is a fixed value indicating a character is a starting mark that reorders before base characters in canonical decomposition.

CCC36
Value
36
Abbr
CCC36
Count
1

CCC36 is a canonical combining class value indicating that a character is a subscript, positioned below the preceding base character.

CCC84
Value
84
Abbr
CCC84
Count
1

CCC84 is a canonical combining class value indicating a character that attaches above or below its base, with a fixed, non-reordering position.

CCC91
Value
91
Abbr
CCC91
Count
1

CCC91 is assigned to Hebrew presentation forms, indicating they decompose with a non-spacing mark that must attach to the preceding base character.

CCC103
Value
103
Abbr
CCC103
Count
2

CCC103 is a fixed value indicating the character's diacritic attaches above the base with a specific combining position, used for precise text reordering.

CCC107
Value
107
Abbr
CCC107
Count
4

CCC107 is a Canonical Combining Class value that marks the addition of a dot above a character, typically for Umlaut or diaeresis marks.

CCC118
Value
118
Abbr
CCC118
Count
2

CCC118 is a canonical combining class value indicating that a character is attached above the preceding base character with a specific fixed position.

CCC122
Value
122
Abbr
CCC122
Count
4

CCC122 is assigned to Hebrew points, marking them as above-base characters that reorder after base letters in canonical decomposition.

CCC129
Value
129
Abbr
CCC129
Count
1

CCC129 is a value marking characters that must not be reordered during canonical composition, like Arabic diacritics.

CCC130
Value
130
Abbr
CCC130
Count
6

CCC130 is a value indicating a character is attached above the base, with a combining class of 130.

CCC132
Value
132
Abbr
CCC132
Count
1

CCC132 is a subclass that includes characters like Arabic vowels and marks, which attach to base letters during text shaping.

CCC133
Value
133
Abbr
CCC133
Count
0

CCC133 is assigned to dotted circles and marks that must not reorder, blocking all canonical combining with preceding characters.

Attached Below Left
Value
200
Abbr
ATBL
Count
0

Attached Below Left is a numeric value (240) indicating a mark sits directly beneath its base character, left-aligned.

Attached Below
Value
202
Abbr
ATB
Count
5

Attached Below is a numeric value indicating that a character visually combines with the mark directly beneath it.

Attached Above
Value
214
Abbr
ATA
Count
1

Attached Above is a numeric value from 0 to 240 indicating how a mark combines with its base character.

Attached Above Right
Value
216
Abbr
ATAR
Count
15

Attached Above Right is a combining class value indicating a mark positioned above and to the right of its base character.

Below Left
Value
218
Abbr
BL
Count
2

Below Left is a numeric priority level used to order diacritical marks, ensuring correct placement when combining characters in text rendering.

Below
Value
220
Abbr
B
Count
197

Below is a numeric code, typically 0, indicating that a character does not combine with preceding characters in a specific way.

Below Right
Value
222
Abbr
BR
Count
4

Below Right is a combining class of 220, indicating the mark attaches below and to the right of the base character.

Left
Value
224
Abbr
L
Count
2

Left is a numeric code indicating diacritic positioning relative to the base character, used to reorder marks during text normalization.

Right
Value
226
Abbr
R
Count
1

Right is a numeric code from 0 to 240 indicating the relative position for diacritic attachment, with higher values meaning closer to the base character.

Above Left
Value
228
Abbr
AL
Count
5

Above Left is a combining class indicating characters that attach above and to the left of their base, used for diacritical marks.

Above
Value
230
Abbr
A
Count
559

Above is a numeric value indicating that a character combines with a preceding base character by placement above it, affecting diacritic ordering.

Above Right
Value
232
Abbr
AR
Count
7

Above Right is a canonical combining class value indicating that a mark is positioned above the preceding base character, slightly to its right.

Double Below
Value
233
Abbr
DB
Count
4

Double Below is a value indicating a canonical combining class of 202, used for marks attached beneath a base character.

Double Above
Value
234
Abbr
DA
Count
6

Double Above is a canonical combining class value indicating a diacritic that attaches above the base character, positioned high, with numeric value 233.

Iota Subscript
Value
240
Abbr
IS
Count
1

Iota Subscript is a numeric value of 240, used to adjust the positioning of diacritical marks in ancient Greek text.