Unicode Blocks

Unicode Version 18.0

Unicode blocks are contiguous ranges of code points that group together characters sharing a common script, set of symbols, or thematic purpose for organized encoding.

Basic Latin
Abbr
ASCII
Range
0000 - 007F
Count
128

Basic Latin is the foundational character set covering ASCII letters, digits, punctuation, and controls, ensuring universal text compatibility.

Latin-1 Supplement
Abbr
Latin_1_Sup
Range
0080 - 00FF
Count
128

Latin-1 Supplement is the part of Unicode covering Western European letters, punctuation, and symbols, plus C1 control characters, extending ASCII’s coverage.

Latin Extended-A
Abbr
Latin_Ext_A
Range
0100 - 017F
Count
128

Latin Extended-A is a set of precomposed letters with diacritics, primarily for Eastern European and Vietnamese languages, enabling full text representation.

Latin Extended-B
Abbr
Latin_Ext_B
Range
0180 - 024F
Count
208

Latin Extended-B is a set of additional Latin letters and symbols, covering phonetic, historical, and African orthographic needs.

IPA Extensions
Abbr
IPA_Ext
Range
0250 - 02AF
Count
96

IPA Extensions is a set of additional phonetic symbols that supplements the International Phonetic Alphabet for representing speech sounds.

Spacing Modifier Letters
Abbr
Modifier_Letters
Range
02B0 - 02FF
Count
80

Spacing Modifier Letters is a set of phonetic symbols that adjust the meaning of preceding letters, including aspiration, palatalization, and tone marks.

Combining Diacritical Marks
Abbr
Diacriticals
Range
0300 - 036F
Count
112

Combining Diacritical Marks is a set of characters that attach to preceding letters, enabling accents and phonetic distinctions across many writing systems.

Greek and Coptic
Abbr
Greek
Range
0370 - 03FF
Count
135

Greek and Coptic is the Unicode segment housing letters for modern Greek, ancient Greek, and Coptic, plus punctuation and variant forms.

Cyrillic
Abbr
Cyrillic
Range
0400 - 04FF
Count
256

Cyrillic is a script block covering standard letters for Russian, Ukrainian, Bulgarian, and other Slavic languages, plus historic and non-Slavic extensions.

Cyrillic Supplement
Abbr
Cyrillic_Sup
Range
0500 - 052F
Count
48

Cyrillic Supplement is a set of letters for minority languages, plus historic forms, extending the main Cyrillic script.

Armenian
Abbr
Armenian
Range
0530 - 058F
Count
94

Armenian is a script block for writing the Armenian language, including its alphabet, punctuation, and ligatures, used across Armenia and the diaspora.

Hebrew
Abbr
Hebrew
Range
0590 - 05FF
Count
90

Hebrew is a contiguous set of characters for writing Hebrew, Yiddish, and Ladino, including letters, niqqud vowel marks, and cantillation signs.

Arabic
Abbr
Arabic
Range
0600 - 06FF
Count
256

Arabic is a script block encoding letters, diacritics, and punctuation for writing Arabic, Persian, Urdu, and other languages using Arabic-derived scripts.

Syriac
Abbr
Syriac
Range
0700 - 074F
Count
77

Syriac is a script block for writing the Syriac language, including its classical, Estrangela, and Eastern variants, plus diacritics and punctuation.

Arabic Supplement
Abbr
Arabic_Sup
Range
0750 - 077F
Count
48

Arabic Supplement is a set of extra letters for early Quranic and African languages, filling gaps in standard Arabic script.

Thaana
Abbr
Thaana
Range
0780 - 07BF
Count
50

Thaana is a script used for writing the Dhivehi language of the Maldives, featuring letters with right-to-left orientation and diacritical vowel marks.

NKo
Abbr
NKo
Range
07C0 - 07FF
Count
62

NKo is a script used for the Manding languages of West Africa, added to encode text in a single, unified writing system.

Samaritan
Abbr
Samaritan
Range
0800 - 083F
Count
61

Samaritan is a script used for writing the Samaritan Hebrew and Aramaic liturgical texts, including the Pentateuch and prayers.

Mandaic
Abbr
Mandaic
Range
0840 - 085F
Count
29

Mandaic is a script block used for writing the Mandaean language, containing letters and diacritics in a right-to-left direction.

Syriac Supplement
Abbr
Syriac_Sup
Range
0860 - 086F
Count
11

Syriac Supplement is a set of nine characters adding letters for Sogdian and additional Syriac orthography needs.

Arabic Extended-B
Abbr
Arabic_Ext_B
Range
0870 - 089F
Count
43

Arabic Extended-B is a set of Arabic script characters covering Quranic annotation marks, added letters for African languages, and editorial symbols.

Arabic Extended-A
Abbr
Arabic_Ext_A
Range
08A0 - 08FF
Count
96

Arabic Extended-A is a character set for additional letters and signs used in Arabic-based orthographies across Africa and Asia, plus Quranic annotation marks.

Devanagari
Abbr
Devanagari
Range
0900 - 097F
Count
128

Devanagari is a script used for Hindi, Marathi, Nepali, and Sanskrit, containing letters, vowels, diacritics, and punctuation marks.

Bengali
Abbr
Bengali
Range
0980 - 09FF
Count
96

Bengali is a script block encoding the Bengali-Assamese alphabet, used for Bangla, Assamese, and related languages, plus digits and diacritics.

Gurmukhi
Abbr
Gurmukhi
Range
0A00 - 0A7F
Count
80

Gurmukhi is a script used for writing Punjabi, containing letters, vowel signs, and punctuation marks integral to the language's religious and literary texts.

Gujarati
Abbr
Gujarati
Range
0A80 - 0AFF
Count
91

Gujarati is a script used for writing the Gujarati language, covering letters, digits, and diacritical marks essential for its phonetics.

Oriya
Abbr
Oriya
Range
0B00 - 0B7F
Count
93

Oriya is a script block for writing the Odia language, containing vowel signs, consonants, digits, and punctuation marks.

Tamil
Abbr
Tamil
Range
0B80 - 0BFF
Count
72

Tamil is an abugida script used for writing the Tamil language, encoded with characters for vowels, consonants, and diacritical marks, plus numerals and signs.

Telugu
Abbr
Telugu
Range
0C00 - 0C7F
Count
101

Telugu is a script used for writing the Telugu language, containing letters, vowels, diacritics, digits, and punctuation marks.

Kannada
Abbr
Kannada
Range
0C80 - 0CFF
Count
92

Kannada is a script block used for writing the Kannada language, including its vowels, consonants, and diacritical marks.

Malayalam
Abbr
Malayalam
Range
0D00 - 0D7F
Count
118

Malayalam is a script block for writing the Dravidian language of Kerala, including vowels, consonants, and conjunct forms.

Sinhala
Abbr
Sinhala
Range
0D80 - 0DFF
Count
91

Sinhala is the script used for writing the Sinhalese language of Sri Lanka, encompassing vowels, consonants, and diacritical marks.

Thai
Abbr
Thai
Range
0E00 - 0E7F
Count
87

Thai is the script for the Thai language, including vowels, tone marks, digits, and punctuation, used in modern and historical texts.

Lao
Abbr
Lao
Range
0E80 - 0EFF
Count
83

Lao is a script block for the Lao language, containing consonants, vowels, tone marks, and digits used in Laos.

Tibetan
Abbr
Tibetan
Range
0F00 - 0FFF
Count
211

Tibetan is a script set for writing the Tibetan language, including letters, numerals, and punctuation marks essential for religious and literary texts.

Myanmar
Abbr
Myanmar
Range
1000 - 109F
Count
160

Myanmar is the script block for writing Burmese, including letters, digits, punctuation, and tone marks used in the Myanmar language.

Georgian
Abbr
Georgian
Range
10A0 - 10FF
Count
88

Georgian is a set of characters for the Georgian script, covering modern letters, Mtavruli capitals, and punctuation.

Hangul Jamo
Abbr
Jamo
Range
1100 - 11FF
Count
256

Hangul Jamo is a set of precomposed letters for the Korean alphabet’s consonant and vowel components, enabling syllable construction.

Ethiopic
Abbr
Ethiopic
Range
1200 - 137F
Count
358

Ethiopic is a script block for Ge’ez and modern Ethiopic languages, used to write Amharic, Tigrinya, and others.

Ethiopic Supplement
Abbr
Ethiopic_Sup
Range
1380 - 139F
Count
26

Ethiopic Supplement is a set of additional syllabic and punctuation marks for medieval and modern Geʽez script usage.

Cherokee
Abbr
Cherokee
Range
13A0 - 13FF
Count
92

Cherokee is a syllabary script used for the Cherokee language, covering syllables, punctuation, and historical letterforms in its encoded character set.

Unified Canadian Aboriginal Syllabics
Abbr
UCAS
Range
1400 - 167F
Count
640

Unified Canadian Aboriginal Syllabics is a script encoding Cree, Inuktitut, and other Indigenous languages, with letters rotated for syllable variations.

Ogham
Abbr
Ogham
Range
1680 - 169F
Count
29

Ogham is a script block encoding the ancient Irish alphabet’s linear inscriptions, used for monumental stones and later manuscripts.

Runic
Abbr
Runic
Range
16A0 - 16FF
Count
89

Runic is a script encoding ancient Germanic alphabets, featuring letters for Elder and Younger Futhark, Anglo-Saxon variants, and punctuation marks.

Tagalog
Abbr
Tagalog
Range
1700 - 171F
Count
23

Tagalog is a script block for precolonial Philippine Baybayin writing, used historically for the Tagalog language and revived in modern cultural contexts.

Hanunoo
Abbr
Hanunoo
Range
1720 - 173F
Count
23

Hanunoo is a script block for writing the Hanunoo language of the Philippines, including syllabic characters and punctuation marks.

Buhid
Abbr
Buhid
Range
1740 - 175F
Count
20

Buhid is a script block used for writing the Buhid language, containing letters, vowels, and punctuation marks for traditional Philippine texts.

Tagbanwa
Abbr
Tagbanwa
Range
1760 - 177F
Count
18

Tagbanwa is a script used for the Tagbanwa language, with letters for syllables and punctuation, encoded for digital text.

Khmer
Abbr
Khmer
Range
1780 - 17FF
Count
114

Khmer is a script block for writing the Cambodian language, including consonants, vowels, diacritics, and punctuation, used across official and religious texts.

Mongolian
Abbr
Mongolian
Range
1800 - 18AF
Count
158

Mongolian is a script for writing the Mongolian language, including letters, digits, punctuation, and marks for historical and modern use.

Unified Canadian Aboriginal Syllabics Extended
Abbr
UCAS_Ext
Range
18B0 - 18FF
Count
70

Unified Canadian Aboriginal Syllabics Extended is a set of supplementary symbols for writing additional Cree, Inuktitut, and other Indigenous languages.

Limbu
Abbr
Limbu
Range
1900 - 194F
Count
68

Limbu is a script block for writing the Limbu language, containing letters, digits, and punctuation used primarily in Nepal and Sikkim.

Tai Le
Abbr
Tai_Le
Range
1950 - 197F
Count
35

Tai Le is a script used for the Tai Nüa language, containing consonants, vowels, and tone marks.

New Tai Lue
Abbr
New_Tai_Lue
Range
1980 - 19DF
Count
83

New Tai Lue is a script used for the Tai Lue language, featuring letters, digits, and punctuation, with modernized forms for vowel and tone marks.

Khmer Symbols
Abbr
Khmer_Symbols
Range
19E0 - 19FF
Count
32

Khmer Symbols is a collection of signs used in lunar calendar calculations, including dates, times, and astronomical notations.

Buginese
Abbr
Buginese
Range
1A00 - 1A1F
Count
30

Buginese is a script used for writing the Bugis language, primarily in South Sulawesi, Indonesia, featuring consonants and vowel signs.

Tai Tham
Abbr
Tai_Tham
Range
1A20 - 1AAF
Count
127

Tai Tham is a script for writing Northern Thai, Lao, and related languages, featuring consonants, vowels, and diacritics in a rounded style.

Combining Diacritical Marks Extended
Abbr
Diacriticals_Ext
Range
1AB0 - 1AFF
Count
65

Combining Diacritical Marks Extended is a set of additional marks for modifying letters, supporting diverse writing systems and scholarly transcription needs.

Balinese
Abbr
Balinese
Range
1B00 - 1B7F
Count
127

Balinese is a script block for writing the Balinese language, containing consonants, vowels, and diacritics used in Bali, Indonesia.

Sundanese
Abbr
Sundanese
Range
1B80 - 1BBF
Count
64

Sundanese is a script block for writing the Sundanese language, containing consonants, vowels, digits, and punctuation marks.

Batak
Abbr
Batak
Range
1BC0 - 1BFF
Count
56

Batak is a script block used for writing the Batak languages of Sumatra, containing consonants, vowels, and punctuation marks.

Lepcha
Abbr
Lepcha
Range
1C00 - 1C4F
Count
74

Lepcha is a script used for writing the Lepcha language of Sikkim, India, including its letters, diacritics, and punctuation.

Ol Chiki
Abbr
Ol_Chiki
Range
1C50 - 1C7F
Count
48

Ol Chiki is a script for the Santali language, featuring letters, digits, and punctuation, designed to write its phonology clearly.

Cyrillic Extended-C
Abbr
Cyrillic_Ext_C
Range
1C80 - 1C8F
Count
11

Cyrillic Extended-C is a small set of letters for early Slavic orthographies, including historic forms of Che and Dzhe.

Georgian Extended
Abbr
Georgian_Ext
Range
1C90 - 1CBF
Count
46

Georgian Extended is a set of additional letters for the Georgian script, mainly used for the historical Mtavruli uppercase forms.

Sundanese Supplement
Abbr
Sundanese_Sup
Range
1CC0 - 1CCF
Count
8

Sundanese Supplement is a set of punctuation marks and signs used for writing the Sundanese script, filling gaps left by the main Sundanese block.

Vedic Extensions
Abbr
Vedic_Ext
Range
1CD0 - 1CFF
Count
43

Vedic Extensions is a set of diacritical and tonal marks for Sanskrit and other Indo-Aryan texts, aiding precise pronunciation and Vedic chanting.

Phonetic Extensions
Abbr
Phonetic_Ext
Range
1D00 - 1D7F
Count
128

Phonetic Extensions is a set of historical and dialectal letters for transcribing sounds not covered by standard alphabets, aiding linguists and phoneticians.

Phonetic Extensions Supplement
Abbr
Phonetic_Ext_Sup
Range
1D80 - 1DBF
Count
64

Phonetic Extensions Supplement is a set of modified Latin letters and symbols for precise phonetic transcription, especially for obscure sounds.

Combining Diacritical Marks Supplement
Abbr
Diacriticals_Sup
Range
1DC0 - 1DFF
Count
64

Combining Diacritical Marks Supplement is a set of symbols for adding phonetic or tone details to letters, spanning medieval and modern linguistic notation.

Latin Extended Additional
Abbr
Latin_Ext_Additional
Range
1E00 - 1EFF
Count
256

Latin Extended Additional is a set of Latin letters with diacritics for Vietnamese and other languages, plus medievalist and phonetic additions.

Greek Extended
Abbr
Greek_Ext
Range
1F00 - 1FFF
Count
233

Greek Extended is a set of precomposed accented Greek letters, used for classical text and polytonic orthography.

General Punctuation
Abbr
Punctuation
Range
2000 - 206F
Count
111

General Punctuation is a set of symbols for spacing, dashes, quotation marks, and invisible format controls, aiding text layout and clarity.

Superscripts and Subscripts
Abbr
Super_And_Sub
Range
2070 - 209F
Count
46

Superscripts and Subscripts is a set of characters for scientific notation, math, and linguistic use, including superscript digits, subscripts, and modifiers.

Currency Symbols
Abbr
Currency_Symbols
Range
20A0 - 20CF
Count
37

Currency Symbols is a set of standardized signs for global currencies, including the euro, rupee, and lira, used in digital text.

Combining Diacritical Marks for Symbols
Abbr
Diacriticals_For_Symbols
Range
20D0 - 20FF
Count
33

Combining Diacritical Marks for Symbols is a set of marks that overlay mathematical, technical, or currency symbols to modify their meaning.

Letterlike Symbols
Abbr
Letterlike_Symbols
Range
2100 - 214F
Count
80

Letterlike Symbols is a set of typographic and mathematical signs that resemble letters, including units, abbreviations, and numerals.

Number Forms
Abbr
Number_Forms
Range
2150 - 218F
Count
60

Number Forms is a set of characters for fractions, roman numerals, and currency symbols, enabling compact representation of numeric concepts in text.

Arrows
Abbr
Arrows
Range
2190 - 21FF
Count
112

Arrows is a collection of symbols for directional indicators, including simple, double, and curved variants, used in math, UI, and logic.

Mathematical Operators
Abbr
Math_Operators
Range
2200 - 22FF
Count
256

Mathematical Operators is a collection of symbols for arithmetic, set theory, logic, and related math notation.

Miscellaneous Technical
Abbr
Misc_Technical
Range
2300 - 23FF
Count
256

Miscellaneous Technical is a collection of symbols for engineering, computing, and control functions, including arrows, keys, and notation like ⌈⌉ and ⌊⌋.

Control Pictures
Abbr
Control_Pictures
Range
2400 - 243F
Count
42

Control Pictures is a set of symbols visually representing ASCII control characters used for diagrams and debugging.

Optical Character Recognition
Abbr
OCR
Range
2440 - 245F
Count
11

Optical Character Recognition is a set of symbols for scanned text correction, including marks like check, erase, and hold.

Enclosed Alphanumerics
Abbr
Enclosed_Alphanum
Range
2460 - 24FF
Count
160

Enclosed Alphanumerics is a set of circled, parenthesized, and full-stop numerals and letters used for lists, footnotes, and decorative numbering.

Box Drawing
Abbr
Box_Drawing
Range
2500 - 257F
Count
128

Box Drawing is a set of characters for creating simple line and box diagrams, including horizontal, vertical, and corner pieces with various line styles.

Block Elements
Abbr
Block_Elements
Range
2580 - 259F
Count
32

Block Elements is a set of 32 graphic symbols for creating shaded, hatched, and partial box-drawing patterns, primarily used in text-based interfaces.

Geometric Shapes
Abbr
Geometric_Shapes
Range
25A0 - 25FF
Count
96

Geometric Shapes is a set of symbols for squares, circles, triangles, and arrows, used in diagrams and text.

Miscellaneous Symbols
Abbr
Misc_Symbols
Range
2600 - 26FF
Count
256

Miscellaneous Symbols is a collection of diverse pictographs, including weather icons, chess pieces, and warning signs, used across digital texts.

Dingbats
Abbr
Dingbats
Range
2700 - 27BF
Count
192

Dingbats is a set of ornamental symbols, including stars, arrows, crosses, and graphic flourishes, used for decoration and visual emphasis.

Miscellaneous Mathematical Symbols-A
Abbr
Misc_Math_Symbols_A
Range
27C0 - 27EF
Count
48

Miscellaneous Mathematical Symbols-A is a collection of operators, angles, and notation for equations, including the lozenge, sum, and integral variants.

Supplemental Arrows-A
Abbr
Sup_Arrows_A
Range
27F0 - 27FF
Count
16

Supplemental Arrows-A is a set of arrow symbols for mathematical notation, including long arrows and bent arrows, designed for technical contexts.

Braille Patterns
Abbr
Braille
Range
2800 - 28FF
Count
256

Braille Patterns is a set of tactile cell combinations, mapping six and eight dot systems to text, math, and music notation.

Supplemental Arrows-B
Abbr
Sup_Arrows_B
Range
2900 - 297F
Count
128

Supplemental Arrows-B is a set of varied arrow symbols, including curved, zigzag, and feathered variants, used for technical notation and mathematics.

Miscellaneous Mathematical Symbols-B
Abbr
Misc_Math_Symbols_B
Range
2980 - 29FF
Count
128

Miscellaneous Mathematical Symbols-B is a set of symbols for advanced math, including operators, arrows, and geometric shapes, aiding technical notation.

Supplemental Mathematical Operators
Abbr
Sup_Math_Operators
Range
2A00 - 2AFF
Count
256

Supplemental Mathematical Operators is a collection of additional symbols for advanced algebra, logic, and set theory, extending standard math notation.

Miscellaneous Symbols and Arrows
Abbr
Misc_Arrows
Range
2B00 - 2BFF
Count
254

Miscellaneous Symbols and Arrows is a set of diverse icons, including arrows, shapes, and weather symbols, for technical and casual use.

Glagolitic
Abbr
Glagolitic
Range
2C00 - 2C5F
Count
96

Glagolitic is a historical script block used for writing Old Church Slavonic, containing uppercase and lowercase letters plus punctuation marks.

Latin Extended-C
Abbr
Latin_Ext_C
Range
2C60 - 2C7F
Count
32

Latin Extended-C is a set of additional Latin letters and symbols used for phonetic transcription and minority languages.

Coptic
Abbr
Coptic
Range
2C80 - 2CFF
Count
123

Coptic is a script block used for writing the Coptic language, derived from Greek with added letters for Egyptian sounds.

Georgian Supplement
Abbr
Georgian_Sup
Range
2D00 - 2D2F
Count
40

Georgian Supplement is a set of lowercase Georgian letters, derived from the ancient Asomtavruli script, used for ecclesiastical texts and historical writing.

Tifinagh
Abbr
Tifinagh
Range
2D30 - 2D7F
Count
59

Tifinagh is a script used for Berber languages, with its block containing letters for Tamazight, Tuareg, and other variants, plus punctuation and digits.

Ethiopic Extended
Abbr
Ethiopic_Ext
Range
2D80 - 2DDF
Count
79

Ethiopic Extended is a set of supplementary syllabic and punctuation signs for Ge'ez, supporting historical and modern liturgical texts.

Cyrillic Extended-A
Abbr
Cyrillic_Ext_A
Range
2DE0 - 2DFF
Count
32

Cyrillic Extended-A is a set of combining marks used for early Slavic orthographies, enabling accurate representation of historical texts.

Supplemental Punctuation
Abbr
Sup_Punctuation
Range
2E00 - 2E7F
Count
98

Supplemental Punctuation is a set of rare punctuation marks, including historic and editorial symbols, added to support specialized text and scholarly use.

CJK Radicals Supplement
Abbr
CJK_Radicals_Sup
Range
2E80 - 2EFF
Count
115

CJK Radicals Supplement is a set of alternate forms of Chinese characters’ radicals, used in dictionaries and indexing, complementing the main Kangxi Radicals.

Kangxi Radicals
Abbr
Kangxi
Range
2F00 - 2FDF
Count
214

Kangxi Radicals is a set of historical Chinese character components, arranged in their traditional dictionary order, used for indexing and referencing.

Ideographic Description Characters
Abbr
IDC
Range
2FF0 - 2FFF
Count
16

Ideographic Description Characters is a set of sixteen symbols used to visually describe the composition of complex Chinese characters from simpler components.

CJK Symbols and Punctuation
Abbr
CJK_Symbols
Range
3000 - 303F
Count
64

CJK Symbols and Punctuation is a set of ideographic punctuation marks, iteration signs, and spacing characters used in Chinese, Japanese, and Korean writing.

Hiragana
Abbr
Hiragana
Range
3040 - 309F
Count
93

Hiragana is a Japanese syllabary of characters used for native words, grammar, and furigana, including voiced marks and the archaic "wi" and "we".

Katakana
Abbr
Katakana
Range
30A0 - 30FF
Count
96

Katakana is a Japanese syllabary of characters used primarily for foreign loanwords, onomatopoeia, and phonetic emphasis.

Bopomofo
Abbr
Bopomofo
Range
3100 - 312F
Count
43

Bopomofo is a set of phonetic symbols for Mandarin Chinese, used to annotate pronunciation, especially in Taiwan, within a dedicated character range.

Hangul Compatibility Jamo
Abbr
Compat_Jamo
Range
3130 - 318F
Count
94

Hangul Compatibility Jamo is a set of precomposed Korean letters for legacy encoding and vertical text, preserving old Hangul syllables and digraphs.

Kanbun
Abbr
Kanbun
Range
3190 - 319F
Count
16

Kanbun is a set of annotation marks used in Japanese texts to indicate reading order for classical Chinese, aiding in translation.

Bopomofo Extended
Abbr
Bopomofo_Ext
Range
31A0 - 31BF
Count
32

Bopomofo Extended is a small set of phonetic symbols adding rare and dialectal Mandarin sounds, plus the unique characters for Hokkien and Hakka.

CJK Strokes
Abbr
CJK_Strokes
Range
31C0 - 31EF
Count
39

CJK Strokes is a set of standardized symbols for writing Chinese, Japanese, and Korean brushstroke components, used in dictionaries and calligraphy references.

Katakana Phonetic Extensions
Abbr
Katakana_Ext
Range
31F0 - 31FF
Count
16

Katakana Phonetic Extensions is a set of small marks for Ainu and other languages, used to modify katakana sounds like p, t, or s.

Enclosed CJK Letters and Months
Abbr
Enclosed_CJK
Range
3200 - 32FF
Count
255

Enclosed CJK Letters and Months is a set of symbols enclosing Korean syllables, Japanese katakana, and month names within circles or parentheses.

CJK Compatibility
Abbr
CJK_Compat
Range
3300 - 33FF
Count
256

CJK Compatibility is a set of precomposed Japanese Katakana squares, like ㌔, and other CJK symbols for legacy character mapping and vertical text layouts.

CJK Unified Ideographs Extension A
Abbr
CJK_Ext_A
Range
3400 - 4DBF
Count
6592

CJK Unified Ideographs Extension A is a collection of rare and archaic Chinese characters, added to support historical texts and lesser-known names.

Yijing Hexagram Symbols
Abbr
Yijing
Range
4DC0 - 4DFF
Count
64

Yijing Hexagram Symbols is a set of monochrome trigram and hexagram glyphs used for casting the I Ching oracle.

CJK Unified Ideographs
Abbr
CJK
Range
4E00 - 9FFF
Count
20992

CJK Unified Ideographs is a set of over 20,000 Chinese characters shared across Chinese, Japanese, and Korean writing systems, covering common Han script usage.

Yi Syllables
Abbr
Yi_Syllables
Range
A000 - A48F
Count
1165

Yi Syllables is a set of characters used for writing the Nuosu language, based on the standardized Liangshan dialect.

Yi Radicals
Abbr
Yi_Radicals
Range
A490 - A4CF
Count
55

Yi Radicals is a set of symbols used to index the traditional Yi script, primarily for dictionary and educational lookup purposes.

Lisu
Abbr
Lisu
Range
A4D0 - A4FF
Count
48

Lisu is a script used for the Lisu language, with letters, digits, and punctuation, designed for tonal and phonetic writing.

Vai
Abbr
Vai
Range
A500 - A63F
Count
300

Vai is a script used for the Vai language of Liberia, with characters for syllables, punctuation, and historical logograms.

Cyrillic Extended-B
Abbr
Cyrillic_Ext_B
Range
A640 - A69F
Count
96

Cyrillic Extended-B is a set of historic letters and signs for Old Church Slavonic and early Cyrillic scripts, including abbreviations and variant forms.

Bamum
Abbr
Bamum
Range
A6A0 - A6FF
Count
88

Bamum is a script for the Bamum language, used for royal decrees and everyday writing in Cameroon's Bamum kingdom.

Modifier Tone Letters
Abbr
Modifier_Tone_Letters
Range
A700 - A71F
Count
32

Modifier Tone Letters is a set of phonetic symbols used for marking tone contours in linguistic transcription, primarily for African and Asian languages.

Latin Extended-D
Abbr
Latin_Ext_D
Range
A720 - A7FF
Count
206

Latin Extended-D is a set of Latin letters and phonetic symbols for medieval and African languages, plus historical additions.

Syloti Nagri
Abbr
Syloti_Nagri
Range
A800 - A82F
Count
45

Syloti Nagri is a script used for the Sylheti language, containing letters, digits, and punctuation, with characters for vowels, consonants, and diacritics.

Common Indic Number Forms
Abbr
Indic_Number_Forms
Range
A830 - A83F
Count
10

Common Indic Number Forms is a set of historical numeral signs used across various Indic scripts, primarily for accounting and monetary values.

Phags-pa
Abbr
Phags_Pa
Range
A840 - A87F
Count
56

Phags-pa is a script used historically for writing Mongolian, Chinese, and other languages, encoded for digital text with letters, marks, and punctuation.

Saurashtra
Abbr
Saurashtra
Range
A880 - A8DF
Count
82

Saurashtra is a script used for writing the Saurashtra language with letters, digits, and diacritics.

Devanagari Extended
Abbr
Devanagari_Ext
Range
A8E0 - A8FF
Count
32

Devanagari Extended is a set of combining marks and signs used to modify vowels and consonants in Sanskrit and other Indic languages.

Kayah Li
Abbr
Kayah_Li
Range
A900 - A92F
Count
48

Kayah Li is a script used for the Kayah language, containing letters, digits, and punctuation with tone marks.

Rejang
Abbr
Rejang
Range
A930 - A95F
Count
37

Rejang is a script used for writing the Rejang language of Sumatra, containing letters for consonants, vowels, and diacritics.

Hangul Jamo Extended-A
Abbr
Jamo_Ext_A
Range
A960 - A97F
Count
29

Hangul Jamo Extended-A is a set of additional consonants and vowel jamo used for writing Middle Korean, filling gaps in the standard Hangul system.

Javanese
Abbr
Javanese
Range
A980 - A9DF
Count
91

Javanese is a script used for writing the Javanese language, with letters, digits, and punctuation, plus historical and modern variant forms.

Myanmar Extended-B
Abbr
Myanmar_Ext_B
Range
A9E0 - A9FF
Count
31

Myanmar Extended-B is a small set of additional letters and signs for minority languages and historical texts, complementing earlier Myanmar encoding.

Cham
Abbr
Cham
Range
AA00 - AA5F
Count
83

Cham is a script used for writing the Cham language of Vietnam and Cambodia, with distinct letters for the Eastern and Western dialects.

Myanmar Extended-A
Abbr
Myanmar_Ext_A
Range
AA60 - AA7F
Count
32

Myanmar Extended-A is a set of additional Tai Laing and other Myanmar script characters for historical and minority language support.

Tai Viet
Abbr
Tai_Viet
Range
AA80 - AADF
Count
72

Tai Viet is a script used for writing the Tai Dam and Tai Don languages, featuring distinct consonants, vowels, and tone marks.

Meetei Mayek Extensions
Abbr
Meetei_Mayek_Ext
Range
AAE0 - AAFF
Count
23

Meetei Mayek Extensions is a supplementary set of characters for the Meitei script, adding letters and digits for historical and modern usage.

Ethiopic Extended-A
Abbr
Ethiopic_Ext_A
Range
AB00 - AB2F
Count
32

Ethiopic Extended-A is a set of additional Gamo, Gofa, and Gurage syllables and punctuation, filling gaps in modern Ethiopic script usage.

Latin Extended-E
Abbr
Latin_Ext_E
Range
AB30 - AB6F
Count
62

Latin Extended-E is a set of characters for transcribing medievalist and dialectal letter forms, including modifier letters and additional vowels.

Cherokee Supplement
Abbr
Cherokee_Sup
Range
AB70 - ABBF
Count
80

Cherokee Supplement is a set of lowercase syllabary characters complementing the main Cherokee block, used for writing the Cherokee language.

Meetei Mayek
Abbr
Meetei_Mayek
Range
ABC0 - ABFF
Count
56

Meetei Mayek is a script used for the Manipuri language, containing letters, digits, and punctuation marks for writing this Tibeto-Burman tongue.

Hangul Syllables
Abbr
Hangul
Range
AC00 - D7AF
Count
11172

Hangul Syllables is a precomposed set of over 11,000 Korean syllable characters enabling direct encoding of each possible syllable block.

Hangul Jamo Extended-B
Abbr
Jamo_Ext_B
Range
D7B0 - D7FF
Count
72

Hangul Jamo Extended-B is a set of precomposed syllables, filling gaps in archaic Hangul orthography for historical and dialectal Korean texts.

High Surrogates
Abbr
High_Surrogates
Range
D800 - DB7F
Count
0

High Surrogates is a reserved range for the first half of surrogate pairs enabling encoding of supplementary characters in UTF-16.

High Private Use Surrogates
Abbr
High_PU_Surrogates
Range
DB80 - DBFF
Count
0

High Private Use Surrogates is a reserved range for private use, enabling applications to assign custom characters without standardized meanings.

Low Surrogates
Abbr
Low_Surrogates
Range
DC00 - DFFF
Count
0

Low Surrogates is a reserved range for the second half of surrogate pairs, enabling encoding of supplementary characters in UTF-16.

Private Use Area
Abbr
PUA
Range
E000 - F8FF
Count
6400

Private Use Area is a reserved zone for characters defined by individual fonts or applications, not standardized for universal exchange.

CJK Compatibility Ideographs
Abbr
CJK_Compat_Ideographs
Range
F900 - FAFF
Count
472

CJK Compatibility Ideographs is a set of mostly duplicate or variant Chinese characters, included for round‑trip compatibility with older standards.

Alphabetic Presentation Forms
Abbr
Alphabetic_PF
Range
FB00 - FB4F
Count
58

Alphabetic Presentation Forms is a set of precomposed ligatures and letter combinations, primarily for Latin and Armenian scripts, aiding typographic rendering.

Arabic Presentation Forms-A
Abbr
Arabic_PF_A
Range
FB50 - FDFF
Count
656

Arabic Presentation Forms-A is a set of contextual letterforms, ligatures, and honorifics used for stylized Arabic writing, spanning over 600 encoded glyphs.

Variation Selectors
Abbr
VS
Range
FE00 - FE0F
Count
16

Variation Selectors is a set of formatting characters that modify the presentation style of preceding ideographs, emoji, or symbols.

Vertical Forms
Abbr
Vertical_Forms
Range
FE10 - FE1F
Count
10

Vertical Forms is a set of ten small-width punctuation marks, including commas and parentheses, used for vertical text layout in East Asian typography.

Combining Half Marks
Abbr
Half_Marks
Range
FE20 - FE2F
Count
16

Combining Half Marks is a set of diacritical signs that attach to preceding letters, enabling precise phonetic and scholarly notation.

CJK Compatibility Forms
Abbr
CJK_Compat_Forms
Range
FE30 - FE4F
Count
32

CJK Compatibility Forms is a small set of vertical presentation forms for East Asian punctuation and symbols, used mainly for legacy vertical text layout.

Small Form Variants
Abbr
Small_Forms
Range
FE50 - FE6F
Count
26

Small Form Variants is a set of compact punctuation and symbols, used mostly for vertical Chinese and Japanese text layout, mirroring fullwidth forms.

Arabic Presentation Forms-B
Abbr
Arabic_PF_B
Range
FE70 - FEFF
Count
141

Arabic Presentation Forms-B is a set of contextual Arabic letter forms and diacritics, including the zero-width no-break space, used for legacy compatibility.

Halfwidth and Fullwidth Forms
Abbr
Half_And_Full_Forms
Range
FF00 - FFEF
Count
225

Halfwidth and Fullwidth Forms is a set of characters used to align East Asian and Latin text, providing fullwidth versions of ASCII and halfwidth katakana.

Specials
Abbr
Specials
Range
FFF0 - FFFF
Count
5

Specials is a set of reserved and noncharacter code points, including the replacement character, used for internal processing and compatibility.

Linear B Syllabary
Abbr
Linear_B_Syllabary
Range
10000 - 1007F
Count
88

Linear B Syllabary is a script encoding ancient Mycenaean Greek syllables, used for administrative records on clay tablets in Bronze Age Crete and Greece.

Linear B Ideograms
Abbr
Linear_B_Ideograms
Range
10080 - 100FF
Count
123

Linear B Ideograms is a set of signs for syllabic and ideographic writing, used to record Mycenaean Greek economic and administrative texts.

Aegean Numbers
Abbr
Aegean_Numbers
Range
10100 - 1013F
Count
57

Aegean Numbers is a set of ancient symbols used for counting in Minoan and Mycenaean scripts, including fractions and numerals.

Ancient Greek Numbers
Abbr
Ancient_Greek_Numbers
Range
10140 - 1018F
Count
79

Ancient Greek Numbers is a set of symbols for acrophonic numerals, including talent signs and fractions, used in classical inscriptions.

Ancient Symbols
Abbr
Ancient_Symbols
Range
10190 - 101CF
Count
14

Ancient Symbols is a collection of glyphs for weights, measures, and astronomical signs, plus historical notations from Greek and Roman antiquity.

Phaistos Disc
Abbr
Phaistos
Range
101D0 - 101FF
Count
46

Phaistos Disc is a historic script symbol set used for encoding the undeciphered ancient Minoan artifact’s markings, including pictographs and dividers.

Lycian
Abbr
Lycian
Range
10280 - 1029F
Count
29

Lycian is a script block encoding the ancient Anatolian alphabet used for the Lycian language.

Carian
Abbr
Carian
Range
102A0 - 102DF
Count
49

Carian is a historical script block encoding the alphabetic writing system of ancient Caria, used from the 7th to 3rd centuries BCE.

Coptic Epact Numbers
Abbr
Coptic_Epact_Numbers
Range
102E0 - 102FF
Count
28

Coptic Epact Numbers is a set of ancient Egyptian numerical symbols used for astronomical and calendrical calculations, encoded in the standard character set.

Old Italic
Abbr
Old_Italic
Range
10300 - 1032F
Count
39

Old Italic is a script encoding ancient Etruscan and related Italic alphabets, including letters for writing early Latin and Oscan texts.

Gothic
Abbr
Gothic
Range
10330 - 1034F
Count
27

Gothic is a script used to write the extinct East Germanic language, comprising letters derived from the Greek and Latin alphabets.

Old Permic
Abbr
Old_Permic
Range
10350 - 1037F
Count
43

Old Permic is a script used for writing the Komi language, containing letters and digits from the historic Abur alphabet.

Ugaritic
Abbr
Ugaritic
Range
10380 - 1039F
Count
31

Ugaritic is a cuneiform script block encoding the ancient Ugaritic alphabet, used for writing the Northwest Semitic language of Ugarit.

Old Persian
Abbr
Old_Persian
Range
103A0 - 103DF
Count
50

Old Persian is a cuneiform script encoding the language of ancient Achaemenid inscriptions, featuring signs for syllables and logograms.

Deseret
Abbr
Deseret
Range
10400 - 1044F
Count
80

Deseret is a historical script block for the phonetic alphabet created by the LDS Church, used mainly in 19th-century Utah documents.

Shavian
Abbr
Shavian
Range
10450 - 1047F
Count
48

Shavian is a phonemic alphabet designed for English, created by Ronald Kingsley Read, with each character representing a distinct speech sound.

Osmanya
Abbr
Osmanya
Range
10480 - 104AF
Count
40

Osmanya is a script invented for Somali, used for writing the language and included in digital text standards.

Osage
Abbr
Osage
Range
104B0 - 104FF
Count
72

Osage is a script used for writing the Osage language, consisting of uppercase and lowercase letters with distinct phonetic forms.

Elbasan
Abbr
Elbasan
Range
10500 - 1052F
Count
40

Elbasan is a historical alphabet block encoding the script used for writing the Albanian language in the 18th century.

Caucasian Albanian
Abbr
Caucasian_Albanian
Range
10530 - 1056F
Count
53

Caucasian Albanian is a script used for an ancient Christian language, containing letters for the now extinct Caucasian Albanian language.

Vithkuqi
Abbr
Vithkuqi
Range
10570 - 105BF
Count
70

Vithkuqi is a historical Albanian script block, encoded for the original alphabet’s 46 letters, used in 19th-century texts.

Todhri
Abbr
Todhri
Range
105C0 - 105FF
Count
52

Todhri is a script block encoding an Albanian alphabet, used historically for writing the Tosk dialect, with characters for letters and punctuation.

Linear A
Abbr
Linear_A
Range
10600 - 1077F
Count
341

Linear A is a proposed script block for the undeciphered Minoan writing system, containing syllabic and ideographic signs from ancient Cretan inscriptions.

Latin Extended-F
Abbr
Latin_Ext_F
Range
10780 - 107BF
Count
62

Latin Extended-F is a set of letters and symbols for medievalist and phonetic notations, including extinct scribal abbreviations and tone marks.

Cypriot Syllabary
Abbr
Cypriot_Syllabary
Range
10800 - 1083F
Count
55

Cypriot Syllabary is a historical script used for ancient Cypriot Greek and Eteocypriot inscriptions with characters for syllabic writing.

Imperial Aramaic
Abbr
Imperial_Aramaic
Range
10840 - 1085F
Count
31

Imperial Aramaic is a right-to-left script used for the Aramaic language of the Achaemenid Empire, containing letters and punctuation.

Palmyrene
Abbr
Palmyrene
Range
10860 - 1087F
Count
32

Palmyrene is a historical script block used for writing the Aramaic dialect of the ancient Syrian city, containing twenty‑seven letters and two numerals.

Nabataean
Abbr
Nabataean
Range
10880 - 108AF
Count
40

Nabataean is a script block for writing the ancient Nabataean language, used in inscriptions from the 2nd century BCE to the 4th century CE.

Hatran
Abbr
Hatran
Range
108E0 - 108FF
Count
26

Hatran is a script used for the Aramaic dialect of the ancient city of Hatra, inscribed right-to-left with letters and numerals.

Phoenician
Abbr
Phoenician
Range
10900 - 1091F
Count
29

Phoenician is an ancient Semitic writing system of 29 letters, used mainly for inscriptions, and encoded as a right-to-left script.

Lydian
Abbr
Lydian
Range
10920 - 1093F
Count
27

Lydian is a historic script block for the extinct Anatolian language, containing letters and a word separator symbol.

Sidetic
Abbr
Sidetic
Range
10940 - 1095F
Count
26

Sidetic is a historical script block used for writing the ancient Sidetic language of Anatolia, containing 14 characters including letters and a numeral sign.

Meroitic Hieroglyphs
Abbr
Meroitic_Hieroglyphs
Range
10980 - 1099F
Count
32

Meroitic Hieroglyphs is a historical script block encoding the ancient Egyptian-derived writing system used for the Meroitic language in Nubia.

Meroitic Cursive
Abbr
Meroitic_Cursive
Range
109A0 - 109FF
Count
90

Meroitic Cursive is a historical script used in ancient Nubia, featuring cursive signs for writing the Meroitic language, with some characters still unresolved.

Kharoshthi
Abbr
Kharoshthi
Range
10A00 - 10A5F
Count
68

Kharoshthi is a historical script from ancient South Asia, used for writing Gandhari and other Prakrits, encoded with letters, numerals, and punctuation marks.

Old South Arabian
Abbr
Old_South_Arabian
Range
10A60 - 10A7F
Count
32

Old South Arabian is a script block encoding the ancient Musnad alphabet, used for inscriptions in the pre-Islamic Yemeni kingdoms.

Old North Arabian
Abbr
Old_North_Arabian
Range
10A80 - 10A9F
Count
32

Old North Arabian is a script block encoding the ancient alphabetic writing system used for inscriptions across northwest Arabia.

Manichaean
Abbr
Manichaean
Range
10AC0 - 10AFF
Count
51

Manichaean is a script used for writing the Manichaean religion’s texts, including letters, punctuation, and abbreviations, derived from Syriac and Estrangela.

Avestan
Abbr
Avestan
Range
10B00 - 10B3F
Count
61

Avestan is a script used for the sacred Zoroastrian texts, encoded with letters, digits, and punctuation marks.

Inscriptional Parthian
Abbr
Inscriptional_Parthian
Range
10B40 - 10B5F
Count
30

Inscriptional Parthian is a script used for Middle Persian inscriptions, containing letters, numbers, and punctuation marks from ancient Parthian texts.

Inscriptional Pahlavi
Abbr
Inscriptional_Pahlavi
Range
10B60 - 10B7F
Count
27

Inscriptional Pahlavi is a script used for Middle Persian inscriptions, encoded with letters and punctuation for historical texts.

Psalter Pahlavi
Abbr
Psalter_Pahlavi
Range
10B80 - 10BAF
Count
29

Psalter Pahlavi is a script used for Middle Persian religious texts, featuring letters and numerals, now encoded digitally for historical scholarship.

Old Turkic
Abbr
Old_Turkic
Range
10C00 - 10C4F
Count
73

Old Turkic is a script used for inscriptions and manuscripts from the 8th to 13th centuries, including runiform letters for several Turkic languages.

Old Hungarian
Abbr
Old_Hungarian
Range
10C80 - 10CFF
Count
108

Old Hungarian is a script block encoding the ancient runiform writing of the Magyars, used for inscriptions and early texts.

Hanifi Rohingya
Abbr
Hanifi_Rohingya
Range
10D00 - 10D3F
Count
50

Hanifi Rohingya is a script used for writing the Rohingya language, featuring letters, digits, and diacritics designed for tonal distinctions.

Garay
Abbr
Garay
Range
10D40 - 10D8F
Count
69

Garay is a script used for the Wolof language, recently encoded with letters, digits, and punctuation for writing in Senegal.

Rumi Numeral Symbols
Abbr
Rumi
Range
10E60 - 10E7F
Count
31

Rumi Numeral Symbols is a set of ancient Maghrebi digits and fractions used historically in North African Arabic manuscripts and mathematical texts.

Yezidi
Abbr
Yezidi
Range
10E80 - 10EBF
Count
47

Yezidi is a script used for writing the Kurdish Yezidi religious language with letters and diacritics.

Arabic Extended-C
Abbr
Arabic_Ext_C
Range
10EC0 - 10EFF
Count
60

Arabic Extended-C is a repertoire of Arabic script characters for Quranic annotation, including additional diacritics and marks used in religious texts.

Old Sogdian
Abbr
Old_Sogdian
Range
10F00 - 10F2F
Count
40

Old Sogdian is a script used for writing the Sogdian language, containing letters and numerals from the ancient Iranian civilization.

Sogdian
Abbr
Sogdian
Range
10F30 - 10F6F
Count
42

Sogdian is a script used for the ancient Iranian language of the Silk Road, with letters written right-to-left and often connected cursively.

Old Uyghur
Abbr
Old_Uyghur
Range
10F70 - 10FAF
Count
26

Old Uyghur is a script used for the Turkic language, featuring cursive letters written vertically, with distinct forms for consonants and vowels.

Chorasmian
Abbr
Chorasmian
Range
10FB0 - 10FDF
Count
28

Chorasmian is a script used for the extinct Eastern Iranian language, encoded with letters, numbers, and punctuation for historical texts.

Elymaic
Abbr
Elymaic
Range
10FE0 - 10FFF
Count
23

Elymaic is a script used for the Aramaic dialect of ancient Elymais, containing letters and numbers for historical inscriptions.

Brahmi
Abbr
Brahmi
Range
11000 - 1107F
Count
115

Brahmi is a historic script block encoding early Indian inscriptions, used for writing Prakrit, Sanskrit, and other ancient languages.

Kaithi
Abbr
Kaithi
Range
11080 - 110CF
Count
68

Kaithi is a historical script from northern India, used mainly for legal and commercial records, now encoded for digital text preservation.

Sora Sompeng
Abbr
Sora_Sompeng
Range
110D0 - 110FF
Count
35

Sora Sompeng is a script used for writing the Sora language, with letters for consonants, vowels, and digits, plus diacritics and punctuation.

Chakma
Abbr
Chakma
Range
11100 - 1114F
Count
71

Chakma is a script used for the Chakma language, containing letters, digits, and diacritics for writing in Bangladesh and India.

Mahajani
Abbr
Mahajani
Range
11150 - 1117F
Count
39

Mahajani is a historical script from northwest India, used for accounting and mercantile records, with its characters encoded in the Unicode standard.

Sharada
Abbr
Sharada
Range
11180 - 111DF
Count
96

Sharada is an ancient script from Kashmir used for Sanskrit and Kashmiri, now encoded for modern digital text preservation and scholarly use.

Sinhala Archaic Numbers
Abbr
Sinhala_Archaic_Numbers
Range
111E0 - 111FF
Count
20

Sinhala Archaic Numbers is a set of ancient numeral symbols used in historical Sri Lankan records, now encoded for digital preservation.

Khojki
Abbr
Khojki
Range
11200 - 1124F
Count
65

Khojki is a script used for the Sindhi language, encoded for writing liturgical and secular texts with distinct letterforms.

Multani
Abbr
Multani
Range
11280 - 112AF
Count
38

Multani is an ancient script used for writing the Saraiki language, with characters for consonants, vowels, and diacritics.

Khudawadi
Abbr
Khudawadi
Range
112B0 - 112FF
Count
69

Khudawadi is a script used for writing Sindhi, encoded with characters for consonants, vowels, and diacritical marks.

Grantha
Abbr
Grantha
Range
11300 - 1137F
Count
86

Grantha is a historic South Indian script block for writing Sanskrit and Tamil, encoded for scholarly and digital preservation.

Tulu-Tigalari
Abbr
Tulu_Tigalari
Range
11380 - 113FF
Count
80

Tulu-Tigalari is a script for the Tulu language, comprising letters, digits, and diacritics for writing ancient inscriptions and modern texts.

Newa
Abbr
Newa
Range
11400 - 1147F
Count
97

Newa is a script used for writing the Nepal Bhasa language, containing letters, digits, and punctuation marks for the Newar people.

Tirhuta
Abbr
Tirhuta
Range
11480 - 114DF
Count
82

Tirhuta is a script block used for writing the Maithili language, primarily in the Mithila region of India and Nepal.

Siddham
Abbr
Siddham
Range
11580 - 115FF
Count
92

Siddham is a historic script block used for Buddhist texts, encoded for digital preservation of medieval Indian and East Asian manuscripts.

Modi
Abbr
Modi
Range
11600 - 1165F
Count
79

Modi is a script used historically for writing the Marathi language, featuring cursive letterforms and distinct vowel signs, now encoded digitally.

Mongolian Supplement
Abbr
Mongolian_Sup
Range
11660 - 1167F
Count
13

Mongolian Supplement is a set of marks for the Mongolian script, used to indicate gender, case, and other grammatical distinctions.

Takri
Abbr
Takri
Range
11680 - 116CF
Count
68

Takri is a script used historically for languages like Dogri and Chambeali, with glyphs for consonants, vowels, and diacritical marks.

Myanmar Extended-C
Abbr
Myanmar_Ext_C
Range
116D0 - 116FF
Count
20

Myanmar Extended-C is a recently added set of characters for minority languages and historical uses, filling gaps in the Myanmar script.

Ahom
Abbr
Ahom
Range
11700 - 1174F
Count
65

Ahom is a script block encoding the historical Ahom language and Tai-Ahom texts, used for writing inscriptions and manuscripts from medieval Assam.

Dogra
Abbr
Dogra
Range
11800 - 1184F
Count
60

Dogra is a script used for the Dogri language, with characters for vowels, consonants, and digits, encoded in the South Asian writing systems section.

Warang Citi
Abbr
Warang_Citi
Range
118A0 - 118FF
Count
84

Warang Citi is a script used for writing the Ho language, with letters and digits for numerals, plus punctuation and spacing marks.

Dives Akuru
Abbr
Dives_Akuru
Range
11900 - 1195F
Count
72

Dives Akuru is a historical script from the Maldives, used for ancient Dhivehi texts, encoded for digital preservation.

Nandinagari
Abbr
Nandinagari
Range
119A0 - 119FF
Count
65

Nandinagari is a script used historically for Sanskrit and Kannada, encoded to preserve ancient inscriptions and manuscripts.

Zanabazar Square
Abbr
Zanabazar_Square
Range
11A00 - 11A4F
Count
72

Zanabazar Square is a historical Mongolian script used for Buddhist texts, with a rounded, calligraphic style and vertical writing direction.

Soyombo
Abbr
Soyombo
Range
11A50 - 11AAF
Count
83

Soyombo is a historic script used primarily for Mongolian and Tibetan Buddhist texts, notable for its distinct vertical style and ceremonial, decorative use.

Unified Canadian Aboriginal Syllabics Extended-A
Abbr
UCAS_Ext_A
Range
11AB0 - 11ABF
Count
16

Unified Canadian Aboriginal Syllabics Extended-A is a small set of characters for writing additional Cree, Ojibwe, and Dene languages.

Pau Cin Hau
Abbr
Pau_Cin_Hau
Range
11AC0 - 11AFF
Count
57

Pau Cin Hau is a historic script used to write the Chin languages of Myanmar, encoded for digital text representation.

Devanagari Extended-A
Abbr
Devanagari_Ext_A
Range
11B00 - 11B5F
Count
11

Devanagari Extended-A is a set of rare historical and ritual characters, including signs for vedic chants and ligatures, used in ancient Sanskrit manuscripts.

Sharada Supplement
Abbr
Sharada_Sup
Range
11B60 - 11B7F
Count
8

Sharada Supplement is a set of vowel signs and marks, used for writing the Sharada script, that follow the main Sharada block.

Sunuwar
Abbr
Sunuwar
Range
11BC0 - 11BFF
Count
44

Sunuwar is a script encoding the Sunuwar language of Nepal for writing its consonants, vowels, and diacritics.

Bhaiksuki
Abbr
Bhaiksuki
Range
11C00 - 11C6F
Count
97

Bhaiksuki is a script used for writing Sanskrit, with historical significance in Buddhist texts, featuring consonants, vowels, and diacritical marks.

Marchen
Abbr
Marchen
Range
11C70 - 11CBF
Count
68

Marchen is a script used for writing the extinct Zhang-Zhung language, with its characters encoded for historical and scholarly use.

Masaram Gondi
Abbr
Masaram_Gondi
Range
11D00 - 11D5F
Count
75

Masaram Gondi is a script used for writing the Gondi language, with letters, digits, and marks for sounds.

Gunjala Gondi
Abbr
Gunjala_Gondi
Range
11D60 - 11DAF
Count
63

Gunjala Gondi is a script used for writing the Gondi language, with characters for consonants, vowels, and diacritics.

Tolong Siki
Abbr
Tolong_Siki
Range
11DB0 - 11DEF
Count
54

Tolong Siki is a recently encoded script used for the Southern Min language, featuring letters, diacritics, and punctuation for tonal writing.

Bengali Supplement
Abbr
Bengali_Sup
Range
11DF0 - 11DFF
Count
2

Bengali Supplement is a small collection of characters for historical documents and Sanskrit texts.

Makasar
Abbr
Makasar
Range
11EE0 - 11EFF
Count
25

Makasar is a historical script from South Sulawesi, Indonesia, used for writing the Makassarese language, encoded for modern digital text support.

Kawi
Abbr
Kawi
Range
11F00 - 11F5F
Count
87

Kawi is a script for Old Javanese, Balinese, and Sundanese texts, used historically across Southeast Asia for inscriptions and literature.

Lisu Supplement
Abbr
Lisu_Sup
Range
11FB0 - 11FBF
Count
1

Lisu Supplement is a set of additional characters for the Lisu script, mainly used for the standardized Naxi language and other minority languages in China.

Tamil Supplement
Abbr
Tamil_Sup
Range
11FC0 - 11FFF
Count
51

Tamil Supplement is a set of characters adding ancient Tamil numerals, symbols, and punctuation for historical and liturgical texts.

Cuneiform
Abbr
Cuneiform
Range
12000 - 123FF
Count
922

Cuneiform is a historic script block encoding Sumerian, Akkadian, and other ancient Middle Eastern wedge-shaped signs for scholarly text processing.

Cuneiform Numbers and Punctuation
Abbr
Cuneiform_Numbers
Range
12400 - 1247F
Count
128

Cuneiform Numbers and Punctuation is a collection of ancient Mesopotamian numeric signs, fractions, and scribal marks for transliterating cuneiform texts.

Early Dynastic Cuneiform
Abbr
Early_Dynastic_Cuneiform
Range
12480 - 1254F
Count
196

Early Dynastic Cuneiform is a script segment housing ancient Sumerian signs from the Early Dynastic period, including logograms and syllabic values.

Archaic Cuneiform Numerals
Abbr
Archaic_Cuneiform_Numerals
Range
12550 - 1268F
Count
311

Archaic Cuneiform Numerals is a collection of ancient numerical signs from fourth and third millennium BCE Mesopotamia.

Cypro-Minoan
Abbr
Cypro_Minoan
Range
12F90 - 12FFF
Count
99

Cypro-Minoan is a script block encoding syllabic signs from ancient Cyprus, used for inscriptions dating to the Late Bronze Age.

Egyptian Hieroglyphs
Abbr
Egyptian_Hieroglyphs
Range
13000 - 1342F
Count
1072

Egyptian Hieroglyphs is the encoded set of over 1,000 ancient signs for digital text, enabling modern use and research of the writing system.

Egyptian Hieroglyph Format Controls
Abbr
Egyptian_Hieroglyph_Format_Controls
Range
13430 - 1345F
Count
38

Egyptian Hieroglyph Format Controls is a set of signs that manage text layout, like joining, overlapping, and insertion, for proper ancient Egyptian writing.

Egyptian Hieroglyphs Extended-A
Abbr
Egyptian_Hieroglyphs_Ext_A
Range
13460 - 143FF
Count
3995

Egyptian Hieroglyphs Extended-A is a set of over 3,900 signs from Old Egyptian, filling gaps in the core block for fuller text encoding.

Anatolian Hieroglyphs
Abbr
Anatolian_Hieroglyphs
Range
14400 - 1467F
Count
583

Anatolian Hieroglyphs is a script encoding the Luwian language’s pictorial signs, used in Bronze Age Anatolia for monumental inscriptions.

Gurung Khema
Abbr
Gurung_Khema
Range
16100 - 1613F
Count
58

Gurung Khema is a script used for the Tamu people’s language, with characters for vowels, consonants, and digits.

Bamum Supplement
Abbr
Bamum_Sup
Range
16800 - 16A3F
Count
569

Bamum Supplement is a set of additional glyphs for the Bamum script, filling gaps in the historical syllabary for modern usage.

Mro
Abbr
Mro
Range
16A40 - 16A6F
Count
43

Mro is a script used for the Mru language of Bangladesh and Myanmar, consisting of letters, digits, and punctuation marks.

Tangsa
Abbr
Tangsa
Range
16A70 - 16ACF
Count
89

Tangsa is a script used for writing the Tangsa languages of Northeast India, with characters for consonants, vowels, and tone marks.

Bassa Vah
Abbr
Bassa_Vah
Range
16AD0 - 16AFF
Count
36

Bassa Vah is a script for writing the Bassa language, with characters for consonants, vowels, and tone marks.

Pahawh Hmong
Abbr
Pahawh_Hmong
Range
16B00 - 16B8F
Count
127

Pahawh Hmong is a script for writing the Hmong language, featuring distinct consonant and vowel signs with unique tonal markers.

Kirat Rai
Abbr
Kirat_Rai
Range
16D40 - 16D7F
Count
58

Kirat Rai is a script used for writing the Kirat-Khambu languages, including its consonants, vowels, and diacritical marks, assigned for modern digital text.

Medefaidrin
Abbr
Medefaidrin
Range
16E40 - 16E9F
Count
91

Medefaidrin is a script used for the West African language of the same name, featuring characters for writing its sounds.

Beria Erfe
Abbr
Beria_Erfe
Range
16EA0 - 16EDF
Count
50

Beria Erfe is a script used for the Beria language including consonants, vowels, and tones.

Miao
Abbr
Miao
Range
16F00 - 16F9F
Count
149

Miao is a script used for writing the A-Hmao language, containing letters, digits, and punctuation, with characters based on the Pollard phonetic system.

Ideographic Symbols and Punctuation
Abbr
Ideographic_Symbols
Range
16FE0 - 16FFF
Count
12

Ideographic Symbols and Punctuation is a small set of historic Chinese characters and marks used for phonetic annotation and textual reference.

Tangut
Abbr
Tangut
Range
17000 - 187FF
Count
6144

Tangut is a historical script block encoding characters from the Tangut language, used in the Western Xia dynasty, with over 6,000 glyphs.

Tangut Components
Abbr
Tangut_Components
Range
18800 - 18AFF
Count
768

Tangut Components is a set of radical-like symbols used to index the historical Tangut script in dictionaries and digital text.

Khitan Small Script
Abbr
Khitan_Small_Script
Range
18B00 - 18CFF
Count
476

Khitan Small Script is a historical writing system for the extinct Khitan language, encoded to preserve its unique logographic and phonetic characters.

Tangut Supplement
Abbr
Tangut_Sup
Range
18D00 - 18D7F
Count
33

Tangut Supplement is a set of rare Tangut ideographs, mainly variant characters and phonetic components, used for scholarly texts and historical research.

Tangut Components Supplement
Abbr
Tangut_Components_Sup
Range
18D80 - 18DFF
Count
115

Tangut Components Supplement is a set of rare radicals and subcomponents used for the Tangut script, aiding in dictionary and character analysis.

Jurchen
Abbr
Jurchen
Range
18E00 - 1919F
Count
914

Jurchen is a historical ideographic script from northeastern China used during the Jin and Ming dynasties.

Jurchen Radicals
Abbr
Jurchen_Radicals
Range
191A0 - 191DF
Count
51

Jurchen Radicals is a set of fifty-one indexing components for the Jurchen script.

Kana Extended-B
Abbr
Kana_Ext_B
Range
1AFF0 - 1AFFF
Count
13

Kana Extended-B is a small set of Hiragana letters with combining marks, used for archaic or dialectal Japanese transcriptions.

Kana Supplement
Abbr
Kana_Sup
Range
1B000 - 1B0FF
Count
256

Kana Supplement is the collection of historical and variant Japanese kana characters, including hentaigana and additional small kana forms.

Kana Extended-A
Abbr
Kana_Ext_A
Range
1B100 - 1B12F
Count
41

Kana Extended-A is a set of additional Hiragana letters, mainly archaic and obsolete forms, used for historical Japanese texts.

Small Kana Extension
Abbr
Small_Kana_Ext
Range
1B130 - 1B16F
Count
10

Small Kana Extension is a set of tiny hiragana and katakana characters used for writing small vowel sounds and glides in modern Japanese.

Nushu
Abbr
Nushu
Range
1B170 - 1B2FF
Count
396

Nushu is a script encoding the syllabic writing system historically used exclusively by women in Hunan, China, preserving their songs and stories.

Duployan
Abbr
Duployan
Range
1BC00 - 1BC9F
Count
143

Duployan is a script encoding shorthand systems, including Duployé, Pernin, and Sloan-Duployan, used for phonetic writing.

Shorthand Format Controls
Abbr
Shorthand_Format_Controls
Range
1BCA0 - 1BCAF
Count
4

Shorthand Format Controls is a set of invisible formatting characters used in shorthand notation to manage overlapping strokes and positioning within text.

Symbols for Legacy Computing Supplement
Abbr
Symbols_For_Legacy_Computing_Sup
Range
1CC00 - 1CEBF
Count
695

Symbols for Legacy Computing Supplement is a set of glyphs preserving obsolete computer symbols, including early terminal graphics and control codes.

Miscellaneous Symbols Supplement
Abbr
Misc_Symbols_Sup
Range
1CEC0 - 1CEFF
Count
53

Miscellaneous Symbols Supplement is a set of additional pictographs, arrows, and technical signs for extended typographic and symbolic use.

Znamenny Musical Notation
Abbr
Znamenny_Music
Range
1CF00 - 1CFCF
Count
185

Znamenny Musical Notation is a set of symbols for Russian Orthodox chant, used to encode neumatic notation in digital texts.

Byzantine Musical Symbols
Abbr
Byzantine_Music
Range
1D000 - 1D0FF
Count
246

Byzantine Musical Symbols is a set of neumes and signs used to notate Byzantine chant, preserving its melodic tradition in digital text.

Musical Symbols
Abbr
Music
Range
1D100 - 1D1FF
Count
256

Musical Symbols is a collection of glyphs for notation, including clefs, notes, rests, accidentals, and dynamics, used in digital music texts.

Ancient Greek Musical Notation
Abbr
Ancient_Greek_Music
Range
1D200 - 1D24F
Count
70

Ancient Greek Musical Notation is a set of symbols for notating melodies and rhythms including vocal and instrumental signs used in classical antiquity.

Musical Symbols Supplement
Abbr
Music_Sup
Range
1D250 - 1D28F
Count
50

Musical Symbols Supplement is a collection expanding notation with new flags

Kaktovik Numerals
Abbr
Kaktovik_Numerals
Range
1D2C0 - 1D2DF
Count
20

Kaktovik Numerals is a set of twenty digit symbols for the base-20 Iñupiaq counting system, designed to aid math education in Alaska.

Mayan Numerals
Abbr
Mayan_Numerals
Range
1D2E0 - 1D2FF
Count
20

Mayan Numerals is a set of twenty glyphs representing the base-20 vigesimal system, including zero, used for historical calendrical and mathematical notation.

Tai Xuan Jing Symbols
Abbr
Tai_Xuan_Jing
Range
1D300 - 1D35F
Count
87

Tai Xuan Jing Symbols is a set of glyphs representing the four binary elements and 81 tetragrams from the ancient Chinese divination text.

Counting Rod Numerals
Abbr
Counting_Rod
Range
1D360 - 1D37F
Count
25

Counting Rod Numerals is a set of ancient Chinese rod-based digits and symbols used for mathematical calculations, including zero and negative values.

Mathematical Alphanumeric Symbols
Abbr
Math_Alphanum
Range
1D400 - 1D7FF
Count
997

Mathematical Alphanumeric Symbols is a set of styled Latin and Greek letters, digits, and symbols for mathematical notation in bold, italic, and script forms.

Sutton SignWriting
Abbr
Sutton_SignWriting
Range
1D800 - 1DAAF
Count
672

Sutton SignWriting is a script for writing sign languages, featuring symbols for handshapes, movements, and facial expressions, enabling precise notation.

Miscellaneous Symbols and Arrows Extended
Abbr
Misc_Arrows_Ext
Range
1DB00 - 1DBFF
Count
29

Miscellaneous Symbols and Arrows Extended is a an extension of symbols and arrows for historical notation including Leibnizian mathematical operators.

Latin Extended-G
Abbr
Latin_Ext_G
Range
1DF00 - 1DFFF
Count
188

Latin Extended-G is a set of characters for medieval and phonetic transcriptions, including letters for African languages and historical scribal abbreviations.

Glagolitic Supplement
Abbr
Glagolitic_Sup
Range
1E000 - 1E02F
Count
38

Glagolitic Supplement is a set of characters used for archaic Glagolitic texts, including combining marks and letters for Old Church Slavonic.

Cyrillic Extended-D
Abbr
Cyrillic_Ext_D
Range
1E030 - 1E08F
Count
63

Cyrillic Extended-D is a set of historic letters and combining marks for early Cyrillic manuscripts, including abbreviations and scribal variants.

Nyiakeng Puachue Hmong
Abbr
Nyiakeng_Puachue_Hmong
Range
1E100 - 1E14F
Count
71

Nyiakeng Puachue Hmong is a script used for writing the Hmong language, designed by Cher Xiong to reflect spoken tones and sounds.

Toto
Abbr
Toto
Range
1E290 - 1E2BF
Count
31

Toto is a script used for the Toto language of northeastern India, encoded with letters and digits for writing that endangered tongue.

Wancho
Abbr
Wancho
Range
1E2C0 - 1E2FF
Count
59

Wancho is a script for the Wancho language of northeastern India containing consonants, vowels, and tone marks.

Nag Mundari
Abbr
Nag_Mundari
Range
1E4D0 - 1E4FF
Count
42

Nag Mundari is a script used for writing the Mundari language, with characters for consonants, vowels, and digits.

Ol Onal
Abbr
Ol_Onal
Range
1E5D0 - 1E5FF
Count
44

Ol Onal is a script used for the Ol Onal language of southern India, comprising letters, digits, and punctuation marks.

Tai Yo
Abbr
Tai_Yo
Range
1E6C0 - 1E6FF
Count
55

Tai Yo is a script for writing the Tai Yo language, encompassing consonants, vowels, tone marks, and punctuation.

Ethiopic Extended-B
Abbr
Ethiopic_Ext_B
Range
1E7E0 - 1E7FF
Count
28

Ethiopic Extended-B is a set of rare and historic Ethiopic syllabic signs, used for transliterating and preserving ancient Geʽez and liturgical texts.

Mende Kikakui
Abbr
Mende_Kikakui
Range
1E800 - 1E8DF
Count
213

Mende Kikakui is a script used for writing the Mende language of Sierra Leone, featuring syllabic characters with distinctive dotted and curved forms.

Adlam
Abbr
Adlam
Range
1E900 - 1E95F
Count
88

Adlam is a modern script for the Fulani language, used to write it from right to left with uppercase and lowercase letters plus digits.

Indic Siyaq Numbers
Abbr
Indic_Siyaq_Numbers
Range
1EC70 - 1ECBF
Count
68

Indic Siyaq Numbers is a set of historical numerals used in Indian accounting, including forms for fractions and units.

Ottoman Siyaq Numbers
Abbr
Ottoman_Siyaq_Numbers
Range
1ED00 - 1ED4F
Count
61

Ottoman Siyaq Numbers is a set of numerals used in Ottoman Turkish financial documents, encoding both digits and accounting-specific notation.

Arabic Mathematical Alphabetic Symbols
Abbr
Arabic_Math
Range
1EE00 - 1EEFF
Count
143

Arabic Mathematical Alphabetic Symbols is a set of distinct letterforms for expressing equations, used in scientific and technical Arabic writing.

Mahjong Tiles
Abbr
Mahjong
Range
1F000 - 1F02F
Count
44

Mahjong Tiles is a set of symbols depicting classic Chinese mahjong tiles, including winds, dragons, and suits, for digital representation.

Domino Tiles
Abbr
Domino
Range
1F030 - 1F09F
Count
100

Domino Tiles is a set of square symbols depicting all standard domino pieces, including blank faces, in a clear, game-ready style.

Playing Cards
Abbr
Playing_Cards
Range
1F0A0 - 1F0FF
Count
82

Playing Cards is a set of symbols depicting playing cards, including suits, courts, jokers, and card backs, used for digital card games.

Enclosed Alphanumeric Supplement
Abbr
Enclosed_Alphanum_Sup
Range
1F100 - 1F1FF
Count
201

Enclosed Alphanumeric Supplement is a set of circled, parenthesized, and squared numbers and letters, plus regional indicator symbols for flag emojis.

Enclosed Ideographic Supplement
Abbr
Enclosed_Ideographic_Sup
Range
1F200 - 1F2FF
Count
64

Enclosed Ideographic Supplement is a set of circled or enclosed Japanese kanji and symbols, used for annotations, ratings, and commercial signage.

Miscellaneous Symbols and Pictographs
Abbr
Misc_Pictographs
Range
1F300 - 1F5FF
Count
768

Miscellaneous Symbols and Pictographs is a collection of emoji, weather icons, and everyday objects, curated for visual communication and playful expression.

Emoticons
Abbr
Emoticons
Range
1F600 - 1F64F
Count
80

Emoticons is a set of pictographic faces and gestures, including smileys, frowns, and winks, used to convey emotions in digital text.

Ornamental Dingbats
Abbr
Ornamental_Dingbats
Range
1F650 - 1F67F
Count
48

Ornamental Dingbats is a set of sixty decorative flourishes, leaf motifs, and curved ornament fragments used for typographic embellishment and page decoration.

Transport and Map Symbols
Abbr
Transport_And_Map
Range
1F680 - 1F6FF
Count
120

Transport and Map Symbols is a set of icons for vehicles, traffic signs, and map-related imagery, used mainly in digital communication and wayfinding.

Alchemical Symbols
Abbr
Alchemical
Range
1F700 - 1F77F
Count
128

Alchemical Symbols is a set of icons for substances, processes, and equipment, based on historical alchemy texts and used in chemistry-related contexts.

Geometric Shapes Extended
Abbr
Geometric_Shapes_Ext
Range
1F780 - 1F7FF
Count
120

Geometric Shapes Extended is a collection of symbols, including squares, circles, triangles, and stars, used for diagrams, games, and visual notation.

Supplemental Arrows-C
Abbr
Sup_Arrows_C
Range
1F800 - 1F8FF
Count
171

Supplemental Arrows-C is a set of mostly heavy and wide arrow symbols, including curved, dashed, and paired variants, for technical and mathematical notation.

Supplemental Symbols and Pictographs
Abbr
Sup_Symbols_And_Pictographs
Range
1F900 - 1F9FF
Count
256

Supplemental Symbols and Pictographs is a set of emoji-like icons covering sports, medical, and fantasy themes, filling gaps beyond earlier symbol sets.

Chess Symbols
Abbr
Chess_Symbols
Range
1FA00 - 1FA6F
Count
102

Chess Symbols is a collection of standardized glyphs for chess pieces, boards, and related notation, supporting digital play and annotation.

Symbols and Pictographs Extended-A
Abbr
Symbols_And_Pictographs_Ext_A
Range
1FA70 - 1FAFF
Count
128

Symbols and Pictographs Extended-A is a collection of symbols supplementing existing emoji, encompassing toys, games, and additional pictographic variants.

Symbols for Legacy Computing
Abbr
Symbols_For_Legacy_Computing
Range
1FB00 - 1FBFF
Count
250

Symbols for Legacy Computing is a set of glyphs for retro terminals, fonts, and early home computer graphics, preserving historical screen output.

CJK Unified Ideographs Extension B
Abbr
CJK_Ext_B
Range
20000 - 2A6DF
Count
42720

CJK Unified Ideographs Extension B is a large set of over 42,000 rare and historical Chinese characters, added to support ancient texts and lesser‑known usage.

CJK Unified Ideographs Extension C
Abbr
CJK_Ext_C
Range
2A700 - 2B73F
Count
4160

CJK Unified Ideographs Extension C is a set of rare and historical Chinese characters, adding over 4,000 entries mainly for uncommon names and classical texts.

CJK Unified Ideographs Extension D
Abbr
CJK_Ext_D
Range
2B740 - 2B81F
Count
223

CJK Unified Ideographs Extension D is a set of 222 rare Chinese characters covering ancient and uncommon usage.

CJK Unified Ideographs Extension E
Abbr
CJK_Ext_E
Range
2B820 - 2CEAF
Count
5774

CJK Unified Ideographs Extension E is a set of over 5,700 rare Chinese characters, primarily for historical and dialectal usage, encoded in Unicode.

CJK Unified Ideographs Extension F
Abbr
CJK_Ext_F
Range
2CEB0 - 2EBEF
Count
7473

CJK Unified Ideographs Extension F is a set of rare and historical Chinese primarily for ancient texts and dialectal usage.

CJK Unified Ideographs Extension I
Abbr
CJK_Ext_I
Range
2EBF0 - 2EE5F
Count
622

CJK Unified Ideographs Extension I is a set of rare Chinese characters for historical and dialectal usage.

CJK Compatibility Ideographs Supplement
Abbr
CJK_Compat_Ideographs_Sup
Range
2F800 - 2FA1F
Count
542

CJK Compatibility Ideographs Supplement is a set of rare and alternate Chinese characters, mainly for historical texts, mapped to unified CJK ideographs.

CJK Unified Ideographs Extension G
Abbr
CJK_Ext_G
Range
30000 - 3134F
Count
4939

CJK Unified Ideographs Extension G is a set of over 4,900 rare Chinese characters, primarily for historical and dialectal usage.

CJK Unified Ideographs Extension H
Abbr
CJK_Ext_H
Range
31350 - 323AF
Count
4192

CJK Unified Ideographs Extension H is a collection of rare and historical Chinese characters, primarily for ancient texts and dialectal use.

CJK Unified Ideographs Extension J
Abbr
CJK_Ext_J
Range
323B0 - 3347F
Count
4298

CJK Unified Ideographs Extension J is a set of rare, historical Chinese characters added to Unicode for encoding ancient and uncommon texts.

Seal
Abbr
Seal
Range
3D000 - 3FC3F
Count
11328

Seal is a historical script that served as a precursor to modern Han ideographs.

Tags
Abbr
Tags
Range
E0000 - E007F
Count
97

Tags is a set of invisible, deprecated format characters used to tag text with language or script metadata, now largely obsolete.

Variation Selectors Supplement
Abbr
VS_Sup
Range
E0100 - E01EF
Count
240

Variation Selectors Supplement is a set of invisible format characters that refine glyph presentation for ideographs, emoji, and other complex scripts.

Supplementary Private Use Area-A
Abbr
Sup_PUA_A
Range
F0000 - FFFFF
Count
65534

Supplementary Private Use Area-A is a reserved region for private character assignments, ensuring no standard Unicode meanings exist there.

Supplementary Private Use Area-B
Abbr
Sup_PUA_B
Range
100000 - 10FFFF
Count
65534

Supplementary Private Use Area-B is reserved for user-defined characters, with no assigned glyphs, allowing private interchange outside standard encoding.

No Block
Abbr
NB
Range
-
Count
0

No Block is a placeholder label for characters not assigned to any named grouping.