Unicode Blocks
Unicode blocks are contiguous ranges of code points that group together characters sharing a common script, set of symbols, or thematic purpose for organized encoding.
- Abbr
- ASCII
- Range
- 0000 - 007F
- Count
- 128
Basic Latin is the foundational character set covering ASCII letters, digits, punctuation, and controls, ensuring universal text compatibility.
- Abbr
- Latin_1_Sup
- Range
- 0080 - 00FF
- Count
- 128
Latin-1 Supplement is the part of Unicode covering Western European letters, punctuation, and symbols, plus C1 control characters, extending ASCII’s coverage.
- Abbr
- Latin_Ext_A
- Range
- 0100 - 017F
- Count
- 128
Latin Extended-A is a set of precomposed letters with diacritics, primarily for Eastern European and Vietnamese languages, enabling full text representation.
- Abbr
- Latin_Ext_B
- Range
- 0180 - 024F
- Count
- 208
Latin Extended-B is a set of additional Latin letters and symbols, covering phonetic, historical, and African orthographic needs.
- Abbr
- IPA_Ext
- Range
- 0250 - 02AF
- Count
- 96
IPA Extensions is a set of additional phonetic symbols that supplements the International Phonetic Alphabet for representing speech sounds.
- Abbr
- Modifier_Letters
- Range
- 02B0 - 02FF
- Count
- 80
Spacing Modifier Letters is a set of phonetic symbols that adjust the meaning of preceding letters, including aspiration, palatalization, and tone marks.
- Abbr
- Diacriticals
- Range
- 0300 - 036F
- Count
- 112
Combining Diacritical Marks is a set of characters that attach to preceding letters, enabling accents and phonetic distinctions across many writing systems.
- Abbr
- Greek
- Range
- 0370 - 03FF
- Count
- 135
Greek and Coptic is the Unicode segment housing letters for modern Greek, ancient Greek, and Coptic, plus punctuation and variant forms.
- Abbr
- Cyrillic
- Range
- 0400 - 04FF
- Count
- 256
Cyrillic is a script block covering standard letters for Russian, Ukrainian, Bulgarian, and other Slavic languages, plus historic and non-Slavic extensions.
- Abbr
- Cyrillic_Sup
- Range
- 0500 - 052F
- Count
- 48
Cyrillic Supplement is a set of letters for minority languages, plus historic forms, extending the main Cyrillic script.
- Abbr
- Armenian
- Range
- 0530 - 058F
- Count
- 94
Armenian is a script block for writing the Armenian language, including its alphabet, punctuation, and ligatures, used across Armenia and the diaspora.
- Abbr
- Hebrew
- Range
- 0590 - 05FF
- Count
- 90
Hebrew is a contiguous set of characters for writing Hebrew, Yiddish, and Ladino, including letters, niqqud vowel marks, and cantillation signs.
- Abbr
- Arabic
- Range
- 0600 - 06FF
- Count
- 256
Arabic is a script block encoding letters, diacritics, and punctuation for writing Arabic, Persian, Urdu, and other languages using Arabic-derived scripts.
- Abbr
- Syriac
- Range
- 0700 - 074F
- Count
- 77
Syriac is a script block for writing the Syriac language, including its classical, Estrangela, and Eastern variants, plus diacritics and punctuation.
- Abbr
- Arabic_Sup
- Range
- 0750 - 077F
- Count
- 48
Arabic Supplement is a set of extra letters for early Quranic and African languages, filling gaps in standard Arabic script.
- Abbr
- Thaana
- Range
- 0780 - 07BF
- Count
- 50
Thaana is a script used for writing the Dhivehi language of the Maldives, featuring letters with right-to-left orientation and diacritical vowel marks.
- Abbr
- NKo
- Range
- 07C0 - 07FF
- Count
- 62
NKo is a script used for the Manding languages of West Africa, added to encode text in a single, unified writing system.
- Abbr
- Samaritan
- Range
- 0800 - 083F
- Count
- 61
Samaritan is a script used for writing the Samaritan Hebrew and Aramaic liturgical texts, including the Pentateuch and prayers.
- Abbr
- Mandaic
- Range
- 0840 - 085F
- Count
- 29
Mandaic is a script block used for writing the Mandaean language, containing letters and diacritics in a right-to-left direction.
- Abbr
- Syriac_Sup
- Range
- 0860 - 086F
- Count
- 11
Syriac Supplement is a set of nine characters adding letters for Sogdian and additional Syriac orthography needs.
- Abbr
- Arabic_Ext_B
- Range
- 0870 - 089F
- Count
- 43
Arabic Extended-B is a set of Arabic script characters covering Quranic annotation marks, added letters for African languages, and editorial symbols.
- Abbr
- Arabic_Ext_A
- Range
- 08A0 - 08FF
- Count
- 96
Arabic Extended-A is a character set for additional letters and signs used in Arabic-based orthographies across Africa and Asia, plus Quranic annotation marks.
- Abbr
- Devanagari
- Range
- 0900 - 097F
- Count
- 128
Devanagari is a script used for Hindi, Marathi, Nepali, and Sanskrit, containing letters, vowels, diacritics, and punctuation marks.
- Abbr
- Bengali
- Range
- 0980 - 09FF
- Count
- 96
Bengali is a script block encoding the Bengali-Assamese alphabet, used for Bangla, Assamese, and related languages, plus digits and diacritics.
- Abbr
- Gurmukhi
- Range
- 0A00 - 0A7F
- Count
- 80
Gurmukhi is a script used for writing Punjabi, containing letters, vowel signs, and punctuation marks integral to the language's religious and literary texts.
- Abbr
- Gujarati
- Range
- 0A80 - 0AFF
- Count
- 91
Gujarati is a script used for writing the Gujarati language, covering letters, digits, and diacritical marks essential for its phonetics.
- Abbr
- Oriya
- Range
- 0B00 - 0B7F
- Count
- 93
Oriya is a script block for writing the Odia language, containing vowel signs, consonants, digits, and punctuation marks.
- Abbr
- Tamil
- Range
- 0B80 - 0BFF
- Count
- 72
Tamil is an abugida script used for writing the Tamil language, encoded with characters for vowels, consonants, and diacritical marks, plus numerals and signs.
- Abbr
- Telugu
- Range
- 0C00 - 0C7F
- Count
- 101
Telugu is a script used for writing the Telugu language, containing letters, vowels, diacritics, digits, and punctuation marks.
- Abbr
- Kannada
- Range
- 0C80 - 0CFF
- Count
- 92
Kannada is a script block used for writing the Kannada language, including its vowels, consonants, and diacritical marks.
- Abbr
- Malayalam
- Range
- 0D00 - 0D7F
- Count
- 118
Malayalam is a script block for writing the Dravidian language of Kerala, including vowels, consonants, and conjunct forms.
- Abbr
- Sinhala
- Range
- 0D80 - 0DFF
- Count
- 91
Sinhala is the script used for writing the Sinhalese language of Sri Lanka, encompassing vowels, consonants, and diacritical marks.
- Abbr
- Thai
- Range
- 0E00 - 0E7F
- Count
- 87
Thai is the script for the Thai language, including vowels, tone marks, digits, and punctuation, used in modern and historical texts.
- Abbr
- Lao
- Range
- 0E80 - 0EFF
- Count
- 83
Lao is a script block for the Lao language, containing consonants, vowels, tone marks, and digits used in Laos.
- Abbr
- Tibetan
- Range
- 0F00 - 0FFF
- Count
- 211
Tibetan is a script set for writing the Tibetan language, including letters, numerals, and punctuation marks essential for religious and literary texts.
- Abbr
- Myanmar
- Range
- 1000 - 109F
- Count
- 160
Myanmar is the script block for writing Burmese, including letters, digits, punctuation, and tone marks used in the Myanmar language.
- Abbr
- Georgian
- Range
- 10A0 - 10FF
- Count
- 88
Georgian is a set of characters for the Georgian script, covering modern letters, Mtavruli capitals, and punctuation.
- Abbr
- Jamo
- Range
- 1100 - 11FF
- Count
- 256
Hangul Jamo is a set of precomposed letters for the Korean alphabet’s consonant and vowel components, enabling syllable construction.
- Abbr
- Ethiopic
- Range
- 1200 - 137F
- Count
- 358
Ethiopic is a script block for Ge’ez and modern Ethiopic languages, used to write Amharic, Tigrinya, and others.
- Abbr
- Ethiopic_Sup
- Range
- 1380 - 139F
- Count
- 26
Ethiopic Supplement is a set of additional syllabic and punctuation marks for medieval and modern Geʽez script usage.
- Abbr
- Cherokee
- Range
- 13A0 - 13FF
- Count
- 92
Cherokee is a syllabary script used for the Cherokee language, covering syllables, punctuation, and historical letterforms in its encoded character set.
- Abbr
- UCAS
- Range
- 1400 - 167F
- Count
- 640
Unified Canadian Aboriginal Syllabics is a script encoding Cree, Inuktitut, and other Indigenous languages, with letters rotated for syllable variations.
- Abbr
- Ogham
- Range
- 1680 - 169F
- Count
- 29
Ogham is a script block encoding the ancient Irish alphabet’s linear inscriptions, used for monumental stones and later manuscripts.
- Abbr
- Runic
- Range
- 16A0 - 16FF
- Count
- 89
Runic is a script encoding ancient Germanic alphabets, featuring letters for Elder and Younger Futhark, Anglo-Saxon variants, and punctuation marks.
- Abbr
- Tagalog
- Range
- 1700 - 171F
- Count
- 23
Tagalog is a script block for precolonial Philippine Baybayin writing, used historically for the Tagalog language and revived in modern cultural contexts.
- Abbr
- Hanunoo
- Range
- 1720 - 173F
- Count
- 23
Hanunoo is a script block for writing the Hanunoo language of the Philippines, including syllabic characters and punctuation marks.
- Abbr
- Buhid
- Range
- 1740 - 175F
- Count
- 20
Buhid is a script block used for writing the Buhid language, containing letters, vowels, and punctuation marks for traditional Philippine texts.
- Abbr
- Tagbanwa
- Range
- 1760 - 177F
- Count
- 18
Tagbanwa is a script used for the Tagbanwa language, with letters for syllables and punctuation, encoded for digital text.
- Abbr
- Khmer
- Range
- 1780 - 17FF
- Count
- 114
Khmer is a script block for writing the Cambodian language, including consonants, vowels, diacritics, and punctuation, used across official and religious texts.
- Abbr
- Mongolian
- Range
- 1800 - 18AF
- Count
- 158
Mongolian is a script for writing the Mongolian language, including letters, digits, punctuation, and marks for historical and modern use.
- Abbr
- UCAS_Ext
- Range
- 18B0 - 18FF
- Count
- 70
Unified Canadian Aboriginal Syllabics Extended is a set of supplementary symbols for writing additional Cree, Inuktitut, and other Indigenous languages.
- Abbr
- Limbu
- Range
- 1900 - 194F
- Count
- 68
Limbu is a script block for writing the Limbu language, containing letters, digits, and punctuation used primarily in Nepal and Sikkim.
- Abbr
- Tai_Le
- Range
- 1950 - 197F
- Count
- 35
Tai Le is a script used for the Tai Nüa language, containing consonants, vowels, and tone marks.
- Abbr
- New_Tai_Lue
- Range
- 1980 - 19DF
- Count
- 83
New Tai Lue is a script used for the Tai Lue language, featuring letters, digits, and punctuation, with modernized forms for vowel and tone marks.
- Abbr
- Khmer_Symbols
- Range
- 19E0 - 19FF
- Count
- 32
Khmer Symbols is a collection of signs used in lunar calendar calculations, including dates, times, and astronomical notations.
- Abbr
- Buginese
- Range
- 1A00 - 1A1F
- Count
- 30
Buginese is a script used for writing the Bugis language, primarily in South Sulawesi, Indonesia, featuring consonants and vowel signs.
- Abbr
- Tai_Tham
- Range
- 1A20 - 1AAF
- Count
- 127
Tai Tham is a script for writing Northern Thai, Lao, and related languages, featuring consonants, vowels, and diacritics in a rounded style.
- Abbr
- Diacriticals_Ext
- Range
- 1AB0 - 1AFF
- Count
- 65
Combining Diacritical Marks Extended is a set of additional marks for modifying letters, supporting diverse writing systems and scholarly transcription needs.
- Abbr
- Balinese
- Range
- 1B00 - 1B7F
- Count
- 127
Balinese is a script block for writing the Balinese language, containing consonants, vowels, and diacritics used in Bali, Indonesia.
- Abbr
- Sundanese
- Range
- 1B80 - 1BBF
- Count
- 64
Sundanese is a script block for writing the Sundanese language, containing consonants, vowels, digits, and punctuation marks.
- Abbr
- Batak
- Range
- 1BC0 - 1BFF
- Count
- 56
Batak is a script block used for writing the Batak languages of Sumatra, containing consonants, vowels, and punctuation marks.
- Abbr
- Lepcha
- Range
- 1C00 - 1C4F
- Count
- 74
Lepcha is a script used for writing the Lepcha language of Sikkim, India, including its letters, diacritics, and punctuation.
- Abbr
- Ol_Chiki
- Range
- 1C50 - 1C7F
- Count
- 48
Ol Chiki is a script for the Santali language, featuring letters, digits, and punctuation, designed to write its phonology clearly.
- Abbr
- Cyrillic_Ext_C
- Range
- 1C80 - 1C8F
- Count
- 11
Cyrillic Extended-C is a small set of letters for early Slavic orthographies, including historic forms of Che and Dzhe.
- Abbr
- Georgian_Ext
- Range
- 1C90 - 1CBF
- Count
- 46
Georgian Extended is a set of additional letters for the Georgian script, mainly used for the historical Mtavruli uppercase forms.
- Abbr
- Sundanese_Sup
- Range
- 1CC0 - 1CCF
- Count
- 8
Sundanese Supplement is a set of punctuation marks and signs used for writing the Sundanese script, filling gaps left by the main Sundanese block.
- Abbr
- Vedic_Ext
- Range
- 1CD0 - 1CFF
- Count
- 43
Vedic Extensions is a set of diacritical and tonal marks for Sanskrit and other Indo-Aryan texts, aiding precise pronunciation and Vedic chanting.
- Abbr
- Phonetic_Ext
- Range
- 1D00 - 1D7F
- Count
- 128
Phonetic Extensions is a set of historical and dialectal letters for transcribing sounds not covered by standard alphabets, aiding linguists and phoneticians.
- Abbr
- Phonetic_Ext_Sup
- Range
- 1D80 - 1DBF
- Count
- 64
Phonetic Extensions Supplement is a set of modified Latin letters and symbols for precise phonetic transcription, especially for obscure sounds.
- Abbr
- Diacriticals_Sup
- Range
- 1DC0 - 1DFF
- Count
- 64
Combining Diacritical Marks Supplement is a set of symbols for adding phonetic or tone details to letters, spanning medieval and modern linguistic notation.
- Abbr
- Latin_Ext_Additional
- Range
- 1E00 - 1EFF
- Count
- 256
Latin Extended Additional is a set of Latin letters with diacritics for Vietnamese and other languages, plus medievalist and phonetic additions.
- Abbr
- Greek_Ext
- Range
- 1F00 - 1FFF
- Count
- 233
Greek Extended is a set of precomposed accented Greek letters, used for classical text and polytonic orthography.
- Abbr
- Punctuation
- Range
- 2000 - 206F
- Count
- 111
General Punctuation is a set of symbols for spacing, dashes, quotation marks, and invisible format controls, aiding text layout and clarity.
- Abbr
- Super_And_Sub
- Range
- 2070 - 209F
- Count
- 46
Superscripts and Subscripts is a set of characters for scientific notation, math, and linguistic use, including superscript digits, subscripts, and modifiers.
- Abbr
- Currency_Symbols
- Range
- 20A0 - 20CF
- Count
- 37
Currency Symbols is a set of standardized signs for global currencies, including the euro, rupee, and lira, used in digital text.
- Abbr
- Diacriticals_For_Symbols
- Range
- 20D0 - 20FF
- Count
- 33
Combining Diacritical Marks for Symbols is a set of marks that overlay mathematical, technical, or currency symbols to modify their meaning.
- Abbr
- Letterlike_Symbols
- Range
- 2100 - 214F
- Count
- 80
Letterlike Symbols is a set of typographic and mathematical signs that resemble letters, including units, abbreviations, and numerals.
- Abbr
- Number_Forms
- Range
- 2150 - 218F
- Count
- 60
Number Forms is a set of characters for fractions, roman numerals, and currency symbols, enabling compact representation of numeric concepts in text.
- Abbr
- Arrows
- Range
- 2190 - 21FF
- Count
- 112
Arrows is a collection of symbols for directional indicators, including simple, double, and curved variants, used in math, UI, and logic.
- Abbr
- Math_Operators
- Range
- 2200 - 22FF
- Count
- 256
Mathematical Operators is a collection of symbols for arithmetic, set theory, logic, and related math notation.
- Abbr
- Misc_Technical
- Range
- 2300 - 23FF
- Count
- 256
Miscellaneous Technical is a collection of symbols for engineering, computing, and control functions, including arrows, keys, and notation like ⌈⌉ and ⌊⌋.
- Abbr
- Control_Pictures
- Range
- 2400 - 243F
- Count
- 42
Control Pictures is a set of symbols visually representing ASCII control characters used for diagrams and debugging.
- Abbr
- OCR
- Range
- 2440 - 245F
- Count
- 11
Optical Character Recognition is a set of symbols for scanned text correction, including marks like check, erase, and hold.
- Abbr
- Enclosed_Alphanum
- Range
- 2460 - 24FF
- Count
- 160
Enclosed Alphanumerics is a set of circled, parenthesized, and full-stop numerals and letters used for lists, footnotes, and decorative numbering.
- Abbr
- Box_Drawing
- Range
- 2500 - 257F
- Count
- 128
Box Drawing is a set of characters for creating simple line and box diagrams, including horizontal, vertical, and corner pieces with various line styles.
- Abbr
- Block_Elements
- Range
- 2580 - 259F
- Count
- 32
Block Elements is a set of 32 graphic symbols for creating shaded, hatched, and partial box-drawing patterns, primarily used in text-based interfaces.
- Abbr
- Geometric_Shapes
- Range
- 25A0 - 25FF
- Count
- 96
Geometric Shapes is a set of symbols for squares, circles, triangles, and arrows, used in diagrams and text.
- Abbr
- Misc_Symbols
- Range
- 2600 - 26FF
- Count
- 256
Miscellaneous Symbols is a collection of diverse pictographs, including weather icons, chess pieces, and warning signs, used across digital texts.
- Abbr
- Dingbats
- Range
- 2700 - 27BF
- Count
- 192
Dingbats is a set of ornamental symbols, including stars, arrows, crosses, and graphic flourishes, used for decoration and visual emphasis.
- Abbr
- Misc_Math_Symbols_A
- Range
- 27C0 - 27EF
- Count
- 48
Miscellaneous Mathematical Symbols-A is a collection of operators, angles, and notation for equations, including the lozenge, sum, and integral variants.
- Abbr
- Sup_Arrows_A
- Range
- 27F0 - 27FF
- Count
- 16
Supplemental Arrows-A is a set of arrow symbols for mathematical notation, including long arrows and bent arrows, designed for technical contexts.
- Abbr
- Braille
- Range
- 2800 - 28FF
- Count
- 256
Braille Patterns is a set of tactile cell combinations, mapping six and eight dot systems to text, math, and music notation.
- Abbr
- Sup_Arrows_B
- Range
- 2900 - 297F
- Count
- 128
Supplemental Arrows-B is a set of varied arrow symbols, including curved, zigzag, and feathered variants, used for technical notation and mathematics.
- Abbr
- Misc_Math_Symbols_B
- Range
- 2980 - 29FF
- Count
- 128
Miscellaneous Mathematical Symbols-B is a set of symbols for advanced math, including operators, arrows, and geometric shapes, aiding technical notation.
- Abbr
- Sup_Math_Operators
- Range
- 2A00 - 2AFF
- Count
- 256
Supplemental Mathematical Operators is a collection of additional symbols for advanced algebra, logic, and set theory, extending standard math notation.
- Abbr
- Misc_Arrows
- Range
- 2B00 - 2BFF
- Count
- 254
Miscellaneous Symbols and Arrows is a set of diverse icons, including arrows, shapes, and weather symbols, for technical and casual use.
- Abbr
- Glagolitic
- Range
- 2C00 - 2C5F
- Count
- 96
Glagolitic is a historical script block used for writing Old Church Slavonic, containing uppercase and lowercase letters plus punctuation marks.
- Abbr
- Latin_Ext_C
- Range
- 2C60 - 2C7F
- Count
- 32
Latin Extended-C is a set of additional Latin letters and symbols used for phonetic transcription and minority languages.
- Abbr
- Coptic
- Range
- 2C80 - 2CFF
- Count
- 123
Coptic is a script block used for writing the Coptic language, derived from Greek with added letters for Egyptian sounds.
- Abbr
- Georgian_Sup
- Range
- 2D00 - 2D2F
- Count
- 40
Georgian Supplement is a set of lowercase Georgian letters, derived from the ancient Asomtavruli script, used for ecclesiastical texts and historical writing.
- Abbr
- Tifinagh
- Range
- 2D30 - 2D7F
- Count
- 59
Tifinagh is a script used for Berber languages, with its block containing letters for Tamazight, Tuareg, and other variants, plus punctuation and digits.
- Abbr
- Ethiopic_Ext
- Range
- 2D80 - 2DDF
- Count
- 79
Ethiopic Extended is a set of supplementary syllabic and punctuation signs for Ge'ez, supporting historical and modern liturgical texts.
- Abbr
- Cyrillic_Ext_A
- Range
- 2DE0 - 2DFF
- Count
- 32
Cyrillic Extended-A is a set of combining marks used for early Slavic orthographies, enabling accurate representation of historical texts.
- Abbr
- Sup_Punctuation
- Range
- 2E00 - 2E7F
- Count
- 98
Supplemental Punctuation is a set of rare punctuation marks, including historic and editorial symbols, added to support specialized text and scholarly use.
- Abbr
- CJK_Radicals_Sup
- Range
- 2E80 - 2EFF
- Count
- 115
CJK Radicals Supplement is a set of alternate forms of Chinese characters’ radicals, used in dictionaries and indexing, complementing the main Kangxi Radicals.
- Abbr
- Kangxi
- Range
- 2F00 - 2FDF
- Count
- 214
Kangxi Radicals is a set of historical Chinese character components, arranged in their traditional dictionary order, used for indexing and referencing.
- Abbr
- IDC
- Range
- 2FF0 - 2FFF
- Count
- 16
Ideographic Description Characters is a set of sixteen symbols used to visually describe the composition of complex Chinese characters from simpler components.
- Abbr
- CJK_Symbols
- Range
- 3000 - 303F
- Count
- 64
CJK Symbols and Punctuation is a set of ideographic punctuation marks, iteration signs, and spacing characters used in Chinese, Japanese, and Korean writing.
- Abbr
- Hiragana
- Range
- 3040 - 309F
- Count
- 93
Hiragana is a Japanese syllabary of characters used for native words, grammar, and furigana, including voiced marks and the archaic "wi" and "we".
- Abbr
- Katakana
- Range
- 30A0 - 30FF
- Count
- 96
Katakana is a Japanese syllabary of characters used primarily for foreign loanwords, onomatopoeia, and phonetic emphasis.
- Abbr
- Bopomofo
- Range
- 3100 - 312F
- Count
- 43
Bopomofo is a set of phonetic symbols for Mandarin Chinese, used to annotate pronunciation, especially in Taiwan, within a dedicated character range.
- Abbr
- Compat_Jamo
- Range
- 3130 - 318F
- Count
- 94
Hangul Compatibility Jamo is a set of precomposed Korean letters for legacy encoding and vertical text, preserving old Hangul syllables and digraphs.
- Abbr
- Kanbun
- Range
- 3190 - 319F
- Count
- 16
Kanbun is a set of annotation marks used in Japanese texts to indicate reading order for classical Chinese, aiding in translation.
- Abbr
- Bopomofo_Ext
- Range
- 31A0 - 31BF
- Count
- 32
Bopomofo Extended is a small set of phonetic symbols adding rare and dialectal Mandarin sounds, plus the unique characters for Hokkien and Hakka.
- Abbr
- CJK_Strokes
- Range
- 31C0 - 31EF
- Count
- 39
CJK Strokes is a set of standardized symbols for writing Chinese, Japanese, and Korean brushstroke components, used in dictionaries and calligraphy references.
- Abbr
- Katakana_Ext
- Range
- 31F0 - 31FF
- Count
- 16
Katakana Phonetic Extensions is a set of small marks for Ainu and other languages, used to modify katakana sounds like p, t, or s.
- Abbr
- Enclosed_CJK
- Range
- 3200 - 32FF
- Count
- 255
Enclosed CJK Letters and Months is a set of symbols enclosing Korean syllables, Japanese katakana, and month names within circles or parentheses.
- Abbr
- CJK_Compat
- Range
- 3300 - 33FF
- Count
- 256
CJK Compatibility is a set of precomposed Japanese Katakana squares, like ㌔, and other CJK symbols for legacy character mapping and vertical text layouts.
- Abbr
- CJK_Ext_A
- Range
- 3400 - 4DBF
- Count
- 6592
CJK Unified Ideographs Extension A is a collection of rare and archaic Chinese characters, added to support historical texts and lesser-known names.
- Abbr
- Yijing
- Range
- 4DC0 - 4DFF
- Count
- 64
Yijing Hexagram Symbols is a set of monochrome trigram and hexagram glyphs used for casting the I Ching oracle.
- Abbr
- CJK
- Range
- 4E00 - 9FFF
- Count
- 20992
CJK Unified Ideographs is a set of over 20,000 Chinese characters shared across Chinese, Japanese, and Korean writing systems, covering common Han script usage.
- Abbr
- Yi_Syllables
- Range
- A000 - A48F
- Count
- 1165
Yi Syllables is a set of characters used for writing the Nuosu language, based on the standardized Liangshan dialect.
- Abbr
- Yi_Radicals
- Range
- A490 - A4CF
- Count
- 55
Yi Radicals is a set of symbols used to index the traditional Yi script, primarily for dictionary and educational lookup purposes.
- Abbr
- Lisu
- Range
- A4D0 - A4FF
- Count
- 48
Lisu is a script used for the Lisu language, with letters, digits, and punctuation, designed for tonal and phonetic writing.
- Abbr
- Vai
- Range
- A500 - A63F
- Count
- 300
Vai is a script used for the Vai language of Liberia, with characters for syllables, punctuation, and historical logograms.
- Abbr
- Cyrillic_Ext_B
- Range
- A640 - A69F
- Count
- 96
Cyrillic Extended-B is a set of historic letters and signs for Old Church Slavonic and early Cyrillic scripts, including abbreviations and variant forms.
- Abbr
- Bamum
- Range
- A6A0 - A6FF
- Count
- 88
Bamum is a script for the Bamum language, used for royal decrees and everyday writing in Cameroon's Bamum kingdom.
- Abbr
- Modifier_Tone_Letters
- Range
- A700 - A71F
- Count
- 32
Modifier Tone Letters is a set of phonetic symbols used for marking tone contours in linguistic transcription, primarily for African and Asian languages.
- Abbr
- Latin_Ext_D
- Range
- A720 - A7FF
- Count
- 206
Latin Extended-D is a set of Latin letters and phonetic symbols for medieval and African languages, plus historical additions.
- Abbr
- Syloti_Nagri
- Range
- A800 - A82F
- Count
- 45
Syloti Nagri is a script used for the Sylheti language, containing letters, digits, and punctuation, with characters for vowels, consonants, and diacritics.
- Abbr
- Indic_Number_Forms
- Range
- A830 - A83F
- Count
- 10
Common Indic Number Forms is a set of historical numeral signs used across various Indic scripts, primarily for accounting and monetary values.
- Abbr
- Phags_Pa
- Range
- A840 - A87F
- Count
- 56
Phags-pa is a script used historically for writing Mongolian, Chinese, and other languages, encoded for digital text with letters, marks, and punctuation.
- Abbr
- Saurashtra
- Range
- A880 - A8DF
- Count
- 82
Saurashtra is a script used for writing the Saurashtra language with letters, digits, and diacritics.
- Abbr
- Devanagari_Ext
- Range
- A8E0 - A8FF
- Count
- 32
Devanagari Extended is a set of combining marks and signs used to modify vowels and consonants in Sanskrit and other Indic languages.
- Abbr
- Kayah_Li
- Range
- A900 - A92F
- Count
- 48
Kayah Li is a script used for the Kayah language, containing letters, digits, and punctuation with tone marks.
- Abbr
- Rejang
- Range
- A930 - A95F
- Count
- 37
Rejang is a script used for writing the Rejang language of Sumatra, containing letters for consonants, vowels, and diacritics.
- Abbr
- Jamo_Ext_A
- Range
- A960 - A97F
- Count
- 29
Hangul Jamo Extended-A is a set of additional consonants and vowel jamo used for writing Middle Korean, filling gaps in the standard Hangul system.
- Abbr
- Javanese
- Range
- A980 - A9DF
- Count
- 91
Javanese is a script used for writing the Javanese language, with letters, digits, and punctuation, plus historical and modern variant forms.
- Abbr
- Myanmar_Ext_B
- Range
- A9E0 - A9FF
- Count
- 31
Myanmar Extended-B is a small set of additional letters and signs for minority languages and historical texts, complementing earlier Myanmar encoding.
- Abbr
- Cham
- Range
- AA00 - AA5F
- Count
- 83
Cham is a script used for writing the Cham language of Vietnam and Cambodia, with distinct letters for the Eastern and Western dialects.
- Abbr
- Myanmar_Ext_A
- Range
- AA60 - AA7F
- Count
- 32
Myanmar Extended-A is a set of additional Tai Laing and other Myanmar script characters for historical and minority language support.
- Abbr
- Tai_Viet
- Range
- AA80 - AADF
- Count
- 72
Tai Viet is a script used for writing the Tai Dam and Tai Don languages, featuring distinct consonants, vowels, and tone marks.
- Abbr
- Meetei_Mayek_Ext
- Range
- AAE0 - AAFF
- Count
- 23
Meetei Mayek Extensions is a supplementary set of characters for the Meitei script, adding letters and digits for historical and modern usage.
- Abbr
- Ethiopic_Ext_A
- Range
- AB00 - AB2F
- Count
- 32
Ethiopic Extended-A is a set of additional Gamo, Gofa, and Gurage syllables and punctuation, filling gaps in modern Ethiopic script usage.
- Abbr
- Latin_Ext_E
- Range
- AB30 - AB6F
- Count
- 62
Latin Extended-E is a set of characters for transcribing medievalist and dialectal letter forms, including modifier letters and additional vowels.
- Abbr
- Cherokee_Sup
- Range
- AB70 - ABBF
- Count
- 80
Cherokee Supplement is a set of lowercase syllabary characters complementing the main Cherokee block, used for writing the Cherokee language.
- Abbr
- Meetei_Mayek
- Range
- ABC0 - ABFF
- Count
- 56
Meetei Mayek is a script used for the Manipuri language, containing letters, digits, and punctuation marks for writing this Tibeto-Burman tongue.
- Abbr
- Hangul
- Range
- AC00 - D7AF
- Count
- 11172
Hangul Syllables is a precomposed set of over 11,000 Korean syllable characters enabling direct encoding of each possible syllable block.
- Abbr
- Jamo_Ext_B
- Range
- D7B0 - D7FF
- Count
- 72
Hangul Jamo Extended-B is a set of precomposed syllables, filling gaps in archaic Hangul orthography for historical and dialectal Korean texts.
- Abbr
- High_Surrogates
- Range
- D800 - DB7F
- Count
- 0
High Surrogates is a reserved range for the first half of surrogate pairs enabling encoding of supplementary characters in UTF-16.
- Abbr
- High_PU_Surrogates
- Range
- DB80 - DBFF
- Count
- 0
High Private Use Surrogates is a reserved range for private use, enabling applications to assign custom characters without standardized meanings.
- Abbr
- Low_Surrogates
- Range
- DC00 - DFFF
- Count
- 0
Low Surrogates is a reserved range for the second half of surrogate pairs, enabling encoding of supplementary characters in UTF-16.
- Abbr
- PUA
- Range
- E000 - F8FF
- Count
- 6400
Private Use Area is a reserved zone for characters defined by individual fonts or applications, not standardized for universal exchange.
- Abbr
- CJK_Compat_Ideographs
- Range
- F900 - FAFF
- Count
- 472
CJK Compatibility Ideographs is a set of mostly duplicate or variant Chinese characters, included for round‑trip compatibility with older standards.
- Abbr
- Alphabetic_PF
- Range
- FB00 - FB4F
- Count
- 58
Alphabetic Presentation Forms is a set of precomposed ligatures and letter combinations, primarily for Latin and Armenian scripts, aiding typographic rendering.
- Abbr
- Arabic_PF_A
- Range
- FB50 - FDFF
- Count
- 656
Arabic Presentation Forms-A is a set of contextual letterforms, ligatures, and honorifics used for stylized Arabic writing, spanning over 600 encoded glyphs.
- Abbr
- VS
- Range
- FE00 - FE0F
- Count
- 16
Variation Selectors is a set of formatting characters that modify the presentation style of preceding ideographs, emoji, or symbols.
- Abbr
- Vertical_Forms
- Range
- FE10 - FE1F
- Count
- 10
Vertical Forms is a set of ten small-width punctuation marks, including commas and parentheses, used for vertical text layout in East Asian typography.
- Abbr
- Half_Marks
- Range
- FE20 - FE2F
- Count
- 16
Combining Half Marks is a set of diacritical signs that attach to preceding letters, enabling precise phonetic and scholarly notation.
- Abbr
- CJK_Compat_Forms
- Range
- FE30 - FE4F
- Count
- 32
CJK Compatibility Forms is a small set of vertical presentation forms for East Asian punctuation and symbols, used mainly for legacy vertical text layout.
- Abbr
- Small_Forms
- Range
- FE50 - FE6F
- Count
- 26
Small Form Variants is a set of compact punctuation and symbols, used mostly for vertical Chinese and Japanese text layout, mirroring fullwidth forms.
- Abbr
- Arabic_PF_B
- Range
- FE70 - FEFF
- Count
- 141
Arabic Presentation Forms-B is a set of contextual Arabic letter forms and diacritics, including the zero-width no-break space, used for legacy compatibility.
- Abbr
- Half_And_Full_Forms
- Range
- FF00 - FFEF
- Count
- 225
Halfwidth and Fullwidth Forms is a set of characters used to align East Asian and Latin text, providing fullwidth versions of ASCII and halfwidth katakana.
- Abbr
- Specials
- Range
- FFF0 - FFFF
- Count
- 5
Specials is a set of reserved and noncharacter code points, including the replacement character, used for internal processing and compatibility.
- Abbr
- Linear_B_Syllabary
- Range
- 10000 - 1007F
- Count
- 88
Linear B Syllabary is a script encoding ancient Mycenaean Greek syllables, used for administrative records on clay tablets in Bronze Age Crete and Greece.
- Abbr
- Linear_B_Ideograms
- Range
- 10080 - 100FF
- Count
- 123
Linear B Ideograms is a set of signs for syllabic and ideographic writing, used to record Mycenaean Greek economic and administrative texts.
- Abbr
- Aegean_Numbers
- Range
- 10100 - 1013F
- Count
- 57
Aegean Numbers is a set of ancient symbols used for counting in Minoan and Mycenaean scripts, including fractions and numerals.
- Abbr
- Ancient_Greek_Numbers
- Range
- 10140 - 1018F
- Count
- 79
Ancient Greek Numbers is a set of symbols for acrophonic numerals, including talent signs and fractions, used in classical inscriptions.
- Abbr
- Ancient_Symbols
- Range
- 10190 - 101CF
- Count
- 14
Ancient Symbols is a collection of glyphs for weights, measures, and astronomical signs, plus historical notations from Greek and Roman antiquity.
- Abbr
- Phaistos
- Range
- 101D0 - 101FF
- Count
- 46
Phaistos Disc is a historic script symbol set used for encoding the undeciphered ancient Minoan artifact’s markings, including pictographs and dividers.
- Abbr
- Lycian
- Range
- 10280 - 1029F
- Count
- 29
Lycian is a script block encoding the ancient Anatolian alphabet used for the Lycian language.
- Abbr
- Carian
- Range
- 102A0 - 102DF
- Count
- 49
Carian is a historical script block encoding the alphabetic writing system of ancient Caria, used from the 7th to 3rd centuries BCE.
- Abbr
- Coptic_Epact_Numbers
- Range
- 102E0 - 102FF
- Count
- 28
Coptic Epact Numbers is a set of ancient Egyptian numerical symbols used for astronomical and calendrical calculations, encoded in the standard character set.
- Abbr
- Old_Italic
- Range
- 10300 - 1032F
- Count
- 39
Old Italic is a script encoding ancient Etruscan and related Italic alphabets, including letters for writing early Latin and Oscan texts.
- Abbr
- Gothic
- Range
- 10330 - 1034F
- Count
- 27
Gothic is a script used to write the extinct East Germanic language, comprising letters derived from the Greek and Latin alphabets.
- Abbr
- Old_Permic
- Range
- 10350 - 1037F
- Count
- 43
Old Permic is a script used for writing the Komi language, containing letters and digits from the historic Abur alphabet.
- Abbr
- Ugaritic
- Range
- 10380 - 1039F
- Count
- 31
Ugaritic is a cuneiform script block encoding the ancient Ugaritic alphabet, used for writing the Northwest Semitic language of Ugarit.
- Abbr
- Old_Persian
- Range
- 103A0 - 103DF
- Count
- 50
Old Persian is a cuneiform script encoding the language of ancient Achaemenid inscriptions, featuring signs for syllables and logograms.
- Abbr
- Deseret
- Range
- 10400 - 1044F
- Count
- 80
Deseret is a historical script block for the phonetic alphabet created by the LDS Church, used mainly in 19th-century Utah documents.
- Abbr
- Shavian
- Range
- 10450 - 1047F
- Count
- 48
Shavian is a phonemic alphabet designed for English, created by Ronald Kingsley Read, with each character representing a distinct speech sound.
- Abbr
- Osmanya
- Range
- 10480 - 104AF
- Count
- 40
Osmanya is a script invented for Somali, used for writing the language and included in digital text standards.
- Abbr
- Osage
- Range
- 104B0 - 104FF
- Count
- 72
Osage is a script used for writing the Osage language, consisting of uppercase and lowercase letters with distinct phonetic forms.
- Abbr
- Elbasan
- Range
- 10500 - 1052F
- Count
- 40
Elbasan is a historical alphabet block encoding the script used for writing the Albanian language in the 18th century.
- Abbr
- Caucasian_Albanian
- Range
- 10530 - 1056F
- Count
- 53
Caucasian Albanian is a script used for an ancient Christian language, containing letters for the now extinct Caucasian Albanian language.
- Abbr
- Vithkuqi
- Range
- 10570 - 105BF
- Count
- 70
Vithkuqi is a historical Albanian script block, encoded for the original alphabet’s 46 letters, used in 19th-century texts.
- Abbr
- Todhri
- Range
- 105C0 - 105FF
- Count
- 52
Todhri is a script block encoding an Albanian alphabet, used historically for writing the Tosk dialect, with characters for letters and punctuation.
- Abbr
- Linear_A
- Range
- 10600 - 1077F
- Count
- 341
Linear A is a proposed script block for the undeciphered Minoan writing system, containing syllabic and ideographic signs from ancient Cretan inscriptions.
- Abbr
- Latin_Ext_F
- Range
- 10780 - 107BF
- Count
- 62
Latin Extended-F is a set of letters and symbols for medievalist and phonetic notations, including extinct scribal abbreviations and tone marks.
- Abbr
- Cypriot_Syllabary
- Range
- 10800 - 1083F
- Count
- 55
Cypriot Syllabary is a historical script used for ancient Cypriot Greek and Eteocypriot inscriptions with characters for syllabic writing.
- Abbr
- Imperial_Aramaic
- Range
- 10840 - 1085F
- Count
- 31
Imperial Aramaic is a right-to-left script used for the Aramaic language of the Achaemenid Empire, containing letters and punctuation.
- Abbr
- Palmyrene
- Range
- 10860 - 1087F
- Count
- 32
Palmyrene is a historical script block used for writing the Aramaic dialect of the ancient Syrian city, containing twenty‑seven letters and two numerals.
- Abbr
- Nabataean
- Range
- 10880 - 108AF
- Count
- 40
Nabataean is a script block for writing the ancient Nabataean language, used in inscriptions from the 2nd century BCE to the 4th century CE.
- Abbr
- Hatran
- Range
- 108E0 - 108FF
- Count
- 26
Hatran is a script used for the Aramaic dialect of the ancient city of Hatra, inscribed right-to-left with letters and numerals.
- Abbr
- Phoenician
- Range
- 10900 - 1091F
- Count
- 29
Phoenician is an ancient Semitic writing system of 29 letters, used mainly for inscriptions, and encoded as a right-to-left script.
- Abbr
- Lydian
- Range
- 10920 - 1093F
- Count
- 27
Lydian is a historic script block for the extinct Anatolian language, containing letters and a word separator symbol.
- Abbr
- Sidetic
- Range
- 10940 - 1095F
- Count
- 26
Sidetic is a historical script block used for writing the ancient Sidetic language of Anatolia, containing 14 characters including letters and a numeral sign.
- Abbr
- Meroitic_Hieroglyphs
- Range
- 10980 - 1099F
- Count
- 32
Meroitic Hieroglyphs is a historical script block encoding the ancient Egyptian-derived writing system used for the Meroitic language in Nubia.
- Abbr
- Meroitic_Cursive
- Range
- 109A0 - 109FF
- Count
- 90
Meroitic Cursive is a historical script used in ancient Nubia, featuring cursive signs for writing the Meroitic language, with some characters still unresolved.
- Abbr
- Kharoshthi
- Range
- 10A00 - 10A5F
- Count
- 68
Kharoshthi is a historical script from ancient South Asia, used for writing Gandhari and other Prakrits, encoded with letters, numerals, and punctuation marks.
- Abbr
- Old_South_Arabian
- Range
- 10A60 - 10A7F
- Count
- 32
Old South Arabian is a script block encoding the ancient Musnad alphabet, used for inscriptions in the pre-Islamic Yemeni kingdoms.
- Abbr
- Old_North_Arabian
- Range
- 10A80 - 10A9F
- Count
- 32
Old North Arabian is a script block encoding the ancient alphabetic writing system used for inscriptions across northwest Arabia.
- Abbr
- Manichaean
- Range
- 10AC0 - 10AFF
- Count
- 51
Manichaean is a script used for writing the Manichaean religion’s texts, including letters, punctuation, and abbreviations, derived from Syriac and Estrangela.
- Abbr
- Avestan
- Range
- 10B00 - 10B3F
- Count
- 61
Avestan is a script used for the sacred Zoroastrian texts, encoded with letters, digits, and punctuation marks.
- Abbr
- Inscriptional_Parthian
- Range
- 10B40 - 10B5F
- Count
- 30
Inscriptional Parthian is a script used for Middle Persian inscriptions, containing letters, numbers, and punctuation marks from ancient Parthian texts.
- Abbr
- Inscriptional_Pahlavi
- Range
- 10B60 - 10B7F
- Count
- 27
Inscriptional Pahlavi is a script used for Middle Persian inscriptions, encoded with letters and punctuation for historical texts.
- Abbr
- Psalter_Pahlavi
- Range
- 10B80 - 10BAF
- Count
- 29
Psalter Pahlavi is a script used for Middle Persian religious texts, featuring letters and numerals, now encoded digitally for historical scholarship.
- Abbr
- Old_Turkic
- Range
- 10C00 - 10C4F
- Count
- 73
Old Turkic is a script used for inscriptions and manuscripts from the 8th to 13th centuries, including runiform letters for several Turkic languages.
- Abbr
- Old_Hungarian
- Range
- 10C80 - 10CFF
- Count
- 108
Old Hungarian is a script block encoding the ancient runiform writing of the Magyars, used for inscriptions and early texts.
- Abbr
- Hanifi_Rohingya
- Range
- 10D00 - 10D3F
- Count
- 50
Hanifi Rohingya is a script used for writing the Rohingya language, featuring letters, digits, and diacritics designed for tonal distinctions.
- Abbr
- Garay
- Range
- 10D40 - 10D8F
- Count
- 69
Garay is a script used for the Wolof language, recently encoded with letters, digits, and punctuation for writing in Senegal.
- Abbr
- Rumi
- Range
- 10E60 - 10E7F
- Count
- 31
Rumi Numeral Symbols is a set of ancient Maghrebi digits and fractions used historically in North African Arabic manuscripts and mathematical texts.
- Abbr
- Yezidi
- Range
- 10E80 - 10EBF
- Count
- 47
Yezidi is a script used for writing the Kurdish Yezidi religious language with letters and diacritics.
- Abbr
- Arabic_Ext_C
- Range
- 10EC0 - 10EFF
- Count
- 60
Arabic Extended-C is a repertoire of Arabic script characters for Quranic annotation, including additional diacritics and marks used in religious texts.
- Abbr
- Old_Sogdian
- Range
- 10F00 - 10F2F
- Count
- 40
Old Sogdian is a script used for writing the Sogdian language, containing letters and numerals from the ancient Iranian civilization.
- Abbr
- Sogdian
- Range
- 10F30 - 10F6F
- Count
- 42
Sogdian is a script used for the ancient Iranian language of the Silk Road, with letters written right-to-left and often connected cursively.
- Abbr
- Old_Uyghur
- Range
- 10F70 - 10FAF
- Count
- 26
Old Uyghur is a script used for the Turkic language, featuring cursive letters written vertically, with distinct forms for consonants and vowels.
- Abbr
- Chorasmian
- Range
- 10FB0 - 10FDF
- Count
- 28
Chorasmian is a script used for the extinct Eastern Iranian language, encoded with letters, numbers, and punctuation for historical texts.
- Abbr
- Elymaic
- Range
- 10FE0 - 10FFF
- Count
- 23
Elymaic is a script used for the Aramaic dialect of ancient Elymais, containing letters and numbers for historical inscriptions.
- Abbr
- Brahmi
- Range
- 11000 - 1107F
- Count
- 115
Brahmi is a historic script block encoding early Indian inscriptions, used for writing Prakrit, Sanskrit, and other ancient languages.
- Abbr
- Kaithi
- Range
- 11080 - 110CF
- Count
- 68
Kaithi is a historical script from northern India, used mainly for legal and commercial records, now encoded for digital text preservation.
- Abbr
- Sora_Sompeng
- Range
- 110D0 - 110FF
- Count
- 35
Sora Sompeng is a script used for writing the Sora language, with letters for consonants, vowels, and digits, plus diacritics and punctuation.
- Abbr
- Chakma
- Range
- 11100 - 1114F
- Count
- 71
Chakma is a script used for the Chakma language, containing letters, digits, and diacritics for writing in Bangladesh and India.
- Abbr
- Mahajani
- Range
- 11150 - 1117F
- Count
- 39
Mahajani is a historical script from northwest India, used for accounting and mercantile records, with its characters encoded in the Unicode standard.
- Abbr
- Sharada
- Range
- 11180 - 111DF
- Count
- 96
Sharada is an ancient script from Kashmir used for Sanskrit and Kashmiri, now encoded for modern digital text preservation and scholarly use.
- Abbr
- Sinhala_Archaic_Numbers
- Range
- 111E0 - 111FF
- Count
- 20
Sinhala Archaic Numbers is a set of ancient numeral symbols used in historical Sri Lankan records, now encoded for digital preservation.
- Abbr
- Khojki
- Range
- 11200 - 1124F
- Count
- 65
Khojki is a script used for the Sindhi language, encoded for writing liturgical and secular texts with distinct letterforms.
- Abbr
- Multani
- Range
- 11280 - 112AF
- Count
- 38
Multani is an ancient script used for writing the Saraiki language, with characters for consonants, vowels, and diacritics.
- Abbr
- Khudawadi
- Range
- 112B0 - 112FF
- Count
- 69
Khudawadi is a script used for writing Sindhi, encoded with characters for consonants, vowels, and diacritical marks.
- Abbr
- Grantha
- Range
- 11300 - 1137F
- Count
- 86
Grantha is a historic South Indian script block for writing Sanskrit and Tamil, encoded for scholarly and digital preservation.
- Abbr
- Tulu_Tigalari
- Range
- 11380 - 113FF
- Count
- 80
Tulu-Tigalari is a script for the Tulu language, comprising letters, digits, and diacritics for writing ancient inscriptions and modern texts.
- Abbr
- Newa
- Range
- 11400 - 1147F
- Count
- 97
Newa is a script used for writing the Nepal Bhasa language, containing letters, digits, and punctuation marks for the Newar people.
- Abbr
- Tirhuta
- Range
- 11480 - 114DF
- Count
- 82
Tirhuta is a script block used for writing the Maithili language, primarily in the Mithila region of India and Nepal.
- Abbr
- Siddham
- Range
- 11580 - 115FF
- Count
- 92
Siddham is a historic script block used for Buddhist texts, encoded for digital preservation of medieval Indian and East Asian manuscripts.
- Abbr
- Modi
- Range
- 11600 - 1165F
- Count
- 79
Modi is a script used historically for writing the Marathi language, featuring cursive letterforms and distinct vowel signs, now encoded digitally.
- Abbr
- Mongolian_Sup
- Range
- 11660 - 1167F
- Count
- 13
Mongolian Supplement is a set of marks for the Mongolian script, used to indicate gender, case, and other grammatical distinctions.
- Abbr
- Takri
- Range
- 11680 - 116CF
- Count
- 68
Takri is a script used historically for languages like Dogri and Chambeali, with glyphs for consonants, vowels, and diacritical marks.
- Abbr
- Myanmar_Ext_C
- Range
- 116D0 - 116FF
- Count
- 20
Myanmar Extended-C is a recently added set of characters for minority languages and historical uses, filling gaps in the Myanmar script.
- Abbr
- Ahom
- Range
- 11700 - 1174F
- Count
- 65
Ahom is a script block encoding the historical Ahom language and Tai-Ahom texts, used for writing inscriptions and manuscripts from medieval Assam.
- Abbr
- Dogra
- Range
- 11800 - 1184F
- Count
- 60
Dogra is a script used for the Dogri language, with characters for vowels, consonants, and digits, encoded in the South Asian writing systems section.
- Abbr
- Warang_Citi
- Range
- 118A0 - 118FF
- Count
- 84
Warang Citi is a script used for writing the Ho language, with letters and digits for numerals, plus punctuation and spacing marks.
- Abbr
- Dives_Akuru
- Range
- 11900 - 1195F
- Count
- 72
Dives Akuru is a historical script from the Maldives, used for ancient Dhivehi texts, encoded for digital preservation.
- Abbr
- Nandinagari
- Range
- 119A0 - 119FF
- Count
- 65
Nandinagari is a script used historically for Sanskrit and Kannada, encoded to preserve ancient inscriptions and manuscripts.
- Abbr
- Zanabazar_Square
- Range
- 11A00 - 11A4F
- Count
- 72
Zanabazar Square is a historical Mongolian script used for Buddhist texts, with a rounded, calligraphic style and vertical writing direction.
- Abbr
- Soyombo
- Range
- 11A50 - 11AAF
- Count
- 83
Soyombo is a historic script used primarily for Mongolian and Tibetan Buddhist texts, notable for its distinct vertical style and ceremonial, decorative use.
- Abbr
- UCAS_Ext_A
- Range
- 11AB0 - 11ABF
- Count
- 16
Unified Canadian Aboriginal Syllabics Extended-A is a small set of characters for writing additional Cree, Ojibwe, and Dene languages.
- Abbr
- Pau_Cin_Hau
- Range
- 11AC0 - 11AFF
- Count
- 57
Pau Cin Hau is a historic script used to write the Chin languages of Myanmar, encoded for digital text representation.
- Abbr
- Devanagari_Ext_A
- Range
- 11B00 - 11B5F
- Count
- 11
Devanagari Extended-A is a set of rare historical and ritual characters, including signs for vedic chants and ligatures, used in ancient Sanskrit manuscripts.
- Abbr
- Sharada_Sup
- Range
- 11B60 - 11B7F
- Count
- 8
Sharada Supplement is a set of vowel signs and marks, used for writing the Sharada script, that follow the main Sharada block.
- Abbr
- Sunuwar
- Range
- 11BC0 - 11BFF
- Count
- 44
Sunuwar is a script encoding the Sunuwar language of Nepal for writing its consonants, vowels, and diacritics.
- Abbr
- Bhaiksuki
- Range
- 11C00 - 11C6F
- Count
- 97
Bhaiksuki is a script used for writing Sanskrit, with historical significance in Buddhist texts, featuring consonants, vowels, and diacritical marks.
- Abbr
- Marchen
- Range
- 11C70 - 11CBF
- Count
- 68
Marchen is a script used for writing the extinct Zhang-Zhung language, with its characters encoded for historical and scholarly use.
- Abbr
- Masaram_Gondi
- Range
- 11D00 - 11D5F
- Count
- 75
Masaram Gondi is a script used for writing the Gondi language, with letters, digits, and marks for sounds.
- Abbr
- Gunjala_Gondi
- Range
- 11D60 - 11DAF
- Count
- 63
Gunjala Gondi is a script used for writing the Gondi language, with characters for consonants, vowels, and diacritics.
- Abbr
- Tolong_Siki
- Range
- 11DB0 - 11DEF
- Count
- 54
Tolong Siki is a recently encoded script used for the Southern Min language, featuring letters, diacritics, and punctuation for tonal writing.
- Abbr
- Bengali_Sup
- Range
- 11DF0 - 11DFF
- Count
- 2
Bengali Supplement is a small collection of characters for historical documents and Sanskrit texts.
- Abbr
- Makasar
- Range
- 11EE0 - 11EFF
- Count
- 25
Makasar is a historical script from South Sulawesi, Indonesia, used for writing the Makassarese language, encoded for modern digital text support.
- Abbr
- Kawi
- Range
- 11F00 - 11F5F
- Count
- 87
Kawi is a script for Old Javanese, Balinese, and Sundanese texts, used historically across Southeast Asia for inscriptions and literature.
- Abbr
- Lisu_Sup
- Range
- 11FB0 - 11FBF
- Count
- 1
Lisu Supplement is a set of additional characters for the Lisu script, mainly used for the standardized Naxi language and other minority languages in China.
- Abbr
- Tamil_Sup
- Range
- 11FC0 - 11FFF
- Count
- 51
Tamil Supplement is a set of characters adding ancient Tamil numerals, symbols, and punctuation for historical and liturgical texts.
- Abbr
- Cuneiform
- Range
- 12000 - 123FF
- Count
- 922
Cuneiform is a historic script block encoding Sumerian, Akkadian, and other ancient Middle Eastern wedge-shaped signs for scholarly text processing.
- Abbr
- Cuneiform_Numbers
- Range
- 12400 - 1247F
- Count
- 128
Cuneiform Numbers and Punctuation is a collection of ancient Mesopotamian numeric signs, fractions, and scribal marks for transliterating cuneiform texts.
- Abbr
- Early_Dynastic_Cuneiform
- Range
- 12480 - 1254F
- Count
- 196
Early Dynastic Cuneiform is a script segment housing ancient Sumerian signs from the Early Dynastic period, including logograms and syllabic values.
- Abbr
- Archaic_Cuneiform_Numerals
- Range
- 12550 - 1268F
- Count
- 311
Archaic Cuneiform Numerals is a collection of ancient numerical signs from fourth and third millennium BCE Mesopotamia.
- Abbr
- Cypro_Minoan
- Range
- 12F90 - 12FFF
- Count
- 99
Cypro-Minoan is a script block encoding syllabic signs from ancient Cyprus, used for inscriptions dating to the Late Bronze Age.
- Abbr
- Egyptian_Hieroglyphs
- Range
- 13000 - 1342F
- Count
- 1072
Egyptian Hieroglyphs is the encoded set of over 1,000 ancient signs for digital text, enabling modern use and research of the writing system.
- Abbr
- Egyptian_Hieroglyph_Format_Controls
- Range
- 13430 - 1345F
- Count
- 38
Egyptian Hieroglyph Format Controls is a set of signs that manage text layout, like joining, overlapping, and insertion, for proper ancient Egyptian writing.
- Abbr
- Egyptian_Hieroglyphs_Ext_A
- Range
- 13460 - 143FF
- Count
- 3995
Egyptian Hieroglyphs Extended-A is a set of over 3,900 signs from Old Egyptian, filling gaps in the core block for fuller text encoding.
- Abbr
- Anatolian_Hieroglyphs
- Range
- 14400 - 1467F
- Count
- 583
Anatolian Hieroglyphs is a script encoding the Luwian language’s pictorial signs, used in Bronze Age Anatolia for monumental inscriptions.
- Abbr
- Gurung_Khema
- Range
- 16100 - 1613F
- Count
- 58
Gurung Khema is a script used for the Tamu people’s language, with characters for vowels, consonants, and digits.
- Abbr
- Bamum_Sup
- Range
- 16800 - 16A3F
- Count
- 569
Bamum Supplement is a set of additional glyphs for the Bamum script, filling gaps in the historical syllabary for modern usage.
- Abbr
- Mro
- Range
- 16A40 - 16A6F
- Count
- 43
Mro is a script used for the Mru language of Bangladesh and Myanmar, consisting of letters, digits, and punctuation marks.
- Abbr
- Tangsa
- Range
- 16A70 - 16ACF
- Count
- 89
Tangsa is a script used for writing the Tangsa languages of Northeast India, with characters for consonants, vowels, and tone marks.
- Abbr
- Bassa_Vah
- Range
- 16AD0 - 16AFF
- Count
- 36
Bassa Vah is a script for writing the Bassa language, with characters for consonants, vowels, and tone marks.
- Abbr
- Pahawh_Hmong
- Range
- 16B00 - 16B8F
- Count
- 127
Pahawh Hmong is a script for writing the Hmong language, featuring distinct consonant and vowel signs with unique tonal markers.
- Abbr
- Kirat_Rai
- Range
- 16D40 - 16D7F
- Count
- 58
Kirat Rai is a script used for writing the Kirat-Khambu languages, including its consonants, vowels, and diacritical marks, assigned for modern digital text.
- Abbr
- Medefaidrin
- Range
- 16E40 - 16E9F
- Count
- 91
Medefaidrin is a script used for the West African language of the same name, featuring characters for writing its sounds.
- Abbr
- Beria_Erfe
- Range
- 16EA0 - 16EDF
- Count
- 50
Beria Erfe is a script used for the Beria language including consonants, vowels, and tones.
- Abbr
- Miao
- Range
- 16F00 - 16F9F
- Count
- 149
Miao is a script used for writing the A-Hmao language, containing letters, digits, and punctuation, with characters based on the Pollard phonetic system.
- Abbr
- Ideographic_Symbols
- Range
- 16FE0 - 16FFF
- Count
- 12
Ideographic Symbols and Punctuation is a small set of historic Chinese characters and marks used for phonetic annotation and textual reference.
- Abbr
- Tangut
- Range
- 17000 - 187FF
- Count
- 6144
Tangut is a historical script block encoding characters from the Tangut language, used in the Western Xia dynasty, with over 6,000 glyphs.
- Abbr
- Tangut_Components
- Range
- 18800 - 18AFF
- Count
- 768
Tangut Components is a set of radical-like symbols used to index the historical Tangut script in dictionaries and digital text.
- Abbr
- Khitan_Small_Script
- Range
- 18B00 - 18CFF
- Count
- 476
Khitan Small Script is a historical writing system for the extinct Khitan language, encoded to preserve its unique logographic and phonetic characters.
- Abbr
- Tangut_Sup
- Range
- 18D00 - 18D7F
- Count
- 33
Tangut Supplement is a set of rare Tangut ideographs, mainly variant characters and phonetic components, used for scholarly texts and historical research.
- Abbr
- Tangut_Components_Sup
- Range
- 18D80 - 18DFF
- Count
- 115
Tangut Components Supplement is a set of rare radicals and subcomponents used for the Tangut script, aiding in dictionary and character analysis.
- Abbr
- Jurchen
- Range
- 18E00 - 1919F
- Count
- 914
Jurchen is a historical ideographic script from northeastern China used during the Jin and Ming dynasties.
- Abbr
- Jurchen_Radicals
- Range
- 191A0 - 191DF
- Count
- 51
Jurchen Radicals is a set of fifty-one indexing components for the Jurchen script.
- Abbr
- Kana_Ext_B
- Range
- 1AFF0 - 1AFFF
- Count
- 13
Kana Extended-B is a small set of Hiragana letters with combining marks, used for archaic or dialectal Japanese transcriptions.
- Abbr
- Kana_Sup
- Range
- 1B000 - 1B0FF
- Count
- 256
Kana Supplement is the collection of historical and variant Japanese kana characters, including hentaigana and additional small kana forms.
- Abbr
- Kana_Ext_A
- Range
- 1B100 - 1B12F
- Count
- 41
Kana Extended-A is a set of additional Hiragana letters, mainly archaic and obsolete forms, used for historical Japanese texts.
- Abbr
- Small_Kana_Ext
- Range
- 1B130 - 1B16F
- Count
- 10
Small Kana Extension is a set of tiny hiragana and katakana characters used for writing small vowel sounds and glides in modern Japanese.
- Abbr
- Nushu
- Range
- 1B170 - 1B2FF
- Count
- 396
Nushu is a script encoding the syllabic writing system historically used exclusively by women in Hunan, China, preserving their songs and stories.
- Abbr
- Duployan
- Range
- 1BC00 - 1BC9F
- Count
- 143
Duployan is a script encoding shorthand systems, including Duployé, Pernin, and Sloan-Duployan, used for phonetic writing.
- Abbr
- Shorthand_Format_Controls
- Range
- 1BCA0 - 1BCAF
- Count
- 4
Shorthand Format Controls is a set of invisible formatting characters used in shorthand notation to manage overlapping strokes and positioning within text.
- Abbr
- Symbols_For_Legacy_Computing_Sup
- Range
- 1CC00 - 1CEBF
- Count
- 695
Symbols for Legacy Computing Supplement is a set of glyphs preserving obsolete computer symbols, including early terminal graphics and control codes.
- Abbr
- Misc_Symbols_Sup
- Range
- 1CEC0 - 1CEFF
- Count
- 53
Miscellaneous Symbols Supplement is a set of additional pictographs, arrows, and technical signs for extended typographic and symbolic use.
- Abbr
- Znamenny_Music
- Range
- 1CF00 - 1CFCF
- Count
- 185
Znamenny Musical Notation is a set of symbols for Russian Orthodox chant, used to encode neumatic notation in digital texts.
- Abbr
- Byzantine_Music
- Range
- 1D000 - 1D0FF
- Count
- 246
Byzantine Musical Symbols is a set of neumes and signs used to notate Byzantine chant, preserving its melodic tradition in digital text.
- Abbr
- Music
- Range
- 1D100 - 1D1FF
- Count
- 256
Musical Symbols is a collection of glyphs for notation, including clefs, notes, rests, accidentals, and dynamics, used in digital music texts.
- Abbr
- Ancient_Greek_Music
- Range
- 1D200 - 1D24F
- Count
- 70
Ancient Greek Musical Notation is a set of symbols for notating melodies and rhythms including vocal and instrumental signs used in classical antiquity.
- Abbr
- Music_Sup
- Range
- 1D250 - 1D28F
- Count
- 50
Musical Symbols Supplement is a collection expanding notation with new flags
- Abbr
- Kaktovik_Numerals
- Range
- 1D2C0 - 1D2DF
- Count
- 20
Kaktovik Numerals is a set of twenty digit symbols for the base-20 Iñupiaq counting system, designed to aid math education in Alaska.
- Abbr
- Mayan_Numerals
- Range
- 1D2E0 - 1D2FF
- Count
- 20
Mayan Numerals is a set of twenty glyphs representing the base-20 vigesimal system, including zero, used for historical calendrical and mathematical notation.
- Abbr
- Tai_Xuan_Jing
- Range
- 1D300 - 1D35F
- Count
- 87
Tai Xuan Jing Symbols is a set of glyphs representing the four binary elements and 81 tetragrams from the ancient Chinese divination text.
- Abbr
- Counting_Rod
- Range
- 1D360 - 1D37F
- Count
- 25
Counting Rod Numerals is a set of ancient Chinese rod-based digits and symbols used for mathematical calculations, including zero and negative values.
- Abbr
- Math_Alphanum
- Range
- 1D400 - 1D7FF
- Count
- 997
Mathematical Alphanumeric Symbols is a set of styled Latin and Greek letters, digits, and symbols for mathematical notation in bold, italic, and script forms.
- Abbr
- Sutton_SignWriting
- Range
- 1D800 - 1DAAF
- Count
- 672
Sutton SignWriting is a script for writing sign languages, featuring symbols for handshapes, movements, and facial expressions, enabling precise notation.
- Abbr
- Misc_Arrows_Ext
- Range
- 1DB00 - 1DBFF
- Count
- 29
Miscellaneous Symbols and Arrows Extended is a an extension of symbols and arrows for historical notation including Leibnizian mathematical operators.
- Abbr
- Latin_Ext_G
- Range
- 1DF00 - 1DFFF
- Count
- 188
Latin Extended-G is a set of characters for medieval and phonetic transcriptions, including letters for African languages and historical scribal abbreviations.
- Abbr
- Glagolitic_Sup
- Range
- 1E000 - 1E02F
- Count
- 38
Glagolitic Supplement is a set of characters used for archaic Glagolitic texts, including combining marks and letters for Old Church Slavonic.
- Abbr
- Cyrillic_Ext_D
- Range
- 1E030 - 1E08F
- Count
- 63
Cyrillic Extended-D is a set of historic letters and combining marks for early Cyrillic manuscripts, including abbreviations and scribal variants.
- Abbr
- Nyiakeng_Puachue_Hmong
- Range
- 1E100 - 1E14F
- Count
- 71
Nyiakeng Puachue Hmong is a script used for writing the Hmong language, designed by Cher Xiong to reflect spoken tones and sounds.
- Abbr
- Toto
- Range
- 1E290 - 1E2BF
- Count
- 31
Toto is a script used for the Toto language of northeastern India, encoded with letters and digits for writing that endangered tongue.
- Abbr
- Wancho
- Range
- 1E2C0 - 1E2FF
- Count
- 59
Wancho is a script for the Wancho language of northeastern India containing consonants, vowels, and tone marks.
- Abbr
- Nag_Mundari
- Range
- 1E4D0 - 1E4FF
- Count
- 42
Nag Mundari is a script used for writing the Mundari language, with characters for consonants, vowels, and digits.
- Abbr
- Ol_Onal
- Range
- 1E5D0 - 1E5FF
- Count
- 44
Ol Onal is a script used for the Ol Onal language of southern India, comprising letters, digits, and punctuation marks.
- Abbr
- Tai_Yo
- Range
- 1E6C0 - 1E6FF
- Count
- 55
Tai Yo is a script for writing the Tai Yo language, encompassing consonants, vowels, tone marks, and punctuation.
- Abbr
- Ethiopic_Ext_B
- Range
- 1E7E0 - 1E7FF
- Count
- 28
Ethiopic Extended-B is a set of rare and historic Ethiopic syllabic signs, used for transliterating and preserving ancient Geʽez and liturgical texts.
- Abbr
- Mende_Kikakui
- Range
- 1E800 - 1E8DF
- Count
- 213
Mende Kikakui is a script used for writing the Mende language of Sierra Leone, featuring syllabic characters with distinctive dotted and curved forms.
- Abbr
- Adlam
- Range
- 1E900 - 1E95F
- Count
- 88
Adlam is a modern script for the Fulani language, used to write it from right to left with uppercase and lowercase letters plus digits.
- Abbr
- Indic_Siyaq_Numbers
- Range
- 1EC70 - 1ECBF
- Count
- 68
Indic Siyaq Numbers is a set of historical numerals used in Indian accounting, including forms for fractions and units.
- Abbr
- Ottoman_Siyaq_Numbers
- Range
- 1ED00 - 1ED4F
- Count
- 61
Ottoman Siyaq Numbers is a set of numerals used in Ottoman Turkish financial documents, encoding both digits and accounting-specific notation.
- Abbr
- Arabic_Math
- Range
- 1EE00 - 1EEFF
- Count
- 143
Arabic Mathematical Alphabetic Symbols is a set of distinct letterforms for expressing equations, used in scientific and technical Arabic writing.
- Abbr
- Mahjong
- Range
- 1F000 - 1F02F
- Count
- 44
Mahjong Tiles is a set of symbols depicting classic Chinese mahjong tiles, including winds, dragons, and suits, for digital representation.
- Abbr
- Domino
- Range
- 1F030 - 1F09F
- Count
- 100
Domino Tiles is a set of square symbols depicting all standard domino pieces, including blank faces, in a clear, game-ready style.
- Abbr
- Playing_Cards
- Range
- 1F0A0 - 1F0FF
- Count
- 82
Playing Cards is a set of symbols depicting playing cards, including suits, courts, jokers, and card backs, used for digital card games.
- Abbr
- Enclosed_Alphanum_Sup
- Range
- 1F100 - 1F1FF
- Count
- 201
Enclosed Alphanumeric Supplement is a set of circled, parenthesized, and squared numbers and letters, plus regional indicator symbols for flag emojis.
- Abbr
- Enclosed_Ideographic_Sup
- Range
- 1F200 - 1F2FF
- Count
- 64
Enclosed Ideographic Supplement is a set of circled or enclosed Japanese kanji and symbols, used for annotations, ratings, and commercial signage.
- Abbr
- Misc_Pictographs
- Range
- 1F300 - 1F5FF
- Count
- 768
Miscellaneous Symbols and Pictographs is a collection of emoji, weather icons, and everyday objects, curated for visual communication and playful expression.
- Abbr
- Emoticons
- Range
- 1F600 - 1F64F
- Count
- 80
Emoticons is a set of pictographic faces and gestures, including smileys, frowns, and winks, used to convey emotions in digital text.
- Abbr
- Ornamental_Dingbats
- Range
- 1F650 - 1F67F
- Count
- 48
Ornamental Dingbats is a set of sixty decorative flourishes, leaf motifs, and curved ornament fragments used for typographic embellishment and page decoration.
- Abbr
- Transport_And_Map
- Range
- 1F680 - 1F6FF
- Count
- 120
Transport and Map Symbols is a set of icons for vehicles, traffic signs, and map-related imagery, used mainly in digital communication and wayfinding.
- Abbr
- Alchemical
- Range
- 1F700 - 1F77F
- Count
- 128
Alchemical Symbols is a set of icons for substances, processes, and equipment, based on historical alchemy texts and used in chemistry-related contexts.
- Abbr
- Geometric_Shapes_Ext
- Range
- 1F780 - 1F7FF
- Count
- 120
Geometric Shapes Extended is a collection of symbols, including squares, circles, triangles, and stars, used for diagrams, games, and visual notation.
- Abbr
- Sup_Arrows_C
- Range
- 1F800 - 1F8FF
- Count
- 171
Supplemental Arrows-C is a set of mostly heavy and wide arrow symbols, including curved, dashed, and paired variants, for technical and mathematical notation.
- Abbr
- Sup_Symbols_And_Pictographs
- Range
- 1F900 - 1F9FF
- Count
- 256
Supplemental Symbols and Pictographs is a set of emoji-like icons covering sports, medical, and fantasy themes, filling gaps beyond earlier symbol sets.
- Abbr
- Chess_Symbols
- Range
- 1FA00 - 1FA6F
- Count
- 102
Chess Symbols is a collection of standardized glyphs for chess pieces, boards, and related notation, supporting digital play and annotation.
- Abbr
- Symbols_And_Pictographs_Ext_A
- Range
- 1FA70 - 1FAFF
- Count
- 128
Symbols and Pictographs Extended-A is a collection of symbols supplementing existing emoji, encompassing toys, games, and additional pictographic variants.
- Abbr
- Symbols_For_Legacy_Computing
- Range
- 1FB00 - 1FBFF
- Count
- 250
Symbols for Legacy Computing is a set of glyphs for retro terminals, fonts, and early home computer graphics, preserving historical screen output.
- Abbr
- CJK_Ext_B
- Range
- 20000 - 2A6DF
- Count
- 42720
CJK Unified Ideographs Extension B is a large set of over 42,000 rare and historical Chinese characters, added to support ancient texts and lesser‑known usage.
- Abbr
- CJK_Ext_C
- Range
- 2A700 - 2B73F
- Count
- 4160
CJK Unified Ideographs Extension C is a set of rare and historical Chinese characters, adding over 4,000 entries mainly for uncommon names and classical texts.
- Abbr
- CJK_Ext_D
- Range
- 2B740 - 2B81F
- Count
- 223
CJK Unified Ideographs Extension D is a set of 222 rare Chinese characters covering ancient and uncommon usage.
- Abbr
- CJK_Ext_E
- Range
- 2B820 - 2CEAF
- Count
- 5774
CJK Unified Ideographs Extension E is a set of over 5,700 rare Chinese characters, primarily for historical and dialectal usage, encoded in Unicode.
- Abbr
- CJK_Ext_F
- Range
- 2CEB0 - 2EBEF
- Count
- 7473
CJK Unified Ideographs Extension F is a set of rare and historical Chinese primarily for ancient texts and dialectal usage.
- Abbr
- CJK_Ext_I
- Range
- 2EBF0 - 2EE5F
- Count
- 622
CJK Unified Ideographs Extension I is a set of rare Chinese characters for historical and dialectal usage.
- Abbr
- CJK_Compat_Ideographs_Sup
- Range
- 2F800 - 2FA1F
- Count
- 542
CJK Compatibility Ideographs Supplement is a set of rare and alternate Chinese characters, mainly for historical texts, mapped to unified CJK ideographs.
- Abbr
- CJK_Ext_G
- Range
- 30000 - 3134F
- Count
- 4939
CJK Unified Ideographs Extension G is a set of over 4,900 rare Chinese characters, primarily for historical and dialectal usage.
- Abbr
- CJK_Ext_H
- Range
- 31350 - 323AF
- Count
- 4192
CJK Unified Ideographs Extension H is a collection of rare and historical Chinese characters, primarily for ancient texts and dialectal use.
- Abbr
- CJK_Ext_J
- Range
- 323B0 - 3347F
- Count
- 4298
CJK Unified Ideographs Extension J is a set of rare, historical Chinese characters added to Unicode for encoding ancient and uncommon texts.
- Abbr
- Seal
- Range
- 3D000 - 3FC3F
- Count
- 11328
Seal is a historical script that served as a precursor to modern Han ideographs.
- Abbr
- Tags
- Range
- E0000 - E007F
- Count
- 97
Tags is a set of invisible, deprecated format characters used to tag text with language or script metadata, now largely obsolete.
- Abbr
- VS_Sup
- Range
- E0100 - E01EF
- Count
- 240
Variation Selectors Supplement is a set of invisible format characters that refine glyph presentation for ideographs, emoji, and other complex scripts.
- Abbr
- Sup_PUA_A
- Range
- F0000 - FFFFF
- Count
- 65534
Supplementary Private Use Area-A is a reserved region for private character assignments, ensuring no standard Unicode meanings exist there.
- Abbr
- Sup_PUA_B
- Range
- 100000 - 10FFFF
- Count
- 65534
Supplementary Private Use Area-B is reserved for user-defined characters, with no assigned glyphs, allowing private interchange outside standard encoding.
- Abbr
- NB
- Range
- -
- Count
- 0
No Block is a placeholder label for characters not assigned to any named grouping.