Word Break - WSegSpace

Unicode Version 18.0

WSegSpace is a specific classification assigned to characters that act as a single, indivisible unit for line breaking and word segmentation, primarily covering the space character (U+0020) and similar whitespace such as tab. When text processing algorithms encounter a WSegSpace character, they treat it as a boundary that can separate words, but they also allow the character itself to be skipped or consumed during segmentation, ensuring that consecutive spaces or spaces adjacent to punctuation do not create empty or unintended word breaks. This value helps maintain consistent behavior when determining where a word starts or ends in languages that use spaces as delimiters, while still permitting fine-grained control over how multiple whitespace characters are handled during analysis and rendering.

1-14 of 14 results

U+0020
SP
Space
U+1680
 
Ogham Space Mark
U+2000
 
En Quad
U+2001
 
Em Quad
U+2002
 
En Space
U+2003
 
Em Space
U+2004
 
Three-Per-Em Space
U+2005
 
Four-Per-Em Space
U+2006
 
Six-Per-Em Space
U+2008
 
Punctuation Space
U+2009
 
Thin Space
U+200A
 
Hair Space
U+205F
MMSP
Medium Mathematical Space
U+3000
 
Ideographic Space