Word Break - WSegSpace

Unicode Version 18.0

WSegSpace is a specific classification assigned to characters that act as a single, indivisible unit for line breaking and word segmentation, primarily covering the space character (U+0020) and similar whitespace such as tab. When text processing algorithms encounter a WSegSpace character, they treat it as a boundary that can separate words, but they also allow the character itself to be skipped or consumed during segmentation, ensuring that consecutive spaces or spaces adjacent to punctuation do not create empty or unintended word breaks. This value helps maintain consistent behavior when determining where a word starts or ends in languages that use spaces as delimiters, while still permitting fine-grained control over how multiple whitespace characters are handled during analysis and rendering.

1-14 of 14 results

U+0020
SP
Space
U+1680
Ogham Space Mark
U+2000
 
En Quad
U+2001
Em Quad
U+2002
En Space
U+2003
Em Space
U+2004
Three-Per-Em Space
U+2005
Four-Per-Em Space
U+2006
Six-Per-Em Space
U+2008
Punctuation Space
U+2009
Thin Space
U+200A
Hair Space
U+205F
MMSP
Medium Mathematical Space
U+3000
 
Ideographic Space