Sentence Break - Sp

Unicode Version 18.0

Sp is a classification used in text segmentation to identify spaces and other whitespace characters that act as sentence boundaries, but only when they appear between sentence-ending punctuation and the start of a new sentence. These characters, such as the regular space, tab, or non-breaking space, are not themselves punctuation but serve as separators that help determine where one sentence ends and another begins. In practice, Sp characters are treated as neutral or boundary-promoting, meaning they can either close a sentence or allow it to continue depending on the surrounding context, such as whether they follow a period, question mark, or exclamation point. Their role is crucial for accurate line breaking and sentence detection in languages that use spaces between words.

1-20 of 20 results

U+0009
HT
CHARACTER TABULATION
U+000B
VT
LINE TABULATION
U+000C
FF
FORM FEED
U+0020
SP
Space
U+00A0
 
No-Break Space
U+1680
Ogham Space Mark
U+2000
 
En Quad
U+2001
Em Quad
U+2002
En Space
U+2003
Em Space
U+2004
Three-Per-Em Space
U+2005
Four-Per-Em Space
U+2006
Six-Per-Em Space
U+2007
Figure Space
U+2008
Punctuation Space
U+2009
Thin Space
U+200A
Hair Space
U+202F
Narrow No-Break Space
U+205F
MMSP
Medium Mathematical Space
U+3000
 
Ideographic Space