Partitioning a string in Python by a regular expression

Question

I need to split a string into an array on word boundaries (whitespace) while maintaining the whitespace.

For example:

'this is  a
sentence'

Would become

['this', ' ', 'is', '  ', 'a' '
', 'sentence']

I know about str.partition and re.split, but neither of them quite do what I want and there is no re.partition.

How should I partition strings on whitespace in Python with reasonable efficiency?

NikitaBaksalyar · Accepted Answer

Try this:

s = "this is  a
sentence"
re.split(r'(\W+)', s) # Notice parentheses and a plus sign.

Result would be:

['this', ' ', 'is', '  ', 'a', '
', 'sentence']

eyquem · Answer

Symbol of whitespace in re is '\s' not '\W'

Compare:

import re


s = "With a sign # written @ the beginning , that's  a
sentence,"\
    '
no more an instruction!,	you know ?? "Cases" & and surprises:'\
    "that will 'lways unknown **before**, in 81% of time$"


a = re.split('(\W+)', s)
print a
print len(a)
print

b = re.split('(\s+)', s)
print b
print len(b)

produces

['With', ' ', 'a', ' ', 'sign', ' # ', 'written', ' @ ', 'the', ' ', 'beginning', ' , ', 'that', "'", 's', '  ', 'a', '
', 'sentence', ',
', 'no', ' ', 'more', ' ', 'an', ' ', 'instruction', '!,	', 'you', ' ', 'know', ' ?? "', 'Cases', '" & ', 'and', ' ', 'surprises', ':', 'that', ' ', 'will', " '", 'lways', ' ', 'unknown', ' **', 'before', '**, ', 'in', ' ', '81', '% ', 'of', ' ', 'time', '$', '']
57

['With', ' ', 'a', ' ', 'sign', ' ', '#', ' ', 'written', ' ', '@', ' ', 'the', ' ', 'beginning', ' ', ',', ' ', "that's", '  ', 'a', '
', 'sentence,', '
', 'no', ' ', 'more', ' ', 'an', ' ', 'instruction!,', '	', 'you', ' ', 'know', ' ', '??', ' ', '"Cases"', ' ', '&', ' ', 'and', ' ', 'surprises:that', ' ', 'will', ' ', "'lways", ' ', 'unknown', ' ', '**before**,', ' ', 'in', ' ', '81%', ' ', 'of', ' ', 'time$']
61

Partitioning a string in Python by a regular expression

Tags:

python

regex

split

whitespace

Trey Hunner

2 Answers

NikitaBaksalyar

eyquem

Recent Activity

Donate For Us

Partitioning a string in Python by a regular expression

Tags:

python

regex

split

whitespace

Trey Hunner

2 Answers

NikitaBaksalyar

eyquem

Related questions

Recent Activity

Donate For Us