How to remove all the escape sequences from a list of strings?

Tags:

python

I want to remove all types of escape sequences from a list of strings. How can I do this? input:

['william', 'short', '\x80', 'twitter', '\xaa', '\xe2', 'video', 'guy', 'ray']

output:

['william', 'short', 'twitter', 'video', 'guy', 'ray']

http://docs.python.org/reference/lexical_analysis.html#string-literals

649

asked Nov 13 '11 22:11

2 Answers

If you want to strip out some characters you don't like, you can use the translate function to strip them out:

>>> s="\x01\x02\x10\x13\x20\x21hello world"
>>> print(s)
 !hello world
>>> s
'\x01\x02\x10\x13 !hello world'
>>> escapes = ''.join([chr(char) for char in range(1, 32)])
>>> t = s.translate(None, escapes)
>>> t
' !hello world'

This will strip out all these control characters:

   001   1     01    SOH (start of heading)
   002   2     02    STX (start of text)
   003   3     03    ETX (end of text)
   004   4     04    EOT (end of transmission)
   005   5     05    ENQ (enquiry)
   006   6     06    ACK (acknowledge)
   007   7     07    BEL '\a' (bell)
   010   8     08    BS  '\b' (backspace)
   011   9     09    HT  '\t' (horizontal tab)
   012   10    0A    LF  '\n' (new line)
   013   11    0B    VT  '\v' (vertical tab)
   014   12    0C    FF  '\f' (form feed)
   015   13    0D    CR  '\r' (carriage ret)
   016   14    0E    SO  (shift out)
   017   15    0F    SI  (shift in)
   020   16    10    DLE (data link escape)
   021   17    11    DC1 (device control 1)
   022   18    12    DC2 (device control 2)
   023   19    13    DC3 (device control 3)
   024   20    14    DC4 (device control 4)
   025   21    15    NAK (negative ack.)
   026   22    16    SYN (synchronous idle)
   027   23    17    ETB (end of trans. blk)
   030   24    18    CAN (cancel)
   031   25    19    EM  (end of medium)
   032   26    1A    SUB (substitute)
   033   27    1B    ESC (escape)
   034   28    1C    FS  (file separator)
   035   29    1D    GS  (group separator)
   036   30    1E    RS  (record separator)
   037   31    1F    US  (unit separator)

For Python newer than 3.1, the sequence is different:

>>> s="\x01\x02\x10\x13\x20\x21hello world"
>>> print(s)
 !hello world
>>> s
'\x01\x02\x10\x13 !hello world'
>>> escapes = ''.join([chr(char) for char in range(1, 32)])
>>> translator = str.maketrans('', '', escapes)
>>> t = s.translate(translator)
>>> t
' !hello world'

183

answered Oct 22 '22 23:10

sarnold

Something like this?

>>> from ast import literal_eval
>>> s = r'Hello,\nworld!'
>>> print(literal_eval("'%s'" % s))
Hello,
world!

Edit: ok, that's not what you want. What you want can't be done in general, because, as @Sven Marnach explained, strings don't actually contain escape sequences. Those are just notation in string literals.

You can filter all strings with non-ASCII characters from your list with

def is_ascii(s):
    try:
        s.decode('ascii')
        return True
    except UnicodeDecodeError:
        return False

[s for s in ['william', 'short', '\x80', 'twitter', '\xaa',
             '\xe2', 'video', 'guy', 'ray']
 if is_ascii(s)]

answered Oct 22 '22 23:10

Fred Foo

Related questions
                            
                                How can I make a video from array of images in matplotlib?
                            
                                dask dataframe how to convert column to to_datetime
                            
                                OpenCV installation stuck at [ 99%] Built target opencv_perf_stitching with no error
                            
                                How to Create Dataframe from AWS Athena using Boto3 get_query_results method
                            
                                Find equal columns between two dataframes
                            
                                Cost of list functions in Python
                            
                                Python find numbers not in set
                            
                                Exclude URLs from Django REST Swagger
                            
                                Processing Large Files in Python [ 1000 GB or More]
                            
                                Django Heroku Error "Your models have changes that are not yet reflected in a migration"
                            
                                How to convert numpy datetime64 into datetime [duplicate]
                            
                                using regex in jinja 2 for ansible playbooks
                            
                                Saving prediction results to CSV
                            
                                How does Python increment list elements?
                            
                                fatal error: ffi.h: No such file or directory on pip2 install pyOpenSSL
                            
                                It is possible to match a character repetition with regex? How?
                            
                                How to set timeout for urlfetch in Google App Engine?
                            
                                In python django how do you print out an object's introspection? The list of all public methods of that object (variable and/or functions)?
                            
                                Serializing a suds object in python
                            
                                Python Equivalent to Ruby's #each_cons?

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With