I have some escaped strings that need to be unescaped. I'd like to do this in Python. For example, in Python 2.7 I can do this: <pre class="prettyprint"><code>>>> "\\123omething special".decode('string-escape') 'Something special' >>> </code></pre> How do I do it in Python 3? This doesn't work: <pre class="prettyprint"><code>>>> b"\\123omething special".decode('string-escape') Traceback (most recent call last): File "<stdin>", line 1, in <module> LookupError: unknown encoding: string-escape >>> </code></pre> My goal is to be able to take a string like this: <pre class="prettyprint"><code>s\000u\000p\000p\000o\000r\000t\000@\000p\000s\000i\000l\000o\000c\000.\000c\000o\000m\000 </code></pre> And turn it into: <pre class="prettyprint"><code>"support@psiloc.com" </code></pre> After I do the conversion, I'll probe to see if the string I have is encoded in UTF-8 or UTF-16.

You'll have to use <code>unicode_escape</code> instead: <pre class="prettyprint"><code>>>> b"\\123omething special".decode('unicode_escape') </code></pre> If you start with a <code>str</code> object instead (equivalent to the python 2.7 unicode) you'll need to encode to bytes first, then decode with <code>unicode_escape</code>. If you need bytes as end result, you'll have to encode again to a suitable encoding (<code>.encode('latin1')</code> for example, if you need to preserve literal byte values; the first 256 Unicode code points map 1-on-1). Your example is actually UTF-16 data with escapes. Decode from <code>unicode_escape</code>, back to <code>latin1</code> to preserve the bytes, then from <code>utf-16-le</code> (UTF 16 little endian without BOM): <pre class="prettyprint"><code>>>> value = b's\\000u\\000p\\000p\\000o\\000r\\000t\\000@\\000p\\000s\\000i\\000l\\000o\\000c\\000.\\000c\\000o\\000m\\000' >>> value.decode('unicode_escape').encode('latin1') # convert to bytes b's\x00u\x00p\x00p\x00o\x00r\x00t\x00@\x00p\x00s\x00i\x00l\x00o\x00c\x00.\x00c\x00o\x00m\x00' >>> _.decode('utf-16-le') # decode from UTF-16-LE 'support@psiloc.com' </code></pre>

The old "string-escape" codec maps bytestrings to bytestrings, and there's been a lot of debate about what to do with such codecs, so it isn't currently available through the standard encode/decode interfaces. BUT, the code is still there in the C-API (as <code>PyBytes_En/DecodeEscape</code>), and this is still exposed to Python via the undocumented <code>codecs.escape_encode</code> and <code>codecs.escape_decode</code>. <pre class="prettyprint"><code>>>> import codecs >>> codecs.escape_decode(b"ab\\xff") (b'ab\xff', 6) >>> codecs.escape_encode(b"ab\xff") (b'ab\\xff', 3) </code></pre> These functions return the transformed <code>bytes</code> object, plus a number indicating how many bytes were processed... you can just ignore the latter. <pre class="prettyprint"><code>>>> value = b's\\000u\\000p\\000p\\000o\\000r\\000t\\000@\\000p\\000s\\000i\\000l\\000o\\000c\\000.\\000c\\000o\\000m\\000' >>> codecs.escape_decode(value)[0] b's\x00u\x00p\x00p\x00o\x00r\x00t\x00@\x00p\x00s\x00i\x00l\x00o\x00c\x00.\x00c\x00o\x00m\x00' </code></pre>

how do I .decode('string-escape') in Python3?

Tags:

python

python-3.x

escaping

I have some escaped strings that need to be unescaped. I'd like to do this in Python.

For example, in Python 2.7 I can do this:

>>> "\\123omething special".decode('string-escape') 'Something special' >>>

How do I do it in Python 3? This doesn't work:

>>> b"\\123omething special".decode('string-escape') Traceback (most recent call last):   File "<stdin>", line 1, in <module> LookupError: unknown encoding: string-escape >>>

My goal is to be able to take a string like this:

s\000u\000p\000p\000o\000r\000t\000@\000p\000s\000i\000l\000o\000c\000.\000c\000o\000m\000

And turn it into:

"[email protected]"

After I do the conversion, I'll probe to see if the string I have is encoded in UTF-8 or UTF-16.

636

asked Feb 11 '13 20:02

vy32

Video Answer

2 Answers

You'll have to use unicode_escape instead:

>>> b"\\123omething special".decode('unicode_escape')

If you start with a str object instead (equivalent to the python 2.7 unicode) you'll need to encode to bytes first, then decode with unicode_escape.

If you need bytes as end result, you'll have to encode again to a suitable encoding (.encode('latin1') for example, if you need to preserve literal byte values; the first 256 Unicode code points map 1-on-1).

Your example is actually UTF-16 data with escapes. Decode from unicode_escape, back to latin1 to preserve the bytes, then from utf-16-le (UTF 16 little endian without BOM):

>>> value = b's\\000u\\000p\\000p\\000o\\000r\\000t\\000@\\000p\\000s\\000i\\000l\\000o\\000c\\000.\\000c\\000o\\000m\\000' >>> value.decode('unicode_escape').encode('latin1')  # convert to bytes b's\x00u\x00p\x00p\x00o\x00r\x00t\x00@\x00p\x00s\x00i\x00l\x00o\x00c\x00.\x00c\x00o\x00m\x00' >>> _.decode('utf-16-le') # decode from UTF-16-LE '[email protected]'

105

answered Oct 10 '22 10:10

Martijn Pieters

The old "string-escape" codec maps bytestrings to bytestrings, and there's been a lot of debate about what to do with such codecs, so it isn't currently available through the standard encode/decode interfaces.

BUT, the code is still there in the C-API (as PyBytes_En/DecodeEscape), and this is still exposed to Python via the undocumented codecs.escape_encode and codecs.escape_decode.

>>> import codecs >>> codecs.escape_decode(b"ab\\xff") (b'ab\xff', 6) >>> codecs.escape_encode(b"ab\xff") (b'ab\\xff', 3)

These functions return the transformed bytes object, plus a number indicating how many bytes were processed... you can just ignore the latter.

>>> value = b's\\000u\\000p\\000p\\000o\\000r\\000t\\000@\\000p\\000s\\000i\\000l\\000o\\000c\\000.\\000c\\000o\\000m\\000' >>> codecs.escape_decode(value)[0] b's\x00u\x00p\x00p\x00o\x00r\x00t\x00@\x00p\x00s\x00i\x00l\x00o\x00c\x00.\x00c\x00o\x00m\x00'

answered Oct 10 '22 10:10

Nathaniel J. Smith

Related questions
                            
                                assigning class variable as default value to class method argument
                            
                                Shortest way to get first item of `OrderedDict` in Python 3
                            
                                Importing installed package from script raises "AttributeError: module has no attribute" or "ImportError: cannot import name"
                            
                                What's the difference between "virtualenv" and "-m venv" in creating Virtual environments(Python)
                            
                                Python: finding uid/gid for a given username/groupname (for os.chown)
                            
                                Difference between the built-in pow() and math.pow() for floats, in Python?
                            
                                Slice indices must be integers or None or have __index__ method
                            
                                Unable log in to the django admin page with a valid username and password
                            
                                f-strings vs str.format()
                            
                                Visual Studio Code: Intellisense not working
                            
                                Parsing files (ics/ icalendar) using Python
                            
                                Best practice for setting the default value of a parameter that's supposed to be a list in Python?
                            
                                matching any character including newlines in a Python regex subexpression, not globally
                            
                                How are exceptions implemented under the hood? [closed]
                            
                                Using moviepy, scipy and numpy in amazon lambda
                            
                                how to merge two data frames based on particular column in pandas python?
                            
                                How to write an inline-comment in Python
                            
                                PIP Install Numpy throws an error "ascii codec can't decode byte 0xe2"
                            
                                Why is range(0) == range(2, 2, 2) True in Python 3?
                            
                                Pandas: sum up multiple columns into one column without last column

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With