What's a Python bytestring? All I can find are topics on how to encode to bytestring or decode to ASCII or UTF-8. I'm trying to understand how it works under the hood. In a normal ASCII string, it's an array or list of characters, and each character represents an ASCII value from 0-255, so that's how you know what character is represented by the number. In Unicode, it's the 8- or 16-byte representation for the character that tells you what character it is. So what is a bytestring? How does Python know which characters to represent as what? How does it work under the hood? Since you can print or even return these strings and it shows you the string representation, I don't quite get it... Ok, so my point is definitely getting missed here. I've been told that it's an immutable sequence of bytes without any particular interpretation. A sequence of bytes.. Okay, let's say one byte: <code>'a'.encode()</code> returns <code>b'a'</code>. Simple enough. Why can I read the a? Say I get the ASCII value for a, by doing this: <code>printf "%d" "'a"</code> It returns 97. Okay, good, the integer value for the ASCII character a. If we interpret 97 as ASCII, say in a C <code>char</code>, then we get the letter <code>a</code>. Fair enough. If we convert the byte representation to bits, we get this: <code>01100001</code> 2^0 + 2^5 + 2^6 = 97. Cool. So why is <code>'a'.encode()</code> returning <code>b'a'</code> instead of <code>01100001</code>?? If it's without a particular interpretation, shouldn't it be returning something like <code>b'01100001'</code>? It seems like it's interpreting it like ASCII. Someone mentioned that it's calling <code>__repr__</code> on the bytestring, so it's displayed in human-readable form. However, even if I do something like: <pre class="prettyprint"><code>with open('testbytestring.txt', 'wb') as f: f.write(b'helloworld') </code></pre> It will still insert <code>helloworld</code> as a regular string into the file, not as a sequence of bytes... So is a bytestring in ASCII?

It is a common misconception that text is ASCII or UTF-8 or Windows-1252, and therefore bytes are text. Text is only text, in the way that images are only images. The matter of storing text or images to disk is a matter of encoding that data into a sequence of bytes. There are many ways to encode images into bytes: JPEG, PNG, SVG, and likewise many ways to encode text, ASCII, UTF-8 or Windows-1252. Once encoding has happened, bytes are just bytes. Bytes are not images anymore; they have forgotten the colors they mean; although an image format decoder can recover that information. Bytes have similarly forgotten the letters they used to be. In fact, bytes don't remember whether they were images or text at all. Only out of band knowledge (filename, media headers, etcetera) can guess what those bytes should mean, and even that can be wrong (in case of data corruption). so, in Python (Python 3), we have two types for things that might otherwise look similar; For text, we have <code>str</code>, which knows it's text; it knows which letters it's supposed to mean. It doesn't know which bytes that might be, since letters are not bytes. We also have <code>bytestring</code>, which doesn't know if it's text or images or any other kind of data. The two types are superficially similar, since they are both sequences of things, but the things that they are sequences of is quite different. Implementationally, <code>str</code> is stored in memory as <code>UCS-?</code> where the ? is implementation defined, it may be UCS-4, UCS-2 or UCS-1, depending on compile time options and which code points are present in the represented string. <hr> "But why"? Some things that look like text are actually defined in other terms. A really good example of this are the many Internet protocols of the world. For instance, HTTP is a "text" protocol that is in fact defined using the ABNF syntax common in RFCs. These protocols are expressed in terms of octets, not characters, although an informal encoding may also be suggested: <blockquote> 2.3. Terminal Values Rules resolve into a string of terminal values, sometimes called characters. In ABNF, a character is merely a non-negative integer. In certain contexts, a specific mapping (encoding) of values into a character set (such as ASCII) will be specified. </blockquote> This distinction is important, because it's not possible to send text over the internet, the only thing you can do is send bytes. saying "text but in 'foo' encoding" makes the format that much more complex, since clients and servers need to now somehow figure out the encoding business on their own, hopefully in the same way, since they must ultimately pass data around as bytes anyway. This is doubly useless since these protocols are seldom about text handling anyway, and is only a convenience for implementers. Neither the server owners nor end users are ever interested in reading the words <code>Transfer-Encoding: chunked</code>, so long as both the server and the browser understand it correctly. By comparison, when working with text, you don't really care how it's encoded. You can express the "Heävy Mëtal Ümlaüts" any way you like, except "Heδvy Mλtal άmlaόts" <hr> The distinct types thus give you a way to say "this value 'means' text" or "bytes".

Python does not know how to represent a bytestring. That's the point. When you output a character with value 97 into pretty much any output window, you'll get the character 'a' but that's not part of the implementation; it's just a thing that happens to be locally true. If you want an encoding, you don't use bytestring. If you use bytestring, you don't have an encoding. Your piece about .txt files shows you have misunderstood what is happening. You see, plain text files too don't have an encoding. They're just a series of bytes. These bytes get translated into letters by the text editor but there is no guarantee at all that someone else opening your file will see the same thing as you if you stray outside the common set of ASCII characters.

As the name implies, a Python 3 <code>bytestring</code> (or simply a <code>str</code> in Python 2.7) is a string of bytes. And, as others have pointed out, it is immutable. It is distinct from a Python 3 <code>str</code> (or, more descriptively, a <code>unicode</code> in Python 2.7) which is a string of abstract Unicode characters (a.k.a. UTF-32, though Python 3 adds fancy compression under the hood to reduce the actual memory footprint similar to UTF-8, perhaps even in a more general way). There are essentially three ways of "interpreting" these bytes. You can look at the numeric value of an element, like this: <pre class="prettyprint"><code>>>> ord(b'Hello'[0]) # Python 2.7 str 72 >>> b'Hello'[0] # Python 3 bytestring 72 </code></pre> Or you can tell Python to emit one or more elements to the terminal (or a file, device, socket, etc.) as 8-bit characters, like this: <pre class="prettyprint"><code>>>> print b'Hello'[0] # Python 2.7 str H >>> import sys >>> sys.stdout.buffer.write(b'Hello'[0:1]) and None; print() # Python 3 bytestring H </code></pre> As Jack hinted at, in this latter case it is your terminal interpreting the character, not Python. Finally, as you have seen in your own research, you can also get Python to interpret a <code>bytestring</code>. For example, you can construct an abstract <code>unicode</code> object like this in Python 2.7: <pre class="prettyprint"><code>>>> u1234 = unicode(b'\xe1\x88\xb4', 'utf-8') >>> print u1234.encode('utf-8') # if terminal supports UTF-8 ሴ >>> u1234 u'\u1234' >>> print ('%04x' % ord(u1234)) 1234 >>> type(u1234) <type 'unicode'> >>> len(u1234) 1 >>> </code></pre> Or like this in Python 3: <pre class="prettyprint"><code>>>> u1234 = str(b'\xe1\x88\xb4', 'utf-8') >>> print (u1234) # if terminal supports UTF-8 AND python auto-infers ሴ >>> u1234.encode('unicode-escape') b'\\u1234' >>> print ('%04x' % ord(u1234)) 1234 >>> type(u1234) <class 'str'> >>> len(u1234) 1 </code></pre> (and I am sure that the amount of syntax churn between Python 2.7 and Python3 around bystestring, strings, and Unicode had something to do with the continued popularity of Python 2.7. I suppose that when Python 3 was invented they didn't yet realize that everything would become UTF-8 and therefore all the fuss about abstraction was unnecessary). But the Unicode abstraction does not happen automatically if you don't want it to. The point of a <code>bytestring</code> is that you can directly get at the bytes. Even if your string happens to be a UTF-8 sequence, you can still access bytes in the sequence: <pre class="prettyprint"><code>>>> len(b'\xe1\x88\xb4') 3 >>> b'\xe1\x88\xb4'[0] '\xe1' </code></pre> And this works in both Python 2.7 and Python 3, with the difference being that in Python 2.7 you have <code>str</code>, while in Python3 you have <code>bytestring</code>. You can also do other wonderful things with <code>bytestring</code>s, like knowing if they will fit in a reserved space within a file, sending them directly over a socket, calculating the HTTP <code>content-length</code> field correctly, and avoiding Python Bug 8260. In short, use <code>bytestring</code>s when your data is processed and stored in bytes.

What is a Python bytestring?

Tags:

python

string

What's a Python bytestring?

All I can find are topics on how to encode to bytestring or decode to ASCII or UTF-8. I'm trying to understand how it works under the hood. In a normal ASCII string, it's an array or list of characters, and each character represents an ASCII value from 0-255, so that's how you know what character is represented by the number. In Unicode, it's the 8- or 16-byte representation for the character that tells you what character it is.

So what is a bytestring? How does Python know which characters to represent as what? How does it work under the hood? Since you can print or even return these strings and it shows you the string representation, I don't quite get it...

Ok, so my point is definitely getting missed here. I've been told that it's an immutable sequence of bytes without any particular interpretation.

A sequence of bytes.. Okay, let's say one byte: 'a'.encode() returns b'a'.

Simple enough. Why can I read the a?

Say I get the ASCII value for a, by doing this: printf "%d" "'a"

It returns 97. Okay, good, the integer value for the ASCII character a. If we interpret 97 as ASCII, say in a C char, then we get the letter a. Fair enough. If we convert the byte representation to bits, we get this:

01100001

2^0 + 2^5 + 2^6 = 97. Cool.

So why is 'a'.encode() returning b'a' instead of 01100001??
If it's without a particular interpretation, shouldn't it be returning something like b'01100001'?
It seems like it's interpreting it like ASCII.

Someone mentioned that it's calling __repr__ on the bytestring, so it's displayed in human-readable form. However, even if I do something like:

with open('testbytestring.txt', 'wb') as f:
    f.write(b'helloworld')

It will still insert helloworld as a regular string into the file, not as a sequence of bytes... So is a bytestring in ASCII?

489

asked Apr 02 '14 22:04

antimatter

3 Answers

It is a common misconception that text is ASCII or UTF-8 or Windows-1252, and therefore bytes are text.

Text is only text, in the way that images are only images. The matter of storing text or images to disk is a matter of encoding that data into a sequence of bytes. There are many ways to encode images into bytes: JPEG, PNG, SVG, and likewise many ways to encode text, ASCII, UTF-8 or Windows-1252.

Once encoding has happened, bytes are just bytes. Bytes are not images anymore; they have forgotten the colors they mean; although an image format decoder can recover that information. Bytes have similarly forgotten the letters they used to be. In fact, bytes don't remember whether they were images or text at all. Only out of band knowledge (filename, media headers, etcetera) can guess what those bytes should mean, and even that can be wrong (in case of data corruption).

so, in Python (Python 3), we have two types for things that might otherwise look similar; For text, we have str, which knows it's text; it knows which letters it's supposed to mean. It doesn't know which bytes that might be, since letters are not bytes. We also have bytestring, which doesn't know if it's text or images or any other kind of data.

The two types are superficially similar, since they are both sequences of things, but the things that they are sequences of is quite different.

Implementationally, str is stored in memory as UCS-? where the ? is implementation defined, it may be UCS-4, UCS-2 or UCS-1, depending on compile time options and which code points are present in the represented string.

"But why"?

Some things that look like text are actually defined in other terms. A really good example of this are the many Internet protocols of the world. For instance, HTTP is a "text" protocol that is in fact defined using the ABNF syntax common in RFCs. These protocols are expressed in terms of octets, not characters, although an informal encoding may also be suggested:

2.3. Terminal Values

Rules resolve into a string of terminal values, sometimes called characters. In ABNF, a character is merely a non-negative integer. In certain contexts, a specific mapping (encoding) of values into a character set (such as ASCII) will be specified.

This distinction is important, because it's not possible to send text over the internet, the only thing you can do is send bytes. saying "text but in 'foo' encoding" makes the format that much more complex, since clients and servers need to now somehow figure out the encoding business on their own, hopefully in the same way, since they must ultimately pass data around as bytes anyway. This is doubly useless since these protocols are seldom about text handling anyway, and is only a convenience for implementers. Neither the server owners nor end users are ever interested in reading the words Transfer-Encoding: chunked, so long as both the server and the browser understand it correctly.

By comparison, when working with text, you don't really care how it's encoded. You can express the "Heävy Mëtal Ümlaüts" any way you like, except "Heδvy Mλtal άmlaόts"

The distinct types thus give you a way to say "this value 'means' text" or "bytes".

157

answered Oct 18 '22 23:10

SingleNegationElimination

Python does not know how to represent a bytestring. That's the point.

When you output a character with value 97 into pretty much any output window, you'll get the character 'a' but that's not part of the implementation; it's just a thing that happens to be locally true. If you want an encoding, you don't use bytestring. If you use bytestring, you don't have an encoding.

Your piece about .txt files shows you have misunderstood what is happening. You see, plain text files too don't have an encoding. They're just a series of bytes. These bytes get translated into letters by the text editor but there is no guarantee at all that someone else opening your file will see the same thing as you if you stray outside the common set of ASCII characters.

answered Oct 19 '22 00:10

Jack Aidley

As the name implies, a Python 3 bytestring (or simply a str in Python 2.7) is a string of bytes. And, as others have pointed out, it is immutable.

It is distinct from a Python 3 str (or, more descriptively, a unicode in Python 2.7) which is a string of abstract Unicode characters (a.k.a. UTF-32, though Python 3 adds fancy compression under the hood to reduce the actual memory footprint similar to UTF-8, perhaps even in a more general way).

There are essentially three ways of "interpreting" these bytes. You can look at the numeric value of an element, like this:

>>> ord(b'Hello'[0])  # Python 2.7 str
72
>>> b'Hello'[0]  # Python 3 bytestring
72

Or you can tell Python to emit one or more elements to the terminal (or a file, device, socket, etc.) as 8-bit characters, like this:

>>> print b'Hello'[0] # Python 2.7 str
H
>>> import sys
>>> sys.stdout.buffer.write(b'Hello'[0:1]) and None; print() # Python 3 bytestring
H

As Jack hinted at, in this latter case it is your terminal interpreting the character, not Python.

Finally, as you have seen in your own research, you can also get Python to interpret a bytestring. For example, you can construct an abstract unicode object like this in Python 2.7:

>>> u1234 = unicode(b'\xe1\x88\xb4', 'utf-8')
>>> print u1234.encode('utf-8') # if terminal supports UTF-8
ሴ
>>> u1234
u'\u1234'
>>> print ('%04x' % ord(u1234))
1234
>>> type(u1234)
<type 'unicode'>
>>> len(u1234)
1
>>>

Or like this in Python 3:

>>> u1234 = str(b'\xe1\x88\xb4', 'utf-8')
>>> print (u1234) # if terminal supports UTF-8 AND python auto-infers
ሴ
>>> u1234.encode('unicode-escape')
b'\\u1234'
>>> print ('%04x' % ord(u1234))
1234
>>> type(u1234)
<class 'str'>
>>> len(u1234)
1

(and I am sure that the amount of syntax churn between Python 2.7 and Python3 around bystestring, strings, and Unicode had something to do with the continued popularity of Python 2.7. I suppose that when Python 3 was invented they didn't yet realize that everything would become UTF-8 and therefore all the fuss about abstraction was unnecessary).

But the Unicode abstraction does not happen automatically if you don't want it to. The point of a bytestring is that you can directly get at the bytes. Even if your string happens to be a UTF-8 sequence, you can still access bytes in the sequence:

>>> len(b'\xe1\x88\xb4')
3
>>> b'\xe1\x88\xb4'[0]
'\xe1'

And this works in both Python 2.7 and Python 3, with the difference being that in Python 2.7 you have str, while in Python3 you have bytestring.

You can also do other wonderful things with bytestrings, like knowing if they will fit in a reserved space within a file, sending them directly over a socket, calculating the HTTP content-length field correctly, and avoiding Python Bug 8260. In short, use bytestrings when your data is processed and stored in bytes.

answered Oct 18 '22 23:10

personal_cloud

Related questions
                            
                                What scalability issues are associated with NetworkX?
                            
                                BdbQuit raised when debugging Python with pdb
                            
                                Is there a way to implement methods like __len__ or __eq__ as classmethods?
                            
                                What does Python optimization (-O or PYTHONOPTIMIZE) do?
                            
                                Get a dict of all variables currently in scope and their values
                            
                                Applying LIMIT and OFFSET to all queries in SQLAlchemy
                            
                                What is the difference between native int type and the numpy.int types?
                            
                                timeit and its default_timer completely disagree
                            
                                Subclassing Python dictionary to override __setitem__
                            
                                Why was PyPI called the cheese shop?
                            
                                How to reference python package when filename contains a period
                            
                                Is it safe to use sys.platform=='win32' check on 64-bit Python?
                            
                                How to get text in QlineEdit when QpushButton is pressed in a string?
                            
                                Keep plotting window open in Matplotlib
                            
                                Is it bad form to call a classmethod as a method from an instance?
                            
                                Using flask inside class
                            
                                When are parentheses required around a tuple?
                            
                                How do I create a date picker in tkinter?
                            
                                Colour chart for Tkinter and Tix
                            
                                How to define free-variable in python?

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With