Logo Questions Linux Laravel Mysql Ubuntu Git Menu
 

Is there a difference between `%`-format operator and `str.format()` in python regarding unicode and utf-8 encoding?

Assume that

n = u"Tübingen"
repr(n) # `T\xfcbingen` # Unicode
i = 1 # integer

The first of the following files throws

UnicodeEncodeError: 'ascii' codec can't encode character u'\xfc' in position 82: ordinal not in range(128)

When I do n.encode('utf8') it works.

The second works flawless in both cases.

# Python File 1
#
#!/usr/bin/env python -B
# encoding: utf-8

print '{id}, {name}'.format(id=i, name=n)

# Python File 2
#
#!/usr/bin/env python -B
# encoding: utf-8

print '%i, %s'% (i, n)

Since in the documentation it is encouraged to use format() instead of the % format operator, I don't understand why format() seems more "handicaped". Does format() only work with utf8-strings?

like image 450
Aufwind Avatar asked Dec 22 '11 11:12

Aufwind


1 Answers

You're using string.format while you don't have a string but an unicode object.

print u'{id}, {name}'.format(id=i, name=n)

will work, since it uses unicode.format instead.

like image 165
Tom van der Woerdt Avatar answered Sep 27 '22 23:09

Tom van der Woerdt