Parsing non-standard XML (CDATA tag)

Question

When I want to parsing XML document in Python using BeautifulSoup library, I faced some problems. The XML document that I want to parse:

<item>
<title><![CDATA[Title Sample]]></title>
<link /><![CDATA[http://banhada.kr/?cateCode=09&viewCode=S0941580]]>
<time_start>2011-10-10 09:00:00</time_start>
<time_end>2011-10-17 09:00:00</time_end>
<price_original>35000</price_original>
<price_now>20000</price_now>
</item>

As you can see above, tag is a little strange. In my opinion, that( tag) is not a stand XML form, right? How can I parse this terrible form?

John Machin · Accepted Answer

You don't need BeautifulStoneSoup or lxml. Python's included batteries do the job just fine, and there doesn't seem to be anything non-compliant about your XML.

>>> content='''\
... <item>
... <title><![CDATA[Title Sample]]></title>
... <link /><![CDATA[http://banhada.kr/?cateCode=09&viewCode=S0941580]]>
... <time_start>2011-10-10 09:00:00</time_start>
... <time_end>2011-10-17 09:00:00</time_end>
... <price_original>35000</price_original>
... <price_now>20000</price_now>
... </item>'''
>>> import xml.etree.cElementTree as et
>>> foo = et.XML(content)
>>> for e in foo:
...     print e.tag, e.text, repr(e.tail)
...
title Title Sample '
'
link None 'http://banhada.kr/?cateCode=09&viewCode=S0941580
'
time_start 2011-10-10 09:00:00 '
'
time_end 2011-10-17 09:00:00 '
'
price_original 35000 '
'
price_now 20000 '
'
>>>

Parsing non-standard XML (CDATA tag)

Tags:

python

xml

beautifulsoup

user513004

1 Answers

John Machin

Recent Activity

Donate For Us

Parsing non-standard XML (CDATA tag)

Tags:

python

xml

beautifulsoup

user513004

1 Answers

John Machin

Related questions

Recent Activity

Donate For Us