Parse text file python and covert to pandas dataframe

Question

I am trying to parse a text file, converting it into a pandas dataframe. The file (inclusive of blank lines):

HEADING1
value 1

HEADING2
value 2

HEADING1,
value 11

HEADING2
value 12

should be converted into a dataframe:

HEADING1, HEADING2
value 1, value 2
value 11, value 12

I have tried the following code. However, I am not sure using converters could work?

df = pd.read_table(textfile, header=None, skip_blank_lines=True, delimiter='
',
                   # converters= 'what should I use?',
                   names= 'HEADING1, HEADING2'.split() )

piRSquared · Accepted Answer

You parse the text yourself and split on ' '

# split file by `'

'` to get rows
# split again by `'
'` to get columns
# `zip` to get convenient lists of headers and values
cols, vals = zip(
    *[line.split('
') for line in open(textfile).read().split('

')]
)

# construct a `pd.Series`
# note: your index contained in the `cols` list will not be unique
s = pd.Series(vals, cols)

# we'll need to enumerate the duplicated index values so that we can unstack
# we do this by creating a `pd.MultiIndex` with `cumcount` then the header values
s.index = [s.groupby(level=0).cumcount(), s.index]

# finally, `unstack`
s.unstack()

   HEADING1  HEADING2
0   value 1   value 2
1  value 11  value 12

Breakdown

list comprehension

[line.split('
') for line in StringIO(txt).read().split('

')]

[['HEADING1', 'value 1'],
 ['HEADING2', 'value 2'],
 ['HEADING1', 'value 11'],
 ['HEADING2', 'value 12']]

with zip

list(zip(*[line.split('
') for line in StringIO(txt).read().split('

')]))

[('HEADING1', 'HEADING2', 'HEADING1', 'HEADING2'),
 ('value 1', 'value 2', 'value 11', 'value 12')]

setting cols and vals

cols, vals = zip(*[line.split('
') for line in StringIO(txt).read().split('

')])

print(cols)
print()
print(vals)

('HEADING1', 'HEADING2', 'HEADING1', 'HEADING2')

('value 1', 'value 2', 'value 11', 'value 12')

Making a series

s = pd.Series(vals, cols)
s

HEADING1     value 1
HEADING2     value 2
HEADING1    value 11
HEADING2    value 12
dtype: object

Enumerating the index values

s.index = [s.groupby(level=0).cumcount(), s.index]
s

0  HEADING1     value 1
   HEADING2     value 2
1  HEADING1    value 11
   HEADING2    value 12
dtype: object

unstack

s.unstack()

   HEADING1  HEADING2
0   value 1   value 2
1  value 11  value 12

Full Demo

import pandas as pd
from io import StringIO

txt = """HEADING1
value 1

HEADING2
value 2

HEADING1
value 11

HEADING2
value 12"""

cols, vals = zip(*[line.split('
') for line in StringIO(txt).read().split('

')])

s = pd.Series(vals, cols)
s.index = [s.groupby(level=0).cumcount(), s.index]

s.unstack()

   HEADING1  HEADING2
0   value 1   value 2
1  value 11  value 12

Parse text file python and covert to pandas dataframe

Tags:

python

pandas

parsing

converters

Andreuccio

1 Answers

piRSquared

Recent Activity

Donate For Us

Parse text file python and covert to pandas dataframe

Tags:

python

pandas

parsing

converters

Andreuccio

1 Answers

piRSquared

Related questions

Recent Activity

Donate For Us