TypeError: first argument must be an iterable of pandas objects, you passed an object of type "DataFrame"

Tags:

I have a big dataframe and I try to split that and after concat that. I use

df2 = pd.read_csv('et_users.csv', header=None, names=names2, chunksize=100000) for chunk in df2:     chunk['ID'] = chunk.ID.map(rep.set_index('member_id')['panel_mm_id'])  df2 = pd.concat(chunk, ignore_index=True)

But it return an error

TypeError: first argument must be an iterable of pandas objects, you passed an object of type "DataFrame"

How can I fix that?

415

asked Sep 16 '16 15:09

Petr Petrov

2 Answers

I was getting the same issue, and just realised that we have to pass the (multiple!) dataframes as a LIST in the first argument instead of as multiple arguments!

Reference: https://pandas.pydata.org/pandas-docs/stable/generated/pandas.concat.html

a = pd.DataFrame() b = pd.DataFrame() c = pd.concat(a,b) # errors out: TypeError: first argument must be an iterable of pandas objects, you passed an object of type "DataFrame"  c = pd.concat([a,b]) # works.

If the processing action doesn't require ALL the data to be present, then is no reason to keep saving all the chunks to an external array and process everything only after the chunking loop is over: that defeats the whole purpose of chunking. We use chunksize because we want to do the processing at each chunk and free up the memory for the next chunk.

In terms of OP's code, they need to create another empty dataframe and concat the chunks into there.

df3 = pd.DataFrame() # create empty df for collecting chunks df2 = pd.read_csv('et_users.csv', header=None, names=names2, chunksize=100000) for chunk in df2:     chunk['ID'] = chunk.ID.map(rep.set_index('member_id')['panel_mm_id'])     df3 = pd.concat([df3,chunk], ignore_index=True)  print(df3)

However, I'd like to reiterate that chunking was invented precisely to avoid building up all the rows of the entire CSV into a single DataFrame, as that is what causes out-of-memory errors when dealing with large CSVs. We don't want to just shift the error down the road from the pd.read_csv() line to the pd.concat() line. We need to craft ways to finish off the bulk of our data processing inside the chunking loop. In my own use case I'm eliminating away most of the rows using a df query and concatenating only the fewer required rows, so the final df is much smaller than the original csv.

112

answered Sep 22 '22 04:09

Nikhil VJ

IIUC you want the following:

df2 = pd.read_csv('et_users.csv', header=None, names=names2, chunksize=100000) chunks=[] for chunk in df2:     chunk['ID'] = chunk.ID.map(rep.set_index('member_id')['panel_mm_id'])     chunks.append(chunk)  df2 = pd.concat(chunks, ignore_index=True)

You need to append each chunk to a list and then use concat to concatenate them all, also I think the ignore_index may not be necessary but I may be wrong

answered Sep 23 '22 04:09

EdChum

Related questions
                            
                                How to make "int" parse blank strings?
                            
                                Python/NumPy first occurrence of subarray
                            
                                Security of Python's eval() on untrusted strings?
                            
                                AttributeError: can't set attribute
                            
                                Django - How to prepopulate admin form fields
                            
                                Converting utc time string to datetime object
                            
                                How to do row-to-column transposition of data in csv table?
                            
                                How can I tell if IPython is running?
                            
                                Feature comparison between npm, pip, pipenv and poetry package managers [closed]
                            
                                How do I split a huge text file in python
                            
                                Changing multiple column names but not all of them - Pandas Python
                            
                                How to convert a list of longs into a comma separated string in python [duplicate]
                            
                                How to install numpy on windows using pip install?
                            
                                Add padding to images to get them into the same shape
                            
                                Pandas Append Not Working
                            
                                Python os.system without the output
                            
                                extract digits in a simple way from a python string [duplicate]
                            
                                Python JSON encoder to support datetime?
                            
                                Python - Way to recursively find and replace string in text files
                            
                                Stack data structure in python

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

TypeError: first argument must be an iterable of pandas objects, you passed an object of type "DataFrame"

Tags:

python

pandas

dataframe

Petr Petrov

People also ask

2 Answers

Nikhil VJ

EdChum

Recent Activity

Donate For Us