Drop all duplicate rows across multiple columns in Python Pandas

Tags:

The pandas drop_duplicates function is great for "uniquifying" a dataframe. However, one of the keyword arguments to pass is take_last=True or take_last=False, while I would like to drop all rows which are duplicates across a subset of columns. Is this possible?

    A   B   C
0   foo 0   A
1   foo 1   A
2   foo 1   B
3   bar 1   A

As an example, I would like to drop rows which match on columns A and C so this should drop rows 0 and 1.

670

asked May 15 '14 00:05

Jamie Bull

Video Answer

6 Answers

This is much easier in pandas now with drop_duplicates and the keep parameter.

import pandas as pd
df = pd.DataFrame({"A":["foo", "foo", "foo", "bar"], "B":[0,1,1,1], "C":["A","A","B","A"]})
df.drop_duplicates(subset=['A', 'C'], keep=False)

135

answered Sep 30 '22 17:09

Ben

Just want to add to Ben's answer on drop_duplicates:

keep : {‘first’, ‘last’, False}, default ‘first’

first : Drop duplicates except for the first occurrence.
last : Drop duplicates except for the last occurrence.
False : Drop all duplicates.

So setting keep to False will give you desired answer.

DataFrame.drop_duplicates(*args, **kwargs) Return DataFrame with duplicate rows removed, optionally only considering certain columns

Parameters: subset : column label or sequence of labels, optional Only consider certain columns for identifying duplicates, by default use all of the columns keep : {‘first’, ‘last’, False}, default ‘first’ first : Drop duplicates except for the first occurrence. last : Drop duplicates except for the last occurrence. False : Drop all duplicates. take_last : deprecated inplace : boolean, default False Whether to drop duplicates in place or to return a copy cols : kwargs only argument of subset [deprecated] Returns: deduplicated : DataFrame

answered Oct 01 '22 17:10

Jake

If you want result to be stored in another dataset:

df.drop_duplicates(keep=False)

df.drop_duplicates(keep=False, inplace=False)

If same dataset needs to be updated:

df.drop_duplicates(keep=False, inplace=True)

Above examples will remove all duplicates and keep one, similar to DISTINCT * in SQL

answered Sep 30 '22 17:09

Ramanujam Allam

use groupby and filter

import pandas as pd
df = pd.DataFrame({"A":["foo", "foo", "foo", "bar"], "B":[0,1,1,1], "C":["A","A","B","A"]})
df.groupby(["A", "C"]).filter(lambda df:df.shape[0] == 1)

answered Sep 30 '22 17:09

HYRY

Try these various things

df = pd.DataFrame({"A":["foo", "foo", "foo", "bar","foo"], "B":[0,1,1,1,1], "C":["A","A","B","A","A"]})

>>>df.drop_duplicates( "A" , keep='first')

>>>df.drop_duplicates( keep='first')

>>>df.drop_duplicates( keep='last')

answered Sep 30 '22 17:09

Priyansh gupta

Actually, drop rows 0 and 1 only requires (any observations containing matched A and C is kept.):

In [335]:

df['AC']=df.A+df.C
In [336]:

print df.drop_duplicates('C', take_last=True) #this dataset is a special case, in general, one may need to first drop_duplicates by 'c' and then by 'a'.
     A  B  C    AC
2  foo  1  B  fooB
3  bar  1  A  barA

[2 rows x 4 columns]

But I suspect what you really want is this (one observation containing matched A and C is kept.):

In [337]:

print df.drop_duplicates('AC')
     A  B  C    AC
0  foo  0  A  fooA
2  foo  1  B  fooB
3  bar  1  A  barA

[3 rows x 4 columns]

Edit:

Now it is much clearer, therefore:

In [352]:
DG=df.groupby(['A', 'C'])   
print pd.concat([DG.get_group(item) for item, value in DG.groups.items() if len(value)==1])
     A  B  C
2  foo  1  B
3  bar  1  A

[2 rows x 3 columns]

answered Sep 29 '22 17:09

CT Zhu

Related questions
                            
                                Setting different color for each series in scatter plot on matplotlib
                            
                                Iterating through directories with Python
                            
                                How can I one hot encode in Python?
                            
                                How to do multiple arguments to map function where one remains the same in python?
                            
                                Pandas dataframe get first row of each group
                            
                                Python - Get path of root project structure
                            
                                Python - Extracting and Saving Video Frames
                            
                                Print list without brackets in a single row
                            
                                How can I filter a date of a DateTimeField in Django?
                            
                                Get MD5 hash of big files in Python
                            
                                How do I get a list of all the duplicate items using pandas in python?
                            
                                psycopg2: insert multiple rows with one query
                            
                                How to split a dos path into its components in Python
                            
                                Python in Xcode 4+?
                            
                                RuntimeError on windows trying python multiprocessing
                            
                                Numpy: Get random set of rows from 2D array
                            
                                How to get the python.exe location programmatically? [duplicate]
                            
                                What is the inverse function of zip in python? [duplicate]
                            
                                How do I raise the same Exception with a custom message in Python?
                            
                                Read specific columns from a csv file with csv module?

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

Drop all duplicate rows across multiple columns in Python Pandas

Tags:

python

pandas

duplicates

drop-duplicates

Jamie Bull

People also ask

Video Answer

6 Answers

Ben

Jake

Ramanujam Allam

HYRY

Priyansh gupta

Edit:

CT Zhu

Recent Activity

Donate For Us