Rename the less frequent categories by "OTHER" python

Tags:

In my dataframe I have some categorical columns with over 100 different categories. I want to rank the categories by the most frequent. I keep the first 9 most frequent categories and the less frequent categories rename them automatically by: OTHER

Example:

Here my df :

print(df)

    Employee_number                 Jobrol
0                 1        Sales Executive
1                 2     Research Scientist
2                 3  Laboratory Technician
3                 4        Sales Executive
4                 5     Research Scientist
5                 6  Laboratory Technician
6                 7        Sales Executive
7                 8     Research Scientist
8                 9  Laboratory Technician
9                10        Sales Executive
10               11     Research Scientist
11               12  Laboratory Technician
12               13        Sales Executive
13               14     Research Scientist
14               15  Laboratory Technician
15               16        Sales Executive
16               17     Research Scientist
17               18     Research Scientist
18               19                Manager
19               20        Human Resources
20               21        Sales Executive


valCount = df['Jobrol'].value_counts()

valCount

Sales Executive          7
Research Scientist       7
Laboratory Technician    5
Manager                  1
Human Resources          1

I keep the first 3 categories then I rename the rest by "OTHER", how should I proceed?

Thanks.

876

asked Dec 06 '18 09:12

Ib D

2 Answers

Convert your series to categorical, extract categories whose counts are not in the top 3, add a new category e.g. 'Other', then replace the previously calculated categories:

df['Jobrol'] = df['Jobrol'].astype('category')

others = df['Jobrol'].value_counts().index[3:]
label = 'Other'

df['Jobrol'] = df['Jobrol'].cat.add_categories([label])
df['Jobrol'] = df['Jobrol'].replace(others, label)

Note: It's tempting to combine categories by renaming them via df['Jobrol'].cat.rename_categories(dict.fromkeys(others, label)), but this won't work as this will imply multiple identically labeled categories, which isn't possible.

The above solution can be adapted to filter by count. For example, to include only categories with a count of 1 you can define others as so:

counts = df['Jobrol'].value_counts()
others = counts[counts == 1].index

177

answered Nov 11 '22 08:11

jpp

Use value_counts with numpy.where:

need = df['Jobrol'].value_counts().index[:3]
df['Jobrol'] = np.where(df['Jobrol'].isin(need), df['Jobrol'], 'OTHER')

valCount = df['Jobrol'].value_counts()
print (valCount)
Research Scientist       7
Sales Executive          7
Laboratory Technician    5
OTHER                    2
Name: Jobrol, dtype: int64

Another solution:

N = 3
s = df['Jobrol'].value_counts()
valCount = s.iloc[:N].append(pd.Series(s.iloc[N:].sum(), index=['OTHER']))
print (valCount)
Research Scientist       7
Sales Executive          7
Laboratory Technician    5
OTHER                    2
dtype: int64

answered Nov 11 '22 10:11

jezrael

Related questions
                            
                                Python wait Slurm job?
                            
                                How do I stagger or offset x-axis labels in Matplotlib?
                            
                                How to plot scipy.hierarchy.dendrogram using polar coordinates?
                            
                                Fatal Python error: init_sys_streams: can't initialize sys standard streams AttributeError: module 'io' has no attribute 'OpenWrapper'
                            
                                LinearConstraint in scipy.optimize
                            
                                matplotlib get_color for subplot
                            
                                how to set label for each subplot in a plot in matplotlib?
                            
                                Python how to remove last comma from print(string, end=“, ”)
                            
                                Get a Discord Role by Id
                            
                                How to remove nan and inf values from a numpy matrix?
                            
                                How to select an inter-year period with xarray?
                            
                                Why opening and iterating over file handle over twice as fast in Python 2 vs Python 3?
                            
                                Reusing Tensorflow session in multiple threads causes crash
                            
                                InvalidArgumentError: input_1:0 is both fed and fetched
                            
                                Moving QSlider to Mouse Click Position
                            
                                Better method to iterate over 3 lists
                            
                                Can static variables be declared as private in python?
                            
                                compare a list with values in dictionary
                            
                                Modify seaborn line relplot legend title
                            
                                Dataflow/apache beam - how to access current filename when passing in pattern?

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

Rename the less frequent categories by "OTHER" python

Tags:

python

pandas

dataframe

counter

categorical-data

Ib D

People also ask

2 Answers

jpp

jezrael

Recent Activity

Donate For Us