Here is my problem, I have a dataframe like this : <pre class="prettyprint"><code> Depr_1 Depr_2 Depr_3 S3 0 5 9 S2 4 11 8 S1 6 11 12 S5 0 4 11 S4 4 8 8 </code></pre> and I just want to calculate the mean over the full dataframe, as the following doesn't work : <pre class="prettyprint"><code>df.mean() </code></pre> Then I came up with : <pre class="prettyprint"><code>df.mean().mean() </code></pre> But this trick won't work for computing the standard deviation. My final attempts were : <pre class="prettyprint"><code>df.get_values().mean() df.get_values().std() </code></pre> Except that in the latter case, it uses mean() and std() function from numpy. It's not a problem for the mean, but it is for std, as the pandas function uses by default <code>ddof=1</code>, unlike the numpy one where <code>ddof=0</code>.

You could convert the dataframe to be a single column with <code>stack</code> (this changes the shape from 5x3 to 15x1) and then take the standard deviation: <pre class="prettyprint"><code>df.stack().std() # pandas default degrees of freedom is one </code></pre> Alternatively, you can use <code>values</code> to convert from a pandas dataframe to a numpy array before taking the standard deviation: <pre class="prettyprint"><code>df.values.std(ddof=1) # numpy default degrees of freedom is zero </code></pre> Unlike pandas, numpy will give the standard deviation of the entire array by default, so there is no need to reshape before taking the standard deviation. A couple of additional notes: <ul> <li>The numpy approach here is a bit faster than the pandas one, which is generally true when you have the option to accomplish the same thing with either numpy or pandas. The speed difference will depend on the size of your data, but numpy was roughly 10x faster when I tested a few different sized dataframes on my laptop (numpy version 1.15.4 and pandas version 0.23.4).</li> <li>The numpy and pandas approaches here will not give exactly the same answers, but will be extremely close (identical at several digits of precision). The discrepancy is due to slight differences in implementation behind the scenes that affect how the floating point values get rounded.</li> </ul>

Pandas : compute mean or std (standard deviation) over entire dataframe

Tags:

python

pandas

numpy

Here is my problem, I have a dataframe like this :

    Depr_1  Depr_2  Depr_3 S3  0   5   9 S2  4   11  8 S1  6   11  12 S5  0   4   11 S4  4   8   8

and I just want to calculate the mean over the full dataframe, as the following doesn't work :

df.mean()

Then I came up with :

df.mean().mean()

But this trick won't work for computing the standard deviation. My final attempts were :

df.get_values().mean() df.get_values().std()

Except that in the latter case, it uses mean() and std() function from numpy. It's not a problem for the mean, but it is for std, as the pandas function uses by default ddof=1, unlike the numpy one where ddof=0.

647

asked Aug 05 '14 14:08

jrjc

1 Answers

You could convert the dataframe to be a single column with stack (this changes the shape from 5x3 to 15x1) and then take the standard deviation:

df.stack().std()         # pandas default degrees of freedom is one

Alternatively, you can use values to convert from a pandas dataframe to a numpy array before taking the standard deviation:

df.values.std(ddof=1)    # numpy default degrees of freedom is zero

Unlike pandas, numpy will give the standard deviation of the entire array by default, so there is no need to reshape before taking the standard deviation.

A couple of additional notes:

The numpy approach here is a bit faster than the pandas one, which is generally true when you have the option to accomplish the same thing with either numpy or pandas. The speed difference will depend on the size of your data, but numpy was roughly 10x faster when I tested a few different sized dataframes on my laptop (numpy version 1.15.4 and pandas version 0.23.4).
The numpy and pandas approaches here will not give exactly the same answers, but will be extremely close (identical at several digits of precision). The discrepancy is due to slight differences in implementation behind the scenes that affect how the floating point values get rounded.

answered Oct 11 '22 09:10

JohnE

Related questions
                            
                                Handle either a list or single integer as an argument
                            
                                How can I create a Word document using Python? [closed]
                            
                                f.write vs print >> f
                            
                                Postgres SSL SYSCALL error: EOF detected with python and psycopg
                            
                                Matplotlib showing x-tick labels overlapping
                            
                                JSON.stringify (Javascript) and json.dumps (Python) not equivalent on a list?
                            
                                How to create conda environment with specific python version?
                            
                                Matplotlib overlapping annotations
                            
                                a good python to exe compiler? [closed]
                            
                                Can I have a Django form without Model
                            
                                Python: Collections.Counter vs defaultdict(int)
                            
                                "ImportError: No module named httplib2" even after installation
                            
                                Python JSON encoder convert NaNs to null instead
                            
                                Total size of serialized results of 16 tasks (1048.5 MB) is bigger than spark.driver.maxResultSize (1024.0 MB)
                            
                                Is there a static constructor or static initializer in Python?
                            
                                Assignment with "or" in python [closed]
                            
                                Python: Make a video using several .png images [closed]
                            
                                How can I accomplish `set_xlim` or `set_ylim` in Bokeh?
                            
                                Tokenize a paragraph into sentence and then into words in NLTK
                            
                                PDB - stepping out of a function

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With