How to vectorize a loop through pandas series when values are used in slice of another series

Tags:

Suppose I have two series of timestamps which are pairs of start/end times for various 5 hour ranges. They are not necessarily sequential, nor are they quantized to the hour.

import pandas as pd

start = pd.Series(pd.date_range('20190412',freq='H',periods=25))

# Drop a few indexes to make the series not sequential
start.drop([4,5,10,14]).reset_index(drop=True,inplace=True)

# Add some random minutes to the start as it's not necessarily quantized
start = start + pd.to_timedelta(np.random.randint(59,size=len(start)),unit='T')

end = start + pd.Timedelta('5H')

Now suppose that we have some data that is timestamped by minute, over a range that encompasses all start/end pairs.

data_series = pd.Series(data=np.random.randint(20, size=(75*60)), 
                        index=pd.date_range('20190411',freq='T',periods=(75*60)))

We wish to obtain the values from the data_series within the range of each start and end time. This can be done naively inside a loop

frm = []
for s,e in zip(start,end):
    frm.append(data_series.loc[s:e].values)

As we can see this naive approach loops over each pair of start and end dates, gets the values from data.

However this implementation is slow if len(start) is large. Is there a way to perform this sort of logic leveraging pandas vector functions?

I feel it is almost like I want to to apply .loc with a vector or pd.Series rather than a single pd.Timestamp?

EDIT

Using .apply is no more/marginally more efficient than using the naive for loop. I was hoping to be pointed in direction of a pure vector solution

411

asked Apr 12 '19 12:04

mch56

1 Answers

Basic Idea

As usual pandas would spend time on searching for that one specific index at data_series.loc[s:e], where s and e are datetime indices. That's costly when looping and that's exactly where we would improve. We would find all those indices in a vectorized manner with searchsorted. Then, we would extract the values off data_series as an array and use those indices obtained from searchsorted with simple integer-based indexing. Thus, there would be a loop with minimal work of simple-slicing off an array.

General mantra being - Do most work with pre-processing in a vectorized manner and minimal when looping.

The implementation would look something like this -

def select_slices_by_index(data_series, start, end):
    idx = data_series.index.values
    S = np.searchsorted(idx,start.values)
    E = np.searchsorted(idx,end.values)
    ar = data_series.values
    return [ar[i:j] for (i,j) in zip(S,E+1)]

Use `NumPy-striding`

For the specific case when the time-period between starts and ends are same for all entries and all slices are covered by that length, i.e. no out-of-bounds cases, we can use NumPy's sliding window trick.

We can leverage np.lib.stride_tricks.as_strided based scikit-image's view_as_windows to get sliding windows. More info on use of as_strided based view_as_windows.

from skimage.util.shape import view_as_windows

def select_slices_by_index_strided(data_series, start, end):
    idx = data_series.index.values
    L = np.searchsorted(idx,end.values[0])-np.searchsorted(idx,start.values[0])+1
    S = np.searchsorted(idx,start.values)
    ar = data_series.values
    w = view_as_windows(ar,L)
    return w[S]

Use this post if you don't have access to scikit-image.

Benchmarking

Let's scale up everything by 100x on the given sample data and test out.

Setup -

np.random.seed(0)
start = pd.Series(pd.date_range('20190412',freq='H',periods=2500))

# Drop a few indexes to make the series not sequential
start.drop([4,5,10,14]).reset_index(drop=True,inplace=True)

# Add some random minutes to the start as it's not necessarily quantized
start = start + pd.to_timedelta(np.random.randint(59,size=len(start)),unit='T')

end = start + pd.Timedelta('5H')
data_series = pd.Series(data=np.random.randint(20, size=(750*600)), 
                        index=pd.date_range('20190411',freq='T',periods=(750*600)))

Timings -

In [156]: %%timeit
     ...: frm = []
     ...: for s,e in zip(start,end):
     ...:     frm.append(data_series.loc[s:e].values)
1 loop, best of 3: 172 ms per loop

In [157]: %timeit select_slices_by_index(data_series, start, end)
1000 loops, best of 3: 1.23 ms per loop

In [158]: %timeit select_slices_by_index_strided(data_series, start, end)
1000 loops, best of 3: 994 µs per loop

In [161]: frm = []
     ...: for s,e in zip(start,end):
     ...:     frm.append(data_series.loc[s:e].values)

In [162]: np.allclose(select_slices_by_index(data_series, start, end),frm)
Out[162]: True

In [163]: np.allclose(select_slices_by_index_strided(data_series, start, end),frm)
Out[163]: True

140x+ and 170x speedups with these ones!

answered Sep 28 '22 03:09

Divakar

Related questions
                            
                                Binning continuous values with round() creates artifacts
                            
                                Django 'SessionStore' object has no attribute '_session_cache'
                            
                                How to create a Dataflow pipeline from Pub/Sub to GCS in Python
                            
                                Export data from Python to Tableau directly
                            
                                Celery task cannot be called (missing positional arguments) from Django app
                            
                                pyspark: Method isBarrier([]) does not exist
                            
                                Error in loading the model with load_weights in Keras
                            
                                How to POST the refresh token to Flask JWT Extended?
                            
                                How does thread pooling works, and how to implement it in an async/await env like NodeJS?
                            
                                Virtualenv python with AWS codebuild: why the deactivate command is not found?
                            
                                Is there a way to log number of retries made
                            
                                dylib cannot load libstd when compiled in a workspace
                            
                                Does cython support dataclasses or something similar
                            
                                Why does python implementation use 9 times more memory than C?
                            
                                TensorFlow 2.0 Keras: How to write image summaries for TensorBoard
                            
                                Inner most dimension of an Array
                            
                                Is there a way to change this nested loop into a recursive loop?
                            
                                How to calculate feature importance in each models of cross validation in sklearn
                            
                                'Tensor' object has no attribute 'numpy' in tf.function in TF 2.0
                            
                                Determine allocation of values - Python

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

How to vectorize a loop through pandas series when values are used in slice of another series

Tags:

python

pandas

vectorization

series

time-series

mch56

People also ask

1 Answers

Basic Idea

Use `NumPy-striding`

Benchmarking

Divakar

Recent Activity

Donate For Us

How to vectorize a loop through pandas series when values are used in slice of another series

Tags:

python

pandas

vectorization

series

time-series

mch56

People also ask

1 Answers

Basic Idea

Use NumPy-striding

Benchmarking

Divakar

Related questions

Recent Activity

Donate For Us

Use `NumPy-striding`