Apply a method to a list of objects in parallel using multi-processing

Tags:

I have created a class with a number of methods. One of the methods is very time consuming, my_process, and I'd like to do that method in parallel. I came across Python Multiprocessing - apply class method to a list of objects but I'm not sure how to apply it to my problem, and what effect it will have on the other methods of my class.

class MyClass():
    def __init__(self, input):
        self.input = input
        self.result = int

    def my_process(self, multiply_by, add_to):
        self.result = self.input * multiply_by
        self._my_sub_process(add_to)
        return self.result

    def _my_sub_process(self, add_to):
        self.result += add_to

list_of_numbers = range(0, 5)
list_of_objects = [MyClass(i) for i in list_of_numbers]
list_of_results = [obj.my_process(100, 1) for obj in list_of_objects] # multi-process this for-loop

print list_of_numbers
print list_of_results

[0, 1, 2, 3, 4]
[1, 101, 201, 301, 401]

438

asked Mar 24 '17 14:03

bluprince13

4 Answers

I'm going to go against the grain here, and suggest sticking to the simplest thing that could possibly work ;-) That is, Pool.map()-like functions are ideal for this, but are restricted to passing a single argument. Rather than make heroic efforts to worm around that, simply write a helper function that only needs a single argument: a tuple. Then it's all easy and clear.

Here's a complete program taking that approach, which prints what you want under Python 2, and regardless of OS:

class MyClass():
    def __init__(self, input):
        self.input = input
        self.result = int

    def my_process(self, multiply_by, add_to):
        self.result = self.input * multiply_by
        self._my_sub_process(add_to)
        return self.result

    def _my_sub_process(self, add_to):
        self.result += add_to

import multiprocessing as mp
NUM_CORE = 4  # set to the number of cores you want to use

def worker(arg):
    obj, m, a = arg
    return obj.my_process(m, a)

if __name__ == "__main__":
    list_of_numbers = range(0, 5)
    list_of_objects = [MyClass(i) for i in list_of_numbers]

    pool = mp.Pool(NUM_CORE)
    list_of_results = pool.map(worker, ((obj, 100, 1) for obj in list_of_objects))
    pool.close()
    pool.join()

    print list_of_numbers
    print list_of_results

A big of magic

I should note there are many advantages to taking the very simple approach I suggest. Beyond that it "just works" on Pythons 2 and 3, requires no changes to your classes, and is easy to understand, it also plays nice with all of the Pool methods.

However, if you have multiple methods you want to run in parallel, it can get a bit annoying to write a tiny worker function for each. So here's a tiny bit of "magic" to worm around that. Change worker() like so:

def worker(arg):
    obj, methname = arg[:2]
    return getattr(obj, methname)(*arg[2:])

Now a single worker function suffices for any number of methods, with any number of arguments. In your specific case, just change one line to match:

list_of_results = pool.map(worker, ((obj, "my_process", 100, 1) for obj in list_of_objects))

More-or-less obvious generalizations can also cater to methods with keyword arguments. But, in real life, I usually stick to the original suggestion. At some point catering to generalizations does more harm than good. Then again, I like obvious things ;-)

141

answered Oct 09 '22 20:10

Tim Peters

If your class is not "huge", I think process oriented is better. Pool in multiprocessing is suggested.
This is the tutorial -> https://docs.python.org/2/library/multiprocessing.html#using-a-pool-of-workers

Then seperate the add_to from my_process since they are quick and you can wait util the end of the last process.

def my_process(input, multiby):
    return xxxx
def add_to(result,a_list):
    xxx
p = Pool(5)
res = []
for i in range(10):
    res.append(p.apply_async(my_process, (i,5)))
p.join()  # wait for the end of the last process
for i in range(10):
    print res[i].get()

answered Oct 09 '22 20:10

Zealseeker

Generally the easiest way to run the same calculation in parallel is the map method of a multiprocessing.Pool (or the as_completed function from concurrent.futures in Python 3).

However, the map method applies a function that only takes one argument to an iterable of data using multiple processes.

So this function cannot be a normal method, because that requires at least two arguments; it must also include self! It could be a staticmethod, however. See also this answer for a more in-depth explanation.

answered Oct 09 '22 22:10

Roland Smith

Based on the answer of Python Multiprocessing - apply class method to a list of objects and your code:

add MyClass object into simulation object

class simulation(multiprocessing.Process):
    def __init__(self, id, worker, *args, **kwargs):
        # must call this before anything else
        multiprocessing.Process.__init__(self)
        self.id = id
        self.worker = worker
        self.args = args
        self.kwargs = kwargs
        sys.stdout.write('[%d] created\n' % (self.id))

run what you want in run function

    def run(self):
        sys.stdout.write('[%d] running ...  process id: %s\n' % (self.id, os.getpid()))
        self.worker.my_process(*self.args, **self.kwargs)
        sys.stdout.write('[%d] completed\n' % (self.id))

Try this:

list_of_numbers = range(0, 5)
list_of_objects = [MyClass(i) for i in list_of_numbers]
list_of_sim = [simulation(id=k, worker=obj, multiply_by=100*k, add_to=10*k) \
    for k, obj in enumerate(list_of_objects)]  

for sim in list_of_sim:
    sim.start()

answered Oct 09 '22 22:10

Huu-Danh Pham

Related questions
                            
                                Python Quicksort Runtime Error: Maximum Recursion Depth Exceeded in cmp
                            
                                Caught exception is None
                            
                                Why is my object properly removed from a list when __eq__ isn't being called?
                            
                                ValueError when using pandas.read_json
                            
                                Why does the escape key have a delay in Python curses?
                            
                                pandas: How to work with _iLocIndexer?
                            
                                Multiline statements using Python doctest
                            
                                flake8/pylint fails in Tox testing environment, raises InvocationError
                            
                                Python install failed windows 8.1- Error 0x80240017: Failed to execute MSU package
                            
                                Why won't dynamically adding a `__call__` method to an instance work?
                            
                                Matplotlib: Center text in its bbox
                            
                                Analysing Time Series in Python - pandas formatting error - statsmodels
                            
                                Python and OpenSSL version reference issue on OS X
                            
                                Formatting Lists into columns of a table output (python 3)
                            
                                argparse: How to make mutually exclusive arguments optional?
                            
                                Huge space between title and plot matplotlib
                            
                                How to return a string from pandas.DataFrame.info()
                            
                                filter pandas dataframe for past x days
                            
                                Invalid and/or missing SSL certificate for URL when calling apiclient.discovery.build
                            
                                How to convert pandas dataframe to nested dictionary

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With