What's the difference between ThreadPool vs Pool in the multiprocessing module?

Tags:

Whats the difference between ThreadPool and Pool in multiprocessing module. When I try my code out, this is the main difference I see:

from multiprocessing import Pool import os, time  print("hi outside of main()")  def hello(x):     print("inside hello()")     print("Proccess id: ", os.getpid())     time.sleep(3)     return x*x  if __name__ == "__main__":     p = Pool(5)     pool_output = p.map(hello, range(3))      print(pool_output)

I see the following output:

hi outside of main() hi outside of main() hi outside of main() hi outside of main() hi outside of main() hi outside of main() inside hello() Proccess id:  13268 inside hello() Proccess id:  11104 inside hello() Proccess id:  13064 [0, 1, 4]

With "ThreadPool":

from multiprocessing.pool import ThreadPool import os, time  print("hi outside of main()")  def hello(x):     print("inside hello()")     print("Proccess id: ", os.getpid())     time.sleep(3)     return x*x  if __name__ == "__main__":     p = ThreadPool(5)     pool_output = p.map(hello, range(3))      print(pool_output)

I see the following output:

hi outside of main() inside hello() inside hello() Proccess id:  15204 Proccess id:  15204 inside hello() Proccess id:  15204 [0, 1, 4]

My questions are:

why is the “outside __main__()” run each time in the Pool?
multiprocessing.pool.ThreadPool doesn't spawn new processes? It just creates new threads?
If so whats the difference between using multiprocessing.pool.ThreadPool as opposed to just threading module?

I don't see any official documentation for ThreadPool anywhere, can someone help me out where I can find it?

762

asked Sep 05 '17 01:09

ozn

1 Answers

The multiprocessing.pool.ThreadPool behaves the same as the multiprocessing.Pool with the only difference that uses threads instead of processes to run the workers logic.

The reason you see

hi outside of main()

being printed multiple times with the multiprocessing.Pool is due to the fact that the pool will spawn 5 independent processes. Each process will initialize its own Python interpreter and load the module resulting in the top level print being executed again.

Note that this happens only if the spawn process creation method is used (only method available on Windows). If you use the fork one (Unix), you will see the message printed only once as for the threads.

The multiprocessing.pool.ThreadPool is not documented as its implementation has never been completed. It lacks tests and documentation. You can see its implementation in the source code.

I believe the next natural question is: when to use a thread based pool and when to use a process based one?

The rule of thumb is:

IO bound jobs -> multiprocessing.pool.ThreadPool
CPU bound jobs -> multiprocessing.Pool
Hybrid jobs -> depends on the workload, I usually prefer the multiprocessing.Pool due to the advantage process isolation brings

On Python 3 you might want to take a look at the concurrent.future.Executor pool implementations.

104

answered Sep 23 '22 07:09

noxdafox

Related questions
                            
                                how to know if a variable is a tuple, a string or an integer?
                            
                                Is there a way to have a conditional requirements.txt file for my Python application based on platform?
                            
                                Check if a program exists from a python script [duplicate]
                            
                                How to implement the ReLU function in Numpy
                            
                                ImportError: No module named 'Queue'
                            
                                Best way to check function arguments? [closed]
                            
                                How do I return JSON without using a template in Django?
                            
                                PIP install unable to find ffi.h even though it recognizes libffi
                            
                                Format certain floating dataframe columns into percentage in pandas
                            
                                Mayavi colorbar in TraitsUI creating blank window
                            
                                How to *actually* read CSV data in TensorFlow?
                            
                                Python Setup Disabling Path Length Limit Pros and Cons?
                            
                                Python PDF library [closed]
                            
                                Should I use np.absolute or np.abs?
                            
                                Example of what SQLAlchemy can do, and Django ORM cannot
                            
                                nose vs pytest - what are the (subjective) differences that should make me pick either? [closed]
                            
                                What is the equivalent of php's print_r() in python?
                            
                                Is there a module for balanced binary tree in Python's standard library?
                            
                                ValueError: Length of values does not match length of index | Pandas DataFrame.unique()
                            
                                Python defaultdict and lambda

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

What's the difference between ThreadPool vs Pool in the multiprocessing module?

Tags:

python

python-3.x

multiprocessing

threadpool

python-multiprocessing

ozn

People also ask

1 Answers

noxdafox

Recent Activity

Donate For Us