My work involves a lot of data processing and streaming and processing data from various sources often times a lot of data. I use Python for everything and was wondering what area of Python should I be researching in order to optimize and build batch processing pipelines? I know there are some open source variations like Luigi which Spotify has created but I'm thinking that is a little bit of overkill for me right now. The only thing I know so far is to study up on generators and lazy evaluations but was wondering what other concepts and libraries I can use for efficient batch processing in python. One example scenario would be reading a ton of json formatted files and convert them into csv before populating into a database by using as little memory as possible. (I need to use SQL standard database as opposed to NoSQL). Any advice would be greatly appreciated.
The example you mentioned, of reading lots of files, translating and then populating a database, reminds me of a signal processing application I wrote.
My application (http://github.com/vmlaker/sherlock) processes large chunks of data (images) in parallel, taking advantage of multi-core CPUs. I used two modules to make a clean implementation: MPipe for assembling the multi-stage concurrent pipeline, and numpy-sharedmem for sharing the NumPy arrays between processes.
If you're trying to maximize runtime performance, and have multiple cores available, you may be able to stage a similar workflow for the example you give:
Read file --> Translate --> Update Database
Reading the json files is I/O bound, but multiprocessing may get you speedups in the translation, as well as database updates.
If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!
Donate Us With