Very basic question about Hadoop and compressed input files

Tags:

I have started to look into Hadoop. If my understanding is right i could process a very big file and it would get split over different nodes, however if the file is compressed then the file could not be split and wold need to be processed by a single node (effectively destroying the advantage of running a mapreduce ver a cluster of parallel machines).

My question is, assuming the above is correct, is it possible to split a large file manually in fixed-size chunks, or daily chunks, compress them and then pass a list of compressed input files to perform a mapreduce?

797

asked Jan 16 '10 20:01

Luis Sisamon

1 Answers

BZIP2 is splittable in hadoop - it provides very good compression ratio but from CPU time and performances is not providing optimal results, as compression is very CPU consuming.

LZO is splittable in hadoop - leveraging hadoop-lzo you have splittable compressed LZO files. You need to have external .lzo.index files to be able to process in parallel. The library provides all means of generating these indexes in local or distributed manner.

LZ4 is splittable in hadoop - leveraging hadoop-4mc you have splittable compressed 4mc files. You don't need any external indexing, and you can generate archives with provided command line tool or by Java/C code, inside/outside hadoop. 4mc makes available on hadoop LZ4 at any level of speed/compression-ratio: from fast mode reaching 500 MB/s compression speed up to high/ultra modes providing increased compression ratio, almost comparable with GZIP one.

answered Oct 14 '22 20:10

Carlo Medas

Related questions
                            
                                What is the canonical method for an HTTP client to instruct an HTTP server to disable gzip responses?
                            
                                How to count the number of files inside a tar.gz file (without decompressing)?
                            
                                Compressing text before storing it in the database
                            
                                Why I should not compress images in HTTP headers?
                            
                                Compressing a directory of files with PHP
                            
                                ZLIB Decompression - Client Side
                            
                                In simple terms, how is compression commonly implemented?
                            
                                Save file from a byte[] in C# NET 3.5
                            
                                Dynamic compression doesn't seem to be used in IIS 7.5
                            
                                how to compress a PNG image using Java
                            
                                Compressing & Decompressing 7z file in java
                            
                                how to extract files from a 7-zip stream in Java without store it on hard disk?
                            
                                Are all PDF files compressed?
                            
                                How is a .zip compressed archive structured?
                            
                                Uglify-js doesn't mangle variable names
                            
                                Python - Compress Ascii String
                            
                                IOS Video Compression Swift iOS 8 corrupt video file
                            
                                Looking for large text files for testing compression in all sizes
                            
                                How to compress files
                            
                                logrotate compress files after the postrotate script

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

Very basic question about Hadoop and compressed input files

Tags:

compression

hadoop

Luis Sisamon

People also ask

1 Answers

Carlo Medas

Recent Activity

Donate For Us