Hadoop gzip compressed files

Tags:

I am new to hadoop and trying to process wikipedia dump. It's a 6.7 GB gzip compressed xml file. I read that hadoop supports gzip compressed files but can only be processed by mapper on a single job as only one mapper can decompress it. This seems to put a limitation on the processing. Is there an alternative? like decompressing and splitting the xml file into multiple chunks and recompressing them with gzip.

I read about the hadoop gzip from http://researchcomputing.blogspot.com/2008/04/hadoop-and-compressed-files.html

Thanks for your help.

514

asked Apr 12 '11 04:04

Boolean

2 Answers

A file compressed with the GZIP codec cannot be split because of the way this codec works. A single SPLIT in Hadoop can only be processed by a single mapper; so a single GZIP file can only be processed by a single Mapper.

There are atleast three ways of going around that limitation:

As a preprocessing step: Uncompress the file and recompress using a splittable codec (LZO)
As a preprocessing step: Uncompress the file, split into smaller sets and recompress. (See this)
Use this patch for Hadoop (which I wrote) that allows for a way around this: Splittable Gzip

HTH

answered Oct 09 '22 09:10

Niels Basjes

This is one of the biggest miss understanding in HDFS.

Yes files compressed as a gzip file are not splitable by MapReduce, but that does not mean that GZip as a codec has no value in HDFS and cannot be made splitable.

GZip as a Codec can be used with RCFiles, Sequence Files, Arvo Files, and many more file formats. When the Gzip Codec is used within these splitable formats you get the great compression and pretty good speed from Gzip plus the splitable component.

answered Oct 09 '22 10:10

Ted Malaska

Related questions
                            
                                Which MySQL connector do I use: mysql-connector-java-5.1.46.jar or mysql-connector-java-5.1.46-bin.jar What is the difference?
                            
                                Failed to find Platform SDK with path: platforms;android-P
                            
                                The minSdk version should not be declared in the android manifest file
                            
                                Why does the compiler allow throws when the method will never throw the Exception
                            
                                How to do Gesture Recognition using Accelerometers
                            
                                Removing DOM nodes when traversing a NodeList
                            
                                Java: Parameterized Runnable
                            
                                Why are some java libraries compiled without debugging information
                            
                                JUnit output in Maven reports
                            
                                Obtaining focus on a JPanel
                            
                                Problems installing Java EE SDK on Linux
                            
                                Java for each vs regular for -- are they equivalent?
                            
                                How do we show the gridline in GridLayout?
                            
                                Why do threads spontaneously awake from wait()?
                            
                                iBatis get executed sql
                            
                                Are entities cached in jpa by default?
                            
                                Difference between System.out.printf and String.format
                            
                                Garbage collection on a local variable
                            
                                How to make the whole line change color in Eclipse when I toggle a breakpoint?
                            
                                java.lang.IllegalStateException: incompatible return value type

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

Hadoop gzip compressed files

Tags:

java

algorithm

data-structures

hadoop

mapreduce

Boolean

People also ask

2 Answers

Niels Basjes

Ted Malaska

Recent Activity

Donate For Us