The Namenode in the Hadoop architecture is a single point of failure. How do people who have large Hadoop clusters cope with this problem?. Is there an industry-accepted solution that has worked well wherein a secondary Namenode takes over in case the primary one fails ?

Yahoo has certain recommendations for configuration settings at different cluster sizes to take NameNode failure into account. For example: <blockquote> The single point of failure in a Hadoop cluster is the NameNode. While the loss of any other machine (intermittently or permanently) does not result in data loss, NameNode loss results in cluster unavailability. The permanent loss of NameNode data would render the cluster's HDFS inoperable. Therefore, another step should be taken in this configuration to back up the NameNode metadata </blockquote> Facebook uses a tweaked version of Hadoop for its data warehouses; it has some optimizations that focus on NameNode reliability. Additionally to the patches available on github, Facebook appears to use AvatarNode specifically for quickly switching between primary and secondary NameNodes. Dhruba Borthakur's blog contains several other entries offering further insights into the NameNode as a single point of failure. Edit: Further info about Facebook's improvements to the NameNode.

High Availability of Namenode has been introduced with Hadoop 2.x release. It can be achieved in two modes - With NFS and With QJM But high availability with Quorum Journal Manager (QJM) is preferred option. <blockquote> In a typical HA cluster, two separate machines are configured as NameNodes. At any point in time, exactly one of the NameNodes is in an Active state, and the other is in a Standby state. The Active NameNode is responsible for all client operations in the cluster, while the Standby is simply acting as a slave, maintaining enough state to provide a fast failover if necessary. </blockquote> Have a look at below SE questions, which explains complete failover process. Secondary NameNode usage and High availability in Hadoop 2.x How does Hadoop Namenode failover process works?

Hadoop namenode : Single point of failure

2 Answers

Yahoo has certain recommendations for configuration settings at different cluster sizes to take NameNode failure into account. For example:

The single point of failure in a Hadoop cluster is the NameNode. While the loss of any other machine (intermittently or permanently) does not result in data loss, NameNode loss results in cluster unavailability. The permanent loss of NameNode data would render the cluster's HDFS inoperable.

Therefore, another step should be taken in this configuration to back up the NameNode metadata

Facebook uses a tweaked version of Hadoop for its data warehouses; it has some optimizations that focus on NameNode reliability. Additionally to the patches available on github, Facebook appears to use AvatarNode specifically for quickly switching between primary and secondary NameNodes. Dhruba Borthakur's blog contains several other entries offering further insights into the NameNode as a single point of failure.

Edit: Further info about Facebook's improvements to the NameNode.

answered Sep 21 '22 09:09

Bkkbrad

High Availability of Namenode has been introduced with Hadoop 2.x release.

It can be achieved in two modes - With NFS and With QJM

But high availability with Quorum Journal Manager (QJM) is preferred option.

In a typical HA cluster, two separate machines are configured as NameNodes. At any point in time, exactly one of the NameNodes is in an Active state, and the other is in a Standby state. The Active NameNode is responsible for all client operations in the cluster, while the Standby is simply acting as a slave, maintaining enough state to provide a fast failover if necessary.

Have a look at below SE questions, which explains complete failover process.

Secondary NameNode usage and High availability in Hadoop 2.x

How does Hadoop Namenode failover process works?

answered Sep 19 '22 09:09

Ravindra babu

Related questions
                            
                                What is "energy" in image processing?
                            
                                Visual Studio 2010 Periodically Hangs for Several Seconds
                            
                                What is the default instance context mode?
                            
                                ios steps to create custom UITableViewCell with xib file
                            
                                Convert datetime column to datetime2 column in SQL Server?
                            
                                LIKE with integers, in SQL
                            
                                Flush just an app not the whole project
                            
                                play MIDI files in python?
                            
                                JUnit assertEquals( ) fails for two objects
                            
                                Why is using DIVs or spans tags "better" than using a table layout? [duplicate]
                            
                                How to verify multiple method calls with Moq
                            
                                Scala Functional Literals with Implicits

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

Hadoop namenode : Single point of failure

Tags:

rakeshr

People also ask

2 Answers

Bkkbrad

Ravindra babu

Recent Activity

Donate For Us