Logo Questions Linux Laravel Mysql Ubuntu Git Menu
 

MapReduce shuffle/sort method

Somewhat of an odd question, but does anyone know what kind of sort MapReduce uses in the sort portion of shuffle/sort? I would think merge or insertion (in keeping with the whole MapReduce paradigm), but I'm not sure.

like image 206
SubSevn Avatar asked Apr 25 '11 15:04

SubSevn


1 Answers

It's Quicksort, afterwards the sorted intermediate outputs get merged together. Quicksort checks the recursion depth and gives up when it is too deep. If this is the case, Heapsort is used.

Have a look at the Quicksort class:

org.apache.hadoop.util.QuickSort

You can change the algorithm used via the map.sort.class value in the hadoop-default.xml.

like image 178
Thomas Jungblut Avatar answered Sep 30 '22 11:09

Thomas Jungblut