Working with Spark – Big Data Hadoop MapReduce
Introduction Before moving to Spark RDD concept, which is the base line of Spark, we need to understand the concept of Hadoop Mapredue. RDD is the improvement of Hadoop Mapreduce to get 100 outputs in memory. In this blog post we are going to discuss about Hadoop Mapredue to understand the concept only. Hope it will be interesting. What is MapReduce MapReduce is a software framework and programming module to handle huge data. Hadoop is capable to run MapReduce programs which are written in different language like Java, Ruby, Python and C++. How MapReduce Works As the name specified, MapReduce is a combination of Mapping and Reducing. We can divide the MapReduc into following Sections. · Input Splits · Mapping · Shuffling · Reducing Before exami...