

{"id":49652,"date":"2019-02-20T18:11:22","date_gmt":"2019-02-20T12:41:22","guid":{"rendered":"https:\/\/data-flair.training\/blogs\/?p=49652"},"modified":"2019-02-22T11:17:55","modified_gmt":"2019-02-22T05:47:55","slug":"hadoop-ecosystem","status":"publish","type":"post","link":"https:\/\/data-flair.training\/blogs\/hadoop-ecosystem\/","title":{"rendered":"Hadoop Ecosystem &#8211; 15 Must Know Hadoop Components"},"content":{"rendered":"<p><span style=\"font-weight: 400\">In this tutorial, we will have an overview of various Hadoop Ecosystem Components. These ecosystem components are actually different services deployed by the various enterprise. We can integrate these to work with a variety of data. Each of the Hadoop Ecosystem Components is developed to deliver explicit function. And each has its own developer community and individual release cycle.<\/span><\/p>\n<p>So, let&#8217;s explore Hadoop Ecosystem Components.<\/p>\n<div id=\"attachment_50496\" style=\"width: 1210px\" class=\"wp-caption aligncenter\"><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hadoop-Ecosystem-01.png\"><img loading=\"lazy\" decoding=\"async\" aria-describedby=\"caption-attachment-50496\" class=\"size-full wp-image-50496\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hadoop-Ecosystem-01.png\" alt=\"Hadoop Ecosystem - 15 Must Know Hadoop Components\" width=\"1200\" height=\"628\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hadoop-Ecosystem-01.png 1200w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hadoop-Ecosystem-01-150x79.png 150w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hadoop-Ecosystem-01-300x157.png 300w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hadoop-Ecosystem-01-768x402.png 768w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hadoop-Ecosystem-01-1024x536.png 1024w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hadoop-Ecosystem-01-520x272.png 520w\" sizes=\"auto, (max-width: 1200px) 100vw, 1200px\" \/><\/a><p id=\"caption-attachment-50496\" class=\"wp-caption-text\">Hadoop Ecosystem &#8211; 15 Must Know Hadoop Components<\/p><\/div>\n<h2>Hadoop Ecosystem Components<\/h2>\n<p><span style=\"font-weight: 400\">Hadoop Ecosystem is a suite of services that work to solve the Big Data problem. The different components of the Hadoop Ecosystem are as follows:-<\/span><\/p>\n<div class=\"df-float-l\">\n<h3>1. Hadoop Distributed File System<\/h3>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/hadoop-HDFS.png\"><img loading=\"lazy\" decoding=\"async\" class=\" wp-image-50445 alignleft\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/hadoop-HDFS.png\" alt=\"Hadoop Ecosystem Component\" width=\"281\" height=\"117\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/hadoop-HDFS.png 300w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/hadoop-HDFS-150x63.png 150w\" sizes=\"auto, (max-width: 281px) 100vw, 281px\" \/><\/a><\/p>\n<p><span style=\"font-weight: 400\"><a href=\"https:\/\/data-flair.training\/blogs\/hadoop-hdfs-tutorial\/\"><strong>HDFS is the foundation of Hadoop<\/strong><\/a> and hence is a very important component of the Hadoop ecosystem. It is Java software that provides many features like scalability, high availability, fault tolerance, cost effectiveness etc. It also provides robust distributed data storage for Hadoop. We can deploy many other software frameworks over HDFS.<\/span><\/p>\n<p><strong>Components of HDFS:-<\/strong><\/p>\n<p><span style=\"font-weight: 400\">There are three major components of Hadoop HDFS are as follows:-<\/span><\/p>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/namenode-datanode-01-1.png\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-50497\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/namenode-datanode-01-1.png\" alt=\"Hadoop Ecosystem\" width=\"1200\" height=\"628\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/namenode-datanode-01-1.png 1200w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/namenode-datanode-01-1-150x79.png 150w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/namenode-datanode-01-1-300x157.png 300w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/namenode-datanode-01-1-768x402.png 768w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/namenode-datanode-01-1-1024x536.png 1024w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/namenode-datanode-01-1-520x272.png 520w\" sizes=\"auto, (max-width: 1200px) 100vw, 1200px\" \/><\/a><\/p>\n<p>&nbsp;<\/p>\n<h4><span style=\"font-weight: 400\">a. DataNode <\/span><\/h4>\n<p><span style=\"font-weight: 400\">These are the nodes which store the actual data. HDFS stores the data in a distributed manner. It divides the input files of varied formats into blocks. The DataNodes stores each of these blocks. Following are the functions of DataNodes:-<\/span><\/p>\n<ul>\n<li><span style=\"font-weight: 400\"> On startup, DataNode does handshake with NameNode. It verifies the namespace ID and software version of DataNode.<\/span><\/li>\n<li><span style=\"font-weight: 400\">Also, it sends a block report to NameNode and verifies the block replicas.<\/span><\/li>\n<li><span style=\"font-weight: 400\"> It sends a heartbeat to NameNode every 3 seconds to tell that it is alive. <\/span><\/li>\n<\/ul>\n<h4><span style=\"font-weight: 400\">b. NameNode<\/span><\/h4>\n<p><span style=\"font-weight: 400\">NameNode is nothing but the master node. The NameNode is responsible for managing file system namespace, controlling the client\u2019s access to files. Also, it executes tasks such as opening, closing and naming files and directories. NameNode has two major files \u2013 FSImage and Edits log<\/span><\/p>\n<p><span style=\"font-weight: 400\"><strong>FSImage \u2013<\/strong> FSImage is a point-in-time snapshot of HDFS\u2019s metadata. It contains information like file permission, disk quota, modification timestamp, access time etc.<\/span><\/p>\n<p><span style=\"font-weight: 400\"><strong>Edits log \u2013<\/strong> It contains modifications on FSImage. It records incremental changes like renaming the file, appending data to the file etc.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Whenever the NameNode starts it applies Edits log to FSImage. And the new FSImage gets loaded on the NameNode. <\/span><\/p>\n<h4><span style=\"font-weight: 400\">c. Secondary NameNode <\/span><\/h4>\n<p><span style=\"font-weight: 400\">If the NameNode has not restarted for months the size of Edits log increases. This, in turn, increases the downtime of the cluster on the restart of NameNode. In this case, Secondary NameNode comes into the picture. The Secondary NameNode applies edits log on FSImage at regular intervals. And it updates the new FSImage on primary NameNode. <\/span><\/p>\n<\/div>\n<div class=\"df-float-l\">\n<h3>2. MapReduce<\/h3>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/hadoop-mapreduce.png\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-50457 alignleft\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/hadoop-mapreduce.png\" alt=\"Hadoop Ecosystem\" width=\"276\" height=\"115\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/hadoop-mapreduce.png 300w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/hadoop-mapreduce-150x63.png 150w\" sizes=\"auto, (max-width: 276px) 100vw, 276px\" \/><\/a><\/p>\n<p><span style=\"font-weight: 400\"><a href=\"https:\/\/data-flair.training\/blogs\/hadoop-mapreduce-tutorial\/\"><strong>MapReduce<\/strong><\/a> is the data processing component of Hadoop. It applies the computation on sets of data in parallel thereby improving the performance. MapReduce works in two phases \u2013 <\/span><\/p>\n<p><span style=\"font-weight: 400\"><strong>Map Phase \u2013<\/strong> This phase takes input as key-value pairs and produces output as key-value pairs. It can write custom business logic in this phase. Map phase processes the data and gives it to the next phase.<\/span><\/p>\n<p><span style=\"font-weight: 400\"><strong>Reduce Phase \u2013<\/strong> The MapReduce framework sorts the key-value pair before giving the data to this phase. This phase applies the summary type of calculations to the key-value pairs.<\/span><\/p>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/MapReduce-Anatomy-01-1.png\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-50498\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/MapReduce-Anatomy-01-1.png\" alt=\"Hadoop Ecosystem Components\" width=\"1200\" height=\"628\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/MapReduce-Anatomy-01-1.png 1200w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/MapReduce-Anatomy-01-1-150x79.png 150w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/MapReduce-Anatomy-01-1-300x157.png 300w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/MapReduce-Anatomy-01-1-768x402.png 768w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/MapReduce-Anatomy-01-1-1024x536.png 1024w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/MapReduce-Anatomy-01-1-520x272.png 520w\" sizes=\"auto, (max-width: 1200px) 100vw, 1200px\" \/><\/a><\/p>\n<p>&nbsp;<\/p>\n<ul>\n<li><span style=\"font-weight: 400\"> Mapper reads the block of data and converts it into key-value pairs.<\/span><\/li>\n<li><span style=\"font-weight: 400\"> Now, these key-value pairs are input to the reducer.<\/span><\/li>\n<li><span style=\"font-weight: 400\"> The reducer receives data tuples from multiple mappers.<\/span><\/li>\n<li><span style=\"font-weight: 400\"> Reducer applies aggregation to these tuples based on the key.<\/span><\/li>\n<li><span style=\"font-weight: 400\"> The final output from reducer gets written to HDFS.<\/span><\/li>\n<\/ul>\n<p><span style=\"font-weight: 400\">MapReduce framework takes care of the failure. It recovers data from another node in an event where one node goes down. <\/span><\/p>\n<\/div>\n<div class=\"df-float-l\">\n<h3>3. Yarn<\/h3>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/hadoop-yarn.png\"><img loading=\"lazy\" decoding=\"async\" class=\" wp-image-50459 alignleft\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/hadoop-yarn.png\" alt=\"Hadoop Ecosystem\" width=\"293\" height=\"122\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/hadoop-yarn.png 300w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/hadoop-yarn-150x63.png 150w\" sizes=\"auto, (max-width: 293px) 100vw, 293px\" \/><\/a><\/p>\n<p><span style=\"font-weight: 400\">Yarn which is short for Yet Another Resource Manager. It is like the operating system of Hadoop as it monitors and manages the resources. Yarn came into the picture with the launch of Hadoop 2.x in order to allow different workloads. It handles the workloads like stream processing, interactive processing, batch processing over a single platform. Yarn has two main components \u2013 Node Manager and Resource Manager.<\/span><\/p>\n<h4><span style=\"font-weight: 400\">a. Node Manager <\/span><\/h4>\n<p><span style=\"font-weight: 400\">It is Yarn\u2019s per-node agent and takes care of the individual compute nodes in a <a href=\"https:\/\/data-flair.training\/blogs\/hadoop-cluster\/\"><strong>Hadoop cluster<\/strong><\/a>. It monitors the resource usage like CPU, memory etc. of the local node and intimates the same to Resource Manager. <\/span><\/p>\n<h4><span style=\"font-weight: 400\">b. Resource Manager <\/span><\/h4>\n<p><span style=\"font-weight: 400\">It is responsible for tracking the resources in the cluster and scheduling tasks like map-reduce jobs.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Also, we have the Application Master and Scheduler in Yarn. Let us take a look at them.<\/span><\/p>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/YARN-working-1.png\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-50499\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/YARN-working-1.png\" alt=\"Hadoop Ecosystem\" width=\"1200\" height=\"628\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/YARN-working-1.png 1200w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/YARN-working-1-150x79.png 150w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/YARN-working-1-300x157.png 300w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/YARN-working-1-768x402.png 768w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/YARN-working-1-1024x536.png 1024w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/YARN-working-1-520x272.png 520w\" sizes=\"auto, (max-width: 1200px) 100vw, 1200px\" \/><\/a><\/p>\n<p>&nbsp;<\/p>\n<p><span style=\"font-weight: 400\">Application Master has two functions and they are:-<\/span><\/p>\n<ul>\n<li><span style=\"font-weight: 400\">Negotiating resources from Resource Manager<\/span><\/li>\n<li><span style=\"font-weight: 400\">Working with NodeManager to monitor and execute the sub-task. <\/span><\/li>\n<\/ul>\n<p><span style=\"font-weight: 400\">Following are the functions of Resource Scheduler:-<\/span><\/p>\n<ul>\n<li><span style=\"font-weight: 400\">It allocates resources to various running applications<\/span><\/li>\n<li><span style=\"font-weight: 400\">But it does not monitor the status of the application. So in the event of failure of the task, it does not restart the same. <\/span><\/li>\n<\/ul>\n<p><span style=\"font-weight: 400\">We have another concept called Container. It is nothing but a fraction of NodeManager capacity i.e. CPU, memory, disk, network etc. <\/span><\/p>\n<\/div>\n<div class=\"df-float-l\">\n<h3>4. Hive<\/h3>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Apache-Hive.png\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-50460 alignleft\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Apache-Hive.png\" alt=\"Hadoop Ecosystem Component\" width=\"159\" height=\"159\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Apache-Hive.png 300w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Apache-Hive-150x150.png 150w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Apache-Hive-160x160.png 160w\" sizes=\"auto, (max-width: 159px) 100vw, 159px\" \/><\/a><\/p>\n<p><span style=\"font-weight: 400\"><strong><a href=\"https:\/\/data-flair.training\/blogs\/apache-hive-tutorial\/\">Hive is a data warehouse project<\/a><\/strong> built on the top of Apache Hadoop which provides data query and analysis. It has got the language of its own call HQL or <strong>Hive Query Language<\/strong>. HQL automatically translates the queries into the corresponding map-reduce job.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Main parts of the Hive are \u2013<\/span><\/p>\n<ul>\n<li><span style=\"font-weight: 400\"><strong>MetaStore \u2013<\/strong> it stores metadata<\/span><\/li>\n<li><span style=\"font-weight: 400\"><strong>Driver \u2013<\/strong> Manages the lifecycle of <strong><a href=\"https:\/\/data-flair.training\/blogs\/hiveql-select-statement\/\">HQL statement<\/a><\/strong><\/span><\/li>\n<li><span style=\"font-weight: 400\"><strong>Query compiler \u2013<\/strong> Compiles HQL into DAG i.e. Directed Acyclic Graph<\/span><\/li>\n<li><span style=\"font-weight: 400\"><strong>Hive server \u2013<\/strong> Provides interface for JDBC\/ODBC server.<\/span><\/li>\n<\/ul>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hive-working.png\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-50500\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hive-working.png\" alt=\"Hadoop Ecosystem\" width=\"1200\" height=\"628\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hive-working.png 1200w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hive-working-150x79.png 150w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hive-working-300x157.png 300w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hive-working-768x402.png 768w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hive-working-1024x536.png 1024w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hive-working-520x272.png 520w\" sizes=\"auto, (max-width: 1200px) 100vw, 1200px\" \/><\/a><\/p>\n<p>&nbsp;<\/p>\n<p><span style=\"font-weight: 400\">Facebook designed <strong>Hive<\/strong> for people who are comfortable in SQL. It has two basic components \u2013 Hive Command Line and JDBC, ODBC. Hive Command line is an interface for execution of HQL commands. And JDBC, ODBC establishes the connection with data storage. Hive is highly scalable. It can handle both types of workloads i.e. batch processing and interactive processing. It supports native <a href=\"https:\/\/data-flair.training\/blogs\/sql-data-types\/\"><strong>data type of SQL<\/strong><\/a>. Hive provides many pre-defined functions for analysis. But you can also define your own custom functions called UDFs or user-defined functions. <\/span><\/p>\n<\/div>\n<div class=\"df-float-l\">\n<h3>5. Pig<\/h3>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Pig.png\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-50461 alignleft\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Pig.png\" alt=\"Hadoop Ecosystem Component\" width=\"156\" height=\"177\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Pig.png 258w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Pig-133x150.png 133w\" sizes=\"auto, (max-width: 156px) 100vw, 156px\" \/><\/a><\/p>\n<p><span style=\"font-weight: 400\">Pig is a SQL like language used for querying and analyzing data stored in HDFS. Yahoo was the original creator of the Pig. It uses pig latin language. It loads the data, applies a filter to it and dumps the data in the required format. Pig also consists of <strong><a href=\"https:\/\/data-flair.training\/blogs\/java-virtual-machine-jvm\/\">JVM<\/a><\/strong> called Pig Runtime. Various <a href=\"https:\/\/data-flair.training\/blogs\/apache-pig-features\/\"><strong>features of Pig<\/strong><\/a> are as follows:-<\/span><\/p>\n<ul>\n<li><span style=\"font-weight: 400\"><strong>Extensibility \u2013<\/strong> For carrying out special purpose processing, users can create their own custom function.<\/span><\/li>\n<li><span style=\"font-weight: 400\"><strong>Optimization opportunities \u2013<\/strong> Pig automatically optimizes the query allowing users to focus on semantics rather than efficiency.<\/span><\/li>\n<li><span style=\"font-weight: 400\"><strong>Handles all kinds of data \u2013<\/strong> Pig analyzes both structured as well as unstructured.<\/span><\/li>\n<\/ul>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/How-pig-works-01-1.png\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-50501\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/How-pig-works-01-1.png\" alt=\"Hadoop Ecosystem\" width=\"1200\" height=\"628\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/How-pig-works-01-1.png 1200w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/How-pig-works-01-1-150x79.png 150w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/How-pig-works-01-1-300x157.png 300w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/How-pig-works-01-1-768x402.png 768w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/How-pig-works-01-1-1024x536.png 1024w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/How-pig-works-01-1-520x272.png 520w\" sizes=\"auto, (max-width: 1200px) 100vw, 1200px\" \/><\/a><\/p>\n<p><strong>a. How does Pig work?<\/strong><\/p>\n<ul>\n<li><span style=\"font-weight: 400\"> First, the load command loads the data.<\/span><\/li>\n<li><span style=\"font-weight: 400\">At the backend, the compiler converts pig latin into the sequence of map-reduce jobs.<\/span><\/li>\n<li><span style=\"font-weight: 400\">Over this data, we perform various functions like joining, sorting, grouping, filtering etc.<\/span><\/li>\n<li><span style=\"font-weight: 400\">Now, you can dump the output on the screen or store it in an HDFS file.<\/span><\/li>\n<\/ul>\n<\/div>\n<div class=\"df-float-l\">\n<h3>6. HBase<\/h3>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hadoop-HBase.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignleft\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hadoop-HBase.png\" alt=\"Hadoop Ecosystem Components\" width=\"258\" height=\"111\" \/><\/a><br \/>\n<span style=\"font-weight: 400\">HBase is a NoSQL database built on the top of HDFS. The various <a href=\"https:\/\/data-flair.training\/blogs\/features-of-hbase\/\"><strong>features of <\/strong><\/a><\/span><a href=\"https:\/\/data-flair.training\/blogs\/features-of-hbase\/\"><b>HBase<\/b><\/a><span style=\"font-weight: 400\"> are that it is open-source, non-relational, distributed database. It imitates <strong>Google&#8217;s Bigtable<\/strong> and written in Java. It provides real-time read\/write access to large datasets. Its various components are as follows:-<\/span><\/p>\n<h4><span style=\"font-weight: 400\">a. HBase Master<\/span><\/h4>\n<p><span style=\"font-weight: 400\">HBase performs the following functions:<\/span><\/p>\n<ul>\n<li><span style=\"font-weight: 400\"> Maintain and monitor the <\/span><span style=\"font-weight: 400\">Hadoop cluster.<\/span><\/li>\n<li><span style=\"font-weight: 400\"> Performs administration of the database.<\/span><\/li>\n<li><span style=\"font-weight: 400\"> Controls the failover.<\/span><\/li>\n<li><span style=\"font-weight: 400\"> HMaster handles DDL operation.<\/span><\/li>\n<\/ul>\n<h4><span style=\"font-weight: 400\">b. RegionServer<\/span><\/h4>\n<p><span style=\"font-weight: 400\">Region server is a process which handles read, writes, update and delete requests from clients. It runs on every node in a Hadoop cluster that is HDFS DataNode.<\/span><\/p>\n<p><span style=\"font-weight: 400\">HBase is a column-oriented database management system. It runs on top of HDFS. It suits for sparse data sets which are common in<a href=\"https:\/\/data-flair.training\/blogs\/big-data-use-cases-case-studies-hadoop-spark-flink\/\"><strong> Big Data use cases<\/strong><\/a>. HBase support writing application in Apache Avro, REST and Thrift. Apache HBase has low latency storage. Enterprises use this for real-time analysis. <\/span><\/p>\n<p><span style=\"font-weight: 400\">The design of HBase is such that to contain many tables. Each of these tables must have a primary key. Access attempts to HBase tables use this primary key.<\/span><\/p>\n<p><span style=\"font-weight: 400\"><strong>As an example<\/strong> lets us consider HBase table storing diagnostic log from the server. In this case, the typical log row will contain columns such as timestamp when the log gets written. And server from which the log originated. <\/span><\/p>\n<\/div>\n<div class=\"df-float-l\">\n<h3>7. Mahout<\/h3>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Mahout.png\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-50463 alignleft\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Mahout.png\" alt=\"Hadoop Ecosystem Components\" width=\"211\" height=\"211\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Mahout.png 300w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Mahout-150x150.png 150w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Mahout-160x160.png 160w\" sizes=\"auto, (max-width: 211px) 100vw, 211px\" \/><\/a><\/p>\n<p><span style=\"font-weight: 400\">Mahout provides a platform for creating machine learning applications which are scalable.<\/span><\/p>\n<p><strong>a. What is Machine Learning?<\/strong><\/p>\n<p><span style=\"font-weight: 400\"><a href=\"https:\/\/data-flair.training\/blogs\/machine-learning-algorithms\/\"><strong>Machine learning algorithms<\/strong><\/a> allow us to create self-evolving machines without being explicitly programmed. It makes future decisions based on user behavior, past experiences and data patterns. <\/span><\/p>\n<p><strong>b. What Mahout does?<\/strong><\/p>\n<p><span style=\"font-weight: 400\">It performs collaborative filtering, clustering, and classification.\u00a0<\/span><\/p>\n<ul>\n<li><span style=\"font-weight: 400\"><strong>Collaborative filtering \u2013<\/strong> Mahout mines user behavior patterns and based on these it makes recommendations to users.<\/span><\/li>\n<li><span style=\"font-weight: 400\"><strong>Clustering \u2013<\/strong> It groups together a similar type of data like the article, blogs, research paper, news etc.<\/span><\/li>\n<li><span style=\"font-weight: 400\"><strong>Classification \u2013<\/strong> It means categorizing data into various sub-departments. For example, we can classify article into blogs, essays, research papers and so on.<\/span><\/li>\n<li><span style=\"font-weight: 400\"><strong>Frequent Itemset missing \u2013<\/strong> It looks for the items generally bought together and based on that it gives a suggestion. For instance, usually, we buy a cell phone and its cover together. So, when you buy a cell phone it will give suggestion to buy cover also. <\/span><\/li>\n<\/ul>\n<\/div>\n<div class=\"df-float-l\">\n<h3>8. Zookeeper<\/h3>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/zookeeper.png\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-50464 alignleft\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/zookeeper.png\" alt=\"Hadoop Ecosystem Component\" width=\"157\" height=\"212\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/zookeeper.png 222w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/zookeeper-111x150.png 111w\" sizes=\"auto, (max-width: 157px) 100vw, 157px\" \/><\/a><\/p>\n<p><span style=\"font-weight: 400\"><strong><a href=\"https:\/\/data-flair.training\/blogs\/zookeeper-tutorial\/\">Zookeeper<\/a><\/strong> coordinates between various services in the Hadoop ecosystem. It saves the time required for synchronization, configuration maintenance, grouping, and naming. Following are the <a href=\"https:\/\/data-flair.training\/blogs\/zookeeper-features\/\"><strong>features of Zookeeper<\/strong><\/a>:-<\/span><\/p>\n<ul>\n<li><span style=\"font-weight: 400\"><strong>Speed &#8211;<\/strong>\u00a0Zookeeper is fast in workloads where reads to data are more than write. A typical read: write ratio is 10:1.<\/span><\/li>\n<li><span style=\"font-weight: 400\"><strong>Organized &#8211;<\/strong> Zookeeper maintains a record of all transactions.<\/span><\/li>\n<li><span style=\"font-weight: 400\"><strong>Simple &#8211;<\/strong> It maintains a single hierarchical namespace, similar to directories and files. <\/span><\/li>\n<li><span style=\"font-weight: 400\"><strong>Reliable<\/strong> &#8211; We can replicate Zookeeper over a set of hosts and they are aware of each other. There is no single point of failure. As long as major servers are available zookeeper is available.<\/span><\/li>\n<\/ul>\n<p><strong>Why do we need Zookeeper in Hadoop?<\/strong><\/p>\n<p><span style=\"font-weight: 400\">Hadoop faces many problems as it runs a distributed application. One of the problems is deadlock. Deadlock occurs when two or more tasks fight for the same resource. For instance, task T1 has resource R1 and is waiting for resource R2 held by task T2. And this task T2 is waiting for resource R1 held by task T1. In such a scenario deadlock occurs. Both task T1 and T2 would get locked waiting for resources. Zookeeper solves Deadlock condition via synchronization. <\/span><\/p>\n<p><span style=\"font-weight: 400\">Another problem is of race condition. This occurs when the machine tries to perform two or more operations at a time. Zookeeper solves this problem by property of serialization. <\/span><\/p>\n<\/div>\n<div class=\"df-float-l\">\n<h3>9. Oozie<\/h3>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Oozie.png\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-50465 alignleft\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Oozie.png\" alt=\"Hadoop Ecosystem Components\" width=\"234\" height=\"50\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Oozie.png 300w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Oozie-150x32.png 150w\" sizes=\"auto, (max-width: 234px) 100vw, 234px\" \/><\/a><\/p>\n<p><span style=\"font-weight: 400\">It is a workflow scheduler systems for managing Hadoop jobs. It supports <strong><a href=\"https:\/\/data-flair.training\/blogs\/hadoop-career\/\">Hadoop jobs<\/a><\/strong> for Map-Reduce, Pig, Hive, and Sqoop. Oozie combines multiple jobs into a single unit of work. It is scalable and can manage thousands of workflow in a Hadoop cluster. Oozie works by creating DAG i.e. Directed Acyclic Graph of the workflow. It is very much flexible as it can start, stop, suspend and rerun failed jobs.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Oozie is an open-source web-application written in Java. Oozie is scalable and can execute thousands of workflow containing dozens of Hadoop jobs. <\/span><\/p>\n<p><span style=\"font-weight: 400\">There are three basic types of Oozie jobs and they are as follows:-<\/span><\/p>\n<ul>\n<li><span style=\"font-weight: 400\"><strong>Workflow \u2013<\/strong> It stores and runs a workflow composed of Hadoop jobs. It stores the job as Directed Acyclic Graph to determine the sequence of actions that will get executed.<\/span><\/li>\n<li><span style=\"font-weight: 400\"><strong>Coordinator &#8211;<\/strong> It runs workflow jobs based on predefined schedules and availability of data. <\/span><\/li>\n<li><span style=\"font-weight: 400\"><strong>Bundle \u2013<\/strong> This is nothing but a package of many coordinators and workflow jobs.<\/span><\/li>\n<\/ul>\n<p><strong>How does Oozie work?<\/strong><\/p>\n<p><span style=\"font-weight: 400\">Oozie runs a service in the Hadoop cluster. Client submits workflow to run, immediately or later.<\/span><\/p>\n<p><span style=\"font-weight: 400\">There are two types of nodes in Oozie. They are action node and control flow node.<\/span><\/p>\n<ul>\n<li><span style=\"font-weight: 400\"><strong>Action Node \u2013<\/strong> It represents the task in the workflow like MapReduce job, shell script, pig or hive jobs etc.<\/span><\/li>\n<li><span style=\"font-weight: 400\"><strong>Control flow Node \u2013<\/strong> It controls the workflow between actions by employing conditional logic. In this, the previous action decides which branch to follow.<\/span><\/li>\n<\/ul>\n<p><span style=\"font-weight: 400\">Start, End and Error Nodes fall under this category.<\/span><\/p>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/MapReduce-Anatomy-01-2.png\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-50504\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/MapReduce-Anatomy-01-2.png\" alt=\"Hadoop Ecosystem\" width=\"1200\" height=\"628\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/MapReduce-Anatomy-01-2.png 1200w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/MapReduce-Anatomy-01-2-150x79.png 150w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/MapReduce-Anatomy-01-2-300x157.png 300w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/MapReduce-Anatomy-01-2-768x402.png 768w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/MapReduce-Anatomy-01-2-1024x536.png 1024w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/MapReduce-Anatomy-01-2-520x272.png 520w\" sizes=\"auto, (max-width: 1200px) 100vw, 1200px\" \/><\/a><\/p>\n<p>&nbsp;<\/p>\n<ul>\n<li><span style=\"font-weight: 400\">Start Node signals the start of the workflow job.<\/span><\/li>\n<li><span style=\"font-weight: 400\">End Node designates the end of job.<\/span><\/li>\n<li><span style=\"font-weight: 400\">ErrorNode signals the error and gives an error message.<\/span><\/li>\n<\/ul>\n<\/div>\n<div class=\"df-float-l\">\n<h3>10. Sqoop<\/h3>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/sqoop-01.png\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-50466 alignleft\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/sqoop-01.png\" alt=\"Hadoop Ecosystem Components\" width=\"185\" height=\"51\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/sqoop-01.png 403w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/sqoop-01-150x41.png 150w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/sqoop-01-300x83.png 300w\" sizes=\"auto, (max-width: 185px) 100vw, 185px\" \/><\/a><\/p>\n<p><span style=\"font-weight: 400\"><strong><a href=\"https:\/\/data-flair.training\/blogs\/sqoop-import\/\">Sqoop imports data<\/a><\/strong> from external sources into compatible Hadoop Ecosystem components like HDFS, Hive, HBase etc. It also transfers data from Hadoop to other external sources. It works with RDBMS like TeraData, Oracle, MySQL and so on. The major difference between Sqoop and Flume is that Flume does not work with structured data. But Sqoop can deal with structured as well as unstructured data. <\/span><\/p>\n<p><strong>Let us see how Sqoop works<\/strong><\/p>\n<p><span style=\"font-weight: 400\">When we submit Sqoop command, at the back-end, it gets divided into a number of sub-tasks. These sub-tasks are nothing but map-tasks. Each map-task import a part of data to Hadoop. Hence all the map-task taken together imports the whole data.<\/span><span style=\"font-weight: 400\"><br \/>\n<\/span><\/p>\n<p><span style=\"font-weight: 400\"><strong><a href=\"https:\/\/data-flair.training\/blogs\/sqoop-export\/\">Sqoop export<\/a><\/strong> also works in a similar way. Only thing is instead of importing, the map-task export the part of data from Hadoop to destination database.<\/span><span style=\"font-weight: 400\"><br \/>\n<\/span><\/p>\n<\/div>\n<div class=\"df-float-l\">\n<h3>11. Flume<\/h3>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Apache-Flume.png\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-50467 alignleft\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Apache-Flume.png\" alt=\"Hadoop Ecosystem Components\" width=\"156\" height=\"156\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Apache-Flume.png 300w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Apache-Flume-150x150.png 150w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Apache-Flume-160x160.png 160w\" sizes=\"auto, (max-width: 156px) 100vw, 156px\" \/><\/a><\/p>\n<p><span style=\"font-weight: 400\">It is a service which helps to ingest structured and semi-structured data into HDFS. <strong><a href=\"https:\/\/data-flair.training\/blogs\/apache-flume-tutorial\/\">Flume<\/a><\/strong> works on the principle of distributed processing. It aids in collection, aggregation, and movement of a huge amount of data sets. Flume has three components source, sink, and channel.<\/span><\/p>\n<p><span style=\"font-weight: 400\"><strong>Source &#8211;<\/strong>\u00a0It accepts the data from the incoming stream and stores the data in the channel<\/span><\/p>\n<p><span style=\"font-weight: 400\"><strong><a href=\"https:\/\/data-flair.training\/blogs\/apache-flume-channel\/\">Channel<\/a> &#8211;<\/strong> It is a medium of temporary storage between the source of the data and persistent storage of HDFS.<\/span><\/p>\n<p><span style=\"font-weight: 400\"><strong>Sink &#8211;\u00a0<\/strong>This component collects the data from the channel and writes it permanently to the HDFS.<\/span><\/p>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Flume-Import-Agent-01-1.png\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-50505\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Flume-Import-Agent-01-1.png\" alt=\"Hadoop Ecosystem Components\" width=\"1200\" height=\"628\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Flume-Import-Agent-01-1.png 1200w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Flume-Import-Agent-01-1-150x79.png 150w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Flume-Import-Agent-01-1-300x157.png 300w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Flume-Import-Agent-01-1-768x402.png 768w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Flume-Import-Agent-01-1-1024x536.png 1024w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Flume-Import-Agent-01-1-520x272.png 520w\" sizes=\"auto, (max-width: 1200px) 100vw, 1200px\" \/><\/a><\/p>\n<\/div>\n<div class=\"df-float-l\">\n<h3>12. Ambari<\/h3>\n<p><span style=\"font-weight: 400\"><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/ambari.png\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-50468 alignleft\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/ambari.png\" alt=\"Hadoop Ecosystem Componnents\" width=\"160\" height=\"160\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/ambari.png 300w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/ambari-150x150.png 150w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/ambari-160x160.png 160w\" sizes=\"auto, (max-width: 160px) 100vw, 160px\" \/><\/a>Ambari is another Hadoop ecosystem component. It is responsible for provisioning, managing, monitoring and securing Hadoop cluster. Following are the different<a href=\"https:\/\/data-flair.training\/blogs\/ambari-features\/\"><strong> features of Ambari<\/strong><\/a>:<\/span><\/p>\n<ul>\n<li><span style=\"font-weight: 400\">Simplified cluster configuration, management, and installation<\/span><\/li>\n<li><span style=\"font-weight: 400\">Ambari reduces the complexity of configuring and administration of Hadoop cluster security.<\/span><span style=\"font-weight: 400\"><br \/>\n<\/span><\/li>\n<li><span style=\"font-weight: 400\">It ensures that the cluster is healthy and available for monitoring.<\/span><\/li>\n<\/ul>\n<p><em>Ambari gives:-<\/em><br \/>\n<strong>Hadoop cluster provisioning<\/strong><\/p>\n<ul>\n<li><span style=\"font-weight: 400\"> It gives step by step procedure for <a href=\"https:\/\/data-flair.training\/blogs\/installation-of-hadoop-3-x-on-ubuntu\/\"><strong>installing Hadoop services<\/strong><\/a> on the Hadoop cluster.<\/span><\/li>\n<li><span style=\"font-weight: 400\"> It also handles configuration of services across the Hadoop cluster.<\/span><\/li>\n<\/ul>\n<p><strong>Hadoop cluster\u00a0management<\/strong><\/p>\n<ul>\n<li><span style=\"font-weight: 400\">It provides centralized service for starting, stopping and reconfiguring services on the network of machines.<\/span><\/li>\n<\/ul>\n<p><strong>Hadoop cluster monitoring<\/strong><\/p>\n<ul>\n<li><span style=\"font-weight: 400\"> To monitor health and status Ambari provides us dashboard.<\/span><\/li>\n<li><span style=\"font-weight: 400\"> Ambari alert framework alerts the user when the node goes down or has low disk space etc.<\/span><\/li>\n<\/ul>\n<\/div>\n<div class=\"df-float-l\">\n<h3>13. Apache Drill<\/h3>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Drill.png\"><img loading=\"lazy\" decoding=\"async\" class=\" wp-image-50475 alignleft\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Drill.png\" alt=\"Hadoop Ecosystem Components\" width=\"242\" height=\"94\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Drill.png 610w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Drill-150x58.png 150w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Drill-300x117.png 300w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Drill-520x202.png 520w\" sizes=\"auto, (max-width: 242px) 100vw, 242px\" \/><\/a><\/p>\n<p>&nbsp;<\/p>\n<p><span style=\"font-weight: 400\">Apache Drill is a schema-free SQL query engine. It works on the top of Hadoop, NoSQL and cloud storage. Its main purpose is large scale processing of data with low latency. It is a distributed query processing engine. We can query petabytes of data using Drill. It can scale to several thousands of nodes. It supports NoSQL databases like Azure BLOB storage, Google cloud storage,<strong> Amazon<\/strong> S3, HBase, <a href=\"https:\/\/data-flair.training\/blogs\/mongodb-tutorial\/\"><strong>MongoDB<\/strong><\/a> and so on.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Let us look at some of the features of Drill:-<\/span><\/p>\n<ul>\n<li><span style=\"font-weight: 400\">Variety of data sources can be the basis of a single query.<\/span><\/li>\n<li><span style=\"font-weight: 400\">Drill follows ANSI SQL.<\/span><\/li>\n<li><span style=\"font-weight: 400\">It can support millions of users and serve their queries over large data sets.<\/span><\/li>\n<li><span style=\"font-weight: 400\">Drill gives faster insights without ETL overheads like loading, schema creation, maintenance, transformation etc.<\/span><span style=\"font-weight: 400\"><br \/>\n<\/span><\/li>\n<li><span style=\"font-weight: 400\">It can analyze multi-structured and nested data without having to do transformations or filtering.<\/span><\/li>\n<\/ul>\n<\/div>\n<div class=\"df-float-l\">\n<h3>14. Apache Spark<\/h3>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Spark.png\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-50471 alignleft\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Spark.png\" alt=\"Hadoop Ecosystem Components\" width=\"144\" height=\"60\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Spark.png 300w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Spark-150x63.png 150w\" sizes=\"auto, (max-width: 144px) 100vw, 144px\" \/><\/a><\/p>\n<p><span style=\"font-weight: 400\"><a href=\"https:\/\/data-flair.training\/blogs\/spark-tutorial\/\"><strong>Apache Spark<\/strong><\/a>\u00a0unifies all kinds of Big Data processing under one umbrella. It has<strong> built-in libraries<\/strong> for streaming, SQL, machine learning and graph processing. Apache Spark is lightening fast. It gives good performance for both batch and stream processing. It does this with the help of <strong>DAG scheduler<\/strong>, query optimizer, and physical execution engine.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Spark offers 80 high-level operators which makes it easy to build parallel applications. Spark has various libraries like <strong>MLlib for machine learning<\/strong>, GraphX for graph processing, SQL and Data frames, and Spark Streaming. One can run Spark in standalone cluster mode on Hadoop, Mesos, or on Kubernetes. One can write Spark applications using SQL, R, Python, Scala, and Java. As such Scala in the native language of Spark. It was originally developed at the University of California, Berkley. Spark does in-memory calculations. This makes <a href=\"https:\/\/data-flair.training\/blogs\/spark-vs-hadoop-mapreduce\/\"><strong>Spark faster than<\/strong> <strong>Hadoop map-reduce<\/strong><\/a>. <\/span><\/p>\n<\/div>\n<div class=\"df-float-l\">\n<h3>15. Solr And Lucene<\/h3>\n<p><a href=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Solr.png\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-50472 alignleft\" src=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Solr.png\" alt=\"Hadoop Ecosystem Components\" width=\"166\" height=\"101\" srcset=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Solr.png 205w, https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Solr-150x91.png 150w\" sizes=\"auto, (max-width: 166px) 100vw, 166px\" \/><\/a><\/p>\n<p><span style=\"font-weight: 400\">Apache Solr and Apache Lucene are two services which search and indexes the Hadoop ecosystem. Apache Solr is an application built around Apache Lucene. Code of Apache Lucene is in Java. It uses <strong>Java libraries<\/strong> for searching and indexing. Apache Solr is an open source, blazing fast search platform. <\/span><\/p>\n<p><span style=\"font-weight: 400\">Various features of Solr are as follows &#8211;<\/span><span style=\"font-weight: 400\"><br \/>\n<\/span><\/p>\n<ul>\n<li><span style=\"font-weight: 400\">Solr is highly scalable, reliable and fault tolerant. <\/span><\/li>\n<li><span style=\"font-weight: 400\">It provides distributed indexing, automated failover and recovery, load-balanced query, centralized configuration and much more. <\/span><\/li>\n<li><span style=\"font-weight: 400\">You can query Solr using HTTP GET and receive the result in JSON, binary, CSV and XML. <\/span><\/li>\n<li><span style=\"font-weight: 400\">Solr provides matching capabilities like phrases, wildcards, grouping, joining and much more. <\/span><\/li>\n<li><span style=\"font-weight: 400\">It gets shipped with a built-in administrative interface enabling management of solr instances. <\/span><\/li>\n<li><span style=\"font-weight: 400\">Solr takes advantage of Lucene\u2019s near real-time indexing. It enables you to see your content when you want to see it.<\/span><\/li>\n<\/ul>\n<\/div>\n<p>So, this was all in the Hadoop Ecosystem. Hope you liked this article.<\/p>\n<h2><span style=\"font-weight: 400\">Summary<\/span><\/h2>\n<p><span style=\"font-weight: 400\">The <strong><a href=\"https:\/\/hadoop.apache.org\/\">Hadoop<\/a><\/strong> ecosystem elements described above are all open system Apache Hadoop Project. Many commercial applications use these ecosystem elements. Let us summarize Hadoop ecosystem components. At the core, we have HDFS for data storage, map-reduce for data processing and Yarn a resource manager. Then we have HIVE a <a href=\"https:\/\/data-flair.training\/blogs\/best-big-data-analytics-tools\/\"><strong>data analysis tool<\/strong><\/a>, Pig \u2013 SQL like a scripting language, HBase \u2013 NoSQL database, Mahout \u2013 machine learning tool, Zookeeper \u2013 a synchronization tool, Oozie &#8211; workflow scheduler system, Sqoop \u2013 structured data importing and exporting utility, Flume \u2013 data transfer tool for unstructured and semi-structured data, Ambari \u2013 a tool for managing and securing Hadoop clusters, and lastly Avro \u2013 RPC, and data serialization framework.<\/span><\/p>\n<p>Since you are familiar with the Hadoop ecosystem and components, you are ready to learn more in Hadoop. Check out the <a href=\"https:\/\/data-flair.training\/big-data-hadoop\/\"><strong>Hadoop training<\/strong><\/a> by DataFlair.<\/p>\n<p>Still, if any doubt regarding Hadoop Ecosystem, ask in the comment section.<span hidden class=\"__iawmlf-post-loop-links\" data-iawmlf-links=\"[{&quot;id&quot;:1163,&quot;href&quot;:&quot;https:\\\/\\\/hadoop.apache.org&quot;,&quot;archived_href&quot;:&quot;http:\\\/\\\/web-wp.archive.org\\\/web\\\/20251008061344\\\/https:\\\/\\\/hadoop.apache.org\\\/&quot;,&quot;redirect_href&quot;:&quot;&quot;,&quot;checks&quot;:[{&quot;date&quot;:&quot;2025-12-09 02:28:24&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2025-12-12 06:49:41&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2025-12-15 09:10:10&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2025-12-18 18:19:24&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2025-12-22 07:02:07&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2025-12-25 14:18:36&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2025-12-28 14:42:46&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2025-12-31 20:25:23&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-01-04 05:35:17&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-01-07 05:38:48&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-01-10 09:31:55&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-01-13 10:17:34&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-01-16 11:17:06&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-01-19 11:27:31&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-01-22 12:37:34&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-01-25 15:42:16&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-01-28 16:04:51&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-01-31 22:35:26&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-02-04 01:35:24&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-02-07 11:50:52&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-02-10 15:00:48&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-02-13 17:30:11&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-02-17 04:31:34&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-02-20 07:27:26&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-02-23 08:58:22&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-02-26 11:58:09&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-03-01 17:13:40&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-03-04 20:02:02&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-03-08 07:00:15&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-03-11 07:24:33&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-03-14 17:13:12&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-03-18 02:37:36&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-03-21 07:22:14&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-03-24 10:20:44&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-03-27 11:15:57&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-03-30 13:36:00&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-04-03 01:50:21&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-04-06 03:38:55&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-04-09 05:27:12&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-04-12 13:40:56&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-04-16 02:05:03&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-04-19 07:29:50&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-04-22 08:32:52&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-04-25 11:03:04&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-04-28 14:02:46&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-05-01 17:27:05&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-05-05 03:36:42&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-05-08 06:22:34&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-05-11 10:16:03&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-05-14 17:21:35&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-05-17 18:15:11&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-05-20 19:19:28&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-05-24 05:01:01&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-05-27 05:22:10&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-05-30 10:25:50&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-06-02 16:48:23&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-06-06 03:05:45&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-06-09 10:29:32&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-06-12 12:41:09&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-06-16 06:31:56&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-06-19 07:49:24&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-06-22 10:11:50&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-06-25 11:23:39&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-06-28 11:52:40&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-07-01 22:40:41&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-07-05 03:45:18&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-07-08 07:15:52&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-07-11 07:55:35&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-07-14 20:33:55&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-07-18 05:00:08&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-07-21 14:05:29&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-07-24 14:44:08&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-07-27 23:42:16&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-08-02 05:29:28&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-08-05 06:32:07&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-08-08 17:25:19&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-08-12 04:14:35&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-08-15 07:29:17&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-08-18 11:20:27&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-08-21 19:35:52&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-08-24 23:49:42&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-08-28 08:09:15&quot;,&quot;http_code&quot;:206},{&quot;date&quot;:&quot;2026-08-31 09:48:57&quot;,&quot;http_code&quot;:206}],&quot;broken&quot;:false,&quot;last_checked&quot;:{&quot;date&quot;:&quot;2026-08-31 09:48:57&quot;,&quot;http_code&quot;:206},&quot;process&quot;:&quot;done&quot;}]\"><\/span><\/p>\n","protected":false},"excerpt":{"rendered":"<p>In this tutorial, we will have an overview of various Hadoop Ecosystem Components. These ecosystem components are actually different services deployed by the various enterprise. We can integrate these to work with a variety&#46;&#46;&#46;<\/p>\n","protected":false},"author":6,"featured_media":50496,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[22],"tags":[18969,18971,762,855,1280,4799,5232,5244,5381,5548,5675,8537,18968,9493,18970,13596,18967,16320,16357],"class_list":["post-49652","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-hadoop","tag-ambari","tag-apache-drill","tag-apache-flume","tag-apache-mahout","tag-avro","tag-flume","tag-hadoop-components","tag-hadoop-ecosystem","tag-hbase","tag-hdfs","tag-hive","tag-mapreduce","tag-oozie","tag-pig","tag-solr","tag-sqoop","tag-what-is-hadoop-ecosystem","tag-yarn","tag-zookeeper"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.3 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Hadoop Ecosystem - 15 Must Know Hadoop Components - DataFlair<\/title>\n<meta name=\"description\" content=\"Hadoop Ecosystem consists of many componets out of which here are 15 trending components - HDFS, Yarn, MapReduce, HBase, Hive, Ambari, Avro, Oozie, Spark\" \/>\n<meta name=\"robots\" content=\"noindex, nofollow\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Hadoop Ecosystem - 15 Must Know Hadoop Components - DataFlair\" \/>\n<meta property=\"og:description\" content=\"Hadoop Ecosystem consists of many componets out of which here are 15 trending components - HDFS, Yarn, MapReduce, HBase, Hive, Ambari, Avro, Oozie, Spark\" \/>\n<meta property=\"og:url\" content=\"https:\/\/data-flair.training\/blogs\/hadoop-ecosystem\/\" \/>\n<meta property=\"og:site_name\" content=\"DataFlair\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/DataFlairWS\/\" \/>\n<meta property=\"article:published_time\" content=\"2019-02-20T12:41:22+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2019-02-22T05:47:55+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hadoop-Ecosystem-01.png\" \/>\n\t<meta property=\"og:image:width\" content=\"1200\" \/>\n\t<meta property=\"og:image:height\" content=\"628\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"DataFlair Team\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:creator\" content=\"@DataFlairWS\" \/>\n<meta name=\"twitter:site\" content=\"@DataFlairWS\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"DataFlair Team\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"15 minutes\" \/>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Hadoop Ecosystem - 15 Must Know Hadoop Components - DataFlair","description":"Hadoop Ecosystem consists of many componets out of which here are 15 trending components - HDFS, Yarn, MapReduce, HBase, Hive, Ambari, Avro, Oozie, Spark","robots":{"index":"noindex","follow":"nofollow"},"og_locale":"en_US","og_type":"article","og_title":"Hadoop Ecosystem - 15 Must Know Hadoop Components - DataFlair","og_description":"Hadoop Ecosystem consists of many componets out of which here are 15 trending components - HDFS, Yarn, MapReduce, HBase, Hive, Ambari, Avro, Oozie, Spark","og_url":"https:\/\/data-flair.training\/blogs\/hadoop-ecosystem\/","og_site_name":"DataFlair","article_publisher":"https:\/\/www.facebook.com\/DataFlairWS\/","article_published_time":"2019-02-20T12:41:22+00:00","article_modified_time":"2019-02-22T05:47:55+00:00","og_image":[{"width":1200,"height":628,"url":"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hadoop-Ecosystem-01.png","type":"image\/png"}],"author":"DataFlair Team","twitter_card":"summary_large_image","twitter_creator":"@DataFlairWS","twitter_site":"@DataFlairWS","twitter_misc":{"Written by":"DataFlair Team","Est. reading time":"15 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/data-flair.training\/blogs\/hadoop-ecosystem\/#article","isPartOf":{"@id":"https:\/\/data-flair.training\/blogs\/hadoop-ecosystem\/"},"author":{"name":"DataFlair Team","@id":"https:\/\/data-flair.training\/blogs\/#\/schema\/person\/2c58ecb4f73a39f0ef993f1ddfcd7b89"},"headline":"Hadoop Ecosystem &#8211; 15 Must Know Hadoop Components","datePublished":"2019-02-20T12:41:22+00:00","dateModified":"2019-02-22T05:47:55+00:00","mainEntityOfPage":{"@id":"https:\/\/data-flair.training\/blogs\/hadoop-ecosystem\/"},"wordCount":3007,"commentCount":2,"publisher":{"@id":"https:\/\/data-flair.training\/blogs\/#organization"},"image":{"@id":"https:\/\/data-flair.training\/blogs\/hadoop-ecosystem\/#primaryimage"},"thumbnailUrl":"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hadoop-Ecosystem-01.png","keywords":["Ambari","Apache Drill","Apache Flume","Apache Mahout","AVRO","Flume","Hadoop Components","Hadoop Ecosystem","hbase","hdfs","hive","MapReduce","Oozie","pig","Solr","Sqoop","What is Hadoop Ecosystem","yarn","Zookeeper"],"articleSection":["Hadoop Tutorials"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/data-flair.training\/blogs\/hadoop-ecosystem\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/data-flair.training\/blogs\/hadoop-ecosystem\/","url":"https:\/\/data-flair.training\/blogs\/hadoop-ecosystem\/","name":"Hadoop Ecosystem - 15 Must Know Hadoop Components - DataFlair","isPartOf":{"@id":"https:\/\/data-flair.training\/blogs\/#website"},"primaryImageOfPage":{"@id":"https:\/\/data-flair.training\/blogs\/hadoop-ecosystem\/#primaryimage"},"image":{"@id":"https:\/\/data-flair.training\/blogs\/hadoop-ecosystem\/#primaryimage"},"thumbnailUrl":"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hadoop-Ecosystem-01.png","datePublished":"2019-02-20T12:41:22+00:00","dateModified":"2019-02-22T05:47:55+00:00","description":"Hadoop Ecosystem consists of many componets out of which here are 15 trending components - HDFS, Yarn, MapReduce, HBase, Hive, Ambari, Avro, Oozie, Spark","breadcrumb":{"@id":"https:\/\/data-flair.training\/blogs\/hadoop-ecosystem\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/data-flair.training\/blogs\/hadoop-ecosystem\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/data-flair.training\/blogs\/hadoop-ecosystem\/#primaryimage","url":"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hadoop-Ecosystem-01.png","contentUrl":"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2019\/02\/Hadoop-Ecosystem-01.png","width":1200,"height":628,"caption":"Hadoop Ecosystem - 15 Must Know Hadoop Components"},{"@type":"BreadcrumbList","@id":"https:\/\/data-flair.training\/blogs\/hadoop-ecosystem\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Blog Home","item":"https:\/\/data-flair.training\/blogs\/"},{"@type":"ListItem","position":2,"name":"Hadoop Tutorials","item":"https:\/\/data-flair.training\/blogs\/category\/hadoop\/"},{"@type":"ListItem","position":3,"name":"Hadoop Ecosystem &#8211; 15 Must Know Hadoop Components"}]},{"@type":"WebSite","@id":"https:\/\/data-flair.training\/blogs\/#website","url":"https:\/\/data-flair.training\/blogs\/","name":"DataFlair","description":"Learn Today. Lead Tomorrow.","publisher":{"@id":"https:\/\/data-flair.training\/blogs\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/data-flair.training\/blogs\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/data-flair.training\/blogs\/#organization","name":"DataFlair","url":"https:\/\/data-flair.training\/blogs\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/data-flair.training\/blogs\/#\/schema\/logo\/image\/","url":"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2016\/07\/Data-Flair.png","contentUrl":"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2016\/07\/Data-Flair.png","width":106,"height":48,"caption":"DataFlair"},"image":{"@id":"https:\/\/data-flair.training\/blogs\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/DataFlairWS\/","https:\/\/x.com\/DataFlairWS","https:\/\/www.linkedin.com\/company\/dataflair-web-services-pvt-ltd\/","https:\/\/www.youtube.com\/user\/DataFlairWS"]},{"@type":"Person","@id":"https:\/\/data-flair.training\/blogs\/#\/schema\/person\/2c58ecb4f73a39f0ef993f1ddfcd7b89","name":"DataFlair Team","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/1ce4a0e3e542444fc73bbebf83e89e8b73e2d95ccb1fcee64da9945f078b97c5?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/1ce4a0e3e542444fc73bbebf83e89e8b73e2d95ccb1fcee64da9945f078b97c5?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/1ce4a0e3e542444fc73bbebf83e89e8b73e2d95ccb1fcee64da9945f078b97c5?s=96&d=mm&r=g","caption":"DataFlair Team"},"description":"The DataFlair Team provides industry-driven content on programming, Java, Python, C++, DSA, AI, ML, data Science, Android, Flutter, MERN, Web Development, and technology. Our expert educators focus on delivering value-packed, easy-to-follow resources for tech enthusiasts and professionals.","url":"https:\/\/data-flair.training\/blogs\/author\/dfteam2\/"}]}},"amp_enabled":true,"_links":{"self":[{"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/posts\/49652","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/users\/6"}],"replies":[{"embeddable":true,"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/comments?post=49652"}],"version-history":[{"count":19,"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/posts\/49652\/revisions"}],"predecessor-version":[{"id":50521,"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/posts\/49652\/revisions\/50521"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/media\/50496"}],"wp:attachment":[{"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/media?parent=49652"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/categories?post=49652"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/tags?post=49652"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}