

{"id":1291,"date":"2016-12-14T18:46:56","date_gmt":"2016-12-14T18:46:56","guid":{"rendered":"http:\/\/data-flair.training\/blogs\/?p=1291"},"modified":"2021-05-09T13:23:42","modified_gmt":"2021-05-09T07:53:42","slug":"hadoop-vs-spark-vs-flink","status":"publish","type":"post","link":"https:\/\/data-flair.training\/blogs\/hadoop-vs-spark-vs-flink\/","title":{"rendered":"Hadoop vs Spark vs Flink \u2013 Big Data Frameworks Comparison"},"content":{"rendered":"<p>In this Hadoop vs Spark vs Flink tutorial, we are going to learn feature wise comparison between Apache Hadoop vs Spark vs Flink. These are the top 3 <strong>Big data technologies<\/strong> that have captured IT market very rapidly with various job roles available for them.<\/p>\n<p>You will understand the limitations of Hadoop for which Spark came into picture and drawbacks of Spark due to which Flink need arose. Here you will learn the difference between Spark and Flink and Hadoop in a detailed manner.<\/p>\n<p>So, let&#8217;s start Hadoop vs Spark vs Flink.<\/p>\n<h2>Comparison between Apache Hadoop vs Spark vs Flink<\/h2>\n<p>Before learning the difference between Hadoop vs Spark vs Flink, let us revise the basics of these 3 technologies:<br \/>\nApache Flink tutorial \u2013 4G of Big Data<br \/>\nApache Spark tutorial \u2013 3G of Big Data<br \/>\nBig data Hadoop tutorial<\/p>\n<p>So let\u2019s start the journey of feature wise comparison between Hadoop vs Spark vs Flink now:<\/p>\n<p><strong>1. Hadoop vs Spark vs Flink &#8211; Data Processing<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong>\u00a0Apache Hadoop built for batch processing. It takes large data set in the input, all at once, processes it and produces the result. Batch processing is very efficient in the processing of high volume data. An output gets delay due to\u00a0the size of the data\u00a0and the computational power of the system.<\/li>\n<li><strong>Spark<\/strong>: Apache Spark is also a part of <strong>Hadoop Ecosystem.<\/strong>\u00a0It is a batch processing System at heart too but it also supports stream processing.<\/li>\n<li><strong>Flink<\/strong>: Apache Flink provides a single runtime for the streaming and batch processing.<\/li>\n<\/ul>\n<p><strong>2.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Streaming Engine <\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> Map-reduce is batch-oriented processing tool. It takes large data set in the input, all at once, processes it and produces the result.<\/li>\n<li><strong>Spark:<\/strong>\u00a0Apache\u00a0<strong>Spark Streaming<\/strong> processes data streams in micro-batches. Each batch contains a collection of events that arrived over the batch period. But it is not enough for use cases where we need to process large streams of live data and provide results in real time.<\/li>\n<li><strong>Flink:<\/strong> Apache Flink is the true streaming engine. It uses streams for workloads: streaming, SQL, micro-batch, and batch. Batch is a finite set of streamed data.<\/li>\n<\/ul>\n<p><strong>3.\u00a0<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong><\/strong><strong>Data Flow<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> <strong>MapReduce<\/strong> computation data flow does not have any loops. It is a chain of stages. At each stage, you progress forward using an output of the previous stage and producing input for the next stage.<\/li>\n<li><strong>Spark:<\/strong> Though Machine Learning algorithm is a cyclic data flow, Spark represents it as\u00a0<strong>(DAG)\u00a0direct acyclic graph.<\/strong><\/li>\n<li><strong>Flink:<\/strong> Flink takes a different approach than others. It supports controlled cyclic dependency graph in run time. This helps to represent the Machine Learning algorithms in a very efficient way.<\/li>\n<\/ul>\n<p><strong>4.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Computation Model<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> MapReduce adopted the batch-oriented model. Batch is processing data at rest. It takes a large amount of data at once, processing it and then writing out the output.<\/li>\n<li><strong>Spark:<\/strong> Spark has adopted micro-batching. Micro-batches are an essentially \u201ccollect and then process\u201d kind of computational model.<\/li>\n<li><strong>Flink:<\/strong> Flink has adopted a continuous flow, operator-based streaming model. A continuous flow operator processes data when it arrives, without any delay in collecting the data or processing the data.<\/li>\n<\/ul>\n<p><strong>5.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Performance <\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong>\u00a0Apache Hadoop supports batch processing only. It doesn&#8217;t process streamed data hence performance is slower when compared Hadoop vs Spark vs Flink.<\/li>\n<li><strong>Spark:<\/strong> Though Apache Spark has an excellent community background and now It is considered as most matured community. But Its stream processing is not much efficient than Apache Flink as it uses micro-batch processing.<\/li>\n<li><strong>Flink:<\/strong>\u00a0Performance of Apache Flink is excellent as compared to any other data processing system. Apache Flink uses native closed loop iteration operators which make machine learning and graph processing more faster when we compare Hadoop vs Spark vs Flink.<\/li>\n<\/ul>\n<p><strong>6.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Memory management<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong>\u00a0It provides configurable Memory management. You can do it dynamically or statically.<\/li>\n<li><strong>Spark:<\/strong>\u00a0It provides configurable memory management. The latest release of Spark 1.6 has moved towards automating memory management.<\/li>\n<li><strong>Flink:<\/strong>\u00a0It provides automatic memory management. It has its own memory management system, separate from Java\u2019s garbage collector.<\/li>\n<\/ul>\n<p><strong>7.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Fault tolerance<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> MapReduce is highly fault-tolerant. There is no need to restart the application from scratch in case of any failure in Hadoop.<\/li>\n<li><strong>Spark:<\/strong>\u00a0Apache Spark Streaming recovers lost work and with no extra code or configuration, it delivers exactly-once semantics out of the box. Read more about Spark Fault Tolerance.<\/li>\n<li><strong>Flink:<\/strong> The fault tolerance mechanism followed by Apache Flink is based on <strong><em>Chandy-Lamport distributed snapshots<\/em><\/strong>. The mechanism is lightweight, which results in maintaining high throughput rates and provide strong consistency guarantees at the same time.<\/li>\n<\/ul>\n<p><strong>8.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Scalability<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> MapReduce has incredible scalability potential and has been used in production on tens of thousands of Nodes.<\/li>\n<li><strong>Spark:<\/strong>\u00a0It is highly scalable, we can keep adding n number of nodes in the cluster. A large known sSpark cluster is of 8000 nodes.<\/li>\n<li><strong>Flink:<\/strong>\u00a0Apache Flink is also highly scalable, we can keep adding n number of nodes in the cluster A large known Flink cluster is of thousands of nodes.<\/li>\n<\/ul>\n<p><strong>9.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Iterative Processing<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> It does not support iterative processing.<\/li>\n<li><strong>Spark:<\/strong>\u00a0It iterates its data in batches. In Spark, each iteration has to be scheduled and executed separately.<\/li>\n<li><strong>Flink:<\/strong>\u00a0It iterates data by using its streaming architecture. Flink can be instructed to only process the parts of the data that have actually changed, thus significantly increasing the performance of the job.<\/li>\n<\/ul>\n<p><strong>10.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Language Support<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong>\u00a0It Supports Primarily <em>Java<\/em>, other languages supported are <em>c, c++, ruby, groovy, Perl, Python.<\/em><\/li>\n<li><strong>Spark:<\/strong>\u00a0It supports <em>Java, Scala, Python<\/em> and <em>R<\/em>. Spark is implemented in Scala. It provides API in other languages like Java, Python, and R.<\/li>\n<li><strong>Flink: <\/strong>It\u00a0Supports <em>Java, Scala, Python and R<\/em>. Flink is implemented in Java. It does provide Scala API too.<\/li>\n<\/ul>\n<p><strong>11.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Optimization<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> In MapReduce, jobs have to be manually optimized. There are several ways to optimize the MapReduce Jobs: Configure your cluster correctly, use a combiner, use LZO compression, tune the number of MapReduce Task appropriately and use the most appropriate and compact writable type for your data.<\/li>\n<li><strong>Spark:<\/strong> In Apache Spark, jobs have to be manually optimized. There is a new extensible optimizer, <strong>Catalyst<\/strong>, based on functional programming construct in Scala. Catalyst\u2019s extensible design had two purposes: First, easy to add new optimization techniques. Second, enable external developers to extend the optimizer catalyst.<\/li>\n<li><strong>Flink:<\/strong>\u00a0Apache Flink comes with an optimizer that is independent with the actual programming interface. The Flink optimizer works similarly to a relational Database Optimizer but applies these optimizations to the Flink programs, rather than SQL queries.<\/li>\n<\/ul>\n<p><strong>12.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Latency<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> The MapReduce framework of Hadoop is relatively slower since it is designed to support the different format, structure and the huge volume of data. That\u2019s why Hadoop has higher latency than both Spark and Flink.<\/li>\n<li><strong>Spark:<\/strong> Apache Spark is yet another batch processing system but it is relatively faster than Hadoop MapReduce since it caches much of the input data on memory by<strong> RDD<\/strong> and keeps intermediate data in memory itself, eventually writes the data to disk upon completion or whenever required.<\/li>\n<li><strong>Flink:<\/strong> With small efforts in configuration, Apache Flink\u2019s data streaming runtime achieves low latency and high throughput.<\/li>\n<\/ul>\n<p><strong>13.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Processing Speed<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> \u00a0MapReduce processes slower than Spark and Flink. The slowness occurs only because of the nature of the MapReduce-based\u00a0execution, where it produces lots of intermediate data, much data exchanged between nodes, thus causes huge disk IO latency. Furthermore, it has to persist much data in disk for synchronization between phases so that it can support Job recovery from failures. Also, there are no ways in MapReduce to cache all subset of the data in memory.<\/li>\n<li><strong>Spark:<\/strong>\u00a0Apache Spark processes faster than MapReduce because it caches much of the input data on memory by RDD and keeps intermediate data <strong>in memory<\/strong> itself, eventually writes the data to disk upon completion or whenever required. Spark is 100 times faster than MapReduce and this shows how Spark is better than Hadoop MapReduce.<\/li>\n<li><strong>Flink:<\/strong>\u00a0It processes faster than Spark because of its streaming architecture. Flink increases the performance of the job by instructing to only process part of data\u00a0that have actually changed.<\/li>\n<\/ul>\n<p><strong>14.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Visualization<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong>\u00a0In Hadoop, data visualization tool is <strong>zoomdata<\/strong> that can connect directly to <strong>HDFS<\/strong> as well on SQL-on-Hadoop technologies such as Impala,<strong> Hive<\/strong>, <strong>Spark SQL,<\/strong> Presto and more.<\/li>\n<li><strong>Spark:<\/strong>\u00a0It offers a web interface for submitting and executing jobs on which the resulting execution plan can be visualized. Flink and Spark both are integrated to Apache <strong>zeppelin<\/strong> It provides data analytics, ingestion, as well as discovery, visualization, and collaboration.<\/li>\n<li><strong>Flink:<\/strong>\u00a0It also offers a web interface for submitting and executing jobs. The resulting execution plan can be visualized on this interface.<\/li>\n<\/ul>\n<p><strong>15.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Recovery<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> MapReduce is naturally resilient to system faults or failures. It is the highly fault-tolerant system.<\/li>\n<li><strong>Spark:<\/strong>\u00a0Apache Spark RDDs allow recovery of partitions on failed nodes by re-computation of the <strong>DAG<\/strong> while also supporting a more similar recovery style to Hadoop by way of checkpointing, to reduce the dependencies of RDDs.<\/li>\n<li><strong>Flink:<\/strong>\u00a0It supports checkpointing mechanism that stores the program in the data sources and data sink, the state of the window, as well as user-defined state that recovers streaming job after failure.<\/li>\n<\/ul>\n<p><strong>16.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Security<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong>\u00a0It supports <strong>Kerberos authentication<\/strong>, which is somewhat painful to manage. But, third party vendors have enabled organizations to leverage Active Directory Kerberos and LDAP for authentication.<\/li>\n<li><strong>Spark:<\/strong>\u00a0Apache Spark\u2019s security is a bit sparse by currently only supporting authentication via shared secret (password authentication). The security bonus that Spark can enjoy is that if you run Spark on HDFS, it can use HDFS ACLs and file-level permissions. Additionally, Spark can run on <strong>YARN<\/strong>\u00a0to use Kerberos authentication.<\/li>\n<li><strong>Flink:<\/strong> There is user-authentication support in Flink via the Hadoop \/ Kerberos infrastructure. If you run Flink on YARN, Flink acquires the Kerberos tokens of the user that submits programs, and authenticate itself at YARN, HDFS, and <strong>HBase<\/strong> with that.Flink&#8217;s upcoming connector, streaming programs can authenticate themselves as stream brokers via SSL.<\/li>\n<\/ul>\n<p><strong>17.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Cost<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> MapReduce can typically run on less expensive hardware than some alternatives since it does not attempt to store everything in memory.<\/li>\n<li><strong>Spark:<\/strong> As spark requires a lot of RAM to run in-memory, increasing it in the cluster, gradually increases its cost.<\/li>\n<li><strong>Flink:<\/strong>\u00a0Apache Flink also requires a lot of RAM to run in-memory, so it will increase its cost gradually.<\/li>\n<\/ul>\n<p><strong>18.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Compatibility<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong>\u00a0Apache Hadoop MapReduce and Apache Spark are compatible with each other and Spark shares all MapReduce\u2019s compatibilities for data sources, file formats, and business intelligence tools via JDBC and ODBC.<\/li>\n<li><strong>Spark:<\/strong>\u00a0Apache Spark and Hadoop are compatible to each other. Spark is compatible with Hadoop data. It can run in Hadoop clusters through YARN or Spark&#8217;s standalone mode, and it can process data in HDFS, HBase, Cassandra, Hive, and any Hadoop InputFormat.<\/li>\n<li><strong>Flink:<\/strong>\u00a0Apache Flink is a scalable data analytics framework that is fully compatible to Hadoop. It provides a Hadoop Compatibility package to wrap functions implemented against Hadoop\u2019s MapReduce interfaces and embed them in Flink programs.<\/li>\n<\/ul>\n<p><strong>19.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Abstraction <\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> In MapReduce, we don\u2019t have any type of abstraction.<\/li>\n<li><strong>Spark:<\/strong> In Spark, for batch, we have <em>Spark RDD<\/em> abstraction and <em>DStream<\/em> for streaming which is internally RDD itself.<\/li>\n<li><strong>Flink:<\/strong> In Flink, we have Dataset abstraction for batch and DataStreams for the streaming application.<\/li>\n<\/ul>\n<p><strong>20.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Easy to use<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> MapReduce developers need to hand code each operation which makes it very difficult to work.<\/li>\n<li><strong>Spark:<\/strong>\u00a0It is easy to program as it has tons of high-level operators.<\/li>\n<li><strong>Flink:<\/strong>\u00a0It also has high-level operators.<\/li>\n<\/ul>\n<p><strong>21.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Interactive Mode<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> MapReduce does not have interactive Mode.<\/li>\n<li><strong>Spark:<\/strong>\u00a0Apache Spark has an interactive shell to learn how to make the most out of Apache Spark. This is a Spark application written in Scala to offer a command-line environment with auto-completion where you can run ad-hoc queries and get familiar with the features of Spark.<\/li>\n<li><strong>Flink:<\/strong>\u00a0It comes with an integrated interactive Scala Shell. It can be used in a local setup as well as in a cluster setup.<\/li>\n<\/ul>\n<p><strong> 22.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Real-time Analysis<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> MapReduce fails when it comes to real-time data processing as it was designed to perform batch processing on voluminous amounts of data.<\/li>\n<li><strong>Spark:<\/strong> It can process real time data ie data coming from the real-time event streams at the rate of millions of events per second.<\/li>\n<li><strong>Flink:<\/strong> It is mainly used for real-time data Analysis Although it also provides fast batch data Processing.<\/li>\n<\/ul>\n<p><strong>23<\/strong><strong>.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Scheduler<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> Scheduler in Hadoop becomes the pluggable component. There are two schedulers for multi-user workload: <em>Fair Scheduler a<\/em>nd <em>Capacity Scheduler<\/em>. To schedule complex flows,\u00a0MapReduce needs an external job scheduler like<strong> Oozie<\/strong>.<\/li>\n<li><strong>Spark:<\/strong> Due to in-memory computation, spark acts its own flow scheduler.<\/li>\n<li><strong>Flink:<\/strong> Flink can use YARN Scheduler but Flink also has its own Scheduler.<\/li>\n<\/ul>\n<p><strong>24<\/strong><strong>.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>SQL support<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> It enables users to run SQL queries using Apache Hive.<\/li>\n<li><strong>Spark:<\/strong> It enables users to run SQL queries using Spark-SQL. Spark provides both Hives like query language and <strong>Dataframe<\/strong> like DSL for querying structured data.<\/li>\n<li><strong>Flink:<\/strong> In Flink, Table API is an SQL-like expression language that supports data frame like DSL and it\u2019s still in beta. There are plans to add the SQL interface but not sure when it will land in the framework.<\/li>\n<\/ul>\n<p><strong>25<\/strong><strong>.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Caching \u00a0<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> MapReduce cannot cache the data in memory for future requirements<\/li>\n<li><strong>Spark:<\/strong>\u00a0It can cache data in memory for further iterations which enhance its performance.<\/li>\n<li><strong>Flink:<\/strong>\u00a0It can cache data in memory for further iterations to enhance its performance.<\/li>\n<\/ul>\n<p><strong>26<\/strong><strong>.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Hardware Requirements<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> MapReduce runs very well on Commodity Hardware.<\/li>\n<li><strong>Spark:<\/strong>\u00a0Apache Spark needs mid to high-level hardware. Since Spark cache data in-memory for further\u00a0iterations which enhance its performance.<\/li>\n<li><strong>Flink:<\/strong>\u00a0Apache Flink also needs mid to High-level Hardware. Flink can also cache data in memory for further iterations which enhance its performance.<\/li>\n<\/ul>\n<p><strong>27.\u00a0<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong><\/strong><strong style=\"font-family: Verdana, Geneva, sans-serif\">Machine Learning<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong>\u00a0It requires machine learning tool like <strong>Apache Mahout<\/strong>.<\/li>\n<li><strong>Spark:<\/strong>\u00a0It has its own set of machine learning MLlib. Within memory caching and other implementation details, it\u2019s really powerful platform to implement ML algorithms.<\/li>\n<li><strong>Flink:<\/strong>\u00a0It has FlinkML which is Machine Learning library for Flink. It supports controlled <em>cyclic dependency graph<\/em> in runtime. This makes them represent the ML algorithms in a very efficient way compared to<em> DAG<\/em> representation.<\/li>\n<\/ul>\n<p><strong>28.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Line of code<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> Hadoop 2.0 has 1,20,000 line of codes. More no of lines produce more no of bugs and it will take much time to execute the program.<\/li>\n<li><strong>Spark:<\/strong> Apache Spark is developed in merely 20000 line of codes. No. of the line of code is lesser than Hadoop. So it will take less time to execute the program.<\/li>\n<li><strong>Flink:<\/strong> Flink is developed in scala and java, so no. of the line of code is lesser than Hadoop. So it will also take the less time to execute the program.<\/li>\n<\/ul>\n<p><strong>29.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>High Availability <\/strong><br \/>\n<em>The High availability<\/em> refers to a system or component that is operational for long length of time.<\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> Configurable in High Availability Mode.<\/li>\n<li><strong>Spark:<\/strong> Configurable in High Availability Mode.<\/li>\n<li><strong>Flink:<\/strong> \u00a0Configurable in High Availability Mode.<\/li>\n<\/ul>\n<p><strong>30.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Amazon S3 connector<\/strong><br \/>\nAmazon Simple Storage Service (Amazon S3) is object storage with a simple web service interface to store and retrieve any amount of data from anywhere on the web.<strong>\u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> Provides Supports for Amazon S3 Connector.<\/li>\n<li><strong>Spark:<\/strong> Provides Supports for Amazon S3 Connector.<\/li>\n<li><strong>Flink:<\/strong> Provides Supports for Amazon S3 connector.<\/li>\n<\/ul>\n<p><strong>31.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Deployment<\/strong><\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> In Standalone mode, Hadoop is configured to run in a single-node, non-distributed mode. In Pseudo-Distributed mode, Hadoop runs in a pseudo distributed mode. Thus, the difference is that each Hadoop daemon runs in a separate Java process in pseudo-distributed mode. Whereas in local mode each Hadoop daemon runs as a single Java process. In a fully-distributed mode, all daemons execute in separate nodes forming a multi-node cluster.<\/li>\n<li><strong>Spark:<\/strong> It also provides a simple standalone deploy mode to running on the <strong>Mesos<\/strong> or YARN cluster managers. It can be launched either manually, by starting a master and workers by hand or use our provided launch scripts. It is also possible to run these daemons on a single machine for testing.<\/li>\n<li><strong>Flink:<\/strong> It also provides standalone deploy mode to running on YARN cluster Managers.<\/li>\n<\/ul>\n<p><strong>32.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Back pressure Handing<\/strong><br \/>\nBackPressure refers to the buildup of data at an I\/O switch when buffers are full and not able to receive more data. No more data packets transfer until the bottleneck of data eliminates or the buffer is empty.<\/p>\n<ul>\n<li><strong>Hadoop:<\/strong>\u00a0It handles back pressure through Manual Configuration.<\/li>\n<li><strong>Spark:<\/strong>\u00a0It also handles back pressure through Manual Configuration.<\/li>\n<li><strong>Flink:<\/strong>\u00a0It handles back pressure Implicitly through System Architecture.<\/li>\n<\/ul>\n<p><strong>33.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Duplication Elimination<\/strong><\/p>\n<ul>\n<li><strong>Hadoop: <\/strong>There is no duplication elimination in Hadoop.<\/li>\n<li><strong>Spark:<\/strong> Spark also processes every record exactly one time hence eliminates duplication.<\/li>\n<li><strong>Flink:<\/strong> Apache Flink processes every record exactly one time hence eliminates duplication. Streaming applications can maintain custom state during their computation. Flink\u2019s checkpointing mechanism ensures exactly once semantics for the state in the presence of failures.<\/li>\n<\/ul>\n<p><strong>34.\u00a0<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong><\/strong><strong style=\"font-family: Verdana, Geneva, sans-serif\">Windows criteria<\/strong><br \/>\nA data stream needs to be grouped into many logical streams on each of which a window operator can be applied.<\/p>\n<ul>\n<li><strong>Hadoop:<\/strong>\u00a0It doesn\u2019t support streaming so there is no need of window criteria.<\/li>\n<li><strong>Spark:<\/strong>\u00a0It has time-based window criteria.<\/li>\n<li><strong>Flink: <\/strong>It has record-based or any custom user-defined Flink Window criteria.<\/li>\n<\/ul>\n<p><strong>35.<strong style=\"font-family: Verdana, Geneva, sans-serif\"> Hadoop vs Spark vs Flink &#8211; <\/strong>Apache License<\/strong><br \/>\nThe Apache License, Version 2.0 (ALv2) is a permissive free software license written by the Apache Software Foundation (ASF). The Apache License requires preservation of the copyright notice and disclaimer.<\/p>\n<ul>\n<li><strong>Hadoop:<\/strong> Apache License 2.<\/li>\n<li><strong>Spark:<\/strong> Apache License 2.<\/li>\n<li><strong>Flink:<\/strong> Apache License 2.<\/li>\n<\/ul>\n<p>So, this is how the comparison is done between the top 3 Big data technologies Hadoop vs Spark vs Flink.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>In this Hadoop vs Spark vs Flink tutorial, we are going to learn feature wise comparison between Apache Hadoop vs Spark vs Flink. These are the top 3 Big data technologies that have captured&#46;&#46;&#46;<\/p>\n","protected":false},"author":6,"featured_media":34408,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[8,10,22],"tags":[750,782,792,793,896,961,4738,5186,5346,5347,5348,13021,13146],"class_list":["post-1291","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-flink","category-spark","category-hadoop","tag-apache-flink","tag-apache-hadoop","tag-apache-hadoop-vs-flink","tag-apache-hadoop-vs-spark","tag-apache-spark","tag-apache-spark-vs-flink","tag-flink","tag-hadoop","tag-hadoop-vs-flink","tag-hadoop-vs-spark","tag-hadoop-vs-spark-vs-flink","tag-spark","tag-spark-vs-flink"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.0 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Hadoop vs Spark vs Flink \u2013 Big Data Frameworks Comparison - DataFlair<\/title>\n<meta name=\"description\" content=\"Hadoop vs Spark vs Flink tutorial-Difference between Spark vs Flink vs Hadoop, how Flink &amp; Spark are better than Hadoop &amp; what to choose Spark,Flink,Hadoop?\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/data-flair.training\/blogs\/hadoop-vs-spark-vs-flink\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Hadoop vs Spark vs Flink \u2013 Big Data Frameworks Comparison - DataFlair\" \/>\n<meta property=\"og:description\" content=\"Hadoop vs Spark vs Flink tutorial-Difference between Spark vs Flink vs Hadoop, how Flink &amp; Spark are better than Hadoop &amp; what to choose Spark,Flink,Hadoop?\" \/>\n<meta property=\"og:url\" content=\"https:\/\/data-flair.training\/blogs\/hadoop-vs-spark-vs-flink\/\" \/>\n<meta property=\"og:site_name\" content=\"DataFlair\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/DataFlairWS\/\" \/>\n<meta property=\"article:published_time\" content=\"2016-12-14T18:46:56+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2021-05-09T07:53:42+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2016\/12\/Feature-Wise-Comparison-of-Hadoop-vs-Spark-vs-Flink-01.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"1200\" \/>\n\t<meta property=\"og:image:height\" content=\"628\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"author\" content=\"DataFlair Team\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:creator\" content=\"@DataFlairWS\" \/>\n<meta name=\"twitter:site\" content=\"@DataFlairWS\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"DataFlair Team\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"14 minutes\" \/>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Hadoop vs Spark vs Flink \u2013 Big Data Frameworks Comparison - DataFlair","description":"Hadoop vs Spark vs Flink tutorial-Difference between Spark vs Flink vs Hadoop, how Flink & Spark are better than Hadoop & what to choose Spark,Flink,Hadoop?","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/data-flair.training\/blogs\/hadoop-vs-spark-vs-flink\/","og_locale":"en_US","og_type":"article","og_title":"Hadoop vs Spark vs Flink \u2013 Big Data Frameworks Comparison - DataFlair","og_description":"Hadoop vs Spark vs Flink tutorial-Difference between Spark vs Flink vs Hadoop, how Flink & Spark are better than Hadoop & what to choose Spark,Flink,Hadoop?","og_url":"https:\/\/data-flair.training\/blogs\/hadoop-vs-spark-vs-flink\/","og_site_name":"DataFlair","article_publisher":"https:\/\/www.facebook.com\/DataFlairWS\/","article_published_time":"2016-12-14T18:46:56+00:00","article_modified_time":"2021-05-09T07:53:42+00:00","og_image":[{"width":1200,"height":628,"url":"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2016\/12\/Feature-Wise-Comparison-of-Hadoop-vs-Spark-vs-Flink-01.jpg","type":"image\/jpeg"}],"author":"DataFlair Team","twitter_card":"summary_large_image","twitter_creator":"@DataFlairWS","twitter_site":"@DataFlairWS","twitter_misc":{"Written by":"DataFlair Team","Est. reading time":"14 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/data-flair.training\/blogs\/hadoop-vs-spark-vs-flink\/#article","isPartOf":{"@id":"https:\/\/data-flair.training\/blogs\/hadoop-vs-spark-vs-flink\/"},"author":{"name":"DataFlair Team","@id":"https:\/\/data-flair.training\/blogs\/#\/schema\/person\/2c58ecb4f73a39f0ef993f1ddfcd7b89"},"headline":"Hadoop vs Spark vs Flink \u2013 Big Data Frameworks Comparison","datePublished":"2016-12-14T18:46:56+00:00","dateModified":"2021-05-09T07:53:42+00:00","mainEntityOfPage":{"@id":"https:\/\/data-flair.training\/blogs\/hadoop-vs-spark-vs-flink\/"},"wordCount":3117,"commentCount":6,"publisher":{"@id":"https:\/\/data-flair.training\/blogs\/#organization"},"image":{"@id":"https:\/\/data-flair.training\/blogs\/hadoop-vs-spark-vs-flink\/#primaryimage"},"thumbnailUrl":"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2016\/12\/Feature-Wise-Comparison-of-Hadoop-vs-Spark-vs-Flink-01.jpg","keywords":["apache flink","apache hadoop","apache hadoop vs flink","apache hadoop vs spark","apache spark","apache spark vs flink","flink","hadoop","hadoop vs flink","hadoop vs spark","hadoop vs spark vs flink","Spark","spark vs flink"],"articleSection":["Apache Flink Tutorials","Apache Spark Tutorials","Hadoop Tutorials"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/data-flair.training\/blogs\/hadoop-vs-spark-vs-flink\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/data-flair.training\/blogs\/hadoop-vs-spark-vs-flink\/","url":"https:\/\/data-flair.training\/blogs\/hadoop-vs-spark-vs-flink\/","name":"Hadoop vs Spark vs Flink \u2013 Big Data Frameworks Comparison - DataFlair","isPartOf":{"@id":"https:\/\/data-flair.training\/blogs\/#website"},"primaryImageOfPage":{"@id":"https:\/\/data-flair.training\/blogs\/hadoop-vs-spark-vs-flink\/#primaryimage"},"image":{"@id":"https:\/\/data-flair.training\/blogs\/hadoop-vs-spark-vs-flink\/#primaryimage"},"thumbnailUrl":"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2016\/12\/Feature-Wise-Comparison-of-Hadoop-vs-Spark-vs-Flink-01.jpg","datePublished":"2016-12-14T18:46:56+00:00","dateModified":"2021-05-09T07:53:42+00:00","description":"Hadoop vs Spark vs Flink tutorial-Difference between Spark vs Flink vs Hadoop, how Flink & Spark are better than Hadoop & what to choose Spark,Flink,Hadoop?","breadcrumb":{"@id":"https:\/\/data-flair.training\/blogs\/hadoop-vs-spark-vs-flink\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/data-flair.training\/blogs\/hadoop-vs-spark-vs-flink\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/data-flair.training\/blogs\/hadoop-vs-spark-vs-flink\/#primaryimage","url":"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2016\/12\/Feature-Wise-Comparison-of-Hadoop-vs-Spark-vs-Flink-01.jpg","contentUrl":"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2016\/12\/Feature-Wise-Comparison-of-Hadoop-vs-Spark-vs-Flink-01.jpg","width":1200,"height":628,"caption":"Hadoop vs Spark vs Flink"},{"@type":"BreadcrumbList","@id":"https:\/\/data-flair.training\/blogs\/hadoop-vs-spark-vs-flink\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Blog Home","item":"https:\/\/data-flair.training\/blogs\/"},{"@type":"ListItem","position":2,"name":"Apache Flink Tutorials","item":"https:\/\/data-flair.training\/blogs\/category\/flink\/"},{"@type":"ListItem","position":3,"name":"Hadoop vs Spark vs Flink \u2013 Big Data Frameworks Comparison"}]},{"@type":"WebSite","@id":"https:\/\/data-flair.training\/blogs\/#website","url":"https:\/\/data-flair.training\/blogs\/","name":"DataFlair","description":"Learn Today. Lead Tomorrow.","publisher":{"@id":"https:\/\/data-flair.training\/blogs\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/data-flair.training\/blogs\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/data-flair.training\/blogs\/#organization","name":"DataFlair","url":"https:\/\/data-flair.training\/blogs\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/data-flair.training\/blogs\/#\/schema\/logo\/image\/","url":"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2016\/07\/Data-Flair.png","contentUrl":"https:\/\/data-flair.training\/blogs\/wp-content\/uploads\/sites\/2\/2016\/07\/Data-Flair.png","width":106,"height":48,"caption":"DataFlair"},"image":{"@id":"https:\/\/data-flair.training\/blogs\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/DataFlairWS\/","https:\/\/x.com\/DataFlairWS","https:\/\/www.linkedin.com\/company\/dataflair-web-services-pvt-ltd\/","https:\/\/www.youtube.com\/user\/DataFlairWS"]},{"@type":"Person","@id":"https:\/\/data-flair.training\/blogs\/#\/schema\/person\/2c58ecb4f73a39f0ef993f1ddfcd7b89","name":"DataFlair Team","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/1ce4a0e3e542444fc73bbebf83e89e8b73e2d95ccb1fcee64da9945f078b97c5?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/1ce4a0e3e542444fc73bbebf83e89e8b73e2d95ccb1fcee64da9945f078b97c5?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/1ce4a0e3e542444fc73bbebf83e89e8b73e2d95ccb1fcee64da9945f078b97c5?s=96&d=mm&r=g","caption":"DataFlair Team"},"description":"The DataFlair Team provides industry-driven content on programming, Java, Python, C++, DSA, AI, ML, data Science, Android, Flutter, MERN, Web Development, and technology. Our expert educators focus on delivering value-packed, easy-to-follow resources for tech enthusiasts and professionals.","url":"https:\/\/data-flair.training\/blogs\/author\/dfteam2\/"}]}},"amp_enabled":true,"_links":{"self":[{"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/posts\/1291","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/users\/6"}],"replies":[{"embeddable":true,"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/comments?post=1291"}],"version-history":[{"count":1,"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/posts\/1291\/revisions"}],"predecessor-version":[{"id":94122,"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/posts\/1291\/revisions\/94122"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/media\/34408"}],"wp:attachment":[{"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/media?parent=1291"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/categories?post=1291"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/data-flair.training\/blogs\/wp-json\/wp\/v2\/tags?post=1291"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}