AWS EMR - Hadoop, Spark, Pesto, and Hive made easy?

3:21 · Watch on YouTube ↗

Transcript 410 words · about 3 min to read

Auto-generated captions from YouTube, not hand-corrected, so names and technical terms may be imperfect. The video is authoritative.

all right to properly understand AWS EMR we have to understand the data Lake and data analytics Journey let's start with the beginning of data Lakes this idea that I have a bunch of different data sources from databases to flat file systems we create a data Lake a file system or file storage mechanism where I can just dump all of that data into the analytics platform so once I have a data Lake I need to transact on that data Lake this is where a product something like

or a project something like Hadoop comes into play with Hadoop I can take my file system hdfs Hadoop file system expand that across multiple nodes and then transact on that data using the power of these nodes it's not just storage it's storage in compute so my data scientist can write software packages that use Hadoop as the interface a problem with Hadoop file system in Hadoop in itself one is complex and then two it's not performing if that's a word compared to something like an in-memory database

we talked about that what Fred is sap Hana Etc so I may want to use something like Spark on top of hdfs to write faster curies real-time curies bash processes Etc in a very simple language compared to what I have to write with Hadoop if you're a it expert you're starting to think and this is the route that I went down how do I put my spark system on top of hdfs without creating a separate spark cluster kind of use the same nodes and this is

where EMR comes into play I can now use EMR to manage all of this layer of infrastructure set up my cluster AWS abstracts away the S3 storage to the compute and I just used the interface from it from either a serverless perspective or if I have to turn the nerd knobs I can do that within a trade additional cluster now I can use this same cluster cluster or technology to run a pesto nodes or set of nodes a hive cluster whatever my framework for analytics is

I can now use that on top of the AWS infrastructure without having to worry about scaling and managing the underlying infrastructure you want to learn more about the CTO advisor you can visit us on the web the ctoadvisor.com AWS everyday.com to subscribe to AWS everyday in your inbox