Apache Hadoop vs. Azure Data Lake Storage

Overview
ProductRatingMost Used ByProduct SummaryStarting Price
Hadoop
Score 7.9 out of 10
N/A
Hadoop is an open source software from Apache, supporting distributed processing and data storage. Hadoop is popular for its scalability, reliability, and functionality available across commoditized hardware.N/A
Azure Data Lake Storage
Score 9.6 out of 10
N/A
Azure Data Lake Storage Gen2 is a highly scalable and cost-effective data lake solution for big data analytics. It combines the power of a high-performance file system with massive scale and economy to help you speed your time to insight. Data Lake Storage Gen2 extends Azure Blob Storage capabilities and is optimized for analytics workloads.N/A
Pricing
Apache HadoopAzure Data Lake Storage
Editions & Modules
No answers on this topic
No answers on this topic
Offerings
Pricing Offerings
HadoopAzure Data Lake Storage
Free Trial
NoNo
Free/Freemium Version
YesNo
Premium Consulting/Integration Services
NoNo
Entry-level Setup FeeNo setup feeNo setup fee
Additional Details
More Pricing Information
Community Pulse
Apache HadoopAzure Data Lake Storage
Best Alternatives
Apache HadoopAzure Data Lake Storage
Small Businesses

No answers on this topic

Backblaze B2 Cloud Storage
Backblaze B2 Cloud Storage
Score 9.6 out of 10
Medium-sized Companies
Cloudera Manager
Cloudera Manager
Score 9.9 out of 10
Azure Blob Storage
Azure Blob Storage
Score 9.7 out of 10
Enterprises
IBM Analytics Engine
IBM Analytics Engine
Score 7.1 out of 10
Azure Blob Storage
Azure Blob Storage
Score 9.7 out of 10
All AlternativesView all alternativesView all alternatives
User Ratings
Apache HadoopAzure Data Lake Storage
Likelihood to Recommend
8.0
(0 ratings)
8.2
(0 ratings)
Likelihood to Renew
9.6
(0 ratings)
-
(0 ratings)
Usability
8.0
(0 ratings)
-
(0 ratings)
Performance
8.0
(0 ratings)
-
(0 ratings)
Support Rating
7.5
(0 ratings)
-
(0 ratings)
Online Training
6.1
(0 ratings)
-
(0 ratings)
User Testimonials
Apache HadoopAzure Data Lake Storage
Likelihood to Recommend
Apache Hadoop (and its subsequent add-ons) are well-suited to larger, unstructured data flows, such as aggregation of web traffic or advertising. Geospatial algorithms and their outputs are well-suited for this kind of aggregation as structuring that data is challenging, but leaving it unstructured and performing queries as-needed is a better fit for most business models. With the advent of data science, I would expect Hadoop fits a LOT of their initial outputs quite well.
Read full review
Azure Data Lake storage is well suited for applications/use cases within organizations where capturing and storing large amounts of data in any format is required, primarily for storing and processing purposes. It's an easy and cost-effective cloud solution for your application data. The ability to integrate with other Azure Services like Azure Databricks and Azure Data Factory is superb.
Read full review
Pros
  • HDFS is reliable and solid, and in my experience with it, there are very few problems using it
  • Enterprise support from different vendors makes it easier to 'sell' inside an enterprise
  • It provides High Scalability and Redundancy
  • Horizontal scaling and distributed architecture
Read full review
  • Azure Data Lake Storage is extremely scalable. It allows us to scale up or down endlessly based on what we need including replication.
  • In terms of security, Azure Data Lake Storage fits our requirements really well as we can monitor and encrypt seamlessly. We can also assign permissions through roles and grant network-level access.
  • Due to the fact that it can scale, we are able to monitor the cost of storage and any given time and make financial decisions about our infrastructure based on how small or big we want to scale.
Read full review
Cons
  • Hadoop is a batch oriented processing framework, it lacks real time or stream processing.
  • Hadoop's HDFS file system is not a POSIX compliant file system and does not work well with small files, especially smaller than the default block size.
  • Hadoop cannot be used for running interactive jobs or analytics.
Read full review
  • I'd like to see a better cross-platform native client. Azure Data Explorer is fine, but it's far from the "SSMS" kind of experience SQL Server users are used to.
  • Listing a large number of file is somewhat problematic and slow. Using the native C# library, running directly on an Azure VM, it can take several hours to list just a couple million files.
  • Switching from V1 to V2 requires the creation of a new Storage Account and that's pretty inconvenient.
Read full review
Likelihood to Renew
Hadoop is organization-independent and can be used for various purposes ranging from archiving to reporting and can make use of economic, commodity hardware. There is also a lot of saving in terms of licensing costs - since most of the Hadoop ecosystem is available as open-source and is free
Read full review
No answers on this topic
Usability
Great! Hadoop has an easy to use interface that mimics most other data warehouses. You can access your data via SQL and have it display in a terminal before exporting it to your business intelligence platform of choice. Of course, for smaller data sets, you can also export it to Microsoft Excel.
Read full review
No answers on this topic
Support Rating
We went with a third party for support, i.e., consultant. Had we gone with Azure or Cloudera, we would have obtained support directly from the vendor. my rating is more on the third party we selected and doesn't reflect the overall support available for Hadoop. I think we could have done better in our selection process, however, we were trying to use an already approved vendor within our organization. There is plenty of self-help available for Hadoop online.
Read full review
No answers on this topic
Online Training
Hadoop is a complex topic and best suited for classrom training. Online training are a waste of time and money.
Read full review
No answers on this topic
Alternatives Considered
I feel that this is a highly reliable and scalable solution computing technology that is highly capable of processing large data sets across multiple servers and thousands of machines in a well-defined and distributed manner. Apache Hadoop can automatically scale up the number of servers and machines that are needed to process, store, and analyze data sets. It also handles explosions in data with big data technology. Apache Hadoop is good at handling all node failures as well.
Read full review
The Azure Data Lake solution is designed for organizations that want to take advantage of big data. It provides a data platform that can help developers, data scientists, and analysts store data of any size and format and perform all types of processing and analytics across multiple platforms and programming languages. It can work with your existing solutions, such as identity management and security solutions. It also integrates with other data warehouses and cloud environments. It can be useful for organizations that need the above softwares.
Read full review
Return on Investment
  • As it was open source makes it popular choice for handling large chuck of datasets
  • It was free earlier but now it’s licensed but still enterprise is a fine tuned version which makes it easier for new users and administrators to use it
  • Our investment is worth every single penny.
  • Initial cost is more as you might need to hire administrators to setup the cluster and make them in scalable. But once done it’s pretty easy
Read full review
  • The cost can be high for more advanced work. In some cases, for instance, time limits and lab runtimes may be too short if you are too slow to learn what is explained as you go along.
  • promote flexible team communication. You can create different spaces for different teams, and share files and tasks.
Read full review
ScreenShots