Azure Databricks is a service available on Microsoft's Azure platform and suite of products. It provides the latest versions of Apache Spark so users can integrate with open source libraries, or spin up clusters and build in a fully managed Apache Spark environment with the global scale and availability of Azure. Clusters are set up, configured, and fine-tuned to ensure reliability and performance without the need for monitoring. The solution includes autoscaling and auto-termination to improve…
N/A
IBM DataStage
Score 7.6 out of 10
N/A
IBM® DataStage® is a data integration tool that helps users to design, develop and run jobs that move and transform data. At its core, the DataStage tool supports extract, transform and load (ETL) and extract, load and transform (ELT) patterns. A basic version of the software is available for on-premises deployment, and the cloud-based DataStage for IBM Cloud Pak® for Data offers automated integration capabilities in a hybrid or multicloud environment.
N/A
Pricing
Azure Databricks
IBM DataStage
Editions & Modules
No answers on this topic
No answers on this topic
Offerings
Pricing Offerings
Azure Databricks
IBM DataStage
Free Trial
No
Yes
Free/Freemium Version
No
No
Premium Consulting/Integration Services
No
No
Entry-level Setup Fee
No setup fee
No setup fee
Additional Details
—
—
More Pricing Information
Community Pulse
Azure Databricks
IBM DataStage
Features
Azure Databricks
IBM DataStage
Platform Connectivity
Comparison of Platform Connectivity features of Product A and Product B
Azure Databricks
8.1
Ratings
3% below category average
IBM DataStage
-
Ratings
Connect to Multiple Data Sources
6.20 Ratings
00 Ratings
Extend Existing Data Sources
9.00 Ratings
00 Ratings
Automatic Data Format Detection
9.00 Ratings
00 Ratings
MDM Integration
8.00 Ratings
00 Ratings
Data Exploration
Comparison of Data Exploration features of Product A and Product B
Azure Databricks
6.4
Ratings
27% below category average
IBM DataStage
-
Ratings
Visualization
5.90 Ratings
00 Ratings
Interactive Data Analysis
6.90 Ratings
00 Ratings
Data Preparation
Comparison of Data Preparation features of Product A and Product B
Azure Databricks
8.0
Ratings
2% below category average
IBM DataStage
-
Ratings
Interactive Data Cleaning and Enrichment
7.00 Ratings
00 Ratings
Data Transformations
9.00 Ratings
00 Ratings
Data Encryption
9.00 Ratings
00 Ratings
Built-in Processors
7.10 Ratings
00 Ratings
Platform Data Modeling
Comparison of Platform Data Modeling features of Product A and Product B
Azure Databricks
8.3
Ratings
1% below category average
IBM DataStage
-
Ratings
Multiple Model Development Languages and Tools
8.10 Ratings
00 Ratings
Automated Machine Learning
9.00 Ratings
00 Ratings
Single platform for multiple model development
8.00 Ratings
00 Ratings
Self-Service Model Delivery
8.00 Ratings
00 Ratings
Model Deployment
Comparison of Model Deployment features of Product A and Product B
Azure Databricks
8.5
Ratings
0% below category average
IBM DataStage
-
Ratings
Flexible Model Publishing Options
8.00 Ratings
00 Ratings
Security, Governance, and Cost Controls
9.00 Ratings
00 Ratings
Data Source Connection
Comparison of Data Source Connection features of Product A and Product B
Azure Databricks
-
Ratings
IBM DataStage
9.5
Ratings
12% above category average
Connect to traditional data sources
00 Ratings
10.00 Ratings
Connecto to Big Data and NoSQL
00 Ratings
9.00 Ratings
Data Transformations
Comparison of Data Transformations features of Product A and Product B
Azure Databricks
-
Ratings
IBM DataStage
8.0
Ratings
2% below category average
Simple transformations
00 Ratings
8.00 Ratings
Complex transformations
00 Ratings
8.00 Ratings
Data Modeling
Comparison of Data Modeling features of Product A and Product B
Azure Databricks
-
Ratings
IBM DataStage
6.3
Ratings
23% below category average
Data model creation
00 Ratings
5.00 Ratings
Metadata management
00 Ratings
5.00 Ratings
Business rules and workflow
00 Ratings
6.00 Ratings
Collaboration
00 Ratings
6.00 Ratings
Testing and debugging
00 Ratings
6.00 Ratings
Data Governance
Comparison of Data Governance features of Product A and Product B
Having access to all databases and tables in one place is what has helped me and my team to function better. The in built functionality/access to SQL and Python is definitely an added bonus! The icing on the cake is the ability to export your data into an Excel spreadsheet for additional analysis. If you have less to no working knowledge of SQL or Python, its better to look at alternatives.
Excellent Cloud data mapping tool and easy creating multiple project data analytics in real-time and the report distribution are excellent via this IBM product. Easy tool to provide data visualization and the integration is effective and helpful to migrating huge amounts of data across other platforms and different websites insights gathering.
Based on my extensive use of Azure Databricks for the past 3.5 years, it has evolved into a beautiful amalgamation of all the data domains and needs. From a data analyst, to a data engineer, to a data scientist, it jas got them all! Being language agnostic and focused on easy to use UI based control, it is a dream to use for every Data related personnel across all experience levels!
Because it is a flexible tool that can manage many flows and create a strong solution with a interesting use of variables. Easy to scale up as you can copy jobs arleady build and modify them. SQL queries allow to be fast in development and have the pushdown feature, but you loose a little of user friendly look. Metadata management is not strong as a visual feature, but can be determine by job codes.
It could load thousands of records in seconds. But in the Parallel version, you need to understand how to particionate the data. If you use the algorithms erroneously, or the functionalities that it gives for the parsing of data, the performance can fall drastically, even with few records. It is necessary to have people with experience to be able to determine which algorithm to use and understand why.
IBM offers different levels of support but in my experience being and IBM shop helps to get direct support from more knowledgeable technicians from IBM. Not sure on the cost of having this kind of support, but I know there's also general support and community blogs and websites on the Internet make it easy to troubleshoot issues whenever there's need for that.
Against all the tools I have used, Azure Databricks is by far the most superior of them all! Why, you ask? The UI is modern, the features are never ending and they keep adding new features. And to quote Apple, "It just works!" Far ahead of the competition, the delta lakehouse platform also fares better than it counterparts of Iceberg implementation or a loosely bound Delta Lake implementation of Synapse
No, it wasn’t my decision to use such an ETL product. I’m just the administrator at this point. I’ve heard there are other products there that are even on cloud support. That is much easier to use, more agile, and user-friendly. That doesn’t have that barrier from user to administrator to the developer standpoint.
Not directly related to ROI or cost figures. Only comment here is that IBM tools tend to be more costly than average ETL tools, but it depends on if the company is an IBM shop.
One positive aspect is the company has had not a need to switch ETL tool for years.
Upgrading to newer versions of the tool brings flexibility in the tool and up-to-date features in relation to other applications.