Apache Spark vs. Informatica PowerCenter (legacy)

Overview
ProductRatingMost Used ByProduct SummaryStarting Price
Apache Spark
Score 9.2 out of 10
N/A
Apache Spark is a multi-language engine for executing data engineering, data science, and machine learning on single-node machines or clusters.N/A
Informatica PowerCenter (legacy)
Score 7.9 out of 10
N/A
Informatica PowerCenter was data integration technology designed to form the foundation for data integration initiatives, application migration, or analytics. It is a legacy product.N/A
Pricing
Apache SparkInformatica PowerCenter (legacy)
Editions & Modules
No answers on this topic
No answers on this topic
Offerings
Pricing Offerings
Apache SparkInformatica PowerCenter (legacy)
Free Trial
NoNo
Free/Freemium Version
NoNo
Premium Consulting/Integration Services
NoNo
Entry-level Setup FeeNo setup feeNo setup fee
Additional Details
More Pricing Information
Community Pulse
Apache SparkInformatica PowerCenter (legacy)
Features
Apache SparkInformatica PowerCenter (legacy)
Data Source Connection
Comparison of Data Source Connection features of Product A and Product B
Apache Spark
-
Ratings
Informatica PowerCenter (legacy)
8.5
Ratings
1% above category average
Connect to traditional data sources00 Ratings9.00 Ratings
Connecto to Big Data and NoSQL00 Ratings8.00 Ratings
Data Transformations
Comparison of Data Transformations features of Product A and Product B
Apache Spark
-
Ratings
Informatica PowerCenter (legacy)
7.5
Ratings
8% below category average
Simple transformations00 Ratings8.00 Ratings
Complex transformations00 Ratings7.00 Ratings
Data Modeling
Comparison of Data Modeling features of Product A and Product B
Apache Spark
-
Ratings
Informatica PowerCenter (legacy)
8.2
Ratings
3% above category average
Data model creation00 Ratings9.00 Ratings
Metadata management00 Ratings8.00 Ratings
Business rules and workflow00 Ratings9.00 Ratings
Collaboration00 Ratings6.10 Ratings
Testing and debugging00 Ratings9.00 Ratings
Data Governance
Comparison of Data Governance features of Product A and Product B
Apache Spark
-
Ratings
Informatica PowerCenter (legacy)
9.0
Ratings
10% above category average
Integration with data quality tools00 Ratings9.00 Ratings
Integration with MDM tools00 Ratings9.00 Ratings
Best Alternatives
Apache SparkInformatica PowerCenter (legacy)
Small Businesses

No answers on this topic

Skyvia
Skyvia
Score 9.9 out of 10
Medium-sized Companies
Cloudera Manager
Cloudera Manager
Score 9.9 out of 10
IBM InfoSphere Information Server
IBM InfoSphere Information Server
Score 8.0 out of 10
Enterprises
IBM Analytics Engine
IBM Analytics Engine
Score 7.1 out of 10
IBM InfoSphere Information Server
IBM InfoSphere Information Server
Score 8.0 out of 10
All AlternativesView all alternativesView all alternatives
User Ratings
Apache SparkInformatica PowerCenter (legacy)
Likelihood to Recommend
9.0
(0 ratings)
8.0
(0 ratings)
Likelihood to Renew
10.0
(0 ratings)
10.0
(0 ratings)
Usability
8.0
(0 ratings)
9.0
(0 ratings)
Performance
-
(0 ratings)
9.4
(0 ratings)
Support Rating
8.7
(0 ratings)
9.0
(0 ratings)
User Testimonials
Apache SparkInformatica PowerCenter (legacy)
Likelihood to Recommend
Apache Spark has rich APIs for regular data transformations or for ML workloads or for graph workloads, whereas other systems may not such a wide range of support. Choose it when you need to perform data transformations for big data as offline jobs, whereas use MongoDB-like distributed database systems for more realtime queries.
Read full review
Informatica Powercenter is the centerpiece of our overall enterprise data warehouse strategy. It's a critical enablement to ensure we can feed in multiple data stream and transform them into digestible data within our data warehouse. With its flexible capabilities and API availability, we were able to feed in industry standard data format as well as home grown data structure. Overall, we are very pleased with their capability and contribution to our data warehouse strategy.
Read full review
Pros
  • It performs a conventional disk-based process when the data sets are too large to fit into memory, which is very useful because, regardless of the size of the data, it is always possible to store them.
  • It has great speed and ability to join multiple types of databases and run different types of analysis applications. This functionality is super useful as it reduces work times
  • Apache Spark uses the data storage model of Hadoop and can be integrated with other big data frameworks such as HBase, MongoDB, and Cassandra. This is very useful because it is compatible with multiple frameworks that the company has, and thus allows us to unify all the processes.
Read full review
  • Informatica has a wide range of support for databases. Pretty much every mainstream DBMS is compatible here.
  • Designing ETL mappings and workflows is a very intuitive process, and takes minimal learning time and effort even for a beginner.
  • Informatica's biggest strength is its sheer performance. It is unmatched in terms of handling large volumes of data.
Read full review
Cons
  • Memory management. Very weak on that.
  • PySpark not as robust as scala with spark.
  • spark master HA is needed. Not as HA as it should be.
  • Locality should not be a necessity, but does help improvement. But would prefer no locality
Read full review
  • One of the challenges of PowerCenter is the lack of integration between the components and functionality provided by PowerCenter. PowerCenter consists of multiple components such has the repository service, integration service, metadata service. Considerable time and resources were required to install and configure these components before PowerCenter was available for use.
  • In order to connect to various data sources such as Netezza database or SAS datasets, PowerCenter requires the installation and configuration of separate plug-ins. We spent considerable time trouble-shooting and debugging problems while trying to get the various plug-ins integrated with PowerCenter and get them up and running as described in the documentation.
  • PowerCenter works well with structured data. That is, it is easy to work with input and output data that is pre-defined, fixed, and unchanging. It is much more difficult to work with dynamic data in which new fields are added or removed ad-hoc or if data format changes during the data ingest process. We have not been as successful in using PowerCenter for dynamic data.
  • One of the challenges of learning PowerCenter is that it is difficult to find documentation or publications that help you learn the various details about PowerCenter software. Unlike SAS Institute, Informatica does not publish books about PowerCenter. The documentation available with PowerCenter is sparse; we have learned many aspects of this technology through trial and error.
Read full review
Likelihood to Renew
Capacity of computing data in cluster and fast speed.
Read full review
Our team enjoys using Informatica and feels that it is one of the best ETL tools on the market.
Read full review
Usability
If the team looking to use Apache Spark is not used to debug and tweak settings for jobs to ensure maximum optimizations, it can be frustrating. However, the documentation and the support of the community on the internet can help resolve most issues. Moreover, it is highly configurable and it integrates with different tools (eg: it can be used by dbt core), which increase the scenarios where it can be used
Read full review
The tool is very flexible and will meet most, if not all, of your data transformation needs. It is an expert-level tool, so building your knowledge-base and user-base (and keeping that base healthy!) is very important. But it will pay off with strong data management and the ability to leverage that data in ways you haven’t thought of yet. Bottom line, data is money, and PowerCenter helps you monetize your data.
Read full review
Performance
No answers on this topic
Positives; - Multi-user development environment. - The speed of transformation. - Seamless integration with other Informatica products. Negatives; - There should be fewer windows, to maintain developers' focus while using. You probably need two big monitors when you start development with Informatica Power Center. - Oracle Analytical functions should be natively used. - E-LT support as well as ETL support.
Read full review
Support Rating
1. It integrates very well with scala or python. 2. It's very easy to understand SQL interoperability. 3. Apache is way faster than the other competitive technologies. 4. The support from the Apache community is very huge for Spark. 5. Execution times are faster as compared to others. 6. There are a large number of forums available for Apache Spark. 7. The code availability for Apache Spark is simpler and easy to gain access to. 8. Many organizations use Apache Spark, so many solutions are available for existing applications.
Read full review
Informatica power center is a leader of the pack of ETL tools and has some great abilities that make it stand out from other ETL tools. It has been a great partner to its clients over a long time so it's definitely dependable. With all the great things about Informatica, it has a bit of tech burden that should be addressed to make it more nimble, reduce the learning curve for new developers, provide better connectivity with visualization tools.
Read full review
Alternatives Considered
We used Surprise Kit for one of the other research works. It is more fine-tuned to Recommendation systems and their algorithms. Apache Spark has MLlib for majority of ML problems. Where as software like Surprse Kit - it suitable for a specific task of Recommendations only
Read full review
Basically the two solutions have, more or less, the same functions and features.The difference, for me, is that ThreatQuotient make more features over the security and I think is oriented to a SOC enviroments. InformaticaExchange Connectors is oriented to the quality, integration and distribution of the data in order to ensure the reliability and access of data from different sources, as well as the integration in a single repository of enterprise data (External/internal)
Read full review
Return on Investment
  • Faster turn around on feature development, we have seen a noticeable improvement in our agile development since using Spark.
  • Easy adoption, having multiple departments use the same underlying technology even if the use cases are very different allows for more commonality amongst applications which definitely makes the operations team happy.
  • Performance, we have been able to make some applications run over 20x faster since switching to Spark. This has saved us time, headaches, and operating costs.
Read full review
  • PowerCenter has been instrumental in being the center of all data movement within the organization.
  • It has also provided a foundation for which re-usability and scalability are the focus.
  • Finding talent with experience and expertise in PowerCenter is far more likely due to its presence and market share.
Read full review
ScreenShots