AWS Glue is a managed extract, transform, and load (ETL) service designed to make it easy for customers to prepare and load data for analytics. With it, users can create and run an ETL job in the AWS Management Console. Users point AWS Glue to data stored on AWS, and AWS Glue discovers data and stores the associated metadata (e.g. table definition and schema) in the AWS Glue Data Catalog. Once cataloged, data is immediately searchable, queryable, and available for ETL.
$0.44
billed per second, 1 minute minimum
Informatica Intelligent Data Management Cloud
Score 7.0 out of 10
N/A
The Informatica® Intelligent Data Management Cloud™ (IDMC) is designed to help businesses efficiently handle the complex challenges of dispersed and fragmented data to innovate with their data on virtually any platform, any cloud, multi-cloud and multi-hybrid.
When the data which requires ETL has different formats, schema, and volume, this service suits them best. So, when the volume is not consistent (typical use-case of healthcare and online shopping), AWS Glue can be the prime choice. When the data is available in both batch and streaming mode, the developer needs to generate a separate codebase. This increases the source code management efforts. So, prefer to go with Glue when the nature of the data is the same (either batched or streamed).
Informatica Cloud is an excellent tool for freeing up trapped data in your internal systems and exporting it into a cloud data warehouse or SQL container. Especially if you have any internal SQL query generators that allow you to copy and paste into the Source setup. For example, we use Business Objects internally, and you can build out the query you need inside the firewall with that, then paste it into the query within Informatica Cloud while establishing the Source connection. This way you are not writing all your source parameters by hand, cutting new jobs setup time by 50%. Where Cloud will not suit you well is if you need significant visibility to the health of the system from a tracking, monitoring or error handling standpoint. Cloud will tell you it failed, maybe give you an example error, but has no further troubleshooting or diagnostic assistance to offer.
After data cleansing, the team also implemented the best practices for using AWS platform services as a Data Lake, such as job bookmarking for AWS Glue jobs, proper delimiter for the AWS Glue crawlers, partitioning in AWS S3, and transformation to parquet file for compression and faster querying time in Amazon Athena.
Data modernization through combining data from multiple sources into a functioning datasets, rebuilding DW, and resctructuring data sources.
Aims to lessen customer complaints, eliminate manual data extraction requests via SR from different data sources, and Increase accuracy, consistency and speed up reconciliation process.
Informatica Cloud is a great tool for automating data imports. Raw data files can be scheduled to run at designated times, and using pre-built mappings ensures that the data goes into the system accurately each time.
It is also a great tool for converting raw data into specific formatting. For example, configuring the mapping so that the first letter of a first, middle or last name is always capitalized.
Additionally, Informatica Cloud has a built-in field-matching tool so that records that already exist are updated rather than having duplicates created.
Amazon responds in good time once the ticket has been generated but needs to generate tickets frequent because very few sample codes are available, and it's not cover all the scenarios.
I've never had trouble getting into contact with Informatica's support for technical help. I give it a nine because it does pretty well for mid to enterprise-scale workflows.
The cataloging of data objects is the best in the case of AWS Glue. We use AWS Glue in all of our data pipelines to sync external and internal data sources and to automatically produce SQL-based ETL based on AWS Glue catalog objects. Integration with Amazon products is the other advantage.
Having used SQL Server Integration Services (SSIS) in the past, Informatica Cloud was a huge-step up in functionality, usability and performance. Hands-down, Informatica Cloud is a much more robust product overall. While SSIS is well-suited for specific instances (i.e. SQL Server-specific implementations), Informatica Cloud is a much better product as an overall integration/transformation tool.
Getting data out of various systems such as Salesforce.com or FTP and consolidating it into tables in one central database.
Having the ability to run data transformation tasks (the equivalent of batch jobs in SF or stored procedures in SQL) on a schedule without consuming those resources on the target or source systems.