IBM Cloud Pak for Data (formerly IBM Cloud Private for Data) provides data management, data governance, and automated data discovery and classification.
N/A
Presto
Score 2.6 out of 10
N/A
Presto is an open source SQL query engine designed to run queries on data stored in Hadoop or in traditional databases.
Teradata supported development of Presto followed the acquisition of Hadapt and Revelytix.
Unlike others analytics tool IBM Cloud Pak for Data provides out-of-the-box privacy, model interpretability and fairness monitoring, along with automatic explanation of data and models written in business language. It's a great tool that all business should emulate. Great user experience because of every feature is functional and improved constantly.
Simple stories & templates work nicely - like for our Insider program. Stories that include a lot of images may be challenging to create & have look appealing.
Linking, embedding links and adding images is easy enough.
Once you have become familiar with the interface, Presto becomes very quick & easy to use (but, you have to practice & repeat to know what you are doing - it is not as intuitive as one would hope).
Organizing & design is fairly simple with click & drag parameters.
Presto was not designed for large fact fact joins. This is by design as presto does not leverage disk and used memory for processing which in turn makes it fast.. However, this is a tradeoff..in an ideal world, people would like to use one system for all their use cases, and presto should get exhaustive by solving this problem.
Resource allocation is not similar to YARN and presto has a priority queue based query resource allocation..so a query that takes long takes longer...this might be alleviated by giving some more control back to the user to define priority/override.
UDF Support is not available in presto. You will have to write your own functions..while this is good for performance, it comes at a huge overhead of building exclusively for presto and not being interoperable with other systems like Hive, SparkSQL etc.
IBM has healing mechanisms when resource usage is high. This platform performs well, but when it runs out of capacity, it has crashed for many clients. This is innate in its original design
I think Presto is one of the best solutions out there today at the cutting edge for interactive query analysis. One of the challenges is presto is a niche tool for the interactive query use case and doesn't have the knobs and whistles as much as Spark. In the foreseeable future if they are able to make presto work without the need for Hive, solving all the gaps it could be game changing and can be a direct threat to spark
can improve readiness for cloud migration, improve licensing flexibility with IBM, and reduce both hardware purchases and infrastructure management efforts.
reduces the expenses of internal resources.
should improve efficiencies, reduce risks, and increase performance