Apache Kafka is an open-source stream processing platform developed by the Apache Software Foundation written in Scala and Java. The Kafka event streaming platform is used by thousands of companies for high-performance data pipelines, streaming analytics, data integration, and mission-critical applications.
N/A
IBM StreamSets
Score 8.2 out of 10
N/A
IBM® StreamSets enables users to create and manage smart streaming data pipelines through a graphical interface, facilitating data integration across hybrid and multicloud environments. IBM StreamSets can support millions of data pipelines for analytics, applications and hybrid integration.
For brokering messages, Confluent Kafka is well suited since it offers a managed solution ready to use. Scenarios where the solution is not very well suited are for example, where pricing is an issue. The solution costs quite a lot for basic usage (for example: for 3 clusters, pricing is above 100k$ a year).
Because real-world sources often change (new fields get added, formats get tweaked, etc.), StreamSets helps detect and adapt to those "schema drifts" or changes automatically, or with minimal manual intervention. That makes pipelines more resilient and significantly reduces the maintenance burden. Therefore, data sets with constantly changing sources/formats are great for StreamSets.
Apache Kafka is able to handle a large number of I/Os (writes) using 3-4 cheap servers.
It scales very well over large workloads and can handle extreme-scale deployments (eg. Linkedin with 300 billion user events each day).
The same Kafka setup can be used as a messaging bus, storage system or a log aggregator making it easy to maintain as one system feeding multiple applications.
The Kafka Tool is a community-made Java application that looks and feels from the past century.
Logging can be confusing. This certainly shows when we have to do troubleshooting.
Hybrid scenarios - pub/sub, but there are services in and outside a Kubernetes cluster. Then there are a ~3 options, but only 2 (the harder ones) are production-safe.
Kafka has suited our use case very well so far. Going forward we are planning to expand our platform manifold so the load on Kafka and our reliance on Kafka is going to increase only.
IBM Stream sets has been a wonderful addition to our technology stack. It has helped in some of our initiatives such as data engineering, data integration for not only external customers but also for internal purposes. The tool has also helped on our use cases related to streaming data. Moving to another tool would require significant amount of work and time.
Apache Kafka is highly recommended to develop loosely coupled, real-time processing applications. Also, Apache Kafka provides property based configuration. Producer, Consumer and broker contain their own separate property file
because i think that overall the solution is having a positive impact on the business, it allows multiple benefits in simplification of the tasks and is capable of doing multiple process that are usually done by a combination of man and systems, reducing the time and effort required to have the data.
Support for Apache Kafka (if willing to pay) is available from Confluent that includes the same time that created Kafka at Linkedin so they know this software in and out. Moreover, Apache Kafka is well known and best practices documents and deployment scenarios are easily available for download. For example, from eBay, Linkedin, Uber, and NYTimes.
Streamsets support has improved a lot in the last couple of years. We had some challenges in the beginning with support, but now the quality of the support and the responsiveness to tickets are better. We have contacted support multiple times when it came to scenarios where the system was slow or the output as not as we expected
Apache Kafka is built for scale. From high throughput and real-time data streaming, it has a strong advantage over RabbitMQ with its low latency. This put Apache Kafka at the forefront as the platform of choice for large datasets messaging and ensuring scalability when data scale up tremendously. RabbitMQ however has its strengths in traditional messaging. Routing and message delivery reliability are the bedrock of RabbitMQ and this is where RabbitMQ excels. In my previous workplace, RabbitMQ was of choice as reliability matters more than scale. In two words. Apache Kafka for scale, RabbitMQ for reliability. And for cloud deployment and large dataset messaging in what I am doing now, Apache Kafka is the default choice.
Before, we were using Informatica since most of our applications were running on on-prem servers. Later, when we started moving to the cloud, we tried Informatica Cloud, but it's more useful for batch-oriented than streaming. That's why one of our tech architects suggested IBM StreamSets for our real-time data streaming. During the POC stage, we were happy that the data streaming was way better with IBM StreamSets compared to the Informatica Cloud way of doing.
Positive: bursts of traffic on special holidays are easy to handle because Kafka can absorb and buffer all the messages we need to process long enough to let an understaffed set of back-end services catch up on processing. Hard to put a number to it but we probably save $5k a month having fewer machines running.
Positive: makes decoupling the web and API services from the deeper back-end services easier by providing topics as an interface. This allowed us to split up our teams and have them develop independently of each other, speeding up software development.
Negative: our engineers have made mistakes such as accidentally dropping a few thousand messages due to the CLI being confusing to use, and as a result a customer lost some of their precious data. I'd say that was more our fault than Kafka's though.