Starburst Announces 100GB/second Streaming Ingest from Apache Kafka to Apache Iceberg Tables
Go from data ingestion to blazing-fast SQL analytics in near real-time with the Starburst Open Hybrid Lakehouse
Businesses that require data to be available for analytics in their cloud data lake with minimal delay traditionally build complex ingestion systems that require cobbling together multiple tools and writing custom software to stream data into cloud data lakes. Alternatively, these organizations may rely on incomplete solutions that only handle the ingestion process. Both approaches tend to be fragile, difficult to scale, costly to maintain, and solve only part of the problem. After the data lands in the lake, it still needs to be transformed and optimized for efficient querying—requiring even more code, pipelines, tools, and added complexity. In addition, the pressure for cost optimization across analytics functions is increasing. CIOs are looking for ways to improve their operational overhead against traditional lakehouses and legacy data warehouses while maintaining control of their data and analytics stack.
"As businesses strive to perform analytics on real-time data, they seek frictionless solutions for continuous data ingestion. They also prioritize open standards like Apache Iceberg to future-proof their environments amid rapidly evolving technologies. Furthermore, reducing complexity and simplifying architectures is critical, helping organizations optimize IT investments and avoid unnecessary costs associated with integrating disparate systems," said
Streaming Ingest from Kafka (general availability) - Starburst now enables the easy creation of fully managed ingestion pipelines for Kafka topics at a verified scale up to 100GB/second, at half the cost of alternative solutions. Configuration is completed in minutes and simply entails selecting the Kafka topic, the auto-generated table schema, and the location of the resulting Iceberg table.
- Starburst Galaxy's streaming ingestion is serverless and does the heavy lifting without any manual configuration, tuning, or additional tools required by the customer. Galaxy automatically ingests incoming messages from Kafka topics into managed Iceberg tables in S3, compacts and transforms the data, applies the necessary governance, and makes it available to query within about one minute.
- Starburst's streaming ingestion can connect to Kafka-compliant systems, which includes Confluent Cloud, Amazon Managed Streaming for Apache Kafka (MSK), and Apache Kafka.
- Starburst guarantees exactly once delivery, ensuring no duplicate messages are read, and no messages are missed to ensure accuracy.
- It is built for a massive scale and has been tested to ingest 100 gigabytes of streaming data per second.
Ingest from Files landing in S3 (public preview) - Additionally, Starburst is expanding its ingestion capabilities by introducing file loading, offering customers a powerful, automated alternative to DIY or off-the-shelf solutions. This feature reads, parses, and writes records from files directly into Iceberg tables, which leverage the new ingestion capabilities to automatically optimize the tables for read performance through capabilities like compaction, snapshot retention, orphaned file removal, and statistics collection. The public preview of file loading will be available in
Enhanced Auto Scaling (general availability) - Starburst makes auto scaling smarter in Starburst Galaxy. In environments with high concurrent users, demand for compute resources can fluctuate dynamically. The enhanced Auto Scaling intelligently monitors both active and pending queries to understand and allocate how much compute resources are needed per query up to 50% faster. Not only does enhanced Auto Scaling provision additional compute resources faster, but it also includes the ability to automatically reactivate draining worker nodes, improving the efficiency of resource utilization.
Next Gen Caching (private preview) - Data engineers undertake various labor-intensive data preparation tasks.
User Role Based Routing (private preview) - Previously, users would spend too much effort determining which queries were appropriate for different cluster types. Also, administrators weren't able to assign groups of users to a cluster via roles and privileges. With User Role Based Routing, Starburst now supports the easy allocation of resources by cluster type. Customers can programmatically route queries to the appropriate Galaxy cluster based on a predefined set of rules. Users can send all queries to a single URL, which will route the queries based on the user's role, minimizing human intervention while improving what is already industry-leading price-performance against other leading cloud data warehouses and lakehouses.
"With our new ingestion capabilities to Iceberg, customers don't have to worry about how fast or how much data they need to land in their data lake. At 100GB/second, Galaxy's ingestion can handle the scale of the most demanding use cases. Because it is so easy to configure and cost-effective to operate, customers don't have to artificially limit the number of up-to-date, fresh tables in their lake, enabling them to make the most informed business decisions," said
Supporting Resources
For more information, read Starburst's Icehouse launch blog.
Download an image of the Starburst Open Data Lakehouse here.
About Starburst
Starburst, the Open Hybrid Lakehouse, is the leading end-to-end data platform to securely access, analyze, and share data for analytics and AI across hybrid, on-premises, and multi-cloud environments. As the leaders in Trino, a modern open-source SQL engine, Starburst empowers the most data-intensive and security-conscious organizations like Comcast, Halliburton, Vectra, EMIS Health, and 7 of the top 10 global banks to democratize data access, enhance analytics performance, and improve architecture optionality. With the Open Hybrid Lakehouse from Starburst, enterprises globally can easily discover and use all their relevant business data to power new applications and analytics across risk mitigation, supply chain, customer experiences, product optimization, streaming, and more.
For additional information, please visit https://www.starburst.io/
View original content to download multimedia:https://www.prnewswire.com/news-releases/starburst-announces-100gbsecond-streaming-ingest-from-apache-kafka-to-apache-iceberg-tables-302285065.html
SOURCE Starburst
Serious News for Serious Traders! Try StreetInsider.com Premium Free!
You May Also Be Interested In
- Boo Bash Fall Festival to Make California Debut September 18, Bringing 100,000+ Square-Foot Immersive Pumpkin Patch to Escondido
- abrdn Asia-Pacific Income Fund, Inc. (FAX) Announces New Managed Distribution Policy and Declares Monthly Distribution
- Scholar Rock Announces FDA Approval of ISEMBYLD™ (apitegromab-mstn), the First and Only Muscle-Targeted Treatment for Children and Adults with Spinal Muscular Atrophy (SMA)
Create E-mail Alert Related Categories
PRNewswire, Press ReleasesRelated Entities
S3Sign up for StreetInsider Free!
Receive full access to all new and archived articles, unlimited portfolio tracking, e-mail alerts, custom newswires and RSS feeds - and more!



Tweet
Share