Ar

Arroyo – Write SQL on streaming data

Hacker News

Arroyo – Write SQL on streaming data

Hey HN, Arroyo is a modern, open-source stream processing engine, that lets anyone write complex queries on event streams just by writing SQL—windowing, aggregating, and joining events with sub-second latency. Today data processing typically happens in batch data warehouses like BigQuery and Snowflake despite the fact that most of the data is coming in as streams. Data teams have to build complex orchestration systems to handle late-arriving data and job failures while trying to minimize latency. Stream processing offers an alternative approach, where the query is compiled into a streaming program that constantly updates as new data comes in, providing low-latency results as soon as the data is available. I started the Arroyo project after spending the past five years building real-time platforms at Lyft and Splunk. I saw first hand how hard it is for users to build correct, reliable pipelines on top of existing systems like Flink and Spark Streaming, and how hard those pipelines are to operate for infra teams. I saw the need for a new system that would be easy enough for any data team to adopt, built on modern foundations and with the lessons of the past decade of research and industry development. Arroyo works by taking SQL queries and compiling them into an optimized streaming dataflow program, a distributed DAG of computation with nodes that read from sources (like Kafka), perform stateful computations, and eventually write results to sinks. That state is consistently snapshotted using a variation of the Chandy-Lamport checkpointing algorithm for fault-tolerance and to enable fast rescaling and updates of the pipelines. The entire system is easy to self-host on Kubernetes and Nomad. See it in action here: https://www.youtube.com/watch?v=X1Nv0gQy9TA or follow the getting started guide ( https://doc.arroyo.dev/getting-started ) to run it locally.

Share card

Actual performance

115points
32comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Hacker NewsStrong engagement from HN community · Strong signals: exist, existing, ide · Missing: https docs, excited, just released
91%91% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
Product HuntOn track for Day 1 leaderboard · Strong signals: user, new, using · Missing: mac, agents, macos
87%87% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Indie HackersFits the IH revenue-focused audience · Strong signals: started · Missing: supports, reddit linkedin, podcasting
73%73% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: users · Missing: mobile apps, ios, personal
49%49% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform, host, soon · Missing: plus, intuitive, reviews
43%43% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: arr · Missing: mrr, revenue, profit
19%19% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Ti
Timeplus Proton 3.0 – First vectorized streaming SQL engine86%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Timeplus Proton 3.0 – First vectorized streaming SQL engine

Hacker News11
Ne
Nextbi.ai – NLP to SQL data querry for SMEs49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Nextbi.ai – NLP to SQL data querry for SMEs

Hacker News1
Eq
Equals – a spreadsheet that you can write SQL in75%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Equals – a spreadsheet that you can write SQL in

Hacker News9
AI2sql
AI2sql21%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Write SQL queries with no knowledge of SQL

AppSumo4
I
I Solved N Queens in SQL56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I Solved N Queens in SQL

Hacker News2
Li
Lithium::SQL C++17, Header Only MySQL and SQLite Connectors and ORM76%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Lithium::SQL C++17, Header Only MySQL and SQLite Connectors and ORM

Hacker News2
Go
Go Fearless SQL67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Go Fearless SQL

Hacker News3
Mo
Mongita is to MongoDB as SQLite is to SQL77%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Mongita is to MongoDB as SQLite is to SQL

Hacker News126
SQ
SQL Basics59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SQL Basics

Hacker News26
ca
can we get rid of SQL?67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

can we get rid of SQL?

Hacker News2