I

I built an open-source data pipeline tool in Go

Hacker News

I built an open-source data pipeline tool in Go

Every data pipeline job I had to tackle required quite a few components to set up: - One tool to ingest data - Another one to transform it - If you wanted to run Python, set up an orchestrator - If you need to check the data, a data quality tool Let alone this being hard to set up and taking time, it is also pretty high-maintenance. I had to do a lot of infra work, and while this being billable hours for me I didn’t enjoy the work at all. For some parts of it, there were nice solutions like dbt, but in the end for an end-to-end workflow, it didn’t work. That’s why I decided to build an end-to-end solution that could take care of data ingestion, transformation, and Python stuff. Initially, it was just for our own usage, but in the end, we thought this could be a useful tool for everyone. In its core, Bruin is a data framework that consists of a CLI application written in Golang, and a VS Code extension that supports it with a local UI. Bruin supports quite a few stuff: - Data ingestion using ingestr ( https://github.com/bruin-data/ingestr ) - Data transformation in SQL & Python, similar to dbt - Python env management using uv - Built-in data quality checks - Secrets management - Query validation & SQL parsing - Built-in templates for common scenarios, e.g. Shopify, Notion, Gorgias, BigQuery, etc This means that you can write end-to-end pipelines within the same framework and get it running with a single command. You can run it on your own computer, on GitHub Actions, or in an EC2 instance somewhere. Using the templates, you can also have ready-to-go pipelines with modeled data for your data warehouse in seconds. It includes an open-source VS Code extension as well, which allows working with the data pipelines locally, in a more visual way. The resulting changes are all in code, which means everything is version-controlled regardless, it just adds a nice layer. Bruin can run SQL, Python, and data ingestion workflows, as well as quality checks. For Python stuff, we use the awesome (and it really is awesome!) uv under the hood, install dependencies in an isolated environment, and install and manage the Python versions locally, all in a cross-platform way. Then in order to manage data uploads to the data warehouse, it uses dlt under the hood to upload the data to the destination. It also uses Arrow’s memory-mapped files to easily access the data between the processes before uploading them to the destination. We went with Golang because of its speed and strong concurrency primitives, but more importantly, I knew Go better than the other languages available to me and I enjoy writing Go, so there’s also that. We had a small pool of beta testers for quite some time and I am really excited to launch Bruin CLI to the rest of the world and get feedback from you all. I know it is not often to build data tooling in Go but I believe we found ourselves in a nice spot in terms of features, speed, and stability. https://github.com/bruin-data/bruin I’d love to hear your feedback and learn more about how we can make data pipelines easier and better to work with, looking forward to your thoughts! Best, Burak

Share card

Actual performance

200points
48comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: supports, ios · Missing: reddit linkedin, podcasting, created
91%91% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, computer, new · Missing: mac, agents, macos
82%82% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: excited, ide, pipe · Missing: https docs, just released, exist
61%61% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: ios, way · Missing: mobile apps, personal, entrepreneurs
41%41% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform · Missing: plus, intuitive, reviews
33%33% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: arr, shopify · Missing: mrr, revenue, profit
11%11% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

I
I made an Open Source Personal Data Pipeline74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made an Open Source Personal Data Pipeline

Hacker News5
Op
Open-source tool to find PII across data silos64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source tool to find PII across data silos

Hacker News3
De
Desbordante – an open-source data profiling tool63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Desbordante – an open-source data profiling tool

Hacker News2
I
I made an open source tool for transfering data63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made an open source tool for transfering data

Hacker News2
Op
Open-source Data Anonymization tool nxs-data-anonymizer62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source Data Anonymization tool nxs-data-anonymizer

Hacker News4
I
I built an open-source data copy tool called ingestr67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I built an open-source data copy tool called ingestr

Hacker News156
I
I built an open-source tool to make on-call suck less54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I built an open-source tool to make on-call suck less

Hacker News319
Si
SirixDB – Storing and Querying of Temporal Data (Java and Open Source)62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SirixDB – Storing and Querying of Temporal Data (Java and Open Source)

Hacker News13
St
Streamdal – an open-source tail -f for your data81%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Streamdal – an open-source tail -f for your data

Hacker News148
Ne
Neosync – Open Source Data Replication and Anonymization73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Neosync – Open Source Data Replication and Anonymization

Hacker News4