I

I built an open-source data copy tool called ingestr

Hacker News

I built an open-source data copy tool called ingestr

Hi there, Burak here. I built an open-source data copy tool called ingestr ( https://github.com/bruin-data/ingestr ) I did build quite a few data warehouses both for the companies I worked at, as well as for consultancy projects. One of the more common pain points I observed was that everyone had to rebuild the same data ingestion bit over and over again, and each in different ways: - some wrote code for the ingestion from scratch to various degrees - some used off-the-shelf data ingestion tools like Fivetran / Airbyte I have always disliked both of these approaches, for different reasons, but never got around to working on what I'd imagine to be the better way forward. The solutions that required writing code for copying the data had quite a bit of overhead such as how to generalize them, what language/library to use, where to deploy, how to monitor, how to schedule, etc. I ended up figuring out solutions for each of these matters, but the process always felt suboptimal. I like coding but for more novel stuff rather than trying to copy a table from Postgres to BigQuery. There are libraries like dlt (awesome lib btw, and awesome folks!) but that still required me to write, deploy, and maintain the code. Then there are solutions like Fivetran or Airbyte, where there's a UI and everything is managed through there. While it is nice that I didn't have to write code for copying the data, I still had to either pay some unknown/hard-to-predict amount of money to these vendors or host Airbyte myself which is roughly back to square zero (for me, since I want to maintain the least amount of tech myself). Nothing was versioned, people were changing things in the UI and breaking the connectors, and what worked yesterday didn't work today. I had a bit of spare time a couple of weeks ago and I wanted to take a stab at the problem. I have been thinking of standardizing the process for quite some time already, and dlt had some abstractions that allowed me to quickly prototype a CLI that copies data from one place to another. I made a few decisions (that I hope I won't regret in the future): - everything is a URI: every source and every destination is represented as a URI - there can be only one thing copied at a time: it'll copy only a single table within a single command, not a full database with an unknown amount of tables - incremental loading is a must, but doesn't have to be super flexible: I decided to support full-refresh, append-only, merge, and delete+insert incremental strategies, because I believe this covers 90% of the use-cases out there. - it is CLI-only, and can be configured with flags & env variables so that it can be automated quickly, e.g. drop it into GitHub Actions and run it daily. The result ended up being `ingestr` ( https://github.com/bruin-data/ingestr ). I am pretty happy with how the first version turned out, and I plan to add support for more sources & destinations. ingestr is built to be flexible with various source and destination combinations, and I plan to introduce more non-DB sources such as Notion, GSheets, and custom APIs that return JSON (which I am not sure how exactly I'll do but open to suggestions!). To be perfectly clear: I don't think ingestr covers 100% of data ingestion/copying needs out there, and it doesn't aim that. My goal with it is to cover most scenarios with a decent set of trade-offs so that common scenarios can be solved easily without having to write code or manage infra. There will be more complex needs that require engineering effort by others, and that's fine. I'd love to hear your feedback on how can ingestr help data copying needs better, looking forward to hearing your thoughts! Best, Burak

Share card

Actual performance

156points
48comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: ios · Missing: supports, reddit linkedin, podcasting
95%95% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: single, coding, code · Missing: mac, agents, macos
86%86% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: ide, io · Missing: https docs, excited, just released
63%63% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: host · Missing: plus, platform, intuitive
32%32% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: ios, way · Missing: mobile apps, personal, entrepreneurs
32%32% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
12%12% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: introduce · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Op
Open-source tool to find PII across data silos64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source tool to find PII across data silos

Hacker News3
De
Desbordante – an open-source data profiling tool63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Desbordante – an open-source data profiling tool

Hacker News2
I
I made an open source tool for transfering data63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made an open source tool for transfering data

Hacker News2
Op
Open-source Data Anonymization tool nxs-data-anonymizer62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source Data Anonymization tool nxs-data-anonymizer

Hacker News4
I
I built an open-source data pipeline tool in Go64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I built an open-source data pipeline tool in Go

Hacker News200
I
I built an open-source tool to make on-call suck less54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I built an open-source tool to make on-call suck less

Hacker News319
Si
SirixDB – Storing and Querying of Temporal Data (Java and Open Source)62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SirixDB – Storing and Querying of Temporal Data (Java and Open Source)

Hacker News13
St
Streamdal – an open-source tail -f for your data81%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Streamdal – an open-source tail -f for your data

Hacker News148
Op
Open-Source Data Replication and Anonymization73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-Source Data Replication and Anonymization

Hacker News24
Ne
Neosync – Open Source Data Replication and Anonymization73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Neosync – Open Source Data Replication and Anonymization

Hacker News4