Pi

PipeRider – open-source Data Impact Analysis for dbt changes

Hacker News

PipeRider – open-source Data Impact Analysis for dbt changes

Hi HN! This is CL and we’re building PipeRider[0]. PipeRider is an open source data impact analysis tool, specifically during pull requests for dbt. Why? In a previous life I worked on distributed version control systems[1] prior to git, weird data systems like in-postgres REST server with plv8, and building civic tech communities with open data. It always startled me when some new characteristics of data were uncovered, and we had to change some data schema & modeling and then recheck all downstream uses of the data (if we even could). Fast forward to the modern era of data systems, the data engineering stack is certainly becoming more like the software engineering stack, with dbt as of one of the main game changers. But the engineering practices and tooling~~s~~ aren’t quite there, yet. We are building the missing pieces of the “data-pipeline-as-code” puzzle, by making pull-requests on data systems more informative and bringing confidence to reviewing and testing the impact of code-change. Here’s an example PR of the campaign finance project[2]. The goal is to augment the CI process to help teams move faster when making data and pipeline logic changes, by being aware of intentional and unintentional downstream models and metrics impact. As a version control nerd, I personally love semantic-aware diffs that can help inform stakeholders, such as diff formatted law amendment proposals in congressional bills. So, what can we do for complex data systems that are now pull-request-able? One thing is the "lineage diff" - A way to visualize the data models you are adding or making changes to, and help you to make sense of the impact. Here's another example based on a pull request in the Danish Parliament data[3] dbt project[4] (Note the lineage diff is part of advanced impact analysis which is not open source) You can try out Lineage Diff in this online viewer[5], by uploading two manifests from your dbt project to see the impact on lineage from your code-changes. I’d love your feedback! [0] https://piperider.io/ [1] https://news.ycombinator.com/item?id=32668334 [2] https://github.com/g0v/tw_campaign_finance/pull/2 [3] https://github.com/bgarcevic/danish-democracy-data/pull/8 [4] https://cloud.piperider.io/clkao/default/runs/33a74e0bef5f48... [5] https://cloud.piperider.io/online-viewer (Edit: link formatting)

Share card

Actual performance

25points
3comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
86%86% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, new, models · Missing: mac, agents, macos
85%85% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: open source, ide, pipe · Missing: https docs, excited, just released
82%82% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: personal, visualize, way · Missing: mobile apps, ios, entrepreneurs
42%42% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
33%33% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
10%10% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Si
SirixDB – Storing and Querying of Temporal Data (Java and Open Source)62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SirixDB – Storing and Querying of Temporal Data (Java and Open Source)

Hacker News13
St
Streamdal – an open-source tail -f for your data81%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Streamdal – an open-source tail -f for your data

Hacker News148
Op
Open-Source Data Replication and Anonymization73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-Source Data Replication and Anonymization

Hacker News24
Ne
Neosync – Open Source Data Replication and Anonymization73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Neosync – Open Source Data Replication and Anonymization

Hacker News4
Li
Lightly Insights – open-source dataset analysis50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Lightly Insights – open-source dataset analysis

Hacker News2
Co
CodeClarity – an open source source code analysis platform48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

CodeClarity – an open source source code analysis platform

Hacker News4
Mu
Multiwoven – Open-Source Data Activation Platform – Ruby on Rails76%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Multiwoven – Open-Source Data Activation Platform – Ruby on Rails

Hacker News1
Op
Open-source Data Anonymization tool nxs-data-anonymizer62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source Data Anonymization tool nxs-data-anonymizer

Hacker News4
Op
Open-source tool to find PII across data silos64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source tool to find PII across data silos

Hacker News3
De
Desbordante – an open-source data profiling tool63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Desbordante – an open-source data profiling tool

Hacker News2