En

Engraph – Automated ETL Pipelines

Hacker News

Engraph – Automated ETL Pipelines

Hey HN, we’re Ross and Javier, co-founders of Engraph (www.engraph.ai). Our goal is to completely automate the process of building ETL pipelines, from ad hoc pipelines for question answering to fully fledged ETL pipelines within large organisations: For ad hoc pipelines, a question answering platform which enables users to ask questions in natural language about their organisation's data. Traditionally, access to data within organisations is limited to a handful of data-engineers. This means that if an employee needs access to some data, they have to go through a lengthy process of requesting it from a data-engineer, who then spends their day dealing with ad hoc requests. This process is time-consuming and inefficient for everyone involved. We solve this problem by passing natural language questions into a planner LLM that decomposes the task given the available data sources. This planner spawns workers that query the appropriate information from each individual data source, whether via SQL or searches across vector embeddings of unstructured information. Once the planner receives the data, it executes Python code to aggregate and operate on the data, and presents the answer to the user. For persistent ETL pipelines, instead of inferring the format of the output data from a natural language question users can provide a data output specification (e.g. YAML) through an API. This gets fed into a planner similar to the above. However, instead of running Python code to return an answer, we load the relevant data into an intermediate data lake. Internally (depending on the user’s preferences), we also use this API for the question answering platform: we learn from recurring patterns in the natural language questions (by clustering question embeddings), and implement persistent data storage that, in expectation, reduces the number of operations that need to be performed across the org’s data sources. For the user there is a trade-off between the cost of the intermediate data storage vs the cost/load of performing the same operations for every question. Surely you’ve been wondering about how we manage security and data access. Our goal is to never have to access company data on our side. Ideally all operations are performed on-prem. We take data privacy very seriously and we are building our platform accordingly. We charge on a per seat licence with lenient fair usage terms. We’d rather do this than charge per usage which we feel could disincentivize the use of the platform. If you're interested in trying out our platform, check out our demo video (https://youtu.be/Q8dNPQ8ofHk) or send us an email at {javier, ross}@engraph.ai. Thanks!

Share card

Actual performance

5points
2comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: user, email, code · Missing: mac, agents, macos
86%86% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
81%81% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: ide, pipe, io · Missing: https docs, excited, just released
73%73% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRFits verified-revenue profile · Strong signals: video, users · Missing: mobile apps, ios, personal
60%60% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform, efficient, users · Missing: plus, intuitive, reviews
27%27% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: recurring · Missing: arr, mrr, revenue
15%15% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

ET
ETL Data Pipelines to Google BigQuery62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ETL Data Pipelines to Google BigQuery

Hacker News5
Ru
Run ETL Pipelines and Store Data in GitHub39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Run ETL Pipelines and Store Data in GitHub

Hacker News1
Pi
Pipelines50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Pipelines

Hacker News1
Pi
Pipelines50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Pipelines

Hacker News5
Au
Automated Pipelines to Your Kubernetes Clusters56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Automated Pipelines to Your Kubernetes Clusters

Hacker News81
Sn
Snapflow – Functional Data ETL60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Snapflow – Functional Data ETL

Hacker News4
Gr
Graphical ETL Based on JSONata54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Graphical ETL Based on JSONata

Hacker News3
Sh
Shadergarden: Create reloadable graphical pipelines with Lisp and GLSL60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Shadergarden: Create reloadable graphical pipelines with Lisp and GLSL

Hacker News47
St
Standardizing NLP for a Modern ETL67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Standardizing NLP for a Modern ETL

Hacker News6
Qu
Quickly build and deploy streaming, batch, CDC, ETL and ML pipelines43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Quickly build and deploy streaming, batch, CDC, ETL and ML pipelines

Hacker News1