We

We open sourced our entire text-to-SQL product

Hacker News

We open sourced our entire text-to-SQL product

Long story short: We (Dataherald) just open-sourced our entire codebase, including the core engine, the clients that interact with it and the backend application layer for authentication and RBAC. You can now use the full solution to build text-to-SQL into your product. The Problem: modern LLMs write syntactically correct SQL, but they struggle with real-world relational data. This is because real world data and schema is messy, natural language can often be ambiguous and LLMs are not trained on your specific dataset. Solution: The core NL-to-SQL engine in Dataherald is an LLM based agent which uses Chain of Thought (CoT) reasoning and a number of different tools to generate high accuracy SQL from a given user prompt. The engine achieves this by: - Collecting context at configuration from the database and sources such as data dictionaries and unstructured documents which are stored in a data store or a vector DB and injected if relevant - Allowing users to upload sample NL <> SQL pairs (golden SQL) which can be used in few shot prompting or to fine-tune an NL-to-SQL LLM for that specific dataset - Executing the SQL against the DB to get a few sample rows and recover from errors - Using an evaluator to assign a confidence score to the generated SQL The repo includes four services https://github.com/Dataherald/dataherald/tree/main/services : 1- Engine: The core service which includes the LLM agent, vector stores and DB connectors. 2- Admin Console: a NextJS front-end for configuring the engine and observability. 3- Enterprise Backend: Wraps the core engine, adding authentication, caching, and APIs for the frontend. 4- Slackbot: Integrate Dataherald directly into your Slack workflow for on-the-fly data exploration. Would love to hear from the community on building natural language interfaces to relational data. Anyone live in production without a human in the loop? Thoughts on how to improve performance without spending weeks on model training?

Share card

Actual performance

464points
144comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agent, model, slack · Missing: mac, agents, macos
94%94% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: including · Missing: supports, reddit linkedin, podcasting
92%92% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, open source, ide · Missing: https docs, excited, just released
58%58% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: users · Missing: mobile apps, ios, personal
42%42% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: interface, users · Missing: plus, platform, intuitive
33%33% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: training · Missing: arr, mrr, revenue
15%15% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: real world · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Pr
Product Analytics in SQL with dbt57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Product Analytics in SQL with dbt

Hacker News8
I
I Solved N Queens in SQL56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I Solved N Queens in SQL

Hacker News2
Li
Lithium::SQL C++17, Header Only MySQL and SQLite Connectors and ORM76%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Lithium::SQL C++17, Header Only MySQL and SQLite Connectors and ORM

Hacker News2
Go
Go Fearless SQL67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Go Fearless SQL

Hacker News3
Mo
Mongita is to MongoDB as SQLite is to SQL77%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Mongita is to MongoDB as SQLite is to SQL

Hacker News126
SQ
SQL Basics59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SQL Basics

Hacker News26
Qu
Querying OpenAPI Definitions with SQL59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Querying OpenAPI Definitions with SQL

Hacker News1
YA
YAML to SQL57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

YAML to SQL

Hacker News1
Ba
Baldrick. Your SQL Dogsbody64%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Baldrick. Your SQL Dogsbody

Hacker News2
ca
can we get rid of SQL?67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

can we get rid of SQL?

Hacker News2