De

DeepFabric – Structured synthetic datasets for model distillation

Hacker News

DeepFabric – Structured synthetic datasets for model distillation

I’ve been working on DeepFabric, an open-source CLI + SDK for generating synthetic datasets using LLMs, based on Topic or Tree Graphs (DAG). The goal is to make it easier to create structured, diverse, domain-specific datasets — especially ones with chain-of-thought (CoT) reasoning — without hand-crafting hundreds of prompts. What it does: Generates datasets via topic graphs/trees to systematically cover a domain and reduce duplication. Supports multiple CoT styles (free-text, structured, hybrid). Works with different LLM providers (OpenAI, Anthropic, local models (Ollama). Configurable via YAML, as a library , or CLI and exports easily for Hugging Face training. Why: Synthetic data is increasingly important for fine-tuning, evaluation, and distillation, aka deepseek and more recently Phi-4 Most existing approaches are ad-hoc; I wanted something systematic and reproducible. Some example data here: https://huggingface.co/datasets/lukehinds/medical_q_and_a https://huggingface.co/datasets/lukehinds/programming-challe... https://huggingface.co/datasets/lukehinds/linux_shell_attack...

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, models, openai · Missing: mac, agents, macos
82%82% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: supports · Missing: reddit linkedin, podcasting, created
80%80% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRFits verified-revenue profile · Missing: mobile apps, ios, personal
51%51% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: exist, lua, existing · Missing: https docs, excited, just released
46%46% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
33%33% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: training · Missing: arr, mrr, revenue
15%15% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

DriftData
DriftData63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Synthetic datasets for AI/ML model training

Indie Hackers1ai
Ho
How to evaluate the quality of synthetic datasets43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

How to evaluate the quality of synthetic datasets

Hacker News1
I
I made an open-source synthetic text datasets generator55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made an open-source synthetic text datasets generator

Hacker News2
I
I made an open-source synthetic text datasets generator55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made an open-source synthetic text datasets generator

Hacker News2
Ca
Cached Datasets53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Cached Datasets

Hacker News4
CertifiedData.io
CertifiedData.io38%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Cryptographic certification for synthetic datasets — prove y

Indie Hackerscommitment-full-time
Tr
Training synthetic models on highly complex datasets49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Training synthetic models on highly complex datasets

Hacker News10
Sy
Synthetic Data Genomics44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Synthetic Data Genomics

Hacker News1
I
I made this tool for navigating pandas datasets50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made this tool for navigating pandas datasets

Hacker News20
Bo
BotMarket — Structured datasets AI agents can query directly44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

BotMarket — Structured datasets AI agents can query directly

Hacker News3