I

I made an open-source synthetic text datasets generator

Hacker News

I made an open-source synthetic text datasets generator

Many LLMs projects suffers due to the lack of custom datasets: - no labelled data at all - lack coverage and diversity in existing data - Data collection and annotation processes are slow and boring - Not enough examples to fine-tune or evaluate LLMs… So I built datafast, an open-source library for synthetic text datasets generation. Right now it supports 5 datasets types: - Text Classification Dataset - Raw Text Generation Dataset - Instruction Dataset (Ultrachat-like) - Multiple Choice Question (MCQ) Dataset - Preference Dataset And more to come. Currently supported LLM providers for generation are: - OpenAI - Anthropic - Google Gemini - Ollama (local LLM server) There is more to come but I am not in a rush for features. I seek data quality, data diversity and reliability over quantity. I don't measure success by shipping more features: I succeed if it works when you try it out, and if you actually use it. Hope you like that!

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: supports, gemini · Missing: reddit linkedin, podcasting, created
89%89% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: google, openai, gemini · Missing: mac, agents, macos
83%83% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: exist, lua, existing · Missing: https docs, excited, just released
55%55% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
44%44% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: google · Missing: mobile apps, ios, personal
43%43% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
15%15% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
1%1% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

I
I made an open-source synthetic text datasets generator55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made an open-source synthetic text datasets generator

Hacker News2
Vi
Vietnam Elections (open, source-linked datasets and site)60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Vietnam Elections (open, source-linked datasets and site)

Hacker News1
De
DeepFabric – Structured synthetic datasets for model distillation45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

DeepFabric – Structured synthetic datasets for model distillation

Hacker News2
Cu
Curator – an open-source library for synthetic data generation67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Curator – an open-source library for synthetic data generation

Hacker News13
Ho
How to evaluate the quality of synthetic datasets43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

How to evaluate the quality of synthetic datasets

Hacker News1
A
A lineage explorer for open source models and datasets60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A lineage explorer for open source models and datasets

Hacker News3
Sy
Synthetic dataset generator for NLP and tabular data44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Synthetic dataset generator for NLP and tabular data

Hacker News3
DriftData
DriftData63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Synthetic datasets for AI/ML model training

Indie Hackers1ai
La
Largest Open Source Synthetic Gun Detection Dataset and Model65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Largest Open Source Synthetic Gun Detection Dataset and Model

Hacker News3
Op
Open-source text-to-geolocation models70%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source text-to-geolocation models

Hacker News46