DeepFabric – Structured synthetic datasets for model distillation
DeepFabric – Structured synthetic datasets for model distillation
I’ve been working on DeepFabric, an open-source CLI + SDK for generating synthetic datasets using LLMs, based on Topic or Tree Graphs (DAG). The goal is to make it easier to create structured, diverse, domain-specific datasets — especially ones with chain-of-thought (CoT) reasoning — without hand-crafting hundreds of prompts. What it does: Generates datasets via topic graphs/trees to systematically cover a domain and reduce duplication. Supports multiple CoT styles (free-text, structured, hybrid). Works with different LLM providers (OpenAI, Anthropic, local models (Ollama). Configurable via YAML, as a library , or CLI and exports easily for Hugging Face training. Why: Synthetic data is increasingly important for fine-tuning, evaluation, and distillation, aka deepseek and more recently Phi-4 Most existing approaches are ad-hoc; I wanted something systematic and reproducible. Some example data here: https://huggingface.co/datasets/lukehinds/medical_q_and_a https://huggingface.co/datasets/lukehinds/programming-challe... https://huggingface.co/datasets/lukehinds/linux_shell_attack...
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
Synthetic datasets for AI/ML model training
How to evaluate the quality of synthetic datasets
I made an open-source synthetic text datasets generator
I made an open-source synthetic text datasets generator
Cached Datasets
Cryptographic certification for synthetic datasets — prove y
Training synthetic models on highly complex datasets
Synthetic Data Genomics
I made this tool for navigating pandas datasets
BotMarket — Structured datasets AI agents can query directly