Sy

SyGra – Graph-oriented Synthetic data generation Pipeline for LLMs

Hacker News

SyGra – Graph-oriented Synthetic data generation Pipeline for LLMs

We're open-sourcing SyGra, a framework for building reproducible synthetic-data pipelines for LLM training and evaluation (SFT, DPO, agent simulation, multimodal). Problem: High-quality datasets are scarce, expensive, and often sensitive. When teams turn to synthetic data, the difficulty isn't single prompts—it's the end-to-end system: designing branching/looping workflows, coordinating multiple inference backends/APIs and tool calls, enforcing validation + schema compliance + quality tagging at scale, and running fault-tolerant jobs with resumability, sharding, and streaming. Ad-hoc notebooks/scripts don't capture that lifecycle. What SyGra is: A graph-oriented framework where you define nodes (LLM calls, samplers, transforms, agents, subgraphs) and edges (conditional / parallel / loops). Author pipelines in low-code YAML (CLI-runnable) or compose in Python. Emphasis on structured outputs and reproducibility. Key capabilities: - Graph model: reusable subgraphs; conditional/parallel edges; loops - Quality: dual-stage quality tagging (heuristics + LLM-based scoring); OASST-style conversation formatting - Backends: vLLM, Hugging Face TGI, Azure OpenAI, Ollama (Triton-compatible) - Data I/O: Hugging Face datasets (read/write, streaming) + local files; schema + metadata tracking - Execution: async runtime; checkpointing/resume; sharding support; multimodal inputs (image/audio/text); agent/tool nodes via LangGraph - Reproducibility: deterministic configs, seeds, artifact paths, and provenance logs - Modes: CLI (execute YAML graphs) or Python APIs (embed in notebooks/apps) - License: Apache-2.0 Links: - Repo & README: https://github.com/ServiceNow/SyGra - PyPI: https://pypi.org/project/sygra/ - Paper (design rationale): https://arxiv.org/abs/2508.15432 Disclosure: I'm part of the team behind SyGra.

Share card

Actual performance

1points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, model · Missing: mac, macos, cursor
93%93% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: para, compatible · Missing: supports, reddit linkedin, podcasting
90%90% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, llama, pipe · Missing: https docs, excited, just released
53%53% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: calls · Missing: plus, platform, intuitive
41%41% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: apps, para · Missing: mobile apps, ios, personal
40%40% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: training · Missing: arr, mrr, revenue
17%17% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: audio · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Ho
How to build a synthetic data pipeline using Gretel and Apache Airflow44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

How to build a synthetic data pipeline using Gretel and Apache Airflow

Hacker News1
Sy
Synthetic Data Studio for LLMs53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Synthetic Data Studio for LLMs

Hacker News4
La
Langchian-Beam,integrates LLMs into Apache beam pipeline with Langchian45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Langchian-Beam,integrates LLMs into Apache beam pipeline with Langchian

Hacker News1
Ya
Yadget Synthetic Data Generation for Testing42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Yadget Synthetic Data Generation for Testing

Hacker News2
Sy
Synthetic Data Genomics44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Synthetic Data Genomics

Hacker News1
St
Streaming LLMs in Data Pipeline55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Streaming LLMs in Data Pipeline

Hacker News5
Th
The Synthetic Data Vault (SDV) – Synthetic Data Generation Ecosystem39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The Synthetic Data Vault (SDV) – Synthetic Data Generation Ecosystem

Hacker News3
A
A poor man's data pipeline56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A poor man's data pipeline

Hacker News3
Go
Goodreads Data Pipeline51%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Goodreads Data Pipeline

Hacker News213
Sy
Synthetic dataset generator for NLP and tabular data44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Synthetic dataset generator for NLP and tabular data

Hacker News3