SyGra – Graph-oriented Synthetic data generation Pipeline for LLMs
SyGra – Graph-oriented Synthetic data generation Pipeline for LLMs
We're open-sourcing SyGra, a framework for building reproducible synthetic-data pipelines for LLM training and evaluation (SFT, DPO, agent simulation, multimodal). Problem: High-quality datasets are scarce, expensive, and often sensitive. When teams turn to synthetic data, the difficulty isn't single prompts—it's the end-to-end system: designing branching/looping workflows, coordinating multiple inference backends/APIs and tool calls, enforcing validation + schema compliance + quality tagging at scale, and running fault-tolerant jobs with resumability, sharding, and streaming. Ad-hoc notebooks/scripts don't capture that lifecycle. What SyGra is: A graph-oriented framework where you define nodes (LLM calls, samplers, transforms, agents, subgraphs) and edges (conditional / parallel / loops). Author pipelines in low-code YAML (CLI-runnable) or compose in Python. Emphasis on structured outputs and reproducibility. Key capabilities: - Graph model: reusable subgraphs; conditional/parallel edges; loops - Quality: dual-stage quality tagging (heuristics + LLM-based scoring); OASST-style conversation formatting - Backends: vLLM, Hugging Face TGI, Azure OpenAI, Ollama (Triton-compatible) - Data I/O: Hugging Face datasets (read/write, streaming) + local files; schema + metadata tracking - Execution: async runtime; checkpointing/resume; sharding support; multimodal inputs (image/audio/text); agent/tool nodes via LangGraph - Reproducibility: deterministic configs, seeds, artifact paths, and provenance logs - Modes: CLI (execute YAML graphs) or Python APIs (embed in notebooks/apps) - License: Apache-2.0 Links: - Repo & README: https://github.com/ServiceNow/SyGra - PyPI: https://pypi.org/project/sygra/ - Paper (design rationale): https://arxiv.org/abs/2508.15432 Disclosure: I'm part of the team behind SyGra.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
How to build a synthetic data pipeline using Gretel and Apache Airflow
Synthetic Data Studio for LLMs
Langchian-Beam,integrates LLMs into Apache beam pipeline with Langchian
Yadget Synthetic Data Generation for Testing
Synthetic Data Genomics
Streaming LLMs in Data Pipeline
The Synthetic Data Vault (SDV) – Synthetic Data Generation Ecosystem
A poor man's data pipeline
Goodreads Data Pipeline
Synthetic dataset generator for NLP and tabular data