Dl

Dlt – Python library to automate the creation of datasets

Hacker News

Dlt – Python library to automate the creation of datasets

Hi HN, We're Anna, Adrian, Marcin and Matt, developers of dlt. dlt is an open source library to automatically create datasets out of messy, unstructured data sources. You can use the library to move data from about anywhere into most of well known SQL and vector stores, data lakes, storage buckets, or local engines like DuckDB. It automates many cumbersome data engineering tasks and can by handled by anyone who knows Python. Here’s our Github: https://github.com/dlt-hub/dlt Here’s our Colab demo: https://colab.research.google.com/drive/1DhaKW0tiSTHDCVmPjM-... — — — In the past we wrote hundreds of Python scripts to fit messy data sources into something that you can work with in Python - a database, Pandas frame or just a Python list. We were solving the same problems and making the similar mistakes again and again. This is why we built an easy to use Python library called dlt that will automate most data engineering tasks. It hides the complexities of data loading and automatically generates a structured and clean datasets for immediate querying and sharing. — — — At its core, dlt removes the need to create the dataset schemas, react to changing data, generate append or merge statements, and to move the data in transactional and idempotent manner. Those things are automated and can be declared right in the Python code, just by decorating functions. Add @dlt.resource decorator, give it a few hints, and convert any data into a simple pipeline that creates and updates datasets. dlt gets the details out of your way: 1. You do not need to worry about the structure of a database or parquet files dlt will create a nice, typed schema out of your data and will migrate it when the data changes. You can put some data contracts and Pydantic models on top to keep your data clean. 2. You do not need to write any INSERT/UPDATE or data copy statements dlt will push the data to DuckDB, Weaviate, storage buckets and many popular SQL stores. It will align the data types, file formats, and identifier names automatically 3. You do not need to worry when you need to add new data or update the changes. dlt lets you declare how to load the data, how to increment it and will keep the loading state together so they are always in sync. 4. You keep how you develop and test your code Iterate and test quickly on your laptop or in a dev container. Run locally on DuckDB and just swap destination name to go to the cloud - your code, schema and data will stay the same. 5. You can work with data on your laptop. Combine dlt with other tools and libraries to process data locally. duckdb, Pandas, Arrow tables and Rust based loading libraries like ConnectorX work nicely with dlt and process data blazingly fast, compared to the cloud. 6. You do not need to worry if your pipeline will work when you deploy it. dlt is a minimalistic Python library, requires no backend and works whenever Python works. You can finetune it to work on constrained environments like AWS Lambda or run with Airflow, GitHub Actions or Dagster. dlt has an Apache 2.0 license. We plan to make money by offering organizations a paid control plane, where dlt users can track and policy what every pipeline does, manage schemas and contracts across organization, create data catalogues, and share them with the team members and customers.

Share card

Actual performance

114points
54comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, google, user · Missing: mac, agents, macos
93%93% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: mistakes, organizations · Missing: supports, reddit linkedin, podcasting
87%87% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: open source, ide, pipe · Missing: https docs, excited, just released
74%74% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Strong signals: users · Missing: plus, platform, intuitive
46%46% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: google, users, way · Missing: mobile apps, ios, personal
42%42% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: arr · Missing: mrr, revenue, profit
12%12% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: make money, paid · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Xr
Xray: N-D labeled arrays and datasets in Python58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Xray: N-D labeled arrays and datasets in Python

Hacker News6
Cr
Create simulated datasets in Python with Simulacrum52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Create simulated datasets in Python with Simulacrum

Hacker News4
A
A Python library for Unicode string collation55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A Python library for Unicode string collation

Hacker News4
A
A Python library for accessing Wunderlist67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A Python library for accessing Wunderlist

Hacker News1
Oh
Ohne I/O, a Python 3.4+ coroutine library without any I/O55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Ohne I/O, a Python 3.4+ coroutine library without any I/O

Hacker News2
Co
Combinatorix, yet another combinator library in Python55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Combinatorix, yet another combinator library in Python

Hacker News4
Tt
Tt Python library67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Tt Python library

Hacker News1
Py
Python library for GoPro cameras67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Python library for GoPro cameras

Hacker News131
B2
B2blaze – A Backblaze B2 library for Python55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

B2blaze – A Backblaze B2 library for Python

Hacker News359
Py
Python library for parsing Dockerfile65%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Python library for parsing Dockerfile

Hacker News2