VD

VDP – open-source unstructured visual data ETL

Hacker News

VDP – open-source unstructured visual data ETL

Hi HN - We're Ping-Lin and Xiaofei from Instill AI ( https://www.instill.tech ). We're building VDP ( https://github.com/instill-ai/vdp ), an open-source ETL tool for unstructured visual data. When people say they are data-driven, most of the time it means they are driven by structured data. I will cut the part where we cite reports claiming that 80% of data are unstructured. The reality is unstructured data are more difficult to analyse and not a lot of companies know or have the resources to deal with them. Before starting Instill AI, we were in a smart video startup dealing with large volumes of visual data every day. Back then, the concept of MLOps was pretty new (2014), every ML company was exploring and building its own stack. We built a battle-proven Vision AI system in-house and had the system running in production for years. What we have learnt from the journey are: 1) buy vs. build: unaffordable high inference cost was the main barrier keeping us from adopting an off-the-shelf solution like Google Vertex AI or Amazon SageMaker, so we went for the "build" route. The truth was the resources we spent on building and maintaining the system were unexpectedly huge, time and money-wise; 2) the Vision AI system we built can actually be modularised and generalised to apply to other industry sectors. We reckon what we had experienced can be a common phenomenon in the industry, and we can help solve the problem. That's why we decide to build VDP, an open-source, general and modularised ETL infrastructure for unstructured visual data for a broader community. Many brilliant MLOps platforms/tools providing computer vision solutions have emerged in the last few years. Most of the tools are built from a model-centric perspective and fall into the following categories: - general ML platforms for model training, experiment tracking, model deployment, etc. - platforms that serve a specific vertical, such as E-commerce, and manufacturing. - platforms that focus on a single component of MLOps, such as data labelling, dataset preparation, and model serving. VDP is built from a data-driven perspective. Although the computer vision model is the most critical component in a visual data ETL pipeline, the ultimate goal of VDP is to streamline the end-to-end visual data flow, with the transform component being able to flexibly import computer vision models from different sources. Today, the early version of VDP supports 2 sources and all Airbyte destination connectors, and it can import computer vision models from various sources including Local, GitHub, DVC, ArtiVC and Hugging Face. Setting up a VDP pipeline is fairly easy via its low-code API and no-code Console. Please take a look at the tutorial: https://www.instill.tech/docs/tutorials/build-an-async-det-p... . VDP can run locally with Docker Compose. We're working on integrating with Kubernetes and a fully managed version in Instill Cloud. We aim to build VDP as the single point of visual data integration, so users can sync visual data from anywhere into centralised warehouses or applications and focus on gaining insights across all data sources, just like how the modern data stack handles structured data. Operation-wise, VDP resources will be managed in a declarative way to fuse them better with the modern cloud-native context. The API-first and microservice design has opened all sorts of possibilities for VDP. Thanks for reading HN! We are first-time open-source project maintainers. There are definitely lots to learn! Let us know what you think in the comments. VDP links: [1] GitHub: https://github.com/instill-ai/vdp [2] Documentation: https://www.instill.tech/docs [3] Demo: https://demo.instill.tech

Share card

Actual performance

11points
1comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: supports, para, including · Missing: reddit linkedin, podcasting, created
92%92% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, google, user · Missing: mac, agents, macos
92%92% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: ide, pipe, io · Missing: https docs, excited, just released
80%80% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: video, google, users · Missing: mobile apps, ios, personal
48%48% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: platform, users · Missing: plus, intuitive, reviews
22%22% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: arr, training · Missing: mrr, revenue, profit
14%14% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: smart · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

We
We made an open-source visual data ETL74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

We made an open-source visual data ETL

Hacker News8
Be
BemiDB – Open-source data warehouse with zero-ETL82%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

BemiDB – Open-source data warehouse with zero-ETL

Hacker News13
VDP
VDP34%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source unstructured data ETL infrastructure

Indie Hackers3ai
Du
Duckle was benchmarked against 4 open source ELT/ETL tools71%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Duckle was benchmarked against 4 open source ELT/ETL tools

Hacker News2
Si
Singer – Simple, Composable Open Source ETL71%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Singer – Simple, Composable Open Source ETL

Hacker News91
CocoIndex
CocoIndex31%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open source ETL to transform data for AI

Indie Hackerscommitment-full-time
Sn
Snapflow – Functional Data ETL60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Snapflow – Functional Data ETL

Hacker News4
Mu
Multiwoven – Open-Source Reverse-ETL71%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Multiwoven – Open-Source Reverse-ETL

Hacker News4
Si
SirixDB – Storing and Querying of Temporal Data (Java and Open Source)62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SirixDB – Storing and Querying of Temporal Data (Java and Open Source)

Hacker News13
St
Streamdal – an open-source tail -f for your data81%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Streamdal – an open-source tail -f for your data

Hacker News148