Co

CocoIndex – Open-Source Data framework for AI, built for data freshness

Hacker News

CocoIndex – Open-Source Data framework for AI, built for data freshness

Hi HN, I’ve been working on CocoIndex, an open-source Data ETL framework to transform data for AI, optimized for data freshness. You can start a CocoIndex project with `pip install cocoindex` and declare a data flow that can build ETL like LEGO - build a RAG pipeline for vector embeddings, knowledge graphs, or extract, transform data with LLMs. It is a data processing framework beyond text. When you run the data flow either with live mode or batch mode, it will process the data incrementally with minimal recomputation and make it super fast to update the target stores on source changes. Get started video: https://www.youtube.com/watch?v=gv5R8nOXsWU Demo video: https://www.youtube.com/watch?v=ZnmyoHslBSc Previously, I’ve worked at Google on projects like search indexing and ETL infra for 8 years. After I left Google last year, I built various projects and went through pivoting hell. In all the projects I’ve built, data still sits in the center of the problem and I find myself focusing on building data infra other than the business logic I need for data transformation. The current prepackaged RAG-as-service doesn't serve my needs, because I need to choose a different strategy for the context, and I also need deduplication, clustering (items are related), and other custom features that are commonly needed. That’s where CocoIndex starts. A simple philosophy behind it - data transformation is similar to formulas in spreadsheets. The ground of truth is at the source data, and all the steps to transform, and final target store are derived data, and should be reactive based on the source change. If you use CocoIndex, you only need to worry about defining transformations like formulas. *Data flow paradigm* came in as an immediate choice - because there’s no side effect, lineage and observability just come out of the box. *Incremental processing* - If you are a data expert, an analogy would be a materialized view beyond SQL. The framework tracks pipeline states in database, and only reprocessing necessary portions. When data has changed, framework handles the change data capture comprehensively and combines the mechanism for push and pull. Then clear stale derived data/versions and re-index data based on tracking data/logic changes or data TTL settings. There’s lots of edge cases to do it right, for example, when a row is referenced in other places, and the row changes. These should be handled at the level of the framework. *At the compute engine level* - the framework should consider the multiple processes and concurrent updates. It should consider how to resume existing states from terminated execution. In the end, we want to build a framework that is easy to build with exceptional velocity, but scalable and robust in production. *Standardized the interface throughout the data flow* - really easy to plugin custom logic like LEGO; with a variety of native built-in components. One example is that it takes a few lines to switch among Qdrant, Postgres, Neo4j. CocoIndex is licensed under Apache 2.0 https://github.com/cocoindex-io/cocoindex Getting started: https://cocoindex.io/docs/getting_started/quickstart Excited to learn your thoughts, and thank you so much! Linghua

Share card

Actual performance

14points
11comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: started, para · Missing: supports, reddit linkedin, podcasting
95%95% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: google, context, using · Missing: mac, agents, macos
93%93% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: excited, exist, existing · Missing: https docs, just released, lua
80%80% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRFits verified-revenue profile · Strong signals: video, google, para · Missing: mobile apps, ios, personal
52%52% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: interface · Missing: plus, platform, intuitive
33%33% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: active · Missing: arr, mrr, revenue
16%16% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

St
Stargate – An open source API framework for data74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Stargate – An open source API framework for data

Hacker News72
Ev
Evvo – an open source framework for distributed evolutionary algorithms75%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Evvo – an open source framework for distributed evolutionary algorithms

Hacker News3
Me
MetricFlow – open-source metric framework81%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

MetricFlow – open-source metric framework

Hacker News98
Di
Dissect – An open source DFIR framework73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Dissect – An open source DFIR framework

Hacker News8
RedBlue
RedBlue48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-Source Hypervideo Framework

Indie Hackers1apis
ZenML
ZenML50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

extensible, open-source MLOps framework

Indie Hackers4$1/moai
Si
SirixDB – Storing and Querying of Temporal Data (Java and Open Source)62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SirixDB – Storing and Querying of Temporal Data (Java and Open Source)

Hacker News13
St
Streamdal – an open-source tail -f for your data81%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Streamdal – an open-source tail -f for your data

Hacker News148
Op
Open-Source Data Replication and Anonymization73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-Source Data Replication and Anonymization

Hacker News24
Ne
Neosync – Open Source Data Replication and Anonymization73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Neosync – Open Source Data Replication and Anonymization

Hacker News4