Pl

Playground for comparing embedding models on Wikipedia+book retrieval

Hacker News

Playground for comparing embedding models on Wikipedia+book retrieval

Introducing embeds.ai: an embedding playground to compare how embedding models work on a real world use case (retrieval augmented generation for Wikipedia articles + Elad Gil's High growth handbook) A few weeks ago, Shreyan and I were looking for an embedding model to use for RAG. We eventually came across the MTEB leaderboard, but we struggled to understand the benchmark scores. We wanted a tool to test various embedding models with example queries on real-world datasets. After unsuccessfully looking for such a “playground”, we decided to just build one ourselves! We embedded HuggingFace’s Simple Wikipedia dataset using @OpenAI, @Cohere, and 2 open-source models via @Baseten. We then stored the embeddings in @Supabase using pgvector. Finally, we built a web app using NextJS and deployed it on @Vercel. Now we’re hosting the playground for anyone to use for free, as well as open-sourcing our work so people can try evaluating other models, datasets, or indexes. Learn more here in our full blog post here: https://shreyanjain.substack.com/p/announcing-embedding-batt... And the repo is here: https://github.com/EGCap/playground If you have other suggestions / pain points from working with embedding models, vector DBs, or RAG, or if you would like to collaborate on any of the above or unrelated projects, please reach out! @shreyanj98 @davidtsong on Twitter

Share card

Actual performance

5points
11comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, models, openai · Missing: mac, agents, macos
88%88% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
75%75% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, ide, io · Missing: https docs, excited, just released
65%65% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoStrong fit for a featured deal · Strong signals: host · Missing: plus, platform, intuitive
51%51% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
27%27% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: growth · Missing: arr, mrr, revenue
14%14% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: collaborate, real world · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

I
I made a dataset for finetuning embedding models61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made a dataset for finetuning embedding models

Hacker News1
Em
EmbedFlow –> Upgrade embedding models without re-embedding your corpus56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

EmbedFlow –> Upgrade embedding models without re-embedding your corpus

Hacker News6
Wo
Word embedding playground (gensim and whatlies)43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Word embedding playground (gensim and whatlies)

Hacker News1
Em
Embedding visualizations for bloggers and journalists – VizFiddle52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Embedding visualizations for bloggers and journalists – VizFiddle

Hacker News1
Vi
Visualizing and Comparing Embedding Vectors as Heatmaps56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Visualizing and Comparing Embedding Vectors as Heatmaps

Hacker News3
Im
Implementing Embedding Gemma in PyTorch28%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Implementing Embedding Gemma in PyTorch

Hacker News3
Ho
How to get around the Wikipedia blackout (if you have to)41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

How to get around the Wikipedia blackout (if you have to)

Hacker News1
Cr
Crowdsourced incremental Wikipedia improvements59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Crowdsourced incremental Wikipedia improvements

Hacker News4
wi
wikiUp - Wikipedia in Tooltips59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

wikiUp - Wikipedia in Tooltips

Hacker News26
Le
Legible Wikipedia59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Legible Wikipedia

Hacker News2