Ze

Zero downtime embedding model upgrades

Hacker News

Zero downtime embedding model upgrades

People use embedding models all the time for rag/semantic retrieval. However, when a newer, more desireable model comes out, there is an expensive (both in time and computational) cost of re-embedding every document in the database. However, I figured out an interesting way to forgo that upfront embedding cost. algo: old model/index -> retrieve top-K docs -> score those docs with the new model -> cache/materialize the new embeddings so instead of rebuilding the entire vector store upfront, the old index keeps getting retrieved from, while the new model reranks those candidates. This works surprisingly well for some model pairs, (i tested 63 source-> target migrations on h100s, on upto 1M documents). For example, on a 1M document Natural Questions dataset, native Qwen3-Embedding-8B: 0.6812 nDCG@10 Qwen3-4B -> Qwen3-8B, K=50: 0.6816 Qwen3-0.6B -> Qwen3-8B, K=50: 0.6638 MiniLM -> Qwen3-8B, K=50: 0.6486 (the hard part is determining k, I held the k constant above to give some sense of migratability). You can install it with pip pip install embedflow and the code is on github https://github.com/arnsri33/embedflow

Share card

Actual performance

6points
3comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, new, models · Missing: mac, agents, macos
88%88% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
84%84% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: io · Missing: https docs, excited, just released
60%60% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: way · Missing: mobile apps, ios, personal
48%48% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
29%29% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
20%20% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Em
Embedding visualizations for bloggers and journalists – VizFiddle52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Embedding visualizations for bloggers and journalists – VizFiddle

Hacker News1
Vi
Visualizing and Comparing Embedding Vectors as Heatmaps56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Visualizing and Comparing Embedding Vectors as Heatmaps

Hacker News3
Im
Implementing Embedding Gemma in PyTorch28%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Implementing Embedding Gemma in PyTorch

Hacker News3
Ar
ArcFont – Font Embedding Model60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ArcFont – Font Embedding Model

Hacker News5
Em
EmbedFlow –> Upgrade embedding models without re-embedding your corpus56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

EmbedFlow –> Upgrade embedding models without re-embedding your corpus

Hacker News6
Gemini Embedding 2
Gemini Embedding 290%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Google's first natively multimodal embedding model

Product Hunt+241Developer Tools
I
I made a dataset for finetuning embedding models61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made a dataset for finetuning embedding models

Hacker News1
Quaere
Quaere33%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Embedding as a Service

Product Hunt+7
Bi
Binarized Attributed Network Embedding (ICDM 2018)50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Binarized Attributed Network Embedding (ICDM 2018)

Hacker News3
As
AskVideos-VideoCLIP: Open-source video-text embedding model58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

AskVideos-VideoCLIP: Open-source video-text embedding model

Hacker News9