Mu

Multilingual Embedding Model for Images, Audio and PDFs

Hacker News

Multilingual Embedding Model for Images, Audio and PDFs

I love building RAG applications and exploring new technologies in this space, especially for retrieval and reranking. Here’s an open source project I worked on previously that explored a RAG application on Postgres and YouTube videos: https://news.ycombinator.com/item?id=38705535 Most RAG applications consist of two pieces: the vector database and the embedding model to generate the vector. A scalable vector database seems pretty much like a solved problem with providers like Cloudflare, Supabase, Pinecone, and many many more. Embedding models, on the other hand, seem pretty limited compared to their LLM counterparts. OpenAI has one of the best LLMs in the world right now, with multimodal support for images and documents, but their embedding models only support a handful of languages and only text input while being pretty far behind open source models based on the MTEB ranking: https://huggingface.co/spaces/mteb/leaderboard The closest model I found that supports multi-modality was OpenAI’s clip-vit-large-patch14, which supports only text and images. It hasn't been updated for years with language limitations and has ok retrieval for small applications. Most RAG applications I have worked on had extensive requirements for image and PDF embeddings in multiple languages. Enterprise RAG is a common use case with millions of documents in different formats, verticals like law and medicine, languages, and more. So, we at JigsawStack launched an embedding model that can generate vectors of 1024 for images, PDFs, audios and text in the same shared vector space with support for over 80+ languages. - Supports 80+ languages - Support multimodality: text, image, pdf, audio - Average MRR 10: 70.5 - Built in chunking of large documents into multiple embeddings Today, we launched the embedding model in a closed Alpha and did up a simple documentation for you to get started. Drop me an email at yoeven@jigsawstack.com or DM me on X with your use case and I would be happy to give you free and unlimited access in exchange for feedback! Some limitations: - While our model does support video, it's pretty expensive to run video embedding, even for a 10 second clip. We’re finding ways to reduce the cost before launching this, but you can embed the audio of a video. - Text embedding is the fastest response, while other modalities might take a few seconds. Which we expected as most other modalities require preprocessing

Share card

Actual performance

1points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Indie HackersFits the IH revenue-focused audience · Strong signals: supports, started, ios · Missing: reddit linkedin, podcasting, created
94%94% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Product HuntOn track for Day 1 leaderboard · Strong signals: model, new, models · Missing: mac, agents, macos
88%88% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: open source, ide, io · Missing: https docs, excited, just released
75%75% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRFits verified-revenue profile · Strong signals: ios, video, way · Missing: mobile apps, personal, entrepreneurs
52%52% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoStrong fit for a featured deal · Missing: plus, platform, intuitive
51%51% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: mrr · Missing: arr, revenue, profit
19%19% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: audio · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Ze
Zero downtime embedding model upgrades60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Zero downtime embedding model upgrades

Hacker News6
Em
Embedding visualizations for bloggers and journalists – VizFiddle52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Embedding visualizations for bloggers and journalists – VizFiddle

Hacker News1
Vi
Visualizing and Comparing Embedding Vectors as Heatmaps56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Visualizing and Comparing Embedding Vectors as Heatmaps

Hacker News3
Im
Implementing Embedding Gemma in PyTorch28%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Implementing Embedding Gemma in PyTorch

Hacker News3
Ar
ArcFont – Font Embedding Model60%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

ArcFont – Font Embedding Model

Hacker News5
Em
EmbedFlow –> Upgrade embedding models without re-embedding your corpus56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

EmbedFlow –> Upgrade embedding models without re-embedding your corpus

Hacker News6
Fu
Fulltext search on PDFs and scanned images63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Fulltext search on PDFs and scanned images

Hacker News3
Gemini Embedding 2
Gemini Embedding 290%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Google's first natively multimodal embedding model

Product Hunt+241Developer Tools
Co
Convert TailwindCSS into Images and PDFs with Tailrender57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Convert TailwindCSS into Images and PDFs with Tailrender

Hacker News5
Phrase Log
Phrase Log37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Multilingual phrases app with audio & pronunciation

Product Hunt+138Android