We

We unified LLMs, vector memory, ranking, pruning models in one process

Hacker News

We unified LLMs, vector memory, ranking, pruning models in one process

There is a lot of latency involved shuffling data for modern and complex ML systems in production. In our experience these costs dominate end-to-end user latency, rather than actual model or ANN algorithms, which unfortunately limits what is achievable for interactive applications. We've extended Postgres w/ open source models from Huggingface, as well as vector search, and classical ML algos, so that everything can happen in the same process. It's significantly faster and cheaper, which leaves a large latency budget available to expand model and algorithm complexity. In addition open source models have already surpassed OpenAI's text-embedding-ada-002 in quality, not just speed. [1] Here is a series of posts explaining how to accomplish the complexity involved in a typical ML powered application, as a single SQL query, that runs in a single process with memory shared between models and feature indexes, including learned embeddings and reranking models. - Generating LLM embeddings with open source models in the database[2] - Tuning vector recall [3] - Personalize embedding results with application data [4] This allows a single SQL query to accomplish what would normally be an entire application w/ several model services and databases e.g. for a modern chatbot built across various services and databases -> application sends user input data to embedding service <- embedding model generates a vector to send back to application -> application sends vector to vector database <- vector database returns associated metadata found via ANN -> application sends metadata for reranking <- reranking model prunes less helpful context -> application sends finished prompt w/ context to generative model <- model produces final output -> application streams response to user [1]: https://huggingface.co/spaces/mteb/leaderboard [2]: https://postgresml.org/blog/generating-llm-embeddings-with-open-source-models-in-postgresml [3]: https://postgresml.org/blog/tuning-vector-recall-while-generating-query-embeddings-in-the-database [4]: https://postgresml.org/blog/personalize-embedding-vector-search-results-with-huggingface-and-pgvector Github: https://github.com/postgresml/postgresml

Share card

Actual performance

4points
1comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, user, models · Missing: mac, agents, macos
94%94% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: including · Missing: supports, reddit linkedin, podcasting
79%79% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: open source, io, including · Missing: https docs, excited, just released
72%72% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
40%40% predicted probability of success on AppSumo, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Strong signals: personal · Missing: mobile apps, ios, entrepreneurs
39%39% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: active · Missing: arr, mrr, revenue
11%11% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
1%1% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Un
Unified memory across all LLMs38%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Unified memory across all LLMs

Hacker News2
Ra
RankFight – We're Ranking Everything50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

RankFight – We're Ranking Everything

Hacker News6
Co
Covid-19 infection ranking adjusted by population45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Covid-19 infection ranking adjusted by population

Hacker News5
MissedQueries
MissedQueries57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

You're ranking for searches you never wrote about.

Indie Hackers1analytics
Ra
Ranking LLMs by Usage over Time67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Ranking LLMs by Usage over Time

Hacker News3
A
A production-style recommender using vector retrieval and re-ranking62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A production-style recommender using vector retrieval and re-ranking

Hacker News2
MemSync
MemSync77%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Unified Memory for all of your apps

Product Hunt+133Chrome Extensions
Ve
Vector Data Marketplace for LLMs68%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Vector Data Marketplace for LLMs

Hacker News14
Pl
Pluggable vector space models57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Pluggable vector space models

Hacker News12
Me
Memleax – detects memory leak of a running process45%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Memleax – detects memory leak of a running process

Hacker News42