Co

Code retrieval findings from a real-world benchmark

Hacker News

Code retrieval findings from a real-world benchmark

While building our OSS chat-with-your-codebase system, we faced so many choices: chunking strategies, embeddings, retrieval algorithms, rerankers, etc. Since we don't want to make decisions based on vibes , and because academic benchmarks are somewhat contrived, we made our own. Our dataset consists of 1,000 questions about Hugging Face's Transformers library, where each question requires 1-3 Python files to be answered correctly. We started by comparing proprietary APIs for the various sub-tasks involved in an AI copilot. Here are our initial learnings: - OpenAI's text-embedding-3-small embeddings perform best. - NVIDIA's reranker outperforms Cohere, Voyage and Jina. - Sparse retrieval (e.g. BM25) is actively hurting code retrieval if you have natural language files in your index (e.g. Markdown). - Chunks of size 800 are ideal; going smaller has very marginal gains. - Going beyond top_k=25 for retrieval has diminishing returns. We're just getting started and plan on continuously sharing our findings with the community. Go OSS!

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: openai, tasks, code · Missing: mac, agents, macos
86%86% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: started · Missing: supports, reddit linkedin, podcasting
66%66% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: ide, 000, io · Missing: https docs, excited, just released
61%61% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRFits verified-revenue profile · Missing: mobile apps, ios, personal
58%58% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
43%43% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: margin, active · Missing: arr, mrr, revenue
18%18% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
1%1% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

CV
CVE-Bench, the first LLM benchmark using real-world web vulnerabilities55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

CVE-Bench, the first LLM benchmark using real-world web vulnerabilities

Hacker News6
Pa
PagerDuty for the real world48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

PagerDuty for the real world

Hacker News1
I
I made Pokémon but with real animals in the real world62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made Pokémon but with real animals in the real world

Hacker News4
Tr
Trying out actioncable in a real world app37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Trying out actioncable in a real world app

Hacker News1
Cr
Crowsnest – API for the Real World52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Crowsnest – API for the Real World

Hacker News61
PharmaSafe
PharmaSafe52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Pharmacovigilance analytics for exploring real-world adverse

Indie Hackers1education
Th
The Whicher: A/B test the Real World40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The Whicher: A/B test the Real World

Hacker News20
NL
NLP algorithms for real-world sentiment analysis48%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

NLP algorithms for real-world sentiment analysis

Hacker News1
Sethco AI
Sethco AI66%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Prepares your team for challenging real-world scenarios

Indie Hackers1$1,200/moai
MemoRep
MemoRep62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Spaced repetition for real-world skills.

Indie Hackers1education