I

I made a website to semantically search ArXiv papers

Hacker News

I made a website to semantically search ArXiv papers

As a grad student (and an ADHDer), I had trouble doing literature review systematically. To combat this, I made a website that finds similar papers using the meaning of the thing I am looking for. I used MixedBread's [^1] embedding model to generate vectors from the abstracts. I store and search similar vectors using Milvus [^2] and finally use Gradio [^3] to serve the frontend. I update the vector database weekly by pulling the metadata dataset from Kaggle [^4]. To speed up the search process on my free oracle instance, I binarise the embeddings and use Hamming distance as a metric. I would love your feedback on the site :) Happy Holidays! [1]: https://www.mixedbread.ai/docs/embeddings/mxbai-embed-large-... [2]: https://milvus.io/ [3]: https://www.gradio.app/ [4]: https://www.kaggle.com/datasets/Cornell-University/arxiv

Share card

Actual performance

324points
104comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: model, using · Missing: mac, agents, macos
68%68% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: adhd, io · Missing: https docs, excited, just released
67%67% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoStrong fit for a featured deal · Missing: plus, platform, intuitive
56%56% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Indie HackersIH features products with proven revenue · Missing: supports, reddit linkedin, podcasting
48%48% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRLess likely to generate early MRR · Missing: mobile apps, ios, personal
38%38% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
11%11% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
2%2% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Th
The “tl;dr” of Recent Transformer Papers58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The “tl;dr” of Recent Transformer Papers

Hacker News5
Th
The Federalist Papers, typeset as the 1787 newspapers they ran in67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The Federalist Papers, typeset as the 1787 newspapers they ran in

Hacker News58
Ba
Back Me Up – Find papers that back your argument63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Back Me Up – Find papers that back your argument

Hacker News2
PagePeek
PagePeek59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Your AI Professor for drafting and evaluating papers in secs

Product Hunt+6
Pl
Plasmyd, a platform for scientists to discuss papers70%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Plasmyd, a platform for scientists to discuss papers

Hacker News6
Ar
Arxiv Reader: Search and Read Arxiv Papers67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Arxiv Reader: Search and Read Arxiv Papers

Hacker News1
Be
Best Rejected Papers78%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Best Rejected Papers

Hacker News1
De
Deep search of all ML papers73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Deep search of all ML papers

Hacker News109
St
StackOverflow for Research Papers74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

StackOverflow for Research Papers

Hacker News2
Ha
HackerNews but for research papers59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

HackerNews but for research papers

Hacker News319