Ge

Getting full-text scientific content into LLMs+Agents is stupidly hard

Hacker News

Getting full-text scientific content into LLMs+Agents is stupidly hard

Most APIs don’t return actual content. You get metadata, maybe an abstract, maybe a snippet...never the thing itself. And if you want proper sources like arXiv, PubMed, or major publishers? Good luck. You’re stuck scraping tens of millions PDFs or semantic scholar and building your own ingestion pipeline. We hit this building agentic workflows and RAG backends. What we needed wasn’t “search”, it was a way to retrieve real, structured full text with enough metadata to plug straight into a reasoning system. So we built a system that could do that: multimodal inputs (text, math, figures), clean citations, reference chaining, and filters that work (by date, by source, etc). The hard part wasn’t retrieval but preprocessing at scale. Figuring out how to analyse, chunk, structure tens of millions of docs without taking months or breaking the bank. Not to mention dealing with licensed content where formats vary wildly or building retrieval systems at this scale. Still a work in progress with more updates on the way. But miles better than duct-taping together PDFs, AI search engines etc. and hoping to find the relevant context you need.

Share card

Actual performance

4points
2comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, agentic · Missing: mac, macos, cursor
90%90% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
70%70% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: pipe, io · Missing: https docs, excited, just released
59%59% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: month, way · Missing: mobile apps, ios, personal
43%43% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
37%37% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
25%25% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

retain.cc
retain.cc

Getting users is hard. Retain them.

BetaList
Op
OpenFaaS – Getting started with minikube and helm46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

OpenFaaS – Getting started with minikube and helm

Hacker News2
Ge
Getting started with OpenFaaS on minikube and helm46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Getting started with OpenFaaS on minikube and helm

Hacker News3
Ge
Getting Started with Cue52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Getting Started with Cue

Hacker News1
I
I wrote Getting Started with ownCloud55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I wrote Getting Started with ownCloud

Hacker News33
Dr
Dropwizard's Getting Started in Kotlin40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Dropwizard's Getting Started in Kotlin

Hacker News2
Ge
Getting started with Ruby and Sinatra49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Getting started with Ruby and Sinatra

Hacker News17
Ge
Getting started with Go and Martini52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Getting started with Go and Martini

Hacker News31
Ge
Getting started with Elixir and Phoenix71%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Getting started with Elixir and Phoenix

Hacker News4
Br
Briefed – Summaries for Hard Paywalled Content46%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Briefed – Summaries for Hard Paywalled Content

Hacker News8