Getting full-text scientific content into LLMs+Agents is stupidly hard
Getting full-text scientific content into LLMs+Agents is stupidly hard
Most APIs don’t return actual content. You get metadata, maybe an abstract, maybe a snippet...never the thing itself. And if you want proper sources like arXiv, PubMed, or major publishers? Good luck. You’re stuck scraping tens of millions PDFs or semantic scholar and building your own ingestion pipeline. We hit this building agentic workflows and RAG backends. What we needed wasn’t “search”, it was a way to retrieve real, structured full text with enough metadata to plug straight into a reasoning system. So we built a system that could do that: multimodal inputs (text, math, figures), clean citations, reference chaining, and filters that work (by date, by source, etc). The hard part wasn’t retrieval but preprocessing at scale. Figuring out how to analyse, chunk, structure tens of millions of docs without taking months or breaking the bank. Not to mention dealing with licensed content where formats vary wildly or building retrieval systems at this scale. Still a work in progress with more updates on the way. But miles better than duct-taping together PDFs, AI search engines etc. and hoping to find the relevant context you need.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
Getting users is hard. Retain them.
OpenFaaS – Getting started with minikube and helm
Getting started with OpenFaaS on minikube and helm
Getting Started with Cue
I wrote Getting Started with ownCloud
Dropwizard's Getting Started in Kotlin
Getting started with Ruby and Sinatra
Getting started with Go and Martini
Getting started with Elixir and Phoenix
Briefed – Summaries for Hard Paywalled Content