Cr

Crustdata (YC F24) – Web Search API for Token-Efficient AI Agents

Hacker News

Crustdata (YC F24) – Web Search API for Token-Efficient AI Agents

Hi HN! We’re Abhilash Chowdhary, Chris Pisarski and Manmohit Grewal. We built Crustdata (YC F24). Today we’re launching our web search API for AI agents, which not only returns the most relevant documents from the web but also maps them to the correct entity (person, company or event). Demo video here https://youtu.be/IouWW97hBN8 If you run agents at scale, tokens become a line item. The web data is the worst input: long pages, repeated content, mixed entities, stale claims. The usual web search -> scrape -> summarize + structure forces the agent to spend tokens doing janitorial work before it can take action. We’re trying to move that work upstream. We keep a canonical graph (ontology) of people and companies: stable internal IDs, aliases, and relationships. Then we continuously index the web and attach each document to the right entity ID. Example: raw web search for "Stripe pricing changes 2026" returns ~10 results across ~4,000 tokens, mostly redundant. We return 6 deduplicated results in ~1,200 tokens. This is not just about saving tokens. It also matters because the common failure isn’t “search missed something.” It’s “search found something about the wrong entity.” Names collide. Companies rebrand. Domains move. Press releases get syndicated and look like independent sources. If you treat strings as IDs, you eventually attach evidence to the wrong person/company and the agent takes a confident action based on that mistake. Under the hood, we run a continuous pipeline that updates the entity-linked index: discover -> fetch -> extract -> dedupe -> entity resolution -> attach -> index . And we serve you this index via our search API. We didn’t start with web search. We spent ~2 years building verified people + company data from higher-trust sources. That forced us to build identity as a system, not a string. When we tried to bolt on web search and started building our integrated index of documents + people + companies, we ended up with a pile of local fixes: parser tweaks, domain rules, prompt hacks. Each fix helped one case and broke another because identity isn’t local. That’s when we committed to an entity-first index: pay the entity resolution cost once, then reuse it everywhere. If you’re building AI agents for sales, recruiting, or investing that do a lot of web searches for people and companies, we’d love for you to try our web search APIs. https://crustdata.com/demo

Share card

Actual performance

10points
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, stripe · Missing: mac, macos, cursor
97%97% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: started · Missing: supports, reddit linkedin, podcasting
96%96% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: ide, pipe, 000 · Missing: https docs, excited, just released
59%59% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: video · Missing: mobile apps, ios, personal
38%38% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: efficient · Missing: plus, platform, intuitive
28%28% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
23%23% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Ti
TinySearch – token-efficient web research for local AI agents44%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

TinySearch – token-efficient web research for local AI agents

Hacker News3
So
Sourceminder.org - token-efficient code search57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Sourceminder.org - token-efficient code search

Hacker News4
Search1API
Search1API37%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Web Search API for AI Agents

Indie Hackerscommitment-full-time
Web search API by Crustdata
Web search API by Crustdata94%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Accurate and the fastest web search API for AI Agents

Product Hunt+334Sales
Platus YC F24
Platus YC F2477%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Your AI Paralegal

Product Hunt+271Legal
Search1API
Search1API30%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

One API to search, crawl, and ingest the live web for AI agents.

AppSumo
YC
YC's Sales Agreement as an API58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

YC's Sales Agreement as an API

Hacker News10
Al
Algolia (YC W14) Presents DocSearch74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Algolia (YC W14) Presents DocSearch

Hacker News9
Wh
Why we killed imwith and launched GIFted (Guggy YC S17)73%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Why we killed imwith and launched GIFted (Guggy YC S17)

Hacker News2
Cr
Crowdfunding by BackerKit (YC W13)63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Crowdfunding by BackerKit (YC W13)

Hacker News2