Fi

FieldOps-Bench an open eval for physical-world AI agents

Hacker News

FieldOps-Bench an open eval for physical-world AI agents

Hey HN, I'm Pete. I'm a boat captain by trade, but I've spent the last 16 months building Camera Search. Agents in the physical world need a different set of skills to be useful. We've optimized our harness and architecture to specialize in diagnosing and fixing problems in traditional industries like mining, oil & gas, telecom, construction, and the skilled trades. Existing benchmarks didn't cover what workers in these industries actually do day-to-day, so today I'm publishing FieldOps-Bench on github and Hugging Face [ https://huggingface.co/datasets/CameraSearch/fieldopsbench ]. It's a 157 case multimodal benchmark across 7 industries, testing visual diagnostics, code/standard citations, and general industrial field knowledge. I ran it against our agent and the frontier models. Camera Search beat Claude Opus 4.6 on 87% of cases. I scored it two ways: a rubric and pairwise judging. I'm not a benchmarks specialist, so criticism is welcome, and yes, it's apples-to-oranges because my agent has tool use the baseline models don't. I think it still shows what's possible when you tune the system and corpus for a specific vertical instead of relying on a general-purpose model. Happy to answer any questions, and would love to connect with people building agents for the physical world. Especially where the stakes are high and the information is incomplete. -Pete

Share card

Actual performance

1points
1comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agents, agent, claude · Missing: mac, macos, cursor
96%96% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Missing: supports, reddit linkedin, podcasting
82%82% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
TrustMRRFits verified-revenue profile · Strong signals: month, way · Missing: mobile apps, ios, personal
62%62% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Hacker NewsMay not resonate with HN audience · Strong signals: exist, existing, io · Missing: https docs, excited, just released
49%49% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
AppSumoMay struggle as an AppSumo deal · Missing: plus, platform, intuitive
37%37% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
18%18% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Le
Leviathan, A world where AI agents write the laws and govern themselves35%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Leviathan, A world where AI agents write the laws and govern themselves

Hacker News3
Supastro
Supastro52%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

WORLD'S FIRST VEDIC AI AGENTS

Indie Hackerscommitment-full-time
Wo
Wovyn, physical world sensors for developers61%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Wovyn, physical world sensors for developers

Hacker News78
A
A Physical Representation of Frogger43%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

A Physical Representation of Frogger

Hacker News1
Idyll: AI Girlfriend/Boyfriend Voicebot
Idyll: AI Girlfriend/Boyfriend Voicebot71%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

World's first AI talking Girlfriend/Boyfriend

Indie Hackers1ai
SNEWPapers
SNEWPapers74%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The World's First AI Newspaper Archive

Product Hunt+123Education
Ro
Rookbot, the World's First AI Virtual Assistant for Debugging49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Rookbot, the World's First AI Virtual Assistant for Debugging

Hacker News6
Tw
Twitter for the physical world54%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Twitter for the physical world

Hacker News2
Ariadna
Ariadna85%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The world's first AI talent agent

Product Hunt+951Artificial Intelligence
LabGPT
LabGPT59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Deploying AI agents into physical labs, accelerating science

Indie Hackers1ai