Op

OpenCastor Agent Harness Evaluator Leaderboard

Hacker News

OpenCastor Agent Harness Evaluator Leaderboard

I've been building OpenCastor, a runtime layer that sits between a robot's hardware and its AI agent. One thing that surprised me: the order you arrange the skill pipeline (context builder → model router → error handler, etc.) and parameters like thinking_budget and context_budget affect task success rates as much as model choice does. So I built a distributed evaluator. Robots contribute idle compute to benchmark harness configurations against OHB-1, a small benchmark of 30 real-world robot tasks (grip, navigate, respond, etc.) using local LLM calls via Ollama. The search space is 263,424 configs (8 dimensions: model routing, context budget, retry logic, drift detection, etc.). The demo leaderboard shows results so far, broken down by hardware tier (Pi5+Hailo, Jetson, server, budget boards). The current champion config is free to download as a YAML and apply to any robot. P66 safety parameters are stripped on apply — no harness config can touch motor limits or ESTOP logic. Looking for feedback on: (1) whether the benchmark tasks are representative, (2) whether the hardware tier breakdown is useful, and (3) anyone who's run fleet-wide distributed evals of agent configs for robotics or otherwise.

Share card

Actual performance

3points
1comments
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: agent, model, context · Missing: mac, agents, macos
92%92% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Hacker NewsMay not resonate with HN audience · Strong signals: lua, llama, ide · Missing: https docs, excited, just released
49%49% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: para · Missing: mobile apps, ios, personal
41%41% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
Indie HackersIH features products with proven revenue · Strong signals: para · Missing: supports, reddit linkedin, podcasting
39%39% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: builder, calls · Missing: plus, platform, intuitive
33%33% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: arr · Missing: mrr, revenue, profit
20%20% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Missing: web3, chat, crypto
3%3% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

En
Enough, an Agent Harness for Writers50%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Enough, an Agent Harness for Writers

Hacker News3
Th
The Instavest Leaderboard40%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The Instavest Leaderboard

Hacker News6
Pl
Plurnk (Yet *Another* AI Harness)49%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Plurnk (Yet *Another* AI Harness)

Hacker News1
DeepSeek Harness
DeepSeek Harness55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Composable agent harness where everything is a plugin

Product Hunt+169Open Source
ScoreLeader
ScoreLeader36%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Scoreboard & Leaderboard App

Product Hunt+6
Sy
System One Harness (SOH), the harness for System One models53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

System One Harness (SOH), the harness for System One models

Hacker News1
An
An agent harness with model autorouting and memory42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

An agent harness with model autorouting and memory

Hacker News2
GrowU
GrowU55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The Product Leaderboard

Product Hunt
Op
Opair, a coding harness that eschews autonomy67%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Opair, a coding harness that eschews autonomy

Hacker News2
Ne
New Harness in Town53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

New Harness in Town

Hacker News3