We

Web-eval-agent – Let the coding agent debug itself

Hacker News

Web-eval-agent – Let the coding agent debug itself

Hey HN! We’ve been building an MCP server to help AI-assisted web app developers by using browser agents to test whether changes made by an AI inside an editor actually work. We've been testing it on scenarios like verifying new flows in a UI, or checking that sending a chat request triggers a response. The idea is to let your coding agent both code and evaluate if what it did was correct. Here’s a short demo with Cursor: https://www.youtube.com/watch?v=_AoQK-bwR0w When building apps, we found the hardest part of AI-assisted coding isn’t the coding—it’s tedious point-and-click testing to see if things work. We got tired of this loop: open the app, click through flows, stare at the network tab, copy console errors to the editor, repeat. It felt obvious this should be AI-assisted too. If you can vibe-code, you should be able to vibe-test! Some agents like Cline and Windsurf have browser integrations, but Cline’s (via Anthropic Computer Use) felt slow and only reported console logs, and Windsurf’s didn’t work reliably yet. We got so tired of manually testing that we decided to fix it. Our MCP server sits between your IDE agent (Cursor/Windsurf/Cline/Continue) and a Playwright-powered browser-use agent. It spins up the browser, navigates your app per instructions from the IDE agent, and sends back steps, console events, and network events so the IDE agent can assess the app’s state. We proxy Browser-use’s original Claude calls and swap in Gemini Flash 2.0, cutting latency from ~8s → ~3s per step. We also cap console/network logs at 10,000 characters to stay within context limits, and filter out irrelevant logs (e.g., noisy XHR requests). At the end, the browser agent outputs a summary like: Web Evaluation Report for http://localhost:5173 Task: delete an API key and evaluate UX Steps: Home → Login → API Keys → Create Key → Delete Key Flow tested successfully; UX had problems X, Y, Z... Console (8)... Network (13)... Timeline of events (57) … This gives the coding agent the ability to recognize the console and network errors, or any issues with clicking around, and have the coding agent fix them before returning back to the user. (There’s a longer example in the README at https://github.com/Operative-Sh/web-eval-agent .) Try it in Cursor / Cline / Windsurf / Claude Desktop: (macOS/Linux): curl -LSf https://operative.sh/install.sh -o install.sh less -N install.sh # inspect if you’d like bash install.sh # installs uv + jq + Playwright + server # then in Cursor/Cline/Windsurf/Continue: craft a prompt using the web_eval_agent tool (For Windows, there’s a 4-line manual install in the README.) What we want to do next: pause/go for OAuth screens; save/load browser auth states; Playwright step recording for automated test creation and regression test creation; supporting Loveable / v0 / Bolt.new sites by offering a web version. We’d love to hear your feedback, especially if you’ve experienced the pain of having to manually test changes happening in your web apps after making changes from inside your IDE, or if you’ve tried any alternative MCP tools for this that have worked well. Try it out if you feel it’d be helpful for your workflow: https://github.com/Operative-Sh/web-eval-agent . (note: the server hits our operative.sh proxy to cover Gemini tokens. The MCP server itself is OSS; Anthropic base-URL support is coming soon. Free tier included; heavy users can grab the $10 plan to offset our model bill.) Let us know what you think! Thanks for reading!

Share card

Actual performance

84points
12comments
Made the leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: mac, agents, macos · Missing: apple, agentic, slack
99%99% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: ios, gemini · Missing: supports, reddit linkedin, podcasting
87%87% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: lua, ide, 000 · Missing: https docs, excited, just released
53%53% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: ios, apps, users · Missing: mobile apps, personal, entrepreneurs
39%39% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: host, soon, users · Missing: plus, platform, intuitive
30%30% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Missing: arr, mrr, revenue
18%18% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Correct prediction on native model

Similar products

Be
Bestie, a coding agent that respects you59%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Bestie, a coding agent that respects you

Hacker News5
Ba
Bailout – The coding agent meant to be deleted53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Bailout – The coding agent meant to be deleted

Hacker News2
Pr
Premortem, a coding-agent-powered airplane blackbox70%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Premortem, a coding-agent-powered airplane blackbox

Hacker News3
Ullbek
Ullbek87%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Website Coding Agent

Product Hunt+3
Ch
Chrome ext to let zot, your terminal coding agent, operate the browser57%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Chrome ext to let zot, your terminal coding agent, operate the browser

Hacker News11
Be
Beating GPT5.5-xhigh for Coding agent security with SLMs and IRM32%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Beating GPT5.5-xhigh for Coding agent security with SLMs and IRM

Hacker News9
Ha
Hazzel – terminal coding agent, BYOK, undo-everything56%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Hazzel – terminal coding agent, BYOK, undo-everything

Hacker News2
Neurogrid TUI
Neurogrid TUI89%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Neurogrid terminal coding agent

Product Hunt+1
SubmitMap
SubmitMap39%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Submit only where you qualify, and let your agent do it

Indie Hackerscommitment-side-project
Compyle
Compyle91%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The coding agent that asks before it builds

Product Hunt+125Developer Tools