Op

Open-Cowork – an open-source, model-agnostic computer-use agent

Hacker News

Open-Cowork – an open-source, model-agnostic computer-use agent

Hi HN, We built Open Cowork, an MIT-licensed desktop agent that can operate a computer through screenshots, mouse input, and keyboard input. We started working on this after using cowork-style computer agents for longer tasks. The interaction model was useful, but computer use burns through model calls much faster than chat: take a screenshot, decide what to do, perform an action, inspect the result, and repeat. We wanted to separate that agent loop from any one model or subscription. Open Cowork currently works with local models through Ollama and hosted models through providers including OpenAI, Anthropic, Google, xAI, Mistral, and OpenRouter. You can watch the agent’s actions as they happen and interrupt it during a run. The basic loop is: Capture the current screen Send the screenshot and task state to the model Parse the proposed action Execute the mouse or keyboard action Capture the resulting state Continue until completion, interruption, or a safety condition Two annoyingly common failures shaped the implementation. The first was selecting a model that appeared correctly configured but could not actually process image input. The agent would start, receive unusable responses, and waste the run. We now test the model’s screenshot capability before beginning the task. The second was agents getting stuck without technically throwing an error. They would click the same area repeatedly, keep observing an unchanged screen, or produce slightly different versions of the same action. We added idle and repetition detection so the run can stop instead of consuming the entire step budget. You can also run the model locally with Ollama. In that setup, the screenshots and model requests remain on your machine. The tradeoff is that computer use is demanding: smaller local models are often noticeably less reliable at understanding cluttered interfaces and recovering from unexpected states. There is a bundled offline mock if you want to inspect the full interaction flow without configuring an API key or spending money on inference. This is still early. The hardest unresolved problem is verification. Seeing the expected pixels after an action does not necessarily mean the underlying operation succeeded. A form can look filled while its internal state is wrong, and some failures are only visible on a later confirmation page or through an external side effect. We’re currently thinking about combining visual checks with task-specific invariants and downstream verification. I’d be particularly interested in how people here handle: detecting progress without relying only on screenshot similarity verifying side effects when DOM or API access is unavailable deciding when to stop a looping agent getting useful computer-use performance from local models The repository includes the setup instructions and offline mock. Happy to discuss the agent loop or any questionable implementation choices.

Share card

Actual performance

2points
Did not reach leaderboard

Launch Intel predictions

Analyze your own launch →
Product HuntOn track for Day 1 leaderboard · Strong signals: mac, agents, agent · Missing: macos, cursor, claude
98%98% predicted probability of success on Product Hunt, based on ML models trained on real launch data.
best fitHighest predicted score across all platforms for this description.
Indie HackersFits the IH revenue-focused audience · Strong signals: started, para, including · Missing: supports, reddit linkedin, podcasting
88%88% predicted probability of success on Indie Hackers, based on ML models trained on real launch data.
Hacker NewsStrong engagement from HN community · Strong signals: llama, ide, io · Missing: https docs, excited, just released
50%50% predicted probability of success on Hacker News, based on ML models trained on real launch data.
nativeThis product was originally launched on this platform.
TrustMRRLess likely to generate early MRR · Strong signals: google, para · Missing: mobile apps, ios, personal
43%43% predicted probability of success on TrustMRR, based on ML models trained on real launch data.
AppSumoMay struggle as an AppSumo deal · Strong signals: host, interface, calls · Missing: plus, platform, intuitive
27%27% predicted probability of success on AppSumo, based on ML models trained on real launch data.
Acquire.comPre-revenue stage for this audience · Strong signals: subscription · Missing: arr, mrr, revenue
16%16% predicted probability of success on Acquire.com, based on ML models trained on real launch data.
BetaListMay not resonate with beta-testers · Strong signals: chat · Missing: web3, crypto, cryptocurrency
0%0% predicted probability of success on BetaList, based on ML models trained on real launch data.

Incorrect prediction on native model

Similar products

Op
Open Computer Use83%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open Computer Use

Hacker News9
CU
CUA-S1 – A System One Model for Computer Use42%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

CUA-S1 – A System One Model for Computer Use

Hacker News90
Ce
Cerebellum – Open-Source Browser Control with Claude 3.5 Computer Use41%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Cerebellum – Open-Source Browser Control with Claude 3.5 Computer Use

Hacker News42
Op
Open-source background computer use for Windows agents53%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source background computer use for Windows agents

Hacker News1
Sp
SpongeCake – open-source SDK for OpenAI computer use agents62%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

SpongeCake – open-source SDK for OpenAI computer use agents

Hacker News13
Op
Open-source framework to deploy personalized computer-use agents58%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source framework to deploy personalized computer-use agents

Hacker News6
OpenAmer
OpenAmer55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

Open-source AI agent with background computer-use

Indie Hackers
I
I built an open source computer-use SDK enabling agents to authenticate63%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I built an open source computer-use SDK enabling agents to authenticate

Hacker News19
I
I made an open-source alternative of computer-use for windows55%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

I made an open-source alternative of computer-use for windows

Hacker News1
CrabTalk
CrabTalk95%Launch Intel prediction score: how likely this product is to succeed on its source platform, based on its name, tagline, and description.

The agent daemon that hides nothing. 5MB. Open Source

Product Hunt+192Developer Tools