I built a tool for mobile and computer use using local and remote LLMs
I built a tool for mobile and computer use using local and remote LLMs
Created a tool that lets you use LLMs to automate task across mobile (android) and computer. Currently, this uses screenshots and LLMs support for extracting screen UI elements effectively. This is still a work in progress and attempting to make this work with local models via Ollama (the code is in place with some issues). As of now, Gemini and GPT 4o works the best for finding UI elements and planning the task. Some examples that work as of now: 1. Use gmail and ask <friend>@example.com for lunch next saturday 2. Start a 3+2 chess game on lichess Working demos: https://github.com/BandarLabs/clickclickclick This improves the cost of one automation task from approx. $0.6 via Claude to: $0.06 - OpenAI 4o mini as planner + free Gemini flash 1.5 (15 calls/min) The Llama vision models will eventually make it 0.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
I built "computer use" before Anthropic, while using their API
Open Computer Use
CUA-S1 – A System One Model for Computer Use
Computer use for automating operations
Agent – A Local Computer-Use Operator for macOS
"Computer use" mcp for webapps and Electron apps
Polaris – Toolbox for AI computer use
AutoBrowser – Automate your browser with Claude Computer Use
Computer use but with OpenAI and Gemini models
Docker for computer-use agents