AgentsReading · ~3 min · 44 words deep

Computer use

Computer use lets a model control a desktop or browser like a person · it sees screenshots and sends mouse and keyboard actions.

Text reviewed October 5, 2026

TL;DR

Computer use lets a model control a desktop or browser like a person · it sees screenshots and sends mouse and keyboard actions.

Level 1

Most tools give a model structured access through APIs. Computer use is the fallback for software without an API: the model looks at the screen, decides what to click or type, and repeats until the task is done. Several labs offer it as a model capability or as a hosted agent.

Level 2

Computer use is slower and less reliable than API calls because every step goes through vision and many actions. Benchmarks such as OSWorld measure success rates on real desktop tasks. Sandboxing matters: an agent with a real screen can click anything a user could.

Level 3

Typical loops combine a vision-capable model, a screenshot tool and an action tool, with step limits and human confirmation for sensitive actions such as payments or sending messages.

The takeaway for you
If you are a
Curious · Normie
  • ·An AI that uses your computer for you
If you are a
Builder
  • ·Prefer APIs or MCP tools when they exist
  • ·Run in a sandbox with step limits
If you are a
Investor
  • ·Opens automation for software without APIs
If you are a
Researcher
  • ·Measured by benchmarks like OSWorld
It is one capability an agent can use. Many agents never touch a screen and only call APIs or tools.