OpenAI Computer-Using Agent
Computer-Using Agent (Operator)
Model Overview
Designed for GUI automation, computer use, and active web navigation, it processes both text and image modalities. As of mid-2026, the standalone "Operator" product has been sunset, and its underlying computer-using capabilities have been integrated directly into ChatGPT as an "agent mode."
Capabilities
- GUI Navigation: Capable of viewing screens, controlling a mouse, and typing to interact with graphical user interfaces just like a human.
- Web Automation: Excels at navigating complex websites, filling out forms, and executing multi-step workflows.
- Adaptive Planning: Uses reinforcement learning to break tasks into plans and self-correct when encountering errors during tool use.
- Safety Guardrails: Includes mechanisms to refuse high-risk tasks and hand control back to the user when sensitive input (e.g., passwords) is required.
Example Use Cases
- Automating repetitive web-based tasks like scheduling appointments, booking flights, or data entry.
- Testing user interfaces and web applications autonomously.
- Assisting users by taking control of the browser to execute complex workflows across multiple software applications.
Performance & Benchmarks
At the time of its initial release in 2025, the model achieved state-of-the-art results on several key benchmarks for computer and web agents:
- OSWorld (GUI Agent): 38.1% success rate
- WebArena (Web Agent): 58.1% success rate
- WebVoyager: 87% success rate
Note: With its integration into newer models (like the GPT-5 family) by 2026, these baseline capabilities have significantly improved.
Intended Use & Limitations
- Intended Use: Designed for users needing advanced, agentic automation of computer interfaces and web browsers. It may struggle with highly dynamic or non-standard GUIs and relies heavily on visual context.
About OpenAI
OpenAI is an AI research and deployment company behind groundbreaking models like ChatGPT, DALL-E, and Sora. Their mission is to ensure that artificial general intelligence benefits all of humanity, pioneering advances in autonomous agents and multi-modal models.
Key Features
GUI screen navigation, mouse control, and typing
Optimized for OSWorld and WebArena tasks
Reinforcement learning for tool use and goal-directed agent workflows
High prompt-adherence for multi-step browser automations
You might also want to compare
Verified Sources
Tags
Model Specs
Parameters
Undisclosed
Context Window
128K
License
Proprietary
Deployment
Cost Tiers
Resources & Links
Lineage
Model Family
Part of the Operator family
Only release in this line currently tracked.
Curator Notes
Launched under the 'Operator' research preview. Focuses on browser-based GUI automation and action planning.
Compare Specs
Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.
Compare Model