// verdict · checked 2026-08-11 · Operator research preview (Jan 2025); folded into ChatGPT agent, standalone site sunset July 2025
Operator
OpenAI
Operator is an agent. It passes 4 of the five criteria in our AI agent test. This is a dated assessment of the shipped product, not a claim about every possible setup.
Clicks through the web like your dad, but it does decide where to click.
- CAPABILITY
- 4/5 AGENT
- PROOF
- OPERATOR INSPECTABLEYou can run it and read the steps yourself — the record exists for whoever runs it
- TRACE ACCESS
- OPERATOR-OWNEDWhoever runs it holds the record; nothing is published
- LEASH
- Not measuredno execution record scored — measure it yourself
- COMMUNITY
- Loading votes…cast yours on the full list
| LOOPSruns until doneyes | CUA repeats a screenshot→chain-of-thought→click/type loop "until a given task is completed or further input is required"; the 400-step eval cap is a ceiling, not a fixed script. |
| CHOOSESpicks its own toolsyes | OpenAI states the model "can choose to open a page using the text browser or visual browser," then run a terminal command — tool and next UI action picked at runtime from the current screen. |
| ACTSchanges the worldyes | Drives a real browser with mouse/keyboard on live sites — clicks, fills and submits forms, orders groceries and tickets — plus terminal commands and file writes on its virtual computer. |
| RECOVERSfixes its own errorsyes | System card p.20: after failing to install a biodesign tool, the agent "researched and wrote substitute scripts" and continued unaided — a different route it chose, not builder-wired retry. |
| UNSUPERVISEDnobody watchingno | Watch mode pauses when you go inactive, and it asks confirmation before purchases or sending mail. |
Capability decisions combine documented behaviour and execution evidence. The proof label on each row shows which of the two decided it, and Leash is reported separately, only when it was measured from an execution record. Primary source: cdn.openai.com. Tested against: Operator research preview (Jan 2025); folded into ChatGPT agent, standalone site sunset July 2025. Our confidence in this dossier: 4/10.
Can you see it work?
the weakest of the five grades aboveOPERATOR INSPECTABLE — You can run it and read the steps yourself — the record exists for whoever runs it. Trace access: OPERATOR-OWNED — Whoever runs it holds the record; nothing is published.
Steps are readable in the ChatGPT UI and kept in the chat; no export, and audit logs omit actions.
Where we looked: Bypassed the 403s via a reader proxy: OpenAI's ChatGPT agent launch post, the ChatGPT agent help-center article (retention + Compliance API sections), Operator launch post. · see the evidence → · re-checked by a second pass built to overturn it
-
VERDICT OVERTURNED
RECOVERS no → yes filed by editorial internal audit · none — filed by the editors against their own verdict · 2026-08-11
Per-action screenshot feedback lets it pick a different next step, and the system card shows it writing substitute scripts after a tool failed — a different action, taken with no human in the loop. Score moved 3/5 to 4/5, WRAPPER WITH AMBITION to AGENT.
read the case →
- 2026-08-113/5 WRAPPER WITH AMBITION → 4/5 AGENTRECOVERS corrected to yes: Per-action screenshot feedback lets it pick a different next step; system card shows it writing substitute scripts after a tool failed.Verdict overturned — case #0002 →
Badges
earned, never for sale
These render from the live dataset, so a badge cannot outlive the verdict it claims. If the verdict changes the badge changes with it, and an unearned one returns 409 rather than an image. No sponsor can buy one, at any price.
Think this is wrong?
bring evidence, not opinionName the criterion and link a trace, a doc, or a recorded run that shows the behaviour. A verdict changed by evidence is the best thing that can happen to this site, and the change goes on the public record with your reason attached. Vendors are welcome, and a vendor's own filing is labelled as such rather than buried.