// verdict · checked 2026-08-11 · deep research API (o3-deep-research / o4-mini-deep-research)
Deep Research
OpenAI
Reads 40 tabs so you don't have to. Doesn't touch anything — the librarian of agents.
- CAPABILITY
- 4/5 AGENT
- PROOF
- OPERATOR INSPECTABLEYou can run it and read the steps yourself — the record exists for whoever runs it
- TRACE ACCESS
- OPERATOR-OWNEDWhoever runs it holds the record; nothing is published
- LEASH
- Not measuredno execution record scored — measure it yourself
| LOOPSruns until doneyes | The model itself plans and extends a multi-step search trajectory over 5-30 minutes; max_tool_calls is only an optional developer ceiling, not a pre-fixed sequence. |
| CHOOSESpicks its own toolsyes | It issues its own queries at runtime: the response output array logs the web_search_call actions it picked (search/open_page/find_in_page) plus code_interpreter, file_search and MCP calls. |
| ACTSchanges the worldno | Output is a cited report; side effects stay read-only: search, browse, sandboxed code, no external writes. |
| RECOVERSfixes its own errorsyes | OpenAI states RL training taught it to backtrack and pivot when a path proves unfruitful, reacting to what it encounters; recovery sits in the model, not in builder-wired retry logic. |
| UNSUPERVISEDnobody watchingyes | No documented per-action approval exists: ChatGPT runs 5-30 min while you step away, and the API forces require_approval "never" for MCP, calling human-in-the-loop unsupported. |
Capability decisions combine documented behaviour and execution evidence. The proof label on each row shows which of the two decided it, and Leash is reported separately, only when it was measured from an execution record. Primary source: developers.openai.com. Tested against: deep research API (o3-deep-research / o4-mini-deep-research). Our confidence in this dossier: 7/10.
Can you see it work?
the weakest of the five grades aboveOPERATOR INSPECTABLE — You can run it and read the steps yourself — the record exists for whoever runs it. Trace access: OPERATOR-OWNED — Whoever runs it holds the record; nothing is published.
API response.output carries web_search_call, code_interpreter_call, reasoning
Where we looked: OpenAI cookbook deep research API intro: response.output items typed web_search_call, code_interpreter_call and reasoning summaries. ChatGPT help pages 403'd. · see the evidence →
No revisions yet. This verdict has stood since 2026-08-11.
Badges
earned, never for sale
These render from the live dataset, so a badge cannot outlive the verdict it claims. If the verdict changes the badge changes with it, and an unearned one returns 409 rather than an image. No sponsor can buy one, at any price.
Think this is wrong?
bring evidence, not opinionName the criterion and link a trace, a doc, or a recorded run that shows the behaviour. A verdict changed by evidence is the best thing that can happen to this site, and the change goes on the public record with your reason attached. Vendors are welcome, and a vendor's own filing is labelled as such rather than buried.