// verdict · checked 2026-08-11 · deep research API (o3-deep-research / o4-mini-deep-research)

Deep Research

OpenAI

4/5
AGENT

Reads 40 tabs so you don't have to. Doesn't touch anything — the librarian of agents.

CAPABILITY
4/5 AGENT
PROOF
OPERATOR INSPECTABLEYou can run it and read the steps yourself — the record exists for whoever runs it
TRACE ACCESS
OPERATOR-OWNEDWhoever runs it holds the record; nothing is published
LEASH
Not measuredno execution record scored — measure it yourself

Compare → Download the card Bring a trace →

01

The five questions

the rules →
LOOPSruns until doneyes The model itself plans and extends a multi-step search trajectory over 5-30 minutes; max_tool_calls is only an optional developer ceiling, not a pre-fixed sequence. OPERATOR INSPECTABLE source · confidence 9/10 · checked 2026-08-11 · ruled by hand
CHOOSESpicks its own toolsyes It issues its own queries at runtime: the response output array logs the web_search_call actions it picked (search/open_page/find_in_page) plus code_interpreter, file_search and MCP calls. OPERATOR INSPECTABLE source · confidence 8/10 · checked 2026-08-11 · ruled by hand
ACTSchanges the worldno Output is a cited report; side effects stay read-only: search, browse, sandboxed code, no external writes. OPERATOR INSPECTABLE source · confidence 7/10 · checked 2026-08-11 · ruled by hand
RECOVERSfixes its own errorsyes OpenAI states RL training taught it to backtrack and pivot when a path proves unfruitful, reacting to what it encounters; recovery sits in the model, not in builder-wired retry logic. OPERATOR INSPECTABLE source · confidence 7/10 · checked 2026-08-11 · ruled by hand
UNSUPERVISEDnobody watchingyes No documented per-action approval exists: ChatGPT runs 5-30 min while you step away, and the API forces require_approval "never" for MCP, calling human-in-the-loop unsupported. OPERATOR INSPECTABLE source · confidence 9/10 · checked 2026-08-11 · ruled by hand

Capability decisions combine documented behaviour and execution evidence. The proof label on each row shows which of the two decided it, and Leash is reported separately, only when it was measured from an execution record. Primary source: developers.openai.com. Tested against: deep research API (o3-deep-research / o4-mini-deep-research). Our confidence in this dossier: 7/10.

02

Can you see it work?

the weakest of the five grades above

OPERATOR INSPECTABLE — You can run it and read the steps yourself — the record exists for whoever runs it. Trace access: OPERATOR-OWNED — Whoever runs it holds the record; nothing is published.

API response.output carries web_search_call, code_interpreter_call, reasoning

Where we looked: OpenAI cookbook deep research API intro: response.output items typed web_search_call, code_interpreter_call and reasoning summaries. ChatGPT help pages 403'd. · see the evidence →

03

Revision history

rss for this product

No revisions yet. This verdict has stood since 2026-08-11.

04

Badges

earned, never for sale

PASSES THE AGENT TEST OPERATOR INSPECTABLE

These render from the live dataset, so a badge cannot outlive the verdict it claims. If the verdict changes the badge changes with it, and an unearned one returns 409 rather than an image. No sponsor can buy one, at any price.

05

Think this is wrong?

bring evidence, not opinion

Name the criterion and link a trace, a doc, or a recorded run that shows the behaviour. A verdict changed by evidence is the best thing that can happen to this site, and the change goes on the public record with your reason attached. Vendors are welcome, and a vendor's own filing is labelled as such rather than buried.