// verdict · checked 2026-08-11 · Cursor 3.6+ run modes

Cursor (agent mode)

Anysphere

5/5
AGENT

An IDE that argues back. Agent mode earns the name, the tab key does the selling.

CAPABILITY
5/5 AGENT
PROOF
OPERATOR INSPECTABLEYou can run it and read the steps yourself — the record exists for whoever runs it
TRACE ACCESS
OPERATOR-OWNEDWhoever runs it holds the record; nothing is published
LEASH
Not measuredno execution record scored — measure it yourself

Compare → Download the card Bring a trace →

01

The five questions

the rules →
LOOPSruns until doneyes Docs state 'There is no limit on the number of tool calls Agent can make during a task', so it keeps going. OPERATOR INSPECTABLE source · confidence 5/10 · checked 2026-08-11 · ruled by hand
CHOOSESpicks its own toolsyes Docs list its tools — file edits, codebase search, terminal execution, browser — and the model picks each. OPERATOR INSPECTABLE source · confidence 5/10 · checked 2026-08-11 · ruled by hand
ACTSchanges the worldyes Edits files without approval (except config files) and executes terminal commands, per the security docs. OPERATOR INSPECTABLE source · confidence 5/10 · checked 2026-08-11 · ruled by hand
RECOVERSfixes its own errorsyes Agent executes terminal commands and monitors their output with no cap on tool calls in a task, so a failed command is read and answered with a different edit inside the same run. OPERATOR INSPECTABLE source · confidence 6/10 · checked 2026-08-11 · ruled by hand
UNSUPERVISEDnobody watchingyes Cursor's Run Modes docs document "Run Everything" mode: "Every tool call runs automatically" with zero prompts. Rule 05: a documented no-approval mode = Yes; the approval default doesn't force No. OPERATOR INSPECTABLE source · confidence 9.7/10 · checked 2026-08-11 · ruled by hand

Capability decisions combine documented behaviour and execution evidence. The proof label on each row shows which of the two decided it, and Leash is reported separately, only when it was measured from an execution record. Primary source: cursor.com. Tested against: Cursor 3.6+ run modes. Our confidence in this dossier: 5/10.

02

Can you see it work?

the weakest of the five grades above

OPERATOR INSPECTABLE — You can run it and read the steps yourself — the record exists for whoever runs it. Trace access: OPERATOR-OWNED — Whoever runs it holds the record; nothing is published.

Headless CLI emits NDJSON tool_call started/completed events; sessions stored locally in SQLite.

Where we looked: Cursor CLI docs (output-format, headless, using); ran cursor-agent --help locally; inspected ~/.cursor/chats/<ws>/<session>/store.db SQLite session store. · see the evidence → · re-checked by a second pass built to overturn it

03

Revision history

rss for this product
  1. 2026-08-11
    4/5 AGENT → 5/5 AGENTUNSUPERVISED corrected to yes: Cursor's Run Modes docs document "Run Everything" mode: "Every tool call runs automatically" with zero prompts. Rule 05: a documented no-approval mode = Yes; the approval default doesn't force No.editorial correction — no case was filed
04

Badges

earned, never for sale

PASSES THE AGENT TEST OPERATOR INSPECTABLE VERDICT CORRECTED

These render from the live dataset, so a badge cannot outlive the verdict it claims. If the verdict changes the badge changes with it, and an unearned one returns 409 rather than an image. No sponsor can buy one, at any price.

05

Think this is wrong?

bring evidence, not opinion

Name the criterion and link a trace, a doc, or a recorded run that shows the behaviour. A verdict changed by evidence is the best thing that can happen to this site, and the change goes on the public record with your reason attached. Vendors are welcome, and a vendor's own filing is labelled as such rather than buried.