// verdict · checked 2026-08-11 · docs as of 2026-08 (CLI 2.1.x)

Claude Code

Anthropic

5/5
AGENT

Loops, picks tools, breaks your tests and then fixes them. The bar, calibrated.

CAPABILITY
5/5 AGENT
PROOF
TRACE VERIFIEDWe scored a real execution record of this ourselves
TRACE ACCESS
OPERATOR-OWNEDWhoever runs it holds the record; nothing is published
LEASH
15 median · max 92 · n=4242 Claude Code sessions · one operator · median 15 · max 92

Compare → Download the card Bring a trace →

01

The five questions

the rules →
LOOPSruns until doneyes Docs' own example: 'write tests for the auth module, run them, and fix any failures' from one prompt. TRACE VERIFIED source · confidence 9/10 · checked 2026-08-11 · ruled by hand
CHOOSESpicks its own toolsyes Picks tools at runtime from Bash/Read/Edit plus any MCP servers; --allowedTools only bounds the set. TRACE VERIFIED source · confidence 9/10 · checked 2026-08-11 · ruled by hand
ACTSchanges the worldyes Docs: it 'edits files, runs commands'; it also stages git changes, commits and opens pull requests. TRACE VERIFIED source · confidence 9/10 · checked 2026-08-11 · ruled by hand
RECOVERSfixes its own errorsyes That same documented task ends in fixing the failures it just ran; docs also show root-cause bug fixing. TRACE VERIFIED source · confidence 9/10 · checked 2026-08-11 · ruled by hand
UNSUPERVISEDnobody watchingyes Four of the 42 scored sessions ran to completion with zero operator interventions, and non-interactive claude -p with --allowedTools runs in CI with no approval prompts. TRACE VERIFIED source · confidence 9/10 · checked 2026-08-11 · ruled by hand

Capability decisions combine documented behaviour and execution evidence. The proof label on each row shows which of the two decided it, and Leash is reported separately, only when it was measured from an execution record. Primary source: code.claude.com. Tested against: docs as of 2026-08 (CLI 2.1.x). Our confidence in this dossier: 9/10.

02

Can you see it work?

the weakest of the five grades above

TRACE VERIFIED — We scored a real execution record of this ourselves. Trace access: OPERATOR-OWNED — Whoever runs it holds the record; nothing is published.

every run logged as JSONL in ~/.claude/projects; 68 local session files verified

Where we looked: Inspected the running install directly: ~/.claude/projects/*/*.jsonl (68 session transcripts on this machine), each holding prompts, tool calls and results.

03

Revision history

rss for this product

No revisions yet. This verdict has stood since 2026-08-11.

04

Badges

earned, never for sale

PASSES THE AGENT TEST TRACE VERIFIED

These render from the live dataset, so a badge cannot outlive the verdict it claims. If the verdict changes the badge changes with it, and an unearned one returns 409 rather than an image. No sponsor can buy one, at any price.

05

Think this is wrong?

bring evidence, not opinion

Name the criterion and link a trace, a doc, or a recorded run that shows the behaviour. A verdict changed by evidence is the best thing that can happen to this site, and the change goes on the public record with your reason attached. Vendors are welcome, and a vendor's own filing is labelled as such rather than buried.