// verdict · checked 2026-08-11 · Amazon Q Developer CLI with custom agents (2025-07+)

Amazon Q Developer

AWS

Amazon Q Developer is an agent. It passes 5 of the five criteria in our AI agent test. This is a dated assessment of the shipped product, not a claim about every possible setup.

5/5
AGENT

Does the work, but you can feel the internal ticketing system through the screen.

CAPABILITY
5/5 AGENT
PROOF
DOCS CLAIMEDThe vendor's documentation states it and we read the page; we did not watch it happen
TRACE ACCESS
OPERATOR-OWNEDWhoever runs it holds the record; nothing is published
LEASH
Not measuredno execution record scored — measure it yourself
COMMUNITY
Loading votes…cast yours on the full list

Compare → Download the card Bring a trace →

01

The five questions

plain-English test · the rules →
LOOPSruns until doneyes Q decomposes the prompt into its own implementation steps and iterates on results - runs a build, reads the errors, fixes, rebuilds - so step count follows the task, not a fixed script. OPERATOR INSPECTABLE source · confidence 8/10 · checked 2026-08-11 · ruled by hand
CHOOSESpicks its own toolsyes At runtime Q picks among native file read/write, shell and AWS-API tools plus any MCP server tools, and judges command risk itself, running ones it deems low-risk on its own. OPERATOR INSPECTABLE source · confidence 8/10 · checked 2026-08-11 · ruled by hand
ACTSchanges the worldyes Agentic coding is on by default and Q "updates your files directly", runs shell commands and calls AWS APIs - real state change on disk and in AWS, not text for a human to apply. OPERATOR INSPECTABLE source · confidence 9/10 · checked 2026-08-11 · ruled by hand
RECOVERSfixes its own errorsyes Transformation docs: as it makes changes it re-builds and runs the existing unit tests to iteratively fix any encountered errors, with review coming only at the diff. DOCS CLAIMED source · confidence 7/10 · checked 2026-08-11 · ruled by hand · overturned in court #0005
UNSUPERVISEDnobody watchingyes Rule 05 counts named flags. `q chat --trust-all-tools --no-interactive` and agent allowedTools `@builtin` run full tasks with no per-action approval; opt-in/discouraged is irrelevant. OPERATOR INSPECTABLE source · confidence 9/10 · checked 2026-08-11 · ruled by hand

Capability decisions combine documented behaviour and execution evidence. The proof label on each row shows which of the two decided it, and Leash is reported separately, only when it was measured from an execution record. Primary source: docs.aws.amazon.com. Tested against: Amazon Q Developer CLI with custom agents (2025-07+). Our confidence in this dossier: 5/10.

02

Can you see it work?

the weakest of the five grades above

DOCS CLAIMED — The vendor's documentation states it and we read the page; we did not watch it happen. Trace access: OPERATOR-OWNED — Whoever runs it holds the record; nothing is published.

Operator can export a full run: /save writes conversation JSON with tool calls and results.

Where we looked: Read the open-source CLI source: /save serializes full ConversationState to JSON; message.rs shows tool_uses (name, args) + tool_use_results (content, Success/Error status). · see the evidence → · re-checked by a second pass built to overturn it

  1. VERDICT OVERTURNED
    RECOVERS no → yes filed by editorial internal audit · none — filed by the editors against their own verdict · 2026-08-11

    Transformation docs: "As it makes changes, it re-builds and runs existing unit tests in your source code to iteratively fix any encountered errors." That is rule 04 met — an error, a different attempt, no human turn. The old NO conceded the loop existed.

    read the case →
04

Revision history

rss for this product
  1. 2026-08-11
    3/5 WRAPPER WITH AMBITION → 4/5 AGENTUNSUPERVISED corrected to yes: Rule 05 counts named flags. `q chat --trust-all-tools --no-interactive` and agent allowedTools `@builtin` run full tasks with no per-action approval; opt-in/discouraged is irrelevant.editorial correction — no case was filed
  2. 2026-08-11
    4/5 AGENT → 5/5 AGENTRECOVERS corrected to yes: transformation docs say it re-builds and runs unit tests to iteratively fix any encountered errors. The old NO conceded the loop existed and answered no anyway.Verdict overturned — case #0005 →
05

Badges

earned, never for sale

PASSES THE AGENT TEST VERDICT OVERTURNED

These render from the live dataset, so a badge cannot outlive the verdict it claims. If the verdict changes the badge changes with it, and an unearned one returns 409 rather than an image. No sponsor can buy one, at any price.

06

Think this is wrong?

bring evidence, not opinion

Name the criterion and link a trace, a doc, or a recorded run that shows the behaviour. A verdict changed by evidence is the best thing that can happen to this site, and the change goes on the public record with your reason attached. Vendors are welcome, and a vendor's own filing is labelled as such rather than buried.