// verdict · checked 2026-08-11 · Copilot cloud agent (formerly 'coding agent'), 2026 docs

Copilot coding agent

GitHub

4/5
AGENT

Assigned an issue like an intern, ships a draft PR like an intern.

CAPABILITY
4/5 AGENT
PROOF
OPERATOR INSPECTABLEYou can run it and read the steps yourself — the record exists for whoever runs it
TRACE ACCESS
OPERATOR-OWNEDWhoever runs it holds the record; nothing is published
LEASH
Not measuredno execution record scored — measure it yourself

Compare → Download the card Bring a trace →

01

The five questions

the rules →
LOOPSruns until doneyes Runs multi-step in its own ephemeral Actions VM — explores repo, plans, edits, runs tests, pushes commits iteratively; step count varies by task, with only a 1-hour session timeout as ceiling. OPERATOR INSPECTABLE source · confidence 8/10 · checked 2026-08-11 · ruled by hand
CHOOSESpicks its own toolsyes Docs state it "will use available tools autonomously, and will not ask for approval before use"; it decides which commands, tests and linters to run as it explores the repo. OPERATOR INSPECTABLE source · confidence 8/10 · checked 2026-08-11 · ruled by hand
ACTSchanges the worldyes It automates branch creation, commit message writing and pushing, and opens a draft pull request — real repository state change, not text for a human to apply. OPERATOR INSPECTABLE source · confidence 9/10 · checked 2026-08-11 · ruled by hand
RECOVERSfixes its own errorsyes Its own test/lint runs plus CodeQL, secret scanning and code review feed back in and it "attempts to resolve issues identified prior to completing the pull request" with no human. OPERATOR INSPECTABLE source · confidence 6/10 · checked 2026-08-11 · ruled by hand
UNSUPERVISEDnobody watchingno Docs: workflows wait for a human to click 'Approve and run workflows'; it can't approve or merge its PR. OPERATOR INSPECTABLE source · confidence 6/10 · checked 2026-08-11 · ruled by hand

Capability decisions combine documented behaviour and execution evidence. The proof label on each row shows which of the two decided it, and Leash is reported separately, only when it was measured from an execution record. Primary source: docs.github.com. Tested against: Copilot cloud agent (formerly 'coding agent'), 2026 docs. Our confidence in this dossier: 6/10.

02

Can you see it work?

the weakest of the five grades above

OPERATOR INSPECTABLE — You can run it and read the steps yourself — the record exists for whoever runs it. Trace access: OPERATOR-OWNED — Whoever runs it holds the record; nothing is published.

PR links to session logs showing Copilot's reasoning and the tools it used

Where we looked: GitHub Docs 'manage and track agents' page and the github/docs source tree (content/copilot/concepts/agents/cloud-agent/about-cloud-agent.md) via the GitHub API. · see the evidence →

03

Revision history

rss for this product

No revisions yet. This verdict has stood since 2026-08-11.

04

Badges

earned, never for sale

PASSES THE AGENT TEST OPERATOR INSPECTABLE

These render from the live dataset, so a badge cannot outlive the verdict it claims. If the verdict changes the badge changes with it, and an unearned one returns 409 rather than an image. No sponsor can buy one, at any price.

05

Think this is wrong?

bring evidence, not opinion

Name the criterion and link a trace, a doc, or a recorded run that shows the behaviour. A verdict changed by evidence is the best thing that can happen to this site, and the change goes on the public record with your reason attached. Vendors are welcome, and a vendor's own filing is labelled as such rather than buried.