// verdict · checked 2026-08-11 · AutoGPT Platform (agpt.co docs, 2026)

AutoGPT

Significant Gravitas

4/5
AGENT

The 2023 fever dream. Looped magnificently. Toward what, nobody ever found out.

CAPABILITY
4/5 AGENT
PROOF
OPERATOR INSPECTABLEYou can run it and read the steps yourself — the record exists for whoever runs it
TRACE ACCESS
OPERATOR-OWNEDWhoever runs it holds the record; nothing is published
LEASH
Not measuredno execution record scored — measure it yourself

Compare → Download the card Bring a trace →

01

The five questions

the rules →
LOOPSruns until doneyes Smart Decision Maker feeds last_tool_output back each cycle; agent_mode_max_iterations = -1 loops until it emits finished, 1+ caps it — the model decides step count. OPERATOR INSPECTABLE source · confidence 9/10 · checked 2026-08-11 · ruled by hand
CHOOSESpicks its own toolsyes Docs: Smart Decision Maker "uses an LLM to decide which tool to use based on the block's prompt"; AutoPilot calls any of ~400 blocks directly with no wiring. OPERATOR INSPECTABLE source · confidence 9/10 · checked 2026-08-11 · ruled by hand
ACTSchanges the worldyes Shipped blocks change external state: Code Execution "executes code in a sandbox environment with internet access", Gmail Send, and Github Make Pull Request. OPERATOR INSPECTABLE source · confidence 9/10 · checked 2026-08-11 · ruled by hand
RECOVERSfixes its own errorsno Docs: implement retry via your own wiring, not automatic retries; a failed block just fills an error pin. OPERATOR INSPECTABLE source · confidence 8/10 · checked 2026-08-11 · ruled by hand
UNSUPERVISEDnobody watchingyes Schedule Task and webhook Triggers run an agent end-to-end with pre-configured inputs; approval exists only as an opt-in Human In The Loop block a builder must add. OPERATOR INSPECTABLE source · confidence 9/10 · checked 2026-08-11 · ruled by hand

Capability decisions combine documented behaviour and execution evidence. The proof label on each row shows which of the two decided it, and Leash is reported separately, only when it was measured from an execution record. Primary source: agpt.co. Tested against: AutoGPT Platform (agpt.co docs, 2026). Our confidence in this dossier: 8/10.

02

Can you see it work?

the weakest of the five grades above

OPERATOR INSPECTABLE — You can run it and read the steps yourself — the record exists for whoever runs it. Trace access: OPERATOR-OWNED — Whoever runs it holds the record; nothing is published.

self-hosted; run view shows per-task inputs/outputs/cost, node logs partial

Where we looked: agpt.co/docs (GitBook) plus its ask endpoint on execution logs, and the AutoGPT GitHub README: task inputs/outputs/cost per run, error pins, stdout only for code blocks. · see the evidence →

03

Revision history

rss for this product
  1. 2026-08-11
    3/5 WRAPPER WITH AMBITION → 4/5 AGENTACTS corrected to yes: Shipped blocks write real external state at runtime: SendEmailBlock calls smtplib sendmail; GithubMakePullRequestBlock POSTs to /pulls.editorial correction — no case was filed
04

Badges

earned, never for sale

PASSES THE AGENT TEST OPERATOR INSPECTABLE VERDICT CORRECTED

These render from the live dataset, so a badge cannot outlive the verdict it claims. If the verdict changes the badge changes with it, and an unearned one returns 409 rather than an image. No sponsor can buy one, at any price.

05

Think this is wrong?

bring evidence, not opinion

Name the criterion and link a trace, a doc, or a recorded run that shows the behaviour. A verdict changed by evidence is the best thing that can happen to this site, and the change goes on the public record with your reason attached. Vendors are welcome, and a vendor's own filing is labelled as such rather than buried.