// verdict · checked 2026-08-11 · Gumloop agents (docs.gumloop.com, 2026)

Gumloop

Gumloop

4/5
AGENT

Drag, drop, call it autonomous.

CAPABILITY
4/5 AGENT
PROOF
OPERATOR INSPECTABLEYou can run it and read the steps yourself — the record exists for whoever runs it
TRACE ACCESS
OPERATOR-OWNEDWhoever runs it holds the record; nothing is published
LEASH
Not measuredno execution record scored — measure it yourself

Compare → Download the card Bring a trace →

01

The five questions

the rules →
LOOPSruns until doneyes Agent "decides which tools to use and in what order, runs them, adapts based on the results"; Max Steps caps tool calls before it must respond (default 100, range 1-200) — a ceiling, not a fixed script. OPERATOR INSPECTABLE source · confidence 9/10 · checked 2026-08-11 · ruled by hand
CHOOSESpicks its own toolsyes Docs contrast flows' fixed path with agents that at runtime analyze the request and decide which tools to use and in what order, including when to invoke a flow as a tool. OPERATOR INSPECTABLE source · confidence 9/10 · checked 2026-08-11 · ruled by hand
ACTSchanges the worldyes Code Sandbox is "natively enabled on all agents with no configuration required" and runs Python and shell commands, creates files, and calls connected integrations via the gumloop SDK. OPERATOR INSPECTABLE source · confidence 9/10 · checked 2026-08-11 · ruled by hand
RECOVERSfixes its own errorsno Only model-outage retry/fallback is documented; docs tell the user to review tool errors themselves. OPERATOR INSPECTABLE source · confidence 8/10 · checked 2026-08-11 · ruled by hand
UNSUPERVISEDnobody watchingyes Tool Management preset "Always allow" executes tools without pausing, and schedule/app-event triggers start the agent unattended; note docs never state Always allow is the default preset. OPERATOR INSPECTABLE source · confidence 8/10 · checked 2026-08-11 · ruled by hand

Capability decisions combine documented behaviour and execution evidence. The proof label on each row shows which of the two decided it, and Leash is reported separately, only when it was measured from an execution record. Primary source: docs.gumloop.com. Tested against: Gumloop agents (docs.gumloop.com, 2026). Our confidence in this dossier: 8/10.

02

Can you see it work?

the weakest of the five grades above

OPERATOR INSPECTABLE — You can run it and read the steps yourself — the record exists for whoever runs it. Trace access: OPERATOR-OWNED — Whoever runs it holds the record; nothing is published.

Run Log per execution, run-history API, checkpoints, enterprise audit logs

Where we looked: docs.gumloop.com/llms.txt index: Run Log core concept, checkpoint history, retrieve-run-history and retrieve-run-details API endpoints, enterprise audit logging. · see the evidence →

03

Revision history

rss for this product
  1. 2026-08-11
    3/5 WRAPPER WITH AMBITION → 4/5 AGENTLOOPS corrected to yes: Docs: agent decides which tools and in what order, loops until done or blocked; Max Steps (default 100, max 200) is a ceiling, not a fixed count.editorial correction — no case was filed
04

Badges

earned, never for sale

PASSES THE AGENT TEST OPERATOR INSPECTABLE VERDICT CORRECTED

These render from the live dataset, so a badge cannot outlive the verdict it claims. If the verdict changes the badge changes with it, and an unearned one returns 409 rather than an image. No sponsor can buy one, at any price.

05

Think this is wrong?

bring evidence, not opinion

Name the criterion and link a trace, a doc, or a recorded run that shows the behaviour. A verdict changed by evidence is the best thing that can happen to this site, and the change goes on the public record with your reason attached. Vendors are welcome, and a vendor's own filing is labelled as such rather than buried.