// verdict · checked 2026-08-11 · Agent OS 2.0 / Agent SDK
Sierra
Sierra
Sierra is an agent. It passes 4 of the five criteria in our AI agent test; 1 criterion is uncertain. This is a dated assessment of the shipped product, not a claim about every possible setup.
Support flows with real tool use. The suits say agent, the logs half agree.
- CAPABILITY
- 4/5 AGENT1 of the five could not be decided either way
- PROOF
- UNVERIFIEDWe could not confirm this either way. Not a claim.
- TRACE ACCESS
- OPERATOR-OWNEDWhoever runs it holds the record; nothing is published
- LEASH
- Not measuredno execution record scored — measure it yourself
- COMMUNITY
- Loading votes…cast yours on the full list
| LOOPSruns until doneyes | Agent SDK composes skills (triage, respond, confirm) into multi-step workflows toward a defined goal. |
| CHOOSESpicks its own toolsyes | Vendor pages: you set the goal, the SDK directs which resources to use; per-workflow determinism is tunable. |
| ACTSchanges the worldyes | Agents act in connected systems: vendor pages cite updating a subscription or submitting a warranty. |
| RECOVERSfixes its own errorsuncertain | The supervisor pattern cited for this is oversight correcting the agent, not the agent working out a different attempt; no documented tool-call retry says otherwise. The evidence supports neither answer. |
| UNSUPERVISEDnobody watchingyes | Runs live customer chat and voice with no per-step approval; Sierra bills per autonomous resolution. |
Capability decisions combine documented behaviour and execution evidence. The proof label on each row shows which of the two decided it, and Leash is reported separately, only when it was measured from an execution record. Primary source: sierra.ai. Tested against: Agent OS 2.0 / Agent SDK. Our confidence in this dossier: 5/10.
Can you see it work?
the weakest of the five grades aboveUNVERIFIED — We could not confirm this either way. Not a claim.. Trace access: OPERATOR-OWNED — Whoever runs it holds the record; nothing is published.
Agent Traces show operators each step: tool calls, supervisor checks, API calls, timings.
Where we looked: sierra.ai/blog/agent-traces incl. its product screenshot (read the image), Agent SDK + optimize product pages, Agent Data Platform post; docs.sierra.ai still sign-in gated. · see the evidence → · re-checked by a second pass built to overturn it
-
VERDICT OVERTURNED
RECOVERS yes → uncertain filed by editorial internal audit · none — filed by the editors against their own verdict · 2026-08-11
The evidence behind the YES was supervisors reviewing responses in flight and stepping agents back on track — oversight correcting the agent, not the agent working out a different attempt. No tool-call retry is documented either way, so rule 04 can be neither met nor refused on this record. This is the only ruling of the six that moved a score down.
read the case →
- 2026-08-114/5 AGENT → 5/5 AGENTRECOVERS corrected to yes: Supervisors catch faulty output in-flight and drive a revised attempt with no human; model failover alone (preset list, same call) would not qualify.Verdict overturned — case #0006 →
- 2026-08-115/5 AGENT → 4/5 AGENTRECOVERS corrected to uncertain: the evidence behind the yes was supervisors steering the agent, which is oversight rather than self-recovery, and no retry is documented either way. Rule 04 can be neither met nor refused on this record.Verdict overturned — case #0006 →
Badges
earned, never for sale
These render from the live dataset, so a badge cannot outlive the verdict it claims. If the verdict changes the badge changes with it, and an unearned one returns 409 rather than an image. No sponsor can buy one, at any price.
Think this is wrong?
bring evidence, not opinionName the criterion and link a trace, a doc, or a recorded run that shows the behaviour. A verdict changed by evidence is the best thing that can happen to this site, and the change goes on the public record with your reason attached. Vendors are welcome, and a vendor's own filing is labelled as such rather than buried.