// bring a trace and a verdict changes

The court

Every challenge to a verdict, the evidence behind it, and the ruling — including the ones where we were wrong. A verdict overturned by evidence is the best thing that happens here.

7cases
6overturned
1upheld
0open
01

Cases

6 shown · every ID is permanent
  1. VERDICT OVERTURNED
    Our own leash headline · LEASH the leash is you → mostly, the agent stops and waits for you filed by editorial internal audit — the site against itself · 2026-08-11 · decided 2026-08-11

    The hero said tool calls run ‘before a human interrupts’ — implying the human ends the runs. We audited why 206 runs actually ended: a human ended 7 of them — five Esc interrupts, two permission denials — 3.4%. The agent stopped itself in 72%. Two scoring bugs flattered the agent along the way: permission denials were counted as agent errors, and the agent’s workaround after a human ‘no’ was counted as an unassisted recovery (138 self-fixed becomes 136). Prior art acknowledged: Karpathy coined the leash; Anthropic measured the fleet in February and found agents hand the leash back more often than humans yank it. Our corpus agrees with them, not with our old headline. The metric survives — the story it told did not. Full method: docs/LEASH-AUDIT.md in the repository.

    the evidence → · case #0007 →
  2. VERDICT OVERTURNED
    Sierra · RECOVERS yes → uncertain 5/5 AGENT → 4/5 AGENT filed by editorial internal audit · 2026-08-11 · decided 2026-08-11

    The evidence behind the YES was supervisors reviewing responses in flight and stepping agents back on track — oversight correcting the agent, not the agent working out a different attempt. No tool-call retry is documented either way, so rule 04 can be neither met nor refused on this record. This is the only ruling of the six that moved a score down.

    the evidence → · case #0006 →
  3. VERDICT OVERTURNED
    Amazon Q Developer · RECOVERS no → yes 4/5 AGENT → 5/5 AGENT filed by editorial internal audit · 2026-08-11 · decided 2026-08-11

    Transformation docs: "As it makes changes, it re-builds and runs existing unit tests in your source code to iteratively fix any encountered errors." That is rule 04 met — an error, a different attempt, no human turn. The old NO conceded the loop existed.

    the evidence → · case #0005 →
  4. VERDICT OVERTURNED
    Lovable · UNSUPERVISED no → yes 4/5 AGENT → 5/5 AGENT filed by editorial internal audit · 2026-08-11 · decided 2026-08-11

    Build mode "handles execution end to end with minimal supervision" and the docs place review after the run, not at each step. Rule 05 turns on capability, and the stop button is an interrupt, not an approval gate. The old NO cited the mode working and answered no anyway.

    the evidence → · case #0004 →
  5. VERDICT OVERTURNED
    ChatGPT Tasks · ACTS no → yes filed by editorial internal audit · 2026-08-11 · decided 2026-08-11

    Live docs state tasks can use apps that take write actions in connected services, which changes state outside the conversation. Score moved 1/5 to 2/5, CRON JOB WITH EXTRA STEPS to WRAPPER WITH AMBITION.

    the evidence → · case #0003 →
  6. VERDICT OVERTURNED
    Operator · RECOVERS no → yes filed by editorial internal audit · 2026-08-11 · decided 2026-08-11

    Per-action screenshot feedback lets it pick a different next step, and the system card shows it writing substitute scripts after a tool failed — a different action, taken with no human in the loop. Score moved 3/5 to 4/5, WRAPPER WITH AMBITION to AGENT.

    the evidence → · case #0002 →

Community votes never move a verdict on their own. Evidence does, and the reasoning that moved it is printed above in full.

02

File a challenge

evidence, not opinion

Name the product and the criterion, and link a trace, a doc page, or a recorded run that shows the behaviour. Vendors are welcome, and a vendor's own filing is labelled as such on the case. Sponsorship does not move a verdict and does not move a case up the queue.