// for anyone writing about this
The finding
Everything a reader, a journalist or a vendor needs in one page: what we found, how we found it, every time we were wrong, and the raw data.
The finding
- Five yes/no questions: loops · chooses · acts · recovers · unsupervised. 4–5 yes = AGENT, 2–3 = WRAPPER WITH AMBITION, 0–1 = CRON JOB WITH EXTRA STEPS.
- Three axes kept apart: capability (yes/no/uncertain), proof (how directly it can be checked), and leash (a measured number). A missing trace lowers the proof grade, never the capability.
- Verdicts come from vendor documentation, read by AI research agents and adjudicated by hand against the published rules.
- Every criterion a verdict depends on was reviewed individually before publication.
- Frameworks are scored on what they do out of the box — a category distinction, not a product failure.
- Anyone can overturn a verdict with evidence. Cases and rulings are public in the court.
Five results that surprised us
Every correction
We publish these because a rating that never moves is a rating nobody checked. Each row is a verdict we got wrong and the evidence that changed it.
Take the data
CC BY 4.0The full dataset — every criterion, its result, reasoning, source, proof grade, confidence and check date — is open and CORS-enabled.
Copy-ready
written by us, posted by nobody automaticallyIf you want to write about this, take the text. Each block is filled from the live dataset, so it cannot quote a number this site no longer stands behind. Nothing here is posted or emailed by any code we run — that is a human action, taken by a human.
Vendors: if your product is in here and you think a criterion is wrong, the court is the front door, and a vendor's own filing is labelled as such rather than buried.