How reviews are collected and coded
We read public reviews by people who used each AI product, dated Oct 1, 2024 or later. Each review is quoted word for word and coded on five points. Shares, intervals and ranks are computed from those codes, never typed in.
The five points
Each is coded positive, negative or not mentioned.
- Does the job
- Agents and voice: resolves customer questions or calls on its own. Copilots: drafts, summaries or suggestions are usable. Auto-QA: AI scores match human judgment and save review time.
- Accuracy
- Negative when the review mentions wrong, invented or off-policy answers, or the AI misunderstanding customers. Positive when it praises accuracy.
- Handoff to a person
- Escalation to a person: whether it happens at the right time and the person gets the context. Not coded for auto-QA.
- Cost
- Value for money, the billing model and surprise charges.
- Setup and upkeep
- Ease of setup, training the AI, and keeping it tuned.
Principles
- Real reviews, quoted
- Every number comes from public reviews written by people who used the AI product. Each one is quoted word for word, with the date, the platform and a link to the original.
- About the AI, not the company
- A review counts only if it talks about the AI product. Reviews of the wider help desk, the sales process or the company in general are left out.
- Coded on five points
- Each review is coded positive, negative or not mentioned on five points. Coding is AI-assisted, and every code is published next to its quote, so you can check it.
- Intervals on every share
- Each share comes with its 95% interval (Wilson) and the number of reviews behind it. Below 10 reviews that mention a point, no share is shown.
- Ranked only when it's clear
- Products are ranked inside a category by the share of reviews that say the AI does its job. Products whose intervals overlap share a rank range.
- Reviews can be wrong too
- Reviewers are a self-selected group: unhappy and delighted users write more often. Shares describe what reviewers say, not how the product performs for everyone.
What “does the job” means in each category
| Category | Does the job |
|---|---|
| AI agents | Resolves customer questions on its own |
| Agent-assist copilots | Drafts and suggestions are usable |
| AI quality assurance | Scores match human judgment and save review time |
| Voice AI agents | Handles calls on its own |
Where reviews come from
Shopify App Store, Trustpilot, TrustRadius, Reddit, G2, Capterra and vendor community forums, wherever the page can be opened and quoted. Each product page lists where we looked, including sources that had nothing or would not load.
Shares need 10 reviews that mention a point. Many enterprise products have few public reviews, so they are listed but not ranked.
Left out
- Reviews dated before 2024-10-01.
- Vendor case studies, testimonials and posts by the vendor's staff, partners or agencies.
- Reviews marked as incentivized or sponsored, and obvious spam.
- Quotes from search-result snippets: every quote comes from a page we opened.
- Duplicates: the same person posting the same text on two platforms counts once.
Independence and money
- No paid placements, and no fee to be listed.
- Vendors cannot add, remove or answer reviews here. Corrections to coding are welcome from anyone, with the review link.
- No affiliate links today. If that changes, paid links will be labelled, marked rel="sponsored", and will never change an order.
- We are not affiliated with any vendor we cover.
Limits
Reviews come from a self-selected group. People with strong views write more often, and some platforms attract one kind of buyer (the Shopify App Store, for example, is mostly ecommerce stores). A share tells you what reviewers say, not what every customer gets.
The interval covers sampling noise only. It does not fix a biased sample.
Corrections
If a quote is wrong, a code looks off, or a review should not count, tell us with the link. Confirmed changes are dated in the changelog, and the shares recompute from the corrected codes.