Tutorial · Research
How to audit an AI product claim before repeating it.
AI product pages move fast, benchmarks are easy to cherry-pick, and social posts routinely turn “can do” into “works reliably.” Before buying or recommending a tool, classify the claim and demand the right kind of evidence.
Start by rewriting the claim precisely
“This agent automates research” is too vague to verify. Rewrite it into measurable statements: “The product can browse these source types,” “the current plan includes X runs,” “it can export citations,” or “on our 20-question test, 18 answers were supported by the cited source.”
Precise claims expose which evidence is missing.
Classify the claim
| Claim | Best evidence | Common mistake |
|---|---|---|
| Price / plan limit | Current official pricing/terms | Using an old review or screenshot |
| Feature availability | Official docs + direct account check | Assuming a launch announcement still reflects the product |
| Speed / accuracy | Reproducible direct test or rigorous independent benchmark | Repeating vendor benchmark as universal performance |
| User satisfaction | Recent reviews/community sample | Presenting a few anecdotes as representative |
| “Saves X hours” | Before/after workflow measurement | Counting generation time but ignoring corrections |
| “Best” | Explicit criteria + relevant alternatives | Ranking by commission or popularity |
Use primary sources for current facts
For pricing, licensing, model availability, usage limits and product policy, begin with the vendor's current documentation. Record the date. SaaS facts decay quickly; a perfectly accurate article can become wrong after one pricing-page update.
Secondary sources are more useful for discovering questions and failure modes than for establishing current product terms.
Use community sources for objections, not truth by vote
Reddit, forums, app reviews and social media are valuable because users reveal what breaks after the demo: billing surprises, difficult onboarding, unreliable integrations, missing exports, quality regressions or support issues. But community comments are anecdotal. Look for repeated patterns across recent independent threads, then test important objections yourself where possible.
Reproduce performance claims on your task
If the purchase depends on speed, accuracy or output quality, vendor benchmarks are not enough. Create a representative test set and define pass/fail before running it. Save inputs, outputs and settings. Repeat enough cases that one lucky result does not dominate the conclusion.
For generative tools, record correction time. A model that produces a better first draft but requires more cleanup may lose the production benchmark.
Look for denominator tricks
Marketing statistics can be technically true but operationally weak:
- “2× faster” — faster than what baseline?
- “90% accuracy” — on what dataset, class balance and metric?
- “used by 10,000 teams” — active paying teams or historical signups?
- “unlimited” — subject to what fair-use/rate limit?
- “free” — free software, free hosted service, or free until you need commercial rights?
Ask for the denominator and the scope.
Separate capability from reliability
A tool that successfully completes a task once has demonstrated capability. Reliability requires repeated success under realistic variation. This is especially important for AI agents: one viral demo can prove possibility while saying almost nothing about unattended production use.
Check the negative space
Product pages emphasize what exists, not what is missing. Search documentation and user reports for:
- export limitations;
- rate limits;
- data retention/privacy constraints;
- unsupported regions or payment methods;
- API access by plan;
- commercial-use restrictions;
- integrations that are read-only rather than read/write;
- features labeled beta/preview.
Maintain a claim ledger
claim_id: c-014 claim: “Starter includes commercial use” type: plan entitlement source: official pricing page checked: 2026-09-08 status: verified-current review_after: 30 days notes: verify again before monetized publication
A ledger looks excessive until you maintain dozens of product pages. Then it becomes the fastest way to know what needs rechecking.
Publication rule
If evidence supports only a narrower statement than the exciting marketing claim, publish the narrower statement. Trust compounds; hype has to be reacquired on every post.