How we test AI tools (plain-language version)
Published · Updated · Sources checked 2026-10-08
This is the plain-language version of our full methodology (version 0.2, dated 2026-10-08). If the two ever disagree, the full methodology wins.
Where things stand today: no tool has finished testing. Every score on the site reads NOT YET TESTED, and stays that way until a real, archived test is behind it. We never estimate, borrow or invent a score.
1. We rank jobs, not tools
"The best video AI" doesn't exist. "The best AI for a 15-second product ad" does. So every ranking on this site is for one specific job, with its own test tasks and its own weights. A tool can rank high for one job and low for another.
2. Same task, same rules, every tool
For each job we write fixed test tasks: exact prompts, exact source files, exact settings. Every tool gets:
- the same input
- 3 runs (we keep the first 3 results and never pick the best-looking one)
- default settings unless the job needs a specific one, which we log
- a plan a normal buyer would use, with the plan name and date recorded
If we use an account a vendor gave us, we say so on the page. Getting free access never buys a better result.
3. Measured vs judged: kept separate
- Measured results come from code or instruments. Examples: how many words a transcript got wrong, whether text in an image is spelled exactly right, how long a tool took, how much a usable output cost.
- Judged results come from people: how good a video looks, how natural a voice sounds. People rate these blind: they don't know which tool made which output. At least 2 raters (aiming for 3), with at least one domain expert per category.
We show the two kinds of result separately on every page, so you can see what's a number and what's an opinion.
4. How a ranking is built
- Each task produces a score from 1 to 10.
- Tasks roll up into dimensions:
- quality
- accuracy
- speed
- ease of use
- features
- value
- reliability
- customisation
- commercial rights
- support
- integrations
- privacy & data
- Each job weights the dimensions differently. For meeting notes, accuracy and privacy carry the most weight. For cinematic video, quality does.
- The result is a lab score from 0 to 100, with a confidence level (high, medium or low) and the share of weights we actually measured.
- If two tools are too close to call, we show a tie. We don't invent a winner.
5. Privacy is scored from the vendor's own words
For privacy and data handling, we read the vendor's official privacy policy, trust page and terms. We record the exact sentence, link and date for each point: training on your data, retention, deletion, audits and GDPR. If something isn't stated, we record "not stated". We never guess.
6. Everything is archived
Every test run gets its own folder with:
- the outputs
- the prompts
- the plan and model version
- a timestamp
- a cryptographic fingerprint (SHA-256 hash) of every file
so a result can be checked later. Outputs are published with the result wherever the vendor's terms allow it. Where they don't, we say so.
7. Money never touches a score
- No paid placement, ever. Rankings, "best for" badges and being included in a test can't be bought.
- Some links on the site are affiliate links. Affiliate status is stored separately and is never an input to scoring code. Our raters don't see it.
- Vendors can correct facts (a price, a feature). They can't change results. Corrections: see our independence page.
8. Re-tests and freshness
AI tools change monthly, so every score carries the date it was measured. We re-test on a fixed schedule:
| Category | Re-test every |
|---|---|
| Video | 60 days |
| Image and voice | 90 days |
| Presentations and transcription | 120 days |
We re-test sooner when a major new version ships, a model is swapped, or the price of the tested plan changes. Prices and plan facts on official pages are re-checked daily, and a fact counts as stale after 30 days without a check.
9. User reviews (later)
User ratings, when we open them, will be verified and shown separately from lab results. They can move a combined score by at most 20%, and only after enough verified reviews exist. Until then they count for nothing.
When will results appear?
Testing starts in October 2026. First results are planned for the second half of October 2026 (meeting notes and image text accuracy), then video (late October), then voice and presentations (early November). Each ranking page shows its own status.
Related guides
Spotted an error or an outdated fact? Email corrections@taskverdict.com. We'll check it and correct it.