Every comparison, ranking, and "best of" post on this site links back to this page. It explains exactly how we test AI models, what we disclose, and what we do when we get something wrong. If you ever find a post of ours that doesn't live up to this page, tell us. Contact details are at the bottom.
The disclosure that matters most: Tusk Central AI is our product. It aggregates the same AI models we review. We think that makes us more qualified to compare models, not less. But you deserve to know it before you read a single ranking, and every post that mentions Tusk says so.
Who does the testing
Real, named people. Every post on this site carries a byline linking to an author page, and the person named actually ran the tests described. We don't publish anonymous "Team" comparisons, and we don't publish AI-generated reviews of AI models without a human running, checking, and standing behind every claim.
What we test
We test models on the tasks our readers actually do, grouped into five categories: writing (drafting, editing, tone-matching), research (sourced answers, summarization, fact-finding), reasoning (math, logic, multi-step problems), everyday work (emails, spreadsheets, documents, planning), and speed (how fast a usable answer arrives). Published benchmark scores don't settle any of this for us. Benchmarks measure benchmark performance. We care about whether the answer helps a real person finish a real task.
How we test
Identical prompts, verbatim. Every model in a comparison receives exactly the same prompt, character for character.
Default settings. We test models the way a normal person encounters them. No custom system prompts, no temperature tuning, no tricks.
Multiple runs. Models are non-deterministic, so we run prompts more than once and note when results vary meaningfully.
Named versions and dates. Every comparison states which model versions were tested and when. An undated AI comparison is a rumor.
Receipts. Where practical we publish the prompts and full outputs as transcripts, screenshots, or video, so you can check our judgment against the raw material.
How models are accessed
We access models through Tusk Central AI, which serves each provider's models via their official APIs. These are the same underlying models you'd get from the providers directly. Where a provider's own app includes features their API doesn't, we say so in the post.

Tusk Central AI model picker
How we judge results
Reviewers score outputs on accuracy, usefulness (did it actually complete the task?), clarity, and honesty about uncertainty. When judgment calls are close, we say they're close. "It depends on your task" is an acceptable conclusion here. Manufactured winners are not.
Freshness: comparisons expire
AI models change monthly. Every comparison displays a "Last tested" date, and we re-test our major comparisons quarterly, or sooner when a significant model version ships. Material updates get noted in an update log on the post. If a post's last-tested date looks stale, treat its conclusions accordingly (and feel free to nudge us).
What we will never do
Accept payment from an AI provider for placement, ranking, or favorable coverage
Rank Tusk-favorable options higher than testing supports
Quietly rewrite conclusions without an update note
Publish claims we haven't tested as if we had
Corrections
When we get something wrong, we correct the post, date the correction, and note what changed. Spotted an error? Email us at [email protected]. Corrections make the site better, and we treat them as a favor, not an attack.
This page was last updated in July 2026 and is reviewed quarterly alongside our comparisons.