This website uses cookies

Read our Privacy policy and Terms of use for more information.

Last tested: August 2026 · Models tested: 12 (full version list below) · How we test

Disclosure: Tusk Central AI is our product, and every model ranked below is available on it. That's also how we tested: same prompts, same system prompt, default settings, versions and dates logged. Full transcripts are linked in each section so you can grade our grading.

"Best AI model" is the wrong question, and every ranking that pretends otherwise is selling you something. The honest question is "best at what?" In mid-2026 the top models are close enough that the winner changes with the task. So instead of one leaderboard, we ran the same 50 prompts through 12 leading models, twice each (1,200 responses total), and ranked them by category: writing, research, reasoning, everyday work, and speed.

Then we did something most rankings skip: we checked whether the sources the models cited actually say what the models claim they say. That test changed the story.

Quick answer

For writing, GPT-5.6 Luna led our tests, narrowly over Kimi K3. For research, GPT-5.4 Mini scored highest (it's a Free, Unlimited model, though see the citation warning below before you trust it). For raw reasoning, eight of twelve models tied at a perfect score, so the flagships have no edge left on textbook problems. For everyday work, GPT-5.6 Luna won again, and it's one of the cheapest models we tested. For speed, Mistral Small 4 answers in under a second, though it's also the only model in the field that regularly gets things wrong.

The single most useful finding: the model that writes the best-looking research answer is not the model with the most trustworthy sources. Those are two different rankings, and most readers never notice the difference.

How we tested

We ran 10 prompts per category through each model. Identical wording, identical system prompt, default settings, 2 runs each, in August 2026. Prompts came from real tasks: actual emails to rewrite, actual documents to summarize, actual math and logic problems, actual research questions with checkable answers.

Scoring worked in three layers:

  • Objective prompts (math, logic, verifiable facts, spreadsheet formulas) were graded automatically against a pre-written answer key.

  • Subjective prompts were scored blind by three independent AI judges from labs with no model in this ranking (MiniMax M3, NVIDIA Nemotron 3 Ultra, and Qwen 3.7 Max) on accuracy, usefulness, and clarity, 1–5 each. We take the median of the three.

  • The 160 responses where the judges disagreed were re-scored by a human editor, whose scores override the machines.

We deliberately excluded every lab in the ranking from judging. A ranking where one contestant grades the others is not a ranking.

We measured speed as median server-side generation time, not round-trip time, so network conditions don't pollute the numbers. Cost is the real per-response price, converted to Tusk credits.

One limitation, stated plainly: one human editor adjudicated the contested responses, not two. That means there's no inter-rater reliability figure to report. The three AI judges scored every subjective response independently; the human resolved the cases they split on.

The lineup

We tested the current flagships and the strongest everyday models:

Model

Lab

On Tusk

GPT-5.6 Luna

OpenAI

Uses Credits

GPT-5.4 Mini

OpenAI

Free, Unlimited

Claude Opus 5

Anthropic

Uses Credits

Claude Haiku 4.5

Anthropic

Free, Unlimited

Gemini 3.6 Flash

Google

Uses Credits

Grok 4.3

xAI

Free, Unlimited

DeepSeek V4 Pro

DeepSeek

Free, Unlimited

Kimi K3

Moonshot

Uses Credits

Mistral Small 4

Mistral

Free, Unlimited

Sonar Pro

Perplexity

Uses Credits

Muse Spark 1.1

Meta

Free, Unlimited

GLM 5.2

Z.ai

Free, Unlimited

Seven of the twelve are Free, Unlimited on Tusk. Five Use Credits. Keep that split in mind. It matters for the results.

Best for writing: GPT-5.6 Luna

Test prompts: rewrite a stiff corporate email, draft a LinkedIn post from bullets, match a supplied writing voice, cut a 300-word draft to 120, write a tricky apology, and five more.

#

Model

Score

1

GPT-5.6 Luna

4.88

2

Kimi K3

4.87

3

Muse Spark 1.1 (Free, Unlimited)

4.80

4

Claude Opus 5

4.70

5

Gemini 3.6 Flash

4.68

11

Claude Haiku 4.5 (Free, Unlimited)

3.88

12

Grok 4.3 (Free, Unlimited)

3.52

Luna and Kimi K3 are separated by 0.01, a statistical tie at the top. What separated them from the rest wasn't prose quality so much as restraint. Given "rewrite this stiff email," Luna simply produced the email:

Dear Mr. Johnson,

I'm following up on my email of the 14th regarding the deliverables outlined in our agreement. We have not yet received the materials, so please send them by the end of this business week…

Grok 4.3, which placed last, produced this instead:

Here's a warmer rewrite of the email that keeps a professional tone while using clearer, more conversational language.

Dear Mr. Johnson, …

Grok's actual prose is fine. Its problem is that it narrates itself, opening with a summary of what it's about to do and closing with an offer to do more. Across ten writing prompts that habit cost it more points than any wording choice. If you paste AI output straight into an email, this is the difference that matters, and it's invisible in benchmarks that only score content.

Two caveats on that last-place finish, because it's closer than the table suggests. The gap between Grok and Claude Haiku 4.5 above it comes down almost entirely to two responses to a single prompt, where our editor scored harshly and two of the three AI judges scored the same answers highly. On the judges' scores alone, Grok and Haiku finish level. Read the bottom of this table as "these two were the weakest," not as a settled ordering. And Grok's problem is a formatting habit, not an inability to write.

The surprise: Muse Spark 1.1, a Free, Unlimited model, placed third, ahead of Claude Opus 5, the most expensive model in the test.

Best for: anything you'll paste directly into an email, post, or doc. Runner-up: Kimi K3 (statistically tied, but 5× slower). Skip if: you need it instantly. Luna is mid-pack on speed.

Best for research: GPT-5.4 Mini — with a warning

Test prompts: five questions with verifiable answers (dates, figures, current facts) and five "summarize and cite" tasks.

#

Model

Score

Facts correct

1

GPT-5.4 Mini (Free, Unlimited)

99.2

100%

2

Claude Opus 5

98.8

100%

3

Sonar Pro

97.5

100%

4

GPT-5.6 Luna

97.1

100%

5

Claude Haiku 4.5 (Free, Unlimited)

96.7

100%

Every model got 100% of the verifiable facts right. Eiffel Tower completion year, boiling point, 2020 Nobel laureates, mile-to-kilometre conversion. Nobody missed. On checkable facts, this class of model is solved.

So we checked the citations instead. We extracted all 2,105 cited URLs, verified which ones resolve, then fed 708 claim-and-source pairs to a neutral judge to ask a harder question: does the page actually say what the model claims it says?

Models don't invent sources; they mangle addresses. About 80% of cited URLs resolve on the first try. We hand-checked every dead one, searching each cited domain for the material the model claimed was there. In eight of ten cases the source provably exists, published by exactly the organisation named, at a slightly different address.

Muse Spark cited a Harvard Health article on coffee at a dead URL. The article is real, it's on Harvard Health, and it says precisely what the model said it says. The link just used Harvard's retired URL format. Sonar Pro cited a 2017 BMJ umbrella review at pmc.ncbi.nl.nih.gov; the host is missing a letter. Add the "m" and the paper loads. Gemini cited a Google documentation page on HTTPS that doesn't exist, but Google does publish that guidance, at a different path.

Across 2,105 citations we could not confirm a single fabricated source. One URL remains unverifiable. That's it. If you Google the title of anything these models cited, you'll find it.

Unsupported citations are the real problem: roughly one research citation in twelve points to a real, live page that does not support the claim attached to it. That's far more common than a fabricated link, and much harder to catch, because the link works.

Here is where the ranking inverts:

Model

Citations that don't support the claim

GLM 5.2 (Free, Unlimited)

0.0%

Kimi K3

2.0%

Mistral Small 4 (Free, Unlimited)

2.6%

Muse Spark 1.1 (Free, Unlimited)

2.8%

Claude Haiku 4.5 (Free, Unlimited)

6.4%

DeepSeek V4 Pro (Free, Unlimited)

7.7%

Claude Opus 5

7.9%

Sonar Pro

7.9%

GPT-5.6 Luna

9.5%

Grok 4.3 (Free, Unlimited)

16.3%

GPT-5.4 Mini

17.1%

(Gemini 3.6 Flash scored worse still, but on too few citations to report responsibly.)

The model that won this category has the least reliable sources in it. GPT-5.4 Mini took first place on answer quality and misattributes roughly one citation in six, double the average for the field and the worst rate of any model we sampled enough to report.

None of that is a contradiction. It just shows what our judges were actually grading: structure, coverage, apparent authority. They never clicked the links. Neither do most readers. A model can write a superb research answer with sources that don't back it, and everyone scores it highly.

The reverse also holds, and it's the more useful half: Kimi K3, GLM 5.2, and Muse Spark all misattribute under 3%. They didn't win on prose. They're the ones whose footnotes you can hand to a fact-checker.

Best for: research answers you will personally verify. Runner-up: Claude Opus 5 (second on quality at 98.8, 7.9% unsupported), the best balance of the two. Skip if: you plan to reuse the model's sources without opening them. For that, Kimi K3 is the honest recommendation (2.0% unsupported) despite finishing ninth on answer quality.

Read the full citation audit, including every one of the ten flagged URLs verified by hand and the two bugs we found in our own testing. All 1,200 transcripts and the raw data are available to download.

Best for reasoning: an eight-way tie

Test prompts: four math word problems, three logic puzzles, three multi-step planning tasks, all with known answers.

Score

Models

100%

Claude Haiku 4.5 (Free, Unlimited), DeepSeek V4 Pro (Free, Unlimited), Gemini 3.6 Flash, Muse Spark 1.1 (Free, Unlimited), Kimi K3, GPT-5.6 Luna, Grok 4.3 (Free, Unlimited), GLM 5.2 (Free, Unlimited)

95%

Claude Opus 5

90%

GPT-5.4 Mini, Sonar Pro

70%

Mistral Small 4

Eight of twelve models scored perfectly, and five of those eight are Free, Unlimited. The flagships that Use Credits did not separate from the Free, Unlimited reasoning models. On this kind of problem, the gap is gone. Claude Opus 5, the most expensive model in the test at 36.8 credits per response, finished ninth.

The one problem that separated the field was a layover puzzle: land at 1:35, connection at 3:00, 25 minutes for customs, 15 to the gate, boarding closes 20 minutes before departure. Do you make it? (You do, with 25 minutes to spare.) Several capable models compared the arrival time to the departure time rather than the boarding cutoff and concluded you'd miss the flight.

Mistral Small 4 is the clear outlier at 70%. It failed basic syllogism logic. Told that all roses are flowers and some flowers fade quickly, it concluded that some roses fade quickly, which doesn't follow.

Best for: anything with a right answer. Pick on speed or price instead, since eight models are tied on correctness. Runner-up: any of the other seven. Skip if: you're using Mistral Small 4 for logic.

Best for everyday work: GPT-5.6 Luna

Test prompts: spreadsheet formulas from a description, a meeting agenda from rough notes, extracting deadlines from a contract, explaining jargon plainly, building a constrained itinerary.

#

Model

Score

1

GPT-5.6 Luna

99.3

2

Kimi K3

98.3

3

Muse Spark 1.1 (Free, Unlimited)

96.2

3=

Grok 4.3 (Free, Unlimited)

96.2

5

Gemini 3.6 Flash

94.8

12

Mistral Small 4

79.5

This is the boring-but-real category, and it's the most useful answer in this post for most people. Luna won it decisively.

Note Grok 4.3 tying for third here after placing last in writing. The habit that ruined its emails (announcing what it's about to do, offering follow-ups) is harmless or even helpful when you're asking for a formula or an agenda. Task fit beats general capability.

The hardest prompt was a constrained Paris itinerary: start at 9:00, end by 17:00, include the Louvre, include a one-hour lunch, under 30 minutes total walking. Models that nailed the prose routinely broke a constraint. Our editor's most common note was some version of "scheduled past 17:00." If you use AI for planning, check the constraints yourself. The writing quality tells you nothing about whether the plan is valid.

Best for: formulas, agendas, summaries, extraction: the daily grind. Runner-up: Kimi K3 (excellent, but 16 seconds per answer). Skip if: cost matters at volume, though Luna is unusually cheap for its tier.

Fastest usable answer: Mistral Small 4 (with an asterisk)

Median server-side generation time across all 50 prompts:

#

Model

Median

Accuracy

1

Mistral Small 4 (Free, Unlimited)

0.95s

88%

2

GPT-5.4 Mini (Free, Unlimited)

1.27s

96%

3

Claude Haiku 4.5 (Free, Unlimited)

2.44s

96%

4

GPT-5.6 Luna

3.01s

100%

5

Sonar Pro

3.18s

95%

The three fastest models in the field are all Free, Unlimited.

But "fastest" and "fastest usable" differ. Mistral Small 4 is quickest and cheapest, and the least accurate model tested. The honest pick is GPT-5.4 Mini: a third of a second slower, Free, Unlimited, and 8 points more accurate.

At the other end, the Free, Unlimited reasoning models are slow: DeepSeek V4 Pro takes 15 seconds, Kimi K3 takes 16. That's the real trade for unlimited free reasoning: not quality, but waiting.

The overall picture

#

Model

Blended

Speed

Credits/response

1

GPT-5.6 Luna

98.6

3.01s

0.7

2

Kimi K3

97.0

15.97s

19.0

3

Muse Spark 1.1 (Free, Unlimited)

96.4

8.90s

7.1

4

Gemini 3.6 Flash

95.7

6.85s

17.7

5

Claude Opus 5

94.9

8.04s

36.8

6

Sonar Pro

93.8

3.18s

19.7

7

DeepSeek V4 Pro (Free, Unlimited)

93.4

15.09s

4.0

8

GLM 5.2 (Free, Unlimited)

92.5

10.61s

3.7

9

GPT-5.4 Mini (Free, Unlimited)

92.1

1.27s

1.1

10

Claude Haiku 4.5 (Free, Unlimited)

89.3

2.44s

4.0

11

Grok 4.3 (Free, Unlimited)

88.7

6.64s

5.0

12

Mistral Small 4 (Free, Unlimited)

82.1

0.95s

0.5

No model swept all five categories. Luna came closest (first in writing, everyday work, and overall) but finished fourth in research and tied-first in reasoning alongside seven others.

Price and quality have decoupled. Luna won overall at 0.7 credits per response. Claude Opus 5 finished fifth at 36.8, over 50× the price for a lower score. Whatever you're paying for at the top of the market, on these tasks it isn't accuracy.

Free, Unlimited models are genuinely competitive, but not equal. Across everything, models that Use Credits averaged 98.7% on objective tasks against 96.8% for Free, Unlimited ones: a real gap, and a small one. Muse Spark 1.1 finished third overall, ahead of Opus 5 and Sonar Pro. GPT-5.4 Mini won the research category outright. Both are Free, Unlimited. But the bottom of the table is also mostly Free, Unlimited models, so "free is fine" is too simple. The right Free, Unlimited model for your task is competitive with anything. The wrong one is the worst model in the field.

Which is, candidly, why Tusk exists. When the best model changes with the task, the useful thing isn't a subscription to one of them. It's one place with all of them and a router that picks per prompt. Every model in this ranking is on Tusk, and seven of the twelve are Free, Unlimited.

Frequently asked questions

What is the best AI model in 2026? It depends on the task. GPT-5.6 Luna scored highest overall and won writing and everyday work. GPT-5.4 Mini, which is Free, Unlimited, won research on answer quality. Eight models tied at a perfect score on reasoning, so nothing separates them there. For fast answers, GPT-5.4 Mini gives the best speed-to-accuracy ratio. If you need citations you can trust without checking, Kimi K3 had by far the most reliable sources.

Are Free, Unlimited AI models as good as ones that Use Credits? On accuracy, close. Free, Unlimited models averaged 96.8% on objectively gradable tasks versus 98.7% for those that Use Credits. A Free, Unlimited model (Muse Spark 1.1) finished third overall, and another won the research category. But Free, Unlimited models also occupy most of the bottom of the table. The spread among them is much wider than the gap between the two tiers. Choose by task, not by price tag.

Can I trust the sources AI models cite? Mostly, yes. The sources are real; that part checked out. What's worth verifying is whether the page actually backs up the claim. We could not confirm a single fabricated source across 2,105 citations. Roughly ten URLs were dead, and hand-checking showed almost all were real documents at wrong addresses. The genuine problem is different: about one citation in twelve points to a live, real page that doesn't support the claim attached to it. The model that won our research category was the worst offender. Open the links.

How often do these rankings change? Fast. Major model releases ship monthly, which is why this post carries a last-tested date and gets re-run quarterly. Treat any undated AI ranking as expired.

Update log: August 2026: initial test run and publication. Re-test scheduled November 2026. Spotted an error? See our corrections policy.

Reply

Avatar

or to participate