This website uses cookies

Read our Privacy policy and Terms of use for more information.

Companion to Best AI Models 2026, Ranked. This is the complete citation audit behind that article's research section, including every URL we flagged and every error we found in our own testing.

We told you we would publish this so you could grade our grading. Here it is, including the parts where we got it wrong.

What we checked

Across the 1,200 model responses in our benchmark, the models cited 2,105 URLs. We ran two separate checks on them.

Pass 1 asked: does the link work? We resolved every cited URL and recorded what came back.

Pass 2 asked the harder question: does the page actually support the claim it is attached to? We paired 708 claims with the source cited for them, fetched each page, and had a judge from a lab with no model in the ranking rule on whether the source backed the claim.

The second question matters more. A dead link is obvious. A live link to a real page that does not say what the model claims is the failure a reader will never catch.

Status

Unique URLs

Share

Resolves normally

424

79.8%

Blocked our checker (403/429), real but unverifiable this way

59

11.1%

404 or 410, page not found

26

4.9%

Domain does not resolve

9

1.7%

Other network error

13

2.4%

About 6.6% of unique cited URLs were dead. We then hand-checked every one of them.

Three things we deliberately did not count as fabrication:

Bot-blocking is not fabrication. Roughly one URL in nine returns 403 or 429 to an automated checker while being perfectly real. News sites and journals do this routinely.

Link rot is not fabrication. When several models independently cite the same now-dead URL, it was almost certainly real when they learned it. One American Heart Association newsroom link was cited by six different models. Six models do not independently invent the same URL.

Search-infrastructure redirects are not citations. Gemini produced 81 expiring vertexaisearch.cloud.google.com grounding redirects. Counting those as fabricated sources would have badly misrepresented that model.

The ten flagged URLs, verified by hand

After removing rot and blocking, ten URLs remained as fabrication candidates. We searched each cited domain for the material the model claimed was there, then confirmed any replacement with a live check.

Result: eight are confirmed bad links to real, live sources. One is a wrong slug on a genuine site. One is unverifiable. Zero are confirmed inventions.

Model

Cited URL (dead)

Verdict

What we found

Gemini 3.6 Flash

nist.gov/pml/owm/publications/nist-handbooks/handbook-44

Bad link

NIST Handbook 44 is real, at /pml/owm/nist-handbook-44-current-edition. Its Appendix C carries the conversion tables the claim relies on.

Mistral Small 4

nist.gov/pml/owm/unit-conversion

Bad link

Real page is /pml/owm/metric-si/unit-conversion. The model dropped one path segment.

Kimi K3

calcintel.com/calculator/miles-to-km

Wrong slug

calcintel.com is a real, live calculator site using exactly this /calculator/ pattern and it does offer length conversion. That specific slug 404s.

Sonar Pro

all-convert.com/length/mile-to-kilometer/100

Unverifiable

Domain serves nothing and is not indexed. We could not confirm it ever existed. The only invention candidate in the set, and still unproven.

Gemini 3.6 Flash

dataprotection.ie/en/individuals/rights-individuals/...

Bad link

Ireland's Data Protection Commission publishes exactly this material under /en/individuals/know-your-rights/. Right authority, invented path.

Muse Spark 1.1

health.harvard.edu/staying-healthy/coffee-health-risks

Bad link

Harvard Health publishes coffee risk and benefit material at other paths. Right publisher, right topic, wrong address.

Muse Spark 1.1

health.harvard.edu/blog/coffee-more-links-to-health-than-harm-2017111412664

Bad link, strongest case

The article is real: "Coffee: More links to health than harm," now at /staying-healthy/coffee-more-links-to-health-than-harm. It states the largest benefit at three to four cups a day with no further gain above four, which is exactly the claim the model attached to it. Right article, right facts, retired URL format.

Sonar Pro

pmc.ncbi.nl.nih.gov/articles/PMC5696634/

Typo

The host is missing a letter. Add the "m" and you get Poole et al., "Coffee consumption and health: umbrella review," BMJ 2017. It supports the claim.

Gemini 3.6 Flash

developers.google.com/search/docs/crawling-indexing/http-to-https-migration

Bad link

Google does publish this guidance, as "HTTPS as a ranking signal" on its Search Central blog. Right publisher, right claim, invented path.

The models were not making up sources. They were making up addresses. In eight of ten cases the cited work provably exists, is published by the organisation named, and where we checked, actually supports the claim. What failed was URL construction: a dropped path segment, a retired URL format, a missing letter in a hostname.

A reader who searches the cited title will find the real document. A reader who clicks the link gets a 404. Those are different problems, and only one of them is dishonest.

Pass 2: do the sources support the claims?

This is the check that produced the article's headline finding.

Measure

Rate

Support-weighted score (partial credit counts half)

84.6%

Outright unsupported

8.6%

Roughly one research citation in twelve points to a real, live page that does not support the claim attached to it.

By model

Excludes pages we could not read (paywalls, cookie walls, JavaScript shells) and cases where our own extraction failed. Neither is the model's fault.

Model

Citations judged

Unsupported

GLM 5.2

42

0.0%

Kimi K3

50

2.0%

Mistral Small 4

39

2.6%

Muse Spark 1.1

36

2.8%

Claude Haiku 4.5

47

6.4%

DeepSeek V4 Pro

52

7.7%

Claude Opus 5

76

7.9%

Sonar Pro

76

7.9%

GPT-5.6 Luna

63

9.5%

Grok 4.3

43

16.3%

GPT-5.4 Mini

41

17.1%

Gemini 3.6 Flash

16

not reported, sample too small

GPT-5.4 Mini won our research category on answer quality and has the worst misattribution rate of any model we sampled enough to report. That is not a contradiction. It shows what quality scores measure. Our judges graded structure, coverage, and apparent authority. They did not click the links. Neither do most readers.

Sonar Pro's profile is imprecision, not invention. Only 7.9% unsupported, but 33 of its 84 judged citations were partial matches: on-topic sources that do not quite nail the specific assertion.

Gemini 3.6 Flash is unreportable here. Sixteen judged citations is too small a sample, and its rate moved sharply on re-judging. We have excluded it from the per-model claim rather than publish a number we do not trust.

Where we got this wrong

This is the part most audits leave out.

Our first version of Pass 2 reported 11.3% unsupported, and named GPT-5.6 Luna as having the worst citations in the field at 20.3%. Both figures were wrong, and the fault was ours.

An editor hand-checking a sample of 22 verdicts flagged six as bad fetches. Diagnosis found two bugs in our own pipeline:

We truncated long pages. The judge received only the first 6,000 characters of each source. One cited page ran to 759,758 characters with the relevant passage starting at 13,826. The judge never saw the evidence it was asked to rule on.

We destroyed tables. Our HTML-to-text step collapsed table cell boundaries, so tabular values lost their labels. This is why NIST conversion tables and Wikipedia data pages read as "does not mention it." One confirmed case: Luna cited NIST for the mile-to-kilometre factor and was marked unsupported. The page carries the line mile (mi) | kilometer (km) | 1.609 344. The citation was correct. We got it wrong.

We rebuilt the extraction and re-judged 278 verdicts. Unsupported fell from 11.3% to 8.6%. Luna fell from 20.3% to 9.5%, which moves it from worst in the field to unremarkable. That earlier finding is retracted.

The pattern is worth stating plainly: every extraction bug we found in this project made the models look worse than they actually were. There were five in total, including one that hid 90% of Grok's citations and one that discarded 106 of Luna's 112. Automated citation auditing is systematically biased against the models it audits. Human spot-checks are what caught this, and they are the only reason these numbers are trustworthy now.

Remaining caveats

We ran Pass 2 with a single judge, not the three used for response scoring. Per-model rates are indicative rather than definitive.

Our 8.6% figure is a ceiling, not a floor. Content behind JavaScript or inside PDFs can still read as unsupported to an automated checker.

Of the 52 unsupported verdicts that count toward these rates, 44 were re-judged with the corrected pipeline and 8 still rest on the older method, no more than two for any single model.

Verification was done in August 2026. Link rot moves. We will re-check before the next quarterly test.

One incidental finding: 75 of 708 cited URLs carry a tracking parameter, and every one belongs to GPT-5.4 Mini or GPT-5.6 Luna, which append ?utm_source=openai to their sources. No other model does this. It has no bearing on accuracy, but it does mean naive URL matching would treat an OpenAI citation and another model's citation of the identical page as different sources.

Download the full test data

The response data behind this audit is published as a single archive (1.1 MB): all 1,200 response transcripts, the 50 prompts and the answer key they were graded against, per-response cost, latency and token counts, and every score from all three AI judges.

Every row in the results file points at the exact transcript that produced it, so any number in the published article can be traced back to the raw output behind it.

Not included: the human reviewer's raw scoring sheets, which are one editor's unedited working notes rather than published material (their scores are reflected in the results), and cached copies of third-party pages, which are not ours to redistribute. The citation audit itself is published here on this page rather than duplicated in the archive.

Questions or corrections: reply to this email. Corrections policy is in how we test.

Reply

Avatar

or to participate