18% of GPT-4's citations are fabricated: verify before you cite
Every article on this site ends with a numbered source list, and each entry costs me a few minutes of searching before it goes live: authors, year, venue, and the sentence that carries the finding. That routine exists because of a number William Walters and Esther Wilder published in Scientific Reports. Across 636 references that ChatGPT produced for 42 literature reviews, 55% of the GPT-3.5 citations and 18% of the GPT-4 citations pointed to papers that do not exist 1.
What the research shows.
Walters and Wilder also counted the ones that survive a search. Among citations that led to a real paper, 43% of GPT-3.5's and 24% of GPT-4's carried a substantive error in the authors, the year, the journal or the volume 1. Finding the paper settles half the question.
Medicine gives a similar reading. Mikaël Chelli's team asked GPT-3.5, GPT-4 and Bard to rebuild the reference lists of 11 published systematic reviews, then scored all 471 returned references against the originals. Hallucination rates came in at 39.6% for GPT-3.5, 28.6% for GPT-4 and 91.4% for Bard. Precision, meaning the share of returned papers that belonged in the review, topped out at 13.4% for GPT-4 2.
Law shows the pattern at its sharpest, because a court case either exists or it does not. Matthew Dahl's group at Stanford asked models specific, verifiable questions about random federal cases. Hallucination rates ran from 58% for GPT-4 to 88% for Llama 2 3. The same models tended to accept a false premise buried in the question, so an invented case name comes back with an invented holding attached.
Retrieval narrows the gap and leaves plenty of room. Varun Magesh and colleagues ran 202 preregistered legal queries through the commercial tools built on retrieval over real case law. Lexis+ AI answered 65% of them correctly and hallucinated on 17%; Westlaw's AI-Assisted Research hallucinated on 33% 4.
Consumer search lands in the same band. Nelson Liu audited four generative search engines and found that 51.5% of their sentences were fully supported by the citations attached, and that 74.5% of citations supported the sentence they sat next to 5. Two years later Klaudia Jaźwińska and Aisvarya Chandrasekar, at Columbia's Tow Center, gave eight AI search tools a quote from a news article and asked for the title, publisher, date and URL. Across 1,600 queries the tools were wrong more than 60% of the time. Grok 3 missed on 94%, Perplexity, the strongest of the eight, on 37%, and more than half of the Gemini and Grok 3 answers pointed at fabricated or broken URLs 6.
The models are graded for guessing.
Adam Kalai's group published the mechanism in Nature this year. Under accuracy-based scoring, an answer of "unknown" earns zero while a guess has some chance of earning a point, so training and benchmarking pay a model to answer everything 7. Their learning-theory argument adds the second half: facts with no repeated support in the training data produce unavoidable errors, while recurring regularities such as grammar do not 7. One paper published once, cited by a handful of others, is the definition of a fact with thin support. The citation format, by contrast, is a regularity the model has seen a million times, which is why a fabricated reference looks flawless.
Asking the model to check itself makes it worse.
Jie Huang's team at Google DeepMind ran the obvious fix and measured the damage. On GSM8K, GPT-4 scores 95.5%. After one round of reviewing and revising its own answer with no external signal, it scores 91.5%, and 89.0% after a second round 8. Across those rounds the model flipped more correct answers to wrong than the reverse. Asking "are you sure about this reference?" adds nothing the model did not already have.
Sampling works where self-review fails. Sebastian Farquhar's group generates several answers to the same question, clusters them by meaning and measures the entropy across clusters. High disagreement between meanings flags the arbitrary answers they call confabulations, and the method needs no task-specific training data or prior knowledge of the task 9.
The protocol.
- Resolve the identifier before you read the claim. Paste the title into Google Scholar or the DOI resolver and confirm authors, year and venue. Budget 2 minutes per reference. Stop reading the summary until the paper exists.
- Open the source and find the sentence. The number you want to quote has to appear in the abstract, the results or a table. Liu's audit puts support for a cited sentence at 51.5% 5, so half of what you skip is unsupported.
- Ask three times, in three fresh chats. Same question, no history. If the three answers disagree on the first author or the year, drop the reference; that disagreement is what Farquhar's method measures 9.
- Write an exit into the prompt. "If you cannot name a source you would bet on, answer: unknown." Models guess because the scoring rewards it 7; a prompt that scores abstention as a valid answer removes part of the incentive.
- Strip the premise out of your question. Ask which papers measured an effect rather than asking a model to summarize a paper you half-remember. Dahl's group showed models building answers on false premises supplied by the user 3.
- Keep checking when the tool cites sources. Retrieval-based legal tools still hallucinate on 17% to 33% of queries 4, and AI search cited fabricated or broken URLs in more than half of some models' answers 6. A visible link is a claim about a source, not a check on it.
- Measure your own rate. Log the next 20 references a model hands you and mark each one verified, wrong in the details, or fabricated. Twenty entries take an hour and give you a number for your field and your tool, which beats the 18% from a 2023 benchmark 1.
For anything with a byline, the whole check runs in the time it takes to read the abstract. I do it for every source in every article here, including the nine below.
Sources.
- Walters, W. H., Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, 14045. doi:10.1038/s41598-023-41032-5
- Chelli, M., Descamps, J., Lavoué, V., Trojani, C., Azar, M., Deckert, M., Raynier, J.-L., Clowez, G., Boileau, P., Ruetsch-Chelli, C. (2024). Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis. Journal of Medical Internet Research, 26, e53164. jmir.org
- Dahl, M., Magesh, V., Suzgun, M., Ho, D. E. (2024). Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models. Journal of Legal Analysis, 16(1), 64-93. academic.oup.com
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., Ho, D. E. (2025). Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. Journal of Empirical Legal Studies, 22(2), 216-242. doi:10.1111/jels.12413
- Liu, N. F., Zhang, T., Liang, P. (2023). Evaluating Verifiability in Generative Search Engines. Findings of the Association for Computational Linguistics: EMNLP 2023, 7001-7025. aclanthology.org
- Jaźwińska, K., Chandrasekar, A. (2025). AI Search Has a Citation Problem. Tow Center for Digital Journalism, Columbia Journalism Review. cjr.org
- Kalai, A. T., Nachum, O., Vempala, S. S., Zhang, E. (2026). Evaluating large language models for accuracy incentivizes hallucinations. Nature, 653, 1047-1051. nature.com
- Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., Zhou, D. (2024). Large Language Models Cannot Self-Correct Reasoning Yet. ICLR 2024. iclr.cc
- Farquhar, S., Kossen, J., Kuhn, L., Gal, Y. (2024). Detecting hallucinations in large language models using semantic entropy. Nature, 630, 625-630. doi:10.1038/s41586-024-07421-0