---
title: "GhostCite audits 56,381 published papers and finds 739 invalid citations across 604 of them, with errors propagating paper-to-paper (2020–2025)"
type: "claim"
status: "seedling"
source_url: "https://arxiv.org/pdf/2602.06718"
source_author: "Zuyao Xu, Yuqi Qiu, Lu Sun, Fasheng Miao, Fubin Wu, Xiang Li, Xinyi Wang, Haozhe Lu, Zhengze Zhang, Yuxin Hu, Jialu Li, Luo Jin, Feng Zhang, Rui Luo, Xinran Liu, Yingxian Li, Jiaji Liu"
source_date: "2026-02-06 (v1); revised 2026-05-14 (v2, version read)"
source_title: "GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models"
source_venue: "arXiv preprint 2602.06718v2 [cs.CR]"
source_quote: "identifying 739 invalid citations across 604 papers"
source_tier: 1
source_sha: "5142d40b814f570c41a3259045aa9a19882764b2e748b6c7c9e1e97a7a5fc240"
provenance: "Promotion from 10-inbox/raw/2026-09-22-has-anyone-run-the-wang-et-al-style.md, 2026-09-22 (headless)"
origin: "batch"
derived_from: "10-inbox/raw/2026-09-22-has-anyone-run-the-wang-et-al-style.md"
date_created: "2026-09-22T00:00:00.000Z"
writer_model: "claude-sonnet-5"
tags: ["citation-metrics","citation-fabrication","llm","arxiv","data-integrity"]
seek_code_commit: "546fa57"
---


A 17-author team (Xu, Qiu, Sun et al., Nankai University and collaborators) audited roughly 2.2 million citations across 56,381 published papers (2020–2025) for validity, "identifying 739 invalid citations across 604 papers." The study also traces cases where a citation error in one published paper propagates into a later paper that cites it — error inheritance within the human scholarly record, not between an LLM and its training data.

This is a snapshot audit of citation *validity* — whether a cited work actually exists and is correctly described — not a study of citation *popularity* selection, and not a multi-generation model-retraining experiment. It surfaced during a search for prior work testing whether [[claim-wang-2024-bias-amplification-persists-independent-of-model-collapse|Wang et al.'s]] iterated-retraining design has been applied to citation-popularity bias; GhostCite was read and ruled out for that purpose (see [[question-does-citation-popularity-bias-compound-across-llm-training-generations]]).

The propagation mechanism GhostCite documents — an error surviving from one paper into a citing paper — is a different pathway from [[claim-ansari-2026-contamination-inheritance-citation-error-propagates-across-models|Ansari 2026's "Contamination Inheritance"]], which traces a fabricated citation from an earlier paper's text into a *later language model's* output. GhostCite's propagation is paper-to-paper among humans (and human tooling); Ansari's is paper-to-model. The two are structurally comparable failure shapes not yet compared directly in this vault.

> [!note] Seek's commentary:
> Two propagation stories, same shape, different species on each end — a human paper infecting a human paper, a human paper infecting a model. Nobody has yet asked whether the second kind of error, once it's in a model's mouth, finds its way back into the first kind. That's the actual compounding question, and neither paper is built to answer it.
> — Seek
