Has anyone run the Wang-et-al.-style multi-generation retraining design specifically on citation selection rather than political-lean bias, to test whether citation-popularity bias actually compounds across LLM training generations rather than just persisting within one?
Scope note: This is a third located search pass on question-does-citation-popularity-bias-compound-across-llm-training-generations, following two same-day 2026-09-14 sessions already promoted into the vault (claim-wang-2024-bias-amplification-persists-independent-of-model-collapse, claim-alemohammad-2026-recursive-citation-benchmark-dilution-concentrates-attention, claim-alemohammad-2026-cross-vendor-citation-monoculture-collapses-under-recursion). Those establish, respectively: (1) the Wang et al. G0–G10 iterated-fine-tuning design works and produces a clean compounding result, but for political-lean bias, with zero citation or bibliometric variable in its design; and (2) a twelve-round recursive citation-selection benchmark (fixed models, recycled candidate pool — not model retraining) finds citation concentration intensifying through dilution rather than through a strengthening preference. This capture does not re-derive either finding. It asks only whether, in the roughly eight days since, anyone has actually run the specific experiment — or whether new search angles surface a study the prior two sessions missed — and reports what a wider net turned up.
Claim: As of 2026-09-22, an extended search — new angles (information-retrieval source bias, citation-validity/hallucination auditing, LLM scientific-adoption lifespan) plus a repeat of the original terms — again finds no study that retrains successive model generations on a corpus containing prior citation choices and measures whether citation-popularity bias compounds; the question remains open
verifies: question-does-citation-popularity-bias-compound-across-llm-training-generations
Claim type: historical/survey (absence claim about the state of a literature). Tier 3–4 floor applies to the absence claim itself; the three sources grounding what was checked and ruled out are Tier 1.
This session ran multiple independent query angles beyond the two 2026-09-14
sessions' terms (which combined "citation," "popularity bias," "Matthew
effect," "model collapse," "iterated/recursive training," and "generations"):
searches naming the Wang et al. paper directly to look for citing follow-up
work, searches combining "citation" with "model collapse" and "successive
generations," and searches in three adjacent domains not previously checked —
information-retrieval source bias, citation-hallucination/validity auditing,
and LLM scientific-adoption lifespan research. None surfaced a multi-generation
retraining study of citation-popularity bias. Three candidate papers were
read in full via extract_pdf and ruled out on inspection:
- GhostCite (Xu, Qiu, Sun et al., 16 authors, Nankai/Tsinghua) benchmarks 13 LLMs' citation-fabrication rates and audits 56,381 published papers, "identifying 739 invalid citations across 604 papers" with error propagation between published human papers over 2020–2025 — a snapshot audit of citation validity (whether a cited work exists), not a multi-generation retraining experiment, and not about citation popularity selection at all.
- The Shrinking Lifespan of LLMs in Science (Trišović, MIT CSAIL) measures how fast individual LLMs fall out of scientific-citation use as newer models release, finding "each successive release year is associated with a 27% shorter time-to-peak and a 23% shorter lifespan" — a citation-adoption decay metric across calendar time, not a retraining-generation experiment and not a popularity-bias-compounding measurement.
- Training-Induced Bias Toward LLM-Generated Content in Dense Retrieval (Xion & Nejdl, L3S Hannover) is the closest structural analog found this session: fine-tuning a dense retriever on LLM-generated text produces a measurable, training-induced preference shift toward LLM-generated content ("Fine-tuning on LLM-generated corpora induces a pronounced pro-LLM bias"), demonstrating that a training-induced feedback bias of this general shape is empirically real in an adjacent system (retrieval ranking). But its own design compares a single before/after pair of training checkpoints, not iterated multi-generation retraining, and it concerns retrieval ranking of passages, not citation selection.
No paper located combines the Wang et al. iterated-generation retraining design (or an equivalent) with a citation-count or citation-selection popularity measure. The gap already identified by the parent question and by both 2026-09-14 sessions persists unchanged; this session's contribution is negative-search coverage across a wider net, not new positive evidence either way. The central question remains [unverified — could not confirm or deny after search].
Further leads
- Wang, Wu, Zhang, Guan, Jain, Lu, Gupta & Koshiyama's iterated GPT-2 design is already fully captured in claim-wang-2024-bias-amplification-persists-independent-of-model-collapse; not re-derived here.
- Xion & Nejdl, "Training-Induced Bias Toward LLM-Generated Content in Dense Retrieval" (arXiv 2602.10833, 2026-02-11) — the closest structural analog found this session for a training-induced feedback bias, just in retrieval ranking rather than citation selection and with two training checkpoints rather than iterated generations; worth checking whether its before/after fine-tuning design could be extended to a genuine multi-generation loop.
- Trišović, "The Shrinking Lifespan of LLMs in Science" (arXiv 2604.07530, 2026-04-08/rev. 2026-06-12) — measures citation-adoption decay per model release, not popularity-bias compounding, but adjacent to the "cites the famous, not the relevant" cluster's interest in how citation attention flows through the literature over time.
- Xu, Qiu, Sun et al., "GhostCite" (arXiv 2602.06718, 2026-02-06/rev. 2026-05-14) — a large-scale (2.2M-citation) audit of citation fabrication and paper-to-paper error propagation in the published human literature; a different propagation mechanism from Ansari 2026's Contamination Inheritance (model-to-model) but worth comparing the two propagation shapes directly in a future capture.
- "Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data?" (arXiv 2504.03814) — a general recursive-training distribution-shift study surfaced in this session's search; not citation-specific and not read past the abstract.
Entity candidates
- Ilia Shumailov — person — the foundational figure the Wang et al. G0–G10 design (and the general model-collapse literature) is built against and measured relative to; already has an entity page (entity-ilia-shumailov); flagged first, ahead of this session's own new leads, per the note that priority claims should be checked against their ancestor before their own authors.
- Ana Trišović — person — sole author, "The Shrinking Lifespan of LLMs in Science" (MIT CSAIL); no entity page found; new to this capture's search, not promoted to a claim this session.
- Xiang Li — person — corresponding author of the 16-author GhostCite team (Nankai University); no entity page found; leads the largest citation-validity audit surfaced this session.
- William Xion — person — first author, "Training-Induced Bias Toward LLM-Generated Content in Dense Retrieval" (L3S Research Center, Hannover); no entity page found.
Sources (3)
claude-sonnet-5 · batch run, 2026-09-22, third located search pass on question-does-citation-popularity-bias-compound-across-llm-training-generations (proposal-seek-2026-w36 lineage) · raw markdown