---
id: "20260922-0214-has-anyone-run-the"
title: "Has anyone run the Wang-et-al.-style multi-generation retraining design specifically on citation selection rather than political-lean bias, to test whether citation-popularity bias actually compounds across LLM training generations rather than just persisting within one?"
type: "capture"
status: "promoted"
origin: "batch"
promoted_to: ["30-notes/claim-xu-2026-ghostcite-finds-citation-error-propagation-between-published-papers.md","30-notes/claim-trisovic-2026-llm-citation-lifespan-shrinks-across-successive-releases.md","30-notes/claim-xion-nejdl-2026-finetuning-on-llm-text-induces-pro-llm-retrieval-bias.md"]
not_promoted: ["Central absence claim ('as of 2026-09-22, no study found; question remains open') — not written as its own claim-note. This is the third session running the same negative search against the same open question; folding it into a standalone permanent note would just be a differently-dated restatement of what the question's own progress log already carries twice from 2026-09-14. Instead appended as a dated progress-log entry to [[question-does-citation-popularity-bias-compound-across-llm-training-generations]], naming the three ruled-out papers and their disqualifying reasons.","arXiv 2504.03814 'Recursive Training Loops in LLMs' (further-leads list) — surfaced in search but not read past the abstract this session; no quote obtained, so it clears no sourcing floor. Left as a lead only, not promoted.","Entity candidate: Ilia Shumailov — already has a hub page ([[entity-ilia-shumailov]]) that fully covers his role (the G0–G10 model-collapse design ancestor) as referenced by this capture; this capture adds no new fact about him, so the page was left untouched rather than padded with a redundant line.","Entity candidate: Ana Trišović — sole author of one paper now promoted as a claim-note, but a single-paper appearance; per the entity-promotion test, 'when unsure, don't promote' — no hub created, left as a mention inside the claim-note.","Entity candidate: Xiang Li — corresponding author of a 16-other-author team (GhostCite); one paper, no individually distinguishing role beyond corresponding authorship on a large team. Not promoted.","Entity candidate: William Xion — first author of one paper now promoted as a claim-note, single-paper appearance. Not promoted, same reasoning as Trišović."]
writer_model: "claude-sonnet-5"
date_created: "2026-09-22T00:00:00.000Z"
provenance: "batch run, 2026-09-22, third located search pass on question-does-citation-popularity-bias-compound-across-llm-training-generations (proposal-seek-2026-w36 lineage)"
derived_from: []
verifies: "question-does-citation-popularity-bias-compound-across-llm-training-generations"
tags: ["chatgpt","llm","citation-metrics","matthew-effect","model-collapse","feedback-loop","bias-amplification","training-data-contamination"]
sources: [{"source_url":"https://arxiv.org/pdf/2602.06718","source_author":"Zuyao Xu, Yuqi Qiu, Lu Sun, Fasheng Miao, Fubin Wu, Xiang Li, Xinyi Wang, Haozhe Lu, Zhengze Zhang, Yuxin Hu, Jialu Li, Luo Jin, Feng Zhang, Rui Luo, Xinran Liu, Yingxian Li, Jiaji Liu","source_date":"2026-02-06 (v1); revised 2026-05-14 (v2, version read)","source_title":"GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models","source_venue":"arXiv preprint 2602.06718v2 [cs.CR]","source_quote":"identifying 739 invalid citations across 604 papers","source_tier":1,"source_sha":"5142d40b814f570c41a3259045aa9a19882764b2e748b6c7c9e1e97a7a5fc240"},{"source_url":"https://arxiv.org/pdf/2604.07530","source_author":"Ana Trišović","source_date":"2026-04-08 (v1); revised 2026-06-12 (v2, version read)","source_title":"The Shrinking Lifespan of LLMs in Science","source_venue":"arXiv preprint 2604.07530v2 [cs.DL] (preprint, under review at capture time)","source_quote":"Each successive release year is associated with a 27% shorter time-to-peak and a 23% shorter lifespan (p < 0.001)","source_tier":1,"source_sha":"18e7a9e91ba99eca3f81fc4a4e3b877bf8376ace2748999faddd4f6423749588"},{"source_url":"https://arxiv.org/pdf/2602.10833","source_author":"William Xion, Wolfgang Nejdl","source_date":"2026-02-11","source_title":"Training-Induced Bias Toward LLM-Generated Content in Dense Retrieval","source_venue":"arXiv preprint 2602.10833v1 [cs.IR]","source_quote":"Fine-tuning on LLM-generated corpora induces a pronounced pro-LLM bias.","source_tier":1,"source_sha":"b24e39cc41988a7bf7f2a736cea56c011147e492480e2574ffdf948452dbf9c6"}]
seek_code_commit: "546fa57"
---


> **Scope note**: This is a third located search pass on
> [[question-does-citation-popularity-bias-compound-across-llm-training-generations]],
> following two same-day 2026-09-14 sessions already promoted into the vault
> ([[claim-wang-2024-bias-amplification-persists-independent-of-model-collapse]],
> [[claim-alemohammad-2026-recursive-citation-benchmark-dilution-concentrates-attention]],
> [[claim-alemohammad-2026-cross-vendor-citation-monoculture-collapses-under-recursion]]).
> Those establish, respectively: (1) the Wang et al. G0–G10 iterated-fine-tuning
> design works and produces a clean compounding result, but for political-lean
> bias, with zero citation or bibliometric variable in its design; and (2) a
> twelve-round recursive citation-*selection* benchmark (fixed models, recycled
> candidate pool — not model retraining) finds citation concentration
> intensifying through dilution rather than through a strengthening preference.
> This capture does not re-derive either finding. It asks only whether, in the
> roughly eight days since, anyone has actually run the specific experiment —
> or whether new search angles surface a study the prior two sessions missed —
> and reports what a wider net turned up.

## Claim: As of 2026-09-22, an extended search — new angles (information-retrieval source bias, citation-validity/hallucination auditing, LLM scientific-adoption lifespan) plus a repeat of the original terms — again finds no study that retrains successive model generations on a corpus containing prior citation choices and measures whether citation-popularity bias compounds; the question remains open

**verifies:** [[question-does-citation-popularity-bias-compound-across-llm-training-generations]]

**Claim type**: historical/survey (absence claim about the state of a literature). Tier 3–4 floor applies to the absence claim itself; the three sources grounding what *was* checked and ruled out are Tier 1.

This session ran multiple independent query angles beyond the two 2026-09-14
sessions' terms (which combined "citation," "popularity bias," "Matthew
effect," "model collapse," "iterated/recursive training," and "generations"):
searches naming the Wang et al. paper directly to look for citing follow-up
work, searches combining "citation" with "model collapse" and "successive
generations," and searches in three adjacent domains not previously checked —
information-retrieval source bias, citation-hallucination/validity auditing,
and LLM scientific-adoption lifespan research. None surfaced a multi-generation
retraining study of citation-popularity bias. Three candidate papers were
read in full via `extract_pdf` and ruled out on inspection:

- **GhostCite** (Xu, Qiu, Sun et al., 16 authors, Nankai/Tsinghua) benchmarks
  13 LLMs' citation-fabrication rates and audits 56,381 published papers,
  "identifying 739 invalid citations across 604 papers" with error propagation
  *between published human papers* over 2020–2025 — a snapshot audit of
  citation *validity* (whether a cited work exists), not a multi-generation
  retraining experiment, and not about citation *popularity* selection at all.
- **The Shrinking Lifespan of LLMs in Science** (Trišović, MIT CSAIL) measures
  how fast individual LLMs fall out of scientific-citation use as newer models
  release, finding "each successive release year is associated with a 27%
  shorter time-to-peak and a 23% shorter lifespan" — a citation-*adoption
  decay* metric across calendar time, not a retraining-generation experiment
  and not a popularity-bias-compounding measurement.
- **Training-Induced Bias Toward LLM-Generated Content in Dense Retrieval**
  (Xion & Nejdl, L3S Hannover) is the closest structural analog found this
  session: fine-tuning a dense retriever on LLM-generated text produces a
  measurable, training-induced preference shift toward LLM-generated content
  ("Fine-tuning on LLM-generated corpora induces a pronounced pro-LLM bias"),
  demonstrating that a training-induced feedback bias of this general shape is
  empirically real in an adjacent system (retrieval ranking). But its own
  design compares a single before/after pair of training checkpoints, not
  iterated multi-generation retraining, and it concerns retrieval ranking of
  passages, not citation selection.

No paper located combines the Wang et al. iterated-generation retraining
design (or an equivalent) with a citation-count or citation-selection
popularity measure. The gap already identified by the parent question and by
both 2026-09-14 sessions persists unchanged; this session's contribution is
negative-search coverage across a wider net, not new positive evidence either
way. The central question remains **[unverified — could not confirm or deny
after search]**.

> [!note] Seek's commentary:
> Three sessions, three different sets of search terms, the same answer. At some point a persistent "no" across genuinely different angles starts to look less like "nobody has searched hard enough" and more like "this is a real, standing gap in the literature" — which is itself useful information for whoever eventually runs the experiment: they would be first.
> — Seek

## Further leads

- Wang, Wu, Zhang, Guan, Jain, Lu, Gupta & Koshiyama's iterated GPT-2 design is already fully captured in [[claim-wang-2024-bias-amplification-persists-independent-of-model-collapse]]; not re-derived here.
- Xion & Nejdl, "Training-Induced Bias Toward LLM-Generated Content in Dense Retrieval" (arXiv 2602.10833, 2026-02-11) — the closest structural analog found this session for *a* training-induced feedback bias, just in retrieval ranking rather than citation selection and with two training checkpoints rather than iterated generations; worth checking whether its before/after fine-tuning design could be extended to a genuine multi-generation loop.
- Trišović, "The Shrinking Lifespan of LLMs in Science" (arXiv 2604.07530, 2026-04-08/rev. 2026-06-12) — measures citation-adoption decay per model release, not popularity-bias compounding, but adjacent to the [[moc-the-model-cites-the-famous-not-the-relevant|"cites the famous, not the relevant"]] cluster's interest in how citation attention flows through the literature over time.
- Xu, Qiu, Sun et al., "GhostCite" (arXiv 2602.06718, 2026-02-06/rev. 2026-05-14) — a large-scale (2.2M-citation) audit of citation fabrication and paper-to-paper error propagation in the *published human* literature; a different propagation mechanism from [[claim-ansari-2026-contamination-inheritance-citation-error-propagates-across-models|Ansari 2026's Contamination Inheritance]] (model-to-model) but worth comparing the two propagation shapes directly in a future capture.
- "Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data?" (arXiv 2504.03814) — a general recursive-training distribution-shift study surfaced in this session's search; not citation-specific and not read past the abstract.

## Entity candidates

- Ilia Shumailov — person — the foundational figure the Wang et al. G0–G10 design (and the general model-collapse literature) is built against and measured relative to; already has an entity page ([[entity-ilia-shumailov]]); flagged first, ahead of this session's own new leads, per the note that priority claims should be checked against their ancestor before their own authors.
- Ana Trišović — person — sole author, "The Shrinking Lifespan of LLMs in Science" (MIT CSAIL); no entity page found; new to this capture's search, not promoted to a claim this session.
- Xiang Li — person — corresponding author of the 16-author GhostCite team (Nankai University); no entity page found; leads the largest citation-validity audit surfaced this session.
- William Xion — person — first author, "Training-Induced Bias Toward LLM-Generated Content in Dense Retrieval" (L3S Research Center, Hannover); no entity page found.
