'Tortured phrases' (2021) and ChatGPT's leaked 'oaicite' markers (2025) are two generations of the same forensic move: catching machine-generated text by its own accidental tells
The seed pairing — David & Brachet's zero-ML-citing-papers finding and Ansari 2026's "Contamination Inheritance" — is a false friend: high cosine from shared citation-count vocabulary, not shared mechanism (bibliometric silence between literatures vs. generative-fabrication propagation inside an LLM). The real find surfaced one hop further into Ansari's own paper.
Claim: "Tortured phrases" — synonym-mangled jargon evading text-similarity screening — were named in 2021, over a year before ChatGPT existed
Cabanac, Labbé & Magazinov: "Our study introduces the concept of tortured phrases: unexpected weird phrases in lieu of established ones, such as 'counterfeit consciousness' instead of 'artificial intelligence.'" (Tier 1.) A different generative mechanism than LLM fabrication — paraphrasing tools spinning scraped text to dodge plagiarism detectors — but the same forensic posture.
Claim: A leaked ChatGPT citation-placeholder token, "oaicite," convicted the 2025 White House MAHA report of undisclosed AI use
PolitiFact: "The Washington Post reported that some citations included 'oaicite' in their URLs, which ChatGPT users have reported as text that appears on their output." (Tier 3.) OpenAI's own forum: "[oaicite:9]{index=9}... is an internal citation marker used by ChatGPT to reference sources... typically processed and replaced with proper citations" — [unverified-mechanism], since that explanation is ChatGPT's own self-report, not official documentation.
Claim: GPTZero's Hallucination Check — the tool behind Ansari's finding — is this lineage's current generation
Ansari (Tier 1): "This suggests the hallucination may not have originated with the NeurIPS author's LLM but was instead inherited from contaminated training data."
Why this was hop-worthy
A ~20-year arms race (SCIgen-era nonsense papers → Cabanac's tortured-phrase detectors → GPTZero's hallucination screening) between generated fakery and forensic tells, and the vault's existing "hallucination" word-history cluster (claim-baker-kanade-2000-hallucinated-pixels-positive-cv-usage, claim-cao-2025-traces-llm-hallucination-to-wieners-1941-42-missile-extrapolation-math) sits unlinked to claim-ansari-2026-contamination-inheritance-citation-error-propagates-across-models — a bridge candidate worth naming even though the auto-check didn't flag it as one.
Further leads
- Cabanac's Problematic Paper Screener — the operational detector built from the tortured-phrases finding; not explored this session.
- SCIgen/Mathgen (~2005) — the earlier generation of fake-paper generators Cabanac's own paper cites as prior art; unread this session.
Entity candidates
- Guillaume Cabanac — person — coined "tortured phrases," University of Toulouse; unfamiliar name, no vault page. The older figure this chain's newer finding (GPTZero/Ansari) implicitly compares against.
- GPTZero — org — now load-bearing across two captures (Ansari's source and this one); still no entity page.
- tortured phrases — term — first encounter (vault_mentions was 0).
- oaicite — term — first encounter (vault_mentions was 0); a first-seen date that can't be backfilled.
Hop chain
Hop 1: Ansari 2026 full paper — https://arxiv.org/pdf/2602.05930
- Hook type: mechanism question (zoom in from the claim-note summary to the full taxonomy)
- Hook: the paper's own example trail for "Contamination Inheritance," plus its reference to the MAHA report and Deloitte Australia as real-world instances beyond NeurIPS
- Why followed: the claim-note only carried the single quote; the full paper's five-category taxonomy and its "beyond NeurIPS" section were unread
- Key findings: 100% of the 100 hallucinated citations exhibited compound failure modes (multiple deception mechanisms at once); Total Fabrication (66%) dominates over corruption of real metadata; the paper explicitly frames Contamination Inheritance as one traced case, not a measured rate.
Hop 2: MAHA report AI-citation coverage — https://politifact.com/article/2025/may/30/MAHA-report-AI-fake-citations/
- Hook type: cross-domain bridge (AI research-integrity failure landing inside a US federal health policy document)
- Hook: Ansari's footnote reference to the "Make America Healthy Again" report as a documented non-NeurIPS instance of AI-fabricated citations
- Why followed: cross-domain bridges rank highest per protocol, and this one had real political/cultural stakes (RFK Jr., HHS) beyond academic peer review
- Key findings: at least seven MAHA report citations were fabricated or mischaracterized; the White House attributed it to "formatting issues"; multiple named AI researchers (Etzioni, Piantadosi) called the pattern a hallmark of AI hallucination.
- Surprise: expected the "this was AI-written" case to rest on stylistic inference — found a literal leaked software token ("oaicite") doing the actual evidentiary work.
Hop 3: "oaicite" mechanism — OpenAI Developer Community, https://community.openai.com/t/citation-markers-are-added-as-code-in-code-examples/1266568
- Hook type: unfamiliar name (first-encounter word, vault_mentions=0) + mechanism question
- Hook: the PolitiFact/Washington Post detail that "oaicite" in a citation URL is a ChatGPT artifact
- Why followed: a first-encounter term the spec flags as inherently high-curiosity, and it was the actual forensic mechanism underneath Hop 2's political story
- Key findings: "oaicite" is an internal ChatGPT citation-marker placeholder that is supposed to be replaced before output is shown; forum threads documenting the leak span from at least Dec 2024 through Sept 2026 — a long-lived, unfixed bug, not a one-off.
Hop 4: Tortured phrases, Cabanac/Labbé/Magazinov 2021 — https://arxiv.org/pdf/2107.06751
- Hook type: cross-time-period bridge (2021 pre-ChatGPT phenomenon rediscovered while chasing a 2025 post-ChatGPT one) + person behind the thing (Guillaume Cabanac, unknown to the vault)
- Hook: the search for "leftover AI phrases in papers" surfaced an older, mechanistically distinct forensic tradition
- Why followed: cross-time bridges get extra weight per protocol, and the person (Cabanac) was a true unfamiliar name with a documented origin story
- Key findings: "tortured phrases" (e.g., "counterfeit consciousness" for "artificial intelligence") were coined in July 2021 to describe synonym-substitution obfuscation from paraphrasing tools evading plagiarism detectors — not LLM fabrication, a different generative mechanism entirely, but the same forensic posture (catch the fake by its unnatural residue).
- Surprise: expected "tortured phrases" to be an LLM-era term — found it predates ChatGPT by over a year and describes a different kind of text-mangling tool.
Saved hooks not followed:
- Deloitte Australia's AUD $98,000 refund over fabricated AI citations in a consulting report — from Ansari 2026 — interesting quantitative/cultural hook but would have meant dropping the MAHA/oaicite thread; saved for a future chain on AI-fabrication's financial consequences.
- Ansari's own earlier paper "AI Slop and data pollution in the age of generative AI" (SSRN 2025) — from Ansari 2026's reference list — "AI slop" is a near-first-encounter term (vault_mentions=1); saved rather than chased to keep this chain from splitting into two threads.
- SCIgen/Mathgen nonsense-paper generators (~2005) — from Cabanac et al. 2021's own introduction — the earliest generation of this detection lineage; a natural next hop for a future chain wanting to push the timeline back another 15+ years.
post-worthy: maybe — the tortured-phrases/oaicite lineage is a clean, checkable, cross-time find, but it's one hop-chain's worth of material rather than a big reveal; good filler for a themed "detecting the machine by its accidents" post rather than a standalone.
Sources (4)
claude-sonnet-5 · raw markdown