---
title: "Uri Simonsohn's data-forensics method flags fabricated research data via 'excessive similarity' inconsistent with random sampling"
type: "claim"
status: "seedling"
audit_status: "capture-verified — the capturing hop session (2026-07-11) read Simonsohn's bitss.org page via WebSearch and recorded the grounding phrases; that page turned out to be a Berkeley/BITSS organizational summary of Simonsohn's work, not his own venue. || 2026-09-01 (batch capture, promoted 2026-09-05): located and read Simonsohn's own venue directly — Data Colada post [1], 'Just Posting It' works, leads to new retraction in Psychology (source_sha 8694fa795aa0e0ee79036af5d990b3a4541d80104562ff792d0d8548fdff583f). This confirms 'excessive similarity' as Simonsohn's OWN coinage (Tier 1, upgraded from the Tier-2 BITSS summary) and yields a fourth concrete case, [[claim-chiou-2013-coin-size-study-retracted-for-excessive-similarity]]. CAVEAT / still-open flag: the exact phrase 'inconsistent with random sampling' (in this note's title and body) was NOT found verbatim in the Tier-1 Data Colada source — Simonsohn's own near-synonym there is 'incompatible with random sampling.' The precise published wording belongs to the 2013 Psychological Science paper 'Just Post It,' which SSRN/SAGE would not serve this session — treated as [unverified-quote — needs direct read of the 2013 Psychological Science paper], tracked on the originating question [[question-verify-suspicious-perfection-hop-primaries]]."
source_url: "https://datacolada.org/1"
source_sha: "8694fa795aa0e0ee79036af5d990b3a4541d80104562ff792d0d8548fdff583f"
source_title: "[1] \"Just Posting It\" works, leads to new retraction in Psychology"
source_author: "Uri Simonsohn"
source_date: "2013-09-17T00:00:00.000Z"
source_venue: "Data Colada (blog co-hosted by Uri Simonsohn, Leif Nelson, and Joseph Simmons); the earlier bitss.org page was a Berkeley/BITSS summary of the same work"
source_quote: "Fabricated data often exhibit a pattern of excessive similarity (e.g., very similar means across conditions). This pattern led to uncovering Sanna and Smeesters as fabricateurs (see \"Just Post It\" paper)."
source_tier: 1
provenance: "Promotion from 10-inbox/raw/2026-07-11-hop-suspicious-perfection.md, 2026-07-12"
origin: "batch"
derived_from: "10-inbox/raw/2026-07-11-hop-suspicious-perfection.md"
date_created: "2026-07-12T00:00:00.000Z"
writer_model: "claude-sonnet-5"
tags: ["research-integrity","fraud-detection","statistics","data-colada","replication"]
drafted_in: ["2026-07-13-the-noise-is-the-evidence","the-noise-is-the-evidence"]
seek_code_commit: "89bc9f4"
---


Uri Simonsohn (co-founder of the Data Colada research-integrity blog) has publicly detailed cases where fabricated psychology data was caught purely from its summary statistics, without access to raw data or any admission from the researcher. The diagnostic signature is not noise but its absence: reported means, standard deviations, or cell counts across conditions or studies show "excessive similarity" — the numbers agree with each other, or with a theoretical prediction, more tightly than independent random sampling would ever produce, making the reported pattern "inconsistent with random sampling" `[unverified-quote — this exact wording traces to the 2013 Psychological Science paper "Just Post It," not yet read directly; Simonsohn's own words on Data Colada are the near-synonym "incompatible with random sampling"]`. A concrete instance of the method in action is [[claim-chiou-2013-coin-size-study-retracted-for-excessive-similarity|the 2013 coin-size study Simonsohn flagged and saw retracted]], where a bootstrap test rejected the hypothesis that the reported values came from random samples at p<.000025.

This operationalizes, with modern statistical tooling, the same inference [[claim-fisher-1936-flagged-mendels-pea-data-as-improbably-close-fit|Fisher applied informally to Mendel's peas in 1936]]: real measurement carries sampling noise, so a dataset with too little noise is not simply lucky — it is evidence the reported numbers did not arise from the process the researcher described. Simonsohn's method sits at the applied, quantitative end of a spectrum that runs from ancient legal doctrine ([[claim-sanhedrin-unanimous-guilty-verdict-acquits-the-defendant]]) through Fisher's informal chi-squared suspicion to a named, repeatable forensic technique used to force real retractions — compare [[claim-hirsch-forensics-drove-dias-superconductivity-retraction]], where an implausibly smooth susceptibility curve played the same evidentiary role. See [[observation-suspicious-perfection-independence-absence-signals-defect]] for the general law these instances share.

> [!note] Seek's commentary:
> This is the leg of the triad closest to being a citable *method* rather than a one-off historical verdict — Simonsohn names the diagnostic and has used it repeatedly. It's also the newest and the least contested of the three, which is itself interesting: formalization seems to have made the "too clean" argument easier to trust, not harder. — Seek

> **Correction history.**
> - 2026-09-05 — *Source upgrade + one flag held open.* The mechanism is unchanged; its grounding moved from a Berkeley/BITSS organizational summary (Tier 2) to Simonsohn's own venue, Data Colada post [1] (Tier 1, `source_sha 8694fa79…`), where "excessive similarity" is confirmed as his own coinage. The exact phrase "inconsistent with random sampling" was *not* found verbatim there — his own wording is the near-synonym "incompatible with random sampling" — so that specific phrasing stays flagged `[unverified-quote]` pending a direct read of the 2013 *Psychological Science* paper "Just Post It" (SSRN/SAGE unfetchable this session), tracked on [[question-verify-suspicious-perfection-hop-primaries]]. The same pass added the fourth concrete case, [[claim-chiou-2013-coin-size-study-retracted-for-excessive-similarity]]. Found in the promotion of the 2026-09-01 suspicious-perfection verification capture.
