Kurisutina

Debunking: a meta-analysis of the psychological efficacy of messages countering misinformation

Psychological Science 28(11): 1531–1546 (online-first pagination 1–16), doi 10.1177/0956797617714579. PMC5673564 (author manuscript). Open Data badge; data at osf.io/9d6t4 (not downloaded). Read by researcher R5b (change with reasons), 5 October 2026. Provenance: papers/change/chan2017_debunking.provenance.json.

What was read

  • Read in full: the publisher's online-first PDF (16 pages) as hosted on the senior author's lab site, converted with pdftotext -layout and folded: abstract, introduction, method, results, Tables 1–6, the Figure 1 flowchart, the Figure 2 caption and axis text, discussion, recommendations, notes, references.
  • Not read: the Supplemental Material (search terms, selection-model and meta-regression bias details) and the OSF data. Figure 2's funnel plots were not inspected as images. Europe PMC and PMC copies sit behind a bot check and were not used.
  • Tables 1, 2 and 5 are wide; their folded rows were re-assembled by reading the column order. Values below were checked against the abstract where it repeats them.

What it is

  • Question: how large are the misinformation effect, the debunking effect and the persistence of misinformation after debunking, and which audience and message factors moderate them?
  • Paradigm (the commonest one, e.g. the warehouse fire): participants read news reports in which a cause is asserted (misinformation), later retracted (debunking) or not; a control group gets neither. After about 10 minutes on an unrelated task they answer causal-inference and factual questions.
  • Three effects, all between-subjects, as Hedges's d:
    • misinformation effect = misinformation group − control;
    • debunking effect = misinformation group − debunked group;
    • misinformation-persistence effect = debunked group − control (what is left after debunking).
  • Studies: 8 reports, 20 experiments, 52 independent samples, N = 6,878; published 1994–2015 (journal articles, theses, working papers). Mean sample 132 (SD 174); 69.2% lab, 30.8% online; 72% female; mean age 20. Topics: robberies, a warehouse fire, traffic accidents and a plane crash, the 2010 ACA "death panels", candidates' policy positions, a candidate taking donations from a felon. The positions were unfamiliar to participants before the experiment.
  • Moderators coded (κ .87–1.00, ICC .90–1.00): (a) likelihood of generating explanations in line with the misinformation (explicit procedure, e.g. repeating it, plus settings that prompt it, e.g. causal-inference questionnaires); (b) likelihood of generating counterarguments to it after the debunking; (c) detail of the debunking (only labelling it incorrect = 1, giving new credible information such as the real cause = 2).

Main results (verified)

Mean effects (Table 2), Hedges's d [95% CI]:

Effect k Fixed Random Three-level Random, weighted by standardized-N residuals
Misinformation, all 16 2.04 [1.93, 2.14] 3.08 [2.02, 4.15] 3.08 [2.00, 4.15] 2.94 [1.80, 4.08]
Misinformation, outliers removed 14 2.01 [1.91, 2.12] 2.46 [1.73, 3.19] 2.49 [1.53, 3.45] 2.41 [1.63, 3.20]
Debunking 30 0.88 [0.81, 0.93] 1.14 [0.68, 1.61] 1.14 [0.68, 1.61] 1.33 [0.62, 2.04]
Persistence, all 42 0.09 [0.07, 0.12] 0.97 [0.60, 1.35] 0.92 [0.40, 1.44] 1.06 [0.68, 1.44]
Persistence, outliers removed 40 0.09 [0.06, 0.12] 0.75 [0.50, 1.00] 0.79 [0.36, 1.23] 0.77 [0.48, 1.05]
  • Outliers were d > 5.50. Heterogeneity is extreme: I² about 98–99.5% throughout; random-effects τ² 4.48 (misinformation), 1.65 (debunking), 1.46 (persistence).
  • The abstract's ranges: misinformation ds 2.41–3.08, debunking 1.14–1.33, persistence 0.75–1.06; "large" by Cohen.
  • The fixed-effects persistence estimate is 0.09, against 0.75–1.06 for random effects. Larger samples showed smaller persistence: unpublished reports (working papers, dissertations) had larger samples than journal articles (r between publication status and sample size .91 for the misinformation effect, .63 for persistence). In the largest, online study (Berinsky 2012, ACA death panels, n 278–618 per condition) the persistence ds were −0.42 to +0.19.
  • Bias checks disagree: funnel asymmetry and trim-and-fill (19 records filled for persistence) and the publication-status meta-regression suggest bias; selection models (except debunking), p-curve and p-uniform do not. The authors re-ran everything with sample-size residual weights; results were similar.

Moderators (Tables 5–6, weighted mixed-effects models):

  • Generating explanations in line with the misinformation raised persistence (b = 2.09 [0.88, 3.30]; unweighted 1.40 [0.26, 2.54]) and weakened debunking (b = −4.08 [−5.50, −2.66]). It also weakened the initial misinformation effect (b = −0.98 [−1.61, −0.33]), which the authors did not expect.
  • Counterarguing the misinformation after the debunking strengthened debunking (b = 0.93 [0.42, 1.44]) and reduced persistence (b = −0.68 [−1.16, −0.20]; unweighted −0.36 [−0.82, 0.11], not significant).
  • Detailed debunking (new information, not just "this was wrong") strengthened debunking (b = 1.82 [0.57, 3.07]) but was also associated with more persistence (b = 1.06 [0.23, 1.90]; unweighted 0.86 [0.04, 1.67]). Post hoc, detail correlated with explanation generation, r(33) = .52, so studies with detailed debunks may have had more detailed misinformation.
  • Effect sizes at moderator levels (Table 6; "high" and "low" = above and below 0.25 SD):
Moderator Level Debunking d (k) Persistence d (k)
Explanations in line with misinformation high 0.62 (18) 1.72 (25)
low 3.77 (3) −0.14 (10)
Counterarguments to misinformation high 1.48 (8) 0.46 (13)
low 0.73 (11) 1.58 (22)
Detail of debunking high 1.25 (18) 1.58 (21)
low 0.16 (3) 0.55 (14)
  • With outliers removed, explanation generation (b = 0.82) and detail (b = 0.52) still predicted persistence; counterarguing did not (b = 0.08). Moderators explained R² .43–.65.

Recommendations (the authors'): reduce arguments that support the misinformation; create conditions for scrutiny and counterarguing; correct with new detailed information, but keep expectations low. They note these ignore the audience's dispositions, ideology and culture.

Limits

  • No time course. Persistence is measured within the session (typically after about 10 minutes). Nothing here says how corrections or residual misinformation evolve over days.
  • Small, mostly lab, mostly student samples; few reports (8), so moderators are confounded with reports and paradigms; the coded moderators are judged from procedures, not measured in participants.
  • Unfamiliar content. Participants had no prior position; worldview-laden misinformation is not what this estimates.
  • Heterogeneity near 99% makes any single mean a poor description; the fixed-effect and random-effect persistence estimates differ by a factor of about 10.
  • Some table cells are missing per record; outlier rules and weighting choices move estimates noticeably.

What it means for Kurisutina (inference)

  • Retraction reduces but does not erase. With content the person had no prior view on, a debunking removes much of a misinformation effect within the session, and a residue remains; its size is uncertain (d 0.09 by fixed effects, 0.75–1.06 by random effects).
  • The person's own reasoning is the strongest moderator. Having generated explanations for a claim makes it resist later correction (persistence 1.72 against −0.14); counterarguing it does the opposite. For a replica, the self-generated explanation is a stored, first-person episode: the rule of change should treat "I reasoned to X" as stronger support than "I was told X" [I].
  • For M1: a "retraction" episode family (observation or hearsay of X, then a correction with or without a replacement cause) has a ground truth here only in sign: correction with a replacement explanation lowers belief more than a bare "that was wrong", and some residue remains [I].
  • For the battery: a correction probe has a human reference only as a standardized difference between groups, not as a per-person switch rate, and only within a session [I].

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.