Journal of Politics 79(4): 1403–1418, 2017, doi 10.1086/692739 (DOI and pages from Crossref). Read as the author's ungated final version (dated 7 August 2017) with its online appendix, both from the author's website. Read by researcher R5b (change with reasons), 5 October 2026. Provenance: papers/change/hill2017_bayesian_learning.provenance.json.
What was read
- Read in full: the 34-page paper (abstract, theory, design, results, robustness, discussion, conclusion, references, Tables 1–4) and the 23-page online appendix (scoring-rule proof, derivation, weighting, the individual examples of Figure A1 as extracted text, the second experiment's details, Tables A1–A10, the measurement-error simulation, the instruction screens).
- Not read: Figure A1 as an image (its four subjects are described in the appendix text); the data. The version of record was not compared line by line with this ungated version.
What it is
- Question: how far do people depart from Bayes' rule when they learn political facts, once every input to Bayes' rule is measured rather than assumed?
- Argument: without measuring each person's prior and how they read a signal (its likelihood ratio), any pattern of updating can be made consistent with Bayes; earlier evidence of partisan bias rests on assumptions about how signals are received.
- Design, experiment 1 (990 US citizens on MTurk, 17–23 September 2015; 50¢ fee plus up to $4.50 in bonuses; about 15 minutes; no deception):
- three contests of five rounds each: two political statements (one favouring each party, drawn from six; e.g. "median household income fell by more than 4 percent from 2009 to 2012", true) and one about the subject's own quiz score (top or bottom half of 50 earlier takers);
- round 1 elicits the prior; rounds 2–5 each add an independent computer signal "TRUE" or "FALSE" that is correct three times in four, so its likelihood ratio is exactly 3 or 1/3; earlier signals stay on screen;
- beliefs (0–100) are elicited with the incentive-compatible crossover scoring rule (10¢ per round won), 20 seconds per answer to prevent searching;
- weights rake the sample to the 2014 Pew Polarization Survey (n = 10,013); unweighted results are reported too.
- Model: logit(posterior) = δ·logit(prior) + β·log(likelihood ratio), plus interactions for signals consistent with the subject's initial belief (β₂, δ₂). Perfect Bayes: δ = β = 1. Subjects at exactly 50 initially are dropped from the consistency analyses; 0 and 100 are recoded to 1 and 99.
- Experiment 2 (395 MTurk, 8–12 September 2016): one political fact (true or false version, random) and an ego-irrelevant fact (day length in Doha on 8 January 2012), order randomized.
Main results (verified)
Political facts, experiment 1 (Table 2, weighted; 7,664 observations, 990 subjects):
| Signals | δ (prior) | β (signal) |
|---|---|---|
| All | 0.61 (SE 0.02) | 0.73 (0.05) |
| Consistent with initial belief | 0.52 (0.04) | 1.06 (0.10), not different from 1 |
| Inconsistent with initial belief | 0.60 (0.03) | 0.59 (0.06) |
| Pooled with interaction | 0.60; δ₂ −0.085 (0.05), n.s. | 0.59; β₂ +0.47 (0.12) |
| Party identifiers only | 0.60; δ₂ −0.078 | 0.59; β₂ +0.45 (0.14) |
- "Cautious Bayesians": people update in the Bayesian direction at about 73% of the Bayesian amount, Anderson and Holt's 73% in non-political cascades being "nearly the same".
- Modest congeniality bias: signals that fit the initial belief are used fully (106%); signals against it at 59%. Prior weight barely differs by consistency: the bias is in discounting unwelcome signals, not in clinging to the prior.
- No polarization: Democrats' and Republicans' average beliefs always moved in the same, correct direction (Table 1). Example: "income fell by more than 4%" (true) moved from 57.8 to 73.5 after three true and one false signal; perfect Bayes from 57.8 gives 92.5.
- Partisanship adds no bias beyond the initial belief: identifiers updated like independents, conditional on where they started.
- Individuals: of 911 subjects with enough variation, 574 (63.0%) had β < 1, and 235 significantly so (one-tailed). About 24% of political contests showed no change over five rounds; 58 subjects (5.9%) never revised any belief. They were kept in the analysis.
- Benchmarks:
- own quiz rank (ego-relevant): δ 0.63, β 0.64; consistency interaction 0.22 (0.18), n.s.;
- experiment 2: political β 0.70 (consistent 0.99, inconsistent 0.55, β₂ 0.44); abstract fact β 0.85; with consistency terms, abstract β 0.76 and β₂ 0.31 (0.14). Political minus abstract on inconsistent signals −0.20 (0.11), n.s.
- Who departs more (Table A3, party identifiers): liberals β 0.54, β₂ 0.50; conservatives 0.47, 0.60; moderates 0.80, 0.25 (n.s.). Primary voters β₂ 0.54 against 0.28; those who like compromise β 0.72 against 0.50.
- By quiz score (Tables A5–A6): low scorers δ 0.53, β 0.67, β₂ 0.48; high scorers δ 0.73, β 0.82, β₂ 0.31.
- By question and party (Table A2): coefficients mostly 0.2–1.3; one negative (Republicans on a false signal about Obama-era income, −0.70, SE 0.50, n.s.); the largest, 1.28, ran against the party's interest. The author: little evidence of biased assimilation.
- Not a tipping-point process: the largest revisions follow mixed signal histories, not runs of consistent signals (Table A7), as a memoryless Bayesian updater predicts. Mean revision to a lone first "false" signal −15.5 (median −4.3), to a lone "true" +12.3 (median 5.0).
- Measurement error explains part of the caution: simulated perfect Bayesians who round to the nearest 0 or 5 give δ 0.82 and β 0.77, against observed 0.62 and 0.70 (all contests pooled).
- Unweighted results are similar: β 0.78, consistent 1.10, inconsistent 0.63, β₂ 0.47.
Limits
- Best case for updating: unambiguous signals of stated reliability, incentives for accuracy, an academic source, no competing information, no choice of what to read. The author argues both ways on whether this is an upper bound.
- Within one sitting of about 15 minutes; nothing on retention.
- MTurk, weighted to Pew; partisanship and other moderators were measured after the task (tests found no post-treatment bias).
- Binary facts with known truth; attitudes and values are not tested. No preregistration is mentioned.
- δ below 1 means stated beliefs drift toward 50 between rounds; part of this is rounding and noise in elicited probabilities.
What it means for Kurisutina (inference)
- A rule of change for evidence can be written in log-odds with three person-level settings: weight on the prior (δ ≈ 0.6), weight on evidence (β ≈ 0.6–0.85 by domain), and a congeniality term (β₂ ≈ 0.3–0.5, larger for partisans and smaller for moderates). People at the population mean update in the right direction, by about three quarters of Bayes, and less for unwelcome evidence. They do not reverse.
- The settings vary between people: 63% under-react, a few per cent never move, some over-react. That is spread for the empty-slot draws, not a single constant [I].
- For M1: with signals of known reliability, the ground truth for a human-like updater is computable exactly: posterior log-odds = δ·prior + β·(1 + β₂·[consistent])·log LR. Setting δ = β = 1, β₂ = 0 gives the ideal Bayesian as a control [I].
- For the battery: this paradigm is a matched human reference for an evidence-updating probe: a stated prior, four signals of 75% reliability, β = 0.73 overall, 1.06 versus 0.59 by consistency, and 63% of individuals below 1.