PNAS 117(15): 8398–8403 (2020), DOI 10.1073/pnas.1915006117, PMC7165437, CC BY-NC-ND 4.0. Also read: the correction,
PNAS 118(50): e2118703118 (2021), PMC8685898. Provenance:
papers/carry_on/salganik2020_predictability.provenance.json.
What was read
- The article: all 667 lines of text converted from the PMC open-access XML by
pmc2txt.py. The converter's self-check passed: nothing lost or added. Coverage:- front matter, significance statement and abstract;
- all text;
- the captions of Figures 1–4;
- data deposition and references.
- The correction: its PMC web page, converted with pandoc, read in full. It is not in the open-access XML subset.
- The corrected Figure 3, from the correction page, was inspected. Figures 1, 2 and 4 were not inspected; only their captions were read.
- Not read: the SI Appendix, which holds the outcome definitions, benchmark details, per-family error analysis and technique survey; the Dataverse submissions; and the commentary "What failure to predict life outcomes can teach us".
Question
How predictable are individual life outcomes from very rich longitudinal data, when hundreds of teams try with the same data and the same metric?
Method
- Data. The Fragile Families and Child Wellbeing Study: births in large US cities around 2000, followed through interviews and home visits at birth and at ages 1, 3, 5, 9 and 15.
- Background data. Waves 1–5 (birth to age 9): 4,242 families and 12,942 variables.
- Outcomes at age 15. GPA, grit and material hardship (continuous); eviction, caregiver layoff and caregiver job training (binary). All are self-reported.
- Design. The common task method: 160 valid team submissions.
- Training outcomes were given for half the families. The rest were split into a leaderboard set and a holdout set kept by the organisers.
- Metric: $R^2_\text{Holdout}$ = 1 − MSE / MSE of predicting the training mean.
- Benchmark. A regression with four predictors: the mother's race/ethnicity, marital status and education at birth, plus a measure of the outcome or a close proxy at age 9.
Results
-
Best predictions are poor (corrected Figure 3; best submission per outcome, $R^2_\text{Holdout}$):
Outcome Best $R^2_\text{Holdout}$ Material hardship .23 GPA .19 Grit .06 Eviction .05 Job training .05 Layoff .03 The text rounds these to "about 0.2" and "about 0.05".
-
The simple benchmark is nearly as good. Four variables, including the earlier measure of the outcome, come "only slightly worse" than the best of 160 teams, and beat many of them.
-
Teams agree with each other more than with the truth.
- "The distance between the most divergent submissions was less than the distance between the best submission and the truth."
- Ensembles barely help.
-
Error belongs to the family, not the method.
- Most families are predicted well by every team, and a few badly by every team.
- The hardest cases are far from the mean (unusually high or low GPA). For binary outcomes the errors are large exactly where the event happened.
-
Predictions are compressed. Figure 3 shows predicted values in a narrow band around the training mean, whatever the truth.
-
Correction (2021). A coding error: the benchmark used linear regression for all outcomes, including the binary ones for which the paper said logistic.
- Figure 3 and SI Figures S6–S9 changed.
- In the SI, layoff's best-to-benchmark ratio changed from four times to three times, and a "100%" became 95%.
- The authors say the conclusions are unaffected.
Limits
- One cohort (families of mostly unmarried parents in US cities), six outcomes and a six-year gap.
- The task allowed training on half of wave-6 outcomes, so it is easier than pure forecasting.
- The authors say predictability will vary with outcome, time gap, data source and group.
- Selecting the best submission on the holdout set and scoring it on the same set is slightly optimistic. The authors judge the bias small.
- Inconsistencies found:
- The published text said the benchmark used logistic regression for binary outcomes. The correction states it did not.
- "About 0.05 for the other four outcomes": the corrected figure gives grit .06 and layoff .03.
- Checked: the correction's changes (Figure 3 values, SI wording) are consistent with its stated cause.
What it means for Kurisutina
- A ceiling for predicting how a specific person turns out. With about 13,000 variables on a family over nine
years, no method explained more than about a fifth of the variance in outcomes six years later, and most explained
under a twentieth.
- A replica predicting a person's later answers after an event works against the same kind of ceiling. Most of what happens to a person, and how they respond, is not in the record.
- Carry-on scores should be judged against a floor and a ceiling, not against perfection:
- the floor is the transition table and no change;
- the ceiling is test–retest noise (Hout) plus unrecorded events.
- The pilot's population baseline is Salganik's benchmark (inferred). P(later | earlier, event) uses the person's
earlier answer plus a few facts. In the Fragile Families challenge, that kind of model was nearly as good as
anything.
- Expect own-condition gains over the transition table to be small in absolute terms. A small gain that survives the other and empty controls is still the result the pilot is looking for.
- The informative cases are the unusual ones. Errors concentrate in the families where the unusual thing
happened, or who ended up far from the mean.
- "React like Alice" is by definition about deviating from the population expectation. Any model regressing to the mean fails there, and average scores hide it.
- Proposal: report pilot and later results separately for people whose later answer changed and those whose did not. Scoring changers separately shows whether a replica predicts change at all, or only stability.
- Agreement among models is not agreement with the person. 160 teams converged on each other, not on the truth.
- For the pilot's arms (inferred): if M1, M2 and M3 agree with each other more than any of them agrees with the person, they share a population prior rather than reading the person. The between-arm distance against the arm-to-truth distance is a cheap homogenisation diagnostic (proposal).
- Compression toward the mean. Every method compressed its predictions. This matches Kim & Lee's homogenisation and Bisbee's small variance: predicted spread across people should be reported alongside accuracy.
Cross-references
summaries/carry_on/kim2024_ai_augmented_surveys.md: homogenisation in a trained slot.summaries/carry_on/kiley2020_personal_culture.md: how rare lasting change is.summaries/carry_on/hout2016_gss_reliability.md: test–retest ceilings for the pilot items.docs/research/gss_pilot_design.md: the transition-table baselines.- Lundberg et al. 2024 (the interview follow-up) is being read by subagent 7.