PNAS 121(24): e2322973121 (online 4 June 2024), DOI 10.1073/pnas.2322973121, PMC11181083 (PMC version 1), CC BY 4.0.
Cornell, Princeton and St. Joseph's University. Contributed by Kathryn Edin; reviewed by Michael Hout and Mario L.
Small. Provenance: papers/carry_on/lundberg2024_unpredictability.provenance.json.
What was read
- Article: all 479 lines of the text derived from the PMC XML. That covers the significance statement, abstract, all sections, the footnote, the data statement, the figure captions and all 49 references. The converter checks that the text keeps every non-whitespace character of the XML.
- SI appendix: all 1,909 lines of
pdftotext -layoutoutput (34 pages), obtained from Europe PMC's supplementary-files service. It covers:- acknowledgements;
- sample selection;
- recruitment and non-response (Table S1);
- the interview procedure;
- the proof of the error decomposition;
- both interview guides (young adult; primary caregiver).
- Figures. Figures 1–5 are conceptual diagrams and were viewed as images. SI pages 5–7 (Figs. S1–S2, Table S1, Fig. S3) were rendered and viewed. Values below taken from Fig. S2 are read off the plot and are approximate.
- Not read:
- Salganik et al. 2020 (the Fragile Families Challenge paper), which the main session read separately;
- the PNAS commentary linked to this article (e2409327121);
- the redacted interview transcripts (available from the authors on request);
- the FFCWS data.
Question
Why are some life outcomes hard to predict, even with rich data and hundreds of teams? The case is one task from the Fragile Families Challenge: GPA at age 15, predicted from 12,942 survey features measured from birth to age 9.
Method
-
Framework. A prediction task has three parts:
- features measured in a "feature observation window";
- an outcome measured later, after an intervening period;
- a process that produces the training sample.
Out-of-sample squared error splits exactly into two parts (Eq. 1; proof in SI section 4):
- irreducible error, E[V(Y | X)]: the spread of outcomes among "observationally identical" people who share all feature values. It is fixed by the task: "Irreducible error cannot be reduced by a new learning procedure; the only way it can be decreased is by changing the task."
- learning error, E[(f̂(X) − E(Y | X))²]: the distance between the prediction and the true conditional mean.
-
The Challenge, as described here.
- 4,242 children from 18 cities: 2,121 families for training, 530 for feedback and 1,591 held out.
- The best holdout R² for GPA was 0.19. On this scale 0 means predicting the training mean and 1 means perfect prediction.
-
Qualitative sample (SI section 2).
- The frame is the 1,591 holdout families in three cities.
- Families were stratified by city and by tercile of the GPA predicted by the best submission (Rigobon's).
- In each of the 9 strata, the families with the most positive and the most negative residual were taken with probability 1 (18 families). One random family was taken from each residual tercile (27 families).
- 45 were sampled and 7 did not take part: 2 refused, 3 were never scheduled and 2 were not reached. 2 were replaced, so 40 families took part, 38 of them from the original draw.
-
Interviews.
- 16 researchers conducted 114 semi-structured interviews with 73 respondents: 66 interviews with 39 young adults and 48 with 34 primary caregivers. They were in English and Spanish, mostly in person.
- Each traced a life history from birth, focused on ages 9–15, then on age 15 to the interview.
- Interviewers worked in pairs. The primary interviewer did not know the predicted or realised GPA. The secondary interviewer knew them and could probe at the end.
-
Analysis. Inductive.
- Several team members answered questions about each case and met to discuss it.
- Later, one researcher wrote a case summary, others read it with the interview, and the team finalised it.
- No coding frequencies are reported.
Results
The paper reports cases, not counts.
- Irreducible error has three sources (Fig. 4).
- Unmeasurable features: consequential intervening events after the window.
- Bella. Married, employed parents who were high-school graduates; by 15 she was fighting and struggling at school, and she later dropped out. Her father died unexpectedly, her mother became depressed, and Bella "found herself effectively parentless". No predicted or actual GPA is given for her.
- Charles. An online charter school. In the term for which GPA was measured, he worked from the basement, where he played video games. He reported a GPA of 1.75 against 3.15 predicted. His mother said "no more downstairs" and his grades recovered: "Ninth grade, I did terrible, then all the other years, I did As."
- The authors: "Consequential intervening events are an important source of irreducible error, particularly for life outcome prediction tasks with long time horizons like that from age 9 to age 15."
- Unmeasured features that could have been measured in the window.
- Lola. Her mother was "dabbling in illegal activities". An elderly neighbour got Lola ready for school each morning. Her grandparents gave health insurance and an address for a better school, and later remodelled a basement for her and her mother. GPA 3.75 against 3.04 predicted.
- Unnamed cases: a wealthy out-of-state family that mentored one boy; a landlord who built a home gym for a tenant's son.
- However: "Our qualitative interviews did not reveal a small set of additional predictors that we think would have greatly improved predictive performance."
- Imperfectly measured features.
- Hennessey. At 9 she chose "not very close" to her mother, the lowest of four options. Her mother ignored her; they fought physically; her mother called the police and told her she "might not live that long".
- GPA 1.25 against 2.71 predicted. The coarse item made her look like any other "not very close" child.
- Unmeasurable features: consequential intervening events after the window.
- Learning error. Tasks have many features, few cases and little expert knowledge.
- With 12,942 binary features there are 2^12,942 possible feature vectors.
- Even 1,000 features with pairwise interactions give 500,500 parameters.
- Generalising (Fig. 5).
- More cases shrink learning error only.
- More features shrink irreducible error but can raise learning error.
- Two exceptions to low predictability: a natural low-dimensional representation, "such as when a lagged outcome is a good predictor of that outcome in the future", and very short horizons, "(e.g., 1 d)".
- Their conjecture for tasks on longitudinal survey data: "low levels of predictability will be the norm."
- Recommendation. Study within-group variability alongside between-group means. Estimate the two error components where possible; they cite a method (ref. 44) but do not apply it.
- Residual sizes (my computation from the text): Charles −1.40, Hennessey −1.46 and Lola +0.71 GPA points.
- In Fig. S2 the respondents' residuals run from about −1.6 to +1.2 (read off the plot).
- Their predicted GPAs span only about 2.5 to 3.7 on a 1–4 scale (read off the plot).
Limits
- Scope. One outcome (self-reported GPA), one study, three cities and 40 families. The three sources are illustrated by four named and two unnamed cases. Nothing says how many of the extreme cases showed each source.
- Hindsight. The accounts are retrospective, given years after the outcome. Blinding is described only for the primary interviewer, not for the analysis. Hindsight explanations cannot be excluded, and the paper does not discuss them.
- The components are not separated. A residual mixes irreducible and learning error, and an interview can only suggest which one dominates a case. The paper estimates neither component.
- Outcome error is not a listed source. Measurement error in the outcome (one term's self-reported GPA) is not among the sources, although Charles's case is partly that (inferred).
- Differential non-response. 5 sampled families were not replaced, 3 of them with the most negative or bottom-third residuals. The SI concedes that "our respondents under-represents those whose GPAs were lower".
- Inconsistencies found:
- Sample size. The main text says: "Our sample size of 40 families was determined by budget constraints." The SI says 45 families were sampled, for budget and time reasons, and 40 took part.
- Replacements. The SI says replacements had similar predicted and residual GPA to the cases they replaced. Table
S1 shows the City C replacement moved from "Extremely high residual" to "High residual".
- My computation from Table S1: of the 18 extreme cases sampled with probability 1, 16 were interviewed, 8 low and 8 high.
- Fig. 2 caption. In the PMC XML it has a "Second," sentence but no "First".
- Title. The SI calls the paper "life trajectory prediction tasks"; the article says "life outcome prediction tasks".
- Fig. S3. It is captioned "Domains of predictors" but shows which respondents were surveyed at which ages, not topics. The main text cites it for "many different topics".
- Checked:
- 66 + 48 = 114 interviews, 39 + 34 = 73 respondents, and 16 interviewers named in the SI;
- 2,121 + 530 + 1,591 = 4,242;
- 18 + 27 = 45, 45 − 7 = 38 and 38 + 2 = 40;
- 1,000 + C(1,000, 2) = 500,500;
- the non-response breakdown by city, tercile and residual in the SI text matches Table S1.
What it means for Kurisutina
- The pilot's contrasts in Lundberg's terms (inferred).
- The population table P(later | earlier, event) estimates E(Y | X) for X = (earlier answer, event).
- The own condition adds the rest of the earlier record to X. So D1 < 0 needs that record to say where this person falls within the cell of people with her earlier answer and event.
- Lundberg's interviews traced within-cell deviations mostly to things no record held: later events, unmeasured networks and coarse answers. They did not find a missing predictor.
- So expect D1 and D2 near zero for questionnaire persons, as the pilot's expectations E4 and E5 already state.
- The Brier score splits the same way (verified by derivation and a simulation).
- For a forecast vector f and a one-hot outcome, expected Brier = E[1 − Σ_k p_k(X)²] + E[Σ_k (f_k − p_k(X))²]. Here p(X) is the true conditional distribution.
- The first term, the expected Gini impurity, is the irreducible error of the task's feature set. The second is learning error.
- Simulation check (mine): with p = (.5, .3, .2) and f = (.7, .2, .1), Brier was 0.681 against 0.62 + 0.06.
- Proposal (data only, no model call, v1 untouched):
- For each item and stratum, compute the expected Gini impurity of the leave-one-out table as that table's own floor.
- The table's Brier minus this floor is its learning error. It is largest in sparse cells such as widowed (73 pairs in the feasibility count).
- A model forecast can go below the floor only with information beyond (earlier answer, event). Report D1 beside the floor.
- The ceiling on predicting individual change (inferred).
- For levels, the pilot is in the paper's easy class: a lagged outcome exists. So raw Brier is dominated by stability.
- The change part is in the hard class: a two-year horizon, intervening events the prompt does not name, and one measurement at one time.
- Charles's case is the GSS problem in miniature: a transient state at the measurement date that later reverts.
- Kiley & Vaisey find most wave-to-wave change reverts.
- Hout & Hastings put the reliability of the pilot's targets at .60–.87.
- That variance is irreducible for any forecaster of a single later answer.
- Timing is also unmeasured. The prompt says only that the event happened between the interviews, not when. Reactions peak and fade within such intervals (Luhmann et al. 2012, 2014).
- Agency (partly verified).
- The paper never uses the word agency (verified by search).
- Its cases include the person's own acts (Charles's gaming) and other people's (a mother's rule, a landlord, a mentor family). All are classed as intervening events or unmeasured features.
- For a replica (inferred): the person's own later choices are intervening events that the replica must generate itself when running open loop. A replica that chooses differently will diverge afterwards, with no fault in how it reacts.
- Proposal: score reactions conditional on the person's real choices (closed loop), and score agreement on the choices separately.
- Q1: react like her, not like the average person with her views (inferred).
- In this framework the average person with her views is E(Y | X). Beating it needs features the record lacks.
- A Kurisutina replica can address all three sources by design:
- intervening events are the input in a carry-on test, not unmeasured. But only the events we describe, and they should include the person's role in them (Brüning);
- unmeasured features: interviews can capture networks and support (Lola);
- coarse measurement: graded, narrative accounts keep intensity that answer categories lose (Hennessey).
- Surveys cannot do this. So a narrative slot should beat a questionnaire slot on reactions. This is untested.
- Their design as a method for us (proposals).
- After the pilot's primary analysis. Take the pairs where the own condition is most wrong and most right relative to the table, stratified by predicted value. Inspect those records for changes between waves that the prompt did not name, such as health, a move or the household. This is a data-only analogue of their residual-stratified interviews.
- With a real participant. After person and replica diverge, interview the person twice: first with an
interviewer who has not been told the replica's prediction, then with one who has. Classify each divergence:
- an intervening event;
- something missing from the slot;
- something too coarse in the slot;
- the model's default (learning error). This is the drift question.
- Pre-registration. The paper names predictions pre-registered before outcomes exist, and common-task designs, as the strongest evidence. The pilot's freeze and declared analyses are of that kind.
Cross-references
summaries/carry_on/salganik2020_predictability.md: the Challenge itself (summarised by the main session; not read for this summary).summaries/carry_on/kiley2020_personal_culture.md: most change reverts; phi per item.summaries/carry_on/hout2016_gss_reliability.md: item unreliability as part of the irreducible error.summaries/carry_on/bruning2022_separation_initiators.md: the person's role splits one event.summaries/carry_on/luhmann2012_adaptation.md,summaries/carry_on/luhmann2014_its_about_time.md: the timing of reactions.summaries/carry_on/fisher2018_group_to_individual.md: group means and individuals.summaries/carry_on/namazova2025_open_loop.md: open-loop divergence.docs/research/gss_pilot_design.md: D1–D3, X1 and the table baselines.