Kurisutina

Leveraging large language models and surveys for opinion prediction

arXiv 2305.09620, version 4 (19 May 2026); v1 was posted 16 May 2023. Junsol Kim (Sociology, University of Chicago) and Byungkyu Lee (Sociology, NYU). A preprint under review ("code and data … [suppressed for peer review]"), distributed under the arXiv non-exclusive licence. Provenance: papers/carry_on/kim2024_ai_augmented_surveys.provenance.json.

What was read

All 4,277 lines of pdftotext -layout output (118 pages):

  • main text, discussion and references;
  • Appendices A–I: prompts, architecture, training, the missing-data simulation, baselines, heterogeneity, practical notes, embedding stability and LoRA;
  • Tables A1–A19;
  • the captions of Figures 1–7 and A1–A22.

Figures are images. Only the values that pdftotext recovered from them were used (Figures A7, A12, A13, A14, A22). Nothing in the figures was inspected visually.

Question

Can a model trained on survey answers predict answers that are missing? The paper tests this in three settings:

  1. Imputation: a person skipped an item that was asked in their wave.
  2. Retrodiction: the item was not asked in that year at all.
  3. Unasked: the item was never asked; it is held out of training entirely.

It also asks whether the individual predictions aggregate to public-opinion trends.

Method

  • Data. GSS 1972–2021 cross-sections: 68,846 respondents and 33 waves.

    • 3,699 items: attitudes 68%, behaviours 27%, knowledge 5%.
    • Every item is binarised into positive against negative, and a neutral midpoint counts as negative. The mapping was drafted by Gemini-2.5-Pro and checked by the authors.
    • Intensity is discarded ("strongly agree" and "agree" become the same answer).
    • Each respondent appears in one wave only.
  • Model. Three 50-dimensional embeddings are concatenated and passed to a Deep Cross Network: three cross layers, then three dense layers, then a sigmoid output.

    • Question: the last-token state of the item text from a frozen Alpaca-7B, through one trainable projection. GPT-J-6B and RoBERTa-large were also tried.
    • Respondent: randomly initialised and learned per person from their answers.
    • Period: randomly initialised and learned per wave.
    • Training data include ideology and party. Demographics are left out because they made no difference (Figure A4, one fold).
  • Evaluation. Ten-fold cross-validation, with folds made of:

    • (respondent, item, year) cells for imputation;
    • (item, year) pairs for retrodiction, split into backcasting, interpolation and forecasting;
    • whole items for unasked prediction.

    Individual-level results are reported as AUC. Aggregate results use weighted means with a logistic recalibration, and are reported as r or ρ, MAE, and the share within ±3 and ±5 points.

  • External check. Roper iPOLL questions were matched to GSS items by embedding similarity (≥ .85), confirmed as the same construct by Gemini, and binarised by Gemini.

  • Baselines. Matrix factorisation (MF), text-aware MF, MICE, and per-item time-series regressions.

    • GPT-4o and Gemini-2.5-flash were also prompted on GSS 2018. Each prompt gave 15 demographics, optionally with the person's answers to 10 random items or to the 10 items most correlated with the target.

Results

  • Individual level (Table A6):

    Task Alpaca-7B MF MICE
    Imputation .868 .856 .780
    Retrodiction .862 .807 .727
    Unasked .701 not applicable not applicable
    • For unasked items, GPT-J scores .661 and RoBERTa .573.
    • Retrodiction barely depends on distance from the nearest observed year (AUC .82–.88 in every bin, Table A8).
  • Who carries the signal (permutation, Table A9). Shuffling one embedding at a time moves retrodiction AUC from .857 to:

    • .509 for the question embedding;
    • .720 for the respondent embedding;
    • .849 for the period embedding.
  • Against prompted frontier models (GSS 2018, Table A15).

    • The trained 7B model reaches accuracy .787 and F1 .813 on imputation, and accuracy .723 on retrodiction.
    • The best prompt (GPT-4o with the 10 most-correlated answers) reaches .703 and .714.
    • Demographics alone reach .638–.655.
    • On unasked items the two are about equal (.672 against .703 accuracy).
  • Aggregate level.

    • Imputation and retrodiction track observed yearly means at r > .98. Unasked items reach r = .68, with only 17.5% of predicted means within ±5 points.
    • Against Roper, where the GSS–Roper agreement itself is ρ = .82 (MAE .112), the model reaches:
      • ρ = .81 in years the GSS asked the item;
      • ρ = .79 in years it did not.
    • Per-item trend regressions match the model when an item has many observed years. The model wins when an item was asked only once or twice.
  • Homogenisation. Between-group spreads are reproduced (ρ .50–.83 across groupings). The authors say within-group spread is underestimated in most categories: the MAE on subgroup SDs is .08–.09.

  • Who is predictable. Higher education and income, White respondents, and strong partisans are more predictable in all three tasks. An example is graduate against less than high school: +.028 in retrodiction AUC. MF shows the same gradient, so the authors attribute it to differences in how coherent belief systems are, not to the training corpus.

  • Which items are predictable. Ideologically loaded items are easier, but only weakly: the correlation of the item's |ρ with ideology| with AUC is .225.

    • Some low-ideology items are predicted well: prayer, gender roles, tolerance.
    • Some ideology-loaded items are predicted poorly (AUC < .65): left–right self-placement (.510), trust in media (.523), police treatment of races, defunding the police.
  • Wording. For the same people, the "should have the right to marry" wording gets 9.4 points less predicted agreement than "have the right" (37.35% against 27.93%).

  • LoRA fine-tuning of the backbone changes AUC by less than .01 and lowers F1 by up to .02.

Limits

  • Binary targets only. The authors call this "a particularly favorable case". Intensity is not modelled.
  • Unasked items are held-out GSS items, which have close relatives in training, not new topics.
  • Cross-sectional data: the respondent embedding describes one person at one time. Nothing in the paper predicts how a person changes.
  • Many settings are evaluated on one fold only: Figures A4, A8 and A22.
  • The code and data are not released.
  • Inconsistencies found:
    • The GSS variable count is 7,136 in the text and 7,735 in Figure A1.

    • The text quotes marhomo1 with "should". Tables A4, A13 and A17 give marhomo1 as "have the right" and marhomo as "should have the right".

    • "N = 16 surveys" for marhomo1 in Roper. Table A12 lists 44 marhomo1 surveys over 16 distinct years.

    • Table A13's 9.4-point wording effect is described in the text as "slightly lower".

    • Alpaca retrodiction AUC appears as:

      • .862 (Table A6);
      • .857 (Table A9);
      • .829 (Table A14; .8287 in Appendix I, fold 0).

      MF retrodiction is .807 in Table A6 and .728 in Table A14. The text compares against Table A14 without explaining the gap.

    • The text puts the strongest regression baseline at ρ = .984, MAE .030. Table A16's "All" row has .981/.033 at best. For items with two observed years, the text's ".948/.055" appears in no Table A16 cell.

    • Table A17 lists six items with retrodiction AUC below .5, i.e. worse than chance, and none is discussed:

      • prolife (.388);
      • litcntrl (.442);
      • prespop (.443);
      • successy (.481);
      • bookmark (.491);
      • hsrespct (.493).
    • The subgroup tables A18 and A19 sum to 67,672 respondents in every grouping, not 68,846. The 1,174 missing are not explained.

    • Inferred: the GPT-4o "correlation-top-k" prompt selects context by correlation with the target in training data. An unasked item has no such data, so the unasked comparison is not like for like.

  • Checked:
    • Table A2's shares (2,521, 993 and 185 of 3,699 give 68.15/26.85/5.00%).
    • Table A16's group counts: 525 + 218 + 574 = 1,317 items, and the variable–year cells sum to 8,858 in all three splits.

What it means for Kurisutina

  • Closest published analogue of the GSS pilot, but for a different question.
    • Kim & Lee fill in other answers of the same person at the same time. The pilot asks for the same answers after time has passed and an event has happened.
    • The paper gives a reference level for cross-item prediction from a person's own record: AUC about .87 with a trained model, on binarised items. It gives none for carrying on.
    • Its AUCs are not comparable with the pilot's Brier scores on full code distributions.
  • Person-specific information carries much of the signal. Shuffling the respondent embeddings drops retrodiction from .86 to .72. That is a rough cross-sectional analogue of own against empty.
    • It is not own against other: the embedding includes the person's ideology and party, so it does not isolate "the rest of her beyond her views".
    • Demographics added nothing once answers were known (one fold). Inferred: that is some support for leaving the pilot's donor unmatched on demographics.
  • Slot-only design, frozen backbone. The person lives entirely in a learned vector; the LLM is frozen, and fine-tuning it did not help.
    • This is an existence proof, for cross-sectional prediction, of the goal's constraint that the person should come from the slot and not from changes to the base model.
    • The period embedding is shared by everyone and adds little (.857 against .849). The analogue is the pilot stating the interview year in every condition, so the model's knowledge of the period cancels in the contrasts.
  • In-context records are used less well than a trained slot (inferred). GPT-4o given ten of a person's most informative answers is well below a trained 7B model (.70 against .79 accuracy).
    • The pilot prompts with the full earlier record, so its gap may be smaller. Still, a trained per-person slot is a candidate arm if the in-context conditions barely separate own from other.
  • Homogenisation survives training. Even a model fitted to 68k people underestimates within-group spread, according to the authors.
    • A replica will tend to be more average than the person. That is the failure behind "react like Alice, not the average person with her views".
    • The pilot's own-against-other contrast and a per-person spread check (the replica's predicted distributions against the person's observed variability) guard against it.
  • Predictability is a property of the person. More coherent belief systems are easier to predict (education, partisanship).
    • This matches Beck & Jackson (consistency differs by person) and Hout (item reliability).
    • Carry-on scores should be read against each person's own noise floor, not a global one.
  • Some items are hard whatever the model. Trust in media and left–right placement are hard even cross-sectionally.
    • The pilot's trust, confinan and polviews should be read as possibly near their ceiling before a replica is blamed.
    • A variant of the GSS trust item (trust5, "people can or cannot be trusted") scores .573 here; whether it is the pilot's trust item is not stated.
  • Wording moves predictions. Changing "have" to "should have" moves predicted agreement by 9 points for the same people. A replica's answers to reworded items are not the same measurement.
    • The pilot uses the exact GSS wording in every condition, which is right.

Cross-references

  • docs/research/gss_pilot_design.md: own, other and empty conditions; transition-table baselines.
  • summaries/carry_on/hout2016_gss_reliability.md: GSS item reliability and person stability.
  • summaries/carry_on/beck2018_idiographic_networks.md: consistency as a person-level trait.
  • summaries/carry_on/fisher2018_group_to_individual.md: population models against individuals.
  • summaries/carry_on/berjawi2026_opinion_twins.md, summaries/carry_on/kinzinger2026_soep_twins.md: other person-level survey replicas.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.