Kurisutina

Reliability of the Core Items in the General Social Survey: Estimates from the Three-Wave Panels, 2006–2014

Sociological Science 3: 971–1002, DOI 10.15195/v3.a43, published 14 November 2016, CC BY. New York University; UC Berkeley. Peer-reviewed journal article; it extends MR119 (summaries/carry_on/gss_mr119.md) from one panel to all three. Provenance: papers/carry_on/hout2016_gss_reliability.provenance.json.

What was read

  • The paper: all 32 pages of the extracted text, including Tables 1–3, the notes, the references and the acknowledgements. Figures 1–10 are images; their item labels and captions were read, the plotted points were not.

  • The online supplement (SocSci_v3_971to1002_supp.xls): all 309 item rows × 29 columns. For each item it gives

    • the three wave-to-wave correlations;
    • Alwin–Heise reliability, polychoric and from simple correlations;
    • both stability coefficients;
    • the random-effects reliability and year effects.

    The per-item values below come from the supplement.

Question

How reliable are the GSS core items, measured by answers to the same question two and four years apart, and how much of the change between waves is measurement error rather than change in the person?

Method

  • Data. All three panels are pooled: 2006–08–10, 2008–10–12, 2010–12–14. Wave 1 had about 2,000 people per panel; 1,276, 1,295 and a similar number were interviewed three times. N = 3,875 persons.
    • Items on the rotating core were asked on four of six ballots, so they have about two-thirds of the cases.
    • Voting items use the first panel only.
  • Two estimators.
    • Alwin–Heise (as in MR119): reliability ρ = r12·r23 / r13, with stability β as the correlation between true scores. Correlations are polychoric (tetrachoric for dichotomies) for items with eight or fewer categories, Pearson otherwise.
    • A multilevel model (new here): y_it = β0 + Y_i + Σ τ_t + e_it, with year dummies to absorb change that all persons share. Reliability is between-person variance ÷ (between + within). For ordered and binary items it uses logit models with within-variance fixed at π²/3.
  • Coding. Nominal variables were split into dichotomies (Table 1): married / never married; employed, unemployed, retired; the religious traditions, current and at 16; the 2004 vote.
  • Scales. Existing ones (abortion, misanthropy, gender roles, vocabulary 10- and 7-word) and new ones (suicide, five civil-liberties scales, social life).
  • Exclusions. Geography coded from the address, household context, interviewer items and ethnic ancestry. For ancestry, 61% gave the same answer in all three waves.
  • Grouping. "Relatively fixed" against "not fixed" items, with 19 topic subtypes. The grouping is used only for summaries, not in estimation.

Results

  • Overall (abstract; the text also says 276 items): of 293 core items, 62 (21%) have reliability above 0.85, 71 (24%) between 0.70 and 0.85, and 15 below 0.40, mostly racial and gender stereotype and discrimination items.

  • Table 2. Median Alwin–Heise / multilevel reliability:

    • all items: 0.74 / 0.68;
    • fixed items: 0.94 / 0.92;
    • not-fixed items: 0.71 / 0.65.

    Subtypes (Alwin–Heise medians):

    Subtype Median
    work 0.91
    religious identity & beliefs 0.85
    trust & misanthropy 0.76
    health & morale 0.73
    politics & government 0.67
    public spending 0.64
    confidence in leadership 0.62
    recession experience 0.59
    gender & family 0.59
    race & immigration 0.51
  • Fixed items are fixed. 71 of 74 stabilities lie between 0.980 and 1.020. The exceptions are siblings (β1 0.924, β2 0.971 in the text; 0.968 in the supplement) and raised without religion (0.948).

  • The pilot's target items, from the supplement. Pooled over the three panels, sorted by reliability:

    Item Alwin–Heise, polychoric From simple correlations Stability 1→2 / 2→3 Multilevel N
    attend 0.870 0.846 0.934 / 0.947 0.815 3,848
    partyid 0.838 0.840 0.954 / 0.960 0.834 3,823
    trust 0.795 0.604 0.955 / 0.969 0.734 2,590
    health 0.779 0.686 0.920 / 0.909 0.699 2,551
    satfin 0.728 0.612 0.881 / 0.887 0.609 3,847
    polviews 0.690 0.663 0.951 / 0.964 0.676 3,629
    confinan 0.666 0.534 0.758 / 0.692 0.452 2,556
    finalter 0.605 0.485 0.661 / 0.714 0.365 3,848
    happy 0.597 0.481 0.864 / 0.936 0.523 3,844

    Event variables:

    • Married: 1.018, an estimate above 1 (sampling); EverMarried 1.005; widowed (ever) 0.97; divorced (ever) 0.999.
    • Employed: 0.871.
    • Unemployed: 0.602 (multilevel 0.524); out of work: 0.573.
  • Volatility reads as unreliability. For unemployment, hours worked and the recession items, "both models mistake accurately reported volatility for unreliability". People laid off in 2008–10 changed status more than the year dummies allow, and most went back to work. The authors conjecture these items would be as reliable as employment status in calmer times.

    • The same applies to confidence in the executive, whose referent flipped with the 2009 change of administration: only the 2010–14 panel is usable (0.68).
    • Confidence in banks is very unstable (β 0.76 / 0.69); the authors call that "appropriate" given the financial crisis.
  • Aggregate change (Table 3; year effects from the multilevel models, relative to 2006):

    • unemployment odds ×3.1 in 2010, ×2.2 in 2012, ×1.8 in 2014;
    • weeks worked −2.2 (2010) to −2.7;
    • financial prospects (goodlife) −0.66 to −0.84;
    • confidence in banks −0.74, −2.10, −2.04, −1.47;
    • party identification toward the Democrats (−0.38 in 2008);
    • support for legal marijuana (+0.21 to +2.01; 37% to 60%) and gay marriage (35% to 59%) rising.

    A check on age recovers +2 and +4 years exactly.

  • Where unreliability sits. Racial stereotype items correlate more within a wave (0.40) than across waves (0.22). Most of their unreliability is in the degree of departure from the midpoint: collapsing 1–7 to three or two categories raises reliability (blacks' intelligence, 0.24 as seven categories, 0.65 as three, 0.78 as two).

  • Voting and church attendance are reliable but biased. The explanation offered is selection: people who do not vote or attend also skip surveys. In the 2006 panel, 68% of three-wave completers said they had voted in 2004, against 58% of those interviewed only once.

Limits

  • Face-to-face interviews with show cards, so reliability may be lower in other survey modes (the authors' caveat).
  • Just-identified three-wave models, with no fit test. Shared change is absorbed by year dummies only, so person-specific volatility is counted as error. That matters for exactly the event-linked items.
  • Pooled estimates; no per-person or per-subgroup reliabilities.
  • Inconsistencies in the text.
    • Item counts: 293 core items (abstract), 283 plus about 20 derived (introduction), 276 (Table 2).
    • Siblings β2: 0.971 in the text, 0.968 in the supplement.
    • A note on the vocabulary words has lost a word ("three … had reliability").
    • Political views are given as 0.66 in the text, which matches the simple-correlation estimate (0.663), not the polychoric one (0.690).

What it means for Kurisutina

  • The pilot's primary items are its four least reliable targets. happy 0.60, finalter 0.60, confinan 0.67 and satfin 0.73 are primary for the event strata; attend, partyid, trust and health are 0.78–0.87.
    • Events carry information on these items (feasibility) because they are the volatile ones.
    • They are also the items where a person's own answer, asked twice with no change, agrees least with itself.
    • A person-specific gain there has the least room, and the noisiest target.
  • The ceiling is not a single number here. On recession-linked items the low reliability is partly real, accurately reported change that the model counts as error.
    • A replica that predicts the change (lost job, then satfin down) is predicting part of what the reliability model calls noise.
    • So the pilot should report accuracy beside the item's reliability, not divided by it.
    • The changed-against-unchanged split (declared analysis X1) is the cleaner way to see whether the model predicts real change.
  • The employment events are themselves noisy. Unemployment status has reliability 0.60 (multilevel 0.52), while the marital events are near 1.
    • Some lost-job and found-job pairs will be reporting artefacts.
    • Expect the marital strata (married, divorced, widowed) to give cleaner contrasts than the employment strata when the per-stratum analysis (X3) is read.
  • Population shifts are large and shared. Confidence in banks fell by about two logits for everyone. This is the change a model can get right from knowing "people in general" in 2010, without the slot. The pilot keeps it out of the person-specific contrasts because every condition states the interview years; the population table also carries it.
  • "I told my story and I'm sticking to it." High reliability does not mean truth: answers can be consistently biased (voting, church attendance). For a replica, agreement with a person's past answers is agreement with their account, the "held by the person" field of brief 5.3, not with what happened.
  • Useful for the drift test. Per-item stability β (e.g. partyid 0.95, confinan 0.69–0.76) says how much a person's true position really moves over two years. A running replica that changes more than β allows on stable items (partyid, attend, polviews) is drifting. One that stays put on volatile items after an event is failing to react.

Cross-references

  • summaries/carry_on/gss_mr119.md: the 2006–10 predecessor; same method, one panel; target values differ slightly (MR119 is one panel with polychoric correlations for fewer than 11 categories).
  • docs/research/gss_panel_feasibility.json: our own Heise estimates (Pearson) for the target items; compare with the "simple correlations" column above.
  • docs/research/gss_pilot_design.md: primary items, and the declared analyses X1 (changed/unchanged) and X3 (per stratum) that these numbers inform.
  • summaries/carry_on/hullman2026_validating.md: report replica accuracy against the human reliability ceiling.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.