GSS Methodological Report 119, NORC. UC Berkeley. Dated 21 June 2012; the August 2014 version corrects a coding error
in the civil liberties scale items. A technical report, not peer reviewed. The same data were later extended to all
three panels (2006–2014) in Hout & Hastings, Sociological Science 3 (2016), which has not been read.
Provenance: papers/carry_on/gss_mr119.provenance.json.
What was read
All 22 pages: text, Tables 1–3, footnotes, figure labels and references. Figures 2A–C plot, for every item, its reliability, the two stability coefficients and the three wave-to-wave correlations. Those values are points on a chart and were not transcribed. Only numbers stated in the text or tables are given below.
Question
How much of the change in a person's answers across the GSS 2006–2008–2010 panel waves is measurement error, and how much is real change in that person, for 265 core items?
Method (Heise 1969, as implemented by Alwin 2007)
- Model. An observed answer = true score + error. The true score at each wave depends only on the previous wave (a Markov chain: no direct 2006→2010 effect). Reliability ρ² is assumed constant across waves.
- Estimation. With three waves the model is just identified:
- reliability ρ² = r₀₁·r₁₂ / r₀₂;
- stability β₁₀ = r₀₁/ρ² = r₀₂/r₁₂ and β₂₁ = r₁₂/ρ² = r₀₂/r₀₁.
- Correlations are polychoric for items with fewer than 11 values, Pearson otherwise.
- Checks and definitions.
- There is no fit test. The check is that fixed items (sex, birth month, parents' education) should show a stability of 1, and they did.
- Stability measures change in people's rank order. A shift of everyone in the same direction does not lower it; change that differs between people does.
- Coverage.
- Unordered categorical variables were split into dichotomies: married, never married; employed, unemployed, retired; no religion now or at 16; religious tradition dummies. Occupations were scored by prestige and SEI.
- Scales were built for vocabulary, abortion, gender roles, suicide, civil liberties and socialising.
- Items were grouped into 5 types and 16 subtypes.
- Sample. 1,276 respondents were interviewed in all three waves, 823 of them in person every time. The vocabulary words had only about 280 cases.
Results
Reliability by type (Table 2; median / mean, number of items).
| Type | Subtype | Median | Mean | Items |
|---|---|---|---|---|
| Facts | all | .918 | .882 | 88 |
| demographic | .958 | .933 | 22 | |
| religious | .964 | .940 | 21 | |
| SES, fixed | .887 | .880 | 10 | |
| SES, can change | .814 | .839 | 25 | |
| behaviours | .754 | .757 | 10 | |
| Words | all | .750 | .782 | 11 |
| Beliefs and values | all | .706 | .690 | 97 |
| sex & sexuality | .839 | .801 | 16 | |
| religious | .811 | .811 | 12 | |
| other social | .757 | .711 | 20 | |
| civil liberties | .709 | .720 | 20 | |
| gender & family | .622 | .594 | 12 | |
| racial | .490 | .515 | 16 | |
| Placements & evaluations | all | .675 | .673 | 22 |
| Attitudes | all | .658 | .664 | 63 |
| institutional | .600 | .598 | 14 | |
| taxes & spending | .681 | .663 | 29 | |
| other social & political | .710 | .710 | 20 |
- Overall. Of 265 items, 80 had reliability above 0.85 and 84 between 0.70 and 0.85. Five came out near 1.1, attributed to sampling error. Confidence in the executive branch (confed) came out at 10.0 and was excluded, because the president changed between waves.
- Beliefs and values.
- Abortion and sexual-behaviour items: 0.76–0.90. Belief in God, the afterlife and the Bible: above 0.75.
- Children thinking for themselves vs obeying: 0.82. Getting ahead by hard work or luck: 0.45.
- Stouffer civil-liberties items: 0.72 on average.
- Racial and gender beliefs had the lowest reliability, "a matter of great concern". Two explanations are left open: the 2008 campaigns changed what the items mean (the model assumes meaning is fixed), or the newer items are poor.
- Stability: beliefs and values are "far more stable than they are reliable". Very few beliefs or values had stability below 0.80. The exceptions were two items on whites' and blacks' intelligence, read as Obama's candidacy shifting a stereotype. After correcting for error, even racial and gender beliefs were steadier than the raw answers suggest.
- Where true individual change showed up: events that hit some people and not others.
- Placements tied to the recession: standard of living vs parents β = 0.87 / 0.89; financial satisfaction 0.85 / 0.86; fear of job loss 0.84 / 0.79; prospect of a better life 0.77 / 0.81; job satisfaction 0.73 / 0.82; finances better or worse 0.68 / 0.71.
- Marital happiness 0.67 / 0.92.
- Employment, hours and income: mostly 0.75–0.85.
- Confidence in the executive, the courts and banks; spending on parks, environment, health, drug rehabilitation and crime.
- The authors take the low stability on recession items as evidence the model is appropriate.
- Shifts of the whole population (distribution changes) are a separate kind of change. Among 821 people who answered all three waves, "hardly any" confidence in financial institutions rose from 15% to 23% to 42%. Yet mean stability was 0.72: the population moved and people's rank order changed only modestly. The authors say this distribution shift "will, in fact, be more interesting than either unreliability or instability" in many applications.
- Direction over time. Instability was greater in 2006–08 than in 2008–10, against expectation, for almost all items, even the vocabulary words.
- Where the model breaks. When earlier status affects later status directly (employment in a recession) or the object of a question changes (the executive after 2008), reliability and stability cannot be separated.
- Interview mode. For 33 of 54 items tested, in-person and phone reliability differed by less than 0.10. There was no consistent advantage for in-person interviews.
Limits
- One panel (2006–10), three waves two years apart. Just identified, so no fit test.
- The Markov assumption and constant reliability are known to fail for some items.
- The results are item averages; nothing is reported per person or per subgroup (left as future work).
- No standard errors for the reliability ratios.
What it means for Kurisutina, especially part (b)
-
It sets the noise floor for "did the person change?" Any claim that an answer changed because of what happened to the person has to beat the item's unreliability. Two answers from the same person to a racial-belief item correlate far below 1 even with no true change (reliability about 0.5). For facts and sexual or religious beliefs the floor is low (0.8–0.96). The feasibility analysis should (1) weight items by reliability or model them as latent, and (2) prefer high-reliability items as targets for predicting change.
-
Real personal change in 2006–2010 was concentrated where events hit people unequally: jobs, finances, marriage, confidence in institutions under new leadership. That is exactly the pattern a carry-on test needs: a life event recorded between waves, and a change in the person's answers that other people did not share.
-
Two kinds of change to separate, which maps onto the drift question.
- The whole population moving (distribution shift, e.g. confidence in banks after 2008) is change a replica could also produce by knowing what "people in general" did, without the person slot.
- A change in the person's position relative to others (instability) is person-specific.
A replica that gets the population shift right but not the person's relative movement is "the average person with her views". Scoring for (b) should report both parts separately.
-
Most beliefs barely move in two years. With stability above 0.80, most people keep their rank. "Predict no change" will be a strong baseline on beliefs, as the exact no-learning control was in the memory work (
summaries/memory/gluck2002_weather_prediction.md). Useful targets will be the items that do move, conditional on an event. -
Question meaning can change under the person. Confidence in "the people running the executive branch" measured something different after the 2008 election. For a replica carrying on, the world changes the referents of its beliefs, not only the evidence for them. That is a case to include in the carry-on tests.
Cross-references
- Eckstein et al. 2022 (
summaries/carry_on/eckstein2022_context.md): individual updating depends on context. Here, change depends on which events reached which people. docs/research/carry_on_goal.md, part (b): the GSS panel feasibility plan this report underpins.