Matthias R. Mehl, James W. Pennebaker (University of Texas at Austin). Journal of Personality and Social Psychology 84(4): 857–870 (2003), DOI 10.1037/0022-3514.84.4.857. Copy read: a course-reading PDF hosted at the University of British Columbia (14 pages, the full article). Read by researcher R5c (voice), 5 October 2026. Provenance: papers/voice/mehl2003_sounds_of_social_life.provenance.json.
What was read
- Read in full: every line of the pdftotext conversion (1,571 lines): abstract, introduction, method, results, discussion, ethics section, future research, footnotes, Tables 1–2, references.
- Figures 1–3 (bar charts of the retest and person–other correlations) exist only as images; they were extracted with
pdfimagesand read visually. Their values are transcribed below. - Not read: nothing else.
What they did
- Participants: 52 introductory psychology students (28 women, 24 men; mean age 19.0).
- Recording: each wore the Electronically Activated Recorder (EAR) for two 48-hour weekday periods, four weeks apart. The EAR records 30-second snippets of ambient sound about every 12.5 minutes during waking hours (about 4% of the day). Participants could not tell when it was recording and could erase anything before handing it in; only one listened, and none erased.
- Footnote 1: three writing sessions (trauma or neutral topic) took place between the two periods; their effects are reported elsewhere.
- Social environment: coded by judges from ambient sound with the SECSI scheme (location, activity, interaction); mean judge reliability α = .94.
- Language:
- all speech was transcribed, keeping repetitions, fillers and non-fluencies, and attributed to the participant (P) or to others nearby (O);
- transcripts were scored with LIWC (word percentages in dictionary categories) on 23 categories: the 16 that Pennebaker & King (1999) found reliable in writing, plus swear words, non-fluencies and fillers (for speech), plus all personal pronouns;
- fillers ("like", "you know", "I mean") were converted to unique tokens by transcribers.
- Analysed samples: 49 people for environments (at least 50 intervals per period) and 47 for language (at least 50 words per 2-day period). The EAR captured on average about 1,064 words per person over the four days. The text gives the range as 164 to almost 3,000; Table 2 prints a minimum of 146.
Main results (verified)
Self-report check. Participants rated the EAR's effect on their behaviour (M = 1.43) and on their talking (M = 1.25) as minimal, on 1–5 scales.
Base rates of everyday speech (Table 2, % of words).
- About one word in seven is a personal pronoun; first-person singular is 6.9%.
- Articles 3.4%, prepositions 8.9%, negations 2.8%, fillers 1.7%, non-fluencies 0.4%, swear words 0.5%, present tense 15.9%, social words 11.1%.
- Variation between people is very large. One person swore as often (3.4%) as another referred to themselves (3.2%). The text says "11 participants (47%)" did not swear at all; 11 of 47 is 23%, so one of the two figures is misprinted.
Four-week retest of spoken style (Figure 2, N = 47; r between two 2-day samples).
| Category | r | Category | r |
|---|---|---|---|
| Swear words | .86 | Positive emotion words | .51 |
| Non-fluencies ("uh", "er") | .62 | Negative emotion words | .50 |
| Fillers ("like", "you know") | .59 | Present tense | .34 |
| Articles | .44 | Social processes | .25 |
| Prepositions | .37 | Causation | .23 |
| "I" | .31 | Exclusive ("but", "except") | .24 |
| "We" | .31 | Past tense | .21 |
| Negations | .31 | Tentative ("perhaps", "maybe") | .16 |
| References to others | .27 | Insight | .10 |
| Word count | .24 | Inclusive | .08 |
| Words over six letters | .24 | Discrepancy ("would", "should") | −.13 |
| "You" | −.10 |
- Category averages: standard linguistic .41, psychological processes .24. Relativity is printed as .29, though the mean of its four bars is about .22 (my computation; the other two averages match Fisher-z means of their bars).
- 16 of 23 correlations were significant (one-tailed p < .05, critical r about .24).
- The authors' reading: the "variety of topics of people's daily conversations suggests that this stability reflects a stability of linguistic style more than linguistic content"; categories unique to speech carry "valuable stylistic information over and beyond written language".
Style shared with conversation partners (Figure 3, correlation between the participant's and their interlocutors' overall profiles).
- Standard linguistic categories average .35. The highest are word count .72, fillers .60, negative emotion .56, non-fluencies .48 and swear words .43.
- Low or absent: "I" .06, "you" −.16; these are used complementarily in dialogue.
- The authors: linguistic synchrony "certainly fosters temporal stability in people's linguistic styles". They could not tell whether people choose similar partners, impose their style on them, or share contexts.
Across everyday contexts (exploratory; 45 people, fewer for some contrasts).
- Location (home, outside, public places): only word count and "we" differed. In public places people said fewer words per 30 seconds, and "we" rose from 0.7% to 1.0%.
- Activity (13 people): only "I" differed, 10.2% in amusement against 3.8% while working.
- Phone against face-to-face (at home): 4 of 23 categories differed. On the phone people said more words, used more non-fluencies (1.0% against 0.4%) and fewer "we" and inclusive words.
- Overall: language was "stable across locations, activities, and modes of interaction".
Social environments (Figure 1). Four-week retest averaged .54 for interactions, .50 for activities and .65 for locations (alone .64, talking .63, amusement .73, working .73, eating .15).
Sex differences in speech. Men used more long words, more articles, fewer "I" (6.2% against 7.5%), four times the swearing (0.9% against 0.2%), fewer fillers (1.3% against 2.1%) and more anger words (1.1% against 0.4%).
Limits
- One cohort (introductory psychology students, Texas), 47–52 people, weekdays only.
- Two 2-day windows only four weeks apart; nothing about years.
- About 500 sampled words per person per period, so the retest values are for small samples. LIWC counts words in fixed dictionaries; it measures word-category style, not syntax, rhythm or discourse.
- The context contrasts are underpowered (13–45 people), so "no difference" is weak evidence.
- Interlocutor speech is pooled across all partners.
What it means for Kurisutina (inference)
- The most stable markers of spoken voice are the ones written text and cleaned transcripts throw away: swearing (r .86 over four weeks), non-fluencies (.62) and fillers (.59). They are more stable than any function-word category (.31–.44). Speech collected for a slot must be transcribed verbatim, with fillers, hesitations and false starts kept and marked. A replica that talks without its person's "erm"s and "like"s has lost the most reliable part of their spoken voice.
- The human retest reference for simple features is modest (r about .3–.6 for most categories from about 500 words per occasion). A voice test built on single feature rates needs much larger samples, or must aggregate many features, to have a reference tight enough to fail anything. Spearman–Brown reasoning says reliability rises with words sampled; this paper does not test that.
- Voice is partly co-produced with the interlocutor. Fillers, swearing, non-fluencies and length match the conversation partners (.43–.72). A test must hold the interlocutor fixed when comparing person and replica (same partner, prompts or script). The slot must carry the person's accommodation, how they adjust to whom they talk with, not only a fixed profile.
- Within everyday talk, context changes little (location, activity, phone). Large shifts come between registers (spoken against written, formal against casual; see PAN 2023). Everyday conversation can be treated as one register for a first voice test.
- A recording design that works exists: unobtrusive random sampling of everyday speech, with consent and erasure rights, captured about 1,000 words per person in four days at 4% duty cycle. The bystander problem is real; the authors limited it with short snippets, transcriber anonymisation and directional microphones in later versions.