Psychological Science 30(3):405–414, doi 10.1177/0956797618818476. Not in PMC. Read from the first author's own copy (jessiesun.me), which is the published article followed by the supplemental material. Read for research direction R2, 4 October 2026. Provenance: papers/amadeus/sun2019_momentary_self_knowledge.provenance.json.
What was read
- Read in full: all 29 pages of the author PDF, converted with
pdftotext:- the article: abstract, introduction, method, results, discussion, limitations, statements and references;
- the supplement: coding-protocol details, the model figure caption, Table S1 (within-person estimates by item), Table S2 (between-person correlations), and Tables S3 to S6 (verbatim transcripts for the 25 largest self–observer discrepancies in agreeableness and neuroticism, in each direction).
- Not read: the OSF data, scripts and the password-protected transcript file. Figure 1 is a plot, read only as extracted text.
- Extraction caveats: the two-column layout extracted in reading order. Table 1 and Table S2 came out as runs of numbers. I re-assembled Table 1 by column order, consistent with the text. I do not quote Table S2's cells because their order is ambiguous.
What they did
- Sample. 248 students (Washington University in St. Louis, 2012–2013; 173 women; mean age 19.2), from the 434-person PAIRS study.
- Self-report. 4 times a day for 15 days (noon, 3, 6 and 9 pm), people rated how they had been in the preceding hour. The 9 items came from the Big Five Inventory and covered state extraversion, agreeableness (only if with others), conscientiousness and neuroticism. Openness was not asked.
- Recorded behaviour. For about a week, the Electronically Activated Recorder (EAR, an iPod app) recorded 30 s of ambient audio every 9.5 minutes from 7 am to 2 am.
- 152,592 usable recordings from 304 people.
- Participants could listen to their files afterwards and delete any. 15 people deleted 99 files.
- Coding. For each hour matched to a self-report, at least 6 of 108 research assistants listened to the 6 or 7 clips (3 to 3.5 minutes). They rated the person on the same items ("seemed ...") or chose "no way to tell". Hours that at least 3 coders found uninformative were dropped: 15% of them.
- Analysis. 2,938 matched hours (2,519 for agreeableness), with at least 5 matched hours per person. Multilevel structural equation models with latent variables (corrected for unreliability) and random slopes per person. The key estimate is the within-person slope: when you say you were more X than usual, did observers hear more X?
- Not preregistered. The transcript analysis was added after seeing the results.
Main results (verified)
- People vary a lot from hour to hour. At least 45% of the variance in each state is within-person, in both self-reports and codings.
- Within-person reliability of self-reports: extraversion .82, agreeableness .26, conscientiousness .47, neuroticism .72.
- Within-person reliability of codings: .93, .62, .76, .73.
- Self-knowledge of states differs by domain. Within-person standardised self–observer agreement:
- extraversion β = 0.63 (95% CI 0.60 to 0.66);
- conscientiousness β = 0.47 (0.40 to 0.55);
- neuroticism β = 0.27 (0.21 to 0.32);
- agreeableness β = 0.20 (0.11 to 0.32). The two weaker intervals do not overlap the two stronger ones. Table S1 gives slightly different composites (agreeableness 0.22, conscientiousness 0.48), presumably a separate Bayesian run.
- The authors' interpretation.
- People know when they are more sociable or more lazy than usual.
- Worry is hard to hear: many hours of high self-reported neuroticism were silent. So the low neuroticism agreement may reflect the observers, not self-ignorance.
- Kindness and rudeness are audible behaviour, so the weak agreeableness agreement likely is real self-ignorance: people do not know when they are being rude or considerate.
- People differ in this self-knowledge (random slopes).
- The transcripts make it concrete. In some hours, people rated themselves as unusually agreeable while complaining bitterly about a sibling's wedding. In another, a person asked a friend to lie for them about the previous night. Some audio shows people talking about the recorder itself ("it's listening to everything we're saying right now").
Limits
- The criterion is imperfect. It is 3 minutes of audio per hour, judged by strangers: a good test for audible behaviour, a poor one for inner states.
- The constructs are narrow: 2 or 3 items per domain, no openness.
- College students over one week; not preregistered.
- Participants knew they were recorded, and some talked about it, so reactivity is possible. They could also delete files, which selects what was observed. In practice little was deleted.
What it means for Amadeus (inference)
- "Say" and "did" disagree in a structured way, as P6 assumes. Where behaviour is visible and not very evaluative (sociability, diligence), momentary self-report tracks recorded behaviour well. Where it is evaluative and interpersonal (kind against rude), people are partly blind to their own behaviour. Inner states (worry) are better reported than observed. So the right source for each field of the slot differs:
- for interpersonal conduct, weight what they did;
- for felt states, weight what they say;
- for how often they are sociable, either works.
- The disagreement itself is personal. Self-knowledge slopes vary between people. How far a person's self-image departs from their recorded conduct is a person-specific variable worth storing, as P6 proposes. It must be measured per domain, not as a single "honesty" score.
- A working design for "did" in daily life. The EAR protocol records passively, samples sparsely (30 s every 9.5 minutes, about 5% of the day), codes with multiple raters, and pairs this with experience sampling. It also lets participants review and delete their own recordings, which is directly reusable for a month of recording, and it was used ethically at scale. A month-long Amadeus protocol could add this channel more cheaply than neural recording. It also captures the third parties a person talks to, which brief 10.2 must cover.