Kurisutina

CFOs meet LLMs

arXiv 2606.13812, version 2 (the PDF is dated 8 September 2026), CC BY 4.0. Duke University and NBER; Georgia State University. The latest version is said to be on SSRN 6920881 (not read). The data are the confidential individual responses to the Duke–Federal Reserve CFO Survey (not available). Provenance: papers/carry_on/graham2026_cfos_llms.provenance.json.

What was read

All 1,537 lines of pdftotext -layout output (41 pages): text, Tables 1–4, Appendices A–F (data construction, the full system prompt, a sample query, stability, a control for any prior answer, and an open-weight replication with Tables A1–A4) and the references. Figure 1 (four scatter plots) was rendered and read.

Question

Given a firm's public profile and a respondent's own past answers, can an LLM forecast what that individual CFO will answer this quarter? Does it do so beyond the respondent's previous answer?

Method

  • Data. 6,075 responses from public-company CFOs to the Duke–Fed CFO Survey, 2002–2025 (93 quarters). The main item is optimism about the US economy, 0–100. Own-firm optimism and expected revenue growth are repeated as checks.

  • Model. The "gpt-5.4" alias through Duke's Azure gateway, unpinned (the authors say re-runs need not reproduce), run from 21 April 2026. Web search was on, with a prompt instruction to use only information dated up to the survey cutoff. Three calls per response, averaged.

  • Prompt. Role-play as this CFO at this firm on this date. The prompt includes:

    • the firm profile (industry, revenue, employees, foreign share, headquarters, rating);
    • up to 12 of the person's past answers from the last 12 quarters, and their average;
    • the model's own earlier forecasts for the same person, labelled a "secondary prior".

    The system prompt says to treat the person's history as "a strong prior on this individual's style, optimism bias, and level".

  • Analysis. OLS of the CFO's answer on the LLM score, with firm and year-quarter fixed effects and a control for the last-quarter answer. Also:

    • splits by amount of history and by number of profile fields;
    • a prompt ablation (profile, web, history);
    • quarterly aggregation against Michigan sentiment and the SPF GDP forecast;
    • a replication with Gemma 4 31B run locally with no web search.

Results

  • Individual-level fit. Without fixed effects, the coefficient is 0.565 (t = 18.0) and R² = .27. With firm and quarter fixed effects it is 0.276 (t = 6.7), R² = .49.

  • Controlling for the last-quarter answer (n = 1,980): coefficients 0.29–0.53. The last answer itself is 0.217 without fixed effects, 0.037 with firm fixed effects, and 0.251 with quarter fixed effects only.

  • Own history is the main input.

    Prompt (history-eligible sample) R²
    Profile only, web only, or both .09–.10
    History only .33
    History + profile .37
    Full prompt .40
    Past answers available R²
    None .10
    1–3 .33
    More than 3 .49
    • Without history, first-timers' scores sit in a narrow band (Figure 1).
    • The LLM spread across CFOs within a quarter is 0.75 of the human spread (12.3 against 16.5).
  • Other items. For own-firm optimism the score works as well (0.604, R² .24), and the prior answer adds nothing once the score is in. For revenue growth it is weaker and not significant in the smallest subsample.

  • Stability. Three calls correlate at .96 (ICC .958); the mean of three has reliability .985.

  • Quarterly. The mean LLM score predicts mean CFO optimism (R² .72). It dominates Michigan sentiment and the SPF forecast.

  • Gemma 4 31B (no web search) gives nearly the same numbers: 0.572 against 0.565 at baseline, 0.282 against 0.286 in the strictest model.

Limits

  • "Beyond persistence" is tested only against one lagged answer, entered linearly.
    • The model also sees up to 12 past answers, their average and its own earlier forecasts. A plain statistical forecaster from the same inputs is never run: the person's recent average plus this quarter's shift in the population, for example.
    • Firm fixed effects use the full-sample mean, including future answers, which is not the rolling mean the model sees.
    • So the paper does not show that the LLM adds anything beyond a smoothed persistence forecast. The within-person-and-quarter coefficient (0.276) could come from the rolling recent mean differing from the lifetime mean (inferred; testable only with their data).
  • Contamination at the aggregate level. The paper says quarterly survey reports are public (footnote 1). So the aggregate CFO Optimism Index for 2002–2025 was probably in both models' training data (GPT-5.4 cutoff August 2025, Gemma January 2025; inferred).
    • The quarterly regressions (Table 4) and the quarter-level part of the individual fit are therefore not out-of-corpus.
    • Only the within-quarter fixed-effect results escape this. The single COVID-quarter check covers one quarter.
  • Other design problems.
    • Web search runs under a prompt-only date restriction.
    • The model is an unpinned alias.
    • There is no human or statistical-forecaster benchmark.
    • Figure 1's per-quarter R² values rest on few respondents per quarter.
  • Inconsistencies found:
    • "Once fixed effects are added, the prior answer is largely subsumed". This holds with firm fixed effects (0.037, t = 0.51), but with year-quarter fixed effects alone the prior answer stays significant (0.251, t = 2.21), and at 0.171 (t = 1.84) with both.
    • Appendix F says the Gemma sample is "the same set of 6,075". Its fixed-effect and last-answer columns have N = 5,623, 1,982 and 1,820, against 5,631, 1,980 and 1,818 for GPT-5.4.
    • Table 2 Panel C's caption doesn't say that its sample is restricted to history-eligible observations. Its N ≈ 3,832 matches 2,236 + 1,602 = 3,838, and its full-prompt R² (.395) differs from Panel A's (.269) for that reason (my inference).
  • Checked:
    • Panel B's Ns sum to 6,075 both ways: 2,237 + 2,236 + 1,602, and 3,102 + 2,973.
    • 6,075 − 444 singletons = 5,631.
    • Spearman–Brown: 3 × .958 / (1 + 2 × .958) = .986, printed as .985.
    • The spread ratio of means is 12.31 / 16.53 = .745 (the paper reports the mean of quarterly ratios, .75).
    • Text figures match Tables 2–3 and A4, and the ablation R² values .095, .325, .374 and .395.

What it means for Kurisutina

  • A third independent result that the person's own record on the same item carries individual prediction.

    • Here, own history triples the fit over profile, web, or both.
    • In Pohl 2026, the real initial stance is what makes updates work.
    • Jia 2026 withheld same-item history and found little gain from the rest.

    For the replica slot, the person's own answers and states come first and descriptive profiles second (verified across the three papers; the synthesis is inferred).

  • The baseline has to use the same record. This paper, like Pohl, never runs a non-LLM forecaster built from the inputs the LLM sees. The GSS pilot's population table conditioned on the earlier answer, and LISS T2's population tables, are that baseline.

    • Proposal for LISS T2: add a person-level statistical forecaster from the same record: the person's rolling mean on the item, plus the population's shift between the two waves. "Own-full beats the table" can then be separated from "own-full beats a rolling mean of her own answers".
  • Feeding the model its own earlier forecasts is a closed loop. Here human answers always outrank the model's own priors. In a replica that carries on after the person's data stop, the model's own outputs become the only history, which is the self-conditioning drift Pohl calls simulation drift (inferred).

    • Any slot design that stores the replica's own answers as memory must keep them apart from the person's real answers and never let them outrank those answers (proposal, consistent with the immutable content layer from the memory papers).
  • Aggregate over-reaction. The quarterly LLM mean is more volatile than the human mean (SD 9.5 against 6.7), and the authors leave over-reaction untested. For the replica, a reaction to world events larger than the person's is a drift risk worth measuring in Q2 (inferred from Table 1).

  • Contamination lesson for LISS and GSS. Published aggregates of a panel (GSS marginals, LISS reports) may be in a model's training data even when individual responses are not. Pilot contrasts that difference out, such as own against other within a stratum and item, are safer than levels (inferred).

Cross-references

  • summaries/carry_on/pohl2026_belief_updates.md: real initial stance needed, no baseline, simulation drift.
  • summaries/carry_on/jia2026_liss_personas.md: own history against demographics on LISS.
  • summaries/carry_on/santurkar2023_opinionqa.md: modal collapse without individual conditioning.
  • summaries/carry_on/chen2023_chatgpt_behavior_drift.md: unpinned hosted models.
  • summaries/carry_on/lundberg2024_unpredictability.md: private information as irreducible error (the CFO's pending orders here).
  • docs/research/liss_q1_design.md: T2 conditions and baselines.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.