Kurisutina

What a fitted strategy can establish

Citation. How do People Solve the “Weather Prediction” Task?: Individual Variability in Strategies for Probabilistic Category Learning. Learning & Memory 9(6), 408–418. DOI. Author-institution PDF; PMC full text. This is a primary human behavioral study with post-hoc computational analysis.

Source findings, concise core. Two groups of thirty young adults completed 200 weather-prediction trials with different cue/outcome schedules. Choices were scored against each pattern's modal outcome, independently of the weather realized on that trial. Experiment 1 combined retrospective open questions, multiple-choice reports, and probability estimates. These measures often disagreed with each other and with fitted response profiles. Candidate profiles were motivated partly by those reports. Across all training, singleton profiles fit most people best; shorter-window fits showed changing classifications. Fifty-nine of sixty profiles met the authors' squared-error tolerance. The authors acknowledged alternative strategies, possible switching or guessing, and that a best fit cannot establish the strategy actually used. Their wider interpretation permits both explicit and implicit implementations, although they also interpret reporting mismatches as supporting nondeclarative learning. The study supplies no held-out strategy prediction or independent manipulation establishing unconscious learning. Main paper.

Methods and interpretive audit

Each task uses fourteen recurring combinations of one to three cards. Outcomes provide feedback on every trial. Experiment 2 uses one randomized order fixed across people; its published probabilities are exact only over the full 200 trials. The late-window model may therefore be judged against a long-run ideal that differs from the local evidence available to the participant. The paper does not report sequence-sensitive alternative models or prospective evaluation.

The post-training questions probe different objects. Describing an approach, selecting verbal options, and estimating the frequency when only one card appears are not interchangeable measurements. The latter concerns a singleton conditional probability, not the marginal outcome frequency whenever that card appears with any others. An inaccurate numerical estimate can coexist with useful ordinal knowledge. Reports may compress a changing strategy, whereas a fit over all trials compresses its resulting choices. Neither is privileged ground truth about a latent process. Prompted reports could also create a new explanation; this experiment does not isolate that contribution.

No full quantitative correspondence analysis, coding manual, or independent report-coding reliability is supplied. A narrative mismatch is weaker evidence than a sensitive, matched awareness test and a preregistered direct contrast. This does not make self-reports reliable by default; it limits what their failure to match these profiles establishes.

The Appendix constructs ideal SUN-response probabilities q_i for each pattern and minimizes Σ_i(N_i q_i−K_i)² / Σ_i N_i², where N_i is its presentation count and K_i its observed SUN count. This gives squared-frequency weighting to errors in response proportion and discards within-window temporal order. The threshold is strictly below .1. One data-driven profile variant was collapsed after fitting very few people. There is no independent calibration of the tolerance against a no-learning generator in the reported analysis.

Model labels and subsequent accuracy comparisons reuse the same choices. A profile defined as optimal responding will naturally be associated with high optimal-response scores; that comparison cannot independently validate a causal explanation. A pattern lookup table can match the multicue ideal on every trained pattern. Neither the winning label nor its accuracy identifies the neural system, awareness, or generalization rule.

Some reporting discrepancies remain local limitations: the introductory text's 75% high-predictive one-cue value differs from Table 4's 87.5% optimal responding; the latter can be reconstructed with Table 3 frequencies after setting aside the two outcome-tied patterns. Small mean differences also occur between Results and Discussion. The behavioral-learning finding need not depend on these discrepancies. No raw data or analysis code was available for this audit.

Original exact negative control

This is our calculation, not a reported experiment. Fix the published 200-trial pattern counts and generate independent fair SUN/RAIN choices, with no omissions or learning. Responses are independent of cues, outcomes, and feedback; their total SUN count is not forced to balance. Experiment 1 counts are derived from Table 3 probabilities times 200; no original trial log is claimed. Compare only the predefined singleton profile: q=1 for A/B, q=0 for D/H, and q=.5 otherwise. Then

K_i ~ Binomial(N_i,.5) independently and

E[score] = Σ_i[N_i/4 + N_i²(q_i−.5)²] / Σ_i N_i².

Published count schedule Expected singleton score Exact probability score < .1
Experiment 1, Table 3-derived counts 95/584 = .162671… .0186625031…
Experiment 2 271/3480 = .0778736… .8767576682…

The calculation convolves the complete binomial count distributions with integer multiplicities. It checks total probability mass, the independent variance/bias expectation, and direct enumeration of all 32 sequences in a smaller fixture. The report retains exact fractions and the probability of equality at the boundary; equality is excluded. The result is invariant to trial order under this specific iid null. Independent review verified the source counts, exact calculations, strict boundary handling, and byte-for-byte agreement between executable output and the saved report.

Consequently, in Experiment 2 an entirely random responder meets the stated tolerance for this single profile about 87.7% of the time. Selecting the best of a set containing it can only increase the acceptance probability. This is a calibration failure if the tolerance is interpreted as sufficient evidence of having acquired a particular strategy. It is not a formal p-value from the original study, proof that participants guessed, or a refutation of their separate evidence of learning. It does not reproduce the shorter-block fits. Changed pattern frequencies alter the score's geometry as well as the evidence available for learning; identical numerical thresholds are not equally diagnostic across the two schedules.

For person-model acquisition, require a no-learning control and competing plausible generators before treating a latent label as extracted knowledge. Fit on an initial sequence, preserve its feedback/order, freeze predictions, and test situations where competing models disagree. Evaluate reports and choice histories as complementary evidence with their own measurement limits.

Follow-up boundary. The later Meeter et al. 2006 paper reports simulation validation and subsequent-response prediction in its abstract. Only metadata/abstract have been screened here. Its method must be read before extending this 2002-specific calculation to later strategy analyses; it is a priority continuation, not evidence already adjudicated by the current audit.

Reading record

Full main text, both experiments, Methods, Discussion, and Appendix read; all eight figures, five tables, and the Appendix equation visually checked. Bibliography retained; cited experiments were not all independently followed. No supplement was identified on the inspected sources. No participant data or original code was obtained. The eleven-page PDF and text already retained for the earlier Shohamy methods reading are reused. The new provenance records their hashes and ownership, new PMC/Crossref records, and the distinction between full reading and the original mathematical analysis. Accessed 22 September 2026.

Subsequent reading update — checkpoint7

Meeter et al.2006 is now fully read. It explicitly acknowledges the older random-to-singleton ambiguity, introduces a likelihood/random comparison, and predicts later choices from personal history. The exact null above remains specific to the 2002 score; it does not estimate the later method's error rate. See the follow-up note for positive evidence and remaining calibration/mechanism limits.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.