Kurisutina

Predicting choices from strategy profiles

Read 2026-09-22. Primary paper, Learning & Memory 13:230–239. Retained PDF, text, provenance. Full main text, Methods, footnotes, nine figures, four tables, and bibliography read; all graphics and the model equation visually inspected. No original code, individual data, or supplementary analysis obtained. Bibliography entries are not additional completed readings.

Core finding

This follow-up materially improves the 2002 strategy analysis. It compares eleven response profiles, including random responding, using likelihoods; investigates recovery and switching in simulations; and predicts later human choices from earlier choices. The authors explicitly identify the older method's tendency to label random responding as singleton use. Their new analysis reports individual predictive information beyond another person's recent history. It describes response dispositions without uniquely identifying their cognitive or neural implementation.

Reading and method audit

Profile selection. Fourteen cue combinations define each profile. Predictions are .95, .05, or .5 with the adopted consistency setting; the random profile predicts .5 throughout. The criterion requires a better AIC than random, treating selection among strategies as one extra parameter. This changes the score and comparator; the earlier exact null calculation is not a false-positive estimate for this method. A discrete model index and an AIC penalty do not by themselves establish a calibrated significance threshold.

Recovery. Simulations generate responses from the candidate set, with execution errors. Table 2's random column has .79 random, .18 singleton, and .03 single-cue classifications. This is informative about confusion under the simulated setting, not a general reliability estimate for arbitrary human learning processes. Forty-trial recovery improves relative to shorter sequences. Whole-series fits can invent a profile when different profiles occur sequentially. Switch detection is imperfect: reported false switches occur in 19% of simulated no-switch sequences; successful detection is distinct from locating the exact transition. The paper reports this limitation rather than concealing it.

Two temporal procedures must remain separate. Retrospective reconstruction fits 24-trial windows overlapping by twelve trials, requires consecutive consistent windows, and uses observations before and after candidate switches. Methods describes three change measures, accepting proximity of two; the introductory illustration simplifies this. By contrast, response prediction fits the twenty trials preceding the target, then predicts its response. Trials 21–200 are evaluated. The future data used to localize retrospective switches are not stated to enter that separate prediction procedure. Do not infer prediction leakage merely from the switch algorithm.

Prediction comparison. Thirty undergraduates' existing 200-trial sessions are reanalyzed. The same randomized cue/outcome order was used across participants. Own-history predictions beat predictions using the previous response or outcome for the same pattern and the tested recency averages. Predictions from a randomly selected other participant's preceding twenty trials perform above chance, but own history is better, with a paired participant-level comparison. This is useful evidence for individual response information within this task. It does not test a novel cohort, a newly frozen method, unfamiliar combinations, or a broad personal identity model. A randomly chosen other person's estimate is also not the strongest possible pooled or hierarchical baseline.

Reported metric needs care. Footnote 10 defines its measure as C/(C+D), for concordant and discordant pairs with ties omitted, although the text calls it a gamma correlation. Standard gamma would instead be (C−D)/(C+D). The printed definition and implementation cannot be reconciled without code/data. Either way, the number near .75 is a pair-ranking measure and must not be reported as 75% next-trial accuracy. Excluding tied predictions can change the set of pairs that contribute for different predictors. No calibration or proper probability-score comparison is reported. The unexplained t(28), rather than t(29), for the half-session comparison is another reporting boundary.

Amnesia sample and units. Results and figure labels pool fifteen patient task records. Methods specifies seven weather participants, eight ice-cream participants, and six people in both: nine distinct patients, contributing fifteen sessions, as implied by those counts. The overlap of the seven and eight controls is not stated. Treating all pooled task records as independent people is unsupported; the reported analyses do not describe participant clustering. Overlapping windows and regressions on counts across those windows add dependence. Radiological confirmation is stated for five of seven weather participants, not every pooled patient record. The primary patient report, Hopkins et al. 2004, remains a separate unread source.

Two internally inconsistent statistic/probability pairs appear on printed p.235: F(1,19)=12.89 with P>.1, and t(270)=.018 with P=.696. These are visible in the PDF, not extraction errors; neither should be silently repaired. The onset comparison also excludes six patient records and one control without an identified strategy and assigns early strategies a fixed onset of trial six. It therefore does not establish equal acquisition onset for the full groups. More generally, a nonsignificant difference does not show equality.

Mechanism and gradual learning. The paper itself declines to map the profiles onto separate brain systems or explicit/implicit mechanisms. Fitting a profile and explaining accuracy from that same window does not independently establish a causal strategy mechanism. A footnote tests particular gradual probability-matching/maximizing generators against observed reversions; this is evidence against those stipulated generators, not every continuous learner. Shared cue weights could change several pattern predictions together, so a global-looking change need not require a discrete conscious rule switch.

Implementation ambiguities. Table 1 repeats the label “Weak sun”; the second such row has the predictions expected from responding rain when the weak-rain cue is present. Preserve the printed row values before making an explicitly documented label correction. The printed indicator expression uses an unusual index arrangement and calls a discrete indicator a Dirac function; Table 1 supplies the usable prediction contract. No original executable implementation was found through the PDF, retained PMC page, or bounded author/paper search; none was run. No supplementary material is linked in the retained PMC main article.

Our mathematical checks and proposed use

These are deductions, not additional experiments or claims of discovering the paper's acknowledged limitation.

For a fixed candidate profile, let n be the number of observations on which it predicts .95 or .05, and x the number agreeing with its preferred response. On the remaining observations it predicts .5, as does the null. Binomial combinatorial factors cancel between models. Therefore:

log(L_profile/L_random) = x log(1.9) + (n−x) log(.1).

Under the stated one-extra-parameter AIC comparison, this profile beats random only when the log ratio exceeds one. This condition is substantively different from a squared count error below .1. Selecting the best of multiple profiles, repeated windows, tie rules, and switch search all require their own null analysis; this equation does not supply those rates.

A rank metric alone cannot verify transferred probability knowledge. In a constructed two-context population with true response probabilities .6 and .4, forecasts (.6,.4) and (.95,.05) produce identical ordering of every context pair. Their expected Brier scores differ: .24 versus .3625. A constant-.5 baseline has Brier score .25, so the calibrated forecast improves on it while the confident forecast is worse. This is a population counterexample, not an estimate of these participants' calibration.

For an Amadeus-like model, this paper supports a narrow but important acquisition target: learn how the particular person tends to respond, then test future predictions with that person's earlier history. Compare against a pooled model and another person's history, use proper scores as well as chosen actions, and separate retrieval of examples from any inferred rule. Record strategy labels as uncertain summaries of behavior. Require novel-context or intervention tests before treating them as identified internal procedures.

Resulting update to our research position

The 2002 adequacy problem remains valid for that score. The later method explicitly addresses it and adds forward prediction within reanalyzed sessions. We should retain this positive evidence while testing its statistical units, probability calibration, and out-of-task scope. The next useful work is a comparison of specified learning mechanisms under matched past information, rather than another unrestricted claim that strategy fitting either succeeds or fails in general.

Checkpoint 8 follow-up. Speekenbrink et al. 2008 is now fully read in its institutional-manuscript version. It contributes an original human experiment with repeated explicit cue ratings and dynamic model fitting, rather than a reanalysis of these participants. Its whole-series parameter fitting and this paper's preceding-window prediction answer different questions. The new synthesis connects both to sequential validation and the distinction between knowledge and response policy.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.