eLife 14:RP107163, version of record 19 January 2026 (reviewed preprint; earlier title "Individuality transfer:
Predicting human decision-making across tasks", bioRxiv 10.1101/2025.03.25.645375). DOI 10.7554/eLife.107163,
PMC12815462. Sole author, University of Osaka. CC BY 4.0. eLife assessment: "Valuable", "Solid". Code, trained models and
the MDP behavioural data: github.com/hgshrs/indiv_trans (not read during the original paper reading; code audit below).
Provenance: papers/carry_on/higashi2026_individuality.provenance.json.
What was read
All of the full-text XML: article, methods, appendix 1 with both coefficient tables, references, the eLife assessment, the three public reviews with their comments on the revision, and the author response to the first round. Figures are images; captions only. The supplementary MDAR checklist was not read.
Question
Can a representation of an individual, extracted from their behaviour in one task condition, predict that same person's behaviour in a different condition they have not been observed in?
Method (EIDT: encoder, individual latent representation, decoder, task solver)
- Architecture.
- An encoder turns a person's behaviour in the source condition into a vector z, the "individual latent representation". It is shown as 2-D in every figure; its dimension M is never stated numerically.
- A decoder, a single linear layer, maps z to all the weights of a small recurrent network, the task solver (a hypernetwork).
- The task solver predicts the person's choices in the target condition.
- Training uses other people who did both conditions, with the loss on their target behaviour only.
- A new person needs only source-condition data: z is the average of the encoder outputs over their source sessions.
- MDP task (value-guided, sequential).
- Binary choices over 2 or 3 steps; reward 0/1 at the end; transition probabilities (0.8/0.2 or 0.6/0.4) switch at random.
- Participants: 123 recruited on Prolific, £4 plus up to £2 bonus. Each did 3 sequences × 50 episodes of each condition, interleaved in random order.
- Exclusions, by interquartile-range rules, if any block fell outside range: 1 for low reward, 23 for action bias, 18 for response times. That is 42 (34%), leaving 81. The first version had excluded 25 by other criteria.
- Task solver: GRU with 4 cells (2-step) or 8 (3-step). Encoder: GRU with 32 cells plus a 4-layer feed-forward head.
- MNIST task (perceptual).
- Rafiei et al. (2024) data: 60 people naming noisy digits in four conditions (easy or difficult, accuracy or speed emphasis), giving 12 ordered transfer pairs.
- Task solver: a pretrained AlexNet with Bayesian (sampled) weights, shared by everyone and not individualised, feeds evidence over 16 time steps to a 4-cell GRU. Choice likelihood is read at the step given by the response time.
- The author's modelling assumption: individuality lives in the decision stage, not in vision.
- Evaluation.
- Leave-one-participant-out cross-validation, added in revision after reviewer 3 noted none was done. Trial-by-trial negative log-likelihood, plus "rate for behavior matched" (the model's most likely action equals the person's).
- MDP prediction is teacher-forced: "the problem ψ is not executed with the task solver; instead, the task solver predicts action probabilities based on the same task state and reward history as in the human behavioral data". An on-policy simulation was added separately in revision.
Results
- Within condition (no transfer). A task solver trained on 80 others predicted a held-out person better than a Q-learning model with population-average parameters (NLL: F(1,80) = 148.8, ηG² = 0.143). In MNIST the task solver matched human accuracy as RTNet does, and predicted choices better than RTNet (NLL ηG² = 0.73; behaviour-matched ηG² = 0.005).
- Across conditions, MDP (2-step ↔ 3-step).
- EIDT beat a Q-learning model fitted to the person's own source data (learning rate, initial value, discount, inverse temperature) and applied to the target condition. NLL: F(1,80) = 95.7, ηG² = 0.142; matched: ηG² = 0.132.
- Fitted discount rates sat near 1 and initial values near 0 for nearly everyone; learning rate and inverse temperature varied.
- Across conditions, MNIST (12 pairs). EIDT beat a "task solver (source)": the same network architecture trained only on the test person's own source-condition trials. Model effect: NLL ηG² = 0.80; matched ηG² = 0.80. Figure 10 shows four chosen participants whose digit-specific errors the model reproduced (e.g. #23 worse on 1, #56 on 6 and 7).
- Another person's representation predicts worse, the farther away it is. A solver built from person l's z was
used to predict person k. Performance got worse with the latent distance d(k,l) (Gamma GLM with a per-person
intercept).
- MDP, NLL: βd = 0.176 (3→2) and 0.316 (2→3). Behaviour-matched: βd = −0.149 (2→3); the 3→2 coefficient is printed as +0.106 although the text calls it significantly negative.
- MNIST, all 12 directions (p < .001): NLL βd from 0.076 to 0.202; matched from −0.053 to −0.270.
- Simulated Q-learning agents showed the same pattern.
- On-policy (MDP only). Individualised solvers run in the participants' environments. Human–model correlations, with one dot per participant per block: total reward R = 0.667 (3→2) and 0.593 (2→3); rate of choosing the high-reward action R = 0.889 and 0.835.
- What z encodes. Run on 1,000 simulated Q-learning agents, both latent dimensions were fitted by log learning rate, log inverse temperature and their interaction, nearly all coefficients significant. So z carries at least those classic parameters.
Inconsistencies in the published text
- The 3→2 behaviour-matched βd sign (above).
- An interaction reported as F(1,80) = 0.008 with p = 0.012.
- MNIST model effects reported with F(3,177) for a two-level model factor.
- "100 participants were sufficient", whereas the samples are 81 and 60.
- The Q-learning GLM formula writes "qrl".
- MNIST pretraining is credited to Krizhevsky 2012.
None of these changes the direction of the main comparisons, but the exact statistics should not be quoted without the code.
Limits
- Conditions, not tasks. Every transfer stays within one paradigm: 2- vs 3-step of the same MDP, or difficulty and speed variants of one digit task. After reviewers 1–3 all raised this, the claim was renamed from "across tasks" to "across task conditions". The author states transfer "would not work" between tasks that need different cognitive functions.
- Baselines. The eLife assessment itself says superiority "could benefit from further validation … against other
baseline models". Both transfer baselines are weak in a specific way:
- Q-learning with two parameters that barely vary.
- A network trained on one person's source trials alone.
- The obvious control is missing from the transfer comparisons: the population task solver, trained on the other 80 people's target-condition data with no individual information. The paper uses it only in the within-condition comparison against Q-learning. How much the person's own z adds over "an average participant" is therefore never reported; the distance analysis shows only that z carries some individual information.
- The mechanism is borrowing from similar people. The author's own explanation for generalising to new people is that "individuals similar to new participants exist in the training participant pool". The person is represented by their location among others.
- Small, filtered samples. 81 and 60 people. 34% of the MDP sample was excluded, so the modelled people are the behaviourally regular ones. There is no measure of how consistent each person is across conditions, which reviewer 1 asked for and did not get directly.
- Evaluation. Mostly teacher-forced next-choice prediction. The on-policy check uses two block-level summary statistics, not choice sequences, and was not done for MNIST.
What it means for Kurisutina
- This is the Kurisutina architecture in miniature. A shared model plus a person-specific vector that reshapes it (here by generating its weights), estimated from one kind of behaviour and used to predict another. It shows the idea works when source and target are close variants of one task.
- Its own analysis is two-thirds of the drift test. Scoring a person with another person's representation (the
cross-individual analysis) is the "other person's slot" condition. The missing third, the population model with no
individual input, is the "empty slot". It is exactly the comparison needed to show that behaviour is "actually
influenced by the person slot and not purely by drifting towards the base model". The replay harness has the
empty-slot reference (
generic, no evidence) but no other-person condition yet; the full own/other/empty design is the one proposed indocs/research/carry_on_goal.md. This paper shows both contrasts are needed, and that published work can omit the empty one. - "React like Alice, not the average person with her views." EIDT represents a person by where they sit among the training people, and generalises because a neighbour exists. That places Alice among similar people rather than capturing what is only hers. A replica needs evidence that its predictions for Alice beat those of her nearest neighbours in the representation, not only those of distant people. The distance-performance curve gives a direct test: prediction with Alice's own slot should be clearly better than with the slot of her closest other person.
- Consistent with Eckstein et al. 2022. Individual transfer works between similar situations, and nobody has shown it across different ones. For carrying on, this again points to acquiring the person in the domains where the replica must act.
- Closed loop vs open loop. Teacher-forced prediction was strong. The open-loop check was weaker (R = 0.59–0.67 on total reward) and coarse. The same gap is what Namazova et al. (queued) are said to address; the Kurisutina harness should keep reporting both.
Cross-references
- Eckstein et al. 2022 (
summaries/carry_on/eckstein2022_context.md): cited here. Parameters generalise only between similar tasks. - Peng et al. 2026 (
summaries/carry_on/peng2025_funhouse.md): twins closer to a generic persona than to their own person, the failure an empty-slot control detects.
Released-code audit, 29 September 2026
The orientation and audit note records a bounded audit at commit
f13b3cf1ba7e713979bc61d13edf05ccdaa902c2, with a separate artifact manifest. At that code-only stage, no human
outcomes or model weights were loaded and no model was run. The subsequent score analysis is recorded below.
- The released MDP entry points use one fixed train/validation/test split, not the leave-one-participant-out loop described in the paper. Their exclusion code also differs from the revised article. The relation between released checkpoints and the reported analysis is unresolved; this does not establish that the paper's analysis was not done.
- Only two transfer checkpoints appear in the repository inventory. Population-solver checkpoints are absent, but saved per-person losses for both model types are listed. A paired saved-score comparison may be possible after coverage, split and provenance checks.
- Source and target blocks were interleaved, and processed blocks are numbered separately by task. Condition transfer cannot be called chronological prediction without reconstructing and restricting the available history.
- A tiny isolated execution reproduced a defect in the supplied Q-agent simulator: a loop overwrites the sampled action before return. Whether any released data were generated by this version is unverified.
These findings further qualify "shows the idea works" above: this is a promising within-domain candidate for testing incremental personal information. The missing strong comparator and release reconciliation remain substantive gaps.
Saved-score continuation, 29 September 2026
The full report records the analysis specifications, artifact provenance, independent checks and limitations. No models or GPU were used.
- The released processed data have 98 people, with a 68/10/20 train/validation/test split; the revised article reports 81 people and leave-one-person-out evaluation. Behavioral coverage and sequence structure pass.
- The transfer-score table mixes 336 rows with human split IDs and 100 rows with 20 unidentified short IDs. Its writer retains historical rows and keeps the last per key. The original whole-table provenance requirement fails; unique saved keys do not establish a common model run.
- A separately specified post-audit comparison uses all 20 exact human test IDs in both directions. Saved own-transfer minus shared-model mean binary cross-entropy is −0.001436 for 2 → 3 (95% participant interval [−0.018069, +0.016198]) and −0.017746 for 3 → 2 ([−0.036473, −0.000262]). Lower is better; 10/20 and 14/20 people respectively have lower transfer loss. These intervals are conditional on the archived score meanings.
- Own-minus-average-donor gaps are larger: −0.091023 and −0.039095. Thus a strong donor comparison does not settle improvement beyond a shared learner given the same target history. The 3 → 2 comparison is a candidate signal, with its interval endpoint close to zero and its generating-run provenance unverified.
This is a comparison of saved artifacts, not a reproduced held-out personal effect. Missing population weights, unknown score/checkpoint/cache history, different sample selection and unresolved chronology remain. The next bounded check is CPU reproduction of the transfer arm; personal bias/initial-state and nearby-donor controls are still needed before attributing any gain to an individual learning rule.
CPU reproduction and nearest-donor control, 29 September 2026
The checkpoint and donor report records two separately specified follow-ups, complete aggregate results and reproduction commands. No GPU or training was used.
- All 40 own-person scores reproduce from the two released checkpoints on freshly reconstructed sequences; maximum per-person loss error is 1.19e-7 against a tolerance fixed at 1e-5. Both full 20 × 20 donor matrices also reproduce all 40 archived average-other values. This establishes inference correspondence for these transfer results, while the missing population weights and unknown actual training membership remain unresolved.
- Checkpoint loss-history fields differ from the pinned training writer's schema. Their list lengths are 32,300 and 110,300, not verified epoch counts. The upstream loader reads those missing generic history fields after installing both weight dictionaries, so that particular reported loading failure does not imply random models.
- The stronger control chooses the closest other source representation by a fixed raw Euclidean metric, without target-loss selection. Own-minus-nearest loss is −0.012204 for 2 → 3, with 12/20 people benefiting, and +0.000140 for 3 → 2, with 11/20 benefiting. The corresponding average-other gaps are −0.091023 and −0.039095. Thus the average-donor comparison substantially overstates the gap to a close substitute.
- Deleting each person from both recipient and donor pools and reselecting neighbors leaves the 2 → 3 mean negative (−0.013895 to −0.006680). The 3 → 2 mean crosses zero (−0.003842 to +0.002663; seven sign-changing omissions). These are descriptive sensitivity ranges, not confidence intervals. No donor serves more than 3/20 recipients, and there are no exact nearest-distance ties.
The 3 → 2 own-person advantage over average donors is absent at the mean under the nearest-donor control. The 2 → 3 own-versus-nearest gap persists, but the earlier comparison with a shared learner was nearly zero at the mean. These findings do not establish a unique-person gain over both stronger controls. They also do not show that source-person information is useless: the neighbor selection uses it. Coarse response preferences could provide useful matching without requiring an exact individual's learning rule. Next, freeze a simple source action-frequency donor control before calculating its contrasts on the existing matrices. Chronology, training exclusions and action-bias/initial-state alternatives still require separate evidence.
Frequency controls and recorded chronology, 29 September 2026
The frequency/chronology report records two separately specified CPU-only studies. Neither required model inference or GPU use.
- Matching donors by overall source action frequency is substantially worse at the mean than matching by the learned representation: learned minus frequency loss is −0.086675 (2 → 3) and −0.024423 (3 → 2). Per-step frequency matching also trails (−0.082994 / −0.026871). All four comparisons remain negative after every deletion from both recipient and donor pools and neighbor reselection. These descriptive results reject competitiveness of these particular coarse selectors in this bank; they do not establish a dynamic mechanism. Selected donors still execute recurrent policies, and other static summaries remain possible explanations.
- Raw-to-processed alignment passes for all 20 people and 18,000 choices. Every test person has recorded depth order [1, 2, 3, 3, 1, 2, 1, 3, 2]. The three 2 → 3 targets have [1, 1, 2] earlier source blocks; the three 3 → 2 targets have [0, 2, 3]. Therefore the existing three-source-block representation includes later-recorded source behavior for every 2 → 3 target and the first two 3 → 2 targets. These are retrospective cross-condition predictions; only the final 3 → 2 target follows all three source blocks. Direction and order are confounded in this cohort. A chronological analysis must separately define source availability/cold starts.
- The 9,000 missing raw initial-state values reproduce the loader's zero substitution, rather than an independently observed state. Numeric timing fields do not decrease within columns, but this does not authenticate dates.
- Independent arithmetic matches all 6,037 frequency fields; separate source/code reviews find no material frequency or chronology deviation. The expanded standard-library suite has 31 passing tests.
Next: a frozen CPU source-order experiment preserves intact episode contents, first/last episodes and the exact encoder-input row multiset while permuting interior episode order. This probes predictor dependence on order, not an individual's learning rule. A shared-history comparator with verified training membership remains needed.
Interior-episode order diagnostic, 29 September 2026
The order report records the frozen design, numerical controls and all outcomes. Thirty-two permutations per person/direction keep source episodes intact, first/last episodes fixed and the exact encoder-input row multiset unchanged; lagged inputs are rebuilt.
- Permuted minus original mean loss is +0.020796 for 2 → 3 (14/20 increase; delete-one-person range [+0.015123, +0.022896]) and +0.001288 for 3 → 2 (11/20 increase; [−0.002001, +0.002947], one sign flip). These descriptive ranges are not intervals or permutation p-values. People remain the units of analysis.
- The 2 → 3 predictor benefits from recorded interior order beyond the preserved static input counts. The 3 → 2 representation moves under perturbation, but its mean predictive benefit from original order is small and sign-sensitive. Fixed endpoints limit the perturbation; a null/small effect cannot establish order independence.
- All original scores reproduce; whole-block reversal and identical-input batch-size controls pass within 1e-5. Numerical loss effects are at most 2.09e-7. Independent arithmetic and reconstruction verify summaries, all 1,280 permutation/input hashes and 3,840 block invariants. All 16 CPU model-check tests pass; total run 11.45 s.
This is model-level order dependence, not proof of a human learning rule. Recency, generic environment tracking, persistence and distribution shift remain alternatives, alongside the already established chronology limitation. Next is complete-cohort raw chronology before a newly fitted chronological shared-history comparator with verified 68/10/20 membership, static source-count control and then incremental order-dependent source summaries.
The subsequent complete-cohort audit verifies all 98 split members and 88,200 choices; every training, validation and test person has the same recorded order. The third target block is preceded by at least two source blocks for everyone in both directions, so a common 100-terminal-choice source budget is feasible without exclusions. The revised article describes randomized task order, adding another unresolved release-versus-publication difference alongside cohort and split changes. This characterizes the release; it does not show the paper's stated procedure was not used in its own analysis. The new comparator draft has not yet been fitted.
Independent reconstruction confirms the full 98-person audit. Explicit numerical replays reproduce every saved direction field in the checkpoint, donor, frequency and order results exactly, while preserving original files; runtime-dependent parent hashes are handled only through separately declared replay modes. Current suites have 43 standard-library and 18 CPU model tests. The living state collects the evidence, remaining interpretation limits and concrete next work for continued research.
Chronological comparator qualification, 29 September 2026
The new frozen design compares shared target history H, additional block-resolved static source information S, and additional ordered source residuals D. It supplies equal earlier-source budgets and verifies new training membership. This is our exploratory comparator, not a reproduction of the paper's model. No human comparator has yet been fitted.
Its synthetic assay failed the required qualification despite all numerical and data controls passing. Shared-feedback D−S is +0.000521119 for the synthetic 2→3 schedule (only 1/4 replicate improvements), and −0.002494008 for 3→2 (4/4 improvements). Both static/unshared comparison gates pass. Eight unique target cohorts are reused across three source policies. The two synthetic directions do not model actual task depths, so their discrepancy cannot establish human directional asymmetry. All 288 candidates converge; H beats a constant on validation in every run.
The failed assay blocks human fitting under this specification. It does not show absence of personal learning transfer; source summaries may be noisy, own-target history may already supply much of the signal, or this model family/finite sample may not recover the increment reliably. Independent saved-loss, sidecar and generator verification agrees in 76,251 checks. Runtime 37.91 seconds on one CPU thread; no GPU. Current fixtures total 50 standard-library and 30 CPU tests. Next is a separate synthetic failure diagnosis, preserving the frozen design and every failed result before considering a revised comparator.
The subsequent no-refit diagnosis evaluates saved predictions under the simulator's true probabilities along observed histories. Shared-feedback expected D−S is −0.000003185 / −0.001008206: sampling changes the observed scores, but the failed schedule's expected gain stays near zero. Source and prior-target reward-conditioned summaries both correlate strongly with the planted coefficient, so useful marginal measurement does not guarantee additional prediction beyond H/S. Expected D−H is +0.002134039 / −0.000403344. No human or psychological conclusion follows from the simulated schedule difference. Independent reconstruction passes 218,739 checks; runtime 16.09 seconds, no refits/GPU. The next diagnostic partitions the saved D prediction into ordered-source terms and its remaining fitted coefficients, preserving the failed qualification and the prohibition on human fitting under this design.
That coefficient diagnosis is complete: in shared-feedback simulations, full D minus D with its ordered-source coefficients zeroed is −0.001372684 / −0.002956874. The ablated model minus separately fitted S is +0.001369498 / +0.001948668. Helpful source terms are therefore largely offset elsewhere in these fitted coefficients. Seven of eight cohorts benefit from the source terms, while all eight ablated models trail S. This parameterization-dependent decomposition does not identify the cause or authorize a new predictor chosen from test diagnostics. Independent verification passes 541,881 checks; runtime 16.51 seconds with no refits/GPU. Current tests: 50 standard-library and 41 CPU. Next is separately specified fresh-cohort replication of the unchanged pipeline to assess stability before changing the estimator. The original failed qualification remains intact; no human comparator has been fitted.
Fresh-cohort stability replication, 29 September 2026
The unchanged pipeline was replicated on 16 new target cohorts paired across three source cases, kept separate from the original eight. All 48 cases, 576 candidate fits and 16 target/H identity controls complete successfully. Shared-feedback observed D−S is +0.000505107 / −0.000191965 and conditional expected D−S is +0.001410755 / −0.000495841. Expected D−H is −0.001082254 / −0.002592578. The 2→3 means change sign under cohort deletion; observed 3→2 D−H is near zero.
The fitted ordered-source contribution helps in 15/16 new shared-feedback cohorts, with opposing differences in the remaining D coefficients in 14/16. Thus the useful-term/weak-total pattern replicates, but stable net advantage over both S and H is not established. One retained large-loss cohort selected a different D penalty; that association does not identify its cause. These are simulations of a common terminal mechanism, not human task-direction effects or evidence for an intrinsic learning rule.
Fresh execution took 108.12 seconds on one CPU thread, plus 15.87 seconds original-generator compatibility. Independent verification passes 1,162,956 checks; 98 tests pass. No GPU, human fits or qualification upgrade. Next is the modular correction draft: hold the selected S model fixed and fit only ordered-source terms, include a separate zero correction and matched-penalty joint control. This is development on inspected synthetic cohorts and still requires a frozen specification and later independent qualification.
Fixed-baseline correction specification, 29 September 2026
The development design is now frozen before new fits. It fixes selected S coefficients/preprocessing and fits only the original ten ordered-source terms, with all ten penalized and no correction intercept. Four positive penalties plus a separate exact-zero candidate are selected by validation loss, preferring zero on ties. Matched-S-penalty modular/joint controls will separate penalty selection from the estimator comparison. All 48 already inspected cases and all 192 positive correction fits are required; a 180-second CPU cap and 48 pretest snapshots are specified. Implementation and independent verification preparation are underway. This remains development, with no new qualification or human fitting; the outcome has not yet been computed at this entry.
Fixed-S correction results, 29 September 2026
The frozen development experiment completed all 48 inspected cases in 34.03 seconds on one CPU thread. Shared-feedback expected M−S is −0.000194089 / −0.001336086 (2→3 / 3→2). The 2→3 mean changes sign under cohort deletion, whereas all eight 3→2 expected contrasts are negative. Observed M−S is −0.000390499 / −0.000970984. Expected M−H is negative in both means, but observed 3→2 M−H remains deletion-sensitive. These are incremental predictions in invented cohorts, not human task-direction effects or intrinsic learning-rule identification.
At S's selected penalty, modular minus joint expected loss is +0.000275267 / −0.000727570; holding the base fixed is not uniformly beneficial. In the earlier large-loss 2→3 shared case, substituting the prescribed archived joint penalty-0.1 candidate for penalty 1 lowers expected loss by 0.017039118. That one substitution changes the eight-cohort joint D−S mean from +0.001410755 to −0.000719134. It establishes sensitivity of these saved predictions to candidate choice, not a generally superior tuning rule. M's mean advantage over original D in 2→3 also changes sign when that large-loss cohort is omitted.
The exact zero correction is selected in 19/48 cases, including one shared-feedback case, while all four static/unshared case means of expected M−S remain slightly positive. The zero option does not guarantee held-out improvement. All 192 fitted candidates converge; 48 pretest snapshots and the one archived joint candidate reproduce exactly. Independent verification passes 865,486 checks with no failures; all 107 tests pass. No GPU, human fitting, new dependencies or qualification upgrade.
The proposed next step at that stage was a no-refit validation-selection stability diagnostic, which has since been frozen and completed, as recorded below. It omits one of ten validation people at a time and reapplies saved candidate selection rules. M remains conditional on its original fixed S base; changed S choices would require new fits. H contributes 16 unique target cohorts, and the deletions are dependent sensitivity checks. All original outcomes and the failed qualification remain preserved.
Validation-selection sensitivity, 29 September 2026
The frozen saved-loss diagnostic is complete: 160 distinct candidate sets, 688 ten-person loss vectors and 1,600 dependent deletion choices, in 1.20 seconds on CPU. All original selections reproduce before deletions, and all four H candidates match across source cases before H is counted once per target cohort. No fitting or additional test prediction was performed.
The earlier large-loss 2→3 shared D case chooses 0.1 instead of 1 in five of ten deletions. Its full validation gap is +0.000310702 for loss(0.1)−loss(1), while deletion gaps range [−0.003481588,+0.005213646]. This directly establishes sensitivity to validation composition in that realization. Most H/S/D sets are stable. Conditional M switches more often: shared-feedback 15/80 / 25/80 decisions for 2→3 / 3→2, across 5/8 / 7/8 cohorts. Even among S-unchanged decisions it changes in 15/76 / 25/80. M retains its original S base throughout.
Choice sensitivity is not prediction harm: the prior 3→2 expected M−S improvements were negative in all eight cohorts. M also has five candidates including exact zero, versus four for H/S/D, so raw switch rates are not directly comparable risk measures. None of the 160 original choices is a numerical tie. The targeted Cawley–Talbot source note supports distinguishing local sensitivity from evaluation of an entire fitting/selection procedure.
Independent verification passes 55,343 checks with zero failures; root ran the independently prepared verifier unchanged after its agent hit a usage limit. All 58 standard-library tests pass; the prior 57 CPU tests' numerical code remains unchanged. Original results and failed qualification are retained.
Next: the prospective procedure comparison
is frozen before generation, SHA-256 0a0735c7824cb90bf8a8e5ad051511edf0e484044fae9902517d2e27192f74c1.
It specifies 64 fresh target cohorts, unchanged H/S/D/M and the existing variants that reuse S's selected
penalty, with all source cases and strong controls. The user then suspended this narrowing of the research
before implementation or fresh generation. The reset note
returns to person-specific learning beyond an adaptive population model. No GPU or new human fitting.