Kurisutina

Selection overfitting and evaluation of the complete procedure

Gavin C. Cawley and Nicola L. C. Talbot, On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation, Journal of Machine Learning Research 11(70):2079–2107. Primary paper; journal record.

Full reading completed 29 September 2026. This replaces the earlier explicitly partial reading of the abstract, Sections 4.1–4.2 and 5.3, and conclusion. Read all 29 pages: Sections 1–6 and every subsection, equations, all 14 figures and eight tables with captions, footnotes, acknowledgment and complete reference list. All figure/table pages and core model-equation pages were also inspected visually. Retained PDF, text and provenance. The journal landing page and article directory expose only the paper and bibliographic record; no separate supplement was located. The referenced GKM toolbox was not audited or executed. No fits or numerical reproduction were performed, and references were not recursively read as full papers.

Central claim and distinction

The paper separates overfitting a selection criterion, which can worsen the chosen model's actual generalization, from selection bias in evaluation, which can make the reported performance optimistic. A held-out criterion can be unbiased for any fixed candidate yet yield poor hyperparameter choices when its sample variation is optimized. The relevant evaluation unit is the combined training-and-selection procedure, with selection repeated inside each outer evaluation split (Sections 1, 4 and 5).

This is a methodological study using kernel classifiers and small benchmark datasets. It supplies no direct evidence about individual human learning, the usefulness of a person embedding, or whether our current candidate grid overfits.

Methods and complete section coverage

Sections 2–3, pp. 2081–2084: The main base learner is kernel ridge regression for binary labels, with an unpenalized intercept and Gaussian RBF kernel. A regularized least-squares system gives the fitted coefficients; an analytic leave-one-out residual formula makes the PRESS selection criterion inexpensive. The isotropic kernel has regularization and kernel-width hyperparameters. The automatic relevance determination (ARD) variant instead has one width per input feature, adding selection freedom. This is not the same optimizer or loss used by our logistic models.

The synthetic benchmark is a mixture of four bivariate Gaussians with known generating distribution and Bayes error approximately 12.38%. Large-data, true-error tuning shows the model class can get close to that boundary, reducing concern that the later illustration is merely misspecification (Fig. 1). Practical demonstrations use 13 benchmark datasets with 100 prespecified train/test partitions each, except image and splice with 20; Table 1 gives their sizes and dimensions. These are repeated partitions, not 100 independent newly collected datasets.

Section 4, pp. 2084–2094: With 1,000 synthetic realizations of 64 examples, optimizing four-fold cross-validation keeps reducing the selection criterion while expected true test error eventually rises (Fig. 2). A conceptual comparison shows why low criterion variance can matter more for locating a good optimum than unbiasedness alone (Fig. 3). With a fixed 256-example training set and varying validation samples, the mean validation surface resembles the true error surface (Fig. 4), but individual surfaces select underfit through overfit models (Figs. 5–6). Increasing validation size from 64 to 256 tightens the distribution of selected models and true errors (Figs. 7–8). Substantially different hyperparameter values can still sit in the same broad, good-performing valley. Figure 9 illustrates that classifier rankings can reverse across data partitions.

Table 2 compares isotropic RBF and more flexible ARD KRR on all benchmarks. The extra flexibility often improves PRESS while worsening test error; nesting the model families does not guarantee that a finite-data selection procedure finds the better member. Table 3 repeats the warning for Gaussian process classifiers tuned by marginal likelihood: optimizing Bayesian evidence is also a finite-data selection procedure, not automatic protection from this problem. Section 4.4 discusses larger samples, regularization, early stopping, averaging, reducing hyperparameter freedom and integrating hyperparameters rather than maximizing them. These are possible remedies, with limitations; this paper does not establish a universal winner or prove they solve every setting.

Section 5.1, pp. 2095–2096: The reference protocol repeats model selection independently within each resampled training set. The three compared procedures are KRR selected by PRESS, kernel logistic regression selected by approximate leave-one-out likelihood, and expectation-propagation Gaussian process classification selected by marginal likelihood. Table 4 and Fig. 10 show similar average ranks across the 13 tasks, with no clear overall superiority.

Section 5.2, pp. 2096–2101: A commonly used alternative tunes only the first five partitions, takes median hyperparameters and reuses them across all partitions. Tables 5–7 and Figs. 11–13 show why its reported performance is not interchangeable with the internally selected procedure: tuning samples can overlap later test sets, and the median also changes the effective procedure by reducing selection variance. A synthetic experiment with entirely disjoint samples retains an advantage for the median protocol, so overlap alone is not the explanation. The bias differs across classifiers and can favor the procedure with less reliable selection. Figure 14 controls selection variance using repeated 9:1 training/validation splits; the median protocol conceals much of the poor performance from using too few repetitions.

Interpretation matters here: variance reduction is not inherently invalid. If averaging hyperparameters is the intended operational procedure, that entire procedure and its required data must be evaluated honestly. The mismatch arises when its performance is treated as the performance of a different, single-sample selection procedure or compared with methods evaluated under different information budgets.

Section 5.3 and conclusion, pp. 2102–2103: Tuning once on all available data before cross-validation leaves the later test folds involved in hyperparameter choice. Table 8 compares this external tuning with independently repeated internal tuning and finds optimistic external estimates across all 13 benchmarks, including a setting with only two hyperparameters. The paper concludes that selection must remain inside the fitting procedure being evaluated, and recommends multiple datasets and partitions to assess variability. The final four pages contain the references; there is no appendix.

Limits and implications for the present research

The experiments establish important failure modes, not a numerical bias estimate for our task. The demonstrated magnitudes depend on sample size, candidate freedom, criterion, learner and data distribution. Synthetic observations are independent examples; repeated human choices, people, sessions and task chronology introduce structure that the paper does not analyze. Ordinary random row splits would therefore not become suitable for human longitudinal prediction merely because they are nested. The outer split must match whether we intend to predict a new person, a later session of a known person, or later trials within the same session. That is our design implication.

The validation deletion diagnostic measures local composition sensitivity among fixed candidates. Its ten overlapping deletions are not fresh independent validation samples, do not refit the whole procedure and do not estimate out-of-sample learning gains. A change of selected penalty alone is not evidence of important prediction instability: the paper's broad error valley is a concrete counterexample. Likewise, the diagnostic cannot establish selection overfitting without evidence that optimization worsened genuine generalization. M remains conditional on its originally selected S base; changing that base would require a different fitted correction.

The useful conclusion for the research reset is procedural: compare a trained shared adaptive population learner with a person-specific alternative using the same authorized histories and outer test units, and evaluate their full fitting/selection pipelines. This paper does not decide which personal mechanism to test, justify continuing synthetic estimator refinements, or substitute for a direct human comparison. The unresolved empirical question remains whether person-specific experience improves later human choice prediction beyond a strong shared learner, and whether that gain can be attributed to learning rather than stable choice preferences.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.