Social Psychological and Personality Science 8(4):355–362, doi 10.1177/1948550617697177; open access (CC BY 4.0), PMC5502906. Read by researcher R4d (the base battery), 5 October 2026. Provenance: papers/base/lakens2017_equivalence_tests.provenance.json.
What was read
- Read in full: the Europe PMC full-text XML, converted to text: abstract, main text, Figure 1 caption, Table 1, equations 1–10 (as linearised text), discussion, notes, author's note and the reference list.
- Not read: the TOSTER spreadsheet and R package (osf.io/q253c; CRAN), and the figure as an image.
What it says
- The problem. A non-significant test does not show that an effect is absent; with small samples it says little, and with huge samples trivial effects become significant. "No effect" appeared in 108 articles of this journal up to August 2016, almost all resting on non-significance. About 38% of articles with non-significant results in the Journal of Applied Psychology accept the null (Finch et al. 2001, cited).
- The method: two one-sided tests (TOST). Specify a lower and an upper equivalence bound (−ΔL, ΔU) from the smallest effect size of interest (SESOI). Test H01: Δ ≤ −ΔL and H02: Δ ≥ ΔU. If both are rejected, the effect lies inside the bounds and is "practically equivalent". Bounds may be symmetric or asymmetric (e.g. −.2 to .4), in raw units or standardised (Cohen's d). Setting one bound to infinity gives a one-sided inferiority test.
- Decision by interval. Equivalence holds when the 90% CI (1 − 2α, since two one-sided tests at α = .05 are run) lies inside the bounds. Combined with an ordinary test there are four outcomes: equivalent and not different from zero; different but not equivalent; different and equivalent (a small but non-zero effect); undetermined.
- Formulas given for independent means (Student and Welch, Satterthwaite df), dependent means (Cohen's dz), one-sample tests, correlations (Fisher z, after Goertzen and Cribbie 2010) and meta-analytic effects (Rogers et al. 1993: the meta-analytic estimate and its SE, or its 90% CI).
- Power. Narrow bounds need large samples. Per group, for 80% power at α = .05 (exact): bound d = 0.1 → 1,713; 0.2 → 429; 0.3 → 191; 0.4 → 108; 0.5 → 70; 0.8 → 28. For 90% power at α = .05 (exact): 0.1 → 2,165; 0.2 → 542; 0.3 → 242; 0.5 → 88 (Table 1). For dependent designs, the number of pairs is half the per-group number when expressed in dz. Equivalence tests need slightly larger samples than ordinary tests.
- Worked example. Moery and Calin-Jageman (2016) replicated Eskine (2013, d = .81). Bound set by Simonsohn's rule (the effect the original had 33% power to detect: d = .48, i.e. .384 points on a 7-point scale). Data: control M 5.25 (SD .95, n 95), organic M 5.22 (SD .83, n 89); tL = 3.14, tU = −2.69, p = .001 and .004: effects larger than .384 scale points are rejected.
How to set the bounds (the paper's recommendations)
- Best: from theory or a cost–benefit argument, in raw units, specified when the effect is first published.
- Otherwise, from resources: the smallest effect the available sample can reliably test. With 100 per group, 80% power and α = .05, bounds of ±0.414 (against a SESOI of 0.389 for an ordinary test).
- For replications: Simonsohn's "small telescopes" — the upper bound is the effect the original study had 33% power to detect (needs about 2.5 times the original sample).
- Last resort: conventional benchmarks (d .2/.5/.8; r .1/.3/.5; for dz either the same as d or scaled by √2 to .14/.35/.57).
- Pre-specify the bounds (as CONSORT expects of trials) and state them wherever "statistically equivalent" is claimed. The author rejects bounds as narrow as ±.1 or ±.05 for replications (Maxwell et al. 2015) as an imbalance with lax original studies.
- Contrast with drug development, where bounds are often set by regulation ("differences up to 20% are not considered to be clinically relevant"); the author thinks such general rules unlikely, and perhaps undesirable, in psychology.
Cautions stated
- Equivalence is shown only for one operationalisation; another manipulation or measure might show an effect. Confounds can produce equivalence.
- Equivalence can come from an incompetently run experiment; the asymmetry with showing an effect is resolved only by independent replication.
- Error rates equal α only when the true effect sits at a bound; small samples may have no power at all to show equivalence.
- Standardised bounds can give different verdicts for identical raw results when SDs differ; use raw bounds if that matters.
- Alternatives: estimation with CIs, a region of practical equivalence (Kruschke), Bayes factors (Dienes; Rouder et al.).
Limits of this reading
- The printed rejection rules for the correlation and meta-analysis tests have sign conventions that do not match the t-test rule given earlier (as converted from the XML: "rejected if ZL ≤ −Zα and ZU ≥ Zα" for correlations; "ZL ≤ −Zα and ZU ≤ Zα" for meta-analyses). The 90% CI criterion is unambiguous and is what should be used. One cross-reference ("Equation 3 can be used for a priori power analyses") appears to mean Equation 5.
- A methods primer: no new data; the software was not inspected.
What it means for the base battery (inference)
- Every pass rule can be an equivalence test against the human reference: the base passes a test when the 90% CI of (base − human) lies inside pre-declared bounds, not when a difference test fails to reach significance. A test the base fails to resolve (CI crosses a bound) is "undetermined", which is not a pass.
- The bounds must be declared before any base exists, in raw units of each measure where possible, and justified by (a) the human spread or between-study heterogeneity of the reference effect, or (b) the precision of the human reference itself (a small-telescopes logic: a base effect the human study could not have told apart from its own estimate is equivalent).
- "Different and equivalent" is a legitimate pass: the base may differ from the human mean by a statistically detectable but inconsequential amount.
- Power sets the run cost: a bound of d = 0.3 needs about 191 simulated participants per arm, d = 0.2 about 429. Because the base can be run on as many simulated people as needed, the human reference's own uncertainty, not the base sample, will usually limit how narrow a bound can be.