A

StatBreak : Identifying “Lucky” Data Points Through Genetic Algorithms

Advances in Methods and Practices in Psychological Science, vol. 3, pp. 216–228

Abstract

Sometimes interesting statistical findings are produced by a small number of “lucky” data points within the tested sample. To address this issue, researchers and reviewers are encouraged to investigate outliers and influential data points. Here, we present StatBreak, an easy-to-apply method, based on a genetic algorithm, that identifies the observations that most strongly contributed to a finding (e.g., effect size, model fit, p value, Bayes factor). Within a given sample, StatBreak searches for the largest subsample in which a previously observed pattern is not present or is reduced below a specifiable threshold. Thus, it answers the following question: “Which (and how few) ‘lucky’ cases would need to be excluded from the sample for the data-based conclusion to change?” StatBreak consists of a simple R function and flags the luckiest data points for any form of statistical analysis. Here, we demonstrate the effectiveness of the method with simulated and real data across a range of study designs and analyses. Additionally, we describe StatBreak’s R function and explain how researchers and reviewers can apply the method to the data they are working with.

Authors 4

  1. Hannes Rosenbusch corresponding

    Tilburg University

    Affiliation as printed

    Department of Social Psychology, Tilburg University

  2. Leiden University

    Affiliation as printed

    Department of Social, Economic and Organisational Psychology, Leiden University

  3. Tilburg University

    Affiliation as printed

    Department of Social Psychology, Tilburg University

  4. Tilburg University · Vrije Universiteit Amsterdam

    Affiliation as printed

    Department of Marketing, VU Amsterdam

    Department of Social Psychology, Tilburg University

Cited by 2 stored of 2

2 results

No patents citing this paper on Lens.org (checked 2026-10-11).

References 34