Statistics have the power to shape how we understand human behavior, inform public policy, and guide clinical decisions. But that same power cuts both ways. When used responsibly, statistics illuminate patterns and help us make sense of a complex world. When misused – whether through carelessness or deliberate manipulation – they can mislead, distort, and erode the public’s trust in science. Understanding why statistics are sometimes distrusted, and how they get misused, is essential for anyone engaging seriously with research in psychology or any other field.
Table of Contents
- Why people distrust statistics
- Common ways statistics get misused
- Correlation vs. causation
- Overgeneralization
- Statistical vs. practical significance
- Data manipulation and selective reporting
- P-hacking and the replication crisis in psychology
- Errors in published psychological research
- The variable-individual confusion in psychology
- The path forward: transparency, ethics, and statistical literacy
Why people distrust statistics
The famous phrase “lies, damned lies, and statistics” has endured for good reason. As Psychology Today notes, those who use or communicate statistics – including scientists, politicians, and journalists – can be biased or even deliberately misleading. This has naturally generated public skepticism. High-profile cases of scientific misconduct, combined with a steady stream of research findings that later fail to hold up, have made many people question whether statistical conclusions can be trusted at face value.
But there is an important distinction to make: statistics themselves are not biased. Numbers don’t lie on their own. The distrust arises from how statistics are collected, analyzed, reported, and communicated – and from the very human motivations of the people doing those things. Recognizing this distinction is the first step toward becoming a more informed consumer of data.
Common ways statistics get misused
Misuse of statistics isn’t always intentional. According to Wikipedia’s overview of statistical misuse, even professional scientists and trained statisticians can be fooled by simple but misleading methods. The consequences, however, can be serious – from distorted public understanding to harmful policy decisions and, in medical research, errors that may take decades to correct.
Correlation vs. causation
One of the most persistent errors in interpreting statistics is confusing correlation with causation. Just because two variables move together doesn’t mean one is causing the other. Misleading conclusions of this kind arise frequently when data is reported through mass media, where the nuance required to properly interpret a correlation is often stripped away in favor of a simpler, more dramatic story.
Overgeneralization
Overgeneralization occurs when a finding from a specific population is incorrectly applied to a much broader group. Research on statistical misuse notes that this often happens when results pass through nontechnical sources – and sometimes the overgeneralization is introduced not by the original researcher, but by others who later reinterpret findings to suit a different purpose. A study on college students in one country, for instance, cannot automatically be taken to describe human behavior universally.
Statistical vs. practical significance
A result can be statistically significant without being practically meaningful. Statistical significance tells you the probability that a finding occurred by chance; it says nothing about how large or important the effect actually is in the real world. Relying solely on statistical significance can create serious misconceptions about the relevance of a finding – a problem especially common when research is summarized for non-specialist audiences.
Data manipulation and selective reporting
Data manipulation – sometimes called “fudging the data” – involves choosing datasets that align with a researcher’s preferred hypothesis while ignoring those that contradict it. This is a serious form of statistical misuse committed by both experienced and inexperienced researchers. Even subtler forms, like analyzing too few variables in a multidimensional dataset, can produce misleading results and leave findings vulnerable to statistical paradoxes.
P-hacking and the replication crisis in psychology
Perhaps no issue has done more to shake confidence in psychological research than the combination of p-hacking and the broader replication crisis. P-hacking refers to the practice of manipulating data analysis – running multiple tests, removing outliers, or adjusting variables – until a p-value below 0.05 is achieved, the conventional threshold for “statistical significance.” As the SAGE Research Methods Community explains, statisticians have a saying for this: “if you torture the data enough, they will confess.”
The scale of the problem came into sharp focus in 2015, when the Open Science Collaboration conducted the landmark Reproducibility Project. Researchers attempted 100 direct replications of published psychology studies and found that while 97% of the originals reported statistically significant results, only 36% could be successfully replicated. This gap between what gets published and what can be independently confirmed struck at the credibility of large swaths of the psychological literature.
Closely related to p-hacking is HARKing – Hypothesizing After Results are Known. This involves presenting a hypothesis in a paper as if it had been formulated before data collection, when in reality it was generated after seeing the results. Like p-hacking, HARKing inflates the risk of false positives and makes replication far less likely.
A contributing factor to both problems is the “publish or perish” culture in academia. The pressure on researchers to produce significant findings pushes some toward questionable research practices as a way to achieve a publishable p-value. Journals, for their part, have historically favored positive results – creating what researchers call a “file drawer problem,” where null results are never published and effectively disappear from the scientific record.
Errors in published psychological research
Misuse of statistics isn’t only about deliberate manipulation. Honest errors are also widespread. A study examining statistical reporting in psychology journals found that approximately 18% of published statistical results were incorrectly reported, and around 15% of articles contained at least one statistical conclusion that changed from significant to non-significant – or vice versa – upon recalculation. Critically, these errors tended to align with what researchers hoped to find, suggesting that even unconscious bias shapes which mistakes go unnoticed.
Research reviewing clinical neuropsychology publications specifically identified four recurring errors: inappropriate use of null hypothesis tests, misuse of p-values, neglect of effect size, and inflation of Type I error rates. These are not obscure technical mistakes – they are fundamental to how findings are interpreted and presented to the scientific community.
The variable-individual confusion in psychology
A deeper, more conceptual form of statistical misuse in psychology involves what some scholars call the variable-individual confusion. Critics argue that mainstream psychological research routinely draws conclusions about individuals from statistical patterns observed in populations – a logical leap that the data simply cannot support. Population-level findings describe averages and trends across groups; they do not necessarily tell us anything meaningful about any specific person within that group. When clinicians or policymakers treat group-level statistics as if they apply uniformly to individuals, the result can be poor decisions and even harm.
The path forward: transparency, ethics, and statistical literacy
Recognizing the problem is the first step. Fixing it requires systemic change – and that change is underway, albeit slowly. A set of reforms recommended in Nature Human Behaviour includes visualizing data, quantifying uncertainty, reporting multiple models, involving multiple analysts, and sharing raw data and code. Each of these practices makes it harder to manipulate findings and easier for others to verify them.
Preregistration has emerged as one of the most promising structural reforms. It requires researchers to specify their hypotheses and planned analyses before collecting data, removing much of the flexibility that enables p-hacking. When preregistration is mandated, the number of null results reported in journals rises significantly – a sign that the file drawer problem is being addressed.
The broader Open Science movement pushes for transparency across the entire research process. This includes making data and materials publicly available, pre-registering analysis plans, and publishing replication attempts and null results – a set of practices collectively aimed at making scientific knowledge more reliable and accessible. With open methods and data, institutions and funders can also more easily detect research misconduct and build genuine public trust in science.
Finally, statistical literacy matters – not just for researchers, but for everyone who encounters data-driven claims in daily life. There are concrete ways to train people to not fall for the misuses of statistics, and fostering this kind of critical thinking is increasingly recognized as a public good. Learning to ask the right questions – Who collected this data? What was the sample? Is this finding practically meaningful? Has it been replicated? – goes a long way toward separating genuine insight from statistical sleight of hand.
Statistics, at their core, are a tool for understanding reality more clearly. When that tool is used with rigor, transparency, and ethical care, it has an extraordinary capacity to illuminate. But when it is applied carelessly or manipulated for convenience, it becomes a source of confusion and distrust. The goal isn’t to be skeptical of all statistics – it’s to be thoughtfully skeptical: asking harder questions, demanding transparency, and holding both researchers and ourselves to a higher standard of statistical honesty.
What do you think? When you come across a striking statistical claim in the news or on social media, what criteria do you use to decide whether to trust it? And do you think the responsibility for preventing statistical misuse falls more on researchers themselves, or on the institutions and journals that publish their work?
References
- https://www.psychologytoday.com/us/blog/bias-fundamentals/202311/can-statistics-ever-be-biased
- https://en.wikipedia.org/wiki/Misuse_of_statistics
- https://www.mastermindbehavior.com/post/lying-statistics-and-facts
- https://www.researchgate.net/publication/362648029_Misuses_of_Statistics_in_Research
- https://researchmethodscommunity.sagepub.com/blog/primer-p-hacking
- https://michael-franke.github.io/intro-data-analysis/app-94-replication-crisis.html
- https://simplifaster.com/articles/p-hacking-harking-scientific-replication/
- https://pmc.ncbi.nlm.nih.gov/articles/PMC3174372/
- https://www.sciencedirect.com/science/article/pii/S0887617705001071
- https://pmc.ncbi.nlm.nih.gov/articles/PMC12239868/
- https://www.nature.com/articles/s41562-021-01211-8
- https://royalsocietypublishing.org/doi/10.1098/rsos.240313
- https://pmc.ncbi.nlm.nih.gov/articles/PMC9283153/
- https://www.thelancet.com/journals/lancet/article/PIIS0140-6736(23)01575-1/fulltext
Leave a Reply