Statistics can be seductive. When two variables move together in a predictable pattern, it’s tempting to conclude that one must be driving the other. But in psychological research – and in everyday reasoning – this leap from correlation to causation is one of the most common and consequential errors a person can make. Understanding exactly why these two concepts are distinct, and what it actually takes to establish a causal claim, is foundational to sound scientific thinking.

Table of Contents

What correlation actually tells us

A correlation is a statistical indicator of the relationship between variables. When two variables are correlated, they covary – as one changes, the other tends to change in a predictable direction. This can be positive (both variables rise together), negative (one rises as the other falls), or absent (no consistent pattern). The strength of that relationship is captured by a correlation coefficient, ranging from -1 to +1.

What correlation does not tell us is why that relationship exists. As one explanation puts it, you may have discovered that a correlation exists between two variables, but you still don’t know why the two things are correlated. That distinction is everything. Knowing that two variables move together allows researchers to make predictions – but prediction is not the same as explanation, and explanation is not the same as causation.

Causation: a much higher bar

Causation means that a change in one variable directly produces a change in another. Establishing cause and effect requires three criteria: association (the variables must be statistically related), temporal precedence (the cause must come before the effect), and non-spuriousness (the relationship cannot be better explained by a third variable). Meeting all three is harder than it sounds – and in psychology, it is often extremely difficult.

The classical formulation from John Stuart Mill, still widely cited in research methodology, lays out these same requirements: the cause must precede the effect, the two must covary, and all plausible alternative explanations must be ruled out. The third criterion is particularly challenging. There will almost always be unmeasured variables lurking in the background, and ruling them all out requires careful study design, not just statistical analysis.

Theory must come first

One often-overlooked requirement is that any causal claim needs a theoretical basis before statistical analysis begins. According to the SAGE Encyclopedia of Communication Research Methods, if a researcher doesn’t specify in advance why variable X should cause variable Y, then finding a strong correlation between them may simply be capitalizing on chance. The practice of trawling through data for any significant associations – known as data mining – can produce correlations that appear meaningful but are statistically coincidental. A solid theoretical rationale is the anchor that prevents this kind of reverse-engineered reasoning.

Why correlation is not causation: two core problems

The third variable problem

The most well-known reason that correlation fails to establish causation is the third variable problem. A confounding variable is one that influences both the variables being studied, making them appear causally related when they are not. A classic example: ice cream sales and drowning deaths are positively correlated, but buying ice cream does not cause drowning. Both rise together because a third factor – hot weather – independently increases both ice cream consumption and the amount of time people spend swimming.

This same logic applies in psychological research. Vitamin D levels and depression are correlated, but it’s unclear whether low vitamin D contributes to depression or whether depression leads to reduced vitamin D intake – or whether some other variable, such as reduced outdoor activity or physical illness, drives both. The correlation alone cannot resolve this ambiguity.

A confounding variable is an unmeasured third factor that influences both the independent and dependent variables, creating the appearance of a relationship that may not be real. When confounders are not controlled for, researchers risk making what is called a spurious association – a false conclusion that two things are causally connected when they are not.

The directionality problem

Even when two variables genuinely influence each other, correlation cannot tell us which direction the causal arrow points. Consider the relationship between drug use and psychiatric disorders: do drugs cause the disorders, or do people with pre-existing conditions use drugs to self-medicate? The data may show a strong association either way, but the correlation alone provides no answer. This is known as the directionality problem or reverse causation – and it is a persistent challenge in psychological and social science research.

Spurious correlations in the wild

Spurious correlations are not just a theoretical concern. They appear regularly in research and media reporting. Studies on cereal consumption and healthy body weight, for example, have been reported as though they demonstrate that eating cereal causes better weight management. But a more cautious reading suggests the opposite may be true: people who already maintain healthy habits are more likely to eat breakfast regularly. The correlation may reflect lifestyle differences rather than any effect of cereal itself.

The problem is compounded by the fact that humans are evolutionarily predisposed to see patterns and psychologically inclined to gather information that supports their existing views – a trait known as confirmation bias. This makes it easy to mistake a coincidence for a connection, and a connection for a cause.

Researchers have also documented how confounders can create the illusion of a direct relationship between two completely unrelated variables. When a hidden variable drives changes in both X and Y simultaneously, X and Y will appear to move together – even if neither has any influence on the other. Without accounting for that hidden variable, the correlation looks real and potentially meaningful.

How researchers approach causal inference

The gold standard for establishing causation in psychology is the randomized controlled experiment. By randomly assigning participants to conditions, researchers can ensure that individual differences are distributed evenly across groups, making it far more likely that any observed difference in outcomes is due to the manipulation itself rather than some extraneous factor. Random assignment is one of the most effective tools for achieving internal validity – the confidence that the independent variable, and not something else, caused the change in the dependent variable.

However, experiments are not always possible in psychology. Many variables of interest – personality traits, early childhood experiences, socioeconomic status – cannot be ethically or practically manipulated. In such cases, researchers rely on observational or correlational data and must be especially careful about causal language. It is impossible to infer causation from correlation without background knowledge about the domain. Statistical controls, longitudinal designs, and theoretical grounding can all strengthen causal arguments – but they do not substitute for experimental evidence.

When experiments are not feasible, methods such as multivariate regression analysis can help researchers statistically control for known confounders. But this approach has limits: it can only account for variables that have been measured and included in the model. Unmeasured confounders remain a persistent threat to causal inference in observational research.

The real-world cost of confusing correlation with causation

The stakes of this distinction are not merely academic. The tobacco industry historically relied on dismissing correlational evidence to deny the link between smoking and lung cancer – a strategy that delayed public health interventions for years. More broadly, when researchers, journalists, or policymakers treat correlational findings as causal, the result can be flawed interventions, misallocated resources, and public misconceptions that persist long after the original claim is corrected.

This is why psychological researchers are trained to describe their findings with precision. A study that finds a correlation between two variables should report an association, not a cause. The language matters because it shapes how findings are interpreted, reported, and acted upon. Correlational research remains crucial – it identifies relationships worth investigating, generates hypotheses, and is often the only option available. But its conclusions must be framed accordingly.

Interpreting correlational findings responsibly

Recognizing the limits of correlation does not diminish its value. Correlational findings are the foundation of much of what psychology knows about human behavior, mental health, personality, and development. The goal is not to dismiss them, but to interpret them cautiously – asking, before concluding that X causes Y, whether a third variable might explain the pattern, whether the causal direction has been established, and whether there is a theoretical framework that supports the claim prior to analysis.

Contemporary causal research requires strong assumptions, subject-matter knowledge, careful statistical analysis, and consideration of alternative explanations. Researchers who approach their data with this level of rigor – rather than jumping from a significant correlation coefficient to a causal conclusion – produce findings that are far more reliable, replicable, and useful.

What do you think? When you read a news headline claiming that a behavior or habit “causes” a health or psychological outcome, what questions would you now ask before accepting that claim at face value? And in fields like psychology, where experiments are often impractical, how confident should we ever be in causal conclusions drawn from observational data?

How useful was this post?

Click on a star to rate it!

Average rating 5 / 5. Vote count: 1

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://www.scribbr.com/methodology/correlation-vs-causation/
  2. https://sites.monroecc.edu/mofsowitz/psychology/correlationcausation/
  3. https://www.statisticssolutions.com/dissertation-resources/research-designs/establishing-cause-and-effect/
  4. https://pmc.ncbi.nlm.nih.gov/articles/PMC2957016/
  5. https://methods.sagepub.com/ency/edvol/the-sage-encyclopedia-of-communication-research-methods/chpt/causality
  6. https://www.simplypsychology.org/confounding-variable.html
  7. https://en.wikipedia.org/wiki/Correlation_does_not_imply_causation
  8. https://pressbooks.openeducationalberta.ca/saitintropsychology/chapter/correlation-vs-causation/
  9. https://science.howstuffworks.com/innovation/science-questions/10-correlations-that-are-not-causations.htm
  10. https://statisticsbyjim.com/basics/spurious-correlation/
  11. https://www.biosourcesoftware.com/post/best-practice-interrogating-causal-claims
  12. https://journals.sagepub.com/doi/10.1177/2515245917745629
  13. https://www.psychologytoday.com/us/blog/all-about-addiction/201003/correlation-causation-and-association-what-does-it-all-mean
  14. https://www.ncbi.nlm.nih.gov/books/NBK606119/

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Statistics in Psychology

1 Introduction to Statistics

  1. Meaning of Statistics
  2. Types of Statistics
  3. Scope and Use of Statistics
  4. Limitations of Statistics
  5. Distrust and Misuse of Statistics

2 Descriptive Statistics

  1. Organising Data
  2. Summarising Data
  3. Use of Descriptive Statistics

3 Inferential Statistics

  1. Concept and Meaning of Inferential Statistics
  2. Inferential Procedures
  3. Hypothesis Testing
  4. General Procedure for Testing Hypothesis

4 Frequency Distribution and Graphical Presentation

  1. Arrangement of Data
  2. Tabulation of Data
  3. Graphical Presentation of Data
  4. Diagrammatic Presentation of Data

5 Concept of Central Tendency

  1. Meaning of Measures of Central Tendency
  2. Functions of Measures of Central Tendency
  3. Types of Measures of Central Tendency
  4. Characteristics of a Good Measures of Central Tendency

6 Mean, Median and Mode

  1. Symbols Used in Calculation of Measures of Central Tendency
  2. The Arithmetic Mean
  3. The Median
  4. The Mode
  5. When to Use the Various Measures of Central Tendency

7 Concept of Dispersion

  1. Concept of Dispersion
  2. Functions of Dispersion
  3. Measures of Dispersion
  4. Significance of Measures of Dispersion
  5. Types of Measures of Variability/Dispersion

8 Range, MD, SD and QD

  1. Range
  2. Quartile Deviation
  3. The Average Deviation
  4. The Standard Deviation
  5. When to Use Different Measures of Dispersion

9 Introduction to Parametric Correlation

  1. Introduction to Correlation
  2. Scatter Diagram
  3. Correlation: Linear and Non-Linear Relationship
  4. Direction of Correlation: Positive and Negative
  5. Correlation: The Strength of Relationship
  6. Measurements of Correlation
  7. Correlation and Causality
  8. Uses of Correlation

10 Product Moment Coefficient of Correlation

  1. Building Blocks of Correlation
  2. Pearsonโ€™s Product Moment Coefficient of Correlation
  3. Interpretation of Correlation
  4. Using Raw Score Method for Calculating r
  5. Significance Testing of r
  6. Other Types of Pearsonโ€™s Correlation

11 Introduction to Non-Parametric Correlation

  1. Parameter Estimation
  2. Parametric and Non-parametric Statistics
  3. Scales of Measurement
  4. Conditions for Rank Order Correlations
  5. Ranking of the Data
  6. Rank Correlations

12 Rank Correlation (rho and Kendall Rank Correlation

  1. Rank-Order Correlations
  2. Spearmanโ€™s rho (rs)
  3. Kendallโ€™s tau (ฯ„)

13 Significance of the Difference of Frequency- Chi-Square

  1. Parametric and Non-Parametric Statistics Tests
  2. Chi-square Test: Definitions
  3. Assumptions for the Application of x2 Test
  4. Properties of the Chi-square Distribution
  5. Application of Chi-square Test
  6. Precautions about Using the Chi-square Test

14 Concept and Calculation of Chi-Square

  1. Application of Chi-square Test
  2. The Chi-square Test when Table Entries are Small (Yateโ€™s Correction)
  3. Chi-square as a Test of Independence
  4. 2 ร— 2 Fold Contingency Tables

15 Significance of the Differences between Means (T-value)

  1. Need and Importance of the Significance of the Difference between Means
  2. Fundamental Concepts in Determining the Significance of the Difference between Means
  3. Methods to Test the Significance of Difference between the Means of Two Independent Groups (t-test)
  4. Significance of the Difference Between two Correlated Means

16 Normal Distribution- Definition, Characteristics and Properties

  1. Definitions of Probability
  2. The Normal Distribution
  3. Deviation from the Normality
  4. Characteristics of a Normal Curve
  5. Properties of the Normal Distribution
  6. Application of the Normal Curve