Every dataset tells a story, but only if you know how to read it. In psychology and statistics, the mean, median, and mode are the three primary tools researchers use to summarize data – yet choosing the wrong one can distort findings, mislead interpretations, and undermine an entire study. The good news? Selecting the right measure isn’t guesswork. It comes down to understanding your data’s level of measurement, its distribution shape, and what your research question actually demands.
Table of Contents
- What are measures of central tendency?
- When to use the mean
- Normal distribution and interval/ratio data
- When not to use the mean
- When to use the median
- Skewed distributions
- Presence of outliers
- Ordinal data
- When to use the mode
- Nominal (categorical) data
- Quick approximation of the most frequent value
- Limitations of the mode
- A practical decision guide
- Why this matters for psychology research
What are measures of central tendency?
Measures of central tendency are summary statistics that identify the central position within a dataset. Rather than presenting every individual data point, they reduce a large collection of numbers into a single representative value. The mean is the arithmetic average, the median is the middle value in an ordered dataset, and the mode is the most frequently occurring value. All three are valid – but they are not interchangeable.
According to Psychology Statistics for Dummies, the most appropriate measure of central tendency depends on two key factors: the level of measurement of the variable (nominal, ordinal, or interval/ratio) and the nature of the distribution of scores – specifically whether outliers or skewness are present. Getting this decision right is foundational to sound data analysis in psychology research.
When to use the mean
The mean is the most widely used measure of central tendency, and for good reason – when conditions are right, it is the most precise and informative summary of a dataset. It incorporates every value in the data, which means it provides the most comprehensive picture of the data as a whole.
Normal distribution and interval/ratio data
The mean works best when your data is measured at the interval or ratio level and follows a roughly normal (bell-shaped) distribution. In a perfectly normal distribution, all three measures of central tendency are identical – the mean, median, and mode all land at the same point. In this scenario, the mean is the preferred choice because it uses all available data and supports further statistical calculations, such as standard deviations, t-tests, and ANOVAs.
For example, if a researcher measures reaction times in milliseconds across a large group of participants and the data is evenly distributed, the mean reaction time is the most stable and representative figure to report. It minimizes the prediction error for any individual score in the dataset, making it the gold standard for normally distributed continuous data.
When not to use the mean
The mean’s key vulnerability is its sensitivity to outliers – extreme values that sit far from the rest of the data. As Laerd Statistics illustrates, consider a small workplace where eight employees earn between ยฃ12,000 and ยฃ18,000 per year, but two senior managers earn ยฃ90,000 and ยฃ95,000. The mean salary comes out at around ยฃ30,700 – a figure that doesn’t accurately represent what any individual employee actually earns. Those two high salaries drag the mean upward, creating a misleading picture of the “average” worker’s pay.
This is why income data – whether in economics or psychological research on socioeconomic status – is almost always reported using the median rather than the mean.
When to use the median
The median is the middle value when data is arranged in ascending or descending order. It divides a dataset exactly in half, meaning 50% of scores fall at or below it. Because it focuses on position rather than value, it is largely immune to the distorting effect of extreme scores.
Skewed distributions
The median is the measure of choice whenever your data is skewed – that is, when the distribution is not symmetrical. In a right-skewed distribution, the mean gets pulled toward the higher end of the scale by extreme values, while the median remains a more accurate reflection of where most of the data actually sits. The more skewed the distribution, the greater the gap between the mean and median, and the stronger the case for reporting the median.
A classic example in psychology research is self-reported income or household wealth. A small number of very high earners pull the mean far above the experience of the typical person in the sample, while the median captures the middle point of the actual distribution far more accurately.
Presence of outliers
Even when data isn’t dramatically skewed, a handful of extreme scores can distort the mean enough to misrepresent the dataset. Consider a memory test score set: 1, 2, 3, 4, 19. The mean would be 5.8 – a number no participant actually scored near – while the median of 3 gives a more accurate reflection of the group’s typical performance, unaffected by that single outlier score of 19.
Ordinal data
The median is also the preferred measure for ordinal data – data where scores can be meaningfully ranked from lowest to highest, but where the intervals between values are not equal. Likert scale responses in psychology surveys are a common example: a response of “4” is not necessarily twice as positive as a “2,” so calculating an arithmetic mean can be misleading. The median preserves the meaningful rank order without making unwarranted assumptions about equal intervals between points.
One important limitation of the median: it does not factor in the precise value of each observation, which means it doesn’t use all the information available in the data. It also cannot be easily used in further algebraic or inferential statistical calculations the way the mean can.
When to use the mode
The mode is the simplest measure of central tendency – it identifies the value that appears most frequently in a dataset. It requires no calculation beyond counting and can be applied to virtually any type of data.
Nominal (categorical) data
The mode is the only appropriate measure of central tendency for nominal data – data that is categorized by label or name without any inherent numerical order. It is the only measure of central tendency that can be used for data measured on a nominal scale. If a researcher asks participants about their preferred therapy type (e.g., CBT, psychodynamic, humanistic), the mode tells you which approach was selected most often. Calculating a mean or median of categorical labels would be statistically meaningless.
Quick approximation of the most frequent value
Even with numerical data, the mode provides a fast, intuitive snapshot of the most common response. In large survey datasets, for instance, knowing the modal response to a question about frequency of anxiety symptoms tells you what experience is most widespread in your sample – without needing to compute anything complex.
Limitations of the mode
The mode has notable weaknesses. A dataset may have no mode (if all values appear with equal frequency), or it may be bimodal or multimodal (with two or more equally frequent values), which complicates interpretation. As Tutor2U notes, the mode is of limited use when a dataset contains many different values appearing with the same frequency. It also completely ignores the majority of scores in the dataset, which makes it a relatively weak descriptor when used in isolation for numerical data.
A practical decision guide
Choosing the right measure becomes straightforward once you apply a few clear criteria. According to Scribbr, the level of measurement of your data is the starting point for this decision.
For nominal data, the mode is the only option. For ordinal data, the median is usually preferable because it respects the rank order without assuming equal intervals. For interval or ratio data, the mean is typically the best choice – unless the data contains significant outliers or is substantially skewed, in which case the median is more appropriate.
It is also worth noting that the three measures work best in combination rather than in isolation, since they have complementary strengths and limitations. Reporting more than one can give a fuller picture – for example, reporting both the mean and median in a dataset with mild skewness allows readers to judge the degree of distributional asymmetry for themselves.
Why this matters for psychology research
In psychological research, the stakes of choosing the wrong measure are real. Misrepresenting the central tendency of a dataset can lead to flawed conclusions about group differences, treatment outcomes, or population trends. A psychologist studying the effectiveness of a new intervention might find that the mean improvement score looks promising – but if a few participants showed dramatic improvement while most showed little change, the median would reveal a more honest picture of the typical patient’s experience.
As published in the Journal of Pharmacology and Pharmacotherapeutics, the relative position of all three measures depends on the shape of the distribution – identical in a normal distribution, but diverging meaningfully as skewness increases. Understanding this relationship helps researchers not only choose the right measure but also interpret what any divergence between them signals about the data’s underlying shape.
In short, no single measure of central tendency is universally superior. The mean is the most powerful when data is clean and normally distributed. The median is the most robust when it isn’t. And the mode is indispensable when working with categories or seeking a quick read of the most common value. Knowing when to reach for each one is what separates rigorous analysis from misleading statistics.
What do you think? When reading a psychology study report, do you pay attention to whether the researchers used the mean or the median – and do you think that choice affected the conclusions they drew? If you were designing a survey to measure anxiety levels on a 5-point Likert scale, which measure of central tendency would you choose to summarize your results, and why?
References
- https://statistics.laerd.com/statistical-guides/measures-central-tendency-mean-mode-median.php
- https://www.dummies.com/article/body-mind-spirit/emotional-health-psychology/psychology/research/choosing-between-mode-median-and-mean-in-psychology-statistics-169545/
- https://pmc.ncbi.nlm.nih.gov/articles/PMC3157145/
- https://www.tutor2u.net/psychology/reference/measures-of-central-tendency
- https://www.scribbr.com/statistics/central-tendency/
Leave a Reply