How do researchers figure out how many people in a population have depression, or what factors might trigger the onset of schizophrenia? The answer lies in epidemiology – the science of studying the distribution, frequency, and causes of health conditions in populations. When applied to mental health, epidemiological methods give researchers the tools to measure how common disorders are, who is most at risk, and what factors drive or protect against illness. Understanding these methods is essential for anyone looking at mental health research critically, from reading a study about anxiety rates to evaluating national mental health policy.
Table of Contents
- Measures of disease frequency
- Incidence
- Prevalence
- What the relationship between incidence and prevalence tells us
- Sampling techniques
- Simple random sampling
- Stratified random sampling
- Cluster sampling
- Types of epidemiological studies
- Observational studies
- Experimental and quasi-experimental studies
- Choosing the right design
Measures of disease frequency
Epidemiology is fundamentally about measuring how often disorders occur – and two core metrics underpin almost everything in mental health research: incidence and prevalence.
Incidence
Incidence refers to the number of new cases of a disorder that develop in a defined population during a specific time period. It answers the question: how fast are new cases emerging? Incidence is typically expressed as a rate that counts new cases per background population per unit of time. Researchers calculate it in two main ways:
- Cumulative incidence: the proportion of people who were initially disease-free and developed the condition within a set period (often expressed as cases per 1,000 people).
- Incidence rate: the number of new cases per population at risk measured in person-years, which accounts for the fact that not all participants are observed for the same length of time.
For example, tracking how many new cases of major depression emerge among college students over one academic year gives public health planners a precise signal about whether student mental health is worsening – and whether stress-heavy exam periods correlate with spikes in onset.
Prevalence
Prevalence captures a broader snapshot. Prevalence rates represent the number of existing cases – both old and new – in a defined population during a specified time period. Researchers distinguish between:
- Point prevalence: the number of cases at one specific moment in time.
- Period prevalence: the number of cases occurring within a defined window, such as over 12 months.
- Lifetime prevalence: the proportion of people who have ever had the disorder at any point in their lives. A landmark U.S. study, the Epidemiological Catchment Area Project, found that roughly a third of Americans experience mental illness at some point in their lives – a finding widely cited as a lifetime prevalence estimate.
What the relationship between incidence and prevalence tells us
The interplay between these two metrics reveals a lot about a disorder’s nature. A condition with low incidence but high prevalence – like schizophrenia – signals a chronic, long-lasting disorder. Conversely, high incidence with relatively lower prevalence may indicate a condition that resolves more quickly. Studies of common mental disorders routinely report both point/period prevalence and lifetime prevalence so that readers can distinguish how many people are currently affected from how many have ever been affected. Getting these numbers right matters for policy – measures of prevalence, incidence, remission, and excess mortality are all required to guide health service delivery and intervention strategies.
Sampling techniques
Measuring disease frequency in an entire population is almost never practical. Instead, researchers select a sample – a representative subset – and use it to draw conclusions about the larger group. In probability (random) sampling, every eligible individual has a known chance of being selected, making it possible to generalize findings to the broader population. The most common probability methods used in mental health research are:
Simple random sampling
Simple random sampling selects participants randomly from a complete list of all members of the target population, meaning every individual has an equal chance of inclusion. It is conceptually clean and avoids systematic bias, but it can be logistically difficult for large or geographically dispersed populations. It also requires a complete and up-to-date sampling frame – a list of all potential participants – which is not always available for marginalized or hard-to-reach mental health populations.
Stratified random sampling
Stratified sampling divides the population into non-overlapping subgroups called strata – such as age groups, gender, income levels, or ethnicity – before randomly sampling within each stratum. This method improves the accuracy and representativeness of results by reducing sampling bias, especially when the outcome of interest is expected to differ across subgroups. For instance, a study of depression rates across the lifespan would stratify by age to ensure that adolescents, working-age adults, and older adults are all adequately represented – not just whichever group happens to be easiest to reach.
Cluster sampling
When a population is large and geographically spread out, cluster sampling offers a practical alternative. Cluster sampling divides the entire population into separate groups (clusters), randomly selects a number of those clusters, and then includes all or a random selection of individuals within the chosen clusters. Clusters are often naturally occurring geographic units – cities, districts, schools, or hospitals. This method was notably used in the WHO’s polio eradication initiative, where specific clusters were chosen within countries for intensive surveillance activities – a logic that translates directly to large-scale mental health surveys. Clusters are commonly based on geographic areas or districts, making this approach more common in epidemiological research than in clinical studies.
The choice of sampling method critically shapes whether study findings can be generalized. Individuals identified in clinical settings often represent only the tip of the iceberg of a disorder and may not be representative of the general population of similarly affected individuals with respect to demographic, social, or clinical characteristics – which is precisely why population-based sampling strategies are essential in mental health epidemiology.
Types of epidemiological studies
Beyond measuring frequency and selecting samples, researchers need to choose an appropriate study design. The design determines what questions can be answered – and how confidently. Epidemiological study designs fall into two broad categories: observational studies, where researchers observe without manipulating exposure, and interventional (experimental) studies, where researchers actively introduce an intervention.
Observational studies
Cross-sectional studies collect data from a population at a single point in time. They provide a snapshot of the frequency of a disease or other health-related characteristics in a population at a given point in time, making them well-suited for estimating prevalence. They are relatively quick and inexpensive, which makes them useful for initial investigations and hypothesis generation. Their main limitation is that because exposure and outcome are measured simultaneously, it is not possible to establish which came first – causality cannot be inferred.
Case-control studies take a retrospective approach. They compare groups retrospectively – seeking to identify possible predictors of outcome – and are particularly useful for studying rare diseases or outcomes. Researchers identify a group of people with the disorder (cases) and a comparable group without it (controls), then look backward to compare their histories of exposure to potential risk factors. For example, a case-control study might compare people diagnosed with PTSD against those without PTSD to identify which traumatic experiences are most strongly associated with the disorder. The key disadvantage is the risk of recall bias – cases may scrutinize their memory more carefully and report past exposures more accurately than controls, which can distort findings.
Cohort studies are prospective by design. Researchers begin with a group of people who do not yet have the disorder, measure their exposure to suspected risk factors, and follow them over time to see who develops the condition. Cohort studies are used to study incidence, causes, and prognosis – and because they measure events in chronological order, they can be used to distinguish between cause and effect. A well-known example in mental health is the Dunedin Multidisciplinary Health and Development Study, a longitudinal cohort that has followed children from birth into adulthood to track how early-life exposures shape psychiatric outcomes. The main drawback is cost and time – following large groups for years or decades is resource-intensive.
Experimental and quasi-experimental studies
In experimental studies, researchers actively intervene – assigning participants to conditions rather than simply observing them. The randomized controlled trial (RCT) is the gold standard in experimental study designs: participants are randomly assigned to receive either an intervention (such as a new therapy or medication) or a control condition, and outcomes are compared. Random assignment minimizes the influence of confounding factors, producing the strongest evidence for causality. However, RCTs are not always feasible in mental health research – it would be ethically impossible to randomly assign people to experience trauma or stressful life events to test their causal effects on depression.
Quasi-experimental studies occupy the middle ground. They resemble experiments in that they compare outcomes between groups exposed to different conditions, but they lack full random assignment. A community intervention study that introduces a school-based mental health program in some schools but not others – and then compares anxiety rates – is a quasi-experimental design. It offers more causal leverage than pure observation, but without the control that randomization provides.
Choosing the right design
No single study design is universally superior. The use of study designs depends on the research questions, the purposes of the study, and the availability of resources. Rare disorders are generally better studied through case-control designs; rare exposures are better addressed with cohort designs. When a condition is poorly understood, a cross-sectional or case-control study is usually the logical starting point. When causal inference is the goal and ethics allow it, experimental designs offer the most robust answers. In practice, surveys based on population-based samples and register-based studies have collectively enriched our understanding of how mental disorders impact society – and the field continues to refine its methods as diagnostic criteria evolve and data linkage technologies improve.
What do you think? Considering that RCTs are often impractical in mental health research for ethical reasons, which study design do you believe offers the best balance between real-world applicability and scientific rigor? And given how much prevalence estimates can vary depending on how a sample is selected, how should policymakers account for sampling limitations when planning mental health services?
References
- https://pmc.ncbi.nlm.nih.gov/articles/PMC8504286/
- https://pmc.ncbi.nlm.nih.gov/articles/PMC2807642/
- https://en.wikipedia.org/wiki/Psychiatric_epidemiology
- https://pmc.ncbi.nlm.nih.gov/articles/PMC3997379/
- https://pmc.ncbi.nlm.nih.gov/articles/PMC3691161/
- https://www.healthknowledge.org.uk/public-health-textbook/research-methods/1a-epidemiology/methods-of-sampling-population
- https://link.springer.com/article/10.1007/s40471-022-00287-8
- https://www.simplypsychology.org/cluster-sampling.html
- https://atlasti.com/research-hub/cluster-sampling
- https://pmc.ncbi.nlm.nih.gov/articles/PMC3271469/
- https://www.healthknowledge.org.uk/public-health-textbook/research-methods/1a-epidemiology/cs-as-is
- https://pmc.ncbi.nlm.nih.gov/articles/PMC1726024/
- https://www.gfmer.ch/Books/Reproductive_health/Cohort_and_case_control_studies.htm
- https://www.vaia.com/en-us/explanations/medicine/public-health/epidemiological-design/
- https://minnstate.pressbooks.pub/hgantunez/chapter/__study-designs-commonly-used-in-epidemiology__/
Leave a Reply