Not all research is created equal. Two studies can investigate the same question and arrive at completely different conclusions – and the difference often comes down to the quality of the research design itself. In psychology, a well-crafted research design is what separates findings that can be trusted and applied from findings that are ambiguous or misleading. But what actually makes a research design “good”? There are three core criteria: the design must be capable of answering the research question, it must adequately control extraneous variables, and its results must be generalisable beyond the original sample. Understanding these criteria helps researchers build studies that are both theoretically sound and practically useful.
Table of Contents
- Why research design criteria matter
- Criterion 1: Capability to answer the research question
- Matching design to purpose
- Criterion 2: Control of variables
- Strategies for controlling variables
- Randomisation: the gold standard of control
- Criterion 3: Generalisability of results
- Population validity and ecological validity
- Representative sampling and the WEIRD problem
- The relationship between internal validity and generalisability
- Putting the criteria together: what makes a design “good”?
Why research design criteria matter
Research design is the overall plan for how a study will be conducted – including who will be studied, how data will be collected, and how variables will be managed. A flawed design can produce results that seem convincing but are actually unreliable, biased, or impossible to apply to the real world. According to research published in PubMed Central, the quality of a study’s design directly determines whether it can answer research questions without bias, and whether its findings can be trusted across different contexts. The criteria of research design are therefore not abstract ideals – they are practical standards that separate robust evidence from weak or misleading conclusions.
Criterion 1: Capability to answer the research question
The most fundamental question about any research design is simple: can it actually answer what it set out to investigate? A design must be structured in a way that directly addresses the research question – choosing the right method (experimental, observational, correlational), the right participants, and the right measures for the specific phenomenon being studied.
This sounds obvious, but it’s easy to get wrong. A researcher studying the long-term effects of therapy on depression cannot rely on a single-session observation. A study examining causal relationships between variables needs an experimental design with manipulation and control groups – a correlational survey simply cannot establish causation. Psychological researchers do not simply assume that their measures work – instead, they must demonstrate that the design and its instruments are genuinely capturing the phenomenon they are supposed to measure. If a study cannot clearly link its methods to its research question, everything downstream – the data, the analysis, the conclusions – becomes suspect.
Matching design to purpose
Different research questions call for different designs. Descriptive questions (“how common is social anxiety among teenagers?”) suit survey-based designs. Causal questions (“does social media use increase anxiety?”) require experimental manipulation. Matching the design to the purpose of the research is the first step toward meeting this criterion. A mismatch here means the study, no matter how carefully executed, will fail to provide the answers researchers and readers need.
Criterion 2: Control of variables
Even a well-designed study can be undermined by variables the researcher didn’t plan for. Extraneous variables are factors outside the main independent variable that could influence the dependent variable and distort results. When these variables are not managed, they become confounding variables – variables that muddy the relationship between what the researcher is manipulating and what they are measuring.
Consider a study on the effect of sleep deprivation on memory performance. If participants in the sleep-deprived group happen to be older on average, or if they consume more caffeine, those factors could independently affect memory. Without controlling for them, it becomes impossible to say whether memory problems stem from sleep deprivation or from these other differences. Controlling for extraneous variables reduces their threats on the research design and gives researchers a better chance to claim that the independent variable causes the changes in the dependent variable – in other words, it strengthens internal validity.
Strategies for controlling variables
Controls to assure internal validity can be accomplished through manipulation, elimination, inclusion, statistical control, and randomisation. Each strategy has its place depending on the design:
- Elimination involves holding a potential extraneous variable constant – for example, testing all participants in the same room at the same time of day to remove environmental variation.
- Inclusion means building the extraneous variable directly into the design as an additional independent variable, allowing its effects to be measured separately rather than ignored.
- Statistical control involves measuring extraneous variables and accounting for them during data analysis, using techniques such as analysis of covariance (ANCOVA).
- Matching pairs participants across groups on key characteristics so that groups are comparable before the study begins.
Randomisation: the gold standard of control
Of all the control techniques available, randomisation stands apart. Random assignment is the only control technique capable of handling both known and unknown confounding variables simultaneously. Every other technique requires the researcher to first identify what the confounding variable is before controlling it. Randomisation does not – it distributes the effects of all extraneous variables, including ones the researcher hasn’t thought of, roughly equally across experimental conditions.
Random assignment is an important part of control in experimental research because it helps strengthen the internal validity of an experiment and avoid biases. When each participant has an equal chance of being placed in any condition, the groups start out comparable. Any differences that emerge after the manipulation can then be attributed with greater confidence to the independent variable itself. This is why randomised controlled designs are considered the gold standard in causal research – they make a strong case for cause-and-effect relationships in a way that other designs cannot.
It is worth noting, however, that randomisation is not a perfect solution. The randomisation process works under the assumption that threats from extraneous variables are equally distributed over all experimental conditions, but threats to internal validity can still arise from study attrition and compensatory reactions among participants. Researchers should therefore combine randomisation with careful study monitoring throughout data collection.
Criterion 3: Generalisability of results
A study might be internally valid – well-controlled, carefully randomised, precisely measured – and still fail on another critical front: can its findings be applied beyond the specific group of people who participated? This is the question of generalisability, also referred to as external validity.
External validity refers to the extent to which the results of a study can be generalised beyond the specific context of the study to other populations, settings, times, and variables. A study finding that a memory intervention improved recall in a group of 20-year-old university students does not automatically mean the same intervention will work for elderly adults, children, or people in clinical settings. External validity determines how broadly the conclusions can travel.
Population validity and ecological validity
Generalisability breaks down into two important subtypes. Population validity refers to how well the findings extend to people beyond the study sample. Ecological validity refers to whether the findings hold up in real-world settings, not just in controlled laboratory conditions.
Ecological validity examines whether the study findings can be generalised to real-life settings, and is therefore a subtype of external validity. A study conducted in a tightly controlled lab may produce very clean data, but if the conditions bear little resemblance to everyday life, its ecological validity is questionable. For instance, testing memory recall in a silent, distraction-free room may not reflect how memory actually operates in a noisy, busy classroom or workplace.
This creates a genuine tension in research design. When conducting experiments in psychology, there is often a trade-off between internal and external validity – having enough control to rule out extraneous variables and randomly assign people to conditions, while also ensuring that results can be generalised to everyday life. Increasing one can sometimes reduce the other. A highly controlled laboratory experiment may maximise internal validity but sacrifice ecological relevance. Field experiments conducted in naturalistic settings may be more ecologically valid but harder to control.
Representative sampling and the WEIRD problem
One significant threat to generalisability in psychology is over-reliance on narrow samples. Samples from Western, Educated, Industrialized, Rich, and Democratic (WEIRD) countries are used in an estimated 96% of psychology studies, even though they represent only 12% of the world’s population. This limits how broadly psychological findings can be applied across different cultures, socioeconomic backgrounds, and life experiences. A research design that draws from a more diverse and representative sample produces findings with stronger generalisability – and more meaningful real-world relevance.
The relationship between internal validity and generalisability
It is tempting to treat internal validity and generalisability as competing goals, but they are better understood as complementary criteria that need to be balanced. Generalisability of findings is not assured even if internal validity is addressed effectively through design – strict controls to ensure internal validity can compromise generalisability. This is why researchers must think carefully about both from the outset, rather than optimising for one at the expense of the other.
Internal validity determines whether the estimate of effect in the study is likely to be accurate, while generalisability determines whether the findings are applicable to people outside the study sample. A study that scores highly on both criteria – producing accurate, unbiased results that apply to a broad population – represents the strongest possible evidence base for psychological knowledge and real-world application.
Putting the criteria together: what makes a design “good”?
A research design that meets all three criteria – capability to answer the research question, rigorous variable control, and strong generalisability – does not emerge by accident. It is the product of deliberate planning, methodological awareness, and honest appraisal of the design’s limitations. Researchers who randomise participants, control for extraneous variables through a combination of strategies, recruit diverse and representative samples, and choose methods that directly match their research questions are far more likely to produce findings that are both scientifically sound and practically useful.
Researchers need to determine the validity and reliability of each assessment to ensure that they are not misleading their readers, and that the data can be trusted based on statistical evidence to support their conclusions. No design is perfect – every study involves trade-offs. But understanding these three criteria gives researchers the tools to make those trade-offs consciously and to build studies that contribute meaningful, trustworthy knowledge to the field of psychology.
What do you think? When researchers face the trade-off between tight experimental control and real-world applicability, which criterion should take priority – and does your answer change depending on the type of research question being asked? Also, given how heavily psychology research has historically relied on WEIRD samples, how might findings from classic studies need to be reconsidered when applied to more diverse global populations?
References
- https://pmc.ncbi.nlm.nih.gov/articles/PMC6149308/
- https://opentextbc.ca/researchmethods/chapter/reliability-and-validity-of-measurement/
- https://stats.libretexts.org/Courses/Kansas_State_University/EDCEP_917:_Experimental_Design_(Yang)/01:_Introduction_to_Research_Designs/1.03:_Threats_to_Internal_Validity
- https://socialsci.libretexts.org/Bookshelves/Social_Work_and_Human_Services/Social_Science_Research_-_Principles_Methods_and_Practices_(Bhattacherjee)/05:_Research_Design/5.02:_Improving_Internal_and_External_Validity
- https://www.scribbr.com/methodology/random-assignment/
- https://www.sciencedirect.com/topics/mathematics/extraneous-variable
- https://www.scribbr.com/methodology/external-validity/
- https://en.wikipedia.org/wiki/External_validity
- https://pubmed.ncbi.nlm.nih.gov/15098414/
- https://www.jospt.org/doi/10.2519/jospt.2020.0701
- https://doi.org/10.11648/j.pbs.20241306.11
Leave a Reply