When a psychologist sits down with a client to administer a psychological test, that test comes with a rulebook – precise instructions about how to ask questions, how to record answers, how much time to allow, and how to calculate the score. These rules exist for a reason. They are the product of rigorous research and validation. But what happens when a psychologist, perhaps with good intentions, decides to bend those rules? They extend the time limit for a nervous client. They rephrase a confusing question. They award partial credit where the scoring manual says none is due. These small, well-meaning deviations seem harmless on the surface – but they are not. They fundamentally compromise the integrity of the assessment, a process known as unstandardising a standardised test.
Table of Contents
- What standardisation actually means
- The two pillars that unstandardisation destroys
- Reliability: consistency under consistent conditions
- Validity: measuring what you claim to measure
- Common ways unstandardisation occurs
- Altering instructions or the way questions are presented
- Modifying time limits
- Changing scoring criteria
- Using a test outside its validated population
- Situational factors: a reason to document, not to deviate
- The ethical and clinical consequences
- Maintaining integrity without ignoring human complexity
What standardisation actually means
Standardisation in psychology is the process of making a test uniform – setting it to a specific standard by administering and scoring it in exactly the same manner for every person who takes it. A standardised test is not just a collection of questions. It includes detailed administration instructions, a validated scoring system, and normative data drawn from large, representative samples. These norms allow a psychologist to compare an individual’s score against the wider population, making it possible to say whether someone’s cognitive performance, emotional profile, or personality traits fall within or outside an expected range.
Before a psychological test reaches clinical use, it goes through extensive development, piloting, and revision by specialists known as psychometricians – professionals who study the science of measuring psychological constructs like intelligence, personality, and mental health. The result of that process is a tool whose results are only meaningful when the tool is used exactly as designed. The Standards for Educational and Psychological Testing, developed jointly by the American Educational Research Association (AERA), the American Psychological Association (APA), and the National Council on Measurement in Education (NCME), represent the authoritative benchmark for best practice in validity, reliability, design, and scoring of psychological tests.
The two pillars that unstandardisation destroys
Every credible psychological test rests on two foundational properties: reliability and validity. When a test is unstandardised, both are placed at risk.
Reliability: consistency under consistent conditions
Reliability refers to the consistency of a measure. A reliable test produces the same results under the same conditions, regardless of when or where it is administered or who administers it. Psychologists consider three types of consistency: consistency over time (test-retest reliability), consistency across items (internal consistency), and consistency across different administrators (inter-rater reliability). When any one of these is undermined by procedural deviation, the test’s scores become less trustworthy.
According to principles for evaluating psychometric tests published by the National Institutes of Health, a test’s results must be consistent across time, across items, and across raters to be considered reliable. Poor reliability does not just reduce the precision of a score – as research published in Advances in Methods and Practices in Psychological Science warns, it has severe implications for validity and interpretation, making it nearly impossible to determine whether results reflect genuine psychological attributes or are simply the product of measurement error.
Validity: measuring what you claim to measure
Validity is traditionally defined as the degree to which a test actually measures whatever it purports to measure. A test can be highly reliable – producing consistent scores every time – and still be completely invalid if it is not measuring the right thing. When a psychologist alters how a test is administered or scored, the scores generated can no longer be compared against the original normative data. The norms become irrelevant because the test is no longer the same test the norms were built for.
Common ways unstandardisation occurs
Unstandardisation does not always look like deliberate misconduct. It often creeps in through seemingly reasonable adjustments.
Altering instructions or the way questions are presented
Test manuals specify that directions be presented verbatim. Changing the wording of instructions – even to make them clearer – introduces a variable that was not present when the test was normed. A client who receives simplified instructions is not responding to the same test that the normative sample responded to. Their score cannot be meaningfully interpreted against those norms.
Modifying time limits
Time constraints are an integral part of many psychological assessments, particularly IQ tests and cognitive assessments. Extending the time allowed because a client appears to struggle introduces an uncontrolled variable that inflates performance beyond what the test was designed to measure, rendering comparisons with standardised norms invalid. This is different from formally approved accommodations for disabilities, which themselves require documented justification and carry specific interpretive caveats.
Changing scoring criteria
When a psychologist awards partial credit for answers that the scoring manual categorises as incorrect, or alters the weight given to specific responses, the resulting score no longer reflects what the test was designed to measure. As noted in guidance from the North Carolina Psychology Board, administering one edition of a test and then scoring it using norms from a different edition is equally problematic – the results of such an assessment carry no valid interpretive weight.
Using a test outside its validated population
Using a cognitive test developed for one age group on clients from a different age group, or applying norms developed in one cultural context to a population from a different cultural background, is a form of unstandardisation with serious consequences. Research on ethical issues in psychological assessment notes that when the selection and use of an instrument deviate from its intended purpose and population, psychologists must acknowledge and communicate the limitations, biases, and errors that may arise – and results should be interpreted with considerable caution.
Situational factors: a reason to document, not to deviate
One of the most common justifications offered for unstandardising a test is that the client’s performance was affected by situational factors – anxiety, fatigue, illness, or environmental disruption. These factors are real. Research consistently shows that test anxiety, defined as a combination of physiological over-arousal, worry, dread, and fear of failure, can create genuine barriers to performance. Studies find that highly test-anxious individuals score meaningfully below their low-anxiety counterparts.
Similarly, research published in AERA Open demonstrates that stress exposure can negatively affect cognitive functioning and test performance through multiple biological pathways, with worse cognitive performance typically occurring at both very low and very high levels of stress arousal.
The correct professional response to situational factors is not to alter the test. It is to document them. The APA’s Ethical Code of Conduct, specifically Section 9.06, explicitly identifies “situational, personal, linguistic and cultural differences” as factors that are imperative to consider in reducing the impact of client-specific variables – not by changing the test, but by accounting for them in the interpretation of results. A note in the assessment report that the client appeared highly anxious or fatigued during testing is far more defensible – and far more ethical – than a quietly modified procedure that produces a score that looks valid but is not.
The ethical and clinical consequences
The consequences of unstandardised assessment stretch beyond the individual test session. Decisions made on the basis of psychological assessments can have profound effects on a person’s life – determining eligibility for educational support, influencing clinical diagnoses, shaping treatment plans, or informing legal proceedings. When a test is unstandardised, those decisions rest on a flawed foundation.
Research in the field of psychometric ethics has found that adherence to ethical standards increases the reliability of test results, reducing the risk of biases and errors that could affect individuals’ opportunities and outcomes. A survey by the British Psychological Society found that 82% of professionals in psychology consider ethical considerations essential in psychometric test administration. Yet the same body of research notes that a significant proportion of psychological assessments contain errors in interpretation – underscoring how often ethical standards around standardisation are not upheld in practice.
The APA Guidelines for Psychological Assessment and Evaluation apply broadly to all aspects of assessment, including test selection, administration, scoring, interpretation, and report writing. Psychologists are expected to adhere to standardised procedures not as bureaucratic compliance but as a professional obligation to the people they assess.
Maintaining integrity without ignoring human complexity
None of this means that psychologists must be robotic in their approach to clients. Good clinical practice demands sensitivity. But sensitivity is expressed through how a psychologist prepares a client for assessment, how they communicate results, and how they contextualise findings – not through altering the test itself. When there is a genuine lack of appropriate standardised instruments for a specific population, the ethical path is to either source a validated alternative, acknowledge the limitations explicitly in any report, or supplement with qualitative clinical information from multiple sources such as interviews, observations, and collateral reports.
As the National Institutes of Health overview of psychological testing makes clear, test publishers provide detailed manuals with verbatim instructions and scoring guidance precisely because consistent, standardised administration is the only basis on which a score can be meaningfully interpreted. Departing from those instructions invalidates the interpretive framework the score depends on.
Standardisation is not about rigidity for its own sake. It is about fairness – ensuring that every person assessed is measured against the same yardstick, under the same conditions, so that their score reflects their actual psychological profile rather than the particular choices a clinician made on a given day. Preserving that standard is not just a technical requirement. It is an ethical one.
What do you think? If a client appears visibly distressed during a psychological assessment, where should the line be drawn between accommodating their emotional state and maintaining the integrity of the standardised procedure? And when situational factors clearly affected a client’s performance, how transparent should assessment reports be in communicating this to those who will act on the results?
References
- https://study.com/academy/lesson/standardization-and-norms-of-psychological-tests.html
- https://en.wikipedia.org/wiki/Standards_for_Educational_and_Psychological_Testing
- https://opentextbc.ca/researchmethods/chapter/reliability-and-validity-of-measurement/
- https://www.ncbi.nlm.nih.gov/books/NBK581902/
- https://journals.sagepub.com/doi/10.1177/2515245919879695
- https://www.britannica.com/science/psychological-testing/Primary-characteristics-of-methods-or-instruments
- https://www.ncpsychologyboard.org/data/documents/NCPB_newsletter_wintwer2022_final-1page.pdf
- https://pmc.ncbi.nlm.nih.gov/articles/PMC11335701/
- https://en.wikipedia.org/wiki/Test_anxiety
- https://journals.sagepub.com/doi/10.1177/2332858417713488
- https://www.linkedin.com/pulse/ethics-technology-trends-psychological-testing-sarah
- https://blogs.psico-smart.com/blog-ethical-considerations-in-psychometric-test-administration-8514
- https://www.apa.org/about/policy/guidelines-psychological-assessment-evaluation.pdf
- https://www.ncbi.nlm.nih.gov/books/NBK305233/
Leave a Reply