How do we measure something as complex and multifaceted as human intelligence? Today, IQ scores are a familiar part of education, clinical psychology, and even popular culture. But the tools we use to assess intelligence didn’t appear overnight. They are the product of over a century of scientific debate, experimentation, and gradual refinement – beginning with a Victorian polymath who believed that faster reflexes meant a sharper mind, and culminating in tests that now assess everything from verbal reasoning to spatial processing. Understanding that history isn’t just an academic exercise – it reveals how our very concept of intelligence has changed, and how those changes shaped the tools clinicians and educators still use today.
Table of Contents
Galton’s groundwork: measuring the mind through the senses
The scientific story of intelligence assessment begins with Sir Francis Galton (1822-1911), a British polymath and cousin of Charles Darwin. Deeply influenced by Darwin’s theory of evolution, Galton believed that intellectual ability was an inherited characteristic – a product of biology as much as breeding. His 1869 book Hereditary Genius was one of the first systematic attempts to study genius scientifically, and it set the stage for decades of debate about whether intelligence is born or made.
Galton’s approach to measurement was built on a specific assumption: that intelligence could be assessed through sensory acuity and reaction time. He reasoned that individuals with superior sensory perception – sharper vision or hearing – would have a more detailed understanding of their environment, contributing to better decision-making and problem-solving. Similarly, he theorized that faster reaction times reflected greater cognitive efficiency.
To test these ideas, Galton established the first Anthropometric Laboratory in the 1880s, where he measured an array of physical and sensory qualities in the general public, including height, visual acuity, reaction time, and the ability to discriminate between sounds and pressures. Visitors paid an admission fee and walked through a series of measurement stations – generating one of the largest datasets on human physical and sensory variation ever assembled at the time.
Despite the ambition behind this project, Galton’s specific hypothesis – that sensory ability directly reflects intelligence – was ultimately disproven by subsequent research. What he left behind, however, was methodologically significant. He pioneered the use of statistical methods such as correlation and regression to the mean in the study of human differences, and introduced the use of questionnaires as research instruments. His student James McKeen Cattell carried these ideas to America, where they laid the groundwork for the next generation of mental testers. Galton is widely regarded today as the founder of psychometrics – the science of measuring psychological attributes.
Alfred Binet and the birth of practical intelligence testing
While Galton was mapping the outer edges of human sensory performance, a different and more consequential approach was taking shape in France. Alfred Binet (1857-1911) started his career as a lawyer before gravitating toward psychology. His interest in intelligence grew partly from observing differences in the intellectual development of his two daughters – observations that informed his belief that intelligence was not a single fixed trait, but something that changed and developed with age.
The immediate trigger for Binet’s most important work was a practical, institutional problem. In 1904, the French Ministry of Public Instruction appointed a commission to address a pressing educational challenge: with compulsory education newly in place, schools faced an influx of students with varying abilities, and teachers struggled to identify which children genuinely required special support versus those who were simply unmotivated or had behavioral difficulties. Binet, along with physician Thรฉodore Simon, was tasked with developing an objective method for this identification.
Binet’s approach was a radical departure from Galton’s. Rather than relying on sensory acuity or skull measurements, Binet proposed that intelligence involved higher-order mental processes including attention, memory, imagination, comprehension, and judgment. Together, Binet and Simon published the Binet-Simon Scale in 1905 – a set of 30 tasks arranged in order of increasing difficulty, designed to distinguish children with normal development from those who needed additional educational support.
The 1905 scale was refined significantly in 1908, when Binet introduced one of his most enduring contributions: the concept of mental age. Binet reasoned that the age at which a child could perform a task could be used as a measure of intelligence – a child who could not accomplish tasks typical of three-year-olds demonstrated subnormal intellectual development for their age. A child’s mental age was thus the most advanced level of tasks they could successfully complete. The full version of the test, with age-appropriate standards ranging from age 3 to 13, was published in 1908 and became the basis for intelligence testing internationally.
Importantly, Binet himself was cautious about the implications of his test. He and Simon stressed that intelligence was not a fixed, single-number concept, that its development was subject to environmental influences, and that the test was only valid for children from comparable backgrounds. Binet worried about the stigma a score might impose on a child – a concern that proved prescient, given what happened when his ideas were adapted and exported abroad.
The IQ formula and the Stanford-Binet scale
Binet’s work attracted attention well beyond France. One of the most consequential readers of the Binet-Simon scales was Lewis Terman, a psychologist at Stanford University. Terman translated and substantially revised the scale for an American population, publishing the result in 1916 as the Stanford-Binet Intelligence Scale. The new scale included detailed administration and scoring instructions, and over a third of the items were new.
Crucially, Terman incorporated a formula developed by the German psychologist William Stern: the Intelligence Quotient, or IQ. It was Stern who transformed Binet’s mental age score into a ratio – dividing mental age by chronological age and multiplying by 100 – creating a single number that could theoretically compare intelligence across different ages. A child with a mental age equal to their chronological age would score exactly 100, representing average intelligence.
Binet himself had resisted applying this kind of ratio to his tests, fearing the stigma that individuals might face as a result of their score. His caution was warranted: the Stanford-Binet, and the IQ concept it popularized, would soon be used in ways that went far beyond educational placement – including to justify racial and ethnic discrimination, and to support the eugenics movement. This is a troubling chapter in the history of psychology that cannot be separated from the history of intelligence testing itself.
World War I and the rise of group intelligence testing
The entry of the United States into World War I in 1917 created an urgent, large-scale challenge: how do you rapidly screen and assign hundreds of thousands of military recruits with widely varying abilities and educational backgrounds? The answer changed the course of intelligence testing permanently.
Under the direction of psychologist Robert Yerkes, who was then president of the American Psychological Association, the U.S. military developed two new group intelligence tests: the Army Alpha, designed for literate English speakers, and the Army Beta, developed for recruits who were non-literate or did not speak English. The Alpha subtests covered arithmetic, analogies, disarranged sentences, and number series, while the Beta relied on memory tasks, picture completion, and geometric construction.
Approximately 1.75 million men were tested through this program, making it the first large-scale intelligence testing initiative in history. Recruits were scored on a scale from A to E, and results were used to inform decisions about job placement and fitness for service. The Army Beta, in particular, was significant because it attempted to assess intelligence independently of literacy – a recognition that verbal ability and general intelligence were not the same thing. Such tests were regarded by many as “culturally fair” because they did not discriminate against those with limited education or language ability.
The wartime testing program had significant drawbacks, however. A major weakness was the failure to consider the impact of cultural differences on test performance – lower scores for foreign-born soldiers were attributed to limited native ability, rather than to limited acculturation or schooling. Nonetheless, the Army Alpha and Beta tests demonstrated that intelligence could be measured at scale, and their format and content would go on to influence developments in both group and individual testing for decades to come.
David Wechsler and a broader vision of intelligence
The next major transformation in intelligence assessment came from David Wechsler (1896-1981), an American psychologist who had gained first-hand experience with the Army Alpha and Beta tests during World War I. That experience shaped his conviction that a single IQ score was insufficient to capture the full range of human cognitive ability.
Wechsler disapproved of the heavily verbal content of the early Stanford-Binet and of its tendency to produce a single global IQ as the only measure of a person’s intellectual level. He believed that intelligence was broader than verbal reasoning – encompassing spatial ability, working memory, processing speed, and practical judgment. In his own formulation, intelligence represented the global capacity to act purposefully, think rationally, and deal effectively with one’s environment.
In 1939, Wechsler published the Wechsler-Bellevue Intelligence Scale, designed specifically for adults – addressing a major gap, since the Stanford-Binet had been developed primarily for children. He designed his test to produce both a verbal and a performance (non-verbal) IQ score, combining the strengths of both the Army Alpha and Beta approaches into a single comprehensive instrument. The structure proved highly effective, and the innovations introduced in the Wechsler-Bellevue have remained largely intact through all subsequent revisions of the scale.
Over time, Wechsler extended his framework to other age groups. Today there are three primary intelligence tests credited to Wechsler: the Wechsler Adult Intelligence Scale (WAIS-IV), the Wechsler Intelligence Scale for Children (WISC-V), and the Wechsler Preschool and Primary Scale of Intelligence (WPPSI-IV). Together, they represent the most widely used family of intelligence assessment instruments in contemporary psychology.
Wechsler also moved away from the ratio IQ formula. Rather than dividing mental age by chronological age, the Wechsler scales calculated IQ as a deviation from the average performance of others in the same age group – with 100 fixed as the mean and scores interpreted in relation to a normal distribution. This approach, known as the deviation IQ, is now the standard across modern intelligence testing, including revised versions of the Stanford-Binet.
What this history tells us
The evolution of intelligence assessment is not a straightforward march from error to truth. It is a story full of genuine scientific advances, serious ethical failures, and gradual conceptual refinement. Galton gave us psychometrics and the tools to measure individual differences, but his underlying theory of intelligence was wrong. Binet gave us the first practical, clinically useful test, but worried – correctly – about how it would be misused. Terman and the Army testers proved intelligence could be measured at population scale, but used those measures to draw discriminatory conclusions. Wechsler broadened the conception of intelligence and gave clinicians far more useful diagnostic tools, but the debate about what IQ scores actually mean has never fully stopped.
Today’s intelligence tests are more sophisticated, more normed, and more sensitive to cultural and demographic factors than anything Galton could have imagined. But they remain imperfect instruments for measuring something that may ultimately resist any single number. The history of intelligence assessment is, in many ways, a history of psychology learning – through its mistakes as much as its breakthroughs – what it means to measure a human mind.
What do you think? Has the shift from measuring sensory abilities to assessing complex cognitive processes brought us closer to understanding what intelligence truly is – or has it simply made the definition more complicated? And given the ethical misuse of early intelligence tests, what responsibilities do psychologists have today when designing or interpreting assessments?
References
- https://en.wikipedia.org/wiki/Francis_Galton
- https://www.cogn-iq.org/galton-intelligence-testing.php
- https://www.cogn-iq.org/sir-francis-galton-intelligence-testing.php
- https://www.cogn-iq.org/learn/history/francis-galton/
- https://www.cogn-iq.org/learn/tests/binet-simon-scale/
- https://en.wikipedia.org/wiki/Binet%E2%80%93Simon_Intelligence_Test
- https://www.ebsco.com/research-starters/history/alfred-binet
- https://en.wikipedia.org/wiki/Alfred_Binet
- https://historyofpsychologicaltesting.wordpress.com/intellectual-assessment/early-intelligence-testing/
- https://pmc.ncbi.nlm.nih.gov/articles/PMC6526414/
- https://www.sciencedirect.com/topics/social-sciences/stanford-binet-intelligence-scales
- https://historyofpsychologicaltesting.wordpress.com/intellectual-assessment/wwi-beyond/
- https://www.sciencedirect.com/topics/medicine-and-dentistry/stanford-binet-intelligence-scale
- https://theconversation.com/show-us-your-smarts-a-very-brief-history-of-intelligence-testing-45444
- https://www.sciencedirect.com/article/abs/pii/S0160289618302411
- https://opentext.wsu.edu/psych105nusbaum/chapter/measures-of-intelligence/
- https://pubmed.ncbi.nlm.nih.gov/11992219/
- https://en.wikipedia.org/wiki/Mental_age
Leave a Reply