Mensa, Wechsler and Stanford-Binet: How Official IQ Tests Work
Search for an assessment and you land in two very different worlds. One is free, instant and takes ten minutes on a phone. The other involves an appointment, a qualified administrator, a fee that runs into the hundreds and a report that arrives days later. Both hand you a number in the same familiar range, which makes it easy to assume they are doing the same job. They are not. Understanding what separates a supervised instrument from a casual online check is the difference between reading your result sensibly and reading far too much into it.
What Makes an Assessment "Official"
There is no global authority that stamps instruments as approved. What professionals mean by an official test is one that satisfies a short list of demanding conditions, and the vast majority of what you find through a search engine satisfies none of them.
- Standardisation. Every person sits the same items, in the same order, with the same instructions and the same time limits. Nothing is improvised.
- Norming. The instrument has been administered to a large, carefully selected sample chosen to mirror the population by age, sex, region and education. Your raw score is converted against that sample.
- Published reliability. The developers report how consistently the test measures, usually as internal consistency and test-retest figures, and they publish the confidence band around a reported score.
- Validity evidence. There is peer-reviewed work showing the score relates to the things it claims to relate to, and does not simply measure reading speed or familiarity with puzzles.
- Restricted access. Materials are sold only to qualified purchasers, precisely so the items do not leak and lose their value.
That last point explains something people find frustrating. If a full instrument were freely available, it would stop working within a year, because performance on a memorised set of items tells you nothing about reasoning.
The Wechsler Scales: The Clinical Workhorse
If a psychologist assesses an adult in a clinical or occupational setting, the instrument is most often a Wechsler scale. David Wechsler's insight in the late 1930s was that a single global figure hides more than it reveals, so his scales break performance into separate domains and report both the parts and the whole.
A modern adult version samples verbal comprehension, perceptual or visual-spatial reasoning, working memory and processing speed. Each domain draws on several subtests, each subtest is scored against age norms, and the domain scores combine into a full-scale figure. Sittings typically run between sixty and ninety minutes and are conducted one-to-one, with the administrator observing how you work as well as recording what you answer. The individual tasks will feel familiar to anyone who has met the standard formats described in the piece on the question types used in these assessments, though the supervised versions run considerably deeper.
The domain profile is frequently more informative than the headline number. Two people can arrive at an identical full-scale score by completely different routes: one strong verbally and slow on timed tasks, the other the reverse. In assessment work that pattern is often the entire point, because an unusually uneven profile can flag something worth investigating that a single figure would smooth away entirely.
Stanford-Binet: The Oldest Name Still in Service
The Stanford-Binet traces directly back to the work Alfred Binet and Théodore Simon carried out in Paris in the early twentieth century, adapted for American use at Stanford by Lewis Terman. It is the instrument that popularised the notion of an intelligence quotient, and it remains in active use more than a century later.
Current editions cover an exceptionally wide age span, from early childhood into late adulthood, and are structured around verbal and non-verbal routes through several reasoning factors. It has a particular reputation at the extremes of the distribution, where it tends to discriminate more finely than alternatives, which matters when an assessment concerns unusually high or unusually low ability. The broader story of how these instruments developed is covered in the history of intelligence testing, and that background explains a good deal about why modern tests look the way they do.
Where Mensa Fits In
Mensa is not a test publisher. It is a membership society with a single entry criterion: evidence of scoring at or above the ninety-eighth percentile on an approved assessment. That distinction causes endless confusion, because a search for the society's name returns dozens of pages promising its test for free, and none of them can offer any such thing.
There are two legitimate routes to membership. The first is sitting a supervised session run by the national society, held at scheduled times and venues, using instruments the society has approved. In Britain that has long meant a combination that includes the Cattell III B, a heavily verbal paper covering vocabulary, comprehension and reasoning under tight time pressure. The second route is submitting prior evidence: a report from a qualified psychologist showing a qualifying score on an accepted instrument, which the society reviews against its list.
What the society does publish openly is practice material and short unsupervised puzzles designed to give a rough indication of whether a formal attempt is worthwhile. These are useful and honestly labelled, but they carry no weight for membership, and a strong result on one is not a qualifying score.
The Norwegian Test and Other Free Alternatives
One free instrument comes up constantly in discussion and deserves separate mention. The test published by the Norwegian society uses matrix reasoning items in the tradition of progressive matrices: each item shows a grid with one cell missing, and you select the piece that completes the pattern. It is untimed in some versions, timed in others, and it produces a figure that people compare enthusiastically with supervised results.
It is a reasonable piece of work by the standards of free material, and a great deal better than the average pop-up quiz. It is still not equivalent to a supervised session. It samples one narrow slice of reasoning, it has no control over your conditions, and its norms rest on self-selected volunteers rather than a representative sample. Treat the result as an indication, not a credential. If you simply want a sense of where you sit before committing time and money to anything formal, taking a straightforward iq test in a quiet room is a perfectly sensible starting point. The nature of what these shorter checks can and cannot do is set out in more depth in the piece on what a ten-minute check actually measures.
Why the Same Person Scores Differently on Different Tests
People routinely sit two respectable instruments and come away with results ten or twelve points apart, then conclude one of them must be broken. Usually neither is. Several ordinary factors produce that spread.
- Different content mixes. A verbally weighted paper and a pure matrix test are sampling different abilities. Someone with a wide vocabulary and average spatial reasoning will do better on the former, and the gap is real rather than an error.
- Different reference samples. A score is a position within a comparison group. Change the group and the position changes, even if performance is identical.
- Different scale conventions. Not every instrument spreads scores by the same amount around the centre. A figure of 130 does not sit at the same rarity on every scale, which quietly breaks a lot of casual comparisons.
- Ceiling effects. Tests built for the general population run out of difficult items near the top. Two very strong performers can hit the same maximum without being equally strong.
- Ordinary day-to-day variation. Sleep, illness, caffeine, anxiety and how recently you last did anything resembling a puzzle all move results by a few points in either direction, as the material on sleep, diet and cognitive performance describes.
What Supervision Actually Buys
The fee for a formal session pays for control. An administrator confirms who is sitting the test, enforces timing exactly, keeps conditions consistent, prevents the use of notes or search engines, and records qualitative observations that never reach an automated report. That control is what makes the resulting number defensible when something rides on it: an educational placement, an occupational decision, a diagnostic question, a society application.
It also buys interpretation. A qualified assessor explains the confidence band, points out where the subtest profile is uneven, notes anything in your circumstances that might have depressed a particular score, and puts the whole thing in context. That interpretive layer is the part free tools cannot replicate at any price, and it is usually worth more than the digits themselves.
Choosing the Right Route for Your Situation
The practical question is not which instrument is best in the abstract, but which is proportionate to what you need. If you are curious, a free reasoning test in decent conditions answers the question at zero cost. If you want society membership, only an approved supervised route or an accepted prior report will do. If a decision about schooling, employment or health depends on the outcome, the assessment should be conducted by a qualified professional using a current instrument, and nothing less will carry weight.
Whatever route you take, hold the result loosely. Formal instruments are the most carefully engineered psychological measures in existence, and even they report a range rather than a point. A number is a snapshot of certain reasoning abilities on one particular morning, and it says nothing at all about persistence, judgement, decency or the willingness to keep learning, all of which shape how a life actually goes. Where the parents' version of this question comes in is covered in the material on assessment in childhood, and the practical business of booking, paying for and sitting a session in Britain is set out in the piece on testing in the UK.
Related Reading
Keep exploring how the mind works with these free articles:
