How IQ Scores Actually Work: Percentiles, Standard Deviation, and What "Average" Really Means
An IQ score isn't a percentage correct or a ranking against everyone alive — it's a position on a specific bell curve, with a margin of error built in. Here's what the number is actually built from, and what it can't tell you.
Two people finish the same 35-item matrix test on the same afternoon. One lands at 112. The other lands at 127. It's tempting to read that fifteen-point gap as settled fact — the second person is measurably sharper, end of discussion. But that's not really what the number is built to say. An IQ score isn't a percentage of questions answered correctly, and it isn't a ranking against every person who has ever taken a test. It's a position on a specific, carefully constructed curve, and understanding how that curve works is most of what you need to read any IQ result honestly, including your own.
That curve is built around one pairing of numbers that quietly does almost all the work: a mean of 100 and a standard deviation of 15. It shows up on the Wechsler scales, the Stanford-Binet, and effectively every mainstream test in use today, including the one on this site. Once you know what those two numbers actually mean, "average," "gifted," and "genius-range" stop being vibes and start being specific, checkable claims.
Mean 100, SD 15: what "average" really means
A score on this scale isn't earned the way a test grade is — it isn't "you got 80% of the matrices right, so your IQ is 80." It's computed by comparing your raw number of correct answers to how everyone else in a large, representative norm group performed, then converting that comparison into a standardized position. Statisticians call that position a z-score: how many standard deviations above or below the average your result sits. That z-score then gets rescaled onto the familiar 100/15 metric using a simple formula — IQ = 100 + 15 × z.
Worked out, that means: if your raw performance is exactly average for your norm group, z is 0, and your score is 100. If you perform one full standard deviation better than average, z is +1, and 100 + (15 × 1) = 115. That 115 doesn't mean you answered "15% more" questions correctly — it means your result sits at roughly the 84th percentile, or better than about 84 out of 100 people in the comparison group. The scale is really a percentile-mapping tool wearing a more familiar-looking number.
- About 68% of scores fall between 85 and 115 — within one standard deviation of average.
- About 95% fall between 70 and 130 — within two standard deviations.
- About 99.7% fall between 55 and 145 — within three standard deviations.
That distribution is also why specific thresholds carry specific weight instead of just sounding impressively large. A score of 130 marks roughly the top 2% of the distribution — the traditional cutoff for Mensa admission — while 145 is closer to one person in a thousand. The jump from 115 to 130 represents a much bigger move across the curve than the jump from 100 to 115, even though both are "15 points."
Why matrices instead of vocabulary
Not every IQ test looks like this one. Full clinical batteries such as the Wechsler scales mix several subtests — verbal comprehension, working memory, processing speed, visual-spatial reasoning — because psychologists treat intelligence as a bundle of related but separable abilities. PsychIQ's test deliberately isolates one specific piece of that bundle: pure abstract, figural reasoning, the kind made famous by Raven's Progressive Matrices.
That format has a real history. Psychologist John C. Raven developed it in the 1930s, publishing the Standard Progressive Matrices in 1938, and stripped the test down deliberately — no words, no numbers, no facts to recall, just a 3×3 grid of shapes with one cell missing and a set of options to complete it. The design worked well enough that by 1942 the British armed forces were already using it to screen recruits regardless of their schooling, in what's generally considered the first large-scale operational use of a nonverbal intelligence test.
"Culture-fair" is the term psychologists use for that design choice, and it means something specific: your result depends on spotting a rule — a shape rotating, a quantity increasing row by row, a Latin-square arrangement, two patterns overlaying by simple set logic (the same rule types this site's 35-item test cycles through) — and extending it, not on vocabulary size, growing up speaking the test's language, or knowing a particular cultural fact. It is a genuinely more level playing field than a test built on word knowledge or arithmetic word problems. It is worth saying plainly, though, that no test is fully culture-free — familiarity with test-taking itself, motivation, and processing speed under time pressure still matter. "Culture-fair" is a real, meaningful, and limited claim, not an absolute one.
Psychologists sometimes describe what matrix tests measure as fluid intelligence — reasoning through a genuinely new problem on the spot — as distinct from crystallized intelligence, the accumulated knowledge and vocabulary you've picked up over a lifetime. A matrix score is a solid proxy for the first. It was never designed to capture the second.
A single number vs. an honest range
Even a professionally administered IQ test — given one-on-one by a trained psychologist, under controlled conditions, using an instrument normed on thousands of people — doesn't report just one number. Every test carries measurement error: day-to-day fluctuation, small guesses, attention lapses, the imperfect reliability of any finite set of questions. Psychologists account for that with something called the standard error of measurement, and for a well-normed individually administered test with reliability around .95, that error typically works out to roughly 3 to 4 IQ points.
Now the part that matters for any test you take on a laptop, including this one: a self-administered online test doesn't have a trained examiner controlling for distractions, screen glare, or a bad night's sleep, and it hasn't been calibrated against the kind of large, demographically representative norm sample that a published clinical instrument carries. That should widen the honest margin of error, not shrink it. It's why this site shows your result as an estimated range — about ±8 points — rather than a bare number. It's a genuine estimate of where you likely sit on the same 100/15 curve a psychologist would use, built from original items with verified answer keys, but it is not the same instrument a licensed examiner would administer, and it isn't a score you'd submit to a gifted program or a Mensa application. Treat it the way it's presented: a well-built estimate with a range attached, not a certificate.
What a matrix score can — and can't — tell you
Abstract reasoning is a real, meaningful thing to measure. Decades of research tie it — modestly but consistently — to outcomes like academic performance and certain kinds of occupational success, because the ability to spot a pattern in something unfamiliar and generalize it turns out to be useful in a lot of situations. That's a fair case for taking a well-built matrix test seriously, and for being genuinely pleased with a strong result.
What it isn't is a measure of your worth, your creativity, your emotional intelligence, or how "smart" you are in any complete sense of that word. It says nothing about whether you can read a room, generate a genuinely original idea, stay steady under pressure, or navigate a messy real-world problem that doesn't come with six answer choices and one correct key. Those are separate, legitimately important capacities — and psychology has entirely different instruments, and entirely different open debates, about how well any of them can be measured at all.
So: know what the number is — a position on a well-defined curve, with an honest margin of error attached — and what it isn't: a verdict on the rest of you. If you're curious where you land on that curve, the 35-item version is a fast, fair way to find out.
For self-reflection and education — not a clinical diagnosis or substitute for professional support.