In short
- An IQ score is normed on humans and floats free of meaning when applied to AI.
- Contamination means high IQ scores may reflect memorization, not reasoning.
- Memorization-resistant benchmarks like ARC-AGI expose the gap between apparent and real ability.
- Honest rulers (GPQA, FrontierMath) are accurate but illegible to non-technical audiences.
- The IQ chart is a door for starting conversations, not a ruler for making decisions.
A chart made the rounds last week: frontier AI models placed on a human IQ scale, side by side, like students on a class ranking. I shared it. It traveled further than most of my posts do.
That reach is worth thinking about, because the chart itself is a bad measurement. And yet it's one of the most useful artifacts in AI communication right now. Both things are true, and the tension between them is the actual story.
Why the number is weak
Start with what an IQ score actually is. It's a statistic normed on human populations, built to predict human outcomes: school performance, job performance, that kind of thing. The number only means something relative to a reference group of people. Apply it to a system with perfect recall, no working-memory limit, and strange blind spots in spatial reasoning, and you get a number floating free of the distribution that gave it meaning. You can compute it. You can't interpret it.
Then there's the contamination problem, which is the sharper critique. The public IQ tests people run these models through (Mensa Norway is the common one) circulate widely online. Which means the questions may sit somewhere in the training data. A high score can reflect memorization rather than reasoning. Projects that track AI "IQ" over time, like Tracking AI, know this, which is why some have moved to offline test versions that models can't have seen. Even then, the construct mismatch stands: you're grading a very different kind of mind on a very human curve.
The cleanest demonstration of the gap comes from benchmarks built specifically to resist memorization. ARC-AGI, and its harder successor ARC-AGI-2, tests novel abstraction with puzzles that don't reward pattern lookup. Models that would post impressive scores on a Mensa-style quiz still lag humans badly there. A system can look brilliant on a human IQ scale and stumble on abstraction tasks a child handles. That combination tells you the IQ number isn't measuring what people assume it measures.
Researchers who need a real signal use instruments designed for the job: GPQA for graduate-level science questions, FrontierMath for research-level mathematics, Humanity's Last Exam for expert-level breadth that models haven't already saturated. These are the honest rulers. They're also, for most people, completely illegible.
Wrong in the specifics, right in the gestalt.
Why the chart works anyway
Here's the thing I keep running into in keynotes and boardrooms: someone always asks a version of "yeah, but how smart are they really?" And I can answer with GPQA percentages and ARC-AGI pass rates, and I'll watch the room glaze over. Those numbers are accurate and meaningless to a non-technical audience, because they have no felt reference point. Nobody knows what 60% on a graduate physics benchmark feels like.
Everyone knows what an IQ scale feels like. Average sits at 100. Gifted starts somewhere north of 130. You've calibrated this scale your whole life, through school, through colleagues, through the people you've worked with. So when a chart places frontier models on that scale, a non-technical reader gets, in one frame, an immediate sense of where the bar sits. Wrong in the specifics, right in the gestalt.
That's why I called it a conversation-opener rather than a benchmark, and why I'm expanding on it here instead of walking it back. The chart is a translation device. It converts something illegible (benchmark scores in unfamiliar units) into something legible (a scale you already carry in your head). Translation always loses precision. The question is whether what survives is worth having. In this case it is, as long as you know what you're holding.
The mental model: rulers versus doors
The distinction I'd offer: some measurements are rulers and some are doors. A ruler has to be accurate, because decisions rest on it. A door just has to get someone into the room where the real conversation happens.
GPQA, FrontierMath, ARC-AGI: rulers. Use them when you're deciding what a model can actually do, whether to deploy it, where it will fail.
The IQ chart: a door. Use it when someone outside the field needs a first foothold, then walk them through to the harder truth on the other side, which is that these systems don't sit anywhere on a human scale. They're superhuman in some directions and subhuman in others, simultaneously. The jagged profile is the real finding. No single number can carry it.
The failure mode is treating a door like a ruler. That's when someone reads "this model has an IQ of X" and starts making hiring or policy decisions on it. The other failure mode is subtler: dismissing the door entirely because it isn't a ruler, and leaving most of your audience outside the room.
Expect more of these charts, with bigger numbers, as models improve. Each one will be technically indefensible and communicatively effective. Knowing which job you're asking the chart to do is the whole skill.