IQ Outcomes

Average IQ by Generation
Rising, Stalling, and What That Actually Means

Raw test performance climbed for most of the twentieth century, fast enough that applying today's norms backward would place the 1930s average near the threshold for intellectual disability. That conclusion is obviously wrong, and working out why is the whole subject.

Illustration of cognitive test score trends across successive birth cohorts

Quick Answer

Updated August 16, 2026 by Structural. Average IQ is always 100 by definition, because scores are renormed against the current population. The real question is what happened to raw performance, and the answer is that it rose substantially across the twentieth century and has stalled or reversed in several countries since.

Direct answer: the phenomenon is called the Flynn effect. Meta-analytic estimates place the rise at roughly three points per decade across the century for which good data exists, which compounds to around thirty points over a hundred years. That is two full standard deviations, an enormous figure that constrains what explanation is possible.

The crucial detail is where the gains happened. They were largest on fluid reasoning and visual processing measures and much smaller on measures of acquired knowledge like vocabulary and general information. A general increase in intelligence would not produce that pattern, which is why almost nobody in the field interprets the effect as one.

What the Flynn Effect Is

The effect is named after James Flynn, whose work in the 1980s established that it was a general phenomenon rather than a curiosity of particular tests.

The evidence came from an unusual source: test publishers' own data. When a battery is revised, the publisher administers both the old and new versions to a sample to establish the relationship between them. Flynn collected these comparisons and found a consistent pattern across instruments and decades. People scored higher on the older version than on the newer one, every time.

The interpretation is straightforward once stated. A person scoring 100 against norms from thirty years earlier scores below 100 against current norms, because the current reference population performs better in raw terms. The score has not changed, the comparison group has.

Flynn then showed the same pattern in national datasets covering entire birth cohorts, particularly military conscription data from countries that tested every young man. Those datasets are exceptionally strong evidence because they cover whole populations rather than samples, remove self-selection entirely, and were collected consistently over decades.

The finding has been replicated extensively. A formal meta-analysis covering a century of data across dozens of countries confirmed both the existence and the approximate magnitude of the gains, and established that they varied by ability domain and by period in ways discussed below.

How Large the Gains Were

The headline figure of roughly three points per decade sounds modest and is not, because it compounds.

IntervalApproximate raw gainIn standard deviations
One decadeAbout 3 points0.2
One generation, about 25 yearsAbout 7 to 8 points0.5
Half a centuryAbout 15 points1.0
One centuryAbout 30 points2.0

The generational row is the one that makes the effect concrete. Somebody tested against their grandparents' norms would score roughly fifteen points higher than against current ones, which is the difference between average and clearly above average. That is a large enough gap to change how a score would be described in any report.

These figures are averages across instruments and periods, and the variation around them is substantial. Gains were faster in some decades than others, faster in some countries than others, and as section 3 explains, dramatically different across ability domains. A single rate applied uniformly is a simplification that the underlying data does not support.

It is also worth noting that the strongest evidence comes from a limited set of countries, mostly in Europe, North America, and parts of East Asia, where either conscription data or repeated standardisation studies exist. Claims about the global trajectory rest on much weaker foundations than claims about, say, Norway or the Netherlands.

Where the Gains Concentrated

This is the single most important fact for interpreting the effect, and it is the one most often omitted when the phenomenon is reported.

The gains were not uniform across what tests measure. Measures of fluid reasoning and visual processing, particularly matrix reasoning tasks, showed very large gains. Measures of acquired knowledge, particularly vocabulary and general information, showed much smaller ones. In some datasets the knowledge measures barely moved at all.

Consider what a genuine increase in general cognitive ability would predict. General ability contributes to performance on every subtest, so an increase in it would raise all of them, roughly in proportion to how strongly each loads on the general factor. Vocabulary loads strongly on general ability, so it should have risen substantially. It did not.

The observed pattern is close to the opposite. Some of the largest gains occurred on subtests with relatively modest general factor loadings, and some of the smallest on subtests with high ones. Research examining this relationship has generally found the gains to be concentrated in ways inconsistent with a rise in the general factor itself.

What the pattern does fit is an increase in familiarity with abstract, rule-based, decontextualised reasoning of the kind that matrix items require and that formal schooling teaches. Flynn's own later interpretation ran along these lines: modern populations have become accustomed to treating hypothetical categories as real objects of reasoning, which is a specific cognitive habit rather than greater processing capacity.

The domains involved are described in Cognitive Domains, and the distinction between fluid and crystallised ability is what makes the differential gain interpretable at all.

Why You Never Noticed

If performance rose by two standard deviations, why does the average IQ remain 100 in every generation?

Because 100 is defined as the average, not measured as one. When a test is revised, the new standardisation sample's mean raw performance is assigned the value 100 by construction. Whatever that sample achieved becomes the new reference, and everybody is scored against it.

This means the scale is silently recalibrated every fifteen to twenty years, and the recalibration absorbs the entire effect. A population performing better in raw terms is scored against a tougher standard, and the reported average stays exactly where it was.

The consequence for individuals is the one covered in WAIS-IV vs WAIS-5. A score obtained against older norms is not comparable to one obtained against newer norms, and the direction of the difference follows from which way the population moved between the two standardisations. This is why norms expire and why clinical practice moves to current editions.

It also means the entire phenomenon is invisible from inside any single generation. You cannot detect it by taking a test, comparing with your peers, or reading your report, because every one of those comparisons is against your own contemporaries. It is visible only in the publisher data comparing editions and in longitudinal population datasets, which is why it took until the 1980s for anybody to establish it as a general fact.

The Paradox That Rules Out the Obvious Reading

The simplest interpretation, that people have become substantially more intelligent, collapses under an argument Flynn himself pressed hardest.

Apply current norms retrospectively. If raw performance rose by about thirty points across the century, then the average person in the 1930s would score around 70 against today's reference sample. On the modern scale, that figure sits at roughly the second percentile and near the conventional threshold used in diagnosing intellectual disability.

That conclusion is plainly false. The 1930s population built industrial economies, ran complex institutions, wrote the literature and mathematics we still teach, and produced the scientific work the present century is built on. Whatever the test scores say, that population was not operating near the threshold for intellectual disability.

Running the argument forward gives the same problem in reverse. Project the gains into the future and the average person in a century would sit around the 98th percentile of today's distribution, which would place the median person at the current threshold for high IQ society membership. Nothing about the trajectory of human capability supports that.

Both directions produce absurdity, and that is the argument's point. When a measurement's implications are absurd at both ends, the measurement is not tracking the thing it appears to track. The tests were measuring something that changed dramatically, and that something is not general intelligence in the sense the score names.

What it tracks is better understood as a shift in how populations engage with the kind of abstract problem tests present. The 1930s population would find matrix reasoning unfamiliar in a way a modern schoolchild does not, and unfamiliarity with a task format is not a deficit in capability.

What Causes It

No single explanation is established, and the leading candidates are not mutually exclusive. Each accounts for part of the pattern and none accounts for all of it.

Education. Both the duration and the character of schooling changed enormously across the century. Years of schooling rose substantially, and the content shifted toward abstract and formal reasoning. The causal effect of education on test performance is established through quasi-experimental designs, with meta-analytic estimates of roughly one to five points per additional year. This is the best evidenced mechanism and it accounts for a substantial share of the gains without covering all of them.

Nutrition and health. Childhood nutrition, infectious disease burden, and access to healthcare all improved markedly. Severe early deprivation demonstrably depresses cognitive development, so reducing it raises population performance. This explanation works well for the earlier part of the century and less well for later decades in countries where deprivation had already become uncommon.

Test familiarity and cognitive habits. Populations became accustomed to formal testing and to the abstract classification that test items require. This is Flynn's own later emphasis and it explains the differential pattern in section 3 better than the alternatives, since it predicts large gains precisely on the most abstract material.

Family size and environment. Smaller families concentrate parental attention and resources per child. Household environments became more cognitively stimulating in terms of print, media, and complexity of everyday technology.

Reduction in severe impairment. Better obstetric care and reduced exposure to environmental toxins, particularly lead, removed causes of impairment that previously pulled the lower tail down. This raises the mean without changing anything about the rest of the distribution.

The strongest evidence for environmental causation of some kind comes from the reversal, discussed next, which occurred too fast for any explanation operating over evolutionary time.

The Reversal

The gains have not continued everywhere. Several countries with long-running population-level data show the rise slowing, stopping, and in some cohorts reversing.

The strongest evidence comes from Nordic conscription data, where entire male birth cohorts were tested consistently for decades. Analysis of Norwegian records found gains turning to declines in cohorts born from the mid-1970s onward.

The most methodologically important feature of that analysis is the within-family comparison. The researchers examined brothers born at different times and found the same pattern of decline within families as between them. That design rules out compositional explanations entirely, since siblings share parents, household, and background. If the decline appeared only between families it could be attributed to changing population composition. Appearing within families, it must be environmental and operating on the birth cohort.

The authors titled their paper accordingly: the Flynn effect and its reversal are both environmentally caused. The reversal is not evidence against the environmental account of the original rise, it is evidence for it, because a rise and fall on this timescale cannot be anything else.

Causes of the reversal are not established. Candidates include changes in educational content and emphasis, shifts in how young people spend time, changes in reading habits, and the exhaustion of the gains available from improvements that have already been made. Any explanation must account for the reversal being observed in some countries and not others, which no current candidate does convincingly.

Evidence from other countries is mixed, with some showing continued gains, some plateaus, and some declines, and the pattern does not resolve into a simple geography.

Generational Labels Are Not Cohorts

Anybody searching for the average IQ of a named generation is asking a question the data cannot answer, and the reason is worth being explicit about.

Labels like Boomer, Millennial, and Gen Z are popular and commercial constructs with no fixed boundaries. Different sources place the cut-offs in different years, they vary by country, and no research protocol samples by them. Nobody has ever administered a battery to a representative sample defined by one of these labels, because the labels do not define a population precisely enough to sample.

The research operates on birth cohorts, meaning people born in a specific year or narrow range, because that is a well defined unit that can be sampled and followed. Cohort effects are gradual and continuous rather than stepping at label boundaries, and a person born in the last year of one label is essentially identical to somebody born in the first year of the next.

There is a second problem specific to comparing generations at a point in time. If you tested a twenty year old and a sixty year old today, the difference between them would combine cohort effects with age effects, and those move in opposite directions on different abilities. Processing speed and fluid reasoning decline gradually across adulthood while crystallised knowledge is stable or rising. Any cross-sectional comparison confounds the two completely.

This is precisely why norms are age banded, as covered in IQ Score vs Percentile. Comparing a sixty year old against twenty year olds would produce a difference driven by age rather than cohort, which is why every properly normed instrument compares people against their own age group.

Anybody claiming a specific IQ figure for a named generation is either reporting a cohort study under a label it did not use or has invented the number.

What This Means for Your Own Score

Three practical consequences follow for anybody holding a test result.

Your score is relative to your contemporaries. It says where you sit among people tested against the same reference sample, and it carries no information about how you would compare against a population from another era. Cross-era comparison is not something a standard score supports.

Older scores are not comparable to newer ones. A result from twenty years ago was measured against a different reference population. The difference is real, its direction follows from population drift, and no conversion formula recovers an individual's equivalent. This is why recency requirements exist independently of everything else.

The norms behind your score matter more than the test's name. An instrument published recently but standardised on an older sample carries older norms than its publication date suggests, and a local adaptation normed years after the original carries the age of its own standardisation. The question to ask is when the norm sample was collected, not when the test was released.

There is one further implication that is easy to miss. Because the gains concentrated on fluid and visual measures and barely touched knowledge measures, the norm drift affects some subtests far more than others. A person compared against newer norms loses more ground on reasoning measures than on vocabulary, which means an older score and a newer one can differ in profile shape as well as in level.

The World Became More Abstract

The explanation that best fits the differential pattern in section 3 deserves setting out properly, because it is more specific and more testable than the general appeal to environment.

Consider what a matrix reasoning item asks. It presents figures with no meaning, governed by a rule that exists only within the puzzle, and requires treating the arrangement as a formal system to be analysed on its own terms. Nothing about it connects to anything real. That is a peculiar mode of thought, and it is one that formal schooling explicitly trains.

Flynn drew on interview material from early twentieth century rural populations who, asked questions requiring hypothetical reasoning, frequently declined to answer in the terms offered. Asked what colour bears are in a place where all bears are white, respondents would answer that they had never been to that place and could not say. That is not a failure of reasoning. It is a refusal to treat a hypothetical as a legitimate object of reasoning, which is a different relationship to abstraction rather than a lesser capacity for it.

Across the century that relationship changed comprehensively. Schooling spread and its content became more formal. Work shifted from concrete manipulation toward symbolic manipulation. Everyday technology came to require operating menus, hierarchies, and abstract categories. Media became structurally more complex, demanding that viewers track multiple threads and infer unstated relationships.

A population trained by all of this to treat abstract systems as natural objects of thought will perform dramatically better on matrix items and only slightly better on vocabulary, because vocabulary depends on exposure to words rather than on comfort with formal abstraction. That is precisely the observed pattern, and it is the strongest reason to prefer this explanation over one framed in terms of general capacity.

It also predicts something useful: the gains should slow once a population is thoroughly saturated with formal schooling and abstract work, because the source of the gain is exhausted. Several countries showing plateaus are exactly those where saturation occurred earliest.

Where Norm Drift Has Real Consequences

Everything above can read as academic. It is not, and the clearest demonstration is what happens where a score determines an outcome at a fixed threshold.

Diagnosis of intellectual disability uses a cognitive criterion set roughly two standard deviations below the mean, alongside adaptive functioning criteria. Which norms produced the score therefore determines whether somebody falls above or below the line. A person tested on an instrument with ageing norms scores higher than they would on current norms, because the ageing reference sample sets a lower bar. That difference can move somebody across a diagnostic threshold without anything about them changing.

The consequences are concrete: eligibility for support services, educational placement, disability benefits, and in some jurisdictions determinations in criminal proceedings where intellectual disability is legally relevant. This is not hypothetical, and the effect of norm obsolescence on such determinations has been argued extensively in the forensic assessment literature.

The problem is symmetric and it is worth being precise about the direction. Obsolete norms overstate a person's standing, so somebody genuinely below a threshold may be scored above it and denied support they qualify for. Practitioners aware of the issue apply corrections or insist on current instruments, which is one of the strongest professional arguments for moving to a new edition promptly rather than continuing with a familiar one.

The general principle applies wherever a score meets a fixed cutoff. Society admission thresholds, gifted programme eligibility, and any selection criterion expressed as a number are all sensitive to which norms produced the score, and none of them is more sensitive than the diagnostic case. This is one of several reasons that applying a hard threshold to a single point value is poor practice, alongside the measurement error covered in IQ Score vs Percentile.

The Height Analogy, and Where It Breaks

Average height rose substantially across the same century, driven by nutrition and reduced childhood disease. The parallel is instructive both for what it explains and for where it fails.

What it explains: a population-level physical characteristic can shift markedly within a few generations through environmental improvement alone, with no genetic change whatsoever. Anybody inclined to treat a large cohort difference as implausible on principle should note that an uncontroversial and directly measurable trait did exactly this.

It also illustrates the renorming point neatly. If height were reported on a scale with mean 100 and standard deviation 15, recalibrated every twenty years, the average would always read 100 and the underlying rise would be invisible in the reported figure. That is precisely what happens with cognitive scores, and the analogy makes it obvious that a fixed average is a property of the scale rather than a finding.

Where the analogy breaks is instructive too. Height is measured on a ratio scale with a physical unit. A metre is a metre, it does not depend on a reference population, and comparing a person from 1920 with a person from 2020 is trivially valid. Cognitive scores have no such unit. There is no cognitive equivalent of a centimetre, only a position within a distribution, which is why cross-era comparison is genuinely problematic in a way that comparing heights is not.

The second break concerns what rose. Height gains appeared across the whole distribution and on the single dimension being measured. Cognitive gains concentrated on some abilities and barely touched others, which is a pattern with no analogue in the height case and is exactly what makes the interpretation contested.

How This Affects an Online Assessment

Everything above applies to any instrument, and there are two specifics worth stating for an online battery.

The first is that norm currency is a question to ask of any test, and one most online products cannot answer because they never state where their norms came from. A product with no stated norm sample cannot tell you how old its reference population is, which means it cannot tell you what its scores mean. ACIS documents the provenance and construction of its norms in technical materials rather than asserting them in marketing copy, which is the minimum standard this section implies.

The second is that a stated reference group is what makes a percentile interpretable at all. A percentile against a self-selected sample of online test takers overstates position substantially compared with a percentile against a general population reference, and the two are frequently presented identically. The reference group is the thing to check.

ACIS reports six domains with percentiles and confidence intervals against a stated reference group. Administration is unsupervised, conditions cannot be verified, and no institution is obliged to accept the result, all of which is stated in the report rather than around it.

What an assessment tells you is where you sit among people measured against the same reference. That is a genuinely useful thing to know and it is a narrower claim than a number floating free of any population, which is what a score without a stated norm sample amounts to.

FAQ: Generational IQ Change

Has average IQ been rising?

Raw test performance rose substantially across the twentieth century. Average IQ stayed at 100 throughout, because scores are renormed against the current population by definition.

How fast did it rise?

Roughly three points per decade on average, which compounds to about fifteen points over fifty years and thirty over a century, or two full standard deviations.

What is the Flynn effect?

The name for this rise, after James Flynn, who established in the 1980s that it appeared consistently across instruments, decades, and countries rather than in isolated datasets.

Does this mean people got smarter?

Almost nobody in the field reads it that way. Applying current norms backward would place the 1930s average near the threshold for intellectual disability, which is plainly false.

Why is that argument decisive?

Because projecting forward produces equal absurdity, placing the median future person near current high IQ society thresholds. Absurdity at both ends means the measure is not tracking what it names.

Where did the gains actually occur?

Largely on fluid reasoning and visual processing measures, particularly matrix reasoning. Vocabulary and general knowledge measures moved much less and in some datasets barely at all.

Why does that pattern matter?

Because a genuine rise in general ability would raise all subtests roughly in proportion to their general factor loadings, and the observed pattern is close to the opposite.

Why does the average always stay 100?

Because 100 is assigned to the mean of each new standardisation sample by construction. The scale is silently recalibrated every revision and the recalibration absorbs the entire effect.

Can I detect this by taking a test?

No. Every comparison available to you is against your own contemporaries, which is why establishing the effect required publisher data comparing editions and population datasets spanning decades.

What causes the rise?

No single explanation is established. Education, nutrition and health, familiarity with abstract reasoning, smaller families, and reduced severe impairment each account for part of it.

Which explanation has the best evidence?

Education, where quasi-experimental designs estimate roughly one to five points per additional year of schooling. It accounts for a substantial share without covering everything.

Have the gains stopped?

In several countries, yes. Nordic conscription data shows the rise slowing, stopping, and reversing in cohorts born from the mid-1970s onward.

Why is the Norwegian data so important?

Because the decline appears within families as well as between them. Siblings share parents and background, so a within-family pattern rules out compositional explanations entirely.

Does the reversal disprove the environmental explanation?

The opposite. A rise and fall on this timescale cannot be anything but environmental, which is why the researchers titled their paper accordingly.

What caused the reversal?

Not established. Candidates include changes in educational emphasis, shifts in how time is spent, and exhaustion of gains from improvements already made. None explains why it appears in some countries and not others.

What is the average IQ of Gen Z?

Unanswerable. Generational labels have no fixed boundaries, vary by source and country, and no study samples by them. Research uses birth cohorts, which are well defined.

Can I compare a young person and an older person today?

Not cleanly. The difference combines cohort effects with age effects, which move in opposite directions on different abilities, and cross-sectional comparison confounds the two.

Why are norms age banded?

Because comparing someone against a different age group would produce a difference driven by age rather than by anything about them. Age banding removes that.

Is my old test score still accurate?

It was accurate against the population it was measured against. It is not numerically comparable to a score from current norms, and no formula recovers an individual equivalent.

Does norm drift affect all subtests equally?

No. Because gains concentrated on fluid and visual measures, drift costs more ground on reasoning than on vocabulary, so profile shape can differ between an old score and a new one.

What should I ask about any test's norms?

When the norm sample was collected, how large it was, and who was in it. Publication date is not the same as norm date, and a test that cannot answer cannot tell you what its scores mean.

Best Next Step

Raw performance rose enormously and average IQ never moved, because the scale is rebuilt to keep it at 100. The gains concentrated where a rise in general ability would not have put them, which is why the effect is a puzzle about measurement rather than a story about people improving.

For why the average is fixed by definition, read Average IQ. For how norm drift affects a specific score, read WAIS-IV vs WAIS-5. For the separate question of whether an individual's ability can be changed, read Can You Improve Your IQ?. For your own position against a stated reference group, take the assessment.

Sources Behind This Page

Magnitude estimates come from formal meta-analyses rather than individual studies. The reversal evidence comes from population-level conscription records.

  • Trahan, L.H., Stuebing, K.K., Fletcher, J.M. & Hiscock, M. (2014). The Flynn effect: a meta-analysis. Psychological Bulletin, 140(5), 1332-1360. Comprehensive meta-analytic estimate of the magnitude of the gains and their variation across instruments and periods.
  • Pietschnig, J. & Voracek, M. (2015). One century of global IQ gains: a formal meta-analysis of the Flynn effect (1909-2013). Perspectives on Psychological Science, 10(3), 282-306. Cross-national synthesis covering a century, including domain differences and period variation.
  • Bratsberg, B. & Rogeberg, O. (2018). Flynn effect and its reversal are both environmentally caused. Proceedings of the National Academy of Sciences, 115(26), 6674-6678. Norwegian conscription data showing decline within families as well as between them.
  • Flynn, J.R. (1987). Massive IQ gains in 14 nations: what IQ tests really measure. Psychological Bulletin, 101(2), 171-191. The paper establishing the effect as a general phenomenon and arguing that the gains cannot represent a rise in general intelligence.
  • Ritchie, S.J. & Tucker-Drob, E.M. (2018). How much does education improve intelligence? A meta-analysis. Psychological Science, 29(8), 1358-1369. Quasi-experimental estimate of the causal effect of schooling, the best evidenced contributing mechanism.
  • McGrew, K.S. (2009). CHC theory and the human cognitive abilities project. Intelligence, 37(1), 1-10. The fluid and crystallised distinction that makes the differential pattern of gains interpretable.
  • Pearson (2008). WAIS-IV Technical and Interpretive Manual. Documents the standardisation procedure by which each revision's sample mean is assigned the value 100.
  • Pearson (2024). WAIS-5, Wechsler Adult Intelligence Scale, Fifth Edition. The most recent renorming of the adult scale, and the reason older scores are not directly comparable.
  • American Educational Research Association, American Psychological Association & National Council on Measurement in Education. Standards for Educational and Psychological Testing. Requirements that norms be current and that the reference population be documented.
  • Voncken, L., Albers, C.J. & Timmerman, M.E. (2019). Improving confidence intervals for normed test scores. Behavior Research Methods. Open access. Why the reference sample determines what any reported score means.
  • Crawford, J.R., Garthwaite, P.H. & Slick, D.J. (2009). On percentile norms in neuropsychology: proposed reporting standards. The Clinical Neuropsychologist, 23(7), 1173-1195. Why a percentile is uninterpretable without a stated reference group.
  • Buros Center for Testing. Mental Measurements Yearbook. Independent evaluation of whether commercial instruments report the age and composition of their norm samples adequately.
Take the assessment

You get a profile, not a number

ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.

Free trial, no card required. Full report from $15.