Your position relative to people your own age is one of the most stable measurements psychology has ever taken. Your raw cognitive performance is not stable at all. It climbs, peaks at different ages for different abilities, and then falls. Both statements describe the same person, and confusing them causes almost every misunderstanding on this topic.
The flat line is your position among age peers. The curved line is what you can actually do. Both belong to the same person.
0 The Short Answer
Your IQ score, meaning your standing relative to people your own age, changes remarkably little across a lifetime. Your raw cognitive performance, meaning what you can actually do, changes a great deal. Those are two different measurements, and the question "does IQ change with age" gets a different answer depending on which one you meant.
The evidence for the first claim is unusually good. In 1932 the Scottish Council for Research in Education tested almost every eleven year old in Scotland, 87,498 children, on a single paper and pencil instrument called the Moray House Test Number 12. Decades later, Ian Deary's group in Edinburgh tracked survivors down and gave them the same test again. Reporting in Psychology and Aging in 2011, Gow and colleagues found stability coefficients of .67 from age 11 to age 70, .66 from 11 to 79, and .51 from 11 to 87. Nothing else in psychology has been measured over a span like that.
Point the same instruments at absolute performance and the picture reverses. Speeded and novel reasoning tasks peak early and decline for decades afterwards, while vocabulary and accumulated knowledge keep climbing into later life. A single composite number hides that split completely, which is why what a score actually reports matters more than the number itself.
The bridge between the two facts is age norming. An IQ score is computed against a reference group of people the same age as you. A 100 at 25 and a 100 at 70 both mean average for that age, and they sit on top of very different raw performance. Most of the confusion on this topic dissolves once that is clear.
.67
Correlation between the same test taken at age 11 and again at age 70, N of 1,017 (Gow and colleagues 2011).
.51
The same correlation stretched to age 87, across 76 years, on a much smaller surviving sample.
Same 100
A score of 100 at 25 and at 70 describe identical standing on very different raw performance.
"Does IQ change with age" is not one question, and the pages that rank for it almost never say so. One version asks whether your position in the distribution moves. The other asks whether your absolute capability moves. They have opposite answers, and a page that gives only one of them is not wrong so much as half finished.
Rank order stability is the technical name for the first. It asks: if you lined up a thousand people by test score at age 11 and lined the same thousand up again at age 70, how similar would the two lines be? The answer, from the Lothian cohorts, is very similar but not identical. A correlation of .67 means the ordering is largely preserved and also means plenty of individuals swapped places.
Absolute change is the second. It asks: on the same items, under the same conditions, does a person get more or fewer right at 70 than at 25? For anything that depends on speed or on reasoning with unfamiliar material, the answer is fewer. For anything that depends on accumulated knowledge, the answer is often more. Nothing about the first answer constrains the second, because a correlation is computed on positions, not on levels. Every person in a cohort can decline by ten raw points and the correlation will not budge.
This is also the boundary between this page and its companion. Why the average stays 100 at every age is a population question about how the scale is constructed and standardized. It is answered on a separate page and is not repeated here. What follows is the within person question: what happens to one individual, measured twice, decades apart. Keeping those separate is the single most useful move available to a reader of this literature, and it depends on knowing what an intelligence test measures in the first place.
Scope of this pageThis page covers change within a single person across the adult lifespan. It does not cover why the population average is pinned at 100 in every age band, which is a question about scale construction and about the 15 point standard deviation rather than about ageing.
2 The Longest Follow Up in Psychology
The Scottish Mental Surveys were an accident of policy that turned into the best dataset on this question anyone will ever have. In 1932 the Scottish Council for Research in Education administered the Moray House Test Number 12 to 87,498 children born in 1921, essentially every eleven year old in the country. The exercise was repeated in 1947 with 70,805 children born in 1936. Both sets of scores sat in an archive for half a century, which is why this evidence exists at all: nobody designed a seventy year follow up, they simply found one waiting.
The test itself matters, because a stability coefficient is only as meaningful as the instrument producing it. The Moray House Test Number 12 is a group administered paper and pencil test with a 45 minute limit and 71 numbered items, 75 in total, scored out of 76. The item mix is broad: 14 following directions items, 11 same opposites, 10 word classification, 8 analogies, 6 practical items, 5 reasoning, 4 proverbs, 4 arithmetic, 4 spatial, 3 mixed sentences, 2 cypher decoding and 4 others. That is a general ability test in the ordinary sense, closer to the item formats used in group tests than to a single puzzle task, and it loads heavily on the general factor those items share.
Recruitment began in 1999. The Lothian Birth Cohort 1921 tested 550 survivors, 234 men and 316 women, in Edinburgh between 1999 and 2001, then again in 2003 to 2005, then a third time in 2007 to 2008. The Lothian Birth Cohort 1936 recruited 1,091 people between 2004 and 2007, tested at a mean age of 69.5 years. Every one of them sat the identical test they had taken as a child.
Interval
Cohort
Raw correlation
Corrected for range restriction
Sample
Age 11 to 70
LBC1936
.67
.78
N of 1,017
Age 11 to 79
LBC1921
.66
.76
542 tested at 79
Age 11 to 87
LBC1921
.51
.61
202 tested at 87
Age 79 to 87
LBC1921
.70
.75
Same third wave group
Age 11 to 90
LBC1921
.54
.67
N of 106
Two notes on reading that table. The sample column gives how many people completed the test at the later occasion, and the correlation itself is computed on the smaller number who have a usable score at both ages, so the pairwise figure is never larger and is usually a little smaller. The corrected column is a modelled adjustment for the fact that survivors span a narrower range of ability than the 1932 population did, not a second measurement.
The last row comes from a separate paper. Deary, Pattie and Starr, writing in Psychological Science in 2013, extended the follow up to age 90 in 106 people and reported a raw coefficient of .54, rising to .67 once range restriction was accounted for. They also checked that the old instrument was still doing its job, reporting strong concurrent and predictive validity for the Moray House Test against standard cognitive tests at both ages. Gow and colleagues described their own 76 year span, from 11 to 87, as believed to be the longest period over which a stability coefficient for intelligence had been computed, and the 2013 paper stretched it by three more years.
3 Who Is Still Alive to Be Measured
Every figure in the table above was computed on survivors, and survivors are not a random sample of the people who took the test in 1932. This is the limitation that most write ups of the Lothian findings skip, and the original papers are explicit about it.
The 2011 paper reported the comparison directly. Among Lothian Birth Cohort 1921 members, those who returned for the third wave of testing had scored higher as children than those who did not: 47.7 against 45.5 on the Moray House Test at age 11, a difference significant at p equal to .014. At age 79 the gap was much wider, 61.9 against 57.2, significant at p below .001. People who score higher live longer and are more likely to keep turning up for a clinic appointment in their late eighties, so the surviving sample gets progressively more selected with every wave.
47.7 vs 45.5
Age 11 Moray House scores of those who returned for the third wave against those who did not.
61.9 vs 57.2
The same comparison at age 79, where the selection is far more pronounced.
550 to 202
Lothian Birth Cohort 1921 members tested at the first wave, and still completing the test at 87.
Selection of this kind narrows the spread of scores, and a narrowed spread mechanically shrinks a correlation. That is why the authors report range restriction corrected values alongside the raw ones, and why the corrected figure for the age 87 comparison rises from .51 to .61. Their own note is worth keeping: the participants tested at 87 were even less representative of the original 1932 population than those tested at 79. Correction is a modelled adjustment, not a measurement, so the honest reading is that true lifelong stability sits somewhere at or above the raw coefficients rather than exactly on them.
Selection does not make the finding fragile. It cuts in a specific direction, and the direction is towards understatement. But it does mean a reader should treat .51 as a floor for the surviving population rather than a precise constant of human nature, in the same way that any reliability figure has to be read alongside the sample that produced it. The general principle is covered under how measurement quality is judged, and the sampling half of it under sampling and fairness problems in testing.
4 What a Stability Coefficient Does Not Promise
A correlation of .67 is a statement about a group, and it makes no promise at all about any particular person in it. This is the point at which the Lothian findings get oversold, usually by writers who are impressed by the number and have not asked what it would look like from the inside.
Start with shared variance. Gow and colleagues put the minimum estimates at 26 to 44 percent of the variance in later life scores accounted for by childhood performance on the same test. They give a range rather than a single figure because there is a genuine methodological argument, going back to Ozer in 1985, about whether squaring a correlation is the right way to express shared variance. Take the more generous end and well over half the variation in old age scores is still unexplained by the childhood score.
Now translate that into people. With a coefficient near .5, a person in the top fifth as a child is much more likely than chance to be in the top fifth in old age, and is still perfectly capable of finishing in the middle. Movement across the distribution is not an anomaly in these data. It is a large fraction of what the data contain. Individual trajectories cross, and the correlation summarises the average tendency of a thousand of them rather than the fate of any one.
There is a second reason the coefficient understates true stability, and it pushes the opposite way from selection. No test is perfectly reliable, so part of the observed disagreement between two administrations is simply measurement error rather than real change. Gow and colleagues note this explicitly. Between selective survival on one side and unreliability on the other, the raw numbers are conservative, and the sensible summary is that lifelong stability is high and clearly incomplete.
Practically, this is why a single score should be read as a band rather than a point, and why the exact percentile for a score and how rare a given score is are more useful framings than a bare number. A childhood score is evidence about an adult, not a verdict on one.
5 Why 100 at 25 and 100 at 70 Are Different Performances
An IQ score is a rank within an age band, and once you understand the arithmetic that produces it, most of the confusion about ageing and intelligence disappears. The score is not a count of what you got right. It is a statement about where your raw performance sits inside the distribution of raw performances produced by people your age.
The Lothian analyses show the procedure in miniature. To model change within old age, Gow and colleagues took the raw Moray House score, regressed it on the participant's age in days at the moment of testing, kept the standardized residual, and converted that residual to a scale with a mean of 100 and a standard deviation of 15. Age was removed by construction. Whatever remained described how a person performed relative to others tested at the same age, which is exactly what a published IQ score does using a normative sample instead of a regression.
The consequence is worth stating plainly. If a 70 year old and a 25 year old both score 100, they have both landed at the midpoint of their own age group, and the 25 year old will almost certainly have answered more items correctly on any speeded task. If the same person scores 100 at 25 and 100 at 70, their raw performance on speeded and novel material has fallen while their position among peers has held. That is not a contradiction. It is the design working as intended.
It also explains a common and reasonable complaint: that age norming flatters older adults. It does not flatter anyone. It answers a specific question, which is how a person compares with their contemporaries, and it deliberately refuses to answer a different one, which is how a person compares with their younger self. If the second question is what you care about, the score is the wrong instrument and raw performance on matched tasks is the right one. The machinery is described in full under the norming process itself, and the historical alternative that it replaced is covered under the mental age idea it replaced.
If you compare a group of 70 year olds with a group of 25 year olds today, every difference you find is a mixture of ageing and of everything else that separates people born in 1956 from people born in 2001. The Seattle Longitudinal Study exists because of that confound, and its results are the strongest demonstration available that the mixture matters.
The study began in 1956 as an ordinary cross sectional survey of adults aged 20 to 70. In 1963 Warner Schaie converted it into a multiple cohort longitudinal design, adding fresh random samples from the membership of a large health maintenance organization at seven year intervals and following participants until death or dropout. That structure lets the same six abilities be examined two ways: as differences between age groups measured once, and as changes within the same individuals measured repeatedly.
The two pictures do not agree, and the disagreement is systematic. Writing in Research in Human Development in 2005, Schaie reported that in the cross sectional data four of the six abilities show consistent negative age differences, reaching significance at age 46 for inductive reasoning, spatial orientation and perceptual speed, and at age 39 for verbal memory. Across the full age span the cross sectional gap averages roughly two standard deviations. In the longitudinal data, reliable within person decline appears decades later.
Ability
Cross sectional: first significant negative age difference
Longitudinal: first reliable within person decline
Verbal memory
Age 39
Age 67
Inductive reasoning
Age 46
Age 67
Spatial orientation
Age 46
Age 67
Perceptual speed
Age 46
Age 60
Numeric facility
Positive until midlife
Age 60
Verbal ability
Positive until midlife
Age 81
Numeric facility is the case that gives the game away. Cross sectionally it looks almost flat, with only minor differences between age groups. Longitudinally it shows early and steep decline. Schaie's explanation is cohort: over the 70 year range his data cover, inductive reasoning and perceptual speed rose by about a standard deviation across successive birth cohorts, while numeric facility drifted slightly downwards. A negative cohort trend makes an ability look stable in a snapshot even when individuals are losing ground fast. The broader version of that cohort story is the secular rise in test scores, and the site keeps a separate page on generational score comparisons.
7 Both Designs Are Biased, in Opposite Directions
The tidy conclusion would be that longitudinal data are correct and cross sectional data are misleading. The literature does not support the tidy conclusion, and the disagreement between two of the field's most serious researchers is the reason.
Timothy Salthouse has argued the opposite case for years. In Neurobiology of Aging in 2009 he concluded that some aspects of age related cognitive decline begin in healthy educated adults in their 20s and 30s, with cross sectional declines of about .02 to .03 standard deviation units per year for adults under 60, steepening to about .04 to .05 per year in a sample of roughly 800 adults aged 61 to 96. His argument is that longitudinal designs mask decline because people improve on a test simply by having taken it before, and that the improvement is large enough to swamp several years of genuine change. He is careful to exempt the crystallized side: measures built on accumulated knowledge, he reports, keep rising until at least age 60 in the same samples.
He put numbers on that in Current Directions in Psychological Science in 2014, using an episodic memory composite from the Swedish Betula project. Among adults born in 1949, tested at 40 in 1989 and again at 45 in 1994, the cross sectional comparisons showed the older group lower by 0.99 and 2.40 T score units on the two occasions, while the longitudinal comparison on the same people showed a gain of 2.08. Two separate methods for estimating change without prior test experience both put the real change at minus 1.26, a gap of 3.34 T score units between what the repeated testing showed and what the experience free estimate showed. Salthouse called the illustration crude, since it rests on two occasions and on group means, but the direction has been replicated in his Virginia Cognitive Aging Project, where more than 2,300 people had completed at least two occasions by December 2013.
Longitudinal designs also lose their weaker performers to attrition, exactly as the Lothian waves did. Putting both together, Salthouse concluded that within cohort comparisons resemble the cross sectional trends more closely than the longitudinal ones. Schaie's reply, published in Neurobiology of Aging in 2009, is that this reifies what he calls the cross sectional fallacy: inferring change within individuals from differences between groups measured once. Age changes and age differences can only coincide, he argues, in a perfectly stable environment with no cohort variation, and no such environment has existed in the last century.
The useful reading is that the truth is bracketed rather than pinned. Cross sectional estimates are inflated by cohort differences. Longitudinal estimates are deflated by practice effects and by the loss of weaker participants. Real decline in speeded and novel reasoning almost certainly begins earlier than the longitudinal curves suggest and later than the cross sectional ones do. Anyone selling a single confident age is selling one design's bias. This is also why preparation and practice before a session is treated carefully on this site rather than dismissed.
8 Fluid Reasoning Falls While Crystallized Knowledge Holds
The abilities inside a composite do not age together, and the gap between the earliest and the latest peak is measured in decades. In CHC terms, Gf, fluid reasoning, peaks early and declines from there. Gc, comprehension knowledge, holds or rises well into later life. Gs, processing speed, peaks earliest of all. A single number averages across all of them and shows you none of them.
The largest convergent test of this comes from Hartshorne and Germine, publishing in Psychological Science in 2015. They combined 48,537 online participants with the normative samples of the WAIS-III (2,450 people aged 16 to 89), the WMS-III (1,225 people) and 26,850 vocabulary responses from the General Social Survey collected between 1974 and 2012. Their finding was not that ability peaks at one age. It was that different abilities peak at wildly different ages, and that no age exists at which a person is at peak on everything.
Late teens
Peak for processing speed measured by digit symbol coding.
About 30
Peak for verbal and visual working memory in the web samples, replicated on two separate datasets.
Ages 40 to 60
Long plateau for emotion recognition, peaking significantly later than either working memory task.
About 50
Vocabulary peak in the WAIS-III normative sample, collected in the mid 1990s.
About 65
Vocabulary peak in the same authors' web sample, collected in 2010.
At least 60
Age to which vocabulary and general information keep increasing in Salthouse's samples.
That last pair is a cohort effect rather than a lifespan effect, and the authors tested it directly. Pooling the General Social Survey series with the Wechsler and web vocabulary data, they estimated that the age of peak vocabulary performance has been arriving about 0.96 years later with each passing year of testing, significant at p equal to .0003. The peak is not just late. It is getting later, for the same reason scores in general have drifted upward across generations.
Salthouse's own data agree on the crystallized side even while he disputes the timing on the fluid side. Measures built on accumulated knowledge, he reports, are consistently found to increase until at least age 60. That is the one point on which the cross sectional and longitudinal camps do not argue.
The consequence for reading your own result is direct. A 60 year old with strong verbal comprehension and slow processing speed is not showing a defect in one domain. That profile is what the lifespan literature predicts, and it is only visible if the abilities were measured separately. The definitions are set out under the Gf and Gc distinction in full, the taxonomy under the CHC structure behind the domains, and the speed domain on its own under processing speed as a measured domain.
9 Why One Number Hides the Pattern
If fluid and crystallized abilities move in opposite directions with age, then a composite is the worst possible summary of an ageing profile, because the two trends partially cancel. A person whose reasoning has slipped and whose vocabulary has grown can produce the same composite at 65 that they produced at 35, having changed substantially in both directions underneath it.
This is the honest case for measuring domains separately rather than reporting one figure. A battery that yields index scores across verbal comprehension, fluid reasoning, quantitative reasoning, visual spatial ability, working memory and processing speed shows the shape of the change. A composite alone shows only its net. Neither is more accurate than the other. They answer different questions, and only the profile answers the one an ageing reader is usually asking.
ACIS is built that way: 20 subtests across six primary cognitive domains, with subtest scaled scores on a mean of 10 and a standard deviation of 3, and composites on a mean of 100 and a standard deviation of 15. The Full Scale composite reaches an omega of .9886 with a g loading of .958 and a standard error of measurement of about 1.60 IQ points, figures drawn from a technical analysis set of 2,750 complete records. The higher order model fit was CFI .9761, TLI .9726, RMSEA .0406 and SRMR .0217, with a chi-square of 916.703 on 166 degrees of freedom.
What those figures do and do not establishThey describe internal structure and precision in the sample that produced them. ACIS is unsupervised, its adult reference frame of 3,243 English speaking records aged 16 to 90 is a modelled frame rather than a census sample, and the people in it selected themselves by choosing to take the test. None of that is fixed by a good fit index, and none of it makes the instrument clinical or diagnostic.
What a profile buys an older adult specifically is interpretive: it separates a low index that reflects the ordinary trajectory of speed from a low index that does not fit the pattern at all. That distinction is invisible in a composite. The domain list is set out under the six primary domains, the argument for breadth under a broad multi domain battery, and the underlying numbers under the published reliability tables.
"Fixed" is the wrong word in both directions, and the Lothian data show why with unusual clarity. Childhood ability turned out to be a powerful predictor of the level a person reached in old age, and no predictor at all of how fast they declined once they got there.
Gow and colleagues modelled Moray House performance between ages 79 and 87 with a growth curve that separated level from rate of change. Age 11 score predicted the level strongly, with a standardized coefficient of .57 at p below .001. It predicted the slope not at all, at minus .03 with a p value of .666. Years of education behaved the same way: it predicted level at .12, p equal to .001, and had no relationship with rate of change. The authors' summary is the part worth carrying away. Higher early life intelligence was protective of intelligence in old age because of stability across the lifespan, not because it slowed the decline experienced within old age.
That result cuts against two popular positions at once. It undermines the idea that a childhood score is destiny, because a large share of variance in late life ability is not accounted for by it. It also undermines the cognitive reserve story in its strongest form, at least for this measure and this age window, because the people who started higher fell at the same rate as everyone else. They simply fell from further up.
None of this is the same question as whether ability can be raised deliberately. That is a separate literature with its own evidence, and this site handles it separately under whether training moves the number. Nor is it a claim about genetics: stability across a lifespan is compatible with a range of causal accounts, which is the subject of what heritability does and does not mean. The mirror question, whether an exposure pushes ability down, has been attacked with the same logic of measuring first: the Dunedin birth cohort tested its members at 13, before anyone had used cannabis, and reported an eight point fall by 38 among persistent adolescent onset users, a figure resting on 23 people that largely disappeared once twins discordant for use were compared, a sequence set out under the replication record on cannabis and measured ability.
The defensible summary is that adult rank order is stable enough to be one of the most reliable individual differences psychology measures, and unstable enough that a substantial minority of people end up somewhere they would not have been predicted to end up. Both halves matter, and the prediction is to a distribution rather than to a point, which is the same caution that applies to school outcomes and measured ability.
11 Comparing Two Scores of Your Own
The practical version of this question is usually personal: someone has a score from school and a score from last week, and wants to know whether the difference means anything. Usually it does not, and there are four reasons that stack.
The first is measurement error. Every score is a point estimate with a band around it. On ACIS the standard error of measurement for the Full Scale composite is about 1.60 IQ points, and comparing two administrations means carrying the error from both, so the interval around a difference is wider than the interval around either score. Two numbers that differ by a few points are consistent with no real change at all.
The second is practice. In Salthouse's Betula illustration the gap between observed longitudinal change and change estimated without prior test experience was 3.34 T score units, which on a scale with a standard deviation of 10 is about a third of a standard deviation. That is not measured in IQ points and should not be converted into them, but it is large enough to reverse the sign of a five year change. If two administrations used the same or similar items, the later one is biased upward for reasons that have nothing to do with ability, and the shorter the gap between them, the worse the bias.
The third is that the tests were probably different instruments with different norms. A school screening from 1994 and an adult battery from today do not share a reference sample, an item pool, or a scaling decision. Comparing them is not comparing like with like, and the mismatch is often larger than any true change. The size of that gap is easy to underestimate, since the Stanford-Binet 5 still in clinical use is anchored to a reference sample collected in 2003 while the Woodcock-Johnson V was normed between 2022 and 2024, and the instrument by instrument inventory prints that norm year next to the age range for every battery, which is the quickest way to see whether two of your own scores were ever on the same footing. The site covers the practical version of this under free quizzes against validated instruments and how the main online batteries compare.
The fourth is age norming, discussed above: the reference group moved between the two occasions, so identical scores can conceal real raw change and different scores can conceal real raw stability.
What survives all four is a profile comparison rather than a composite comparison. If the same battery is administered twice with a meaningful gap, a shift in the shape of the profile, verbal holding while speed falls, is more interpretable than a shift in the single number. Even then, an unsupervised administration adds testing conditions to the list of things that differ between occasions, which is why the standard error of measurement is the first thing to read before either score is taken at face value.
12 What This Evidence Cannot Tell You
Nothing on this page is a way of finding out whether your own decline is normal. The literature describes averages and distributions in cohorts. It does not diagnose individuals, and an online score is not the instrument for that question even in principle.
The Lothian and Seattle findings describe healthy community dwelling adults. Pathological decline is a different phenomenon with a different trajectory, and it is identified by clinical assessment that includes history, informant report, medical investigation and supervised neuropsychological testing, not by comparing two numbers. A person worried about memory or reasoning needs a clinician, and the cost of an unnecessary reassurance from a self administered test is that they do not get one.
Not a clinical instrumentACIS is an unsupervised online assessment for adults aged 16 to 90. It is not a clinical or diagnostic instrument, and it is not suitable for detecting cognitive impairment, dementia, or any medical condition. It cannot be used for diagnosis, for hiring decisions, for accommodation requests, or for high IQ society admission. Concerns about cognitive change belong with a qualified clinician.
There are also limits inside the research itself that a careful reader should hold onto. The Lothian cohorts are Scottish, born in 1921 and 1936, educated in a specific system, and tested on an instrument designed in the 1930s. The Seattle sample was drawn from the membership of one health maintenance organization in one American city. Salthouse's participants were recruited by advertisement and scored roughly half a standard deviation to a full standard deviation above the Wechsler normative means. None of these is a representative sample of humanity, and the fact that they agree with one another on the broad pattern is more persuasive than any one of them alone.
Finally, the group statistics here are not individual predictions. A stability coefficient of .67 tells you about the average behaviour of a thousand trajectories. It tells you nothing about which of them is yours. That distinction matters as much here as it does in measured ability and later health outcomes, and it is the main reason a supervised assessment is a different product from an online one, as set out under supervised testing against unsupervised testing.
13 How to Read Your Own Score Across a Lifetime
The reader who came here asking whether their intelligence has changed is entitled to a usable answer, and it has three parts.
Your standing among people your own age has probably moved less than you think. That is the plainest implication of correlations of .67 and .66 measured across six and seven decades, and it holds even after allowing for the fact that survivors are a selected group. If you were a strong test taker at school, the base rate says you are still above average now, and the base rate also says a meaningful minority of people in your position are not. What you think is a poor instrument for settling it either way, since a self estimate pooled over 115 samples and nearly 37,000 people correlates about .30 with a measured score and predicts a range some 55 points wide, which is why the impression of having slipped carries far less information than a childhood coefficient does, as how poorly people place themselves on the scale makes clear.
Your raw performance has certainly moved, and it has moved in different directions in different domains. Speed and novel reasoning are lower than they were at 25. Vocabulary and accumulated knowledge are probably higher. Neither of those is visible in a composite, which is why a domain profile is the more informative output for anyone past their thirties and why the shape of the profile is worth more attention than the headline figure.
And the number itself is a comparison, not a quantity. It was computed against people your own age, which is what makes it interpretable and also what makes it useless for the specific question of how you compare with your younger self. If that is the question, matched raw performance is the measurement you want, and no standardized score will substitute for it. For interpretation of the number you do have, what counts as a good score and score interpretation and routing are the practical starting points, and the current Wechsler adult scale is the supervised reference the online instruments are measured against.
All of this sits inside a professional framework rather than outside one. The Standards for Educational and Psychological Testing, published in 2014 by the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education, require that a test be interpreted only for the population and purpose on which its evidence was gathered, that normative samples be described so users can judge their relevance, and that reliability evidence be reported alongside any score. APA testing standards make the same demand of anyone reporting a result. Applied to ageing, they say something specific and useful: an age normed score answers an age normed question, and a claim about what a person has lost or gained since their youth needs evidence the score was never designed to supply.
14 Frequently Asked Questions
Does IQ change with age?
Where you sit among same age peers moves very little. What you can actually do moves a great deal, climbing into early adulthood and then falling on speeded and unfamiliar material while knowledge based abilities keep growing. Two measurements, two answers, one person.
Does IQ increase with age?
Standardized IQ generally does not, because the comparison group ages with you. Certain underlying abilities do increase: vocabulary and general knowledge keep rising into at least the sixties in Salthouse's samples, and peaked around 65 in the web sample studied by Hartshorne and Germine.
Does IQ decrease with age?
Standardized IQ usually holds because it is scored against same age peers. Raw ability on speed dependent and novel reasoning tasks does decrease. In the Seattle Longitudinal Study, reliable within person decline in perceptual speed appeared by age 60.
Is IQ fixed for life?
No. Rank order among peers is highly stable but not perfect, with correlations near .67 across six decades leaving well over half the variance in later scores unexplained. A substantial minority of people finish somewhere their childhood score would not have predicted.
What is the Lothian Birth Cohort study?
Two Edinburgh follow up studies of people born in 1921 and 1936 who had taken the Moray House Test as eleven year olds in Scotland's national mental surveys. Researchers traced survivors decades later and readministered the identical test.
How stable is IQ from childhood to old age?
Gow and colleagues reported raw correlations of .67 from age 11 to 70, .66 to age 79 and .51 to age 87. Deary, Pattie and Starr later reported .54 from age 11 to age 90 in 106 people.
Why do the Lothian correlations get smaller with age?
Partly real change, partly measurement error accumulating, and partly selection. The people still returning at 87 had already scored higher at 11 and 79 than those who did not, which narrows the score range and mechanically shrinks the correlation.
What does a correlation of .51 actually mean for one person?
It describes how well a group keeps its ordering, not what happens to any individual in it. A high scoring child is more likely than chance to remain high scoring, and could still end up anywhere across the adult distribution.
What is the Moray House Test?
The general ability instrument used in Scotland's national mental surveys of 1932 and 1947. Pupils worked through it on paper in groups, within three quarters of an hour, across a broad mix of verbal, numerical, spatial and reasoning item types.
Why is IQ normed by age?
Because the question a score answers is how you compare with your contemporaries. Raw performance differs systematically by age, so without age based reference groups a fifty year old would be penalised on speeded tasks and rewarded on vocabulary for reasons unrelated to standing.
If I score 100 at 25 and 100 at 70, did nothing change?
Your position among peers did not change. Your raw performance almost certainly did, falling on speeded and novel tasks and possibly rising on knowledge based ones. The constant score conceals both movements because the reference group moved with you.
At what age does fluid reasoning peak?
Estimates differ by design. Cross sectional work places the turn in the twenties or thirties, while longitudinal work in the Seattle study places reliable decline in inductive reasoning around 67. The true figure is bracketed by those two biased estimates.
Why do cross sectional and longitudinal studies disagree?
Cross sectional comparisons confound ageing with cohort, since older groups were educated differently. Longitudinal comparisons are inflated by practice on repeated tests and by losing weaker participants to attrition. Each design errs in the opposite direction.
What is the cross sectional fallacy?
Schaie's term for inferring change within individuals from differences between age groups measured at one time. He argues the two coincide only in a perfectly stable environment without cohort variation, a condition never met in the last century.
Which cognitive abilities hold up best with age?
Crystallized abilities. Vocabulary, general information and accumulated expertise keep increasing to at least 60 and often longer. In the Seattle data, verbal ability was the last of six abilities to show reliable within person decline, at age 81.
Does a high childhood IQ protect against decline?
Not the rate of decline. In the Lothian growth curve model, age 11 score strongly predicted the level of ability at 79 and 87 but had no relationship with how fast people declined between those ages.
Can I compare an old school test result with a recent online score?
Not reliably. The instruments differ, the norm samples differ, practice inflates any repeat, and both scores carry measurement error. Differences of a few points between two different tests taken decades apart carry almost no information.
Does taking the same test twice raise your score?
Usually yes. Salthouse's Betula worked example separates what a repeated test showed from what an experience free estimate implied, and the two differ by 3.34 T score units, roughly a third of a standard deviation. Enough to hide real decline.
Can an online IQ test detect cognitive decline or dementia?
No. ACIS is unsupervised and is not a clinical or diagnostic instrument. Detecting pathological decline requires clinical assessment including history, informant report and supervised testing. Concerns about memory or reasoning belong with a qualified clinician.
How does a domain profile help an older test taker?
It separates change that follows the expected lifespan pattern from change that does not. A composite averages a falling speed index against a rising verbal one, so the net figure can look unchanged while both components have moved substantially.
Has ACIS studied how its own scores change with age?
No longitudinal ACIS evidence is published. The reference frame of 3,243 records aged 16 to 90 is cross sectional and self selected, so it cannot show within person change over time. The lifespan findings on this page come from external cohort studies.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.