Professional Testing

WAIS-IV vs WAIS-5
What Actually Changed, and Whether It Matters to You

Sixteen years separate the two editions. The headline change is that one index became two, but the change that affects more people is invisible: the norms were rebuilt, and that alone moves scores.

Illustration comparing the WAIS-IV and WAIS-5 index structures

Quick Answer

Updated August 16, 2026 by Structural. The Wechsler Adult Intelligence Scale, Fourth Edition was published in 2008. The Fifth Edition, published by Pearson in 2024, replaced it. Both cover ages 16 through 90.

Direct answer: the structural headline is that the Perceptual Reasoning Index was split into two separate indices, one for visual spatial ability and one for fluid reasoning. This brings the adult scale into line with the child scale, which made the same change a decade earlier, and with the theoretical framework the research literature had already adopted.

The practical headline is different and affects more people. The norms were rebuilt on a new standardisation sample. Because population performance drifts over time, the same raw performance generally maps to a somewhat lower standard score against newer norms, which means a WAIS-IV score and a WAIS-5 score are not directly comparable even when nothing about the person changed.

Why Tests Get Revised at All

Cognitive batteries are revised roughly every fifteen to twenty years, and the reasons are worth understanding because they explain what a revision does and does not fix.

The first reason is norm obsolescence. A standard score is a comparison against a reference sample, and that sample was drawn at a point in time. Population performance on cognitive measures has drifted upward across the twentieth century, a phenomenon documented across many countries and instruments. If the reference sample ages while the population moves, the norms progressively overstate everybody, and a score of 100 against a thirty year old sample no longer means average.

The second is theoretical. Understanding of cognitive ability structure has changed substantially, and a battery designed around a two-factor verbal and performance split reflects a model the field has moved past. Revisions are the mechanism by which instruments catch up with the evidence about what they are measuring.

The third is content. Items age badly. Vocabulary words fall out of use, general knowledge questions become obscure or trivially searchable, and pictured objects become unrecognisable to younger examinees. An instrument whose items reference technology from two decades ago is measuring exposure as much as ability.

The fourth is practical. Batteries get long, and revisions typically prune redundant subtests and rebalance administration time. Reducing burden matters because fatigue in the last third of a session degrades the scores obtained there.

A revision addresses all four at once, which is why the resulting instrument is not simply the old one with fresh norms and why scores across editions require care to compare.

The Index Structure Change

The clearest difference between the editions is in how many indices they report and what those indices represent.

The WAIS-IV reported four: Verbal Comprehension, Perceptual Reasoning, Working Memory, and Processing Speed. Perceptual Reasoning combined two things that the evidence had come to treat as separable, namely reasoning about relationships between abstract elements and constructing or manipulating visual representations.

Splitting them separates Visual Spatial ability from Fluid Reasoning, which is the same change the child scale made in 2014 and which brings the adult instrument into line with the Cattell-Horn-Carroll framework described in The CHC Model.

Broad abilityWAIS-IV indexCurrent structure
Crystallised knowledge (Gc)Verbal ComprehensionVerbal Comprehension
Visual processing (Gv)Perceptual ReasoningVisual Spatial
Fluid reasoning (Gf)Fluid Reasoning
Working memory (Gwm)Working MemoryWorking Memory
Processing speed (Gs)Processing SpeedProcessing Speed

Why this is more than cosmetic: the two abilities dissociate in real people. Somebody can reason well about abstract relationships while finding mental rotation and visual construction difficult, or the reverse. Under a combined index those two profiles produce the same number, and the information distinguishing them is discarded before it reaches the report.

The clinical consequence is largest where the dissociation is diagnostically meaningful. Certain developmental and acquired conditions affect visual spatial processing selectively, and a combined index can average a real deficit into an unremarkable score. Reporting the two separately is what makes that pattern visible.

The cost is that indices built from fewer subtests are less reliable than one built from more. Splitting an index in two means each half rests on a narrower base, and that is a genuine trade rather than a free improvement. It is the reason the split was made only after factor analytic evidence had accumulated that the two components were separable enough to justify it.

The Norms Change, and Why It Moves Scores

This is the change that affects the most people and gets the least attention, because it is invisible in the structure of the report.

Every revision is standardised on a new sample, drawn to represent the current population by age, sex, education, region, and ethnicity. The new norms replace the old ones entirely. That means the same raw performance is now being compared against a different reference group.

The direction of the effect follows from the drift in population performance. Where the population has improved on the material a test measures, a newer reference sample sets a higher bar, so identical raw performance yields a lower standard score. This is a well documented phenomenon in the intelligence literature and is the reason clinical guidance warns against using an instrument whose norms have aged substantially.

The size of the effect is not uniform across abilities. The drift has historically been largest on measures of fluid reasoning and visual processing and considerably smaller on measures of crystallised knowledge, which means a person retested on a newer edition may see their reasoning indices fall relative to their verbal indices for reasons entirely unrelated to them.

Two practical implications follow. A score obtained on an older edition should not be compared numerically to one obtained on a newer edition without accounting for the norm change, and any interpretation that treats the difference as change in the person is likely to be wrong. And an assessment conducted on an obsolete edition when a current one exists is a defensible criticism of the assessment, which is why professional practice moves to new editions rather than continuing with the familiar one.

The related and separate issue of confidence intervals still applies on top of this. A revision changes which reference group you are compared against. It does not make any individual score precise, and both editions report scores with error attached for the reasons in IQ Score vs Percentile.

Comparing a Score Across Editions

People with an older score naturally want to know what it would be now. The honest answer has three parts and none of them is a conversion factor.

First, there is no exact conversion. The instruments differ in norms, in subtest composition, and in index structure, so a mapping from one to the other would require assuming those differences away. Publishers conduct linking studies in which a sample takes both editions, and those studies establish the average relationship, not a formula that applies to an individual.

Second, the expected direction is downward for the reason in section 3, and the magnitude depends on how much time separates the standardisations and which abilities are involved. Somebody whose strength is crystallised knowledge will see less movement than somebody whose strength is fluid reasoning.

Third, the index structure change means some comparisons have no counterpart at all. A WAIS-IV Perceptual Reasoning Index has no single equivalent in the current structure, because the ability it represented is now reported as two numbers. If those two numbers differ substantially, no single figure summarises them honestly, and the old index was averaging a difference that the new structure exposes.

What to do with an older score in practice depends on the purpose. For personal interest it remains informative about relative standing, with the caveat that it was measured against a population that has since moved. For any formal purpose, recency requirements usually apply independently, and an assessment more than a few years old will need repeating regardless of edition.

The exception worth noting is that an older score is highly valuable as a comparison point when a new assessment is conducted. A clinician evaluating possible cognitive decline wants the earlier result precisely because the change over time is the finding, and they will account for the edition difference rather than being confused by it. Keep old reports.

There is a further complication in that comparison which is easy to miss. Retesting on any cognitive measure produces some gain from practice, and that gain works in the opposite direction to the norm change. Somebody retested on a newer edition is subject to a downward pressure from newer norms and an upward pressure from having done something similar before, and the two partly cancel by an amount nobody can determine for an individual. This is why clinicians interpret retest differences against published expectations for change rather than by simple subtraction, and why a difference of a few points across editions and years is close to uninterpretable on its own.

Working Out Which One You Were Given

If you have a report and want to know which edition produced it, several markers settle it quickly.

The most direct is that reports state the instrument and edition by name in the identifying section, usually on the first page. This is required practice rather than a courtesy, and its absence is itself a finding about the report's quality.

Failing that, the index structure is diagnostic. A report showing four indices with one named Perceptual Reasoning is from the Fourth Edition or earlier. A report showing separate Visual Spatial and Fluid Reasoning indices is from the current structure. This distinction is unambiguous and does not require knowing subtest names.

The date is a weaker signal than it appears. A report from 2023 is necessarily Fourth Edition since the Fifth had not been published, but a report from 2025 could be either, because practitioners do not switch on publication day for reasons covered in the next section.

If you were assessed in a language other than English, the relevant question is which local adaptation was used and when it was standardised. Adaptations are standardised separately on local samples, sometimes years after the original, so the edition number alone does not tell you how old the norms behind your score are. That is the number that matters for section 3, and a report should state it.

Why the Transition Takes Years

A new edition does not replace the old one immediately, and understanding why explains why both remain in circulation.

Cost is the first reason. A full test kit represents a substantial investment, and practices replace them on a schedule rather than on publication. Institutions with multiple clinicians face a multiple of that cost.

Training is the second. Administration and scoring must be learned, and a clinician administering a new instrument for the first time is slower and more error-prone until fluent. Practices commonly run a deliberate transition period rather than switching abruptly.

Continuity is the third and most legitimate. A clinician following a patient over years may deliberately continue with the older edition to preserve comparability, because the value of a directly comparable retest can exceed the value of current norms for that specific question. This is a defensible clinical judgment rather than inertia.

Localisation is the fourth. Non-English adaptations follow the original by a period that is often measured in years, since each requires its own translation, item review, and local standardisation. In many countries the current edition of record is the previous one simply because the newer adaptation does not exist yet.

The practical upshot for anybody being assessed is that it is reasonable to ask which edition will be used and why. A clinician using an older edition should have a reason, and continuity for a follow-up assessment is a good one while unfamiliarity with the new one is not.

When the Difference Actually Matters

The edition question ranges from decisive to irrelevant depending on what the assessment is for.

Matters a great deal

Threshold decisions: eligibility determinations, placement cutoffs, and diagnostic criteria expressed as a score. A few points of norm drift can cross a boundary.

Society thresholds
Matters moderately

Profile interpretation where visual spatial and fluid reasoning dissociate. The older structure cannot show a difference the newer one reports.

Cognitive domains
Matters little

General personal interest in approximate standing. Both editions place you in roughly the same part of the distribution.

Score vs percentile

The threshold case deserves emphasis because it is where people are most affected and least aware. Any decision made by comparing a score to a fixed cutoff is sensitive to which norms produced the score, and norm drift can move somebody across a boundary without anything about them changing. This is one of several reasons that using a single point value against a hard threshold is poor practice, the others being measurement error and the base rate issues covered in The Subtests Inside an IQ Test.

The general-interest case is worth stating plainly too. If you took a WAIS-IV a decade ago and want to know roughly where you stand, the answer is still roughly where that score put you. The edition difference is real and it is smaller than the confidence interval on most individual scores, so it does not justify discarding what you know.

The middle case is where most curiosity actually sits, and it deserves a concrete illustration. Somebody whose old report gave a single Perceptual Reasoning figure of 112 has no way to know whether that reflected balanced ability across both components or a strong reasoning score averaged with a weaker spatial one. Those are meaningfully different people with the same number, and the difference is the kind that shows up in real life as being good at abstract problems while struggling to assemble furniture from a diagram. A current assessment would separate them. That is a genuine reason to consider being reassessed even where nothing is wrong, and it is a better reason than wanting a fresher number.

The Broader Pattern Across Instruments

The WAIS revision is one instance of a change that has run through the whole field, and seeing the pattern makes the specific change easier to understand.

The original Wechsler structure split ability into Verbal and Performance, a division that survived for decades and reflected how the material looked rather than what factor analysis found. Batteries then moved to four-index structures separating verbal knowledge, perceptual work, working memory, and speed. The current generation splits perceptual work again into reasoning and visual processing, producing structures with five or more indices.

Every step of that sequence moved in the same direction: from divisions based on the surface form of the task toward divisions based on the ability being measured. Verbal versus performance is a statement about item presentation. Fluid reasoning versus visual processing is a statement about cognition.

The CHC framework is where that convergence has landed, and it is why instruments from different publishers with different histories now report structurally similar sets of broad abilities. The Woodcock-Johnson battery was built on the framework from the outset, and the Wechsler scales have arrived at a similar structure by successive revision from a different starting point.

The direction of travel suggests where subsequent revisions will go. The framework identifies broad abilities that current clinical batteries still do not report, including auditory processing, long-term retrieval, and reaction time. Whether those enter mainstream batteries depends on whether adding them improves clinical decisions enough to justify the administration time, which is an empirical question rather than a theoretical one.

What Did Not Change

Revisions are described by what moved, which gives a misleading impression of how much is stable. The continuity across editions is substantial and is itself informative about which measurement decisions have held up.

The scaling convention is unchanged. Subtests use a mean of 10 and a standard deviation of 3, indices and the full scale use 100 and 15, and results are reported with percentile ranks and confidence intervals. Anybody comparing reports across editions is at least reading the same units.

Several subtest concepts have survived every revision since the scales were created. Defining words, repeating digit sequences, reproducing a pattern with blocks, and copying symbols against a clock all appear in the earliest Wechsler forms and in the current one. Items have been replaced many times, presentation has changed, and the underlying tasks have not, because they measure what they were built to measure and nothing better has displaced them.

Administration principles are also stable. Difficulty-graded items, start points that depend on age or estimated ability, discontinue rules after consecutive failures, and standardised wording that the examiner may not improvise are all common to both editions. These are the parts of the procedure that make norms meaningful, and revising them would break comparability without gaining anything.

The full scale score itself persists, despite recurring argument about whether a single composite should be the headline figure. It survives because it is the most reliable number a battery produces and because it carries most of the predictive validity the instrument has, which is a claim about evidence rather than about tradition.

Seeing what stayed put helps calibrate what the revision means. A person assessed on either edition sat similar tasks under similar rules and received scores on the same scale. What changed is which population they were compared against and how the resulting information was grouped.

The Drift Behind the Norms, in More Detail

Section 3 asserted that population performance drifts and that this is why norms expire. The phenomenon is worth setting out properly because it is frequently overstated in both directions.

The observation is that scores on cognitive tests rose substantially across the twentieth century in every country with adequate data. A meta-analysis covering many samples and instruments confirmed the effect and estimated its magnitude across the period studied. The rise is large enough that norms from several decades earlier misstate a person's standing noticeably.

What the rise is not is straightforward evidence that people became more intelligent. The gains were concentrated on measures of fluid reasoning and visual processing and were far smaller on measures of crystallised knowledge, which is a strange pattern for a general improvement in ability. Explanations proposed include increased schooling, better nutrition and health, smaller family sizes, and growing familiarity with the abstract, rule-based reasoning that test items require. None of these is settled, and the pattern of gains constrains which explanations are plausible.

More recently the picture has become complicated. Evidence from several countries with long-running population-level data indicates that the rise slowed and in some cohorts reversed. Analysis of Norwegian conscript data found gains turning to declines within families as well as between them, which points toward environmental causes rather than compositional change in who was being tested.

For the practical question of norms, none of the interpretive dispute matters. What matters is that the distribution moves, which means a reference sample has a shelf life regardless of why it moves and in which direction. A revision resets that clock, and the direction of the score change for an individual follows from which way the population moved between the two standardisations.

This is also why the age of the norms, rather than the edition number, is the quantity to ask about. An edition published recently but standardised on a sample collected years earlier, or a local adaptation normed long after the original, both carry older norms than their edition suggests.

Reading a Report From Either Edition

Whichever edition produced it, the same reading order gets the most out of a report and avoids the common errors.

Start with the identifying section. Instrument, edition, date of assessment, date of norms if stated, and who administered it. This establishes what you are looking at and, per section 5, is where a poor report reveals itself by omission.

Read the validity statement before any score. The examiner's judgment about whether the results are a fair estimate governs everything after it. A report noting fatigue, anxiety, or inconsistent effort is telling you how much weight the numbers carry, and reading the scores first inverts the intended order.

Then read the indices with their intervals, not the full scale alone. The composite is the most reliable single number and the least informative about how somebody works, because it averages exactly the pattern that distinguishes one profile from another with the same total.

Compare indices against each other only where the difference exceeds the intervals. A gap smaller than the combined error is not a finding. A large one is worth discussing, and a good report will say how common a difference of that size is in the reference sample rather than presenting it as remarkable by assertion.

Treat subtest-level narrative with suspicion. This holds across editions and is covered in The Subtests Inside an IQ Test. Short measures are unreliable, subtest-specific reliable variance is small, and a battery with many subtests always contains a large-looking gap somewhere.

Read the recommendations last and check they follow. A recommendation should trace back to something in the scores or the observations. Recommendations that would fit any profile are filler, and their presence is a reasonable signal about how much of the rest was written for this specific person.

Where ACIS Sits Relative to This

ACIS reports six domains: verbal comprehension, fluid reasoning, visual spatial, working memory, processing speed, and quantitative reasoning.

The first five correspond to the current clinical structure, with visual spatial and fluid reasoning reported separately for the reasons in section 2. The sixth reports quantitative ability in its own right rather than distributing it across other indices, which follows the framework-aligned tradition rather than the Wechsler one, for the reasons set out in The Subtests Inside an IQ Test.

Two things should be said plainly about the comparison. ACIS is not a Wechsler scale, is not equivalent to one, and produces scores that are not interchangeable with WAIS scores. It is an independent instrument with its own items and its own norms, and any resemblance in structure reflects both being built against the same evidence about ability organisation rather than any relationship between them.

And ACIS is unsupervised. Everything in Professional IQ Test vs Online IQ Test about what supervision provides applies, and no structural similarity to a clinical battery changes it. The domain structure is the part that can be matched. The administration conditions are the part that cannot.

The provenance and construction of the ACIS norms are documented in its technical materials rather than asserted, which is the same standard section 3 applies to any instrument: the question is not whether a test has norms but whether it says where they came from.

FAQ: WAIS-IV and WAIS-5

When was each edition published?

The WAIS-IV in 2008 and the WAIS-5 in 2024, both by Pearson. Both cover ages 16 through 90.

What is the main structural change?

The Perceptual Reasoning Index was split into separate Visual Spatial and Fluid Reasoning indices, matching the change the child scale made in 2014.

Why does splitting that index matter?

Because the two abilities dissociate in real people. A combined index gives the same number to two very different profiles and discards the information that distinguishes them.

Is my WAIS-IV score still valid?

For personal interest, yes, with the caveat that it was measured against a population that has since moved. For formal purposes, recency requirements usually apply independently of edition.

Will my score be lower on the newer edition?

Probably somewhat, because population performance drifts upward and newer norms therefore set a higher bar for the same raw performance.

Is there a conversion formula between editions?

No. Publishers conduct linking studies establishing average relationships, but the instruments differ in norms, subtests, and structure, so no formula applies to an individual.

Why do newer norms produce lower scores?

Because a standard score is a comparison. If the population has improved on what the test measures, the reference group is stronger and identical performance ranks lower against it.

Does the drift affect all abilities equally?

No. It has historically been larger on fluid reasoning and visual processing measures and smaller on crystallised knowledge measures.

How do I tell which edition I took?

The report should name it. Failing that, four indices including Perceptual Reasoning indicates the Fourth Edition, while separate Visual Spatial and Fluid Reasoning indices indicate the current structure.

Why are clinicians still using the older edition?

Kit cost, training time, deliberate continuity for patients being followed over years, and the fact that non-English adaptations lag the original by years.

Is using an older edition bad practice?

Not automatically. Continuity for a follow-up assessment is a legitimate reason. Unfamiliarity with the current edition is not, and it is reasonable to ask which will be used and why.

How often are batteries revised?

Roughly every fifteen to twenty years, driven by norm obsolescence, theoretical developments, item ageing, and administration burden.

Should I keep my old report?

Yes. An earlier assessment is highly valuable as a comparison point, particularly where change over time is the clinical question, and a clinician will account for the edition difference.

Does the edition matter for a society application?

Yes, potentially decisively. Societies publish lists of accepted instruments, older editions are removed as norms age, and thresholds are sensitive to which norms produced the score.

What happened to the Perceptual Reasoning Index?

It has no single counterpart in the current structure. The ability it represented is now reported as two numbers, and if those differ, no single figure summarises them honestly.

Are non-English versions on the same schedule?

No. Adaptations require separate translation, item review, and local standardisation, so they follow the original by years and the edition of record differs by country.

Is the newer edition more reliable?

Not uniformly. Splitting an index means each half rests on fewer subtests, which reduces reliability at that level even as it improves construct clarity. It is a trade rather than a free gain.

What is the CHC framework and why does it keep coming up?

It is the ability taxonomy the research literature converged on, and successive revisions have moved clinical batteries toward reporting the broad abilities it identifies.

Will future editions add more indices?

The framework identifies broad abilities current batteries do not report, including auditory processing and long-term retrieval. Whether they are added depends on whether they improve decisions enough to justify the time.

Is ACIS comparable to a WAIS score?

No. It is an independent instrument with its own items and norms. The structural similarity reflects both being built against the same evidence, not any relationship between them.

Does structural similarity make an online test equivalent?

No. Domain structure can be matched. Supervision, clinical observation, and verified conditions cannot, and those are what separate the two categories.

Best Next Step

The structural change is the interesting one and the norms change is the consequential one. If you hold an older score, it still tells you roughly where you stand, and it is not numerically comparable to a score from the current edition.

If you are arranging an assessment, read The Adult IQ Testing Process and ask which edition will be used. For the ability structure both editions are converging on, read Cognitive Domains and The CHC Model. For a six-domain profile now, take the assessment.

Sources Behind This Page

Edition details come from the publisher. Claims about norm drift and structural change come from the peer reviewed literature and from the technical manuals documenting each revision.

  • Pearson (2024). WAIS-5, Wechsler Adult Intelligence Scale, Fifth Edition. Publication details, age range of 16:0 to 90:11, and the current index structure.
  • Pearson (2008). WAIS-IV Technical and Interpretive Manual. The four-index structure including Perceptual Reasoning, and the standardisation sample behind its norms.
  • Pearson (2014). WISC-V Technical and Interpretive Manual. Documents the separation of perceptual reasoning into Visual Spatial and Fluid Reasoning indices in the child scale a decade earlier.
  • Canivez, G.L., Watkins, M.W. & Dombrowski, S.C. (2016). Factor structure of the Wechsler Intelligence Scale for Children, Fifth Edition. Psychological Assessment. Independent factor analytic evaluation of the revised index structure.
  • McGrew, K.S. (2009). CHC theory and the human cognitive abilities project. Intelligence, 37(1), 1-10. The framework successive revisions have been converging toward.
  • Trahan, L.H., Stuebing, K.K., Fletcher, J.M. & Hiscock, M. (2014). The Flynn effect: a meta-analysis. Psychological Bulletin, 140(5), 1332-1360. Quantifies population-level score drift, which is the mechanism behind norm obsolescence.
  • Bratsberg, B. & Rogeberg, O. (2018). Flynn effect and its reversal are both environmentally caused. Proceedings of the National Academy of Sciences, 115(26), 6674-6678. Norwegian conscript data showing gains turning to declines within families as well as between them.
  • American Educational Research Association, American Psychological Association & National Council on Measurement in Education. Standards for Educational and Psychological Testing. Requirements that norms be current and that the instrument and edition be documented in reporting.
  • Pearson (2008). WAIS-IV Score Report sample. Shows how instrument, edition, index structure, and confidence intervals appear in a report, which is how to identify which edition produced one.
  • Voncken, L., Albers, C.J. & Timmerman, M.E. (2019). Improving confidence intervals for normed test scores. Behavior Research Methods. Open access. Why the edition difference should be weighed against the error already attached to any single score.
  • Calamia, M., Markon, K. & Tranel, D. (2012). Scoring higher the second time around: meta-analyses of practice effects in neuropsychological assessment. The Clinical Neuropsychologist, 26(4), 543-570. Why retest comparisons require accounting for practice as well as for edition.
  • Riverside Insights. Woodcock-Johnson IV. The battery built on the CHC framework from the outset, which the Wechsler scales have converged toward from a different starting point.
  • Buros Center for Testing. Mental Measurements Yearbook. Independent reviews evaluating whether a revision's norms, structure, and evidence base justify replacing the prior edition.
Take the assessment

You get a profile, not a number

ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.

Free trial, no card required. Full report from $15.