An IQ score is not a percentage and not a quotient. It is the output of a chain of conversions that starts with credit on a single item and ends with a range rather than a point. Nine steps sit between those two ends, and each one throws away information the previous step contained.
No step in the calculation divides one number by another to produce an IQ. The arithmetic that people imagine was retired in 1960.
0 Quick Answer
An IQ score is calculated by converting item credit into a raw total, converting that total into a scaled score against people in the test taker's own age band, summing the scaled scores, converting the sum into an index on a mean of 100 and a standard deviation of 15, combining the indexes into a composite, and finally attaching a confidence interval built from the standard error of measurement. Nothing in that chain divides one age by another. The formula most people picture, mental age divided by chronological age multiplied by 100, has not produced a score on a major published test since the 1960 revision of the Stanford-Binet, and David Wechsler had already replaced it on his own scale in 1939.
The chain matters because each link changes what the number means. A raw total is a count of credit and is meaningless on its own. A scaled score is a rank against an age group. An index is that rank re-expressed on a familiar metric. A composite averages several indexes and discards the shape of the profile. The confidence interval at the end is the only part of the report that tells you how much of the observed score is signal.
What follows walks each step in order with real numbers from documents you can open. Pearson's published WAIS-IV sample score report supplies a complete worked case from raw scores to a 95 percent interval. The ACIS technical manual supplies the reliability coefficients that set the width of those intervals on this instrument. The point of the exercise is that a score with no norm sample behind it and no interval around it is not a smaller version of an IQ. It is a different kind of object.
Nine steps
Conversions between a single item response and a reported Full Scale composite with an interval around it.
13 age bands
Separate normative groups spanning ages 16:0 to 90:11 in the WAIS-IV standardization, per Pearson's published brochure.
2.40 to 5.22
The range of standard errors of measurement across the six ACIS primary indexes, from VCI to PSI, in the published technical manual.
Every modern cognitive battery runs the same sequence, and the differences between instruments are differences in the detail of each step rather than in the shape of the chain. Publishers vary in how many subtests feed an index, in whether they smooth the norm tables by age or by continuous regression, and in whether they report a 90 percent or a 95 percent interval. None of them skips a link.
Reading the chain in order is also the fastest way to see where a score can go wrong. A misapplied discontinue rule corrupts step three. A norm table built on the wrong age group corrupts step four. A missing confidence interval leaves step nine undone, which is the most common failure in the free end of the online market.
Step
What happens
What the number means after it
1. Item credit
Each response is scored against a key or a graded rubric.
Points on one item, comparable to nothing yet.
2. Raw total
Item credit is summed across the subtest.
A count. Interpretable only against a reference group.
3. Stopping rules applied
The subtest ends early on a discontinue criterion, and items never presented are entered as zero.
A raw total that includes assumed failures.
4. Age band lookup
The raw total is located in the norm table for the test taker's own age group.
A rank within an age cohort.
5. Scaled score
That rank is expressed on a metric with mean 10 and standard deviation 3.
Distance from the age group average, in thirds of a standard deviation.
6. Sum of scaled scores
Scaled scores for the subtests in one index are added.
A composite raw quantity on the scaled score metric.
7. Index conversion
The sum is converted through a published table onto mean 100, standard deviation 15.
Standing on one broad ability, in IQ units.
8. Composite
Indexes are combined into a Full Scale score.
Overall standing, with the profile shape averaged out.
9. Confidence interval
The standard error of measurement is multiplied by a normal deviate and applied either side of the score.
A band of plausible true standing rather than a point.
The table is the whole article in compressed form. The sections below take each row and show the arithmetic, the conventions that govern it, and the specific ways it can mislead. Where a step is already covered in depth elsewhere on this site, it is linked rather than repeated: the Full Scale composite page covers what step eight produces and how to read it, and the page on what IQ measures covers what the whole chain is built to capture.
2 Step One and Two: Credit on an Item, and the Raw Total
The first number in the chain is credit on a single item, and it is the only number in the entire process that describes what a person actually did. Everything after it is a comparison. On a multiple choice reasoning item the credit is binary: the response matches the key or it does not. On a produced response the credit is graded by rubric, which is why the examiner's judgment enters the calculation long before any table is opened.
Speed changes the shape of item credit. On ACIS Symbol Search the examinee sees two target symbols and scans an array to decide whether either appears, with at most one match present per item and incorrect responses penalized across 80 trials in 120 seconds. Credit there is not simply a count of correct responses, because guessing under time pressure would otherwise be rewarded. On a power subtest with a generous time allowance, an item is worth what the rubric says it is worth and nothing about the clock enters the score.
Summing item credit gives the raw total, and the raw total is the point at which most casual reasoning about IQ breaks down. Twenty-two correct out of thirty is a fact about a person and a test form. It is not a percentage of intelligence, it is not convertible into a percentile without a table, and it does not transfer between forms of the same subtest. When ACIS replaced a broken item in Logic Grid in September 2026, every raw total on that subtest gathered before the replacement stopped being comparable to totals gathered after it, which is why the instrument records form epochs rather than treating a subtest as a fixed object.
Where a raw total stops and a score beginsA raw total carries no information about difficulty. Two people can reach the same total by passing entirely different items, and on an adaptively ordered subtest the person who reached further into the form has demonstrated more. The norm table is what converts a count into standing, which is also why a public test with no conversion table cannot honestly report an IQ at all.
This is the step where the difference between a battery and a quiz is already visible. A battery defines the rubric, fixes the time allowance, records which items were presented, and publishes what each of those decisions does to the total. A quiz counts correct answers. The distinction is developed on the comparison of free and validated instruments and again on the page on whether online tests are accurate.
3 Item Scoring Rules and Partial Credit
Not every correct answer is worth the same, and the graded rubric is doing psychometric work rather than being generous. On produced response verbal tasks the classic convention awards 2 points for a response that captures the underlying category, 1 point for a response that is true but concrete, and 0 for a response that is wrong or unclassifiable. The gradation exists because the difference between those two kinds of correct answer is the construct.
ACIS Similarities uses that structure directly. The examinee is given two words naming objects or concepts and explains how the two are alike, and the response is produced rather than selected. A superordinate category that captures both members earns full credit while a concrete or functional resemblance earns partial credit, across 31 pairs with 90 seconds per item and 0 to 2 points available on each. The published technical manual states the reason plainly: the scoring distinction is what separates induction from lexical knowledge alone. Collapse the rubric to right or wrong and the subtest becomes a vocabulary check.
Partial credit also changes the arithmetic of the raw total in a way that matters at the top of the range. A 31 item subtest scored 0 to 2 has a maximum of 62 points rather than 31, so the raw scale has twice the resolution and can separate people who passed the same items with different quality of response. That extra resolution is one of the reasons produced response verbal subtests carry high general factor loadings. In the ACIS published figures, Similarities loads .865 on g and Vocabulary .848, against .752 for Symbol Search, which is scored as a rapid binary decision.
0 to 2 points
Credit available per item on ACIS Similarities, across 31 pairs with 90 seconds allowed for each.
.865 against .752
Published ACIS g loadings for Similarities and Symbol Search, the graded verbal task and the speeded binary one.
80 trials in 120 seconds
The Symbol Search administration, where incorrect responses are penalized so that speed alone cannot buy credit.
The consequence for anyone reading a report is that raw totals are not comparable across subtests even within the same battery. A raw 45 on a two point rubric and a raw 45 on a one point rubric describe entirely different performances. That is one reason reports print scaled scores rather than raw totals, and it is why the page on subtest formats is worth reading alongside any score profile. The individual ACIS tasks are documented on their own pages, including Similarities and Symbol Search.
4 The Discontinue Rule, and Why Unreached Items Score Zero
Most subtests stop before the end, and every item the examinee never saw is entered into the raw total as a zero. That convention is old, it is universal across the Wechsler tradition, and it is the single least intuitive part of the calculation. It is also the part that most often makes a legitimate short administration look to an untrained reader like a broken one.
The logic is that items are ordered by difficulty. If a person fails a run of consecutive items at one difficulty level, the probability that they would pass a harder item further down the form is treated as low enough to be worth nothing against the cost of asking. Pearson shortened these rules deliberately when the WAIS-IV replaced the WAIS-III. The published WAIS-IV brochure documents the changes item by item: Block Design moved from three consecutive scores of zero to two, Similarities, Matrix Reasoning, Arithmetic and Comprehension moved from four to three, Vocabulary and Information moved from six to three, and Picture Completion moved from five to four. The core battery dropped from 13 subtests to 10 and administration time fell by close to 15 percent.
ACIS applies the same principle with a uniform rule: subtests discontinue after three consecutive errors or timeouts, subject to time caps. Not every subtest carries one. Spatial Comprehension presents all 35 items within a 30 minute global allowance with no discontinue rule at all, because the items are not strictly ordered by a single difficulty dimension. Digit Span ends a part after two consecutive errors, matching the span task convention. A short administration on a subtest that discontinues is a valid administration, not a failed one.
The published limit of this conventionScoring unseen responses as wrong is not free. Von Davier, Cho and Pan reported in Psychometrika in 2019, in volume 84 at pages 147 to 163, that across the scoring rules they compared, ability estimates are biased most when the not observed responses are scored as wrong, and that this is the scoring used operationally. They list the Stanford-Binet Fifth Edition, the KABC-II, the Kaufman Adolescent and Adult Intelligence Test and the Universal Nonverbal Intelligence Test as examples of tests using the rule. The convention is defensible for administration time and it is known to cost accuracy, and both facts belong in the same sentence.
Two practical readings follow. The first is that a raw total is partly an assumption, not purely an observation, and the size of that assumption grows the earlier a subtest stops. The second is that comparing raw totals between a person who discontinued at item 14 and a person who reached item 28 compares two quantities built from different amounts of evidence. Both convert to a scaled score through the same table, which is exactly why the table, and not the raw total, is the object that carries meaning. Time limits on individual subtests are covered separately on the page on test time limits, and the Digit Span page sets out how a span task terminates.
5 Step Four: Your Raw Total Against Your Own Age Band
The raw total is looked up in a table built from people of the test taker's own age, and this is the step that makes an identical performance worth different scores at different ages. It is not a correction applied afterward. It is the definition of the score. A scaled score answers the question of where this performance sits among people born around the same time, and nothing else.
The WAIS-IV standardization sample makes the machinery concrete. Pearson's published brochure states that the test was standardized on 2,200 individuals divided into 13 age bands spanning ages 16:0 to 90:11, stratified against United States population figures on age, sex, education level, race or ethnicity, and geographic region. Thirteen bands across 75 years means the reference group for a 22 year old and the reference group for an 82 year old are different sets of people, and the conversion tables differ accordingly.
Pearson's public WAIS-IV sample score report shows the size of the effect, because it prints both conversions side by side for the same person. The sample examinee is a man aged 70 years and 7 months. His subtest summary lists each raw score twice: once as a scaled score against his own age band, and once as a reference group scaled score against the young adult anchor group.
Subtest
Raw score
Scaled score for his age
Reference group scaled score
Gap
Block Design
42
13
9
4 points higher against his own cohort
Visual Puzzles
13
12
8
4 points higher against his own cohort
Coding
65
13
9
4 points higher against his own cohort
Symbol Search
36
15
11
4 points higher against his own cohort
Matrix Reasoning
23
17
14
3 points higher against his own cohort
Vocabulary
55
17
18
1 point lower against his own cohort
Information
24
16
17
1 point lower against his own cohort
Similarities
36
19
19
No difference
Read the last three rows against the first four. On the speeded and visual spatial tasks the same raw performance is worth four extra scaled score points when the comparison group is his own age, because those abilities decline across the adult range and his cohort has declined with him. On the acquired knowledge tasks the direction reverses, because vocabulary and general information keep accumulating and his cohort has accumulated too. On Similarities the two conversions agree exactly.
Four scaled score points is not a rounding artifact. The scaled score metric has a standard deviation of 3, so four points is roughly 1.3 standard deviations on that subtest, and a shift of that size across several subtests moves an index by well over ten IQ points. The same raw performance, scored honestly under two defensible reference frames, produces two genuinely different scores. Which one is correct depends entirely on the question being asked, and the age related pattern behind it is covered on the page on whether IQ changes with age and on the average IQ by age page.
6 Step Five: What a Scaled Score of 13 Actually Says
The scaled score metric places the average of the age group at 10 and one standard deviation at 3, so every subtest in a battery becomes comparable to every other one. That is the entire purpose of the transformation. Raw totals from a 31 item verbal task and an 80 trial speeded task cannot be compared. Scaled scores of 13 and 13 can.
Reading the metric takes one piece of arithmetic. A scaled score of 13 is one standard deviation above the age group mean, which places the person near the 84th percentile of that group. A scaled score of 7 is one standard deviation below, near the 16th percentile. A score of 16 is two standard deviations above, near the 98th. The steps are one third of a standard deviation each, which is why a two point difference between subtests is a smaller thing than it looks and a six point difference is a large one.
ACIS uses the same metric for the same reason, with subtest scaled scores on a mean of 10 and a standard deviation of 3. Each subtest also carries its own standard error on that metric, computed as three multiplied by the square root of one minus the reliability coefficient for that task. In the published manual, non speeded subtests use McDonald's omega estimates and the working memory and speeded subtests use test retest reliability, because temporal stability is the cleaner quality estimate for a span or a speeded efficiency task.
10 is the middle. Half the age reference group scores at or below it, by construction.
Each point is a third of a standard deviation. Moving from 10 to 13 covers the same distance as moving from 100 to 115 on the IQ metric.
The useful range runs roughly 1 to 19. Beyond that, the norm table has too few people to anchor a conversion.
Subtest scores carry their own error. In the Pearson sample report the printed subtest standard errors run from 0.67 on Vocabulary to 1.37 on Digit Span Backward, so a one point difference between two subtests is inside the noise.
The last point is the one clinicians act on and casual readers skip. A profile that shows 13, 12, 14, 12 across four subtests is a flat profile with noise on top, not evidence of four different abilities. Pearson's sample report formalizes this by printing critical values: on that examinee, the Digit Span against Arithmetic difference of four scaled score points exceeded the .05 critical value of 2.57 and was flagged as significant, while the Symbol Search against Coding difference of two points did not exceed its critical value of 3.41 and was not. Two apparently similar gaps, two different verdicts, both from the same table.
7 Step Six: Summing Scaled Scores Into an Index
An index is built by adding the scaled scores of its subtests and converting the sum, not by averaging the subtests and rescaling the average. The distinction sounds pedantic and is not. Summing preserves the number of indicators, and the conversion table for a sum of five subtests is a different table from the conversion table for a sum of two, because a five subtest index is measured more precisely.
The Pearson sample report shows the sums explicitly. Verbal Comprehension summed to 52 across three subtests. Perceptual Reasoning summed to 42 across three. Working Memory summed to 32 across two. Processing Speed summed to 28 across two. The Full Scale sum of 154 is the total of the ten core subtests. Each of those sums enters a different published table.
Precision rises with the number of indicators, and the ACIS figures show the pattern cleanly. The published standard errors of measurement run 2.40 for VCI on five subtests, 2.48 for FRI on five, 3.39 for VSI on three, 4.02 for QRI on two, 4.12 for WMI on three, and 5.22 for PSI on two. A two subtest index carries roughly twice the measurement error of a five subtest index, so the same nominal ten point gap between two people means something quite different on VCI than on PSI.
ACIS index
Indicators
Omega
SEM
g loading
VCI, Verbal Comprehension
5
.9745
2.40
.864
FRI, Fluid Reasoning
5
.9727
2.48
.922
VSI, Visual Spatial
3
.9488
3.39
.906
QRI, Quantitative Reasoning
2
.9283
4.02
.882
WMI, Working Memory
3
.9247
4.12
.788
PSI, Processing Speed
2
.8790
5.22
.648
The g loading column explains why indexes are not interchangeable inputs. PSI at .648 and WMI at .788 carry substantially less of the general factor than FRI at .922, so an index built to summarize reasoning behaves differently from one built to summarize efficiency. That asymmetry is the reason a separate reasoning and knowledge composite exists at all, which is the subject of the General Ability Index page. Which subtests feed which index on this instrument is set out on the cognitive domains page, and the reasoning domain in particular on the fluid intelligence page.
8 Step Seven: The Sum Becomes an Index on Mean 100
The conversion from a sum of scaled scores to an index is a table lookup, and the table is the most valuable object a test publisher owns. It is not a formula anyone can apply from outside. Building it requires the distribution of sums observed in the reference sample, smoothing across age, and decisions about where to truncate at the extremes, all of which are documented in a technical manual or not documented at all.
The output metric is the familiar one: mean 100, standard deviation 15. An index of 115 is one standard deviation above the reference mean and sits near the 84th percentile. An index of 130 is two standard deviations above and near the 98th. The reasons this particular metric was chosen, and why some instruments use a standard deviation of 16 or 24 instead, are set out on the page on the 15 point standard deviation.
Two properties of the conversion catch people out. The first is that it is nonlinear in percentile terms. Moving from 100 to 105 crosses about 13 percentile points; moving from 130 to 135 crosses about one. The second is that the mapping is not one to one across the whole range. Near the middle of the distribution several adjacent sums can map to the same index because so many people are packed there, while at the tails a single point of sum can move the index by several points because so few people are.
The Pearson sample report shows the whole step in one line. A Verbal Comprehension sum of 52 became an index of 145 at the 99.9th percentile. A Processing Speed sum of 28 became an index of 122 at the 93rd. The two sums are 24 points apart on the scaled score metric across a different number of subtests, and the tables handle that difference invisibly, which is precisely why the tables have to exist.
What is missing when the table is missingA public item set with no published conversion table cannot produce an index no matter how carefully the items were written. The ICAR battery has been item analyzed and validated in the literature since 2014 and still has no norm table, so a raw total on it stays a raw total. Any site that prints an IQ number from an unnormed item set has invented the last two steps of this chain. Converting a real score into a rank is a separate operation, available on the percentile calculator.
The practical test for a reader is simple. Ask what reference group the index was computed against, how many people were in it, when it was collected, and whether the manual states it. If those four answers are not available, step seven did not happen in any meaningful sense. The norming article covers how a defensible reference frame is built, and the score versus percentile page covers what the resulting rank does and does not say.
9 Step Eight: From Indexes to the Full Scale Composite
The Full Scale score is a composite of the indexes, and the information it discards is exactly the information the profile contained. That trade is deliberate. A single figure is what a composite is for, and the six index scores stay on the report precisely because the composite cannot carry them.
ACIS builds its Full Scale composite from all 20 subtests across the six CHC domains. The published reliability figures for that composite are stratified alpha of .9886, composite omega total of .9886, hierarchical omega of .9164, and a general factor loading of .958, with a standard error of measurement of about 1.60 IQ points. Those coefficients come from a technical analysis set of 2,750 complete records, inside an adult reference frame of 3,243 English speaking records aged 16 to 90. The full derivation is in the technical manual.
Two features of composite construction are worth stating because they are counterintuitive. The first is that a composite is more reliable than any of its parts. Twenty indicators average out the idiosyncratic error attached to each one, which is how the Full Scale standard error of 1.60 ends up smaller than the 2.40 on VCI and the 5.22 on PSI. The second is that a composite is not more valid for every question. Averaging a Verbal Comprehension index of 145 with a Processing Speed index of 122, as in the Pearson sample case, produces a Full Scale figure that describes neither.
The disagreement between indexes is measurable rather than a matter of impression. The ACIS published intercorrelations show VCI and FRI at .801 and FRI and VSI at .813, while WMI and PSI correlate .475 and VCI and PSI correlate .577. Two indexes correlating .475 are carrying substantially different information, so a composite that includes both is averaging across a real seam. That is the seam that the General Ability Index page examines from the composition side, and it is the reason a serious report prints six index scores next to the composite rather than the composite alone.
10 Step Nine: The Confidence Interval, With the Arithmetic
The last step converts a point into a band, and it takes two numbers: the standard error of measurement and a multiplier from the normal distribution. Nothing about it is difficult, which makes its absence from most consumer score reports harder to excuse rather than easier.
The standard error of measurement is computed from the reliability of the score and the standard deviation of the scale. The ACIS technical manual states the formula directly: SEM equals the standard deviation of the scale multiplied by the square root of one minus the reliability coefficient. On the IQ metric the standard deviation is 15, so a reliability of .960 produces a standard error of about 3 IQ points and a reliability of .860 produces about 5.6.
The interval is then the observed score plus or minus the standard error multiplied by the normal deviate for the confidence level. For a 68 percent band the multiplier is 1. For a 90 percent band it is 1.645. For a 95 percent band it is 1.96. Two worked examples on the ACIS index standard errors show how much the width depends on which score is being reported.
VCI of 122
SEM 2.40. Multiply by 1.96 to get 4.7, giving a 95 percent interval of roughly 117 to 127, about 9 points wide.
PSI of 122
SEM 5.22. Multiply by 1.96 to get 10.2, giving a 95 percent interval of roughly 112 to 132, about 20 points wide.
Full Scale of 122
SEM 1.60. Multiply by 1.96 to get 3.1, giving a 95 percent interval of roughly 119 to 125, about 6 points wide.
The same observed number, 122, carries three different amounts of certainty depending on how many indicators produced it. That is the honest content of a score report and it is why a two subtest index should never be quoted with the same confidence as a twenty subtest composite. Pearson prints the same structure on the sample report: a Full Scale IQ of 139 with a 95 percent interval of 134 to 142, a Verbal Comprehension index of 145 with an interval of 138 to 149, and a Processing Speed index of 122 with an interval of 111 to 128, all noted as based on the overall average standard errors.
The interval published is narrower than the real uncertaintyVoncken, Albers and Timmerman reported in Behavior Research Methods in 2019, in volume 51 at pages 826 to 839, that the confidence intervals test publishers provide reflect only the unreliability of the test, and that the uncertainty caused by sampling variability during norming is ignored. Their proposed method incorporates the second source. Until publishers adopt something like it, every printed interval on every instrument, including this one, understates the total uncertainty by an amount that depends on how large the norm sample was.
The rule that follows is short. Treat a reported score as the center of a band, treat the band as narrower than the truth, and never compare two scores that sit inside each other's intervals as if the difference were real. How rare a given band is, once you have one, is on the rarity calculator, and the general treatment of measurement error is on the reliability and validity page.
11 The Formula Nobody Uses Anymore
Mental age divided by chronological age, multiplied by 100, has not produced a score on a major published intelligence test since 1960, and the man who built the dominant adult scale abandoned it in 1939. The formula survives in general circulation because it is short and can be done in your head. Neither property is evidence.
Wechsler dismantled it in print in The Measurement of Adult Intelligence, published in 1939, the year the Wechsler-Bellevue appeared. His argument runs across a few pages of chapter three and it is worth quoting because it is more direct than most modern restatements. Mean scores on almost every intelligence scale, he wrote, cease to increase with age at some point that varies by task: digit span stops improving at 14 and vocabulary at about 22. Beyond an age varying from 15 to 22, "all scores of mental ability, far from remaining constant, start to fall off."
What follows from that observation is the part people miss. If mental age stops growing but chronological age keeps growing, the quotient falls for reasons that have nothing to do with the person. Wechsler's response was to reject the fixed denominator entirely. "To calculate an I.Q. for a man of 60 by dividing his M.A. score by 15," he wrote, "is as incorrect as to obtain an I.Q. for a boy of 12 by dividing his M.A. score by 15." And then the conclusion: "As soon as the age factor is discarded, the I.Q. ceases to be an I.Q. in the original sense of the term," so that all adult figures obtained by dividing a mental age score by any fixed denominator "are not I.Q.'s at all."
The Stanford-Binet followed twenty-one years later. The 1960 third revision, published by Houghton Mifflin as Terman and Merrill's manual for the third revision form L-M, carried revised IQ tables by Samuel R. Pinneau on its title page, and those tables replaced the quotient with a deviation score referenced to the test taker's own age group. After 1960 the arithmetic that had defined the term for four decades was no longer performed by either of the two instruments that had made it famous. The formula itself, and the history of how Stern and Terman arrived at it, is set out on the mental age page.
The ratio has since been tested empirically rather than merely argued against, and it failed. Ostrolenk and Courchesne reported in Acta Psychologica in 2023, in volume 240, article 104054, on 16,751 autistic participants aged 2 to 18 drawn from four databases and tested on the MSEL, the DAS-II early years and school age forms, the WISC-IV and the WISC-V. They computed ratio IQs for every participant and compared them with the standard Full Scale scores. Participants at the extremes of the distribution showed the largest discrepancies, and age predicted the direction: the ratio ran higher than the Full Scale for younger participants and lower for older ones. Their conclusion was that the practice is questionable precisely for the individuals it is most often used on.
12 Why the Same Person Gets Different Numbers
Two well built tests administered to the same person weeks apart do not return the same figure, and the size of the gap is documented rather than speculative. Nothing in the chain above guarantees agreement between instruments, because each instrument runs the chain against its own reference sample, its own subtest mix and its own conversion tables.
The clearest recent evidence comes from the counterbalanced studies Pearson ran for the WAIS-5. Winter, Trudel and Kaufman summarized them in the Journal of Intelligence in 2024, in volume 12, article 118. In the first study 186 adolescents and adults aged 16 to 90, mean age 47.8, took the WAIS-IV and the WAIS-5 in counterbalanced order with a mean interval of 28.2 days. In the second, 98 sixteen year olds, mean age 16.4, took the WISC-V and the WAIS-5 with a mean interval of 26.9 days. Full Scale IQs correlated .92 between the two Wechsler adult versions and .87 between the fifth editions of the child and adult scales.
The mean differences were small: 1.9 points between the WAIS-IV and the WAIS-5, and 1.2 points between the WISC-V and the WAIS-5. Winter and colleagues note the result was unexpected, describing it as a reduction in the Flynn effect from an expected 3 IQ points to 1.2. A small mean difference and a correlation of .92 are not the same claim as agreement for an individual, though. With two scales sharing a standard deviation of 15, a correlation of .92 implies a standard deviation of about six points for the difference between them, which is arithmetic from the published correlation rather than a figure Pearson printed. Under that arithmetic a gap of six points or more between two administrations is ordinary rather than alarming.
Source of disagreement
What causes it
Typical size
Measurement error
Item sampling, attention, fatigue, and day to day variation.
Set by the SEM. Roughly 3 points either side on a strong composite at 95 percent confidence.
Different norm samples
Each instrument ranks against its own reference group, collected in a different year.
Reported as 1.9 points on average between WAIS-IV and WAIS-5 by Winter and colleagues, 2024.
Different subtest mix
Batteries weight verbal, spatial, memory and speed differently.
Largest for uneven profiles. Grows with the gap between a person's own indexes.
Age band choice
Age adjusted against reference group conversion of the same raw score.
Up to 4 scaled score points per subtest in Pearson's own sample report.
Practice effect
Re-exposure to item formats within a short interval.
Direction is upward. Pearson's counterbalancing exists to remove it from the group comparison.
None of this makes a score worthless. It makes the score a measurement with a known error term, which is what every measurement in every field is. The failure mode is not disagreement between good instruments. It is a number reported with no reference group, no reliability coefficient and no interval, because there is then nothing to disagree with. That number is not an IQ that happens to be less precise. It is a raw total wearing the vocabulary of one. The market that produces those is surveyed on the online test comparison, and what a proctored administration buys instead is on the professional versus online page.
Every figure above traces to a document you can open, and the ones that cannot be opened are labeled as such in the sentence that uses them. The list below is the working bibliography rather than a decorative reference section.
Pearson. WAIS-IV Score Report, sample report. The worked case used throughout this page: composite sums, index scores, percentile ranks, 95 percent confidence intervals, age adjusted and reference group scaled scores, subtest standard errors and discrepancy critical values. Published PDF, copyright 2008, sample dated 2016.
Pearson. WAIS-IV brochure. Standardization sample of 2,200 individuals across 13 age bands from 16:0 to 90:11, stratification variables, the table of discontinue rule changes from WAIS-III to WAIS-IV, and the reduction of the core battery from 13 subtests to 10. Published PDF.
Von Davier, M., Cho, Y., and Pan, T. (2019). Effects of Discontinue Rules on Psychometric Properties of Test Scores. Psychometrika, 84(1), 147 to 163. Source for the finding that ability estimates are biased most when unobserved responses are scored as wrong. DOI.
Wechsler, D. (1939). The Measurement of Adult Intelligence. Baltimore: Williams and Wilkins. Source for every quotation about the failure of the mental age quotient in adulthood. Full text scan.
Terman, L. M., and Merrill, M. A. (1960). Stanford-Binet Intelligence Scale: Manual for the Third Revision Form L-M, with revised IQ tables by Samuel R. Pinneau. Boston: Houghton Mifflin. The edition that replaced the quotient with deviation tables. Archive record, borrowing access only.
Ostrolenk, A., and Courchesne, V. (2023). Examining the validity of the use of ratio IQs in psychological assessments. Acta Psychologica, 240, 104054. N of 16,751 across five instruments. DOI.
Winter, E. L., Trudel, S. M., and Kaufman, A. S. (2024). Wait, Where's the Flynn Effect on the WAIS-5? Journal of Intelligence, 12(11), 118. Counterbalanced samples of 186 and 98, correlations of .92 and .87, mean differences of 1.9 and 1.2 points. Open access.
Voncken, L., Albers, C. J., and Timmerman, M. E. (2019). Improving confidence intervals for normed test scores: Include uncertainty due to sampling variability. Behavior Research Methods, 51(2), 826 to 839. Source for the claim that published intervals understate total uncertainty. Open access.
ACIS technical manual. Every ACIS figure on this page: index omegas and standard errors, subtest g loadings, index intercorrelations, the Full Scale reliability model, the standard error formula, the reference frame of 3,243 records and the technical analysis set of 2,750. Published on this site.
What this page does not haveNo conversion table appears here, from either ACIS or any other publisher. Norm tables are the part of an instrument that has to stay closed, and publishing one would let a reader convert a raw total without an administration. Every reliability coefficient quoted describes consistency within the documented ACIS adult reference frame.
The professional framework that governs all of it is written down. The Standards for Educational and Psychological Testing, published jointly in 2014 by the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education, require that a score be reported with evidence for the specific interpretation being made, that measurement error accompany the score, and that the composition and limits of the norm sample be disclosed. The International Test Commission guidelines on test use make parallel demands of the person administering the instrument. APA standards on test use add the requirement that the person reporting a score state what it supports and what it does not, which is the reason this page ends with a bibliography instead of a number.
One boundary belongs here rather than in a footer. ACIS is a self-administered online assessment. It is not a clinical instrument, it is not diagnostic, and it is not appropriate for hiring decisions, accommodation requests, or high IQ society admission. What it does produce is the full chain described above, run against a documented adult frame, with the reliability coefficients and standard errors published rather than asserted.
14 Frequently Asked Questions
How is IQ calculated in simple terms?
Item credit is summed into a raw total, the raw total is ranked against people in your own age band to give a scaled score on a mean of 10, the scaled scores are summed and converted onto a mean of 100 with a standard deviation of 15, and the resulting score is reported with a confidence interval. No division of one age by another occurs anywhere in the sequence.
What is the IQ calculation formula?
There is no single formula, because the conversion from a sum of scaled scores to an index is a published lookup table rather than an equation. The only genuine formula in the chain is the one for measurement error: the standard error equals the standard deviation of the scale multiplied by the square root of one minus the reliability coefficient.
Is IQ still mental age divided by chronological age?
No. Wechsler rejected that arithmetic in 1939 and the Stanford-Binet replaced it in its 1960 third revision, which carried Samuel R. Pinneau's revised deviation tables. Both of the instruments that made the quotient famous stopped computing it more than sixty years ago.
Why does the ratio formula break down for adults?
Because mental age stops rising while chronological age does not. Wechsler documented that mean scores on digit span stop increasing at 14 and on vocabulary at about 22, so dividing by an ever larger denominator drives an adult's quotient downward for no reason connected to the person.
Has anyone tested whether ratio IQ actually works?
Yes. Ostrolenk and Courchesne reported in Acta Psychologica in 2023 on 16,751 autistic participants aged 2 to 18 across the MSEL, both DAS-II forms, the WISC-IV and the WISC-V. Discrepancies between ratio and Full Scale scores were largest at the extremes of the distribution, which is exactly where the ratio is most often used.
How are IQ tests scored on individual items?
Selected response items are scored against a key. Produced response items are scored by rubric, most commonly on a 0, 1, or 2 scale where a superordinate category earns full credit and a concrete resemblance earns partial credit. ACIS Similarities uses that structure across 31 pairs with 90 seconds per item.
Why does partial credit exist at all?
Because the difference between two correct answers can be the construct. On a similarities task a categorical answer and a functional answer are both defensible, and only the graded rubric separates induction from vocabulary recall. Collapsing the rubric to right or wrong would halve the resolution of the raw scale.
What is a discontinue rule?
A stopping criterion that ends a subtest once a set number of consecutive items are failed, on the reasoning that harder items further down an ordered form would also be failed. ACIS discontinues after three consecutive errors or timeouts, and Pearson shortened the WAIS-IV rules relative to the WAIS-III, moving Vocabulary and Information from six consecutive zeros to three.
Why are unreached items scored as zero?
It is the Wechsler convention, and it keeps the raw total on the same scale for everyone regardless of where they stopped. Von Davier, Cho and Pan reported in Psychometrika in 2019 that this specific choice biases ability estimates more than the alternatives they compared, so the convention is a documented trade of accuracy for administration time.
Does stopping early mean the test failed?
No. A short administration on a subtest with a discontinue rule is a valid administration by design. Not every subtest carries one either: ACIS Spatial Comprehension presents all 35 items within its global time allowance because the items are not ordered on a single difficulty dimension.
Why does the same raw score give different scores at different ages?
Because the conversion table is age specific. In Pearson's published WAIS-IV sample report, a 70 year old man's raw Block Design score of 42 converts to a scaled score of 13 against his own age band and 9 against the young adult reference group, a gap of four points on a metric whose standard deviation is 3.
Does the age adjustment always work in the same direction?
No, and the same report shows the reversal. On Vocabulary the same examinee scores 17 against his own age band and 18 against the reference group, because acquired knowledge keeps accumulating while speed and visual spatial performance decline. On Similarities the two conversions agreed exactly at 19.
What does a scaled score of 13 mean?
One standard deviation above the mean of the age reference group, since the scaled score metric has a mean of 10 and a standard deviation of 3. That places the performance near the 84th percentile of that group, the same relative position that an index of 115 marks on the mean 100 metric.
How do you convert a scaled score to an IQ?
Not directly, and not with arithmetic. Scaled scores for the subtests in an index are summed, and the sum is looked up in a published table specific to that index and that number of subtests. A two subtest sum and a five subtest sum use entirely different tables even on the same instrument.
Why do some indexes have wider error than others?
Because precision rises with the number of indicators. The published ACIS standard errors run 2.40 for the five subtest VCI and 5.22 for the two subtest PSI, so the same nominal ten point difference between two people means considerably less on PSI than on VCI.
How is a confidence interval calculated on an IQ score?
Multiply the standard error of measurement by the normal deviate for the confidence level, then apply the result either side of the observed score. For 95 percent the multiplier is 1.96, so an ACIS VCI of 122 with a standard error of 2.40 gives an interval of roughly 117 to 127.
Is the printed confidence interval the whole uncertainty?
No. Voncken, Albers and Timmerman reported in Behavior Research Methods in 2019 that publisher intervals reflect only the unreliability of the test and ignore the sampling variability introduced during norming. Every printed interval on every instrument is therefore narrower than the real uncertainty.
How is IQ calculated for adults specifically?
The same nine steps run, with an adult reference frame. ACIS uses a frame of 3,243 English speaking records aged 16 to 90, and the WAIS-IV standardization divided 2,200 people into 13 age bands from 16:0 to 90:11. The adult case is where the ratio formula fails and the deviation approach was invented.
Why do two tests give me different IQ scores?
Because each instrument runs the chain against its own norm sample, subtest mix, and tables. Winter, Trudel and Kaufman reported in the Journal of Intelligence in 2024 that in Pearson's counterbalanced study of 186 people, WAIS-IV and WAIS-5 Full Scale scores correlated .92 with a mean difference of 1.9 points.
Is a five point difference between two of my scores meaningful?
Usually not, and the report should tell you. Pearson prints critical values for exactly this: on its sample examinee a four point scaled score gap between Digit Span and Arithmetic exceeded the .05 critical value of 2.57 and was flagged, while a two point gap between Symbol Search and Coding did not exceed its value of 3.41 and was not.
What is a score worth with no norm sample and no interval?
It is a raw total using the vocabulary of an IQ. Without a reference group there is nothing to rank against, and without a standard error there is no way to know whether a difference between two such numbers is real. Both are checkable in a technical manual, and a score whose manual does not exist cannot be checked at all.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.