IQ levels: how classification systems name the same score differently
An IQ classification is the word beside a score, such as Average, Superior or Highly Gifted. That word belongs to a particular classification system. It can change when the test edition changes even if the number does not. This guide compares the labels, separates them from percentiles and eligibility rules, and explains why a category boundary should be read with the uncertainty of the measurement.
Classification divides a continuous score into named steps. The boundary between two steps does not establish a sudden change in ability.
0 The short answer
IQ levels are named score bands, and their labels depend on the test, edition and reporting system. On the common mean 100, standard deviation 15 scale, a score of 100 is central and 130 is two standard deviations above the mean. Neither number dictates a universal label. Pearson's WAIS-5 sample calls 111 Above average and 124 Very high; ACIS uses High Average and Superior for the bands containing those scores. Use the score to percentile table to look up a number. Use this page to understand the words beside it. For a decision about one person, read the original test name, reference group, interval and purpose before treating a classification as an identity.
100
The mean on the common deviation IQ scale; it is a reference point, not a percentage correct.
15
The standard deviation on the scale discussed here; other reporting scales require their own conversion.
111
A score labelled Above average in Pearson's official WAIS-5 sample report, illustrating why edition matters.
3
Separate questions: where a score stands, what label a system assigns, and whether an institution accepts the evidence.
1 What does an IQ classification actually classify?
A classification places a reported score into a named range; it does not divide people into naturally separate kinds. Suppose a reporting system assigns one label to scores from 110 through 119 and another to scores from 120 through 129. A score of 119 and a score of 120 receive different words because a rule has been applied. The rule does not demonstrate a sudden change in reasoning, learning or everyday functioning between those two observations. Conversely, 110 and 119 receive the same word despite a nine point difference. Classification compresses information in both directions.
There are three layers to keep separate. The measurement layer concerns the tasks, the scoring procedure and the reference group. The description layer adds a convenient term to a score range. The decision layer applies a rule for a particular purpose, perhaps admission to an organization or consideration for a support service. A test publisher can supply the first two without deciding the third. A school can use a criterion without creating a new cognitive category. An online chart can explain a descriptor without establishing whether any submitted score is valid for a decision.
The familiar number itself is not an answer count. On a deviation scale, performance is compared with an appropriate norm group and expressed in standardized units. The guide to how IQ scores are normed covers that process. Two tests may put their scores on the same mean and standard deviation while differing in the abilities they sample, the age groups they represent, the conditions of administration and the precision of their scores. A shared unit makes comparison intelligible; it does not make the instruments interchangeable.
Consider two hypothetical reports that both print 120. One summarizes a broad individually administered battery. The other summarizes a short set of visual problems completed at home. Before asking whether both deserve Superior, ask what each instrument measured and what evidence supports its interpretation. Relabelling the second score with the first instrument's vocabulary adds no missing validity. The same caution applies to a professionally obtained index score: a visual spatial index and a Full Scale IQ can share a numerical scale while summarizing different sets of performances.
This distinction also explains why a page about labels has a different job from a page about percentiles. A percentile reports relative standing. A label reports a convention. An eligibility result reports a decision under a rule. The word high can be useful conversationally, but a report becomes more informative when it names the test, score type and reference group. A statement such as “the FSIQ was 120 on this edition, with this interval” preserves information that “I am in the Superior level” leaves out. The relevant follow-up is what that result helps the reader understand or decide.
2 Which labels are associated with the WAIS-IV?
The conventional WAIS-IV classification scheme uses seven broad descriptors, but an original report remains the authority for a particular assessment. The familiar sequence is shown below as a reading aid for fourth edition terminology. It should not be silently carried into a fifth edition report, a different Wechsler instrument or a translated edition. The dedicated WAIS-IV guide explains the battery, its score types and the distinction between a Full Scale result and its component indices.
Composite score range
Conventional WAIS-IV descriptor
130 and above
Very Superior
120 to 129
Superior
110 to 119
High Average
90 to 109
Average
80 to 89
Low Average
70 to 79
Borderline
69 and below
Extremely Low
The table is a conventional vocabulary guide, not a newly verified reproduction of the complete manual. Pearson's public WAIS-IV sample directly illustrates Superior at 123 and Very Superior at 139; it does not display all seven endpoints. The original manual and examiner's interpretation govern the actual report. Pearson's WAIS-IV product information identifies the instrument and its intended assessment context. The terms above describe score ranges. They do not independently establish giftedness, a diagnosis or the level of assistance someone needs. In particular, a lower score descriptor is not the clinical name of a person's condition.
A common reading error is to apply the descriptor of the strongest index to the whole battery. If an individual has a verbal index in a high band, that is a statement about that verbal composite. It is not permission to replace the Full Scale score with the same label. The reverse error is to assume that a Full Scale classification describes each index equally. A broad average can coexist with meaningful variation across verbal comprehension, reasoning, working memory and processing speed. Each number should keep its own name.
Another error arises when a report rounds a score or a percentile. The printed composite is an integer, while the underlying scoring and its uncertainty do not become perfectly sharp at that integer. Someone should not infer that moving from the last point of one band to the first point of the next proves a new kind of ability. Nor should they assume that all scores sharing a descriptor are equally close to the boundary. A score of 120 is at an edge in this scheme; 125 lies farther inside it.
For a comparison across reports, copy the score and interval first, then the edition and descriptor. This order makes it possible to notice that a wording change may be administrative rather than substantive. If a later report uses Very high where an earlier one uses Superior, compare the actual measurements before attributing a gain to the word. The article on High Average IQ discusses that particular conventional band, while the next section shows why it is unsafe to assume that High Average is the wording in every current Wechsler report.
3 What do the public WAIS-5 examples actually verify?
Pearson's public WAIS-5 reports verify specific score and descriptor pairs, which is narrower evidence than a complete classification table. The 2024 sample score report labels 111 Above average, 124 Very high and 103 Average. A second sample with demographically referenced scores supplies further examples. These are direct publisher evidence for the wording, not an invitation to reconstruct unseen boundaries from memory.
Score shown in a public WAIS-5 example
Printed qualitative description
What the example establishes
79
Very low
This score and descriptor occur together in the publisher's example
82
Below average
The fifth edition uses this wording for the displayed score
103
Average
The descriptor refers to the reported composite
111
Above average
High Average should not be substituted as a verbatim fifth edition label
124
Very high
Superior should not be substituted as the printed fifth edition wording
The useful conclusion is that the fifth edition language differs from a table copied from the fourth edition. This page does not use those examples to assert every lower or upper endpoint of each band, or to fill the outermost bands that the examples do not display. For an individual assessment, the original report and the applicable manual settle those questions. The broader WAIS-5 overview describes the instrument itself.
It helps to see why this evidence boundary matters. Imagine finding a sample that labels 111 Above average. The observation verifies that example. It does not by itself reveal whether a neighboring reporting policy changes at 110, whether an examiner may select alternative descriptors, or how a localized edition presents the same score. Treating a sample as a complete manual would turn a small piece of reliable evidence into a larger claim the source did not supply. Accurate comparison can be useful without pretending that every missing detail is known.
The same principle applies to intervals. Pearson's first sample explicitly identifies its confidence interval method as using the Standard Error of Estimation. A generic worksheet that multiplies an SEM by 1.96 therefore should not be described as reproducing that report's exact calculation. The broad lesson is uncertainty around scores; the particular interval method belongs to the instrument. Copying a word from one edition and an interval formula from another creates a report that neither publisher issued.
For a reader holding an older result, the changed terminology is usually a reason to be precise, not a reason to seek another test just for a different adjective. First identify what practical question remains unanswered. If the question concerns a current clinical or educational decision, a qualified professional can determine whether the existing assessment remains suitable. If it concerns the language of the report, a clarification of the terms may answer it without any new measurement. The descriptor is the summary at the end of an interpretation process; it is not a replacement for that process.
4 How does the Stanford-Binet illustrate a different system?
Stanford-Binet terminology should be read from Stanford-Binet documentation, rather than assumed from a Wechsler chart. The publisher's SB5 assessment service bulletin reproduces a table containing the following middle bands. This is a verified extract of that table, not a claim to reproduce every classification in the complete manual. The Stanford-Binet 5 guide covers its structure and the importance of identifying the exact edition.
SB5 range in the publisher's table
Label in that table
70 to 79
Borderline Delayed
80 to 89
Low Average
90 to 109
Average
110 to 119
High Average
120 to 129
Superior
The source presents proposed instructional approaches alongside ability bands. Its recommendations are not individual prescriptions, and the existence of a table does not show that everyone within one row needs the same classroom experience. The appropriate use of the extract here is to establish what the publisher called the ranges. It does not authorize a diagnosis or a complete educational plan from an FSIQ alone.
Two distinctions are particularly useful when reading a Stanford-Binet result. First, an edition number matters. A historical Form L-M score and a fifth edition deviation score do not become equivalent merely because both are called Stanford-Binet IQs. Second, the score type matters. A total, a verbal score and a nonverbal score answer related but different questions. When a number is repeated without its edition and score type, an apparently precise classification may rest on incomplete information.
Suppose a parent has a childhood report and later receives an adult assessment on a different instrument. It can be tempting to arrange the two descriptors as steps on a single ladder and infer a gain or loss from their wording. A more useful comparison begins with the reference ages, norms, task coverage and intervals. A change in language can occur without a meaningful numerical difference. A numerical difference can occur without a stable change in underlying standing. Neither possibility can be resolved by selecting the more flattering classification chart.
The situation is similar at the upper end. Some systems subdivide high scores into several giftedness categories, while others use one broad descriptor for a larger span. A finely divided naming scheme does not automatically demonstrate equally fine measurement precision. If a table has more named rows than the evidence can distinguish reliably, the names can create an impression of detail that the assessment does not support. Readers should look for the range the instrument can actually report and for how its manual treats unusually high performance.
A useful classification comparison therefore preserves missing information instead of filling it with familiar terms. The table above deliberately ends where its directly reviewed extract ends. That is enough to show that systems can share some boundaries while assigning different wording to others. It also models the question worth asking about any chart: which document, edition and score type supports each row? If those details are absent, the attractive layout of the chart cannot supply them.
5 Which labels does ACIS use on its public chart?
ACIS has its own public classification bands, which should be identified as ACIS terminology rather than a universal clinical standard. The following ranges are taken from the site's public score classification chart. They describe the published guides. They do not independently specify the lowest or highest score reportable by every assessment form, and they do not convert the online assessment into a diagnostic instrument.
Public score range
ACIS public label
Below 40
Profound Impairment
40 to 54
Severe Impairment
55 to 69
Mild Impairment
70 to 79
Borderline
80 to 89
Low Average
90 to 109
Average
110 to 119
High Average
120 to 129
Superior
130 to 134
Moderately Gifted
135 to 144
Highly Gifted
145 to 159
Exceptionally Gifted
160 to 174
Profoundly Gifted
175 to 177
Profoundly Gifted
The lower labels deserve particular care. In this table, Impairment is a public chart term. It is not a diagnosis of intellectual disability and is not a clinical assignment of severity. Those decisions involve functioning and assessment context that a score band does not contain. The focused page on what a low IQ means explains the distinction. A reader should never infer a person's capacity for daily life, legal status or support needs from one row of a public chart.
At the upper end, ACIS divides some territory more finely than the broad conventional Wechsler table. An ACIS score of 132 falls in Moderately Gifted; 140 falls in Highly Gifted. The IQ 140 guide discusses that single score, while the broader gifted IQ range article explains why identification conventions differ. These names do not mean that a score crossing 134 to 135 has established a qualitatively new type of person. The score, its precision and the purpose of interpretation remain more informative than the boundary alone.
Keeping the two top rows separate preserves the actual published chart. Both carry the same wording, and their presence is not evidence that the assessment precisely distinguishes every extreme percentile suggested by a mathematical normal curve. When an article displays an exceptionally rare theoretical score, a reader still needs evidence about the instrument's ceiling, the applicable norm tables and the number of observations supporting that region. Numerical possibility and demonstrated measurement resolution are different things.
ACIS is a paid, self-administered online assessment, and this page is published by its seller. Its technical manual describes the reference frame, analysis set, score construction and interpretation limits. Those are the sources for statements about ACIS measurement, including any edition-specific qualifications. The public labels are a communication layer over that information. They cannot repair an unsuitable test sitting, establish a diagnosis, or satisfy an institution's documentation requirements by themselves. The most useful product question is whether the available assessment answers the reader's actual purpose.
6 Where did older terms such as genius come from?
Historical IQ classifications belong to the scoring systems and social assumptions of their time. Lewis Terman's 1916 The Measurement of Intelligence includes a classification with language that modern readers will recognize from recycled online charts. The small extract below is presented as historical evidence. Its terminology should not be adopted as a contemporary way to describe people, and its ratio IQ categories should not be treated as identical to current deviation scores.
Range written in Terman's 1916 classification
Historical wording
Above 140
Near genius or genius
120 to 140
Very superior intelligence
110 to 120
Superior intelligence
90 to 110
Normal, or average, intelligence
Notice the first boundary: above 140, not a verified modern diagnostic threshold that begins exactly at 140. Also notice how the historical intervals are written with shared endpoints. Turning that table into nonoverlapping modern integer bins would require an editorial choice. A historical quotation should preserve the source rather than quietly making that choice and attributing it to the author.
Terman's system used the relationship between mental age and chronological age. Modern deviation IQ instead expresses performance relative to an age reference group. Those are different operations. The page on mental age explains why a ratio based on childhood age levels cannot simply be carried through adult life. When a biography supplies an unusually high historical IQ without a test edition or score method, a modern rarity calculator cannot restore the missing context.
The word genius adds another problem. In everyday use it can refer to exceptional creativity, achievement or influence. A test score measures performance on particular cognitive tasks under specified conditions. An eminent composer's work and a child's performance on an intelligence scale are not interchangeable records. A person may have a high measured score without producing an historically influential body of work, and important achievement can occur without a publicly documented IQ. The genius IQ discussion focuses on this difference between a cultural label and a test result.
Historical labels also show why terminology deserves scrutiny. Descriptive words can become stigmatizing when they are treated as a person's whole identity. Replacing an old label with a newer one is useful only if the interpretation also respects what was and was not measured. A modern sounding category attached to an unsupported number still leaves the central evidence problem unresolved. The meaningful improvement is a better account of the tasks, reference group, uncertainty and purpose.
For an old family report, preserve its wording when documenting history, but translate its practical meaning cautiously. Identify the instrument, year, age and original reporting conventions before comparing it with today's categories. If those details are unavailable, say that the exact modern equivalent is unknown. That is more informative than assigning a precise current label to a number produced under an unidentified historical system.
7 Why can the same standing have different IQ numbers?
A number changes when the reporting unit changes, even if the standardized distance from the mean stays the same. For a scale defined to have mean 100 and standard deviation 15, the standardized distance is the score minus 100, divided by 15. To express that same distance on a hypothetical mean 100, standard deviation 24 scale, multiply the distance by 24 and add 100. This is arithmetic for specified scales, not proof that two different instruments measure exactly the same thing.
Standardized distance above the mean
Mean 100, SD 15 scale
Hypothetical mean 100, SD 24 scale
Zero standard deviations
100
100
One standard deviation
115
124
Two standard deviations
130
148
Three standard deviations
145
172
All entries in this table are calculated from the definitions. A score of 148 on the specified SD 24 scale expresses the same standardized distance as 130 on the SD 15 scale. It is not an extra 18 points of ability. The guide to the standard deviation of 15 explains the common unit. Before using a conversion, verify the actual mean and standard deviation in the relevant test documentation.
The name Cattell is not sufficient to establish an SD of 24. Different instruments, versions and reported scores associated with that name have different conventions. The Culture Fair test discussion concerns the instrument family; a generic scale conversion should not be described as its official scoring table. Intertel's own admissions list, for example, lists different qualifying numbers for different Cattell entries. That alone shows why a reader needs the complete test name rather than an inherited assumption about a surname.
Even after the unit is known, the arithmetic is only a scale conversion. Imagine converting two temperatures correctly but discovering that one thermometer was used outdoors and the other inside a heated room. Matching units did not make the measurements equivalent. In assessment, differences in task coverage, norms and administration conditions can likewise remain after the numbers are expressed on a shared scale. An empirically supported comparison between tests requires more than rescaling their standard deviations.
A percentile can help communicate relative standing across different reporting units, provided the percentile actually comes from the relevant norm tables. It should not be manufactured from a convenient normal curve when an institution requires the publisher's score conversion. The score versus percentile guide explains the distinction. Two reported numbers that look far apart may reflect similar standing, while two identical numbers can summarize different abilities or reference groups.
For practical reading, record four details before comparing: the exact test and edition, the type of score, the scale and the reference population. If one is missing, describe the comparison as incomplete. This is especially useful for old school records and high IQ society documents, where different units and score types can coexist. A careful comparison often resolves an apparent contradiction without deciding that one assessment must have been wrong.
8 How much of a normal distribution lies in a band?
The proportion in a band depends on its numerical boundaries, not on the adjective assigned to it. For an ideal normal distribution with mean 100 and standard deviation 15, the area between two scores is calculated from their standardized distances. The table below uses standard deviation intervals to explain the geometry, rather than reproduce the detailed IQ percentile chart. These are theoretical proportions, not observed counts from every population or test administration.
Continuous interval on the SD 15 scale
Standardized interval
Approximate share under a normal model
Below 70
Below minus two standard deviations
2.28 percent
70 to below 85
Minus two to minus one
13.59 percent
85 to below 100
Minus one to zero
34.13 percent
100 to below 115
Zero to one
34.13 percent
115 to below 130
One to two
13.59 percent
130 and above
Two or more
2.28 percent
This table is our arithmetic from the standard normal model. It deliberately does not call 85 to below 115 the official Average band of every test. Statistical centrality and a publisher's descriptor boundaries are separate conventions. A range extending one standard deviation on either side of the mean contains about 68.27 percent under the model, while a narrower descriptor band can contain a smaller share. Neither choice changes the definition of a standard deviation.
Integer scores introduce an additional detail. A chart that says 90 to 109 may treat each displayed integer as representing an interval from half a point below to half a point above. Another informal chart may calculate directly from 90 and 110. These conventions produce slightly different band percentages. An article should state which arithmetic it uses instead of presenting one decimal place as though it came from the test publisher. Actual manual percentile tables can include rounding and score-specific conventions as well.
A percentile is cumulative: it describes standing relative to the comparison group below a score. A band proportion is the difference between two cumulative areas. The two quantities answer different questions. The percentile calculator helps with a single modeled score; the rarity calculator expresses an upper tail as a proportion or a one in X figure. Neither calculation proves that an extreme reported score is supported by a sufficiently informative test ceiling.
The normal curve is also not a census. A class selected for advanced mathematics, a clinic receiving referrals and an online group interested in cognitive testing need not display the same score distribution as the reference population. It would be a mistake to predict exactly how many people in one of those groups fall in each band by multiplying its head count by this table. Selection, measurement conditions and random variation can all matter.
The useful reading is modest: the model gives a common language for relative distances and areas. It explains why the same point difference changes percentiles more near the middle than near the tails. It does not supply a clinical judgment, describe an entire country, or guarantee that every observed sample follows a smooth bell curve. A classification chart becomes clearer when it keeps those limits visible.
9 What happens when measurement error crosses a boundary?
A category can change even when the evidence does not support a meaningful difference between two measured scores. Take a hypothetical instrument with a standard error of measurement of 1.6 points. Under a simple symmetric normal-error approximation, multiplying 1.6 by 1.96 gives about 3.1 points. An observed 129 would then have an illustrative interval of about 125.9 to 132.1, and an observed 130 an interval of about 126.9 to 133.1. Both cross a boundary at 130, and the intervals overlap extensively.
This calculation is an illustration, not a reproduction of a particular clinical report. An instrument may calculate intervals differently or use error estimates that depend on score, age or other features. Pearson's WAIS-5 sample, for example, identifies a Standard Error of Estimation method. For a real result, use the printed interval and the manual's interpretation. The reliability and validity guide explains why score precision is one part of evaluating a test rather than a guarantee of every possible use.
The hypothetical 129 and 130 make an important point without requiring a claim about an individual's true score. The one point difference changes a label under some schemes. It does not, by itself, establish a meaningful change in standing. A clear report can say that the point estimate falls in one band and that the uncertainty extends into an adjacent band. That wording preserves the reporting convention while avoiding false certainty.
Confidence language should also be precise. In the usual repeated-sampling interpretation, a procedure designed for 95 percent coverage would cover the relevant true score in about 95 percent of comparable repetitions under its assumptions. A completed interval is not a guarantee about the person, and it does not include every source of uncertainty. Unusual administration conditions, invalid responding or an inappropriate reference group may create problems that the ordinary interval does not repair.
Repeated testing adds more complications. If someone takes a similar test again, the difference can reflect practice, item familiarity, temporary conditions, measurement error or a change in the ability sampled. Selecting the highest score from many attempts does not remove uncertainty; it adds a selection process. Moving into a new descriptor band on a retest should therefore be interpreted with the test's retest guidance and the reason for repeating it. A label is an especially poor endpoint for judging whether a training program changed broad ability.
A decision maker may nevertheless need a rule. An organization can require a qualifying result on an accepted instrument, while a clinician may need to integrate a range of evidence. Those are different tasks. The existence of measurement error does not let a reader rewrite an organization's admissions policy, and an admissions cutoff does not justify ignoring uncertainty in clinical interpretation. The right question is how the particular decision process handles borderline evidence.
For a personal report, keep the category in perspective by recording the estimate, interval and date together. If the original report contains qualifications about validity or testing conditions, keep those with the number as well. This compact record is more useful than remembering a single adjective, especially when later comparing results that use different descriptors.
10 Are high IQ society thresholds another set of IQ levels?
Society thresholds are membership rules, not a universal taxonomy of intelligence. Three organizations illustrate the distinction. Mensa specifies the upper two percent on an approved, properly administered and supervised test. Intertel specifies the 99th percentile on a qualifying supervised test. The Triple Nine Society lists accepted test scores for its 99.9th percentile standard. Their current documentation governs eligibility.
Organization
Stated percentile criterion
What else the applicant must check
Mensa
98th percentile
Accepted tests, local procedures and documentation
Intertel
99th percentile
Qualifying supervised test and the applicable score listing
Triple Nine Society
99.9th percentile
Listed instrument, test date restrictions and submission rules
These rows summarize the organizations' own criteria as reviewed on September 24, 2026. They do not promise admission from an online result or from an arithmetic conversion. The focused high IQ society requirements guide compares the evidence applicants are asked to provide; the Mensa testing guide addresses that organization's testing route.
A threshold expressed as a percentile is useful because accepted tests do not all use identical reporting units. However, an organization may list a particular integer score for a particular instrument. A reader should use that published requirement rather than calculating a theoretical decimal IQ and rounding it in the direction they prefer. Accepted editions, test dates and score types can matter as much as the nominal percentile.
Membership also does not create a new measurement. A person who qualifies for one organization has provided evidence satisfying its rule. The membership itself does not produce a more precise estimate than the original assessment, nor does it show that a person just below the cutoff differs categorically from someone just above it. The social purpose of an organization and the psychometric interpretation of a score can coexist without being confused.
It is equally unhelpful to use a society's cutoff as a definition of a good IQ. The question what counts as a good IQ depends on what someone wants to understand, and a membership criterion answers only one narrow practical question. A score can be informative about an individual's profile without qualifying for any organization. A high accepted score can qualify someone without answering a clinical question, an educational planning question or a question about everyday adaptation.
If eligibility is the goal, start with the organization before purchasing an assessment. Check the accepted instrument and required evidence, then follow that process. If personal understanding is the goal, compare the measurement's scope and limitations instead of choosing it for the most impressive possible label. This prevents a common disappointment: obtaining a score that looks numerically sufficient but is not accepted for the purpose that motivated the testing.
11 Can changing norms move someone into a different level?
A score is relative to a reference group, so a change in norms can change the reported number and its descriptor without proving an equivalent change in the person. The historical rise in test performance across generations is often called the Flynn effect. Flynn's 1984 paper examined substantial gains across American test comparisons. Pietschnig and Voracek's 2015 meta-analysis reviewed a century of evidence and documented variation across contexts and abilities. The literature does not justify a universal annual adjustment for every current individual score.
An illustration makes the issue clear. Suppose two hypothetical norm tables map the same raw performance to 121 and 118. A classification with a boundary at 120 assigns different labels, although the answers did not change. The example is not an estimate for a real edition transition. It shows how a reference system can affect a category. In actual assessment, revisions can change content, administration and scoring as well as the norms, so a difference between editions may have several contributors.
The age of the examinee is another reference issue. Age-standardized scores compare performance with the applicable age group, not with a permanent personal baseline fixed in childhood. A person can learn more in absolute terms while keeping a similar relative standing among peers. Someone can also retain substantial knowledge while performance on some speeded tasks changes. A single descriptor hides the difference between absolute task performance and relative position.
The Flynn effect guide covers the historical findings and their limits. The practical lesson for classification is to record the edition and norm context. Do not take an old score, subtract a fixed number for every elapsed year and present the result as a newly measured IQ. That procedure cannot reconstruct the individual's current performance, determine whether the old administration was valid, or replace a professional judgment about the suitability of the evidence.
For comparing reports, a useful worksheet has separate columns for date tested, test edition, age, score type, point estimate, interval and descriptor. Once those are visible, an apparent contradiction can become understandable. The score may have been produced by another instrument, the wording may have changed, or the interval may show that the numerical difference is modest. Where an important decision depends on the comparison, the qualified professional interpreting the reports should examine those possibilities.
Classification is therefore best understood as a statement attached to a particular assessment, rather than a lifetime certificate with an unchanging adjective. A well documented result can remain informative, but its use in a new context should be justified. The question is not whether the old label has expired as a word. The question is whether the underlying evidence remains suitable for the new inference or decision.
12 How should you read your own level without overreading it?
Begin with the purpose of the assessment, then read the number, interval and profile before the classification. If the purpose is curiosity, a well explained score may help someone understand how they performed on a range of tasks. If the purpose is support, diagnosis or formal eligibility, the assessment must be suitable for that purpose and satisfy the relevant professional or institutional requirements. The same numerical result cannot be assumed to serve all of those uses. A level is only as sound as the score behind it, so it should come from a broad full scale test that combines many subtests across domains rather than a single task quiz.
The first practical step is to identify what the result actually is. A Full Scale IQ summarizes a broad set of tasks; an index describes a more specific domain; a subtest score uses its own scale. The Full Scale IQ guide and cognitive domains overview explain those distinctions. Treating an index as a total is one of the easiest ways to attach a correct label to the wrong claim. Treating a short practice result as an interchangeable broad assessment makes the same mistake at a different level.
Next read the interval and the conditions. A score close to a boundary carries no special exemption from error. If the assessment was interrupted, completed in an unsuitable language or administered outside its intended conditions, those facts matter before any descriptor is assigned. A narrow printed interval describes precision under a measurement model; it does not prove that every condition needed for a valid interpretation was met.
Then examine the profile. A broad score provides information about overall performance, while the component scores can show relative strengths and weaknesses. Variation does not automatically invalidate the total, and a striking pair of indices does not automatically establish a disorder. Differences should be interpreted with the instrument's comparison procedures, reliability evidence and the practical question being asked. A profile should add detail rather than become a collection of speculative diagnoses.
At the lower end, clinical interpretation is especially important. The AAIDD definition of intellectual disability requires significant limitations in intellectual functioning and adaptive behavior with developmental onset. A chart label alone cannot establish those elements. The American Psychiatric Association's explanation likewise describes an assessment that extends beyond a single IQ number. Public terms such as Extremely Low or Mild Impairment should not be used to assign clinical severity from a chart.
At the upper end, a high descriptor does not explain a person's personality, mental health, relationships or achievements. The gifted adults guide reviews evidence about adult giftedness without treating a trait checklist as an IQ test. A person can find a high score useful while resisting the temptation to explain every success or difficulty through it. The same restraint applies when interpreting someone else's result.
A practical written summary can be simple: name the instrument and edition, identify the score, give the interval, state the comparison group, and explain the purpose for which the result is being used. Add the descriptor as a convenience. For example, a hypothetical summary might say that a particular composite was 118 with the interval printed on the report, and that the system calls that score High Average. It should not claim that the descriptor determines a career, a diagnosis or a person's value.
ACIS sells a paid online assessment for personal cognitive measurement. It is self-administered and is not a clinical or diagnostic instrument. Its documentation is available for readers to inspect before choosing it. A reader who needs an evaluation for diagnosis, accommodations or an institutional decision should establish the accepted assessment route first. A reader seeking personal understanding should expect a transparent explanation of what the score can support, including the limits that remain after a label is applied.
The source boundary is explicit: public examples establish the labels they display, while unverified endpoints are not reconstructed as official tables. Continuous distribution proportions and hypothetical scale conversions are our arithmetic. ACIS terminology is taken from its public score chart. The page was reviewed on September 24, 2026; institutional requirements should be checked at the time they are used.
American Psychological Association. Testing and assessment. Resources on appropriate test use and assessment standards.
The interpretive standard is the joint 2014 Standards for Educational and Psychological Testing together with APA guidance on appropriate assessment use. Descriptors should communicate supported conclusions, preserve uncertainty and remain tied to the intended use of the score.
14 Frequently Asked Questions
What are IQ levels?
IQ levels are named ranges in a classification system. They summarize reported scores with terms such as Average or Superior. The words depend on the test, edition or chart, so they should be read alongside the original score, reference group and interval.
Is there one official IQ classification chart?
No single chart governs every instrument. Publishers and other reporting systems use different descriptors and sometimes different boundaries. For a particular assessment, the original report and its applicable manual are more informative than a generic table copied without a named source.
What IQ level is average?
Many familiar classification systems use Average for scores from 90 to 109 on a mean 100, standard deviation 15 scale. That descriptor band is a convention. It is different from defining the statistical middle as everyone within one standard deviation of the mean.
Is 110 called High Average on every test?
No. High Average is familiar in the conventional WAIS-IV and ACIS classifications, but wording varies by edition. Pearson's public WAIS-5 sample labels a score of 111 Above average. Use the terminology of the actual report rather than silently replacing it.
What does Superior mean on an IQ report?
In the conventional WAIS-IV and ACIS schemes, Superior refers to the 120 to 129 band. It is a description of measured standing under that system. It does not establish superior judgment, character, expertise or performance in every part of a person's life.
Does 130 always mean gifted?
A score near 130 is a common convention for intellectual giftedness on the SD 15 scale, but definitions and educational eligibility rules vary. A publisher may instead use a broad high-score descriptor. Identify the purpose and system before treating the word as an official status.
Is an IQ of 140 a genius score?
Genius is not a universal modern test classification. Terman's historical table used near genius or genius above 140, under an older scoring framework. Modern interpretation should identify the instrument, scale and uncertainty rather than equating a number with exceptional historical achievement.
Why do WAIS-IV and WAIS-5 labels differ?
The editions use different reporting language. Public WAIS-5 examples show Above average and Very high where an older conventional chart uses different wording for those scores. A change of adjective alone does not establish a change in the person's measured ability.
Are the WAIS-5 examples a complete classification table?
No. A sample report verifies the score and descriptor pairs it displays. It does not automatically disclose every band endpoint or outer category. This page preserves that limit and directs individual interpretation to the original report and applicable publisher documentation.
Are Stanford-Binet labels identical to Wechsler labels?
Some middle bands share familiar wording, but the systems should not be assumed identical. Edition, score type and manual terminology matter. This page shows a limited publisher-verified SB5 extract and avoids filling unverified extremes with terms taken from another test.
Why can 148 on one scale correspond to 130 on another?
If both scales have mean 100 but one has standard deviation 24 and the other 15, both numbers sit two standard deviations above the mean. That is a scale conversion, not evidence that two different tests are perfectly equivalent measures.
Do all Cattell tests use a standard deviation of 24?
No such assumption should be made from the name alone. Identify the complete instrument, version and reporting scale. A hypothetical SD 24 conversion is useful arithmetic, but it should not be presented as the official scoring rule for every Cattell instrument.
Is an IQ percentile the same thing as an IQ level?
A percentile describes relative standing in a comparison group. A level is a label assigned by a classification rule. Different systems can give the same score different labels while using similar percentiles, and an eligibility rule can add further requirements beyond either quantity.
How many people score between 85 and 115?
About 68.27 percent lie between those continuous boundaries under an ideal normal distribution with mean 100 and standard deviation 15, our arithmetic. That modeled proportion is not a guaranteed count in every selected group or a universal definition of the Average descriptor.
Does moving from 129 to 130 prove a higher ability level?
It changes the classification in some systems, but the difference may be smaller than the uncertainty of the measurement. Compare the intervals, test conditions and retest guidance. A one point threshold crossing does not by itself demonstrate a meaningful change in standing.
Should I use the highest result from several tests?
Selecting the highest result introduces a selection process and may also reflect practice or differences between instruments. Keep the test names, dates and conditions with the scores. A qualified interpreter can assess discrepant results when a consequential decision depends on them.
Do Mensa and Intertel define the same IQ level?
They set different membership criteria: Mensa uses the upper two percent and Intertel the 99th percentile on qualifying evidence. These are institutional rules with accepted tests and documentation requirements. They are not interchangeable clinical categories or guarantees based on any online score.
Can a low classification diagnose intellectual disability?
No. The assessment also requires evidence about adaptive functioning and developmental onset, interpreted in a clinical framework. Public score labels cannot establish those elements or determine severity. A low online result should not be used to diagnose yourself or another person.
Can new norms change a person's classification?
Yes, because the reference system helps determine the reported score. Revisions can also change tasks and scoring. That possibility does not justify applying a universal yearly correction to an old result; compare the actual editions and seek professional interpretation when needed.
What should I keep from my IQ report?
Keep the test name and edition, date, age at testing, score type, point estimate, interval and relevant conditions or qualifications. The descriptor is useful additional information. Those details make later comparisons more meaningful than remembering only an adjective or an isolated number.
What does an ACIS level establish?
It identifies the public ACIS band containing a reported score, subject to the assessment's documentation and limits. ACIS is self-administered and nonclinical. Its labels do not diagnose a condition or establish eligibility for an organization that requires a different assessment route.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.